Text to Figure文生图poster

FastGen: Adaptive KV Cache Compression for LLMs

A conference poster presenting FastGen, an adaptive KV cache compression method for LLMs. It details motivation based on attention structure diversity, the method involving policy selection, and experimental results showing memory reduction and performance trade-offs across LLaMA models.

论文上下文

Paper title: Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs Abstract: A conference poster presenting FastGen, an adaptive KV cache compression method for LLMs. It details motivation based on attention structure diversity, the method involving policy selection, and experimental results showing memory reduction and performance trade-offs across LLaMA models. Paper body (method & results): <!DOCTYPE html><html lang="en"> <head> <meta http-equiv="content-type" content="text/html; charset=UTF-8"> <title>Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs</title> <!--Generated on Tue Oct 3 05:14:52 2023 by LaTeXML (version 0.8.7) http://dlmf.nist.gov/LaTeXML/.--> <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" type="text/css"> <link rel="stylesheet" href="https://browse.arxiv.org/latexml/ar5iv_0.7.4.min.css" type="text/css"> <link rel="stylesheet" href="https://browse.arxiv.org/latexml/styles.css" type="text/css"> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="https://browse.arxiv.org/latexml/addons.js"></script> <script src="https://browse.arxiv.org/latexml/feedbackOverlay.js"></script> </head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a href="#S1" title="1 Introduction ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"><a href="#S2" title="2 Related Work ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related Work</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a href="#S3" title="3 Adaptive KV Cache Compression ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Adaptive KV Cache Compression</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S3.SS1" title="3.1 Generative Inference of Autoregressive LLMs ‣ 3 Adaptive KV Cache Compression ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.1 </span>Generative Inference of Autoregressive LLMs</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S3.SS2" title="3.2 FastGen Framework ‣ 3 Adaptive KV Cache Compression ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.2 </span>FastGen Framework</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S3.SS3" title="3.3 Model Profiling ‣ 3 Adaptive KV Cache Compression ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3 </span>Model Profiling</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S3.SS4" title="3.4 KV Cache Compression Policies ‣ 3 Adaptive KV Cache Compression ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.4 </span>KV Cache Compression Policies</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a href="#S4" title="4 Diversity and Stability of Attention Structures ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>Diversity and Stability of Attention Structures</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S4.SS1" title="4.1 Head Distinctive Attention Structure ‣ 4 Diversity and Stability of Attention Structures ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Head Distinctive Attention Structure</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S4.SS2" title="4.2 Profile Tends to Be Consistent in One Sequence ‣ 4 Diversity and Stability of Attention Structures ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Profile Tends to Be Consistent in One Sequence</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a href="#S5" title="5 Experiment ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Experiment</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S5.SS1" title="5.1 Trade-off between performance and memory reduction ‣ 5 Experiment ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.1 </span>Trade-off between performance and memory reduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S5.SS2" title="5.2 Memory Footprint Reduction Analysis ‣ 5 Experiment ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.2 </span>Memory Footprint Reduction Analysis</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a href="#S5.SS3" title="5.3 Ablations ‣ 5 Experiment ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.3 </span>Ablations</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"><a href="#S6" title="6 Conclusion ‣ Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs" class="ltx_ref"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">6 </span>Conclusion</span></a></li> </ol></nav> </nav> <div class="ltx_page_main"> <div class="ltx_page_content"> <article class="ltx_document ltx_authors_1line"> <h1 class="ltx_title ltx_title_document">Model Tells You What to Discard: <br class="ltx_break">Adaptive KV Cache Compression for LLMs</h1> <div class="ltx_authors"> <span class="ltx_creator ltx_role_author"> <span class="ltx_personname">Suyu Ge<math id="id1.1.m1.1" class="ltx_Math" alttext="{}^{1}" display="inline"><semantics id="id1.1.m1.1a"><msup id="id1.1.m1.1.1" xref="id1.1.m1.1.1.cmml"><mi id="id1.1.m1.1.1a" xref="id1.1.m1.1.1.cmml"></mi><mn id="id1.1.m1.1.1.1" xref="id1.1.m1.1.1.1.cmml">1</mn></msup><annotation-xml encoding="MathML-Content" id="id1.1.m1.1b"><apply id="id1.1.m1.1.1.cmml" xref="id1.1.m1.1.1"><cn type="integer" id="id1.1.m1.1.1.1.cmml" xref="id1.1.m1.1.1.1">1</cn></apply></annotation-xml><annotation encoding="application/x-tex" id="id1.1.m1.1c">{}^{1}</annotation><annotation encoding="application/x-llamapun" id="id1.1.m1.1d">start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT</annotation></semantics></math>, Yunan Zhang<math id="id2.2.m2.2" class="ltx_Math" alttext="{}^{1*}" display="inline"><semantics id="id2.2.m2.2a"><msup id="id2.2.m2.2.2" xref="id2.2.m2.2.2.cmml"><mi id="id2.2.m2.2.2a" xref="id2.2.m2.2.2.cmml"></mi><mrow id="id2.2.m2.2.2.2.4" xref="id2.2.m2.2.2.2.3.cmml"><mn id="id2.2.m2.1.1.1.1" xref="id2.2.m2.1.1.1.1.cmml">1</mn><mo lspace="0.222em" id="id2.2.m2.2.2.2.4.1" xref="id2.2.m2.2.2.2.3.cmml">⁣</mo><mo id="id2.2.m2.2.2.2.2" xref="id2.2.m2.2.2.2.2.cmml">*</mo></mrow></msup><annotation-xml encoding="MathML-Content" id="id2.2.m2.2b"><apply id="id2.2.m2.2.2.cmml" xref="id2.2.m2.2.2"><list id="id2.2.m2.2.2.2.3.cmml" xref="id2.2.m2.2.2.2.4"><cn type="integer" id="id2.2.m2.1.1.1.1.cmml" xref="id2.2.m2.1.1.1.1">1</cn><times id="id2.2.m2.2.2.2.2.cmml" xref="id2.2.m2.2.2.2.2"></times></list></apply></annotation-xml><annotation encoding="application/x-tex" id="id2.2.m2.2c">{}^{1*}</annotation><annotation encoding="application/x-llamapun" id="id2.2.m2.2d">start_FLOATSUPERSCRIPT 1 * end_FLOATSUPERSCRIPT</annotation></semantics></math>, Liyuan Liu<math id="id3.3.m3.2" class="ltx_Math" alttext="{}^{2*}" display="inline"><semantics id="id3.3.m3.2a"><msup id="id3.3.m3.2.2" xref="id3.3.m3.2.2.cmml"><mi id="id3.3.m3.2.2a" xref="id3.3.m3.2.2.cmml"></mi><mrow id="id3.3.m3.2.2.2.4" xref="id3.3.m3.2.2.2.3.cmml"><mn id="id3.3.m3.1.1.1.1" xref="id3.3.m3.1.1.1.1.cmml">2</mn><mo lspace="0.222em" id="id3.3.m3.2.2.2.4.1" xref="id3.3.m3.2.2.2.3.cmml">⁣</mo><mo id="id3.3.m3.2.2.2.2" xref="id3.3.m3.2.2.2.2.cmml">*</mo></mrow></msup><annotation-xml encoding="MathML-Content" id="id3.3.m3.2b"><apply id="id3.3.m3.2.2.cmml" xref="id3.3.m3.2.2"><list id="id3.3.m3.2.2.2.3.cmml" xref="id3.3.m3.2.2.2.4"><cn type="integer" id="id3.3.m3.1.1.1.1.cmml" xref="id3.3.m3.1.1.1.1">2</cn><times id="id3.3.m3.2.2.2.2.cmml" xref="id3.3.m3.2.2.2.2"></times></list></apply></annotation-xml><annotation encoding="application/x-tex" id="id3.3.m3.2c">{}^{2*}</annotation><annotation encoding="application/x-llamapun" id="id3.3.m3.2d">start_FLOATSUPERSCRIPT 2 * end_FLOATSUPERSCRIPT</annotation></semantics></math>, Minjia Zhang<math id="id4.4.m4.1" class="ltx_Math" alttext="{}^{2}" display="inline"><semantics id="id4.4.m4.1a"><msup id="id4.4.m4.1.1" xref="id4.4.m4.1.1.cmml"><mi id="id4.4.m4.1.1a" xref="id4.4.m4.1.1.cmml"></mi><mn id="id4.4.m4.1.1.1" xref="id4.4.m4.1.1.1.cmml">2</mn></msup><annotation-xml encoding="MathML-Content" id="id4.4.m4.1b"><apply id="id4.4.m4.1.1.cmml" xref="id4.4.m4.1.1"><cn type="integer" id="id4.4.m4.1.1.1.cmml" xref="id4.4.m4.1.1.1">2</cn></apply></annotation-xml><annotation encoding="application/x-tex" id="id4.4.m4.1c">{}^{2}</annotation><annotation encoding="application/x-llamapun" id="id4.4.m4.1d">start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT</annotation></semantics></math>, Jiawei Han<math id="id5.5.m5.1" class="ltx_Math" alttext="{}^{1}" display="inline"><semantics id="id5.5.m5.1a"><msup id="id5.5.m5.1.1" xref="id5.5.m5.1.1.cmml"><mi id="id5.5.m5.1.1a" xref="id5.5.m5.1.1.cmml"></mi><mn id="id5.5.m5.1.1.1" xref="id5.5.m5.1.1.1.cmml">1</mn></msup><annotation-xml encoding="MathML-Content" id="id5.5.m5.1b"><apply id="id5.5.m5.1.1.cmml" xref="id5.5.m5.1.1"><cn type="integer" id="id5.5.m5.1.1.1.cmml" xref="id5.5.m5.1.1.1">1</cn></apply></annotation-xml><annotation encoding="application/x-tex" id="id5.5.m5.1c">{}^{1}</annotation><annotation encoding="application/x-llamapun" id="id5.5.m5.1d">start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT</annotation></semantics></math>, Jianfeng Gao<math id="id6.6.m6.1" class="ltx_Math" alttext="{}^{2}" display="inline"><s

完整 Prompt

Above I've shared:
(1) the full paper text,
(2) all paper figures labeled by figure number,
(3) the caption for the central poster figure I'm building.

TASK: This is a CONFERENCE POSTER. **NOT** an academic-paper figure.
Style requirements:

  - Multi-section layout with a clear poster structure: large title banner
    at the top with the paper title + author/affiliation strip, then 3-6
    distinct content panels arranged in columns or a grid.
  - Large legible fonts (text must be readable at 2 m viewing distance) —
    headings ≥ 60 pt visual size in the final image.
  - Use colour blocks / panel backgrounds to delineate sections (this is
    what makes it a poster, not a single-figure diagram).
  - Aspect ratio: portrait or landscape rectangle, NOT square.

If your output looks like a standard academic-paper figure (single panel,
no title banner, dense small text, no colour blocks), you've failed the
task. Render the COMPLETE poster, not just the central figure.

Just give me the final poster image.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt