Text to Figure텍스트 → 이미지poster

Snap Video Text-to-Video Generation Poster

A conference poster for Snap Video, a text-to-video model. It demonstrates high-motion, out-of-domain, and 3D-consistent generation. It details a scalable FiT architecture and video-first diffusion process, with evaluations against Gen-2 and PikaLab.

논문 컨텍스트

Paper title: Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis Abstract: A conference poster for Snap Video, a text-to-video model. It demonstrates high-motion, out-of-domain, and 3D-consistent generation. It details a scalable FiT architecture and video-first diffusion process, with evaluations against Gen-2 and PikaLab. Paper body (method & results): <!DOCTYPE html> <html lang="en"> <head> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <title>Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis</title> <!--Generated on Thu Feb 22 10:37:20 2024 by LaTeXML (version 0.8.7) http://dlmf.nist.gov/LaTeXML/.--> <meta content="width=device-width, initial-scale=1, shrink-to-fit=no" name="viewport"/> <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv_0.7.4.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/latexml_styles.css" rel="stylesheet" type="text/css"/> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="/static/browse/0.3.4/js/addons.js"></script> <script src="/static/browse/0.3.4/js/feedbackOverlay.js"></script> <base href="/html/2402.14797v1/"/></head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S1" title="1 Introduction ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S2" title="2 Related Work ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related Work</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3" title="3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Method</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS1" title="3.1 Introduction to EDM ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.1 </span>Introduction to EDM</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS2" title="3.2 EDM for High-Resolution Video Generation ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.2 </span>EDM for High-Resolution Video Generation</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS3" title="3.3 Image-Video Modality Matching ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3 </span>Image-Video Modality Matching</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS4" title="3.4 Scalable Video Generator ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.4 </span>Scalable Video Generator</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS5" title="3.5 Training ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.5 </span>Training</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS6" title="3.6 Inference ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.6 </span>Inference</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4" title="4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>Evaluation</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS1" title="4.1 Datasets ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Datasets</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS2" title="4.2 Evaluation Protocol ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Evaluation Protocol</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS3" title="4.3 Ablations ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.3 </span>Ablations</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS4" title="4.4 Quantitative Evaluation ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.4 </span>Quantitative Evaluation</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS5" title="4.5 Qualitative Evaluation ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.5 </span>Qualitative Evaluation</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S5" title="5 Conclusions ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Conclusions</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S6" title="6 Acknowledgements ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">6 </span>Acknowledgements</span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A1" title="Appendix A Architecture details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref"><span class="ltx_text" style="font-size:144%;">A</span> </span><span class="ltx_text" style="font-size:144%;">Architecture details</span></span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A2" title="Appendix B Training details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref"><span class="ltx_text" style="font-size:144%;">B</span> </span><span class="ltx_text" style="font-size:144%;">Training details</span></span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"> <a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3" title="Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref"><span class="ltx_text" style="font-size:144%;">C</span> </span><span class="ltx_text" style="font-size:144%;">Additional Evaluation Results and Details</span></span></a> <ol class="ltx_toclist ltx_toclist_appendix"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS1" title="C.1 Sampler Parameters Ablations ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.1 </span>Sampler Parameters Ablations</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS2" title="C.2 User Studies ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.2 </span>User Studies</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS3" title="C.3 Qualitative Results Against Baselines ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.3 </span>Qualitative Results Against Baselines</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS4" title="C.4 Additional Qualitative Results ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.4 </span>Additional Qualitative Results</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS5" title="C.5 Hierarchical Video Generation ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.5 </span>Hierarchical Video Generation</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS6" title="C.6 Zero-Shot UCF101 Evaluation ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.6 </span>Zero-S

프롬프트 본문

Above I've shared:
(1) the full paper text,
(2) all paper figures labeled by figure number,
(3) the caption for the central poster figure I'm building.

TASK: This is a CONFERENCE POSTER. **NOT** an academic-paper figure.
Style requirements:

  - Multi-section layout with a clear poster structure: large title banner
    at the top with the paper title + author/affiliation strip, then 3-6
    distinct content panels arranged in columns or a grid.
  - Large legible fonts (text must be readable at 2 m viewing distance) —
    headings ≥ 60 pt visual size in the final image.
  - Use colour blocks / panel backgrounds to delineate sections (this is
    what makes it a poster, not a single-figure diagram).
  - Aspect ratio: portrait or landscape rectangle, NOT square.

If your output looks like a standard academic-paper figure (single panel,
no title banner, dense small text, no colour blocks), you've failed the
task. Render the COMPLETE poster, not just the central figure.

Just give me the final poster image.

지금 이 프롬프트 시도하기

생성기에 자동으로 채워진 상태로 열립니다.

이 프롬프트 시도

관련 프롬프트