A conference poster for Snap Video, a text-to-video model. It demonstrates high-motion, out-of-domain, and 3D-consistent generation. It details a scalable FiT architecture and video-first diffusion process, with evaluations against Gen-2 and PikaLab.
Paper title: Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis Abstract: A conference poster for Snap Video, a text-to-video model. It demonstrates high-motion, out-of-domain, and 3D-consistent generation. It details a scalable FiT architecture and video-first diffusion process, with evaluations against Gen-2 and PikaLab. Paper body (method & results): <!DOCTYPE html> <html lang="en"> <head> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <title>Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis</title> <!--Generated on Thu Feb 22 10:37:20 2024 by LaTeXML (version 0.8.7) http://dlmf.nist.gov/LaTeXML/.--> <meta content="width=device-width, initial-scale=1, shrink-to-fit=no" name="viewport"/> <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv_0.7.4.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/latexml_styles.css" rel="stylesheet" type="text/css"/> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="/static/browse/0.3.4/js/addons.js"></script> <script src="/static/browse/0.3.4/js/feedbackOverlay.js"></script> <base href="/html/2402.14797v1/"/></head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S1" title="1 Introduction ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S2" title="2 Related Work ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related Work</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3" title="3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Method</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS1" title="3.1 Introduction to EDM ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.1 </span>Introduction to EDM</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS2" title="3.2 EDM for High-Resolution Video Generation ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.2 </span>EDM for High-Resolution Video Generation</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS3" title="3.3 Image-Video Modality Matching ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3 </span>Image-Video Modality Matching</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS4" title="3.4 Scalable Video Generator ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.4 </span>Scalable Video Generator</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS5" title="3.5 Training ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.5 </span>Training</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S3.SS6" title="3.6 Inference ‣ 3 Method ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.6 </span>Inference</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4" title="4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>Evaluation</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS1" title="4.1 Datasets ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Datasets</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS2" title="4.2 Evaluation Protocol ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Evaluation Protocol</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS3" title="4.3 Ablations ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.3 </span>Ablations</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS4" title="4.4 Quantitative Evaluation ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.4 </span>Quantitative Evaluation</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S4.SS5" title="4.5 Qualitative Evaluation ‣ 4 Evaluation ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.5 </span>Qualitative Evaluation</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S5" title="5 Conclusions ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Conclusions</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#S6" title="6 Acknowledgements ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">6 </span>Acknowledgements</span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A1" title="Appendix A Architecture details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref"><span class="ltx_text" style="font-size:144%;">A</span> </span><span class="ltx_text" style="font-size:144%;">Architecture details</span></span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A2" title="Appendix B Training details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref"><span class="ltx_text" style="font-size:144%;">B</span> </span><span class="ltx_text" style="font-size:144%;">Training details</span></span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"> <a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3" title="Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref"><span class="ltx_text" style="font-size:144%;">C</span> </span><span class="ltx_text" style="font-size:144%;">Additional Evaluation Results and Details</span></span></a> <ol class="ltx_toclist ltx_toclist_appendix"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS1" title="C.1 Sampler Parameters Ablations ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.1 </span>Sampler Parameters Ablations</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS2" title="C.2 User Studies ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.2 </span>User Studies</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS3" title="C.3 Qualitative Results Against Baselines ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.3 </span>Qualitative Results Against Baselines</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS4" title="C.4 Additional Qualitative Results ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.4 </span>Additional Qualitative Results</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS5" title="C.5 Hierarchical Video Generation ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.5 </span>Hierarchical Video Generation</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.14797v1#A3.SS6" title="C.6 Zero-Shot UCF101 Evaluation ‣ Appendix C Additional Evaluation Results and Details ‣ Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">C.6 </span>Zero-S