Text to FigureText → Imageposter

FlowVid Consistent Video-to-Video Synthesis — Poster

This poster presents FlowVid, a method for consistent video-to-video synthesis. It edits the first frame and propagates changes using optical flow. Results show it outperforms baselines like TokenFlow and Rerender in user preference and runtime.

Paper context

Paper title: FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis Abstract: This poster presents FlowVid, a method for consistent video-to-video synthesis. It edits the first frame and propagates changes using optical flow. Results show it outperforms baselines like TokenFlow and Rerender in user preference and runtime. Paper body (method & results): <!DOCTYPE html> <html lang="en"> <head> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <title>FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis</title> <!--Generated on Fri Dec 29 16:50:30 2023 by LaTeXML (version 0.8.7) http://dlmf.nist.gov/LaTeXML/.--> <meta content="width=device-width, initial-scale=1, shrink-to-fit=no" name="viewport"/> <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv_0.7.4.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/latexml_styles.css" rel="stylesheet" type="text/css"/> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="/static/browse/0.3.4/js/addons.js"></script> <script src="/static/browse/0.3.4/js/feedbackOverlay.js"></script> <base href="/html/2312.17681v1/"/></head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="#S1" title="1 Introduction ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S2" title="2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related Work</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS1" title="2.1 Image-to-image Diffusion Models ‣ 2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.1 </span>Image-to-image Diffusion Models</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS2" title="2.2 Video-to-video Diffusion Models ‣ 2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.2 </span>Video-to-video Diffusion Models</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS3" title="2.3 Optical flow for video-to-video synthesis ‣ 2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.3 </span>Optical flow for video-to-video synthesis</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S3" title="3 Preliminary ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Preliminary</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S3.SS0.SSS0.Px1" title="Latent Diffusion Models ‣ 3 Preliminary ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Latent Diffusion Models</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S3.SS0.SSS0.Px2" title="ControlNet ‣ 3 Preliminary ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">ControlNet</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S4" title="4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>FlowVid</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S4.SS1" title="4.1 Inflating image U-Net to accommodate video ‣ 4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Inflating image U-Net to accommodate video</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S4.SS2" title="4.2 Training with joint spatial-temporal conditions ‣ 4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Training with joint spatial-temporal conditions</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S4.SS3" title="4.3 Generation: edit the first frame then propagate ‣ 4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.3 </span>Generation: edit the first frame then propagate</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S5" title="5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Experiments</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S5.SS1" title="5.1 Settings ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.1 </span>Settings</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS1.SSS0.Px1" title="Implementation Details ‣ 5.1 Settings ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Implementation Details</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS1.SSS0.Px2" title="Evaluation ‣ 5.1 Settings ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Evaluation</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S5.SS2" title="5.2 Qualitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.2 </span>Qualitative results</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S5.SS3" title="5.3 Quantitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.3 </span>Quantitative results</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS3.SSS0.Px1" title="User study ‣ 5.3 Quantitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">User study</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS3.SSS0.Px2" title="Pipeline runtime ‣ 5.3 Quantitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Pipeline runtime</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S5.SS4" title="5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.4 </span>Ablation study</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS4.SSS0.Px1" title="Condition combinations ‣ 5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Condition combinations</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS4.SSS0.Px2" title="Different control type: edge and depth ‣ 5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Different control type: edge and depth</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS4.SSS0.Px3" title="𝑣-prediction and ϵ-prediction ‣ 5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><math alttext="v" class="ltx_Math" display="inline"><semantics><mi>v</mi><annotation-xml encoding="MathML-Content"><ci>𝑣</ci></annotation-xml><annotation encoding="application/x-tex">v</annotation><annotation encoding="application/x-llamapun">italic_v</annotation></semantics></math>-prediction and <math alttext="\epsilon" class="ltx_Math" display="inline"><semantics><mi>ϵ</mi><annotation-xml encoding="MathML-Content"><ci>italic-ϵ</ci></annotation-xml><annotation encoding="application/x-tex">\epsilon</annotation><annotation encoding="application/x-llamapun">italic_ϵ</annotation></semantics></math>-prediction</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S5.SS5" title="5.5 Limitations ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.5 </span>Limitations</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="#S6" title="6 Conclusion ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">6 </span>Conclusion</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="#S7" title="7 Acknowledgments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">7 </span>Acknowledgments</span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"><a class="ltx_ref" href="#A1" title="Appendix A Webpage Demo ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_

The prompt

Above I've shared:
(1) the full paper text,
(2) all paper figures labeled by figure number,
(3) the caption for the central poster figure I'm building.

TASK: This is a CONFERENCE POSTER. **NOT** an academic-paper figure.
Style requirements:

  - Multi-section layout with a clear poster structure: large title banner
    at the top with the paper title + author/affiliation strip, then 3-6
    distinct content panels arranged in columns or a grid.
  - Large legible fonts (text must be readable at 2 m viewing distance) —
    headings ≥ 60 pt visual size in the final image.
  - Use colour blocks / panel backgrounds to delineate sections (this is
    what makes it a poster, not a single-figure diagram).
  - Aspect ratio: portrait or landscape rectangle, NOT square.

If your output looks like a standard academic-paper figure (single panel,
no title banner, dense small text, no colour blocks), you've failed the
task. Render the COMPLETE poster, not just the central figure.

Just give me the final poster image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts