This poster presents FlowVid, a method for consistent video-to-video synthesis. It edits the first frame and propagates changes using optical flow. Results show it outperforms baselines like TokenFlow and Rerender in user preference and runtime.
Paper title: FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis Abstract: This poster presents FlowVid, a method for consistent video-to-video synthesis. It edits the first frame and propagates changes using optical flow. Results show it outperforms baselines like TokenFlow and Rerender in user preference and runtime. Paper body (method & results): <!DOCTYPE html> <html lang="en"> <head> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <title>FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis</title> <!--Generated on Fri Dec 29 16:50:30 2023 by LaTeXML (version 0.8.7) http://dlmf.nist.gov/LaTeXML/.--> <meta content="width=device-width, initial-scale=1, shrink-to-fit=no" name="viewport"/> <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv_0.7.4.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/latexml_styles.css" rel="stylesheet" type="text/css"/> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="/static/browse/0.3.4/js/addons.js"></script> <script src="/static/browse/0.3.4/js/feedbackOverlay.js"></script> <base href="/html/2312.17681v1/"/></head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="#S1" title="1 Introduction ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S2" title="2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related Work</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS1" title="2.1 Image-to-image Diffusion Models ‣ 2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.1 </span>Image-to-image Diffusion Models</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS2" title="2.2 Video-to-video Diffusion Models ‣ 2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.2 </span>Video-to-video Diffusion Models</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S2.SS3" title="2.3 Optical flow for video-to-video synthesis ‣ 2 Related Work ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.3 </span>Optical flow for video-to-video synthesis</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S3" title="3 Preliminary ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Preliminary</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S3.SS0.SSS0.Px1" title="Latent Diffusion Models ‣ 3 Preliminary ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Latent Diffusion Models</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S3.SS0.SSS0.Px2" title="ControlNet ‣ 3 Preliminary ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">ControlNet</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S4" title="4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>FlowVid</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S4.SS1" title="4.1 Inflating image U-Net to accommodate video ‣ 4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Inflating image U-Net to accommodate video</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S4.SS2" title="4.2 Training with joint spatial-temporal conditions ‣ 4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Training with joint spatial-temporal conditions</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S4.SS3" title="4.3 Generation: edit the first frame then propagate ‣ 4 FlowVid ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.3 </span>Generation: edit the first frame then propagate</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="#S5" title="5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Experiments</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S5.SS1" title="5.1 Settings ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.1 </span>Settings</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS1.SSS0.Px1" title="Implementation Details ‣ 5.1 Settings ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Implementation Details</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS1.SSS0.Px2" title="Evaluation ‣ 5.1 Settings ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Evaluation</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S5.SS2" title="5.2 Qualitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.2 </span>Qualitative results</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S5.SS3" title="5.3 Quantitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.3 </span>Quantitative results</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS3.SSS0.Px1" title="User study ‣ 5.3 Quantitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">User study</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS3.SSS0.Px2" title="Pipeline runtime ‣ 5.3 Quantitative results ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Pipeline runtime</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="#S5.SS4" title="5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.4 </span>Ablation study</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS4.SSS0.Px1" title="Condition combinations ‣ 5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Condition combinations</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS4.SSS0.Px2" title="Different control type: edge and depth ‣ 5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title">Different control type: edge and depth</span></a></li> <li class="ltx_tocentry ltx_tocentry_paragraph"><a class="ltx_ref" href="#S5.SS4.SSS0.Px3" title="𝑣-prediction and ϵ-prediction ‣ 5.4 Ablation study ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><math alttext="v" class="ltx_Math" display="inline"><semantics><mi>v</mi><annotation-xml encoding="MathML-Content"><ci>𝑣</ci></annotation-xml><annotation encoding="application/x-tex">v</annotation><annotation encoding="application/x-llamapun">italic_v</annotation></semantics></math>-prediction and <math alttext="\epsilon" class="ltx_Math" display="inline"><semantics><mi>ϵ</mi><annotation-xml encoding="MathML-Content"><ci>italic-ϵ</ci></annotation-xml><annotation encoding="application/x-tex">\epsilon</annotation><annotation encoding="application/x-llamapun">italic_ϵ</annotation></semantics></math>-prediction</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="#S5.SS5" title="5.5 Limitations ‣ 5 Experiments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5.5 </span>Limitations</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="#S6" title="6 Conclusion ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">6 </span>Conclusion</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="#S7" title="7 Acknowledgments ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">7 </span>Acknowledgments</span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"><a class="ltx_ref" href="#A1" title="Appendix A Webpage Demo ‣ FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_