A conference poster presenting DAS (Diffusion Alignment as Sampling), a method using tempered Sequential Monte Carlo to align diffusion models without fine-tuning, avoiding reward over-optimization while maintaining diversity and unseen rewards.
Paper title: Test-time Alignment of Diffusion Models without Reward Over-optimization Abstract: A conference poster presenting DAS (Diffusion Alignment as Sampling), a method using tempered Sequential Monte Carlo to align diffusion models without fine-tuning, avoiding reward over-optimization while maintaining diversity and unseen rewards. Paper body (method & results): <!DOCTYPE html> <html lang="en"> <head> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <title>Alignment without Over-optimization: Training-Free Solution for Diffusion Models</title> <!--Generated on Fri Jan 10 08:10:58 2025 by LaTeXML (version 0.8.8) http://dlmf.nist.gov/LaTeXML/.--> <meta content="width=device-width, initial-scale=1, shrink-to-fit=no" name="viewport"/> <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv.0.7.9.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv-fonts.0.7.9.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/latexml_styles.css" rel="stylesheet" type="text/css"/> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="/static/browse/0.3.4/js/addons_new.js"></script> <script src="/static/browse/0.3.4/js/feedbackOverlay.js"></script> <base href="/html/2501.05803v1/"/></head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S1" title="In Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S2" title="In Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related Work</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S2.SS1" title="In 2 Related Work ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.1 </span>Fine-tuning Diffusion Models for Alignment</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S2.SS2" title="In 2 Related Work ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.2 </span>Guidance Methods</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3" title="In Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Diffusion Alignment as Sampling (DAS)</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3.SS1" title="In 3 Diffusion Alignment as Sampling (DAS) ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.1 </span>Problem Setup: Aligning Diffusion Models with Rewards</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3.SS2" title="In 3 Diffusion Alignment as Sampling (DAS) ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.2 </span>Limitations of Existing Methods</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3.SS3" title="In 3 Diffusion Alignment as Sampling (DAS) ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3 </span>Sampling from Reward-aligned Target Distribution via Tempered SMC</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_subsubsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3.SS3.SSS1" title="In 3.3 Sampling from Reward-aligned Target Distribution via Tempered SMC ‣ 3 Diffusion Alignment as Sampling (DAS) ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3.1 </span>Backward Kernel</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsubsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3.SS3.SSS2" title="In 3.3 Sampling from Reward-aligned Target Distribution via Tempered SMC ‣ 3 Diffusion Alignment as Sampling (DAS) ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3.2 </span>Intermediate Targets: Approximate Posterior with Tempering</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsubsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3.SS3.SSS3" title="In 3.3 Sampling from Reward-aligned Target Distribution via Tempered SMC ‣ 3 Diffusion Alignment as Sampling (DAS) ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3.3 </span>Proposal: Approximating Locally Optimal Proposal</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsubsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S3.SS3.SSS4" title="In 3.3 Sampling from Reward-aligned Target Distribution via Tempered SMC ‣ 3 Diffusion Alignment as Sampling (DAS) ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3.4 </span>Asymptotic Behavior</span></a></li> </ol> </li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S4" title="In Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>Experiments</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"> <a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S4.SS1" title="In 4 Experiments ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Single Reward</span></a> <ol class="ltx_toclist ltx_toclist_subsection"> <li class="ltx_tocentry ltx_tocentry_subsubsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S4.SS1.SSS1" title="In 4.1 Single Reward ‣ 4 Experiments ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1.1 </span>Experiment Setup</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsubsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S4.SS1.SSS2" title="In 4.1 Single Reward ‣ 4 Experiments ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1.2 </span>Results</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S4.SS2" title="In 4 Experiments ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Multi Rewards</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S4.SS3" title="In 4 Experiments ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.3 </span>Online Black-box Optimization</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#S5" title="In Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Conclusions</span></a></li> <li class="ltx_tocentry ltx_tocentry_appendix"> <a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#A1" title="In Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">A </span>Pseudocodes</span></a> <ol class="ltx_toclist ltx_toclist_appendix"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#A1.SS1" title="In Appendix A Pseudocodes ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">A.1 </span>Pseudocode for Full Algorithm of DAS</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#A1.SS2" title="In Appendix A Pseudocodes ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">A.2 </span>Pseudocode for DAS with Adaptive Tempering</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#A1.SS3" title="In Appendix A Pseudocodes ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">A.3 </span>Pseudocode for Online black-box optimization</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_appendix"> <a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#A2" title="In Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">B </span>Introdution to SMC</span></a> <ol class="ltx_toclist ltx_toclist_appendix"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#A2.SS1" title="In Appendix B Introdution to SMC ‣ Alignment without Over-optimization: Training-Free Solution for Diffusion Models"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">B.1 </span>Feynman-Kac Models and Particle Filtering</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2501.05803v1#