Text to FigureText → Imageposter

Pose-Guided Person Image Synthesis — CVPR 2024 Poster

A CVPR 2024 poster presenting a novel training paradigm for pose-guided person image synthesis using a perception-refined decoder and hybrid-granularity attention to decouple appearance and pose information.

Paper context

Paper title: Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis Abstract: A CVPR 2024 poster presenting a novel training paradigm for pose-guided person image synthesis using a perception-refined decoder and hybrid-granularity attention to decouple appearance and pose information. Paper body (method & results): <!DOCTYPE html> <html lang="en"> <head> <meta content="text/html; charset=utf-8" http-equiv="content-type"/> <title>Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis</title> <!--Generated on Wed Feb 28 06:04:29 2024 by LaTeXML (version 0.8.7) http://dlmf.nist.gov/LaTeXML/.--> <meta content="width=device-width, initial-scale=1, shrink-to-fit=no" name="viewport"/> <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/css/bootstrap.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/ar5iv_0.7.4.min.css" rel="stylesheet" type="text/css"/> <link href="/static/browse/0.3.4/css/latexml_styles.css" rel="stylesheet" type="text/css"/> <script src="https://cdn.jsdelivr.net/npm/bootstrap@5.3.0/dist/js/bootstrap.bundle.min.js"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/html2canvas/1.3.3/html2canvas.min.js"></script> <script src="/static/browse/0.3.4/js/addons.js"></script> <script src="/static/browse/0.3.4/js/feedbackOverlay.js"></script> <base href="/html/2402.18078v1/"/></head> <body> <nav class="ltx_page_navbar"> <nav class="ltx_TOC"> <ol class="ltx_toclist"> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S1" title="1 Introduction ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">1 </span>Introduction</span></a></li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S2" title="2 Related Work ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2 </span>Related Work</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S2.SS1" title="2.1 Pose-Guided Person Image Synthesis ‣ 2 Related Work ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.1 </span>Pose-Guided Person Image Synthesis</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S2.SS2" title="2.2 Controllable Diffusion Models ‣ 2 Related Work ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">2.2 </span>Controllable Diffusion Models</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S3" title="3 Method ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3 </span>Method</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S3.SS1" title="3.1 Preliminary ‣ 3 Method ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.1 </span>Preliminary</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S3.SS2" title="3.2 Coarse-to-Fine Latent Diffusion ‣ 3 Method ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.2 </span>Coarse-to-Fine Latent Diffusion</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S3.SS3" title="3.3 Optimization ‣ 3 Method ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">3.3 </span>Optimization</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"> <a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S4" title="4 Experiments ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4 </span>Experiments</span></a> <ol class="ltx_toclist ltx_toclist_section"> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S4.SS1" title="4.1 Setup ‣ 4 Experiments ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.1 </span>Setup</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S4.SS2" title="4.2 Quantitative Comparison ‣ 4 Experiments ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.2 </span>Quantitative Comparison</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S4.SS3" title="4.3 Qualitative Comparison ‣ 4 Experiments ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.3 </span>Qualitative Comparison</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S4.SS4" title="4.4 User Study ‣ 4 Experiments ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.4 </span>User Study</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S4.SS5" title="4.5 Ablation Study ‣ 4 Experiments ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.5 </span>Ablation Study</span></a></li> <li class="ltx_tocentry ltx_tocentry_subsection"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S4.SS6" title="4.6 Appearance Editing ‣ 4 Experiments ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">4.6 </span>Appearance Editing</span></a></li> </ol> </li> <li class="ltx_tocentry ltx_tocentry_section"><a class="ltx_ref" href="https://arxiv.org/html/2402.18078v1#S5" title="5 Conclusion ‣ Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis"><span class="ltx_text ltx_ref_title"><span class="ltx_tag ltx_tag_ref">5 </span>Conclusion</span></a></li> </ol></nav> </nav> <div class="ltx_page_main"> <div class="ltx_page_content"> <div aria-label="Conversion errors have been found" class="package-alerts ltx_document" role="status"> <button aria-label="Dismiss alert" onclick="closePopup()"> <span aria-hidden="true"><svg aria-hidden="true" focusable="false" height="20" role="presentation" viewbox="0 0 44 44" width="20"> <path d="M0.549989 4.44999L4.44999 0.549988L43.45 39.55L39.55 43.45L0.549989 4.44999Z"></path> <path d="M39.55 0.549988L43.45 4.44999L4.44999 43.45L0.549988 39.55L39.55 0.549988Z"></path> </svg></span> </button> <p>HTML conversions <a href="https://info.dev.arxiv.org/about/accessibility_html_error_messages.html" target="_blank">sometimes display errors</a> due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on.</p> <ul arial-label="Unsupported packages used in this paper"> <li>failed: epic</li> </ul> <p>Authors: achieve the best HTML results from your LaTeX submissions by following these <a href="https://info.arxiv.org/help/submit_latex_best_practices.html" target="_blank">best practices</a>.</p> </div><div class="section" id="target-section"><div id="license-tr">License: arXiv.org perpetual non-exclusive license</div><div id="watermark-tr">arXiv:2402.18078v1 [cs.CV] 28 Feb 2024</div></div> <script> function closePopup() { document.querySelector('.package-alerts').style.display = 'none'; } </script> <article class="ltx_document ltx_authors_1line"> <h1 class="ltx_title ltx_title_document">Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis</h1> <div class="ltx_authors"> <span class="ltx_creator ltx_role_author"> <span class="ltx_personname"> Yanzuo Lu<sup class="ltx_sup" id="id1.1.id1">1</sup> </span></span> <span class="ltx_author_before">  </span><span class="ltx_creator ltx_role_author"> <span class="ltx_personname">Manlin Zhang<sup class="ltx_sup" id="id2.1.id1">1</sup> </span></span> <span class="ltx_author_before">  </span><span class="ltx_creator ltx_role_author"> <span class="ltx_personname">Andy J Ma<sup class="ltx_sup" id="id3.1.id1">1,2,3 </sup> </span><span class="ltx_author_notes">Corresponding author.</span></span> <span class="ltx_author_before">  </span><span class="ltx_creator ltx_role_author"> <span class="ltx_personname">Xiaohua Xie<sup class="ltx_sup" id="id4.1.id1">1,2,3</sup> </span></span> <span class="ltx_author_before">  </span><span class="ltx_creator ltx_role_author"> <span class="ltx_personname">Jianhuang Lai<sup class="ltx_sup" id="id5.1.id1">1,2,3,4</sup> </span></span> <span class="ltx_author_before">  </span><span class="ltx_creator ltx_role_author"> <span class="ltx_personname"><sup class="ltx_sup" id="id6.1.id1">1</sup>School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China <br class="ltx_break"/><sup class="ltx_sup" id="id7.2.id2">2</sup>Guangdong Province Key Laboratory of Information Security Technology, China <br class="ltx_break"/><sup class="ltx_sup" id="id8.3.id3">3</sup>Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China <br class="ltx_break"/><sup class="ltx_sup" id="id9.4.id4">4</sup>Pazhou Lab (HuangPu), Guangzhou, China <br class="ltx_break"/><span class="ltx_text ltx_font_typewriter" id="id10.5.id5" style="font-size:90%;">{luyz5, zhangmlin3}@mail2.sysu.edu.cn, {majh8, xiexiaoh6, stsljh}@mail.sysu.edu.cn</span> </span></span> </div> <div class="ltx_abstract"> <h6 class="ltx_title ltx_title_abstract">Abstract</h6> <p class="ltx_p" id="id11.id1">Diffusion model is a promising approach to image generation and has been employed for Pose-Guided Person Image Synthesis (PGPIS) with competitive performance. While existing methods simply align the person appearance to the target pose, they are prone to overfitting due to the lack of a high-level semantic understanding on the source person image. In this paper, we propose a novel Coarse-to-Fine Latent Diffusion (CFLD) method for PGPIS. In the absence of image-caption pairs and te

The prompt

Above I've shared:
(1) the full paper text,
(2) all paper figures labeled by figure number,
(3) the caption for the central poster figure I'm building.

TASK: This is a CONFERENCE POSTER. **NOT** an academic-paper figure.
Style requirements:

  - Multi-section layout with a clear poster structure: large title banner
    at the top with the paper title + author/affiliation strip, then 3-6
    distinct content panels arranged in columns or a grid.
  - Large legible fonts (text must be readable at 2 m viewing distance) —
    headings ≥ 60 pt visual size in the final image.
  - Use colour blocks / panel backgrounds to delineate sections (this is
    what makes it a poster, not a single-figure diagram).
  - Aspect ratio: portrait or landscape rectangle, NOT square.

If your output looks like a standard academic-paper figure (single panel,
no title banner, dense small text, no colour blocks), you've failed the
task. Render the COMPLETE poster, not just the central figure.

Just give me the final poster image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts