Text to Figure文生图poster

Total Selfie Full-Body Selfie Generation Poster

A conference poster for 'Total Selfie', a method generating full-body selfies from partial inputs via a selfie-conditioned inpainting model trained on synthetic data and fine-tuned per capture.

论文上下文

Paper title: Total Selfie: Generating Full-Body Selfies Abstract: A conference poster for 'Total Selfie', a method generating full-body selfies from partial inputs via a selfie-conditioned inpainting model trained on synthetic data and fine-tuned per capture. Paper body (method & results): Total Selfie: Generating Full-Body Selfies Bowei Chen Brian Curless Ira Kemelmacher-Shlizerman Steven M. Seitz University of Washington {boweiche, curless, kemelmi, seitz}@cs.washington.edu Input Selfies Background Two Examples of Total Selfie Automatic Target Pose Selection Pose Guidance … Reference for Target Pose Selection Figure 1. We generate full-body selfies of you (right), from self-captured images of your face and body (top left) and background. You can choose any target pose from a reference photo — we auto-select a set of good candidates from your photo collection (bottom). Abstract We present a method to generate full-body selfies from photographs originally taken at arms length. Because self- captured photos are typically taken close up, they have lim- ited field of view and exaggerated perspective that distorts facial shapes. We instead seek to generate the photo some one else would take of you from a few feet away. Our ap- proach takes as input four selfies of your face and body, a background image, and generates a full-body selfie in a de- sired target pose. We introduce a novel diffusion-based ap- proach to combine all of this information into high-quality, well-composed photos of you with the desired pose and background. 1. Introduction The prevalence of selfies has skyrocketed in recent years, with an estimated 93 million taken each day. Despite their popularity, they suffer from multiple shortcomings: (1) they capture only the upper portion of the subject, (2) the close- up camera viewpoint distorts faces and requires awkward poses (e.g., with arm reaching out), and (3) it is difficult to compose a shot that optimally captures both the subject and arXiv:2308.14740v2 [cs.CV] 3 Apr 2024 the scene. Instead, what if you could capture the full-body image that someone else would take of you in the scene? We call this a total selfie. As input, we require four selfies to cover different parts of your body, and a photo of the background that you would like to be composited into (Fig. 1). Based on this information, we generate convincing full-body photos of you in a specified target pose in the desired scene (Fig. 1 right). In practice, we automatically select candidate ref- erence photos from your photo collection, allowing you to choose one or more of them to determine the target pose. Solving this problem requires addressing a number of challenges. First, we must render a complete and accurate image of your body, piecing together separate close-up im- ages of your face, upper torso, legs, and shoes. Second, we must reproject you to a virtual viewpoint from several feet away – far enough to compose your full body within the scene. And third, we need to render you with a desired tar- get pose (where you’re not holding the camera), which can be completely different from the one you used to take the selfies. The target pose can be specified by any full-body image from your photo collection. To facilitate target pose selection, we auto-detect photos from your collection where you are wearing similar clothing to the input selfie set, lead- ing to results with more accurate body shapes for a given type of clothing. Most importantly, the resulting composite must retain your identity, expression, and clothing, but be composited realistically into the target scene with the de- sired pose and appearance. One approach to this problem would be to collect a paired dataset of selfies and full-body images of many peo- ple, and train a generative model on it. However, acquiring such a dataset would be time and cost intensive. Instead, we train a selfie to full-body model using a paired syn- thetic dataset, and further perform per-capture fine-tuning to narrow the gap between real and synthetic data. Specifi- cally, we first introduce a diffusion-based inpainting model trained in a self-supervised manner. This model takes four selfies as input and inpaints a full-body subject into a masked background. Given a set of input images, we first remove perspective distortion of the face selfie, and then fine-tune the trained model on these images to enhance the fidelity of generated full-body photo further. Our contributions can be summarized as follows: • We introduce a novel type of self-captured photo – total selfie – that captures your entire body, as if some one was taking a photo of you in the scene. • We propose a diffusion-based full-body generation model, followed by per-capture fine-tuning techniques, to generate total selfie from four selfies covering the body, a background image, and an auto-selected reference image as target pose. • We demonstrate results for twelve individuals in various scenes (e.g., indoor and sunny outdoor) and clothing (e.g., skirts) with a wide range of poses and expressions. Our experimental results outperform existing methods in gen- erating realistic and accurate full-body images. 2. Related Work Full-Body Image Generation. Extensive research has been dedicated to generating full-body images, either with- out specific conditions [15–17, 56] or with conditions such as pose [62], shape [49], or text prompts [22]. One related area of research to our task is human reposing [3, 10, 13, 14, 19, 21, 24, 26, 29, 32, 34, 42–44, 48, 53, 66, 68]. These methods transform a full-body (or partial-body) human im- age from one pose to another, with a target pose provided. For example, DisCo [53] designed a diffusion-based frame- work to achieve this by using a person image (in any pose), a background image, and a target pose. However, these meth- ods are tailored for single image input and fall short when applied to our task due to the inherent limitation of a single selfie in capturing the full view of the person. Another related line of work is Virtual Try-On [9, 11, 12, 20, 28, 31, 57, 59, 69], where the goal is to generate a visualization of how a garment might appear on a person, given the person image (in different clothing) and the gar- ment image. For instance, LaDI-VTON [36] introduced the first Latent Diffusion textual Inversion-Enhanced model to synthesize an image of the person wearing a specified cloth- ing item. However, these approaches usually assume sim- ple backgrounds, posing challenges in generating realistic shading within complex backgrounds. They also cannot al- ter footwear and facial expressions, crucial for our task. Despite these challenges, both streams of work assume input from third-person view images, not selfies. Selfie-Related Techniques. Numerous studies have ex- plored selfies for applications like reposing [33], face recog- nition [6, 27], style transfer [30, 52], novel view synthe- sis [1, 4, 23, 37, 38], relighting [7], and video stabiliza- tion [63, 64]. For example, [33] proposed a coordinate- based method to transform a typical selfie, mainly focusing on the face, into a neutral-pose portrait. Nevertheless, their approach was limited to upper body selfies and could not generate full-body selfies capturing both the subject and the surroundings. Our work is the first to propose and generate total selfie from arm-captured selfies. Diffusion Models. Diffusion models have recently demon- strated their success in various tasks such as text-to- image [41, 45, 47] and image-to-image translation [18, 60]. DreamBooth [46] was proposed to personalize a text-to- image model by fine-tuning the model on a few reference images. RealFill [51] introduced an image completion tech- nique to outpaint an input photo using reference images from the same scene. It fine-tuned a text-to-image inpaint- ing model using reference images and applied it to com- GT Generated by SD Selfie Simulation Face Upper Body Lower Body Shoes Selfie- Conditioned Inpainting Model Feature Embedding Extractor Face Undistortion, Augmentation Appearance Refinement Fine-Tune Selfie-Conditioned Inpainting Model Training Stage Per-Capture Fine-Tuning Stage I0 gt I0 gt · M 0 zt I0 f I0 u I0 s zt−1 Synthetic Selfies Photo Collection Face vh+Hbmt5+o0iyWj2aS0EDgoWQRI9hYya/e96Nqv1xa+4caJV4OalAjma/NUbxCQVBrCsdZdz01MkGFlGOF0WuqlmiaYjPGQdi2VWFAdZPNjp+jMKgMUxcqWNGiu/p7IsNB6IkLbKbAZ6WVvJv7ndVMTXQcZk0lqCSLRVHKkYnR7HM0YIoSwyeWYKYvRWREVaYGJtPyYbgLb+8Slr1mndZu3ioVxo3eRxFOIFTOAcPrqABd9AEHwgweIZXeHOk8+K8Ox+L1oKTzxzDHzifP89Ijgk=</latexit>If Upper Body EJ3FL8Qv1KOXRjDhRHaJUY8kXvSGiQsksCHd0oWGtrtpuyZkw2/w4kFjvPqDvPlvLAHBV8yct7M5mZFyacaeO6305hY3Nre6e4W9rbPzg8Kh+ftHWcKkJ9EvNYdUOsKWeS+o YZTruJoliEnHbCye3c7zxRpVksH80oYHAI8kiRrCxkl+9H6TVQbni1t0F0DrxclKBHK1B+as/jEkqDSEY617npuYIMPKMLprNRPNU0wmeAR7VkqsaA6yBbHztCFVYoipUt adBC/T2RYaH1VIS2U2Az1qveXPzP6UmugkyJpPUEmWi6KUIxOj+edoyBQlhk8twUQxeysiY6wMTafkg3BW315nbQbde+qfvnQqDRreRxFOINzqIEH19CEO2iBDwQYPMrvDn SeXHenY9la8HJZ07hD5zPH95Bjf4=</latexit>Iu Scene Lower Body Ib Shoes =">AB7HicbVBNS8NAEJ3Ur1q/qh69LaCp5IUY8FL3qrYGqhDWznbRLN5uwuxFK6W/w4kERr/4gb/4bt20O2vpg4PHeD DPzwlRwbVz32ymsrW9sbhW3Szu7e/sH5cOjlk4yxdBniUhUO6QaBZfoG24EtlOFNA4FPoajm5n/+IRK80Q+mHGKQUwHkecU WMlv3rX09VeueLW3DnIKvFyUoEczV75q9tPWBajNExQrTuem5pgQpXhTOC01M0pSN6A7lkoaow4m82On5MwqfRIlypY0Z K7+npjQWOtxHNrOmJqhXvZm4n9eJzPRdTDhMs0MSrZYFGWCmITMPid9rpAZMbaEMsXtrYQNqaLM2HxKNgRv+eV0qrXvMvax X290iB5HEU4gVM4Bw+uoAG30AQfGHB4hld4c6Tz4rw7H4vWgpPHMfOJ8/2M+N9A=</latexit>Is Mask M=">AB6nicbVDLSgNBEOyNrxhfUY9eBhPBU9gNoh4DXrwIEc0DkiXMTnqTIbOzy8ysEI+wYsHRbz6Rd78GyfJHjSxoKGo 6qa7K0gE18Z1v53c2vrG5lZ+u7Czu7d/UDw8auo4VQwbLBaxagdUo+ASG4Ybge1EIY0Cga1gdDPzW0+oNI/loxkn6Ed0IHnI GTVWeijflXvFkltx5yCrxMtICTLUe8Wvbj9maYTSMEG17nhuYvwJVYzgdNCN9WYUDaiA+xYKmE2p/MT52SM6v0SRgrW9K Qufp7YkIjrcdRYDsjaoZ62ZuJ/3md1ITX/oTLJDUo2WJRmApiYjL7m/S5QmbE2BLKFLe3EjakijJj0ynYELzl1dJs1rxLis X9VSjWRx5OETuEcPLiCGtxCHRrAYADP8ApvjnBenHfnY9Gac7KZY/gD5/MHVK+NEg=</latexit>M Augmented Data Final Output Reference Photo Yh+Ad410=">AB7HicbVBNS8NAEJ3Ur1q/qh69LaCp5IUY8FL3qrYGqhDWznbRLN5uwuxFK6W/w4k ERr/4gb/4bt20O2vpg4PHeDPzwlRwbVz32ymsrW9sbhW3Szu7e/sH5cOjlk4yxdBniUhUO6QaBZfoG24 EtlOFNA4FPoajm5n/+IRK80Q+mHGKQUwHkecUWMlv3rXU9VeueLW3DnIKvFyUoEczV75q9tPWBajNExQ rTuem5pgQpXhTOC01M0pSN6A7lkoaow4m82On5MwqfRIlypY0ZK7+npjQWOtxHNrOmJqhXvZm4n9eJ zPRdTDhMs0MSrZYFGWCmITMPid9rpAZMbaEMsXtrYQNqaLM2HxKNgRv+eV0qrXvMvaxX290iB5HEU4gV M4Bw+uoAG30AQfGHB4hld4c6Tz4rw7H4vWgpPHMfOJ8/10qN8w=</latexit>Ir Ia Pose-Guided Generation Inference Stage Fine-Tuned Output Io Target Pose aCp5IUY8FL3qrYGqhDWz3bRLN5uwOxFK6W/w4kERr/4gb/4bt20O2vpg4PHeDPzwlQKg67RTW1jc2t4rbpZ3dvf2D8uFRySZtxniUx0O6SGS6G4jwIlb6ea0ziU/DEc3cz8xyeujUjUA45THsR0oEQkGEUr +dW7HlZ75Ypbc+cgq8TLSQVyNHvlr24/YVnMFTJjel4borBhGoUTPJpqZsZnlI2ogPesVTRmJtgMj92Ss6s0idRom0pJHP198SExsaM49B2xhSHZtmbif95nQyj62AiVJohV2yxKMokwYTMPid9oTlDObaEMi3srY QNqaYMbT4lG4K3/PIqadVr3mXt4r5eaZA8jiKcwCmcgwdX0IBbaIPDAQ8wyu8Ocp5cd6dj0VrwclnjuEPnM8f2lSN9Q=</latexit>It … Φ I` I0 ` Automatic Target Pose Selection Figure 2. Overview of Total Selfie. First, we train a selfie-conditioned inpainting model based on a synthetic selfie to full-body dataset (blue box). Second, we fine-tune the trained model on a specific capture (orange box), and use it to produce a full-body self

完整 Prompt

Above I've shared:
(1) the full paper text,
(2) all paper figures labeled by figure number,
(3) the caption for the central poster figure I'm building.

TASK: This is a CONFERENCE POSTER. **NOT** an academic-paper figure.
Style requirements:

  - Multi-section layout with a clear poster structure: large title banner
    at the top with the paper title + author/affiliation strip, then 3-6
    distinct content panels arranged in columns or a grid.
  - Large legible fonts (text must be readable at 2 m viewing distance) —
    headings ≥ 60 pt visual size in the final image.
  - Use colour blocks / panel backgrounds to delineate sections (this is
    what makes it a poster, not a single-figure diagram).
  - Aspect ratio: portrait or landscape rectangle, NOT square.

If your output looks like a standard academic-paper figure (single panel,
no title banner, dense small text, no colour blocks), you've failed the
task. Render the COMPLETE poster, not just the central figure.

Just give me the final poster image.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt