Text to Figure文生图poster

Social Reward Image Scoring Model Poster

A conference poster presenting 'Social Reward,' a scoring model trained on the Picsart Image-Social dataset. It compares Social Reward against HPS v2, ImageReward, and PickScore using quantitative tables, qualitative image grids, and win-rate charts, demonstrating superior alignment with human preferences in text-to-image generation.

论文上下文

Paper title: Social Reward: Evaluating and Enhancing Generative AI through Million-User Feedback from an Online Creative Community Abstract: A conference poster presenting 'Social Reward,' a scoring model trained on the Picsart Image-Social dataset. It compares Social Reward against HPS v2, ImageReward, and PickScore using quantitative tables, qualitative image grids, and win-rate charts, demonstrating superior alignment with human preferences in text-to-image generation. Paper body (method & results): SOCIAL REWARD: EVALUATING AND ENHANCING GENERATIVE AI THROUGH MILLION-USER FEED- BACK FROM AN ONLINE CREATIVE COMMUNITY Arman Isajanyan1∗, Artur Shatveryan1∗, David Kocharyan1, Zhangyang Wang1,2, Humphrey Shi1,3 1Picsart AI Research (PAIR), 2UT Austin, 3Georgia Tech {arman.isajanyan, artur.shatveryan, david.kocharyan, atlas.wang, humphrey.shi}@picsart.com ABSTRACT Social reward as a form of community recognition provides a strong source of motivation for users of online platforms to engage and contribute with content. The recent progress of text-conditioned image synthesis has ushered in a collab- orative era where AI empowers users to craft original visual artworks seeking community validation. Nevertheless, assessing these models in the context of collective community preference introduces distinct challenges. Existing evalu- ation methods predominantly center on limited size user studies guided by im- age quality and prompt alignment. This work pioneers a paradigm shift, un- veiling Social Reward - an innovative reward modeling framework that lever- ages implicit feedback from social network users engaged in creative editing of generated images. We embark on an extensive journey of dataset curation and refinement, drawing from Picsart: an online visual creation and editing plat- form, yielding a first million-user-scale dataset of implicit human preferences for user-generated visual art named Picsart Image-Social. Our analysis exposes the shortcomings of current metrics in modeling community creative preference of text-to-image models’ outputs, compelling us to introduce a novel predic- tive model explicitly tailored to address these limitations. Rigorous quantita- tive experiments and user study show that our Social Reward model aligns bet- ter with social popularity than existing metrics. Furthermore, we utilize Social Reward to fine-tune text-to-image models, yielding images that are more favored by not only Social Reward, but also other established metrics. These findings highlight the relevance and effectiveness of Social Reward in assessing com- munity appreciation for AI-generated artworks, establishing a closer alignment with users’ creative goals: creating popular visual art. Codes can be accessed at https://github.com/Picsart-AI-Research/Social-Reward. 1 INTRODUCTION Social reward mechanisms play a pivotal role in incentivizing and modulating human behavior. Grounded in neurobiology and psychology, positive social feedback, such as approval, validation, and recognition, are essential for maintaining social cohesion and individual well-being (Rudolph, 2021; Baumeister & Leary, 1995). This reward-driven behavior extends to online social platforms, where users seek satisfaction via the accumulation of their network’s peer engagement with shared content in forms such as likes or views (Deters & Mehl, 2013; Lemai Nguyen & Nallaperuma, 2023). Recently, the field of text-conditioned image synthesis has witnessed remarkable advancements, leading to the development of generative algorithms capable of producing high-fidelity images that closely adhere to textual descriptions. This technological breakthrough has significantly impacted the realm of online social networks, as it empowers users with a novel and creative means of content creation and sharing. As users leverage this technology to craft and post compelling visual content, they simultaneously tap into the well-established reward mechanisms of social validation and recog- nition. Given the pace with which the number of synthetic images grows filling the digital spaces ∗Equal contribution 1 arXiv:2402.09872v1 [cs.CV] 15 Feb 2024 Prompt Social Reward HPS v2 Image Reward PickScore Digital anime art of mattress-man with a serious expression in an empty warehouse, highly detailed energizing morning routines princess crown A beautiful woman standing in a dystopian city Figure 1: Best image out of 20 generations as chosen by different scoring models, including ours. of online creative communities (as of August 2023 more than 15 billion synthetic images had been generated globally (Valyaeva, 2023)) evaluating the performance of generative models within the context of social network popularity emerges as an important challenge. While social network popularity can be defined in many ways, with likes, views, and other similar types of user interactions traditionally serving as popularity estimates (McParlane et al., 2014; Ding et al., 2019), the nature of text-to-image technology, adopted by industry largely as co-editing tool integrated into creative platforms (Weisz et al., 2023; Huang & Grady, 2022), introduces another dimension to social network content popularity measurement, namely the frequency of synthetic image reuses for editing purposes by community members. This metric resonates with the popula- tion of artists and creators receiving social rewards when their synthetic images are being leveraged in the editing process by network peers. The central question then can be summarized as, to what ex- tent text-to-image models can produce visual content aligned with social popularity, which is defined as community preference for creative editing purposes? Recently, researchers have increasingly turned to human preferences as a guiding beacon, inspired by the transformative impact of human feedback in the realm of Large Language Models (LLM) (Ouyang et al., 2022; Nakano et al., 2022). In the domain of text-to-image generation, reward models have been harnessed to channel human feedback into the learning process Xu et al. (2023a); Wu et al. (2023a); Kirstain et al. (2023). These works, have sought to leverage human preferences to construct reward models that facilitate generative model evaluation and fine-tuning. Despite the commendable efforts, the existing reward models have notable limitations in our domain of interest. Some of them rely on limited-size data annotation process, as presented in Table 1, which cannot be deemed the equivalent of the “community-scale” feedback. Moreover, prompts utilized for dataset creation (collected from COCO Captions dataset (Chen et al., 2015) and open source prompt dataset DiffusionDB (Wang et al., 2023)) along with the moderation process and guidelines that adhere mainly to image fidelity and textual alignment as annotation criteria, do not emphasize creative purpose and hence, potentially, are limiting in expressing collective community creative preference. On the other hand, some other approaches collect explicit organic user feedback, but as a downside, exhibit a relatively small scale of collected user preference and absence of ”collective feedback” (when more than one user engages with a given image) as an important indicator of social popularity. While these approaches do capture a broad spectrum of user preferences, they never- theless showcase insufficiency to model social popularity in the context of community-level editing preference. These limitations are substantiated by our extensive quantitative and qualitative analysis. To bridge this gap and address the unique demands of text-to-image in the framework of creative social popularity, we introduce a novel concept: Social Reward. This paradigm shift in reward modeling leverages collective implicit feedback from social network users who specifically employ generated images for creative purposes. This distinction underscores the relevance and applicability of Social Reward to the creative editing process, providing a more faithful estimation of alignment between AI-generated images and community-level creative preference. However, the collection of 2 Social Reward data presents its own set of challenges, notably the inherent noise stemming from the implicit nature of user feedback, the absence of a formal annotation process with precise guidelines, and the unequal content exposure caused by social network-specific factors (such as some content being surfaced more frequently than the others, etc). In this work we embark on a comprehensive exploration, starting with data curation sourced from Picsart (https://picsart.com/): one of the world’s leading online visual creation and edit- ing platforms. Due to the inherent noise and subjectivity in individual user choices, the collective feedback, which implies multiple users’ editing choices for the given content item, is leveraged as a cleaning mechanism of organic implicit user behavior. Several more data collection techniques have been utilized for addressing such biases as caption bias, content exposure time, and user follower base biases. Our analysis reveals the shortcomings of existing metrics in evaluating text-to-image models’ fitness for generating popular art, which motivates us to introduce a new model explic- itly designed to address these limitations. Moreover, we demonstrate the potential of our model in fine-tuning text-to-image models to better align with community-level creative preference. Our contributions are outlined as follows: • We identify an unexplored, but extremely relevant dimension in human preference reward modeling for text-to-image models: evaluating the performance within the context of social network popularity for creative editing. Our analysis provides compelling evidence that existing reward models are ill-suited to capture this dimension. • We embark on a journey of dataset curation and leveraging Picsart’s creative community data. We build a large scale dataset of implicit human preferences motivated by creative editing intent over synthetic images, named Picsart Image-Social dataset. Contrary to existing methods, we utilize social network user feedback and curate dataset relying on editing community collective behavior. • Building upon this curated dataset, we develop and validate our Social Reward Model, showcasing its superiority for the given task, as evidenced in Table 4 and Figure 5. Further- more, our model captures distinct image attributes that go beyond mere aesthetics (see Figure 1), demonstrating its potential to enhance text-to-image model performance for community-level creative preference (see Table 5 and Figure 7). 2 RELATED WORK 2.1 TEXT-TO-IMAGE GENERATION Text-to-image generative models allow synthesizing images conditioned on the text input. GANs (Goodfellow et al., 2014) allowed for the first successful results in this realm (Zhang et al., 2017; Xu et al., 2017; Zhu et al., 2019; Liao et al., 2022). Transformer-based (Ramesh et al., 2021; Chang et al., 2023) models have also yielded great improvements. Recently diffusion-based architectures showed great ability of producing high fidelity images (Nichol et al., 2022; Ramesh et al., 2022; Xu et al., 2023b; Lu et al., 2023). LDM (Rombach et al., 2022) employs a diffusion process in the underlying latent space rather than directly in the pixel space. This approach delivers notable performance gains while also enhancing processing speed. 2.2 POPULARITY PREDICTION The domain of predicti

完整 Prompt

Above I've shared:
(1) the full paper text,
(2) all paper figures labeled by figure number,
(3) the caption for the central poster figure I'm building.

TASK: This is a CONFERENCE POSTER. **NOT** an academic-paper figure.
Style requirements:

  - Multi-section layout with a clear poster structure: large title banner
    at the top with the paper title + author/affiliation strip, then 3-6
    distinct content panels arranged in columns or a grid.
  - Large legible fonts (text must be readable at 2 m viewing distance) —
    headings ≥ 60 pt visual size in the final image.
  - Use colour blocks / panel backgrounds to delineate sections (this is
    what makes it a poster, not a single-figure diagram).
  - Aspect ratio: portrait or landscape rectangle, NOT square.

If your output looks like a standard academic-paper figure (single panel,
no title banner, dense small text, no colour blocks), you've failed the
task. Render the COMPLETE poster, not just the central figure.

Just give me the final poster image.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt