Text to FigureText → Imageposter

LOKI Synthetic Data Detection Benchmark — Poster

This ICLR Spotlight poster introduces LOKI, a benchmark for detecting synthetic data across video, image, 3D, text, and audio. It details dataset construction, annotation examples, evaluation tasks like judgement and explanation, and experimental results comparing LMMs against human performance.

Paper context

Paper title: LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models Abstract: This ICLR Spotlight poster introduces LOKI, a benchmark for detecting synthetic data across video, image, 3D, text, and audio. It details dataset construction, annotation examples, evaluation tasks like judgement and explanation, and experimental results comparing LMMs against human performance. Paper body (method & results): Published as a conference paper at ICLR 2025 LOKI: A COMPREHENSIVE SYNTHETIC DATA DE- TECTION BENCHMARK USING LARGE MULTIMODAL MODELS Junyan Ye1,2∗, Baichuan Zhou2* , Zilong Huang1* , Junan Zhang6,2* , Tianyi Bai2,5* , Hengrui Kang2 , Jun He1 , Honglin Lin2 , Zihao Wang1 , Tong Wu4 , Zhizheng Wu6,2 , Yiping Chen1 , Dahua Lin2,4 , Conghui He2,3†, Weijia Li1 † 1 Sun Yat-sen University, 2 Shanghai AI Laboratory, 3 SenseTime Research, 4 The Chinese University of Hong Kong, 5 The Hong Kong University of Science and Technology, 6 SDS, SRIBD, The Chinese University of Hong Kong, Shenzhen Video Synthesis or Real Label: Abnormal Details Selection 3D Text Audio These are the Rights … which make the Essence of … and which are the markes, whereby a man … , or Assembly of men, the Soveraign Power is placed … and resideth. For these are … Image … Image Video Text Audio Mainstream models (> 25) … Accuracy Recall GPT-Eval Score Scenery Human Video Singing Voice Music Environmental Sound … Audio News Scientific Papers Philosophy Wikipedia … Text Satellite Portrait Scenery … Image Documents Yes No Judgement <A> <B> Multiple Choice Abnormal Explanation <Answer> ... 3D <A> ... <B> … <C> … <D> … “Synthesis Audio” Diverse Modalities Heterogeneous Domain Multi-level Annotations Synthetic Detection Evaluation Framework 3D Nerf-based Others Gaussian-based Animals Fine-grained Anomaly Annotations: “Is the given audio a generated audio? / Please select a real audio.” … Metric Figure 1: Overview of LOKI benchmark. LOKI possesses four key characteristics: 1) Diverse modalities (video, image, 3D, text and audio); 2) Heterogeneous categories (26 detailed subcate- gories); 3) Multi-level annotations; 4) Multimodal synthetic data evaluation framework. ABSTRACT With the rapid development of AI-generated content, the future internet may be inundated with synthetic data, making the discrimination of authentic and credi- ble multimodal data increasingly challenging. Synthetic data detection has thus garnered widespread attention, and the performance of large multimodal models (LMMs) in this task has attracted significant interest. LMMs can provide natural language explanations for their authenticity judgments, enhancing the explain- ability of synthetic content detection. Simultaneously, the task of distinguish- ing between real and synthetic data effectively tests the perception, knowledge, and reasoning capabilities of LMMs. In response, we introduce LOKI, a novel benchmark designed to evaluate the ability of LMMs to detect synthetic data across multiple modalities. LOKI encompasses video, image, 3D, text, and audio modalities, comprising 18K carefully curated questions across 26 subcategories with clear difficulty levels. The benchmark includes coarse-grained judgment and multiple-choice questions, as well as fine-grained anomaly selection and expla- nation tasks, allowing for a comprehensive analysis of LMMs. We evaluated 22 open-source LMMs and 6 closed-source models on LOKI, highlighting their po- tential as synthetic data detectors and also revealing some limitations in the de- velopment of LMM capabilities. More information about LOKI can be found at https://opendatalab.github.io/LOKI/. ∗These authors contributed equally to this work. †Corresponding author(s). E-mail(s): liweij29@mail.sysu.edu.cn, heconghui@pjlab.org.cn 1 arXiv:2410.09732v2 [cs.CV] 21 Apr 2025 Published as a conference paper at ICLR 2025 1 INTRODUCTION With the rapid development of diffusion models (Rombach et al., 2022; Dhariwal & Nichol, 2021b) and large language models (Abdullah et al., 2022; Brown, 2020), AI-generated content (AIGC) technology has increasingly integrated synthetic multimodal data into our daily lives. For instance, tools like SORA (Brooks et al., 2024) can produce highly realistic video, while Suno (Shulman et al., 2022) enables the creation of music at a level comparable to professional artists. However, synthetic multimodal data also brings significant risks, including potential misuse and societal dis- ruption (Cooke et al., 2024; Ju et al., 2022). For example, the risks include generating fake news using large language models (LLMs), synthesizing fraudulent faces with diffusion models for scams, and potential contamination of internet training data. Due to the convenience of artificial intelligence synthesis, the future Internet may be saturated with AI-generated content, making the task of dis- cerning the authenticity and trustworthiness of multimodal data increasingly challenging. To address such threats, the field of synthetic data detection has garnered widespread attention in recent years (Barni et al., 2020; Frank et al., 2020; Gragnaniello et al., 2021; Shao et al., 2023; 2024).However, most current synthetic data detection methods are primarily focused on authentic- ity evaluation, with certain limitations regarding the human interpretability of the prediction results (Li et al., 2024b). The recent rapid advancement of large multimodal models (LMMs) has sparked curiosity about their performance in detecting synthetic multimodal data (Ku et al., 2023; Wu et al., 2024b). On one hand, for synthetic data detection tasks, LMMs can provide reasoning behind au- thenticity judgments in natural language, paving the way for enhanced explainability. On the other hand, the task of distinguishing between real and synthetic data involves the perception, knowledge, and reasoning abilities of multimodal data, serving as an excellent test of LMM capabilities. There- fore, the focus of this paper is to evaluate the performance of LMMs in synthetic data detection tasks. However, traditional synthetic data detection benchmarks, such as Fake2M (Lu et al., 2023b) and ASVSpoof 2019 (Wang et al., 2020b), primarily assess conventional detection methods, and eval- uations of LMMs in detecting multimodal synthetic data are still lacking. These benchmarks often miss fine-grained anomaly annotations represented in natural human language, making it difficult to transparently analyze the explainability capabilities of LMMs. FakeBench (Li et al., 2024a) aligns more closely with our objectives, but it only evaluates the performance of LMMs within a single standard image modality, lacking both breadth and depth. Specifically, FakeBench fails to explore other modalities such as audio and 3D data, focusing primarily on general image types and not conducting thorough tests on expert domain images like satellite imagery. To bridge this gap, we introduce LOKI, a comprehensive benchmark for evaluating the performance of LMMs on synthetic data detection. The key highlights of the LOKI benchmark include: • Diverse Modalities. LOKI includes high-quality multimodal data generated by recent popular synthetic models, covering video, image, 3D data, text, and audio. • Heterogeneous Categories. Our collected dataset includes 26 detailed categories across different modalities, such as specialized satellite and medical images; texts like philosophy and ancient Chinese; and audio data like singing voices, environmental sound and music. • Multi-level Annotations. LOKI includes basic ”Synthetic or Real” labels, suitable for fundamental question settings like true/false and multiple-choice questions. It also incorporates fine-grained anomalies for inferential explanations, enabling tasks like abnormal detail selection and abnormal explanation, to explore LMMs’ capabilities in explainable synthetic data detection. • Multimodal Synthetic Evaluation Framework. We propose a comprehensive evaluation framework that supports inputs of various data formats and over 25 mainstream multimodal models. On the LOKI benchmark, we evaluated 22 open-source LMMs, 6 advanced proprietary LMMs, and several expert synthetic detection models. Our key findings are summarized as follows: For synthetic data detection tasks we find: (1) LMMs exhibit moderate capabilities in synthetic data detection tasks, with certain levels of explainability and generalization, but there is still a gap compared to human performance; (2) Compared to expert synthetic detection models, LMMs ex- hibit greater explainability and, compared to humans, can detect features invisible to the naked eye, demonstrating promising developmental prospects. 2 Published as a conference paper at ICLR 2025 For LMMs capabilities we find: (1) Most LMMs exhibit certain model biases, tending to favor syn- thetic or real data in their responses; (2) LMMs lack of expert domain knowledge, performing poorly on specialized image types like satellite and medical images; (3) Current LMMs show unbalanced multimodal capabilities, excelling in image and text tasks but underperforming in 3D and audio tasks; (4) Chain-of-thought prompting enhances LMMs’ performance in synthetic data detection, whereas simple few-shot prompting falls short of providing the necessary reasoning support. These findings highlight the challenging and comprehensive nature of the LOKI task and the promis- ing future of LMMs in synthetic data detection tasks. 2 RELATED WORK 2.1 SYNTHETIC DATA DETECTION Currently, synthetic data detection has garnered widespread attention to prevent the misuse of mul- timedia synthetic data (Gragnaniello et al., 2021; Hou et al., 2023). The detection of synthetic data in image and audio has long been a popular research (Barni et al., 2020; Frank et al., 2020), while methods for synthetic video detection have recently emerged, such as DuB3D(Ji et al., 2024) and AIGVDet(Bai et al., 2024a). However, most work primarily focuses on the binary distinction be- tween authentic and synthetic data, resulting in poor interpretability. Some studies aim to enhance the interpretability of synthetic detection by providing latent representations(Dong et al., 2022), fea- ture explanations(Chai et al., 2020), and artifact localization (Zhang et al., 2023a; Shao et al., 2023; 2024); however, most research remains limited to the interpretability of abstract symbols, leaving a significant gap in alignment with human understanding. In practice, current AI-generated syn- thetic data still exhibits noticeable flaws, such as discontinuities in synthetic videos and insufficient geometric accuracy in 3D data. These shortcomings can be effectively captured and perceived by hu- man users(Tariang et al., 2024), who can provide reasonable explanations. However, existing expert synthetic data detection methods fail to provide human-interpretable bases for their judgments. 2.2 LARGE MULTIMODAL MODELS Recently, the rapid development of multimodal large models (LMMs) has been notable, with models like GPT-4o (OpenAI, 2024) and Claude 3.5 (Anthropic, 2024) excelling in various tasks such as scientific questioning (Lu et al., 2022; Yue et al., 2024) and commonsense reasoning (Talmor et al., 2018), showcasing exceptional perceptual and reasoning abilities (Bai et al., 2024b). Research has also applied LMMs to evaluate AIGC synthetic results, utilizing GPT to

The prompt

Above I've shared:
(1) the full paper text,
(2) all paper figures labeled by figure number,
(3) the caption for the central poster figure I'm building.

TASK: This is a CONFERENCE POSTER. **NOT** an academic-paper figure.
Style requirements:

  - Multi-section layout with a clear poster structure: large title banner
    at the top with the paper title + author/affiliation strip, then 3-6
    distinct content panels arranged in columns or a grid.
  - Large legible fonts (text must be readable at 2 m viewing distance) —
    headings ≥ 60 pt visual size in the final image.
  - Use colour blocks / panel backgrounds to delineate sections (this is
    what makes it a poster, not a single-figure diagram).
  - Aspect ratio: portrait or landscape rectangle, NOT square.

If your output looks like a standard academic-paper figure (single panel,
no title banner, dense small text, no colour blocks), you've failed the
task. Render the COMPLETE poster, not just the central figure.

Just give me the final poster image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts