Text to FigureText → Imageacademic

SurFITR MLLM Tampering Generation Pipeline

Figure 4. Overview of the SurFITR MLLM-powered tampering generation and verification pipeline. MLLMs are used to select forensically valuable frames from source corpora and, with assistance from grounding models, generate semantically and visually grounded tampering instructions. Localised, mask-guided manipulation is then applied, followed by verification to ensure tampering quality.

Paper context

Paper title: SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation Abstract: We present the Surveillance Forgery Image Test Range (SurFITR), a dataset for surveillance-style image forgery detection and localisation, in response to recent advances in open-access image generation models that raise concerns about falsifying visual evidence. Existing forgery models, trained on datasets with full-image synthesis or large manipulated regions in object-centric images, struggle to generalise to surveillance scenarios. This is because tampering in surveillance imagery is typically localised and subtle, occurring in scenes with varied viewpoints, small or occluded subjects, and lower visual quality. To address this gap, SurFITR provides a large collection of forensically valuable imagery generated via a multimodal LLM-powered pipeline, enabling semantically aware, fine-grained editing across diverse surveillance scenes. It contains over 137k tampered images with varying resolutions and edit types, generated using multiple image editing models. Extensive experiments show that existing detectors degrade significantly on SurFITR, while training on SurFITR yields substantial improvements in both in-domain and cross-domain performance. SurFITR is publicly available on GitHub. Passages referencing this figure: types, generated using multiple image editing models. Extensive experiments show that existing detectors degrade significantly on SurFITR, while training on SurFITR yields substantial improvements in both in-domain and cross-domain performance. SurFITR is publicly available on GitHub. 1 1 1 https://github.com/mike-qz-wang/SurFITR . † † conference: Preprint; Apr 08, 2026; Arxiv † † copyright: none Figure 1. Visualisations from SurFITR showing realistic, fine grained tampering across diverse surveillance scenes. Top row: original images, middle row: tampered images (yellow boxes indicate manipulated regions), bottom row: edited pixel masks. 1. Introduction Recent publicly available image generation models (Ramesh et al. , 2022 ; Esser et al. , 2024 ; Labs, 2024 ; Wu et al. , 2025 ; Saharia a

The prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts