Text to FigureText → Imageacademic

MCSC Multimodal Context-to-Script Task Overview

Figure 1: An overview of our Multimodal Context-to-Script Creation (MCSC) task. Models should comprehend the multimodal long contexts, create the plot, and output the structured script, which includes material-based shots and newly planned shots.

Paper context

Paper title: MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production Abstract: Real-world video creation often involves a complex reasoning workflow of selecting relevant shots from noisy materials, planning missing shots for narrative completeness, and organizing them into coherent storylines. However, existing benchmarks focus on isolated sub-tasks and lack support for evaluating this full process. To address this gap, we propose Multimodal Context-to-Script Creation (MCSC), a new task that transforms noisy multimodal inputs and user instructions into structured, executable video scripts. We further introduce MCSC-Bench, the first large-scale MCSC dataset, comprising 11K+ well-annotated videos. Each sample includes: (1) redundant multimodal materials and user instructions; (2) a coherent, production-ready script containing material-based shots, newly planned shots (with shooting instructions), and shot-aligned voiceovers. MCSC-Bench supports comprehensive evaluation across material selection, narrative planning, and conditioned script generation, and includes both in-domain and out-of-domain test sets. Experiments show that current multimodal LLMs struggle with structure-aware reasoning under long contexts, highlighting the challenges posed by our benchmark. Models trained on MCSC-Bench achieve SOTA performance, with an 8B model surpassing Gemini-2.5-Pro, and generalize to out-of-domain scenarios. Downstream video generation guided by the generated scripts further validates the practical value of MCSC. Datasets will be public soon. Passages referencing this figure: qual contribution. , Zihui Ren 2 1 1 footnotemark: 1 , Dingyi Yang 3 1 1 footnotemark: 1 , Liangyu Chen 1 , Qixiang Gao 2 , Tiezheng Ge 2 , Qin Jin 1 † † thanks: Corresponding author. 1 Renmin University of China 2 Alibaba Group 3 Nanyang Technological University {huanranhu, liangyuchen, qjin}@ruc.edu.cn {renzihui.rzh, gaoqixiang.gqx, tiezheng.gtz}@taobao.com dingyi.yang@ntu.edu.sg 1 Introduction Figure 1: An overview of our Multimodal Context-to-Script Creation (MCSC) task. Models should comprehend the multimodal long contexts, create the plot, and output the structured script, which includes material-based shots and newly planned shots. Automated video creation has become an important tool for modern content production across advertising, entertainment, and social media (Chi et al., 2020 l., 2025 ; Mu et al., 2026 ) . While recent advances in video creation have shown promising results, most existing systems assume simplified settings, such as relying on short text prompts or pre-curated multimodal inputs (Hong et al., 2022 ; Kondratyuk et al., 2023 ; He et al., 2023 ; Chen et al., 2024a ; Qian et al., 2025 ) , which do not reflect real-world creative workflows. As illustrated in Figure 1 , real-world video creation typically begins with pre-existing textual materials and redundant video materials Dancyger ( 2018 ); Pearlman ( 2018 ); Wang et al. ( 2019 ) . Guided by specific user instructions, creators must identify useful video shots, discard irrelevant content, and plan additional shots to bridge narrative gaps Bourgeois-Bougrine et al. ( 2014 ); Zhang et al. ( 2022 ); ) . This process makes it a long-context multimodal reasoning and structured planning problem, rather than a pure generation problem Wang et al. ( 2024a ); Ji et al. ( 2022 ); Han et al. ( 2023 ); Ch

The prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts