SketchImage → Imageacademic

MINT-Bench Taxonomy and Evaluation Pipeline Overview

Figure 1. Overview of MINT-Bench, consisting of a hierarchical multi-axis taxonomy, a three-stage data construction pipeline, and a hierarchical hybrid evaluation protocol.

Input image
Generated result

Paper context

Paper title: MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Abstract: Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due to limited benchmark coverage, weak diagnostic granularity, and insufficient multilingual support. We present \textbf{MINT-Bench}, a comprehensive multilingual benchmark for instruction-following TTS. MINT-Bench is built upon a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol that jointly assesses content consistency, instruction following, and perceptual quality. Experiments across ten languages show that current systems remain far from solved: frontier commercial systems lead overall, while leading open-source models become highly competitive and can even outperform commercial counterparts in localized settings such as Chinese. The benchmark further reveals that harder compositional and paralinguistic controls remain major bottlenecks for current systems. We release MINT-Bench together with the data construction and evaluation toolkit to support future research on controllable, multilingual, and diagnostically grounded TTS evaluation. The leaderboard and demo are available at https://longwaytog0.github.io/MINT-Bench/ Passages referencing this figure: with a scalable multi-stage data construction pipeline, enabling semantically clear, structurally controlled, and extensible benchmark construction across diverse control settings and languages. • We design a hierarchical hybrid evaluation protocol that progressively assesses content consistency, instruction following, and perceptual quality by combining objective tools with LALM-based judgments. Figure 1. Overview of MINT-Bench, consisting of a hierarchical multi-axis taxonomy, a three-stage data construction pipeline, and a hierarchical hybrid evaluation protocol. 2. Related Work 2.1. Controllable and Instruction-Following TTS Controllable TTS has evolved from structured attribute manipulation toward open-ended natural-language-guided generation. Early controllable systems typically reli ird, many existing resources remain limited in multilingual extensibility. MINT-Bench is designed to address these gaps through a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol, thereby providing a more comprehensive, diagnostic, and multilingual benchmark for instruction-following TTS. 3. MINT-Bench 3.1. Overview Figure 1 presents the overall framework of MINT-Bench. Rather than treating instruction-following TTS evaluation as a flat collection of prompts, MINT-Bench formulates it as a structured benchmark construction and evaluation problem. The framework consists of three tightly coupled components: a Hierarchical Multi-axis Taxonomy that organizes benchmark coverage, a controlled data construction pipel interpretation by modern scholars. 3.2. Hierarchical Multi-axis Taxonomy MINT-Bench is designed to cover a broad range of control cases in instruction-following TTS, while keeping the benchmark spac

The prompt

A reference image is attached above. It is my rough sketch of what I
want my final figure to look like — sometimes hand-drawn, sometimes
an AI quick-draft. The quality is rough; details may be wrong; some
elements may be missing — but it shows the STRUCTURE / SPATIAL LAYOUT
I'm going for.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Refine my rough sketch into a polished publication-quality figure.

  - Preserve the SPATIAL STRUCTURE of the sketch: where the boxes are,
    how they connect, the overall reading order, the rough proportions.
  - You may correct details: better text labels (use the paper context
    to get the right component names), cleaner shapes, real icons
    instead of stick-figure placeholders.
  - Do NOT regenerate from scratch with a different layout. The
    finished figure must be visibly the same composition as the sketch.

If your output bears no spatial resemblance to the reference sketch,
you've failed the task. Refine the sketch — don't replace it. Just
give me the polished figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts