Figure 1. Overview of MINT-Bench, consisting of a hierarchical multi-axis taxonomy, a three-stage data construction pipeline, and a hierarchical hybrid evaluation protocol.
Paper title: MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Abstract: Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due to limited benchmark coverage, weak diagnostic granularity, and insufficient multilingual support. We present \textbf{MINT-Bench}, a comprehensive multilingual benchmark for instruction-following TTS. MINT-Bench is built upon a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol that jointly assesses content consistency, instruction following, and perceptual quality. Experiments across ten languages show that current systems remain far from solved: frontier commercial systems lead overall, while leading open-source models become highly competitive and can even outperform commercial counterparts in localized settings such as Chinese. The benchmark further reveals that harder compositional and paralinguistic controls remain major bottlenecks for current systems. We release MINT-Bench together with the data construction and evaluation toolkit to support future research on controllable, multilingual, and diagnostically grounded TTS evaluation. The leaderboard and demo are available at https://longwaytog0.github.io/MINT-Bench/ Passages referencing this figure: with a scalable multi-stage data construction pipeline, enabling semantically clear, structurally controlled, and extensible benchmark construction across diverse control settings and languages. • We design a hierarchical hybrid evaluation protocol that progressively assesses content consistency, instruction following, and perceptual quality by combining objective tools with LALM-based judgments. Figure 1. Overview of MINT-Bench, consisting of a hierarchical multi-axis taxonomy, a three-stage data construction pipeline, and a hierarchical hybrid evaluation protocol. 2. Related Work 2.1. Controllable and Instruction-Following TTS Controllable TTS has evolved from structured attribute manipulation toward open-ended natural-language-guided generation. Early controllable systems typically reli ird, many existing resources remain limited in multilingual extensibility. MINT-Bench is designed to address these gaps through a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol, thereby providing a more comprehensive, diagnostic, and multilingual benchmark for instruction-following TTS. 3. MINT-Bench 3.1. Overview Figure 1 presents the overall framework of MINT-Bench. Rather than treating instruction-following TTS evaluation as a flat collection of prompts, MINT-Bench formulates it as a structured benchmark construction and evaluation problem. The framework consists of three tightly coupled components: a Hierarchical Multi-axis Taxonomy that organizes benchmark coverage, a controlled data construction pipel interpretation by modern scholars. 3.2. Hierarchical Multi-axis Taxonomy MINT-Bench is designed to cover a broad range of control cases in instruction-following TTS, while keeping the benchmark spac