Inpaint图生图academic

HalluAudio Framework for LALM Hallucination Analysis

Figure 1: Overview of the HalluAudio framework. HalluAudio combines controlled multi-domain audio construction, unified prompting, and structured output validation to systematically analyze hallucination in LALMs.

输入图
生成结果

论文上下文

Paper title: HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models Abstract: Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain. Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth. We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music. HalluAudio comprises over 5K human-verified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA. To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions. Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes. We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music. Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs. Passages referencing this figure: improved instruction following (Frieske and Shi, 2024 ) . Overall, LALMs are rapidly evolving from transcription-focused systems to general audio reasoning models. Although several benchmarks have been proposed to evaluate various capabilities of LALMs (Yang et al. , 2024 ; Chen et al. , 2024 , 2026b , 2026a ) , their reliability, particularly with respect to hallucination, remains underexplored. Figure 1: Overview of the HalluAudio framework. HalluAudio combines controlled multi-domain audio construction, unified prompting, and structured output validation to systematically analyze hallucination in LALMs. Hallucinations in Audio Tasks Hallucination has been extensively studied in text Bang et al. ( 2025 ) and vision Cao et al. ( 2024 ) domains, revealing systematic failures in grounding, ill lacks a unified taxonomy, controlled contrastive audio pairs, multi-format evaluation tasks, and large-scale human-annotated assessments. HalluAudio addresses these gaps by introducing the first large-scale, multi-domain, and multi-dimensional benchmark for hallucination in LALMs. 3 HalluAudio Benchmark 3.1 Overview of HalluAudio The overall evaluation pipeline of HalluAudio is illustrated in Figure 1, which follows a modular, end-to-end process designed to systematically elicit, measure, and analyze hallucinations in LALMs. Specifically, given curated audio inputs from three domains: speech, sound, and music, we construct diverse task instances via template-based, parameterized prompts with domain-aware instantiation and controlled answer formats. These audio–prompt pairs are evaluate

完整 Prompt

A reference image is attached above. It is a partial figure for an
academic paper — most of it is already drawn, but ONE region has been
masked out / left blank. The masked region is described as:

  'The Evaluation Setup panel.'

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: This is an INPAINT EDIT, not a fresh generation.

  - The reference image IS your starting point. Preserve every pixel
    that is NOT in the masked region. Layout, colours, components,
    labels, typography of the un-masked region must be identical.
  - ONLY modify the masked region: fill it in with content matching
    the description above, in a style consistent with the rest of the
    figure.
  - Do NOT redraw the figure from scratch. Do NOT change the layout.
    Do NOT restyle the un-masked region.

If your output bears no resemblance to the reference image except in
the masked area, you've failed the task. Just give me the inpainted
final figure.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt