Text to Figure文生图academic

HyPeR Framework for Audio Reasoning Overview

Figure 2: An overview of our framework HyPeR. First, we construct the PAQA dataset with complex audio collection and hierarchical audio augmentation. Then each training example is converted into a grounded reasoning Chain-of-Thought containing verifiable acoustic evidence for explicit reasoning. In Stage 1, a pretrained audio language model is optimized on PAQA for structured reasoning that explicitly links perception to reasoning. To handle acoustic cues, HyPeR further introduces implicit reasoning with <PAUSE>, which allows HyPeR to listen, pause, and then reason. In Stage 2, the SFT-i

论文上下文

Paper title: Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding Abstract: Recent Large Audio Language Models have demonstrated impressive capabilities in audio understanding. However, they often suffer from perceptual errors, while reliable audio reasoning is unattainable without first grounding the model's perception in structured auditory scenes. Inspired by Auditory Scene Analysis, we first introduce a Perception-Aware Question Answering (PAQA) dataset. PAQA implements a hierarchical decoupling strategy that separates speech from environmental sound and distinguishes multiple speakers, providing explicit perceptual reasoning for training. Building on this, we propose HyPeR, a two-stage Hybrid Perception-Reasoning framework. In Stage I, we finetune the model on PAQA to perceive acoustic attributes in complex audio. In Stage II, we leverage GRPO to refine the model's internal deliberation. We also introduce PAUSE tokens to facilitate latent computation during acoustically ambiguous phases and design perceptual consistency reward to align reasoning rationales with raw audio. Experiments across benchmarks demonstrate that HyPeR achieves absolute improvements over the base model, with performance comparable to large-scale models, stressing the effectiveness of hybrid perception-grounded reasoning for robust and multi-speaker audio understanding. Passages referencing this figure: nd reinforcement-learning (RL) post-training (Li et al. , 2025 ; Wu et al. , 2025 ) , the reasoning paths produced upon unreliable perceptions may hallucinate evidence and bring about bad comprehension in Audio Question-Answering (QA) (Yue et al. , 2025 ) . Moreover, current models often derive answers primarily from text-based reasoning without acoustic evidence, leading to weak audio grounding. Figure 1: ASA-inspired layered decoupling for perception-grounded audio reasoning. Rather than directly mapping audio to text, we separate background sound from speech and distinguish multiple speakers to construct verifiable acoustic evidence, and then perform grounded reasoning on top of this evidence. Previous research on audio grounding centered on Sound Event Detection with on- and off-set ti es to improve audio grounding. Drawing inspiration from Auditory Scene Analysis (ASA) , the human brain processes complex soundscapes through layered decoupling pathways Bregman ( 1994 ); Michelsanti et al. ( 2021 ) , effectively segregating the background sound (ENV) from the foreground one (SPEECH) and distinguishing multiple speakers before performing high-level semantic synthesis, as shown in Figure 1 . However, directly applying LALMs to background sound recognition remains unsatisfactory in practice. Specialized audio–text alignment models (e.g., CLAP Elizalde et al. ( 2023 , 2024 ); Ghosh et al. ( 2025 ); Niizumi et al. ( 2024 ) ) report mean Average Precision (mAP) values below 50% on FSD50K, a multi-label audio tagging dataset, while Qwen2-Audio-7B-Instruct only achieves 15% mAP i

完整 Prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt