Text to Figure文生图academic

TeMuDance Text and Music Driven Dance Generation

Figure 2. An overview of TeMuDance. We learn a motion-centred bank by contrastively aligning disjoint text–motion and music–dance datasets in a shared motion space, enabling similarity-filtered modality completion to form pseudo triplets. For generation, a pretrained diffusion Transformer is frozen as the backbone, while a text control branch steers denoising and produces rhythm-aligned, semantically controllable dances.

论文上下文

Paper title: TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation Abstract: Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making it difficult to guide specific movements through natural language descriptions. This limitation primarily stems from the absence of large-scale datasets that jointly align music, text, and motion for supervised learning of text-conditioned control. To address this challenge, we propose TeMuDance, a framework that enables text-based control for music-conditioned dance generation without requiring any manually annotated music-text-motion triplet dataset. TeMuDance introduces a motion-centred bridging paradigm that leverages motion as a shared semantic anchor to align disjoint music-dance and text-motion datasets within a unified embedding space, enabling cross-modal retrieval of missing modalities for end-to-end training. A lightweight text control branch is then trained on top of a frozen music-to-dance diffusion backbone, preserving rhythmic fidelity while enabling fine-grained semantic guidance. To further suppress noise inherent in the retrieved supervision, we design a dual-stream fine-tuning strategy with confidence-based filtering. We also propose a novel task-aligned metric that quantifies whether textual prompts induce the intended kinematic attributes under music conditioning. Extensive experiments demonstrate that TeMuDance achieves competitive dance quality while substantially improving tex Passages referencing this figure: ning strategy with confidence-based filtering. We also propose a novel task-aligned metric that quantifies whether textual prompts induce the intended kinematic attributes under music conditioning. Extensive experiments demonstrate that TeMuDance achieves competitive dance quality while substantially improving text-conditioned control over existing methods. Preprint. Under review. 1. Introduction Figure 1. The proposed TeMuDance method is able to generate dances conditioned on music and text jointly, producing sequences that are both rhythmically aligned and semantically controllable. In recent years, the media production industry has fueled a strong demand for automated, high-fidelity character animation (Mourot et al. , 2022 ; Zhu et al. , 2023 ) . As a complex form of expressive motion, alignment but lack textual annotations, while text–motion datasets provide language supervision but without accompanying music. The absence of music–text–motion triplets therefore prevents existing models from jointly learning rhythmic coherence and semantic control. To bridge this gap, we introduce TeMuDance, which enables text-based control for music-conditioned 3D dance generation, as shown in Figure 1 . Specifically, our model learns text controllability without requiring any paired music–text–motion supervision. At its core, TeMuDance introduces a motion-centred bridging mechanism that leverages motion as a shared semantic anchor to align separate music–dance and text–motion datasets within a unified embedding space. This unified representation enables cross-modal retrieval of missing

完整 Prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt