Text to FigureText → Imageacademic

Universal Pose Denoiser Transformer Architecture

Figure 3 : The architecture of our universal pose denoiser f θ f_{\theta} . Given a mesh M M and its skeleton S S in rest pose, we sample plausible poses using multiple iterations of the denoising process illustrated here. The rest pose is concatenated as reference to the noisy pose, and to capture skeleton node semantics, we compute per-node semantic features from the textured mesh. This information is aggregated into a set of per-node tokens and used to estimate a denoised pose with a transformer. Skeleton edges are used to guide attention using a specialized Skeletal Attention block [ 6 ] .

Paper context

Paper title: ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes Abstract: Kinematic rigs provide a structured interface for articulating 3D meshes, but they lack an inherent representation of the plausible manifold of joint configurations for a given asset. Without such a pose space, stochastic sampling or manual manipulation of raw rig parameters often leads to semantic or geometric violations, such as anatomical hyperextension and non-physical self-intersections. We propose Video-informed Pose Spaces (ViPS), a feed-forward framework that discovers the latent distribution of valid articulations for auto-rigged meshes by distilling motion priors from a pretrained video diffusion model. Unlike existing methods that rely on scarce artist-authored 4D datasets, ViPS transfers generative video priors into a universal distribution over a given rig parameterization. Differentiable geometric validators applied to the skinned mesh enforce asset-specific validity without requiring manual regularizers. Our model learns a smooth, compact, and controllable pose space that supports diverse sampling, manifold projection for inverse kinematics, and temporally coherent trajectories for keyframing. Furthermore, the distilled 3D pose samples serve as precise semantic proxies for guiding video diffusion, effectively closing the loop between generative 2D priors and structured 3D kinematic control. Our evaluations show that ViPS, trained solely on video priors, matches the performance of state-of-the-art methods trained on synthetic artist-created 4D data in both plaus Passages referencing this figure: d/or parameter coupling, does not specify where in this control space an asset can plausibly go. Unfortunately, a plausible pose space (in short, pose space ) – the distribution of joint configurations that are valid and semantically consistent for any specific shape – is not straightforward to define. We ask: given an auto-rigged 3D mesh, can we automatically discover its (plausible) pose space? Figure 1 : Overview. We introduce ViPS, a universal feed-forward model that lifts static, auto-rigged meshes into a plausible and editable pose manifold. ViPS leverages the rich priors of foundational video models to automatically reveal a pose space that enables (a) manifold-constrained editing; (b) smooth pose-space interpolation, and (c) pose-guided video synthesis by using 3D proxies as struct

The prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts