Text to FigureText → Imageacademic

PASR Architecture for 3D Shape Pose Retrieval

Figure 2 : Overall architecture of PASR. (a) We distill semantic and spatial knowledge from a 2D foundation model into a 3D point cloud encoder. (b) Given a query image, we first perform an initial search by rendering the features of each 3D shape from a set of predefined poses and comparing them to the query image features. (c) Then, for the top-k candidates, we iteratively refine their poses to find the shape and pose that best reconstruct the query’s feature map via analysis-by-synthesis.

Paper context

Paper title: PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views Abstract: Single-view 3D shape retrieval is a fundamental yet challenging task that is increasingly important with the growth of available 3D data. Existing approaches largely fall into two categories: those using contrastive learning to map point cloud features into existing vision-language spaces and those that learn a common embedding space for 2D images and 3D shapes. However, these feed-forward, holistic alignments are often difficult to interpret, which in turn limits their robustness and generalization to real-world applications. To address this problem, we propose Pose-Aware 3D Shape Retrieval (PASR), a framework that formulates retrieval as a feature-level analysis-by-synthesis problem by distilling knowledge from a 2D foundation model (DINOv3) into a 3D encoder. By aligning pose-conditioned 3D projections with 2D feature maps, our method bridges the gap between real-world images and synthetic meshes. During inference, PASR performs a test-time optimization via analysis-by-synthesis, jointly searching for the shape and pose that best reconstruct the patch-level feature map of the input image. This synthesis-based optimization is inherently robust to partial occlusion and sensitive to fine-grained geometric details. PASR substantially outperforms existing methods on both clean and occluded 3D shape retrieval datasets by a wide margin. Additionally, PASR demonstrates strong multi-task capabilities, achieving robust shape retrieval, competitive pose estimation, and accurate categ Passages referencing this figure: st to partial occlusion and sensitive to fine-grained geometric details. PASR substantially outperforms existing methods on both clean and occluded 3D shape retrieval datasets by a wide margin. Additionally, PASR demonstrates strong multi-task capabilities, achieving robust shape retrieval, competitive pose estimation, and accurate category classification within a single framework. 1 Introduction Figure 1 : PASR achieves competitive clean, occluded, and unseen 3D shape retrieval from single images, while also enabling multi-tasking with accurate category classification and pose estimation. A fundamental challenge in computer vision is enabling machines to perceive the 3D world from 2D images. Single-view 3D shape retrieval [ 34 , 41 ] directly addresses this problem, requiring a model to r odels. Furthermore, both lines of work fail to sufficiently explore rich 3D information, like pose, that is available during training, precluding their use for fine-grained, multi-task applications. To address these limitations of prior work, we propose Pose-Aware 3D shape retrieval (PASR), a novel framework that reformulates retrieval as a feature-level analysis-by-synthesis problem. As shown in Figure 1 , our framework is designed for high robustness, particularly against partial occlusions, and exhibits strong generalization to unseen object meshes. PASR operates in two stages. During training, we distill fine-grained knowledge from a 2D foundation model into our 3D encoder. This design of obtaining an explicit 3D representation is the key to its strong generalization capabilities to no

The prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts