Sketch图生图academic

VistaBot Architecture for 4D Geometry and Policy Execution

Figure 2: Architecture of VistaBot . (1) 4D Geometry Estimation with VGGT for pose and depth prediction; (2) View Synthesis via a video diffusion model with memory to generate spatiotemporal-consistent latent features; (3) Policy Execution using a Transformer that fuses scene and robot state features for closed-loop manipulation under unseen views.

输入图
生成结果

论文上下文

Paper title: VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis Abstract: Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based ($π_0$) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79$\times$ and 2.63$\times$ over ACT and $π_0$, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available. Passages referencing this figure: hile VLAs aspire to unify instructions and control across diverse tasks. However, a fundamental bottleneck undermines both paradigms: poor generalization across camera viewpoints. Unlike appearance shifts such as lighting or texture, viewpoint variations disrupt the spatial grounding between perception and action. As a result, even slight changes in perspective can cause dramatic policy failures (Fig. 1 ), often forcing practitioners to re-collect demonstrations or retrain models—an outcome directly at odds with the very scalability that end-to-end frameworks claim to deliver. To address cross-view generalization, prior efforts have primarily fallen into two directions: (1) Reconstruction-based methods [ 47 , 33 ] attempt to recover the underlying 3D geometry from synchronized multi-view s cess of multi-camera calibration further complicates real-world training and deployment. (2) Video generation-based methods [ 22 , 24 ] leverage multi-view-consistent video generative models and derive actions from predicted future frames. However, these typically lack integrated action learning and suffer from low inference efficiency, making them unsuitable for closed-loop robotic manipulation. Figure 1: Our proposed VistaBot demonstrates superior cross-view generalizability compared with SOTA visuomotor policy ( π 0 \pi_{0} and ACT). As shown in the figure, when the camera observation angle undergoes significant changes, VistaBot consistently maintains a high average success rate even under substantial camera viewpoint changes, whereas the success rates of baseline policies drop to near

完整 Prompt

A reference image is attached above. It is my rough sketch of what I
want my final figure to look like — sometimes hand-drawn, sometimes
an AI quick-draft. The quality is rough; details may be wrong; some
elements may be missing — but it shows the STRUCTURE / SPATIAL LAYOUT
I'm going for.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Refine my rough sketch into a polished publication-quality figure.

  - Preserve the SPATIAL STRUCTURE of the sketch: where the boxes are,
    how they connect, the overall reading order, the rough proportions.
  - You may correct details: better text labels (use the paper context
    to get the right component names), cleaner shapes, real icons
    instead of stick-figure placeholders.
  - Do NOT regenerate from scratch with a different layout. The
    finished figure must be visibly the same composition as the sketch.

If your output bears no spatial resemblance to the reference sketch,
you've failed the task. Refine the sketch — don't replace it. Just
give me the polished figure.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt