Text to Figure文生图academic

Hi-WM Human-in-the-World-Model Pipeline

Figure 2 : Human-in-the-World-Model (Hi-WM) pipeline. (a) In the pre-training stage, the policy network is learned from real-world data. (b) The policy is rolled out in closed-loop inside the world model. (c) When the policy exhibits suboptimal behavior, a human operator intervenes through hardware-agnostic input devices, such as a robot arm, keyboard, or VR controllers, and the corrective actions are executed directly in the world model. (d) The collected intervention segments are then added to the training set for post-training.

论文上下文

Paper title: Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training Abstract: Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipelines remain tied to physical execution: each correction requires robot time, scene setup, resets, and operator supervision in the real world. Meanwhile, action-conditioned world models have been studied mainly for imagination, synthetic data generation, and policy evaluation. We propose \textbf{Human-in-the-World-Model (Hi-WM)}, a post-training framework that uses a learned world model as a reusable corrective substrate for failure-targeted policy improvement. A policy is first rolled out in closed loop inside the world model; when the rollout becomes incorrect or failure-prone, a human intervenes directly in the model to provide short corrective actions. Hi-WM caches intermediate states and supports rollback and branching, allowing a single failure state to be reused for multiple corrective continuations and yielding dense supervision around behaviors that the base policy handles poorly. The resulting corrective trajectories are then added back to the training set for post-training. We evaluate Hi-WM on three real-world manipulation tasks spanning both rigid and deformable object interaction, and on two policy backbones. Hi-WM improves real-world success by 37.9 points on average over the base policy and by 19.0 points over a world-model closed-loop baseline, while world-model evaluation correlates strongly with real-world p Passages referencing this figure: , while world-model evaluation correlates strongly with real-world performance (r = 0.953). These results suggest that world models can serve not only as generators or evaluators, but also as effective corrective substrates for scalable robot post-training. keywords: Scalable Post-Training and Interactive World Model and Human-in-the-loop \checkdata [Project Page] https://hi-wm.github.io/ \teaser Figure 1 : Hi-WM overview. A policy is rolled out in a learned world model. When the rollout becomes failure-prone, a human provides short corrective actions. Cached states support rollback and branching, producing multiple continuations for post-training. 1 Introduction Large-scale robot pretraining has produced generalist policies that transfer across tasks, objects, and language instructions. Y

完整 Prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

立即试用此 Prompt

在生成器中自动预填此 prompt。

试用此 Prompt

相关 Prompt