Key ElementsImage → Imageacademic

JEPA-WM World Model Training and Planning Pipeline

Figure 1: Left: Training of JEPA-WM: the encoder E ϕ , θ E_{\phi,\theta} embeds video and optionally proprioceptive observation, which is fed to the predictor P θ P_{\theta} , along with actions, to predict (in parallel across timesteps) the next state embedding. Right: Planning with JEPA-WM: sample action sequences, unroll the predictor on them, compute a planning cost L p L^{p} for each trajectory, and use this cost to iteratively refine the action sampling. The action encoder A θ A_{\theta} and proprioceptive encoder E θ p ​ r ​ o ​ p E_{\theta}^{prop} are not explicitly displayed in this figure for readability.

Input image
Generated result

Paper context

Paper title: What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? Abstract: A long-standing challenge in AI is to develop agents capable of solving a wide range of physical tasks and generalizing to new, unseen tasks and environments. A popular recent approach involves training a world model from state-action trajectories and subsequently use it with a planning algorithm to solve new tasks. Planning is commonly performed in the input space, but a recent family of methods has introduced planning algorithms that optimize in the learned representation space of the world model, with the promise that abstracting irrelevant details yields more efficient planning. In this work, we characterize models from this family as JEPA-WMs and investigate the technical choices that make algorithms from this class work. We propose a comprehensive study of several key components with the objective of finding the optimal approach within the family. We conducted experiments using both simulated environments and real-world robotic data, and studied how the model architecture, the training objective, and the planning algorithm affect planning success. We combine our findings to propose a model that outperforms two established baselines, DINO-WM and V-JEPA-2-AC, in both navigation and manipulation tasks. Code, data and checkpoints are available at https://github.com/facebookresearch/jepa-wms. Passages referencing this figure: Figure 1: Left: Training of JEPA-WM: the encoder E ϕ , θ E_{\phi,\theta} embeds video and optionally proprioceptive observation, which is fed to the predictor P θ P_{\theta} , along with actions, to predict (in parallel across timesteps) the next state embedding. We summarize JEPA-WM training and planning in figure 1 .

The prompt

A reference image is attached above. It contains ONLY the visual elements
I want to use in my final figure — icons, charts, photos, illustrations —
arranged roughly in the layout I'm imagining. There are no panel borders,
arrows, text labels, or section titles in the reference yet.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Build the finished publication-quality figure USING the specific
visual elements I've already provided.

  - The icons / charts / photos / illustrations in the reference image
    are the ONES I want in the final figure. Use them. Don't substitute
    different icons. Don't pick generic stock visuals.
  - Their rough positions in the reference are my intended layout —
    keep them roughly where they are unless a small adjustment clearly
    helps composition.
  - Add the connecting structure: panel borders, arrows, text labels,
    section titles, captions — whatever is needed to make the figure
    coherent and publication-quality.
  - Do NOT generate the figure from scratch with different elements.
    Do NOT replace my icons with new ones.

If your output uses different icons / charts / illustrations from the
reference, you've failed the task. Just give me the final figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts