Key ElementsImage → Imageacademic

Spatial Intelligence Pre-Training Survey Structure

Figure 1 : Overview of the paper structure. We systematically structure the landscape of multi-modal data pre-training for forging Spatial Intelligence. This work is organized into four key pillars: (1) Background , introducing onboard sensors and foundational learning paradigms; (2) Platforms & Datasets , analyzing benchmarks across autonomous vehicles, drones, and other robotic systems; (3) Pre-Training Methodologies , categorized into single-modality, cross-modal (Camera/LiDAR-centric), and unified frameworks; and (4) Applications , highlighting downstream tasks from 3D perception to open-world planning.

Input image
Generated result

Paper context

Paper title: Forging Spatial Intelligence: A Roadmap of Multi-Modal Data Pre-Training for Autonomous Systems Abstract: The rapid advancement of autonomous systems, including self-driving vehicles and drones, has intensified the need to forge true Spatial Intelligence from multi-modal onboard sensor data. While foundation models excel in single-modal contexts, integrating their capabilities across diverse sensors like cameras and LiDAR to create a unified understanding remains a formidable challenge. This paper presents a comprehensive framework for multi-modal pre-training, identifying the core set of techniques driving progress toward this goal. We dissect the interplay between foundational sensor characteristics and learning strategies, evaluating the role of platform-specific datasets in enabling these advancements. Our central contribution is the formulation of a unified taxonomy for pre-training paradigms: ranging from single-modality baselines to sophisticated unified frameworks that learn holistic representations for advanced tasks like 3D object detection and semantic occupancy prediction. Furthermore, we investigate the integration of textual inputs and occupancy representations to facilitate open-world perception and planning. Finally, we identify critical bottlenecks, such as computational efficiency and model scalability, and propose a roadmap toward general-purpose multi-modal foundation models capable of achieving robust Spatial Intelligence for real-world deployment. Passages referencing this figure: Figure 1 : Overview of the paper structure. As depicted in Fig. 1 , these strategies form the core techniques to forge what we define as Spatial Intelligence – a capability that transcends simple detection to encompass holistic scene understanding, reasoning, and future prediction [ 142 , 270 , 126 ] .

The prompt

A reference image is attached above. It contains ONLY the visual elements
I want to use in my final figure — icons, charts, photos, illustrations —
arranged roughly in the layout I'm imagining. There are no panel borders,
arrows, text labels, or section titles in the reference yet.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Build the finished publication-quality figure USING the specific
visual elements I've already provided.

  - The icons / charts / photos / illustrations in the reference image
    are the ONES I want in the final figure. Use them. Don't substitute
    different icons. Don't pick generic stock visuals.
  - Their rough positions in the reference are my intended layout —
    keep them roughly where they are unless a small adjustment clearly
    helps composition.
  - Add the connecting structure: panel borders, arrows, text labels,
    section titles, captions — whatever is needed to make the figure
    coherent and publication-quality.
  - Do NOT generate the figure from scratch with different elements.
    Do NOT replace my icons with new ones.

If your output uses different icons / charts / illustrations from the
reference, you've failed the task. Just give me the final figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts