SketchImage → Imageacademic

Language-Embedded 3D Gaussians for Large-Scale Scenes

Figure 1. Language Embedded 3D Gaussians for Large-scale Scenes . Given Internet collections of images depicting large-scale scenes (left), we augment a 3D Gaussian representation with a learnable semantic bottleneck (center). Our approach enables interactive text-based virtual exploration (right), enabling users to zoom-in and view semantic regions of interests, such as the spires and lintels ∗ in the Milano Catedral depicted above. ∗ A spire refers to a tall, slender, pointed structure on top of a roof of a building or tower, and a lintel is a horizontal structural element that spans openings such as doors and windows.

Input image
Generated result

Paper context

Paper title: Lang3D-XL: Language Embedded 3D Gaussians for Large-scale Scenes Abstract: Embedding a language field in a 3D representation enables richer semantic understanding of spatial environments by linking geometry with descriptive meaning. This allows for a more intuitive human-computer interaction, enabling querying or editing scenes using natural language, and could potentially improve tasks like scene retrieval, navigation, and multimodal reasoning. While such capabilities could be transformative, in particular for large-scale scenes, we find that recent feature distillation approaches cannot effectively learn over massive Internet data due to challenges in semantic feature misalignment and inefficiency in memory and runtime. To this end, we propose a novel approach to address these challenges. First, we introduce extremely low-dimensional semantic bottleneck features as part of the underlying 3D Gaussian representation. These are processed by rendering and passing them through a multi-resolution, feature-based, hash encoder. This significantly improves efficiency both in runtime and GPU memory. Second, we introduce an Attenuated Downsampler module and propose several regularizations addressing the semantic misalignment of ground truth 2D features. We evaluate our method on the in-the-wild HolyScenes dataset and demonstrate that it surpasses existing approaches in both performance and efficiency. Passages referencing this figure: 3763973 † † isbn: 979-8-4007-2137-3/2025/12 † † ccs: Computing methodologies Rendering † † ccs: Computing methodologies Computer graphics † † ccs: Computing methodologies Machine learning † † ccs: Computing methodologies Scene understanding † † ccs: Human-centered computing Visualization Figure 1. Lang3D-XL unlocks interactive text-guided exploration over large-scale scenes, allowing users to zoom-into regions containing the queried text and view them from various viewpoints; see Fig. 1 for an illustration.

The prompt

A reference image is attached above. It is my rough sketch of what I
want my final figure to look like — sometimes hand-drawn, sometimes
an AI quick-draft. The quality is rough; details may be wrong; some
elements may be missing — but it shows the STRUCTURE / SPATIAL LAYOUT
I'm going for.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Refine my rough sketch into a polished publication-quality figure.

  - Preserve the SPATIAL STRUCTURE of the sketch: where the boxes are,
    how they connect, the overall reading order, the rough proportions.
  - You may correct details: better text labels (use the paper context
    to get the right component names), cleaner shapes, real icons
    instead of stick-figure placeholders.
  - Do NOT regenerate from scratch with a different layout. The
    finished figure must be visibly the same composition as the sketch.

If your output bears no spatial resemblance to the reference sketch,
you've failed the task. Refine the sketch — don't replace it. Just
give me the polished figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts