Text to Figure텍스트 → 이미지academic

GeoLink 3D Structure Anchors for Cross-View Geo-Localization

Figure 3. The overview of the proposed framework. The key idea of GeoLink is to use scene-level 3D structure as a stable structural anchor for generalizable CVGL. Multi-view drone images are first converted into offline scene point clouds, which provide view-consistent structural priors during training. These 3D anchors guide 2D retrieval learning at two complementary levels: they regularize 2D features to reduce potentially view-biased dependencies, and they transfer instance-level structural relations from 3D space to 2D image features.

논문 컨텍스트

Paper title: GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization Abstract: Generalizable cross-view geo-localization aims to match the same location across views in unseen regions and conditions without GPS supervision. Its core difficulty lies in severe semantic inconsistency caused by viewpoint variation and poor generalization under domain shift. Existing methods mainly rely on 2D correspondence, but they are easily distracted by redundant shared information across views, leading to less transferable representations. To address this, we propose GeoLink, a 3D-aware semantic-consistent framework for Generalizable cross-view geo-localization. Specifically, we offline reconstruct scene point clouds from multi-view drone images using VGGT, providing stable structural priors. Based on these 3D anchors, we improve 2D representation learning in two complementary ways. A Geometric-aware Semantic Refinement module mitigates potentially redundant and view-biased dependencies in 2D features under 3D guidance. In addition, a Unified View Relation Distillation module transfers 3D structural relations to 2D features, improving cross-view alignment while preserving a 2D-only inference pipeline. Extensive experiments on multiple benchmarks show that GeoLink consistently outperforms state-of-the-art methods and achieves superior generalization across unseen domains and diverse weather environments. Passages referencing this figure: † copyright: acmlicensed † † doi: XXXXXXX.XXXXXXX † † isbn: 978-1-4503-XXXX-X/2018/06 1. Introduction Cross-view geo-localization (CVGL) is formulated as a retrieval task that geolocates multi-view imagery (e.g., drone and satellite) in GPS-denied environments (Zeng et al. , 2023 ) , aiming to identify geographically corresponding images from different views using a single-view query. As shown in Figure 1 , although existing methods (Li et al. , 2025b ; Shi et al. , 2019 ) have achieved promising performance in the same-area setting, they mainly rely on location labels for training. And they often degrade noticeably when deployed in unseen regions due to different scene layouts and building styles, limiting practical applications such as autonomous driving (Häne et al. , 2017 ) and robotic ade noticeably when deployed in unseen regions due to different scene layouts and building styles, limiting practical applications such as autonomous driving (Häne et al. , 2017 ) and robotic navigation (McManus et al. , 2014 ) . Therefore, CVGL tasks face two core challenges: maintaining semantic consistency under drastic viewpoint changes and achieving robust generalization under domain shift . Figure 1. Performance comparisons are conducted on SUES-200@(150m). Existing methods perform well in the same-area setting but degrade noticeably in cross-area evaluation, revealing limited generalization under domain shift. To address the first challenge, existing CVGL methods often align cross-view features directly in the 2D feature space. Although such alignment helps reduce the view gap, it m

프롬프트 본문

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

지금 이 프롬프트 시도하기

생성기에 자동으로 채워진 상태로 열립니다.

이 프롬프트 시도

관련 프롬프트