ArXiv

GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation

Authors
Kaizhen Tan, Hanzhe Hong, Siru Tao
Categories
cs.CV
arXiv
https://arxiv.org/abs/2606.23669v1
PDF
https://arxiv.org/pdf/2606.23669v1

Brief

GeoFidelity-Bench introduces a 7,117-image benchmark across 109 OSM road segments in 25 cities to measure segment-conditioned geographic fidelity of text-to-image street-view generation. Evaluating six open-weight generators with city-only, street+neighborhood, and GPS-augmented prompts, authors find a 5.5 pp top-1 accuracy boost from local names (95% CI 3.4–7.7) yet near-zero segment discrimination margin and no clear benefit from GPS text, revealing a gap between local plausibility and true segment-level fidelity.

Why it matters

GeoFidelity-Bench comprises 7,117 curated Mapillary images covering 109 named OpenStreetMap road segments in 25 cities across six continents; each generated panel is ranked against nearest-segment, other segments in the same city, and segments from other cities to test local discrimination.

Key details

  • Appending street and neighborhood names to city-only prompts increases top-1 retrieval accuracy by 5.5 percentage points (95% CI: 3.4–7.7) over city-only prompts, but the similarity margin between the target and the nearest same-city segment remains near zero; appending raw GPS coordinates as ordinary text shows no statistically clear benefit, and using incorrect local names captures only part of the gain.
  • Six open-weight text-to-image generators were evaluated; held-out real-image queries successfully recover segment identity (so the references contain usable segment-level signal), and results indicate a persistent gap between generating city-/neighborhood-plausible street views and faithfully reproducing a specific road segment.
Source evidence

Abstract

Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generic city prior remains unclear. We introduce GeoFidelity-Bench, a reference-panel benchmark for segment-conditioned geographic fidelity in street-view generation. It contains 7,117 curated Mapillary images covering 109 named OpenStreetMap road segments in 25 cities across six continents. For each generated panel, the benchmark ranks the target reference panel against panels from the nearest segment in the same city, other segments in the same city, and segments from other cities, making local discrimination rather than absolute target similarity the primary test. We evaluate six open-weight text-to-image generators under city-only, street-and-neighborhood, and GPS-augmented prompts. Adding street and neighborhood names is associated with an increase of 5.5 percentage points in top-1 retrieval accuracy over city-only prompts, with a 95% confidence interval from 3.4 to 7.7 percentage points. However, the similarity margin between the target and the nearest segment in the same city remains near zero, indicating that local names improve broad local plausibility more than exact segment identity. Prompts that keep the city fixed but use incorrect street or neighborhood names further show that only part of the gain depends on the correct local names, while appending raw GPS coordinates as ordinary text yields no statistically clear additional benefit. Held-out real-image queries successfully recover segment identity, showing that the curated references contain usable segment-level signal. GeoFidelity-Bench thus reveals a persistent gap between city- or neighborhood-plausible street-view generation and faithful generation for a specific road segment.