ArXiv

Visual Grounding in Zero-Shot Vision-Language Control

Authors
J. de Curtò, Dayani Plasencia, Diego Sánchez...
Categories
cs.RO, cs.AI, cs.CV
arXiv
https://arxiv.org/abs/2608.06154v1
PDF
https://arxiv.org/pdf/2608.06154v1

Brief

The paper probes whether zero-shot vision-language models (VLMs) genuinely ground decisions in images, using an input-ablation battery and 32,874 scored calls across two embodiments and three simulators. Most direct-control and local VLMs proved image-invariant or failed lateral/longitudinal grounding, yet an image-only controller achieved 0.090 m MAE and mirror equivariance. A symmetry-consensus guardian reached 0.954 balanced accuracy on 272 held-out frames. Full text not available; summary based on the abstract.

Why it matters

Across nine direct-action models, six structured local VLMs and an exploratory VLM-MPC hierarchy—evaluated with 32,874 scored calls over two embodiments and three simulators—no local VLM met the joint longitudinal-and-lateral grounding criterion; several models were image-invariant and a constant-SLOW policy outperformed a scripted geometric controller.

Key details

  • An image-only deterministic positive control estimated the lead gap with 0.090 m MAE and showed exact mirror equivariance, demonstrating that the visual stimuli contained sufficient information even though most models failed to use it.
  • A post-hoc leakage-controlled symmetry-consensus guardian picked two models (from 16 calibration frames) and froze a 2-of-4 hazard vote; on 272 held-out frames it achieved 0.954 balanced accuracy (episode-cluster bootstrap 95% CI [0.895,0.990]); abstaining on ties raised committed balanced accuracy to 0.973 at 0.824 coverage, and offline modular replay produced 0.934 action agreement with exact mirror equivariance.
Source evidence

Abstract

Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks. Across nine direct-action models, six structured local VLMs, and an exploratory VLM-MPC hierarchy, we analyse 32,874 scored calls over two embodiments and three simulators. The direct-control results are largely negative: a constant-SLOW policy outperforms a scripted geometric controller, several models are image-invariant or nearly constant, and models that recognize longitudinal hazards still fail to transform LEFT and RIGHT under reflection. No local VLM meets the joint longitudinal and lateral grounding criteria. However, an image-only deterministic positive control estimates the lead gap with 0.090 m MAE and exact mirror equivariance, confirming the stimuli carry sufficient visual information; the failures are modular, not universal. A post-hoc, leakage-controlled symmetry-consensus guardian selects two models from 16 calibration frames and freezes a 2-of-4 hazard vote across original and reflected views. On 272 held-out frames it reaches 0.954 balanced accuracy (episode-cluster bootstrap 95% CI [0.895,0.990]); nested leave-one-episode-out recovers the same pair and threshold in all 12 folds. Abstaining on ties raises committed balanced accuracy to 0.973 at 0.824 coverage. With deterministic perception retaining lateral authority, offline modular replay achieves 0.934 action agreement and exact mirror equivariance. These results support current VLMs as bounded, selective hazard assistants, not monolithic zero-shot controllers.