ArXiv

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

Authors
Siddharth Damodharan, Radhika Gupta, Ali Alshami...
Categories
cs.AI, cs.CV
arXiv
https://arxiv.org/abs/2607.08745v1
PDF
https://arxiv.org/pdf/2607.08745v1

Brief

AUTOPILOT-VQA is an incident-centric visual question answering benchmark for dashcam videos that tests vision-language and multimodal LLMs on safety-critical, temporally grounded reasoning. Introduced by Damodharan et al. (arXiv:2607.08745v1, 2026-07-09), it poses structured questions across weather, road state, signage, participants, accident occurrence, impact location, and avoidability. Released with the AUTOPILOT CVPR 2026 competition; summary based on the abstract.

Why it matters

AUTOPILOT-VQA is an incident-centric VQA benchmark for dashcam video understanding introduced by Damodharan et al. (arXiv:2607.08745v1; published 2026-07-09) to evaluate safety-critical reasoning about real-world driving incidents and near-incidents.

Key details

  • The benchmark uses structured questions across safety-relevant categories—weather & lighting, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability—requiring temporally grounded, event-level reasoning beyond object recognition.
  • AUTOPILOT-VQA was released as part of the AUTOPILOT CVPR 2026 competition to standardize assessment of interpretable, robust, and safety-conscious vision-language systems for autonomous driving.
Source evidence

Abstract

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving incidents and near-incidents. The benchmark covers diverse safety-relevant categories, including weather and lighting conditions, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability-related reasoning. By requiring models to answer grounded questions about both contextual scene properties and event-level incident details, AUTOPILOT-VQA moves beyond object recognition toward temporally grounded, safety-aware reasoning. The dataset is released as part of the AUTOPILOT CVPR 2026 competition and provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios. Our benchmark support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.

Comment: CVPR Autopilot Workshop