ArXiv

VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model

Authors
Diyari Mohammed Salih, Lingxiang Hu, Naima AitOufroukh-Mammar...
Categories
cs.RO
arXiv
https://arxiv.org/abs/2607.17171v1
PDF
https://arxiv.org/pdf/2607.17171v1

Brief

VIDAR is a visual-inertial dense reconstruction framework that anchors metric scale by coupling SVO+IMU odometry with the Depth Anything 3 (DA3) foundation model. It studies pose-conditioned DA3 and a decoupled alignment strategy; on EuRoC pose injection cuts scale error to ≈1% with mean F@0.10=0.463, while a decoupled hybrid reaches 0.676 without ground-truth poses. Summary is based on the paper abstract.

Why it matters

VIDAR couples SVO+IMU odometry with the Depth Anything 3 (DA3) foundation model to provide a metric anchor for dense monocular reconstruction, enabling fusion of detailed local geometry into a consistent global model.

Key details

  • On EuRoC, injecting poses reduces scale error to ≈1% and yields mean F@0.10 = 0.463; a decoupled hybrid alignment improves mean F@0.10 to 0.676 without using ground-truth poses. Evaluations also include TUM RGB-D.
Source evidence

Abstract

Monocular foundation models provide dense geometry but usually lack a stable metric scale. This paper presents VIDAR, a visual-inertial dense reconstruction framework that couples SVO+IMU odometry with Depth Anything 3. VIDAR uses the visual-inertial front end as a metric anchor: it provides camera poses, scale, and a consistent world frame for aligning dense foundation-model predictions across time. The foundation model then contributes detailed local geometry that is fused into a global reconstruction. We study both pose-conditioned DA3 and a decoupled alignment strategy. On EuRoC, pose injection reduces scale error to about 1\% and reaches 0.463 mean F@0.10; the decoupled hybrid improves this to 0.676 without ground-truth poses. Results on EuRoC and TUM RGB-D show that VIDAR is a practical route to metric dense monocular reconstruction.