ArXiv

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation

Authors
JoungBin Lee, Jaewoo Jung, Jongmin Lee...
Categories
cs.CV
arXiv
https://arxiv.org/abs/2606.26087v1
PDF
https://arxiv.org/pdf/2606.26087v1

Brief

MVTrack4Gen (Lee et al., arXiv 2026-06-24) augments camera-conditioned novel-view video diffusion models with multi-view point-tracking supervision. It routes attention-layer correspondence features into an auxiliary tracking head and jointly trains a point-tracking loss, improving motion fidelity and cross-view geometric consistency; authors report state-of-the-art geometric consistency and competitive camera accuracy across benchmarks.

Why it matters

MVTrack4Gen (Lee et al., 2026-06-24) augments camera-conditioned novel-view video diffusion models by routing attention-layer correspondence features into an auxiliary multi-view point-tracking head and jointly training a point-tracking loss, explicitly supervising geometry and motion to reduce cross-view and temporal misalignment.

Key details

  • The authors report state-of-the-art geometric consistency and competitive camera accuracy across diverse benchmarks; project page: https://cvlab-kaist.github.io/MVTrack4Gen/, arXiv: https://arxiv.org/abs/2606.26087v1.
Cleaned source text

Abstract

Comment: Project Page : https://cvlab-kaist.github.io/MVTrack4Gen/