ArXiv

Objects as Audio-Visual Modal Sound Fields

Authors
Zisen Shao, Zihao Wei, Derong Jin...
Categories
cs.CV
arXiv
https://arxiv.org/abs/2608.05145v1
PDF
https://arxiv.org/pdf/2608.05145v1

Brief

AV-MSF (Audio-Visual Modal Sound Field) targets object impact-sound modeling, addressing limits of expensive physics simulation and data-hungry learned models by reconstructing a compact, physically grounded modal sound field from multi-view imagery plus only a few impact recordings. It fuses 3D Gaussian Splatting with dense 3D visual features to supply a geometry-aware prior, yields state-of-the-art rendering on two real-world datasets versus physics-based and data-driven baselines, and enables contact localization and sound editing. (Summary based on the provided abstract; full paper not included here.)

Why it matters

AV-MSF (Audio-Visual Modal Sound Field) reconstructs object-level impact-sound fields from multi-view images plus only a few impact sound recordings by combining 3D Gaussian Splatting with dense 3D visual features and compact, physically meaningful modal parameters.

Key details

  • On two real-world datasets the method achieves state-of-the-art impact-sound rendering, outperforming both physics-based simulators and purely data-driven baselines in few-shot reconstruction scenarios.
  • The representation enables downstream tasks such as contact localization and object sound editing; authors Zisen Shao, Zihao Wei, Derong Jin, and Ruohan Gao presented the work (arXiv 2608.05145v1) and it appears at ECCV 2026 (posted 2026-08-05).
Source evidence

Abstract

While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.

Comment: ECCV 2026, Project page: $\href{https://zisenshao.github.io/AV-MSF/}{\text{this https URL}}$