ArXiv

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Authors
Yukang Cao, Haozhe Xie, Beichen Wen...
Categories
cs.CV
arXiv
https://arxiv.org/abs/2607.28625v1
PDF
https://arxiv.org/pdf/2607.28625v1

Brief

ACE introduces a human-centric Ambient Capture Engine that converts real homes into calibrated, synchronized studios at table and room scales to record egocentric/exocentric video, body and hand kinematics, object geometry and 6-DoF trajectories, audio, and touch. ACE-Data-0 (150h, 17M frames, 200 tasks, 75k episodes) exposes gaps in current models under contact, occlusion, egomotion, and long horizons, providing a scalable foundation for imitation learning, world models, and vision–language–action systems.

Why it matters

ACE-Data-0 is a multisensory human-centric dataset built with the Ambient Capture Engine: 150 hours, 17M video frames, 200 task categories, 50 participants, 2 environments, and ~75,000 interaction episodes spanning atomic manipulation to long-horizon household activity.

Key details

  • The ACE capture system operates at table-scale and room-scale and records synchronized egocentric + multi-view exocentric video, full-body and articulated-hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals; a hierarchical benchmark shows current SOTA methods struggle under contact, occlusion, egomotion, and long temporal horizons.
Source evidence

Abstract

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

Comment: Project Page: https://ace-data-engine.github.io/ACE-Data-0/