ArXiv

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Authors
Zhe Li, Zhenzhe Zhang, Yangyang Wei...
Categories
cs.RO
arXiv
https://arxiv.org/abs/2608.06375v1
PDF
https://arxiv.org/pdf/2608.06375v1

Brief

ω-0 is a latent predictive world-action model for concurrent humanoid loco-manipulation that maps language, vision, and proprioception to controller-compatible whole-body action latents. Instead of reconstructing video, it predicts compact future observation embeddings and couples visual foresight with diffusion-based action generation. Trained with controller-based simulation replay and the 40+ hour ω-HOME dataset, it outperforms several baselines on 11 real-world household tasks (authors: Zhe Li et al.; arXiv 2608.06375v1).

Why it matters

ω-0 is a latent predictive whole-body world-action model that, given a language instruction, current visual observation, and robot proprioception, directly predicts controller-compatible whole-body action latents using compact future observation embeddings and diffusion-based action generation; it supports egocentric RGB, exocentric RGB, and exocentric depth inputs.

Key details

  • The authors collected ω-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents; real-robot experiments on 11 household tasks (ArXiv: 2608.06375v1, published 2026-08-06) show a single ω-0 model produces smooth manipulate-while-moving behaviors and consistently outperforms imitation learning, VLA, humanoid, and WAM baselines.
Source evidence

Abstract

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present $ω$-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, $ω$-0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, $ω$-0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect $ω$-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single $ω$-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.