ArXiv

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

Authors
Chenghao Gu, Hanyang Yu, Jingbo Zhang...
Categories
cs.RO
arXiv
https://arxiv.org/abs/2608.06332v1
PDF
https://arxiv.org/pdf/2608.06332v1

Brief

GeniWorld is an interactive world model for robotic manipulation that uses URDF-based rendering and pretrained video generative models to convert numeric actions into spatially grounded visual action representations. By decoupling robot kinematics from environmental dynamics and using an autoregressive video-prediction model with high-frequency kinematic control, it achieves strong in-domain performance and zero-shot generalization from limited fixed-scene training, and serves as a scalable policy evaluator. (Abstract only.)

Why it matters

GeniWorld achieves robust zero-shot generalization to highly randomized, unseen environments despite being trained solely on limited fixed-scene data (Gu et al., 2026; arXiv:2608.06332v1; posted 2026-08-06).

Key details

  • The method converts numerical robot actions into spatially grounded visual-action inputs via URDF-based rendering and leverages pretrained video generative models while explicitly decoupling embodiment kinematics from environmental dynamics to mitigate scene overfitting.
  • GeniWorld implements an autoregressive video-prediction model integrated with high-frequency robot kinematic control for closed-loop interaction, enabling scalable policy evaluation and generation of diverse manipulation trajectories that improve downstream policy performance and robustness with limited real-world demonstrations.
Source evidence

Abstract

Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, but they often suffer from limited action controllability and poor generalization to out-of-distribution (OOD) scenarios. To this end, we present GeniWorld, an interactive world model for robots that generalizes robustly across unseen scenarios. Building on pretrained video generative models, we use URDF-based rendering to transform numerical actions into visual action representations, enabling spatially grounded action control. By explicitly decoupling embodiment kinematics from environmental dynamics, our model mitigates scene overfitting and facilitates modeling of robot-environment interactions. To achieve closed-loop control, we construct an autoregressive video prediction model integrated with high-frequency robot kinematic control, enabling interaction with both robot policies and human teleoperators. In our experiments, even when trained solely on limited fixed-scene data, our model achieves superior in-domain performance and robust zero-shot generalization to highly randomized, unseen environments. For downstream applications, GeniWorld serves as a scalable policy evaluator that remains reliable under environmental perturbations. Furthermore, even with limited real-world demonstrations, GeniWorld generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in complex environments.