ArXiv

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

Authors
Rongxu Cui, Zongzheng Zhang, Jingrui Pang...
Categories
cs.RO
arXiv
https://arxiv.org/abs/2606.23686v1
PDF
https://arxiv.org/pdf/2606.23686v1

Brief

LIBERO-Safety introduces a parametric, procedurally generated benchmark and a keypose-driven pipeline to scale safety-critical robot manipulation data collection. The authors curate 19,664 strictly collision-free demonstrations with heavy domain randomization, then systematically evaluate eight vision-language-action and two embodied foundation models. Results show high-diversity training improves trajectory safety but task success is constrained by poor trajectory synthesis and semantic misalignment, framing future work on safe VLA models. (Accepted to ECCV 2026.)

Why it matters

LIBERO-Safety delivers a parametric safety benchmark plus a novel keypose-driven data-generation pipeline that procedurally creates stochastic, safety-critical scenarios and produced a large dataset of 19,664 strictly collision-free manipulation demonstrations with extensive domain randomization.

Key details

  • A cross-paradigm evaluation of eight VLA and two embodied foundation models revealed a generalization–safety tension: high-diversity training yields safer trajectories, but overall task success is bottlenecked by sub-optimal trajectory synthesis and semantic misalignment.
  • Paper accepted to ECCV 2026 (arXiv 2026-06-22); project page: https://libero-safety.github.io/
Source evidence

Abstract

Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address this, we introduce a parametric safety benchmark to procedurally generate safety-critical scenarios with comprehensive stochasticity. To overcome the scalability bottlenecks of human teleoperation, we develop a novel keypose-driven data generation pipeline. Leveraging this infrastructure, we curate a large-scale dataset of 19,664 strictly collision-free demonstrations with extensive domain randomization. We then conduct a systematic cross-paradigm evaluation of eight VLA and two embodied foundation models. Our analysis reveals a critical generalization-safety tension: although high-diversity training fosters safer trajectories, task success remains fundamentally bottlenecked by sub-optimal trajectory synthesis and semantic misalignment. By providing a scalable pipeline, a robust dataset, and profound failure-mode insights, LIBERO-Safety establishes a crucial foundation for developing safe and reliable VLA models.

Comment: Accepted by ECCV 2026, Project Page: https://libero-safety.github.io/