ArXiv

Safe Vision Language Action Models via Barrier Enhanced Flow Matching

Authors
Kasra Sinaei, Hung-Chieh Wu, Donald Ebeigbe
Categories
cs.RO, eess.SY
arXiv
https://arxiv.org/abs/2607.29569v1
PDF
https://arxiv.org/pdf/2607.29569v1

Brief

The paper presents a modular inference technique that integrates Flow Matching generative models with Control Barrier Function (CBF) guarantees by modifying the denoising process to enforce a smooth Log-Sum-Exponential aggregate barrier across action chunks. The authors prove a bounded 2-Wasserstein distance to the target distribution, avoid retraining or safety datasets, and validate on two manipulation platforms and a 2D navigation benchmark; full text was not available, summary based on the abstract.

Why it matters

Kasra Sinaei, Hung-Chieh Wu, and Donald Ebeigbe (arXiv:2607.29569v1, 2026-07-31) propose a modular inference framework that gives Flow Matching generative models formal safety guarantees via Control Barrier Functions (CBFs).

Key details

  • The method alters the Flow Matching denoising process (not post-hoc filtering) using a smooth Log-Sum-Exponential aggregate barrier applied over entire action chunks, adding minimal computational overhead while preserving the model's semantic intent.
  • The paper proves the 2-Wasserstein distance between the generated and target distributions remains bounded, requires no safety-specific datasets or retraining, and empirically verifies reliable safety without degrading success rate on two robotic manipulation platforms plus a 2D navigation benchmark.
Source evidence

Abstract

This article presents a modular inference framework that integrates Flow Matching generative models with formal Control Barrier Function (CBF) safety guarantees. Unlike existing methods that apply external safety filters to a model's final output, our approach modifies the Flow Matching denoising process within the model to inherently generate safe trajectories. By employing a smooth Log-Sum-Exponential aggregate barrier, we enforce safety over entire action chunks. This aggregate barrier ensures a minimal increase in computational overhead and does not alter the semantic intent of the model. We show that, within the proposed framework, the 2-Wasserstein distance between the generated distribution and the target distribution remains bounded. Our method eliminates the need for safety-specific datasets or costly model retraining, providing a versatile solution for safe inference. We validate the approach on two robotic manipulation platforms and a 2D navigation benchmark, verifying that our framework achieves reliable safety without degrading the success rate of the model.