バリア強化フローマッチングによる安全な視覚言語行動モデル
Safe Vision Language Action Models via Barrier Enhanced Flow Matching
フローマッチング生成モデルに制御バリア関数の安全性保証を統合し、モデル内部のノイズ除去過程を変更することで安全な軌道を生成する手法を提案。外部フィルタ不要で、操作タスクとナビゲーションで安全性と成功率を両立。
著者: Kasra Sinaei, Hung-Chieh Wu, Donald Ebeigbe
分類: cs.RO, eess.SY
原文アブストラクト
This article presents a modular inference framework that integrates Flow Matching generative models with formal Control Barrier Function (CBF) safety guarantees. Unlike existing methods that apply external safety filters to a model's final output, our approach modifies the Flow Matching denoising process within the model to inherently generate safe trajectories. By employing a smooth Log-Sum-Exponential aggregate barrier, we enforce safety over entire action chunks. This aggregate barrier ensures a minimal increase in computational overhead and does not alter the semantic intent of the model. We show that, within the proposed framework, the 2-Wasserstein distance between the generated distribution and the target distribution remains bounded. Our method eliminates the need for safety-specific datasets or costly model retraining, providing a versatile solution for safe inference. We validate the approach on two robotic manipulation platforms and a 2D navigation benchmark, verifying that our framework achieves reliable safety without degrading the success rate of the model.