日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.01083

WBAG: 視覚言語行動マニピュレーションのための全身・把持物体ジオメトリ安全フレームワーク

WBAG: A Whole-Body and Attached-Geometry Safety Framework for Vision-Language-Action Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

VLAポリシーの実行時に、ロボット全身と把持物体の形状を考慮した安全制約を微分可能CBFとして組み込み、衝突を回避しながらタスクを遂行する枠組みを提案。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) ポリシーの実世界展開における衝突回避を目的とした安全フレームワーク WBAG を提案。 - ロボットの whole-body と把持依存の attached geometry をモデル化する。 - 把持条件付き安全集合を構築し、微分可能な CBF 制約に変換する。 - VLA の native 6次元 operational-space action を最小限修正して衝突回避を実現。

2. 先行研究と比べてどこがすごい?

- 既存の inference-time VLA 安全フレームワークは simplified end-effector-centered 表現に依存。 - それらは full articulated robot や attached-object geometry を明示的にモデル化しない。 - WBAG は whole-body と grasp-dependent attached geometry を明示的に扱う点で異なる。 - 把持に応じて保護形状を適応させる grasp-conditioned safe set を構築する。

3. 技術・手法の肝は?

- 把持条件付き安全集合を構築し、物体把持に応じて保護形状を適応。 - その進化する形状を微分可能な CBF 制約に変換。 - 制約は VLA の native 6次元 operational-space action を最小限修正。 - ロボット、シーン、attached geometry 全体の衝突回避を統合的に扱う。

4. どうやって有効だと検証した?

- SafeLIBERO ベンチマークで評価。 - SafeLIBERO は LIBERO に障害物を追加した安全評価用変種。 - scene-level safety evaluator が全 eligible non-task objects を監視。 - 評価手法中で最高の総合安全性と safe task success を達成。 - 97.38% aggregate Scene Safety、59.38% Safe Success を記録。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、一般化性に関する議論は記述されていない。

6. 次に読むべき論文は?

- SafeLIBERO ベンチマーク(LIBERO の安全評価変種)。 - VLA ポリシー(Vision-Language-Action policies)。 - CBF (Control Barrier Function) を用いた安全制約。 - inference-time VLA safety frameworks。 - これらは要旨で参照・比較されている。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Samuel Zhen, Siwon Jo, Yanze Zhang, Wenhao Luo

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) policies have demonstrated impressive capabilities in generalizable robotic manipulation, but their deployment in the real world remains challenging due to potential collisions involving different parts of the robot, manipulated objects, and the surrounding environment. Existing inference-time VLA safety frameworks typically rely on simplified end-effector-centered representations that do not explicitly model the full articulated robot and attached-object geometry. In this paper, we present WBAG, a safety framework that models the robot's whole-body and grasp-dependent attached geometry. WBAG constructs a grasp-conditioned safe set that adapts the protected geometry as objects are grasped, then converts this evolving geometry into differentiable CBF constraints that minimally modify the VLA's native six-dimensional operational-space action for collision avoidance across robot, scene, and attached geometry. On the SafeLIBERO benchmark, a variant of LIBERO augmented with obstacles for safety evaluation, WBAG achieves the best overall safety and safe task success among the evaluated methods under a scene-level safety evaluator that monitors all eligible non-task objects, reaching 97.38\% aggregate Scene Safety and 59.38\% Safe Success.

関連論文

PR本紙発行元 EmplifAI