日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.21223

SafeStage: 視覚言語条件付きロボットマニピュレーションの実行前・実行中・実行後の安全性評価

SafeStage: Evaluating Safety Before, During, and After Vision-Language-Conditioned Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語条件付きロボット操作の安全性を、タスク実行の前・中・後の3段階で評価するベンチマークSafeStageを提案し、VLAポリシーがタスク成功と安全性違反を同時に起こしうることを示した。

詳しい要約

1. どんなもの?

- Vision-language-conditioned robot manipulation の安全性を評価する benchmark「SafeStage」を提案。 - タスク実行の前・中・後の3段階で安全性を診断する lifecycle-structured な枠組み。 - 97個の purpose-built risk scenarios を含む。 - 段階は Initial-State Hazards、Execution-Time Safety、Final-State Hazards。 - タスク成功率と段階別の安全結果を独立に報告する。

2. 先行研究と比べてどこがすごい?

- 既存評価は task success、isolated physical constraints、semantic refusal、realized physical damage に偏る。 - closed-loop manipulation 中にどこで安全が破綻するかの洞察が限られていた。 - SafeStage は安全性を実行前・実行中・実行後に分解し、違反発生の時点を局所化できる。 - task success と safety を分離して評価する点が新しい。

3. 技術・手法の肝は?

- 3段階の risk scenarios を設計: Initial-State Hazards、Execution-Time Safety、Final-State Hazards。 - Execution-Time Safety では unsafe contacts、trajectories、region entries、object interactions を対象。 - Final-State Hazards では nominal task completion 後に残る不安定・不安全状態を対象。 - event-based と state-based の checks で realized interactions を評価。 - 共通の closed-loop protocol で direct-action VLA policies と world-model-based policies を比較。

4. どうやって有効だと検証した?

- representative direct-action Vision-Language-Action (VLA) policies と world-model-based policies を共通 closed-loop protocol で評価。 - nominal task completion が safety violations と頻繁に共存することを示した。 - 異なる policies が3段階で distinct failure profiles を示すことを確認。 - これにより task success と safety の分離評価と違反時点の局所化が有効と検証。

5. 議論はある?

- nominal task completion と safety violations が共存しうる点を指摘。 - policies ごとに failure profiles が異なるため、一律の安全評価では不十分。 - 安全性を段階別に局所化することで診断と改善に資する。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: direct-action Vision-Language-Action (VLA) policies、world-model-based policies。 - 関連手法として Vision-Language-Action (VLA) モデル、world model ベースのロボット政策。 - 同分野の定番として robot manipulation safety、closed-loop evaluation、semantic refusal に関する研究。 - 個別論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinzhu Luo, Qi Zhang, Wei Wang, Wei Jiang

分類: cs.RO

原文アブストラクト

Vision-language-conditioned robot policies integrate perception, language understanding, and control for general-purpose manipulation. However, existing evaluations often focus on task success, isolated physical constraints, semantic refusal, or realized physical damage, providing limited insight into where safety fails during closed-loop manipulation. We introduce SafeStage, a lifecycle-structured benchmark for evaluating manipulation safety before, during, and after task execution. SafeStage contains 97 purpose-built risk scenarios organized into three stages. Initial-State Hazards captures safety-relevant relations that must be resolved before manipulating the target. Execution-Time Safety evaluates unsafe contacts, trajectories, region entries, and object interactions during execution. Final-State Hazards capture unstable or otherwise unsafe conditions remaining after nominal task completion. The benchmark evaluates realized interactions using event-based and state-based checks and reports native task success independently from stage-specific safety outcomes. We evaluate representative direct-action Vision-Language-Action (VLA) policies and policies with world-model-based policies under a common closed-loop protocol. Our results demonstrate that nominal task completion frequently coexists with safety violations and that different policies exhibit distinct failure profiles across the three stages. By separating task success from safety and localizing when violations occur, SafeStage provides a unified diagnostic testbed for evaluating and improving vision-language-conditioned robot manipulation policies.

関連論文

PR本紙発行元 EmplifAI