日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
身体AI安全性arXiv:2610.09294

RT-Safe: 実時間身体環境におけるエージェント安全性のベンチマーク

RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment

シェア:XThreadsFacebookLINEはてブBluesky

実時間制約下で動く都市環境のシミュレーションベンチマークRT-SAFEを提案し、VLMエージェントが高いタスク達成率を示す一方で安全性イベントをほぼ回避できないことを明らかにした。

詳しい要約

1. どんなもの?

- 本論文は RT-SAFE を提案する。 - 実時間制約下で embodied agent の安全性を評価する simulated urban benchmark。 - navigation タスクに moving actors、environmental hazards、traffic rules を組み合わせる。 - inference と action execution の間も world が進展する。 - 8 つの VLM を評価。 - タスク完了率は高いが、安全に完了する episode はほぼない。 - 最も難しい設定では安全に終わる episode は 0.7% のみ。 - 静的評価と実時間評価を比較。 - タスク完了率は 91.3% と 94.1% で同程度。 - 実時間実行では衝突が 12.3 倍に増加。 - 結論として、標準的なタスク成功は安全性の失敗を隠し、decision latency 自体が物理的リスク源になり得る。 - RT-SAFE は offline RL training を支援し、衝突率を大幅に減らしつつ高いタスク完…

2. 先行研究と比べてどこがすごい?

- 先行研究は digital environments での agent safety 評価が中心。 - 本論文は physical world の embodied safety に焦点。 - 失敗が human injury や hardware damage につながる点を重視。 - 実時間制約を明示的に組み込む点が新しい。 - 推論中も world が動き続け、観測時に安全な行動が実行時には危険になり得る。 - 安全性を decision quality と decision latency の両面から評価する枠組みを提示。 - 静的評価と実時間評価を matched で比較し、タスク成功率が同程度でも衝突が 12.3 倍増えることを示した。

3. 技術・手法の肝は?

- simulated urban benchmark の RT-SAFE を構築。 - navigation タスク、moving actors、environmental hazards、traffic rules を統合。 - inference と action execution の間も world が evolve する設計。 - 8 つの VLM を agent として評価。 - 静的評価と実時間評価を matched で比較する実験設計。 - offline RL training を RT-SAFE 上で行い、collision rate と task completion を評価。 - 詳細なアルゴリズムや実装は要旨からは不明。

4. どうやって有効だと検証した?

- 8 つの VLM を RT-SAFE 上で評価。 - 最も難しい設定で安全に完了した episode は 0.7%。 - 静的評価と実時間評価のタスク完了率は 91.3% と 94.1%。 - 実時間実行で衝突が 12.3 倍に増加。 - offline RL training により衝突率を大幅に削減し、高いタスク完了を達成できることを示した。 - これにより、タスク成功率が安全性を反映しないこと、decision latency がリスク源になることを検証。

5. 議論はある?

- 標準的なタスク成功指標は安全性の失敗を隠蔽し得る。 - decision latency 自体が physical risk の源になり得る。 - 実時間 embodied safety には decision quality と decision latency の両方が重要。 - RT-SAFE が offline RL training を支援し、衝突率削減とタスク完了の両立が可能であることを示す。 - 限界や今後の課題は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明記されていない。 - 関連手法として embodied agent safety、real-time decision making、offline RL、VLM-based navigation の文献が挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tianruo Rose Xu, Jiawei Ren, Yichi Yang, Zhaoxu Zheng, Lianhui Qin

分類: cs.AI, cs.LG

原文アブストラクト

Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures can cause human injury and costly hardware damage. Beyond selecting safe actions, embodied agents must also operate under real-time constraints: the physical world does not pause while an agent reasons. As pedestrians move and vehicles approach during inference, an action that appears safe at observation time may become unsafe before execution. Real-time embodied safety therefore depends on both decision quality and decision latency. We introduce RT-SAFE, a simulated urban benchmark for evaluating embodied-agent safety under real-time constraints. RT-SAFE combines navigation tasks with moving actors, environmental hazards, and traffic rules, while allowing the world to evolve throughout inference and action execution. Across eight VLMs, agents achieve high task completion yet almost never complete safely: in the hardest setting, only 0.7% of episodes finish without a safety event. More strikingly, matched static and real-time evaluations yield task completion rates of 91.3% and 94.1%, respectively, while real-time execution increases collisions by $12.3\times$. These results reveal that standard task success can mask substantial safety failures, and that decision latency itself can become a source of physical risk. Finally, we show that RT-SAFE can support offline RL training and substantially reduce collision rates while achieving strong task completion.

PR本紙発行元 EmplifAI