Inspect Robots:身体性AIの能力と安全性を評価する
Inspect Robots: Evaluating the Capabilities and Safety of Embodied AI
身体性エージェントの評価を開発・実行するためのモジュール式オープンソースフレームワーク「Inspect Robots」を提案し、最先端言語モデルに基づく6つのポリシーの能力と安全性を評価した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Christopher Leet, Achu Menon, Sravanthi Machcha, Sabrina Zou, Aayushya Patel, Aditya Kumar Singh, Anish Kr Singh, Galaba Vamsi, Javin Ahuja, Sai Asish Yamani, Tushar Anand, Vedang Alle, Zihan Jack Zhang, Tzu Kit Chan, Jay Chooi
分類: cs.RO
原文アブストラクト
General purpose language models are increasingly able to control robotic hardware. Understanding the capabilities and safety of these models when embodied is therefore increasingly important for understanding their societal impact and risks. To this end, we introduce Inspect Robots, a modular, open-source framework for developing and running evaluations of embodied agents. Inspect Robots pairs customizable, reusable abstractions for specifying physical evaluations and analyzing their results with infrastructure that automates evaluation setup, execution and termination. We demonstrate Inspect Robots by using it to evaluate the capabilities and safety of six policies based on frontier language models. Inspect Robots has seen significant early uptake, receiving nearly 100,000 downloads in the three months since its release.
関連論文
- DeepInsight II: ベンチマークからロボットへの一つのトレース評価フレームワーク
- RoboPlayground: 構造化物理ドメインによるロボット評価の民主化評価フレームワーク
- AGI-Elo:タスクを極めるまであとどれくらい?評価フレームワーク