ArtifactArena:物理世界で何を作れるかでモデルを評価する
ArtifactArena: Evaluating Models by What They Build in the Physical World
モデルが物理シミュレーション環境でロボットを設計・構築し、対戦トーナメントのEloで能力を評価するオープンなベンチマークを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Kushagra Tiwary*, David Mayo*, Nikhil Behari, Xiangzhou Sun, Abdulrahman Alabdulkareem, Isaac Galatzer-Levy, Boris Katz, Brian Cheung
分類: cs.RO, cs.AI
原文アブストラクト
To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit \href{https://artifactarena.ai}{https://artifactarena.ai} for more information.
関連論文
- 低コストロボットナビゲーションにおける効率的なSim-to-Real転移のためのデュアル変分オートエンコーダsim2real
- SimForcing: シミュレーションの運動事前分布を実世界ロボット世界モデルへ蒸留sim2real
- 人間の動画を物理的に整合したロボット操作データに変換sim2real
- 視覚ベースUAV着陸における二項結果を伴うDNN再学習のためのベイズデータ拡張sim2real
- AffordCraft: 単一画像からタスク対応シミュレーション資産をスケーラブルに構築sim2real
- サンプリングベース外乱オブザーバによるSim-to-Realギャップの克服:解析モデルから学習型世界モデルまでsim2real