日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.06511

ArtifactArena:物理世界で何を作れるかでモデルを評価する

ArtifactArena: Evaluating Models by What They Build in the Physical World

シェア:XThreadsFacebookLINEはてブBluesky

モデルが物理シミュレーション環境でロボットを設計・構築し、対戦トーナメントのEloで能力を評価するオープンなベンチマークを提案。

詳しい要約

1. どんなもの?

- 物理世界でモデルが何を構築できるかを評価するオープンエンドなプラットフォーム - モデルはハードウェア・ソフトウェアの協調設計課題に挑戦し、シミュレーションアリーナで競う完全機能のロボットを設計 - 3つのハーネスでzero-shot、verifier guided refinement、open-ended physical design能力を評価 - モデルが生成したbot artifactsをEloランキングでベンチマーク - コミュニティ投稿を受け付けるlivingで非飽和なテストベッドを提供

2. 先行研究と比べてどこがすごい?

- 従来の評価はモデルの発言やテキスト出力に基づくが、本手法は物理環境で構築・設計する能力を測定 - 物理的に接地されたハードウェア・ソフトウェア協調設計課題を課す点が新しい - テキスト記述、物理シミュレータフィードバック、ゲームプレイデータに基づく改良を組み合わせる - 対戦トーナメントによるEloランキングでフロンティアモデルを比較 - 継続的なコミュニティ投稿を受け付ける非飽和テストベッドを確立

3. 技術・手法の肝は?

- 物理的に接地されたハードウェア・ソフトウェア協調設計チャレンジを提供 - モデルはシミュレーションアリーナで競う完全機能ロボットを設計 - 3つのハーネス: テキスト記述、物理シミュレータフィードバック、ゲームプレイデータに基づく改良 - zero-shot、verifier guided refinement、open-ended physical design能力を評価 - 生成されたbot artifacts間の対戦トーナメントからEloランキングを導出

4. どうやって有効だと検証した?

- フロンティアモデルのzero-shot、verifier guided refinement、open-ended physical design能力を3つのハーネスで評価 - モデルが生成したartifacts間のhead-to-headトーナメントを実施 - トーナメント結果からEloランキングを算出しベンチマーク - コミュニティ投稿を受け付けるトーナメントインフラを公開し、継続的評価を可能に

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 同分野の関連手法として、物理シミュレーション環境でのロボット設計評価、ハードウェア・ソフトウェア協調設計、オープンエンド進化、Eloレーティングを用いたモデル評価などが挙げられる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kushagra Tiwary*, David Mayo*, Nikhil Behari, Xiangzhou Sun, Abdulrahman Alabdulkareem, Isaac Galatzer-Levy, Boris Katz, Brian Cheung

分類: cs.RO, cs.AI

原文アブストラクト

To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit \href{https://artifactarena.ai}{https://artifactarena.ai} for more information.

関連論文

PR本紙発行元 EmplifAI