日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2608.22301

模倣者ゲーム:行動予測を超えたロボットの模倣能力のベンチマーク

The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

シェア:XThreadsFacebookLINEはてブBluesky

人間のデモンストレーションから意図を推論して実行するロボットの能力を評価するため、4段階のベンチマーク「The Imitator Game」と大規模データセット「IG-10K」を構築し、既存モデルが意図レベルの模倣で失敗することを示した。

詳しい要約

1. どんなもの?

The Imitator Gameは、ロボットの意図レベルでの模倣能力を評価するための4段階(L0-L3)のベンチマークである。人間のデモンストレーションとロボットのシーンとのギャップを段階的に広げ、軌道再生が不十分でタスク理解が必要になる点を特定する。また、最大規模の環境整合ペア型ヒューマン・ロボットデータセットIG-10K(20,000以上のペアエピソード、50以上のタスク、6ドメイン、実環境とシミュレーションの両方で全4レベルを実装)と、ブラインドA/B人間評価のためのオープンプラットフォームImitator Arenaを提供する。

2. 先行研究と比べてどこがすごい?

既存のロボットポリシーは視覚入力と言語指示から観測-行動マッピングを学習し、デモンストレーションの意図を明示的に推論しないため、人間のビデオからの学習は軌道レベルに留まり、ほぼ同一シーンでの動作再生はできるが、意図の模倣は困難である。The Imitator Gameは、デモとロボットシーンのギャップを段階的に広げる4レベルを導入し、軌道再生が不十分になる点を特定することで、意図レベルの模倣の評価を可能にする点が新しい。また、IG-10Kは環境整合ペア型データセットとして最大規模であり、全4レベルを実環境とシミュレーションで実装した唯一のデータセットである。

3. 技術・手法の肝は?

手法の肝は、4レベル(L0-L3)のベンチマーク設計にある。L0からL3へと、人間のデモとロボットのシーンとのギャップ(オブジェクトの配置、ツール、レイアウトなど)を段階的に広げ、特にL3では機能代替(異なるオブジェクトのアフォーダンスを通じて同じ意図を達成)を必要とする。これにより、軌道再生が機能しなくなり、タスク理解が必要になる点を隔離する。また、IG-10Kデータセットは環境整合ペア型で、人間とロボットのデモを同一環境で収集し、全レベルで実環境とシミュレーションを提供する。評価にはImitator ArenaでのブラインドA/B人間評価を用いる。

4. どうやって有効だと検証した?

9つの最先端モデルを評価し、L0からL2までは性能が安定しているが、L3で性能が急落することを示した。これにより、機能代替が意図レベル模倣の決定的な障壁であると特定した。また、人間ビデオ条件付けモデルがキャプション条件付けモデルよりも優れているが、未見タスクでのゼロショット成功率は全モデルで13%未満であることを示した。さらに、IG-10Kで事前学習したモデルを、わずか10ペアの人間-ロボットデモでファインチューニングすると、大きな改善が見られ、その改善は事前学習のスケールに応じて増大することを検証した。

5. 議論はある?

議論としては、L3での性能崩壊が機能代替の難しさを示しており、現在のモデルは意図レベルの模倣に必要なタスク理解を欠いていることが示唆される。また、人間ビデオ条件付けがキャプション条件付けより優れている点は、視覚的手がかりの重要性を示すが、それでも未見タスクでの成功率が低いことから、汎化にはさらなる研究が必要である。ファインチューニングの効果が事前学習スケールに依存することは、スケール則の可能性を示唆するが、10ペアという少数での改善がどの程度持続するかは不明である。

6. 次に読むべき論文は?

要旨で参照されている研究は、人間のビデオからの学習、軌道レベルの模倣、キャプション条件付けモデル、人間ビデオ条件付けモデルなどである。具体的な論文名は挙げられていないが、関連する分野の定番として、Learning from Demonstration (LfD)、Behavior Cloning、Video-conditioned Policy Learning、Task Understanding in Roboticsなどの研究が考えられる。次に読むべき論文としては、これらの分野の代表的な論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xunzhe Zhou, Yiyang Cai, Fengyi Wang, Ran Ju, Hanxiang Ren, Ruizhe Liu, Yu Zhang, Qian Luo, Feng Chen, Pei Zhou, Yi Ma, Yanchao Yang

分類: cs.RO, cs.AI

原文アブストラクト

Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and language instructions, without explicitly inferring the demonstrated task. Learning from human video thus remains largely trajectory-level: models can replay motions in near-identical scenes, but still struggle to imitate what the demonstrator intends rather than merely what they do. We introduce The Imitator Game, a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, isolating where trajectory replay ceases to suffice and task understanding becomes necessary. We pair it with IG-10K, the largest environment-aligned paired human-robot dataset to date and the only one instantiated across all four levels in both real and simulated settings (20,000+ paired episodes, 50+ tasks, 6 domains), and Imitator Arena, an open platform for blind A/B human evaluation. Across nine state-of-the-art models, performance is stable from L0 to L2 but collapses at L3, identifying functional substitution - achieving the same intent through a different object affordance - as the decisive barrier to intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet every model falls below 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models with only $10$ paired human-robot demonstrations yields large gains that grow with pretraining scale. The project website and access to Imitator Arena are available at https://imitator-game.github.io.

関連論文