日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.27449

X2Real:実世界汎用ポリシー評価のための拡張可能なシミュレーションベンチマーク

X2Real: an eXtensive simulation benchmark for real-world generalist policies

シェア:XThreadsFacebookLINEはてブBluesky

実機との相関0.84を達成したIsaac Labベースの進化型シミュレーションベンチマークで、10能力次元・44タスクにより汎用マニピュレーションポリシーを忠実・多様・公平に評価する。

詳しい要約

1. どんなもの?

- 汎用ロボットマニピュレーション政策の実世界性能を忠実に評価するための進化可能なシミュレーションベンチマーク。 - Nvidia Isaac Lab-Arenaを基盤とし、faithfulness, diversity, fairnessの3原則に従う。 - 10の能力次元と44の階層的長地平線タスクからなる分類法を特徴とする。 - 視覚的・物理的特性を実機に合わせて調整し、シミュレーションと実機の評価結果間に0.84の線形相関を達成。 - カスタム物理ドメイン特化言語Manaにより、モジュラーなタスク設計と反復性能分析を支援。 - 約300時間の注釈付きシミュレーション軌道データセットを提供。

2. 先行研究と比べてどこがすごい?

- 既存のシミュレーションベンチマークはsim-to-realギャップ、タスク範囲の狭さ、曖昧な訓練-テストパイプラインによる不公平な評価という根本的欠陥を持つ。 - 先行研究はこれらの問題を部分的にしか解決せず、faithfulness, diversity, fairnessを同時に満たすものはない。 - 静的なベンチマーク設計は長期的な政策開発を維持できない。 - X2Realはこれらを同時に解決し、進化可能な評価基盤を提供する点が優れている。

3. 技術・手法の肝は?

- Nvidia Isaac Lab-Arenaを基盤としたシミュレーションベンチマーク。 - シミュレーションの視覚的・物理的特性を実機に合わせてキャリブレーション。 - 10の能力次元と44の階層的長地平線タスクを含む包括的分類法。 - 多軸ドメインランダム化と厳密に分離された訓練-評価パイプラインを採用。 - カスタム物理ドメイン特化言語Manaによるモジュラーなタスク設計と反復性能分析。 - 約300時間の注釈付きシミュレーション軌道データセットを提供。

4. どうやって有効だと検証した?

- シミュレーションと実機の評価結果間に0.84の線形相関を達成したと報告。 - これによりシミュレーション評価の実機性能への忠実性を検証。 - 具体的な検証方法の詳細は要旨からは不明。

5. 議論はある?

- 既存ベンチマークの根本的欠陥(sim-to-realギャップ、タスク範囲の狭さ、不公平な評価)を指摘。 - 先行研究がfaithfulness, diversity, fairnessを同時に解決できていないことを議論。 - 静的なベンチマーク設計の限界を述べ、進化可能な評価基盤の必要性を主張。 - ベンチマークの悪用を防ぐための厳密な訓練-評価分離の重要性を議論。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法としてNvidia Isaac Lab-Arena、Mana simulation ecosystemが挙げられる。 - 同分野の定番として、RoboSuite、Meta-World、RLBench、CALVINなどのシミュレーションベンチマークが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lian Ruan, Jade Yang, Sherphylan Gao, Felix Gao, Kyson Liang, Galen Liu, Ligo Wu, Lane Jin, Guu Gu, Bevan Xie, Cloud Yan, Zongzi Yuan, Kino Luo, Emma Chen, Shuwen Chen, Yang Ping, Miles Guo, Rain Sun, Kayden Zhang, Alex Du, Ruihai Wu, Liang Hao, Zhaoshuo Li, Roy Gan, Hao Wang, Qian Wang

分類: cs.RO

原文アブストラクト

Generalist robot manipulation policies have developed rapidly, yet their reliable evaluation remains challenging due to fundamental flaws in existing simulation benchmarks: prominent sim-to-real gaps, narrow task coverage, and unfair evaluation caused by ambiguous training-test pipelines. Prior works only partially resolve these issues and lack simultaneous faithfulness, diversity, and fairness, while static benchmark designs fail to sustain long-term policy development. We present X2Real, an evolvable simulation benchmark for faithfully evaluating the real-world performance of robotic manipulation policies based on Nvidia Isaac Lab-Arena. Following three core principles (faithfulness, diversity, and fairness), X2Real calibrates simulation visual and physical properties to align with real hardware, achieving a 0.84 linear correlation between simulated and real-robot evaluation results. It features a comprehensive taxonomy with 10 capability dimensions and 44 hierarchical long-horizon tasks, covering basic manipulation skills and advanced capacities such as visual grounding, language understanding, and bimanual control. We further adopt multi-axis domain randomization and strictly disjoint training-evaluation pipelines to mitigate benchmark exploitation and ensure credible evaluation. Powered by a custom physical domain-specific language, the Mana simulation ecosystem supports modular task design and iterative performance analysis, alongside a nearly 300-hour annotated simulation trajectory dataset. X2Real offers a faithful, diverse, and fair evolving evaluation infrastructure, effectively bridging the sim-to-real evaluation gap and supporting the advancement of generalist robotic manipulation policies.

関連論文

PR本紙発行元 EmplifAI