日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデル評価arXiv:2608.28718v1

RoboPhys-3D: 3D再構成による包括的な具現化世界モデル評価

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

シェア:XThreadsFacebookLINEはてブBluesky

ロボット操作タスクのビデオ世界モデルを3D再構成に基づいて評価するベンチマークを提案し、生成ビデオの3Dシーン整合性とタスク実行可能性を測定する。

詳しい要約

1. どんなもの?

RoboPhys-3Dは、3D再構成に基づく包括的な具現化世界モデル(EWM)評価ベンチマークである。RoboTwin 2.0上に構築され、50の操作タスク、4つのレジーム、5,000エピソード、25,000の多視点ground-truthビデオを含む。生成ビデオとground-truthビデオを同一の3D再構成パイプラインで処理し、再構成起因の誤差と生成起因の誤差を区別する。50の補完的メトリクスを18のサブ次元、4つのレベル(ピクセル忠実度、3D幾何整合性、状態理解、タスク完全性)に整理する。さらに、全50メトリクスの階層的平均であるAverage Full Scoreと、タスク成功と最も相関するメトリクスの平均であるRoboPhyscoreを導入する。

2. 先行研究と比べてどこがすごい?

従来のEWMベンチマークは3D基盤の統一プロトコルを欠き、生成ロールアウトが3Dシーン状態を保持するか、実行可能なアクションに変換されるかを評価できなかった。RoboPhys-3Dは、生成ビデオとground-truthビデオを同一の3D再構成パイプラインで処理することで、再構成誤差と生成誤差を分離し、3D幾何整合性や状態理解、タスク完全性を評価する点が革新的である。また、タスク成功と相関するメトリクスに基づくRoboPhyscoreを導入し、人間評価との強い一致を示す。

3. 技術・手法の肝は?

手法の肝は、生成ビデオとground-truthビデオを同一の3D再構成パイプラインに通すことで、再構成起因の誤差を分離し、生成起因の誤差を正確に評価できる点にある。ベンチマークは50のメトリクスを4レベル(ピクセル忠実度、3D幾何整合性、状態理解、タスク完全性)に階層化し、18のサブ次元に整理する。さらに、全メトリクスの平均であるAverage Full Scoreと、タスク成功と最も相関するメトリクスを選択して平均するRoboPhyscoreを導入する。

4. どうやって有効だと検証した?

4つの代表的なビデオ世界モデル(Cosmos 3を含む)を評価し、Cosmos 3が最高のRoboPhyscore(0.6330、ground truthの92.7%)を達成した。状態・実行基盤のメトリクスが、知覚的・VLMベースの判断では捉えられない重大な失敗を明らかにした。また、RoboPhyscoreは人間評価と強い一致を示した(Pearson r = 0.9761、Spearman ρ = 0.8962)。

5. 議論はある?

要旨からは、RoboPhys-3Dが知覚的・VLMベースの評価では捉えられない失敗を検出できることが示唆されるが、具体的な議論や限界については不明。また、RoboPhyscoreがタスク成功と強く相関する一方で、他のメトリクスとの関係や、異なるタスクや環境への一般化については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照されているRoboTwin 2.0、および比較対象のビデオ世界モデル(Cosmos 3を含む)に関する論文が次に読むべきである。また、関連する具現化世界モデルの評価手法や3D再構成技術の論文も有用である。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel

分類: cs.RO, cs.AI, cs.CV, cs.ET, eess.SY

原文アブストラクト

Video world models increasingly serve as data engines, action planners, and simulators for embodied AI, but conventional embodied world model (EWM) benchmarks lack a unified 3D-grounded protocol for establishing whether generated rollouts preserve the underlying 3D scene state or translate into executable actions. We introduce RoboPhys-3D, a 3D-grounded EWM benchmark built on RoboTwin 2.0, covering 50 manipulation tasks across four regimes, with 5,000 episodes and 25,000 multi-view ground-truth videos. A defining feature of RoboPhys-3D is that generated and ground-truth videos are processed through the same 3D reconstruction pipeline, enabling reconstruction-induced error to be distinguished from generation-induced error. The RoboPhys-3D benchmark organizes 50 complementary metrics into 18 sub-dimensions across four levels: pixel-level fidelity, 3D geometry consistency, state-level understanding, and task-level completeness. We further introduce Average Full Score, a hierarchical score averaging all 50 metrics for comprehensive evaluation, and RoboPhyscore, a compact task-aligned score averaging the metrics most strongly correlated with task success. Among the four representative video world models, Cosmos 3 achieves the highest RoboPhyscore (0.6330, 92.7% of ground truth), while state- and execution-grounded metrics reveal substantial failures that perceptual and vision-language model-based judgments fail to capture. RoboPhyscore further exhibits strong agreement with human evaluation (Pearson r = 0.9761 and Spearman \r{ho} = 0.8962), demonstrating the importance of grounded, execution-aware evaluation for EWM capability.

関連論文