日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.26314

TriWorldBench: 三視点整合性で評価する身体性世界モデル

TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models

シェア:XThreadsFacebookLINEはてブBluesky

頭部と両手首のカメラ映像を同期させた二腕操作タスクのベンチマークを提案し、視点間の整合性や物理的・時間的一貫性を19指標で評価する。

詳しい要約

1. どんなもの?

- 本論文は TRIWORLDBENCH を提案する。 - Embodied world models を評価するための benchmark。 - head, left-wrist, right-wrist の同期動画を用いる。 - 500 episodes, 50 bimanual manipulation tasks を含む。 - 19 metrics で tri-view consistency 等を評価。 - TWB-Score と per-view results を提供。

2. 先行研究と比べてどこがすごい?

- 従来は各視点を独立評価していた。 - それでは同一 action と object state を記述しているか判定できない。 - 本 benchmark は cross-view checks を導入。 - 単一視点の visual quality を超えた評価を可能にする。 - 各カメラ向けの測定と組み合わせる点が新しい。

3. 技術・手法の肝は?

- head, left-wrist, right-wrist の同期動画を使用。 - 19 metrics で以下を評価。 - tri-view consistency - task alignment - physical and 3D coherence - motion quality - temporal consistency - visual quality - cross-view checks と各カメラ特化の測定を統合。 - TWB-Score で全体性能を要約。 - per-view results で失敗箇所を特定。

4. どうやって有効だと検証した?

- 500 episodes, 50 bimanual manipulation tasks で構成。 - 19 metrics を用いて評価。 - tri-view consistency 等を測定。 - TWB-Score と per-view results を算出。 - 個別動画が意図したタスクの一貫した予測か検証。

5. 議論はある?

- 単一視点の visual quality 評価の限界を指摘。 - 複数視点の一貫性評価の重要性を議論。 - 失敗箇所を per-view results で特定可能。 - 具体的な議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明記されていない。 - 関連手法として embodied world models の評価研究が挙げられる。 - 同分野の定番として robot learning や manipulation の world models 研究。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xuanyi Liu, Haofeng Wang, Ruiqi Li, Danni Yu, Rui Wan, Ruixu Zhang, Siyu Tao, Xue Yang, Shaofeng Zhang, Zicheng Zhang, Jiaqi Zhang, Siwei Ma

分類: cs.RO, cs.AI

原文アブストラクト

Embodied world models predict the outcomes of robot actions to support learning and planning. For robots equipped with head and wrist cameras, this requires complementary views: the head view captures the overall task, while wrist views reveal local gripper-object interactions. However, evaluating these views independently cannot determine whether they describe the same action and object state. We introduce TRIWORLDBENCH, a benchmark for evaluating embodied world models through synchronized head, left-wrist, and right-wrist videos. It contains 500 episodes across 50 bimanual manipulation tasks and uses 19 metrics to assess tri-view consistency, task alignment, physical and 3D coherence, motion quality, temporal consistency, and visual quality. By combining cross-view checks with measurements tailored to each camera, the benchmark evaluates whether plausible individual videos also form a consistent prediction of the intended task. We summarize overall performance with TWB-Score and retain per-view results to identify where predictions fail. This extends world-model evaluation beyond single-view visual quality. Code, data, and metric definitions are available at https://github.com/TriWorldBench/TriWorldBench.

関連論文

PR本紙発行元 EmplifAI