日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.04432

Video2World:身体性動画から対話型世界モデルを構築するコーディングエージェントのベンチマーク

Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos

シェア:XThreadsFacebookLINEはてブBluesky

ロボットや人間の実演動画から、コーディングエージェントが自動でシミュレーション環境とロボット行動を構築できるかを評価するベンチマークVideo2Worldを提案し、最先端エージェントの性能と限界を分析した。

詳しい要約

1. どんなもの?

- 実世界の観測から対話型シミュレータを構築する研究 - 自律的な video-to-simulation をソフトウェアエンジニアリング課題として定式化 - エージェントが embodied video を観察し、シミュレーション環境とロボット挙動を構築 - 実行フィードバックで反復的に改善 - 評価用ベンチマーク Video2World を導入 - 189 本のロボット・人間デモ動画から得た 222 の再構成インスタンス - 幾何忠実度、動的忠実度、機能的正当性で評価

2. 先行研究と比べてどこがすごい?

- 従来のパイプラインは手動の環境構築とキャリブレーションに大きく依存 - 本研究は frontier foundation models と coding agents による end-to-end 自動化の可否を検討 - 9 つの frontier coding-agent システムを評価 - Claude Opus 5 以降で Task success が 5% 未満から 15% 超へ急改善 - ただし人間支援の再構成との間には依然大きなギャップ - 見た目が良い world が必ずしも機能するとは限らない点を発見

3. 技術・手法の肝は?

- autonomous video-to-simulation をソフトウェアエンジニアリングタスクとして定式化 - エージェントが embodied video を観察 - 対応する simulated environment と robot behavior を構築 - 実行フィードバックを通じて反復的に refine - Video2World は再構成 world を次の軸で測定 - geometric fidelity - dynamic fidelity - functional correctness - これらは spatial perception、physical reasoning、executable interaction を捉える

4. どうやって有効だと検証した?

- Video2World ベンチマークで 9 つの frontier coding-agent システムを評価 - 189 本のロボット・人間デモ動画から 222 の再構成インスタンスを使用 - Task success を指標に評価 - Claude Opus 5 以降で 5% 未満から 15% 超への改善を確認 - 人間支援の再構成との比較も実施 - 視覚的忠実度とタスク成功率の関係を分析

5. 議論はある?

- 視覚的忠実度が高い world が必ずしもタスク成功率が高いとは限らない - 知覚的リアリズムと事実的正しさの間のギャップを示唆 - 生成モデルで観察される同様のギャップと一致 - 人間支援の再構成との間には依然として大きな隔たり - 自動化の限界と今後の改善余地が議論の対象

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない - 関連手法として frontier foundation models、coding agents、generative models が挙げられる - 同分野の定番として embodied AI、video-to-simulation、interactive world modeling に関する研究が次に読むべき候補

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jinzhou Tang, Zijun Zhang, Jing Yang, Yuchen Yan, Kun Zhou, Lingjun Mao, Ruobing Han, Jinglin Cao, Wenpeng Xu, Lukun He, Minghao Fu, Fan Feng, Biwei Huang

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5\% to over 15\%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.

関連論文

PR本紙発行元 EmplifAI