日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.38059

WorldLine: ロボットマニピュレーションのための行動駆動型視覚シミュレーション

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

行動なしのロボット動画1万時間以上から操作ダイナミクスを学習し、10種類以上の実機の行動軌跡で接地することで、異なる機体間で共有可能な行動駆動型視覚シミュレータを構築した研究。

詳しい要約

1. どんなもの?

WorldLineは、ロボット操作のためのaction-drivenなvisual simulator。物理実行前にactionの結果を予測する。 - 目的: 実世界ロボット学習の経験収集・候補行動評価コストを削減 - 構成: 転移可能なdynamics学習とheterogeneousなaction groundingを分離 - 学習データ: action-freeなrobot動画10,000時間以上 + 10種類以上のembodimentにわたるaction trajectory 2,000時間以上 - 応用: policy evaluationとembodied planning

2. 先行研究と比べてどこがすごい?

先行のvideo generation modelやaction-conditioned simulatorの課題を克服。 - 従来video生成: 視覚的もっともらしさを優先し、action追従やrobot-object dynamicsの整合性が不十分 - 従来action-conditioned: 希少でembodiment固有のデータに依存し、制御空間が異なると共有困難 - 本研究: action-free動画からdynamicsを学び、image-space action representationでembodiment間の共有制御interfaceを実現 - 結果: 失敗trajectoryでrobot-mask IoUを最強baseline比+0.1626改善

3. 技術・手法の肝は?

技術の肝は以下。 - 転移可能なdynamics学習とheterogeneous action groundingの分離 - image-space action representationによるembodiment間の共有control interface - multi-viewおよびfailure-enriched trainingとrelational regularizationでinteraction-sensitive predictionを改善 - robot-focused few-step distillationで効率的なcausal rolloutを実現し、action-critical motionを保持

4. どうやって有効だと検証した?

held-outおよびout-of-domain設定で検証。 - 視覚品質とrobot-motion agreementを維持 - 失敗trajectoryでrobot-mask IoUが最強baseline比+0.1626 - RoboTwinとAgiBotでtrajectory成功を平均74%精度で予測(最強baseline比+1ポイント) - RoboTwin訓練・適応なしで、rolloutが直接policy実行比でtask成功を最大21.4ポイント改善

5. 議論はある?

要旨からは不明。 - 限界や失敗事例、計算コスト、汎化範囲の議論は明示されていない - ただし、scalableで効率的なvisual simulatorとしてpolicy evaluationとembodied planningに有用と主張

6. 次に読むべき論文は?

要旨で参照・比較されている研究や関連手法。 - video generation models - action-conditioned simulators - RoboTwin - AgiBot - 同分野の定番としてrobot manipulationのvisual simulatorやworld model関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia, Haoyang Huang, Nan Duan, Jiaya Jia

分類: cs.RO, cs.CV

原文アブストラクト

Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \href{https://zhengsh123.github.io/WorldLine/}{project page}.

関連論文

PR本紙発行元 EmplifAI