日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.02368

世界行動モデルを再考する:構成的・文脈内ロボットマニピュレーションに向けて

Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

操作タスクを視覚サブゴール計画と実行に階層化し、事前学習済み世界モデル表現を共有するViGARを提案。長期構成的タスクと文脈内学習で高い成功率を示した。

詳しい要約

1. どんなもの?

- 長期的な compositional manipulation のための階層的フレームワーク ViGAR を提案。 - 操作を visual subgoal planner と subgoal executor に分解。 - planner は現在の観測と global instruction から次の subtask の visual subgoal を予測。 - executor は予測された subgoal を条件に future visual trajectories と actions を同時生成。 - 両コンポーネントは pretrained world-model representation を共有。 - in-context learning を自然にサポートし、global goal image を context として与えるとパラメータ更新なしで異なる subtask 分解や行動を誘導できる。

2. 先行研究と比べてどこがすごい?

- 既存の world-action models (WAMs) は short-horizon の visual futures と actions を jointly 予測するが、subtask レベルの明示的推論を欠く。 - ViGAR は階層化により subtask レベルの planning と action generation を分離。 - 両者で pretrained world-model representation を共有し、共通の物理知識を活用。 - in-context learning をパラメータ更新なしで実現する点が新しい。 - RoboTwin Clean2Random benchmark で Clean 82.00%、Random 67.02% の成功率を達成し、最強 baseline を平均成功率で 12.86 ポイント上回る。

3. 技術・手法の肝は?

- 階層的フレームワーク ViGAR (Visual Goal-conditioned Action Reasoning)。 - visual subgoal planner: 現在の observation と global instruction から次の subtask の visual subgoal を予測。 - subgoal executor: 予測された subgoal を条件に future visual trajectories と actions を jointly 生成。 - 両コンポーネントが pretrained world-model representation を共有。 - global goal image を context として与える in-context learning をサポート。 - これにより task-level planning と action generation が共通の物理知識から恩恵を受ける。

4. どうやって有効だと検証した?

- RoboTwin Clean2Random benchmark で評価。 - Clean 設定で 82.00%、Random 設定で 67.02% の成功率。 - 最強 baseline を平均成功率で 12.86 ポイント上回る。 - 実世界ロボット実験を 5 つの compositional タスクと 2 つの in-context learning タスクで実施し有効性を確認。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、スケーラビリティに関する議論は要旨に記載なし。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: world-action models (WAMs)、RoboTwin Clean2Random benchmark。 - 関連手法として visual subgoal planning、hierarchical manipulation、in-context learning for robotics が挙げられる。 - 具体的な論文名は要旨に記載なし。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shukai Gong, Xuanran Zhai, Yintianrun Zhang, Ruopeng Cui, Ye Huang, Yiyang Fu, Dexuan Lyu, Chaojie Li, Xinyi Song, Peiwen Lin, Chuang Wang, Mingyuan Jia, Yufan Deng, Jiaxin Fang, Bo Liang, Jiaxin Li, Yuxiang Gao, Hao Liu, Daquan Zhou

分類: cs.RO

原文アブストラクト

Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We propose Visual Goal-conditioned Action Reasoning (ViGAR), a hierarchical framework that factorizes manipulation into a visual subgoal planner and a subgoal executor. Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal. Both components share a pretrained world-model representation, enabling task-level planning and action generation to benefit from common physical knowledge. Moreover, our framework naturally supports in-context learning: using a global goal image as context can induce different subtask decompositions and behaviors without parameter updates. On the RoboTwin Clean2Random benchmark, ViGAR achieves 82.00% and 67.02% success rates under the Clean and Random settings, respectively, surpassing the strongest baseline by 12.86 percentage points in average success rate. Real-world robot experiments on five compositional and two in-context learning tasks further confirm the effectiveness of ViGAR.

関連論文

PR本紙発行元 EmplifAI