日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.09309

予測された未来だけでは不十分:ロボットマニピュレーションのための実行可能な目標の学習

Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

3Dトレース世界モデルから実行可能な終端目標を明示的に出力する学習インターフェースを提案し、5つのマニピュレーションタスクで平均79.69%の成功率を達成した。

詳しい要約

1. どんなもの?

- 3D trace world model から実行可能な終端ゴールを明示的に出力する学習型インターフェース『Entity-Level Goal Readout』を提案 - object-centric pose prediction と観測 depth に基づく translation を組み合わせ、SE(3) のコンパクトなゴールを生成 - 共有された Pose-Native Executor がこの固定ゴールを online object-pose feedback と共に消費し、world model を再実行せず closed-loop control を実現 - 5つの manipulation タスクで平均成功率 79.69% を達成

2. 先行研究と比べてどこがすごい?

- 従来の generative world model は未来予測のみを監督し、terminal goal accuracy が明示的な学習目的になっていなかった - 幾何学的復元が可能でも、制御に必要なコンパクトなタスク変数を直接露出しない問題があった - 本研究は prediction-to-execution interface を world-model planning の明示的な学習コンポーネントとして扱う点が新しい - 制御パイプラインにおける偶発的な後処理ではなく、学習対象として扱うことを主張

3. 技術・手法の肝は?

- Entity-Level Goal Readout: 3D trace world model の出力から実行可能な終端ゴールを明示的に読み出す学習型インターフェース - object-centric pose prediction と観測 depth に基づく translation を統合し、SE(3) のコンパクトなゴールを生成 - Pose-Native Executor: 固定ゴールと online object-pose feedback を消費する共有エグゼキュータ - world model を再実行せずに closed-loop control を実現

4. どうやって有効だと検証した?

- 5つの manipulation タスクで平均成功率 79.69% を達成 - Goal diagnostics により terminal goal accuracy を直接測定 - 制御された translation perturbations でゴール誤差下の実行劣化を特性評価 - Franka arm への zero-shot 展開: nominal StackCube 73.33%、distractors あり 66.67%、policy training 中に未見のターゲットを含む PickPlate 75.00%

5. 議論はある?

- prediction-to-execution interface を world-model planning の明示的な学習コンポーネントとして扱うべきと主張 - 制御パイプラインにおける偶発的な後処理ではないと位置づけ - 要旨からは限界や失敗事例、計算コスト、一般化範囲に関する詳細な議論は不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な先行研究は明示されていない - 関連手法として generative world models、object-centric pose prediction、SE(3) ゴール表現、closed-loop control、Franka arm を用いた manipulation 研究が挙げられる - 同分野の定番として world model ベースの planning や model-based reinforcement learning の文献を読むべき

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tzu-Yu Chuang, Ching-Hsiang Chang, Yi-Hsiu Lee, Yi-Ting Chen, Min Sun, YuanFu Yang

分類: cs.RO

原文アブストラクト

Generative world models provide rich predictions of how manipulation scenes may evolve toward task objectives, yet those futures do not directly expose the compact task variables required by control. When training supervises future prediction alone, terminal goal accuracy is not an explicit learning objective, even when geometric recovery is available. We present Entity-Level Goal Readout, a learned prediction-to-execution interface that makes the executable terminal goal an explicit output of a 3D trace world model. It combines object-centric pose prediction with translation grounded in observed depth to produce a compact goal in SE(3). A shared Pose-Native Executor consumes this fixed goal with online object-pose feedback for closed-loop control without rerunning the world model. Across five manipulation tasks, the pipeline achieves a mean success rate of 79.69%. Goal diagnostics directly measure terminal goal accuracy, while controlled translation perturbations characterize how execution degrades under goal error. Zero-shot deployment on a Franka arm achieves 73.33% success on nominal StackCube, 66.67% with distractors, and 75.00% on PickPlate with a target unseen during policy training. These results support treating the prediction-to-execution interface as an explicit learned component of world-model planning rather than incidental post-processing in the control pipeline itself. Project page: https://claire0730.github.io/executable-goals/

関連論文

PR本紙発行元 EmplifAI