日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04607

ForeAct3D: VLAポリシーのためのポリシー接地型未来世界モデリング

ForeAct3D: Policy-Grounded Future World Modeling for VLA Policies

シェア:XThreadsFacebookLINEはてブBluesky

VLAポリシー内で、行動に条件付けられた未来の3Dシーン状態を予測し、物理整合性で拘束することで、操作性能を向上させるフレームワークを提案した。

詳しい要約

1. どんなもの?

- VLA policies のための policy-grounded future world modeling フレームワーク ForeAct3D を提案 - 未来の観測予測を、policy が実際に実行する action と結びつける - 学習可能な geometric queries が depth, semantic segmentation, camera pose を policy 表現から復号 - 現在と未来の semantic 3D scene states を生成 - 未来の query は policy が生成した action chunk に条件付けられる - 推論時には未来予測を必要としない

2. 先行研究と比べてどこがすごい?

- 既存の VLA policies は共有特徴から未来観測を予測するが、予測が実行 action と切り離されている - また scene の進化に物理的制約を課していない - ForeAct3D は予測を planned interaction に接地し、物理的一貫性を導入 - robot pretraining なしで LIBERO 98.3% 平均成功、CALVIN 平均タスク長 3.73 を達成 - 全 suite で base policy を上回る - 実世界実験で平均成功率を 6.7% から 37.8% に向上

3. 技術・手法の肝は?

- 学習可能な geometric queries が depth, semantic segmentation, camera pose を復号 - 現在と未来の semantic 3D scene states を policy 表現から生成 - 未来の query を policy 生成 action chunk に条件付け - physical-consistency closure が background staticity と instance-level rigidity で2状態を関連付け - wrist-camera pose を end-effector kinematics に固定 - これらの目的が action 生成に使われる共有表現を訓練中に形成 - 推論時に未来予測は不要

4. どうやって有効だと検証した?

- LIBERO で 98.3% 平均成功 - CALVIN で平均タスク長 3.73 - 全 suite で base policy を上回る - Ablations で semantic 3D supervision, physical consistency, action conditioning がそれぞれ manipulation 性能を改善 - action conditioning が未来の object localization を大幅改善 - 実世界で spatial placement, object insertion, sequential manipulation を検証 - 平均成功率 6.7% から 37.8% へ向上

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- LIBERO - CALVIN - Vision-Language-Action (VLA) policies - 関連する future world modeling 研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhe Tao, Feiran Wang, Gaowen Liu, Ramana Rao Kompella$, Yan Yan

分類: cs.RO, cs.CV

原文アブストラクト

Robots need to anticipate how their actions will change the world, since manipulation success hinges on the resulting contacts and object motions. However, existing Vision-Language-Action (VLA) policies that predict future observations from shared features leave the forecast decoupled from the actions the policy will actually execute, and impose no physical constraints on how the scene may evolve. We introduce ForeAct3D, a framework for policy-grounded future world modeling within VLA policies. Learnable geometric queries decode depth, semantic segmentation, and camera pose from the policy representation into current and future semantic 3D scene states, and the future queries are conditioned on the policy-generated action chunk to ground the forecast in the planned interaction. A physical-consistency closure relates the two states through background staticity and instance-level rigidity, and anchors the wrist-camera pose to end-effector kinematics. These objectives shape the shared representation used for action generation during training, and no future prediction is required at inference. Without robot pretraining, ForeAct3D achieves 98.3\% average success on LIBERO and an average task length of 3.73 on CALVIN, outperforming its base policy on every suite. Ablations show that semantic 3D supervision, physical consistency, and action conditioning each improve manipulation performance, and that action conditioning substantially improves future object localization. Real-world experiments on spatial placement, object insertion, and sequential manipulation further raise average success from 6.7\% to 37.8\% over the base policy. The project page and code are available at https://github.com/anthonytao80-crypto/ForeAct3D.

関連論文

PR本紙発行元 EmplifAI