日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12026

関節状態と行動を同時生成するヒューマノイド・ワールドアクションモデル

Humanoid World Action Model With Joint State--Action Generation

シェア:XThreadsFacebookLINEはてブBluesky

ヒューマノイドの参照行動と実際に実行される身体状態を同時に生成・予測することで、行動と実行のギャップを埋めるWorld Action Model「HWAM」を提案。

詳しい要約

1. どんなもの?

- ヒューマノイドロボットの汎用マニピュレーション向けに、HWAM (Humanoid World Action Model) を提案。 - 階層型システムにおける action-execution gap を解決するため、実行後の固有受容状態 (post-execution proprioceptive state) を明示的な予測対象とする。 - 参照行動と実現された身体状態を同時生成する joint state-action generation を特徴とする。

2. 先行研究と比べてどこがすごい?

- 従来の VLA や WAM は参照行動を出力するが、whole-body control 等により実行結果と乖離する action-execution gap があった。 - HWAM は実行後の身体状態を予測に組み込むことで、参照行動と物理的結果の関連付けを改善。 - 実機タスクで最高成功率を達成し、Candy Picking では 70.6% 対 Fast-WAM の 43.3%。

3. 技術・手法の肝は?

- 三つの条件付きパスで学習:Policy path は現在の観測のみから state-action 軌道を joint denoising。 - Forward Dynamics Modeling (FDM) は行動と実行後状態から将来の視覚観測を予測。 - Inverse Dynamics Modeling (IDM) は視覚遷移から joint trajectory を再構成。 - これらが policy references、realized body motion、visual outcomes を結びつける。

4. どうやって有効だと検証した?

- LimX OLI ヒューマノイド上で 3 つの実ロボットタスクを評価。 - 評価ベースラインの中で最高成功率を達成。 - Candy Picking で 70.6% の成功率(Fast-WAM は 43.3%)。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Fast-WAM (比較対象として言及) - Vision-Language-Action (VLA) policies - World Action Models (WAMs)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yan Yang, Jikun Rong, Minzhao Zhu, Zheyi Zhao, Qirui Hu, Zihan Lan, Weixin Mao, Yinhao Li, Zhen Fu, Hua Chen

分類: cs.RO, cs.AI

原文アブストラクト

Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action (VLA) policies learn actions directly from multimodal observations, while World Action Models (WAMs) further incorporate future visual prediction to improve action generation. However, in hierarchical humanoid systems, VLA and WAM policies output reference actions that are subsequently realized through whole-body control, robot dynamics, balance, and contact. This hierarchy creates an action--execution gap: the reference produced by the policy can differ from the motion realized by the robot. Without explicitly modeling the realized body state, future visual prediction must jointly explain scene evolution and discrepancies between reference actions and executed motion, making it difficult to associate an action with its physical outcome. We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. By jointly generating reference actions and their realized body states, HWAM directly incorporates supervision of executed motion into action learning. HWAM is trained through three complementary conditional paths. The Policy path jointly denoises state--action trajectories conditioned only on current observations, matching deployment conditions. Forward Dynamics Modeling (FDM) predicts future visual observations conditioned on actions and post-execution states, while Inverse Dynamics Modeling (IDM) reconstructs the joint trajectory from visual transitions. Together, these paths connect policy references, realized body motion, and visual outcomes. HWAM achieves the highest success rate among evaluated baselines on three real-robot tasks on the LimX OLI humanoid. On Candy Picking, HWAM achieves a 70.6% success rate, compared with 43.3% for Fast-WAM.

関連論文

PR本紙発行元 EmplifAI