日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ヒューマノイド制御arXiv:2609.24840

PredActor: 予測的行動拡散による操縦可能なオンボードヒューマノイド制御

PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control

シェア:XThreadsFacebookLINEはてブBluesky

固有感覚のみを用いて行動と将来状態軌道を同時生成し、テスト時にも目標へ操縦できるヒューマノイド制御ポリシーを提案。

詳しい要約

1. どんなもの?

- 提案手法: PredActor - 予測的行動拡散ポリシー - 人間型ロボットのオンボード制御 - 特徴: - 自己受容感覚のみで直接実行 - 将来状態軌道を内部生成 - テキスト条件付け行動選択 - テスト時目標への誘導 - 応用: - シミュレーションと実機Unitree G1 - テキスト条件運動、外乱応答、ジョイスティック操作、意味補間

2. 先行研究と比べてどこがすごい?

- 先行研究との比較: - 階層型: 参照がトラッカー能力を超える可能性 - 行動のみ拡散: 将来状態軌道なし - 状態行動拡散: 特権的全身体状態依存、行動選択と誘導が断片的 - 優位性: - 自己受容感覚のみで直接実行 - 行動と将来状態軌道を同時生成 - テキスト条件付けとテスト時誘導を統合 - 別トラッカーや外部全身体状態不要 - シミュレーションで全15目標到達、テキスト検索スコア0.580(条件付き行動拡散0.373)

3. 技術・手法の肝は?

- 核心: - 予測的行動拡散ポリシー - 自己受容感覚履歴と任意タスクコンテキストで条件付け - 実行可能行動と内部将来状態軌道を共同生成 - 分類器なしガイダンスでテキスト条件付け強化 - 分類器ガイダンスで予測状態をテスト時目標へ誘導 - 行動のみ実行、別トラッカーや外部全身体状態不要 - オンボード実用化: - ローリングデノイジング - 計算保存ランタイム最適化 - Jetson Orin NXでコールバック中央値16.790ms、p95 19.383ms(制御周期20ms未満)

4. どうやって有効だと検証した?

- 検証方法: - シミュレーション: - 全15目的地到達 - テキスト検索スコア0.580(条件付き行動拡散0.373) - 外乱生存率は同程度 - 実機: - Unitree G1に展開 - シミュレーションと物理ハードウェアで評価 - テキスト条件運動、外乱応答、ジョイスティック制御、意味補間を実証 - 計算性能: - Jetson Orin NXでコールバック中央値16.790ms、p95 19.383ms

5. 議論はある?

- 議論点: - 階層型や行動のみ拡散の限界を克服 - 状態行動拡散の特権状態依存と断片的誘導を改善 - 自己受容感覚のみで直接実行可能 - オンボード計算制約下での実用性 - 外乱生存率は同程度 - 詳細な議論や限界は要旨からは不明

6. 次に読むべき論文は?

- 次に読むべき論文: - 階層型人間型制御 - 行動のみ拡散ポリシー - 状態行動拡散ポリシー - 分類器なしガイダンス - 分類器ガイダンス - ローリングデノイジング - 関連手法: Diffusion Policy, Decision Diffuser, Diffuser

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lei Ye, Haibo Gao, Yitang Li, Peng Xu, Zetong Jing, Junhan Sun, Fanrong Dong, Ziqi Han, Xue Wang, Jianhua Sun, Cewu Lu, Hao Zhao, Liang Ding

分類: cs.RO, cs.LG

原文アブストラクト

Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.

関連論文

PR本紙発行元 EmplifAI