PredActor: 予測的行動拡散による操縦可能なオンボードヒューマノイド制御
PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
固有感覚のみを用いて行動と将来状態軌道を同時生成し、テスト時にも目標へ操縦できるヒューマノイド制御ポリシーを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Lei Ye, Haibo Gao, Yitang Li, Peng Xu, Zetong Jing, Junhan Sun, Fanrong Dong, Ziqi Han, Xue Wang, Jianhua Sun, Cewu Lu, Hao Zhao, Liang Ding
分類: cs.RO, cs.LG
原文アブストラクト
Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.
関連論文
- ResSafe: 残差強化学習によるヒューマノイドの安全フィルタリングヒューマノイド制御
- ViBe: 知覚型ヒューマノイド全身制御のための視覚行動適応ヒューマノイド制御
- GLoRI: グローバル・ローカル参照相互作用によるヒューマノイド移動操作のための閉ループ全身追跡ヒューマノイド制御
- 学習された停止可能性値によるヒューマノイドの安全停止ヒューマノイド制御
- ADAPT: 俊敏な拡散行動事前分布による堅牢で操縦可能なオンライン文章駆動ヒューマノイド制御ヒューマノイド制御
- StableMimic: 人間らしいスムーズな復帰を実現するヒューマノイド動作追跡 - 追跡分布を超えた学習による構造化された転倒後行動ヒューマノイド制御