日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2609.21447

FootQuery: 予測接地位置に基づく深度履歴検索による知覚型ヒューマノイド歩行

FootQuery: Future-Touchdown-Guided Retrieval from Depth History for Perceptive Humanoid Locomotion

シェア:XThreadsFacebookLINEはてブBluesky

各足の次回接地位置予測を用いて過去の深度フレームから必要な地形情報を検索し、複雑地形を歩行するヒューマノイド制御を実現した研究。

詳しい要約

1. どんなもの?

- 複雑地形を歩行する humanoid のための知覚・運動フレームワーク FootQuery を提案。 - 各足の次回 touchdown 予測と不確実性から depth history を query し、接地時に見えない地形情報を取得。 - proprioception と onboard depth images のみで展開可能。

2. 先行研究と比べてどこがすごい?

- 従来の知覚歩行は現在の視覚や global visual memory に依存し、自己遮蔽や視野制限で touchdown 時の地形が見えない問題があった。 - FootQuery は未来の接地位置を手がかりに過去の depth frames を検索し、接地時に不可視な領域の情報を補完。 - 単一 policy で屋外階段・屋内経路を連続踏破できる点が先行研究と比べて優れる。

3. 技術・手法の肝は?

- policy が proprioception から touchdown locations と uncertainty を予測。 - その分布と per-foot features を用いて、sparsely sampled historical depth frames を query。 - 訓練時は realized contacts を過去画像に投影し、接触が見えていた領域の retrieval を監督。 - 取得した per-foot features を global visual memory と融合して制御行動を生成。 - progressive force-assistance curriculum と event-consistent tread-midline shaping を併用。

4. どうやって有効だと検証した?

- シミュレーションで最も難しい stairs, gaps, platforms において完全なフレームワークが component ablations を上回る。 - 実世界では Unitree G1 を用い、単一 policy で屋外階段と屋内経路(階段昇降、platforms、gaps)を連続踏破。

5. 議論はある?

- 結果は、知覚 humanoid locomotion において視覚履歴を予期される接地周辺で組織化することの有効性を支持。 - 限界や失敗事例、計算コスト、汎化性に関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている研究は明示されていない。 - 関連手法として、humanoid locomotion における perceptive locomotion、depth-based terrain representation、visual memory、touchdown prediction を扱う研究が次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tao Dong, Jia Yu, Yuxuan Fan, Linna Zhao, Jiaqi Gong, Andong Yang, Chao Gao, Guyue Zhou

分類: cs.RO, cs.LG, eess.SY

原文アブストラクト

Humanoid locomotion over complex terrain requires anticipating footholds that may no longer be visible at touchdown. Limited camera coverage and self-occlusion make it necessary to retrieve relevant terrain information from earlier observations. We present FootQuery, a perceptive locomotion framework that queries depth history using each foot's predicted next touchdown. The policy predicts touchdown locations and uncertainty from proprioception and uses these distributions, together with per-foot features, to query sparsely sampled historical depth frames. During training, realized contacts are projected into historical images to supervise retrieval at the regions where those contacts were visible. The retrieved per-foot features are fused with global visual memory to generate control actions. A progressive force-assistance curriculum supports early exploration, while event-consistent tread-midline shaping encourages coordinated stair contacts. Deployment requires only proprioception and onboard depth images. In simulation, the complete framework outperforms its component ablations on the most challenging tested stairs, gaps, and platforms. Real-world experiments on a Unitree G1 demonstrate continuous traversal with a single policy across outdoor stairs and indoor routes combining stair ascent and descent, platforms, and gaps. These results support organizing visual history around anticipated contacts for perceptive humanoid locomotion.

関連論文

PR本紙発行元 EmplifAI