日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2609.18732

PASSAGE: 雑然環境における知覚型ヒューマノイド移動のためのシーン整合運動学習のスケーリング

PASSAGE: Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments

シェア:XThreadsFacebookLINEはてブBluesky

VRと慣性モーションキャプチャで収集した1500の雑然シーンにおける100時間の人間動作データを用い、条件付きフローマッチングプランナと全身トラッカーを組み合わせて、未知の障害物環境でもヒューマノイドが知覚に基づき踏み越え・すり抜け・くぐり抜けを選択・実行できる枠組みを提案。

詳しい要約

1. どんなもの?

PASSAGEは、雑然とした環境を移動するhumanoid robotのための、perception-conditionedなplanner--trackerフレームワークである。 - VRとinertial motion captureで1,500のcluttered scenesにわたる100時間のscene-aligned human motionを収集 - conditional flow-matching plannerがmotion history、local destination、robot-centric multi-layer elevation mapから短horizonのreferenceを生成 - perceptive whole-body trackerが50 Hzでgeometric feedbackを用いて実行 - 実時間chunkingとplanner-side RL post-trainingを統合 - 一つのplanner--tracker pairでskill annotationやobstacle-specific policyなしに未見geo…

2. 先行研究と比べてどこがすごい?

既存手法はtask-specific reinforcement-learning objectivesやcurated motion librariesに依存し、広範なbehavioral coverageの獲得コストが高い。 - PASSAGEはskill annotationやobstacle-specific policyを必要とせず、単一のplanner--tracker pairで複数のtraversal behaviorを選択・合成 - 大規模なscene-aligned human motionデータを活用し、データスケーリングの効果を実証 - 完全onboardシステムとして、事前地図やoffboard計算なしで実環境移動を実現

3. 技術・手法の肝は?

技術の肝は、perception-conditionedなplannerとtrackerの統合、および大規模データ収集と学習パイプラインにある。 - VRとinertial motion captureによるscene-aligned human motion収集 - conditional flow-matching plannerがmotion history、local destination、robot-centric multi-layer elevation mapから短horizon referenceを生成 - perceptive whole-body trackerが50 Hzでgeometric feedbackを実行 - 実時間chunkingによるinter-chunk consistencyの促進 - frozen tracker下でのplanner-side RL post-trainingによるclosed-loop性能向上

4. どうやって有効だと検証した?

シミュレーションと実機で検証されている。 - シミュレーションでcomponent ablationを実施し、各stageの寄与を定量化 - 3つの独立したtraining seedで、収集データを6時間から100時間にスケールすると、held-out scenesでのmean contact-free successが48.1%から68.9%に向上 - validated scene augmentationを加えた最終モデルは70.3%に到達 - 50の未見physical layoutで、事前地図やoffboard計算なしの移動を実証 - 完全onboardシステムはegocentric 3D LiDAR perception、online occupancy mapping、6.25 Hz planning、50 Hz controlをJetson AGX Orin上で統合

5. 議論はある?

要旨からは、限界や議論の詳細は不明。 - データスケーリングの効果とablationの寄与は示されている - 実環境での未見layoutへの汎化が示唆される - しかし、失敗事例や安全性、計算負荷、他のhumanoidプラットフォームへの転移性などについての議論は要旨からは不明

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。 - 関連手法として、task-specific reinforcement-learning objectivesやcurated motion librariesを用いる既存のhumanoid traversal研究が挙げられる - 同分野の定番として、humanoid locomotionのためのreinforcement learning、whole-body control、perceptive locomotion、flow matching、motion captureからの学習などが次に読むべき候補

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuxuan Ma, Zicheng Zeng, Chunlin Peng, Zhoujian Li, Zetong Zhao, Zhikai Zhang, Yunrui Lian, Han Xue, Sikai Liang, Weiyi Zhu, Mulin Chen, Chenghuai Lin, Jiayu Zeng, Yanwei An, Songan Zhang, Jiayuan Gu, Jilong Wang, Jingbo Wang, He Wang, Li Yi

分類: cs.RO

原文アブストラクト

Humanoid robots can step over, squeeze past, and duck under obstacles, but learning to select and coordinate these behaviors from onboard perception remains challenging. Many existing approaches rely on task-specific reinforcement-learning objectives or curated motion libraries, making broad behavioral coverage costly. We present PASSAGE, a perception-conditioned planner--tracker framework for humanoid traversal. Using virtual reality and inertial motion capture, we collect 100 h of scene-aligned human motion across 1,500 cluttered scenes. A conditional flow-matching planner generates short-horizon references from motion history, a local destination, and a robot-centric multi-layer elevation map, while a perceptive whole-body tracker executes them at 50 Hz with geometric feedback. Real-time chunking promotes inter-chunk consistency, and planner-side RL post-training under the frozen tracker further improves closed-loop performance. Without skill annotations or obstacle-specific policies, one planner--tracker pair selects and composes traversal behaviors across unseen geometries. In simulation, component ablations quantify the contribution of each stage. Across three independent training seeds, scaling captured data from 6 to 100 h increases mean contact-free success from 48.1% to 68.9% on held-out scenes, while the final model with validated scene augmentation reaches 70.3%. The fully onboard system integrates egocentric 3D LiDAR perception, online occupancy mapping, 6.25 Hz planning, and 50 Hz control on a Jetson AGX Orin; tests across 50 unseen physical layouts demonstrate traversal without prebuilt maps or offboard computation.

関連論文

PR本紙発行元 EmplifAI