UniPoint:多様な地形を歩行するヒューマノイドのための統合ポイントレベルセンサフュージョン
UniPoint: Unified Point-Level Sensor Fusion for Humanoid Locomotion Across Challenging Terrains
LiDARと2台の深度カメラの点群を早期融合し、固定トークン数の自己注意で処理するヒューマノイド歩行ポリシーを提案。8種類の地形を単一ポリシーで走破し、実機で70cm段差や100cmギャップなどを検証した。
著者: Sicen Li, Zhen Chu, Chao Li, Qiuguo Zhu, Jun Wu
分類: cs.RO
原文アブストラクト
Open-world deployment requires humanoid robots to cross highly heterogeneous terrain safely, with perception that simultaneously provides wide coverage, local accuracy, and redundancy against sensor failure. Existing approaches struggle to satisfy all three: one forward depth camera or nearby height sampling covers too little; odometry-corrected elevation maps drift under aggressive motion and miss thin vertical structures; image-level encoding costs grow with camera count. We present UniPoint, a humanoid whole-body locomotion framework built on multi-source point-level sensor fusion. Measurements from a 360° light detection and ranging (LiDAR) sensor and two depth cameras are early-fused into one base-frame point set. Voxelization resamples it to a fixed number of tokens encoded by linear self-attention and proprioception-queried cross-attention, decoupling forward cost from sensor count. The point set retains standing thin barriers; a single-modality failure removes only part of the tokens, so the policy degrades gracefully. A single training run with terrain-aware rewards, perception-degradation injection, and domain randomization produces one policy for all eight terrain types, deployed on an onboard RK3588 without fine-tuning. On a DR02 humanoid, 20 trials at each of nine real-world settings over seven terrain types validate the policy on 70-cm-high platforms, 100-cm gaps, thin barriers, and sparse or narrow footholds; it also generalizes zero-shot outdoors.