局所全身VLA行動からシーン規模の空中マニピュレーションへ
From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation
関節型無人航空マニピュレータ向けに、合成データでVLAを訓練し、遅延に頑健な行動実現とシーングラフによる指示接地でシーン規模の空中作業を可能にする統合フレームワークを提案。
著者: Weixiang Guo, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, Lihua Xie
分類: cs.RO
原文アブストラクト
Vision-language-action (VLA) models enable task-conditioned interaction, but extending them to scene-scale aerial manipulation remains challenging due to costly whole-body demonstrations, latency-induced action-state misalignment, and cross-site behavior composition. We present a unified framework for synthetic policy training and scene-scale execution on articulated uncrewed aerial manipulators (UAMs). A scene-reconfigurable pipeline synthesizes task-conditioned, kinodynamically feasible trajectories and synchronized multiview observations for VLA training without physical-platform demonstrations. Measured-progress-aligned realization (MPAR) aligns asynchronously returned action chunks with measured execution progress and realizes them as continuous, dynamically feasible trajectories. A relational Scene Graph grounds language goals to object instances and feasible interaction regions, while topology-guided transfer connects local behaviors across sites. Local VLA skills achieve 39/60 successes (65.0%) in simulation under oracle target and feasible-handoff conditions. Under 500-ms added latency, with and without a transient command-update stall, MPAR reduces median takeover phase error by 0.212 s over nominal-time alignment. The complete system completes 21/50 simulated multi-site missions (42.0%) and is further validated on a physical articulated UAM.
関連論文
- 部分的な視覚観測による全身空中把持と持ち上げ空中マニピュレーション
- 可動コンプライアントアンカーによるケーブル吊り下げ空中マニピュレーションの受動剛性成形空中マニピュレーション
- 空中グリッパー:勾配ベースのリアルタイム逆ゲーム予測・計画フレームワーク空中マニピュレーション
- QuadHand:MRC-SDFを用いた全身運動計画によるコンパクトなクアッドロータ空中マニピュレータ空中マニピュレーション
- オンボード知覚・方策学習・全身制御による屋外空中マニピュレーション空中マニピュレーション
- テザー吊り下げ型空中放射線センシングペイロードの視覚ベース制御空中マニピュレーション