日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00878

UniTrackPLA:指示に基づくナビゲーションと動的人物追跡のための統合パノラマ・言語・行動モデル

UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

シェア:XThreadsFacebookLINEはてブBluesky

全方位パノラマ観測を活用し、言語指示ナビゲーションと人物追跡を単一モデルで行う統合VLAを提案。行動予測の一貫性検証により再計画を可能にし、シミュレーションと実環境で性能を向上させた。

詳しい要約

1. どんなもの?

- 指示に基づくナビゲーションと動的な人物追跡を統合したモデル - 全方位パノラマ知覚を活用するUniTrackPLAを提案 - 言語指示とパノラマコンテキストを接地し、連続的なwaypointを予測 - 2つのタスクを単一のポリシーで閉ループ制御 - OmniTrackNav-Benchというベンチマークも導入

2. 先行研究と比べてどこがすごい?

- 既存手法は前方視点に依存し、タスクごとに別ポリシー - 全方位知覚と統合閉ループ制御を実現 - 追跡SRを23.50%から35.00%に改善 - Omni-VLN SR/SPLを13.00%/12.77%から19.75%/19.29%に改善 - 実世界ルート追加でEP@0.2mを42.92%から92.08%に向上

3. 技術・手法の肝は?

- Panoramic-Aware Encoding (PAE)でパノラマから投影した視点の時間・方位構造を保持 - 視点事前学習済み視覚エンコーダで全方位観測を処理 - 共有vision-languageバックボーンが指示をパノラマ文脈に接地 - 連続的なrobot-centric waypointチャンクを予測 - World-Action Consistency (WAC)が行動条件付き未来視覚状態を予測し、waypointプレフィックスをオンライン検証 - 一貫性があれば行動を再利用、不一致なら再計画

4. どうやって有効だと検証した?

- OmniTrackNav-Benchを構築 - 5,000のシミュレーション追跡軌道、10,000のシミュレーションVLNルート、96の検証済み実世界ルート - 919,978のwaypoint監督インスタンスを提供 - 追跡SR、Omni-VLN SR/SPL、EP@0.2mで評価 - Go2-Wロボットでの閉ループ実験で屋内・屋外環境での統合パノラマ追跡とナビゲーションを実証

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 同分野の定番としてVision-and-Language Navigation (VLN)、Person Tracking、Panoramic Perception、Waypoint Prediction、World Modelに基づく手法が挙げられる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Pengfei Qi, Haoran Lin, Sizhuang Chen, Kai Luo, Sirui Zhang, Xinqi Liu, Fei Cheng, Wenrui Chen, Liming Yin, Kailun Yang

分類: cs.RO, cs.CV, eess.IV

原文アブストラクト

General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perception and unified closed-loop control. We present UniTrackPLA, a unified panorama-language-action model for instruction-guided navigation and dynamic person tracking. Its Panoramic-Aware Encoding (PAE) preserves the temporal and azimuthal structure of perspective views projected from each panorama, enabling perspective-pretrained visual encoders to process omnidirectional observations. A shared vision-language backbone grounds instructions in the panoramic context and predicts continuous robot-centric waypoint chunks for both tasks. World-Action Consistency (WAC) further predicts action-conditioned future visual states and verifies waypoint prefixes online, allowing reliable actions to be reused while triggering replanning upon inconsistency. We also introduce OmniTrackNav-Bench, comprising 5,000 simulated tracking trajectories, 10,000 simulated VLN routes, and 96 verified real-world routes, providing 919,978 waypoint-supervision instances. UniTrackPLA improves overall tracking SR from 23.50% to 35.00% and Omni-VLN SR/SPL from 13.00%/12.77% to 19.75%/19.29%. Incorporating 76 real-world routes further improves held-out EP@0.2m from 42.92% to 92.08%. Closed-loop experiments on a Go2-W robot demonstrate unified panoramic tracking and navigation across indoor and outdoor environments. The project page is at https://tw5775.github.io/UniTrackPLA.

関連論文

PR本紙発行元 EmplifAI