日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーションarXiv:2609.19824

TADreamer: ビデオ想像による陸空二モードロボットのゼロショット言語誘導3Dナビゲーション

TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination

シェア:XThreadsFacebookLINEはてブBluesky

生成ビデオから3Dウェイポイントを復元し、幾何情報でキャリブレーションすることで、陸空両用ロボットの言語指示ナビゲーションを学習なしで実現した。

詳しい要約

1. どんなもの?

- 陸空両用の bimodal robot 向けに、言語指示で3Dナビゲーションをゼロショットで行う TADreamer を提案。 - 生成 video を想像し、その video から3D waypoint と terrestrial/aerial の locomotion mode を復元。 - タスク特異的な training や fine-tuning を必要としない。 - VLM が onboard 観測と指示を navigation prompt に変換し、有効な生成 video を選択。 - 必要に応じて再生成のための corrective feedback も与える。 - 選択 video を3D waypoint に再構成し、mode ラベルを付与。 - 2段階 calibration で scale と geometry を実測に合わせ、planner が実行。

2. 先行研究と比べてどこがすごい?

- 生成 video を用いる navigation は scale ambiguity と軸依存の幾何歪みのため、metric に一貫した参照を得るのが困難だった。 - TADreamer は video imagination を measured geometry に接地し、zero-shot で実現。 - タスク特異的な training/fine-tuning が不要。 - 比較対象の NavDreamer に対し、calibration 観測で mean absolute depth error を87.7%、mean absolute relative depth error を86.3%削減。 - 7つの indoor/outdoor シナリオで、1ラウンドあたり5候補から2ラウンド以内に使用可能な video を取得。

3. 技術・手法の肝は?

- VLM が onboard 観測と指示から navigation prompt を生成し、生成 video を選択。 - 不適切な場合は corrective feedback を与えて再生成。 - 選択 video を3D waypoint に再構成し、terrestrial/aerial の mode を注釈。 - 2段階 calibration: まず field-of-view 制約で scale 推定を初期化。 - 次に再構成 point cloud を measured geometry に registration し、軸依存 scale、rotation、translation を refinement。 - 校正済み waypoint と mode ラベルを planner に渡し、measured geometry を組み込んで robot が実行。

4. どうやって有効だと検証した?

- 実世界実験で7つの indoor/outdoor シナリオのナビゲーションを実施。 - 各ラウンド5候補の video から、全7シナリオで2ラウンド以内に使用可能な video を取得。 - calibration 観測において、NavDreamer と比較して mean absolute depth error を87.7%、mean absolute relative depth error を86.3%削減。 - これにより video 想像から metric に一貫した navigation 参照を得る有効性を検証。

5. 議論はある?

- 生成 video の scale ambiguity と軸依存の幾何歪みが課題として挙げられている。 - 2段階 calibration でこれに対処しているが、限界や失敗ケースの詳細は要旨からは不明。 - 7シナリオで2ラウンド以内に使用可能 video が得られたが、候補数や再生成回数の影響は要旨からは不明。 - 実世界実験の規模や多様性、計算コスト、リアルタイム性は要旨からは不明。

6. 次に読むべき論文は?

- NavDreamer(要旨で比較対象として明示) - vision-language model を用いた language-guided navigation 関連研究 - video generation を用いた navigation 関連研究 - point cloud registration や calibration 関連研究 - terrestrial-aerial bimodal robot のナビゲーション関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiangyu Li, Tiancheng Lai, Xijie Huang, Ruitian Pang, Siqi Shen, Juncheng Chen, Zaisheng Pan, Chao Xu, Fei Gao, Yanjun Cao

分類: cs.RO

原文アブストラクト

Language-guided navigation for terrestrial-aerial bimodal robots requires selecting routes and locomotion modes that match scene context and task intent. Generated videos can represent such motion sequences, but recovering metrically consistent navigation references from them is challenging because of scale ambiguity and axis-dependent geometric distortions. We present TADreamer, a zero-shot framework that grounds video-imagined navigation in measured geometry without task-specific training or fine-tuning. A vision-language model translates onboard observations and instructions into navigation prompts, selects valid generated videos, and provides corrective feedback when regeneration is needed. The selected video is reconstructed into 3D waypoints annotated with terrestrial or aerial modes. A two-stage calibration procedure uses field-of-view constraints to initialize scale estimation, then refines axis-dependent scales, rotation, and translation by registering the reconstructed point cloud to measured geometry. The calibrated waypoints and mode labels guide a planner that incorporates measured geometry for robot execution. Real-world experiments demonstrate navigation across seven indoor and outdoor scenarios. With five candidates per round, usable videos are obtained within two rounds in all seven scenarios. On the calibration observations, our method reduces mean absolute depth error by 87.7% and mean absolute relative depth error by 86.3% compared with NavDreamer.

関連論文

PR本紙発行元 EmplifAI