日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーション/世界モデルarXiv:2609.19863

渡る前に地形を感じる:オフロードナビゲーションのための世界モデル

Feeling Terrain Before Crossing: World Models for Off-Road Navigation

シェア:XThreadsFacebookLINEはてブBluesky

固有感覚を入力に用いて、カメラ映像だけでなくロボットが感じる物理的未来(滑り・傾き・振動・失敗リスク)を予測するオフロードナビゲーション世界モデルFeel-WMを提案し、実機とシミュレーションで視覚のみのモデルを上回る性能を示した。

詳しい要約

1. どんなもの?

- オフロードナビゲーションのためのWorld Model「Feel-WM」を提案。 - 自己受容感覚(proprioception)を条件に、ロボットが感じる物理的未来(将来のproprioceptive stateとfailure risk)とカメラが見るシーンを同時に予測。 - プランナは物理的未来とシーンをロールアウトし、予測failure riskとgoal similarityを分離可能なスコアで重み付け。 - 実オフロードデータとシミュレーションで、視覚のみのナビゲーションWorld Modelを上回る。 - Huskyに搭載し山道でオンライン計画、荒れた地面を予測して回避、end-to-end policyが失敗するコースを完走。

2. 先行研究と比べてどこがすごい?

- 既存のシーンフォーカスなWorld Modelは、計画軌道に沿ったロボットの滑り・傾き・振動を予測しない。 - 都市環境では予測シーンが十分な代理となるが、オフロードではロボット-地形相互作用が重要。 - Feel-WMは初めてproprioceptionを条件に、物理的未来を予測するオフロードナビゲーションWorld Model。 - 視覚のみのナビゲーションWorld Modelと比較して、オープンループ計画とクローズドループの粗地形ナビゲーションで優位。 - 車輪型と脚型の両プラットフォームで有効性を確認。

3. 技術・手法の肝は?

- proprioceptionを入力として条件付け、将来のproprioceptive stateとfailure riskを予測。 - 物理的未来はロボット自身の経験から人間ラベルなしで学習。 - プランナは物理的未来とシーンをロールアウトし、予測failure riskとgoal similarityを分離可能なスコアで評価。 - カメラが見るものとロボットが感じるものを同時に予測するWorld Model。 - 計画候補の行動系列ごとに未来を予測し最良を選択するforesightベースの計画。

4. どうやって有効だと検証した?

- 実オフロードデータとシミュレーションで実験。 - オープンループ計画とクローズドループの粗地形ナビゲーションで評価。 - 車輪型と脚型のプラットフォームで検証。 - Huskyを山道に展開し、オンラインで計画、荒れた地面を予測して回避、end-to-end policyが失敗するコースを完走。 - 視覚のみのナビゲーションWorld Modelと比較して性能向上を確認。

5. 議論はある?

- 要旨からは不明。 - ただし、proprioceptionが物理的未来の予測を改善することが示唆されている。 - 予測failure riskとgoal similarityの分離可能なスコアの重み付けに関する議論は要旨からは不明。 - 実世界展開の限界や計算コストについては要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 視覚のみのナビゲーションWorld Model、end-to-end policy。 - 関連手法: シーンフォーカスなWorld Model、オフロードナビゲーションのためのWorld Model。 - 同分野の定番: モデルベース強化学習、予測制御、ナビゲーションWorld Model。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: E-In Son, Dong-Wook Kim, Ji-Hoon Hwang, Kangsun Lee, Jisung Bae, Jung-Taak Kim, Seung-Woo Seo

分類: cs.RO, cs.CV

原文アブストラクト

Navigation world models plan by foresight, predicting the future that each candidate action sequence produces and selecting the best, rather than mapping observations to actions directly. Unlike urban settings where a predicted scene is a sufficient proxy, off-road navigation hinges on the robot--terrain interaction, so the prediction must cover not only what the camera will see but what the robot will feel. However, existing scene-focused models do not predict how much the robot will slip, tilt or shake along a planned trajectory. Proprioception captures these dynamics directly and, when used as input, improves the prediction of the physical future. We present Feel-WM, the first off-road navigation world model that conditions on proprioception and predicts what the robot will feel alongside what the camera will see. The physical future takes the form of a future proprioceptive state and a failure risk, both learned from the robot's own experience without human labels. The planner rolls out the physical future alongside the scene and weighs the predicted failure risk against goal similarity in a separable score. Experiments on real off-road data and in simulation demonstrate that Feel-WM outperforms visual-only navigation world models in open-loop planning and closed-loop rough-terrain navigation across wheeled and legged platforms. Deployed on a Husky on mountain trails, Feel-WM plans onboard, predicts rough ground ahead and steers around it, completing courses that an end-to-end policy fails.

PR本紙発行元 EmplifAI