日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ナビゲーションarXiv:2610.08640

RIWANav:都市ナビゲーションのための自己改善型再帰的世界行動モデル

RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation

シェア:XThreadsFacebookLINEはてブBluesky

都市ナビゲーション向けに、行動条件付き世界モデルと方策を交互に更新する再帰的自己改善フレームワークを提案し、想像上のフィードバックと自己カリキュラムで性能を向上させた。

詳しい要約

1. どんなもの?

- 都市ナビゲーションのための長期ホライズンタスクを対象としたフレームワーク。 - 行動条件付き世界モデルと行動モデル(ポリシー)を再帰的に自己改善する。 - 各サイクルで世界モデルとポリシーを交互に更新し、想像上の結果を利用する。 - 実世界実験も行い、実用性を示す。

2. 先行研究と比べてどこがすごい?

- 従来の模倣学習(IL)は失敗から学びにくく、物理的な試行錯誤強化学習(RL)はコストが高い。 - 固定された世界モデルはポリシーの進化に伴い信頼性が低下する問題があった。 - 本手法は世界モデルとポリシーを再帰的に自己改善し、この問題に対処する。 - 実験で訓練ベースラインや先行手法を上回る性能を示す。

3. 技術・手法の肝は?

- 世界モデルとポリシーの結合適応をタスク固有の再帰的自己改善(RSI)として定式化。 - 各サイクルで2つの更新を交互に行う。 - 世界モデルが想像上の結果を通じてポリシーの行動を評価し、グループ相対ポリシー最適化(GRPO)の比較フィードバックを提供。 - 改善されたポリシーは、行動的新規性と予測誤差に基づいて専門家と整合する行動-ビデオペアを選択し、自己カリキュラムを構築。 - 洗練された世界モデルが次のポリシー更新のフィードバックを提供し、ループを閉じる。

4. どうやって有効だと検証した?

- 実験により、RIWANAVが訓練ベースラインや先行手法を上回ることを示す。 - ポリシーと世界モデル間の再帰的自己改善ループの有効性を検証。 - 実世界トライアルで実用性を実証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法:模倣学習(IL)、強化学習(RL)、行動条件付き世界モデル、グループ相対ポリシー最適化(GRPO)。 - 同分野の定番:都市ナビゲーションのための長期ホライズン計画手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jing Xie, Shouwei Ruan, Yubin Wang, Yuxiang Zhang, Haitao Yang, Songchang Jin, Dianxi Shi

分類: cs.RO

原文アブストラクト

Long-horizon urban navigation requires sequential local decisions whose errors can compound over time. Imitation learning (IL) rarely learns from failures, while physical trial-and-error reinforcement learning (RL) is costly. Action-conditioned world models can provide imagined feedback by predicting visual consequences for candidate actions. However, a frozen world model may become less reliable as the policy evolves. In this paper, we introduce RIWANAV, a post-training framework that casts the coupled adaptation of a world model and an action model (policy) as task-specific recursive self-improvement (RSI). Each cycle alternates two updates. The world model evaluates policy actions through imagined outcomes, providing comparative feedback for group-relative policy optimization (GRPO). The improved policy then constructs a grounded self-curriculum, selecting expert-consistent action-video pairs by behavioral novelty and prediction error. The refined world model supplies feedback for the next policy update, closing the recursive self-improvement loop. Experiments show that RIWANAV outperforms training baselines and prior methods, validating the proposed recursive self-improvement loop between the policy and world model. Real-world trials further demonstrate its practical applicability.

関連論文

PR本紙発行元 EmplifAI