日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
プランニングarXiv:2610.10274

コスト勾配による視覚ワールドモデルのスパースプランニング

Sparse Planning in Visual World Models via Cost Gradients

シェア:XThreadsFacebookLINEはてブBluesky

ワールドモデル内の空間トークンをプランニングコストの勾配ノルムでランク付けし、重要トークンだけを使って高速に計画する学習不要の手法を提案。

詳しい要約

1. どんなもの?

トークンベースの world model における空間トークン選択手法 COSTGRAD を提案する研究。 - 目的: 大規模な spatial token grid を毎回処理する action search の計算コストを削減する。 - 特徴: training-free かつ goal-conditioned な selector で、planning cost の各入力トークンに対する gradient norm でトークンをランク付けする。 - 対象: 連続制御ベンチマークでの latent planning。

2. 先行研究と比べてどこがすごい?

予測ではなく planning に効くトークンを選ぶ点が新しい。 - 従来のトークン選択は prediction 目的の重要度に基づくことが多いが、COSTGRAD は下流の control objective から重要度を導出する。 - 50% sparsity の AdaLN-conditioned predictor で、4 ベンチマーク中 3 つで full-token planning と同等以上。 - 環境あたりの planning step で 2.6× の wall-clock speedup を実測。 - token sparsity と CEM search 削減を組み合わせると約 5× の total speedup を達成し、full-token baseline を上回る。

3. 技術・手法の肝は?

planning cost の入力トークンに対する gradient norm を重要度として使う。 - training-free で goal-conditioned な selector。 - 各 spatial token を cost gradient norm でランク付けし、上位のみを planning に用いる。 - AdaLN-conditioned predictor を主対象とし、token sparsity と CEM search 削減を組み合わせ可能。 - 詳細なアルゴリズムやハイパーパラメータは要旨からは不明。

4. どうやって有効だと検証した?

4 つの連続制御ベンチマークで評価。 - AdaLN-conditioned predictor、50% sparsity で full-token planning と比較し、4 中 3 で同等以上。 - planning step あたり 2.6× の wall-clock speedup を測定。 - token sparsity と CEM search 削減の併用で約 5× の total speedup を確認。 - AdaLN vs concat の比較で、concat では pure COSTGRAD が random selection に対する優位を失う failure mode を同定。 - action-pathway drift を指標に、AdaLN では gradient-selected removal が random removal より drift が小さいが、concat では逆に大きいことを確認。

5. 議論はある?

selector と architecture の相性が sparse world-model planning の設計軸であると指摘。 - AdaLN では COSTGRAD が有効だが、concat では full-token 性能は同等でも COSTGRAD の優位が失われる。 - この差は action-pathway drift と対応する。 - 限界や一般化可能性、他の architecture への適用性は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている研究や関連手法。 - AdaLN-conditioned predictors と concat の比較。 - CEM search。 - token-based world models。 - 関連する sparse planning や token selection の研究(具体的な論文名は要旨からは不明)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yingchen Xu, Edward Grefenstette

分類: cs.LG

原文アブストラクト

Token-based world models enable fine-grained latent planning, but repeatedly processing large spatial token grids makes action search expensive. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at $50\%$ sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured $2.6\times$ wall-clock speedup per environment planning step. Combining token sparsity with reduced CEM search increases this to a $\sim 5\times$ total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This difference tracks action-pathway drift: gradient-selected removal produces less drift than random removal on AdaLN, but more on concat. These results highlight selector-architecture compatibility as a design axis for sparse world-model planning. Project page and demos: https://ycxuyingchen.github.io/costgrad/

関連論文

PR本紙発行元 EmplifAI