日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.21740

サンドイッチ残差:世界モデルのパラメータ効率的なテスト時適応

Sandwich-Residuals: Parameter-Efficient Test-time Adaptation of World Models

シェア:XThreadsFacebookLINEはてブBluesky

事前学習済み世界モデルを凍結し、予測器周辺に小さな残差補正のみをオンライン学習させることで、テスト時適応をパラメータ効率的に実現する手法を提案。

詳しい要約

1. どんなもの?

- テスト時適応のための軽量な手法である Sandwich-Residuals を提案。 - 学習済みの world model を凍結し、predictor の周囲に小さな residual 補正のみを学習。 - オンラインで自己教師あり予測誤差を用いて最適化し、報酬・ラベル・ソースドメインデータは不要。 - AdaJEPA ベンチマークの21条件で評価し、凍結モデルの1.3倍の成功率、最強の AdaJEPA 変種の95%の性能を維持しつつ、97-99%少ないパラメータで適応。 - 複合シフト下では凍結モデルの1.9倍の成功率で、内部ブロック適応に匹敵。 - DINO-WM モデルでの3次元マニピュレーションでも同様の適応原理を実証。

2. 先行研究と比べてどこがすごい?

- 既存のテスト時適応手法は、事前学習モデルの一部を更新し、しばしば数百万のパラメータを変更し、どの内部コンポーネントを適応するかの選択が必要。 - 提案手法は事前学習 world model を凍結し、predictor 周囲の小さな residual 補正のみを学習するため、適応パラメータ数が97-99%少ない。 - 報酬・ラベル・ソースドメインデータを必要としない点も既存手法と異なる。 - 性能面では、凍結モデルより成功率が高く、最強の AdaJEPA 変種の95%の性能を維持し、複合シフト下では内部ブロック適応に匹敵する。

3. 技術・手法の肝は?

- 事前学習済み world model を凍結し、predictor の周囲に小さな residual 補正を学習する。 - residual はオンラインでモデルの自己教師あり予測誤差を用いて最適化される。 - 報酬、ラベル、ソースドメインデータは不要。 - これにより、内部重みを変更せずにテスト時適応を実現する。

4. どうやって有効だと検証した?

- AdaJEPA ベンチマークの21条件で評価。 - 凍結モデルの1.3倍の成功率を達成し、最強の AdaJEPA 変種の95%の性能を維持。 - 適応パラメータ数は97-99%少ない。 - 複合シフト下では凍結モデルの1.9倍の成功率で、内部ブロック適応に匹敵。 - DINO-WM モデルを用いた3次元マニピュレーションでも同様の適応原理を実証。

5. 議論はある?

- 有効なテスト時適応は必ずしも事前学習済み内部重みの変更を必要としないことを示唆。 - ただし、要旨からは具体的な限界や議論の詳細は不明。

6. 次に読むべき論文は?

- AdaJEPA ベンチマークおよびその最強変種(内部ブロック適応)。 - DINO-WM モデル。 - 関連するテスト時適応手法(例:パラメータ効率的適応、residual 学習)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Krishnam Soni, Aditya Sehgal, Vedant Dave, Elmar Rueckert

分類: cs.RO

原文アブストラクト

Latent world models enable planning by predicting the effects of actions in a learned representation space, but their predictions can become unreliable when test-time conditions differ from training. Existing test-time adaptation methods address this by updating parts of the pretrained model, often modifying millions of parameters and requiring a choice of which internal components to adapt. We introduce Sandwich-Residuals, a lightweight alternative that keeps the pretrained world model frozen and learns only small residual corrections around the predictor. The residuals are optimized online using the model's self-supervised prediction error and require no rewards, labels, or source-domain data. Across 21 conditions on the AdaJEPA benchmark, our method achieves $1.3\times$ the success rate of the frozen model while retaining 95% of the performance of the strongest AdaJEPA variant and adapting 97-99% fewer parameters. Under compound shifts, this advantage increases to $1.9\times$ the success rate of the frozen model, while remaining comparable to internal block adaptation. We further demonstrate the same adaptation principle on a DINO-WM model for 3-D manipulation. These results suggest that effective test-time adaptation of world models does not necessarily require modifying their pretrained internal weights.

関連論文

PR本紙発行元 EmplifAI