日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2610.01224

成功基準を監督せよ:潜在世界モデル計画のための基準整合補助損失

Supervise What Decides Success: Criterion-Aligned Auxiliary Losses for Latent World-Model Planning

シェア:XThreadsFacebookLINEはてブBluesky

潜在世界モデルに成功基準量を回帰する補助損失を加え、テスト時は変更なしでPushTとcubeの成功率を改善した。

詳しい要約

1. どんなもの?

- Latent world model の計画における成功基準のずれを補正する手法。 - 成功基準量(例:end-effector position)を latent state が十分な精度で保持できていない問題を指摘。 - 成功基準量を訓練目標として使う補助損失(criterion-aligned auxiliary loss)を提案。 - テスト時には補助ヘッドを捨て、モデル・コスト・入力は不変。

2. 先行研究と比べてどこがすごい?

- 既存の latent world model は成功基準量を入力としてのみ扱う。 - 提案手法は成功基準量を訓練目標として直接用いる点が新しい。 - 4つの latent world model で end-effector position の符号化誤差が成功基準を超えることを示した。 - 補助損失のみで PushT と cube の成功率をそれぞれ 3.5% と 3.4%(絶対値)改善し、統計的有意。

3. 技術・手法の肝は?

- encoder と predictor の出力に線形ヘッドを付け、成功基準量を回帰。 - 回帰誤差を訓練損失に加える。 - 訓練後はヘッドを破棄し、テスト時のモデル・コスト・入力は変更なし。 - 成功基準が world model の latent state に保持すべき情報を指定する。

4. どうやって有効だと検証した?

- PushT と cube のタスクで成功率を評価。 - 補助損失のみで成功率がそれぞれ 3.5% と 3.4%(絶対値)向上。 - 両改善は統計的に有意。 - 4つの latent world model で end-effector position の符号化誤差が成功基準を超えることを確認。

5. 議論はある?

- 成功基準が world model の latent state に保持すべき情報を指定できることを示唆。 - 成功基準を直接訓練目標として使えることを実証。 - 他の成功基準量やタスクへの一般性は要旨からは不明。 - 補助損失の設計や線形ヘッドの詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として latent world model の planning 手法(例:PlaNet, Dreamer, TD-MPC)や PushT ベンチマーク関連の研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Takumi Hara, Kanata Suzuki

分類: cs.LG, cs.RO

原文アブストラクト

Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at test time are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), respectively, and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.

関連論文

PR本紙発行元 EmplifAI