CEMにおける世界モデルは提案メカニズムでもある
In CEM, a World Model Is Also a Proposal Mechanism
クロスエントロピー法(CEM)において、世界モデルが行動選択と次回の候補分布の両方に影響することを実験的に分離し、モデルのスコア誤差が提案メカニズムを通じて将来の候補にも影響を与えることを示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Oliver Obst, Frieder Stolzenburg
分類: cs.LG, cs.RO
原文アブストラクト
The cross-entropy method (CEM) uses world-model scores to select action sequences and fit the distribution sampled in its next iteration. A scoring error can therefore change both the present decision and the candidates considered later. We evaluate these two roles separately. Four types of predictive model generate CEM traces, and every model rescores every saved candidate pool. Executing the same candidates in the environment provides a reference elite set and proposal update. Across twelve independently trained task-seed units on Walker and Cheetah, the pre-specified proposal distance falls from the first to the final CEM iteration in every unit. Proposal widths contract and fitted means separate relative to the remaining search width. Pairwise ranking agreement stays near chance on Walker and declines on Cheetah; elite-set agreement does not improve. This comparison shows greater variation between scorers than between pool sources on Cheetah; Walker has variation in both and in their pairings. We use the original six units to select Random nonlinear for a one-update intervention, without inspecting intervention outcomes. Replacing its first model-ranked update with an environment-ranked update lowers final realised selected-sequence cost in those six units and in six further units held out from the selection.
関連論文
- 計画と学習のループを閉じる:学習済み世界モデルによるロボット制御モデルベース強化学習
- 速度と精度の両立:油圧ショベル制御のためのサンプル効率の高いオンラインモデルベース強化学習モデルベース強化学習
- 表現世界モデル:表現空間における状態・遷移・実行可能計画の学習モデルベース強化学習
- CAST: 交互状態価値目標と拡張方策勾配によるモデルベース強化学習モデルベース強化学習
- ニューロシンボリック世界モデルによるゼロショットタスク転送に向けてモデルベース強化学習
- BRICKS-WM: インターフェース合成力学による構造化世界モデルの再利用性構築モデルベース強化学習