CURIO: 好奇心駆動によるテスト時学習でオープンエンドな発見を実現
CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery
強化学習に好奇心ベースの世界モデルを組み合わせ、報酬の低い有望な探索方向を早期に切り捨てずにテスト時に学習する手法を提案。数学的発見タスクと単一細胞ノイズ除去で性能を改善。
著者: Tao Feng, Fangxu Yu, Zijie Lei, Jiaru Zou, Changjiang Jiang, Yi Yan, Jiaxuan You, Pan Lu
分類: cs.LG, cs.CL
原文アブストラクト
Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement learning (RL) enables such adaptation; however, strongly favoring high-reward trajectories may suppress low-reward yet potentially promising directions too early. We introduce CURIO, a curiosity-driven test-time learning framework that complements task feedback with an Intrinsic Curiosity World Model (ICWM). The ICWM learns transitions in the policy's hidden-state representation and supplies prediction-error bonuses at sampled tokens outside the policy's top-k choices. Epoch normalization and an annealed weight regulate their contribution to the policy update. On six mathematical discovery tasks and single-cell denoising with Qwen3 backbones from 8B to 235B, three-run means improve over a matched task-only RL control on five mathematical objectives, match the best reported performance on Circle Packing, and improve denoising Score and mean squared error (MSE) on both held-out corpora at every tested scale. Relative gains reach 18.3% on Hadamard and 10.8% on denoising Score. Code-diversity measurements show greater structural variation among generated programs, supporting curiosity as a complementary exploration signal for learning in open-ended discovery.
関連論文
- プロンプト駆動探索強化学習/探索
- 報酬なし事前学習:占有被覆最大化による強化学習強化学習/探索
- DF-ExpEnse: 拡散フィルタ探索によるサンプル効率的なファインチューニング強化学習/探索
- 表形式基盤モデルはロボット方策学習の探索を導けるか?強化学習/探索
- どこで学ぶか:オンポリシーロボット強化学習のための解析的方針勾配による指向的探索強化学習/探索
- 価値誘導フローによる高次元連続制御のスケーラブルな探索強化学習/探索