日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
生涯強化学習arXiv:2610.03119

生涯強化学習における継続的適応のためのポリシー探索と再利用

How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

オンライン経験からタスク類似度を推定し、過去のポリシーを選択・重み付けして新タスクの事前知識として再利用する手法AMSCを提案。

詳しい要約

1. どんなもの?

- lifelong reinforcement learning における新タスクへの適応手法。 - 過去の policy を保持するだけでは不十分で、複数の policy に分散した知識を選択・統合する必要がある。 - Adaptive Mask Selection and Composition (AMSC) を提案。 - オンライン経験から task similarity を推定し、prior を形成する。

2. 先行研究と比べてどこがすごい?

- 従来の modular composition baseline と比較。 - CT-graph と MiniGrid で平均性能と forward transfer が向上。 - forgetting が生じない。 - 単純な policy 保持や固定構成より、類似度に基づく選択・重み付けが有効。

3. 技術・手法の肝は?

- 非パラメトリック Wasserstein task embeddings を state-action-reward サンプルから算出。 - 類似度スコアに z-score 正規化 sparsemax を適用。 - 可変サイズの support を導出し、定期的に policy を選択・重み付け。 - 新タスク学習時の prior を形成。

4. どうやって有効だと検証した?

- CT-graph と MiniGrid で評価。 - 平均性能と forward transfer を modular composition baselines と比較。 - Continual World でも結果を報告。 - Ablations で関連ソースの選択と再利用強度の決定が重要と確認。 - 独立に測定した pairwise transfer と task-embedding similarity の正の関連を確認。

5. 議論はある?

- Continual World では層ごとの関連知識の特定と構成に追加の層別チューニングが必要かもしれない。 - タスク類似性が知識の選択・重み付けの有効な基準となり得る。 - ただし、層別の調整や一般化には議論の余地。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: modular composition baselines, CT-graph, MiniGrid, Continual World。 - 関連手法: Wasserstein task embeddings, sparsemax, z-score normalization。 - 同分野の定番: lifelong reinforcement learning, continual learning, forward transfer, policy reuse。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio

分類: cs.LG, cs.AI, cs.RO

原文アブストラクト

In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and Composition (AMSC) is designed to estimate similarity from online experience via non-parametric Wasserstein task embeddings from state-action-reward samples. The z-score-normalized sparsemax of the similarity scores are used to derive a variable-size support to periodically choose and weight policies to form a prior when learning a new task. On CT-graph and MiniGrid, AMSC achieves higher mean performance and forward transfer than the evaluated modular composition baselines while exhibiting no forgetting. Results on Continual World suggest that identifying relevant prior knowledge and determining its layer-specific composition may require additional layer-specific tuning. Ablations show that selecting relevant sources and determining how strongly to reuse them are central to these gains. Independently measured pairwise transfer is also positively associated with task-embedding similarity. These results indicate that task similarity can be an effective criterion to select and weight specific knowledge for reuse in lifelong reinforcement learning.

PR本紙発行元 EmplifAI