日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチタスク強化学習arXiv:2505.23150

大きく、正則化し、カテゴリカルに:高容量価値関数は効率的なマルチタスク学習者

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

シェア:XThreadsFacebookLINEはてブBluesky

高容量な価値モデルをクロスエントロピーで学習し、学習可能なタスク埋め込みで条件付けることで、オンライン強化学習におけるタスク干渉を解決し、マルチタスク学習をスケーラブルにする手法を提案。

著者: Michal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar, Pieter Abbeel

分類: cs.LG

原文アブストラクト

Recent advances in language modeling and vision stem from training large models on diverse, multi-task data. This paradigm has had limited impact in value-based reinforcement learning (RL), where improvements are often driven by small models trained in a single-task context. This is because in multi-task RL sparse rewards and gradient conflicts make optimization of temporal difference brittle. Practical workflows for generalist policies therefore avoid online training, instead cloning expert trajectories or distilling collections of single-task policies into one agent. In this work, we show that the use of high-capacity value models trained via cross-entropy and conditioned on learnable task embeddings addresses the problem of task interference in online RL, allowing for robust and scalable multi-task training. We test our approach on 7 multi-task benchmarks with over 280 unique tasks, spanning high degree-of-freedom humanoid control and discrete vision-based RL. We find that, despite its simplicity, the proposed approach leads to state-of-the-art single and multi-task performance, as well as sample-efficient transfer to new tasks.

関連論文

PR本紙発行元 EmplifAI