日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2601.10905

Action Shapley:強化学習における世界モデルのための訓練データ選択指標

Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

世界モデルの訓練データ選択にShapley値を応用した指標Action Shapleyを提案し、計算を効率化するランダム動的アルゴリズムで80%以上の高速化を実現、データ制約下の実世界5事例で有効性を示した。

著者: Rajat Ghosh, Debojyoti Dutta

分類: cs.LG, stat.ME

原文アブストラクト

Numerous offline and model-based reinforcement learning systems incorporate world models to emulate the inherent environments. A world model is particularly important in scenarios where direct interactions with the real environment is costly, dangerous, or impractical. The efficacy and interpretability of such world models are notably contingent upon the quality of the underlying training data. In this context, we introduce Action Shapley as an agnostic metric for the judicious and unbiased selection of training data. To facilitate the computation of Action Shapley, we present a randomized dynamic algorithm specifically designed to mitigate the exponential complexity inherent in traditional Shapley value computations. Through empirical validation across five data-constrained real-world case studies, the algorithm demonstrates a computational efficiency improvement exceeding 80\% in comparison to conventional exponential time computations. Furthermore, our Action Shapley-based training data selection policy consistently outperforms ad-hoc training data selection.

関連論文