日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2610.00729

報酬を観測として:迅速な適応のための報酬ベース方策の学習

Reward as Observation: Learning Reward-Based Policies for Rapid Adaptation

シェア:XThreadsFacebookLINEはてブBluesky

報酬と行動のみを条件とする方策を提案し、観測空間が全く異なる環境へのゼロショット適応を実現した。

著者: Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha

分類: cs.LG, cs.RO

原文アブストラクト

This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerable amount of samples for successful transfer. Instead, we propose a novel reward-based policy only conditioned on rewards and actions, enabling zero-shot adaptation to new environments with completely different observations. We discuss the challenges and feasibility of a reward-based policy and then propose a practical algorithm for training. We demonstrate that a reward policy can be trained within three different environments, Pointmass, Cartpole, and 2D Car Racing, and transferred to completely different observations, such as different color palettes or 3D rendering, or Stretch robot navigation in Habitat-Sim, in a zero-shot manner. We also demonstrate that a reward-based policy can further guide the training of an observation-based policy in the target environment.

関連論文

PR本紙発行元 EmplifAI