日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2406.08315

ε-再訓練による方策最適化の改善

Improving Policy Optimization via $\varepsilon$-Retrain

シェア:XThreadsFacebookLINEはてブBluesky

方策最適化において、エージェントが望ましい行動を取れなかった状態領域を再訓練し、性能とサンプル効率を向上させる手法を提案。

著者: Luca Marzari, Priya L. Donti, Changliu Liu, Enrico Marchesini

分類: cs.AI, cs.LG

原文アブストラクト

We present $\varepsilon$-retrain, an exploration strategy encouraging a behavioral preference while optimizing policies with monotonic improvement guarantees. To this end, we introduce an iterative procedure for collecting retrain areas -- parts of the state space where an agent did not satisfy the behavioral preference. Our method switches between the typical uniform restart state distribution and the retrain areas using a decaying factor $\varepsilon$, allowing agents to retrain on situations where they violated the preference. We also employ formal verification of neural networks to provably quantify the degree to which agents adhere to these behavioral preferences. Experiments over hundreds of seeds across locomotion, power network, and navigation tasks show that our method yields agents that exhibit significant performance and sample efficiency improvements.

関連論文

PR本紙発行元 EmplifAI