日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.03620

UniIntervene++:実世界強化学習を効率化する適応的介入エージェント

UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

強化学習方策の能力変化に応じて、自律実行と多様な支援行動の間で制御を適応的に配分する介入エージェントを提案し、実世界のマニピュレーション5タスクで成功率89.67%・人的介入0.77%を達成した。

詳しい要約

1. どんなもの?

- 実世界のオンライン強化学習(RL)において、自律実行と異種の支援行動の間で制御を割り当てる適応的介入エージェント UniIntervene++ を提案。 - 進化する RL ポリシー、軌道修正、タスク構造化 CodePolicy を Options として統一半マルコフ決定過程(semi-Markov decision process)で定式化し、オンラインで相対価値を学習。 - 能力適応型介入により、RL ポリシーを非支援実行で定期的にプローブし、制御割り当てを能力変化に応答させる。 - 結合経験学習により、支援行動が RL ポリシーを改善し、その進化した結果が将来の介入決定を再形成する。 - いつ介入するか、どのように介入するか、いつ制御を戻すかを RL ポリシーの改善に合わせて共同決定する。

2. 先行研究と比べてどこがすごい?

- 既存の介入戦略はオフライン推定や固定決定ルールに基づくため、現在のポリシーと不一致になる可能性がある。 - UniIntervene++ は適応的介入エージェントとして、オンライン RL 中に制御を動的に割り当てる。 - 5つの実世界マニピュレーションタスクで平均成功率 89.67% を達成し、全ベースラインを少なくとも 6 パーセントポイント上回る。 - 人間の介入を 0.77% に削減し、最良ベースラインから少なくとも 94.6% の相対的削減を実現。

3. 技術・手法の肝は?

- 進化する RL ポリシー、軌道修正、タスク構造化 CodePolicy を Options として統一半マルコフ決定過程で定式化し、オンラインで相対価値を学習。 - 能力適応型介入:RL ポリシーを非支援実行で定期的にプローブし、制御割り当てを能力変化に応答させる。 - 結合経験学習:支援行動が RL ポリシーを改善し、その進化した結果が将来の介入決定を再形成する。 - これにより、いつ介入するか、どのように介入するか、いつ制御を戻すかを共同決定する。

4. どうやって有効だと検証した?

- 5つの実世界マニピュレーションタスクで評価。 - 平均成功率 89.67% を達成し、全ベースラインを少なくとも 6 パーセントポイント上回る。 - 人間の介入を 0.77% に削減し、最良ベースラインから少なくとも 94.6% の相対的削減を達成。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として、オンライン強化学習、実世界ロボットマニピュレーション、人間介入を用いた強化学習(例:HG-DAgger、Interactive Imitation Learning)などが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang

分類: cs.LG, cs.RO

原文アブストラクト

Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or fixed decision rules can therefore become mismatched to the current policy. To address this, we propose UniIntervene++, an adaptive intervention agent that learns to allocate control between autonomous execution and heterogeneous assisted behaviors during online RL. Specifically, UniIntervene++ first formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a unified semi-Markov decision process and learns their relative values online. Building on this, competence-adaptive intervention periodically probes the RL policy through unassisted execution, keeping control allocation responsive to its evolving capability. Finally, coupled experience learning allows assisted behaviors to improve the RL policy, whose evolving outcomes in turn reshape future intervention decisions. In this way, UniIntervene++ jointly determines when to intervene, how to intervene, and when to return control as the RL policy improves. Across five real-world manipulation tasks, UniIntervene++ achieves an average success rate of 89.67%, outperforming all baselines by at least 6 percentage points, while reducing human intervention to 0.77%, a relative reduction of at least 94.6% from the best baseline. Code is available in our \href{https://github.com/dannyyudong/An-Adaptive-Intervention-Agent-for-Efficient-Real-World-Reinforcement-Learning}{GitHub repository}.

関連論文

PR本紙発行元 EmplifAI