日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2610.03395

強化学習のための双方向ボロノイバイアス探索カリキュラム

Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

目標条件付き強化学習において、開始状態と目標を両端から同時に拡張し、未探索領域にバイアスをかけて互いに近づけることで、スパース報酬の長期的タスクを効率的に学習するカリキュラムを提案。

詳しい要約

1. どんなもの?

- 強化学習のための双方向Voronoiバイアス探索カリキュラム(BVER)を提案。 - 長期的タスクで報酬が疎な場合の探索ボトルネックを解決。 - 開始状態と目標を両端から同時に拡張する。 - 目標条件付きポリシーを訓練。

2. 先行研究と比べてどこがすごい?

- 参照動作、手設計カリキュラム、整形報酬はデモやタスク固有のエンジニアリングが必要。 - 自動開始状態・目標カリキュラムは片側からのみ拡張するため、目標までの全距離を片側からカバーする必要がある。 - BVERは両端から同時に拡張し、参照不要で高速学習を実現。 - 0.4mボックスクライミングで95%成功を約65%少ないイテレーションで達成。 - 0.7mボックスを学習した唯一の手法。 - 開始、目標、ヨー変化にロバストなポリシーを生成。

3. 技術・手法の肝は?

- 双方向RRTプランニングに着想。 - 目標から開始状態を外側に成長させ、初期状態分布から目標を外側に成長させる。 - 両者を未探索タスク空間にバイアスし、互いに誘導。 - 単一の目標条件付きポリシーを両方で訓練。

4. どうやって有効だと検証した?

- ポイントマス迷路、四足歩行ボックスクライミング、ロボットアームリングオンペグ転送で評価。 - 参照不要カリキュラム全てより高速学習。 - 0.4mボックスで95%成功、0.7mボックスを学習した唯一の手法。 - 開始、目標、ヨー変化にロバスト。 - デモなしで参照ベースカリキュラムのサンプル効率に接近。 - アブレーションで両端拡張が片側より優れることを確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 双方向RRTプランニング、参照ベースカリキュラム、自動開始状態・目標カリキュラム。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter

分類: cs.LG, cs.RO

原文アブストラクト

Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space, and steers them toward each other, training one goal-conditioned policy on both. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. On box climbing, it reaches 95% success on a 0.4 m box in roughly 65% fewer iterations than the best of them, is the only one of them to learn to climb a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Without a demonstration, it approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. Ablations show that expanding from both ends outperforms either direction alone.

関連論文

PR本紙発行元 EmplifAI