報酬ハッキングを超えて:段階的人型学習パイプラインの4層におけるプロキシ乖離
Beyond Reward Hacking: Proxy Divergence Across Four Layers of a Staged Humanoid Learning Pipeline
脚式ロボットの強化学習パイプラインを構成する4つのプロキシ(報酬・カリキュラム・評価・参照動作)がそれぞれ目標から乖離する仕組みを分析し、各層の代替定式化を提案。シミュレーション人型ロボットのPPO学習で実測例を示す。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Arunabh Bora
分類: cs.RO
原文アブストラクト
A reinforcement-learning (RL) pipeline for a legged robot is assembled from proxies. A reward stands in for intended behaviour, a curriculum gate stands in for competence, an evaluation statistic stands in for robustness, and a reference motion stands in for an achievable skill. The traditional view treats only the first of these as optimised against, and so locates specification failure (reward hacking) in the reward alone. I argue that all four are proxies in the same formal sense, that each has a characteristic divergence mechanism, and that each admits a reformulation that closes it. For every layer I state the traditional formulation, derive the condition under which it diverges from its target, and give the alternative: first-order (L1) costs where quadratic kernels are flat, peak and outcome statistics where curriculum gates average, gate reachability and information checks, deterministic and phase-desynchronised evaluation, curriculum state treated as part of the model, feasibility-first reference design with residual feed-forward, and function-preserving input widening that lets one policy grow instead of being retrained. The arguments are illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid, grown over four stages and 13,500 iterations on a single laptop GPU. Among them, a curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent, and a batched push test whose synchronised resets aliased the gait phase ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one.
関連論文
- SceneFactory-3D:2D交通シーンを3D物理的反実世界へ持ち上げ、スケーラブルな物理基盤の安全評価を実現sim2real
- Skill2Real:ゼロショットSim-to-Realロボットマニピュレーションのためのエージェント型スキル学習sim2real
- RoboBridge:シミュレーションから実世界への転移のための自己進化型具現化エージェントフレームワークsim2real
- I2CD: 単一画像から凸分解された衝突形状を直接生成sim2real
- 大規模ロボット学習のためのGPUバッチ5Gシミュレーションによるネットワーク・イン・ザ・ループsim2real
- Awomo-SimDataEngine: エージェント型シミュレーション対応世界生成sim2real