日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習/マニピュレーションarXiv:2609.36250

フィードバック補正を備えたアクションチャンキングPPO

Action Chunking Proximal Policy Optimization with Feedback Correction

シェア:XThreadsFacebookLINEはてブBluesky

PPOにアクションチャンキングを導入し、チャンク単位の計画とステップごとのフィードバック補正を組み合わせることで、高次元ロボット制御タスクの性能を向上させた研究。

詳しい要約

1. どんなもの?

- 高次元ロボティック制御のための強化学習手法 - Action Chunking PPO (ACPPO) を提案 - PPO を拡張し、chunked actor と標準的な state-value critic を使用 - chunked Q-function を回避 - ACPPO-Corr も提案 - chunk planner に stepwise feedback corrector を追加 - chunk 内で計画行動をオンライン調整 - 対象: locomotion, arm manipulation, dexterous hand-object interaction

2. 先行研究と比べてどこがすごい?

- 既存の action chunking 手法の2つの限界に対処 - 第1: action chunk 上の value function は action 次元と chunk 長が増えると学習困難 - 第2: chunk を open-loop で実行すると chunk 内 feedback が失われ、contact-rich タスクでの反応性が制限 - ACPPO は chunked Q-function を避け、標準 state-value critic を保持 - ACPPO-Corr は closed-loop correction を導入し、chunk-level planning と local feedback を両立 - 評価した手法の中で最も強い総合性能を達成

3. 技術・手法の肝は?

- PPO を基盤とした拡張 - chunked actor を使用しつつ、標準的な state-value critic を維持 - chunked Q-function を学習する必要がない - ACPPO-Corr では chunk planner に stepwise feedback corrector を追加 - 各 chunk 内で計画された action をオンラインで調整 - corrector regularization が重要 - chunk-level planning と local feedback のバランスを取る - 適度な chunk 長が最良

4. どうやって有効だと検証した?

- IsaacGym と Bi-DexHands の25のシミュレーションロボティクスタスクで評価 - locomotion, arm manipulation, dexterous hand-object interaction を含む - ACPPO-Corr は評価した手法の中で最も強い総合性能 - decision-frequency-sensitive と decision-frequency-neutral の両タスクサブセットで最良 - Ablations により、適度な chunk 長が最良であること、corrector regularization が重要であることを確認

5. 議論はある?

- action chunking は online PPO で有効となり得る - chunk-level planning と closed-loop correction を組み合わせた場合 - 限界や課題については要旨からは不明 - 今後の方向性については要旨からは不明

6. 次に読むべき論文は?

- PPO (Proximal Policy Optimization) - Action Chunking 関連手法 - IsaacGym - Bi-DexHands - その他、要旨で参照/比較されている具体的な研究は不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sanghyun Hahn, Jonghyun Choi

分類: cs.LG, cs.RO

原文アブストラクト

Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online within each chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregate performance among evaluated methods and performs best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that corrector regularization is important for balancing chunk-level planning with local feedback. These results suggest that action chunking can be effective in online PPO when chunk-level planning is paired with closed-loop correction. The code is available at: https://github.com/hshhahn/ACPPO.

関連論文

PR本紙発行元 EmplifAI