日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2610.08789

QF3: フィルタリングされたQ勾配による高速フロー強化学習

QF3: Fast Flow RL with Filtered Q-Gradients

シェア:XThreadsFacebookLINEはてブBluesky

フローポリシーを強化学習で訓練する際、批評家の行動勾配を1ステップ予測に逆伝播し、信頼できる行動次元のみに適用するオフポリシー手法を提案。ヒューマノイド歩行をゼロから学習し実機にゼロショット転移、既存手法より10倍高速化を実現。

詳しい要約

1. どんなもの?

- 強化学習(RL)アルゴリズムの一種で、flow policyを学習するためのオンラインoff-policy手法。 - 名前はQF3 (Fast Flow RL with Filtered Q-Gradients)。 - flow matchingとcriticのaction gradientを組み合わせ、flowの出力のone-step predictionを通して逆伝播する。 - 事前学習済みflow policyの改善や、デモンストレーションからの学習、ゼロからの学習に適用可能。 - 人型ロボットのlocomotion policyをゼロから学習し、ハードウェアにzero-shot転移した初のoff-policy flow RL手法と主張。

2. 先行研究と比べてどこがすごい?

- 従来のon-policy flow RL手法であるFPO++と比較して、wall-clockで10倍の高速化を実現。 - 人型locomotionとmotion-trackingのpolicyを学習可能。 - off-policy flow RLとして初めて、人型locomotion policyをゼロから学習し、実機にzero-shot転移したと主張。 - 事前学習済みflow-based manipulation policyのfine-tuningにも適用可能で、ABC-SimとRobomimicタスクで有効性を示す。

3. 技術・手法の肝は?

- flow matchingに加えて、criticのaction gradientを利用する。 - criticの勾配は、flowの出力のone-step predictionを通して逆伝播される。 - 更新を信頼できる領域に保つため、critic勾配をreplay actionの近傍に留まるaction次元にのみ適用する(filtered Q-gradients)。 - 高スループットのoff-policy学習レシピと組み合わせる。

4. どうやって有効だと検証した?

- 人型locomotion policyをゼロから学習し、実機にzero-shot転移して検証。 - FPO++と比較して10倍のwall-clock高速化を確認。 - 事前学習済みflow-based manipulation policyをABC-SimとRobomimicタスクでfine-tuningし、有効性を検証。 - 詳細な実験設定や評価指標は要旨からは不明。

5. 議論はある?

- 要旨からは、手法の限界や議論についての明示的な記述はない。 - 今後の課題や制約については要旨からは不明。

6. 次に読むべき論文は?

- FPO++ (on-policy flow RL手法) - flow matching - off-policy RL - ABC-Sim - Robomimic - 人型locomotion - motion-tracking

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chung Min Kim, Brent Yi, David McAllister, Hongsuk Choi, Himanshu Gaurav Singh, Jinkun Cao, Ken Goldberg, Pieter Abbeel, Carmelo Sferrazza, Angjoo Kanazawa

分類: cs.RO, cs.LG

原文アブストラクト

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with a high-throughput off-policy training recipe, it trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++, a recent on-policy flow RL method. We further apply QF3 to fine-tune pretrained flow-based manipulation policies on both ABC-Sim and Robomimic tasks. These results suggest that QF3 can both learn robot policies from scratch and refine those acquired from demonstrations. Website: https://qf3-rl.github.io/

関連論文

PR本紙発行元 EmplifAI