日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
歩行arXiv:2610.10453

RFPO: 身体性制御のための整流フローポリシー最適化

RFPO: Rectified Flow Policy Optimization for Embodied Control

シェア:XThreadsFacebookLINEはてブBluesky

フローポリシーの少ステップ推論時の性能劣化を解消するため、オンラインリフローとPPO監督を組み合わせたRFPOを提案し、1ステップ実行で64ステップと同等の制御性能を維持しつつ推論を54.9倍高速化した。

詳しい要約

1. どんなもの?

RFPOは、flow-based policyをembodied control向けにfew-step実行で信頼できるよう最適化する枠組み。 - 対象は連続ロボット制御で、flow policyの反復ODE積分による推論コストを削減。 - 学習時はfull-stepで最適化されるが、粗い数値積分では性能が落ちるfew-step discretization gapを問題視。 - 提案はReward-aware online Reflowとfrozen Gaussian PPO controllerを組み合わせ、展開時は1 Euler stepの単一flow student。 - Unitree Go2, Boston Dynamics Spot, Unitree H1, Unitree G1で検証。

2. 先行研究と比べてどこがすごい?

flow-based policyは表現力が高いが反復ODE積分で推論コスト大。 - 単純に積分予算を減らすと、full-step前提の最適化のため制御性能が大きく劣化する。 - RFPOはfew-step discretization gapに対処し、one-step実行でもfull-step性能を維持。 - 具体的にはone-step returnsが64-step値の2.4%以内、Unitree Go2で98.5%の報酬を保持しつつ推論遅延を4.39 msから0.08 msへ54.9x高速化。

3. 技術・手法の肝は?

中核はReward-aware online Reflow。 - on-policy学習中にstudent-induced transport pathsをrectifyし、粗い積分への堅牢性を高める。 - frozen Gaussian PPO controllerがfullおよびintermediate integration budgetsで補完的なaction-space supervisionを供給。 - 展開ポリシーは単一flow studentで、1 Euler stepで実行。

4. どうやって有効だと検証した?

Unitree Go2, Boston Dynamics Spot, Unitree H1, Unitree G1で評価。 - zeroおよびrandom initializationの両方で、one-step returnsが64-step値の2.4%以内。 - Unitree Go2ではone-step実行が64-step報酬の98.5%を保持し、オンボード平均推論遅延を4.39 msから0.08 msへ短縮、54.9x高速化。 - 実機実験で安定したone-step locomotionを確認。

5. 議論はある?

要旨からは不明。 - 限界や失敗事例、計算コスト、ハイパーパラメータ感度などは記述されていない。 - 実機実験の詳細条件や他手法との統計比較も要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。 - 関連手法としてflow-based policy、Gaussian PPO、Rectified Flow、Euler integration、on-policy learningが挙げられる。 - 同分野の定番としてPPOやdiffusion policy、flow matching policyの文献を次に読むべき。 - コードとウェブサイトが示されているため、実装と実験詳細を確認するのが有用。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ting Huang, Lisiyu Pan, Haoyu Wang, Zeyu Zhang, Siyuan Qian, Yanjun Li, Yandong Guo, Boxin Shi, Hao Tang

分類: cs.RO

原文アブストラクト

Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost. Naively reducing the integration budget can severely degrade control, since policies optimized under full-step execution are not explicitly constrained to remain reliable under coarse numerical integration. We refer to this mismatch as the few-step discretization gap. To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution. Reward-aware online Reflow rectifies student-induced transport paths during on-policy learning, making the resulting policy more robust to coarse integration. A frozen Gaussian PPO controller supplies complementary action-space supervision at full and intermediate integration budgets, while the deployed policy remains a single flow student executed with one Euler step. Across Unitree Go2, Boston Dynamics Spot, Unitree H1, and Unitree G1, RFPO consistently preserves full-step control performance under one-step execution, with one-step returns remaining within 2.4% of their corresponding 64-step values across both zero and random initialization. On Unitree Go2, one-step execution retains 98.5% of the 64-step reward while reducing onboard mean inference latency from 4.39 ms to 0.08 ms, yielding a 54.9x speedup. Real-robot experiments further validate stable one-step locomotion. Code: https://github.com/AIGeeksGroup/RFPO. Website: https://aigeeksgroup.github.io/RFPO.

関連論文

PR本紙発行元 EmplifAI