日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.00864

キネマティックMeanFlow:ロボット基盤モデルのための一段階行動生成ポリシー

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

シェア:XThreadsFacebookLINEはてブBluesky

多段階フローマッチングの推論遅延を解消するため、速度場のダイナミクスを運動学的に分解した一段階行動生成ポリシーK-MFを提案し、GR00T-N1.6の遅延を大幅に削減した。

詳しい要約

1. どんなもの?

- ロボティック基盤モデル(RFM)における1ステップ行動生成を目指す研究。 - 多段階のflow matchingの推論遅延を克服するため、MeanFlowを直接適用すると性能が崩壊する問題を発見。 - その原因をRFMの速度場に見られる2つの動態(局所加速度の後半急増、サンプル間の大きさのばらつき拡大)に特定。 - これに対処するKinematic MeanFlow(K-MF)を提案し、1ステップ行動生成を実現。

2. 先行研究と比べてどこがすごい?

- 従来のMeanFlowをRFMに直接適用すると性能崩壊が生じることを明らかにした。 - 提案するK-MFは、多段階flow matchingをほとんどの設定で上回る性能を達成。 - ゼロからの学習とファインチューニングの両方で1ステップ生成を可能にした。 - 推論効率では、GR00T-N1.6のaction-head遅延を67.5%~74.4%削減し、エンドツーエンドで30.3%~54.9%の遅延削減を実現。

3. 技術・手法の肝は?

- 運動学的恒等式に基づき、MeanFlowの時間微分項を中間点で2つの部分区間項に分離。 - 分離された2項がそれぞれ初期段階と後期段階のdenoising動態を捉える。 - これによりプロセス全体での誤差増幅を緩和。 - 結果として、RFMの1ステップ行動生成を可能にする。

4. どうやって有効だと検証した?

- 多様なタスクにおいて、ゼロからの学習とファインチューニングの両パラダイムで評価。 - 多段階flow matchingと比較し、ほとんどの設定で優位性を確認。 - 推論効率をL40とJetson Orin上でeagerモードとcompiledモードで測定。 - GR00T-N1.6のaction-head遅延とエンドツーエンド遅延の削減率を報告。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- MeanFlow - flow matching - GR00T-N1.6 - Robotic Foundation Models (RFMs)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao

分類: cs.RO, cs.AI, cs.CL, cs.CV

原文アブストラクト

In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at https://github.com/IntelChina-AI/K-MF.

関連論文

PR本紙発行元 EmplifAI