日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2609.27070

ガウスで十分:大規模行動モデルのファインチューニングにフローマッチング事前分布は効かない

The Gaussian Is Enough: Flow-Matching Priors Do Not Help When Fine-Tuning Large Behavior Models

シェア:XThreadsFacebookLINEはてブBluesky

大規模行動モデルのファインチューニングでは、目標に近い非ガウス事前分布を使っても標準ガウスと性能が変わらないことを大規模実験で示し、その理由を分析した論文。

詳しい要約

1. どんなもの?

大規模行動モデル(LBM)のファインチューニング時における事前分布選択の影響を検証した研究。 - 対象は diffusion や flow-matching ベースの生成的政策。 - 標準 Gaussian 事前分布と、目標に近い非 Gaussian 事前分布を比較。 - LBM 1.0、$π_{0.5}$、GR00T N1.5 の3モデルを対象。 - 2つのシミュレーションプラットフォーム、40以上のタスク、10万回以上のロールアウト、実機5タスク1250回のロールアウトで評価。 - 結論として、非 Gaussian 事前分布はファインチューニング性能を統計的に有意に改善しないか、むしろ悪化させる。

2. 先行研究と比べてどこがすごい?

先行研究では、ゼロから学習する場合に非 Gaussian 事前分布が性能を大幅に改善することが示されていた。 - 本研究は、その利得がファインチューニングに転移するかを検証。 - 直感に反し、事前学習済み LBM のファインチューニングでは非 Gaussian 事前分布の利点がほぼ消失。 - 非常に低いファインチューニングデータ割合でのみ、わずかな利点の可能性があるが、明確ではない。 - 大規模な実験により、事前分布選択よりもエンコーダの学習が支配的であることを示した点が新しい。

3. 技術・手法の肝は?

技術の肝は、事前分布の選択がファインチューニング性能に与える影響を大規模に比較した点。 - 標準 Gaussian 事前分布と、目標分布に近い非 Gaussian 事前分布を用意。 - 3つの LBM(LBM 1.0、$π_{0.5}$、GR00T N1.5)をファインチューニング。 - シミュレーションと実機の両方でロールアウトを実施。 - 診断分析として、ファインチューニング後の政策が事前分布によらず類似の行動予測に収束することを確認。 - エンコーダ埋め込みは事前学習時から大きく乖離し、互いに異なることも示す。 - 学習率アブレーションにより、エンコーダの学習が性能を支配する要因であると確認。

4. どうやって有効だと検証した?

有効性の検証は、大規模なシミュレーションと実機実験で実施。 - シミュレーション:2つのプラットフォーム、40以上のタスク、10万回以上のロールアウト。 - 実機:5つの両手マニピュレーションタスクで1250回のロールアウト。 - 非 Gaussian 事前分布が目標に近いことを示した上で、ファインチューニング性能を Gaussian 事前分布と比較。 - 結果、統計的に区別できないか、非 Gaussian の方が悪い場合が多い。 - 診断分析と学習率アブレーションにより、エンコーダ学習の重要性を確認。

5. 議論はある?

議論の焦点は、なぜ非 Gaussian 事前分布がファインチューニングで効かないか。 - ファインチューニング後の政策は、事前分布によらず類似の行動予測に収束する。 - 一方で、ファインチューニングされたエンコーダ埋め込みは事前学習時から大きく乖離し、事前分布間でも異なる。 - 学習率アブレーションから、エンコーダの学習がファインチューニング性能の支配的要因であり、事前分布選択の影響を大きく上回る。 - 非常に低いデータ割合でのみ事前分布が影響する可能性があるが、要旨からは明確でない。 - 今後の研究方向として、いつなぜ学習された事前分布がファインチューニングで重要になるかを探る必要がある。

6. 次に読むべき論文は?

要旨で参照・比較されている研究や関連手法を挙げる。 - LBM 1.0 - $π_{0.5}$ - GR00T N1.5 - diffusion policy - flow-matching - 非 Gaussian 事前分布を用いたゼロからの学習に関する先行研究(具体的な論文名は要旨からは不明) - 同分野の定番として、imitation learning、behavior cloning、generative policies に関する基礎文献。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chen Xu, Rishi Shah, Hadas Kress-Gazit, Haruki Nishimura, Masha Itkina

分類: cs.RO, cs.AI

原文アブストラクト

Modern robot imitation learning increasingly relies on generative policies based on diffusion or flow-matching models, which generate actions by transforming samples from a prior distribution. A key question is whether the choice of prior matters. Replacing the standard Gaussian with a closer-to-target, non-Gaussian prior has been shown to substantially improve performance when training from scratch. A natural next step is to ask whether these gains transfer to fine-tuning pretrained Large Behavior Models (LBMs) such as LBM 1.0, $π_{0.5}$, and GR00T~N1.5, where one might expect even larger gains. Surprisingly, we find that this is not the case, except possibly at very low fine-tuning data fractions. Across over 100K simulation rollouts spanning all three aforementioned LBMs on 40+ tasks in two simulation platforms, and 1250 hardware rollouts on five bimanual manipulation tasks, non-Gaussian priors that are demonstrably closer to the target yield statistically indistinguishable or worse fine-tuning performance than a standard Gaussian prior. Diagnostic analyses suggest why: fine-tuned imitation learning policies converge to similar action predictions across priors, despite their fine-tuned encoder embeddings diverging substantially from the pretrained embeddings and each other. A learning-rate ablation further confirms that encoder training is the dominant factor in fine-tuning performance, substantially outweighing the effect of prior choice. We conclude with concrete directions for future research on when and why learned priors might still matter in fine-tuning. Project page: https://cxu-tri.github.io/non_gaussian_FT/

関連論文

PR本紙発行元 EmplifAI