日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2610.12231

残差モデリングでロボット学習の回帰と生成ポリシーのギャップを埋める

Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの模倣学習において、MSE回帰とフローマッチングの性能差を残差の統計的性質から分析し、ヘテロスケダスティックなStudent-t回帰により生成ポリシーに匹敵する性能をより高速に実現した。

詳しい要約

1. どんなもの?

- ロボット学習における policy learning の手法。 - 従来の MSE-Policies と Flow-Policies の性能差を統計モデリングの観点から再検討。 - 実世界のロボットデモデータの action-prediction residuals を分析。 - 残差の state-dependent なスケール変動と heavy-tailed 性を発見。 - この知見に基づき heteroscedastic Student-t action regression (HT-Policies) を提案。 - 単一の feed-forward pass で action chunks を予測し、pretrained flow-matching-based policy networks を backbone として再利用可能。

2. 先行研究と比べてどこがすごい?

- 従来、Flow-Policies が MSE-Policies より優れる理由は multimodal demonstrations とされてきた。 - 本研究は残差の統計的性質に着目し、MSE が大きな action residuals を持つ観測に過大な勾配を割り当て最適化を損なうことを示した。 - HT-Policies は heavy tails の影響を低減し、生成モデルベースラインと競合する成功率を達成。 - 訓練と推論がより高速である点が優位。 - 事前学習済み vision-language-action や world-action モデルからも訓練可能。

3. 技術・手法の肝は?

- heteroscedastic Student-t action regression (HT-Policies) を導入。 - 入力依存の residual scales を学習。 - Student-t 分布により heavy tails の影響を低減。 - 単一の feed-forward pass で action chunks を予測。 - pretrained flow-matching-based policy networks を backbone として再利用可能。

4. どうやって有効だと検証した?

- 4つの simulation benchmarks と real-robot evaluations で評価。 - ゼロから訓練した場合と、pretrained vision-language-action および world-action モデルから訓練した場合の両方で検証。 - 生成モデルベースラインと競合する成功率を達成。 - 訓練と推論がより高速であることを確認。

5. 議論はある?

- 生成目的関数の実用的優位性に光を当てた。 - 効率的な直接回帰の代替手段を提供。 - 多様なアーキテクチャとタスクに適用可能。 - 具体的な限界や議論の詳細は要旨からは不明。

6. 次に読むべき論文は?

- Flow-Policies (diffusion or flow matching) - MSE-Policies - pretrained vision-language-action models - world-action models - heteroscedastic Student-t action regression (HT-Policies)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuchen Zhou, Jiacheng You, Weikang Wan, Weijun Dong, Yang Gao, Jiayuan Mao

分類: cs.RO

原文アブストラクト

Learning from demonstration has enabled impressive robot behaviors. A common choice for policy learning is to use diffusion or flow matching (Flow-Policies), which often outperforms direct action regression trained with mean squared error (MSE-Policies). This gap is commonly attributed to multimodal demonstrations. We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization. Our analysis of real-world robot demonstration data reveals substantial state-dependent variation in residual scales and heavier-than-Gaussian tails. While both MSE-Policies and Flow-Policies exhibit heavy-tailed action residuals, their training gradients behave differently: MSE allocates more gradient magnitude to observations with large action residuals, which hurts optimization. Motivated by these findings, we introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails. HT-Policies predict action chunks with a single feed-forward pass and can reuse pretrained flow-matching-based policy networks as the backbone. Across four simulation benchmarks and real-robot evaluations, HT-Policies achieves success rates competitive with generative policy baselines, both when trained from scratch and from pretrained vision-language-action and world-action models, despite being faster in training and inference. Together, these findings shed light on the practical advantages of generative objectives in robot learning from demonstrations and offer an efficient direct-regression alternative for a range of architectures and tasks. Project page: https://the-labone.github.io/regression-policy-project/

関連論文

PR本紙発行元 EmplifAI