日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34061

Vision-Language-Actionモデル向け分位点ヘッド

Quantile Head for Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルの行動ヘッドとして、回帰とフローマッチングを統一的に扱う分位点目的を導入し、1回の順伝播で分位点を予測するQuantile Headを提案。LIBERO系ベンチマークと実機タスクで成功率と効率を改善した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデル向けの新しい action head である Quantile Head を提案する研究。 - 従来の action head の限界(point regression は点推定のみ、standard flow-matching は反復サンプリングが必要)を解決する。 - regression と flow matching を共通目的関数で統合し、quantile 目的関数を導出する。 - 1回の forward pass で median と正の gap を予測し、順序付きの marginal action quantiles を構成する。 - 再学習なしで複数のサンプリング戦略をサポートし、joint supervision でデフォルトの median policy を訓練する。

2. 先行研究と比べてどこがすごい?

- 従来の point regression は action 分布の点推定しか得られない。 - standard flow-matching sampler は計算コストの高い反復サンプリングを必要とする。 - 提案手法は regression と flow matching を統合した quantile 目的関数により、1回の forward pass で ordered marginal action quantiles を予測できる。 - 再学習なしで複数のサンプリング戦略を利用可能。 - 実験では LIBERO、LIBERO-Plus、LIBERO-Pro、実機2タスクで比較手法中最高の平均成功率を達成し、matched LIBERO baselines 中で最短の平均 episode time を示す。

3. 技術・手法の肝は?

- regression と flow matching を共通の目的関数の下で統合し、quantile 目的関数を導出する。 - Quantile Head は median と正の gap を予測し、順序付きの marginal action quantiles を1回の forward pass で構成する。 - これらの quantiles は再学習なしで複数のサンプリング戦略をサポートする。 - joint supervision によりデフォルトの median policy を訓練する。 - 局所解析により、calibrated nearby quantiles、fixed gaps、matched correction speed の条件下で、direct median updates が median-only supervision より低分散であることを示す。

4. どうやって有効だと検証した?

- LIBERO、LIBERO-Plus、LIBERO-Pro、および実機2タスクで実験を実施。 - 比較手法の中で最高の平均成功率を達成。 - matched LIBERO baselines の中で最短の平均 episode time を示す。 - コードは https://github.com/xwangrs/Quantile-Head-for-VLA で公開。

5. 議論はある?

- 局所解析により、calibrated nearby quantiles、fixed gaps、matched correction speed の条件下で direct median updates が median-only supervision より低分散であることを示す。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: point regression、standard flow-matching samplers。 - 関連手法: Vision-Language-Action (VLA) モデル、Vision-Language Models (VLMs)、flow matching、quantile regression。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xuan Wang, Yinan Wu, Haoran Duan, Jungong Han

分類: cs.RO, cs.CV

原文アブストラクト

Vision-Language-Action (VLA) models integrate pretrained Vision-Language Models (VLMs) with action heads for robot control. Common action heads have distinct limitations: point regression provides only a point estimate of the action distribution, while standard flow-matching samplers require costly iterative sampling. To address these limitations, we unify regression and flow matching under a shared objective and extend it to derive a quantile objective. This quantile objective guides the design of our Quantile Head, which predicts a median and positive gaps to form ordered marginal action quantiles in one forward pass. These quantiles support multiple sampling strategies without retraining and are jointly supervised to train the default median policy. Our local analysis of this joint supervision shows that, with calibrated nearby quantiles, fixed gaps, and matched correction speed, direct median updates have lower variance than under median-only supervision. Experiments show that this jointly supervised median policy achieves the highest average success rates among the compared methods on LIBERO, LIBERO-Plus, LIBERO-Pro, and two real-robot tasks, together with the shortest mean episode time among matched LIBERO baselines; code is available at https://github.com/xwangrs/Quantile-Head-for-VLA.

関連論文

PR本紙発行元 EmplifAI