日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04659

PermVLA: 分解順序を正則化として活用するVLA学習

PermVLA: Factorization Order as a Regularizer for VLA Learning

シェア:XThreadsFacebookLINEはてブBluesky

VLAポリシーの学習において、行動チャンクの分解順序を正則化として利用するCAPを提案し、LIBEROやCALVINで性能向上を実証した論文。

詳しい要約

1. どんなもの?

- Vision-language-action (VLA) ポリシーの学習において、action chunk の factorization order を正則化の選択肢として捉える研究。 - 従来は固定の left-to-right (LTR) factorization が一般的だが、同じ expert trajectory 分布は複数の有効な chain-rule factorization を持つ点に着目。 - causally anchored permutation (CAP) を導入し、chronological prefix を調整可能な形で action の reveal order をサンプリング。 - 展開時は決定的な LTR 制御を維持する。

2. 先行研究と比べてどこがすごい?

- 従来の VLA 学習は固定 LTR factorization による teacher forcing が標準であり、factorization order は見過ごされてきた正則化の選択肢であると指摘。 - CAP は追加の demonstration なしに、1つの expert chunk から複数の条件付き予測問題を生成する conditional-set augmentation を実現。 - 単一の chronological prefix への依存を減らす点が先行研究と異なる。 - LIBERO と LIBERO-Plus で標準 LTR 学習を一貫して上回り、cross-dataset CALVIN 評価でも同様の優位性を示す。

3. 技術・手法の肝は?

- causally anchored permutation (CAP): 調整可能な chronological prefix を持つ action reveal order をサンプリング。 - 補助目的関数により、同一 expert chunk の異なる既知部分集合から action を予測するよう単一の共有ポリシーを訓練。 - conditional-set augmentation: 1つの expert chunk から複数の条件付き予測問題を生成し、追加デモンストレーションを不要にする。 - 展開時は決定的な LTR 制御を保持。 - diffusion action generators への拡張も示す。

4. どうやって有効だと検証した?

- LIBERO と LIBERO-Plus での制御実験により、CAP が標準 LTR 学習を一貫して上回ることを確認。 - cross-dataset CALVIN 評価でも同様の優位性を確認。 - 2つの reveal order 間での chunk の joint log likelihood の期待二乗差を測る diagnostic により、CAP 訓練が reveal order 間の一致を内部化することを検証。

5. 議論はある?

- factorization order が VLA 学習における見過ごされた正則化選択肢であると位置づけ。 - サンプリングされた subset-conditioned 補助目的が VLA 正則化を構築する一般的なレシピとなり得ることを示唆。 - diffusion action generators への拡張例を提示。 - 具体的な限界や失敗ケース、計算コストなどの議論は要旨からは不明。

6. 次に読むべき論文は?

- LIBERO - LIBERO-Plus - CALVIN - diffusion action generators - 関連手法として VLA ポリシー、teacher forcing、chain-rule factorization に関する研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yanqiao Chen, Yuhan Rui, Dongsheng Hou, Zijie Nie, Yutong Wan, Qi Hao

分類: cs.AI, cs.RO

原文アブストラクト

Vision-language-action (VLA) policies commonly learn action chunks through a fixed left-to-right (LTR) factorization, although the same expert trajectory distribution admits many valid chain-rule factorizations. We identify factorization order as an overlooked regularization choice and introduce causally anchored permutation (CAP), which samples action reveal orders with a tunable chronological prefix. Its auxiliary objective trains one shared policy to predict actions from different known subsets of the same expert chunk, while deployment retains deterministic LTR control. We call this conditional-set augmentation: it creates multiple conditional prediction problems from one expert chunk without adding demonstrations. This discourages reliance on the single chronological prefix used by ordinary teacher forcing. Controlled experiments show that CAP consistently outperforms standard LTR training on LIBERO and LIBERO-Plus, with the same advantage appearing in cross-dataset CALVIN evaluation. A diagnostic that measures the expected squared difference between a chunk's joint log likelihood under two reveal orders verifies that CAP training internalizes agreement across reveal orders. These findings position sampled subset-conditioned auxiliary objectives as a general recipe for constructing VLA regularizers, illustrated by an extension to diffusion action generators.

関連論文

PR本紙発行元 EmplifAI