日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.39971

指示が軌道を呼び出すとき:VLAモデルの汎化失敗の診断と緩和

When Instructions Retrieve Trajectories: Diagnosing and Mitigating Generalization Failures in VLA Models

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルが言語と視覚を組み合わせて行動を選べず失敗する「指示-行動バインディング」問題を分析し、同指示で異なる行動を要するデータと損失で訓練するECTを提案した。

詳しい要約

1. どんなもの?

- 本論文は、Vision-Language-Action (VLA) モデルにおける新たな失敗モード「instruction-action binding」を診断し、その緩和手法を提案する。 - この失敗は、言語と視覚の両方に反応するが、それらを組み合わせてタスクに必要な行動を選択できない現象である。 - 具体的には、指示が既知のtrajectory familyを手がかりにし、視覚フィードバックがその実行を調整するが、反事実的変化(異なる行動を要求する変化)に対応できない。 - 集約的なロバスト性スコアでは隠蔽される、より特定的な失敗である。 - 微調整されたπ_{0.5}とGR00T-N1.7ポリシーの行動分析により、失敗したロールアウトは元の行動を保持するか、別の実演タスクに切り替わることが明らかになった。

2. 先行研究と比べてどこがすごい?

- 従来のVLAモデルは、分布内タスクで90%以上の成功率を達成し、必要な行動を保存するnuisance変化には耐性があるとされてきた。 - しかし、異なる行動を要求する反事実的変化には失敗する。 - 集約的なロバスト性スコアはこの特定的失敗を隠蔽するため、先行研究では見過ごされてきた。 - 本研究は、この失敗を「instruction-action binding」と名付け、そのメカニズムを明らかにし、対策を提案する点で新しい。 - また、模倣目的の分析から、狭い条件付き行動サポートが実演上でgrounded solutionとinstruction-keyed solutionを区別不能にすることを示した。

3. 技術・手法の肝は?

- 技術の肝は、Equivariant Counterfactual Training (ECT) である。 - ECTは二つのレベルで作用する:データレベルと損失レベル。 - ECTデータは、同じ指示が異なる行動を要求する識別可能なシーンでの有効な実演を提供する。 - ECT損失は、各実演を同じ更新内でそのカウンターパートと共に訓練する。 - これにより、モデルは指示と行動の結合を学習し、反事実的変化に対応できるようになる。

4. どうやって有効だと検証した?

- 制御されたLIBERO-PRO比較において、完全なECTはπ_{0.5}の平均position-swap成功率を36%から59%に向上させた。 - CALVINでは、カウンターパートが元のデータに既に存在する場合、ECT損失は新しい実演なしで5タスク完了を改善した。 - 実機UR5eで固定実演予算の下、完全なECTは未見位置での成功率を8%から88%に向上させた。 - これらの実験により、ECTの有効性が検証された。

5. 議論はある?

- 行動分析と介入により、失敗したロールアウトが元の行動を保持するか別の実演タスクに切り替わることが示され、言語が単に無視されているわけではないことが明らかになった。 - 読み出しと介入は、これらの選択をタスク条件付き内部状態に結びつける。 - 模倣目的の分析は、狭い条件付き行動サポートが実演上でgrounded solutionとinstruction-keyed solutionを区別不能にすることを示す。 - これにより、instruction-action bindingの失敗が説明される。 - 議論の詳細や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照されている研究:π_{0.5}、GR00T-N1.7、LIBERO-PRO、CALVIN、UR5e。 - 関連手法:Equivariant Counterfactual Training (ECT)。 - 同分野の定番:Vision-Language-Action (VLA) モデル、模倣学習、ロバスト性評価。 - 次に読むべき論文としては、これらのモデルやベンチマークに関する原論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee

分類: cs.RO, cs.LG

原文アブストラクト

Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fail under counterfactual changes that demand a different action. Aggregate robustness scores can therefore conceal a more specific failure, in which a policy responds to both language and vision yet does not combine them to select the action the task requires. We call this failure instruction-action binding. Instructions cue familiar trajectory families, and visual feedback adjusts their execution. Behavioral analyses of fine-tuned $π_{0.5}$ and GR00T-N1.7 policies reveal that failed rollouts often retain the source behavior or switch to another demonstrated task. These switches show that language is not simply ignored. Readouts and interventions connect these choices to task-conditioned internal states. Our analysis of the imitation objective shows how narrow conditional action support can leave grounded and instruction-keyed solutions indistinguishable on the demonstrations. This motivates Equivariant Counterfactual Training (ECT), which acts at two levels. ECT data supply valid demonstrations in which the same instruction requires different actions in distinguishable scenes, while the ECT loss trains each demonstration with its counterpart in the same update. In a controlled LIBERO-PRO comparison, full ECT raises $π_{0.5}$'s mean position-swap success from 36% to 59%. On CALVIN, where counterparts already occur in the original data, the ECT loss improves five-task completion without new demonstrations. On a real UR5e under a fixed demonstration budget, full ECT raises unseen-position success from 8% to 88%.

関連論文

PR本紙発行元 EmplifAI