日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04616

PerturBot: 摂動的学習で視覚言語行動モデルのショートカット事前分布を打破する

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

シェア:XThreadsFacebookLINEはてブBluesky

VLAポリシーがタスクに無関係な手がかり(モダリティショートカット)に依存する問題に対し、手首視点の摂動や指示文の拡充、失敗軌跡の再ラベル付けで学習データを再構成する手法と、ショートカット依存度を測る指標GroundingFscoreを提案した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) ポリシーが、行動を決めるべき証拠を無視して成功デモの規則性(modality shortcuts)に依存する問題を扱う。 - 視覚・語彙・運動の各モダリティでショートカットが生じ、成功デモを増やしても解消しないことを指摘。 - 提案手法 Perturbot は、タスク保存の wrist-view 摂動、decision-relevant captions による指示の強化、ランダム/失敗軌跡の再ラベル付けを組み合わせる。 - 併せて GroundingFscore を提案し、ポリシーがショートカットに依存する度合いをオフラインで診断する。

2. 先行研究と比べてどこがすごい?

- 従来は同一種類のデモを増やす scaling が主流で、タスク成功率は上がっても modality shortcuts は残ると指摘。 - Perturbot は『何をスケールするか』を変えることで scaling を補完し、推論時は変更しない点が新しい。 - 成功率だけでなく GroundingFscore で健全なスケーリング(タスク証拠への依存)を評価する枠組みを提供。 - 要旨からは、比較対象となる具体的な先行研究名は不明。

3. 技術・手法の肝は?

- タスクを保存する wrist-view 摂動を加え、タスク関連証拠を使いやすくする。 - 指示文に decision-relevant captions を付与して言語的手がかりを強化。 - ランダムおよび失敗軌跡のセグメントを、その中に含まれる振る舞いで再ラベル付けして学習に加える。 - これによりショートカットだけでは不十分にし、推論時の変更は不要。 - GroundingFscore はオフラインでショートカット依存度を診断する指標。

4. どうやって有効だと検証した?

- 要旨からは、具体的な実験設定・データセット・ベースライン・定量結果は不明。 - 提案手法 Perturbot と GroundingFscore を組み合わせた訓練・評価フレームワークを提示している。 - タスク成功率と GroundingFscore の両面から有効性を検証する方針が述べられている。

5. 議論はある?

- modality shortcuts は成功デモの規則性に起因し、デモを増やしても残る可能性がある。 - タスク成功率だけではポリシーの健全なスケーリングを判断できない。 - GroundingFscore によりショートカット依存を診断できるが、要旨からは限界や失敗事例の議論は不明。 - 推論時は変更しないため、既存の VLA パイプラインに適用しやすい可能性がある。

6. 次に読むべき論文は?

- 要旨で参照・比較されている具体的な研究は不明。 - 同分野の定番として、Vision-Language-Action (VLA) モデル(例: RT-1, RT-2, OpenVLA など)や、imitation learning、behavior cloning、distribution shift に関する研究が次に読むべき候補。 - また、shortcut learning や causal confusion に関する一般文献も関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mingyu Liu, Chonghao Sima, Tianjian Feng, Hanqing Wang, Cong Chen, Hao Chen, Chunhua Shen

分類: cs.RO, cs.CV

原文アブストラクト

A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training even after the verb changes, and a gripper that closed on nothing may lift anyway. We call these dependencies modality shortcuts: regularities in successful demonstrations make visual, lexical, or motor cues sufficient to predict expert actions without the task evidence needed for the underlying decision. More demonstrations of the same kind can raise task success while leaving these shortcuts intact. We propose Perturbot which makes task-relevant evidence easier to use and shortcuts insufficient on their own: it applies task-preserving wrist-view perturbations, enriches instructions with decision-relevant captions, and adds random and failed trajectory segments relabeled with the behavior they contain. It complements scaling by changing what is scaled, and leaves inference unchanged. Moreover, we propose GroundingFscore, an offline score that diagnoses how severely a policy relies on modality shortcuts. Task success rate shows whether a policy improves, while GroundingFscore reveals whether the policy scales healthily, relying on task evidence rather than shortcuts. Together, Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.

関連論文

PR本紙発行元 EmplifAI