日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10178

Vision-Language-Actionモデルは指示を理解しているのか?言語グラウンディングに関するメカニスティック解釈可能性研究

Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルが言語指示に実際に依存しているかを調べるため、π0.5とGR00T N1.7の残差ストリームに活性化・帰属パッチングを適用し、指示を様々に破壊した際の感度を分析した。

詳しい要約

1. どんなもの?

- Vision-Language-Action models (VLAs) の language grounding を調べる研究。 - 対象は $π_{0.5}$ と GR00T N1.7 の2つの state-of-the-art VLAs。 - 行動生成が言語指示に依存するか、視覚的手がかりや表面的相関に依存するかを検証。 - LIBERO benchmark の入力サンプルの task instruction を5戦略で corrupt。 - activation patching と attribution patching を action generation modules の residual stream に適用。

2. 先行研究と比べてどこがすごい?

- VLAs は明示的な grounding module を持たず、Vision-Language model backbone の内在的 grounding 能力に依存。 - 従来はこの grounding 能力のメカニズム解明が不十分。 - 本研究は mechanistic interpretability を VLA の language grounding に適用。 - 2モデルを制御実験で比較し、感度の locus や因果効果の違いを明らかに。 - attribution patching の信頼性がモデル依存であることを示す。

3. 技術・手法の肝は?

- activation patching と attribution patching を action generation modules の residual stream に適用。 - task instruction を5戦略で corrupt: synonym replacement, semantic scaling, directional corruption, random object substitution, empty string。 - LIBERO benchmark の入力サンプルを使用。 - 各モデルの感度を層ごとに分析。 - GR00T N1.7 では early, periodic cross-attention layers に集中、$π_{0.5}$ では earliest と selected later layers に分散。

4. どうやって有効だと検証した?

- LIBERO benchmark 上で2モデルの language grounding を系統的に評価。 - 5つの corrupt 戦略に対する感度を比較。 - 抽象的な言い換えや存在しない物体への参照には比較的鈍感。 - 空の task description や directional language には強く反応。 - GR00T N1.7 では directional perturbations が最大の因果効果を生むが内部表現幾何は比較的変化しない dissociation を観察。 - attribution patching は GR00T N1.7 では activation patching とよく一致するが $π_{0.5}$ では一致しない。

5. 議論はある?

- 両モデルは抽象的言い換えや非存在物体参照に鈍感。 - 空文字列や directional language に強く反応。 - 感度の locus はモデル間で異なる。 - GR00T N1.7 で見られる因果効果と表現幾何の dissociation は $π_{0.5}$ では明確でない。 - attribution patching の信頼性はモデル依存。 - 実世界展開には視覚・言語観測空間の分散への頑健性が重要。

6. 次に読むべき論文は?

- $π_{0.5}$ の論文 - GR00T N1.7 の論文 - LIBERO benchmark の論文 - activation patching の原論文 - attribution patching の原論文 - Vision-Language-Action models の survey

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Theodor Wulff, Angelo Cangelosi

分類: cs.RO

原文アブストラクト

Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations. Robustness to variance in the visual and linguistic observation space is critical for real-world deployment, yet VLAs lack explicit grounding modules and instead rely on the intrinsic language grounding capabilities of their Vision-Language model backbones. For this reason, we conduct a controlled mechanistic interpretability study on the language grounding capabilities of two state-of-the-art Vision-Language-Action models, $π_{0.5}$ and GR00T N1.7, by applying activation and attribution patching to the residual stream of the action generation modules. We systematically corrupt the task instruction of input samples of the LIBERO benchmark following five strategies: synonym replacement, semantic scaling, directional corruption, random object substitution, and empty string. Our experiments find that both models are comparatively insensitive to abstract rephrasing and to referencing non-existent objects, but react strongly to empty task descriptions and, especially, to directional language. During action generation, this sensitivity is concentrated in different loci for each model: mainly in the early, periodic cross-attention layers for GR00T N1.7, versus distributed across the earliest and selected later layers for $π_{0.5}$. For GR00T N1.7, directional perturbations drive some of the largest causal effects while leaving the internal representational geometry comparatively unchanged, a dissociation we do not observe clearly for $π_{0.5}$. Finally, the reliability of attribution patching is model-dependent: it closely tracks activation patching for GR00T N1.7 but not for $π_{0.5}$.

関連論文

PR本紙発行元 EmplifAI