日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.06235

言語は符号化されるが制御に至らない:視覚言語ロボット政策における接地ギャップの解明

Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語行動政策が指示を符号化しても、シーンに基づく元の対象選択を変えられない「接地ギャップ」を、指示介入実験と注意分析により明らかにした。

詳しい要約

1. どんなもの?

- 言語条件付きロボット政策における指示追従の曖昧さを研究 - scene-preserving instruction interventions を用いて、シーンを固定したまま有効な対象置換、任意名詞、無関係文を提示 - VLA policies と WAMs をシミュレーションと実世界で評価 - タスク成功が指示追従を意味するか、言語感度、符号化された言語が行動を制御しない理由を分析

2. 先行研究と比べてどこがすごい?

- 従来はタスク成功率で政策を評価しがちだが、成功が指示追従を保証しないことを示す - シーン固定の指示介入により、シーン由来の推論と指示追従を分離 - 中間行動予測や線形プローブ、注意分析、UMAP、NMF を組み合わせ、grounding gap を可視化 - 名目上の成功に隠れた問題を診断する枠組みを提供

3. 技術・手法の肝は?

- scene-preserving instruction interventions: 有効な対象置換、任意名詞、無関係文をシーン固定で提示 - VLA policies と WAMs をシミュレーションおよび実世界で評価 - layerwise action lens で中間行動予測の指示摂動への応答を観察 - linear probes で指示対象の復元を試み、符号化の有無を検証 - attention analysis で指示トークンの行動生成への寄与を評価 - UMAP と shared non-negative matrix factorization で表現構造を分析

4. どうやって有効だと検証した?

- シミュレーションと実世界の両方で VLA policies と WAMs を評価 - 指示が異なる可視物体を要求する場合、全政策がシーンに関連する元の対象に接近・把持 - 指示摂動がタスク性能に影響し、中間行動予測が応答 - linear probes が指示対象を正確に復元 - attention 分析で指示トークンの寄与が弱く、対象トークンの注意が元の物体に留まることを確認 - UMAP と NMF で対象情報がシーン同一性に組織化された表現内に残ることを示す

5. 議論はある?

- タスク成功が指示追従を意味しないという grounding gap を明示 - 言語は符号化されるが行動選択を支配しない - 対象符号化と視覚的 grounding の不一致が原因 - 名目上の成功に隠れた問題を診断する枠組みを提供 - 進歩の基準として、シーンに好まれる行動と衝突しても有効な意図変更に確実に従うことを提案

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として vision-language-action (VLA) policies、world-action models (WAMs)、linear probes、attention analysis、UMAP、shared non-negative matrix factorization が挙げられる - 同分野の定番として language-conditioned robot policies、instruction following、visual grounding に関する研究を読むべき

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Shaohan Jiang, Jiahang Cao, Qiduo He, Fengting Deng, Kun Wu, Jingkai Sun, Jiaxu Wang, Qiang Zhang, Qihao Zheng, Chunfeng Song, Ping Luo, Andrew F. Luo

分類: cs.RO, cs.LG

原文アブストラクト

Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through scene-preserving instruction interventions, using valid target substitutions, arbitrary nouns, and unrelated sentences while holding the scene fixed. We evaluate vision-language-action (VLA) policies and world-action models (WAMs) in simulation and in real-world experiments. Our analysis addresses three questions: (a) Does task success imply instruction following? When instructions request a different visible object, all evaluated policies predominantly approach and pick up the incorrect original target associated with the scene. (b) Is this failure caused by language insensitivity? Instruction perturbations affect task performance. A layerwise action lens shows intermediate action predictions respond to these perturbations. Linear probes accurately recover instructed targets, indicating modified instructions are encoded despite rarely determining target selection. (c) Why does encoded language fail to control action? Attention analysis indicates weak instruction-token contributions to action generation. Target-token attention can remain focused on the original object, revealing a mismatch between target encoding and visual grounding. UMAP and shared non-negative matrix factorization show target information remains accessible within representations increasingly organized by scene identity. Our findings expose a grounding gap concealed by nominal success and provide a diagnostic framework. They further establish a concrete criterion for progress: policies should reliably follow valid changes in user intent, even when they conflict with scene-favored behavior.

関連論文

PR本紙発行元 EmplifAI