日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09496

スパース特徴のポリシー忘却による視覚言語行動モデルの状態幻覚の軽減

Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルが未達成の状態を達成したかのように振る舞う「状態幻覚」を、幻覚に関連するスパース特徴を選択的に忘却させることで軽減する手法SOULを提案。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルにおける state hallucination という失敗パターンを研究 - state hallucination とは、実現していない robot-object state が達成されたかのように VLA が行動し続ける現象 - この失敗は task-relevant visual regions への attention 低下と関連 - sparse autoencoders による解釈で、hallucination 関連の sparse features が失敗時に活性化することを発見 - 提案手法 SOUL (Sparse feature pOlicy UnLearning) は、state hallucination に関連する policy knowledge を選択的に unlearn する - hallucination 失敗から同定した sparse features を forgetting target、成功行動の features を retention target として使用 - シミュ…

2. 先行研究と比べてどこがすごい?

- 従来の VLA モデルは pretrained vision-language models の豊富な表現を活用しロボット manipulation で強い汎化を示すが、実環境展開は繰り返す信頼性の低い行動に制限される - 本研究は state hallucination という特定の失敗パターンに焦点を当て、そのメカニズムを attention と sparse autoencoders で解明 - 先行研究と比べ、解釈可能な feature 分析に基づき、望ましくない知識を選択的に修正する実用的基盤を提供 - 具体的な比較対象は要旨からは不明

3. 技術・手法の肝は?

- sparse autoencoders を用いた mechanistics 解釈により、hallucination 関連の sparse features を同定 - 同定した features を forgetting target として、state hallucination 行動に関連する policy knowledge を選択的に unlearn - 同時に、成功行動から得た sparse features を retention target として設定し、既存の manipulation 能力を保持 - SOUL (Sparse feature pOlicy UnLearning) と名付けたこの手法で、VLA の policy を選択的に修正

4. どうやって有効だと検証した?

- シミュレーション環境と実世界環境の両方で、複数の VLA アーキテクチャを対象に実験 - 結果、hallucinated failures が大幅に減少し、タスク成功率が向上 - 既存の manipulation 能力に大きな悪影響がないことを確認 - 具体的な評価指標やデータセットは要旨からは不明

5. 議論はある?

- 解釈可能な feature 分析が、ロボット policy における望ましくない知識を選択的に修正する実用的基盤となることを示唆 - state hallucination が task-relevant visual regions への attention 低下と関連することを発見 - 限界や今後の課題については要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として sparse autoencoders を用いた解釈可能性研究や VLA モデル全般が挙げられる - 同分野の定番として Vision-Language-Action (VLA) モデル、sparse autoencoders、policy unlearning に関する論文を読むべき

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiho Lee, Jeongeun Park, Heayoun Choi, Taekyung Kim, Eunwoo Kim

分類: cs.RO, cs.AI, cs.LG

原文アブストラクト

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.

関連論文

PR本紙発行元 EmplifAI