日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30996

視覚言語行動モデルにおける線形表現仮説

The Linear Representation Hypothesis for Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語行動モデル(VLA)において、物理量の未来の変化を線形に読み出し・操作できる表現が存在することを理論的に示し、ナビゲーション実験で検証した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルに対する Linear Representation Hypothesis (LRH) の理論的定式化を提案。 - 身体性相互作用の動的性質を考慮し、物理量 (QoI) の表現と政策を統合。 - 表現側では線形プロービングによる将来の QoI 予測、政策側では signature generalized linear model による線形ステアリングを可能にする。

2. 先行研究と比べてどこがすごい?

- 従来の LRH は LLM の意味的属性(性別や言語など)に焦点を当てていた。 - 本研究は VLA における動的な QoI に拡張し、表現と政策の統合を図る点が新しい。 - 身体性相互作用の課題に対処する理論的枠組みを提供。

3. 技術・手法の肝は?

- signature-based 定式化により、表現と政策を統合。 - 表現側: 候補行動軌跡下での QoI の将来進化を線形プロービングで復元可能な表現の存在を確立。 - 政策側: 確率的行動チャンクに対する signature generalized linear model を導入。 - 自然パラメータ空間の線形経路に沿って期待将来 QoI が単調変化し、線形ステアリングを実現。

4. どうやって有効だと検証した?

- 平面制御アフィン航法実験において、明示的なオラクル表現を構築。 - 予測された線形プロービングとステアリングメカニズムを検証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として Linear Representation Hypothesis (LRH) や Vision-Language-Action (VLA) モデル、signature generalized linear model が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han

分類: cs.LG, cs.AI

原文アブストラクト

The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in a VLA evolves jointly with the system dynamics: the representation influences the actions selected by the policy, which alter the physical state and, in turn, the next representation. In this paper, we develop a theoretical, signature-based formulation of the LRH for VLA that unifies representations and policies. On the representation side, we establish the existence of representations from which the future evolution of a QoI under a candidate action trajectory can be recovered via linear probing. On the policy side, we introduce a signature generalized linear model for stochastic action chunks. This structure yields a monotonic change in the expected future QoI along linear paths in natural parameter space, enabling linear steering. We construct an explicit oracle representation in a planar control-affine navigation experiment and verify the predicted linear probing and steering mechanisms.

関連論文

PR本紙発行元 EmplifAI