日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLMarXiv:2609.35486

誰が誰の左にいるのか?相対位置推論における空間的証拠と役割バインディングの追跡

Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語モデルが物体の相対位置を推論する際、入力内の位置情報の追跡とクエリ内の役割表現という2つの要素を活性化パッチングで分析し、介入により精度と一貫性を改善できることを示した。

詳しい要約

1. どんなもの?

- 相対位置推論における物体の位置追跡とクエリ役割の内部表現を調査 - 3つのVLM(視覚・テキスト入力)とその言語モデルバックボーンを対象 - 高精度でも位置交換や役割反転で一貫性が崩れる問題を扱う - activation patchingと介入で因果的役割を解明

2. 先行研究と比べてどこがすごい?

- 従来はインスタンス精度が重視され、内部表現の理解が不十分 - 位置情報とクエリ役割の2成分を分離して追跡 - 段階的処理(初期→中間→後期)を因果的に示した点が新しい - 合成シーンでの介入が自然画像ベンチマークに汎化

3. 技術・手法の肝は?

- activation patchingで層ごとの表現寄与を分析 - 初期層のsource表現→中間層のquery-object表現→後期層のanswer状態へ - source側操作で位置情報と関係予測が変化する因果リンクを確立 - クエリ側の役割方向を推定し、steeringで一貫性改善

4. どうやって有効だと検証した?

- 3つのVLMとLMバックボーンでactivation patchingを実施 - 合成シーンで推定した方向を自然画像ベンチマークに適用 - 再学習なしで精度とペア一貫性の両方を改善 - 因果的介入により段階的進行を検証

5. 議論はある?

- 位置追跡と役割表現が相補的であることを示唆 - 視覚・テキスト設定で共通のメカニズムが存在 - 介入による一貫性改善の可能性を提示 - 限界や未解決点は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法としてactivation patching, steering, VLM, 相対位置推論 - 同分野の定番としてVisual Question Answering, spatial reasoning benchmarks

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yingjin Song, Denis Paperno, Albert Gatt

分類: cs.CL, cs.CV

原文アブストラクト

High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models' behavior.

関連論文

PR本紙発行元 EmplifAI