日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
視覚的注意arXiv:2609.31364

OpenVAM: VLMによるオープンワールド視覚的注意モデリング

OpenVAM: Open-World Visual Attention Modeling with VLMs

シェア:XThreadsFacebookLINEはてブBluesky

視覚的注意の予測において、密な顕著性マップだけでなく、注視点が何を指すか・なぜ注目されるかを説明可能にするVLMベースの統合フレームワークを提案。

詳しい要約

1. どんなもの?

- 人間の視線予測(visual attention modeling)を行うフレームワーク - 従来はdense saliency mapのみを出力 - OpenVAMはsaliencyに加え、注視ピークに対応する離散要素(what)とその理由(why)を説明 - 自然画像、商業画像、UI/webレイアウトなど異種ドメインを統合 - VLM(vision-language model)を活用し、汎用性と説明可能性を両立

2. 先行研究と比べてどこがすごい?

- 従来手法はdense saliency mapのみで、actionに結びつくwhat/whyが得られない - ドメインシフト(自然画像→商業→UI/web)への頑健性が不十分 - OpenVAMは異種ドメインと多様な監督モダリティを統合し、汎用性と説明可能性を同時に実現 - 説明を画像にグラウンディングすることでsaliency予測の解釈性を向上

3. 技術・手法の肝は?

- decoupled-but-aligned設計:dense visual pathwayが空間的に正確な定位を提供 - instruction-following vision-language semantic headが同一画像とdata-typeプロンプトに条件づけられ、what/why説明を生成 - 3段階学習戦略:定位priorを保持しつつ、言語グラウンディングを段階的に導入 - parameter-efficient adaptationでsaliency branchを乱さず説明アラインメントを改善 - 多ドメインsaliency-reasonアノテーションを生成するスケーラブルパイプラインを提案

4. どうやって有効だと検証した?

- 多様なデータセットでの実験を実施 - ドメインシフト下での頑健性向上を確認 - 画像にグラウンディングされた説明がsaliency予測の解釈性を高めることを示す - 具体的なデータセット名や評価指標は要旨からは不明

5. 議論はある?

- 異種ドメインと監督モダリティの統合における課題 - 説明の質と定位精度のトレードオフ - parameter-efficient adaptationの限界 - 多ドメインアノテーション生成パイプラインのスケーラビリティ - 具体的な議論の詳細は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 同分野の定番としてsaliency prediction(例:SalGAN, DeepGaze)、VLM(例:CLIP, BLIP)、visual grounding(例:GLIP)などが挙げられる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati

分類: cs.CV

原文アブストラクト

Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elements in the scene (what) and understand the drivers of those peaks in context (why), while remaining robust to domain shift across natural images, commercial content, and UI/web layouts. We, therefore, introduce OpenVAM (Open-world Visual Attention Modeling with VLMs), a unified framework that jointly addresses universality and explainability across heterogeneous domains (natural scenes, commercial imagery, and UI/web layouts) and supervision modalities. OpenVAM adopts a decoupled-but-aligned design: a dedicated dense visual pathway provides stable, spatially precise localization, while an instruction-following vision--language semantic head generates grounded what/why explanations conditioned on the same image and data-type prompts. A three-stage training strategy preserves strong localization priors while progressively introducing language grounding and improving explanation alignment via parameter-efficient adaptation without perturbing the saliency branch. We further propose a scalable pipeline to generate multi-domain saliency-reason annotations for training and systematic evaluation. Experiments across diverse datasets show that OpenVAM improves robustness under domain shift while producing image-grounded explanations that make saliency predictions more interpretable.

関連論文

PR本紙発行元 EmplifAI