日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30783

会話は不要、視覚に再注力:マルチモーダル大規模言語モデルにおける推論セグメンテーションのための潜在推論

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

シェア:XThreadsFacebookLINEはてブBluesky

明示的なChain-of-Thoughtの代わりに学習可能な潜在トークンを用いて推論セグメンテーションを行うLIRSegを提案し、セグメンテーション精度と推論効率を向上させた。

詳しい要約

1. どんなもの?

- 暗黙的なテキストクエリを解釈し、きめ細かい視覚知覚を可能にするReasoning Segmentationのための手法。 - 既存手法はMLLMsによる明示的なChain-of-Thought (CoT)を生成してから対象を特定するが、冗長なテキストトークンが注意を妨害し、視覚トークン間の実効距離を増大させる問題がある。 - 提案手法LIRSegは、明示的CoTをコンパクトな学習可能なlatent tokensの集合に完全に置き換える。 - 2段階で訓練:空間アラインメントでlatent tokensを物体関連の視覚証拠に接地し、GRPOでセグメンテーション報酬を用いて最適化。 - 情報理論的観点から3つのメカニズム(extreme-advantage sampling、decoupled exploration-stability updates、latent diversity amplification)を導入。

2. 先行研究と比べてどこがすごい?

- 既存手法は明示的CoTを生成するため、冗長なテキストトークンが注意を妨害し、視覚トークン間の実効距離を増大させる問題があった。 - LIRSegは明示的CoTを完全に排除し、コンパクトなlatent tokensで推論することで、注意干渉を回避し、推論効率を大幅に向上。 - VisionReasonerベースラインと比較して、ReasonSegでgIoUが4.9%絶対改善、MUSEで7.1%、MMRで4.7%改善。 - 推論トークン数を約16倍削減。

3. 技術・手法の肝は?

- 明示的CoTをコンパクトな学習可能なlatent tokensの集合に置き換える。 - 2段階訓練: - 空間アラインメント:latent tokensを物体関連の視覚証拠に接地。 - GRPO:セグメンテーション報酬でlatent tokensを最適化。 - 情報理論的観点から3つのメカニズム: - extreme-advantage sampling:情報量の多い訓練信号を選択。 - decoupled exploration-stability updates:補完的な表現を学習。 - latent diversity amplification:表現の崩壊を防止。

4. どうやって有効だと検証した?

- ベンチマーク(ReasonSeg、MUSE、MMR)での広範な実験。 - VisionReasonerベースラインと比較して、gIoUがReasonSegで4.9%、MUSEで7.1%、MMRで4.7%絶対改善。 - 推論トークン数を約16倍削減。 - セグメンテーション精度と推論効率の両方を一貫して改善することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- VisionReasoner(ベースラインとして比較) - GRPO(最適化手法として使用) - Chain-of-Thought (CoT) を用いた既存のReasoning Segmentation手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan

分類: cs.CV, cs.AI

原文アブストラクト

Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before localizing the target. Although intuitive, such explicit verbal reasoning introduces substantial attention interference: redundant textual tokens disrupt attention during perception-token generation and also increase the effective distance between visual tokens. To address this issue, we propose LIRSeg, which fully replaces explicit CoT with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages: spatial alignment grounds the latent tokens in object-relevant visual evidence, and GRPO further optimizes them with segmentation rewards. To make these compact latent tokens more informative, we introduce three complementary mechanisms from an information perspective: extreme-advantage sampling for selecting informative training signals, decoupled exploration-stability updates for learning complementary representations, and latent diversity amplification for preventing representational collapse. Extensive experiments on benchmarks demonstrate that LIRSeg consistently improves both segmentation accuracy and reasoning efficiency. Compared with the VisionReasoner baseline, LIRSeg achieves absolute gIoU improvements of 4.9% on ReasonSeg, 7.1% on MUSE, and 4.7% on MMR, while achieving a approximately 16x reduction in reasoning tokens. Code is available in supplementary materials.

関連論文

PR本紙発行元 EmplifAI