画像条件付きインスタンスプロンプトネットワークによる参照リモートセンシング画像セグメンテーション
Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation
リモートセンシング画像の参照セグメンテーションにおいて、画像条件付きインスタンスプロンプトと双方向情報融合を導入し、クロスモーダル特徴融合のボトルネックを解消する新しいネットワークを提案した。
著者: Biaoyu Ren, Qingsheng Wang, Cun Xu, Dingkang Yang, Wenxuan Wang
分類: cs.CV
原文アブストラクト
Referring Remote Sensing Image Segmentation (RRSIS) is a situated, task-driven cross-modal task related to the embodied perception paradigm, requiring models to align visual-spatial features with linguistic intentions for precise target perception. Recent research has focused on refining the granularity of textual features and optimizing image-text feature fusion to better guide target feature representations. However, insufficient descriptive granularity and sensitivity to semantic shifts can cause bottlenecks in cross-modal feature fusion. To address these issues, we propose the Image-Conditioned Instance Prompt Network (ICIPNet) with Bilateral Information Fusion, which is designed to alleviate bottlenecks in cross-modal feature fusion. ICIPNet introduces an Image-Conditioned Instance Prompt (ICIP) module to generate self-adaptive visual and semantic representations without external knowledge. The Bilateral Information Fusion (BIF) module enhances feature fusion along the token and channel dimensions. Experiments demonstrate that the proposed ICIPNet outperforms existing RRSIS models.
関連論文
- DropClick: 農業ロボットデータのための半自動ワンクリックセグメンテーションセグメンテーション
- UAV画像の雑然シーンにおける通信鉄塔部品のゼロショットセグメンテーションのための顕著性-深度条件付けセグメンテーション
- SOS!:モデルフリーセグメンテーションのための合理化されたオブジェクト条件付きトランスフォーマーセグメンテーション
- VespaSeg: リソースを考慮したグラウンディング→セグメンテーションのパイプラインによる参照表現セグメンテーションセグメンテーション
- アフォーダンスセグメンテーションのための軽量ニューラルネットワーク:デコーダモジュールの改良セグメンテーション
- DA-Fusion: 変形可能アテンションに基づくRGB-D融合トランスフォーマーによる未知物体のインスタンスセグメンテーションセグメンテーション