適応すべき時: オープンボキャブラリセグメンテーションにおける効率的な訓練不要適応のためのマルチシグナルドメインシフト検出
When to Adapt: Multi-Signal Domain Shift Detection for Efficient Training-Free Adaptation in Open-Vocabulary Segmentation
連続フレームの時間的一貫性を利用し、視覚変化・アダプタ不一致・意味ドリフトを組み合わせてドメインシフトを検出し、必要な時だけ訓練不要の適応をトリガーすることで、オープンボキャブラリセグメンテーションの精度を維持しつつ適応回数を大幅に削減する手法を提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Michele Antonazzi, Alejandra C. Hernandez, José Araujo, Olov Andersson, Patric Jensfelt
分類: cs.RO, cs.CV
原文アブストラクト
Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapters. Recent approaches apply this on a per-frame basis, which is impractical for deployments on resource-constrained robotic hardware. To tackle this, we propose a multi-signal domain shift detection method for training-free continual test-time adaptation (TF-CTTA) in open-vocabulary segmentation. Our method leverages temporal coherence across consecutive frames by monitoring and combining complementary aspects of domain shift (visual change, adapter mismatch, and semantic drift) to trigger adaptation only when needed. We validate our approach on a benchmark including indoor and outdoor environments and using real robotic data. We demonstrate that our approach maintains segmentation accuracy while substantially reducing adaptations, making training-free adaptation practical and feasible for long-term, real-world robotic deployments.
関連論文
- SAM-V: マルチビューインスタンスセグメンテーションのための幾何認識型Segment Anythingセグメンテーション
- DropClick: 農業ロボットデータのための半自動ワンクリックセグメンテーションセグメンテーション
- エンコーダは実際に何を決めているのか?樹木のジョイントセグメンテーションとステレオ深度における視覚バックボーンの制御比較セグメンテーション
- UAV画像の雑然シーンにおける通信鉄塔部品のゼロショットセグメンテーションのための顕著性-深度条件付けセグメンテーション
- SOS!:モデルフリーセグメンテーションのための合理化されたオブジェクト条件付きトランスフォーマーセグメンテーション
- VespaSeg: リソースを考慮したグラウンディング→セグメンテーションのパイプラインによる参照表現セグメンテーションセグメンテーション