S4VY: フィードフォワード4D視覚幾何におけるセグメントエニシング
S4VY: Segment Anything in Feed-Forward 4D Visual Geometry
RGB観測からフィードフォワードな4D視覚幾何特徴を抽出し、時空間クエリデコーダで永続的な4Dインスタンスマスクを生成するセグメンテーションモデルを提案。言語指示に基づく4Dグラウンディングも可能。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Jingdong Zhang, Xin Li, Jan Kautz, Wenping Wang, Chris Choy
分類: cs.CV
原文アブストラクト
Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.
関連論文
- シーン手がかりからの内在的4Dガウスセグメンテーション4Dセグメンテーション