日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
4DセグメンテーションarXiv:2609.36875

S4VY: フィードフォワード4D視覚幾何におけるセグメントエニシング

S4VY: Segment Anything in Feed-Forward 4D Visual Geometry

シェア:XThreadsFacebookLINEはてブBluesky

RGB観測からフィードフォワードな4D視覚幾何特徴を抽出し、時空間クエリデコーダで永続的な4Dインスタンスマスクを生成するセグメンテーションモデルを提案。言語指示に基づく4Dグラウンディングも可能。

詳しい要約

1. どんなもの?

- 動的シーンにおける4DインスタンスセグメンテーションのためのSegment Anythingモデル「S4VY」を提案。 - feed-forwardな4D視覚幾何に基づき、RGB観測からclass-agnosticな4Dインスタンスマスクを生成。 - 各persistent object queryが全観測を通じて1つのentityを束縛。 - prompt-independentなセグメンテーションと、point/box条件付き選択をサポート。 - seed maskやtemporal orderingを必要としない。 - 自然言語グラウンディングのためのagentic harnessも開発。

2. 先行研究と比べてどこがすごい?

- 既存のSegment Anythingモデルは主に2D画像/ビデオマスクで、sequential memoryによりidentityを保持。 - feed-forward視覚幾何に基づくpromptable 4Dインスタンスセグメンテーションは未探索。 - S4VYは4D視覚幾何を基盤とし、space-time query decoderで4Dマスクを生成。 - 静的・動的シーンを統合評価でSOTAの4Dインスタンスセグメンテーションと言語誘導グラウンディング性能を実証。

3. 技術・手法の肝は?

- feed-forward 4D視覚幾何モデル上に構築。 - 共有視覚幾何特徴をspace-time query decoderでclass-agnosticな4Dインスタンスマスクに変換。 - 各persistent object queryが全観測を通じて1つのentityを束縛。 - agentic harness: active tree searchで関連フレームを同定(全固定ウィンドウ走査不要)。 - dual-stream grounder: 細粒度VLM視覚priorと幾何整合インスタンス特徴を組み合わせ、bounding-box予測とobject-queryマッチングを補完。 - independent criticが最終4Dインスタンスマスクを選択。

4. どうやって有効だと検証した?

- 静的・動的シーンにわたる統合評価を実施。 - 4Dインスタンスセグメンテーションと言語誘導グラウンディングの両方でSOTA性能を実証。 - 詳細な実験設定やデータセットは要旨からは不明。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: Segment Anything models(2D画像/ビデオマスク、sequential memory)。 - 関連手法: feed-forward visual geometry, promptable 4D instance segmentation, VLM visual priors, active tree search, dual-stream grounder, independent critic。 - 同分野の定番: 4D instance segmentation, language-guided grounding。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jingdong Zhang, Xin Li, Jan Kautz, Wenping Wang, Chris Choy

分類: cs.CV

原文アブストラクト

Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.

関連論文

PR本紙発行元 EmplifAI