日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3DトラッキングarXiv:2609.30222

TrackEverything: 3Dシーン表現の重複排除による長距離高密度トラッキング

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

シェア:XThreadsFacebookLINEはてブBluesky

動画を3Dシーンの持続的なトラックとして表現し、重複排除により1000フレーム超の動画でも全可視点を追跡可能にした3Dポイントトラッカー。

詳しい要約

1. どんなもの?

- 3D point tracker: TrackEverything - 動画をworld coordinatesのpersistent 3D scene tracksとして表現 - 長horizonで全可視点を追跡可能 - 1000フレーム超の動画を40GB GPUメモリで追跡 - 従来のsparse long-termとdense short-termのトレードオフを解消

2. 先行研究と比べてどこがすごい?

- 従来: sparse pointsを長期間追跡 or 全点を短期間追跡のトレードオフ - TrackEverything: 全可視点を長期間追跡可能 - TAPVid-3Dでopen-source all-frame dense 3D trackersより20%以上APD向上 - 長シーケンスでstate-of-the-art sparse trackersと競合 - 追跡点数がはるかに多い

3. 技術・手法の肝は?

- 動画を3D worldの2D投影と捉え、model complexityをvideo durationから分離 - voxelization-based de-duplication: sliding-window境界でco-located tracksをマージ - trackingをendpoint refinerとtrajectory refinerに分解 - endpoint refiner: 各点のdestinationとstatic/dynamic分類を予測 - trajectory refiner: dynamic pointsのみdense trajectoryをデコード - 3D WAFT: 4D correlation volumesを効率的feature sampling in scene cloudに置換

4. どうやって有効だと検証した?

- TAPVid-3Dベンチマークで評価 - 短いクリップでopen-source all-frame dense 3D trackersより20%以上APD向上 - 長いシーケンスでstate-of-the-art sparse trackersと競合 - 1000フレーム超の動画を40GB GPUメモリで追跡可能なことを確認

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- TAPVid-3D - open-source all-frame dense 3D trackers - state-of-the-art sparse trackers

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt, Jakob Engel, Katerina Fragkiadaki, Adam W. Harley

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordinates. Grounded in the insight that videos are 2D projections of an underlying 3D world, TrackEverything decouples model complexity from video duration, allowing it to scale with unique physical scene geometry instead. Our approach introduces three key innovations. First, we employ a voxelization-based de-duplication mechanism at sliding-window boundaries to merge co-located tracks, preventing repeated observations of the same surface from redundantly accumulating. Second, we decompose tracking into an endpoint refiner that predicts each point's destination and static-versus-dynamic classification, followed by a lightweight trajectory refiner that decodes dense trajectories exclusively for dynamic points. Third, we propose 3D WAFT, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud. To the best of our knowledge, TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory. On TAPVid-3D, TrackEverything outperforms all open-source all-frame dense 3D trackers by more than 20% APD on short clips, while remaining competitive with state-of-the-art sparse trackers on long sequences, despite tracking far more points.

PR本紙発行元 EmplifAI