日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
オドメトリ/多カメラarXiv:2609.40244

StreamRig: リグ内幾何を活用したストリーミング多カメラオドメトリ

StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry

シェア:XThreadsFacebookLINEはてブBluesky

凍結した多視点3D基盤モデル上に、キャリブレーション済みカメラリグの幾何を活かした因果的ストリーミングオドメトリを構築するフレームワークを提案。

詳しい要約

1. どんなもの?

- モバイルロボットや車両に搭載された同期 multi-camera rig 向けの streaming odometry フレームワーク。 - 凍結した multi-view 3D foundation model を front-end として利用し、rig calibration を活用して同期視点を jointly に知覚。 - Rig-Resampler で特徴を圧縮、CausalBridge で causal attention と key-value cache を適用、軽量 head で rig pose を回帰。 - 定期的な re-anchoring protocol により長いシーケンスでも安定した pose 推定を実現。 - 学習対象はこれらのモジュールのみで合計 74.6M parameters、相対 pose のみを supervision とする。

2. 先行研究と比べてどこがすごい?

- 多くの streaming 3D foundation model は monocular 入力向けに設計されており、rig geometry の効率的利用が課題だった。 - StreamRig は frozen multi-view 3D foundation model 上で calibrated rig の causal streaming odometry を構築。 - 評価した non-oracle monocular streaming モデルや rig-aware offline モデルよりも、4 データセットすべてで translation/rotation drift が低い。 - 推論コストも低く抑えている点が先行研究と比べて優位。

3. 技術・手法の肝は?

- freeze-and-stream フレームワーク:frozen front-end が rig calibration を用いて同期視点を jointly に知覚。 - Rig-Resampler が特徴を圧縮し、CausalBridge が causal attention と key-value cache を適用。 - 軽量 head が rig pose を回帰し、periodic re-anchoring protocol が長期安定性を支える。 - 学習はこれらのモジュールのみ(74.6M parameters)で、相対 pose のみを supervision に使用。 - 二段階学習戦略:group relocalization pretraining と causal rig training を組み合わせ、frozen front-end の幾何 prior と pretrained modules の alignment 能力を streaming odometry に転移。

4. どうやって有効だと検証した?

- NCLT、TartanGround、KITTI-360、および自収集 humanoid-robot dataset ZJH で評価。 - 学習は simulation のみで行い、実世界評価は zero-shot。 - 4 データセットすべてで、評価した non-oracle monocular streaming モデルおよび rig-aware offline モデルより translation/rotation drift が低い。 - 推論コストが低いことも確認。 - Ablations と controlled camera-count 実験で性能向上の要因を特定し、長い training window が長期間推論に与える影響も検討。

5. 議論はある?

- Ablations と controlled camera-count 実験により、性能向上の要因を分析。 - 長い training window が長期間推論に与える影響を検討。 - 具体的な議論の内容や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:non-oracle monocular streaming モデル、rig-aware offline モデル。 - 関連手法:multi-view 3D foundation model、causal attention with key-value cache、group relocalization pretraining。 - 同分野の定番:visual odometry、multi-camera SLAM、streaming 3D foundation model。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang

分類: cs.CV, cs.RO

原文アブストラクト

Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model. The frozen front-end jointly perceives the synchronized views using rig calibration. A Rig-Resampler compresses their features, a CausalBridge applies causal attention with a key-value cache, and a lightweight head regresses rig poses. A periodic re-anchoring protocol supports stable pose estimation over long sequences. Only these modules are trained, 74.6M parameters in total, with relative poses as the sole supervision. Our two-stage training strategy combines group relocalization pretraining with causal rig training to transfer the geometric priors of the frozen front-end and the alignment ability of the pretrained modules to streaming odometry. We evaluate on NCLT, TartanGround, KITTI-360, and our self-collected humanoid-robot dataset ZJH, where training uses only simulation and real-world evaluation is zero-shot. Across all four datasets, StreamRig achieves lower translation and rotation drift than the evaluated non-oracle monocular streaming and rig-aware offline models, while maintaining low inference cost. Ablations and controlled camera-count experiments identify the sources of these gains. We further examine how longer training windows affect inference over longer horizons. Code has been released at https://github.com/WeiYuFei0217/StreamRig.

PR本紙発行元 EmplifAI