日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
カメラ軌道生成arXiv:2610.09513

OmniCam: 幾何学に基づくポーズトークン学習による全方位カメラ軌道生成

OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning

シェア:XThreadsFacebookLINEはてブBluesky

単一のパノラマ画像と言語による軌道記述から、自己回帰モデルでカメラポーズ列を生成する手法を提案し、軌道誤差と衝突率を大幅に削減した。

詳しい要約

1. どんなもの?

- 単一のpanoramaとテキストによるtrajectory記述からcamera pose列を生成するautoregressiveモデルOmniCamを提案。 - 3つの要素からなるgeometry-grounded pose token learningを採用。 - panoramic point-cloud encoderによる全方向幾何文脈 - hybrid absolute-rotation and relative-translation tokenizationと時間的一貫性のあるquaternion符号 - 幾何・意味の分離conditioning streamと明示的3D target anchor - 4つのcamera behaviorを含む267,700 trajectoryのデータセットOmniCaTも構築。

2. 先行研究と比べてどこがすごい?

- OmniCaT評価でGenDoPをOmniCaTで再学習したものと比較し、trajectory誤差を28–47%、collision rateを65.8%削減。 - 各指標の最良baselineと比較してATEを43.0%、collisionを62.3%削減。 - 単一panoramaとテキストからの生成という設定で、幾何とtarget-aware framingを両立。

3. 技術・手法の肝は?

- autoregressiveモデルでcamera pose列を生成。 - geometry-grounded pose token learningの3要素が肝。 - panoramic point-cloud encoderで全方向幾何文脈を取得 - hybrid absolute-rotation and relative-translation tokenizationと時間的一貫性のあるquaternion符号 - 幾何・意味の分離conditioning streamと明示的3D target anchor - これらによりscene geometryとtarget-aware framingを同時に扱う。

4. どうやって有効だと検証した?

- OmniCaT評価でtrajectory誤差とcollision rateをGenDoP再学習版および最良baselineと比較。 - component ablationで幾何conditioningとtarget-aware conditioningの有効性を支持。 - downstream実験でcamera-controlled video generationとrobotic active perceptionを検討。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- GenDoP(要旨で比較されている再学習baseline) - camera-controlled video generationおよびrobotic active perceptionの関連研究(要旨で言及)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhenyang Liu, Chenjie Cao, Yisu Zhang, Xuhui Zuo, Xiangyang Xue, Yanwei Fu, Tengfei Wang, Chunchao Guo

分類: cs.CV, cs.AI

原文アブストラクト

Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.

PR本紙発行元 EmplifAI