日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.22757

Object-Uni: オブジェクト中心の空間理解と制御可能な生成のための統一モデル

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

シェア:XThreadsFacebookLINEはてブBluesky

オブジェクトの姿勢を明示的な幾何変数として扱い、姿勢認識・空間推論・姿勢条件付き生成・新規視点合成を統一したモデルを提案。視点ベースの姿勢抽象化とオブジェクトトークンに基づく姿勢アンカーにより、連続的な姿勢の理解と生成を実現する。

詳しい要約

1. どんなもの?

Object-Uniは、オブジェクト中心の空間理解と制御可能な生成を統合したモデルです。オブジェクトのポーズを明示的な幾何学的変数として扱い、ポーズ知覚、空間推論、ポーズ条件付き生成、オブジェクト中心の新規ビュー合成を統一問題として扱います。視点ベースの方向抽象化により、ポーズをマルチモーダル大規模言語モデルで扱えるようにし、オブジェクトトークンに基づくポーズアンカーを用いて各インスタンスとポーズ状態を関連付けます。

2. 先行研究と比べてどこがすごい?

既存の統合モデルはオブジェクトを自然言語で記述できますが、連続的なオブジェクトポーズを正確に表現したり、目標視点で幾何学的に一貫した画像を生成することが困難でした。Object-Uniは、ポーズを単なる予測ラベルや制御信号ではなく、理解と生成で共有される明示的な幾何学的変数として扱う点で優れています。これにより、ポーズ理解とポーズ制御生成の両方を改善し、統合モデルをオブジェクトの記述から空間状態の操作へと進化させます。

3. 技術・手法の肝は?

手法の核は、ポーズを視点ベースの方向抽象化によって構造化された視点記述にマッピングしつつ、連続的な幾何学的監督を保持することです。また、オブジェクトトークンに基づくポーズアンカーを導入し、各インスタンスをそのポーズ状態に関連付けます。さらに、オブジェクト中心の空間ベンチマークUniSpatial-80Kを構築し、統合モデルを訓練します。

4. どうやって有効だと検証した?

実験により、オブジェクトレベルのポーズ理解とポーズ制御生成の改善が示されました。具体的な評価指標や比較対象は要旨からは不明ですが、提案モデルが空間状態の操作において有効であることを確認しています。

5. 議論はある?

要旨からは、ポーズ表現の抽象化と連続的監督のバランス、UniSpatial-80Kの構築方法、モデルの汎用性や限界についての詳細な議論は不明です。また、他のタスクへの拡張や実世界応用に関する考察も要旨には含まれていません。

6. 次に読むべき論文は?

要旨で参照されている関連研究は明示されていませんが、同分野の定番として、視覚理解と生成の統合モデル(例:Unified Model for Visual Understanding and Generation)、オブジェクト中心の表現学習(例:Object-Centric Learning)、およびポーズ推定と新規ビュー合成(例:Novel View Synthesis)に関する論文が挙げられます。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen, DanDan Zheng, Libin Wang, Weiming Dong

分類: cs.CV, cs.AI

原文アブストラクト

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.

関連論文