日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.22757v1

Object-Uni: オブジェクト中心の空間理解と制御可能な生成のための統一モデル

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

シェア:XThreadsFacebookLINEはてブBluesky

オブジェクトの姿勢を明示的な幾何変数として扱い、姿勢認識・空間推論・姿勢条件付き生成・新規視点合成を統一したモデルを提案。視点ベースの姿勢抽象化とオブジェクトトークンに基づく姿勢アンカーにより、連続的な姿勢の理解と生成を実現する。

著者: Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen, DanDan Zheng, Libin Wang, Weiming Dong

分類: cs.CV, cs.AI

原文アブストラクト

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.

関連論文