日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
3Dシーン理解arXiv:2606.06485v1

PAR3D: 部品認識表現を用いた統合3Dマルチモーダル大規模言語モデルによるシーン理解

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

シェア:XThreadsFacebookLINEはてブBluesky

3Dシーン理解のための部品レベルの認識を可能にする統合3D-MLLMフレームワークPAR3Dを提案し、部品注釈付き合成データセットScenePartと階層的セグメンテーションクエリ生成により、部品レベルの質問応答と参照セグメンテーションを改善した。

著者: Shaohui Dai, Yansong Qu, You Shen, Shengchuan Zhang, Liujuan Cao

分類: cs.CV

原文アブストラクト

Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual question answering, captioning, and referring segmentation. However, existing 3D-MLLMs remain largely object-centric, limiting their ability to model fine-grained part structures that are essential for embodied interaction with 3D environments. In this work, we present PAR3D, a unified part-aware 3D-MLLM framework that enables models to understand, reason about, and ground both objects and their parts in 3D scenes. To enable training and evaluation of part-aware 3D scene understanding, we introduce ScenePart, a synthetic 3D scene dataset with part-level annotations and language instructions. We further develop Part-Aware 3D Representation Learning to enrich 3D visual representations with fine-grained part-level semantics, and propose Hierarchical Segmentation Query Generation to ground part targets via hierarchical object-part queries. Extensive experiments show that our method substantially improves part-level question answering and referring segmentation, while also achieving strong performance across object-level vision-language tasks.

関連論文