日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
関節物体認識arXiv:2609.20673

FunArt: 生成3D潜在表現から機能構造と関節運動を解読

FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

シェア:XThreadsFacebookLINEはてブBluesky

単一の静的RGB-D観測から、生成3Dモデルの潜在表現を構造事前分布として活用し、物体の可動部・機能要素のセグメンテーションと関節運動(軸・原点・範囲)を推定するフレームワーク。

詳しい要約

1. どんなもの?

- 静的な単一配置のRGB-D観測から、関節を持つ物体の機能的な3Dシーングラフを構築するFunArtを提案。 - 物体インスタンスを再構成し、融合形状をTRELLIS.2のO-Voxel表現に変換。 - 凍結したsparse-compression VAEを構造事前として活用。 - 軽量なquery-based decoderで、可動部と機能的インタラクティブ要素を同時にセグメンテーションし、運動タイプ・軸・原点・範囲を推定。

2. 先行研究と比べてどこがすごい?

- 既存の関節シーン表現は観測されたインタラクションから運動学を復元するものが多い。 - 静的スキャンに基づく手法は、関節運動と機能的インタラクティブ要素を分離しがち。 - FunArtは単一の静的配置から、関節認識と機能要素を統合的に推定し、物理的インタラクション前にロボット知覚・計画を初期化可能。 - Articulate3Dで可動部セグメンテーション、関節推定、機能要素セグメンテーションのSOTAを達成。

3. 技術・手法の肝は?

- ポーズ付きRGB-D観測から物体インスタンスを再構成。 - 融合形状をTRELLIS.2のO-Voxel表現に変換。 - 凍結したsparse-compression VAEを構造事前として利用。 - 軽量query-based decoderが、コンパクトな物体レベル潜在と密な表面整合特徴を組み合わせ、可動部と機能要素を同時セグメンテーション。 - 運動タイプ、軸、原点、範囲を推定。

4. どうやって有効だと検証した?

- Articulate3Dデータセットで評価。 - 可動部セグメンテーション、関節推定、機能要素セグメンテーションにおいてSOTAを達成。 - 真の物体入力あり・なしの両設定で評価。 - エンドツーエンド設定で、最強ベースラインを可動部で1.5 AP_{50}、関節原点・軸制約下で2.8 AP_{50}、機能要素で6.7 AP_{50}上回る。

5. 議論はある?

- 生成3D潜在が、物理的インタラクション前にロボット知覚・計画を初期化できる実用的な構造手がかりを符号化することを示す。 - 単一静的配置からの関節・機能推定の有効性を実証。 - 限界や失敗事例、計算コスト、一般化性に関する議論は要旨からは不明。

6. 次に読むべき論文は?

- TRELLIS.2 - Articulate3D - O-Voxel - sparse-compression VAE - query-based decoder - 関節物体認識・運動学推定の関連手法(例:articulated object perception, kinematic model estimation)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dennis Rotondi, Abdelrhman Werby, Kai O. Arras

分類: cs.CV, cs.RO

原文アブストラクト

To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_{50} points for movable parts, 2.8 AP_{50} points under joint origin-and-axis constraints, and 6.7 AP_{50} points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.

関連論文

PR本紙発行元 EmplifAI