日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.01758

GenCOPE: ロボットピッキングのための合成から実への汎化カテゴリレベル物体姿勢推定

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

シェア:XThreadsFacebookLINEはてブBluesky

合成データのみで学習し実世界に直接汎化するカテゴリレベル物体姿勢推定手法を提案。2D/3D意味的一貫性制約と2D-3D相互融合によりドメインギャップを克服し、軽量なアーキテクチャでロボット操作に応用可能。

詳しい要約

1. どんなもの?

- カテゴリーレベルの物体姿勢推定(COPE)を、合成データのみで学習し実世界に直接汎化するSyn2Real設定で実現する手法。 - ロボットの3Dシーン理解やピッキングへの応用を目的とする。 - レンダリングされた合成データのみで訓練し、実世界展開に汎化する。 - ドメインギャップ、特にテクスチャ外観の差異が中心課題。 - 2D/3Dの意味的一貫性制約と2D-3Dクロス一貫性学習を導入。 - グローバル特徴のみで動作する軽量・効率的なアーキテクチャ。

2. 先行研究と比べてどこがすごい?

- 既存COPEは新カテゴリごとに実世界訓練データの再収集が必要でスケーラビリティに欠ける。 - 本研究は合成データのみで訓練し実世界に直接汎化するSyn2Real一般化COPEを目指す。 - ドメイン不変表現学習により、テクスチャ外観のドメインギャップに対処。 - 2D/3D意味的一貫性制約で特徴エンコーダのドメイン固有特徴への感度を低減。 - 2D-3Dクロス一貫性学習と密なクロスモダリティ融合で姿勢推定を精緻化。 - グローバル特徴のみで軽量・効率的なアーキテクチャを実現。

3. 技術・手法の肝は?

- ドメイン不変表現の学習により、同カテゴリ内の意味的共通性を捉える。 - 2Dおよび3Dの意味的一貫性制約を導入し、特徴エンコーダのドメイン固有特徴への感度を低減。 - 2D-3Dクロス一貫性学習を行うエンドツーエンド姿勢回帰フレームワークを提案。 - 密なクロスモダリティ融合を活用し姿勢推定をさらに精緻化。 - 実世界ロボット展開のため簡潔さと有効性を重視し、グローバル特徴のみで動作。 - 軽量かつ効率的なアーキテクチャを実現。

4. どうやって有効だと検証した?

- REAL275およびWild6Dベンチマークで広範な実験を実施。 - 実世界のロボットマニピュレーションシーンでも評価。 - これらの実験により、提案パラダイムの優れたSyn2Real汎化性能を示す。 - コードとデモは https://paperreview99.github.io/GenCOPE/ で公開。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- REAL275およびWild6Dベンチマークに関連するCOPE研究。 - カテゴリーレベル物体姿勢推定(COPE)の既存手法。 - Syn2Realやドメイン汎化に関する研究。 - 2D-3Dクロスモダリティ融合や意味的一貫性制約を用いた手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao

分類: cs.CV

原文アブストラクト

Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm. Code and demos are released at https://paperreview99.github.io/GenCOPE/.

関連論文

PR本紙発行元 EmplifAI