日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.03453

I2CD: 単一画像から凸分解された衝突形状を直接生成

I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry

シェア:XThreadsFacebookLINEはてブBluesky

単一のRGB画像から物理シミュレーションでそのまま使える凸形状の集合を直接予測する手法を提案し、従来の再構成→分解パイプラインより高精度かつ6〜37倍高速に生成できることを示した。

詳しい要約

1. どんなもの?

- 単一の RGB 画像から直接 convex decomposition を予測する手法 I2CD を提案。 - 出力は K 個の convex polytope の halfplane parameters で、compact かつ構成上 convex。 - physics engine に post-processing なしでそのままロード可能、1 画像あたり約 0.5 秒。 - 対象は physics simulator や motion planner が必要とする collision geometry。

2. 先行研究と比べてどこがすごい?

- 従来は reconstruct-then-decompose pipeline(repair, decimation, approximate convex decomposition)で遅く脆い。 - I2CD は image-to-3D モデルを新規学習せず、pretrained Hunyuan3D-2 を凍結し軽量 head のみ学習。 - 8 つの reconstruct-then-decompose pipeline 中で最高の volumetric IoU を達成しつつ、end-to-end で 6–37 倍高速。 - 生成 mesh は多くの engine で黙って別 collision shape に置換されるが、I2CD はそのまま使われる。

3. 技術・手法の肝は?

- pretrained Hunyuan3D-2 の image-conditioned diffusion transformer と shape decoder を凍結。 - 軽量 cross-attention head(38M parameters、10 GPU-hours 未満)のみを学習。 - learned "convex-slot" tokens が K 個の convex polytope の halfplane parameters を出力。 - 出力は構成上 convex で compact、post-processing 不要。

4. どうやって有効だと検証した?

- 227 の held-out OmniObject3D と Google Scanned Objects インスタンスで評価。 - 8 つの reconstruct-then-decompose pipeline と比較し、volumetric IoU 最高、6–37 倍高速。 - MuJoCo, PyBullet, Genesis, Isaac Sim での cross-simulator study を実施。 - 物理 xArm7 で 20-object cluttered scene の planner-ready geometry を 11 秒で生成(最強 baseline は 328 秒)、pick-and-place 成功率は 85/100 対 90/100。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Hunyuan3D-2(pretrained image-conditioned diffusion transformer) - approximate convex decomposition 関連手法 - OmniObject3D, Google Scanned Objects データセット - MuJoCo, PyBullet, Genesis, Isaac Sim の physics simulator

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Qian Wang, Liam Merz Hoffmeister, Brian Scassellati, Daniel Rakita

分類: cs.RO, cs.CV

原文アブストラクト

Physics simulators and motion planners require convex collision geometry, yet image-to-3D generative models output dense, frequently non-manifold visual meshes. Bridging the two today takes a slow, brittle reconstruct-then-decompose pipeline of repair, decimation, and approximate convex decomposition. We present I2CD, which predicts a convex decomposition directly from a single RGB image. Rather than train a new image-to-3D model, I2CD freezes the pretrained Hunyuan3D-2 image-conditioned diffusion transformer and shape decoder and trains only a lightweight cross-attention head (38M parameters, under ten GPU-hours) whose learned "convex-slot" tokens emit the halfplane parameters of $K$ convex polytopes. The output is compact, convex by construction, and loads into physics engines without any post-processing, in ${\sim}0.5$s per image. On $227$ held-out OmniObject3D and Google Scanned Objects instances, I2CD attains the highest volumetric IoU among eight reconstruct-then-decompose pipelines while running $6$-$37\times$ faster end-to-end. In a cross-simulator study in MuJoCo, PyBullet, Genesis, and Isaac Sim, every engine uses I2CD geometry as delivered, whereas raw generated meshes "load" everywhere but are silently replaced by a different collision shape in most cases or need seconds to minutes of per-object preprocessing. On a physical xArm7, I2CD produces planner-ready geometry for a $20$-object cluttered scene in $11$s versus $328$s for the strongest baseline, at comparable pick-and-place execution success ($85$ vs. $90$ of $100$ trials).

関連論文

PR本紙発行元 EmplifAI