I2CD: 単一画像から凸分解された衝突形状を直接生成
I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry
単一のRGB画像から物理シミュレーションでそのまま使える凸形状の集合を直接予測する手法を提案し、従来の再構成→分解パイプラインより高精度かつ6〜37倍高速に生成できることを示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Qian Wang, Liam Merz Hoffmeister, Brian Scassellati, Daniel Rakita
分類: cs.RO, cs.CV
原文アブストラクト
Physics simulators and motion planners require convex collision geometry, yet image-to-3D generative models output dense, frequently non-manifold visual meshes. Bridging the two today takes a slow, brittle reconstruct-then-decompose pipeline of repair, decimation, and approximate convex decomposition. We present I2CD, which predicts a convex decomposition directly from a single RGB image. Rather than train a new image-to-3D model, I2CD freezes the pretrained Hunyuan3D-2 image-conditioned diffusion transformer and shape decoder and trains only a lightweight cross-attention head (38M parameters, under ten GPU-hours) whose learned "convex-slot" tokens emit the halfplane parameters of $K$ convex polytopes. The output is compact, convex by construction, and loads into physics engines without any post-processing, in ${\sim}0.5$s per image. On $227$ held-out OmniObject3D and Google Scanned Objects instances, I2CD attains the highest volumetric IoU among eight reconstruct-then-decompose pipelines while running $6$-$37\times$ faster end-to-end. In a cross-simulator study in MuJoCo, PyBullet, Genesis, and Isaac Sim, every engine uses I2CD geometry as delivered, whereas raw generated meshes "load" everywhere but are silently replaced by a different collision shape in most cases or need seconds to minutes of per-object preprocessing. On a physical xArm7, I2CD produces planner-ready geometry for a $20$-object cluttered scene in $11$s versus $328$s for the strongest baseline, at comparable pick-and-place execution success ($85$ vs. $90$ of $100$ trials).
関連論文
- SceneFactory-3D:2D交通シーンを3D物理的反実世界へ持ち上げ、スケーラブルな物理基盤の安全評価を実現sim2real
- Skill2Real:ゼロショットSim-to-Realロボットマニピュレーションのためのエージェント型スキル学習sim2real
- RoboBridge:シミュレーションから実世界への転移のための自己進化型具現化エージェントフレームワークsim2real
- 報酬ハッキングを超えて:段階的人型学習パイプラインの4層におけるプロキシ乖離sim2real
- 大規模ロボット学習のためのGPUバッチ5Gシミュレーションによるネットワーク・イン・ザ・ループsim2real
- Awomo-SimDataEngine: エージェント型シミュレーション対応世界生成sim2real