日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
変形物体生成arXiv:2609.18620

DeformSmith: 物理ハーネス誘導によるロボットマニピュレーション用変形アセットの階層的生成

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

テキストや1枚の画像から、物理的に妥当な変形物体アセットを階層的エージェント構築と物理ベースの検証ループで自動生成し、ロボット操作のシミュレーションやデータ合成を可能にするフレームワーク。

詳しい要約

1. どんなもの?

- テキストまたは単一画像から、ロボットマニピュレーション用の変形可能アセットを自動生成するフレームワーク DeformSmith を提案。 - 生成対象は geometry, appearance, physical properties を統合した interactive かつ physically credible な deformable assets。 - 階層的 agentic construction と shared physics-grounded harness により、構築・テスト・改良を反復する。 - 生成ループを robot interaction で閉じ、manipulation feedback と replayable interaction data を活用する。

2. 先行研究と比べてどこがすごい?

- 従来の text/image ベース生成では、変形や接触応答の物理的証拠が乏しく、interaction 適性を担保しにくい。 - DeformSmith は物理要件と相互作用証拠を生成プロセスに統合する点が異なる。 - 結果として、visual quality と physical plausibility の両面で PhysGen3D, PhysGM, PhysX-Omni などの state-of-the-art baselines を上回る。 - 変形物体のロボットマニピュレーション用データ合成も可能にする。

3. 技術・手法の肝は?

- 階層的 agentic construction により、geometry, physical models, material behavior, robot interaction を段階的に構築。 - shared physics-grounded harness が生成物をテストし、物理的妥当性を評価・改良する共通基盤として機能。 - robot interaction が generation loop を閉じ、manipulation feedback と replayable interaction data を生成・洗練に還元。 - 最終的に simulation と manipulation に供するアセットへ収束させる。

4. どうやって有効だと検証した?

- DeformSmith が生成したアセットの visual quality と physical plausibility を評価。 - 比較対象は PhysGen3D, PhysGM, PhysX-Omni などの state-of-the-art baselines。 - これらのベースラインを上回る結果を示したと報告。 - 変形物体のロボットマニピュレーション用データ合成のサポートも示す。 - 具体的な評価指標・データセット・被験者実験の詳細は要旨からは不明。

5. 議論はある?

- 変形物体では text や image が変形・接触応答の証拠を十分に与えないという課題を指摘。 - 自動生成には coupled physical requirements の解決と interaction evidence による構築・洗練の誘導が必要と論じる。 - robot interaction を生成ループに組み込む意義を強調。 - 限界や失敗事例、計算コスト、sim-to-real の議論は要旨からは不明。

6. 次に読むべき論文は?

- PhysGen3D - PhysGM - PhysX-Omni - 関連する deformable object manipulation や physics-grounded asset generation の研究(要旨では個別名の参照なし)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Can Li, Jie Gu, Zishun Deng, Jingmin Chen, Lei Sun

分類: cs.RO

原文アブストラクト

Creating deformable assets for robot manipulation requires jointly specifying their geometry, appearance, and physical properties. This is especially challenging for deformable objects, since text and images provide limited evidence about how they deform and respond to contact, yet these responses directly affect their suitability for interaction. Automated generation therefore needs to resolve coupled physical requirements and use interaction evidence to guide construction and refinement. We present DeformSmith, a framework that enables automated generation of interactive, physically credible deformable assets from text or a single image. Through hierarchical agentic construction and a shared physics-grounded harness, it progressively builds, tests, and refines geometry, physical models, material behavior, and robot interaction until the resulting asset is ready for simulation and manipulation. Robot interaction closes the generation loop through manipulation feedback and replayable interaction data. Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects. Project page: https://can-lee.github.io/deformsmith-web/

PR本紙発行元 EmplifAI