日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.18504

InterMASH: 把持合成のための統一幾何表現

InterMASH: A Unified Geometric Representation for Grasp Synthesis

シェア:XThreadsFacebookLINEはてブBluesky

球面固定アンカーと球面調和関数で手と物体の局所形状・接触を統一的に符号化し、拡散Transformerで手形状と接触を同時生成する把持合成手法を提案。

詳しい要約

1. どんなもの?

- 手と物体の安定かつ物理的に妥当な把持を生成するgrasp synthesisのための統一表現InterMASHを提案。 - 人間とロボットの手の形態・表面モデリングの違いを超え、sphere-fixed anchorsでcross-embodiment対応を確立。 - 各anchorで低次spherical harmonicsにより局所的な手形状・物体形状・接触をコンパクトに符号化し、明示的で解釈可能なtoken sequenceを形成。 - このtoken化構造上でconditional Diffusion Transformerを動作させ、手形状と接触を同時生成。

2. 先行研究と比べてどこがすごい?

- 従来はcontact mapsやdense implicit descriptorsで相互作用を表現していたが、不完全または計算コストが高く冗長という問題があった。 - InterMASHは統一幾何表現により、これらよりコンパクトで明示的・解釈可能な表現を実現。 - 大規模ShadowHandベンチマークで主要な物理的実現可能性指標においてstate-of-the-artと競合する性能を達成。 - 複数の手でのjoint trainingをサポートし、人間の把持データによるcross-embodiment fine-tuningがロボット把持の成功率と多様性を改善することを示した。

3. 技術・手法の肝は?

- sphere-fixed anchorsを介してcross-embodiment correspondenceを確立。 - 各anchorで低次spherical harmonicsを用いて局所的な手形状・物体形状・接触をコンパクトに符号化し、明示的で解釈可能なtoken sequenceを構成。 - このnatively tokenized構造に基づき、conditional Diffusion TransformerをInterMASH表現空間で直接動作させ、手形状と接触をjointlyに生成。 - これにより一貫性と物理的妥当性を向上。

4. どうやって有効だと検証した?

- 大規模ShadowHandベンチマークにおいて、主要な物理的実現可能性指標でstate-of-the-art手法と競合する性能を達成。 - 複数の手でのjoint trainingが可能であることを示した。 - 人間の把持データによるcross-embodiment fine-tuningがロボット把持の成功率と多様性を改善することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な先行研究は明示されていない。関連手法としてcontact maps、dense implicit descriptors、conditional Diffusion Transformer、spherical harmonics、ShadowHandベンチマークなどが挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xuanze Yang, Yumeng Liu, Haiyang Xin, Changhao Li, Haowei Shen, Kai Xu, Ligang Liu, Ruizhen Hu

分類: cs.RO, cs.GR

原文アブストラクト

Grasp synthesis aims to generate stable and physically plausible hand--object interactions, and has become a fundamental problem in both human hand modeling and robotic manipulation. However, a unified representation across human and robotic hands is still lacking, mainly due to differences in hand morphology and surface modeling. Prior methods typically rely on either contact maps or dense implicit descriptors to represent interaction, but these representations are often incomplete or computationally expensive and redundant. We propose InterMASH, a unified geometric representation that establishes cross-embodiment correspondence using sphere-fixed anchors. At each anchor, low-degree spherical harmonics compactly encode local hand geometry, object geometry, and contact, forming an explicit and interpretable token sequence. Building on this natively tokenized structure, we introduce a conditional Diffusion Transformer that operates directly in the proposed InterMASH representation space and jointly generates hand geometry and contact, improving consistency and physical plausibility. Our method achieves competitive performance with state-of-the-art methods on key physical feasibility metrics in a large-scale ShadowHand benchmark, supports joint training across multiple hands, and shows that cross-embodiment fine-tuning with human grasp data can improve robotic grasp success and diversity. Project page is available at https://inter-mash.github.io/.

関連論文

PR本紙発行元 EmplifAI