日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.03333

接触の多い操作のための同変視覚触覚拡散ポリシー

Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

視覚と触覚の観測を球面トークンに射影し、同変拡散ポリシーで空間的に一貫した行動を予測することで、接触の多い模倣学習のデータ効率を大幅に向上させるVISTAを提案。

詳しい要約

1. どんなもの?

接触の多い操作のための模倣学習において、高品質な専門家データの取得コストが高い問題に対処するため、サンプル効率の良いポリシー学習を目指す研究。提案手法 VISTA は、workspace-level の equivariant な visuotactile diffusion policy であり、視覚と触覚の観測を球面トークンに投影し、触覚の接触キューを視覚球面方向に注入し、end-effector の向きで融合表現を回転させる。これにより空間的に一貫した行動を予測する。

2. 先行研究と比べてどこがすごい?

要旨では、強力な visuotactile imitation learning ベースラインと比較して、シミュレーションと実世界の両方でデータ効率を大幅に改善したと述べられている。具体的なベースライン名や先行研究の詳細は要旨からは不明。

3. 技術・手法の肝は?

- VISTA は workspace-level の equivariant visuotactile diffusion policy - 視覚と触覚の観測を spherical tokens に投影 - permutation-equivariant spherical fusion により触覚の接触キューを視覚球面方向に注入 - 融合された harmonic representation を end-effector の向きで回転 - 得られた表現が equivariant diffusion policy の条件付けとなり、空間的に一貫した行動を予測

4. どうやって有効だと検証した?

要旨によれば、シミュレーションと実世界のロボティクス設定の両方で広範な実験を行い、強力な visuotactile imitation learning ベースラインと比較してデータ効率を大幅に改善した。具体的な評価指標やタスクの詳細は要旨からは不明。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究や関連手法の具体名は記載されていない。同分野の定番として、visuotactile imitation learning、equivariant policy、diffusion policy に関する論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lik Hang Kenny Wong, Yiyao Ma, Xiu-Shen Wei, Zelong Tan, Zhuheng Song, Dongsheng Xie, Kai Chen, Qi Dou

分類: cs.RO, cs.AI

原文アブストラクト

Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: https://vista-paper.github.io/

関連論文

PR本紙発行元 EmplifAI