日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マルチモーダル模倣学習arXiv:2609.29760

PolyUMI: 視覚・触覚・音声を統合した物体推論と操作のためのアクセシブルなデータ収集

PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

無線ハンドヘルドグリッパで視覚・触覚・音声・固有感覚を同期収集し、ロボットに展開できるオープンソース基盤と、異種センサ情報を統合するマルチモーダル方策VisTAを提案。物体推論や滑り制御、接触の多い操作で有効性を示した。

詳しい要約

1. どんなもの?

- PolyUMIは、visual・tactile・audioを統合したdemonstration収集とロボット展開のためのオープンソースplatform。 - 軽量・無線のhandheld gripperで、wrist-camera、optical tactile、contact-audio、proprioceptiveの観測を同期記録する。 - 有線workstationを必要とせず、同じsensing fingerをrobot end effectorに移せるため、demonstration収集とpolicy実行でsensing geometryを保てる。 - 異種観測を扱うtoken-level multimodal policy「VisTA」も提案し、contact-awareなrobot actionを予測する。

2. 先行研究と比べてどこがすごい?

- 多くのimitation-learning systemはvisionとproprioception中心で、視覚的に推論しにくいcontact情報へのアクセスが限られていた。 - PolyUMIはvisual・tactile・audio・proprioceptionを同期収集でき、tether不要でscalableな点が異なる。 - demonstration収集時とpolicy実行時でsensing geometryを保つ設計は、既存手法に対する利点として示されている。 - VisTAは既存のmultimodal policyと比較してcompetitiveまたは上回ると報告されている。

3. 技術・手法の肝は?

- 軽量・無線のhandheld gripperで、wrist-camera、optical tactile、contact-audio、proprioceptive observationを同期記録する。 - 同じsensing fingerをrobot end effectorへ移すことで、demonstrationと実行間のsensing geometryを維持する。 - VisTAはtoken-level multimodal policyであり、sensor間および時間方向の情報を統合してcontact-aware actionを予測する。 - 詳細なnetwork構造や学習手順は要旨からは不明。

4. どうやって有効だと検証した?

- object inference、slip control、contact-rich manipulationを含む実験で検証している。 - touchとaudioがvisionを超えてtask-relevantな情報を明らかにすることを示した。 - VisTAが既存のmultimodal policyとcompetitiveまたはそれを上回ることを示した。 - 具体的なdataset規模、baseline名、評価指標は要旨からは不明。

5. 議論はある?

- touchとaudioが視覚では得にくいcontact情報を補完し、contact-rich manipulationに有効である点が議論の中心。 - PolyUMIとVisTAを組み合わせ、視覚を超えて物理的interactionを知覚するpolicy学習のaccessibleなpipelineを提供する。 - 限界、失敗事例、汎化性、コストなどの議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている既存のmultimodal policy(具体的名称は要旨からは不明)。 - imitation learning、visual-tactile learning、optical tactile sensing、contact-audio sensingに関する同分野の定番研究。 - Project Page: https://polyumi-vista.github.io で関連研究が参照されている可能性がある。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Conor W. Hayes, Rickmer Krohn, Aravind Ramaswami, Anunth Ramaswami, Nils Dengler, Kevin M. Lynch, J. Edward Colgate, Georgia Chalvatzaki, Matthew L. Elwin

分類: cs.RO

原文アブストラクト

Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io

PR本紙発行元 EmplifAI