日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚/ワールドモデルarXiv:2609.20649

DexTouch-WM: 人間の触覚から器用なロボット操作のための行動条件付き触覚ワールドモデルを学習

DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

人間とロボットの手に同じ触覚センサ配置を装着し、人間の触覚データからロボットの視覚・触覚の未来を予測するワールドモデルを学習。人間のデータを増やすほどロボット領域の予測精度が向上し、政策評価や合成データ生成にも使えることを示した。

詳しい要約

1. どんなもの?

- 接触の多いdexterous manipulationの予測モデルを学習するための、action-conditioned world modelであるDexTouch-WMを提案する。 - 人間の触覚データから学習し、将来のRGB観測と両側のtactile dynamicsを同時に予測する。 - 人間とロボットの手に同じsensing layoutのflexible piezoresistive arraysを装着し、人間の動きをロボットのaction spaceにretargetする。 - これにより人間のinteractionが実ロボット予測と同じdynamics modelを監督できる。

2. 先行研究と比べてどこがすごい?

- 従来の接触リッチなdexterous manipulationの予測モデルは、実ロボットでのtactile interactionデータのスケールが高コストで、embodiment-specific sensorsに縛られていた。 - 本研究は、人間とロボットのmanipulationがtactile observationsとaction spacesを互換にすればtransferable contact dynamicsを共有するという洞察に基づく。 - 人間の触覚データをスケーラブルに活用し、実ロボット監督を固定したまま人間interactionを増やすことで、held-out robot-domainのvisual, geometric, contact predictionが大幅に改善することを示した。

3. 技術・手法の肝は?

- 人間とdexterous robot handsの両方に、shared sensing layoutを持つflexible piezoresistive arraysを配置する。 - 人間のmotionをrobot action spaceにretargetし、人間interactionが実ロボット予測と同じdynamics modelを監督できるようにする。 - DexTouch-WMは、pretrained video expertと軽量なtactile expertを、anatomy-aware tactile tokensとaligned action conditioningで結合する。 - 将来のRGB観測とbilateral tactile dynamicsをjointlyに予測するaction-conditioned world modelである。

4. どうやって有効だと検証した?

- human-to-robot scaling experimentsを実施し、実ロボット監督を5時間に固定したまま、人間interactionを0から100時間に増やした。 - 人間とロボットのタスクセットがdisjointであるにもかかわらず、held-out robot-domainのvisual, geometric, contact predictionが大幅に改善した。 - 予測以外に、world modelsをpolicy evaluationのsurrogate environmentsとして、また実ロボットpolicy learningのためのsynthetic trajectoriesのgeneratorとして評価した。

5. 議論はある?

- スケーラブルな人間interactionが、dexterous robot world modelsを学習するための補完的なdata axisを提供することを示した。 - 人間とロボットのタスクセットがdisjointでも予測改善が見られた点が議論の対象となる。 - 限界や今後の課題についての具体的な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、action-conditioned world models、video expert、tactile expert、human-to-robot transfer、dexterous manipulation、tactile sensingが挙げられる。 - 同分野の定番として、world models、model-based reinforcement learning、tactile representation learning、human-to-robot imitation learningに関する論文を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yan Qin, Yue Chen, Wenwei Lin, Shujia Liu, Chuqiao Lyu, Kailun Su, Chenze Yu, Ping Luo, Wenbo Ding, Tianxing Chen, Renjing Xu

分類: cs.RO, cs.CV

原文アブストラクト

Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.

PR本紙発行元 EmplifAI