日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚/マニピュレーションarXiv:2609.21449

ME-Dex 1.0:異種触覚センシングを世界行動モデリングへ統合

ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling

シェア:XThreadsFacebookLINEはてブBluesky

視覚・触覚・行動を統合的に学習するWorld Action Tactile Modelを提案し、異なる触覚センサや実装を共通空間に写像する仕組みとデータ生成基盤を構築した。

詳しい要約

1. どんなもの?

- 視覚・触覚・行動を統合的に学習する World Action Tactile Model である ME-Dex-1.0 を提案。 - Mixture-of-Transformers 構成で Video Expert, Tactile Expert, Action Expert を持ち、flow matching で学習。 - 触覚を未来観測として扱い、映像と同様に世界状態の観測としてモデル化する点が特徴。 - 異種触覚入力を扱うため Canonical Hand Model と Unified Tactile Autoencoder を導入。 - データ不足に対応する Agentic Tactile Data Engine も開発。

2. 先行研究と比べてどこがすごい?

- 既存手法は触覚特徴を条件入力として使うが、未来の触覚状態・視覚観測・行動を同時に予測しない。 - 本研究は触覚を映像と同様に未来観測として扱い、視覚・触覚・行動を統合的に学習する点が新しい。 - 異種触覚入力を共有空間に写像する仕組みにより、異なる embodiment や sensing layout に対応。 - データ不足を補う agent-based データ生成プラットフォームを構築。

3. 技術・手法の肝は?

- Mixture-of-Transformers アーキテクチャで Video Expert, Tactile Expert, Action Expert を構成。 - 各 Expert は flow matching で訓練され、中間層で shared attention により接続。 - 共同 denoising 中に行動生成が視覚・触覚ダイナミクスの表現を利用可能。 - Canonical Hand Model と Unified Tactile Autoencoder で異種触覚入力を共有空間・潜在空間に写像。 - Agentic Tactile Data Engine がシミュレーション軌道再生中に力センサから触覚データを記録。

4. どうやって有効だと検証した?

- RoboTwin, DexJoCo, ManiFeel のシミュレーションプラットフォームで実験。 - 実ロボット評価も実施。 - 触覚センシングを備えた gripper と dexterous hand の両方で manipulation 性能の向上を確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として World Action Models, flow matching, Mixture-of-Transformers, RoboTwin, DexJoCo, ManiFeel が挙げられる。 - 同分野の定番として tactile sensing を用いた manipulation 研究や vision-language-action モデルが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xuancheng Zhang, Xuetao Liu, Qianying Tang, Jizhe Wang, Zhijing Cheng, Bochen Lin, Haoran Wen, Ming Li, Kun Zhan, Yu Liu

分類: cs.CV

原文アブストラクト

World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.

関連論文

PR本紙発行元 EmplifAI