日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.24976

DexTacWAM: 巧みな操作のための視触覚ワールドアクションモデル

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

各指先の触覚を個別に符号化し、指・姿勢を考慮した圧縮器で統合して動画拡散ワールドモデルに注入する視触覚ワールドアクションモデルを提案し、接触の多い両手巧み操作タスクで大幅な性能向上を達成した。

詳しい要約

1. どんなもの?

- 視覚と触覚を統合した World-Action Model (WAM) である DexTacWAM を提案。 - 各指先の触覚を独立にエンコードし、finger- and pose-aware tactile compressor で集約。 - 触覚潜在を video diffusion world model に注入し、視触覚の世界モデリングを共同で行う。 - 22-DoF bimanual platform 上の6つの contact-rich dexterous manipulation タスクを対象とする。

2. 先行研究と比べてどこがすごい?

- 従来の WAM は vision-centric で、接触ダイナミクスを直接モデル化できなかった。 - DexTacWAM は触覚を world state の一部として予測に組み込む点が新しい。 - 全6タスクで最高スコア、平均 70.6 対 最強 baseline の 38.0。 - 触覚 conditioning のみでは得られない利得を、contact evolution の world modeling で実現。

3. 技術・手法の肝は?

- 各 fingertip を独立にエンコードし、finger- and pose-aware tactile compressor で特徴集約。 - 触覚 latent を video diffusion world model に注入し、視触覚の joint world modeling を実施。 - frozen pretrained vision VAE を用い、tactile-encoder adaptation を4時間行う continual vision-to-touch learning。 - タスクあたり約100 demonstrations で触覚へ拡張、tactile midtraining は不要。

4. どうやって有効だと検証した?

- 22-DoF bimanual platform 上の6つの contact-rich タスクで評価。 - 全タスクで最高スコア、平均 70.6 対 baseline 38.0。 - Ablation: tactile world modeling を除くと4タスク平均が 74.7 から 26.6 に低下。 - 視覚予測品質は vision-only 比 0.5 dB 以内を維持。 - compressor は pre-fusion contact recall の 89.4% を保持し、訓練 2.26x、推論 1.29x 高速化。

5. 議論はある?

- 触覚 conditioning だけでなく、contact evolution を predicted world state としてモデル化することが性能向上に寄与。 - pretrained video priors を distributed multi-finger contact dynamics へ data- and compute-efficient に拡張可能。 - 限界や失敗事例、一般化性、実世界展開の議論は要旨からは不明。

6. 次に読むべき論文は?

- World-Action Models (WAMs) の vision-centric な先行研究。 - video diffusion world model を用いた行動生成手法。 - tactile sensing を活用した dexterous manipulation 研究。 - 要旨で参照/比較されている具体的な論文名は不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan

分類: cs.RO, cs.AI, cs.CV

原文アブストラクト

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.

関連論文

PR本紙発行元 EmplifAI