巧みな触覚ワールドモデル
Dexterous Tactile World Model
手袋型触覚センサと映像を組み合わせ、一人称視点の操作を予測する映像拡散トランスフォーマーベースのワールドモデルを提案。触覚情報により手の動きの過小評価を23%から9%に低減し、手領域の知覚誤差を7.4%改善した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan, Zhengxiang Yu, Fengyu Yang, Tianyu Liu, Zhiwen Fan, Daniel Rakita
分類: cs.CV, cs.LG, cs.RO
原文アブストラクト
World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.