日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2609.34286

巧みな触覚ワールドモデル

Dexterous Tactile World Model

シェア:XThreadsFacebookLINEはてブBluesky

手袋型触覚センサと映像を組み合わせ、一人称視点の操作を予測する映像拡散トランスフォーマーベースのワールドモデルを提案。触覚情報により手の動きの過小評価を23%から9%に低減し、手領域の知覚誤差を7.4%改善した。

詳しい要約

1. どんなもの?

- ビデオと触覚信号から未来のフレームを予測する世界モデル。 - 各手に装着したグローブからの触覚信号を利用。 - 一人称視点の操作を対象。 - ビデオ拡散トランスフォーマーをベースに、触覚信号を条件付け。 - 接触の生成・解放など視覚的に観測しにくい事象を触覚で補完。

2. 先行研究と比べてどこがすごい?

- 視覚のみのモデルと比較して、手の動きの過小評価を23%から9%に低減。 - 手領域の知覚誤差を7.4%削減(3回の訓練実行平均)。 - 予測ホライズンが長くなるほど改善が増大し、後半のチャンクでは最初の約4.1倍の改善。 - 同じ設定下で他の視覚-触覚世界モデルよりも優れる。 - 推論時に触覚が利用できない場合でも、訓練に触覚を用いることで未来フレーム予測が改善。

3. 技術・手法の肝は?

- 事前訓練済みビデオ拡散トランスフォーマーをベースに使用。 - 各手の触覚信号を、ビデオトークン内の対応する手の位置にゼロ初期化残差で条件付け。 - 因果マスクにより、予測フレームが未来の情報にアクセスするのを防止。 - 触覚信号の大きさと空間位置の両方を活用。 - 二値接触状態(手ごとまたは位置ごと)に置き換えると予測誤差が増加。

4. どうやって有効だと検証した?

- 視覚のみのモデルとアーキテクチャ、パラメータ数、訓練を一致させて比較。 - 各モデル3回の訓練実行で評価。 - 手の動きの過小評価率と手領域の知覚誤差を指標に使用。 - 予測ホライズンごとの改善を分析。 - 他の視覚-触覚世界モデルとの比較、および触覚信号のアブレーションを実施。

5. 議論はある?

- 触覚信号の大きさと空間位置の両方が重要であることをアブレーションで示唆。 - 力の時間的変化が相互作用の持続か変化かを示す。 - 訓練時に触覚があれば、推論時に触覚がなくても予測が改善することを確認。 - 予測ホライズンが長いほど触覚の利点が増大。 - 具体的な限界や今後の課題については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:視覚のみの世界モデル、他の視覚-触覚世界モデル。 - 関連手法:ビデオ拡散トランスフォーマー、触覚グローブを用いた操作学習。 - 同分野の定番:視覚-触覚融合によるロボット操作、触覚に基づく世界モデル。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan, Zhengxiang Yu, Fengyu Yang, Tianyu Liu, Zhiwen Fan, Daniel Rakita

分類: cs.CV, cs.LG, cs.RO

原文アブストラクト

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.

関連論文

PR本紙発行元 EmplifAI