日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2610.10288

TouchScale:視覚と触覚の大規模ヒューマンインタラクションデータセット

TouchScale: 500 Hours of Human Vision and Touch for Visual-Tactile Learning

シェア:XThreadsFacebookLINEはてブBluesky

単一のウェアラブル装置で収録した500時間の接触豊富な人間の行動データを構築し、視覚-触覚学習の事前学習に用いることで、未知の触覚センサへのゼロショット予測やロボットの接触を伴う操作タスクの成功率を大幅に改善した。

詳しい要約

1. どんなもの?

- 500時間の接触豊富な人間の視覚・触覚データセットTouchScaleを提案。 - 単一の統一ウェアラブル装置で収録。 - 約2Kの事前定義タスク記述が日常活動と構造化操作を網羅。 - 各記録はegocentric RGB-D video、wrist RGB video、両手の密な触覚計測を時間同期。 - 公開予定で、同期視覚触覚記録と再構成物体モデルを含む。

2. 先行研究と比べてどこがすごい?

- 従来の視覚触覚データセットは同期触覚データ量が人間動画よりはるかに小さい。 - 最大規模の資源は異なるセンサや注釈手順を統合し、データ規模の効果を分離しにくい。 - TouchScaleは単一の統一ウェアラブル設定で収録し、センサと収集プロトコルを固定。 - 全TouchScaleで訓練すると、未見触覚センサのzero-shot contact IoUが0.134から0.383に向上。 - 視覚エンコーダの事前学習で3ベンチマークの行動認識精度が比較視覚触覚データセット中最高。

3. 技術・手法の肝は?

- 単一の統一ウェアラブル収録装置で接触豊富な人間相互作用を記録。 - egocentric RGB-D video、wrist RGB video、両手密触覚計測を時間同期。 - 約2Kの事前定義タスク記述を付与。 - 視覚エンコーダの事前学習と、ロボットポリシーの視覚触覚mid-trainingに利用。 - データ規模増加に伴うzero-shot触覚予測とロボット成功率の傾向を検証。

4. どうやって有効だと検証した?

- 全TouchScaleで訓練し、未見触覚センサのzero-shot contact IoUを0.134から0.383へ改善。 - 視覚エンコーダ事前学習で3ベンチマークの行動認識精度が比較視覚触覚データセット中最高。 - ロボットポリシーの視覚触覚mid-trainingで、4つの接触豊富な実世界操作タスクの平均成功率が22.5%から57.5%に向上。 - センサと収集プロトコル固定で、データ増加に伴いzero-shot触覚予測とロボット成功率が全体的に上昇傾向。

5. 議論はある?

- 一貫したセンシングで大規模収集した人間視覚触覚データが知覚とロボット操作の両方に利益をもたらす可能性を示唆。 - データ規模の効果を分離できる設計が利点。 - 限界や失敗事例、計算コスト、一般化範囲などは要旨からは不明。 - 倫理・プライバシー面の議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている先行視覚触覚データセット(具体的名称は要旨からは不明)。 - 同分野の定番として、大規模egocentric human interaction videoデータセット、視覚触覚学習、robot manipulationのvisual-tactile policy学習に関する研究。 - 具体的な論文名は要旨に無いため、関連手法・データセット名を挙げることはできない。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Dayou Li, Hao Wang, Qianqian Yang, Zihao Zhu, Haoquan Fang, Ziyao Zeng, Yan Han, Zihan Wang, Yan Wang, Baoru Huang, Dilin Wang, Kenji Shimada, Yiyue Luo, Manling Li, Teresa Lv, Mustafa Mukadam, Rakesh Ranjan, Ruohan Zhang, Qi He, Changliu Liu, Xu Chen, Marco Pavone, Bangya Liu, Jiachen Li, Masayoshi Tomizuka, Zhiwen Fan

分類: cs.CV

原文アブストラクト

Large-scale egocentric human interaction data is becoming an important source of physical supervision for embodied learning, yet video alone leaves the contact and pressure that characterize physical interaction unrecorded. Recent visual-tactile datasets provide this missing supervision, but their synchronized tactile data remain far smaller in volume than human video. Moreover, the largest resources often merge recordings from different sensors or annotation procedures, which makes the effect of data scale difficult to isolate. We therefore introduce TouchScale, a 500-hour dataset of contact-rich human interaction recorded with a single unified wearable setup. Its approximately 2K predefined task descriptions span everyday activities and structured manipulation, and each recording temporally aligns egocentric RGB-D video with wrist RGB video and dense full-hand bimanual tactile measurements. Compared with prior tactile data, training on the full TouchScale raises zero-shot contact IoU on data from an unseen tactile sensor from 0.134 to 0.383. Pretraining a visual encoder on TouchScale also yields the highest action recognition accuracy on three benchmarks among the compared visual-tactile datasets. Used for visual-tactile mid-training of a robot policy, TouchScale improves the average real-world success rate across four contact-rich manipulation tasks from 22.5% to 57.5%. With the sensor and collection protocol held fixed, both zero-shot tactile prediction and robot success show an overall upward trend as more TouchScale data is used. These results suggest that human visual-tactile data collected at scale with consistent sensing benefits both perception and robot manipulation. We will publicly release TouchScale, including all synchronized visual-tactile recordings and reconstructed object models, to support future research on scalable visual-tactile learning.

関連論文

PR本紙発行元 EmplifAI