日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2608.15060v1

EgoTac: 自己中心視覚からの実環境触覚予測

EgoTac: In-the-wild Tactile Prediction from Egocentric Vision

シェア:XThreadsFacebookLINEはてブBluesky

自己中心視覚映像から触覚情報を予測するモデルEgoTacを提案し、570万以上の画像-触覚ペアで学習して高精度な触覚予測を実現した。

詳しい要約

1. どんなもの?

EgoTacは、人間のEgocentricビデオから触覚情報(連続的な力計測と二値の接触)を直接予測する、一般化可能なモデルである。5.7M以上の画像-触覚ペアからなる統一コーパスで訓練され、多様なインタラクションにおける微妙な触覚ダイナミクスを捉える。ロボット学習のための触覚プリヤを大規模に抽出するスケーラブルな経路を提供する。

2. 先行研究と比べてどこがすごい?

従来の触覚予測は、センサデータや限定された設定に依存することが多く、大規模な人間のビデオデータを活用していなかった。EgoTacは、豊富でスケーラブルな人間のEgocentricビデオから触覚を予測する点で新規性が高い。また、連続的な力と二値の接触の両方を扱う統一モデルであり、ドメイン外の接触予測ベンチマークで最先端の接触推定器を一貫して上回る。

3. 技術・手法の肝は?

手法の肝は、大規模な画像-触覚ペアのコーパスを構築し、Egocentricビデオから触覚を予測するモデルを訓練することにある。具体的なアーキテクチャや損失関数は要旨からは不明だが、データの多様性と量が性能向上に寄与することが示されている。

4. どうやって有効だと検証した?

有効性は、ドメイン内予測で平均力誤差0.06N未満を達成し、ドメイン外の接触予測ベンチマークで最先端の接触推定器を上回ることで検証された。また、実世界のビデオでのゼロショット予測が可能であり、スケーリング解析によりデータの多様性と量が性能を着実に向上させることが示された。

5. 議論はある?

要旨からは、モデルの限界や倫理的な考察、実世界での応用における課題などは不明である。また、触覚予測の精度が実際のロボット操作にどの程度寄与するかについての議論は要旨に含まれていない。

6. 次に読むべき論文は?

要旨で参照されているのは、state-of-the-art contact estimatorと比較されているが、具体的な名称は不明。次に読むべき論文としては、触覚予測やEgocentricビジョンと触覚の関連研究、およびロボット学習における触覚活用の研究が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenkang Zhang, Chengbo Yuan, Zicheng Zhang, Zhengxue Cheng, Yang Gao

分類: cs.CV, cs.RO

原文アブストラクト

Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce EgoTac, a generalizable model that predicts rich tactile information directly from egocentric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.