本論文は、Haptic Foundation Models (HFMs) の可能性と発展経路を探る展望論文である。言語や視覚におけるfoundation modelsの成功を背景に、embodied AIへの拡張が触覚の汎用性不足によって制限されていると指摘する。特に、スマートフォン、ウェアラブル、VRコントローラ、家庭用ロボット、健康モニタリング機器など、安全で適応的な物理的相互作用を必要とする民生電子機器において重要である。現在の触覚モデルはハードウェアの不均一性と能動的な物理データ収集の必要性により、タスク特化型に留まっている。本論文は、受動的なLarge Language Models (LLMs) やVision Language Models (VLMs) から能動的なHFMsへのパラダイムシフトを、action coupling、physical dynamical representation space、continuous time-series data granularity、action-conditioned future state predictionの…
Despite the success of foundation models in language and vision, their expansion into embodied AI is bottlenecked by a lack of generalized touch sensing. This limitation is especially relevant to consumer electronics, where smartphones, wearables, VR controllers, home robots, and health monitoring devices require safe and adaptive physical interaction. Constrained by hardware heterogeneity and the necessity of active physical data collection, current haptic models remain rigidly task-specific. To overcome these limitations, this article explores the transformative potential and developmental trajectory of Haptic Foundation Models (HFMs). We detail the paradigm shift required to transition from passive Large Language Models and Vision Language Models into active HFMs across four core dimensions: action coupling, physical dynamical representation space, continuous time-series data granularity, and action-conditioned future state prediction. Furthermore, we synthesize existing large-scale tactile datasets and benchmark UniTouch, AnyTouch, T3, and Sparsh on TacBench for force estimation, slip detection, and relative pose estimation.