日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2609.24385

Tactile-JEPA: 分散型触覚センサのためのトポロジー考慮型自己教師あり表現学習

Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors

シェア:XThreadsFacebookLINEはてブBluesky

触覚センサの空間配置をグラフで表現し、マスク予測による自己教師あり学習でトポロジーを考慮した表現を獲得する手法を提案。力推定誤差を6.3%、手内姿勢誤差を20.8%削減した。

詳しい要約

1. どんなもの?

- 分散型の触覚センサ(electronic skin)向けの自己教師あり事前学習手法。 - 視覚ベースの触覚センサではなく、sparseで不規則配置のセンサ群を対象。 - センサの空間配置(connectivity graph)を利用し、topology-awareな表現を学習。 - masked sensing elementsのembeddingをunmaskedから予測するJEPA型。 - 力推定やin-hand orientation、policy learningなど下流タスクに適用。

2. 先行研究と比べてどこがすごい?

- 従来の触覚encoderはrawでnoisyな信号からscratch学習が一般的。 - 既存SSLは主にvision-based tactile sensorsに集中し、distributed electronic skinsは未開拓。 - 視覚SSLの直接流用は、sparseで不規則な配置のため最適でないと指摘。 - Tactile-JEPAはtopology-awareな事前学習で、force estimation誤差を6.3%、in-hand orientation誤差を20.8%削減。 - policy learningを含む下流タスクで一貫した改善。

3. 技術・手法の肝は?

- センサのconnectivity graphを用いてspatial maskingを誘導。 - masked sensing elementsのembeddingをunmasked remainderから予測する自己教師あり学習。 - 局所的なcontact detailsと触覚面全体のglobal stateの両方を捉える必要があると分析。 - dual-scale maskingにより、局所と全体の両スケールの表現を獲得。 - magneticおよびpiezoresistiveセンサ、単一・ペア構成に対応。

4. どうやって有効だと検証した?

- magneticおよびpiezoresistiveセンサ、異なるrobot embodiments、single- and paired-sensor configurationsを含む3つの多様なデータセットで評価。 - force estimation誤差を6.3%削減、in-hand orientation誤差を20.8%削減(prior state-of-the-art比)。 - policy learningを含む他の下流アプリケーションでも一貫した改善を確認。 - 触覚センシングの利点はencoder事前学習の品質に強く依存することを示す。

5. 議論はある?

- 触覚センシングの効果はencoder pre-trainingの品質に決定的に依存する問題を直接扱う。 - 既存SSLがvision-based tactile sensorsに偏り、distributed electronic skinsが未開拓である点を指摘。 - sparseで不規則なセンサ配置が視覚SSLの直接流用を最適でなくする要因と議論。 - 局所と全体の両方を捉える必要性を分析し、dual-scale maskingの有効性を主張。 - 具体的な限界や失敗事例については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されているprior state-of-the-artの触覚SSL手法(具体的名称は要旨からは不明)。 - vision-based tactile sensors向けの既存SSLアプローチ。 - JEPA(Joint-Embedding Predictive Architecture)関連の自己教師あり学習。 - masked modelingを用いた触覚・ロボット表現学習の関連研究。 - distributed tactile sensorsやelectronic skinの表現学習に関する同分野の定番研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Elizaveta Kovtun, Matvey Konovalov, Andrey Sakhovskiy, Semen Budennyy

分類: cs.RO, cs.AI

原文アブストラクト

Tactile sensing is an essential modality for robots performing contact-rich, dexterous manipulation, particularly under visual occlusion. While pre-trained image encoders are standard in robot learning pipelines, tactile encoders are still commonly trained from scratch from raw, noisy signals, which might limit their expressivity. Existing self-supervised learning (SSL) approaches focus predominantly on vision-based tactile sensors, leaving distributed electronic skins largely unaddressed. These sensors, however, have a distinctive property: their sensing elements are sparse and irregularly arranged over the surface they cover, which makes direct reuse of visual SSL methods suboptimal. We present Tactile-JEPA, an efficient self-supervised pre-training method that uses the spatial arrangement of tactile sensors to learn topology-aware representations. Specifically, it is trained to predict the embeddings of masked sensing elements from the unmasked remainder, using the sensor connectivity graph to guide spatial masking. Our analysis shows that effective tactile representations require capturing both local contact details and the global state of the tactile surface, which we achieve through dual-scale masking. Across three diverse datasets spanning magnetic and piezoresistive sensors, different robot embodiments, and single- and paired-sensor configurations, Tactile-JEPA reduces force estimation error by 6.3% and in-hand orientation error by 20.8% over the prior state-of-the-art, with consistent gains in other downstream applications, including policy learning. Overall, our results demonstrate that the benefit of tactile sensing depends critically on the quality of encoder pre-training, a problem which Tactile-JEPA addresses directly. Code is available at https://github.com/E-Kovtun/tactile.

関連論文

PR本紙発行元 EmplifAI