日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2610.02784

SimpleTouch: 触覚ポリシー事前学習なしでVLAモデルは接触の多いマニピュレーションを習得できるか

SimpleTouch: Can Vision-Language-Action Models Master Contact-Rich Manipulation Without Tactile Policy Pretraining?

シェア:XThreadsFacebookLINEはてブBluesky

触覚エキスパートを追加したVLA拡張SimpleTouchを提案し、触覚ポリシーの大規模事前学習や視触覚アライメントなしでも、タスク実演のみの単段階学習で接触の多い操作タスクを高成功率で達成できることを示した。

詳しい要約

1. どんなもの?

ビジョン・言語・行動(VLA)モデルに触覚を組み込む手法「SimpleTouch」を提案する研究。 - 対象は接触を伴うロボットマニピュレーション。 - 既存の$\pi_{0.5}$を拡張し、触覚エキスパートを追加。 - 触覚ポリシーの大規模事前学習や視触覚アラインメントを不要にし、単一ステージのファインチューニングのみで学習。 - 6つのUniVTACタスクと4つの実世界タスクで評価。

2. 先行研究と比べてどこがすごい?

触覚をVLAに統合するには大規模な触覚ポリシー事前学習や視触覚アラインメントが必要とされてきた。 - 本研究はそれらが必須でないことを示す。 - タスクごとに50デモのみで、UniVTAC全6タスクで最高成功率。 - 平均成功率77.5%はFTP-$\pi_{0.5}$の45.2%を32.3ポイント、FTP-1の66.7%を10.8ポイント上回る。 - 実世界4タスクでも平均71.3%でFTP-1を8.8ポイント上回る。

3. 技術・手法の肝は?

SimpleTouchは$\pi_{0.5}$に触覚エキスパートを追加したVLA拡張。 - 凍結した事前学習済み触覚エンコーダの全トークンを利用。 - エキスパートは行動教師信号と、将来の触覚潜在表現のマルチホライズン予測から学習。 - 学習はタスクデモのみを用いる単一ステージで、追加の触覚ポリシー事前学習やアラインメントは行わない。

4. どうやって有効だと検証した?

6つのUniVTACタスクと4つの実世界タスクで検証。 - 各タスク50デモで学習。 - UniVTAC全6タスクで評価手法中最高の成功率、平均77.5%。 - 比較対象はFTP-$\pi_{0.5}$(45.2%)とFTP-1(66.7%)。 - 実世界4タスクでは平均71.3%でFTP-1を8.8ポイント上回る。

5. 議論はある?

事前学習済みVLAと触覚表現があれば、追加の触覚ポリシー事前学習はこれらのタスクで強性能を得る前提条件ではないと主張。 - 接触リッチなマニピュレーションへのより単純な道筋を提示。 - 限界や失敗事例、他のタスクへの一般化については要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている研究:FTP-$\pi_{0.5}$、FTP-1、$\pi_{0.5}$、UniVTAC。 - 関連手法として触覚ポリシー事前学習や視触覚アラインメントを用いる既存研究が挙げられるが、個別の論文名は要旨からは不明。 - 同分野の定番としてVision-Language-Actionモデル、触覚エンコーダ、模倣学習に関する研究を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chen Yang, Linzhe Shi, Changjie Wu, Hang Zhang, Ronghan Chen, Lingjun Zhang, Xu Hu, Mu Xu, Jiansheng Fan, Chen Wang

分類: cs.RO

原文アブストラクト

Tactile sensing provides essential contact information for robotic manipulation, yet incorporating it into pretrained vision-language-action (VLA) models remains challenging. A common concern is that simply introducing touch during task-specific fine-tuning may fail to bridge the cross-modal gap, yielding limited gains or even reduced success. Consequently, existing methods often rely on large-scale tactile policy pretraining or separate visuotactile alignment, adding data requirements and training stages. We introduce SimpleTouch, a simple VLA extension that augments $π_{0.5}$ with a tactile expert, to test whether these additional stages are necessary. Leveraging all tokens from a frozen pretrained tactile encoder, the expert learns from action supervision and multi-horizon prediction of future tactile latents. This single-stage training uses only task demonstrations, without additional tactile policy pretraining or separate alignment. With 50 demonstrations per task, SimpleTouch achieves the highest success rate among evaluated methods on all six UniVTAC tasks. Its average success rate reaches 77.5%, compared with 45.2% for FTP-$π_{0.5}$ and 66.7% for FTP-1, corresponding to gains of 32.3 and 10.8 percentage points, respectively. Across four real-world tasks, it averages 71.3%, exceeding FTP-1 by 8.8 percentage points. These results demonstrate that, given pretrained VLA and tactile representations, additional tactile policy pretraining is not a prerequisite for strong performance on these tasks, offering a simpler route to contact-rich manipulation. Project page: https://simpletouch-robot.github.io/

関連論文

PR本紙発行元 EmplifAI