日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.20980

ForeTac-VLA: 接触の多いロボットマニピュレーションのための触覚予測に基づく視覚-言語-行動モデル

ForeTac-VLA: A Forecasting-Based Tactile-Vision-Language-Action Model for Contact-Rich Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

未来の触覚状態を予測して行動生成を導く触覚・視覚・言語融合モデルを提案し、接触の多い実世界タスクで高い成功率を達成した。

詳しい要約

1. どんなもの?

- 接触の多いロボットマニピュレーション向けのVLAモデル。 - 視覚だけでなく触覚を統合し、将来の触覚状態を予測して行動生成を導く。 - 名前はForeTac-VLA。 - 視覚言語特徴と触覚時系列を双方向cross-attentionで融合。 - transformerベースのforecasting moduleで多段先の触覚を予測。 - 予測と融合表現をVLA backboneに入れ行動を条件付け。

2. 先行研究と比べてどこがすごい?

- 既存のtactile-enhanced VLAは観測触覚を使うが反応的で、接触の進展を明示的にモデル化しない。 - 本手法は将来触覚を予測し、観測と予期の接触を同時に推論。 - 4つの実世界contact-richタスクで平均成功率95%。 - fine-tuned VLAを36.25ポイント上回る。 - state-of-the-art tactile-enhanced VLAベースラインを22ポイント以上上回る。 - 低照度・視覚的 clutter 下でも性能維持。

3. 技術・手法の肝は?

- 直近の触覚観測をtemporal representationに符号化。 - 視覚言語特徴とbidirectional cross-attentionで統合。 - transformer-based forecasting moduleがmulti-step future tactile statesを予測。 - 融合multimodal representationと予測触覚をVLA backboneへ入力し行動生成を条件付け。 - 訓練安定化のためground-truth-to-prediction curriculumを採用。 - 初期予測が不安定な段階では正解触覚を利用。

4. どうやって有効だと検証した?

- 4つの実世界contact-rich manipulationタスクで評価。 - 平均成功率95%を達成。 - fine-tuned VLAモデルを36.25 percentage points上回る。 - state-of-the-art tactile-enhanced VLAベースラインを22 percentage points以上上回る。 - 低照度および視覚的clutter条件下でも強い性能を維持。 - ビデオデモは https://foretac-vla.github.io/ で公開。

5. 議論はある?

- 要旨からは不明。 - 限界、失敗事例、計算コスト、触覚センサ依存性、一般化性についての議論は要旨に記載なし。 - ground-truth-to-prediction curriculumの詳細な設計や感度分析も要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: fine-tuned VLA model、state-of-the-art tactile-enhanced VLA baselines。 - 関連手法としてtactile-enhanced VLA、vision-language-action (VLA) models、bidirectional cross-attention、transformer-based forecasting。 - 具体的な論文名は要旨に記載がないため、同分野の定番としてtactile-enhanced VLAやcontact-rich manipulation向けVLAを挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhengyu Tao, Xin Li, Xin Wang

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, yet their reliance on visual perception limits robustness in contact-rich environments, where critical physical interaction states may not be visually observable. Existing tactile-enhanced VLA methods improve physical grounding using observed tactile feedback, but most remain largely reactive rather than explicitly modeling how contact may evolve. Therefore, we propose ForeTac-VLA, a forecasting-based tactile-vision-language fusion model that predicts future tactile states to guide action generation. Specifically, ForeTac-VLA encodes recent tactile observations into temporal representations and integrates them with vision-language features through bidirectional cross-attention. Further, a transformer-based forecasting module predicts multi-step future tactile states, enabling the model to reason jointly over observed and anticipated contact. Finally, the fused multimodal representations and predicted future tactile states are fed into the VLA backbone to condition action generation. To stabilize training, a ground-truth-to-prediction curriculum is employed when early forecasts are unreliable. Across four real-world contact-rich manipulation tasks, ForeTac-VLA achieves an average success rate of 95%, outperforming the fine-tuned VLA model by 36.25 percentage points and state-of-the-art tactile-enhanced VLA baselines by over 22 percentage points. ForeTac-VLA also maintains strong performance under low-illumination and visually cluttered conditions. Video demonstrations can be found on https://foretac-vla.github.io/

関連論文

PR本紙発行元 EmplifAI