日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.34319

テキスト・視覚相乗トークンキャッシュ:効率的なVLA推論のための学習不要フレームワーク

Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルの推論を高速化するため、テキストと視覚の相乗効果を利用して注意ヘッドを選別し、キャッシュ再利用層を選択する学習不要のトークンキャッシュ手法を提案。

詳しい要約

1. どんなもの?

Vision-Language-Action (VLA) モデルの推論を高速化する training-free な token caching フレームワーク TVCache を提案する研究。 - VLA は汎用ロボット制御を可能にするが計算コストが高い。 - token caching は training-free で plug-and-play な加速手段。 - 既存 VLA caching は text-vision synergy という inductive bias を十分活用していない。 - TVCache は attention head の選択と cache 再利用層の選択を text-vision 情報に基づき行う。

2. 先行研究と比べてどこがすごい?

既存 VLA caching と比べて以下の点が優れる。 - 既存手法は attention 集約における head-wise reliability を十分考慮していない。 - 既存手法は cache 再利用における layer-wise stability を十分考慮していない。 - TVCache は text-vision 情報に基づき attention head をフィルタし、task-relevant で物理的に整合した visual grounding を改善。 - text-vision entropy 差に基づく reuse-layer selection で不安定な表現の caching を回避し、cache 資源配分を改善。 - 同一 token-retention ratio で既存 VLA caching より task success を一貫して改善。

3. 技術・手法の肝は?

技術の肝は text-vision synergy を活用した 2 つの選択機構。 - attention head を text-vision information focus に基づきフィルタし、task-relevant かつ物理的に整合した visual grounding を実現。 - text-vision entropy 差に基づく reuse-layer selection 機構を導入。 - 不安定な表現の caching を避け、cache 資源配分を改善。 - training-free で plug-and-play な枠組み。

4. どうやって有効だと検証した?

以下の実験で有効性と汎用性を検証。 - 4 つの代表的な VLA モデル。 - 2 つの simulation benchmark。 - 実世界ロボットタスク。 - 同一 token-retention ratio で既存 VLA caching より task success を一貫して改善。 - OpenVLA-OFT では 12.5% retention 時に VLA-Cache 比で平均 success を最大 14.5 percentage points 改善。 - full-token inference 比で FLOPs を 2.45x 削減。

5. 議論はある?

要旨からは不明。 - 限界や失敗ケース、計算コストの詳細な内訳、他手法との理論的比較などは記述されていない。

6. 次に読むべき論文は?

要旨で参照・比較されている研究として VLA-Cache が挙げられる。 - また OpenVLA-OFT が評価対象モデルとして言及されている。 - 関連手法として token caching 全般、VLA モデル、Vision-Language-Action 推論高速化の研究が次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Qianer Li, Chengjie Zhang, Jingwen Chen, Zanjia Tong, Jiyuan Zhang, Hong Zhang

分類: cs.CV, cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.

関連論文

PR本紙発行元 EmplifAI