日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.05273

いつ何を刈り込む?VLA効率化のための段階認識型視覚トークンプルーニング

When and What to Prune? Stage-Aware Visual Token Pruning for Efficient VLA

シェア:XThreadsFacebookLINEはてブBluesky

VLA推論で視覚トークンを削減する際、層ごとの注意パターンを校正して信頼できる段階でのみ刈り込み、注目トークンと周辺文脈を保護する学習不要の手法を提案。

著者: Tianjun Shi, Haotian Xiong, Ziyu Gong, Qi Lu, Li Li

分類: cs.CV, cs.RO

原文アブストラクト

Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.

関連論文

PR本紙発行元 EmplifAI