日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.08133

VLA-ACL: 行動整合性に基づく視覚トークン刈り込みによる効率的なVision-Language-Actionモデル

VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

ベースのVLAモデルを凍結したまま、行動レベルの教師信号で軽量な視覚トークン選択ポリシーを学習し、最大87.5%の視覚トークンを削減しつつ操作性能を維持する手法を提案。

詳しい要約

1. どんなもの?

VLA-ACLは、Vision-Language-Action (VLA) モデルの推論コストを削減するための、Action Consistency Learningに基づくvisual token pruning手法である。 - 対象: ロボット操作を行うVLAモデル。 - 問題: 各制御ステップで長いtoken系列を処理するため計算コストが高く、リアルタイム展開を制限。 - 着眼: 入力の大部分を占めるvisual patchには冗長性がある。 - 提案: 軽量なvisual token pruning policyをaction-level supervisionで学習。 - 特徴: base VLAモデルは完全にfrozenのまま。

2. 先行研究と比べてどこがすごい?

既存のvisual token pruningと比べて、以下の点が異なる。 - 既存手法はattention scoreやmotion thresholdなどの間接的なtraining-free heuristicに依存。 - あるいはbase VLAモデルの高コストなfine-tuningを必要とする。 - 提案手法はbase VLAをfrozenに保ちつつ、action-level supervisionでpruning policyを学習。 - token選択を下流の制御出力への影響に直接結びつける。 - 結果として、既存のfrozen-VLA pruning手法より強い性能-効率トレードオフを実現。

3. 技術・手法の肝は?

技術の肝はAction Consistency Learningによる軽量pruning policyの学習。 - base VLAモデルは完全にfrozen。 - 学習目的: pruned visual contextから生成されるactionが、full-context teacherのactionと一致するように促す。 - 補助監督としてground-truth actionも利用。 - これによりtoken選択が下流のcontrol outputに与える影響を直接反映。 - 軽量なvisual token pruning policyをaction-level supervisionで学習する点が中心。

4. どうやって有効だと検証した?

LIBEROとreal-world manipulationタスクで検証。 - visual tokenを最大87.5%削減。 - 競争力のある性能を維持。 - 計算量を最大75%削減。 - 推論速度を1.5倍高速化。 - 既存のfrozen-VLA pruning手法より強い性能-効率トレードオフを示す。

5. 議論はある?

要旨からは不明。 - 限界や失敗ケース、pruning policyの汎化性、real-worldタスクの詳細条件などは記述されていない。 - 議論の有無や具体的な論点は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究や関連手法を挙げる。 - 既存のtraining-free heuristicによるvisual token pruning(attention score、motion thresholdなど)。 - base VLAモデルをfine-tuningするpruning手法。 - frozen-VLA pruning methods。 - 評価に用いられたLIBERO。 - 関連するVLAモデルやvisual token pruningの一般的な手法。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Owen Du, Yang Yue, Jie Zhang, Jiaqi Pi, Chi Bene Chen, Gao Huang

分類: cs.RO, cs.CV

原文アブストラクト

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.

関連論文

PR本紙発行元 EmplifAI