日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLA量子化arXiv:2610.02666

CHASE-VLA: チャンク認識型スケール推定による視覚-言語-行動モデルの訓練後量子化フレームワーク

CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルの拡散ベース行動エキスパートをW4A4に量子化する際、生成済み行動チャンクとノイズ除去ステップ情報を活用して活性化スケールを適応させ、精度を維持しつつメモリを大幅削減する手法を提案。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルを低ビットで量子化する Post-Training Quantization (PTQ) フレームワーク。 - 特に diffusion-based action expert (AE) の W4A4 量子化を対象。 - 生成済み action chunk を活用する chunk-aware なスケール推定が特徴。 - 事前学習済みポリシーを変更せずに適用可能。

2. 先行研究と比べてどこがすごい?

- 従来の PTQ は固定 calibration scale を用いるため、denoising 進行や意図された動作で変動する AE の activation range とミスマッチが生じる。 - CHASE-VLA は VLA 固有の信号である action chunk (未実行の future suffix を含む) を利用。 - これにより AE の MLP と attention projection の両方を W4A4 量子化しても FP16 相当の性能を回復。 - 従来手法との直接比較の詳細は要旨からは不明。

3. 技術・手法の肝は?

- 生成済み action chunk を causal action context として利用。 - denoising step group 情報と組み合わせ、AE の activation scale を適応的に調整。 - これにより AE 層の静的スケールマッチングのみに依存しない。 - 事前学習済みポリシーを変更せず、MLP と attention projection の W4A4 量子化を実現。

4. どうやって有効だと検証した?

- LIBERO ベンチマークで評価。 - π_{0.5} において AE の MLP と attention projection を W4A4 量子化した際、平均成功率 97.3% を達成し FP16 レベルの性能を回復。 - 量子化された AE 線形層の重みストレージを 73.4% 削減。 - 単一 chunk のメモリトラフィックを π_{0.5} で 70.9%、GR00T N1.6 で 71.2% 削減。 - predictor のオーバーヘッドは最大で削減ストレージの 1.26%。

5. 議論はある?

- 要旨からは不明。 - 想定される議論点: 他 VLA モデルやタスクへの汎化性、量子化時の精度劣化、predictor の追加計算コストなど。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: 明示的な参照はない。 - 関連手法: Post-Training Quantization (PTQ)、Vision-Language-Action (VLA) モデル、diffusion-based action expert (AE)、W4A4 量子化。 - 同分野の定番: LIBERO ベンチマーク、π_{0.5}、GR00T N1.6 などの VLA モデル。 - 具体的な次に読むべき論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jin Hyun, Jung Gyu Min, Gyuhyun Jung, Youngjoo Lee

分類: cs.CV

原文アブストラクト

Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, where fixed calibration scales can be mismatched with activation ranges that vary with denoising progress and intended motion. We propose CHASE-VLA, a chunk-aware PTQ method that exploits a VLA-specific signal readily available from the policy: the generated action chunk, including its unexecuted future suffix. Rather than relying only on static scale matching for AE layers, CHASE-VLA combines the previously generated chunk as causal action context with denoising step group information to adapt AE activation scales. This enables W4A4 quantization of both MLP and attention projections in the repeated AE without modifying the pretrained policy. On LIBERO, CHASE-VLA achieves 97.3% average success rate on $π_{0.5}$ when both MLP and attention projections in the AE are quantized to W4A4, restoring FP16-level performance. CHASE-VLA also reduces the weight storage of the quantized AE linear layers by 73.4% and their single-chunk memory traffic by 70.9% and 71.2% on $π_{0.5}$ and GR00T N1.6, respectively, with a predictor overhead of at most 1.26% of the saved storage.

関連論文

PR本紙発行元 EmplifAI