日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.24433

FoldQuantVLA: 一貫した折り畳みによる視覚-言語-行動モデルのネイティブローbit量子化

FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding

シェア:XThreadsFacebookLINEはてブBluesky

視覚-言語-行動モデルを再学習なしで4bit量子化し、TensorRTプラグインで高速化する手法を提案。実機タスクで成功率を維持・向上させた。

詳しい要約

1. どんなもの?

- FoldQuantVLAは、Vision-Language-Action (VLA) モデル向けの post-training quantization (PTQ) フレームワーク。 - 低ビット推論により observation-to-action レイテンシを削減しつつ、ロボットの挙動を保持することを目的とする。 - キャリブレーション、重み丸め、ネイティブ整数実行を通じて一貫した活性化表現を維持する。 - ポリシーの再学習なしで、W4A4 (4-bit weights and activations) を実現。 - TensorRT プラグインにより、言語バックボーンと反復 action expert の投影を Ada GPUs および Jetson AGX Orin 上で実行。

2. 先行研究と比べてどこがすごい?

- 従来の PTQ と異なり、キャリブレーションから実行まで一貫した活性化表現を維持する点が新しい。 - ポリシー再学習なしで W4A4 を達成し、浮動小数点 TensorRT 比で Orin 上 1.20–1.33×、デスクトップ上 1.25–1.52× の高速化を実現。 - 言語 attention-output と feed-forward down 投影を 8-bit (W8A8) に保持することで、4 つのチェックポイントすべてで held-out action fidelity が向上。 - 実ロボットタスクで GR00T N1.7 の成功率を 80.0% (uniform W4A4) から 92.5% に改善。

3. 技術・手法の肝は?

- channel scaling と block Hadamard transforms を dynamic per-token quantization と組み合わせる。 - キャリブレーション、重み丸め、ネイティブ整数実行を通じて一貫した活性化表現を維持。 - カスタム TensorRT プラグインにより、言語バックボーンと反復 action expert の投影を W4A4 で実行。 - 一部の投影 (language attention-output と feed-forward down) を W8A8 に保持する混合精度戦略を採用。

4. どうやって有効だと検証した?

- LIBERO、SimplerEnv、および 2 つのロボットプラットフォームで評価。 - 3 つの GR00T チェックポイントと π_{0.5} で W4A4 の速度向上を測定。 - 4 つの実ロボットタスクで、W8A8 保持構成が GR00T N1.7 の成功率を 80.0% から 92.5% に向上 (各構成 80 試行)。 - Orin 上での追加レイテンシは 1 ms と測定。

5. 議論はある?

- 要旨からは、限界や議論点についての明示的な記述は不明。 - ただし、W4A4 の uniform 適用よりも一部を W8A8 に保持する方が action fidelity に有効であることが示唆される。 - 実ロボットでの成功率向上とレイテンシ増加 (1 ms) のトレードオフが議論の余地として考えられるが、要旨では深く触れられていない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: GR00T (N1.7 を含む)、π_{0.5}、LIBERO、SimplerEnv。 - 関連手法: post-training quantization (PTQ)、TensorRT、Hadamard transform、per-token quantization。 - 同分野の定番: Vision-Language-Action (VLA) モデル、量子化手法 (W4A4, W8A8)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen, Thanh Q. Duong, Ngan Le, Meng Guo, Vien A. Ngo, An T. Le

分類: cs.RO, cs.AI, eess.SY

原文アブストラクト

Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and $π_{0.5}$, W4A4 achieves $1.20$ to $1.33\times$ speedups over floating-point TensorRT on Orin and $1.25$ to $1.52\times$ on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from $80.0\%$ with uniform W4A4 to $92.5\%$ over 80 trials per configuration, with a measured additional Orin latency of 1 ms.

関連論文

PR本紙発行元 EmplifAI