日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
量子化arXiv:2610.06260

CentriQ: 平均中心化による拡散トランスフォーマーのキャリブレーション不要量子化

CentriQ: Calibration-Free Quantization of Diffusion Transformers via Exact Mean Centering

シェア:XThreadsFacebookLINEはてブBluesky

拡散トランスフォーマーを4bit量子化する際、適応的レイヤー正規化が生むトークンごとの平均がHadamard回転で均一化できない問題を指摘し、回転前に中心化してランク1の全精度ブランチで平均を復元するキャリブレーション不要手法を提案した。

詳しい要約

1. どんなもの?

本論文は、Diffusion Transformer (DiT) の推論コストを削減するための、キャリブレーション不要な 4-bit 量子化手法 CentriQ を提案する。 - 対象: DiT の weights と activations の両方を 4-bit に量子化 - 課題: 既存の calibration-based 手法は checkpoint や prompt 分布に依存し、data-free Hadamard rotation は DiT で品質が落ちる - 提案: token ごとに mean centering を行い、rank-1 の full-precision branch で mean を厳密に復元する - 特徴: データ不要で per-token scale を閉形式で決定し、weights は robust な ℓp 目的関数でフィットする

2. 先行研究と比べてどこがすごい?

先行研究と比べて、以下の点が優れている。 - calibration-based 手法 (SVDQuant など) は特定の checkpoint と prompt 分布に縛られるが、CentriQ は calibration-free である - data-free Hadamard rotation は LLM では有効だが DiT では品質が落ちる。本論文はその構造的原因 (Adaptive layer-norm conditioning による per-token mean が回転後も dominant direction として残る) を明らかにした - 3 つの DiT で、CentriQ は 4-bit において calibrated SVDQuant と同等の品質を達成 - 2-bit weights では、報告されている中で最強の calibration-free 手法を上回る - 2-bit activations で実用的な画像品質を保持した初の calibration-free 手法である

3. 技術・手法の肝は?

技術の肝は以下の通り。 - Adaptive layer-norm conditioning が加える per-token mean に着目し、Hadamard rotation 前に各 token を centering する - mean を rank-1 の full-precision branch で厳密に復元することで、per-token scale をデータなしで閉形式に導出する - weights は robust な ℓp 目的関数でフィットし、各グループの dense mode を追跡しつつ heavy tails を割り引く - これにより calibration 不要で 4-bit 量子化を実現する

4. どうやって有効だと検証した?

3 つの DiT を用いて検証している。 - 4-bit において、CentriQ は calibrated SVDQuant と同等の品質を達成 - 一方、plain per-token activation quantization を用いた calibration-free weight quantizer は崩壊または大幅に劣化する - 2-bit weights では、報告されている最強の calibration-free 手法を上回る - 2-bit activations では、calibration-free 手法として初めて実用的な画像品質を保持することを示した

5. 議論はある?

要旨からは不明。 - 限界や失敗事例、計算コスト、他の DiT への一般化、rank-1 branch のオーバーヘッドなどについての議論は要旨に記載されていない

6. 次に読むべき論文は?

要旨で参照・比較されている研究や関連手法を挙げる。 - SVDQuant (calibrated な量子化手法) - data-free Hadamard rotation (LLM 向けの calibration-free 手法) - Adaptive layer-norm conditioning を用いる DiT - plain per-token activation quantization を用いた calibration-free weight quantizer - 同分野の定番として、Diffusion Transformer (DiT) や量子化一般 (weights/activations の 4-bit 量子化) に関する文献

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Nataša Jovanović, Mathieu Salzmann, Saqib Javed

分類: cs.CV, cs.AI, cs.RO

原文アブストラクト

Diffusion transformers (DiTs) achieve state-of-the-art image generation, but their sampling cost limits deployment. Quantizing both weights and activations to 4 bits reduces this cost, yet existing methods fall short in one of two ways. Calibration-based methods are tied to a specific checkpoint and prompt distribution, whereas data-free Hadamard rotation, effective for LLMs, loses quality on DiTs. We show that this loss has a structural cause. Adaptive layer-norm conditioning adds a per-token mean to the activations, and at the widths of the evaluated DiTs, the Hadamard rotations used by data-free methods cannot spread this mean uniformly across coordinates. A single dominant direction therefore survives the rotation and sets the quantization range. We introduce CentriQ, a calibration-free quantizer that centers each token before rotation and restores the mean exactly through a rank-1 full-precision branch, so that per-token scales follow in closed form without data. Weights are fitted under a robust $\ell_p$ objective that tracks the dense mode of each group and discounts heavy tails. Across three DiTs, CentriQ matches the quality of calibrated SVDQuant at 4 bits, whereas calibration-free weight quantizers with plain per-token activation quantization collapse or degrade substantially. CentriQ outperforms the strongest calibration-free method reported to date at 2-bit weights. It is also the first calibration-free method to retain usable image quality at 2-bit activations.

関連論文

PR本紙発行元 EmplifAI