サポートを分割し、残差を再構成する:ビデオ生成とワールドモデルのための学習不要スパースアテンション
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
ビデオ生成やワールドモデルのトランスフォーマーを高速化するため、学習不要のブロックスパースアテンション手法SparsePRを提案。応答結合分割とプローブフィット残差再構成を組み合わせ、アテンション再構成誤差を低減しつつ、実行ペア密度22-26%で1.48-2.61倍の高速化を達成。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Pardis Taghavi, Reza Langari, Gaurav Pandey
分類: cs.CV, cs.AI, cs.LG
原文アブストラクト
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post-softmax error from skipped interactions. We show that partition geometry affects both pooled support and the predictability of the remaining residual from the sparse output. We introduce SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. Sampled-query key responses form paired K/V groups, whose centroids induce query-response coordinates for shared routing. A small set of exact query rows then calibrates a call-specific affine correction from the sparse output within the output subspace observed in the probe residuals. Across four heterogeneous video generation and world models, SparsePR consistently reduces attention-reconstruction error. Ablations show that probe fitting accounts for most of this reduction, while response-coupled partitioning lowers hard-drop error and improves reconstruction under a finite probe budget. SparsePR preserves generation quality at 22.0-26.0% realized executed-pair density while achieving 1.48x-2.61x end-to-end speedups. Project page: https://pardistaghavi.github.io/SparsePR-website/