日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
学習最適化arXiv:2605.17923

AdaptiveLoad: 効率的なビデオ拡散トランスフォーマー学習に向けて

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training

シェア:XThreadsFacebookLINEはてブBluesky

ビデオ生成モデルの学習におけるGPU負荷不均衡を解消するため、メモリと計算量の二重制約による適応的負荷分散と、融合LayerNorm-Modulate CUDAカーネルを提案し、スループットを27.2%向上させた。

著者: Yucheng Guo, Yongjian Guo, Zhong Guan, Haoran Sun, Wen Huang, Wanting Xu, Jing Long, Shuai Di, Junwu Xiong

分類: cs.DC, cs.AI, cs.LG

原文アブストラクト

In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to the extreme variance in sequence lengths within mixed-mode datasets. Existing bucket-based data loading strategies typically rely on "equal token length" constraints. This approach fails to account for the quadratic complexity of self-attention mechanisms, leading to severe load imbalance and underutilization of GPU resources. This paper proposes \textit{AdaptiveLoad}, an integrated optimization framework consisting of two core components: (1) A dual-constraint adaptive load balancing system, which eliminates long-sequence bottlenecks by simultaneously limiting memory consumption and computational load ($B \times S^p \le M_{\text{comp}}$); (2) A fused LayerNorm-Modulate CUDA kernel, which utilizes a D-tile coalesced reduction strategy to increase throughput and alleviate memory pressure. Experimental results on the Wan 2.1 world model demonstrate that our method reduces the computational imbalance rate from 39\% to 18.9\%, improves peak VRAM utilization efficiency by 22.7\%, and achieves an overall training throughput increase of 27.2\%.