日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
コンテキスト圧縮arXiv:2610.05885

コンパクト化されたコンテキストモデル学習の最適化景観

The Optimization Landscape of Learning Compacted Context Models

シェア:XThreadsFacebookLINEはてブBluesky

凍結したベースモデル上でKVキャッシュを圧縮する最適化問題の難しさを分析し、簡略化したPerceiverアーキテクチャが完全なPerceiverと同等以上の圧縮性能を達成することを示した。

著者: Thomas Villeneuve, Alex Sandomirsky, Charles O'Neill, Max Kirkby, Michael Psenka

分類: cs.LG

原文アブストラクト

Many works approach continual learning through the lens of infinite context windows. As an agent puts more observation into context (concretely the KV cache), compacting said context is akin to direct memory manipulation, without affecting the base model's weights. Many works pose KV compaction as an optimization problem: learn a smaller set of KV vectors that matches the behavior of the full KV cache. While this preserves base model behavior, optimizing through a frozen base model results in a highly nontrivial optimization problem with a brittle and flat loss landscape. In this paper, we characterize what makes these optimization problems difficult and demonstrate that a heavily simplified Perceiver-based architecture not only matches performance of a full Perceiver transformer in continuous context compaction, but outperforms baselines on compaction utility. Results are presented on MCQ tasks across Finance, Legal, Gutenberg, and Code.

PR本紙発行元 EmplifAI