Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
1ByteDance Seed2Princeton University3Stanford University4University of California, Berkeley
*Core contributors. Work done during internship at ByteDance Seed.†Corresponding authors.
Summary
- Chunked KV compression gives every token a phase, its position modulo the compression stride \(S\). A model is phase sensitive when its accuracy depends on the phase of the information it has to retrieve.
- DeepSeek-V4 and DeepSeek-V4.1, which compress their KV caches in chunks, are phase sensitive. At 128K tokens, needle retrieval accuracy differs by up to 40 points across target positions, with a period equal to the stride. When the text before a line of code from DeepSeek-V4's own FP8 kernel is lengthened one token at a time, DeepSeek-V4-Flash-Base alternates between two completions of that line, with a period of four tokens, equal to its stride.
- Small models pretrained from scratch, which differ from full-attention baselines substantially only in their chunked compression, are also phase sensitive, while the baselines are not.
- In the trained models, knocking out some KV heads lowers accuracy at only a few phases, and some learned compression gates put most of their weight on the same slots of every window. In an idealized model, theory shows that such concentration is optimal and that training makes it the same for every input. The theory also suggests why the heads may fail to cover every phase.
A coordinate created by compression
Many long-context designs keep a compressed memory of past tokens [1, 2, 3, 4], including the compressed sparse attention of DeepSeek-V4 [5]. A chunked compressor reads the context in windows of \(W\) tokens that start every \(S\) tokens and writes one cache entry per window. Write \(\mathbf h_{j,i}\) for the hidden state at offset \(i\), its slot \(0,\dots,W-1\) within window \(j\). A gate turns logits into scores over the window, and the scores weight payloads \(\mathbf C_i\mathbf h_{j,i}\). Each offset \(i\) has its own gate parameters \(\mathbf Z_i\) and \(\mathbf b_i\) and its own linear payload map \(\mathbf C_i\):
The score \(\boldsymbol\alpha_{j,i}\) can be one scalar per offset or one score per channel, hence the elementwise product. Values are compressed the same way. Later queries attend over these entries together with a short local window of exact recent tokens. An entry is written before any later query exists, so it must retain whatever a later query might need. Because windows start every \(S\) tokens, each token at position \(t\) also has a phase:
Three related terms recur below. The offset is a token's slot within one window, the phase is its position modulo the stride, and a residue is its position modulo a multiple of \(S\), which lets models with different strides share one axis. With non-overlapping windows (\(S=W\)) offset and phase coincide. With overlapping windows (\(S<W\)) a token appears at several offsets, but its phase is still fixed. With non-overlapping windows, a token at phase \(S-1\) is the last of its window, so the token after it falls into the next entry, and we call \(S-1\) the boundary phase. Shifting the input by a few tokens changes which tokens share an entry without changing their content or order, so if retrieval depends on \(\phi\), that dependence is a property of the model, not of the data.
Weak spots in DeepSeek-V4
DeepSeek-V4 comes in a smaller model, Flash, and a larger one, Pro. The suffix -Base marks a pretrained checkpoint, and -0731 and -0813 mark post-trained ones.
A flipping completion
We ask DeepSeek-V4-Flash-Base to complete a line of its own FP8 quantization kernel, with the cursor after T.Cast(FP. The correct continuation is 8, since 32 would duplicate an outer cast. We prepend a docstring and lengthen it one token at a time. The code does not change, but the top prediction switches between 8 and 32 every two tokens, so the pattern repeats every four tokens (Figure 1). Four is the stride of DeepSeek-V4's compressed sparse attention (\(W=8, S=4\)) [5]. Across four families of filler text with 16 lengths each, 60 of 64 rankings follow this pattern. Under several filler sweeps, the post-trained DeepSeek-V4-Flash-0731 (\(S=4\)) and DeepSeek-V4.1-Flash (\(S=2\)) [6] also switch periodically between 8 and 32. DeepSeek-V3.1-Base, which has no chunked compression, prefers 32 on only 4 of the same 64 inputs, each by a probability margin below 0.07.
T.Cast(FP in its own FP8 kernel. The filler length \(L\) counts every docstring token before the code. The docstring's opening sentence is 15 tokens, and each added = sign adds one token, so 15 signs give \(L=30\). The code below never changes. It only shifts by one position per added sign. The model prefers the correct 8 when \(L \bmod 4\in\{2,3\}\) and 32 otherwise, a period of four that matches its compression stride. Probabilities are measured values, and dashed lines in the strip separate cycles of four.A position-matched needle test
Each prompt holds 128K tokens with 16,000 four-token key–value records, and asks for the value of one target key near the middle of the context. Let \(t_K\) be the position of the target key. We group prompts by the residue \(r=t_K \bmod 8\), which covers two stride cycles for DeepSeek-V4 and four for DeepSeek-V4.1, and write \(A(r)\) for the answer accuracy on group \(r\). Filler lengths are chosen so that the query position is fixed and
so the groups differ in phase but not in these position statistics. Each group has 256 prompts for each of the four DeepSeek-V4 models (Flash-Base, Flash-0731, Pro-Base, and Pro-0813). DeepSeek-V4.1-Flash was evaluated in a separate, larger run with 2,560 prompts per group.
| Model | \(S\) | mean | worst | best | gap |
|---|
The accuracy pattern repeats with the stride. For the four DeepSeek-V4 models, the mean of \(|A(r)-A(r+4)|\) over \(r\in\{0,1,2,3\}\) is 1.1 to 2.0 points, small next to gaps of 15 to 40 points. DeepSeek-V4.1-Flash compresses with \(S=2\), and all four of its even residues stay above all four odd ones. Post-training raises accuracy and narrows the gap, but it keeps the period.
Controlled pretraining
DeepSeek-V4 differs from a plain transformer in many ways besides compression, so we pretrain our own models. The backbone is Qwen3-0.6B [7] (28 layers, width 1,024, 16 query heads, head dimension 128). Every model is trained from scratch on 100B tokens, with chunked compression as the only substantial architectural change from the full-attention baselines. We name each model by its signature attributes. W8/S8 · 1 KV has window \(W=8\), stride \(S=8\), and one KV head per layer, and anything else that differs is appended, as in W8/S8 · 1 KV · no RoPE. Unless its name says otherwise, a compressed model has learned scalar gates, separate key and value branches, and RoPE, and it attends jointly to compressed history and a 16-token local window of exact tokens. W8/S8 · 1 KV also has three extra seeds (43, 44, and 45), and without a seed the name means seed 42. Across the sweep, we vary \(W\) and \(S\) independently, the number of KV heads, and these attributes:
- vector gates: one gate score per channel rather than one scalar per offset.
- tied K/V: values reuse the normalized, rotated keys instead of their own branch.
- no gate bias: the offset biases \(\mathbf b_i\) are removed.
- no Q/K norm: queries and keys skip RMSNorm.
- no RoPE or RoPE on 16 dims: no rotary embedding, or one on 16 of the 128 head dimensions.
- uncompressed tail: raw tokens are kept only for the current incomplete window, instead of a fixed 16-token window.
- DeepSeek-style kernel: W8/S4 with vector gates, tied K/V, and RoPE on 16 dimensions, which matches DeepSeek-V4's compression kernel. A variant with scalar gates gives each token one score instead of one per channel.
- uniform averaging: every offset gets weight \(\boldsymbol\alpha_{j,i}=1/W\), with no learned gate.
- post-norm: a normalization after the payload projection.
The evaluation is an animal-to-name lookup. Each prompt lists 64 records, each a single-token animal key followed directly by a single-token name, and ends with eight demonstrations and a target animal. The model must output that animal's name, the token right after its key. We report full-vocabulary top-1 accuracy, grouped by the target key's position modulo 24, a common multiple of all tested strides.
Prefix padding shifts all records together. In-sequence padding varies the gaps between records, so that distractor phases do not move in lockstep with the target. Both hold the query position and the first two moments of the target position fixed.
Full attention
1, 4, and 8 KV heads · prefix padding
| Model | val. loss | mean | worst | gap |
|---|---|---|---|---|
| Full attention · 1 KV | 2.466 | 18.8 | 17.2 | 3.1 |
| Full attention · 8 KV | 2.437 | 61.1 | 58.4 | 4.3 |
| W8/S8 · 1 KV · seed 42 | 2.478 | 50.6 | 5.7 | 75.3 |
| W8/S8 · 8 KV | 2.431 | 59.4 | 9.9 | 71.1 |
| W8/S4 · 1 KV | 2.467 | 66.9 | 46.2 | 39.8 |
| W12/S12 · 1 KV | 2.487 | 33.0 | 5.0 | 63.2 |
| W8/S8 · 1 KV · no RoPE | 2.504 | 29.4 | 5.0 | 40.9 |
| W8/S8 · 1 KV · uniform averaging | 2.497 | 12.5 | 4.7 | 19.2 |
Selected runs, prefix padding, final checkpoint. Accuracies are in percent, and gap is best minus worst residue in percentage points.
- Every compressed model's accuracy is periodic in the target key's position, and the period follows the stride: \(S\in\{4,6,8,12\}\) give periods of about 4, 6, 8, and 12, whether or not windows overlap (Figure 2).
- Under prefix padding, full-attention baselines stay within 6.1 points across residues, while gaps in compressed models reach 78 points across the whole sweep (all 23 runs are in Figure 2).
- Phase sensitivity persists without RoPE, and with uniform averaging in place of learned gates.
- Mean accuracy can hide phase sensitivity. W8/S8 · 8 KV nearly matches Full attention · 8 KV in mean accuracy and validation loss, yet its worst residue is 9.9% against 58.4%.
- The window boundary alone does not explain phase sensitivity. In W8/S8 · 1 KV, the boundary phase is \(S-1=7\), where the name after a key lands in the next entry. Yet accuracy varies by more than 50 points among phases 0 to 6, where a key and its name share an entry.
Phase specialization
Knockouts
To see which heads matter at which phase, we knock out one KV head at a time [8, 9] by mean replacement [10]. At the final position \(T_x\) of prompt \(x\), the outputs \(\mathbf z_{\ell,a,T_x}\) of the query heads \(a\in\mathcal Q(h)\) that read KV head \(h\) in layer \(\ell\) are replaced by their mean over \(N_{\mathrm{cal}}\) prompts from a disjoint calibration set, and nothing else changes:
A drop in accuracy at phase \(r\) means that the head's prompt-specific output contributes to retrieval at that phase. In full attention, knocking out an important layer lowers accuracy at all phases about equally (Figure 3, left). Under compression, some heads' losses concentrate on a few phases, and different heads cover different phases. With several KV heads per layer, knocking out some single heads still lowers accuracy at only one or two phases, so phase specialization occurs in single heads, not only in whole layers. In DeepSeek-V4-Flash-Base, we replace a whole layer's attention output at the query tokens at the end of the prompt, instead of one head at one position. Its knockout map shows phase-dependent losses repeating with period four.
Full attention · 1 KV
position mod 8 · prefix padding
accuracy change relative to clean
Static and input-dependent gates
We next examine the gate, which sets how much each offset contributes to an entry. A head can read back from compressed memory only what its gate kept. Because gate scores depend on the input, a gate can concentrate in two ways. It can favor the same offsets in every window, or offsets that change with content. Only the first ties a head to fixed phases. With \(H(\cdot)\) the Shannon entropy and \(\boldsymbol\alpha\in\mathbb R^W\) one window's gate scores, we measure the two kinds of concentration with effective numbers of offsets, from 1 (one-hot) to \(W\) (uniform), called static and per window:
layer
training paths
In Figure 4, we place every key gate in this plane. A gate that gives each offset the same score in every window lies on the dashed diagonal, where the two measures agree. The farther below the diagonal a gate falls, the more its scores change from window to window. The three traced gates start near the upper right at step 1, with little concentration of either kind. By the end of pretraining, the key gates of W8/S8 · 1 KV have split. Most sit at the right edge. Averaged over windows their scores are uniform, yet a single window can still concentrate on a few offsets, with per-window values anywhere from about 1.4 to 7.6. Such gates favor offsets that change with the content, so they are not tied to any phase. A few, including those in layers 9, 10, and 14, instead move down along the diagonal and become statically concentrated, favoring the same offsets in every window.
The same split appears across the sweep. With more KV heads per layer, the most concentrated heads reach further toward the lower left. With overlapping windows, including the DeepSeek-style kernel, many gates sit left of the right edge but below the diagonal: they favor some offsets across windows while still varying with the input.
Static preferences line up with the knockouts (Figure 5). In W8/S8 · 1 KV, the key gates of layers 9, 10, and 14 peak late, in the middle, and early in the window, respectively, and each value gate peaks about one offset later, so the head keeps a token together with its successor. Their knockout losses concentrate on those phases. With four KV heads per layer, many of the most concentrated heads narrow to a single offset and lose accuracy at a single phase.
Moving the weak spots
If static gate preferences help decide which phases are retrieved well, moving them should move the accuracy pattern. In models with \(W=S\), we cycle (circularly shift) the offset-specific gate parameters by \(\delta\) offsets and keep the payload maps fixed. With \(A_\delta(r)\) the accuracy at phase \(r\) after a cycle by \(\delta\), the prediction is
We test five models, W12/S12 · 1 KV, W12/S12 · 4 KV, and W8/S8 · 1 KV with seeds 42, 43, and 44, with \(\delta\in\{\pm1,\pm2,\pm3\}\). Accuracy at phases other than the boundary phase largely follows the prediction (Figure 6). Excluding the boundary phase, the \(R^2\) of observed against predicted accuracy changes lies between 0.66 and 1.00, closest for the two W12/S12 models and lowest for seeds 43 and 44 of W8/S8 · 1 KV. The boundary phase \(S-1\) stays weak under every cycle, so cycling relocates the weak within-window phases but does not remove the weakness at the boundary.
Accuracy by phase (%)
Predicted vs observed change (pp)
R² excluding the boundary phase
We conjecture that the observed phase sensitivity combines two asymmetries. The window boundary separates key–name pairs that share an entry from pairs that straddle two. Within-window asymmetries separate phases inside a window, and static gates can shape them when the parameterization allows it. The boundary cannot account for the spread among within-window phases. Learned gates do not account for all of it either: uniform averaging, which has no gate preferences, still varies across within-window phases, possibly because the payload maps \(\mathbf C_i\) differ by offset, so even equal weights need not preserve every offset equally well.
Why compression concentrates
We study an induction task in an idealized model. A context holds \(L\) chunks of \(S\) distinct tokens drawn from a vocabulary of \(N\) embeddings. In this section \(L\) counts chunks, \(\ell\) indexes a chunk rather than a layer, and \(r\in\{1,\dots,S\}\) is a position within a chunk, so position \(r\) is offset \(r-1\). After compression, a query repeats a token at a nonfinal position and asks for its successor [8, 11]. Each of \(m<S\) heads compresses chunk \(\ell\) with coefficient vectors \(p^K_{h,\ell}\) and \(p^V_{h,\ell}\), probability distributions over the \(S\) positions that are chosen before the query is seen:
The last condition makes embeddings nearly equiangular: two distinct tokens have inner product \(\kappa\), and a token has \(\kappa+1\) with itself, up to an error \(\varepsilon\). The shared component \(\kappa\) reflects the anisotropy of token representations [12, 13, 14]. Heads attend over chunks with a softmax at scale \(\gamma\), their outputs are summed and unembedded at scale \(\beta\), and the loss is the population cross-entropy of predicting the successor. Under exact geometry (\(\varepsilon=0\)), a head's evidence for the target is its attention on the right chunk times its value weight on the answer, which favors putting key weight on some position \(r\) and value weight on \(r+1\).
Theorem 1informal
Optimal compression is concentrated.
For sufficiently large \(N\) and small \(\varepsilon\), every global minimizer over all query-independent coefficient rules sets \(p^K_{h,\ell}=e_r\) and \(p^V_{h,\ell}=e_{r+1}\), where \(e_r\) is the one-hot vector at position \(r\) and \(r\) may depend on the context, chunk, and head. At \(\varepsilon=0\), the optimal assignments are exactly those with distinct positions across heads.
Theorem 2informal
Gradient flow makes the concentration static.
Take log-linear gates \(p^B_{h,\ell}(r)\propto\exp\langle\mathbf c^B_{h,r},\mathbf x_{\ell,r}\rangle\) for \(B\in\{K,V\}\), shared across inputs, with \(\varepsilon=0\) and a random initialization in a small ball of radius \(\rho\). For fixed scales \(\gamma\) and \(\beta\) and \(N\) sufficiently large relative to \(\log(e/\rho)\), with high probability gradient flow gives each head a single position \(r_h\) with \(p^K_{h,\ell}(r_h),\,p^V_{h,\ell}(r_h+1)\to1\) uniformly over contexts and chunks. Different heads may select the same position.
In the language of the experiments, both effective numbers of offsets tend to one, the lower-left corner of Figure 4. The theory also suggests why coverage can fail. Distinct positions lower the loss only through the output normalizer, by at most an amount of order \(1/N\), and our dynamics analysis has no term that pushes heads toward distinct positions. A position that no head selects receives no target evidence in any window, so in this model retrieval depends on phase unless the heads cover all \(S-1\) nonfinal positions. The model queries only pairs within a chunk, so it addresses phase sensitivity among within-window phases, not at the boundary phase.
Related work
Compressed memories. Compressive Transformer summarizes past activations [1], LoMA trains models to use compressed KV representations [2], Activation Beacon replaces raw activations with the KV entries of regularly spaced learned tokens [3], and CAT lets later chunks attend to compressed vectors of earlier chunks [4]. Sparse attention [15, 16, 17, 5] and cache eviction [18, 19] impose related block or token structure on memory. Sparse Transformer already defines a strided head whose visibility depends on position modulo a stride [15]. Hybrids that keep an exact branch over distant tokens, such as NSA [16], could mask a failure of the compressed branch alone. DeepSeek-V4 keeps exact tokens only in its local window [5].
Position-dependent failures. Retrieval quality depends on where relevant information appears in a long context [20], easy needle tests can overstate usable context length [21], and aggregate scores can hide method-specific failures of KV cache compression [22]. In gist-based compression, quality degrades near generation-segment boundaries [23]. Phase is a different coordinate: it recurs every compression stride and is distinct from depth in the context or the query's position within a segment.
Heads and training dynamics. Retrieval heads are heads whose ablation impairs long-context retrieval [9], and cache policies such as RazorAttention and DuoAttention keep full caches only for such heads [24, 25]. Under compression, a head's contribution also depends on phase. On the theory side, gradient descent can turn softmax attention into hard token selectors [26, 27], and multi-head training can allocate tasks across heads [28, 29]. The analysis here instead concerns compression chosen before the query is known.
Limitations
The mechanistic evidence is strongest in the controlled models. The DeepSeek knockouts act on whole layers rather than single heads. Gate cycling moves the within-window pattern but does not repair the boundary phase. The theory is an idealized model that explains why phase sensitivity arises, not how large it is in trained networks.
References
- Rae et al. Compressive Transformers for Long-Range Sequence Modelling. International Conference on Learning Representations, 2020.
- Wang and Xiao. LoMA: Lossless Compressed Memory Attention. arXiv preprint arXiv:2401.09486, 2024.
- Zhang et al. Long Context Compression with Activation Beacon. International Conference on Learning Representations, 2025.
- Prakash et al. Controllably Efficient Language Models. arXiv preprint arXiv:2511.05313, 2025.
- DeepSeek-AI. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv preprint arXiv:2606.19348, 2026.
- DeepSeek-AI. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression. arXiv preprint arXiv:2609.19969, 2026.
- Qwen Team. Qwen3 Technical Report. arXiv preprint arXiv:2505.09388, 2025.
- Olsson et al. In-Context Learning and Induction Heads. Transformer Circuits Thread, 2022.
- Wu et al. Retrieval Head Mechanistically Explains Long-Context Factuality. International Conference on Learning Representations, 2025.
- Wang et al. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. International Conference on Learning Representations, 2023.
- Bietti et al. Birth of a Transformer: A Memory Viewpoint. Advances in Neural Information Processing Systems, 2023.
- Gao et al. Representation Degeneration Problem in Training Natural Language Generation Models. International Conference on Learning Representations, 2019.
- Razzhigaev et al. The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models. Findings of the Association for Computational Linguistics: EACL 2024, 2024.
- Queipo-de-Llano et al. Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin. International Conference on Learning Representations, 2026.
- Child et al. Generating Long Sequences with Sparse Transformers. arXiv preprint arXiv:1904.10509, 2019.
- Yuan et al. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025.
- Lu et al. MoBA: Mixture of Block Attention for Long-Context LLMs. Advances in Neural Information Processing Systems, 2025.
- Zhang et al. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models. Advances in Neural Information Processing Systems, 2023.
- Li et al. SnapKV: LLM Knows What You Are Looking for Before Generation. Advances in Neural Information Processing Systems, 2024.
- Liu et al. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 2024.
- Hsieh et al. RULER: What's the Real Context Size of Your Long-Context Language Models? First Conference on Language Modeling, 2024.
- Chen et al. The Pitfalls of KV Cache Compression. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026.
- Deng et al. A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025.
- Tang et al. RazorAttention: Efficient KV Cache Compression Through Retrieval Heads. International Conference on Learning Representations, 2025.
- Xiao et al. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads. International Conference on Learning Representations, 2025.
- Tarzanagh et al. Transformers as Support Vector Machines. arXiv preprint arXiv:2308.16898, 2023.
- Tarzanagh et al. Max-Margin Token Selection in Attention Mechanism. Advances in Neural Information Processing Systems, 2023.
- Chen et al. Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality (extended abstract). Conference on Learning Theory, 2024.
- Yüksel et al. Incremental Learning of Sparse Attention Patterns in Transformers. International Conference on Machine Learning, 2026.