Xingyu Zhu

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

1ByteDance Seed2Princeton University3Stanford University4University of California, Berkeley

*Core contributors. Work done during internship at ByteDance Seed.†Corresponding authors.

Summary

A coordinate created by compression

Many long-context designs keep a compressed memory of past tokens [1, 2, 3, 4], including the compressed sparse attention of DeepSeek-V4 [5]. A chunked compressor reads the context in windows of \(W\) tokens that start every \(S\) tokens and writes one cache entry per window. Write \(\mathbf h_{j,i}\) for the hidden state at offset \(i\), its slot \(0,\dots,W-1\) within window \(j\). A gate turns logits into scores over the window, and the scores weight payloads \(\mathbf C_i\mathbf h_{j,i}\). Each offset \(i\) has its own gate parameters \(\mathbf Z_i\) and \(\mathbf b_i\) and its own linear payload map \(\mathbf C_i\):

\(\displaystyle \mathbf s_{j,i}=\mathbf Z_i\mathbf h_{j,i}+\mathbf b_i,\)\(\displaystyle \boldsymbol\alpha_{j,i}=\frac{\exp(\mathbf s_{j,i})}{\sum_{u=0}^{W-1}\exp(\mathbf s_{j,u})},\)\(\displaystyle \widehat{\mathbf K}_j=\sum_{i=0}^{W-1}\boldsymbol\alpha_{j,i}\odot\mathbf C_i\mathbf h_{j,i}.\)

The score \(\boldsymbol\alpha_{j,i}\) can be one scalar per offset or one score per channel, hence the elementwise product. Values are compressed the same way. Later queries attend over these entries together with a short local window of exact recent tokens. An entry is written before any later query exists, so it must retain whatever a later query might need. Because windows start every \(S\) tokens, each token at position \(t\) also has a phase:

\[\phi(t)=t \bmod S.\]

Three related terms recur below. The offset is a token's slot within one window, the phase is its position modulo the stride, and a residue is its position modulo a multiple of \(S\), which lets models with different strides share one axis. With non-overlapping windows (\(S=W\)) offset and phase coincide. With overlapping windows (\(S<W\)) a token appears at several offsets, but its phase is still fixed. With non-overlapping windows, a token at phase \(S-1\) is the last of its window, so the token after it falls into the next entry, and we call \(S-1\) the boundary phase. Shifting the input by a few tokens changes which tokens share an entry without changing their content or order, so if retrieval depends on \(\phi\), that dependence is a property of the model, not of the data.

prepend
Illustration with \(W=S=4\), so a token's offset \(i\) in its window is also its phase, and colors mark offsets. 1 Each token in the window gets a logit: a content term \(\mathbf Z_i\mathbf h_i\) that depends on the token, plus a bias \(\mathbf b_i\) that belongs to the offset, not the token. 2 A softmax over the window turns the four logits into scores that sum to one. 3 Each token's payload \(\mathbf C_i\mathbf h_i\) is scaled by its score and summed into a single cache entry. Each stacked bar shows how much of one channel comes from each offset. This gate's bias favors offset 2. Click a token to follow it (outlined). Prepending tokens moves every token to a new offset without changing any content, and the followed token's score changes with its new offset. The small bars on the prepend buttons show that score for each shift, and dashed chips are padding. With a uniform gate every offset gets \(1/4\). The numbers are illustrative, and only keys are drawn.

Weak spots in DeepSeek-V4

DeepSeek-V4 comes in a smaller model, Flash, and a larger one, Pro. The suffix -Base marks a pretrained checkpoint, and -0731 and -0813 mark post-trained ones.

A flipping completion

We ask DeepSeek-V4-Flash-Base to complete a line of its own FP8 quantization kernel, with the cursor after T.Cast(FP. The correct continuation is 8, since 32 would duplicate an outer cast. We prepend a docstring and lengthen it one token at a time. The code does not change, but the top prediction switches between 8 and 32 every two tokens, so the pattern repeats every four tokens (Figure 1). Four is the stride of DeepSeek-V4's compressed sparse attention (\(W=8, S=4\)) [5]. Across four families of filler text with 16 lengths each, 60 of 64 rankings follow this pattern. Under several filler sweeps, the post-trained DeepSeek-V4-Flash-0731 (\(S=4\)) and DeepSeek-V4.1-Flash (\(S=2\)) [6] also switch periodically between 8 and 32. DeepSeek-V3.1-Base, which has no chunked compression, prefers 32 on only 4 of the same 64 inputs, each by a probability margin below 0.07.


    
next-token probability
8
32
\(P(8)\) at all 16 lengths \(L\) (click to jump)
8 on top32 on topheight = \(P(8)\)
Figure 1. DeepSeek-V4-Flash-Base completes T.Cast(FP in its own FP8 kernel. The filler length \(L\) counts every docstring token before the code. The docstring's opening sentence is 15 tokens, and each added = sign adds one token, so 15 signs give \(L=30\). The code below never changes. It only shifts by one position per added sign. The model prefers the correct 8 when \(L \bmod 4\in\{2,3\}\) and 32 otherwise, a period of four that matches its compression stride. Probabilities are measured values, and dashed lines in the strip separate cycles of four.

A position-matched needle test

Each prompt holds 128K tokens with 16,000 four-token key–value records, and asks for the value of one target key near the middle of the context. Let \(t_K\) be the position of the target key. We group prompts by the residue \(r=t_K \bmod 8\), which covers two stride cycles for DeepSeek-V4 and four for DeepSeek-V4.1, and write \(A(r)\) for the answer accuracy on group \(r\). Filler lengths are chosen so that the query position is fixed and

\[\mathbb E[t_K\mid r]\ \text{and}\ \operatorname{Var}[t_K\mid r]\ \text{are the same for every } r,\]

so the groups differ in phase but not in these position statistics. Each group has 256 prompts for each of the four DeepSeek-V4 models (Flash-Base, Flash-0731, Pro-Base, and Pro-0813). DeepSeek-V4.1-Flash was evaluated in a separate, larger run with 2,560 prompts per group.

x axis

Model\(S\)meanworstbestgap
Needle accuracy by residue group. Each group holds 256 prompts at 128K tokens (2,560 for DeepSeek-V4.1-Flash), and bands are 95% Wilson intervals. Select a model in the table to plot it against the others in gray. Dashed vertical lines mark the boundaries between its stride periods, and model lines are dotted where they cross one. “Fold” overlays the periods, the first solid and later ones dashed. Accuracy is in percent, and gap is best minus worst group, in percentage points.

The accuracy pattern repeats with the stride. For the four DeepSeek-V4 models, the mean of \(|A(r)-A(r+4)|\) over \(r\in\{0,1,2,3\}\) is 1.1 to 2.0 points, small next to gaps of 15 to 40 points. DeepSeek-V4.1-Flash compresses with \(S=2\), and all four of its even residues stay above all four odd ones. Post-training raises accuracy and narrows the gap, but it keeps the period.

Controlled pretraining

DeepSeek-V4 differs from a plain transformer in many ways besides compression, so we pretrain our own models. The backbone is Qwen3-0.6B [7] (28 layers, width 1,024, 16 query heads, head dimension 128). Every model is trained from scratch on 100B tokens, with chunked compression as the only substantial architectural change from the full-attention baselines. We name each model by its signature attributes. W8/S8 · 1 KV has window \(W=8\), stride \(S=8\), and one KV head per layer, and anything else that differs is appended, as in W8/S8 · 1 KV · no RoPE. Unless its name says otherwise, a compressed model has learned scalar gates, separate key and value branches, and RoPE, and it attends jointly to compressed history and a 16-token local window of exact tokens. W8/S8 · 1 KV also has three extra seeds (43, 44, and 45), and without a seed the name means seed 42. Across the sweep, we vary \(W\) and \(S\) independently, the number of KV heads, and these attributes:

The evaluation is an animal-to-name lookup. Each prompt lists 64 records, each a single-token animal key followed directly by a single-token name, and ends with eight demonstrations and a target animal. The model must output that animal's name, the token right after its key. We report full-vocabulary top-1 accuracy, grouped by the target key's position modulo 24, a common multiple of all tested strides.

Prefix padding shifts all records together. In-sequence padding varies the gaps between records, so that distractor phases do not move in lockstep with the target. Both hold the query position and the first two moments of the target position fixed.

padding

Full attention

1, 4, and 8 KV heads · prefix padding

​

​

Figure 2. Accuracy by target-key position mod 24 at the final checkpoint. Left: full-attention baselines with 1, 4, and 8 KV heads stay flat. Right: any of the 23 compressed runs, under either padding protocol. Dashed vertical lines mark the period boundaries every \(S\) positions, and small triangles under the axis mark the boundary phase \(S-1\). Each compressed line is solid within a period and dotted where it crosses a period boundary. The dashed gray line is full attention with the same number of KV heads, where such a run exists. Bands are 95% intervals: Wilson intervals for full attention, and for compressed runs a bootstrap over key–name mappings, each evaluated at every position. Full-attention runs exist only under prefix padding.
Modelval. lossmeanworstgap
Full attention · 1 KV2.46618.817.23.1
Full attention · 8 KV2.43761.158.44.3
W8/S8 · 1 KV · seed 422.47850.65.775.3
W8/S8 · 8 KV2.43159.49.971.1
W8/S4 · 1 KV2.46766.946.239.8
W12/S12 · 1 KV2.48733.05.063.2
W8/S8 · 1 KV · no RoPE2.50429.45.040.9
W8/S8 · 1 KV · uniform averaging2.49712.54.719.2

Selected runs, prefix padding, final checkpoint. Accuracies are in percent, and gap is best minus worst residue in percentage points.

Phase specialization

Knockouts

To see which heads matter at which phase, we knock out one KV head at a time [8, 9] by mean replacement [10]. At the final position \(T_x\) of prompt \(x\), the outputs \(\mathbf z_{\ell,a,T_x}\) of the query heads \(a\in\mathcal Q(h)\) that read KV head \(h\) in layer \(\ell\) are replaced by their mean over \(N_{\mathrm{cal}}\) prompts from a disjoint calibration set, and nothing else changes:

\(\displaystyle \tilde{\mathbf z}_{\ell,a,T_x}(x)=\boldsymbol\mu_{\ell,a}\)\(\displaystyle {}=\frac1{N_{\mathrm{cal}}}\sum_{n=1}^{N_{\mathrm{cal}}}\mathbf z_{\ell,a,T_{x^{(n)}}}\big(x^{(n)}\big),\)\(\displaystyle a\in\mathcal Q(h).\)
knock out
Illustration of the knockout in W8/S8 · 1 KV. Each row of the grid is a layer and each column a token position. The rightmost column is the final position, where the answer is predicted. Knocking out a layer's KV head replaces the outputs of the query heads that read it at the final position, \(\mathbf z\), with their calibration mean \(\boldsymbol\mu\) (red cell). Every other computation runs as usual, so the change reaches the answer only through the later layers at that position (tinted). The bars are the measured changes in accuracy relative to the clean run, by phase (prefix padding, the same values as this model's map in Figure 3): layer 0 matters at every phase, while layers 9, 10, and 14, examined below, each matter at a few adjacent phases. Select any layer: click it, use the arrow keys, or use the buttons.

A drop in accuracy at phase \(r\) means that the head's prompt-specific output contributes to retrieval at that phase. In full attention, knocking out an important layer lowers accuracy at all phases about equally (Figure 3, left). Under compression, some heads' losses concentrate on a few phases, and different heads cover different phases. With several KV heads per layer, knocking out some single heads still lowers accuracy at only one or two phases, so phase specialization occurs in single heads, not only in whole layers. In DeepSeek-V4-Flash-Base, we replace a whole layer's attention output at the query tokens at the end of the prompt, instead of one head at one position. Its knockout map shows phase-dependent losses repeating with period four.

position padding

Full attention · 1 KV

position mod 8 · prefix padding

accuracy change relative to clean

 

 

 

Figure 3. Knockout maps. Each row knocks out one KV head (one row per layer when a layer has a single KV head), each column is a target-key position, and color is the accuracy change relative to clean accuracy: red is accuracy lost, blue gained, clipped at ±100%. Left: full attention with one KV head, where knocking out an important layer hurts every position alike. Right: any compressed run under prefix or in-sequence padding, folded to one stride (position mod \(S\)) or shown across all 24 positions, or DeepSeek-V4-Flash-Base, where a whole layer's attention output is replaced (positions mod 8). Click a row, or focus the map and use ↑/↓, to see its changes by position below.

Static and input-dependent gates

We next examine the gate, which sets how much each offset contributes to an entry. A head can read back from compressed memory only what its gate kept. Because gate scores depend on the input, a gate can concentrate in two ways. It can favor the same offsets in every window, or offsets that change with content. Only the first ties a head to fixed phases. With \(H(\cdot)\) the Shannon entropy and \(\boldsymbol\alpha\in\mathbb R^W\) one window's gate scores, we measure the two kinds of concentration with effective numbers of offsets, from 1 (one-hot) to \(W\) (uniform), called static and per window:

\(\displaystyle \underbrace{\exp\!\big(H(\mathbb E[\boldsymbol\alpha])\big)}_{\text{static}}\)\(\displaystyle \text{and}\qquad\underbrace{\mathbb E\big[\exp(H(\boldsymbol\alpha))\big]}_{\text{per window}}.\)

layer

training paths

Figure 4. Two ways a gate can concentrate. Each point is one KV head's key gate at the end of pretraining, placed by its static effective number of offsets \(\exp(H(\mathbb E[\boldsymbol\alpha]))\) (horizontal) and its per-window effective number \(\mathbb E[\exp(H(\boldsymbol\alpha))]\) (vertical). Both run from 1 (a single offset) to \(W\) (uniform). On the dashed diagonal, a gate gives each offset the same score in every window. At the right edge, its average over windows is uniform, however concentrated single windows are. For W8/S8 · 1 KV · seed 42, lines trace layers 9, 10, and 14 from step 1 (open circle) to step 47,518 (filled), and replay animates them. Color marks layer depth. Uniform averaging has no learned gate, and vector gates score every channel separately, so those runs are omitted.

In Figure 4, we place every key gate in this plane. A gate that gives each offset the same score in every window lies on the dashed diagonal, where the two measures agree. The farther below the diagonal a gate falls, the more its scores change from window to window. The three traced gates start near the upper right at step 1, with little concentration of either kind. By the end of pretraining, the key gates of W8/S8 · 1 KV have split. Most sit at the right edge. Averaged over windows their scores are uniform, yet a single window can still concentrate on a few offsets, with per-window values anywhere from about 1.4 to 7.6. Such gates favor offsets that change with the content, so they are not tied to any phase. A few, including those in layers 9, 10, and 14, instead move down along the diagonal and become statically concentrated, favoring the same offsets in every window.

The same split appears across the sweep. With more KV heads per layer, the most concentrated heads reach further toward the lower left. With overlapping windows, including the DeepSeek-style kernel, many gates sit left of the right edge but below the diagonal: they favor some offsets across windows while still varying with the input.

Static preferences line up with the knockouts (Figure 5). In W8/S8 · 1 KV, the key gates of layers 9, 10, and 14 peak late, in the middle, and early in the window, respectively, and each value gate peaks about one offset later, so the head keeps a token together with its successor. Their knockout losses concentrate on those phases. With four KV heads per layer, many of the most concentrated heads narrow to a single offset and lose accuracy at a single phase.

map
1 = uniform8×
−100%+100%

in the text

Figure 5. Gate preferences line up with knockout losses. Each row is one KV head at the final checkpoint. Left: the static concentration of its key gate, the effective number of offsets \(\exp(H(\mathbb E[\boldsymbol\alpha]))\), from 1 (a single offset) to \(W\) (uniform). Middle: the mean gate score by phase, relative to a uniform share, for the key or the value gate. Right: the head's knockout effect by phase (prefix padding, relative accuracy change, red lost, blue gained). Select a head to compare its key and value scores with its losses. The value scores are shifted back by one offset, so a head that keeps a key together with its successor peaks at the same phase in both. For W8/S4 · 1 KV and the DeepSeek-style kernel, windows overlap and each phase pools two offsets, so the left column counts effective phases, from 1 to \(S\).

Moving the weak spots

If static gate preferences help decide which phases are retrieved well, moving them should move the accuracy pattern. In models with \(W=S\), we cycle (circularly shift) the offset-specific gate parameters by \(\delta\) offsets and keep the payload maps fixed. With \(A_\delta(r)\) the accuracy at phase \(r\) after a cycle by \(\delta\), the prediction is

\(\displaystyle \tilde{\mathbf Z}_i=\mathbf Z_{(i-\delta)\bmod W},\quad \tilde{\mathbf b}_i=\mathbf b_{(i-\delta)\bmod W}\)\(\displaystyle \Longrightarrow\qquad A_\delta(r)\approx A_0\big((r-\delta)\bmod S\big).\)

We test five models, W12/S12 · 1 KV, W12/S12 · 4 KV, and W8/S8 · 1 KV with seeds 42, 43, and 44, with \(\delta\in\{\pm1,\pm2,\pm3\}\). Accuracy at phases other than the boundary phase largely follows the prediction (Figure 6). Excluding the boundary phase, the \(R^2\) of observed against predicted accuracy changes lies between 0.66 and 1.00, closest for the two W12/S12 models and lowest for seeds 43 and 44 of W8/S8 · 1 KV. The boundary phase \(S-1\) stays weak under every cycle, so cycling relocates the weak within-window phases but does not remove the weakness at the boundary.

Accuracy by phase (%)

 

Predicted vs observed change (pp)

 

R² excluding the boundary phase

Figure 6. Gate cycling moves the weak spots. Cycling the offset-specific gate parameters by \(\delta\) predicts that the accuracy at phase \(r\) becomes the uncycled accuracy at phase \((r-\delta)\bmod S\) (dashed). Left: observed accuracy after the cycle (solid, with its 95% bootstrap band) and the uncycled model (gray). Right: predicted against observed change for each phase. \(R^2\) is their squared correlation, excluding the boundary phase \(S-1\) (hollow), with a 95% bootstrap interval. The grid gives that \(R^2\) for every model and cycle. Select a cell, or drag the slider to set \(\delta\) (0 shows the uncycled model).

We conjecture that the observed phase sensitivity combines two asymmetries. The window boundary separates key–name pairs that share an entry from pairs that straddle two. Within-window asymmetries separate phases inside a window, and static gates can shape them when the parameterization allows it. The boundary cannot account for the spread among within-window phases. Learned gates do not account for all of it either: uniform averaging, which has no gate preferences, still varies across within-window phases, possibly because the payload maps \(\mathbf C_i\) differ by offset, so even equal weights need not preserve every offset equally well.

Why compression concentrates

We study an induction task in an idealized model. A context holds \(L\) chunks of \(S\) distinct tokens drawn from a vocabulary of \(N\) embeddings. In this section \(L\) counts chunks, \(\ell\) indexes a chunk rather than a layer, and \(r\in\{1,\dots,S\}\) is a position within a chunk, so position \(r\) is offset \(r-1\). After compression, a query repeats a token at a nonfinal position and asks for its successor [8, 11]. Each of \(m<S\) heads compresses chunk \(\ell\) with coefficient vectors \(p^K_{h,\ell}\) and \(p^V_{h,\ell}\), probability distributions over the \(S\) positions that are chosen before the query is seen:

\(\displaystyle \mathbf k_{h,\ell}=\sum_{r=1}^{S}p^{K}_{h,\ell}(r)\,\mathbf x_{\ell,r},\)\(\displaystyle \mathbf v_{h,\ell}=\sum_{r=1}^{S}p^{V}_{h,\ell}(r)\,\mathbf x_{\ell,r},\)\(\displaystyle \big|\mathbf x^\top\mathbf x'-\kappa-\mathbf 1\{\mathbf x=\mathbf x'\}\big|\le\varepsilon .\)

The last condition makes embeddings nearly equiangular: two distinct tokens have inner product \(\kappa\), and a token has \(\kappa+1\) with itself, up to an error \(\varepsilon\). The shared component \(\kappa\) reflects the anisotropy of token representations [12, 13, 14]. Heads attend over chunks with a softmax at scale \(\gamma\), their outputs are summed and unembedded at scale \(\beta\), and the loss is the population cross-entropy of predicting the successor. Under exact geometry (\(\varepsilon=0\)), a head's evidence for the target is its attention on the right chunk times its value weight on the answer, which favors putting key weight on some position \(r\) and value weight on \(r+1\).

Theorem 1informal

Optimal compression is concentrated.

For sufficiently large \(N\) and small \(\varepsilon\), every global minimizer over all query-independent coefficient rules sets \(p^K_{h,\ell}=e_r\) and \(p^V_{h,\ell}=e_{r+1}\), where \(e_r\) is the one-hot vector at position \(r\) and \(r\) may depend on the context, chunk, and head. At \(\varepsilon=0\), the optimal assignments are exactly those with distinct positions across heads.

Theorem 2informal

Gradient flow makes the concentration static.

Take log-linear gates \(p^B_{h,\ell}(r)\propto\exp\langle\mathbf c^B_{h,r},\mathbf x_{\ell,r}\rangle\) for \(B\in\{K,V\}\), shared across inputs, with \(\varepsilon=0\) and a random initialization in a small ball of radius \(\rho\). For fixed scales \(\gamma\) and \(\beta\) and \(N\) sufficiently large relative to \(\log(e/\rho)\), with high probability gradient flow gives each head a single position \(r_h\) with \(p^K_{h,\ell}(r_h),\,p^V_{h,\ell}(r_h+1)\to1\) uniformly over contexts and chunks. Different heads may select the same position.

In the language of the experiments, both effective numbers of offsets tend to one, the lower-left corner of Figure 4. The theory also suggests why coverage can fail. Distinct positions lower the loss only through the output normalizer, by at most an amount of order \(1/N\), and our dynamics analysis has no term that pushes heads toward distinct positions. A position that no head selects receives no target evidence in any window, so in this model retrieval depends on phase unless the heads cover all \(S-1\) nonfinal positions. The model queries only pairs within a chunk, so it addresses phase sensitivity among within-window phases, not at the boundary phase.

Compressed memories. Compressive Transformer summarizes past activations [1], LoMA trains models to use compressed KV representations [2], Activation Beacon replaces raw activations with the KV entries of regularly spaced learned tokens [3], and CAT lets later chunks attend to compressed vectors of earlier chunks [4]. Sparse attention [15, 16, 17, 5] and cache eviction [18, 19] impose related block or token structure on memory. Sparse Transformer already defines a strided head whose visibility depends on position modulo a stride [15]. Hybrids that keep an exact branch over distant tokens, such as NSA [16], could mask a failure of the compressed branch alone. DeepSeek-V4 keeps exact tokens only in its local window [5].

Position-dependent failures. Retrieval quality depends on where relevant information appears in a long context [20], easy needle tests can overstate usable context length [21], and aggregate scores can hide method-specific failures of KV cache compression [22]. In gist-based compression, quality degrades near generation-segment boundaries [23]. Phase is a different coordinate: it recurs every compression stride and is distinct from depth in the context or the query's position within a segment.

Heads and training dynamics. Retrieval heads are heads whose ablation impairs long-context retrieval [9], and cache policies such as RazorAttention and DuoAttention keep full caches only for such heads [24, 25]. Under compression, a head's contribution also depends on phase. On the theory side, gradient descent can turn softmax attention into hard token selectors [26, 27], and multi-head training can allocate tasks across heads [28, 29]. The analysis here instead concerns compression chosen before the query is known.

Limitations

The mechanistic evidence is strongest in the controlled models. The DeepSeek knockouts act on whole layers rather than single heads. Gate cycling moves the within-window pattern but does not repair the boundary phase. The theory is an idealized model that explains why phase sensitivity arises, not how large it is in trained networks.

References

  1. Rae et al. Compressive Transformers for Long-Range Sequence Modelling. International Conference on Learning Representations, 2020.
  2. Wang and Xiao. LoMA: Lossless Compressed Memory Attention. arXiv preprint arXiv:2401.09486, 2024.
  3. Zhang et al. Long Context Compression with Activation Beacon. International Conference on Learning Representations, 2025.
  4. Prakash et al. Controllably Efficient Language Models. arXiv preprint arXiv:2511.05313, 2025.
  5. DeepSeek-AI. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv preprint arXiv:2606.19348, 2026.
  6. DeepSeek-AI. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression. arXiv preprint arXiv:2609.19969, 2026.
  7. Qwen Team. Qwen3 Technical Report. arXiv preprint arXiv:2505.09388, 2025.
  8. Olsson et al. In-Context Learning and Induction Heads. Transformer Circuits Thread, 2022.
  9. Wu et al. Retrieval Head Mechanistically Explains Long-Context Factuality. International Conference on Learning Representations, 2025.
  10. Wang et al. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. International Conference on Learning Representations, 2023.
  11. Bietti et al. Birth of a Transformer: A Memory Viewpoint. Advances in Neural Information Processing Systems, 2023.
  12. Gao et al. Representation Degeneration Problem in Training Natural Language Generation Models. International Conference on Learning Representations, 2019.
  13. Razzhigaev et al. The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models. Findings of the Association for Computational Linguistics: EACL 2024, 2024.
  14. Queipo-de-Llano et al. Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin. International Conference on Learning Representations, 2026.
  15. Child et al. Generating Long Sequences with Sparse Transformers. arXiv preprint arXiv:1904.10509, 2019.
  16. Yuan et al. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025.
  17. Lu et al. MoBA: Mixture of Block Attention for Long-Context LLMs. Advances in Neural Information Processing Systems, 2025.
  18. Zhang et al. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models. Advances in Neural Information Processing Systems, 2023.
  19. Li et al. SnapKV: LLM Knows What You Are Looking for Before Generation. Advances in Neural Information Processing Systems, 2024.
  20. Liu et al. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 2024.
  21. Hsieh et al. RULER: What's the Real Context Size of Your Long-Context Language Models? First Conference on Language Modeling, 2024.
  22. Chen et al. The Pitfalls of KV Cache Compression. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026.
  23. Deng et al. A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025.
  24. Tang et al. RazorAttention: Efficient KV Cache Compression Through Retrieval Heads. International Conference on Learning Representations, 2025.
  25. Xiao et al. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads. International Conference on Learning Representations, 2025.
  26. Tarzanagh et al. Transformers as Support Vector Machines. arXiv preprint arXiv:2308.16898, 2023.
  27. Tarzanagh et al. Max-Margin Token Selection in Attention Mechanism. Advances in Neural Information Processing Systems, 2023.
  28. Chen et al. Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality (extended abstract). Conference on Learning Theory, 2024.
  29. Yüksel et al. Incremental Learning of Sparse Attention Patterns in Transformers. International Conference on Machine Learning, 2026.