T²MLR: Transformer with Temporal Middle-Layer Recurrence
1Princeton Language and Intelligence, Department of Computer Science, Princeton University
*Joint first authors, equal contribution.†Core contribution.LIT Workshop @ ICLR 2026. This post follows the extended arXiv version.
Summary
- At each decoding step, a Transformer sees the past only through tokens and through key–value caches at the same depth. The shallow layers of token \(t\) cannot read the deep representation that was computed for token \(t-1\). T²MLR adds one recurrent vector \(\mathbf R_t\). It is taken after layer \(\ell_{\text{end}}\) at token \(t-1\) and fused, through a gated module, into the input of layer \(\ell_{\text{start}}\) at token \(t\).
- Inference still runs each of the \(L\) layers once per token, plus one small fusion. In our measurements, generation takes 1.041 to 1.082 times the wall-clock time of a parameter-matched Transformer. For training, a fixed number of Jacobi iterations approximates the recurrence in parallel, and in a 1B-token comparison the result matches training with the exact recurrence.
- On \(S_5\)-Retrieval, which needs both group state tracking and in-context lookup, a 4-layer T²MLR succeeds where parameter-matched Transformers and LSTMs fail.
- Recurrence over a middle block usually works better than recurrence over the whole stack. At 135M parameters, recurring over layers 13–18 (20% of 30 layers) gives the best zero-shot average, 44.14 against 42.83 for the baseline, and a block of the same width placed early or late helps much less or not at all. After finetuning, Variable Assignment accuracy is 0.494 for the baseline, 0.773 with full recurrence, and 0.945 with layers 13–18.
- The gains persist at about 370M and 1B parameters and with 50B training tokens. Retrofitting a pretrained SmolLM2-1.7B-Instruct raises GSM8K from 35.78 to 39.88 and MATH500 from 12.80 to 18.00. The cost is training time: 2 to 4 times the wall-clock of a Transformer.
The depth–time barrier
An autoregressive Transformer projects its last hidden state onto the vocabulary, emits one token, and starts the next forward pass from that token's embedding. Whatever the network computed for the previous position reaches the next one only through this unembed–decode–embed bottleneck, or through attention. Attention, in turn, connects positions only at equal depth: layer \(\ell\) of token \(t\) reads the keys and values that layer \(\ell\) wrote for earlier tokens.
Mechanistic analyses place much of the abstract computation of a Transformer in its middle layers [1, 2, 3, 4, 5]. The middle-layer state of token \(t-1\) is already in memory when token \(t\) is processed, but the early layers of token \(t\) have no path to it. To use it, the model has to rebuild it by traversing the same depth again and reading it indirectly through attention, or recover it from the emitted token.
Two lines of work relax this. Latent-reasoning methods such as Coconut and CODI feed a last-layer hidden state back as the next input, and soft-token methods feed mixtures of token embeddings [6, 7, 8, 9, 10]. In both cases the recurrent signal typically enters at the embedding, outside the middle layers. Looped Transformers run a block several times per token [11, 12, 13], which adds depth but multiplies the cost of every decoding step by the number of loops.
T²MLR adds a pathway from deep to shallow layers across time. Decoding stays standard autoregressive decoding, training keeps dense teacher-forced supervision on every token, and the per-token inference cost stays that of a Transformer.
Middle-layer recurrence
Write the \(\ell\)-th block of an \(L\)-layer decoder as a map that takes the current token's representation and the layer's past key–value cache, and returns the new representation and the extended cache. Here \(\mathbf h_t^{(0)}\) is the input embedding of token \(x_t\):
T²MLR picks two layers \(1\le\ell_{\text{start}}\le\ell_{\text{end}}\le L\) and keeps a recurrent cache \(\mathbf R_t\in\mathbb R^d\) of constant size. Only layer \(\ell_{\text{start}}\) changes. Its input is the current representation fused with the cache from the previous step:
The fusion \(\Phi\) is gated. With \(\mathbf h=\mathbf h_t^{(\ell_{\text{start}}-1)}\) and \(\mathbf R=\mathbf R_{t-1}\),
where \(f_{\text{cur}},f_{\text{rec}}:\mathbb R^{2d}\to\mathbb R^d\) are learned linear layers applied to the concatenation, \(\mathbf W_{\text{rec}}\) is a learned \(d\times d\) projection, and \(\sigma\) is the elementwise sigmoid. The scalars \(\tanh(\gamma_{\text{cur}})\) and \(\tanh(\gamma_{\text{rec}})\) are input-independent gates, and the sigmoids are input-dependent, per-channel gates. After layer \(\ell_{\text{end}}\), the cache is updated with a temporal residual:
Both \(\gamma\) start at zero, so at initialization \(\Phi(\mathbf h,\mathbf R)=\mathbf h\) and the model is exactly a standard Transformer. It departs from one only as far as training moves the gates. Across all tasks we studied, \(\gamma_{\text{rec}}\) converges to positive values and \(\gamma_{\text{cur}}\) to negative ones. We read this as the model treating the recurrent term as an additive refinement of the residual stream rather than a replacement for it.
Special casesAppendix E.1
T²MLR contains both a Transformer and a recurrent network.
With \(\gamma_{\text{cur}}=\gamma_{\text{rec}}=0\), the fusion is the identity and T²MLR is a standard Transformer. If the attention pattern reduces to the identity, so that each position attends only to itself, the only link between tokens is \(\mathbf R_t\) and the model is a fully recurrent network.
We write \(D=\ell_{\text{end}}-\ell_{\text{start}}+1\) for the number of recurrent layers and name a model T²MLR\((\ell_{\text{start}},\ell_{\text{end}})\). For the 30-layer backbone used below, T²MLR(1,30) recurs over the whole stack and T²MLR(13,18) over six middle layers.
The cost per token does not change: every token runs each of the \(L\) layers once, plus one fusion. What changes is the longest chain of computation behind an output. The recurrent block of token \(t-1\) feeds the recurrent block of token \(t\), so the output at token \(t\) sits at the end of a chain of \(L+(t-1)D\) layer applications, while in a Transformer that chain has length \(L\) at every position. A looped Transformer lengthens the chain too, but only by running more layers at every token. Figure 1 draws these data paths for the configurations studied in the paper.
data flow: one column per token, one cell per layer
longest chain vs token position
Training in parallel
A Transformer trains on all positions of a sequence at once, since with teacher forcing every position's input is known in advance. The recurrence breaks this: \(\mathbf R_{t-1}\) must be known before token \(t\) passes layer \(\ell_{\text{start}}\), and it depends on the forward pass of token \(t-1\). Running the positions one after another gives up the sequence parallelism that makes pretraining on 2048-token sequences scalable.
We instead approximate the caches of all positions at once with Jacobi fixed-point iterations, following the parallel continuous chain of thought of Wu et al. [14]. Let \(\mathbf H^{(\ell)}\) stack the layer-\(\ell\) states of all positions, and let \(\mathcal S\) (ShiftRight) move every row one position later and put a zero vector at position 1. Layers \(1,\dots,\ell_{\text{start}}-1\) run once. The middle block \(\mathcal F_{\ell_{\text{start}}:\ell_{\text{end}}}\) then runs repeatedly:
- Run the middle block once without any cache and set \(\mathbf R^{\langle 1\rangle}=\mathcal S\big(\mathbf H^{(\ell_{\text{end}})}_{\langle 0\rangle}\big)\).
- For \(k=2,\dots,d_{\text{forward}}\): fuse the current cache into the block's input and rerun only the middle block,
\(\displaystyle \mathbf H^{(\ell_{\text{end}})}_{\langle k-1\rangle}\)\(\displaystyle {}=\mathcal F_{\ell_{\text{start}}:\ell_{\text{end}}}\big(\Phi(\mathbf H^{(\ell_{\text{start}}-1)},\mathbf R^{\langle k-1\rangle})\big),\)\(\displaystyle \mathbf R^{\langle k\rangle}=\operatorname{RMSNorm}\big(\mathcal S(\mathbf H^{(\ell_{\text{end}})}_{\langle k-1\rangle}+\mathbf R^{\langle k-1\rangle})\big).\)
- Fuse \(\mathbf R^{\langle d_{\text{forward}}\rangle}\) once more and run layers \(\ell_{\text{start}},\dots,L\) to the output.
Every step is differentiable. Gradients pass through at most \(d_{\text{backward}}\) of the iterations, analogous to truncated backpropagation through time [15]. All experiments use \(d_{\text{forward}}=16\) and \(d_{\text{backward}}=4\) unless stated otherwise.
ObservationAppendix B.3
Exactness spreads one position per iteration.
Because of the shift, the cache at position \(t\) depends only on positions before \(t\). After \(k\) iterations, the caches read by the first \(k\) positions equal those of the exact sequential recurrence, whatever the weights.
With 16 iterations and 2048 tokens, this guarantee covers 16 positions. The rest of the sequence relies on the approximation being good long before it is exact. We do not measure the cache error of trained models directly; we check the approximation through gradients and end-to-end training below. Figure 2 runs the scheme on a small random middle block, with the paper's fusion and cache update, against the exact recurrence.
Relative error of the parallel cache
per position \(t\), after \(k\) iterations
Worst position vs iterations
\(\max_t\) of the relative error, log scale
The toy shows the mechanism. Whether 16 iterations suffice for a trained language model is an empirical question. We took an intermediate checkpoint of T²MLR(5,26) from the pretraining runs below and compared its gradients on one random batch of 2048-token sequences under different \((d_{\text{forward}},d_{\text{backward}})\) with the gradients under \((32,32)\) (Figure 3). With \(d_{\text{forward}}\ge 8\) and \(d_{\text{backward}}\ge 4\) the gradients stay close to that anchor: cosine similarity is at least 0.97 for all parameters, the fusion module, and layer 15, and 1.00 over all parameters at the training setting (16, 4). With four iterations or fewer they are far off, for example 0.48 at (4, 4) over all parameters and 0.26 for the fusion module.
training setting (16, 4)
anchor (32, 32)
The end-to-end check is to train the same models with the exact recurrence. Exact training on 2048-token sequences is expensive, so this comparison uses the first 1B tokens of FineWeb-Edu, \(\ell_{\text{start}}=8\), and the same truncation \(d_{\text{backward}}=4\) for both methods. The Jacobi-trained models reach the same validation loss as the exactly trained ones, marginally lower at every size, and evaluating them with Jacobi iterations instead of the exact rollout changes the loss by at most 0.0001.
| Size | Transformer | exact | Jacobi | Jacobi, Jacobi eval. |
|---|---|---|---|---|
| 135M | 3.9514 | 3.9083 | 3.9073 | 3.9073 |
| 360M | 3.3803 | 3.3246 | 3.3204 | 3.3205 |
| 1B | 3.1562 | 3.1148 | 3.1103 | 3.1102 |
Validation loss after the first 1B FineWeb-Edu tokens. The T²MLR columns name the training method. T²MLR models are evaluated with the exact sequential rollout, except in the last column. Values from Table 6 of the paper.
The price is paid in training. Per sequence, a Transformer costs \(O\big(L(N^2d+Nd^2)\big)\) for length \(N\) and width \(d\), while T²MLR costs \(O\big((L+d_{\text{forward}}D)(N^2d+Nd^2)\big)\), since the middle block runs \(d_{\text{forward}}\) times. Generation needs no iterations: each new token reads the exact \(\mathbf R_{t-1}\) left by the previous one. The recurrent state adds less than 0.1% to peak GPU memory, and the measured slowdown is at most about 8%:
| Tokens | 135M (1,30) | 135M (5,26) | 135M (9,22) | 361M (9,24) | 1B (9,24) |
|---|---|---|---|---|---|
| 512 | 1.070 | 1.076 | 1.082 | 1.065 | 1.067 |
| 1024 | 1.064 | 1.068 | 1.078 | 1.057 | 1.057 |
| 2048 | 1.067 | 1.076 | 1.078 | 1.056 | 1.041 |
Wall-clock time to generate the given number of tokens with T²MLR, relative to a parameter-matched Transformer. Values from Table 7 of the paper.
The prompt can also be prefilled with Jacobi iterations instead of token by token. For T²MLR(5,26) with 16 iterations, prefill is 11.1 times faster than exact prefill, and the zero-shot average is 0.435 against 0.437.
State tracking with retrieval
Before language modeling, we test the inductive bias on a synthetic task that separates Transformers from recurrent networks in both directions. In \(S_5\) state tracking [16], the input is a sequence of permutations \(a_1,\dots,a_N\) of five elements and the output is the sequence of running products \(a_1,\ a_1\circ a_2,\ \dots,\ a_1\circ\cdots\circ a_N\), the states. A Transformer needs \(\Omega(\log N)\) layers to solve this exactly [17], and the shallowest solution known to be learned, a parallel associative scan, uses \(\lceil\log_2 N\rceil\) layers [18]. A 4-layer scan therefore covers at most 16 states. A recurrent network tracks the state with constant depth, since its depth grows along time, but it has to store any key–value associations in its fixed-size state.
\(S_5\)-Retrieval needs both abilities. Each sequence comes with a fresh random dictionary \(\mathcal D\) that maps each of the 120 permutations to a four-character string, written out in the context. After each state, the model must also output that state's entry:
That is five tokens per state, or 240 tokens at \(N=48\). Figure 4 generates instances of the task.
The sequence
Computing all states in parallel
prefix scan: layer j composes each node with the node 2j−1 to its left
Carrying the state instead
T²MLR(1,4): the state after layer 4 at one step feeds layer 1 at the next, so each state is one composition on top of the last. Illustrative, not a trained model; in the model the state is passed at every token, five tokens per state.
We compare a 4-layer LSTM, a 4-layer, 6-head Llama-style Transformer, and T²MLR(1,4) built on that Transformer, so that all four layers recur. The three have matched parameter counts. Training uses \(N\) drawn uniformly from 1 to 32, and evaluation covers \(N=1,\dots,48\). T²MLR trains for 150k steps and the baselines for 400k.
T²MLR reaches near-perfect exact match on whole sequences within the training range, beyond the 16 states that a 4-layer scan can cover, and keeps non-trivial per-token accuracy at lengths beyond 32. Both baselines collapse on exact match at every length and keep only modest per-token accuracy. These descriptions come from the curves in the paper's Figure 4, which print no values. This matches the intended inductive bias: the recurrent pathway can carry the state from step to step, while attention keeps access to the dictionary for the lookup.
Where to put the loop
For language modeling we start from SmolLM2-135M [19], a 30-layer Llama-style model, and build T²MLR variants with different recurrence boundaries on the same backbone. The fusion module adds parameters, so we widen the baseline from 576 to 584 dimensions. Every model then has about 136.4M parameters, with the baseline slightly larger. All models train for one epoch on the 10B-token sample of FineWeb-Edu [20] with identical hyperparameters and are evaluated zero-shot on seven benchmarks.
Every variant improves on the baseline's zero-shot average of 42.83, though T²MLR(15,16) only barely, at 42.89. The best is not full recurrence. T²MLR(13,18), with \(D=6\), reaches 44.14, followed by (9,22) at 44.01, (5,26) at 43.73, and (1,30) at 43.36. Validation loss follows a different trend. Apart from full recurrence and the 2-layer block, which do not beat the baseline, loss after 16k steps mostly falls as more layers recur. Saunshi et al. observed a similar gap for stacking middle layers, which helps reasoning in ways that perplexity does not show [4].
Size is not the only variable. Holding the block width fixed and moving the block isolates location. With six layers, the middle block 13–18 averages 44.14, against 42.74 for layers 1–6 and 42.87 for layers 25–30. With fourteen layers, the middle block 9–22 averages 44.01, against 42.21 for layers 1–14 and 43.14 for layers 17–30. Both early blocks end below the baseline, and the late blocks gain at most 0.31 points over it.
We also compared against other ways of spending more compute at the same parameter count and data: a 2× pause-token model [21] (43.31), a Transformer that loops its full stack twice (42.99), and one that loops layers 9–22 three times (42.68). All three pay for the extra compute at every decoding step. T²MLR adds only the fusion step (at most about 8% in wall-clock), and every variant of the first sweep with at least six recurrent layers averages higher (43.36 to 44.14). Figure 5 gathers these comparisons.
Zero-shot averages differ from the baseline by at most 1.31 points. The larger differences appear after finetuning on reasoning tasks: Variable Assignment from the reasoning primitives of Saunshi et al. [4] at depth 5, ProsQA-Hard [6], the easy subset of HotpotQA [22] with chain-of-thought traces, and GSM-Aug in symbolic and natural-language form [23]. Every finetuned T²MLR variant beats the baseline on every task (bar labels of the paper's Figure 6), and on every task the best score comes from a middle variant rather than full recurrence. The gap is largest on Variable Assignment, where full recurrence reaches 0.773 and T²MLR(13,18) reaches 0.945. On GSM-Aug the best middle variant is 0.045 above full recurrence in the symbolic form (0.436 against 0.391) and 0.046 above it in natural language (0.314 against 0.268). The middle variants are not uniformly better, though: on HotpotQA, T²MLR(13,18) scores 0.216, barely above the baseline's 0.214 and below full recurrence at 0.246.
Scale and retrofitting
The gains persist at larger scale. With \(\ell_{\text{start}}=8\) and 10B training tokens, the zero-shot average rises from 48.48 to 49.22 at about 370M parameters and from 51.38 to 52.37 at about 1B. Extending training to 50B tokens, it rises from 48.35 to 49.42 at 135M with T²MLR(9,22), and from 52.83 to 54.78 at 370M with T²MLR(9,24). Finetuned reasoning improves more:
| Model, tokens | HotpotQA-Easy | GSM-Aug-NL | GSM-Aug-Sym |
|---|---|---|---|
| 370M, 10B | 24.43 → 28.28 | 31.08 → 34.12 | 39.35 → 42.46 |
| 1B, 10B | 23.28 → 26.52 | 32.37 → 36.69 | 43.97 → 44.96 |
| 135M, 50B | – | 26.46 → 30.48 | 38.67 → 41.55 |
| 370M, 50B | – | 40.18 → 44.35 | 42.53 → 46.63 |
Accuracy in percent, parameter-matched Transformer → T²MLR after finetuning. A dash means not evaluated. Values from Tables 9 and 10 of the paper.
T²MLR does not have to be pretrained from scratch. We inserted the recurrent pathway, as T²MLR(5,28), into the pretrained SmolLM2-1.7B-Instruct (32 layers) and finetuned for one epoch on OpenMathReasoning [24]. Against the same model finetuned in the same way without the pathway, GSM8K rises from 35.78 to 39.88 and MATH500 from 12.80 to 18.00.
One consequence of the recurrence is visible in probes. The loss at token \(t+1\) backpropagates through \(\mathbf R_t\) into the computation of token \(t\), which rewards states that anticipate what comes next. Two-layer MLP probes trained to predict the next token from the intermediate layers of token \(t\) reach lower loss on T²MLR than on the baseline. With the recurrent block ending three layers below the output head, probes one and two layers below the head reach losses of 5.091 and 5.088, against 5.135 and 5.131 for the baseline.
These comparisons fix parameters, data, and inference compute, not training compute. A Transformer trained for 2.24 epochs, which matches the training wall-clock of T²MLR(13,18), averages 45.30 zero-shot and beats T²MLR's 44.14. The contribution of T²MLR is architectural, and it shows at inference time, not as a saving in training compute.
Related work
Latent reasoning. Coconut [6] and CODI [7] replace parts of a textual chain of thought with last-layer hidden states that are fed back after the embedding. Soft Thinking, Multiplex Thinking, and related methods feed back mixtures of token embeddings [8, 9]. The closest prior work is HRPO [10], which also uses gated fusion but feeds its latent through the embedding and trains with reinforcement learning, without gradients through past tokens. T²MLR injects a middle-layer state into an earlier middle layer and trains through time with dense supervision.
| Method | fed back | enters at | fusion | BPTT |
|---|---|---|---|---|
| Coconut | last-layer state | after embedding | identity | yes |
| CODI | last-layer state | after embedding | 2-layer MLP | yes |
| Multiplex Thinking | soft token mixture | after embedding | identity | no |
| HRPO | soft token mixture | after embedding | gated | no |
| T²MLR | middle-layer state | earlier middle layer | gated | yes |
After Table 4 of the paper. For HRPO and T²MLR, the discrete token still enters through the input embedding as usual.
Looped Transformers. Universal Transformers [25] and later looped models [26, 11, 12, 13] reuse a block several times per token, and their inference cost grows with the number of loops. Geiping et al. [12] and Ouro [13] let earlier loops read key–value caches written by later loops, which is another shortcut from deep to shallow representations.
State-space, hybrid, and memory models. Mamba, RetNet, Griffin, and DeltaNet replace or interleave attention with recurrent sequence mixers [27, 28, 29, 30], usually for long-context efficiency. Transformer-XL, the Compressive Transformer, the Recurrent Memory Transformer, and Titans carry memory across segments for long contexts [31, 32, 33, 34]. T²MLR leaves attention unchanged and adds one residual pathway across decoding steps, aimed at reasoning within the context rather than at extending it.
Limitations
- Training takes 2 to 4 times the wall-clock of a Transformer, depending on the block size, and a Transformer trained for the same time wins on zero-shot benchmarks.
- Models trained from scratch have 135M to 1B parameters, and the retrofit uses 1.7B. Multi-seed variance estimates are not reported.
- Which fusion parameterization and recurrence boundaries work best is still open, and what information travels through \(\mathbf R_t\) is not yet understood mechanistically.
- The \(S_5\)-Retrieval comparison covers one small configuration per architecture, and its results are reported only as curves.
References
- Tenney et al. BERT Rediscovers the Classical NLP Pipeline. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
- Geva et al. Transformer Feed-Forward Layers Are Key-Value Memories. Conference on Empirical Methods in Natural Language Processing, 2021.
- Meng et al. Locating and Editing Factual Associations in GPT. Advances in Neural Information Processing Systems, 2022.
- Saunshi et al. On the Inductive Bias of Stacking Towards Improving Reasoning. Advances in Neural Information Processing Systems, 2024.
- Atanas and Liu. A Modular Dataset to Demonstrate LLM Abstraction Capability. arXiv preprint arXiv:2503.17645, 2025.
- Hao et al. Training Large Language Models to Reason in a Continuous Latent Space. arXiv preprint arXiv:2412.06769, 2024.
- Shen et al. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation. arXiv preprint arXiv:2502.21074, 2025.
- Zhang et al. Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space. Advances in Neural Information Processing Systems, 2025.
- Tang et al. Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge. arXiv preprint arXiv:2601.08808, 2026.
- Yue et al. Hybrid Latent Reasoning via Reinforcement Learning. arXiv preprint arXiv:2505.18454, 2025.
- Saunshi et al. Reasoning with Latent Thoughts: On the Power of Looped Transformers. International Conference on Learning Representations, 2025.
- Geiping et al. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. Advances in Neural Information Processing Systems, 2025.
- Zhu et al. Scaling Latent Reasoning via Looped Language Models. arXiv preprint arXiv:2510.25741, 2025.
- Wu, Teng, and Tu. Parallel Continuous Chain-of-Thought with Jacobi Iteration. Conference on Empirical Methods in Natural Language Processing, 2025.
- Williams and Zipser. Gradient-Based Learning Algorithms for Recurrent Networks and Their Computational Complexity. In Backpropagation, pp. 433–486, 2013.
- Liu et al. Transformers Learn Shortcuts to Automata. International Conference on Learning Representations, 2023.
- Merrill, Petty, and Sabharwal. The Illusion of State in State-Space Models. International Conference on Machine Learning, 2024.
- Li, Guo, and Andreas. (How) Do Language Models Track State? International Conference on Machine Learning, 2025.
- Allal et al. SmolLM2: When Smol Goes Big – Data-Centric Training of a Small Language Model. arXiv preprint arXiv:2502.02737, 2025.
- Penedo et al. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv preprint arXiv:2406.17557, 2024.
- Goyal et al. Think Before You Speak: Training Language Models with Pause Tokens. International Conference on Learning Representations, 2024.
- Yang et al. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. Conference on Empirical Methods in Natural Language Processing, 2018.
- Deng et al. Implicit Chain of Thought Reasoning via Knowledge Distillation. arXiv preprint arXiv:2311.01460, 2023.
- Moshkov et al. AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning Dataset. arXiv preprint arXiv:2504.16891, 2025.
- Dehghani et al. Universal Transformers. International Conference on Learning Representations, 2019.
- Giannou et al. Looped Transformers as Programmable Computers. International Conference on Machine Learning, 2023.
- Gu and Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv preprint arXiv:2312.00752, 2023.
- Sun et al. Retentive Network: A Successor to Transformer for Large Language Models. arXiv preprint arXiv:2307.08621, 2023.
- De et al. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models. arXiv preprint arXiv:2402.19427, 2024.
- Yang et al. Parallelizing Linear Transformers with the Delta Rule over Sequence Length. Advances in Neural Information Processing Systems, 2024.
- Dai et al. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context. arXiv preprint arXiv:1901.02860, 2019.
- Rae et al. Compressive Transformers for Long-Range Sequence Modelling. International Conference on Learning Representations, 2020.
- Bulatov, Kuratov, and Burtsev. Recurrent Memory Transformer. arXiv preprint arXiv:2207.06881, 2022.
- Behrouz, Zhong, and Mirrokni. Titans: Learning to Memorize at Test Time. arXiv preprint arXiv:2501.00663, 2024.