Writing
Blog
Notes on machine learning, theory, and whatever else. RSS
2026
- On the Power of Context-Enhanced Learning in LLMsTraining with helpful text in the context but no loss on it can be exponentially more sample-efficient than plain fine-tuning, and that text is hard to recover afterwards.
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleA four-parameter scalar network reproduces edge-of-stability training, and a two-step parabola explains why gradient descent ends just below the stability threshold 2/η.
- T²MLR: Transformer with Temporal Middle-Layer RecurrenceT²MLR feeds a middle-layer state from the previous token into an earlier layer of the next, adding latent recurrence at almost no inference cost.
- Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache CompressionChunked KV-cache compression gives every token a phase, and models that use it retrieve worse at some phases than others.