Skip to content

PhD Candidate in Computer Science · Princeton University

Xingyu Zhu朱星宇

Portrait of Xingyu Zhu

I am a fourth year PhD Candidate in Computer Science at Princeton University working on theoretical machine learning and language modeling. I am fortunate to be advised by Professor Sanjeev Arora. I did my undergrad at Duke, where I was fortunate to be advised by Professor Rong Ge.

01Research

My interest spans across both theoretical and empirical ML. Recently I am especially interested in understanding the power of large language models through a semi-theoretical lens. I have also worked on optimization dynamics of deep neural nets. Besides, I am also broadly interested in theoretical computer science and algorithmic fairness.

Interests

  • Machine Learning Theory
  • Language Modeling
  • Theoretical Computer Science

Education

  • Ph.D. in Computer SciencePrinceton University · 2023 –
  • B.S. in Math & Computer ScienceDuke University · 2018 – 2022

02News

03Writing

04Selected work

05Publications

2026

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

Xingyu Zhu*, Pu (Luke) Yi*, Ziheng Cheng*, Ang Lv*, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong

arXiv

T²MLR: Transformer with Temporal Middle-Layer Recurrence

Ziyang Cai*, Xingyu Zhu*, Yihe Dong, Yinghui He, Sanjeev Arora

LIT Workshop @ ICLR 2026

Contextual Drag - How Errors in the Context Affect LLM Reasoning

Yun Cheng, Xingyu Zhu, Haoyu Zhao, Sanjeev Arora

RSI Workshop @ ICLR 2026Best Paper

arXiv

2025

On the Power of Context-Enhanced Learning in LLMs

Xingyu Zhu*, Abhishek Panigrahi*, Sanjeev Arora

ICML 2025Spotlight

arXivInteractive post

2023

Understanding Edge-of-Stability Training Dynamics with a Minimalist Example

Xingyu Zhu*, Zixuan Wang*, Xiang Wang, Mo Zhou, Rong Ge

ICLR 2023

arXivInteractive post

Fairness in the Assignment Problem with Uncertain Priorities

Zeyu Shen, Zhiyi Wang, Xingyu Zhu, Brandon Fain, Kamesh Munagala

AAMAS 2023

arXiv

2022

Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks

Yikai Wu*, Xingyu Zhu*, Chenwei Wu, Annie Wang, Rong Ge

arXiv

arXiv

* Equal contribution

Publications page →

06Talks