School of Mathematics and Statistics Xi'an Jiaotong University
Hi, I am Jiaxuan Zou (邹嘉轩).
I am an undergraduate student in Mathematics and Statistics at Xi’an Jiaotong University and currently a research intern in the ByteDance Seed Pre-training team. Previously, I was a research intern at the Gaoling School of Artificial Intelligence, Renmin University of China, advised by Prof. Yong Liu.
My research focuses on mechanistic interpretability, deep learning theory, optimization, and scaling laws. I am especially interested in turning empirical training phenomena into first-principles mechanisms: why models train, where behaviors emerge, and when scaling or optimization rules break.
I also work as an AI Technical Consultant for a Tsinghua-affiliated AI startup.
I write research notes on my blog and welcome conversations on these topics. I am interested in future research opportunities in LLM pre-training and AI theory. More background: English / 中文.
Research Interests
- Mechanistic interpretability of LLMs
- Training dynamics of finite-width networks
- Optimizer design for LLM pre-training
- Scaling laws and their failure modes
- Deep learning theory and optimization
News
| Jul 01, 2026 | I joined the ByteDance Seed Pre-training team as a research intern. |
|---|---|
| May 25, 2026 | Our Nora optimizer has been included in the ScalingOpt optimizer library. |
| Apr 18, 2026 | I joined the ScalingOpt project as a co-maintainer, working on optimizer design for large language model training. |
Latest Posts
| Sep 20, 2026 | 如何搭建一个科学的 Scaling Ladder |
|---|---|
| Sep 19, 2026 | 当代知识分子是否仍有责任参与公共治理 |
| Sep 07, 2026 | pretrain 和 scaling 作为一种方法论和科学观 |
| Aug 25, 2026 | Hyperball、effective lr 与峰值加衰减的形状 |
| Jul 04, 2026 | 本科两年,我最深的感悟:认知复利 |
| Jun 20, 2026 | DASF:一种闭环的 batch size schedule-free 方法 |
| Jun 16, 2026 | 为什么 LLM pretrain 过程中途要把 batch size 翻倍 |
| Jun 03, 2026 | 不要只学习 19 世纪的西方:文明中心论、世界主义与青年领袖的公共责任 |