These are English translations of posts originally written in Chinese. Each one links back to its source. The originals live on the main blog page.
-
How to Build a Scientific Scaling Ladder
A practical engineering guide to designing and running a Scaling Ladder: covering evaluation protocols, measurement conventions, Dense and MoE scaling rules, experimental grids, hyperparameter transfer, data mixing and multi-epoch repetition, Loss Scaling Law fitting, downstream task prediction, and extrapolation validation, synthesizing public literature from Chinchilla, DeepSeek, StepFun, Cerebras, Llama 3, Delphi, and others.
-
Do Contemporary Intellectuals Still Have a Responsibility to Participate in Public Governance?
From ancient Greek poleis and the Chinese literati tradition to the postmodern dismantling of grand narratives, questioning where intellectuals' mandate to intervene in public governance truly comes from. This essay records how my initial assumptions collapsed, proposing an engaged posture on the ruins of postmodernism that acknowledges contingency yet still chooses commitment.
-
Pretraining and scaling as a methodology and scientific perspective
Starting from GEN-1.5's long-term pretraining, this article discusses general learning frameworks, training-time and test-time scaling, and how efficiency, training stability, and predictability guide research decisions.
-
Hyperball, effective lr, and the shape of peak-then-decay
Addressing two core questions in pretraining learning rate scheduling: the effective learning rate directly governing optimization is dynamically determined by weight norms, and Hyperball eliminates this implicit schedule by constraining weight norms; peak-then-decay corresponds to the optimal solution of the bias–variance trade-off, where shapes satisfying this balance form a set not limited to a specific analytic form.
-
Two years of undergrad: my deepest takeaway is cognitive compounding
Many people tend to make long-term linear plans based on limited present cognition, easily falling into anxiety when short-term efforts yield no immediate return. Drawing from my two years of undergraduate experience, this essay reflects on cognitive compounding: consistent public writing, immersing in high-density environments, and frequently iterating cognition—and how this yields nonlinear growth over the long run.
-
DASF: A Closed-Loop Schedule-Free Method for Batch Size
This post proposes DASF (Drift-Aware Schedule-Free): based on the duality of Schedule-Free (iterate averaging ↔ learning rate schedule, gradient averaging ↔ batch size schedule), it uses on-the-fly measured gradient statistics to set the effective batch size online, with no schedule and no tuning, eliminating the cost of training proxy models and fitting scaling laws for batch calibration. On real transformers it matches or exceeds tuned baselines, and provides a falsifiable negative result: under compute-optimal conditions, the optimal effective batch is approximately constant, not growing as √t.
-
Why does the batch size need to be doubled midway through LLM pretraining?
Starting from the Double GBS phenomenon of Apertus 70B, we use the gradient noise scale, critical batch size, and calculus of variations to derive the optimal schedule for increasing the batch size midway through LLM pretraining, and validate it on the noisy quadratic model.
-
Don't Just Learn from the 19th-Century West: Civilizational Centrism, Cosmopolitanism, and the Public Responsibility of Young Leaders
After reading Liu Qing's article on tianxia and new cosmopolitanism, here are some of my thoughts on civilizational centrism, 19th-century-style imaginings of great power, and the public responsibility of young technologists.
-
Re-listening to Yang Zhilin: Bet on Scaling, First Principles, and Long-termism
At lunch, I re-listened to the conversation between Yang Zhilin and Zhang Xiaojun from January 2024, jotting down some thoughts on long context, scaling laws, agents, AGI organizational forms, and long-termism.
-
μP Map
A reading guide and thread overview for the μP-related blog posts.
-
In the LLM context, how does noise in the gradient affect training dynamics?
A discussion of gradient noise in the late stages of LLM pretraining, and why blockwise normalized updates look more like constraining the update magnitude than correcting the gradient direction.
-
The Trade-off Between Parallelism and Expressiveness: Theoretical Boundaries from $AC^0$/$TC^0$ to Linear Attention
From the perspective of circuit complexity, this post gives a unified explanation of why constant-depth Transformers cannot exactly perform integer multiplication of arbitrary length, and why stronger linear attention variants often fail to maintain full token parallelism.
-
Bias and Fluctuations of the Spectral Norm of Random Gaussian Matrices at Finite Width
Starting from Wishart random matrix theory, this article derives the expansion of the spectral norm of a Gaussian matrix with element variance 1/n at finite width, showing that it not only converges to the macroscopic limit 2, but also carries a bias of order $n^{-2/3}$ and Tracy-Widom type random fluctuations.
-
Frobenius Norm Estimation of the Update Matrices for Adam and Muon Optimizers
This article rigorously derives and estimates the Frobenius norm of the update matrices of the Adam and Muon optimizers in a single iteration step, and explores the influence of matrix shape on the order of magnitude of the norm.
-
On the Hypersphere: μP Scaling of Optimizers with the Hyperball Mechanism
Starting from the first principles of continuous-time spherical dynamics, this article explores how the intrinsic dependence of weight norms disrupts hyperparameter alignment, and rigorously derives the underlying mathematical mechanisms by which various Hyperball optimizer variants achieve feature-space alignment.
-
On the Hypersphere: From Spherical Dynamics to μP
This article departs from the probabilistic framework of Tensor Programs and, from the perspective of continuous-time spherical dynamics, rigorously derives how to achieve alignment of networks of different sizes by aligning the dynamics on the hypersphere in network architectures that apply RMSNorm.
-
On the Loose Use and Methodological Boundaries of the "Manifold" Concept in Current AI
This article discusses the loose use of the "manifold" concept in AI theory research and draws the boundaries between engineering terminology, geometric intuition, and rigorous mathematical argument.
-
Tensor Programs (Part 2): From Tensor Programs to μP
This article systematically reviews the core theoretical derivations of the maximal update parameterization (μP) obtained from Tensor Programs. The most fundamental and central insight of Tensor Programs theory in deriving neural network scaling laws is that one must strictly distinguish and apply the law of large numbers (LLN) and the central limit theorem (CLT) according to the different generation mechanisms of weight tensors.
-
Tensor Programs (Part 1): From the Spectral Condition of Feature Learning to μP
This article introduces the introductory paper of Greg Yang's Tensor Programs series—A Spectral Condition for Feature Learning—which derives the scaling conditions required for feature learning from the perspective of the spectral norm, and then re-derives the Maximal Update Parametrization (μP) from them.
-
From Gated DeltaNet to Kaczmarz
This article starts from the online learning formulation of Gated DeltaNet, introduces the Kaczmarz algorithm as an alternative to SGD, and analyzes its geometric meaning and its connection to Longhorn.
-
How to align data scaling curves under different initialization scales
We study the relationship between the empirical slope of data scaling and the initialization std, and propose a simple method to align data scaling curves under different initialization scales.