Re-listening to Yang Zhilin: Bet on Scaling, First Principles, and Long-termism

阅读中文原文

During lunch today, with nothing else to do, I listened again to the conversation between Yang Zhilin and Zhang Xiaojun from January 2024.

What surprised me most was that he had already thought many things through clearly more than two years ago. Long context, scaling laws, agents—these directions were not yet so clearly defined at the time, but he had already bet on them firmly, and he repeatedly emphasized long-termism and first principles. Two years later, these directions have almost all become unavoidable issues today.

What moved me most was this sentence:

Find a non-consensus, and within that non-consensus, find the one fundamental principle—the first principle—that actually works for AI today.

I have always believed that to do long-term and correct things, you cannot just pick easy tasks. So hearing him repeatedly talk about long-termism, what to do in the next decade, and first principles resonated with me.

In the end, whether it is entrepreneurship or research, everyone is betting. We draw on our own priors, taste, and judgment to bet on certain directions, and then try to stick with them for the long run.

Below are a few points that struck me after listening. In some places I will paraphrase as closely as I can to Yang Zhilin’s original wording.

More breakthroughs will emerge in industry

Yang Zhilin mentioned a trend: in the future, more valuable breakthroughs will emerge in industry.

This is not to say research is unimportant. More precisely, many breakthroughs are already difficult to achieve through pure research alone. What needs to be built now is often a huge system: the algorithm must be new, the engineering must be solid, and product and commercialization must also be considered.

Such things are hard to complete purely in a laboratory. The role of the research and education system will also change, taking on more of the function of training and developing talent.

AGI organizations need new organizational forms

He also said that AGI organizations need new organizational forms.

Existing organizational forms, such as companies, big tech departments, and research institutes, may not be entirely suitable for AGI. Moonshot itself is also continuously iterating, trying to find better ways to let the right people do the right things.

Nor will AGI be something built purely behind closed doors inside a company. It is likely to co-work and co-evolve with users, involving many collaborations and organizational issues beyond technology.

I agree with this judgment. AGI has already exceeded the scope of ordinary software projects and ordinary research projects. It requires long-term technical vision, capital, talent density, engineering execution, product feedback, and organizational iteration. In a sense, the AGI company itself must also be redesigned.

Freeing yourself from endless polishing

Yang Zhilin said that his biggest takeaway from working on Transformer language models at Google was freeing himself from endless polishing and learning to see the big directions and the big gradients.

Often, the gap in the field is not whether some technical detail is sophisticated enough, but which problem you are actually looking at. For example, repeatedly optimizing perplexity or loss on very small datasets, or designing complex architectures and regularization tricks, may improve metrics in the short term without necessarily getting at the essence of the problem.

A better question might be: what structure is general enough and scalable enough?

If a framework can accommodate many problems and keeps improving with sustained compute investment, it is closer to the underlying mechanism. Scaling laws are such an example.

So I strongly agree with what he said:

For problems that can be solved by scale, don’t solve them with new algorithms; the greatest value of new algorithms is to help the system scale better.

Entrepreneurship requires capital, talent, and timing

Yang Zhilin has always felt that creating an independent organization genuinely built for AGI makes sense. But this is not something you can do just because you want to; it requires certain variables to mature.

He mentioned two factors of production: capital and talent.

After deciding to start a company, the hardest part is assembling capital and talent, and both are highly dependent on timing. At that time, the window for the first round of funding was very short, possibly only one month. Too early, and the market has not caught up yet; too late, and the opportunity may already be gone.

He said that when he was in the United States, one night he did a precise calculation and concluded that he needed to raise at least $100 million within a few months. At that time, many people doubted whether such an amount could be raised, but it later proved to be the right judgment.

Hiring followed a similar logic. By March or April 2023, a lot of top talent began to realize that general AI might be the direction most worth investing in over the next decade, and the talent market began to move. At that point, what was needed was to reach the right people quickly at the right time. The industry circle is fairly tight-knit, which also gave them some advantage.

So the most interesting thing about 2023 was that capital began to enter in a concentrated way, talent began to aggregate rapidly, and general AI, which had not existed as an industry before, became a new direction that everyone bet on together.

Maintain long-termism while keeping a sufficiently keen instinct.

Technical vision determines many things

Yang Zhilin believes that the biggest advantage of companies like Moonshot is that top-level decisions are guided by technical vision.

For example, long context was something they basically decided to do when the company was founded. This was not a reactive move after spotting a hot trend, but a first-principles judgment: which capability is fundamental, such that solving it makes many existing problems disappear on their own.

This is not entirely the same as general product judgment. Ordinary product judgment asks: What do users want now? Is there PMF in the current market? These questions are certainly important, but AGI companies also need to ask longer-term questions: What capabilities will become infrastructure in the next 10 to 20 years? Which direction, once it holds, will rewrite the shape of the products that follow?

Long context is such a direction. Once the context length is sufficiently long, many problems that originally required complex external systems will be absorbed by the model capability itself.

AGI is what to bet on for the next decade

Yang Zhilin holds a strong view: AGI is the only meaningful thing in the next decade.

AI will not end in the next year or two because of some PMF. It concerns how the world will change in the next 10 to 20 years. Short-term PMF is certainly important, but moving too hastily makes it easy to be rendered obsolete by stronger model capabilities.

Many past customer service systems and dialogue systems are examples. They looked clearly valuable at one stage, but once the underlying model capabilities took a leap, the original product value was quickly absorbed.

So long-termism in AI is not a slogan. The underlying capabilities change too fast; if the direction you choose is too short-term, you may not even finish before the next generation of model capabilities subsumes it.

A few personal thoughts

After listening to the entire conversation, what I ultimately remember are a few phrases: bet on scaling, first principles, long-termism, and AGI.

These phrases cannot be understood in isolation.

First principles keep directional judgment from being pure guesswork. Scaling is a verifiable and extensible underlying mechanism. Long-termism determines whether you are willing to sustain investment during the non-consensus phase. AGI provides a sufficiently large goal that makes these bets meaningful.

That’s why I listened to this conversation again. The whole thing is really about one thing: taste is all you need.

The hard part is judging, when a direction is not yet consensus, whether it rests on a sufficiently fundamental principle. If it does, you should bet on it for the long term.

This holds for entrepreneurship, and it also holds for research.

So the judgment I want to leave at the end of this note is simple:

bet on scaling, bet on first principles, bet on long-term AGI.

References