AI

"Different Organizations Allow Different Things to Emerge": Yang Zhilin's First In-Depth Interview

Editor's note: This conversation, recorded in November 2023 shortly after the launch of Kimi Chat, was the first time Moonshot.AI founder and CEO...

Share: Share on X LinkedIn

Editor's note: This conversation, recorded in November 2023 shortly after the launch of Kimi Chat, was the first time Moonshot.AI founder and CEO Yang Zhilin shared his thinking publicly at length. With Kimi K3 drawing global attention to Moonshot, we are publishing English editions of his two long dialogues with @GeekParkHQ founder Jack Zhang, starting with this one.

At the time, Moonshot had one product, a chatbot built on a 100B-plus parameter model, and three clear labels: long context, proprietary and closed-source, consumer-first. Read it as a time capsule. The transcript has been edited and condensed for clarity.

Highlights

  • The necessary path to AGI is a new kind of organization, not any single technique. OpenAI's success is, at its core, organizational innovation. "Only if you get the organization right can you actually walk the AGI road."

  • Large model innovation cannot be planned in advance. Mobile-era requirements were deterministic; AGI innovation is post-hoc. You have to try before you know.

  • Different organizations allow different things to emerge. Google's environment could produce a scientific result, the Transformer, but not ChatGPT. OpenAI invented nothing new, yet an industrial masterpiece emerged from combining three factors: the Transformer, 10^25 FLOPs of compute, and twenty years of internet data.

  • The Transformer is a new computer. Parameter count is its CPU, context length is its memory. Forty years ago people thought 500K of memory was plenty. The same fallacy is being repeated about context.

  • The ultimate case for long context is a lifelong AI companion. Trust and complex emotion only show their power over decades, and "an AI that has to reset its context window every day cannot do that."

  • The super app entry point will most likely be closed-source, because owning the model's evolution creates product differentiation from day one. This is the bet Moonshot has since reversed with the open-weight K series.

  • An AI-native product is defined by two datasets. Training data determines what the model can do, test data determines whether it is actually usable. Define the datasets and the product is defined.

  • The classical product manager pointed at one spot on the map and planted a tree. In the AGI era you mark out the whole plot and let the model sweep the field. Not design first, then build; you complete the design through making.

  • The metric that matters is a Moore's Law of use cases: the number of usable scenarios must double every N months, exponentially, not one scenario and one dataset at a time.

  • The greatest companies of the next decade will fuse two cultures: Silicon Valley's technical idealism as the drive, and the Chinese emphasis on usefulness and business models as the fuel.

Why "the dark side of the moon"

Jack Zhang: People say your company is a bit mysterious, starting with the name. What is the story behind "Moonshot AI," or in Chinese, "the dark side of the moon"?

Yang Zhilin: It comes from the Pink Floyd album, The Dark Side of the Moon. The founders all love rock music. We used to play in bands. And 2023 happened to be the album's 50th anniversary.When you look at the moon, you only ever see the lit side. You never see the back, and that gives you a strong urge to explore it. Large models feel similar. You want to explore something mysterious and unknown. It is hard, and it connects to the spirit of rock music: keep innovating, keep challenging the existing shape of things, keep imagining what comes next.We settled on the pairing of "dark side of the moon" in Chinese and "Moonshot" in English. It reflects our commitment to AGI. It might also define what kind of people we are.

Jack Zhang: I'm a Pink Floyd fan too. The album has an astronomical name, but it is really about the human subconscious.

Yang Zhilin: Right. The lead track is Brain Damage, about a person experiencing hallucination. Fifty years later, here we are trying to fix hallucination in large models.

Jack Zhang: What did you play in the band?

Yang Zhilin: Drums. The drummer keeps time and gives the whole band a frame to play inside.

Building a new kind of organization is the necessary path to AGI

Jack Zhang: How did you decide to go all in, and specifically to build a company around it?

Yang Zhilin: My understanding shifted enormously over the past few years. At first I thought language models were a tool that could improve results across scenarios. In the second stage, I thought they might be useful for many tasks. Eventually the view became: language modeling may be the only problem AI needs to solve. Everything can be addressed by making the model better, by making next-token prediction better.

In 2018 and 2019 at Google, we started training language models on thousands of chips with Transformers. You would observe so many phenomena along the way, and they kept adding evidence that this path was correct. Just keep walking it, keep finding more efficient ways to scale, and you get remarkable results on problems that used to be very hard: memory, reasoning, common sense, even complex multi-step problems. That experience left a deep mark on me. It was the runway toward founding a company.From 2020 I worked with many institutions to train large models. I was involved in some of the earliest large-scale efforts in China, including Pangu and Wudao. Through that process, I saw the challenges up close. Some were technical. Others were organizational. We found that if you use a traditional organizational structure, training frontier models is very hard to pull off. OpenAI's success is, at its core, also a success of radical organizational innovation.

So you could say I had been searching for one opportunity the whole time: the chance to build a new organization from zero. I believe this is the necessary path to AGI. It matters even more than the technical details we touch every day, because the organization is the deeper layer. Only if you get the organization right can you actually walk the AGI road.

By 2023, both the capital market and the talent market had changed dramatically. The timing was finally right.

Innovation in the large model era cannot be planned in advance

Jack Zhang: What convinced you that the organization is the core problem?

Yang Zhilin: Practice, mostly. Before this year I tried many modes: working inside a large company, working with independent research institutes, other supposedly efficient setups. None of them could produce fundamental organizational innovation.

Here is a simple example. You cannot innovate on large models through planning. In the mobile internet era, I could plan the features I wanted to build. Once a requirement was defined, it could be deterministically produced. You rarely heard of an app that a team suddenly did not know how to build. It was a deterministic event: humans encode the logic, computers execute it.AGI is different. I cannot plan that today we will fulfill a certain requirement to a certain degree, because it cannot be hard-coded or expressed as rules. AGI-style innovation is not front-loaded planning. It is post-hoc. You have to try before you know. You need an underlying machine that does many things in a systematic way.That is a fundamental difference.

Your organization has to match how you do things, so when the underlying logic changes, you need a new organizational form. The internet era produced excellent organizations, superb at things like recommendation-driven products. The new era will likely produce organizations that are superb at AGI. I think that is very likely to happen.

Jack Zhang: Is OpenAI a good template in your eyes? What did they get right, and what might not be optimal?

Yang Zhilin: On results alone, OpenAI made an enormous breakthrough. Without that company, the trajectory of humanity might be different.Going deeper: a good organization needs very high talent density, a shared vision, and the ability to focus efficiently on one goal. They did all of that extremely well.But the most essential point, and the one hardest to see from outside, is this: once you have those preconditions, how do you find a systematic way of doing things? That is the precondition for all the technology, and it is what we most want to iterate on.

Jack Zhang: By systematic, do you mean it can be replicated and scaled?

Yang Zhilin: Yes, but replication in the sense of reusing it across different problems, not copying it to another company. It forms inside one company and is very hard to transplant. But that company can use the same system again and again. Today I can use it to crack long context. Tomorrow, autonomous AI capabilities. The day after, multimodality. It is a reusable system that accumulates into your core asset. Every AGI company should spend serious time polishing this.

Google's organization let Transformer emerge. OpenAI's let ChatGPT emerge.

Jack Zhang: When people discuss OpenAI and organizational design, the debate is often bottom-up versus top-down. How do you frame it?

Yang Zhilin: The top-down frame still applies, especially for large models. Top-down is about leadership vision: can you judge what is right and worth doing, and what you should not do right now? AGI is like a lunar program, a giant system that takes a long time and many tightly coupled people. That kind of top-down design is necessary.Under that frame, what matters is whether the system lets many small units each produce efficiently, and then the top-down frame integrates the output.

Jack Zhang: So in a sense, you are running a big frame within which different directions can produce emergent innovation, because nobody can precisely define the one path to AGI today. The organization has to support emergence.

Yang Zhilin: Exactly. Look at how the Transformer came to be. Google gave those researchers an environment for emergence. Before the Transformer, the pieces already existed: attention, residual networks, LayerNorm, SGD and the training toolchain, learning rate schedules. Everything was prepared. Google provided the environment where people could freely combine them, and the emergence happened.But different environments allow different things to emerge. Google's environment could produce a scientific result. It could not produce a great work of systems engineering like ChatGPT, an epoch-defining product that pushed execution to the extreme while precisely capturing demand. Google's organization and ChatGPT simply do not match.OpenAI invented nothing new.

What emerged from it was an industrial masterpiece. It combined three things on a different dimension from Google: the Transformer architecture, compute centers capable of 10^25 floating point operations, and twenty years of data accumulated by the entire internet, which may be the internet's greatest value. OpenAI saw those three factors and provided the environment for them to combine. Out came a milestone on the way to AGI.That is my point. Different organizations allow different things to emerge. Whatever you want to emerge, that is the direction you should tune the organization toward.

The technical path to AGI is set. The product path is not.

Jack Zhang: A popular line in 2023 goes: the cards are on the table, now bet big. Meaning the technical path to AGI is settled and the game is now about who pours in more resources. Do you agree?

Yang Zhilin: The first principle of AGI is now clear: keep improving lossless compression and you get higher degrees of intelligence, eventually beyond human. There is abundant evidence. Like Newton's laws for classical mechanics, the big direction is basically settled. So that popular line has some validity.What remains is the second layer: under the big principle, how exactly do you do specific things?

For example, how do you build genuinely lossless compression over a long context? That is not simple. Even OpenAI has only taken the first step. Every subsequent step still carries uncertainty.Then, beyond technology, there is the product layer. We are still far from the superintelligent AI of science fiction, and today's products are not necessarily heading in the right direction.Every era has its greatest people and its next-greatest people. The greatest discover the correct first principles. Then a cohort of people, slightly less great but still great, solve the technical, product, and commercial challenges. That second layer is a vast open sea. How to play in it is still full of unknowns.

Transformer is the new computer. Context length is its memory.

Jack Zhang: Long context is Moonshot's specialty. With so many possible directions, why this one?

Yang Zhilin: Everyone should ask themselves what they want AI to do for them, what the human-AI relationship should be. In the ultimate form, one big capability is missing today: a much longer input window.The gap between a long window and a short one is more fundamental than I once thought. One ultimate form of AI is building long-term emotional value with a person, a lifelong companion over nearly unlimited time. Time is a critical dimension. Only over long stretches do trust, complex emotion, and decades-spanning interaction show their power. That is when AI can offer deep value to the human spirit. An AI that has to reset its context window every day cannot do that.

Jack Zhang: Tomorrow it forgets what you did today.

Yang Zhilin: Right. So think of the Transformer as a new computer with two critical dimensions. Parameter count determines computational complexity, like the CPU in the old computer. Context length is the new computer's memory. It determines how much can participate in the computation.Given sufficiently complex computation, the bigger the memory, the bigger the unlocked application space. Look at computing history. Forty or fifty years ago, everyone thought 500K of memory was plenty. Today that is obviously absurd. The same thing will happen with this new computer system. Long context, as the memory of the new computer, is absolutely essential.

Closed source, in service of a super app

Jack Zhang: This wave of model startups includes plenty of open-source players. Moonshot is closed-source with no plans to open up. What is the thinking?

Yang Zhilin: We strongly support open source. Open and closed will be complementary in this field. Open source lets developers try all kinds of applications, with stronger compliance control over data, training, and deployment, and more flexible scenarios.Closed source has its own value. The future's super app entry points, whether in productivity or consumer entertainment, will likely be built around closed models. The two approaches complement rather than conflict. The choice depends on each company's strategy. Ours is to build a super app. That is where all our time goes.

Jack Zhang: If someone wants to build a super app on an open model, buy the engine and modify it, why insist on building the engine end to end?

Yang Zhilin: Building applications on open source is a real opportunity, and the two do not conflict. But if the endgame is a super entry point, it will most likely be closed, because closed source lets you differentiate the product from day one of building the model. When you control the model's long-term evolution, you have full room to create a decisive product advantage.Applications on open models may not become the super entry point, but they can create incremental value: productization, proprietary data, fine-tuned capabilities others do not have. Both paths will exist in the ecosystem.

Also, right now we are in a technology-driven phase. A better foundation model converts into product advantage, so your base capabilities need to stay ahead of the commodity level. In ten or twenty years, when the technology commoditizes, you will need to convert first-mover advantage into more durable moats, like stronger network effects.

Begin with the end in mind: consumer is the only mode that matches AGI

Jack Zhang: So no API business, no helping enterprises deploy models. You are going consumer. The last AI wave produced almost nothing on the consumer side. Everyone ended up doing B2B, which at least reliably generates revenue. Why so firmly consumer this time?

Yang Zhilin: We are not refusing B2B entirely, but the focus and the push are consumer.For a long time, AI technology had no consumer success stories. With the new technical variable, AI can achieve results that were previously impossible, and those results can appear as new applications and new entry points, with exponential revenue growth and fast-growing users.

Midjourney, Character AI, and ChatGPT have all largely proven that AI-native apps have a real shot.And if you are doing AGI, you must choose a business mode that matches it, one that demands extreme innovation efficiency. Only consumer lets you close the loop fast. Only there can the organization form a culture of rapid iteration: updating models, adjusting the organization, and meeting user needs on a daily cadence, everything revolving around data at high speed. Only the consumer side generates that kind of energy, the kind that matches a company whose goal is AGI.This is thinking from the end backward. Co-creating with consumer users is itself doing AGI. It may even be a necessary precondition. AGI cannot be built behind closed doors.

The core is data. Without co-creating with users, you cannot get enough high-quality data, you cannot know what problems the model produces in real use, and you cannot dig deeper into scenarios together with users.

Jack Zhang: So it comes back to first principles about the goal. Without a consumer super app, enough users, and enough data, you cannot actually reach AGI. A company that does not want to build a super app is arguably not a true AGI believer.

What super app means in the AGI era

Jack Zhang: How do you define a super app here? WeChat is one because users do everything on it. Taobao is one by scale and value.

Yang Zhilin: The definition itself is not new. What changes is the value delivered. Only by providing value that could not be provided before does a new entry point appear. In the end you still need many users, high frequency, and large value created in use.But AGI has one property that makes a super app possible: the G, generality. It does not solve one class of problems. The set of problems it can solve keeps growing, so the product boundary keeps expanding. AI penetrates deeper into every part of life, the application's value keeps strengthening, and it fits the definition of a super app more and more. Generality and super app status are compatible.

Jack Zhang: That is a key point. In the mobile era, an app first reached scale, then expanded its service boundary to become general. In the new paradigm, you start with a general productivity engine, and it naturally becomes a super app. Different eras, different genes. Last era was a land grab for scale. This era, the technology engine gives you super app genes from birth.

Yang Zhilin: Well put. One addition: however general the technology, you still start from a subset of scenarios and generalize outward. And the generalization speed can be exponential rather than linear.

AI-native development: define two datasets and you have defined the product

Jack Zhang: Under the old paradigm, product managers, frontend, and backend collaborated toward a defined target, shipped in cycles, ran A/B tests. What does development look like for an AGI-era super app?

Yang Zhilin: Product development changes with the underlying technology. Mobile-era development meant clear requirements mapping to deterministic operations and fully deterministic events, built on the old computer and deterministic code. Deterministic logic gave rise to deterministic graphical interaction. Deterministic GUI plus deterministic systems: that was twenty years of internet product development.

Today the paradigm has shifted. The frontend becomes conversational UI, and more products will adopt it. The backend has been unified, to an extreme degree, into one language model. And it handles more than language. It processes all the world's information. It is encoding and losslessly compressing everything.With both ends settled, most application-layer development no longer touches backend architecture or frontend framework. There may be blends of GUI and conversational UI, but the overall architecture is basically fixed.

Most of what we call development today happens in the middle layer: data. Same interaction, same model class, different data, different product. ChatGPT, GitHub Copilot, Midjourney: essentially the same thing, differing mainly in how the data is defined.This is a massive paradigm shift. What product managers increasingly need to think about is how to build a product out of two datasets. Define the datasets and the product is defined. The training data determines what capabilities the model provides. The test data determines how usable it actually is.Before, there were no AI-native products, only AI features, so this way of working was rare. Many people with strong product sense do not know how to apply it.

Say you want to build something like Character AI, or improve on it. How do you define your two datasets? You need strong data production and processing techniques, ways to acquire data, judgments about which data is effective. We are still making these workflows concrete through exploration. The new development method of AGI probably requires a new organizational form to pull off.

The product manager of the new era

Jack Zhang: Early mobile internet ran on product managers with imagination, people who defined future scenarios by instinct. I once joked with Allen Zhang, the creator of WeChat, and he called himself a classical product manager. I like to call that quality a touch of the divine: they could not fully explain their convictions, but the convictions turned out right. Later, product management became scientific, data-driven, A/B tested. For a super app, do you want more of the divine or more of the scientific?

Yang Zhilin: You look for the balance, but long term, I believe the system will win overwhelmingly and become the mainstream development paradigm. That does not make the divine instinct unimportant. It just needs a strong system underneath. The system should be the main force.

Here is my usual metaphor. Allen Zhang pointed at one spot on a giant map and said, plant a tree here. He turned out to be a god, because the spot was right and the tree grew into a forest. A sharpshooter who hits whatever he aims at.AGI does not work that way. The industrial designer Sori Yanagi liked to say you do not manufacture according to a design, you complete the design through making. Build the thing, and the design is done, rather than designing first and building after.The old way was like scouting forever for the one spot, planting one tree, and celebrating when it survived.

Now, you glance around, say this patch of land looks decent, and the sharpshooter marks out the whole plot. AGI is your main force. It rolls across the entire field and finds every opportunity, every place a tree could grow.With a strong system, you realize that having talent-and-luck-driven judgment pick individual planting spots, out of tens of millions of possible scenarios, is extremely inefficient. The search for product-market fit in the AGI era should be done the AGI way: use its generality, use the user ecosystem, use the system, and push across the whole front at once. That is the biggest difference from the classical product manager.

Jack Zhang: So instinct concentrates at the front, on problem definition and goal selection. Getting there faster and better is the system's job.

Watch the delta, not the balance

Jack Zhang: More concretely, the product people you have hired, what do they share? Reverse-engineer the traits for us.

Yang Zhilin: An open mind and the ability to learn. Together they point to one thing: can this person iterate fast? That is the trait we value most. Not just product managers. Every role, every person, AI or not. Change is too fast now. You basically cannot predict what AI will offer by the end of next year, let alone three, five, ten years out. So everyone needs to learn fast with an open mind, then go execute concretely.For product specifically, you need strong consumer sense. Shipping version one is the easy part, because you have not yet gone through making it continuously better, or the process of defining ever more precisely what product you want. "Better" is deeply abstract. How do you make ChatGPT better? What counts as better? In which direction, and by how much? All hard to define.

A common trap for product managers is defining a pile of features, the old way. That is probably wrong now, because your features are defined through data. That is the AI-native way. And this is not static theory. You learn, then you try. What I said today might be wrong. Fine. Try it, gain something, take another gradient step, deepen your understanding. Out of that process, this era's Allen Zhang will appear.

Jack Zhang: Or not a new Allen Zhang. He was the classic of his era. The next era will produce a completely new figure, whose touch of the divine will look different. That is the most exciting part of progress. And one thing is certain about talent now: watch the delta, not the balance. Maybe ignore someone's history, but look at the increment between their past and their present, their present and their tomorrow. In an uncertain era, a large enough delta means a lot.

Yang Zhilin: Yes. Often, carrying no historical baggage is an advantage.

Jack Zhang: People in the comments are asking whether you are hiring.

Yang Zhilin: We are. Full-time and interns. We are quite open about backgrounds. AGI is a comprehensive undertaking. The technology alone is full-stack: NLP, computer vision, RL, alignment, infrastructure, kernel engineering. Beyond technology, product, operations, commercialization. Ideally, very diverse backgrounds, but one shared vision. We welcome anyone with passion for AGI, for the super app, and for the global market.

The metric that matters: a Moore's Law of use cases

Jack Zhang: Sam Altman has written about Moore's Law for intelligence, and Moore's Law for everything: a Moore's Law relationship between the cost and capability of intelligence. Do you buy it?

Yang Zhilin: Yes, and I think Moore's Law itself is going through a paradigm shift. The original version: transistor count doubles every N months. Then model parameters and compute followed it: FLOPs double every N months.For us, the Moore's Law of intelligence ultimately means this: every N months, the number of usable use cases doubles. It is an extension of scaling laws. Standard scaling laws describe pre-training: add compute, data, and parameters, and watch how training loss changes.But the metric that matters most in the end is the Moore's Law of use cases. How many scenarios reach usability? It has to rise exponentially, not linearly, doubling every N months.

You cannot use the traditional AI method of adding one scenario and one dataset at a time to make it work in that scenario. You will never get exponential growth that way.Measure how many scenarios get unlocked. With that, the search for product-market fit accelerates enormously, and you can try many things at once. Do not plant one tree. Mark out the land, and with a Moore's Law of use cases, you test the entire plot in one pass. A superfast tree-planting machine that knows whether a tree will live or die without planting it. When the scenarios multiply past a certain point, you become a super entry point.

Jack Zhang: The Moore's Law of intelligence and the Moore's Law of use cases should form a double helix. Costs keep falling, capabilities keep unlocking, and scenarios multiply.

The greatest companies of the next decade will fuse two cultures

Jack Zhang: You worked at Meta and Google. Compare Silicon Valley's engineering culture with China's. What is each good at?

Yang Zhilin: Silicon Valley's engineer culture is distinctive. Take Noam Shazeer at Character AI. They built a product with a real degree of product-market fit, and in the early days they barely had dedicated product managers. Engineers with their own ideas, fusing technology with those ideas, taking one step forward on their own. That is worth borrowing, especially in the AI-native paradigm where technology and demand have to move toward each other. Think of constructing those two datasets. Without engineers willing to take that step forward, much of it simply does not get done.

At the bottom layer, we want to absorb the best of both East and West. OpenAI runs on strong technical idealism. "I want to build AGI," business model unclear, heavily funded from the start. The recent effective accelerationism wave is another expression of it. Google and Microsoft were also born of a degree of technical idealism.Chinese culture emphasizes usefulness, thinking within the premise of a business model.

In the next ten years of the AGI era, the greatest companies will probably combine the two. The pragmatic side finds a genuinely good business model, the fuel that keeps you burning. The idealistic side drives you beyond money and usefulness, toward the simple desire to see what the far side of the moon actually looks like.

Jack Zhang: We have always been good at goal-driven usefulness. But making something useful and universal may now require some moonshot spirit: aiming at something high, hit or miss, and moving toward the deep of the universe. Exciting goals are what gather truly excellent people. Thank you, Zhilin.

If you read both feeds, follow @GeekParkHQ.

Get the China angle before it hits the global feed.

@GeekParkHQ

Subscribe by email

Get a note when GeekPark publishes a new English story.