Inside China's AI Labs: Nathan Lambert on Open Models, Compute Constraints, and US-China Divergence
打开互动全文版(中英对照 + 朗读 + 问答)→机器学习研究员 Nathan Lambert 分享其走访中国顶尖 AI 实验室的一手见闻,探讨文化组织差异、算力限制与开源生态。
Machine learning researcher Nathan Lambert shares firsthand insights from his trip to China's leading AI labs, discussing cultural and organizational differences, compute constraints, and the open-source ecosystem.
欢迎回到《差异化理解》的另一期节目。我是主持人 Grace Shiao。你们很多人也知道,我还在 Substack 上写行业通讯《AI Pro》,请关注一下。今天的嘉宾让我非常兴奋。如果你密切关注开源 AI 领域,Nathan Lambert 几乎不需要介绍。Nathan 是一位机器学习研究员,专注于构建、理解和倡导开放语言模型以及负责任的自主系统。他曾是艾伦人工智能研究所的后训练负责人,还在 Hugging Face、DeepMind 和 Facebook AI 工作过。他也是通讯《Interconnects》的作者。Nathan 最近刚从中国旅行回来,在那里他会见了许多领先的 AI 实验室和研究人员。他的反思是我读过的关于中国 AI 最深刻的文章之一,不仅因为涵盖了模型能力,还捕捉了这些模型背后的文化、激励和组织差异。他的文章还触及了一个我反复思考的更大问题:我们真的在看一场全球性的 AI 竞赛吗?还是美国和中国开始围绕截然不同的生态系统进行优化?Nathan,非常感谢你今天加入我们。
Welcome back to another episode of Differentiating Understanding. I'm your host Grace Shiao. As many of you know, I also write the industry newsletter AI Pro on Substack. So please do give that a follow. I'm very excited for today's guest. For those of you who follow the open source AI space closely, Nathan Lambert needs little introduction. Nathan is a machine learning researcher focused on building, understanding, and advocating for open language models and responsible autonomous systems. He was a post-training lead at the Allen Institute for AI and has also worked at Hugging Face, DeepMind, and Facebook AI. He's also the author of the newsletter Interconnects. Nathan recently came back from a trip to China where he met many of the leading AI labs and researchers across the ecosystem. His reflections were some of the most thoughtful pieces I've read on Chinese AI. Not just because they cover model capabilities, but because they capture the culture, incentives, and organizational differences behind how these models are being built. His piece also gets at a bigger question I keep coming back to: Are we really looking at one global AI race? Or are the US and China starting to optimize around very different ecosystems? Nathan, thank you so much for joining us today.
是的,谢谢邀请。我真的很兴奋终于能聊聊你的中国之行,以及中国 AI 实验室和美国 AI 实验室之间的情况。你认为潜在的算力限制对这些实验室及其未来表现意味着什么,当然还有开源生态系统。所以在我们深入讨论之前,你能简单介绍一下你是如何最终从事后训练和开放语言模型工作的吗?就说说你自己。好的。我实际上是在 2017 年开始在伯克利读博,当时并不是做 AI 相关的工作。我本科是电气工程,回想起来很有趣,因为那一年正是 Transformer 论文发表的时候,我当时想我应该做 AI 这件事,并试图找著名的导师指导我,但他们说不能收我。所以我的博士生涯是一条曲折的进入 AI 的道路,之后我去了 Hugging Face,那实际上是我唯一一份行业研究工作,但也是一个非常热门的初创公司,学习起来很有趣,处于人们大量用于 AI 的工具和研究的交叉点。然后当 ChatGPT 出现时,RLHF 在技术方面成了热词。我的博士研究方向是强化学习,这正好是基于人类反馈的强化学习的前半部分。所以很自然地转向了那个方向,Hugging Face 是一个很好的地方,因为整个公司都支持这个方向,即弄清楚如何在这个热门领域支持社区并构建平台。他们对此非常高兴,我帮助在 Hugging Face 建立了一个团队。后来我厌倦了远程工作的时差问题,发现艾伦研究所也在做类似的事情,我想,哇,我可以有面对面的朋友做类似的事情。我觉得生活质量需要这个,所以几年后我构建了一系列模型。我认为在非营利组织工作打开了这个生态系统的信息真空,因为没多少人能谈论他们在做什么。然后凭借一些运气和每周坚持写作,我觉得我的影响力填补了没人说合理事情的真空,并且我的写作和日常工作之间形成了很好的协同效应。它以一种非常有趣的方式越变越大。我认为在最高层面上,我的动力是希望 AI 沿着这条轨迹良好发展。我担心很多近期的事情,比如美国的社会动荡以及对 AI 的巨大仇恨,我认为这是一个非常大的近期问题;中期则是权力更加集中,因为我认为 AI 会以人们意想不到的方式变得超级强大。所以总的来说,开放模型是一种很好的方式,通过对人们更加透明来缓和这两个问题,它自然是对抗权力集中的对冲,虽然原因各不相同,但这在过去几年里一直是我生活中的一个反复出现的主题。
Yeah, thanks for having me. Yeah, really really excited to finally hear about your thoughts on your big China trip on what's happening between the China AI labs and the US AI labs. What you think of the potential compute constraint might mean for these labs and their performance in the future and obviously the open source ecosystem. So before we get into all of that, could you briefly tell us about like how you ended up actually working on post-training and open language models? Just a bit about yourself. Yeah. So I actually started my PhD at Berkeley in 2017 not working on AI things. I was an electrical engineer my training in undergrad which is funny looking back because that's the same year that the transformer paper came out and I was like I think I should do this AI thing and tried to get the famous advisers to mentor me and they were like we can't take you. So I had my PhD as this wandering path to become in AI and then I ended up at Hugging Face after that which was like a it was like the realistically the only industry research job that I had but also a very hot startup and very fun to learn kind of at the intersection of these tools that people use a lot for AI and research is what I was doing. And then when ChatGPT hit the kind of RLHF thing blew up as the hot word on the technical side of things. my PhD had ended up being in reinforcement learning which is just like the first half of reinforcement learning from human feedback. So it was kind of a natural pivot to be like well I might just do that and Hugging Face was a good place for doing that cuz they were the whole company is kind of all for that which is like figure out how to support the community on the hot thing and build platforms there. So they were very happy about that and I helped build a team at Hugging Face and then I was kind of burnt on the remote work time zone thing and found out that the Allen Institute was doing such similar stuff and I was like wow I have people that could be in-person friends and do similar things. I was like quality of life I need to do this and a few years later I ended up building a bunch of models and I think the being at a nonprofit kind of opened this ecosystem vacuum of information where there aren't many people I can talk about what they're doing. So then with some luck and committing to write every week, I just feel like my influence just filled the vacuum of nobody is saying reasonable things and is this nice synergy between what I write about and what I work on in my day job. And it kind of got bigger and bigger in a very fun way. And I think that generally at the highest level I'm motivated by wanting AI to go well on this trajectory. that I worry about a lot of near near-term things whether it's social unrest in the US and just kind of the massive hatred for AI I think is a very big near-term problem and then like medium-term more concentration of power because I think AI will be super powerful in ways that people don't expect. So generally open models is a nice way to curve on both of them by being a bit more transparent to people and it naturally is a hedge against concentration of power and there's been different reasons throughout that but that's kind of a a recurring theme in my life in the last few years.
当然。我喜欢你的工作,因为我认为你帮助像我这样的非技术人员更好地理解这些实验室背后发生的事情。实际上我上周刚和你的前同事 Tia Jin Wang 聊过,他曾在 Hugging Face 工作,他也说了同样的话。开源在很多方面是前进的最佳方式,因为我们知道这项技术不会停止进化,但它是为垄断设置护栏和制衡的最佳方式。好的,我今天不想在这方面花太多时间,因为我们的重点是你的中国之行。所以在我们深入细节之前,我想听听你这次旅行本身。大多数写中国 AI 的人都是二手信息。你真的去了那里。你和研究人员待在一起。你会见了构建模型的人。告诉我们,当你说你带着极大的谦逊回来时,你是什么意思?你的眼界更开阔了,无论是好是坏。跟我们说说你的旅行。
Definitely. And I love your work because I think you help non-technical people like myself really understand what's behind what's happening in these labs a lot better. And then I actually just spoke to your former colleague Tia Jin Wang, he was with Hugging Face just last week, and you know he was saying the same thing. It's just open source in many ways is kind of the best way to go forward as we know that this technology will not stop in evolution but it's the best way to have kind of put up guardrails and like checks and balances for the monopolies. Okay, I don't want to take up too much time on that side of things today because our focus really is about your China trip. So before we get into the weeds of it all, I want to hear about the trip yourself. Most people who are writing about Chinese AI are getting their information secondhand. You really went there. You spent time with researchers. You met with people who are building the models. Tell us about what you meant when you said you came back with great humility, right? Like your eyes were a bit more open whether it's the good or the bad. Like tell us about your trip.
我觉得我进去的时候,我的文章里有一个糟糕的英文短语,就是‘我知道我对中国一无所知’,这试图表明我在旅行前就知道自己一无所知,而且在我现在的写作中仍然是事实。这是一个写得很糟糕的句子,我之所以提到它是因为有人指出了这一点。我当时想,‘这是什么?’但离开时,我只是知道这是一个如此大的国家,有如此多的人才在研究这些问题,而且作为人类,要模拟具有截然不同世界观、成长经历和训练体系的人是多么不可预测。实际上,中国培养人的方式非常不同,我认为即使在那里,你也不能完全把握住那些三到六人的研究小组在做什么,即使他们在追求同样的目标,也可能与西方有所不同。我认为你可以通过社会学研究深入到那个粒度,真正看到他们在研究什么方面的差异,而这总会改变输出。
I feel like I kind of went in. I mean, I had this horrible English phrase in my writing, which was like 'I knew I knew nothing about China', which is like kind of tried to indicate that I knew going into the trip that I knew nothing and it was still the fact in my current writing. This is like a horribly written sentence that I had in there and I only talk about it because somebody called me out on it. I was like, 'What is this?' And it's like leaving but just knowing that it's such a big country and there are just like such vast amounts of talent working on these problems and how unpredictable it is as a human to model people with very different worldviews and upbringings and training systems realistically like the way that people are trained in China is very different and I just think that even being there you can't fully grasp just like what are the pockets of three to six researchers doing that is actually a bit different than in the west even if they're working on the same goal. Like I think you could get down to that level of granularity in a sociological study and actually like see differences in what they're working on and that it'll always change the output.
而且我并没有深入到那个粒度,但这只是为了开始获得真实经验,理解人们如何解释他们解决这些问题的方式。对我来说,实际上很多都是建立联盟,我只是希望技术公司在国际机构中的行为不要充满敌意。所以,与双方的实验室会面是很好的,因为你需要这样做,他们未来才会和你讨论更敏感的问题。有些人批评我的文章,说我不应该那样访问中国,但如果你要正式访问一堆公司,你还能怎么做呢?不友好怎么能进门呢?你必须从某个地方开始。我认为保持尊重很重要。坦率地说,我认为那些批评并不公平,因为我觉得你非常透明地表明了你并不是中国专家,你没有去那里把一切都异国情调化。相反,很多有中国背景的人喜欢用龙啊虎啊来形容事物,而你实际上非常谦逊,只是作为一个技术人员去和这些实验室讨论他们的技术研究。因为你亲自在那里,你对文化和人有了观察。所以,我认为你的文章相当不错。
And it's like I didn't get to that level of granularity but it's just to start having real experiences and understanding how people explain how they work on these problems. And for me realistically a lot of it is coalition building which is just like I want there to not be vitriol at the level of the technical companies doing things in international bodies. So just like meeting all the labs on both sides is really nice because you need to do that for them to get to talk to you about more sensitive issues in the future. So some people I got some criticism on this on the piece which is like this is how you shouldn't visit China and it's like what else are you going to do if you're going on an official visit to a bunch of companies. It's like how do you expect to get in the door without being nice? It's like you got to start you have to start somewhere. And I think that it's important to be respectful. piece was frankly like I don't think you you I don't think the criticism was like fair to be honest because like I think you were really transparent with the fact you're not a China person right it's not like you're like going there and exoticizing everything and if anything a lot of people even with China background like to use certain dragons and you know tigers to describe things I feel like you actually were really humbled going be like I'm just a technical dude meeting with these labs talking about their technical research right and then because you were physically there you had observations of the culture and the people. So yeah, I I actually thought your piece was quite good and yeah,
我同意。我愿意让那些批评过去。但我认为听众需要意识到这些公司多么积极地讨好西方受众,这就是为什么我们能进门。我们这次行程有一些知名人士,但正因如此,我们才能在想要的日子见到所有这些人。基本上每个人,比如和我一起做互连的 Katherine Rantel 和其他一些创作者。他以前住在中国,有中国的人脉。所以他协调了他的关系,并利用了我与实验室的联系,还有一些更大牌的人也在行程中。把这些整合起来,让所有实验室都到位,需要几个月的社交网络工作,确保行程与那些有既定网络和实验室联系人的人对接。但这些公司希望给西方受众留下好印象,所以他们只会对合适的研究人员说“是”。研究人员知道房间里有两到四个公关人员在确保一切顺利。公司越大,公关人员越多。你去阿里巴巴,会有三到五个人,从公关主管到一些特别办公室的人。不接受这些“陪同人员”的成本,你就根本见不到这些人。在美国也一样,你不能随便把一个高级主管按在椅子上。
I agree. I was willing to I let that sail bass. But I think it's it's important for people that listen to realize how actively these companies are trying to court western audiences, which is why we could get in the door. I mean like we had some prominent people on this trip, but that's why like we got all of them in the days that we wanted them. So essentially everyone some like Katherine Rantel who works with me on interconnects and some other creators. He used to live in China and has connections in China. So he kind of orchestrated the mix of his connections and leveraging like my connections to labs and some like we had some bigger names on the trip as well. Just like stringing all these together to get all the various labs in place is like a few months of like networking to make sure the trip lines up with people with established networks and contacts with the various labs. But like these people want to look good to western audiences. So they're going to only say yes to the right researchers. And the researchers know that there's two to four comms people in their room hanging out making sure that it goes well. Like especially the bigger the company, the more they're comms people. It's like you go to Alibaba and there's three to five various people from the head of comms to some special offices and it's like you're not going to get these people in in the office or at all without accepting the cost of these these types of handlers. It's the same thing in the US. It's like you're not going to just plop a a senior executive into a chair.
就像
It's like
当然。
of course.
所以这也是好事,因为现在我有了很多中国研究人员的微信,可以随时发消息。比如“嘿,恭喜发布新模型。”这个 Lee 在小米工作。我和他在商场里聊了一个小时,我不记得那个海产品店的名字了,但现在我们有了这些关系,这非常有用,有助于信息在生态系统中传播到这些可信的各方,这种关系以前并不存在。我认为这样的人不多。反向的行程非常困难,因为中国研究人员很难进入美国。签证炼狱太复杂了,而我们行程中的很多人要么是加拿大人,要么是免签入境,这使得美国技术人才现在很容易去中国,这就是为什么我认为会有这么多行程,而且会有更多。我们收到了很多来自美国风投和开源实验室的咨询,他们想与这些实验室建立合作,因为它们是开放权重模型最好的,他们想为美国构建开放权重模型的公司建立技术栈。所以我认为会有更多知名但并非巨头的美国初创公司尝试建立这些关系,我认为这是一个非常有趣的技术发展,因为我们从未见过美国科技公司在中国进行这种专业工作旅行。大多数科技公司,你带设备到中国,它会自动变砖,你必须上交。所以主动以专业身份派人是一个很大的变化。有很多角度可以看待这件事,但我认为看到它如何展开很酷。这甚至不完全是关于这次行程,而是后续我们听到人们说:“嘿,你是怎么做到的?我们也想做这样的行程。”
So it's also like good because now I have we chat of a bunch of researchers from China that I could just text about things. It's like, "Hey, congrats on the new model release." This like Lee works at Xiaoi Xiai Mimo. It's like talked to this guy for an hour at a at a mall at like a I don't remember the name of the sea store, but it's like now we have these relationships which is very useful and that helps information spread across the ecosystem to these trusted parties which doesn't really exist. There's not that many I think. And the opposite direction of the trip is very hard because Chinese researchers can't really enter the US. the visa purgatory is too complicated where a lot of us on the trip were either Canadian or enter on a travel without a transit without a visa entry which makes it very easy for American technical talent to go to China right now which is why I think there are so many trips I think there's a lot of there'll be more of them so we've got a lot of inbound from VCs and like open source labs in the US that want to establish collaborations with these various labs because they're the best openweight models and they want to build a stack for companies in the US building open weight models so I think there's going to be more prominent but not gigantic US startups going to try to build these relationships which I think is a really interesting technological development because we've never seen this type of pro work trip in China from US tech companies. So like most of the tech companies have a you bring a device to China, it auto bricks itself and you have to hand it into it. So to like actually proactively send people in a professional capacity is a really big change. And I I kind of like there's a lot of angles who take this but I think it's cool to see how it unfolds. This isn't even really about the trip. This is like follow on that we're hearing from people that are like, "Hey, how'd you do this? We want to do this trip."
是的,确实如此。实际上,从我这边听到的,我认为风投或投资者一直都很活跃地去中国,因为显然在互联网时代,美国基金在中国非常活跃,人们总是试图找到进入这些好交易或保持脉搏的方式。但我认为对整个 AI 生态系统来说,有这种某种程度上的公平透明交流是非常积极的。但正如你所说,明星研究人员不可能在没有合规的情况下出来私下谈话,因为在美国也不会发生。那只是公司保护自己。我只是觉得你的行程非常有意义,我想回到你的观察上。你谈了很多文化方面的事情。你谈到在中国,明星研究人员的名人光环较少。人们更谦逊,更注重执行。你认为中国实验室特别适合当前的语言模型构建游戏,因为他们非常注重细致的堆栈级工作,而且有时更少自我,愿意做脏活或不那么光鲜的工作。那么请为我们解析一下,为什么你认为如此?你提到他们成长方式不同。
Yeah, definitely. Actually, from my end, I hear about like I think VCs or investors have always been quite active going to China because obviously previously American funds were very very active in China during the internet era and then people were kind of always trying to find a way to either get into these good deals or potentially keep it, you know, their pulse on it. But I think it's really really positive for the whole AI ecosystem to have this kind of you know I guess like fair transparent you know exchange in some capacity. But to your point it's there's no way that like this researcher star researchers can come out and talk to you know off the record without any compliance because that doesn't happen in the US either. Like that's just like companies protecting themselves. I just think your trip was quite meaningful and I want to bring it back to your observations actually. You talked a lot about, you know, the cultural aspects of it. You talked about how you felt like in China there was less of this like star researcher of celebrity status around on people. People were more like either humble or there's more humility. It was very focused on execution. You argue that Chinese labs particularly well suited to the current LM building game because they're very focused on meticulous stack level work. Um, and there's less ego sometimes to work on kind of the dirty work or you know the non-s sexy work. So kind of unpack that for us like why do you think that is? You kind of touched on you said they were brought up differently.
他们被教导的方式不同,但到底有什么不同?所以,这次旅行中一个有趣的协同点是,我们拜访了一些学术机构,比如新加坡的 AIR 之类的,你听到这些学术领袖都在说他们正在努力推动改变。所以,中国——是的,他们知道中国发表的论文比任何地方都多,但他们仍然认为这些研究不够具有变革性,他们正在努力培育国内的学术生态系统,以改变工作的类型、分布,并承担更多风险。然后,你会和一些行业领袖私下交谈,听到他们说:‘哦,这永远不会改变,因为教育体系太僵化了,漏斗的很多层都奖励死记硬背之类的东西,这种研究文化根本不会出现。’接着,AI 实验室的交叉点在于,这些实验室在做快速跟进,他们已经有了概念验证,知道应该是什么样子,所以在这个领域,你不是在试图发明新范式,不是试图做出第一个在云端代码中工作的 01 或 03 模型,而是‘我看到了,我要尝试去做,并把它做到最好,努力让它更便宜,最大化那个目标’。我认为很多公司不需要发明新范式。OpenAI 已经做过很多次了。那是他们的拿手好戏——永远不要怀疑 OpenAI 发布一篇博客文章和一张图表就能改变人们对 AI 看法的能力。我仍然认为,在未来四年的这场大繁荣中,这种情况还会发生几次。OpenAI 就是有一种直觉,知道什么东西可以稍微提前推动,从而改变一切。但我不指望,其他人也不指望中国公司能做到那么多,因为那是一种……我不知道怎么描述它的积极面——一种建设的文化,但可能更务实一些,就是‘你的工作就是建造这个东西’。很多研究人员,也许因为他们认识他们的经理——有些经理就在房间里——他们觉得自己在公司的角色就是让模型变得优秀。特别是对于学生,我和学生一起工作,他们就是这么说的。我在艾伦研究所工作,我们有学生会共同领导我们的语言模型,这并不奇怪,因为如果你在美国做行业研究工作,很多导师会告诉你,你基本上没有官僚和政治的负担。所以,学生的天真和单纯实际上非常有助于完成大量技术工作。生活方面也是如此,如果你更年轻,没有那么多家庭负担,通常也没有养成太多习惯和其他生活琐事。语言模型非常复杂,你需要吸收大量上下文才能理解瓶颈在哪里。信息太多了,你必须能够找出瓶颈并打破它。如果你没有足够的心理空间去吸收所有上下文,你最终会做一些表面功夫,但不会在模型上取得突破。这就是我看到的区别,有些人在语言模型出现之前在学术上非常成功,他们中的一些人能够转向这种务实的心态——‘系统的状态是什么?我如何改进它?’而另一些人则试图构建抽象的框架,像做学术一样处理问题,这通常不会对模型有太大改进。所以,我认为如果学术体系更务实一些,更结构化一些,而你在语言模型上的工作也是结构化的——比如让这个内核实现更快,让这个想法奏效——那么也许可以。我认为这是过度简化。我在文章中稍微强调了这一点,只是为了对比你对美国实验室可能有的印象。我有一些轶事,比如我听说一个美国实验室付钱让一个研究人员闭嘴,因为他的东西不在瓶子里。所有这些一次性事件更多是讲故事的手段,因为大多数一次性事件根本不重要。但 Long Before 崩溃了,那是因为它被描述为一种‘权力的游戏’式的政治环境,所有副总裁都在争相展示他们的东西让基准测试上升,很多人都这么告诉你。我们也有 Quen 的人员更替,但似乎不像 Llama 那样,或者 XAI 现在几乎不存在了。美国也有一些戏剧性的事情,这些公司如何来来去去。
They were taught differently but what's so different? So essentially the interesting part of that synergizes on the trip is we stopped by some academic institutions like I think it's like AIR in Singwa and stuff and you hear all of these academic leadership talk about how they're pushing hard to try to change it. So China like yes they know China is producing more papers than anything else but they still see think that it's not as transformative of research and they think that they're trying to cultivate the academic domestic ecosystem to change like the type of work it works on and the distribution and take more risk and then I would talk you would talk to some industry leaders like off the record behind closed doors and you would hear things like oh it's never going to change because the education system is so structured and there are so many layers of the funnel that reward things like memorization and stuff that they they're just like this research culture is not going to emerge and then the kind of follow on is with the AI labs crossover is that these labs are doing fast following and there kind of is a proof of concept and they know what it needs to look like and therefore that domain you're not trying to invent the new paradigm you're not trying to make the model that is 01 or 03 the first model to work in Claude Code you like I see it and I'm going to try to do that and make it the best thing and it's going to be try to make it cheaper and just like maximize that goal and I like I don't think that like a lot of companies don't need to invent the new paradigm. Like OpenAI has done this so many times. That's like their bread and butter is just like never doubt OpenAI's ability to release a blog post and a plot that changes how people think about AI. Like it's I still think it's going to happen a few times in this massive boom over the next four years. It's like OpenAI just kind of has that sense of what is the thing that you can push on a bit earlier and just like transform things. But I don't expect and other people wouldn't expect the Chinese companies to do that as much because it's just like such a culture of I guess that's like I don't know how to describe the the positive version of this as a culture of building but it's just like maybe slightly more practical-minded in terms of just like it's your job to build this thing. A lot of the researchers maybe because they knew their manager like some of them had managers in the room or so on but it's just like they see their role in the company as being to make the models excellent and especially for students like I work with students and that's what they say like I work at the Allen Institute and we have students that will co-lead our language models and it's not that surprising because if you do an industry research job in the US like a lot of mentors will tell you that you're kind of free of the the burden of bureaucracy and politics. So, it's like the naivity of students and the simple-mindedness is actually so good at just getting a lot of technical work done. There's also the life side if you don't have like you're you're younger, you don't have as much family, you normally haven't built up as many habits and other things you do with your life. So, it's just like language models are so complex in the amount of context that you need to absorb to understand what the bottleneck is. There's just so much information and you have to be able to pick what the bottleneck is and break it. And if you just don't have the mental space to absorb all the context, you kind of end up doing things that are acute but don't make breakthroughs on the model. So that's kind of like a difference that I've seen in people who are both very successful academically before language models. Like some of them are able to pivot to this practical mind which is just like what is the state of the system? How do I improve it? And then some try to make kind of these abstract frames of what's happening and approach it like an academic and it normally doesn't improve the model as much. So I just kind of see like if the academic system is a bit more practical-minded, a bit more like structured and the work you're doing is structured in the language model which is like make this kernel implementation faster, make this idea work, then maybe it can be. I think it's oversimplification. I push on that a bit in the piece just to really contrast like what you could think a US lab will look like. And I have a few anecdotes like I've heard a US lab paying off a researcher to be quiet about their thing not being in the bottle. Like all of these one-off things are more storytelling devices than anything because like most one-off things don't matter at all. But also Long Before Imploded and like that was because it was described as a Game of Thrones political style environment with all of the VPs vying for influence on showing that their thing made the benchmarks go up and it kind of felt like many many people will tell you that. And we've had the Quen turnover, but it doesn't seem like it was quite the the same type of thing as Llama or like XAI barely exists now. Like there's been some dramatic things in the US with how these companies have kind of come and gone out of the fold.
是的,我有点同意你的观点,但我也想反驳一下。我认为东亚地区显然有一个更僵化、更具竞争性的学术体系,这默认导致学生更遵循官僚和权威的文化。所以我同意他们非常务实,专注于给定的任务。然而,我想知道随着 AI 颠覆教育,事情是否会改变。此外,我今天合作的许多年轻研究人员看起来非常不同。我遇到的很多企业家出生在 80 年代和 90 年代,有些甚至更年轻,在 2000 年代,他们身上有一种自信的光环。他们更有个体主义思想。上海街头的人们穿着非常独特,有夸张的服装,寻求个性化的方式展示自己。所以我想知道这是否会改变。但对于像清华和北大这样的学术机构,它们仍然非常老派。我想说这在西方的一些学术机构可能也是如此。
Yeah, I kind of agree with you, but also I would push back. I think there's obviously a more rigid and competitive academic system which by default in East Asia results in a culture of students following bureaucracy and authority more. So I agree they are very pragmatic, focusing on the task given. However, I wonder if things will change with how AI disrupts education. Also, many young researchers I work with today seem quite different. A lot of entrepreneurs I meet are born in the 80s and 90s, some even younger in the 2000s, and there's an aura of confidence coming from them. They are more individualistically minded. People in Shanghai dress very uniquely with outrageous outfits, seeking individual ways to showcase personality. So I wonder if that will shift. But for academic institutions like Tsinghua and Beida, they are still very old school. I'd say the same might be true for some academic institutions in the West.
让我们回到中国的生态系统。当 DeepSeek V4 出来时,我们私下讨论了我写的一篇文章,说 DeepSeek 如何开始看起来像中国的基础层。一些实验室承认他们资源非常有限,人员有限,实验室很小,最多 100-200 人,资本有限,算力有限。从这个意义上说,生态系统看起来不那么零和,更像是不同玩家优化自己的优势。如果我错了请纠正我,但 DeepSeek 提供了一个基础层,许多实验室迅速跟进并采用他们的工程突破。然后 Julli CI 专注于编码,Minia Max 专注于多模态,等等。字节跳动专注于视频模型,快手仍然是超大规模领域的领导者。所以每个人都在做自己的事情,而不是直接竞争。
Let's go back to China's ecosystem. When DeepSeek V4 came out, we talked about it offline about a piece I wrote saying how DeepSeek is starting to look like a base layer for China. Some labs admitted they have very limited resources, limited people, labs are tiny with 100-200 people max, limited capital, limited compute. In that sense, the ecosystem looks less zero-sum and more like different players optimizing their own strengths. Correct me if I'm wrong, but DeepSeek provides a base layer where many labs quickly follow and adopt their engineering breakthroughs. Then Julli CI focuses on coding, Minia Max on multimodality, etc. ByteDance focuses on video models, and Kuaishou is still the leader in hyperscalers. So everyone is doing their own thing instead of competing directly.
我同意专业化,我认为这是正常的商业演变。你找出自己擅长的领域,机会很多,所以他们追随那个。我只是对 DeepSeek 作为基础层更持怀疑态度,因为我不知道 DeepSeek 在做什么。我们当时在那里时,一些实验室因为 DeepSeek V4 刚出来,说他们看了 DeepSeek 做的事情,但似乎比需要的更复杂。如果你读论文,这个模型里有太多东西。作为研究人员,有些东西看起来有点假,或者依赖于他们的设置,不一定能在其他模型中工作。构建语言模型是路径依赖的:你有你的 GPU、预训练数据集、预期的部署统计等。所以你根据约束做决定。DeepSeek 有这些约束,最终得到了他们的模型。但 Moonshot 和 Jippu 有不同的约束,可能更灵活,他们构建了不同的模型。他们会测试 DeepSeek 的创新,有些会说某个创新没有改进他们的模型。所以这些组织处于不同的发展路径,有核心相似性,比如大型混合专家模型,但很多部分最终不同。这就是为什么我不确定 DeepSeek 是否是基础层。如果是,你会看到中国实验室只做后训练,拿基础模型,然后适应他们的领域。我在考虑创办一个后训练实验室,以及如何更好地组织后训练研究。所以我思考共享基础会是什么。一些实验室在创建基础模型上投入了巨大成本,如果不需要,他们不会这么做。一个实验室告诉我们他们的预训练运行有多长,我惊掉了下巴。那太长了。任何美国公司的顾问都会说你承担了太多风险。如果他们没能从过去的大实验室那里获得那个预训练运行,我不知道公司会不会死。大多数美国公司现在知道,你不希望你的大预训练运行超过几个月,因为风险太大。这表明他们没有那么大的峰值规模集群。预训练时间可以通过更大的整体集群大大缩短,获得更多吞吐量。但如果你的最大集群较小,更难获得吞吐量,所以你会用更长时间。这是算力约束。我认为专业化是真实的,但我不知道 DeepSeek 在做什么。我知道他们现在在融资。我不知道计划是什么。
I agree with the specialization, which I think is normal business evolution. You figure out what you're good at, and there's so much opportunity that they follow that. I'm just more skeptical of DeepSeek as a base because I have no idea what DeepSeek is doing. Some labs when we were there, because DeepSeek V4 had just come out, said they look at the things DeepSeek does but they seem more intricate than needed. If you read the paper, there's just so much going on in this model. As a researcher, some of it seems a little fake or dependent on their setup and not necessarily going to work in other models. Building an LM is path-dependent: you have your GPUs, pre-trained dataset, intended deployment stats, etc. So you make decisions based on your constraints. DeepSeek has these constraints and ends up with their model. But Moonshot and Jippu have different constraints, maybe more flexibility, and they build a different model. They will test DeepSeek's innovations and some will say X innovation doesn't improve their model. So these organizations are on different development paths with core similarities like large mixture of expert models, but many parts end up different. That's why I'm not sure if DeepSeek is a base. If it were, you'd see Chinese labs just do post-training, take the base model, and adapt it to their domain. I think about starting a post-training lab and how to format post-training research better. So I think about what a shared base would be. Some labs put an extreme cost on creating their base model, and if they didn't need to, they wouldn't. One lab told us how long their pre-training run was, and I was jaw-dropped. That's way too long. Any US company adviser would say you're taking too much risk. If they didn't land that pre-training run from one of the past big labs, I don't know if the company would be dead. Most US companies now know you don't want your big pre-training run to be more than a few months because it's too much risk. That's a sign that they don't have as big of a peak-size cluster. Pre-training time can come down a lot with a bigger overall cluster, giving more throughput. But if your biggest cluster is smaller, it's harder to get throughput, so you use it longer. That's a compute constraint. I think specialization is real, but I have no idea what DeepSeek is doing. I know they're raising money now. I don't know what the plan is.
不过没人知道。没人知道。
No one knows though. No one knows.
他们似乎是中文生态系统中唯一没有专长的。
They seem the most without a specialty in the Chinese ecosystem.
对。我觉得他们有点被国有化了,不管是否自愿,因为他们拿了中国政府的钱。他们变得神秘了。他们偏爱中国教育背景的研究人员,这不是秘密。他们从人才到资本到整个栈都非常本土化。
Right. I feel like they've been kind of nationalized, whether willingly or not, because they're taking the Chinese government's money. They've gone secretive. It's not a secret that they prefer Chinese educated researchers. They're keeping it very domestic, from talent to capital to the whole stack.
所以感觉他们在某种程度上被华为化了,因为他们做得好,名声传遍全球,然后不管愿不愿意,他们就成了下一个华为。
So it seems like they're being Huaweied in some ways because they did well and got their name globally, and then by default they're becoming the next Huawei, willingly or not.
我不认为国有化会让你成为其他公司的基础,至少现阶段不是。可能有一些激励,但如果你能通过带动一个团队来推动整个行业,这可能会成为你的 KPI。
I don't think nationalization makes you a base for other companies, at least not at this stage. There could be some incentive, but if you can propel the whole industry by taking one of the team, it could be in your KPI.
协调问题非常困难。本质上,在美国和中国,即使是开放实验室也会分叉开源代码并将其匹配到内部系统。每家公司都这样做,因此所有可能回馈到开源代码并形成更高效基础的改进都没有完成反馈循环。我认为中国可能更接近这一点。如果人们真的投入,以 DeepSeek 作为标准架构,他们分享训练代码和具体方法,从中国经济角度看,这将是一个巨大的胜利,因为你节省了算力乘数。但这太分散、竞争太激烈,难以实现。美国也不会发生,尽管为了让开放模型更接近前沿,这会是更好的方式。美国的开放模型需要一个联盟,但即使有足够的资金组建联盟,也会失败,因为模型会因为太多需求而表现不佳,尽管这是创建共享基础的唯一途径。
The coordination problem is so hard. Essentially, both in the US and China, even open labs fork open source code and match it to their internals. Every company does this, so all the improvements that could go to the open code and form a more efficient base are not completing the feedback loop. I think China could be closer to it. If people really leaned in, with DeepSeek as a standard architecture and they shared their training code and specifics, from a Chinese economic perspective that would be a huge win because you're saving compute multipliers. But it's too decentralized and too competitive to have that happen. It wouldn't happen in the US either, even though for open models to be closer to the frontier it would be better. Open models in the US need a consortium, but there's enough money to make one, yet you fail because the model won't be good due to too many asks on the model, even though that's the only way to create a shared base.
是的。
Yeah.
没有商业理由。
There's not a commercial reason.
好的。如果你必须对每个主要实验室给出高层次的评论,会是什么?比如字节跳动、阿里巴巴、腾讯,如果相关的话,DeepSeek、月之暗面、智谱、MiniMax,还有现在的面壁智能成为其中的一部分……
Okay. So if you had to give a high-level commentary on each of the major labs, what would it be? Like ByteDance, Alibaba, Tencent, if they're relevant, DeepSeek, Moonshot, Zhipu, MiniMax, and now Xiaoi being part of the kind of...
你可能需要提示我多说,但我可以随便聊聊,这挺有意思的。阿里巴巴以云为中心,理解开放模型可以增加平台的使用。所以我会说阿里巴巴非常非常以云为中心。字节跳动的主要特点是其他人都怕他们,而且非常以用户为中心,包括多模态。月之暗面的办公室氛围很好;你会觉得这是中美之间最好的创业公司氛围之一。智谱非常 AGI 导向,令人惊讶地谨慎地对自己被列入实体清单感到兴奋,尽管他们不知道为什么,因为这给他们贴上了重要角色的标签。
You might have to prompt me to say more, but I could just ramble through them, which is kind of fun. Alibaba is cloud-focused, understands that open models can enable more usage of the platform. So I would say Alibaba is very, very cloud-focused. ByteDance is mostly characterized by everybody else being intimidated by them, and very user-focused, including multimodal. Kimi's office vibes were great; it would be one of the best startup vibes you'd visit among US or China. Zhipu is very AGI-pilled, surprisingly cautiously excited about being entity-listed even though they have no idea why, because it stamps them as a big deal.
我认为是因为他们之前与……合作过,这是主要原因。
I think because they previously worked with... that's like the main reason.
是的。或者他们仍然有合作,但那是他们主要的收入来源之一。不幸的是,因为很多实验室是从清华分拆出来的,而清华在北京离政府很近。但问题是,离政府近可能意味着在真正的政府之下有三层代理机构。但人们喜欢将其与接受政府资金联系起来,所以他们会受到怀疑。这非常不幸,但很多公司都被归入这一类。甚至像联想和其他一些中国公司也曾被美国参议员指责接受中国政府资金,但实际上他们的科学家或研究实验室是从政府附属或资助的学术机构分拆出来的。就是这样。
Yeah. Or they still do, but that was like one of their main sources of income. And unfortunately, because a lot of these labs spun out of Tsinghua, and Tsinghua is close to the government in Beijing. But the thing is, when it's close to government, it could mean there are three layers of agency underneath the actual government. But then people like to link it to taking government money, so they are suspicious. It's very unfortunate, but a lot of companies get thrown into that category. Even companies like Lenovo and a few other Chinese companies have been called out by US senators for taking Chinese government money, but really their scientists or research labs spun out of a government-affiliated or funded academic institution. That's what it is.
是的。还有一些,比如面壁智能,对于一个新团队在一家普通公司来说,研究氛围出奇地好。他们似乎做得很好。
Yeah. Some more would be like Xiaoi, surprisingly great research vibes for a new team at a random company. And they seem to be crushing it.
哦,我没见到她。我认为她是你现在最接近明星研究员的人。有明星 CEO 的层级,比如 Dario 和 Sam,类似的情况,但明星研究员……她是最接近这个的。我需要多看一些采访。但她没参加会议。他们似乎在做正确的事情,构建通用模型,还没有专业化。Florian,他帮我写 Interconnects 上关于开放模型的文章,和我绕道去看了 MiniMax,因为我们好奇他们为什么构建这些模型,他们对此非常务实。那是一次不那么光鲜的访问,在一个普通的科技办公室,他们说:‘是的,我们是一个主要的在线平台,我们将在各处使用 LLM,我们需要构建自己的 LLM 并针对我们的产品进行专业化。’惊讶吧,非常务实。我猜中国还有很多这样的公司。
Oh, I didn't get to meet her. I think she's like as close to a star researcher as you have right now. There's the tier of star CEO, like Dario and Sam, the analogies are there, but star researchers like... she's the closest you have to this. I need to watch more interviews. But she wasn't at a meeting. They just seem to be doing the right thing, making general models, no specialization yet. Florian, who helps me write about open models on Interconnects, and I took a detour to see MiniMax because we were wondering why they're building these models, and they were very practical about it. It was a less glamorous visit at a normal tech office, and they said, 'Yeah, we're a major online platform, we're going to use LLMs everywhere, we need to build our own LLM and specialize it to our products.' Surprise, it's very practical-minded. I'm guessing there are many more companies in China like that.
腾讯也是这么说的。因为他们想服务现有消费者,并针对自己的分发和界面或活动循环优化 LLM。
That's what Tencent is saying too. It's because they want to serve their existing consumers and optimize their LLM for their own distribution and interface or activity loop.
是的。我离开后,组里的一些人去了小红书,他们发布了一些语言模型,是多模态数据处理的东西。
Yeah. After I left, some people in the group went to Xiaohongshu, like Red Note, and they've released some language models that are multimodal data processing things.
很多并不令人惊讶。创业公司只是有不同的文化。我之前见过一些 MiniMax 的人。这次旅行我在去 MiniMax 之前就提前离开了。但 MiniMax 很古怪。他们公司有很多女性,这很有趣。他们有产品,可能稍微更注重产品,但公司的古怪之处与西方对他们产品用途的困惑相匹配,也匹配他们更高效的语言模型。
A lot of them are not that surprising. The startups just have different cultures. I have met some MiniMax people before. I left the trip early before MiniMax on this one. But MiniMax was quirky. They have a ton of women in their company, which is very fun. They have products, maybe slightly more product-focused, but the quirkiness of the company matches the western confusion over what their products are doing, and it matches their language models that are a bit more efficient.
他们推出了很多非常面向消费者的应用,对吧?他们之前有像 Xiao 和 Taki 这样的角色陪伴机器人产品。
They came out with a lot of very consumer-focused applications, right? They had like Xiao and Taki, all these character companion bot products before.
是的。
Yeah.
然后我最后去的是蚂蚁集团,它也非常企业化,但没那么激烈,因为我觉得他们把它看作是为自己的产品服务。阿里云则是‘这是我们必须要赢的金矿’。对蚂蚁来说,这比蚂蚁集团本身重要得多。但当你列出这些公司,大概八到十家,它们都相当合理,考虑到公司的年龄和它最擅长的领域。是的,顺便说一句,阿里现在在医疗聊天机器人方面低调地做得最好。我想这说得通,因为每个人都有支付宝,对老年人来说,这可能是除了微信之外他们唯一经常使用的应用。所以它成了默认的医疗咨询应用,这真的很随机,但现在是他们的 niche。
And then the last one I went to was Ant Group, which is also very corporate but in a less intense way because I think they see it as serving their own products. Alibaba Cloud is like 'this is the gold mine we have to win.' It's a much bigger deal for them than Ant Group. But when you list them, maybe eight to ten companies, they're all pretty reasonable with respect to the age of the company and what it does best. Yeah, not to mention, Alibaba is low-key best at medical chatbot right now. I guess it made sense because everyone has access to Alipay, and for seniors, it might be the only application they use regularly aside from WeChat. So it became the default medical consultation app, which is really random but it's their niche now.
是的,我觉得你说得很准。你能在见面后得到这些见解真的很酷。我的意思是,我读关于他们的文章已经很久了。所以很多这些先验知识当它们与你看到的东西吻合时就很容易被证实。
Yeah, I think you're pretty spot on. It's pretty cool that you got those takeaways even just meeting. I mean, I've been reading about them for so long. So it's like a lot of these priors are easy to confirm when they kind of fit with things you have seen.
中国的展厅文化非常有趣,也是软件公司最令人惊讶的事情之一。这很有趣,而且他们显然在吸引西方观众。比如智谱有翻译得很差的周边商品。有些东西……我喜欢其中一些,但在美国可能算是边缘不当的翻译。比如‘ship big go hard’之类的,就是一些非常奇怪的翻译。而且他们在展厅里有实时的 API 统计数据。所以智谱说‘我们每天服务 5.5 万亿个 token’。所有美国公司宣布 token 统计数据时都会被密切关注。我觉得……我知道至少其中一个数字是错的。比如 Fireworks 每天处理 30 或 300 万亿个 token,或者 Together 是那个数字,然后 Fireworks 和 Together 中的一家以及另一家每天处理大约 100 万亿个 token。别把这些当来源,去查一下。最近有一些公开声明,但那是人们第一次获得美国主要基础设施公司的更新。比如,推理是一个巨大的市场。你听不到 Fireworks 的任何消息,因为他们正在努力满足需求,而且他们赚得盆满钵满,因为推理比裸金属好卖得多。所以本质上,推理是销售软件实现来更高效地提供 token,当你为固定模型改进堆栈时,你可以获得更多利润。所以一个模型发布后,你托管它,然后你可以让你的堆栈在该模型上越来越高效。这样你就获得更多利润,并希望使用量增长。这与 GPU 完全不同,GPU 的最佳情况是你锁定一个长期的大承诺。我想就这些。而且能够走进办公室了解他们的 API 很有趣,因为他们还有地理分布,比如中国大概占三分之二,美国大概 20%,剩下的百分比是新加坡、韩国、日本在智谱的 API 上。这很酷。这是我一直想了解的公司信息,但我毫无头绪。我认为我一直想了解的事情之一是,开放模型在美国和中国之外是如何被使用的,以及这个长达数十年的技术扩散过程是否已经开始以任何公司都能衡量的方式启动?我认为还没有人掌握好的数据,但我认为很明显,在某个时候,运行成本低的开放模型将在全球范围内对工业化程度较低的国家、长尾国家产生一些有趣的玩法。我当时想,也许我会走进一家中国开源权重公司的大门,然后得到答案。
The Chinese showroom culture is so interesting and also one of the most surprising things to have for software companies. It's so funny and they're definitely appealing to Western audiences. Like Zhipu had poorly translated merch. It was something like... I love some of it, it would be borderline inappropriate translation in the US. It was like 'ship big go hard' or something, just some really weird translations. And they have live API statistics in their showroom. So Zhipu was like 'we're serving 5.5 trillion tokens a day.' All the US companies are so closely watched for when they announce token statistics. I think something like... I know at least one of these numbers is wrong. It's something like Fireworks does either 30 or 300 trillion tokens a day, or Together for that one, and one of Fireworks and Together and another one is at like 100 trillion tokens a day. Don't take these as source, go look them up. There were some public announcements recently, but those were like the first updates that anyone has on major infra companies in the US. Like, inference is a huge market. You don't hear anything from Fireworks because they're just struggling to meet demand and they're making bank because inference is a much better thing to sell than bare metal. So essentially, inference is selling the software implementation to serve tokens more efficiently, and you could just get more margin when you improve the stack for a fixed model. So a model comes out and you host it, and then you can make your stack more and more efficient on that model. So you just get more margin and hopefully growing usage. And that's way different than GPUs, where the best case is you just lock in a huge commitment for a long term. I think that's all. And just being able to walk into an office and learn about their API is interesting because they also had distribution of geography, which is like China was I don't know like two-thirds, US say 20%, and then the last percent was like Singapore, Korea, Japan on the Zhipu API. Like that's cool. This is what I always want to know about the companies and I have no idea. I think one of the things I always want to know is how are open models being used outside of US and China, and has this decades-long process of technological diffusion started to kick in in a way that any company can measure? I don't think anyone has good data on it yet, but I think it's obvious that at some point open models that are cheap to run are going to have some interesting playbook across the globe for less industrialized countries, the long tail of countries. And I was like, maybe I'll just walk into the front door of a Chinese open-weight company and I'll get my answer.
但实际上,我认为这些实验室的文化,很多因为是由非常年轻、充满激情的人运营的,你会觉得它们商业化程度低得多,或者不那么企业化,至少不那么光鲜。你知道,他们不擅长,你可以说他们在资本市场方面不那么老练,但你也可以说他们只是非常天真、思想开放,对他们正在做的产品充满热情。他们周围的企业护栏较少。
But actually, I think the culture of these labs, a lot of them because they're run by really young, passionate people, you would feel like they're a lot less commercialized or less corporate, or at least less sleek. You know, they're not sophisticated with, you can say they're less sophisticated with say the capital market side of things, but you can also say that they're just really naive and open-minded about and passionate about the product they're working on. Less of a corporate guard rail built around them.
是的。比如智谱有个人在 X 上挺有名,大概有 9000 粉丝。叫 Lou。她走过来打招呼说‘嗨,我是学生,20 岁,我是 X 上的 Lou。’我当时觉得这太搞笑了。很多都是这样。比如……我觉得负责 Moonshot 开发者生态的人就是个刚毕业的女孩,对吧?她整天发搞笑的表情包。她的社交媒体上没有任何过滤。很有趣。
Yeah. It's like one of the people at Zhipu who's known on X is like, I don't know, 9,000 followers. It's like Lou. She came up and said 'hi I'm a student. I'm 20. I'm Lou from X.' And I was like that's hilarious. It's like a lot of that. It's like okay... I think the one that runs Moonshot's developer ecosystem or something is literally a girl fresh out of school, right? And she just posts hilarious memes all day long. There's no filter on her social media. It's funny.
好了,我们跑题了。Nathan,我们需要回到正轨。开源开放权重。为什么?你觉得中国实验室为什么采用或拥抱它,尤其是在你拜访他们之后?是因为他们不得不这样做,就像我们讨论过的,他们因为各种限制而互相依赖,还是你认为哲学驱动力在那个生态系统中实际上更大,或者这是为了长期扩散而进行的更大战略思考?
Okay, we go on these tangents. Nathan, we need to come back on track. Open source open weight. Why? Like, why do you think that Chinese labs are adopting it or embracing it, however you want to put it, especially after visiting them? Is it because they simply have to because what we talked about they are leaning on each other because of all the constraints they have, or do you think the philosophical drive is actually bigger in that ecosystem, or is this a bigger strategic thinking for diffusion in the long run?
我实际上不觉得这在意识形态上有什么特别。我认为当你这样做时,很容易说出意识形态的路线。现在你可以看看扎克伯格。他当时也说了意识形态的路线,然后他停止了。我认为这主要是,第一,在美国生态系统中分发,特别是面向企业,是最高价值的市场,但他们签不了很多企业合同。最接近的好事是像 Cursor 采用 Kimi 的模型,即使 Kimi 没有因此得到报酬,他们也很高兴。这对他们来说是最大的可信度标志,他们将来会想办法销售 token 或其他东西。所以,我认为实际上,第一,影响美国市场的唯一途径是发布这些模型。第二,似乎他们觉得如果发布和分享东西,他们不会损失太多。比如,如果模型是封闭的,他们只会认为影响力会变小。他们会被更少人看到,更少人使用模型,他们实际的付费产品也会被更少采用。这几乎太明显了,因为有所有这些好处,而没有明显的坏处。
I actually don't feel like it's that special ideologically. I think it's easy to say the ideological line when you are doing it. Now you can look at Zuckerberg. It's like he said the ideological line when he was doing it and then he stopped. And I think it's mostly just like for one, distributing within the US ecosystem especially to enterprises is the highest value market and they can't sign many enterprise deals. The closest best thing is things like Cursor adopts Kimi's model, and even if Kimi doesn't get paid for that, they're happy. That's the biggest sign of credibility for them, and they'll figure out selling tokens or whatever in the future. So, I think practically speaking, one, the only way to influence the US market is by releasing these models. And then two, it's just like it seems like they don't feel like they're losing as much if they release and share things. Like if the model was closed, they just think they would get less influence. They would be seen less, less people would use the model, their actual paid offerings would be adopted less. It just seems kind of almost just like so overwhelmingly obvious because there are all these benefits and not as obvious of a drawback.
而且总会有更好的模型,就这样一直下去。但我认为每个科学家……很多美国实验室都反对开源,因为他们不开源也能赚同样多的钱。比如 Anthropic 和 OpenAI 不开源反而赚得更多,他们能赚那么多钱,何必费心去搞一个不赚钱的开源模型呢?所以这其实就是不同层面的影响力问题。谷歌也一样,谷歌赚得盆满钵满。我认为 Meta 会通过在其产品中植入好的 AI 模型来赚大钱。如果他们能整合好,甚至谷歌也可以发布更多模型。除了 Gemini,他们还有很多其他服务需要 AI 商品化,比如云服务等等。Meta 也可以发布他们的模型,但对某些公司来说,这不值得费劲。他们需要达到那么高的收入目标,以至于觉得走法律流程、准备发布太麻烦了。我不知道,也许这有点愤世嫉俗,但我认为微软和 Meta 本可以公开他们最好的模型,因为如果 AI 成为商品层,这对他们是有好处的。但我不指望他们会这么做,因为专注的好处太大了,他们只是觉得没必要这么做。
And there will always be better models and just kind of keep going. But I think every scientist... so many US labs are against it because they can make as much money without it. Like Anthropic and OpenAI make more money by not releasing them, and they could just make so much money that why bother thinking about an open model that doesn't make money. So it's just kind of like different scales of influence. Same with Google. Google's making so much money. I think Meta will make a lot of money by having good AI models in their products. If they get their act together, even Google could release more models. They have so many services other than Gemini that need AI to be commoditized, like the cloud and all of this. Meta could release their models, but it's just not worth the effort for some of them. They need to hit such a high revenue target that they think it's too much of a pain to go through legal and make it ready to release. I don't know, maybe it's a cynical take, but I think Microsoft and Meta could release their best models openly because it's a good benefit if it's a commodity layer. But I don't expect them to because the benefits of focus are so high and they just see it as something they don't have to do.
但最终我们也会看到市场整合,因为现在每个主导国家不可能都有 10 个实验室。
But then eventually we will see some consolidation in the market as well, assuming you can't really have 10 labs in each dominant country right now.
我确实预期会有整合。我认为这可能是一个微妙的文化差异:美国实验室更倾向于接受‘我们很特别,需要快速前进并保持封闭’的说法,而中国实验室则不然。这可能有些道理。这也取决于决策最终由谁做出。我和阿里巴巴做决策的人聊过,但我不能透露他们说的所有话。有些是一对一、不公开的。我不能说这些。在其他实验室,都有一个做决定的人,我猜。我不认为……我觉得那些是我们没接触过的高层领导。所以很难确切知道他们真正的想法。我肯定预期会有整合。我的想法是,我原本以为中国会更快整合,因为资本市场不如美国强劲。但我没有模型来预测这个。我认为你可以建模:你认为收入增长会是多少?他们需要做什么来融资以继续训练更大的模型?算力成本是多少?然后看看潜在的融资,想想哪个国家会先无法完成融资。但还有一件疯狂的事:OpenAI 筹集了 1200 亿美元。你在开玩笑吗?美国的估值现在其他人真的无法理解。在中国,关于你的观点,我一直在写文章。我不知道,腾讯收购一家实验室是合理的。他们有钱,需要能力,而且坦率地说,他们在 LLM 竞争中一直很挣扎。他们的所有许可证都很糟糕。他们发布的所有模型许可证都很糟糕。模型不够好,许可证也很糟糕。所以我觉得像腾讯这样的公司从财务角度优化一下,直接收购一个实验室是合理的,然后实验室也可以利用他们的分发渠道。因为归根结底,在中国,当市场被阿里巴巴、字节跳动和腾讯主导时,他们如何赢得消费者的心智或分发渠道?但这是我的看法。我认为大公司有很多惯性,高层领导做决定,他们也有惯性。我仍然认为苹果最终会花 250 到 500 亿美元收购某个实验室。这不是最坏的事。只是给研究人员戴上金手铐。但我认为现在实验室还有梦想。一些研究人员还有梦想。我和其中几个聊过,他们说:‘不,我们不想那样做。我们想致力于自己的前沿研究。’如果他们想加入大科技公司,他们早就可以去了。所以他们为什么要卖呢?这是研究人员的想法。但正如你所说,我们不知道最高层的一两个人怎么想,尤其是当他们仍然能获得算力和资本的时候。
I do expect consolidation. I think this is potentially a subtle cultural point: US labs are more likely to buy into the 'we're special, we need to go fast to keep it closed' narrative, and Chinese labs are not. There could be something there. That's just who the decisions funnel up to as well. I talked to the Alibaba people that make these decisions. I can't say all the things they say about them. Some of these were one-on-one and off the record. I can't say all these things. At all the other labs, there is a person that makes the call, guessing. I don't expect... I think those are senior leadership that we're not talking to. So it's kind of hard to know exactly what they really think. I definitely expect consolidation. My thinking is that I expected it in China faster because the capital markets aren't as strong as in the US. But I don't have a model for that. I think you can model it: what do you think the revenue growths would be? What do they need to do to raise to keep training bigger models? What is the compute cost? Then you look at the potential raises and think about which country would just not be able to do that raise first. But also it's wild: OpenAI raised $120 billion. Are you kidding me? The valuation in the US is not really understandable by anyone else right now. I think in China, on your point, I've been writing about this. I don't know, it would make sense for Tencent to just buy out one of the labs. They have the money, they need the capabilities, and frankly they've really been struggling to compete with their LLM. All their licenses are so bad. They release all these models that have horrible licenses. They're not that good and the licenses are just horrible. So I feel like it financially makes sense for a company like that to optimize and just buy out a lab, and then the labs can also lean in on their distribution. Because end of the day, how are they going to win consumer mind share or distribution in China right now when it's really just dominated by Alibaba, ByteDance, and Tencent. But that's my spill. When I think big companies have a lot of inertia, and the senior leadership has the call and they can have inertia. I still think Apple ended up just buying some lab for $25 to $50 billion. It's not the worst thing. Just golden handcuffs the researchers. But I think right now the labs still have a dream. Some of the researchers still have a dream. When I spoke to a few of them, they said, 'No, we don't want to do that. We want to commit to our own frontier research.' If they wanted to join one of the big tech, they could have. So why would they want to sell? That's what the researchers think. But to your point, we don't know what the one person or two people at the very top think, especially if they continue to have compute access and capital access.
这也引出了我的一个问题。
It also brings me to a question.
是的,你可以问下一个问题。我不需要打断你。这取决于你对推理的看法。如果这些智能体需要大量推理,我确实认为会形成寡头垄断,而不是垄断市场。那么从财务角度看,两家和四五家拥有优秀模型的大公司有什么区别?如果需求以微妙不同的方式如此之大,这难道不可持续吗?有很多例子是两三家,比如云服务,两三家。但什么阻止了这种情况?
Yeah, you can ask your next question. I don't need to cut you off. It's like it depends on your view of inference. If these agents are just so much inference, I do think it's going to be an oligopoly, not a monopoly style market. And what's the difference financially between two and four or five big companies with great models? Is that actually not sustainable if there's so much demand in subtly different ways? There's a lot of cases where we have two or three, like the cloud, two or three. But what's stopping that?
它们会成为基础设施提供商。
They would be the infrastructure providers.
是的。
Yeah.
是的。它们会利用各自现有的生态系统或分发渠道,为特定用途提供特定模型,这样企业也可以选择最符合他们需求的模型。我能看到这一点。是的。我确实想把话题引到一个更有争议的话题上,那就是蒸馏和模型趋同。你知道,你提出了一个问题:中国模型在结构上是否不同。我们经常听到一些说法,说很多这些实验室比美国实验室落后大约三到六个月或六到九个月。显然有很多噪音或指控,某些美国实验室指责中国实验室在蒸馏它们。
Yeah. And they would kind of lean into each of their existing ecosystems or distribution, whatever you want to call it, and serve certain specific models for specific use, so enterprises can choose what matches their needs the best as well. I can see that. Yeah. I do want to bring the conversation to a more contentious topic which is on distillation and model convergence. You know, you raised a question of whether Chinese models are structurally different. And often we are hearing claims saying a lot of these labs are about three to six months or six to nine months behind US labs. There's obviously a lot of noise or allegations, accusations from certain US labs saying Chinese labs are distilling them.
你实际上怎么看这种指责或这种动态?我最大的未知数,也是影响很大的,就是中国公司到底是在积极尝试破解 API,还是只是作为客户付费使用。如果他们试图破解 API 以获取推理轨迹,从而创建与目标模型相似的推理基础,这与标准的 API 使用方式——只获取模型输出——非常不同,后者是一个更间接的学习过程。我不知道规模有多大。如果更像是他们直接使用 Anthropic 的 API,按预期使用但为了构建竞争模型,我对 Anthropic 并不太同情。他们可以禁止,如果他们想的话。我认为影响只是标准做法。你可以对很多不同模型这样做。如果 Anthropic 提供的证据规模不够大,不足以让我称之为全天候的行业级知识产权盗窃,那么在蒸馏实际发生的情况中肯定存在灰色地带。这就是为什么在政策方面,我想推动人们不要把所有情况都混为一谈。本质上,使用任何 API 端点生成合成数据来训练你的模型都是某种形式的蒸馏。但如果你试图破解模型以获得对训练极其有用的不同行为且不被发现,那就非常不同了。这些是完全不同的行为,但现在都被归入“蒸馏”这个通用术语。我认为这是我最大的问题:学术研究人员和小公司广泛使用蒸馏作为其业务和研究方法的核心。如果美国政府将其视为 AI 生态系统中不可为之事,那主要对小玩家、中美关系和学术界不利。所以这是我的主要担忧,我试图让实验室多说一些。在蒸馏方面,然后是性能,在基准测试上,中国实验室似乎确实落后六到九个月。在通用使用上,我一直觉得闭源模型在难以衡量的方面更好。我在闭源模型是否更好这个问题上反复摇摆。我认为我们尤其会看到 Anthropic 和 OpenAI 在知识工作任务上领先,比如法律、医疗、金融服务,因为我不认为中国实验室会为这些数据付费——所有这些数据都来自每小时收费数百美元的人来标注和创建这些环境。这是一个全新的资本建设。如果训练一个模型需要数十亿美元的数据、算力和人才,我不认为他们有这个能力。所以我认为那里有更大的差距。这很难。这是一个非常有趣的问题。再次,Florian 和帮我的人对此有分歧。这是一条细线:在编码等评估以及许多其他事情上,甚至是中国实验室肯定没有训练过的随机评估,开源模型确实取得了令人难以置信的惊人分数。我认为还存在测试者偏差,我不怎么使用开源模型,而且很难在脑海中确定六个月到九个月前我用 AI 在做什么——我当时甚至没有广泛使用 Claude Code。
How do you actually see that accusation or that dynamic? The biggest unknown I don't have an answer to, which actually has a lot of sway, is how much are the Chinese companies actively trying to hack APIs versus just show up as a customer and pay. If they're trying to hack the APIs normally to get reasoning traces out so that they can create a reasoning foundation similar to the model they're targeting, that's very different from the standard API form, which is just the output of the model—a much less direct process for learning. I don't know the magnitudes. If it's more like they just walk up to the Anthropic API and use it as intended but to build a competitive model, I'm not very sympathetic to Anthropic. They could ban it if they want to. I think the impacts are just standard practice. You can do it with many different models. If the evidence Anthropic provided is not large-scale enough for me to call it industry-wide IP theft 24/7 365, there's definitely some gray area in what actually happens in distillation. That's why on the policy side, I want to push people not to call all of it the same thing. Essentially, using any API endpoint to make synthetic data to train your model is some form of distillation. But it's very different if you're trying to break the model to get a different behavior that's hyper useful for training and not get caught. Those are pretty different actions, and they're all lumped into this common phrase of distillation right now. I think that's my biggest problem: academic researchers and small companies use distillation extensively as the core of their business and research method. If the US government nukes that as something that can be done in the AI ecosystem, it's mostly bad for small players, bad for US-China tensions, and bad for academics. So that's my primary concern, and I'm trying to get the labs to actually say more. On the distillation side and then performance, on benchmarks, it does seem like the Chinese labs tend to be six to nine months behind. When it comes to general use, I've always found the closed models to be better in ways that are hard to measure. I go back and forth on whether closed models are better. I think we will especially see Anthropic and OpenAI pull ahead on knowledge work tasks like legal, healthcare, financial services, because I just don't see the Chinese labs paying for that data—all that data comes from people charging hundreds of dollars an hour to annotate and create these environments. It's a whole new capital buildout. If it's going to be billions of dollars for data, compute, and talent to train a model, I don't think they have that. So I think there is a bigger gap there. It's so hard. This is a very interesting question. Again, Florian and the guy who helps me disagree on it. It's a fine line: on evals like coding and many other things, even random evals that surely the Chinese labs aren't training on, the open models really are genuinely crazy impressive scores. I think there's also a tester bias where I don't use the open models as much, and it's hard to ground in my head what I was doing with AI six to nine months ago—I wasn't even using Claude Code as extensively.
所以,就像……
So, it's like...
我想问题是:到今年年底,我能否使用一个开源模型,而像 Claude Code 这样的东西感觉完全不行?这就是对性能差距的考验。从六月到八月开始,以及是否达到那个点……我不认为开源模型已经达到了。如果所有在 Claude 上花费数十亿美元的公司都说,‘哦,我们可以只花 1%就用 DeepSeek’,那就会成为更大的话题。大公司的 CIO 们——有些公司在员工 token 上的花费比人力成本还多。这些通常是初创公司,但如果真的那么相似,他们会很乐意将 token 成本降到 1%,因为那样就可以使用 10 倍的 token。但我不期望那会发生,我预计像最新的 Claude 和 GPT-5.5 这样的东西——我预计今年会有更多这样的东西。我们会看到我是否正确。这正好处于我们整个世界对它们越来越清晰的中间点,而且它们是长达 18 个月的故事正在展开。我觉得我们正处于性能差距和蒸馏以及了解更多信息的中间阶段。
I guess the question is: at the end of this year, could I use an open model and something like Claude Code didn't feel like it works at all? That's the test on the performance gap. Starting in June to August, and whether or not that hits... I don't think the open models have hit that yet. I think it would be way more of a narrative if all the companies spending billions of dollars on Claude were like, 'Oh, we can spend 1% and just use DeepSeek.' The CIOs at all the big companies—some companies spend more on tokens for employees than on headcount. These are normally startups, but they would happily reduce token cost to 1% if it really was that similar, because then you could just use 10x the tokens. But I don't expect that to happen, and I expect things like the latest Claude and GPT-5.5—I expect more of these things through the year. We'll see if I end up being right. It's both right in the middle of us as a world getting more clarity on them, and they're 18-month-long stories unfolding. I feel like we're just in the middle of performance gap and distillation and learning more.
是的。我觉得想想很有趣——你提到,帮我回忆一下我和其他人的对话——关于蒸馏的观点是,我刚刚和你以前在 Hugging Face 的同事、亚太区负责人王天骄聊过,他说,看,蒸馏的指责其实没有意义,因为我们都在互相蒸馏。就像,我在向你学习,你向我学习——我们都在蒸馏。用这么模糊的术语来指责我们各种行为是不对的。所以你的观点是,我认为技术界了解情况的人实际上希望更清楚地知道什么是灰色地带,什么是非黑即白,什么是不合适或不道德的。而这需要,我认为,整个行业团结起来,真正制定护栏和规则。现在就是这个情况。
Yeah. I think it's interesting to think—you mentioned, help me recall a conversation I had with other people as well—the point on distillation is that I just had a conversation with your former colleague at Hugging Face, lead for APAC, called Wang Tiaojin, and he was just saying, look, the distillation accusations don't really make sense because we're all distilling off of each other as we speak. Even like, you know, I'm learning from you, you learn from me—we're distilling. And it's so vague of a terminology to just use that to accuse us of all these various behaviors. So to your point, I think people in the technical world who understand what's happening actually want more clarity on what is the gray area, what is actually black and white, and not appropriate or unethical. And that needs, I think, the industry to come together to really put together guardrails and rules. Now that's on that.
第二点,关于算力和数据方面。我觉得有个轶事你会感兴趣:今年二月春节前后,我和北京的一位实验室研究员聊过,他们说想要更好的数据但拿不到,因为通常美国实验室会花几千万甚至上亿美元买一套非常冷门或 niche 的数据集,并签独家合同。而中国实验室会等独家期结束,两三个月后以十分之一或二十分之一的价格买同一套数据。所以当他们开始用那套数据做后训练时,就会产生三到六个月或六到九个月的滞后。
Now number two on the compute side and the data side. I think something anecdotally will be interesting to you is that when I spoke to one of the lab researchers in Beijing in February around Chinese New Year, they were saying they want to get better data but they can't because usually American labs would pay tens of millions, if not a hundred million dollars, for a set of very obscure or niche data set with an exclusivity contract. Then the Chinese labs would wait out the exclusivity contract and pay one-tenth or one-twentieth of the price for that same data set two or three months later. So when they start post-training on that data set, that's where the three to six months or six to nine months come in as well.
对,关于这点我想说,美国的数据行业有两方面。一是实验室向数据供应商提出特定数据需求,数据供应商是一个连接人员和实验室的网络。二是数据供应商知道什么重要,所以他们尝试为特定可用人群制作适合 hill climbing 的好数据,但成本更低,因为他们只做一次并期望从中获利。所以可能存在一个流程:一旦 OpenAI 处于前沿并创造了像 deep research 这样的新东西,数据行业就会说‘我们来做一些更便宜的东西来卖’。所以存在时间差。但我也听到了同样的实地情况:他们受到数据行业的负面影响——质量差、没有真正获取渠道、自己做了一些内部工作。这和今天美国数据行业的成熟度相比有很大差距。
Yeah. On that I want to say the data industry in the US has two things. One, the lab asks the data vendor for a specific type of data, and the data vendor is a network that connects people to the lab. The other thing is the data vendors know what's important, so they try to create good data for hill climbing on specific available people, but it's less expensive because they make it once and expect to take margin on it. So there could be a pipeline where once OpenAI is at the cutting edge and creates something new like deep research, the data industry says 'let's make things that are a little bit cheaper to sell.' So there is a time lag. But yeah, I heard the same thing on the ground: they have a negative impact or influence from the data industry—quality is bad, they don't really have access, they do some in-house. That's a very big difference from today, where you have data in the US which is insane.
数据公司非常成熟。是的,它本身就是一个复杂的生态系统。在深入数据之前,我想问你一个问题。最近很多说法认为,Anthropic 和 OpenAI 已经证明预训练的缩放定律仍然成立,尤其是最近的模型。我们讨论过中国方面明显的算力限制,而且随着未来几个月 Blackwell 的缺失,这种限制可能会更加严重。那么在这场竞赛中,如果比较中国和美国,我们是否会看到中国实验室和美国实验室在性能和基准测试上的差距扩大到 12 个月或 24 个月,因为中国实验室在预训练突破方面算力非常受限?
Data companies are so mature. Yeah. It's like its own sophisticated ecosystem. Before we get into data, I actually want to ask you this question. I think recently a lot of the narrative is saying look, Anthropic and OpenAI have proven that pre-training scaling laws continue to hold, especially with the recent models. There's an obvious compute constraint on the China side that we talked about, and it will likely be more amplified with the absence of Blackwells in the coming months. So as we move forward in this race, if you have to put it in China versus US, will we see a wider gap in performance and benchmarks between Chinese labs and US labs, going to 12 months or 24 months, as Chinese labs are very constrained on compute for pre-training breakthroughs?
我认为这更多是关于预训练作为一个可以真正完成的事情——你能预训练一个多大且能完成并服务的模型。所以我认为中国实验室可以训练出类似 GPT-4.5 的巨型模型,但你无法服务它。他们最终会训练一个 2.5 万亿参数的模型并发布,但没多少人——没人能用。他们几乎无法通过 API 提供服务,因为他们没有 Blackwell 或 VL72 机架或任何用于服务这些大模型的机架;他们就是没有足够的数量。所以你能构建的模型和实际有用的模型之间存在差异。一些中国实验室明确表示我们不需要发布巨型模型,因为没人会用,而开源模型最终通过 API 提供服务。所以市场上可能会有一些细分。但我确实认为推理以及你服务客户所需的经济资源正在成为决定构建什么模型的因素。这就是为什么我认为差距会继续扩大。所有迹象都表明 GPT-5.5 会是一个更大的模型,我不认为这会停止。其经济学基础很简单:你需要一定的规模才能有利润来支持研究,因为你不能永远进行这些荒谬的融资轮次。我认为 OpenAI、Anthropic 和 Google 是唯一拥有足够 AI 使用量来继续沿着缩放定律前进到另一个 10 倍训练算力的公司,这需要惊人的模型投资。所以这就是为什么我认为当经济市场融资放缓时,差距会显现出来——这三巨头之间的模型差距会大得多。这就是我的预测:情况会变得不同。这些实验室无法融资,他们上市,无法通过付费服务产生更多收入,然后就看有多少训练算力可以分配或不能分配。
I think it's more about pre-training as a thing that you could actually finish—how big can you pre-train a model that you can finish and serve. So I think Chinese labs could train models that look like GPT-4.5, which is a giant model, but you can't serve it. They would end up training a model that is two and a half trillion parameters and release it, but not many people—no one can use it. They can barely serve it on their API because they don't have Blackwell or VL72 racks or whatever racks that are serving these large models; they just don't have the quantity. So there's a difference between models you can build and models that are actually useful. Some Chinese labs definitely say we don't need to release the gigantic models because nobody is going to use them, and open-weight models end up being served via API. So there might be some segmentation in that market. But I do think that inference and the amount of economic resources you have to serve your customers is becoming a thing that dictates what models are built. That's why I think the gap will continue to rise. All signs point to GPT-5.5 being a bigger model, and I don't expect that to stop. The economics of it is just the basics: you need a certain volume to have the margin to support the research because you can't keep raising these ridiculous rounds forever. I think OpenAI, Anthropic, and Google are the only people with that AI usage volume to keep marching down the scaling laws to another 10x of training compute, which is just mind-boggling amounts of investment in a model. So that's why I think it will show up when the economic markets slow for fundraising—the model gap between these big three will show a lot more. So that's my prediction: when things will look different. It's like these labs can't fundraise, they go public, they can't generate revenue more on their paid services, and then it's just about how much training compute can be allocated or can't be allocated.
所以我基本同意你的看法。你认为未来几个月差距会更大。那么什么可以弥补呢?比如国产芯片或更好的数据?为什么有时人们认为中国有非常强大的数据生态系统或数据产品,但实际上中国的数据供应商生态系统非常薄弱?数据方面我不清楚,但国产芯片可能帮助的方式是:如果华为芯片在推理方面没问题,并且有足够的数量来支持推理经济,从而反哺收入。我的看法是,他们就是没有足够的芯片数量,尤其是分散到那么多公司。本质上,华为芯片的总算力被分配到各个地方,根本不够大。像字节跳动和阿里巴巴这样拥有离岸数据中心的公司可能能坚持更久,因为他们能通过离岸方式获得英伟达算力,而且已经持续很久了。也许这能稳定生态系统。我们看看像 Kimi 和智谱这样的 AI 初创公司最终会怎么做,因为如果他们集中资源,可以多撑一年——如果他们都联合起来,能再提升一个数量级,但我不认为他们会这么做。
So I generally agree with what you said. We'll see a bigger gap you think in the coming months. Then what can make up for that? Like domestic chips or better data? And why is it that sometimes people assume China has a very strong data ecosystem or data products, but actually the data vendor ecosystem is very weak in China? I don't know on the data side, but the way that domestic chips could help is that if Huawei chips are fine for inference and if they have sufficient volume to support the inference economics, which then trickles back into revenue. My read is that they just don't have the volume of the chips, especially spread out across the amount of companies. Essentially, the total flops of Huawei chips produced are going to all these different places; it's just not big enough. It could be something like ByteDance and Alibaba with offshore data centers can keep up a lot longer because they have access to Nvidia compute and have for a long time through offshoring. Maybe that stabilizes the ecosystem. We'll see what the AI startups like Kimi and Zhipu end up doing, because if they pull resources together, they could last an extra year—you get another order of magnitude if they all pull together, but I don't see them doing it.
但这就是我们刚刚讨论的问题,对吧?比如 MiniMax 和 DeepSeek,如果需要离岸数据中心,他们怎么可能与超大规模公司竞争?而且智谱被列入实体清单也无济于事。他们获取资源不会容易。
But that's the thing we're just talking about, right? Like MiniMax and DeepSeek, how can they possibly compete with the hyperscalers at this point if you need offshore data centers? And the fact that Zhipu is on the entity list doesn't help. It's not going to be easy for them to access.
我认为他们不能。我认为他们不会。我认为人性会让他们不合作。他们只会做更小的事情。他们只会拥有不同的成功企业。
I think they can't. I think they won't. I think human nature will make it so they won't collaborate. They'll just do something smaller. They'll just have successful businesses that are different.
所以你写过类似的话,不是什么秘密,但每个人都想要英伟达的芯片。他们想要,但不知道怎么得到。
So you wrote something like, nothing a secret, but everyone wants Nvidia chips. They want it. They don't know how to get it.
唯一能用于训练的东西,所有模型都在英伟达上训练。我不相信 DeepSeek 的宣传说它是在华为上训练的。唯一在华为上训练的模型是……声称是在华为上推理,但推理在华为上并不好用。每个实验室都说推理在华为上可行。那些没有实际推理需求的实验室会说,我们被告知要买华为,所以我们买了,但不用。这就像早期研究实验室说我们没有推理需求,也不需要华为。任何有实际模型使用的公司都搞清楚了如何在华为上运行推理,这正如黄仁勋所说会发生的事,但令人惊讶。
The only thing that works for training, all the models are trained on Nvidia. I don't believe the DeepSeek propaganda that it's trained on Huawei. The only models that are trained on Huawei are... claim that it was inference on Huawei, not inference on Huawei works. Every lab is like inference on Huawei works. The labs that don't have meaningful inference are like we are told to get Huawei, so we buy them, but we don't use them. Which is like earlier research labs are like we don't have any inference and we don't have a need for Huawei. Any company that has meaningful use of their models has figured out how to run them on Huawei for inference, which to Jensen's credit is like it's happening what he said was going to happen, but it's surprising.
是的。我不太明白他为什么因此招致那么多仇恨,因为即使不考虑你的政治立场,他说的在逻辑上其实很有道理:如果你不卖给他们我们产品的低配版,他们就会自己搞出同样低配的版本。
Yeah. It wasn't like I don't actually understand why he got so much hate for it because even without your political stance it's like what he said actually made sense logically by saying if you don't sell them the kind of crappier versions of what we have they will have an equally crappy version.
我认为他们会两者都买。我认为两者都买是事实,比如你要卖给中国多少英伟达芯片才能让他们停止购买英伟达或华为?因为华为几乎肯定便宜得多,英伟达的利润率太高了。所以他们什么时候才会停止购买呢?
I think they would buy both. I think buying both is actually true like the amount of Nvidia chips that you would have to sell to China for them to stop buying Nvidia or stop buying Huawei because Huawei is almost surely way cheaper because Nvidia margins are insane. So it's like when would they actually stop buying?
而且你必须继续,你能把一切重新路由吗?开发者生态不存在,这是黄仁勋的观点,对吧?或者习惯不存在,所以我认为……
And you have to go on, can you have to reroute everything back on, you know, you have to the developer ecosystem is not there though that's Jensen's point right or the habits are not there so I think...
但我是说他们也会用华为。我认为他们供应如此受限,他们会两者都用。我的意思是,Anthropic 什么都用。美国很多公司都会用多平台。比如 Meta 是 AMD 的大买家。需求如此之高,任何在几代内可能用于模型的芯片都非常有价值。而你能在任何华为芯片上运行一个相当大的模型,这对华为来说是一个重大突破。我不知道他们能否快速生产出那么多芯片并扩大规模,尤其是当他们试图转向更小制程时。这是标准的半导体辩论。但问题是华为能否扩大生产?我认为这是唯一的问题。如果华为能设法扩大生产,黄仁勋就会显得非常正确。如果华为不能扩大生产,黄仁勋会看起来有点像个疯子,但这不在他的掌控之中。
But I'm saying they would also use Huawei. I think they are so supply limited they would use both. I mean like Anthropic uses everything. Like a lot of companies in the US will use multiplatform. Like Meta is a huge buyer of AMD. Like it's just like demand is so high that any chip that is potentially viable on the models within a few generations is very valuable. And the fact that you can run a reasonably large model on any Huawei chip is a big line cross for Huawei. I don't know if they could produce the volume of chips and scale that quickly, especially as they try to move to lower nodes. That's like the standard semi-debate. But that's like the question is like can Huawei scale production? That I think that's like the only question. And if Huawei can manage to scale production, Jensen will just look really right. And if Huawei can't scale production, Jensen will look a little bit like a lunatic, but it'll be outside of his hands.
我们不太清楚这次行程发生了什么。似乎在这个大型特朗普代表团之后,没有什么实质性的进展。
And we don't really know what happened during this trip. It seemed like nothing really substantial really happened after this big Trump delegation.
看起来更像是一次高调的旅游行程,而不是真正的交易行程。
It seems like it was more like a high-profile tourism trip versus an actual deal trip.
好的。我想问你一个你写过的小众话题,不是你通常写的那种。是关于 SaaS 方面的,但你说了……好吧。有一个常见的论点:中国难以从 AI 中盈利,因为他们不愿意为企业软件付费。你知道,我们看了很多中国如何尝试从消费者 AI 中盈利,但显然这还没有被证实。在你的文章中,你反驳了这个说法,并指出 SaaS 支出与云或推理支出之间存在区别。请谈谈你对这种生态系统的看法,以及中国 AI 实验室在赚钱方式上可能与美国 AI 实验室有何不同。
Okay. I want to ask you on something you wrote about a bit niche, like not something you usually write about. It's on the SaaS side of things, but you said that. Okay. So there's a common argument that China struggled to monetize AI because they're unwilling to pay for enterprise software. You know, we've looked a lot at how China tries to monetize on consumer AI, but clearly that's not really been proven yet. In your piece, you push back on the claim and say that there's a distinction between SaaS spend and cloud or inference spend. Tell us about what you think about that kind of ecosystem and how Chinese AI labs are trying to make money maybe a bit differently from American AI labs.
我不知道是否一定不同,但我问了很多研究人员,他们说每个人都会尝试新的 AI 工具,如果不喜欢就停止使用。如果喜欢,就会在消费者端继续使用。比如 OpenClaw 就是一个例子。很多人试过。我打赌在中国和美国一样有很多人尝试,但消费者很快会采纳和尝试新东西,但如果不能真正服务他们,就不会坚持。而在企业端,云服务确实存在,数字服务非常庞大,他们基本上认为 AI 模型赚钱的跑道更长,属于这一类。他们都使用编码智能体。他们都用 Claude。这很有趣。他们都很喜欢 Claude,没人提 Codex,而在西方媒体中 Claude 对 Codex 是个大事。他们都用 Claude,这显然是一个付费服务。所以我认为有一些类似的事情是论点的裂缝,我预计 AI 模型会被视为云的一部分,但可能正是它改变了一些预期,因为它如此具有变革性,竞争如此激烈,可以被视为一个阶段转变。
I don't know if it's necessarily different, but I ask a lot of researchers about this, and they say that everybody is trying the new AI tools when they come out and if they don't like them, they stop using them. If they like them, they keep using them on the consumer side. So something like OpenClaw would be an example. Like tons of people tried it. I bet lots of them turn in China just like in the US, but the consumers are very quick to adopt and try new things, but won't stick if it's not actually serving them. And then the enterprise is like there's definitely cloud that exists like digital services are gigantic and they're essentially think that there's more runway for making money on AI models that falls into that and like they all use coding agents. They all use Claude. It's a hilarious thing. They're all very Claude, no mention of Codex where in the Western media Claude versus Codex is this whole thing. They all use Claude and that is obviously a paid service. So I think there's some things like that that are cracks in the argument and that I expect AI models to be seen as a bit of cloud but potentially that it is the thing that changes some of the expectations where it's just so transformative because they're so competitive and that it could be seen as a bit of a phase shift.
是的。我认为这是一个代际转变、阶段转变,而且实际上最近 Dobbala 提高了价格,我想是在 sea dance 使用等方面。这就像转向试图抓住专业消费者市场。你可以说街上的普通大叔可能仍然不愿意为消费者应用付费。但我认为在中国有更多的专业消费者市场份额可以捕获。也许不完全是企业市场。
Yeah. And I think it's a generational shift, phase shift, but also actually recently Dobbala raised their prices on I think sea dance usage and whatnot. And it's like a shift into trying to capture the prosumer market. You can say the average uncle NT maybe on the street still don't want to pay for an app for a consumer app. But I think there's more prosumer market share that could be captured in China. Maybe not fully enterprise either.
我想问你关于政府角色和地缘政治的问题。我知道有一种常见的说法,人们通常认为中国的 AI 实验室得到了大量补贴。实际上,我三月份在旧金山时,和几个投资者(主要是公共投资者)共进晚餐,有一个人问我,他说,嘿,所有实验室基本上都是由政府补贴的吗?我说,哦,绝对不是,你知道大多数都不是,坦白说他们甚至不想拿政府的钱。他很难理解这一点,因为我认为误解在于所有中国实验室或中国科技都是由政府资助的。有点像我们之前说的,任何与政府机构的关联都被默认意味着得到了支持。首先,我甚至不知道政府是否有那么多钱可以给。
I want to ask you about government roles and geopolitics. I know there is a common narrative that usually people assume Chinese AI labs are heavily subsidized. Actually when I was in San Fran in March I was at a dinner with a couple investors mostly public investors and one guy asked me he said hey like are all labs just basically subsidized by the government. I was like oh like definitely not you know majority of them are not if not they frankly don't want to take money from the government. And it was really like hard for him to understand that because I think the misconception is all Chinese labs or Chinese tech are just funded by the government. Kind of to our point earlier where like any affiliation to any government agency just by default is assumed to be therefore backed. First of all, the government I don't even know if they have that much money to give out.
第二,我不认为竞争是那样运作的,对吧?那么你对这一切有什么看法?
Number two, I don't think that's how competition works, right? So what's your thought on all of this?
这更像是地方政府试图帮助企业获得办公室、人才之类的东西。我不知道地方政府能做什么。在北京有北京智源人工智能研究院之类的,那是一个由北京某个区资助的真正的研究机构。我想,好吧,美国也可以这样做,但远没有蚂蚁集团那种风格——政府在投资轮中持有主要股权。而且我认为也许 Kimi 的最新一轮融资有政府背景的风险投资参与。我不知道这种中间机构是如何运作的。所以我仍然认为这是非常间接的。因为政府体系在不同层级之间竞争激烈,每个层级都在竞争帮助企业,但他们没有大把现金来购买 GPU。坦率地说,他们不知道自己在做什么。有一半的时间,我认为这个论点——就像你刚才说的——北京的海淀区或朝阳区会资助一个研究院,该研究院将努力帮助所谓的走向 AGI,但现实是他们在努力遵循这个高层次的 KPI,即'让我们实现 AI',他们只想在报告中写道他们资助了 AI 相关的东西,这样他们就完成了指标。我不认为这像人们假设的那样亲力亲为。如果你读过——大多数美国科技人士没有读过《苹果在中国》或《急速》——你只需要读这些书,了解一下科技与中国的接口,理解他们也对 AI 感到兴奋,然后你就会明白这是一个混乱的涓滴过程。而中国政府——如果他们国有化一个实验室,那会非常明显,就像在美国一样明显。这并没有发生。
It seemed more like a provincial government trying to help the companies do stuff like get offices, get talent. I don't know what the provincial government can do. In Beijing there's Beijing Academy for AI or whatever, which is a real research institute funded by a certain neighborhood in Beijing. I thought, okay, the US could do that, but much less of the Ant Group style thing where government takes a major ownership stake in an investment round. And I think maybe Kimi's latest round had mentions of government-backed VCs. I don't know how that intermediary works. So I still think it's very indirect. Because the government system is so competitive across the different layers, each of those layers are competing to help the companies, but they don't have piles of cash sitting around to buy GPUs. And frankly, they don't know what they're doing. Half the time, I think this argument—like what you just said—Haidian district or Chaoyang district of Beijing will fund an academy, and the academy will be in the effort to help so-called go towards AGI, but the reality is they're trying to follow this high-level KPI of 'let's make AI happen,' and all they want to do is write in their report that they funded something about AI, so they've hit their quota. I don't think it's as hands-on as people assume. If you read—most US tech people haven't read 'Apple in China' or 'Breakneck'—all you need to do is read these books and learn a little bit about the interface between tech and China, and understand that they are also hyped about AI, and then you'll kind of understand that it's a messy trickle-down process. And the Chinese government—it would be very obvious if they were nationalizing a lab, as obvious as if it were in the US. That hasn't happened.
是的。那么最后,美国 AI 生态系统目前对中国的看法、对中国 AI 的看法,与你实地所见之间最大的脱节是什么?还有什么是我们今天没有谈到但你想分享的?
Yeah. So to close, what is the biggest disconnect between how the US AI ecosystem right now thinks of China, Chinese AI, and what you saw on the ground? And what is something you think we didn't touch on today that you want to share?
这是每个人都会问我的问题。他们通常在我下飞机时问的第一件事就是:'有什么大事?' 而我觉得,没有什么特别令人震惊的。我认为很多人只是没有读过关于科技如何与政府互动的基本书籍,不了解这些事情,或者他们听到的是非常地缘政治的叙述,针对政府体系的顶层以及美国如何参与。这带来了很多碎片。我的意思是,Anthropic 推动非常激进的中国叙事,而 Anthropic 是一家在科技界备受关注的公司。只是在美国生态系统中,大多数人没有花时间在这上面,也没有深入研究。我认为就是这样。所以,我没有什么令人震惊的事情。只是鼓励人们去做一些了解是好的,因为这些动态影响着像中国开源模型这样非常有影响力的事情,而现在硅谷正在构建 AI。所以这对很多人来说都很重要,但他们不研究他们为什么会这样做的原因。他们只是认为,'哦,它在这里,我不需要考虑中国。'
This is the thing everybody asks me. They normally ask me the first thing I get off the plane: 'What's the big thing?' And it's like, I don't think there's anything that shocking. I think many people just haven't read basic books about how tech has interfaced with the government and know these things, or they hear narratives that are very geopolitical, which is targeting the top end of the government system and how the US engages. And there's a lot of shrapnel from that. I mean, Anthropic pushes very aggressive Chinese narratives, and Anthropic is a very followed company in tech. It's just that most people don't spend time on this in the US ecosystem and just don't go deep on it. And I think that's it. So it's like, I don't have anything shocking. It's just good to encourage people to do some of that because these dynamics impact things like the Chinese open models are really influential, and now Silicon Valley is building AI. So it matters to a lot of people, but they don't study the causes of why they might do this. They just think, 'Oh, it's here, I don't need to think about China.'
是的。而且我认为,我不知道是什么原因。所以甚至比尔·格利在他的一些公开露面中也说,似乎中国的研究人员、科技人士、CEO 等,对美国领导人和思想领袖、科技领袖、商业领袖更加了解或关注得更紧密,而不是反过来。这其中有某种原因。我不知道是不是更容易忽视它,或者更容易不必学习新东西。
Yeah. And I think, I don't know what it is. So even Bill Gurley was saying in a couple of his public appearances that it seems like Chinese researchers, tech people, CEOs, whoever, are a lot more aware or following more closely of US leaders and thought leaders, tech leaders, business leaders than vice versa. And there's something about that. I don't know if it's just easier to dismiss it or easier to not have to learn something new.
我认为这是美国文化。美国文化非常封闭。
I think it's American culture. American culture is very insular.
作为一个加拿大人,我有同感。你说得对。
As a Canadian, I have that. You said it.
是的,这很搞笑。美国文化很荒谬。太荒谬了。
Yeah, it's hilarious. American culture is ridiculous. It's so ridiculous.
但我认为,看,我要无耻地自我推销一下。作为一个华裔加拿大人,我在 AI Pro 这里的人生目标就是弥合这个差距。归根结底,我认为正如你所说,顶层会有地缘政治的叙述和言论,但对于普通人甚至只是建设者、科技人士等,了解另一边正在发生的事情,停止疏远它或停止把它说得那么不同,可能对每个人都有好处。因为我认为通过我们的对话,真的只是想说,看,很多方面是如此相似,但很多方面又略有不同,但这种不同并不是政府指令与也许是文化差异或资源限制差异,尤其是在构建技术方面。但这是我的看法。最后一个问题,这是我在节目中问每个人的:你持有一个什么不同的观点?给我一些疯狂的东西。
But I think, look, I'll shamelessly self-plug here. As a Chinese Canadian, my life goal here with AI Pro is really just to bridge that gap. At the end of the day, I think to your point, there's going to be geopolitical narratives and rhetoric at the very top, but for the average person or even for just builders, tech people, whatever, it's probably in everyone's benefit to understand what's happening on the other side and stop alienating it or stop making it as if it's so different. Because I think through our conversation, it's really just to say, look, so much of it is so similar, but so much of it is slightly different, but the difference is not really a government mandate versus maybe a cultural difference or a resource constraint difference, especially in building technology. But that's kind of my view. One last question for you, which is a question I ask everyone on the show: what is one differential view you hold? Throw me something crazy.
我不知道。我一直有点是一个开源模型的悲观者,尽管我在它们之上构建。我只是认为这太不可持续了,而且通过构建闭源软件可以赚很多钱,所以我一直对开源模型满足需求的前景感到悲观。
I don't know. I've always been kind of an open models doomer, even though I build on them. I just think it's so unsustainable, and there's so much money to be made with building closed software that I'm constantly doomy about the prospects of open models satisfying.
像 Hugging Face 这样的公司实际上如何赚钱?
How does a company like Hugging Face actually make money?
我不知道。你可以查一下他们实际赚了多少钱。不幸的是,并不多。
I don't know. You can look into how much money they actually make. It's not very much, unfortunately.
是的,我认为这就是我们所处的资本方面的不幸现实。尽管它激励了竞争和突破,但它对我们刚才讨论的事情没有帮助。好了,Nathan,非常感谢你的时间。非常感谢你的见解和分享。
Yeah, I think that's the unfortunate reality of the capital side we live in. As much as it incentivizes competition and breakthroughs, it doesn't help with what we just talked about earlier. All right, Nathan, thank you so much for your time. Really appreciate your insights and your sharing.
是的,谢谢你邀请我。嘿,很高兴见到你。
Yeah, thanks for having me. Hey, good to see you.
感谢收听 AI Pro 的不同理解。我是主持人 Grace Show。如果你喜欢这期节目,请在 Spotify、YouTube 或任何你收听播客的地方关注和点赞。感谢收听。下次见。
Thank you for listening to AI Pro's Different Understanding. This is your host, Grace Show. If you enjoyed the episode, please give it a follow and a thumbs up on Spotify, YouTube, wherever you get your podcast. Thanks for listening. See you next time.