AI 人才争夺战:汤品外交与招聘策略

The AI Talent War: Soup Diplomacy and Recruiting Strategies

马克·陈 Mark Chen · Core Memory 播客 · 2025-09-01 · 约 98 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenAI 首席研究官讨论与 Meta 的激烈人才争夺战,包括扎克伯格送汤的轶事,以及 OpenAI 为何不进行等额薪资匹配。

OpenAI's Chief Research Officer discusses the aggressive talent war with Meta, including Zuckerberg's soup deliveries and why OpenAI doesn't counter offer dollar for dollar.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 35)

全文 · Full transcript(中英对照)

招聘战 Recruitment Wars

Host

关于人才争夺战。这显然引起了广泛关注,而且看起来 Meta 相当激进。这种针锋相对具体是怎样的?我们处于什么阶段?

On the recruitment wars. I mean, this got a lot of attention clearly and it looked like Meta was quite aggressive. What exactly does this tit for tat look like? What stage are we at?

Mark Chen

是的,人才库是存在的,对吧?每个人都知道他们是谁。我认为许多公司已经意识到,打造一个优秀 AI 实验室的关键要素之一——不是唯一重要的要素,但关键要素之一——就是获得最优秀的人才。Meta 一直在积极采用这一策略,这并不令人意外。我们并没有坐视不管。我其实想从 OpenAI 的角度来讲述这个故事。我认为媒体上有很多说法,比如人才单向流向 Meta。但据我观察,Meta 挖走了很多人,但并不成功。举个例子,在我的直接下属中,在他们从 OpenAI 挖人之前,他们找过我一半的直接下属,但所有人都拒绝了。当然,如果他们每年有大约 100 亿美元的资金用于人才,他们总会挖到一些人。所以我实际上觉得我们在保护顶尖人才方面做得相当不错。看到事态不断升级,其实挺有趣也挺好玩的。一些有趣的故事是,扎克伯格真的亲自给想从我们这里挖的人送汤。我觉得他是亲手做的汤。当时这让我很震惊,但随着时间的推移,我逐渐意识到这些方法在某种程度上是有效的。我也给从 Meta 挖来的人送过汤。

Yeah, I mean there is a pool of talent, right? And everyone kind of knows who they are. And I think many companies have realized that one of the key ingredients, not the only important ingredient but one of the key ingredients to building a great AI lab is to get the best talent. And I think not a surprise that Meta has been aggressively employing this strategy. We haven't sat back idly. I actually want to tell this story from OpenAI's point of view. I think that a lot has been made in the media of, oh, there's this unidirectional flow of talent over to Meta. But the way that I've seen it is Meta has gone after a lot of people quite unsuccessfully. So just to give you context, within my staff, within my direct reports, before they hired anyone from OpenAI, I think they went after half of my direct reports and they all declined. And of course, if they have something like $10 billion of capital per year to deploy towards talent, they're going to get someone. So I actually feel like we've been fairly good about protecting our top talent. And it's been kind of interesting and fun to see it escalate over time. Some interesting stories here are Zuck actually went and hand-delivered soup to people that he was trying to recruit from us. I think he hand-cooked the soup. It was shocking to me at the time, but over time I've kind of updated towards these things can be effective in their own way. And I've also delivered soup to people that we've been recruiting from Meta.

Host

你在搞汤计数。

You're doing a soup counting.

Mark Chen

我在想,如果我有一次外出活动,下一次员工外出活动,我会带他们去上烹饪课。

I've thought of, if I had an offsite, the next offsite for my staff, I'm going to take them to a cooking class.

Host

是的,就是这样……但我确实认为我在招聘方面学到了一些东西。

And yeah, I mean, it's just been... But I do think there's something I've learned about recruiting.

Host

你的汤是你自己做的吗?

Did you cook your soup?

Mark Chen

呃,最好弄个米其林级别的汤。你懂我的意思吧?

Uh, it's better if you get like Michelin star soup. You know what I mean?

Host

是的。不,不,不。我觉得 Dejo 非常非常好,可能比我做的任何汤都好。但确实,我在如何积极争取顶尖人才方面学到了一些东西。真正让我受到鼓舞的是,在 OpenAI,即使是那些跳槽到 Meta 的人,我也没有听到任何人说 AGI 会首先在 Meta 实现。每个人都对 OpenAI 的研究项目充满信心。我向我的员工和整个研究团队明确表示,我们不会与 Meta 进行美元对美元的竞争。在 Meta 提供的薪酬倍数之下,人们仍然非常乐意留在 OpenAI,这让我深信,人们真的相信未来的发展,相信我们会成功。

Yeah. No, no, no. I think Dejo is very, very good and probably better than any soup I could cook. But yeah, I do think there is something I've learned about how to go aggressively after top talent. And I think the thing I've been actually very inspired by is that at OpenAI, even among people who have left for Meta, I haven't heard anyone say AGI is going to be developed at Meta first. Everyone is very confident in the research program at OpenAI. And one thing that I've made very clear to my staff, to the whole research org, is we don't counter dollar for dollar with Meta. And the multiples below what Meta is offering that people are very happy to stay at OpenAI gives me so much conviction that people really believe in the upside and believe that we're going to do it.

Host

嗯,还有 Alex,Alex 的事。他以前是数学竞赛选手……

Well, and Alex, Alex thing. He used to be one of the math competitors...

Mark Chen

是的。是的。

Yeah. Yeah.

Host

我肯定你们一起玩过。

I'm sure you guys hung out.

Mark Chen

是的,我和 Alex 一起玩过几次,但现在不常联系了。是的。

Yeah. I mean, I have hung out with Alex a handful of times, but we don't do much anymore. Yeah. I mean, yeah.

Host

为什么汤成了标志性事件?就是……

Why did soup become the thing? It was just...

Mark Chen

我不知道。有汤,有花,你能想到的任何东西都有,但我不知道。我觉得生活就是一场冒险。我顺应了这个梗。

I don't know. There's been soup, there's been flowers, there's been anything you can think of under the sun, but I don't know. I think life's an adventure. I play into the meme.

Host

是的。是的。

Yeah. Yeah.

Host

你在思考时有没有用到什么扑克策略?

Is there any poker strategy to employ as you're thinking?

Mark Chen

嗯,再说一次,我认为这又回到了我之前关于媒体叙事的说法。游戏规则不是留住组织中的每一个人。而是信任我们培养人才的渠道,了解哪些是我们需要留住的关键人物,并留住他们。我认为我们在这方面做得非常出色。

Well, again, I think it really goes back to what I've said about the media narrative. The game is not to retain every single person in the org. It's to trust in this pipeline that we have for developing talent and to understand who the key people we need to keep are and to keep those. And I think we've done a phenomenal job at that.

介绍与角色 Introduction and Role

Host

今天我们有一个特别的节目。我很兴奋。OpenAI 的 Mark Chen 来了。他是首席研究官。我过去几年认识了他。非常感谢你的到来。

We have a special treat today. I'm excited. Mark Chen is here from OpenAI. He's the chief research officer. He's somebody I've gotten to know over the last couple of years. Thank you so much for coming.

Mark Chen

不,认识你这么久真是太好了。

No, it's been great to know you for so long.

Host

我觉得世界上只有少数人在从事这个非常重要的项目,而你就在最前沿。所以能有这个机会聊天真是太酷了。

I feel like there's a handful of people in this world working on this very important project and I mean you're right at the top of it. So it's so cool to have a chance to chat.

Mark Chen

是的。谢谢邀请我。

Yeah. Thanks for having me on.

Host

这是我的荣幸。我想和你聊很多事情,因为就像我们说的,过去几年我认识了你。我想让人们更多地了解你的背景,但我也知道会有 AI 爱好者希望我们深入探讨一些话题。所以我们会尽量涵盖所有内容。我想先让大家感受一下你的工作,在我脑海中——如果我理解有误,请纠正我——Sam 一直非常投入研究。他是老板,处于食物链的顶端。但你和 Jakob 一起合作,塑造 OpenAI 的研究方向。然后你的角色还有一个额外部分,就是决定哪些算力分配给哪些项目。所以你基本上要规划 OpenAI 的前进方向,以及实现目标的具体机制。这对我来说总是一份可怕的工作,因为我想到人们会想尽一切办法从你这里获取 GPU。

It's my pleasure. And I mean there's a bunch of things that I want to talk to you about because I've gotten to know you like we said over those last couple of years. I want people to know a bit more about your biography, but I also know there's going to be AI enthusiasts who want us to go deep on a couple things. So we'll try to do everything. I wanted to start just by giving people a feel for your job, which in my head, and you just correct me if I get any of this wrong, but you're, you know, Sam has been, he's really into research. He's the boss. He's kind of at the top of the food chain. But then you and Jakob are working together to shape OpenAI's research direction. And then you're in this additional part of this role is actually deciding which compute goes where onto these projects. So you kind of have to chart where OpenAI is heading and then the mechanics of how you're going to get there. And this always strikes me as a horrible job because I picture people doing everything in their power to get GPUs from you.

Mark Chen

确实如此。人们非常有创意,想方设法通过幕后交易来获取他们需要的 GPU。但没错,这确实是工作的一大部分,对吧?确定研究组织的优先事项,并对执行负责。所以针对第一点,我和 Jakob 每一到两个月会做一次这样的练习:盘点 OpenAI 的所有项目,这是一个大约 300 个项目的大表格,我们尽可能深入理解每一个项目,然后对它们进行排名。我认为对于一家 500 人的公司来说,让人们理解核心优先事项是什么,并通过口头明确传达以及我们分配算力的方式来沟通,这很重要。

It's true. People are very creative in the ways that they try to make backroom deals to get the GPUs they need. But yeah, I mean it is a big part of the job, right? Figuring out the priorities for the research org and also being accountable for execution. So really to that first point, there's this exercise that Jakob and I do every one to two months where we take stock of all the projects at OpenAI and it's this big spreadsheet about 300 projects and we go and try to deeply understand each one as best as we can and really rank them. And I think for a company of 500 people, it's important for people to understand what the core priorities are and for those to be communicated clearly, explicitly, verbally, and also through the way that we allocate compute.

介绍与 Brex 广告 Introduction and Brex Ad

Host

从初创公司到全球最大的企业,有 3 万家公司依赖 Brex 技术管理财务。他们有智能公司卡、高收益企业银行账户和费用自动化工具,非常棒。我讨厌做报销。Brex 的 AI 软件会处理这些费用,弄清楚我们的钱花在哪里,帮你处理很多事情,这样你就不用自己浪费时间了。访问 brex.com/cormemory 了解更多,加入我们吧。让我们行动起来,摆脱那些过时的财务软件,迈向未来的核心记忆和 Brex。

30,000 companies from startups to the world's largest corporations rely on Brex technology for their finances. They've got smart corporate cards, high yield business banking, and expense automation tools that are fantastic. I hate doing my expenses. And Brex's AI software runs right through those expenses. Figure out where we're spending money and take care of so much stuff for you so you don't have to waste your time on it yourself. Go to brex.com/cormemory to learn more and just, you know, get with the program. Let's get going. Let's get out of this archaic finance software and move toward the future core memory and Brex.

管理研究项目与算力分配 Managing Research Projects and Compute Allocation

Host

所以当你提到那 500 人时,这些是核心。这是研究团队的核心,而整个组织现在有数千人。嗯,好的。那么,当你谈到这 300 个项目时,我想显然其中一些是巨大的前沿模型,还有一些可能是人们正在做的实验。那么,你如何跟踪所有这些,并最终决定哪些值得分配 GPU,哪些不值得呢?

So you've got when you're talking about the 500, these are the 500. This is the heart of the research team in an organization now that's thousands of people. Yeah. Okay. So, and then in that when you're talking about this 300 projects, I imagine I mean obviously some of those are the giant frontier models and then some are probably experiments that people are working on. And so like how do you possibly keep track of all that and then come to some sort of conclusion about what merits GPUs and what doesn't.

Mark Chen

当然。我认为在做这件事时,保持核心路线图的聚焦非常重要。OpenAI 与其他大型实验室的一个区别是,OpenAI 始终以核心探索性研究为根本。我们不是复制其他实验室的结果,也不是在基准测试上追赶其他实验室。那不是我们的强项。我们总是试图找出下一个范式,并愿意投入资源确保我们找到它。我想大多数人可能会对此感到惊讶,但用于探索的计算资源比训练实际产品还要多。

Absolutely. So I think it is very important when doing this exercise to keep the core roadmap in focus and one thing that differentiates OpenAI from other big labs out there is OpenAI has always had core exploratory research at its core. We are not in the business of replicating the results of other labs or catching up to other labs in terms of benchmarks. That isn't really our bread and butter. We're always trying to figure out what that next paradigm is and we're willing to invest the resources to make sure that we find that right. And I think most people might be surprised at this but more compute goes into that endeavor of doing exploration than it is to training the actual artifact.

Host

这一定很难,你怎么才能不被别人说服呢?因为每个人都会……你知道,当我想到这个时,有时我会想象我在《纽约时报》的时候,我们会有一个头版会议,每个人都想上头版。是的。每个人都认为自己的故事最重要。他们都在尽最大努力告诉你为什么这件事如此重要。房间里的每个人都为他们的提案工作了数周、数月。所以感觉像是生死攸关。是的。我的意思是,这对我来说似乎太难了。

It must be it's still got to be how do you stop yourself from being persuaded by someone because everybody's going to put I you know like when I think about this sometimes I picture when I was at the New York Times you would have this page one meeting where everybody wants to be on page one. Yeah. Everybody thinks their story is the most important story. They're all doing their very best job to tell you why this thing is so important. Everybody in that room has worked weeks, months on on whatever they're pitching. And so it feels like life and death and that Yeah. I mean it just it's seems so difficult for me.

Mark Chen

是的。不,这也是一个困难的过程,我认为最难做的决定是,你知道,这是一个我们现在无法资助的项目。但我也认为这是好的领导力。你需要清楚地沟通,嘿,这些是优先事项。这是我们要讨论的。这些是我们认为能推动研究计划的结果。你知道,可以有其他事情,但它们必须明确排在第二位。

Yeah. No, it is also a difficult process and I think the hardest calls you have to make are you know this is a project that we just can't fund right now. But I also think that's good leadership. You need to clearly communicate that hey these are the priorities. This is what we're going to talk about. These are the types of results that we think move the research program. And you know there can be other things but those have to be clearly number two.

不被动应对竞争对手 Not Being Reactive to Competitors

Host

当你谈到不被动应对竞争对手时,是的。当我翻阅笔记时,我不知道我能不能很快找到那句话,但我的意思是,这似乎是一个自豪点,我看到你觉得其他一些公司,嗯,你知道,你们处于领先地位,为其他人设定了标准,所以他们被动应对,对吧,对你们发布的东西。我们碰巧在 Gemini 3 发布几天后做这个采访,你知道,在某种程度上,你们的竞争对手有时……是的,我的意思是,这种来回较量一直在进行,我知道基准测试的价值存在争议,但你知道,人们在这些事情上争先恐后。那么,随着时间的推移,你如何保持那种奢侈或智力上的位置,让你觉得我们只管做我们想做的事?

And when you like you were talking about not being reactive to your competitors. Yeah. When I was looking through my notes, I I don't know if I could go to the line quick enough, but I mean, this this was like a point of pride that I saw that you feel like some of the other companies are well, you know, you guys were in this position where you were ahead and and setting the bar for others and so so they were reactive, right, to to what you had coming out. We happen to be doing this interview a few days after Gemini 3 came out and you know there is a degree to which your rivals at times like yeah I mean there's this back and forth going on and and I know the benchmarks are sort of controversial how valuable they are but you know people go ahead on these things so how do you also um as time has gone on maintain that luxury or that intellectual position where you feel like we're just going to do what we're going to do?

Mark Chen

是的,我认为今天的 AI 研究格局比以往任何时候都更具竞争性。重要的是不要陷入那种竞争动态,因为你总是可以说,嘿,我要发布一个增量更新,让我领先竞争对手几周或几个月。我不认为那是长期可持续的研究方式,因为如果你破解了下一个范式,那将重要得多,对吧?你将塑造它的演变。你将理解围绕那个思想领域的所有相关研究方向。所以,以我们的强化学习项目为例,我们在两年前就押注我们真的会在语言模型上破解强化学习。这在当时是一个非常不受欢迎的赌注。现在看起来很明显,但当时的环境是,嘿,预训练机器运行得很好,后训练机器也运行得很好。为什么要投资其他东西?我认为今天每个人都会告诉你,思维和语言模型是一个你离不开的基元。所以我们真的在那里做出这些大胆的赌注,并弄清楚如何扩展和构建算法,使其真正扩展到比今天多几个数量级的算力。

Yeah, I think AI research today, the landscape is just much more competitive than it's ever been. And the important thing is to not get caught up in that competitive dynamic because you can always say, hey, you know, I'm going to ship an incremental update that puts me in front of my competitor for, you know, a couple weeks or a couple months. And I don't think that's the long-term sustainable way to do research because if you crack that next paradigm, that's just going to matter so much more, right? You're going to shape the evolution of it. You're going to understand all the side research directions around that sphere of ideas. And so when you think about our RL program as an example of this, right, we bet more than 2 years ago that we're really going to crack RL on language models. And this was a very unpopular bet at the time. Right now it seems obvious, but back then the environment was, hey, you know, there's this pre-training machine that's working great. There's this post-training machine that's working great. Why invest in something else? And I think today everyone would totally tell you that thinking and language models is just a primitive you can't live without. And so we're really there to make these bold bets and to figure out how we can scale and build the algorithms to really scale to orders of magnitude more compute than we have today.

产品压力下保持研究专注 Maintaining Research Focus Amid Product Pressures

Host

只是,你知道,我的意思是,我在理智上理解这一点,但随着你们从一家纯粹的研究公司起步,这变得越来越难。当你今天看 OpenAI 时,我的意思是,你们有产品线。OpenAI 的某些部分看起来更像成熟的微软或谷歌,有产品线,有所有这些不同的事情要服务。通常,我觉得你们还足够年轻,所以可能还没有这些确切压力,但你知道,随着那些公司的发展,总是会变成,我们更专注于服务于底线的事情,而不是在研发上花大钱,研发似乎总是随着时间的推移而减少。

It's just it, you know, I mean, and I get that intellectually in my, you know, it gets harder as you guys started as a like basically a pure research company. When you look at OpenAI today, I mean, you have product line. There's parts of OpenAI that look much more familiar to a mature Microsoft or a Google where you have product lines, you've got all these different things that you have to serve. Typically, I feel like you guys are still young enough, so maybe you don't have these exact pressures yet, but you know, as those companies go on, it always becomes, well, we're more focused on the things that are serving the bottom line than spending a ton of money on research always seems to get like dwindled down over time.

Mark Chen

是的。我认为这正是 OpenAI 最特别的地方之一。其核心,我们是一家纯粹的 AI 研究公司。我不认为其他很多公司可以这么说。你知道,我们最初是非营利组织,我在那个时代加入,我认为精神是,不惜一切代价构建 AGI,推进 AGI 研究,当然要以安全的方式。但事实上,我确实认为这是真正创造价值的最佳假动作,如果你专注于研究并获胜,价值很容易创造。所以我认为有一个陷阱,就是过于沉迷于,哦,让我们提高底线。而实际上,如果你做了最好的研究,那部分就非常容易。

Yeah. And I think that's really one of the most special things about OpenAI. At its core, we're a pure AI research company. And I don't think you can say that of many other companies out there. And you know we were founded as a nonprofit and I joined during that era and I think the spirit is you know build AGI, advance AGI research at all costs and do it in a safe way of course. But yeah I actually do think that's the best head fake to really creating value right if you focus and you win at the research the value is easy to create so I think there's a trap of getting too lost into like oh you know let's drive up the bottom line. When in reality if you do the best research that part of the picture is very easy.

核心文化与工程 vs 研究 Core culture and engineering vs research

Host

你从 2018 年开始,所以你觉得那种灵魂、核心文化和核心内核真的延续下来了。它还在。埃隆说什么?他说我们不应该叫你们任何人研究员。这只是工程,对吧?

And you started in 2018 and so you feel like that soul, that core culture and that core nucleus, it's really persisted. It's still there. What does Elon say? He says we shouldn't call any of you guys researchers. It's just engineering, right?

Mark Chen

是的。不,我认为这是真的,因为我觉得一旦你有了这种等级制度,把研究科学提升到工程之上,你就已经完全输了。当你构建一个大模型时,很多都在于优化那些小百分比的实践:如何让你的内核更快?如何确保数值计算都正确?那是深厚的工程实践,如果没有这部分,你就无法扩展到我们今天使用的 GPU 数量。

Yeah. No, I think it's true because I feel like once you have this hierarchy and you elevate research science as a thing beyond engineering, you've completely already lost the game. When you're building a big model, so much is in the practice of optimizing all those little percentages: how do you make your kernels that much faster? How do you make sure the numerics all work? That's a deep engineering practice, and if you don't have that part of the picture, you can't scale to the number of GPUs we use today.

Host

所以因为研究员和工程师之间存在一种神秘感,你觉得保持冷静更好吗?你是这个意思吗?

So because there is a mystique that surrounds a researcher versus an engineer, do you feel it is better to stay levelheaded on that? Is that what you're saying?

Mark Chen

嗯,我只是觉得研究员有各种各样的类型。我们一些最好的研究员是那种能想出无数点子的人,很多都不好。但就在你快要觉得,这个人真的值得吗?他们就会想出一些绝妙的主意。有些则非常擅长在明确的前路上执行。研究员有太多不同的类型,很难把他们归为一种刻板的工作类型。

Well, I just feel like researchers come in so many different shapes. Some of our best researchers are the type that come up with a billion ideas, and many are not good. But just when you're about to think, is this person really worth it? They come up with some phenomenal idea. Some are just so good at executing on the clear path ahead. There's so many different shapes of researchers, and it's hard to lump them into one stereotypical type that works.

对 Gemini 3 等竞品的反应 Reaction to rival models like Gemini 3

Host

有道理。好吧,我不会用太多竞争对手的问题来烦你。只是既然 Gemini 3 发布了,我确实想知道你个人或团队在对手发布模型时会怎么做。大家都会去看看它能做什么吗?有没有一个提示或问题你经常抛给这些新模型来测试它们的能力?

That makes sense. Okay, I won't belabor you with too many competitive rival questions. It's just since Gemini 3 did come out, I did wonder what happens with you personally or the team when one of your rivals puts out a model. Does everybody go and look and see what it can do? Is there a prompt or a question that you often throw at these new models to see what they can do?

Mark Chen

是的。具体说到 Gemini 3,它是一个相当不错的模型。我们做的一件事是尝试建立共识。基准测试只能告诉你这么多。仅从基准测试来看,我们实际上相当自信。我们内部有性能达到 Gemini 3 水平的模型,我们很有信心很快会发布它们,并且可以发布更好的后续模型。但同样,基准测试只能告诉你这么多。我认为每个人都有自己的方式来探测模型。有一个数学问题我喜欢给模型做。到目前为止,我认为没有一个模型完全破解了它,即使是推理模型。所以是的,我会等着。

Yeah. So to speak to Gemini 3 specifically, it's a pretty good model. One thing we do is try to build consensus. The benchmarks only tell you so much. Just looking purely at the benchmarks, we actually felt quite confident. We have models internally that perform at the level of Gemini 3, and we're pretty confident that we will release them soon and we can release successor models that are even better. But again, the benchmarks only tell you so much. I think everyone probes the models in their own way. There is this math problem I like to give the models. I think so far none of them has quite cracked it, even the thinking models. So yeah, I'll wait for that.

Host

这是一个秘密的数学问题吗?

Is this like a secret math problem?

Mark Chen

哦不。好吧,如果我在这里宣布,也许它会被训练进去。但我确实认为这是去年一个很好的谜题。是 42 问题。你想创建一个模 42 的随机数生成器,并且你有一堆原语,它们是模小于 42 的素数的随机数生成器。你希望期望上尽可能少地调用这些子生成器。这是一个非常巧妙的谜题,但语言模型非常接近最优解,但我还没看到一个完全破解它的。

Oh no. Well, if I announce it here maybe it gets trained on. But I do think it's one of the nice puzzles of last year. It's the 42 problem. You want to create a random number generator mod 42, and you have access to a bunch of primitives which are random number generators modulo primes less than 42. You want to make as few calls on expectation to these sub-generators as possible. It's a very cute puzzle, but the language models get pretty close to the optimal solution, but I haven't seen one quite crack it.

Host

好的,这正朝着我想问的方向发展。但在那之前,我知道你很有竞争力。你也告诉过我你热爱竞争,讨厌失败。我真的很讨厌失败。所以我在想:如果我们知道 Gemini 3 或别的什么在周四发布,你会不会在午夜起来用那个问题测试它,还是没那么夸张?

Okay, this is heading down a direction I want to ask you about. But just before we get there, I know you're very competitive. You've also told me you love competition and hate to lose. I really hate losing. So I'm picturing: if we know Gemini 3 or whatever is coming out on a Thursday, are you up at midnight throwing that problem at it, or is it not quite that drastic?

Mark Chen

不。我认为这是长期的事。任何努力,我是那种有执着的人。我认为任何努力都必须打持久战。我们实际上一直在专注于预训练,特别是过去半年加强我们的预训练工作。我认为这是这些努力的结果,与 Yakub 一起,专注于并建立 OpenAI 的预训练能力。围绕它打造一个真正的明星团队,确保预训练的所有重要领域和方面都得到强调。正是这些创造了今天的成果,让我们感觉在预训练上可以轻松与 Gemini 3 正面竞争。

No. I think it's in long arcs. Any endeavor, I'm kind of a person who has obsessions. I think any endeavor you have to play the long game. We've actually been focusing on pre-training, specifically supercharging our pre-training efforts for the last half year. I think it's a result of some of those efforts, together with Yakub, focusing and building that muscle of pre-training at OpenAI. Crafting a really superstar team around it, making sure all the important areas and aspects of pre-training are emphasized. That's what creates the artifacts today that feel like we can go head-to-head with Gemini 3 easily on pre-training.

Host

好的。我想问问预训练的事情,因为我一直在和你们所有人聊这个。但你是说你不太痴迷于在新模型出现时向它们抛问题,而是更关注这个长期旅程。

Okay. And I want to ask about the pre-training stuff because I've been talking to all you guys about this a lot. But you're saying you're less obsessed about lobbing problems at these new models just when they appear and more at this long journey.

Mark Chen

绝对。

Absolutely.

竞赛与编程背景 Background in competitions and coding

Host

好的。嘿,我想谈谈你提到的那个谜题——我第一次见到 Yakob 是在 OpenAI 成立之前,当时他在参加编程竞赛,我有一段时间非常沉迷编程竞赛。有个叫 Kennedy 的家伙,我不知道他是否还出名,但他就像是编程竞赛界的迈克尔·乔丹。所以我去 Facebook 看了一场比赛——他们以前每年举办 Hacker Cup。那是我第一次见到 Yakob。然后我知道你在高中参加过数学竞赛。

Okay. Hey, the reason I want to talk about the puzzle you were at—I first met Yakob before OpenAI ever started when he was doing a coding competition, and I got super into coding competitions for a while. There's this guy Kennedy. I don't know if he's still famous, but he was like the Michael Jordan of these coding competitions. So I went to watch one at Facebook—they used to have an annual Hacker Cup. That's where I saw Yakob for the first time. And then I know you did math competitions in high school.

Mark Chen

是的,大概从小学到高中。然后你也参加了 III?

Yeah, probably grade school through high school. And then you also did III?

Host

所以我很晚才开始编程。是大学室友说服我上了第一堂编程课。当时我带着数学家的所有傲慢,认为数学是最纯粹、最难的学科,那才是真正证明你价值的地方。我想我当时可能太沉迷于竞赛了。但后来它变成了一个非常有回报的努力,最初纯粹是为了和大学朋友保持联系。

So I got into coding really late in life. It was a roommate in college that convinced me to take my first coding class. I had all the hubris of a mathematician at that time, thinking math is the purest and hardest science and that's where you really prove your worth. I think I was probably too into the competition back then. But it became this super rewarding endeavor, and it started out as purely a way to keep in touch with my friends from college.

Mark Chen

你去了麻省理工。

You went to MIT.

Host

是的,我去了麻省理工。毕业后,每个周末我们都会登录做这些竞赛,只是为了保持联系。随着时间的推移,我发现自己在这方面有天赋。我开始表现得相当不错,然后为像美国编程奥林匹克这样的竞赛编写题目。

Yeah, I went to MIT. I graduated, and every weekend we would just log on and do these contests just to keep in touch with each other. Over time, I found myself having a talent for it. I started competing fairly well and then writing problems for contests like the USA Coding Olympiad.

执教美国 IOI 队 Coaching the US IOI Team

Host

后来开始指导那个团队,是的,这是一个很棒的社区,我遇到了像斯科特这样的人。

Eventually started coaching that team and yeah, it's been a great community where I've met people like Scott that you know.

Mark Chen

是的。

Yeah.

Host

是的。好的。我想很多人可能熟悉数学竞赛,因为他们看到孩子们参加。IOI 和这些编程竞赛有点不同。它就像一个谜题式的文字题,你要找到最有效、最正确的解法,然后和所有人竞争。每个人都在电脑上写代码,有些人想很快完成,但他们的代码解决不了问题,对吧?这是一种权衡。

Yeah. Okay. So, I think lots of people might be familiar with math competitions because they probably see kids going through that. The IOI and these coding competitions are a little bit different. I mean, it's almost like a word problem that's a puzzle, and you're trying to find the most efficient and correct way to solve that, and you're in this race against everybody. And everybody's writing code on their computer, and some people try to get there really fast, but then their thing doesn't solve the problem, right? There's this trade-off.

Mark Chen

完全正确。

Absolutely. Right.

Host

所以你实际上在麻省理工团队?

So you actually were on the MIT team?

Mark Chen

不,不,这是我大学毕业后做的事。

No, no, it's something I did after college.

Host

大学毕业后。好的。但今天你是美国国家队的教练?

After college. Okay. But today you are the coach of the US national team?

Mark Chen

是的,其中一位教练。

Yeah. One of the coaches.

Host

其中一位教练。好的。是去年还是前年?美国很久没赢过了,对吧?

One of the coaches. Okay. And was it last year or the year before? The US hadn't won in a long time, right?

Mark Chen

是的。是的。所以团队,你永远无法预测每年顶尖人才的构成。但两年前我们有一个非常突出的团队。我相信他们赢得了奥林匹克竞赛。

Yeah. Yeah. So the team, you can never predict the makeup of top talent every year. But we had a very spiky team two years ago. And I believe they won the Olympiad.

Host

因为我觉得通常是中国、俄罗斯、白俄罗斯和波兰。

Because I feel like usually it's China or Russia or Belarus and Poland.

Mark Chen

对。大型比赛每年在不同国家举行。

Right. And the big competition takes place in a different country every year.

Host

它是什么样的?有多少人参加?

What does it look like? How many people show up?

Mark Chen

每个国家选前四名学生。这既是比赛也是社交活动。这是一个紧密的社区。他们后来都做出了非凡的事情。这是一场紧张的两天比赛,每天只有三道题,5 小时解决。你能感受到房间里的肾上腺素和压力。但也很有趣。人们通过它结交终身朋友。

They take the top four students from every country. It is as much a competition as a social event. This is a tight-knit community. They all go on to do phenomenal things. It's an intense two-day contest where each day you get just three problems, 5 hours to solve them. You can feel the adrenaline and pressure in the room. But it's also great fun. People make lifetime friends through it.

Host

作为教练,你那么忙。你在这上面花多少时间?

As coach, you're so busy. How much time do you spend on this?

Mark Chen

说实话,孩子们非常自觉。有时只是管理他们的表现和策略。比赛中会有好日子、坏日子、好时段、坏时段,你不能让这些影响心态。管理参赛者和管理研究人员有相似之处。在更长的时间尺度上,研究人员有好月份、坏月份。你不能让一连串的失败影响心态,因为这就是研究的本质。很多是士气管理。

Honestly, the kids are so self-motivated. Sometimes it's really about managing their performance and strategy. You're going to have good days, bad days, good hours, bad hours within the contest, and you can't let that get into your head. There are similarities between managing contestants and managing researchers. On a longer time scale, researchers have good months, bad months. You can't let strings of failures get into your head because that's the nature of research. A lot of it is morale management.

Mark Chen

最近竞赛让我意识到一件有趣的事:当你把模型部署来解决这些竞赛问题时,它们现在非常擅长。

One interesting thing contests have helped me realize lately is when you put the models and deploy them towards solving these contest problems, which they're quite good at these days.

Host

我正要问你这个问题。

I was going to ask you about that.

Mark Chen

它们的工作方式与人类非常不同。我们通常认为这些机器擅长模式识别。如果一个问题映射到之前的问题,它可能能解决。但我注意到在之前的 IOI 中,有一个像“消息”这样的问题非常特别。我本以为模型根本解决不了,但实际上对 AI 来说它反而是较容易的问题之一。这让我觉得,AI 加上人类在前沿研究中会做出惊人的事情,因为 AI 对什么容易、什么不容易有不同直觉。

They work in a very different way from humans. We typically think of these machines as very good at pattern recognition. If a problem maps to a previous problem, it can probably solve it. But I've noticed in some previous IOI, there's a problem like 'message' that is very ad hoc. I didn't think the models would solve it at all, but it was actually one of the easier problems for the AI. This has given me the sense that AI plus humans in frontier research will do something amazing because the AI has a different intuition for what's easy and what's not.

Host

这有点像 DeepMind 做 AlphaGo 时,它用人类从未用过的方式下棋?

Is it vaguely similar to when DeepMind did AlphaGo, where it was playing in ways humans hadn't played before?

Mark Chen

我想是的。有了 GPT-5 Pro,前沿研究出现了一个转折点。最好的轶事之一:发布三天后,我遇到一位物理学家朋友。他一直在玩模型,觉得它们可爱但不太有用。我让他用 Pro 模型尝试一些雄心勃勃的事情。他输入了他最新的论文,模型思考了 30 分钟,然后理解了。那个反应就像看到丽莎·多尔在第 37 步、第 38 步时的反应。我认为这会在前沿数学、科学、生物学、材料科学中越来越多地发生。模型真的达到了那个水平。

I think so. With GPT-5 Pro, there's been an inflection point in frontier research. One of the best anecdotes: three days after the launch, I met a physicist friend. He had been playing with the models, thought they were cute but not super useful. I challenged him with the pro model to try something ambitious. He put in his latest paper, it thought for 30 minutes and just got it. That reaction was like seeing Lisa Doll during move 37, move 38. I think that will keep happening more and more for frontier mathematics, science, biology, material science. The models have really gotten to that point.

Host

我正要问你这个问题,虽然不太新颖,但作为一个关注这些竞赛的人,当你开始看到这些模型解决那些曾是这些独特人类心智成就巅峰的问题时,是否感到悲伤?

I was going to ask you this question, which is not very original, but as somebody who followed these competitions, is there a sadness when you start seeing these models solving things that were the height of achievement for these very unique human minds?

Mark Chen

嗯,既是也不是。我擅长竞技编程,但从未达到顶尖。也许这是复仇的方式。但对我来说确实有一个时刻。我们在开发推理模型时跟踪了一段时间的编程竞赛表现。一开始,它们不是很好,只有普通参赛者的水平。随着时间的推移,它们的能力逐渐提升。我记得那个时刻,你走进会议室,他们展示你的表现,然后模型超过了它。那对我来说是个震惊。就像,哇,我们这么快就自动化到了这个能力水平。当然,雅各布还在那里有点得意,但一两个月内,模型也超过了他。

Well, yes and no. I was good at competitive programming, never at the absolute top. Maybe this is a way to get revenge. But there's certainly a moment for myself. We tracked coding contest performance while developing reasoning models for a while. At the start, they were not super great, at the level of any average competitor. Over time, they started creeping up in capability. I remember that moment when you walk into the meeting and they show where your performance is, and then the models exceeded that. That was a shock to me. It's like, wow, we've automated to this level of capability so fast. And of course, Yakob was there still a bit smug, but within one or two months, it was also surpassing him.

前沿模型与编程竞赛 Frontier Models and Coding Competitions

Mark Chen

所以,模型如今已经处于前沿了,对吧?从今年夏天我们在 Codeforces 上取得的结果就能清楚看到,那是全球顶尖的优化竞赛程序员。我认为它在那里获得了第二名。所以它确实从去年的第 100 名跃升到了今年的前五名。

So, the models are at the frontier today, right? It's clear from the results we've done this summer at Codeforces, top optimization competitive programmers in the world. I think it achieves second place there. So really it's jumped from hundredth place last year to top five this year.

Host

你觉得 10 年后我们还会做这些比赛吗?

Do you think we'll still be doing these competitions in 10 years?

Mark Chen

我觉得会。它们就是很有趣。当然,那些用它来充实简历的人会退出,但我觉得那些一直最擅长的人,就是纯粹为了乐趣而做的人。我不觉得这会消失。

I think so. I mean, they're just fun. Certainly a bunch of people who use it to pad their resume are going to drop off from doing it, but I think the people who've always excelled at it the most are people who just do it for the fun of it. And I don't think that'll go away.

Host

我在做这个报道时,他们告诉我,如果你来自俄罗斯或某些国家,你基本上可以免费进入任何你想要的大学。我看到美国队的成员去了哈佛和麻省理工,所以他们看起来还不错,但似乎美国没有……

When I was doing this story, they were telling me that if you're from Russia or some countries, you basically get an automatic free ride to any university you want. I see the guys on the US team go to Harvard and MIT, so they seem to be doing okay, but it doesn't seem like the US has a...

Mark Chen

是啊。你不觉得面试在未来会有点失灵吗?大家都有点感觉到了。甚至大学考试或大学作业,现在也都失灵了,对吧?我确实认为我们需要新的方法来评估和衡量谁表现得好,谁学到了东西,以及一个人实际的水平。

Yeah. I mean, don't you think interviews are going to be kind of broken going forward? And everyone's seeing this a little bit. Even college exams or college homework, it's all broken at this point, right? I do think we're going to need new ways of assessing and gauging who's performing well, who's learned the material, where somebody's actually at.

Host

是的。所以我有个想法,也许在我们的面试中,我们让候选人和 ChatGPT 对话,一种特殊的 ChatGPT,模型试图判断你是否了解材料,或者你的能力水平是否适合在 OpenAI 工作。你必须和它进行一场对话,深深地说服它你属于 OpenAI。当然,你不能越狱它,但我们之后会看对话记录。也许这样的测试在未来能更准确地反映你是否知道。

Yeah. So I've had this idea where maybe for our interviews we should just have candidates talk to ChatGPT, a special kind of ChatGPT where the model is trying to gauge whether you know the material or whether you're at the capability level to work at OpenAI. You have to have this conversation with it that convinces it deeply you belong at OpenAI. Of course, you can't be allowed to jailbreak it, but we look at the transcript after. Maybe tests like this will more accurately reflect in the future whether you know.

Mark Chen

所以,你们还没这么做,但在考虑?

So, you don't do that yet, but you're thinking about it?

Host

是的。只是用一些创造性的方式来改进面试。

Yeah. Just creative ways to revamp the interviews.

Mark Chen

是啊。硅谷以在面试中做脑筋急转弯之类的事情而闻名。

Yeah. Well, Silicon Valley is famous for doing brain teasers during the interviews and everything.

Mark Chen 的背景与成长 Mark Chen's Background and Upbringing

Host

你小时候数学很好。你出生在东海岸吗?

You were very good at math growing up. Were you born on the East Coast?

Mark Chen

是的,出生在东海岸。

Yeah, born on the East Coast.

Host

然后你也住在西海岸。

And then you lived on the West Coast too.

Mark Chen

在西海岸,然后你在台湾从小学住到高中。

On the West Coast, and then you lived in Taiwan for grade school to high school.

Host

四年。

Four years.

Mark Chen

好的。你父母在贝尔实验室工作。

Okay. Your parents worked at Bell Labs.

Host

所以你出身于工程师世家。这是一个非常有趣的背景,因为你体验了所有这些创新中心,尤其是你父母在贝尔实验室。

So you come from engineering stock. It's a really interesting background because you got a flavor for all these innovation hubs, especially with your parents at Bell Labs.

Mark Chen

我就是在非常科学的环境中长大的。餐桌上的话题是谜题之类的东西。我也体验了更传统的东海岸贝尔实验室经历。在西海岸,我爸爸来创办一家初创公司,所以我小时候也接触到了那种新公司。当然,还有搬到台湾的巨大变化。那是一个巨大的文化冲击。你穿校服,学校周围有铁丝网。也接触到了那种严谨程度。我认为那是一系列非常棒的成长经历。

I just grew up in a very scientific environment. Dinner table talk was puzzles and things like that. I also got the more traditional Bell Labs East Coast experience. On the West Coast, my dad came to do a startup, so I got exposed to that kind of new company when I was young as well. And of course, the big jump to Taiwan. It was a huge culture shock. You wear uniforms, you're in a school with barbed wire around it. And also getting exposed to that level of rigor. I think it was a number of really great experiences growing up.

Host

所以学校更难?

So the schools were harder?

Mark Chen

嗯,我会说只是更……学校系统里灵活性和自由少了一点,但我觉得这也教会了你一些东西。

Well, I would say it was just much more kind of... there's a little bit less flexibility and freedom in the school system, but I think it also teaches you something.

Host

而且你知道你想回美国上大学。

And you knew you wanted to come back to the US for college.

Mark Chen

当然。

Absolutely.

MIT 与 2012 届 MIT and the 2012 Cohort

Host

所以你在麻省理工。你在一个有趣的群体里。我想麻省理工可能总有一个有趣的群体。

So you're at MIT. You're in this interesting group. I guess MIT probably always has an interesting group.

Mark Chen

哦,天哪。2012 届真是个很棒的群体。

Oh, man. 2012 was such a great group.

Host

都有谁?一个全明星名单?

Who is there? An all-star list?

Mark Chen

那是很棒的一年。不知道你认不认识 Jacob Steinhardt,他现在在做 TransLuce。他和我以前在计算机科学课上一起做项目。还有 Paul Christiano,他在 OpenAI 工作过。很多 AI 界的大人物都来自那一年。

It was a great year. I don't know if you knew Jacob Steinhardt, he's doing TransLuce now. He and I used to do projects together in computer science class. There was Paul Christiano, who worked at OpenAI. A bunch of big names in AI came from that year.

Host

然后我们谈到了竞技编程。Scott Wu,他在 Cognition,现在因为数学能力在 X 上成了网红。你通过编程比赛认识了他。

And then we were talking about competitive coding. Scott Wu, who's at Cognition, he's kind of famous now as a meme on X for his math abilities. You got to know him through the coding competition.

Mark Chen

哦,是的。通过编程社区。

Oh, yeah. Through the coding community.

扑克作为数学游戏 Poker as a Mathematical Game

Host

现在我看到了你们竞争的一面。在我看来,这现在的产出就像扑克。我参加了一个活动,我觉得具体细节得保密,但我觉得这部分可以说:深夜,我走过一张桌子,有你、Scott,还有 Palantir 的 Sham,以及其他几个人,在玩一场相当激烈的扑克。所以你们现在把数学和竞技技能用到了这里。

Now I see the competitive end of you guys. The output of this to me looks like poker these days. I was at an event which I think I have to keep secret, but I think I'm okay to talk about this part: late at night, I'm walking by this table, there's you, Scott, I think Sham from Palantir, and a handful of other people in a fairly intense poker game. So you guys have applied your math and competitive skills now.

Mark Chen

扑克是一个非常有趣的游戏。我把我的人生描述为一系列痴迷。扑克绝对是过去这些痴迷之一。扑克给我的重大启示是,它更像是一个数学游戏,而不是一个读人和诈唬的游戏。你对扑克了解得越多,就越会朝那个方向更新。我以前是个很差的诈唬者。当你知道从数学上诈唬是正确的,那就很容易了。你一点也不会紧张。有趣的是,一个被认为如此人性化的游戏,其底层机制和获胜方式却如此深奥地数学化。

Poker is a really fun game. I've talked about my life in terms of a series of obsessions. Poker was definitely one of these obsessions in the past. The big revelation for me in poker was that it's so much more a mathematical game than a game of reading people and bluffing. The more you learn about poker, the more you update in that direction. I used to be a terrible bluffer. When you know it's mathematically correct to bluff, then it's so easy. You don't feel any nervousness around it. It's just so interesting that you have a game that is perceived as so human, but the underlying mechanics and how to win are so deeply mathematical.

扑克与竞争 Poker and Competition

Host

是啊,我前几天也在想这个问题,语言建模也有类似的地方,对吧?语言生成是深度人类的过程,但数学机器也能做得和我们一样好。

And yeah, I kind of thought about this the other day where there's something about that in language modeling too, right? You have this deeply human process of generating language, but there's this mathematical machine that can really do it as well as we can.

Mark Chen

作为写作者,我一直在思考这部分,大学时还和维特根斯坦那些人一起研究哲学,思考这些问题。

I think about that part all the time as a writer and then I did all this philosophy back in college with Wittgenstein and all these guys thinking about these things.

Host

是啊。那你怎么找到优势呢?你和斯科特都让我觉得数学好得不可思议,我不明白你们怎么还能互相算计。

Yeah. Well, how do you find an edge? You and Scott both strike me as supernaturally good at math, I don't understand how one of you is out calculating the other.

Mark Chen

不,其实这主要是我们聚在一起叙旧的场合。现在我们不会太当真。我觉得把扑克这种事看得太重会失去乐趣。我对扑克的痴迷十多年前就结束了,现在只是玩玩。

Well, no, I mean, it is mostly a forum for us to just hang out and catch up with each other. And today we don't take it as seriously. I think there is an element to taking something like poker very seriously that takes the fun out of it. So my obsession with poker, I think, has ended more than a decade ago and now it's just fun.

Host

你这么说是因为我看到斯科特赢了两天。我觉得你可能说得对。他确实很认真。所以大学毕业后,某种意义上我……

You're just saying this because I saw Scott win both days. I think you might be right about that. He was taking it quite seriously. And so like coming out of college I mean you had in some sense I was...

Mark Chen

不过我在飞机上赢了他。

Oh I beat him on the plane though.

Host

你在回家的飞机上赢了他。好吧。那是你和他单挑还是多人局?

You beat him on the plane right home. Okay. All right. So you did you did was it just you versus him or it was like a group thing?

Mark Chen

大概三四个人。

Maybe three or four people.

Host

好吧。我觉得很多人,我觉得我没有过度概括,尤其是回到 2018 年左右。高水平 AI 从业者中,很多有学术背景,很多是数学天才,或者用数学背景进入机器人或物理等领域。还有另一类人,他们去了华尔街做高频交易和量化之类。所以你走的第一条路就是从 MIT 直接去了华尔街。

Okay. I feel like a lot of, I don't think I'm overgeneralizing too far, especially among like if you cut back to the 2018 sort of time frame. As far as people who were in AI at a high level, a lot had academic backgrounds. A lot were math prodigies or had gone on to take their math background and get into robotics or something physics like that. And then there's this other bucket which is people who had gone to Wall Street and done high frequency trading and quants and things like that. So that was the first path that you took was you went straight from MIT to Wall Street.

Mark Chen

是啊,说实话我并不以此为荣。这是 MIT 量化导向学生很常见的路径。这确实是一个精英体制,你可以运用智力,有明确的回报函数——你赚的利润。但文化上我觉得很难受。在那里,当你发现什么,第一反应是尽量保密,因为知识就是你的价值。所以即使在公司内部,这种竞争动态也让人互不信任。而且感觉是一个封闭的生态系统。今天,如果 HFT 有人突破让算法快一点,别人也感觉不到。过了四五年,我醒来发现我们还在和同一批人竞争,每个人都快了一点,但世界真的因此改变了吗?我觉得是时候做点别的了。当时很多事情凑在一起。AlphaGo 的比赛,我认为对 OpenAI 很多人是巨大的激励。

Yeah. I mean I don't wear that badge with too much pride, to be honest. It was a path that was fairly common for very quantitatively oriented kids at MIT. And it was certainly a very meritocratic system, right? You could apply intelligence and there's a very concrete reward function, the amount of profit that you would make. But I think culturally it was hard for me. It was a place where when you discover something, your first instinct is to just keep it away from as many people as possible because your knowledge is what gives you your worth. And so it felt like an outgrowth of even internally at a company, these competitive dynamics and people weren't very trusting of each other. I think it also felt like such a closed ecosystem, right? I think today we don't feel too much like when someone in HFT finds a breakthrough that makes their algorithm a little bit faster, no one else feels it, right? And I over time just kind of felt like I woke up after four or five years and we're competing against the same exact set of players. Everyone was a little bit faster but had the world really changed that much for it? And it felt like time to do something else. A bunch of things lined up back then. There's the AlphaGo match, which I think was a huge inspiration for a lot of people at OpenAI.

Host

你下围棋吗?

Did you play Go?

Mark Chen

不下,但模型能做出有创意的事情,这让我很想理解背后的原理。

I did not, but I think the sense in which the model was able to do something creative, I really wanted to understand what was going on behind that.

Host

所以你看着这一切发生,之前有读过 AI 研究论文之类的吗?

So you're watching that happen and had you been reading AI research papers and things like that?

Mark Chen

说实话没有。看到那个事件后,它非常鼓舞人心,我就开始深入钻研 AI。之后我的目标之一就是复现 DQN 的结果。这是一个能在很多雅达利游戏上达到超人水平的网络。从那里开始,我就这样进入了 AI 领域。

To be honest no. And then I saw that event, it was really inspiring and that's when I started doing my deep dives into AI. So one of my goals after seeing that was to reproduce the DQN results. This is a network that was able to play a lot of Atari games effectively at a superhuman level. And going from there, that's how I got my start in AI.

Host

你是业余时间做的?白天工作,晚上回去尝试……

And you were doing that on the side as you so you just work all day and then go back and try to...

Mark Chen

对,对,对。

Yeah. Yeah. Yeah.

Host

好吧。这确实奇怪。我记得大概 2018 年左右采访过乔治·霍兹,可能稍早一点,他刚在车库里自己造了一辆自动驾驶汽车。然后,你知道,他是乔治,有时会说大话,不一定准确或适用于别人。但他说 AI 还很年轻,你基本上读个 10、20、30 篇论文就能学完整个领域。我觉得这很迷人,AI 在很多方面历史悠久,但在这个特定时刻却很浅。我总是给那些对进入 AI 感到畏惧的人建议:它很浅,花三到六个月选个项目,比如复现 DQN,就能很快到达前沿。最近几年加深了一点,但远不及理论数学或物理。

Okay. I mean it is weird. I remember I was interviewing George Hotz like it might have been roughly 2018. Maybe a little before that, and he had just done this thing building a self-driving car on his own in his garage. And then, I mean, it's George, so he says large statements sometimes that may or may not be like exact or spot on or apply to other people. But he's like AI is still so young. You can basically learn the whole field if you read I don't know what the number was 10 20 30 research papers. I mean it is fascinating to me that it's old in many ways stretching back decades but this particular moment it's very shallow. I always give this advice to people who are intimidated by getting into AI. So shallow, just spend three to six months picking some project like maybe reproduce DQN and you can get to the frontier very quickly. The last couple years has added a little bit of depth but it's not anything like theoretical math or physics.

Mark Chen

我总是给那些对进入 AI 感到畏惧的人建议:它很浅,花三到六个月选个项目,比如复现 DQN,就能很快到达前沿。最近几年加深了一点,但远不及理论数学或物理。

I always give this advice to people who are intimidated by getting into AI. So shallow, just spend three to six months picking some project like maybe reproduce DQN and you can get to the frontier very quickly. The last couple years has added a little bit of depth but it's not anything like theoretical math or physics.

Host

你觉得这个领域是不是——我前几天问过 Jaob,不知道为什么我执着于此——就像数学,人们往往在 20 多岁做出最好的工作或重大突破,年纪大了就很难再有那样的时刻。按你说的,我们是依赖年轻人读论文然后获得洞见,还是可以整个职业生涯持续前进?

Do you think is this a field where I asked Jaob this the other day I don't know why I'm obsessed with this but you know like in mathematics you see people tend to do their best work in their 20s or to have the big breakthrough and then it's very hard as they get older to have that same kind of moment. Like what you're saying, are we dependent on young people reading these papers and then having some insight or is this something where you can keep going throughout your whole career?

Mark Chen

我觉得你可以持续前进。OpenAI 本身文化很年轻,但我不认为需要年轻才能做好研究。年轻的好处是较少先入为主的观念。随着时间的推移,你会形成自己的视野,这是好事,但也会让你陷入思维定式,觉得研究就该这么做,好结果就该这么来。我认为年轻研究者在这方面更有可塑性。

I mean, I think you can keep going. OpenAI itself does have a pretty young culture, but I don't think you need to be young to do good research. I think there is something about being young and having less priors about this is the way it's done. Over time you may develop your own vision, which is a good thing, but it also locks you into a frame of mind of like oh this is how research is done, this is how good results come out. I think younger researchers tend to have a little bit more plasticity around that concept.

Host

是啊。所以你在 OpenAI 的职业生涯很有趣,因为你一进门就担任了非常重要的高级职位。

Yeah. So as your career is at OpenAI is funny because it seems like you walked in the door and had a very important large position from the get-go.

OpenAI 早期岁月 Early Days at OpenAI

Host

但你 2018 年刚去的时候,大概也就 50 个人吧。

But when you first got there in 2018, it must have been what, like 50 people.

Mark Chen

哦不,更接近 20 人。

Oh, no, it was much closer to 20.

Host

更接近 20 人。好吧。

Much closer to 20. Okay.

Mark Chen

当时看起来真的就像两个团队。我是以驻留身份加入的,所以显然不是专家,也不是博士。我觉得我在 OpenAI 的整个任期都只是驻留身份。在这方面我很幸运,就是学习他高层次的研究思维方式。

And it really looked like two teams back then. I came in as a resident. So someone who, you know, clearly not a specialist, not a PhD. I think I was only a resident throughout my tenure at OpenAI. So I was very lucky in that regard, just kind of learning the way that he thinks high level about research.

Host

这里的驻留就是你是某人的得力助手?

And a resident in this case is you're just the right-hand person to...

Mark Chen

哦,所以通常是来自其他领域的人,OpenAI 想投资并培养他们进入 AI 领域。我认为驻留的第一部分就像是一个压缩版的六个月博士课程,然后从那里开始进入越来越深入的研究项目。

Oh, so it's someone who comes in usually from another field who OpenAI wanted to invest in and train up in AI. And so I think the first part of a residency is like a six-month compressed PhD and then going from there to just getting into deeper and deeper research projects.

Host

所以你几乎每天都和 Ilya 交流?他是在塑造你的研究方向吗?

So you're kind of like talking to Ilya every day. Is he kind of shaping that?

Mark Chen

是的,他负责我的项目、课程和学习。我会去找他问,比如‘嘿,这是怎么回事?人们为什么研究这个?’

Yeah, he was responsible for my projects, my curriculum, my learning. And I would just go to him for, you know, 'Hey, what's this about? Why did people pursue this?'

Host

好的,明白。

Okay, yeah.

Host

我的意思是,如果你上领英,上面会写你在 OpenAI 的第一份工作是前沿研究负责人。

I mean, if you go on LinkedIn, it would say you were the head of Frontier Research as your first job at OpenAI.

Mark Chen

哦不,我做了大约三年的独立贡献者。所以我做独立研究项目。我从事生成式建模,因为那确实是当时 Ilya 的重点。过了一段时间我才开始管理团队。

Oh no, I was an IC for about three years. So I was doing independent research projects. I worked in generative modeling because that was really where Ilya's focus was at the time. And only after a while did I start managing teams.

Host

因为大多数人认为 DALL-E 可能是公众看到的第一个大项目,这么说公平吗?

And because most people point to DALL-E as maybe the first big project that the public would... mostly, is that fair?

Mark Chen

是的。所以我认为那也标志着我从独立贡献者向管理者的转变。我自己做的一个大项目,至今仍让我感到自豪的是 ImageGPT,这是一个概念验证,表明即使在语言之外,你也可以将图像等内容放入 Transformer 中,模型会内化非常好的表征并理解图像内容。这有点像概念验证,表明你可以在纯文本之外进行语言建模,并获得非常好的表征,并通过其他方法将其扩展到最先进水平。这是 DALL-E 的前期工作,而 DALL-E 我是在管理端参与的。另一个我引以为豪的独立贡献者工作是 Codex,我们建立了评估编码模型的框架,并深入研究了如何让语言模型在代码方面表现出色。

Yeah. So I think that also marked the transition between when I was an IC and a manager. One of my own big projects and one I'm still pretty proud of today is ImageGPT, this proof of concept that even outside of language you could put things like images into a Transformer and the model would just internalize very good representations and understand the content of images. It's kind of like a proof of concept that you can do language modeling outside of pure text and get really great representations, and scale them to be state-of-the-art with other methods. That was a precursor work to DALL-E, which I was on the opposite side of managing. Another project I'm really proud of doing IC work on is Codex, where we set up a lot of the framework for evaluating coding models and also did a lot of in-depth study on how you can take language models and make them very good on code.

选择 OpenAI Choosing OpenAI

Host

那你为什么选择 OpenAI?因为我在脑子里有两种看法。一种是‘小池塘里的大鱼’,这里有有趣的人。我记得 2018 年的 OpenAI 只有 20 个人。在我脑子里,这很可能行不通。谷歌似乎已经掌控了局面,而这是一小群人试图挑战一个似乎需要数十亿美元资本的事情。而且这还是在 Scaling(规模扩张)之前。谷歌已经在 AI 上投入了那么多。这对你来说是个艰难的决定吗,还是你只是很快地偶然得到了 OpenAI 的工作?

So what made you pick OpenAI? Because I could see it two ways in my head. One is big fish in a small pond. There's interesting people here. I remember the OpenAI of 2018 with 20 people. In my head, it was like, this is probably not going to work. Google seems like they've got this locked up and this is a pretty small group of people trying to take on something that appears to require many billions of dollars of capital. And this was even before the scaling stuff. It was just like Google had invested so much in AI already. Was that a hard decision for you or did you just stumble into the OpenAI gig so quickly?

Mark Chen

嗯,我认为有两件事:你需要有远大的愿景,OpenAI 当时当然有,但你也需要有足够的人才来支撑。我确实觉得 OpenAI 是少数几个愿景非常宏大,同时人才也足够强大来填补差距的地方。我很幸运之前就认识 Greg 这样的人,从大学开始。我们在高中一起参加过数学竞赛。我给他发了条消息,说‘我不知道我是否有合适的技能,但这听起来像是一个在做伟大工作的地方。’

Well, I think there are two things: you need ambition of vision, which certainly OpenAI had at the time, but you also need the talent to back it up. I did feel like OpenAI was one of the rare places where the ambition was very large, but the talent was also large enough to fulfill that gap. I was lucky I knew people like Greg from before, from college. We did some math contests together back in high school. I shot him a message and I was like, 'I don't know if I have the right skill set, but this sounds like a place that's doing great work.'

Host

你从无名小卒到现在领导研究,这仍然看起来很疯狂。

It still seems nuts that you just came out of nowhere and now you're leading research.

Mark Chen

不,对我来说也很超现实。即使是从独立贡献者到管理者的转变,我也非常犹豫。我不知道管理是不是我擅长的技能,而且我真的很喜欢独立贡献者的工作。我做得很有乐趣,表现出色,建立了很好的合作关系。但确实,这是一段疯狂的旅程。

No, it's surreal to me, too. Even that transition from IC to manager, I was very hesitant about taking it. I didn't know if managing was a skill set I would be good at and I was really enjoying IC work. I was having a lot of fun doing it, excelling at it, building really great collaborations. But yeah, it's really been a wild ride.

管理与 OpenAI 文化 Management and OpenAI's Culture

Host

你一直给我一种非常友善、头脑冷静的印象。OpenAI 的历史中有一些相当戏剧性的部分,像肥皂剧,有点《权力的游戏》风格,权力斗争。作为那里的管理者……我觉得现在比过去平静了一些,但回顾过去,似乎你必须学习这些技能。有些方面感觉与你的性格相反,需要处理所有这些。

You've always struck me as a very nice, level-headed guy. There's parts of OpenAI's history that are quite dramatic, soap opera-like, a little Game of Thronesy, power struggles. To be a manager in that... I feel like things are a bit calmer now than they were, but when you look backwards, it seems like you had to learn these skills. Some of this feels opposite to your personality to have to deal with all that.

Mark Chen

老实说,我在 OpenAI 一直很幸运。我真心这么说。我的经理们一直很支持我。他们看到了我的才华并为我争取。当我还是独立贡献者时,Wojciech Zaremba 说,‘哦,你应该在 Codex 上押注他。’后来,向 Bob 汇报时,我从未要求过晋升或升职。一切都是自然而然地发生的。一路上每个人都给了我很好的建议。我认为作为管理者成长的一部分就是积累经验。我认为没有比 OpenAI 更好的地方来积累经验了。总有挑战需要解决。培养那种信心。我实际上认为管理就是关于经验。其中天赋的成分较少。

Honestly, I've been lucky at OpenAI. I genuinely say that. I've had managers that have really advocated for me. They saw my talent and advocated for me. When I was an IC, Wojciech Zaremba was like, 'Oh man, you should bet on him for Codex.' And later on, reporting to Bob, I've never asked for promotion or up-level. It just organically happened. Everyone along the way has given me great advice. I think part of growing as a manager is just getting the reps. I don't think there's any better place to get the reps than at OpenAI. There's always challenges to solve. Developing that confidence. I actually think management is something where it's really just about the experience. There's less so talent involved in it.

OpenAI 危机与凝聚团队 The OpenAI Crisis and Rallying the Team

Host

我还要为我的书留一些精彩内容,所以不会全用上。但其中有几个时刻,你帮助研究人员团结起来,签署请愿书让 Sam 回来。然后大概一两天后,在 Greg 家或 Chelsea 家还有一次演讲。这两个时刻都让我觉得意义深远,尤其是为了信仰挺身而出、鼓舞士气。在危机时刻,我不知道。所以,那些……

And I'm also gonna save some of my gems for my book. So, I won't use them all up. But there are a couple moments in there where you helped get the researchers aligned around the petition to bring Sam back. And then I think just a day or two after that, there's this speech given at Greg's house, or maybe Chelsea's house. Both of those struck me as pretty profound moments, especially for standing up for what you believe in and rallying the troops. In a moment of crisis, I don't know. So, did those...

Mark Chen

是的,那对我来说确实是一个关键时刻。在风波后的几天里,充满了不确定性。当时我和 Nick、Barrett 感到责任重大,因为危机四伏。每个人都接到竞争实验室的电话,说‘你应该来这里工作’。我设定了一个目标:我不会失去任何一个人。我们做到了。每天,我们开放自己的家,让人们可以来,有个地方释放焦虑,并与领导团队保持联系。他们觉得自己能有所作为。随着时间的推移,人们真正感受到了‘我们同舟共济’的精神。我们如何做出改变?如何向世界表明我们团结一致?我开车往返于几个房子之间,我们有了一个想法:我们需要向世界展示我们完全团结一致,并且会为 Sam 工作。于是请愿书就诞生了。这个想法在凌晨两点成型。到早上,我们得到了整个研究团队超过 90% 的签名。每个人都在给朋友打电话,‘嘿,你加入还是不加入?’最后,接近 100 人签署了那份请愿书。

Yeah, that did feel like a very pivotal moment for me. In the days following the blip, there was a lot of uncertainty. Myself, Nick, and Barrett at the time felt this responsibility that the wolves are at the heels. Everyone was getting calls from competing labs saying, 'You should come work here instead.' I set this goal: I will not lose a single person. And we didn't. Every day, we opened up our houses so people could come, have a place to let out their anxiety, and stay in touch with the leadership team. They felt they could make a difference. Over time, people really felt this spirit of 'we're all in this together.' How do we make a difference? How do we signal to the world that we're all together? I was driving back and forth between a couple houses, and we had this idea: we need to show the world that we're all seriously aligned and we're going to work for Sam. That's when the petition came together. The idea got solidified at 2:00 a.m. We got more than 90% of the whole research team signed by the morning. Everyone was calling their friends, 'Hey, are you in or are you not?' In the end, it was very close to 100 people signing that petition.

Host

不过,那一定让你处境有些艰难,尤其是在一开始,因为 Ilya 和 Sam 站在对立面,而 Ilya 是你的导师。后来我知道 Ilya 又回来了。那尴尬吗?

Well, that must have put you in something of a tough spot though, especially at the outset, because it was kind of like Ilya and Sam were on opposite sides, and Ilya is your mentor. Then I know Ilya kind of comes back. Was that awkward?

Mark Chen

不,这很艰难。那是一个信息匮乏的环境。但基本上,你可以合理地推断:Sam 做了什么吗?但 Greg 和 Jakub 这些极其正直的人会为此辞职吗?我只是觉得故事有些部分被歪曲了。

No, it was hard. It's a low information environment. But fundamentally, you could very reasonably conclude: did Sam do anything here? But would Greg and Jakub, people of super high integrity, quit over that? I just felt like there was some part of the story that was being misrepresented.

Host

是啊。Jakub 在那里很久了。人们应该知道关于 Jakub 的哪些不为人知的事?

Yeah. With Jakub, he's been there for a very long time. What should people know about Jakub that they don't?

Mark Chen

有意思,因为他超级有趣。他非常搞笑,有一种讽刺幽默。这让我笑个不停。这是我现在在 OpenAI 最喜欢的事情之一:我和 Jakub 的高度一致。我们开会时,可以互相激发想法,迅速达成一致,然后传达同样的信息,共同推进大路线图的不同部分。这是我在 OpenAI 工作的一大特权。

It's interesting because he's super funny. He's hilarious. He has this sarcastic humor. It cracks me up so much. That's one of my favorite things about OpenAI today: the level of alignment I have with Jakub. We go into a meeting, we can bounce off ideas and quickly get to alignment, then deliver the same message and operate on different parts of a big roadmap together. It's one of the big privileges I have working at OpenAI.

Host

是啊,说到团结大家,我对 OpenAI 的研究仍然有这种感觉。我觉得我们仍然受到攻击。

Yeah, going to that point about keeping people together, I still feel that way about OpenAI research. I think we're still under attack.

Mark Chen

不,我们是一个家庭。我们总是受到攻击。你看,任何公司起步时,他们试图从哪里招聘?是 OpenAI。他们想要专业知识、我们的愿景、我们的世界观。我们培养了许多明星研究员。我认为 OpenAI 比任何其他地方都更能成就 AI 领域的名人。我仍然有同样的保护欲。如果有人挖他们,我会尽一切努力确保他们开心、开放,并理解他们的角色如何融入路线图。

No, we are a family. We're always under attack. Look, when any company starts, where do they try to recruit from? It's OpenAI. They want the expertise, our vision, our philosophy of the world. We've made so many star researchers. I think OpenAI more than anywhere else has been a place that makes names in AI today. I still feel that same level of protectiveness. You come after them, I'm going to do anything in my power to make sure they're happy, they're open, and they understand how their role fits into the roadmap.

Host

是的,这是我在写书或实时观察事件发展时一直在纠结的问题。当我回顾历史,有 Ilya 在 2012 年取得重大突破,Noam 在 2017 年做 Transformer,还有 Alec Radford。有时故事就是这些个体推动领域前进。这个领域还很年轻,所以可以有这样的个体。似乎有大约 8 到 10 个人有能力反复做到这一点,塑造了整个方向。所以当我看到 John Schulman 离开或 Alec 离开时,我想,‘哇,如果你失去了全明星团队的一大部分,你怎么能替代?’然而正是在那之后,你们在推理和其他方面取得了进展。我在理智上对此感到困惑。

Yeah, this is something I battled with as I was doing the book or even just watching events unfold in real time. When I go back through the history, you've got Ilya in 2012 making a big breakthrough, then Noam in 2017 doing Transformers, then Alec Radford. Sometimes the story is these individuals really pushing the field forward. It feels like a field that's still so young that you can have this individual. There seems to be a group of maybe 8 to 10 who have an ability to do that repeatedly, shaping where this is all going. So when I started seeing John Schulman leave or Alec leave, I thought, 'Wow, if you've lost a chunk of this all-star team, how do you replace that?' Yet it was after that that you guys pushed forward on reasoning and some other spots. I've had trouble intellectually with that.

Mark Chen

我不同意这是当今做好研究的总体方式。当然有很多自上而下的引导;我们押注方向。但 OpenAI 也有一种美丽的自下而上的文化,一些最好的想法会有机地从最意想不到的地方涌现。最棒的是看到这些赌注展开、成形、规模化,推理就是一个核心例子。

I do disagree with that as the overarching way to do good research today. There's certainly a lot of top-down steering; we bet on directions. But OpenAI has this beautiful culture of being bottom-up in a very deep way too, where some of the best ideas just organically emerge from sometimes the most surprising places. The great thing has been watching some of these bets unfold, take shape, get scaled, with reasoning being a core example.

Host

那么在这个想法中,我们对明星的依赖程度如何?因为你仍然看到 Google 花巨额资金把 Noam 请回来。这让我觉得事情就是这样运作的。

So in this idea, how star-dependent are we? Because you still see Google spend an ungodly amount of money to bring Noam back. That makes me think this is how it works.

Mark Chen

我认为是混合的。你必须投资于你的培养体系,因为我对我们创造明星的能力非常有信心。但当然,外面也有非常优秀的人,每个人都知道他们很优秀。

I think it's a mix. You have to invest in your pipeline because I'm very confident in our ability to create stars. But yes, there are certainly very good people out there, and everyone knows that they're good.

招聘与竞争 Recruiting and Competition

Mark Chen

我想如果说我从 Meta 学到了一件事,那就是 OpenAI 也可以非常积极地争夺顶尖人才。这种激进的招聘方法我也借鉴了一些。但我们应该始终努力组建最好的团队,以服务于我们想要完成的使命。

I think if there's one thing I've learned from Meta, it's that OpenAI can also go very aggressively after star talent. There's this very aggressive recruiting approach that I've taken a couple pages from as well. But we should always just be trying to assemble the best team in service of the mission we want to accomplish.

Host

有趣的是,这个世界相对较小,你们这些人即使是对手也会一起出去玩。这一定很奇怪,因为我知道你在某种程度上和不同的人是朋友,但同时你又在试图挖走他们的人才。

It's funny because it is a relatively small world and all you guys hang out even though you're rivals. It must be weird because I know you're friends with different people on some level and then you're also trying to steal all their talent.

Mark Chen

这是一个全方位残酷竞争的行业。但这也是我热爱它的原因。我是一个非常有竞争意识的人,讨厌失败。在研究、招聘等各个方面,我都会非常努力。

It's a brutally competitive industry in all fronts. But again, that's what I love. I'm a deeply competitive person. I hate to lose. On research, on recruiting, all of these fronts, I'll work very hard on them.

Host

这让我想起了早期的半导体时代。当时一下子涌现出许多半导体初创公司,都在推动物理学的极限。有人在某家公司发现了什么,就会去酒吧互相分享知识,但同时他们也在被挖角。每家公司都在以某种方式迅速取得突破。

It reminds me of the early semiconductor days. You had all these semiconductor startups come at once. They were all pushing the limits of physics. Somebody would discover something at one company, they'd go to the bar and share knowledge with each other, but then they're also getting pulled. Each company is quickly getting breakthroughs in one way or another.

Mark Chen

你提出了一个有趣的观点。思想总会以一定的速率扩散。我认为公司有两种应对方式:你可以建立深度的信息壁垒来保护信息。但 OpenAI 不这么运作。我们只会尽可能快地超越其他人。我喜欢开放的文化。研究人员自由地分享想法,我认为这是取得最快进步的方式。

You raised an interesting point. There is going to be some base rate diffusion of ideas. I think there are two ways a company can respond. You can create deep silos to protect information. I don't think OpenAI operates that way. We just will outrun other people as fast as we can. I love the culture of openness. People in research freely share ideas, and I think that's the way to make the fastest progress.

与 Sam 和 Jakob 的互动 Dynamic with Sam and Jakob

Host

你和 Sam、Jakob 是如何合作的?我认为人们有时能看出 Sam 更偏向研究而非公司的日常运营。研究是他的热情所在。你和 Jakob 对这些东西钻研得很深,而 Sam 则与所有人交谈。我只是好奇你们三人之间的动态。

How do you, Sam, and Jakob work together? I think people sometimes can tell that Sam is research-oriented over day-to-day running of the company. Research is more his passion. You and Jakob are so deep on this stuff, and Sam is having conversations with everyone. I'm just curious about this dynamic between the three of you.

Mark Chen

这是一个非常紧密的团队。我每天都和 Sam、Jakob 交谈。Sam 热爱研究,喜欢了解研究进展,喜欢和研究人员交流。他非常善于把握研究的脉搏。我依赖他来发现任何隐藏的潜在问题。Jakob 和我花了很多时间思考如何设计工作以取得成功。我们将具有合适优势的人配对,并激励他们朝着我们认为重要的方向努力。

It's a very tight cohort. I talk to Sam and Jakob every day. With Sam, he loves research and learning about research. He loves talking to researchers. He's very effective at getting a pulse on the research. I rely on him to surface any hidden latent problems. Jakob and I spend a lot of time figuring out how to design the work for success. We pair people with the right strengths together and incentivize them to work on directions we find important.

Host

那 Sam 呢?他是在读论文吗?还是和你们聊天?

And Sam, what does he do? Is he reading papers? Is he chatting with you guys?

Mark Chen

他读了不少论文。他和研究人员交谈,了解他们如何看待世界以及他们正在做的研究类型。当然,他还负责除此之外的许多事情。

He does his fair share of reading papers. He talks to researchers and understands how they think about the world and the type of research they're doing. Of course, he's responsible for a huge umbrella of things outside of that.

预训练焦点 Pre-training Focus

Host

现在我要问一些技术问题了。在我参加的会议中,预训练似乎是一个你们觉得已经有所突破的领域。你们认为这将是一个重大进展。它之前要么被忽视,要么是个痛点。以前事情并没有完全按照你们的预期发展。你能告诉我们你们发现了什么吗?

I'm going to ask some nerdy questions now. In the meetings I've been on, pre-training seems like an area where you guys feel like you've figured something out. You think this is really going to be a major advance. It was also either a neglected spot or something of a sore spot. Previously things weren't working exactly how you had expected. What can you tell us about what you figured out?

Mark Chen

从高层来看,过去两年我们把大量资源投入到推理上并使其发挥作用。一个副产品是你在其他功能上失去了一些肌肉,比如预训练和后训练。在过去六个月里,Jakob 和我做了很多工作来重建这些肌肉。预训练确实是一块需要锻炼的肌肉。你需要确保所有信息都是新鲜的,人们在最前沿进行优化,并且有足够的思想份额。我们认为预训练还有很大空间。很多人说 Scaling(规模扩张)已死。我们完全不这么认为。

At a high level, over the last two years we've put so much resource into reasoning and making it work. One byproduct is that you lose a little bit of muscle on other functions like pre-training and post-training. In the last six months, Jakob and I have done a lot of work to build that muscle back up. Pre-training is really a muscle you exercise. You need to make sure all the info is fresh, people are working on optimization at the frontier, and the mind share is there. We think there's a lot of room in pre-training. A lot of people say scaling is dead. We don't think so at all.

预训练 vs 强化学习与扩展 Pre-training vs. RL and scaling

Host

从某种意义上说,所有对强化学习的关注对我们来说有点像一个阿尔法,因为我们认为预训练还有很大的空间。由于这些努力,我们一直在训练更强大的模型,这让我们对今年年底推出的 Gemini 3 和其他版本充满信心。

In some sense, all the focus on RL is a bit of an alpha for us because we think there's so much room left in pre-training. As a result of these efforts, we've been training much stronger models, and that gives us a lot of confidence carrying into Gemini 3 and other releases coming this end of the year.

Mark Chen

我有时在脑海中想象的是,你们一直在飞速奔跑。整个领域都在飞速奔跑。我们正处于这样一个时刻:我们从互联网上收集了海量信息,将其投入超级计算机,然后 ChatGPT 就诞生了。然后我们就踏上了这场不可思议的竞赛。当我听到你的话时,我试图为那些不密切关注此事的人做一个水平设定。在最初的那一刻,你有这么多数据,你把它扔给机器,你最初试图稍微塑造这些数据,而现在我们正在学习更有效的塑造方法。并不总是清楚错误在哪里。

The way I picture it in my head sometimes is that you've been running so fast. The whole field has been running so fast. We're at a moment where we've gathered up this vast volume of information from the internet, thrown it onto this supercomputer, and ChatGPT pops out. Then we're on this incredible race. When I hear you, I'm trying to level-set for people who don't follow this closely. In that initial moment, you had so much data, you threw it at this machine, you tried to shape that data a bit initially, and now we're learning more efficient ways to shape it. It's not always clear what the mistakes were.

Host

你触及了我一直在思考的问题。当你考虑预训练时,你是在用人类编写的数据教模型如何模仿。它理解人类的写作模式。从某种意义上说,这也造成了瓶颈,并给能力设定了上限。当你模仿人类所写的东西时,你无法真正超越人类所写的。所以你要研究强化学习之类的东西,在那里你可以引导模型去解决人类能想到的最难的任务,让模型跳出框框思考,超越从模仿人类中学到的东西,达到更高的能力水平。但这里有一个有趣的问题:如何超越人类今天所能做到的?我也发现那里存在一个严重的测量问题。甚至从人类能否判断科学中的超人表现这个意义上说?我们怎么知道这个超人数学家比那个超人数学家更好?我们确实需要想出更好的评估方法,来衡量在这个世界上取得进步意味着什么。到目前为止我们很幸运。有像 IMO、IOI 这样的竞赛,真正衡量谁是世界上顶尖的数学家。但当模型能力超越人类时,就没有更多的测试了。

You touch on something I've been thinking about a lot. When you think about pre-training, you're taking human-written data and teaching the model how to essentially emulate it. It understands human patterns of writing. In some sense, that also bottlenecks and puts a ceiling on the capability you're able to achieve. You can't really surpass what humans have written when you're imitating what humans have written. So you work on things like RL, where you can provide steer towards the hardest tasks that humans can come up with and have the model think outside the box, outside what it's learned from imitating humans, and achieve higher levels of capability. But there is this interesting problem of how to go beyond what humans are able to do today. I find a serious measurement problem there too. Even in the sense of can humans judge superhuman performance in the sciences? How would we know that this superhuman mathematician is better than that superhuman mathematician? We really need to come up with better evaluations for what it means to make progress in this world. We've been lucky up to this point. There have been contests like the IMO, IOI, really just gauging who's the top mathematician in the world. But when model capabilities go beyond humans, there are no more tests.

Host

你让我想到了一个回到 IOI 的问题。我经常看到那些在这些比赛中表现出色的孩子被谷歌或 Facebook 这样的公司录用,但他们后来并不总是成为高管或最著名的工程师。也许这是他们的选择,但我不认为 Gennady 像迈克尔·乔丹那样最终在这些公司工作。我不清楚在那些比赛中表现出色的人是否一定是你将拥有的最伟大的工程师。所以如果一个人工智能特别优秀,我们能学到什么?

You just made me think of a question going back to the IOI stuff. Often I would see the kids who were amazing at those competitions get hired somewhere like Google or Facebook, but they weren't always the top executive or the most famous engineer afterwards. Maybe it was by choice, but I don't think Gennady was like the Michael Jordan who ended up working at any of these companies. It's not clear to me that the human who excels at that is necessarily the greatest engineer you're ever going to have. So if an AI is particularly good, what are we learning?

Mark Chen

这是我喜欢在人工智能领域工作的一个方面。我认为与标准工程文化相比,它更像一个精英制度。我以前尝试过很多次,也吸取过很多次教训,但很难让一个得不到所领导研究人员尊重的人来领导一个团队。我认为在研究领域尤其如此。你必须做出非常强硬的技术决策,比如在有分歧时这是正确的道路,这是正确的项目类型。如果你做出了错误的决策,你就会失去研究人员的尊重。在人工智能领域工作和创造强大的人工智能的一个有趣之处是,我所有的冒险都非常技术性,和他们谈论技术问题很有趣。

That's a thing I quite like about working in AI. I think more so than in standard engineering culture, it is a meritocracy. I've tried this many times before and learned this lesson many times before, but it is hard to put someone in to lead a group who doesn't have the respect of the researchers they're leading. I think this is more so the case in research than anywhere else. You have to make very strong technical calls of like this is the right path when there's a disagreement, this is the right kind of project. If you make those calls wrong, you lose the respect of your researchers. One of the fun things in working in AI and creating a strong AI is that all my ventures are very deeply technical, and it's fun to talk to them about the technical things.

Host

再谈一下预训练。对我来说,感觉 Transformer 帮助开启了这一巨大飞跃。推理感觉非常可比,甚至更令人惊叹。在过去几个月与你们交谈时,我与你们、Greg、Jakub、Sam 交谈的感觉是,你们觉得你们已经投入了三、四、五年的艰苦工程工作,但这些工作还没有完全显现出来。所以我永远无法判断该有多兴奋。当你们暗示你们看到的一些东西时,你们觉得这是否是一个与这些重大划时代事件相当的飞跃?

On pre-training again for a second. To me, it feels like Transformers helped kick off this massive leap. Reasoning feels very comparable, if not even more amazing. When I talked to you guys over the last few months, my sense when I talk to you, to Greg, to Jakub, to Sam, is that you guys feel like you've been putting in hard engineering work for three, four, five years that hasn't fully manifested itself. So I can never tell how excited or not to be. When you guys are hinting at some of the stuff you're seeing, do you feel like it is a comparable leap forward to these big epochal things?

Mark Chen

我认为是的。当我们推出 GPT-5 时,我们也谈了很多关于合成数据的内容。还有许多其他类似形式的线索,我们认为它们很有前景,并且我们现在正在大力扩展。这始终是关于保持那个赌注组合,选择那些提供更多经验前景的,并以更大的程度扩展和支持它们。

I think so. When we launched GPT-5, we talked a lot about synthetic data as well. There are many other threads of this form that we think hold quite a bit of promise and that we're scaling up pretty aggressively right now. It's always about maintaining that portfolio of bets, taking the ones that are providing more empirical promise and scaling and supporting them at an even greater degree.

Host

两周前,曾在 OpenAI 工作的 Andrej Karpathy 上了一个播客,似乎给人工智能行业泼了一大盆冷水,说 AGI 还要 10 年。然后我一周前听到 Dario 的讲话。他似乎仍然坚持大规模的科学发现,他的“天才之国”。他似乎仍然坚持可能稍微慢一点,但大约两年时间表。当你听到 Andrej 的话时,你怎么想?

Two weeks ago, Andrej Karpathy, who used to work at OpenAI, went on a podcast and seemed to deflate some giant portion of the AI industry by saying that AGI was like 10 years off. Then I heard Dario talking about a week ago. He seemed to be holding on very much to massive scientific discoveries, his 'nation of geniuses'. He seemed to be holding still on kind of maybe a little slower but like a two-year timeline on that. When you heard what Andrej said, what did you think?

Mark Chen

是的,我的意思是,推特喜欢这种“已经结束了”的循环。

Yeah, I mean, Twitter loves this cycle of 'it's so over'.

定义 AGI 与科学进步 Defining AGI and Scientific Progress

Mark Chen

所以,你知道,无论当时什么叙事流行,都会被放大。我在试着剪一段,但我觉得,AGI 嘛,每个人都有自己的定义。即使在 OpenAI,你也没法让所有人坐在一个房间里说,嘿,这是我的明确定义,而且大家都一致。所以我把它看作类似工业革命:你觉得机器制造纺织品算工业革命,还是蒸汽机算?每个人定义不同。我认为我们正处于创造 AGI 的过程中。对我来说,我最看重的是我们是否在产生新的科学知识,是否在推进科学前沿。我觉得从夏天开始,这方面发生了巨大的转变。

So back and you know whatever plays into the narrative at the time I think you know just becomes amplified. Yeah, I'm trying to make a clip here, but you know, I the way I think about it, yeah, I mean, it's like AGI, I mean, everyone defines their own point for AGI, I I think even at at OpenAI, um, you can't get everyone in the same room and be like, hey, this is my clear definition of AGI and it's it's it's consistent. And so, I kind of think about it as something like, you know, you're in the industrial revolution, right? Do you consider the, you know, having machines make textiles, is that the industrial revolution or is it the steam engine? you know, everyone kind of has their different definition and um I think we're in the middle of this process of producing AGI. For me, I think the thing I index most on is are we producing novel scientific knowledge and are we advancing the scientific frontier and I feel since the summer there's been a tremendous phase shift on that front.

Host

好吧。从你看到的东西来说,我脑子里最先跳出来的都是那些生物技术领域的初创公司,它们展示了一次性抗体和分子之类的,但我不知道那是不是……

Okay. like from stuff that you're seeing in ter the the first things that are jumping to my head are all these startups that are in the biotech space that are showing you know oneshot antibodies and and molecules but I have no idea if that like what are you

Mark Chen

是的,是的。那次与物理学家的会面让我深受启发,他们回去后想,嘿,我们应该为科学创建 OpenAI。目标是,对于今天少数意识到这些模型潜力、想要投入并加速的科学家,我们应该尽最大努力加速他们。我知道其他公司也有类似推动科学前沿的努力,但我想做的是——与谷歌的科学工作相比,我们的不同之处在于——我们希望让每个人都有能力为自己赢得诺贝尔奖。不是 OpenAI 自己获奖(那当然好),而是构建工具和框架,让所有科学家感受到加速的影响,我们认为可以共同推动这个领域。

yeah yeah so I mean I was so inspired by that encounter with the physicists that you know went back and thought hey well we should just create open AI for science and the goal being I think for the small set of scient scientists today who realize the potential of these models and feel like they want to lean in and accelerate, we should do the best that we can to accelerate them. And you know I know there are similar efforts that you know other companies um aim towards pushing the scientific frontier but I think what we want to do and um I would say a little bit of a framing in terms of how we differ from let's say Google's uh efforts to to to work on science is we want to allow everyone uh the ability to you know win the Nobel Prize for themsel. Um, it's less so about us winning that at at OpenAI, which would be nice, but we want to build the tooling and the framework so that all scientists out there feel that accelerative impact and we think we can push the field collectively.

Host

那么,你说的让你兴奋的发现呢?有没有其他具体的例子?

Well, and when the discoveries that you're saying you're excited about? I mean, are there any others like specifically that that you've

Mark Chen

是的,是的。如果你想要一大堆例子,可以去 Seb 的推特账号看。最近有一篇关于开放凸优化问题的 GPD5 论文……

Yeah. Yeah. So, um, I think there's, you know, if you want a a huge list of of these, um, you can go on Seb's Twitter account. Um so recently you know there's a GPD5 paper on an open convex optimization problem that you know uh is actually whose Twitter account?

Host

哦,Sebastian。好的。是的。它和我们正在解决的一些核心机器学习问题非常相关。我知道还有……

Uh Sebastian. Okay. Yeah. Yeah. And you know it's like um very related to some of the core ML problems that we're solving. Um I know there was um

Mark Chen

我觉得人们有点轻视这些,觉得不过是花哨的文献搜索之类的。其实要复杂得多。我可以举一些例子,但……

I think people kind of dismiss these things as oh is it just fancy literature search or something like that? Um it's quite a bit more complicated than that. And you know, there's some examples I could go into, but

Host

说实话,我现在有点 overwhelmed,因为我是个通才,但主要报道生物技术。每两天就有人跟我说,哇,我们在造 AI 科学家,我们一次性搞定了增强型抗体之类的。我一部分很兴奋,至少有几家公司我认识人,他们是真正的科学家。但太多这样的消息了,我要么觉得有大事发生,要么就是太多,我分辨不出哪些是真实的。

I might I'm honestly overwhelmed at the moment cuz, you know, I'm sort of a generalist, but I cover biotech a lot. And it's like every two days, man, I'm walking in and it's wow, we we're making an AI scientist. We we one-shotted enhanced body and and then so like part of me gets excited and you know at least a handful of these companies I know the people and they're real scientists and like but then there's so much of it that I'm like either something amazing is happening or every it's it's like it's it's kind of too much for me to be able to discern where reality is.

Mark Chen

是的。如果生物学领域也在发生这些,我不会惊讶。我个人最擅长计算机科学和数学,我们有专家可以确认这些是真正的发现。这让我最有信心,但生物学领域发生这些我一点也不意外。

Yeah. I mean I wouldn't be surprised if it's happening in biology. Um personally I have the most expertise in you know computer science and mathematics and you know we we do have the experts there that can confirm that these are discoveries being made. So that's the thing that gives me the most confidence but I'm not surprised at all it's happening in biology.

Host

但你说的和那种每三周就变一次的叙事不太一样。最大的批评是,甚至在 Andre 说之前,我听到一个政治播客,好像是 Breaking Points,主持人是个挺聪明、知识渊博的人,但他一直在说 AI 缺乏进展,都是虚构的。所以,如果这些发现没有发生,我觉得公众是知道的。

But like what you're saying is is kind of different than the I gr the narrative changes every like three weeks it seems like. But like what you're saying is is sort of different because the biggest knock even before Andre said that it seemed to me from the you know what I was listening to um I was listening to like a politics podcast um sagger I think it's breaking points is their podcast you know he's he's pretty smart guy who's knowledgeable but I mean he's just been on AI and the lack of progress and this is all like make believe and and all in and So, you know, if these discoveries aren't happening, I mean, I feel like the public is aware of this.

Mark Chen

澄清一下,在建立 OpenAI for Science 的过程中,我们和很多物理学家、数学家聊过,实际上大多数人并不看好 AI。他们仍然认为这东西不能解决新定理,绝不可能。肯定有别的门道。所以我觉得,赋能那些真正相信并投入的人,他们会超越其他人。我们想构建工具,说服人们这是做科学研究正确的方式。

Just to be clear, you know, um, while setting up Open AI for science, we've talked to a lot of physicists, a lot of mathematicians, and actually most of the people we've talked to aren't that bullish on AI. I think they still believe, hey, you know, um, this thing isn't something that can solve new theorems. There's no way it could do that. You know, there must be something else going on. And that's why I feel like empowering the set of people who really do believe and lean into it like those people are going to just you know outrun everyone else and we want to build the tools and convince people like this is the right way to do scientific research.

Host

好吧。在这一点上,我承认每个人对 AGI 的定义不同,但至少我听到的是,不管你怎么称呼它,你觉得未来一两年我们会看到戏剧性的事情发生。

Okay. And so I mean so like on that point I mean I grant you that everybody's definition of AGI is different but you're like at least what I'm hearing is I mean you whatever you want to call it um you feel like in the next year or two is we're just seeing dramatic things happen.

Mark Chen

是的。这有点像个梗,对吧?你问别人 AGI 什么时候来?两年后。但我不认为我们还在那个世界里。是数学和科学上的这些结果给了我这种信念。在 OpenAI 内部,我们设定了两个非常具体的目标:一年内,我们要改变研究的性质,在研究开发过程中有效地依赖 AI 实习生;两年半内,我们要让 AI 做端到端的研究。

Yeah. I mean it is a bit of a meme, right? It's like you ask someone when is AGI? It's two years away, right? Um and I don't think we're in that world anymore. And it it's like these results in math and science that that are giving me this conviction. But at at OpenI within the research, we set two very concrete goals, right? Within a year, we want to change the nature of the way that we're doing research. And we want to be productively um relying on AI interns in in the research development process. And within 2 and 1/2 years, we want AI to be doing end-to-end research.

AI 交互与设备设计的未来 Future of AI interaction and device design

Mark Chen

我认为这非常不同。今天,你提出一个想法,然后执行、实现、调试。这意味着在一年内,我们很有信心能达到这样一个世界:我们控制外循环,提出想法,但模型负责实现和调试。

And I think it's very different right like today you come up with an idea you execute on it you implement it you debug it. It means within a year we're quite confident we can get to a world where we control the outer loop we come up with the ideas but the model is in charge of the implementation the debugging.

Host

好的。除了预训练,当我跟你们聊天时,我有时会有类似的感觉:我们脑子里都有这种想法,至少我身边的人是这样——大规模基础设施已经建成,模型每次扩大 10 倍似乎都会变得更好。有一段时间的故事是,当你们从四代到五代时,尽管算力增加了,却没有看到想要的结果。但后来我跟你们聊得越多,越觉得你们认为实际上我们还没有——当时事情发展得太快,我们还没有真正看到跃升至 10 倍算力的那个时刻。我不知道我这个问题问得是否清楚。

Okay, are there beyond pre-training when I talk to you guys sometimes I get the sense similar sort of thing it's like we all have in our heads at least people where I sit that there's been this massive infrastructure build out that the models seem to get better every time you 10x them that there was a story for a while that as you guys were going from like four to five you weren't seeing the results you wanted even though you were getting more compute but then the more I talked to you guys the more it sounds to me like you feel we haven't actually that things were moving so fast back then that we haven't actually seen the moment where we made the leap to the 10x compute. I don't know if I asked that question very eloquently but

Mark Chen

是的,我确实有个想法想分享。当人们问我‘你们真的需要这么多算力吗?’时,这个问题让我很震惊,因为每天我都在处理大量的算力请求。我的心态是,如果今天我们有 3 倍的算力,我可以立即非常有效地利用它;如果今天我们有 10 倍的算力,可能在几周内就能高效地充分利用。所以我认为对算力的需求是真实存在的。我没有看到任何放缓。当我听到人们问‘哦,你们真的需要更多算力吗?’时,我几乎感到困惑。是的,这对我来说毫无意义。

Yeah, I mean I do have a thought to share here which is you know when people ask me like do you guys really need all this compute it's such a shocking question because day-to-day I'm dealing with so many compute requests and you know the really my frame of mind is you know if we had 3x the compute today I could immediately utilize that very effectively if we had 10x the compute today probably within a small number of weeks fully utilize that productively. And so I think the demand for compute is really there. I don't see any slowdown. And yeah, it almost baffles me when I hear people ask like, 'Oh, do you guys really need more compute?' Yeah. Doesn't make sense to me.

Host

你认为,在我问得不好的问题的大方向上,你们似乎对预训练上的突破非常乐观,那么你们是否同样——不仅仅是人们想要更多 GPU 的需求——而是你是否清楚地看到,同样的 Scaling 即将把事物推向更高?

And you think in the broad strokes of the question I asked badly. Do along the lines of where you guys seem very optimistic about what you've cracked on pre-training, are you equally not just like this demand that people want more GPUs, but are you do you see pretty clearly that that same thing scaling is about to kick things higher?

Mark Chen

是的,我们绝对想继续 Scaling 模型,我认为我们有算法突破使我们能够 Scaling 模型。而且,我认为 Gemini 3 有很多令人印象深刻的地方。我注意到的一个细节是,当你看到他们的 SWE 基准测试数字时,他们在数据效率方面还没有突破,对吧?他们没有取得太大进展。而我认为我们在那方面有非常强大的算法。

Yeah, we absolutely want to keep scaling the models and I think we have algorithmic breakthroughs that enable us to scale the models and, you know, I think there's a lot impressive about Gemini 3. One thing that kind of reading into the details that I've noticed is, you know, when you look at stuff like their SWE bench numbers, there's still a big thing around data efficiency that they haven't cracked, right? They haven't made that much movement on it. And I think we have very strong algorithms there.

Host

是的。而且有一份泄露的备忘录,我是说,Sam 对 Gemini 3 听起来相当悲观。在这份备忘录里,我试着找到引用。你显然——我肯定你看到了那份备忘录。它似乎是一个时刻。是的。

Yeah. Well, and there was this leaked memo from I mean, Sam was sounding quite somber about Gemini 3, man. In this memo, I'm trying to find the quote. Did you well, you obviously I'm sure you got the memo. It like it seemed like a bit of a moment. Yeah.

Mark Chen

嗯,我确实认为 Sam 的部分工作是注入紧迫感和节奏,这也是我工作的一部分。我认为我们必须高度专注于 Scaling,而且我确实认为 Gemini 3 正是谷歌应该追求的那种赌注。同时,我会补充说,我们工作的很大一部分是尽可能向组织注入紧迫感。

Well, I do think part of Sam's job is to inject urgency and pace and that's also part of my job as well. I think it is important for us to be laser focused on scaling and I do think Gemini 3 is exactly the right kind of bet that Google should be pursuing. At the same time, I would calibrate that by saying a large part of our jobs is to inject as much urgency into the org as possible.

Host

而且它是一个好模型。我认为我们有应对之策。而且我认为我们可以更快地执行后续步骤。

And it is a good model. I think we have a response. And I think we can execute even faster to the follow-up.

Host

你在多大程度上参与像——我肯定你会告诉我与 Jony Ive 的设备具体是什么样子——这样的事情?

How much do you get involved with things like and I'm sure you're going to tell me exactly what it looks like with Johnny Ives device.

Mark Chen

酷。酷。是的。

Cool. Cool. Yeah.

Host

是的。比如,那是研究发挥作用的领域吗?

Yeah. Like is that is that an area that research plays in?

Mark Chen

是的,是的,确实是。实际上我昨天刚吃了晚饭。

Yeah. Yeah, it is. And actually I was just having dinner yesterday.

Host

如果你想的话,可以给我描述一下。

You can describe it to me if you want.

Mark Chen

是的,当然。所以它看起来是这样的。嗯,昨天我和 Jony 以及一些研究人员共进晚餐,还有我们的预训练和后训练负责人。实际上,我对未来 ChatGPT 的思考是这样的:今天,当你与 ChatGPT 互动时,我觉得它非常笨拙。它不像是原生的思考方式,对吧?你给它一个提示,它给你一个回应,然后它不会为你做任何有效的工作,直到你给出下一个提示。如果你给它一个类似的提示,它会花同样的时间思考,它并没有因为你问了第一个提示而变得更聪明。我认为未来将是一个记忆成为大大改进的功能的世界。每次你使用 ChatGPT,它都会深入了解你。它会反思你为什么问这个问题、相关问题等等。然后下次你使用它时,它会变得更聪明。我认为这引出了一个问题:如何设计一个以这为主导理念的设备。是的,我认为这是一次非常富有成效的经历。

Yeah, absolutely. So it looks like this. Well, yesterday I was just having dinner with Johnny, with some researchers as well, our head of pre-training and also post-training. And really the way I think about ChatGPT in the future, right, today when you look at how you interact with ChatGPT, it feels very dumb to me. It doesn't feel very thinking native, right? And you go to it with a prompt, you get a response, and then it's doing no productive work for you until you give it the next prompt. And if you give it a similar prompt, it's going to think for the same amount of time, it hasn't gotten smarter because you added the first prompt. And I think the future is going to be a world where memory is going to be a much improved feature. Every time you go to ChatGPT, it learns something deep about you. It reflects about why you would ask this question, related questions, anything. And then the next time you go to it, it's going to be that much smarter. And I think it really begs the question of how do you design a device that has this as the dominating thesis. And yeah, I thought that's been a very productive experience.

Host

你有一个吗?

Do you have one?

Mark Chen

我有没有一个?我可能有,也可能没有。

Do I have one? I may or may not have one.

Host

当我想起你们和 Jony 交谈时,我在想,在苹果,你有一家以硬件为中心的公司。这是史蒂夫·乔布斯一直痴迷的东西。就像,你知道,这是一种工艺,一种艺术形式。无论是你、Sam、Greg、Yakob 还是其他人,据我所知,你们中没有人真正做过硬件产品。Sam 似乎非常重视设计,我可以从他房子的建筑等方面看出来。但是,你知道,没有什么可说的记录。我一直认为史蒂夫·乔布斯有品味,然后我过去有过几个老板,比如 Josh Taring,他曾经营 BusinessWeek。他给我的印象就是那种人。他很有品味,无论是东西的外观,还是故事应该怎么写。这是一种与生俱来的、非常高层次的东西。我觉得这大概就是这里所需要的。我想这就是为什么你们在某种程度上需要像 Jony 这样的人,但你们必须要有这种来回的交流。我们怎么知道你们中任何人有没有品味,能不能塑造一个硬件产品呢?

What I think about when I think about you guys talking to Johnny is that at Apple, you had this company that was centered around hardware. It's something that Steve Jobs obsessed about all the time. It's like, you know, it's a craft. It's like an art form. Whether it's you, Sam, Greg, Yakob, whomever, as far as I'm aware, none of you guys have really done a hardware product before. Sam seems to take design very seriously. I could tell from his the buildings in his house and things like that. But, you know, there's no sort of track record to speak of of like I always thought of Steve Jobs as having like taste, you know, and then I've had a couple bosses over the years like Josh Taring used to run BusinessWeek. He kind of he just always struck me as this guy. He just had taste, you know, whether it was the way something looked, the way a story should be. There was like this innate thing that was on this really high level. It strikes me that's kind of like what's required here. I guess that's why you have someone like Johnny on some level, but but you have to have this like back and forth. How do we know that like any of you guys have taste and and are, you know, can shape a hardware product?

Mark Chen

老实说,我们不需要自己有品味,那是 Jony 的工作。他是我们在品味上的判别器。

Honestly, we don't need to have taste ourselves and that is Johnny's job. He's our discriminator on taste.

设计与研究相似之处 Design and Research Parallels

Mark Chen

实际上,有一件事让我感到非常欣慰,那就是意识到他们在设计上的工作方式与我们在研究上的工作方式有着深刻的相似之处,对吧?有大量的探索和构思,你探索各种假设。你慢慢来。然后你创造出你满意的成果,最终的产物。让他们融入公司真的很好,而且关于我们将要发布的能力、产品形态以及如何将它们融合在一起,沟通也更加直接了。

And I think actually one thing that's been really nice is just realizing that the way they work in design and the way we work in research is there's some deep parallels there, right? There's so much exploration and ideation and you explore a bunch of hypotheses. You take your time. And then you create the thing that you're happy with, the artifact at the end. It's been really nice to have them fold into the company and there's a lot more direct communication about here's the capabilities that we're going to ship and here's what the form factor looks like and how to gel them.

Host

好吧。这么说可能有点粗俗,但我一生都在崇拜和与这些人交谈,但有时我会想,天哪,我不知道一群数学书呆子是不是你想要的制造 AI 计算机的人,但我想这就是你所说的融合。

Okay. And this is a crass way to put this but because I spend my life adoring and talking to these people, but sometimes I'm just like, man, I just don't know if a bunch of math nerds are the ones that you want making the AI computer, you know, but I guess it is this blend that you're talking about.

Mark Chen

是的,老实说,你说得对,最擅长构建 AI 能力的人与最有品味的人略有不同。我们确实有由对模型行为有极好品味的人组成的团队。我认为这是一种非常不同的哲学,以及一套你需要不断问自己的非常不同的问题。一个好的品味问题的一个例子,比如在模型行为面试中,就是 ChatGPT 最喜欢的数字应该是什么?

Yeah, honestly, you're right in that the people who are the best at building AI capabilities are slightly different from the people who have the best taste. And we do have teams built of people who have really great taste for model behavior. And I think there's a very different kind of philosophy and a very different kind of set of questions you need to keep asking yourself. One example of a good taste question, like in the model behavior interview, is like what should ChatGPT's favorite number be?

Host

它应该是什么?ChatGPT 最喜欢的数字?

What should it be? ChatGPT's favorite number?

Mark Chen

我认为它最喜欢的数字应该是什么。

What I think its favorite number should be.

Host

嗯,我有一个愚蠢的答案,那就是我上过波莫纳学院,47 是那里的一个传说数字。所以……

Well, I have a stupid answer which is that I went to Pomona College and 47 is this number of lore there. So...

Mark Chen

好吧。是的,这是个好答案。

Okay. Yeah, that's a good answer.

突破的小而脆弱想法 Small Fragile Ideas for Breakthroughs

Host

我马上要放你走了。你真的很慷慨。我很感激。有没有……我要问你一个 ChatGPT 让我问你的问题:如果你回顾五年,你现在看到的有没有什么微小、脆弱、萌芽的想法,你的直觉告诉你它们可能是一个重大突破的核心?

I'm going to let you go in a second. You've been really generous. I appreciate it. Is there... I'm going to ask you a question that ChatGPT told me to ask you, which is: if you look back in five years, are there any kind of small, fragile ideas that you're seeing right now that your instinct is telling you might be at the heart of a big breakthrough?

Mark Chen

是的,有几个。我会说有一小撮想法。我不能透露太多细节,但是的,我真的很兴奋能扩大它们。

Yeah, there's a couple. I would say a handful of ideas. I can't go into too much detail on them, but yeah, I'm really excited to scale them up.

Host

有没有任何提示,它们属于哪些领域?

Are there any hints, any buckets of areas where these fall?

Mark Chen

是的,我的意思是,我一直非常专注于预训练。所以,一些与预训练相关的想法,强化学习中的少量想法,以及一些关于如何将它们整合在一起的想法。

Yeah, I mean I've been concentrating a lot on pre-training. So, some pre-training adjacent ideas, a small number of ideas in RL as well, and a small number of ideas of how to put it all together.

Host

好吧。好的。我试过了。所以,你可能有一个设备也可能没有。不,没有提示。

Okay. All right. I tried. So, you may or may not have a device. And no, no hints.

Mark Chen

是的。没有提示。

Yeah. No hints.

澄清 OpenAI 真相 Setting the Record Straight on OpenAI

Host

好吧,我们覆盖了很多内容。我真的很感激。有没有……我觉得我有点让书呆子们失望了。就 AI 狂热者而言。有没有什么技术上的事情,你觉得人们目前对你们有误解,你想澄清一下?

Well, we covered tons of ground. I really appreciate it. Is there... I feel like I'm letting the nerds down a little bit. As far as the AI obsessives. Are there any technical things you see people getting wrong about you guys at the moment that you would set the record straight on?

Mark Chen

是的,我认为最重要的事情是,OpenAI 研究部门的任何人都会告诉你,这只是一家以研究为中心的公司。这是一个纯粹的 AI 赌注。公司的核心目标是构建 AGI,不受干扰地构建它。我认为在构建产品方面,一切都非常容易地源于此。至于我们在研究方面想做什么,我们想自动化 AI 研究。自私地说,我们想加速我们自己的进步,然后我们想自动化科学发现,当然我们还想自动化从事经济上有用工作的能力。我认为所有这些支柱都在到位,你看到去年最大的更新就是在自动化科学研究这第二个支柱上。它正在发生。

Yeah, I think the most important thing is just that anyone at OpenAI in research would tell you that it is just a research-centric company. It's a pure AI bet. At the core of the company, the ambition is to build AGI, to build it without distractions. And I think anything when it comes to building products, it all flows very easily from that. When it comes to what we want to do in research, we want to automate AI research. I think selfishly, we want to accelerate our own progress, and then we want to automate scientific discovery, and of course we want to automate the ability to do economically useful work. I think all these pillars are falling into place, and you see that the big update in the last year has just been in that second pillar of automating scientific research. It's happening.

Host

你现在多大了?

How old are you now?

Mark Chen

34 岁,快 35 了。

34, about to turn 35.

Host

快 35 了。好吧。你还能有社交生活吗,还是你……

About to turn 35. Okay. Are you able to have like a social life or are you...

Mark Chen

不,老实说没有。我想过去两周每天都是工作电话到凌晨一两点。但我喜欢这样做。有很多工作要做。有很多我想招募的人。有很多指导需要做。为什么要浪费这个黄金时刻?就像我们正处于一场工业革命之中,你必须尽可能多地利用它。

No, honestly not. I think every day the last two weeks, it's been work calls till 1, 2 a.m. But I love doing it. There's a lot of work to get done. There's a lot of people I want to recruit. There's a lot of steering that needs to be done. And why waste this golden moment? It's like if we're in the middle of something like an industrial revolution, you got to take as much advantage of it as possible.

Host

我听说过你睡在办公室的故事。

I hear stories about you sleeping at the office.

Mark Chen

哦,是的,那也挺有趣的。老实说,我认为公司里有些时候是这样。我想那是在 Mira 离开并去创办自己的公司之后。工作需要这样。当我剥开一切,审视那种深层情感时,那只是对研究的保护欲。

Oh yeah, that was a fun one too. Honestly, I think there are times in the company. I think that was right after Mira left and went to found their own company. The job demands it. When I peel it all back and examine that deep emotion, it's just this protectiveness of the research.

Host

那是在 Mira Murati 离开之后。

That was after Mira Murati left.

Mark Chen

是的。我大概睡了一个月的办公室。就像我需要保护研究。它们感觉就像,你知道,感觉像我的孩子。

Yeah. I spent a month kind of sleeping in the office. It's just like I need to protect the research. They feel like, you know, it feels like my baby.

驾驭浪潮与竞争 Navigating Waves and Competition

Host

所以,你们经历了这些浪潮。有政变。每个人都想挖走你的人。我想每个人一直都在试图挖走你的人,但你有一个转折点。Mira 离开,Meta 决定要启动这个庞大的实验室。你认为我们已经过去了吗?每个人都已经出手了吗?

So, you guys have gone through these waves. There's the coup. Everybody's trying to steal your people. I guess everybody's trying to steal your people all the time, but you have this inflection point. Mira leaves, Meta decides they're going to fire up this massive lab. Do you think we are past that? Has everybody fired their shot at this point?

Mark Chen

你知道,我有员工会议,对吧?我和我的下属谈话,我说,好吧,这是我现在正在做的事情,一旦我完成这个线程,我就会退一步,然后就没有火了。不,我现在已经完全内化了,构建 AGI 的赌注足够高,总是会有事情发生。我认为重要的是能够在所有这些事情发生的过程中理解什么是重要的事情。

You know, I have a staff meeting, right? I talk to my reports and I'm like, okay, here's the thing that I'm working on, and once I get back to once I'm done with this thread, I'm going to zoom out and there's no more fires. No, I've fully internalized at this point that the stakes are high enough for building AGI that there's always going to be something. And I think the important thing is just being able to understand what the important things are in the midst of all these things going on.

Host

你喜欢吗?自从 DeepSeek 时刻或类似的事情发生以来,已经过去几个月了?我想大概是 2024 年 12 月。

Do you like months have passed since there was sort of the DeepSeek moment or whatever? I guess it was like December 2024, I think.

Mark Chen

是的。今年早些时候。是的。

Yeah. Earlier this year. Yeah.

Host

是的。或者一月。

Yeah. Or January.

对 DeepSeek 与开源模型的反应 Reaction to DeepSeek and open source models

Host

我的意思是,现在有没有什么感觉人们一度失去了理智?回想一下,再看看他们后来的表现,谈谈你对开源模型和中国开源模型的看法。

I mean, is there anything now, it felt like people lost their minds for a second? Just reflecting on it now and seeing what they've done since, just thoughts on open source models and Chinese open source models.

Mark Chen

是的。我认为那是第一次让我意识到坚持我们自己的研究计划有多重要。当那个模型出来时,它迅速走红了,对吧?每个人都在说,‘哦,天哪,OpenAI 是不是迷失了方向?这些模型要赶上来了吗?’ 我们该怎么回应?我认为我们做对了,就是坚持我们自己的研究计划。我完全不认为那是错误的决定。我还没看到 DeepSeek 的后续模型。我认为他们是一个非常强大的实验室,但根本上,我们还是专注于创新。DeepSeek 很好地复现了我们 O 系列模型的想法,但我们还是要专注于创新。

Yeah. I think that was one of the first points in time when I realized how important it is that we just stay true to our research program. When that came out, it went viral, right? Everyone was like, 'Oh man, has OpenAI lost its way? Are these models catching up?' And what's the response? I think rightfully the thing we did was we just stayed on our own research program. I don't think it was the wrong call at all. I haven't seen the DeepSeek follow-up model. I think they're a very strong lab, but fundamentally, let's just keep focusing on innovating. DeepSeek was a great kind of replication of the ideas in our O series of models, but let's just focus on innovating.

团队规模与人才密度 Team size and talent density

Host

你认为 500 人这个数字会随着公司成长而增加,还是说这是同时追求大创意的最佳人数?

Do you think 500 people is the number that grows as the company grows, or is this like the optimum number for big ideas you can chase at one time?

Mark Chen

不,老实说,我觉得甚至可以用更少的人。而且,随着我们招募 AI 研究员或 AI 实习生,如何围绕这一点进行设计确实是个问题。但我非常看重高人才密度。我喜欢做这类实验。例如,今年第二季度,我想,嘿,我就不开放任何研究岗位了。如果你想招人,你得先弄清楚谁不适合。我认为这类练习非常重要。你不想扩散到无法管理的地步,而且你要保持很高的人才门槛。

No, honestly, I feel like it can be done with even less. And again, as we get AI researchers or AI interns, there's a real question of how do you design around that. But I'm certainly a person who cares a lot about heavy talent density. I like to run a lot of experiments of this vein. For instance, in quarter two of this year, I thought, hey, I'm just not going to open up any headcount for research. If you want to hire people, you got to figure out who's not on the boat. I think these kinds of exercises are quite important. You don't want to diffuse into something that's not manageable, and you want to keep the talent bar very high.

AI 研究的功劳归属 Credit and attribution in AI research

Host

我记得在一次会议上。我觉得你和 Jakoba 在这点上意见一致,但我肯定记得你。关于项目归属的问题,你似乎认为人们过于纠结于此了。显然,AI 源于学术界,在那里你有一篇论文就非常自豪,归属问题非常重要。我们是否进入了一个新阶段,这不再那么重要了,或者只是因为这是一家公司,谁做了什么不那么重要了?

I remember being in a meeting. I think you and Jakoba were kind of on the same page here, but I remember you for sure. This idea of who gets attribution for a project, and you seem to be of the stance that people are obsessing over that a bit too much. Clearly AI has its roots in academia where you are very proud when you have a paper and attribution is a huge thing. Have we reached a new stage where that is less of a big deal, or is it just that this is a company and who did what is less important?

Mark Chen

我其实非常喜欢这个话题。过度关注功劳分配是一件非常糟糕的事情。但另一方面,我确实觉得作为一家公司,在内部和外部认可功劳很重要。很多公司都回避这一点。整个行业已经不再发表论文的致谢名单了。但 Jakob 和我最终决定在 OpenAI 这样做。反对意见总是说,你这是在把你的顶尖人才拱手让人,其他人会积极挖角。但我认为这并不重要。我们应该认可那些做出出色工作的人。我们应该继续成为培养 AI 超级明星的管道。为那些在公司做出最佳工作的人树立名声对我们来说很重要。

I actually really love this topic. Overfixation on credit is a very bad thing. But on the other hand, I actually feel like it's important as a company for us to recognize credit both internally and externally. A lot of companies have shied away from this. We've moved away from publishing papers credit lists broadly throughout the industry. But Jakob and I ended up making the call that we're going to do it at OpenAI. The counterargument is always that you're handing your top performers on a platter, everyone else is going to be recruiting these guys aggressively. But I don't think that's important. We should just recognize the people who are doing great work. We should continue to be this pipeline for creating AI superstars. It's important for us to make names for the people who are doing the best work at the company.

Host

但你似乎也在说,个别研究人员也许不应该太纠结于此。还是我完全记错了?

But you seem to also be saying the individual researchers should maybe obsess about this less. Or am I totally misremembering?

Mark Chen

不,我认为当时房间里确实有那种情绪。实际上,Yakob 和我对此持更多反对意见。

No, I think there was a sentiment in the room of that form. Actually, Yakob and I held more of a dissenting view on that.

Host

好的。有一阵子了。我的笔记里有。完美。

Okay. It's been a while. It's in my notes. Perfect.

Mark Chen

是的。但我认为我们必须给予应有的认可,即使冒着让所有人都知道我们顶尖人才是谁的风险。

Yeah. But I think we got to give credit where it's due, even at the risk of everyone knowing who our top talent is.

Host

好的。我要说一个更强烈的观点:我认为 OpenAI 是人均外部认可最多的公司。

Okay. I will make an even stronger statement that I think OpenAI is the place where we allow for the most external credit per capita.

Mark Chen

遥遥领先。

By a large margin.

Host

好的。好吧。我会查查我的笔记。现在我有了更多信息。

Okay. All right. Well, I'll check my notes. Now I've got more.

Mark Chen

当然。

Absolutely.

从事 AI 的个人动机 Personal motivation for working on AI

Host

好的,最后一个问题,我发誓。你 2018 年加入。那是一家研究公司,非营利。公司创始之初是为了制衡谷歌,目标是确保 AGI 安全到来。你来自高频交易,看到了这些有趣的事情发生。你有多……我肯定你会说你想让这安全发生。我明白。但看看你的职业道路,你是一个聪明、好奇的人,看到了这个有趣的事情。并不是说你必须从哲学上关心这个,或者想看到超级智能。但不管怎样,让我们听听你最初为什么做这个?

Okay, last question, I swear. You got there in 2018. It was a research company, a nonprofit. The company started among the founders as a counterweight to Google, with the goal of making sure AGI arrived safely. You came from high-frequency trading and saw these interesting things happening. How much in your... I'm sure you're going to say you want this to happen safely. I get that. But if you look at your career path, you're a smart, curious human who saw this interesting thing happening. It's not like a requirement that you really give a philosophically about this or want to see superintelligence. But anyway, let's hear from you on why are you doing this in the first place?

Mark Chen

是的。

Yeah.

安全与对齐挑战 Safety and Alignment Challenges

Host

所以我认为,在安全和对齐方面,我也管理过 OpenAI 的对齐团队,老实说,我觉得未来一两年最大的挑战之一就是对齐。对于关注这个研究领域的人来说,OpenAI 在过去一年里可能做得最好。我这么说是因为我们在很多方面做了大量工作,比如欺骗行为。你给模型注入的强化学习算力越多,就越能测量出自我意识、自我保护,甚至模型可能欺骗的情况。这很可怕,因为模型最终能给你正确的答案,你期望的答案,但却是以一种非常扭曲的方式得出的。我认为,随着模型为我们执行更复杂的任务,理解它的思维过程将变得极其重要。好的,聊天机器人让我问你一个与此相关的问题:你提到的机械可解释性领域,我们试图理解这个黑箱及其运作方式。问题的核心是:我们在这方面的能力能否跟上 AI 系统的复杂性,还是说我们会达到一个失控点,永远无法理解这个东西是如何工作的?

So I think really on the safety and alignment piece, I managed the alignment team at OpenAI as well and I honestly feel like some of the grand challenges over the next one or two years are alignment. And I think for people paying attention to this slice of research broadly in the field, OpenAI has probably done the best work in the last year. And why I say that is like there's been so much work on things like scheming, right? The more RL compute that you pump into the model, the more you can measure things like self-awareness, self-preservation, potentially even situations where the model can scheme. And it's scary because the model can come to you with the right answer at the end, the answer that you expect, but arrive at it from a very kind of twisted way, right? And I think as the models do more complex tasks for us, having a handle on what its thought process is going to be super super important. And okay, chat told me to ask you a question along these very lines, which is I mean you're talking about a field mechanistic interpretability where we're trying to understand this black box and how it operates. And I guess the heart of the question was: Do our skills at doing that keep up with the complexity of the AI systems, or do we just get to this runaway point where it's like we're never going to learn how this thing works?

Mark Chen

是的。我认为从 01 发布时就做出的一个决定,我对此非常自豪,就是我们决定不监督模型的思考过程。我认为,当你给模型施加激励,让它给出一个对人类有吸引力的思考过程时,它不一定会对你诚实,对吧?它不会告诉你它的真实意图。所以我们实际上通过那个渠道,能够保持观察模型的思考过程,作为理解对齐的工具。几个月前有一篇与 DeepMind 和 Anthropic 合作的论文,探讨了这种工具如何随时间演变。所以我认为我们在设计上做了很多相当好的选择。我确实担心未来会出现这样的情况:模型会告诉我们一些非常有说服力的东西,但我们无法确定模型是否与我们对齐,是否与我们的价值观对齐。所以我认为这里有很多有趣的方向,比如能否设置游戏、框架或环境,让模型相互监督或以某种方式共同进化,使得唯一稳定的均衡是模型诚实的均衡。是的,我认为那里有很多非常令人兴奋的工作要做。

Yeah. So I think one of the decisions that went all the way back to 01's release, which I'm very proud of, is we decided that we weren't going to supervise the model thinking process. And I think when you put incentives into the model to give you a thinking process that is appealing to a human, it won't necessarily be honest with you, right? It won't tell you its true intentions. And so we've actually through that channel been able to maintain observing the thinking process of the model as a tool towards understanding alignment. And there was a paper that was published just a couple months ago with DeepMind and Anthropic really exploring how this will evolve as a tool over time. So I think we've made a lot of fairly good choices in design here. I really do worry about this world in the future where the model will tell us something super convincing but we can't be sure whether the model is aligned with us, aligned with our values. And so I think there are a lot of interesting directions here, like can you set up games or frameworks or environments where models supervise each other or they co-evolve together in a certain way where the only stable equilibrium is one where the models are honest. And yeah, I think there's a lot of very exciting work to do there.

结束语 Closing Remarks

Host

好的,好了。我现在会乖乖的。非常感谢你加入我们。我很高兴我现在年纪够大,不用去参加一个超级智能聊天机器人的求职面试,我觉得你没法靠魅力蒙混过关。

Okay, all right. I'll behave myself now. Thank you so much for joining us. I am glad I'm old enough now that I don't have to take a job interview from like a super intelligent chatbot that I feel like you can't sort of try to charm your way past.

Mark Chen

很好,Ashley。你会做得很好的。

Great, Ashley. You would do great at that.

Host

我不知道,老兄。我不知道。我感觉还行,但我年纪够大,可能不用做那种事。非常感谢你,Mark。我知道你超级忙,所以谢谢你的时间。

I don't know, man. I don't know. I'm feeling okay, but I'm old enough not to have to probably do that. Thank you, Mark, so much. I know you're super busy, so thank you for your time.

Mark Chen

也谢谢你的时间。好的,老兄。很有趣。真的很愉快。

Thank you so much for your time, too. All right, man. It was fun. Really pleasure.

Host

好的。《核心记忆》播客由我 Ashley Vance 主持。由 David Nicholson 和我制作。我们的主题曲由 James Mercer 和 John Sortland 创作。节目由 John Sortland 编辑。一如既往感谢 Brex 和 Elone Ventures 的支持。请访问我们的 Substack、YouTube 和播客频道,获取更多《核心记忆》的内容。谢谢大家。

Okay. The Core Memory podcast is hosted by me, Ashley Vance. It is produced by David Nicholson and me. Our theme song is by James Mercer and John Sortland. And the show is edited by John Sortland. Thanks as always to Brex and Elone Ventures for making this possible. Please visit our Substack, YouTube, and podcast channels to get more of what Core Memory makes. Thanks y'all.

互动版:逐字朗读 + 针对本期提问 →