Gemini 3.0 编写 Gemini 4.0:AI 里程碑与测试时计算

Gemini 3.0 Writes Gemini 4.0: AI Milestones and Test-Time Compute

诺姆·沙泽尔 Noam Shazeer · Unsupervised Learning · 2025-03-17 · 约 70 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Noam Shazeer 和 Jack Rae 讨论测试时计算、模型惊喜以及 AI 中有意义的里程碑。

Noam Shazeer and Jack Rae discuss test-time compute, model surprises, and meaningful milestones in AI.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 38)

全文 · Full transcript(中英对照)

0. 引言 Introduction

Host

有没有一些对你来说有意义的里程碑?我会说当 Gemini 3.0 写出 Gemini 4.0 的时候。几天前我给我妈看了这个。妈妈们是检验某样东西是否从推特圈突破到现实世界的终极测试。妈妈 vibe 检查很重要。你觉得当今 AI 世界里什么被过度炒作,什么又被低估了?我有几个想法,但我觉得都太劲爆了。

Are there a set of milestones that are meaningful to you? I'd say when Gemini 3.0 writes Gemini 4.0. I was showing this to my mum a couple of days ago. Mums are the ultimate test of whether something has broken the barrier from the Twitter sphere to the real world. The mum vibe check is a big deal. What do you feel is overhyped in the AI world today and what's underhyped? I have a few, but I think they're all too spicy.

Host

在我的播客里没有太劲爆这回事。这就像打地鼠一样,总是在区分哪些是真正在插值已知想法,哪些是创造全新想法。关于这一点,我要请出 Noam、Gome Shazir 和 Jack。他们真的不需要介绍。这两位处于 Google Gemini 大语言模型工作的前沿,并且参与了过去十年中 AI 领域一些最重要的发现。Noam 是 Transformer 和混合专家模型的共同发明者之一。Jack 是许多 DeepMind 突破的关键成员。能和他们两位坐在一起,问他们当今 AI 领域所有最热门的问题,真是莫大的荣幸。我们讨论了测试时计算能走多远,以及它在哪些领域有效、哪些无效。我们讨论了测试时计算与大规模预训练范式相比,基础设施需求会有何不同。我们还谈到了开源模型追赶闭源同行的惊人速度,以及他们对 DeepSeek 的反应。我们讨论了他们对 Ilya 说测试时计算范式无法让我们完全达到 AGI 的看法,以及 Yann LeCun 说当前这一代模型实际上无法产生任何新颖想法的观点。我们讨论了当今进行前沿 AI 研究实际是什么样子,他们的日常工作是什么,以及对他们来说真正重要的未来模型里程碑。然后我们还听到了 Noam 对性格的反思,以及他们两人对 AGI 对人类角色意味着什么的回应。我想大家会喜欢这个的。能和 Noam 和 Jack 交谈真的非常愉快。闲话少说,他们来了。

There's no such thing as too spicy on my podcast. It's always going to be this kind of whack-a-mole of what is considered actually interpolating known ideas versus creating a completely novel idea. And for that one, I'm going to go to Noam, Gome Shazir, and Jack. They really need no introduction. The two are at the forefront of Google's Gemini LLM efforts and have been involved in some of the most important discoveries in AI in the last decade. Noam is one of the co-inventors of the Transformer and mixture of experts. Jack was a key part of many DeepMind breakthroughs. It's a real privilege to get to sit with these two and ask them literally every top-of-mind question in AI today. We talked about how far test-time compute will get us and the spaces where it will and won't work. We talked about how the infrastructure needs will be different for test-time compute versus the large pre-training paradigm. We also hit on the impressive pace at which open-source models have caught up with closed-source peers and their reactions to DeepSeek. We talked about their reception and reaction to Ilya saying that this test-time compute paradigm won't get us all the way to AGI, as well as Yann LeCun saying this current generation of models can't actually have any novel thoughts. We talked about what it actually looks like to do cutting-edge AI research today and what their day-to-days look like, as well as the future model milestones that actually matter to them. And then we also got Noam's reflections on character, as well as both of their responses to what AGI means for the role of humanity. I think folks are going to love this. It was really just a pleasure to get to speak with both Noam and Jack. Without further ado, here they are.

Host

那么,Noam,Jack,非常感谢你们来做客播客。

Well, Noam, Jack, thanks so much for coming on the podcast.

Noam

哦,谢谢。

Oh, thank you.

Host

我们现在坐的这间办公室就是 Transformer 被发明的地方吗?

Are we sitting in the very office where the Transformer was invented?

Noam

不,这是新楼。我们以前在查尔斯顿 196 号,大概半英里远。

No, this is a new building. We were in 196 Charleston, I think, about half a mile away.

Host

所以空气中还弥漫着那种气息,很近。

So it's in the air, pretty close.

Noam

很近,是的。

Pretty close, yeah.

1. Gemini 2.0与测试时计算 Gemini 2.0 and Test-Time Compute

Host

今天有很多话题要深入探讨。显然,我想从最新的 Gemini 2.0 模型开始,以及你们围绕测试时计算和 Gemini 2.0 Flash Thinking 所做的所有工作。我想对听众来说,从最高层面来说,你如何描述这些模型今天在哪里有效,在哪里不太有效,当你在实验它们时,哪些结果最让你惊讶?

Well, many things to dive into today. I mean, obviously want to start with some of the latest Gemini 2.0 models and obviously all the work you've been doing around test-time compute and Gemini 2.0 Flash Thinking. I guess just at the highest level for our listeners, how do you characterize where these models work today, where they don't work as well, and as you were experimenting with them, what surprised you most about those results?

Noam

一件令人惊讶的事情是,当我们开始特别努力地将大量关于测试时计算的研究整合到 Gemini 中,然后考虑发布它时,我们最初专注于推理任务,所以数学和代码是重点领域。当时并不清楚,虽然我们在那个领域冲刺,我们显然想随着时间的推移自然扩展,但不太清楚这会如何运作。会有某种泛化吗?如果我们作为研究人员只专注于这些推理任务,思考能力是否会对这些任务之外也有用?我认为看到早期的一个模型很有趣,这个模型我们觉得已经训练到匹配 Gemini Flash 的风格,所以它经过了思考训练,但也经过了一些训练,使其实际上成为一个普遍友好的风格,一个很好对话的模型。看到思考能力与创造性任务互动并改进它们,实际上非常有趣。你可以要求模型就某个主题写一篇文章,思考内容读起来很有趣,它会经历各种不同的想法,然后对想法进行修改或删减,这很有趣。然后输出感觉非常好。所以这是让我惊讶的一件事。

One surprising thing is, when we started the particular concerted effort to build a lot of research into test-time compute into Gemini and then think about shipping it, we were really focused on starting out with reasoning tasks, so math and code were big areas of focus. It wasn't really clear, while we were sprinting in that domain, we obviously want to broaden it naturally over time, but it wasn't really clear how that would work. Would there be any sense of generalization? Would thinking be useful beyond those reasoning tasks if we're just concentrating on those as researchers? And I think it was pretty fun to see one of the early models that we felt like had been trained to try and match the style of Gemini Flash, so it had been trained with thinking but then it had also undergone some kind of training to actually be just like a generally nice style, a nice model to talk with. It was actually very fun seeing thinking interact and improve creative tasks as well. You could ask the model to compose an essay on a particular topic, and actually the thought content was very fun to read, and it would go through various different ideas and then revisions of the idea or things that it should cut, and that was kind of fun. And then the output felt really nice. So that was one thing that surprised me.

Host

Jack,你有什么惊讶的吗?

Any surprises for you, Jack?

Jack

嗯,是的。我的意思是,总的来说,我支持通用性。让我们训练一个在各方面都出色的东西。这很重要。我一开始对如此专注于数学之类的东西持怀疑态度,但拥有好的基准非常重要,这些基准会鼓励你能够推理困难的任务。因为很多事情会降低困惑度,增加模型参数,并记住更多。所以拥有能够更好地区分一些更困难问题的评估是很好的。

Well, yeah. I mean, in general, I'm all for generality. Let's train something that's great at everything. It is important. I was skeptical at first of this intense focus on things like math, but it is very important to have good benchmarks that are going to encourage you to be able to reason about difficult tasks. Because a lot of things will drop perplexity, add more parameters to the model, and memorize more. So it's nice to have the evals that can distinguish better some of the more difficult problems.

2. 有意义的评估与直觉检查 Meaningful Evals and Vibe Checks

Host

我的意思是,在这一点上,什么评估对你来说甚至有意义?我觉得人们都在试图爬同一组评估的山,这些评估感觉越来越与日常工作无关。你们在测试这些模型时做什么,或者你们如何对它们进行 vibe 检查?

I mean, what evals are even meaningful to you at this point? I feel like people are trying to hill-climb the same set of evals that feel increasingly less relevant to day-to-day work. What do you guys do when you're testing these models, or how do you vibe check them?

Jack

感觉我们总是找到一个评估,我们喜欢它,哦实际上我们忽略了另一个评估。即使是在数学领域,比如我们做了一堆数学评估,但也许,我不知道,只给出答案,它们仍然被认为具有挑战性,然后突然就完成了,它们完全饱和了,我们真的不在乎了,而且它们很小,我们几乎在想我们当初为什么要研究它们?很容易忘记,六个月前我们认为很难的东西,几个月前被认为非常难,也许对模型来说太难了,然后突然就变得微不足道了。所以现在,我确实觉得 DeepMind 和整个 Google 内部一直有很多协同努力来开发有用的评估,但看到这也是许多不同 AI 实验室的共同责任,甚至 Scale AI 的想法也在真正加强,开发像“人类最后一次考试”这样的东西。这是一个危险的标题,你知道,如果每六个月我们都觉得它不太难。这是一个危险的标题。真正的最后一次考试。总的来说非常非常具有挑战性,因为评估会被泄露。一旦人们开始谈论这些评估,那么关于这些评估的所有文本就都出来了,它们就不再好了,因为每个人都知道。

It feels like we keep landing on an eval, we like it, oh actually we overlooked this eval. Even if it's in math, it's like okay we've done a bunch of math evals, but maybe like I don't know, put them answers only, they're still considered challenging, and then it's like done, they're completely saturated, and we really don't care, and they're small, and we almost think why did we ever even work on them? And it's easy to forget that was so easy what we thought six months ago, a couple of months ago it was considered really hard, maybe too hard for the model, and then it just snaps to being trivial. So right now I do feel like there's always been a lot of concerted effort within DeepMind and within Google as a whole to develop useful evals, but it's kind of very nice to see also this is like a shared responsibility across many different AI labs and even Scale AI's thoughts is really stepping up and developing like Humanity's Last Exam. It was a dangerous title, you know, if every six months we think it wasn't too hard. It's a dangerous title. The really, really last exam. It's very, very challenging in general because evals get leaked. Once people start talking about the evals, then there's all this text out there about the evals, and they're no good anymore because everyone knows.

3. 评估与里程碑 Evaluation and Milestones

Host

问题在于,除非你非常非常小心,否则所有模型都会知道这些评估问题。所以,在拥有私有的评估方面还有很多工作要做。对你来说,有没有一组有意义的里程碑?比如当 Gemini 3.0 能做某件事时,那会是一个非常令人兴奋的里程碑,无论是评估还是你尝试过但模型还做不到的事情。

Problems and you know all the models will know the problems unless you're very very careful. So there's still a lot of work that goes into having evals that are private. Are there a set of milestones that are meaningful to you? Like when Gemini 3.0 can do X, that's a really exciting milestone, whether it's an eval or just something you've tried with these models that they can't quite do yet.

Noam

我会说,当 Gemini 3.0 写出 Gemini 4.0 时,或者应该说 Gemini X 写出 Gemini X+1 时。我认为这些强化循环可能是最值得关注的事情。目前有几个强化循环在发生。我刚才提到的可能是最重要的一个:我们可以利用我们正在构建的 AI 作为工具,让我们在构建 AI 时更高效。但还有其他强化循环,比如数据飞轮:让人们使用这些模型并提供反馈,使它们在人们关心的事情上变得更好。我认为我们会看到由此带来的巨大加速。还有全球的兴奋情绪和资金飞轮,这在过去几年似乎也在升温。

I'd say when Gemini 3.0 writes Gemini 4.0, or I should say Gemini X writes Gemini X plus one. I think these reinforcement loops are probably the most important thing to pay attention to. There are several reinforcement loops going on. The one I just mentioned is probably the most important one: we can actually use the AI we're building as a tool to make ourselves more productive at building AI. But then there are other reinforcement loops around data flywheels, like you have people use these models and provide feedback and make them better at the things people care about. I think we'll see huge acceleration from that. And then there's the global excitement and funding flywheel, which seems to be kicking up in the last few years.

4. 谷歌的AI辅助开发 AI-Assisted Development at Google

Host

确实如此。说到你提到的身边有一千个 AI 工程师让你更高效,我们现在处于什么阶段?你们现在在研究中使用这些 Gemini 模型,是否相当于有了 0.1 个全职员工?

Certainly. To your point of having a thousand AI engineers by your side making you more productive, where are we? Do you have the equivalent of a 0.1 FTE today using some of these Gemini models alongside your research?

Noam

我们在 Google 的一个优势是,我们在一个结构化的单一代码库中工作,而且我们已经有很多出色的工具来贡献代码。因此,AI 在很多方面被引入作为我们自身开发的工具。Jeff 引用过一个统计数据,虽然我不知道具体数字,是关于有用的拉取请求数量,比如 AI 修复 bug 或代码审查。这些有一天就被引入了;我注意到它,心想:‘哦,太酷了,我现在可以应用它已经发现的修复了。’但这只是我们已经在将 AI 引入自身编码开发的一个方面。我们对智能体式编码感到非常兴奋;这绝对非常重要。让模型能够处理更开放和困难的任务是我们非常兴奋的事情。在某些方面,当我们有非常明确的库定义方式(比如构建规则等)时,协调起来更容易。整个代码库的各个方面都很好地融合在一起。我可以想象,随着进展的继续,会有一个非常明确的时刻,突然之间很多库可以在我们的代码库中快速迭代。

One benefit we have at Google is that we work in a very structured monorepo, and we have a lot of amazing tooling around contributing to the code base already. So there are lots of angles where AI is being pulled in as tooling for our own development. Jeff quoted a statistic, though I don't know the exact number, about the number of pull requests that are useful, like AI bug fixes or code reviews. That already gets pulled in one day; I notice it there and think, 'Oh, that's cool, I can now apply fixes that it already spots.' But that's just one element where we're already pulling in AI towards our own coding development. We're incredibly excited about agentic coding; it's definitely very important. Trying to get the model to tackle more open-ended and difficult tasks is something we're very excited about. In some ways, it's easier to orchestrate when we have a very defined way of defining libraries, like build rules and things. Everything gels together in terms of the whole code base very well. I can imagine that as progress continues, there will be a very discrete moment where suddenly lots of libraries can be very quickly iterated on within our code base.

5. 扩展到低验证领域 Scaling to Less Verifiable Domains

Host

你的 AI 工程师会向你提出实验建议,比如什么值得尝试。这些模型在编码和数学等易于验证的领域表现非常好。你如何看待这些模型在那些不太容易验证的领域最终如何扩展并变得有用?我的意思是,它们在这些方面也在变得更好,但这肯定更难。在这些领域,我们需要更好的验证方法或更多的人类反馈循环。

You get your AI engineers proposing experiments to you, like what's good to try. These models work super well in easily verifiable domains like coding and math. How do you think about how for some of these less easily verifiable domains these models may end up scaling and being useful? I mean, they're getting better at that stuff too, but it's definitely harder. In those domains, we'll need either better ways of verifying or more human feedback loops.

Noam

好的一面是,我们看到即使是 Gemini 模型系列,它们也能遵循更抽象的指令。所以,能够尝试对一项定性工作提供奖励信号——如果人类要给出反馈信号,可能会有一套相当广泛的评分标准或评分准则,甚至可能是一种简单的感觉,比如什么是好的风格,什么是有趣的。我认为问题的一部分实际上是训练模型接受这些非常广泛的标准,然后应用奖励信号。一旦我们有了奖励信号,我们就可以用强化学习来训练。我认为我们已经看到这开始变得合理了;它不再是一个抽象的东西。一两年前它似乎还非常抽象。

What's good is that we're seeing even with the Gemini model series, they're able to follow much more abstract instructions. So being able to try and provide a reward signal over a qualitative piece of work, which maybe if a human were going to try and give a feedback signal, there would be a quite broad set of rubrics or grading criteria, or maybe even a simplistic sense of what is good style, what is interesting. I think part of the problem is really training models to take on these very broad criteria for how to take in a very broad set of criteria and then apply a reward signal. Once we have the reward signal, we can train with reinforcement learning against it. I think we're already seeing that kind of makes sense; it's not an abstract thing anymore. It seemed like a very abstract thing maybe a year or two ago.

6. 研究路径中的意外 Surprises in Research Path

Host

一两年前,你预料到这会奏效吗?还是你认为整个研究路径更倾向于这些更容易验证的领域?

A year or two ago, did you expect that to work? Or did you think this whole path of research was really more toward these more easily verifiable domains?

Noam

我觉得你期待它有一天能奏效,然后感觉有一堆非常复杂的事情需要解决才能达到目标。但通常的情况是,哦,原来有一条更简单的路。我对此就是这种感觉。通常会有好的惊喜,好的惊喜多于坏的惊喜。

I feel like you expect it to work one day, and then it feels like there's a very complicated stack of things we need to solve to get there. And then usually the case is, oh, it turns out there was a much simpler path. That's how I feel about it anyway. Usually there are good surprises, more good surprises than bad surprises.

7. 用户采纳与Gemini应用更新 User Adoption and Gemini App Update

Host

你们已经将这些模型发布到外界。你最喜欢看到人们使用它们的哪些方式?你希望看到更多人尝试和构建哪些东西?

You've released these models out in the wild. What are some of your favorite ways you've seen people using them? And what would you like to see more people trying and building with these things?

Noam

实际上,今天 Gemini 应用正好有一个更新,我们加入了一个更强大的模型,并且它与 Gemini 应用工具的完整套件集成在一起。所以它支持 Gemini 应用的所有功能,从地图到搜索集成,现在还有很长的上下文。所有这些使它成为应用中一个功能齐全的 Gemini 思维版本。我认为这是一种非常愉快的体验。到目前为止,即使是日常使用,人们——我不确定他们是否愿意为了模型在回应前思考而多等一点延迟。但事实似乎是,如果人们要掏出手机输入内容,我们以为很长的几秒钟时间,实际上对于他们认为更高质量的答案来说是非常小的代价。而且他们有时还可以查看模型的思考过程并检查它们。几天前我给我妈妈展示了这个——这是检验某样东西是否从 Twitter 圈子突破到主流的终极测试。

Actually, there's just an update today happening in the Gemini app where we're putting in a much stronger model and it's being integrated with basically the full suite of Gemini app tools. So it's like all of the apps that Gemini app supports, from things like Maps, also search integration, and now has very long context. All these things make it a fully featured Gemini thinking release in the app. It is a very enjoyable experience, I think. So far, even for their day-to-day stuff, people—it wasn't clear to me whether they would like to pay a bit more latency to get the model to think about stuff before responding. It does seem to be the case that if people are going to pull their phone out of their pocket and then type something in, what we thought was a very long time, maybe a couple of seconds, is actually a very small price to pay for something that they feel like is a better quality answer. And also maybe they can sometimes look at the thoughts and inspect them. I was kind of showing this to my mom a couple of days ago—the ultimate test of whether something has broken the barrier from the Twitter sphere to the mainstream.

8. 多模态能力与智能体任务 Multimodal capabilities and agentic tasks

Host

现实世界,是的,氛围检查很重要。对,我猜她问了很多我认为非常通用的问题,比如‘生命的意义是什么?’她问了这个问题,然后她真的坐着读了很久,还读了思考过程,然后她沉思了。她似乎很欣赏有思考过程来指导如何回答这种开放性问题。所以越来越多的人在用这些模型进行哲学对话。

The real world, yeah, like the vibe check is a big deal. Yeah, and I guess she asked a lot of what I would consider very generic questions to ask a model, like 'What is the meaning of life?' She went for that, and she really sat and read it for a long time, and then she read the thoughts as well, and she contemplated on it. She seemed to appreciate the presence of thoughts on how to approach such an open-ended question. So more folks are building philosophical conversations with these models.

Noam

嗯,我觉得它们非常令人印象深刻的一点是多模态能力,但从应用角度来看,这些能力似乎还没有被充分探索。我不太确定原因。但我觉得,目前我们在某种程度上对 Gemini 的多模态能力比较谦虚。我认为这个模型在图像输入方面一直非常强大。图像输入加上思考实际上效果非常好。我看到很多人在 X 上对模型进行红队测试,尝试困难或具有挑战性的视觉推理问题。我认为效果相当不错。也有一些玩具评估,但将多模态与智能体任务结合起来也非常有趣。所以我们去年 12 月推出了 Mariner,这是一个使用浏览器的智能体,内置了很多多模态功能。让模型不仅能够扫描屏幕,还能真正理解屏幕内容并知道如何在各种不同类型的网站上行动,这一点非常重要。所以将这种能力——智能体式、相当开放、可能需要视觉理解混乱场景——与思考结合起来,是我感到非常兴奋的事情。

Well, one thing I think is super impressive about them is the multimodal capabilities, but it still feels like those are very underexplored from an application perspective. Not sure totally why that is. But yeah, I think right now, in some ways, we're quite modest about the multimodal capabilities of Gemini. I feel like the model has always been incredibly strong at image input. Image input plus thinking is actually remarkably good. I see a lot of people red-teaming the model on things like X and trying difficult or challenging visual reasoning problems. I think that is working pretty well. There are also some toy evaluations, but pulling in multimodal with agentic tasks is super interesting as well. So we launched Mariner last December, which is an agent that uses a browser and has a lot of multimodal aspects built in. It was super important to get the model to be incredibly strong at not only scanning a screen but really understanding it and knowing how to act on many different types of websites. So pairing that kind of capability—agentic, quite open-ended, maybe requiring visual understanding of potentially messy scenes—with thinking is something I'm feeling very excited about.

Host

这很有趣。我的意思是,从文本开始是有道理的。一图胜千言,但一张图有百万像素,所以文本的信息密度仍然高出一千倍。而且主要的训练数据中文本也更多。我们有很多文本示例,代表了人类通过语言接收和产生信息的方式。但像图像生成这样的任务,我们拥有的示例较少,因为没有人们生成图像的现成例子。所以事情更具挑战性,但正如 Jack 所说,进展很棒。

That's fun. I mean, it made sense to start with text. A picture is worth a thousand words, but it's a million pixels, so text is still a thousand times more information dense. And there's also a lot more text in the main training data. We have so many examples of text that represent the way humans receive and produce information through language. But we have fewer examples for something like image generation because you don't have examples of people generating images. So things are a little more challenging, but as Jack said, great stuff happening.

15. 推理成本与人力成本对比 Cost of inference vs human labor

Noam

你每美元能获得超过一百万个 token。你可以查一下 Gemini 或其他家的价格。所以每美元能获得数百万个 token,这比你能想到的大多数其他东西便宜好几个数量级。比如,一个非常便宜的消遣——买本平装书来读——大约每美元一万个 token。所以推理比读书便宜两三个数量级,比雇人做任何事大概便宜四个数量级,比雇一个软件工程师便宜六到八个数量级。所以有巨大的空间可以投入更多算力让模型更聪明。如果价值存在——我认为确实存在——你愿意为差劲的工程师付每小时 5 美分,还是为好的工程师付每小时 10 美分?所以有大量未开发的算力等待我们去利用。一个直接的方法就是训练更大更好的模型。我们已经在做了,但模型训练成本随模型规模呈二次方增长。所以如果做得好,推理仍然相对便宜。现在大家都在做的就是在推理时投入更多算力,通过思维链或我们能想到的任何其他优秀算法。我认为我们也会开始看到那里的 Scaling 曲线,就像我们在很多地方看到的那样。

You're getting like over a million tokens per dollar. You know, you can just check the prices on Gemini or anybody else. So you're getting millions of tokens per dollar, which is orders of magnitude below the cost of most other things you can think of. For example, a really cheap pastime like buying a paperback book and reading it costs about 10,000 tokens per dollar. So inference is a couple of orders of magnitude cheaper than reading a book, probably four orders of magnitude cheaper than paying anybody to do anything, and six to eight orders of magnitude cheaper than paying a software engineer. So there's a huge margin to apply more compute and make the thing smarter. If the value is there, which I think it is, would you pay 5 cents an hour for a bad engineer or 10 cents an hour for a good engineer? So there's a huge amount of unexploited flops to use if we can find ways to use them. One straightforward way is to just train a bigger, better model. We're already doing that, but model training costs tend to go up quadratically with the size of the models. So you still end up with relatively cheap inference if you do it right. So what everyone is doing now is applying more compute at inference time through chain of thought thinking or any other brilliant algorithms we can come up with. I think we're just going to start seeing a scaling curve there as well, as we're seeing in a lot of places.

16. 扩展至AGI与自主性需求 Scaling to AGI and need for agency

Host

但这是否能一直 Scaling 到人们设想的 AGI 未来,还是需要某种完全相邻的东西?

But does that scale all the way to the AGI future people envision, or is there some completely adjacent thing required?

Noam

我想我们会看到是人类还是 AI 发明下一个突破。我已经放弃整理车库之类的事了,因为我在等机器人。你觉得它们快来了吗?我猜你们昨天发布了重大的机器人产品。那太棒了。不过它能清理你的车库吗?据我所知还不能。但我认为,测试时计算范式还能贡献多少——它能一路走到 AGI 吗?我不认为它能一路走到 AGI。我觉得我们已经基本确定还有其他组成部分:能够在复杂环境中行动。行动非常重要。对行动智能体的研究是一项明确的投资。还有很多其他方面。

I guess we'll see whether it's the humans that invent the next breakthrough or the AI. I've given up on organizing my garage and stuff like that because I just wait for the robots. You think they're coming? I guess you guys had a big robotics release yesterday. That's awesome. Is it ready to clean your garage though? Not that I know of. But I think the question of how much more there is to give in this test-time compute paradigm—is it all the way to AGI? I don't think it's all the way to AGI. I think we already kind of established that there are other components: being able to act in a complex environment. Acting is very important. Research into acting agents is a definite investment. And there are many other aspects.

17. 深度思考与数据效率 Deep thinking and data efficiency

Host

我们是否看到测试时计算在趋近极限?

Are we seeing test-time compute asymptoting?

Noam

我脑子里一直有一个典型的例子:我不希望测试时计算只是针对任何给定问题,想得更久,最终得出一个解决方案。我们还希望它能够非常深入地思考,并实际创造有用的知识,然后将其融入思考中以解决后续任务,从而大幅提升数据效率。如果你只有一本数学教科书,大部分时间只是真正地思考和摆弄这些想法,然后就能成为世界级的数学家——这就是我们应该努力通过深度思考模型实现的目标。我们做到了吗?没有。我们有路径吗?我认为有很多梯度方向指向这样的模型。我们已经看到,通过用强化学习训练模型在解决问题时深入思考,数据效率有了惊人的提升。所以即使我们有一堆强化学习数据,不再增加任何数据,仅仅采用测试时计算范式就能让我们从这些数据中学到更多。我认为我们可能还能进一步推动。我们感兴趣的研究角度有很多:这个模型如何思考——不是仅仅吐出几千个 token 来解决问题,而是可能更深入地思考,更像一个研究者思考一个困难且开放的问题。

I think the running example is always in my mind here: I don't want test-time compute to just, for any given problem, think longer and eventually arrive at a solution. We also want it to be able to think very deeply and actually create useful knowledge that it will incorporate to solve further tasks in its thoughts, thus dramatically improving data efficiency. If you can just have one math textbook and spend most of your time really just thinking and playing around with the ideas, and then you can become a world-class mathematician—that would be the kind of thing we should strive to achieve with a very deep thinking model. Are we there yet? No. Do we have a path there? I think there are many directions of gradient towards such a model. We've already been seeing amazing improvement in data efficiency by training the model to think deeply with reinforcement learning when solving tasks. So even if we have a bunch of RL data and we're not going to add any more, just going the test-time compute paradigm is allowing us to learn a lot more from that data. I think we could probably push that much more. There are many particular research angles we're interested in: how this model could think not just by spitting out a couple thousand tokens to solve a task, but maybe think much more deeply, much more like a researcher might think about a hard and open-ended problem.

18. 模型作为研究者与数学基准 Models as researchers and math benchmarks

Host

你提到的很多内容都是这些模型越来越像研究者。你会寻找哪些早期迹象?

So much of what you've touched on is these models increasingly acting like researchers. What are the early signs of that you would look for?

Noam

我举一个例子:数学。目前数学被当作基准测试和考试,甚至可能是数学竞赛。将会有一个非常重要的转变:从数学基准测试转向真正开始生成有用的数学,开始解决我们真正关心的重要问题。我认为 FrontierMath 评估的创建非常酷,它试图提供一条通向那个目标的梯度。较难的类别几乎像是未发表的数学发现,较容易的类别只是更难或更刁钻。我不知道那个特定的评估是否是完美的方式,但拥有某种从我们目前所处位置通向实际有用科学贡献的评估斜坡,是我特别兴奋的事情。邀请教授和研究者,说:‘以非讽刺的方式使用这个工具来真正加速这个领域的研究。你会怎么做?缺少什么?我们需要什么来进步?’我认为这种逐步变难的基准测试的想法听起来可能像是 AI 研究者总在说的,但我认为这基本上将成为我的进步指标。数学是一个很好的例子,因为这是一个你实际上不需要更多数据的领域。人们在没有外部输入的情况下发明了所有这些数学,常常是在一个房间里只是思考。艾萨克·牛顿拿了一堆天文观测数据,然后就从那里开始了。

I'll give one example: math. Right now math is being treated as benchmarks and exams, maybe even math competitions. There's going to be a very important pivot from math benchmarks to really starting to become about actually generating useful math, starting to solve important problems we really care about. I think it was very cool that the FrontierMath eval was created, which tries to provide a gradient towards that. The harder category is almost like unpublished math findings, the easier category is just harder or trickier. I don't know whether that particular eval is the perfect way, but having some kind of ramp of evals that bridges from where we are now to actually useful scientific contributions is something I'm especially excited about. Bringing in professors and researchers and saying, 'Use this tool in a non-ironic way to actually accelerate research in this area. What would you do? What's missing? What do we need to advance?' I think this notion of incrementally harder benchmarks might sound like what AI researchers always say, but I think that's going to basically be my metric of progress. Math is a great example because it's a field where you don't actually need more data. People invented all this math without external input, often in a room just thinking. Isaac Newton took a bunch of astronomical observations and went from there.

19. 新发现与AI Novel Discovery and AI

Host

这可以反驳那种说法,即 AI 只是学习模仿人类,我们最多只能让 AI 重新学习已知的东西。这种“学习模仿”的批评已经存在一段时间了。你觉得现在这已经被完全推翻了吗?即认为在这些架构上不可能有新的发现和思考?

That could provide a counterproof to the assertion that this is just learning to mimic people, that the most we can do with AI is to relearn what everybody knows. The learning-to-mimic critique has been around for a while. Do you feel like that's entirely disproven at this point, that novel discovery and thinking is impossible on these architectures?

Noam

有一类科学发现几乎没人能反驳:很多科学就是把不相关的信息联系起来。比如材料科学,更好地关联事物意味着你能更了解什么新材料可能是光伏的。所以光是插值就能极大加速科学。但总会有一个打地鼠式的争论:什么算是插值已知想法,什么算是创造全新想法。这个我抛回给 Yann,让他证明自己生成了一个全新的想法。我其实不在乎争论这些;我们就造 AI,大幅提升技术水平,帮助人类。这已经够好了。

There's definitely one class of scientific discovery that almost no one could argue against: a lot of science is about connecting disjoint pieces of information. For example, in material science, associating things better means you know more about what new material may be photovoltaic. So interpolation alone would completely accelerate science. But there will always be a whack-a-mole of what is considered interpolating known ideas versus creating a completely novel idea. For that, I'll throw it back to Yann to prove he generated a completely novel idea. I don't actually care about arguing this stuff; let's just build AI, greatly increase the level of technology, and help people. That seems good enough.

Host

回到数学的例子,也许数学的现状就像 15 世纪的地理探索。有一些已知的东西,模糊的区域,少数精英数学教授在推动边界。如果我们能训练一个模型,能在所有有用数学的空间中提出正确的问题,并且它非常强大能解决这些问题,也许有一天我们会有上百万个陶哲轩级别的模型。那时我们有望完成地图。那将是科学最伟大的贡献之一。但提出问题似乎是最难的部分;解决问题我相信我们能达到。

Going back to the math example, maybe the state of mathematics is like geographic exploration in the 15th century. There are known things, fuzzy areas, and a small number of elite math professors pushing the boundary. If we can train a model that can pose the right questions in the space of all useful mathematics, and if it's very strong and can solve them, maybe one day we'll have a million Terry Tao-level models. Then we could hope to complete the map. That would be one of the greatest contributions to science. But the question-posing seems the hardest part; solving I'm confident we'll get there.

Noam

数学是无限的,所以我们可以做得更好。

Mathematics is infinite, so we can do much better.

20. AI研究文化 Culture of AI Research

Host

我们来谈谈 AI 研究的文化。你是原始 Transformer 论文的作者之一,你们两位都参与了许多突破。从推动这些创新的文化和团队结构中,你学到了什么?

Let's talk about the culture of AI research. You were part of the original Transformer paper, and both of you have been part of many breakthroughs. What lessons do you take from the culture and team structure that drove these innovations?

Noam

AI 研究就像 15 世纪的炼金术。我们不知道为什么它有效;它高度实验性。你有个想法,但证明在于尝试。然后你有观察和假设。有时你对,有时错。通常需要更多实验。研究人员喜欢分享并对自己做的事感到兴奋。慷慨地给予他人认可很重要,因为往往很难知道哪个想法导致了什么。在某个时候,我们得暂时放弃功劳分配,采取一种基于超级智能的方法:等超级智能来理清一切。它会写出关于推动这些变革的文化的伟大文章。但在谷歌和这群才华横溢的人一起工作非常有趣。

AI research is like alchemy in the 15th century. We don't know why it works; it's highly experimental. You get an idea, but the proof is in trying it out. Then you have observations and hypotheses. Sometimes you're right, sometimes wrong. It usually takes more experimentation. Researchers love to share and get excited about what they're doing. It's important to credit people liberally because it's often complicated to know which idea led to what. At some point, we'll have to give up on credit assignment temporarily and take a superintelligence-based approach: wait for the superintelligence to sort it all out. It will write great thought pieces on the culture that drove these transformations. But it's super fun working at Google with this group of brilliant people.

21. Transformer故事:偶然与必然 The Transformer Story: Serendipity and Inevitability

Host

与大家合作时,Transformer 的故事让我印象深刻。你讲这个故事时,听起来像是偶然发生的——如果你那天没走过那条走廊,可能就不会发生。你觉得这是否不可避免,六个月后也会有人想出来?还是说,这种随机性让我们走上了不同的道路?

Collaborating with everybody and one thing I'm struck by by the Transformer story I think is like you know it involved I think you were like you heard randomly that some people were working on you know when you tell the story it sounds like wow that could have easily not happened like I don't know if you weren't walking down that one day does it feel like inevitably someone 6 months later would have figured that out or like how much just random happenstances there that kind of leads us toward a different path from just you know people in this building colliding in different ways.

Noam

哦,有意思。是啊,我们可能还在用 LSTM 之类的。我觉得这就像问:如果没人发明内燃机,我们现在还在用蒸汽机吗?我想总会有人想出类似 Transformer 的东西。从我在伦敦的角度看,感觉你们一直在围绕这个想法打转。有 Neural GPU,关键是不再用 RNN,而是并行化,但也不只是卷积网络,而是让深度成为序列长度的函数。这个想法在 Neural GPU 里没完全成功,但去掉 RNN 的核心思想已经存在了。当时大家都在搞卷积网络,都想干掉 LSTM。而注意力机制已经在翻译模型里流传了。所以这些想法最终需要融合在一起。

Oh interesting yeah we would have all been using LSTMs or something. I mean I guess maybe it's like okay you asked the same question about like okay if somebody hadn't invented like an internal combustion engine would we still be using steam engines at this point. I mean I think someone would have come up with something like Transformer. From my vantage point like over in London it did feel like you were circling around the idea. There was Neural GPU which was also like okay the key thing is we're not going to have an RNN anymore we're going to parallelize but it's not just going to be a convnet we're going to have like a notion of depth being some kind of function of your sequence length. That was an idea in Neural GPU it didn't quite work out but it was like the key idea of get rid of RNNs felt like that vibe was there. Yeah, because there was all this work on convnets running around, everyone wanted to kill the LSTM. And attention had been floating around from the translation model. So yeah it kind of needed to come together.

22. 自底向上与自顶向下的计算分配 Bottom-Up vs Top-Down Compute Allocation

Host

我喜欢这个。另外在文化方面,我理解 Google 的工作方式之一是自下而上的算力分配——人们做不同项目,说服别人分配算力。显然这是一种模式。其他地方则是一股脑投入一件事,算力分配更自上而下。你怎么看这两种模式的权衡?

I love that. Also kind of on the culture side, I think one of the interesting things of how I understand Google works is basically there's this bottoms-up compute allocation, of folks getting to do different projects and convincing other people to allocate compute for that. Obviously that's one model. There's other places that are like we're going to go all in on one thing and they're much more top-down on the compute side. How do you think about the trade-offs between those two models?

Noam

是的,我们两种都经历过,或者说我都经历过。我在 Google 时,Google Brain 主要是自下而上的,而 DeepMind 则主要是自上而下的。自上而下有助于合作和进行大规模训练。但自下而上也有利于合作,因为如果你带新人加入项目,人均资源不会减少,总资源反而增加。这很好。而且有很多打破抽象的想法,很难分类。如果你说这是预训练的算力,那是后训练的算力,但有些东西完全不属于这些类别,就会掉进缝隙。所以我们正在重新引入相当程度的自下而上,因为我认为这非常重要。平衡。

Yeah, I mean we've been through both, or I've been through both. I guess Google Brain when I was at Google previously was mostly bottom-up as you describe, and then DeepMind has been mostly top-down. It was a bit more top-down. Different philosophies, and there are pluses and minuses to both. I think top-down can be good for getting people to collaborate and for getting larger training runs working. But bottom-up I think is also great for collaboration because then if you bring someone new onto your project, it doesn't mean you have fewer resources per person, you have more total resources. So that's great. And there are so many abstraction-breaking ideas that there's no great way to categorize them. So if you're saying okay this is the compute for pre-training and this is the compute for post-training, well okay you've got something completely different that doesn't fall into those things nicely, and then it falls between the cracks. So we are bringing back a good measure of bottoms-up because I think that's super important. Balance, yeah.

23. DeepMind的集中押注与愿景 Concentrated Bets and Vision at DeepMind

Host

是的,我在 OpenAI 短暂工作过,我喜欢他们集中押注的方式,在某些领域确实回报不错。我也喜欢 DeepMind 集中押注的方式,尤其是……我确实认为有时需要远见。比如 AlphaGo:我一开始就被邀请加入 AlphaGo 团队,因为他们需要工程师,而我是研究工程师,但我完全不明白。我想,谁会关心一个棋盘游戏?我不懂。所以你真的需要有远见的研究带头人。你不能指望每个人都拥有最佳视角,知道影响会出现在哪里。现在,比如思考这个研究领域,我们投入了相当多的自下而上研究,完全不指定方向,这是一个有趣的元研究过程:如何让这些人效率最大化,如何让基线轻量化,如何让信号有效,如何让人尽可能快地行动。这非常重要。但不能全是这样。还需要自上而下的指令性押注。但总是令人谦卑。我认为我们对全局有很好的把握,知道哪些领域会变得重要。但通常自下而上的研究会让你谦卑——你没想到的东西最终比预想的影响更大。所以保持这种机制运行非常重要。

Yeah, like I guess I worked at OpenAI briefly and I liked the way concentrated bets were made and that paid off well for them in certain areas. I also liked the way concentrated bets were made at DeepMind, especially kind of like... I do think there was a bit of a vision sometimes. For example, AlphaGo: I was asked to join the AlphaGo team right at the beginning because they needed an engineer and I was a research engineer, and I just didn't get it. I was like why would anyone care about a board game? I don't get it. So you do really need good research leads that have vision to drive these things. You can't expect everyone from a like that doesn't have always the best vantage point to know where the impact is going to be. Right now within, for example, thinking as an area of research that we kind of work with, it's incredibly important. We have a reasonable non-trivial investment in just bottom-up research where we're really not dictating anything, and it's just a fun process of... it's kind of a fun meta research process of how can we make these people maximally efficient, how do you make your baselines lightweight, how do you make these signal bearing, how can people move as fast as possible. And that's very important. And then it's just very important that it can't all be that. It needs to be and here is a mandate of top-down bets that we have to deliver on. But it's always humbling. I think we have a very good scope of what everything is happening and we get a sense of here are the areas that are going to be really important. And usually you get humbled by the bottom-up research of a thing that you weren't even thinking about ends up being way more impactful than you thought. So just always keeping that running is super impactful.

Noam

是的,回顾过去十年,有没有一些转折点或决策点,艰难的决定最终产生了巨大影响?我认为对我们俩来说,押注大语言模型是正确的。是的,那可能是个好决定。现在看起来很明显,但当时对大多数人来说并不明显,只有少数合作者觉得这肯定会成为大事。而且一度需要逆流而上。

Yeah, I mean reflecting back on the past decade, are there some of these inflection points or decision points where maybe hard decisions ended up being super impactful? I mean I think for both of us, going in on large language models was good. Yeah, I'd say that was probably a good call. And it kind of seems obvious now but it kind of was at least I mean I found it was like in a state of being not obvious to most people but obvious I felt like to a small number of collaborators that this was definitely going to be a big thing. And kind of had to go against the grain for a while.

Host

是的,没错。有一段时间人们对语言模型并不兴奋。现在很难记起来了。对我来说,这始终是地球上最好的问题,但当时的深度学习从业者可能觉得机器翻译更酷,或者计算机视觉。视觉一度很令人兴奋。我不知道为什么大家都搞视觉。大概是因为 ImageNet,有张猫的图片之类的。哦,那是个好例子。是的,那只猫。你做过猫的项目吗?没有,从来没有。我其实没怎么做过视觉。现在你绕了一圈,又回到 Gemini 2.0 模型,我看到很多基于视觉的演示,你有点……

Yeah, that's true. There was definitely a time people were not excited about language models. Hard to remember now. I mean it always seemed like the best problem on Earth to me, but you know, the deep learning people at the time maybe thought machine translation was still a bit cooler or computer vision. Yeah, vision was exciting for a while. I don't know why everyone was in vision. I guess like the ImageNet thing, there's like a picture of a cat or something like that. Oh that was a good one. Yeah, the cat. Did you work on the cat thing? No, no, no, never. Never actually did much in vision. And now you come full circle with these Gemini 2.0 models. I see all these demos that are vision-based and you're kind of...

24. 愿景与世界模型 Vision and World Models

Host

现在视觉领域做的事情比识别猫酷多了。我觉得几乎每个早期的 LLM 研究者都在思考世界模型。实际上,早期的 LLM 研究者甚至不一定是语言背景出身,他们不完全是语言学家。更像是做大规模的无人监督学习,先做语言因为它是知识压缩度最高的,然后把所有东西都吞进一个大的生成模型里,理解一切。看到这一点不断被验证真的很酷。昨天我们推出了原生图像生成,太棒了。我觉得现在很多图像生成只专注于生成极致美观的图像,但原生图像生成能让你做更多事情:理解、编辑、图像和文本序列。再次强调,就是在大量数据上训练一个生成模型来实现。

Doing cooler things in vision now than certainly identifying cats. I mean, it felt like almost every early LLM person had world models on their mind. I actually don't feel like the early LLM researchers were even from a language-oriented background; they weren't really linguists. It was more like train unsupervised learning at scale, do language first because it's the most knowledge-compressed, but then gobble everything up into a big generative model and understand everything. It's very cool to see that just continually proving out. Yesterday we launched native image generation; it's amazing. I think a lot of image generation right now is focused purely on getting maximally aesthetic images, but having native image generation allows you to do a lot more with images: understanding, editing, sequences of images and text. Once again, it's just train a generative model on lots of data to arrive at that.

25. 专用模型与通用模型 Specialized vs. General Models

Host

你们显然都坚信这些模型会变得越来越通用。我想有些人会问:想想医疗这样的领域。最终成为我们 AI 医生的模型——只是我们正在做的某个巨型模型的延续吗?还是会有医疗专用版本,只输入某些数据,或者加一堆护栏?请为我描绘一下,你们最终认为我们的 AI 医生或 AI 生物学研究者会是什么样子。

You guys are obviously both big believers in these models becoming more and more general purpose. I guess a question some folks are asking is: you think about domains like healthcare. What the model that ultimately is going to be our AI doctor—is that just a continuation of what we're doing in some giant model? Is there a healthcare-specific version that is released with only some set of data inputted, or just a bunch of guardrails? Paint that picture for me of what you ultimately think our AI doctor or AI biology researcher looks like.

Noam

我不认为对于这么高价值的东西需要非常特定任务的模型,因为你跟医生聊天可能每 token 要花一美元。所以 LLM 目前要便宜得多。任务特定模型的唯一理由是价格。如果你不愿意为某些东西付每 token 一美元——比如分析大量数据以获得边际价值——那么你可能需要更针对性的模型。一直存在负迁移的概念,所以要把东西分开。但我并不觉得实际情况是这样。如果可以衡量,那才是分开的好理由。如果没有负迁移,只有正迁移,那就用一个大的模型。我个人的理念——尽管不是所有人都同意,这是一个活跃的研究领域——是你想多大程度地专业化并衍生出专家模型。但在我看来很简单:如果有正迁移,就放在同一个模型里,只要它不会变得太贵而无法服务。

I don't really think you would need very specific task-specific models for something that high value, because you probably pay like a dollar a token for talking to your doctor. So the LLM is way cheaper at this point. The only reason for task-specific models is price. If there are things you wouldn't pay a dollar a token for—like analyzing vast quantities of data for marginal value—then maybe you want something more targeted. There's always this notion of negative transfer, so compartmentalize things. I don't really feel like that has ended up being the case. If it can be measured, that's a good reason to compartmentalize. If there's no negative transfer and there's positive transfer, then just have one big model. My personal philosophy—though not everyone agrees, it's an active area of research—is how much you want to specialize and spin off expert models. But from the way I see it, it's simple: if there's positive transfer, put it in the same model as long as it doesn't become too expensive to serve.

26. 对时间线与进展的看法转变 Changed Minds on Timelines and Progress

Host

显然你们已经在最前沿一段时间了。过去一年里,你们有什么看法改变了?

Obviously you guys have been at the cutting edge for a while. What's one thing you've changed your mind on in the past year?

Noam

我觉得时间线提前了。我不是在模糊地说。我认为现在的进步速度要快得多。一年前,这个领域在进步,但每当出现新的范式转变,就会带来突然的加速。我改变看法的一件事是我对信息传播和人们如何采纳科学进步的心智模型。当 Transformer 论文出来时,我在伦敦的 DeepMind。人们觉得这是一篇很酷的论文,但有点怀疑。我最终在假期里把它实现到我们的代码库中,大约在论文发表三个月后,并尝试用于语言建模,但并没有真正被采用。后来我和一个想把它用于强化学习的人合作。我想说,从论文发表到你在 DeepMind 各个领域看到 Transformer,大约花了六到九个月,而且这还是在 Alphabet 内部,信息传播更容易。这个领域接受测试时计算范式的速度让我非常惊讶。许多实验室已经训练并发布了看起来非常好的模型,探索这个空间。如果你发布一个公告说这很重要,哪怕只是一篇博客文章,然后人们就能在几个月内取得突破并发布模型——这对我来说是一个警钟。现在有更多的算力和更多的聪明人在从事 AI 工作。我经常戴着玫瑰色眼镜看事情,回想 2016 年,觉得我们当时很聪明、很有创造力。现在的人非常聪明、非常有创造力,而且他们拥有多得多的算力,人数也多得多。所以如果有什么东西会产生巨大影响,它就能传遍全世界,人们会采取行动,这有点疯狂。现在的算力——一个车库里的小孩拥有的算力比发明 Transformer 所需的还要多。人们总是担心大实验室有多少算力,但用比想象中少得多的算力取得突破是完全可能的。

I feel like timelines have shifted forward. I don't mean that in a vague sense. I think the rate of progress is much faster right now. A year ago, the field was advancing, but whenever you have a new paradigm shift, it creates a sudden acceleration. One thing I've changed my mind on is my mental model of the propagation of information and how people adopt a scientific advance. When the Transformer came out, I was at DeepMind in London. People thought it was a cool paper but were a bit suspicious. I eventually implemented it in our codebase over the holiday break, about three months after the paper came out, and tried it for language modeling, but it wasn't really getting picked up. I eventually collaborated with someone who wanted to use it for reinforcement learning. I'd say it was about six to nine months from the paper coming out before you saw Transformers dotted around all areas of DeepMind, and that's within Alphabet where information propagates more easily. The speed at which the field has picked up this test-time compute paradigm is very surprising to me. Many labs have already trained and released models that look very good, exploring the space. The fact that if you make an announcement and say this is important, and it's just a blog post, and then people can make breakthroughs and release models in a matter of months—that was a wake-up call. There's a lot more compute and a lot more smart people working in AI. I often think of things with rose-tinted glasses, thinking about 2016, thinking we were very smart and creative then. People are very smart and creative now, and they have way more compute, and there are way more of them. So if anything is going to be very impactful, it can spread all across the world and people act on it, which is kind of crazy. The amount of compute out there—a kid in a garage has more compute than was necessary to invent the Transformer. People always worry about how much compute the big labs have, but it is definitely possible to make breakthroughs with way less compute than you would imagine.

Host

是啊,你过去一年有什么看法改变了吗?

Yeah, anything you've changed your mind on in the past year?

Noam

我一直对强化学习的成功印象深刻。我以前没怎么接触过它,但它确实很厉害。

I mean, I've been continuously impressed with the success of RL. I'd never really worked with it much before, but it's actually pretty good.

27. 开源与测试时计算 Open Source and Test-Time Compute

Host

嗯,你刚才提到了这一点,Jack。显然,在 DeepSeek 以及所有在测试时计算领域快速跟进的那些模型的反应中,展望未来,你期望开源模型能够跟上这些模型的每一代吗?这显然比很多人预期的要快。

Well, you kind of alluded to this, Jack. Obviously, in the reaction to DeepSeek and all these models that fast-followed in the test-time compute space, going forward, do you expect the open-source models to be able to keep up with each subsequent generation of these models? It obviously seemed to happen faster than a lot of people expected.

28. 开源模型与前沿模型 Open source vs frontier models

Host

我本来也这么想。这其实是我正在改变看法的一件事。我确实觉得开源模型保持与前沿非常接近和竞争的能力正在持续。我实际上觉得我们可能有一种虚假的安心感,认为这正在发生,因为也许感觉在收敛,但这些东西可能会再次拉开差距。但实际上,这似乎非常令人印象深刻。我对昨天刚发布的 Gemma 3 的表现感到非常惊讶,太棒了,完全不可思议。团队做得非常好。还有其他的开源模型,DeepSeek V3 发布时就是一个非常好的模型。所以看起来人们在开源领域非常有热情,非常有创造力和聪明,而且他们有算力。所以我不太明白为什么他们不能持续创新。你现在怎么看?

Would have expected. Yeah, that's actually something I'm changing my mind on. I do feel like the open source, the ability for open source models to stay very close and competitive with the frontier is persisting. I actually felt we were getting a maybe false sense of assurance that it's happening because maybe it felt like it was converging, but then these things can pull away again. But actually, that seems to have been very impressive. I'm really impressed that the performance of Gemma 3, which just got released yesterday, is amazing. It's completely incredible. The team did a really good job. And other open source models, they've... DeepSeek V3 was a very good model when they released it. So it seems like people are very passionate in the open source space, and they're very creative and smart, and they have compute. So I don't really see why they wouldn't be able to continually innovate. What do you think now?

Noam

是的,我的意思是,闭源和开源之间的时间差距似乎一直在缩小。我认为技术将继续加速。所以可能是质量差距会很大,但时间差距会非常小。但我们必须看看情况如何发展。看到所有这些公司都取得巨大成果,真是令人兴奋。

Yeah, I mean, it seems like the time gap between closed source and open source has been shrinking. I think that the technology will continue to accelerate. So it could be that the quality gap will be large and the time gap will be very small. But we'll have to see how it plays out. It's super exciting to see all of these companies getting great results.

29. 更广泛的影响与个人变化 Broader implications and personal changes

Host

转向我们一直在讨论的 AI 进步对社会的一些更广泛的影响。我很好奇,显然你们俩在过去一年里都对强化学习的力量印象深刻,对我们在测试时算力扩展的速度感到惊讶。基于这种更清晰的认识或信念,即很多 AI 驱动的未来正在到来,你们在自己的生活中有什么改变吗?你们都有孩子。有什么调整吗?

Switching into some of the broader implications of all the AI progress we've been talking about for society. I'm curious, obviously both of you in the last year have been impressed with the power of RL, you've been surprised by the kind of pace at which we've scaled a lot of this test-time compute. Have you changed anything in your own lives based on this probably more clarity or belief that a lot of this AI-driven future is coming? You both have kids. Anything that you've adapted?

Noam

听起来没有。我知道你不清理车库,但我想那是今年之前的事。实际上,那确实是个好播客素材。你考虑过不清理它。是的,我不太担心全球变暖。反正很快 AI 就会处理碳排放问题。但我想,你在自己的生活中有什么不同的想法,或者你如何看待你孩子未来的生活?AI 和教育就像……我认为人们还没有充分讨论这一点。我儿子在监督下,但他喜欢和 Gemini 聊天。它强大得不可思议,尤其是当他去花园,拍植物的照片,拍蜥蜴的照片,他现在有了一个非常准确的个性化百科全书,可以给他信息并适应他。我不确定那会如何运作,但确实如此。我四岁的儿子走来走去,非常详细地谈论植物,他会用拉丁名。他们吸收这么多东西,就像海绵。我觉得我看到了一种我认为人类从未有过的教育正在发生。AI 和教育将会非常不可思议。他去学校,他说,‘哦,是的,我向老师叫了一只蜥蜴。’老师就说,‘哦,那很酷。’然后他说,‘哦,那看起来像一只大蜥蜴。’他说,‘不,那不是大蜥蜴,那是西部篱笆蜥蜴。’他对蜥蜴的种类非常讲究。他还说,‘我还看到了一只蓝尾石龙子。’老师问,‘那是什么?’他说,‘那是一种两栖蜥蜴,……’然后他就开始滔滔不绝。当你看到它时,这很明显。孩子们非常好奇,他们就像海绵一样吸收信息。如果你能将其与 AI 有效地结合起来,我认为那将会非常不可思议。我确实觉得下一代看起来会更聪明。这就是我现在感到充满希望的地方。

It sounds like no. I know you don't clean your garage, but I guess that was prior to this year. Actually, that does make for a good podcast. You thought about not cleaning it. Yeah, I don't worry too much about global warming. You know, have AI to take care of the carbon stuff soon enough anyway. But I guess anything that you've thought about differently in your own lives, or how you think about the life your kids will have. AI and education is like... I don't think people are really talking about it enough yet. My son, under supervision, but he likes to talk to Gemini. It's actually insane how powerful it is, especially if he can go out to the garden, take pictures of plants, take pictures of lizards, and he now has this very accurate personalized encyclopedia which can give him information and adapt to that. I wasn't sure how that would work, but they do. My four-year-old son walks around talking very detailed about the plants, he'll use the Latin name. They absorb so much stuff, they're sponges. I feel like I'm seeing a type of education that I don't really think has ever existed for humanity happening. AI and education is going to be incredible. He went to school and he was like, 'Oh yeah, I called a lizard to his teacher.' And his teacher was like, 'Oh, that's cool.' And he's like, 'Oh, that looks like a big lizard.' He's like, 'No, it's not a big lizard, it's a western fence lizard.' Very particular about his type varieties of lizards. And he's like, 'I also saw a blue-tailed skink.' She's like, 'What's that?' He's like, 'It's an amphibious lizard that...' So he starts reeling these things out. It's just obvious when you see it. Children are very curious, they're like sponges for information. If you can combine that productively with AI, I think that's going to be really incredible. I do feel like the next generation will just seem like smarter people. That's what I'm feeling hopeful about now.

Host

你有什么改变吗?

Anything you've changed?

Noam

是的,预测未来会是什么样子极其困难。我们都会尽最大努力确保 AI 安全且有益。但这确实让你思考,我现在做的事情真的很重要。我们不知道未来人类劳动是否在物质上必不可少,但这只意味着如果你想做一些在物质上重要的事情,现在就去做。除此之外,就努力做个好人。无论你发现什么在精神上有意义,就去做,因为这可能是未来人类的目的。可能不再是满足物质需求。所以我们必须弄清楚未来我们在哪里找到意义,但我们会有很多时间。

Yeah, it is extremely hard to predict what the future will be like. We will all do our best to make sure AI will be safe and beneficial. But it does make you think that what I do now really matters. We don't know if human labor will be materially necessary in the future, but that just means it makes more difference if you want to do something that matters materially, go do it now. And other than that, just try to be a good person. Whatever you find spiritually meaningful, go do it, because that may be the purpose of humanity in the future. It may not be about providing for physical needs. So we've got to figure out where we find meaning in the future, but we'll have plenty of time.

Host

嗯,尤其是你妈妈已经在和模型进行深度哲学对话了。也许我们能够推理出答案。但我很震惊。我觉得其他一些人来过这个播客,首席研究官 Bob McGrew 说,‘听着,人类在提问方面总是有作用的。模型会去执行任务。’但我认为,回到我们之前的对话,大问题是:人类是否永远是提问的最佳人选,还是模型会随着时间的推移提出更好的问题?显然,这有大量的影响。我觉得每一代人都认为他们生活在历史上最重要的时刻,但确实感觉我们正处于其中。

Well, especially your mom's already having the deep philosophical conversations with the models. Maybe we'll be able to reason our way to that. But I am struck. I feel like some other people have come on this podcast, Bob McGrew, chief research officer, was like, 'Look, humanity is always going to have a role in asking the questions. The models will go off and do things.' But I think to our conversation earlier, the big question is: will people always be the best people to ask questions, or will the models actually ask better questions over time? Obviously, that has a ton of implications. I feel like every generation thinks they're living through the most important moment in history, but it does feel like we are certainly in that.

Noam

是的,我认为有时候技术进步会吓到人,人们在这个阶段也有权感到不安。但比如说,即使电视的引入,人们也会说,‘哦,这会不会让我们都失去注意力?我们会不会完全失去出门散步和交朋友的意愿?’这是人们曾经恐慌的事情。显然,现在我们知道了那是没必要的。那只是一项小技术。

Yeah, and I think sometimes a technological advancement scares people, and people have a right to feel trepidation at this stage as well. But there is, I mean, this isn't such a good example, but even the introduction of television, people were like, 'Oh, is this going to make us all lose our attention span? Are we going to completely lose the will to go out and walk outside and have friends?' That was a thing people freaked out about. Obviously, now we know that was unnecessary. It was a small piece of technology.

30. AI风险与社会影响 AI Risks and Societal Impact

Host

你们俩对 AGI 风险有多担心?

How worried are you both about AGI risks?

Noam

我会说中等程度。很难找到这样的例子:创造出比创造者聪明得多的东西,却仍然以可预测且有用的方式为创造者服务。这类论点令人担忧。还有更实际的 AI 与社会影响,比如确保 AI 对经济有建设性,以及就业格局不会发生剧烈变化。这两点我经常思考。还有一些务实的事情:当我们推出能力更强的模型时,作为技术人员,我们很兴奋地开发和发布产品,但我们内部也有另一个团队在更全面地思考这些问题,比如如何安全发布、有哪些意想不到的后果。这非常重要,我很高兴有这样的机制。

I would say moderately. It is hard to find examples of creating something which becomes far more intelligent than its creator but still acts in predictable and useful ways for its creator. That class of argument is concerning. There are also practical AI and society implications, like making sure AI is constructive to the economy and that we don't have sharp changes in the employment landscape. Both are often on my mind. And there are pragmatic things: when we put out more capable models, we are excited as technologists to develop and ship things, but we also have a good balance of having another group internally that thinks about this more holistically, like how to make the launch safe and what unintended consequences might arise. That is super important, and I'm glad it happens.

Host

我同意我们需要在所有的安全方面努力。有些例子表明,我们创造出了更聪明、更强大的东西——我们有比我们更聪明的孩子,然后他们变成了青少年,然后你解决了对齐问题。所以希望如果我们尊重父母、善待他们,AI 会从互联网上那些尊重父母的人的 token 中学到这一点。

I agree we need to work on all the safety aspects. There are examples of creating something smarter and more powerful—we have kids who are smarter than us, and then they become teenagers, and then you solve the alignment problem. So hopefully if we respect our parents and treat them well, the AI will learn from the internet tokens of people being respectful toward their parents.

Noam

是的,我们必须停止推倒机器人。尊重创造者,没错。那感觉挺危险的。

Yeah, we have to stop pushing the robots over. Respect for creators, exactly. That feels dangerous.

31. Character.AI与AI伴侣的未来 Character.AI and the Future of AI Companions

Host

有一件事我想问你,Noam:显然在你之前在 Google 之间的那段时间,你花了很多时间打造 Character.AI,并深入思考了 AI 伴侣的产品空间,以及让人们与各种不同类型的人聊天的能力。你觉得这个领域现在怎么样?我们还需要解决哪些问题?

One thing I wanted to ask you about, Noam: obviously in your previous stint between Google, you spent a lot of time building out Character.AI and thought a lot about the product space of AI companions and the ability for folks to chat with all sorts of different kinds of people. What do you feel about where that space is today? What kind of problems do we still need to solve there?

Noam

我离开 Google 创办 Character 的主要原因是我认为 LLM 行业最需要的是一个应用,让任何人都可以与之互动并发现对自己有用的用例。那是在 ChatGPT 和 Gemini 推出之前。所以任务完成了——现在每个人都在和 LLM 聊天。这和以前的情况不同。另一件事是,在打造 Character 时,我们并没有专注于把它做成一个娱乐产品或其他什么。我们以开放的心态进入,以通用的方式推出它,并帮助人们将其概念化为一种可以扮演不同角色的通用技术。我们确实发现很多人把它用于娱乐,部分原因是那时没人知道如何让这个东西不产生幻觉,所以人们把它用在幻觉实际上是特性的应用上,比如娱乐。这效果很好。我知道很多人喜欢使用 Character。

The main reason I left Google to start Character was I thought the biggest thing the LLM industry could use was an application where anybody can go and interact with LLMs and discover use cases that were good for them. This was before ChatGPT launched, before Gemini launched. So mission accomplished—everyone is out there talking to LLMs now. That is different from how things were. The other thing was, going into Character, we were not focused on it being an entertainment product or something else. We went in with an open mind, put it out there in a general way, and helped people conceptualize it as a very general technology that can take on different personas. We definitely found a lot of people using it for entertainment, partly because at that point, nobody had figured out how to make the thing not hallucinate, so people used it for applications where hallucination is actually a feature, like entertainment. That worked pretty well. I know a lot of people like using Character.

Host

你认为未来 5 到 10 年这个领域会是什么样子?

What do you think the future of that looks like 5 to 10 years from now?

Noam

我不知道。我认为人们永远会想要与人类的关系,因为那在精神上更有意义。但我认为人们会拥有更接近人类形态的 AI,用于他们想要的事情。想象一下你刚当选总统,你有一个 AI 内阁,无论你去哪里都能给你建议。或者也许你是自己 AI 公司的 CEO,所以你的生产力大大提高。在很多这些情况下,与其说是关于个性,不如说是关于生产力。但就人们喜欢与感觉像人类的东西互动而言,我们可能会看到很多在各种方面更人性化的 AI。

I do not know. I think people will always want relationships with humans because it is spiritually more meaningful. But I think people will have AIs that are more in human form for things they want. Imagine you just get elected president and you get your AI cabinet to advise you wherever you go. Or maybe you are CEO of your own AI company, so you get a lot more productivity. In many of those cases, it is less about personality and more about productivity. But to the degree that people like interacting with something that feels human, we will probably see a lot of AI that feels more human in various ways.

Host

这个领域的进步仅仅是模型变得更好,还是有一整套人机交互的挑战?

Is progress in that space just about the models getting better, or is there a whole other set of human-computer interaction challenges?

32. 产品问题与用户界面 Product questions and user interfaces

Host

你知道,需要解决的产品问题。这是个好问题。我的意思是,我认为模型变得更好是件大事,但部分原因在于,对于运行应用程序的人来说,要决定我们允许人们做什么。我认为用户会很擅长指定他们想要什么样的界面,而这主要取决于我们是否想让他们完全指定这一点。

You know, product questions that need to be solved. It's a good question. I mean, I think the models getting better is pretty big, but then part of it is for whoever is running the application to decide what we are going to let people do. I think users will be pretty good at specifying what they want for an interface, and it'll be mostly about whether we want to let them specify that totally.

33. AI中的过度炒作与低估 Overhyped and underhyped in AI

Host

好了,两位,精彩的对话。我们总是喜欢以快速问答结束,听听你们对一些过于宽泛问题的看法。听起来不错。那么,也许先开始,你觉得当今 AI 世界什么被过度炒作,什么被低估了?

Well, look, both of you, fascinating conversation. We always like to end with a quick fire to get your take on some overly broad questions. Sounds good. Okay, so maybe to start, what do you feel is overhyped in the AI world today and what's underhyped?

Noam

我个人觉得 ARC AGI 评估被过度炒作了。嗯,这很劲爆。是的,我认为实际上进展相当缓慢,因为很多研究人员并不特别受启发去做这些非常特定类型的谜题。我们在 2015 年、2016 年做了很多这类工作,然后我们觉得,是的,如果你了解谜题领域,你会花很多时间修复所有实际瓶颈,比如在这些大网格上操作可能有点棘手,然后你取得了很多进展,但你不一定会继续构建真正 AGI 且有用的东西。所以我觉得我个人经历了一个转变,从所有这些合成任务转向只建模自然语言,我觉得从长远来看,朝那个方向推进是更 AGI 的事情。

I personally feel like the ARC AGI eval is overhyped. Well, that's very spicy. Yeah, I think actually the progress has been quite slow because a lot of researchers just don't feel particularly inspired to do these very specific types of puzzles. We did a lot of that in like 2015, 2016, and then we kind of felt like, yeah, you can, if you know the puzzle domain, you spend a lot of time fixing all the actual bottlenecks, like maybe acting on these large grids is a little bit finicky, and then you make a lot of progress, but then you don't necessarily continue on building something that's really AGI and useful. So I feel like I personally had a transition from going to all these synthetic tasks to then going to just model natural language, and I felt that dragging in that direction was a much more AGI thing in the long run.

Host

还有什么想到的,关于过度炒作或低估的?

Anything else comes to mind for overhyped or underhyped?

Noam

我不知道。我认为 AGI 被低估了。AI 大语言模型仍然被严重低估。我认为人们仍然认为它只会是一些愚蠢的万亿美元产品。是的,我听说你在另一个播客上说万亿已经不酷了,千万亿才酷。是的,没错。你得把邪恶博士搬出来。

I don't know. I think AGI is underhyped. AI LLMs are still massively underhyped. I think people are still thinking about it like it's only going to be about some silly trillion-dollar products. Yeah, I heard you said on another pod that a trillion wasn't cool anymore, quadrillion was. Yeah, exactly. You got to put the Dr. Evil up.

34. 在模型之上构建应用 Building applications on top of models

Host

显然,把你从构建模型的工作中抽出来去构建应用程序,将是对社会资源的严重错配。但我很好奇,如果我们今天说去构建一个应用程序,你认为最有趣的是什么?你之前谈过教育,但还有其他你觉得在模型之上构建应用会很有趣的吗?

Obviously, it would be a gross misallocation of societal resources to take you away from building models to build applications. But I am curious, if we were to say today go build an application, what do you think are the most interesting? You've talked about education before, but any others that come to mind that you think would be fun to build apps on top of these models?

Noam

嗯,我确实认为有很多应用试图进入这个智能体领域,然后人们揭露说‘哦,这只是一个已知模型的包装器’,这其实很酷。但如果你想让模型为你做一些有用的事情,比如行动并做有用的事情,那么拥有正确的应用体验似乎很有价值。所以我觉得,在智能体领域,我认为代码现在非常拥挤,但我确实认为还有很多其他事情,我觉得让模型为我自动化会很有用,这些可能超越聊天体验,实际上是走出去做有用的事情。

Well, I do think it's actually very cool how many apps have been trying to break into this agentic space, and people then expose it and say, 'Oh, this is just a wrapper around a known model.' But there seems to be a lot of value in actually having the right app experience if you want to have a model do something useful for you, like act and do something useful. So I feel like, okay, in the agentic space, I think code is very crowded now, but I do think there's a lot of other things that I would find it useful for a model to automate for me that goes beyond maybe a chat experience, but it's actually going out and doing useful things.

Host

是的,我会说代码被低估了。我认为它很巨大,因为首先,人类甚至不擅长它。像代码和数学这样的东西并不是为此而设计的。然后它是会自我加速的事情之一。如果你构建一个自动化的软件工程师研究员,那么它会构建下一个更好的 AI。所以工程和智能体的结合,能够控制足够广泛的界面来完成工程师的工作。如果我要专注于应用程序,那就是我会专注的。

Yeah, I'd say code is underhyped. I think it's huge because, for one, humans aren't even that good at it. Things like code and math were not super designed for that. And then it's one of these things that will help self-accelerate. If you build an automated software engineer researcher, then it'll build the next better AI. So the combination of engineering and agentic, something that can control surfaces broad enough to do the job of an engineer. If I were to focus on applications, that is what I would be focused on.

35. 测试时计算模型的基础设施需求 Infrastructure needs for test-time compute models

Host

测试时计算模型的基建需求与大规模预训练相比会有多大不同?在硬件要求、分布式数据中心等方面。

How different will the infrastructure needs look for test-time compute models versus these massive pre-training? In terms of hardware requirements, distributed data centers, all that.

Noam

是的,我的意思是,这是一个相当乐观的故事,我会说。对吧?所以如果构建 AI 变成主要是推理问题,一个可以比预训练中的大批量训练更分布式的推理问题,我认为这意味着我们可以更灵活地使用算力。但是的,这意味着我们可能不那么在意模型跨数据中心训练。也许可以分散智能体,它们出去获取经验,并从许多不同的数据中心发回经验,因为它们不需要都有非常强的快速互连。所以这也会推动价格下降,因为那时我们可能会开始真正优化这样的设置,这本质上是更便宜的。

Yeah, I mean, it's a pretty rosy story, I would say. Right? So if it turns out that building AI becomes mostly an inference problem, an inference problem that can be much more distributed than maybe large batch training that happens in pre-training, I think that means we can be much more flexible with our compute. But yeah, it's going to mean we maybe don't mind the model training across data centers as much. Maybe it can be spreading out actors that are going to go off and get experience and send that experience back from many different data centers, because they don't all need to have very strong fast interconnects. So that is also going to drive price down as well, because then we might start to really optimize towards such a setup, which is intrinsically cheaper.

Host

酷的是,谷歌,我们与 TPU 团队有这种协同设计联系,所以我们总是向他们反馈我们如何花费算力的概况,这使他们能够在几年内调整芯片设计和数据中心设计,我认为这真的很激励人。正如 Jack 所说,你可以分布式,这变得更好。推理比训练更糟糕的一点是,你失去了 Transformer 中的很多并行性。天真地使用 Transformer,你会受限于内存,为你生成的每个 token 查看注意力键和值。所以有很多伟大的工作要做,从模型架构角度和硬件角度来攻击这个问题,坦率地说,让我们更接近那个点,即利用我们拥有的巨大计算能力,并使我们能够完全将其应用于推理。

The cool thing is that Google, we kind of have this codesign link with the TPU team, so we're always feeding them our profile of how we're spending our compute, which allows them to tweak the chip design and the data center design within a couple of years timeframe, which I think is really motivating. The fact that you can be distributed, as Jack said, gets better. The thing that gets worse about inference than training is that you lose a lot of the parallelism in the Transformer. Naively using Transformer, you end up memory bound, looking at your attention keys and values for every token that you're generating. So there's a lot of great work to do in both attacking this from a model architecture perspective and from a hardware perspective, frankly, to get ourselves closer to the point where we're taking the massive computational power of the sh we have and making ourselves able to fully apply that to inference.

36. 结语与学习资源 Final words and where to learn more

Host

我想把最后一句话留给你。他们可以去哪里了解更多关于你和你们正在做的事情?

I want to leave the last word to you. Where can they go learn more about you and what you guys are doing?

Noam

是的,我们有一个新的更新版 flash 模型,应用了思考能力,比我们 1 月份发布的上一款模型强得多。它已经在 Gemini 应用上发布了。我绝对鼓励人们去试试,并给我们反馈。我们一直在将开发者和用户的反馈纳入每个模型系列,所以我认为这是一件好事。

Yeah, well, we have a new and updated flash model that applies thinking, which is considerably stronger than the last model we released in January. It's out on the Gemini app. I would definitely encourage people to try that out and give us feedback. We have been incorporating developer feedback and user feedback into each model series, so I think that's a good thing.

37. 闭幕词 Closing remarks

Host

我会鼓励大家去做杰克所说的事。好了,非常感谢两位。说真的,能和你们一起做这个节目真是太愉快了。真的很愉快。谢谢。

I would encourage people to do what Jack said. Yeah, well, thank you both so much. Seriously, it's such a pleasure to be able to do this with you. Real pleasure. Yeah, thanks.

Host

嘿,大家好,我是雅各布。在你们离开之前还有一件事。如果你喜欢这次对话,请考虑给节目打个五星评分。这样做有助于播客触达更多听众,并帮助我们邀请到最优秀的嘉宾。这里是《无监督学习》的一期节目,一个由 Redo Ventures 制作的 AI 播客,我们在这里探讨 AI 领域最敏锐的头脑,了解什么是今天真实的,什么是未来会真实的,以及这对商业和世界意味着什么。随着 AI 的快速发展,我们旨在帮助你解构和理解最重要的突破,看清更清晰的现实图景。感谢收听,下期再见。

Hey guys, this is Jacob. Just one more thing before you take off. If you enjoyed that conversation, please consider leaving a five-star rating on the show. Doing so helps the podcast reach more listeners and helps us bring on the best guests. This has been an episode of Unsupervised Learning, an AI podcast by Redo Ventures, where we probe the sharpest minds in AI about what's real today, what's going to be real in the future, and what it means for businesses in the world. With the fast-moving pace of AI, we aim to help you deconstruct and understand the most important breakthroughs and see a clearer picture of reality. Thank you for listening, and see you next episode.

互动版:逐字朗读 + 针对本期提问 →