Organizing Information: From Trillion to Quadrillion Dollar Opportunity
打开互动全文版(中英对照 + 朗读 + 问答)→Jeff Dean 和 Noam Shazeer 讨论谷歌的成长、AI 使世界 GDP 倍增的潜力,以及他们从早期到领导 Gemini 的历程。
Jeff Dean and Noam Shazeer discuss Google's growth, AI's potential to multiply world GDP, and their journey from early days to leading Gemini.
组织信息显然是一个万亿美元的机会,但万亿美元已经不酷了。酷的是千万亿美元。世界 GDP 几乎肯定会大幅增长,比今天高出几个数量级,因为我们有所有这些人工智能工程师。如今我们代码库中检查的字符有 25%是由基于 AI 的编码模型生成的。我们将需要一百万个自动化研究人员来发明所有这些。是的,如果事情朝这个方向发展,这就像在 2018 年邀请 Noam 上播客,然后说‘是的,所以我认为语言模型会成为一个东西。我猜用于帮助每个人的 AI 算力将是天文数字。’今天我有幸与 Jeff Dean 和 Noam Shazeer 聊天。Jeff 是谷歌的首席科学家,在公司的 25 年里,他参与了现代计算中几乎所有最具变革性的系统,从 MapReduce、Bigtable、TensorFlow、AlphaFold,名单真的没完,现在还有 Gemini。而 Noam 是当前 AI 革命中最重要的个人。他是现代 LLM 使用的所有主要架构和技术的发明者或共同发明者,从 Transformer 本身到混合专家模型,再到 mesh tensor flow 等等。他们是谷歌 DeepMind Gemini 的三位联合负责人中的两位。太棒了,非常感谢你们的到来。
Organizing information is clearly like a trillion dollar opportunity, but a trillion dollars is not cool anymore. What's cool is a quadrillion dollars. The world GDP is almost certainly going to go way way up to like orders of magnitude higher than it is today due to the fact that we have all of these artificial engineers. 25% of the characters that we're checking into our codebase these days are generated by our AI based coding models. We're going to need like a million automated researchers to invent all of this stuff. Yeah, if this is where things go, this is actually like getting like Noam on a podcast in 2018 and being like 'Yeah, so I think like you know language models will be a thing. I'm guessing that the amount of compute being used for AI to help each person will be astronomical.' Today I have the honor of chatting with Jeff Dean and Noam Shazeer. Jeff is Google's Chief Scientist and through his 25 years at the company he has worked on basically the most transformative systems in modern computing, from MapReduce, Bigtable, TensorFlow, AlphaFold, genuinely the list doesn't end, Gemini now. And Noam is the single person most responsible for the current AI revolution. He has been the inventor or the co-inventor of all the main architectures and techniques that are used for modern LLMs, from the Transformer itself to mixture of experts and to mesh tensor flow to many other things. And these are two of the three co-leads of Gemini at Google DeepMind. Awesome, thanks so much for coming on.
谢谢邀请。非常兴奋能来这里。
Thanks for having us. Super excited to be here.
好的,第一个问题。你们俩都在谷歌工作了 25 年或接近 25 年。在公司早期,你可能了解一切是如何运作的。什么时候不再是这种情况了?你觉得有一个明确的时刻吗?
Okay, first question. Both of you have been at Google for 25 or close to 25 years. At some point early on in the company, you probably understood how everything worked. When did that stop being the case? Do you feel like there was a clear moment that happened?
我的意思是,我知道我是在 2000 年底加入的,他们有一个每个人都有一个导师的制度。所以我什么都不知道,我就问我的导师所有问题,我的导师什么都知道。结果我的导师是 Jeff,并不是谷歌的每个人都什么都知道。只是 Jeff 什么都知道,因为他基本上写了所有东西。
I mean, I know I joined at the end of 2000 and they had this thing where everybody gets a mentor. So I knew nothing, I would just ask my mentor everything, and my mentor knew everything. It turned out my mentor was Jeff, and it was not the case that everyone at Google knew everything. It was just the case that Jeff knew everything because he had basically written everything.
你太客气了。
You're very kind.
我的意思是,我认为随着公司的发展,你会经历这些阶段。当我加入时,我们有 25 人,26 人左右,所以你认识每个人的名字。即使我们在成长,你也跟踪所有加入的人。在某个时候,你开始记不清公司里每个人的名字,但你仍然认识所有从事软件工程工作的人。然后你开始记不清软件工程组里所有人的名字,但你至少知道每个人在做的所有不同项目。然后到了某个时候,公司变得足够大,你收到一封邮件说‘Project Platypus 周五上线’,你会想‘Project Platypus 是什么鬼?’所以我认为这通常是一个很好的惊喜,就像‘哇,我不知道我们在做那个。’而且我认为即使只是非常高层地了解公司里发生的事情也是好的,即使你不知道每一个细节。并且认识公司里的很多人也很好,这样你可以去问某人更多细节或找出该和谁谈。我认为通过一层间接关系,你通常能在公司里找到合适的人,如果你有一个随着时间的推移建立起来的好人脉网络。
I mean, I think as companies grow, you kind of go through these phases. When I joined, we were 25 people, 26 people, something like that, and so you learned everyone's name. Even though we were growing, you kept track of all the people who were joining. At some point, you kind of lose track of everyone's name in the company, but you still know everyone working on software engineering things. Then you sort of lose track of all the names of people in the software engineering group, but you know at least all the different projects everyone's working on. And then at some point the company gets big enough that you get an email that 'Project Platypus is launching on Friday' and you're like 'What the heck is Project Platypus?' So I think usually it's a very good surprise, like you're like 'Wow, I had no idea we were doing that.' And it turns out I think it is good to keep track of what's going on in the company even at a very high level, even if you don't know every last detail. And it's good to know lots of people throughout the company so that you can go ask someone for more details or figure out who to talk to. I think with one level of indirection, you can usually find the right person in the company if you have a good network of people that you built up over time.
顺便问一下,谷歌是怎么招募你的?
How did Google recruit you, by the way?
实际上是我联系他们的。
I kind of reached out to them, actually.
Noam,你是怎么被招募的?也是一样吗?
And Noam, how did you get recruited? Was it the same?
我实际上是在 1999 年的一个招聘会上看到谷歌的,我以为它已经是一家大公司了,加入没有意义,因为我认识的每个人都用谷歌。我想那是因为我当时是伯克利的研究生。我退学了几次。但结果发现它实际上并没有那么大。所以我没有在 1999 年申请,而是在 2000 年一时兴起给他们发了简历,因为我觉得我应该申请多个工作。它是我最喜欢的搜索引擎。但结果发现真的很有趣。看起来是一群聪明人在做很棒的事情,他们墙上有一张很好的蜡笔图表,显示每天的搜索查询数量,是有人一直在维护的。它看起来非常指数增长。这些人会非常成功,而且看起来他们有很多好问题要解决。所以我想‘好吧,也许我去那里工作一段时间,然后有足够的钱之后就去搞 AI,想搞多久就搞多久。’
I actually saw Google at a job fair in like 1999 and I assumed that it was already this huge company, no point in joining, because everyone I knew used Google. I guess that was because I was a grad student at Berkeley at the time. I've dropped out of grad programs a few times. But it turned out that actually it wasn't really that large. So I did not apply in 1999, but just kind of sent them a resume on a whim in 2000 because I figured I should apply to multiple places for a job. It was my favorite search engine. But then it turned out to be really fun. Looked like a bunch of smart people doing good stuff, and they had this really nice crayon chart on the wall of the daily number of search queries that somebody had just been maintaining. And it looked very exponential. These guys are going to be very successful, and it looks like they have a lot of good problems to work on. So I was like 'Okay, maybe I'll go work there for a little while and then have enough money to just go work on AI for as long as I want after that.'
是啊,某种程度上你做到了,对吧?
Yeah, in a way you did that, right?
是的,是的,完全按计划实现了。
Yeah, yeah, it totally worked out exactly according to plan.
你在 1999 年就在考虑 AI 了?
You were thinking about AI in 1999?
是的,那是 2000 年左右。我记得在研究生院,当时的一个朋友告诉我,他 2000 年的新年决心是活着看到公元 3000 年,他打算通过发明 AI 来实现这个目标。所以我想‘哦,这听起来是个好主意。’但那时我没想到可以在大公司做这件事。但我想,嘿,很多人似乎在初创公司赚大钱,也许我赚点钱,然后就有足够的钱生活,长期搞 AI 研究。但结果发现谷歌是一个搞 AI 的绝佳地方。我的意思是,我喜欢谷歌的一点是,我们的使命一直需要相当先进的 AI:组织世界信息,使其普遍可访问和有用。这里面有一个非常广泛的授权。所以公司不会只做这一件小事然后停滞不前。而且你可以看到我们最初做的事情就是朝着那个方向,但你可以在这个方向上做得更多。
Yeah, this was like 2000. I remember in grad school, a friend of mine at the time had told me that his New Year's resolution for 2000 was to live to see the year 3000, and that he was going to achieve this by inventing AI. So I was like 'Oh, that sounds like a good idea.' But then I didn't get the idea at the time that you could go do it at a big company. But I figured, hey, a bunch of people seem to be making a ton of money at startups, maybe I'll just make some money and then I'll have enough to live on and just work on AI research for a long time. But it actually turned out that Google was a terrific place to work on AI. I mean, one of the things I like about Google is our mission has always been something that would kind of require pretty advanced AI: organizing the world's information and making it universally accessible and useful. There's a really broad mandate in there. So it's not like the company was going to do this one little thing and stay doing that. And also you could see that what we were doing initially was in that direction, but you could do so much more in that direction.
在过去的二三十年里,摩尔定律如何改变了你在设计新系统、确定哪些项目可行时必须考虑的方面?哪些保持不变?仍然有哪些限制?你现在能做哪些以前显然不能做的事情?
How has Moore's law over the last two or three decades changed the kinds of considerations you have to take on board when you design new systems, when you figure out what projects are feasible? What has stayed? What are still the limitations? What are things you can now do that you obviously couldn't do before?
我的意思是,我认为在过去几十年里它实际上变化很大。比如从二十年前到十年前,那很棒,因为你只要等待,18 个月后你就能得到更快的硬件,你什么都不用做。而最近,我觉得基于通用 CPU 的机器在性能提升方面确实慢了下来。所以我们不得不更多地考虑专用硬件,比如 TPU,以及如何设计能够利用并行和分布式计算的算法和系统。所以约束的性质已经从仅仅等待更快的 CPU 转变为需要构建能高效扩展到许多处理器和加速器的系统。
I mean, I think of it as actually changing quite a bit in the last couple decades. So like two decades ago to one decade ago, it was awesome because you just wait and 18 months later you get much faster hardware and you don't have to do anything. And then more recently, I feel like the general purpose CPU based machines have really slowed down in terms of performance improvements. So we've had to think more about specialized hardware, like TPUs, and also about how to design algorithms and systems that can take advantage of parallelism and distributed computing. So the nature of the constraints has shifted from just waiting for faster CPUs to needing to architect systems that scale efficiently across many processors and accelerators.
Scaling 的效果不如从前。制程工艺的改进现在需要三年而不是每两年一次。多核处理器等架构改进带来的性能提升也不如 10 到 20 年前。但与此同时,我们看到了更多专用计算设备,比如机器学习加速器、TPU 和专注于 ML 的 GPU,这些设备让我们能够为现代计算任务获得高性能和高效率,这些任务与用 C++ 代码运行 Microsoft Office 不同。
Scaling has not been as good. Fabrication process improvements now take three years instead of every two years. Architectural improvements in multi-core processors are not giving the same boost as 10 to 20 years ago. But at the same time, we're seeing much more specialized computational devices like machine learning accelerators, TPUs, and ML-focused GPUs, which allow us to get high performance and good efficiency for modern computations that are different from running Microsoft Office with C++ code.
感觉算法在跟随硬件。算术运算非常便宜,而数据搬运相对昂贵得多。几乎整个深度学习都因此起飞——你可以用矩阵乘法来构建,其计算复杂度为 O(N^3),数据通信量为 O(N^2) 字节。
It feels like the algorithms are following the hardware. Arithmetic is very cheap, and moving data around is comparatively much more expensive. Pretty much all of deep learning has taken off because of that—you can build it out of matrix multiplications that are O(N^3) operations and O(N^2) bytes of data communication.
转向围绕这一点的硬件是一个重要的转变。在此之前,CPU 并不特别适合深度学习。然后我们开始在 Google 构建 TPU,它们实际上只是降低精度的线性代数机器。
The pivot to hardware oriented around that was an important transition. Before that, CPUs were not especially well suited for deep learning. Then we started building TPUs at Google, which were really just reduced precision linear algebra machines.
一旦你有了这个,你就必须看到洞察:这一切都关乎机会成本。Larry Page 曾经说过,我们的第二大成本是税收,最大的成本是机会成本。在这种情况下,你拥有这么多芯片面积,却只放了很少的算术单元。把芯片填满算术单元——你可以获得数量级更多的算术运算。那么还有什么需要改变?算法、数据流等等。而且,算术运算可以非常低精度,这样你就能塞进更多的乘法器单元。
Once you have that, you have to see the insight that it's all about opportunity cost. Larry Page used to say our second biggest cost is taxes and our biggest cost is opportunity cost. In this case, you have all this chip area and you're putting a very small number of arithmetic units on it. Fill the thing with arithmetic units—you could have orders of magnitude more arithmetic getting done. Now what else has to change? The algorithms, the data flow, and everything else. And by the way, the arithmetic can be really low precision, so you can squeeze even more multiplier units in.
如果你想象一个反事实的世界,其中内存成本下降得比算术运算更多,因此数据流极其便宜而算术运算不便宜,那么今天的 AI 会是什么样子?
If you imagine a counterfactual world where the cost of memory had declined more than arithmetic, so data flow is extremely cheap and arithmetic is not cheap, what would AI look like today?
你会更多地查找非常大的内存。我认为它可能看起来更像 20 年前的 AI,但方向相反。
You'd have a lot more lookups into very large memories. I think it might look more like AI looked 20 years ago, but in the opposite direction.
我于 2012 年加入 Google Brain。我离开 Google 几年,碰巧回去吃午饭看望妻子,然后坐在 Jeff 和早期 Google Brain 团队旁边。我想,‘那是一群聪明的人。’他们进展不错,所以我又跳了回来。加入 Jeff 很棒。那是 2012 年。我似乎每 12 年加入一次 Google——2012 年和 2024 年。2036 年会发生什么?我不知道,我们拭目以待。
I joined Google Brain in 2012. I left Google for a few years, happened to go back for lunch to visit my wife, and we sat down next to Jeff and the early Google Brain team. I thought, 'That's a smart group of people.' They were making pretty good progress, so I jumped back in. It was great to join Jeff. That was 2012. I seem to join Google every 12 years—2012 and 2024. What's going to happen in 2036? I don't know, we shall see.
你在考虑为未来版本的 TPU 改变哪些权衡?你对算法的思考有何不同?
What are the trade-offs you're considering changing for future versions of TPU? How are you thinking about algorithms differently?
一个总体趋势是我们在量化或使用更低精度模型方面做得越来越好。我们从 TPU v1 开始;我们甚至不太确定能否用 8 位整数量化模型用于服务,但我们有一些早期证据表明可能可行,所以我们围绕这一点构建了整个芯片。随着时间的推移,人们也能够在训练中使用更低的精度。推理精度已经降到 INT4 或 FP4,这在 20 年前对超级计算浮点专家来说听起来很疯狂。有些人甚至将模型量化到 2 位或 1 位。这是一个值得关注的趋势。
One general trend is we're getting better at quantizing or having much more reduced precision models. We started with TPU v1; we weren't even quite sure we could quantize a model for serving with 8-bit integers, but we had some early evidence that it might be possible, so we built the whole chip around that. Over time, people have been able to use much lower precision for training as well. Inference precision has gone to INT4 or FP4, which sounded crazy to a supercomputing floating-point person 20 years ago. Some people are quantizing models to 2 bits or 1 bit. That's a trend to pay attention to.
这真的需要协同设计。如果算法设计师没有意识到通过低精度可以大幅提高吞吐量,他们当然会说不需要低精度,因为这会带来风险和麻烦。然后如果你问芯片设计师要构建什么,他们会问算法设计师,而算法设计师会拒绝量化。所以你需要看到全局:量化很烦人,但你的模型会快三倍,所以你必须接受它。
It really has to be a co-design thing. If the algorithm designer doesn't realize that he can get greatly improved throughput with lower precision, of course they'll say they don't want low precision because it introduces risk and irritation. Then if you ask the chip designer what to build, they'll ask the algorithm designer, who will say no to quantization. So you need to see the whole picture: quantization is irritating, but your model is going to be three times faster, so you'll have to deal with it.
在你的职业生涯中,你多次从事与我们现在用于生成式 AI 的东西惊人相似的工作。1990 年,你的毕业论文是关于反向传播。2007 年,你训练了一个两万亿 token 的 n-gram 模型用于语言建模。请带我回顾一下:当你开发那个模型时,你认为自己在做什么?
At various times in your career, you've worked on things that have an uncanny resemblance to what we're using now for generative AI. In 1990, your senior thesis was about backpropagation. In 2007, you trained a two trillion token n-gram model for language modeling. Walk me through that: when you were developing that model, what did you think you were doing?
让我从论文开始。我在大四的一节并行计算课上接触到了神经网络。我需要完成一篇荣誉论文才能毕业,所以我找到教授说,围绕神经网络做些事情会很有趣。我们决定让我在 1990 年实现几种不同的并行化反向传播训练神经网络的方法。我在论文里给它起了个有趣的名字,比如‘Pattern...’
Let me start with the thesis. I got introduced to neural nets in one section of one class on parallel computing in my senior year. I needed to do an honors thesis to graduate, so I approached the professor and said it would be fun to do something around neural nets. We decided I would implement a couple of different ways of parallelizing backpropagation training for neural nets in 1990. I called it something funny in my thesis, like 'Pattern...'
我在一台 32 处理器的超立方体机器上实现了模型并行和数据并行。数据并行是把所有样本分成不同的批次,每个 CPU 都有一份模型副本;模型并行则是把样本流水线式地传给持有模型不同部分的处理器。我比较了这两种方法。我对这种抽象感到非常兴奋,因为神经网络能解决当时其他方法无法解决的小型玩具问题。天真地,我以为 32 个处理器就能训练出非常棒的神经网络,但结果发现,在它们真正解决实际问题之前,我们需要大约一百万倍的算力。从 2008 年到 2010 年左右,得益于摩尔定律,我们终于有了足够的算力让神经网络真正发挥作用,那是我重新关注神经网络的时期。
I implemented model parallelism and data parallelism on a 32-processor hypercube machine. In data parallelism, you split all the examples into different batches and every CPU has a copy of the model. In model parallelism, you pipeline examples along processors that have different parts of the model. I compared and contrasted them. I was really excited about the abstraction because neural nets could solve tiny toy problems that no other approach could at the time. Naively, I thought 32 processors would let us train really awesome neural nets, but it turned out we needed about a million times more compute before they really worked for real problems. Starting in the late 2008-2010 timeframe, thanks to Moore's law, we had enough compute to make neural nets work for real things, and that was when I re-entered looking at neural nets.
首先,与其他学术成果不同,它真的只有四页纸,你可以直接读完,然后还有 30 页的 C 代码,但这是一个制作精良的成果。跟我说说 2007 年的那篇论文是怎么诞生的吧。
First of all, unlike other academic artifacts, it's really like four pages and you can just read it, and then 30 pages of C code, but it's a well-produced artifact. Tell me about how the 2007 paper came together.
我们在谷歌有一个机器翻译研究团队,由 Franz Och 领导,他大约一年前加入谷歌。每年他们都会参加一个 DARPA 竞赛,将几种语言翻译成英语,我记得是中文到英文和阿拉伯语到英文。谷歌团队提交了参赛作品。比赛方式是周一拿到 500 个句子,周五必须提交答案。我看到了结果:我们以相当大的优势赢得了比赛,以 BLEU 分数衡量,这是翻译质量的指标。我联系了获胜团队的负责人 Franz,说:“太棒了,我们什么时候发布它?”他说:“哦,我们不能发布,这不实用,因为翻译一个句子需要 12 个小时。”我说:“这时间太长了,我们怎么解决?”结果发现他们并没有为高吞吐量设计。它在一个大型语言模型中进行 10 万次磁盘寻道,该模型对语料库计算统计信息——我不认为那是真正的训练。对于每个要翻译的词,它都要进行 10 万次寻道,这并不快。我说:“好吧,我们深入研究一下。”我花了大约两三个月的时间和他们一起设计了一种内存中的 n-gram 数据压缩表示。n-gram 基本上是每个 n 词序列在大语料库中出现频率的统计信息。我们有两万亿个词,当时大多数 n-gram 模型使用二元或三元模型,但我们决定使用五元模型:每个五词序列在我们能处理的尽可能多的网页中出现的频率。然后我们构建了一个数据结构,可以在 200 台机器的内存中存储所有这些信息,并提供一个批量 API,你可以说:“这一轮我需要查找这 10 万个东西,”然后它会并行返回所有结果。这使我们从花一个晚上翻译一个句子变成了大约 100 毫秒。
We had a machine translation research team at Google led by Franz Och, who had joined Google maybe a year before. Every year they competed in a DARPA contest translating a couple of languages to English, I think Chinese to English and Arabic to English. The Google team submitted an entry. The way it works is you get 500 sentences on Monday and you have to submit the answer on Friday. I saw the results: we won the contest by a substantial margin measured in BLEU score, which is a measure of translation quality. I reached out to Franz, the head of the winning team, and said, "This is great, when are we going to launch it?" He said, "Oh, we can't launch this, it's not very practical because it takes 12 hours to translate a sentence." I said, "That seems like a long time, how could we fix that?" It turned out they hadn't designed it for high throughput. It was doing 100,000 disk seeks in a large language model that computed statistics over a corpus—I wouldn't say trained really. For each word it wanted to translate, it did 100,000 seeks, which is not super speedy. I said, "Okay, let's dive into this." I spent about two or three months with them designing an in-memory compressed representation of n-gram data. An n-gram is basically statistics for how often every n-word sequence occurs in a large corpus. We had two trillion words, and most n-gram models of the day used two-grams or maybe three-grams, but we decided to use five-grams: how often every five-word sequence occurs in as much of the web as we could process. Then we built a data structure that let us store all those in memory on 200 machines, with a batched API where you could say, "Here are the 100,000 things I need to look up in this round for this word," and it would give them all back in parallel. That enabled us to go from taking a night to translate a sentence to doing it in about 100 milliseconds.
有一个 Jeff Dean 事实列表,就像 Chuck Norris 事实一样。例如,“对 Jeff 来说,NP 等于没问题。”其中一个很有趣,因为我听你说过:“光速是每小时 35 英里,直到 Jeff Dean 决定在一个周末优化它。”从 12 小时到 100 毫秒,我得算算数量级。所有这些都非常恭维,很有趣,就像我的同事们搞砸的愚人节玩笑。
There's a list of Jeff Dean facts, like Chuck Norris facts. For example, "For Jeff, NP equals No Problemo." One of them is funny because I hear you say it: "The speed of light was 35 miles per hour until Jeff Dean decided to optimize it over a weekend." Going from 12 hours to 100 milliseconds, I got to do the orders of magnitude there. All of these are very flattering, pretty funny, like an April Fool's joke gone awry by my colleagues.
回想起来,这种仅通过考虑词之间的关系就能开发出整个互联网的潜在表示的想法,就像大型语言模型,比如 Gemini。当时,这只是一个翻译的想法,还是你把它看作一种不同范式的开端?
In retrospect, this idea that you can develop a latent representation of the entire internet through just considering relationships between words is like large language models, like Gemini. At the time, was it just a translation idea, or did you see it as the beginning of a different kind of paradigm?
我认为一旦我们为翻译构建了那个系统,大型语言模型的服务就开始用于其他事情,比如补全:你开始打字,它会建议合理的补全。所以这绝对是谷歌语言模型多种用途的开始。Noam 在谷歌还做了其他一些事情,比如使用语言模型的拼写纠正系统。那是在 2000-2001 年左右,我想那完全是在一台机器的内存中运行的。他在 2001 年构建的拼写纠正系统非常棒。他向整个公司发送了演示链接,我尝试了各种我能想到的拼写错误的查询。我把“scrambled eggs benedict”拼成“Uggs bundick”,它每次都完美纠正。那就是语言建模。
I think once we built that for translation, the serving of large language models started to be used for other things, like completion: you start to type and it suggests completions that make sense. So it was definitely the start of a lot of uses of language models at Google. Noam has worked on a number of other things at Google, like spelling correction systems that use language models. That was around 2000-2001, and I think it was all in memory on one machine. His spelling correction system he built in 2001 was amazing. He sent out this demo link to the whole company, and I tried every butchered spelling of every few-word query I could get. I scrambled "Uggs bundick" instead of "scrambled eggs benedict," and it just nailed it every time. That was language modeling.
当你开发这些系统时,你有没有感觉到,如果你让它们越来越大,不只是五个词,而是 100 个词、一千个词,那么潜在表示就是智能?那个洞察是什么时候出现的?
When you were developing these systems, did you have the sense that if you make them more and more, not just five words but 100 words, a thousand words, then the latent representation is intelligence? When did that insight hit?
并没有。我不认为我曾觉得 n-gram 模型会席卷世界。当时,很多人对贝叶斯网络感到兴奋;那看起来令人兴奋。看到那些早期的神经语言模型当然很酷。但其中的魔力——好吧,这做了一些非常酷的事情——而且它让我觉得这是世界上最好的问题,因为它表述非常简单:给我下一个词的概率分布。
Not really. I don't think I ever felt like n-gram models were going to sweep the world. At the time, a lot of people were excited about Bayesian networks; that seemed exciting. Definitely seeing those early neural language models was cool. But both the magic in that—okay, this is doing something extremely cool—and also it struck me as the best problem in the world because it is very simple to state: give me a probability distribution over the next word.
此外,训练数据几乎是无限的,比如网络文本。你有数万亿个无监督、自监督的训练样本。这很好,因为你有了正确答案,你可以训练模型预测当前词,而用其他所有词。这是一种惊人的能力,仅通过观察世界来学习。如果你能在这方面做得很好,那么你几乎可以做任何事情,这就是 AI 完备的。
Also, there's roughly infinite training data out there, like the text of the web. You have trillions of training examples of unsupervised, self-supervised data. It's nice because you then have the right answer, and you can train on all but the current word and try to predict the current word. It's this kind of amazing ability to just learn from observations of the world. And then it's AI-complete: if you can do a great job of that, then you can pretty much do anything.
我很高兴介绍我们的新赞助商 Meter,一家网络公司,它支撑着全球互联网基础设施中越来越大的份额。有趣的是:大约 3 到 4 年前,在这个播客的早期,我通过 Meter CEO Anil 的捐赠来运营这个播客,直到今天,我仍然从他的建议中受益匪浅。现代世界运行在网络之上。从自动驾驶汽车到大型语言模型训练,再到像这样全球播客的广播,各个领域的进步都受限于设计和调试大型复杂网络。Meter 希望通过训练一个大型端到端基础模型,利用时间序列数据包数据、支持工单、网络教科书以及他们内部构建网络每一层所拥有的所有专有数据,为网络工程师提供 100 倍的效率提升。Meter 刚刚宣布与微软建立长期计算合作伙伴关系,以获得数万个 GPU 的访问权限。他们正在招募世界一流的 AI 研究团队。他们的目标是构建自主网络,从根本上改善我们习以为常的数字世界。了解更多信息,请访问 meter.com/aresh。
I'm excited to introduce our new sponsor, Meter, a networking company that is behind a growing fraction of the world's internet infrastructure. Fun fact: about 3 to 4 years ago, in the very early days of the podcast, I ran this podcast from a donation from Meter CEO Anil, and I continue to benefit enormously from his advice to this day. The modern world runs on networks. Progress in fields as diverse as self-driving cars to giant LLM training runs to even broadcasting a podcast like this around the world is bottlenecked on designing and debugging large complex networks. Meter wants to give network engineers a 100x multiplier by training a large end-to-end foundation model using time series packet data, support tickets, networking textbooks, and all the other proprietary data they have as a result of building every layer of the networking stack in-house. Meter just announced a long-term compute partnership with Microsoft for access to tens of thousands of GPUs. They're currently recruiting a world-class AI research team. Their goal is to build autonomous networks that radically improve the digital world that we take for granted. To learn more, go to meter.com/aresh.
好了,回到 Jeff 和 Noam。科学史上有一个有趣的讨论:想法是否只是存在于空气中,大想法具有某种必然性,还是从某个切线方向中摘取出来的?在这种情况下,我们非常逻辑地阐述它,这是否意味着它有多必然?
All right, back to Jeff and Noam. There's this interesting discussion in the history of science about whether ideas are just in the air and there's a sort of inevitability to big ideas, or whether it's sort of plucked out of some tangential direction. In this case, the way we're laying it out very logically, does that imply how inevitable this is?
确实感觉像是存在于空气中。肯定有一些关于神经图灵机的想法,所以是的,围绕注意力机制/键值存储的一堆想法,这些在神经网络中可以用来聚焦事物。所以我认为在某种意义上,它存在于空气中,在某种意义上,你需要某个团队去实现它。我喜欢把很多想法看作部分存在于空气中,有几个不同的、独立的研究想法,当试图解决一个新问题时,你会眯着眼看它们,并从中汲取灵感。然后有一些方面尚未解决,你需要弄清楚如何解决。现有事物的某种变形和新事物的结合,导致了以前不存在的新突破或研究成果。
It does feel like it's in the air. There were definitely some ideas around this neural Turing machine, so yeah, a bunch of ideas around this attention slash key-value stores that could be useful in neural networks to kind of focus on things. So I think in some sense it's in the air, and in some sense you need some group to go do it. I like to think of a lot of ideas as partially in the air, where there are a few different separate research ideas that one is kind of squinting at when trying to solve a new problem, and you draw on those for inspiration. Then there's some aspect that is not solved, and you need to figure out how to solve that. The combination of some morphing of existing things and some new things leads to a new breakthrough or research result that didn't exist before.
有没有一些关键时刻让你印象深刻,你看着一个研究领域,想出了这个想法,然后有那种‘天哪,我不敢相信这居然成功了’的感觉?
Were there key moments that stand out to you where you looked at a research area and came up with this idea and had this feeling of like, 'Holy, I can't believe that worked'?
我记得的一件事是在大脑团队的早期。我们专注于看看能否构建一些基础设施,让我们训练非常非常大的神经网络。那时,我们的数据中心没有 GPU,只有 CPU,但我们知道如何让大量 CPU 协同工作。所以我们构建了一个系统,使我们能够通过模型并行和数据并行来训练相当大的神经网络。我们有一个系统,用于对实际随机选择的 1000 万个 YouTube 帧进行无监督学习。它是一种空间局部表示,因此它会基于从高层表示重建事物的尝试来构建无监督表示。我们让它在 2000 台计算机上使用 16000 个核心进行训练。过了一段时间,那个模型实际上能够在最高层构建一个表示,其中一个神经元会被猫的图片激活。它从未被告知猫是什么,但它在训练数据中看到了足够多的猫正面脸部视图的例子,以至于那个神经元会为此激活,而对其他东西则不然。类似地,你会有其他神经元用于人脸和行人的背部。这很酷,因为从无监督学习原理出发,它构建了这些非常高级的表示。然后我们能够在有监督的 ImageNet 20000 类挑战中获得非常好的结果,将最先进水平提高了约 60%的相对改进,这在当时相当不错。那个神经网络可能比之前训练过的网络大 50 倍,并且取得了良好的结果。这告诉我:‘嘿,实际上扩大神经网络规模似乎是个好主意,而且看起来有效,所以我们应该继续推进。’
One thing I remember was in the early days of the brain team. We were focused on seeing if we could build some infrastructure that lets us train really, really big neural nets. At that time, we didn't have GPUs in our data centers, just CPUs, but we knew how to make lots of CPUs work together. So we built a system that enabled us to train pretty large neural nets through both model and data parallelism. We had a system for unsupervised learning on actually 10 million randomly selected YouTube frames. It was a spatially local representation, so it would build up unsupervised representations based on trying to reconstruct the thing from the high-level representations. We got that working and training on 2,000 computers using 16,000 cores. After a little while, that model was actually able to build a representation at the highest level where one neuron would get excited by images of cats. It had never been told what a cat was, but it had seen enough examples in the training data of head-on facial views of cats that that neuron would turn on for that and not for much else. Similarly, you'd have other ones for human faces and backs of pedestrians. That was kind of cool because from unsupervised learning principles, it built up these really high-level representations. Then we were able to get very good results on the supervised ImageNet 20,000-category challenge, advancing the state-of-the-art by like 60% relative improvement, which was quite good at the time. That neural net was probably 50x bigger than one that had been trained previously, and it got good results. That said to me, 'Hey, actually scaling up neural nets seems like a good idea, and it seems to be working, so we should keep pushing on that.'
这些例子说明了这些 AI 系统如何融入你刚才提到的:谷歌本质上是一家组织信息的公司。在这种情况下,AI 所做的是在信息之间、概念之间寻找关系,以帮助你更快地获得想法,更快地获得你想要的信息。现在我们正在使用当前的 AI 模型。显然,它们非常擅长信息检索,但更根本的是,它们可以为你编写整个代码库,做更多像实际工作者一样的事情,这超越了单纯的信息检索。那么你的想法改变了吗?如果你在构建 AGI,谷歌还是一家信息检索公司吗?AGI 可以做信息检索,但它也可以做很多其他事情,对吧?
These examples illustrate how these AI systems fit into what you were just mentioning: that Google is sort of a company that organizes information fundamentally. What AI is doing in this context is finding relationships between information, between concepts, to help get ideas to you faster, information you want faster. Now we're moving with current AI models. Obviously they're very good at information retrieval, but more fundamentally, they can write your entire code base for you and do more like an actual worker, which is going beyond just information retrieval. So has your thinking changed? Is Google still an information retrieval company if you're building an AGI? AGI can do information retrieval but it can do many other things as well, right?
我认为我们是一家‘组织世界信息’的公司,这比信息检索更广泛。这可能是根据你给出的指导来组织和创造新信息。‘你能帮我给我的兽医写一封关于我的狗的信吗?它有这些症状。’然后它会起草那封信。或者‘你能输入这段视频,并每隔几分钟生成一段视频中发生事情的摘要吗?’我认为我们的多模态能力表明,这不仅仅是文本;它是关于理解信息存在的所有不同模态的世界,既包括面向人类的模态,也包括非面向人类的模态。
I think we're an 'organize the world's information' company, and that's broader than information retrieval. That's maybe organizing and creating new information from some guidance you give it. 'Can you help me write a letter to my veterinarian about my dog? It's got these symptoms.' And it'll draft that. Or 'Can you feed in this video and produce a summary of what's happening in the video every few minutes?' I think our multimodal capabilities are showing that it's more than just text; it's about understanding the world in all the different modalities that information exists in, both human-oriented ones but also non-human-oriented ones.
比如自动驾驶汽车上的奇怪传感器,或者基因组信息、健康信息,然后如何提取并转化为对人们有用的见解,帮助他们做各种想做的事情?有时是我想和聊天机器人聊天娱乐,有时是我想要一个非常复杂问题的答案,没有单一来源可以检索,你需要从一百个网页中提取信息,弄清楚发生了什么,并整理合成这些数据,然后处理多模态或编码相关问题。我认为这些模型的能力非常令人兴奋,而且它们进步很快,所以我很期待看到未来的发展。
Weird sensors on autonomous vehicles, or genomic information, or health information, and then how do you extract and transform that into useful insights for people and make use of that in helping them do all kinds of things they want to do? Sometimes it's I want to be entertained by chatting with a chatbot, sometimes it's I want answers to this really complicated question where there is no single source to retrieve from, you need to pull information from like a hundred web pages and figure out what's going on and make an organized synthesized version of that data, and then dealing with multimodal things or coding related problems. I think it's super exciting what these models are capable of and they're improving fast, so I'm excited to see where we go.
我也很期待看到未来的发展。我认为组织信息显然是一个万亿美元的机会,但万亿美元已经不够酷了,酷的是千万亿美元。显然,想法不是堆砌一大堆钱,而是在世界上创造价值。当这些系统能够真正为你做事、编写代码或解决你自己无法解决的问题,并且大规模地做到这一点时,就能创造更多的价值。随着我们提高这些模型的能力,我们必须非常灵活和动态。我对许多基础研究问题感到非常兴奋,因为你会发现如果我们尝试这种方法或这个大致方向,我们正在做的事情可以得到显著改善。也许行得通,也许不行。但我也认为,看看我们能为用户实现什么,然后从那里倒推来构建能够做到这一点的系统,是有价值的。举个例子,组织信息应该意味着世界上的任何信息都应该能被任何人使用,无论他们说什么语言。我们已经做了一些,但远未达到完整的愿景:无论你说数千种语言中的哪一种,我们都能让你获得任何内容并为你所用。任何视频都可以用任何语言观看。我认为那会很棒。我们还没有完全实现,但这绝对是我看到的地平线上应该可能的事情。
I am also excited to see where we go. I think organizing information is clearly a trillion dollar opportunity, but a trillion dollars is not cool anymore, what's cool is a quadrillion dollars. Obviously the idea is not to pile up some giant pile of money, but to create value in the world. So much more value can be created when these systems can actually go and do something for you, write your code, or figure out problems that you wouldn't have been able to figure out yourself, and to do that at scale. We're going to have to be very flexible and dynamic as we improve the capabilities of these models. I'm pretty excited about a lot of fundamental research questions that come about because you see something we're doing could be substantially improved if we tried this approach or things in this rough direction. Maybe that'll work, maybe it won't. But I also think there's value in seeing what we could achieve for end users and then working backwards from that to actually build systems that are able to do that. As one example, organizing information should mean any information of the world should be usable by anyone regardless of what language they speak. We've done some amount of that, but it's not nearly the full vision of no matter what language you speak out of thousands of languages, we can make any piece of content available to you and make it usable by you. Any video could be watched in any language. I think that would be pretty awesome. We're not quite there yet, but that's definitely things I see on the horizon that should be possible.
说到你可能尝试的不同架构,我知道你现在正在做的一件事是更长的上下文。如果你把谷歌搜索看作它的上下文中有整个互联网的索引,但它就像一个非常浅层的搜索。而显然语言模型现在有有限的上下文,但它们可以真正思考,就像黑魔法一样,上下文学习,它们可以真正思考它们看到的东西。你认为将谷歌搜索和上下文学习之类的东西合并起来会是什么样子?
Speaking of different architectures you might try, I know one thing you're working on right now is longer context. If you think of Google search as it's got the entire index of the internet in its context, but it's like a very shallow search. And then obviously language models have limited context right now, but they can really think, it's like dark magic, in-context learning, they can really think about what they're seeing. How do you think about what it would be like to merge something like Google search and something like in-context learning?
我先迈出第一步,因为我对这个问题思考过一段时间。我认为这些模型的一个特点是它们相当不错,但有时会产生幻觉和事实性问题。部分原因是你在数万亿个词元上训练,并将所有这些混合在数千亿个参数中,但这一切都有点模糊,因为你把所有这些词元搅在了一起。模型对这些数据有相当清晰的视图,但有时会混淆,给出错误的日期。而上下文窗口中的信息,即模型的输入,非常清晰和明确,因为我们在 Transformer 中有非常好的注意力机制,模型可以关注事物,并且知道正在处理的精确文本或视频的精确帧或音频等。现在我们有一些模型可以处理数百万个词元的上下文,这相当多,比如数百页的 PDF,或 50 篇研究论文,或数小时的视频,或数十小时的音频,或这些的组合,这很酷。但如果模型能够关注数万亿个词元,能够关注整个互联网并为你找到正确的东西,那将非常好。它能否关注你所有的个人信息?我希望有一个模型可以访问我所有的电子邮件、所有文档和所有照片,当我要求它做某事时,它可以在我的许可下利用这些来帮助解决我想要它做的事情。但这将是一个巨大的计算挑战,因为朴素的注意力算法是二次的,你几乎只能在相当多的硬件上处理数百万个词元,但不可能天真地扩展到数万亿个词元。所以我们需要大量有趣的算法近似来实现你真正想要的,让模型在概念上能够关注更多更多的词元,达到数万亿个词元。也许我们可以把谷歌的所有代码库放在每个谷歌开发者的上下文中,把世界上所有的源代码放在任何开源开发者的上下文中。那将是惊人的,难以置信的。
I'll take a first step at it because I've thought about this for a bit. I think one of the things you see with these models is they're quite good but they do hallucinate and have factuality issues sometimes. Part of that is you've trained on say tens of trillions of tokens and you've stirred all that together in your tens or hundreds of billions of parameters, but it's all a bit squishy because you've churned all these tokens together. The model has a reasonably clear view of that data but it sometimes gets confused and will give the wrong date for something. Whereas information in the context window, in the input of the model, is really sharp and clear because we have this really nice attention mechanism in Transformers that the model can pay attention to things and it knows the exact text or the exact frames of the video or audio or whatever it's processing. Right now we have models that can deal with millions of tokens of context, which is quite a lot, like hundreds of pages of a PDF, or 50 research papers, or hours of video, or tens of hours of audio, or some combination of those things, which is pretty cool. But it would be really nice if the model could attend to trillions of tokens, could it attend to the entire internet and find the right stuff for you? Could it attend to all your personal information for you? I would love a model that has access to all my emails and all my documents and all my photos, and when I ask it to do something it can make use of that with my permission to help solve what I want it to do. But that's going to be a big computational challenge because the naive attention algorithm is quadratic and you can barely make it work on a fair bit of hardware for millions of tokens, but there's no hope of making that just naively go to trillions of tokens. So we need a whole bunch of interesting algorithmic approximations to what you would really want, to make a way for the model to attend conceptually to lots and lots more tokens, to trillions of tokens. Maybe we can put all of the Google codebase in context for every Google developer, all the world's source code in context for any open source developer. That would be amazing, it would be incredible.
模型参数的美妙之处在于它们在记忆事实方面非常高效。你可能每个模型参数大约能记住一个事实。而如果你在上下文中有一些词元,每一层都有键和值,每个词元可能占用千字节或兆字节的内存。你拿一个词,把它放大到 10 千字节。所以实际上有很多创新围绕如何最小化这一点,以及你需要哪些词,是否有更好的方式来访问这些信息。杰夫似乎是解决这个问题的合适人选,比如我们的内存层次结构从 SRAM 一直到全球数据中心级别是什么样子。
The beautiful thing about model parameters is they are quite memory efficient at memorizing facts. You can probably memorize on the order of one fact per model parameter. Whereas if you have some token in context, there are keys and values at every layer, could be a kilobyte or a megabyte of memory per token. You take a word and you blow it up to 10 kilobytes. So there is actually a lot of innovation going on around how to minimize that, and what words do you need to have, and are there better ways of accessing bits of that information. Jeff seems like the right person to figure this out, like what does our memory hierarchy look like from SRAM all the way up to data center worldwide level.
如果你只考虑这一个用例及其含义。你有像 Google 的单一代码库这样的东西,如果你解决了长上下文的问题,你可以把整个代码库放入上下文,或者在上面进行微调。基本上,为什么这还没有做到?因为你可以想象 Google 拥有专有访问权限的代码量,即使只是内部使用,也能让你的开发者更高效、更有生产力。
If you just think about that one use case and what that implies. So you've got like the Google monorepo, and if you maybe figure out the long context thing, you can put the whole thing in context or fine-tune on it. Basically, why hasn't this been already done? Because you can imagine the amount of code that Google has proprietary access to, just to make even if you're just using it internally to make your developers more efficient and productive.
哦,澄清一下,我们实际上已经在内部代码库上对 Gemini 模型进行了进一步训练,供内部开发者使用。但这与关注所有代码不同,对吧?因为它把代码库混合成了一堆参数。我认为将其放在上下文中会让事情更清晰。但即使是内部进一步训练的模型也非常有用。我记得 Sundar 说过,我们最近提交到代码库的字符中有 25% 是由我们的 AI 编码模型生成的,有人类参与...。
Oh, to be clear, we have actually already done further training on a Gemini model on our internal code base for internal developers. But that's different than attending to all of it, right? Because it sort of stirs together the code base into a bunch of parameters. And I think having it in context makes things clearer. But even the further trained model internally is incredibly useful. I think Sundar has said that 25% of the characters that we're checking into our code base these days are generated by our AI-based coding models, with human kind of...
根据你看到的前沿能力,你想象一两年后你自己的个人工作会是什么样子?作为 Google 的研究员会是什么样?你有了一个新想法,一年后你与这些模型互动的方式会是什么样?
How do you imagine in a year or two, based on the capabilities you see on the horizon, your own personal work? What will it be like to be a researcher at Google? You have a new idea or something, with the way in which you're interacting with these models in a year. What does that look like?
嗯,我认为我们会拥有好得多的模型,希望能变得更有生产力。我认为除了研究背景之外,任何时候你看到这些模型被使用,它们都能让软件开发者更高效,因为它们可以接受一个高层规格或一句话描述你想要做的事情,并给出一个相当合理的初稿。所以从研究的角度来看,也许你可以说,‘我真的希望你探索这种想法,类似于这篇论文中的,但也许我们试试让它变成卷积的或类似的。’如果你能做到这一点,让系统自动生成一堆实验代码,然后你看着它说,‘嗯,看起来不错,运行它。’这似乎是一个很好的梦想方向,而且在一两年内取得很大进展似乎是可行的。这似乎被低估了,因为你实际上可以有数百万的额外员工,你可以立即检查他们的输出。员工可以互相检查输出,他们可以立即流式传输 token。抱歉,我不是故意炒作。我认为这非常令人兴奋,我只是不喜欢炒作尚未完成的事情。
Well, I assume we will have these models a lot better and hopefully be able to be much more productive. I think one of the things, in addition to researchy context, anytime you're seeing these models used, they're able to make software developers more productive because they can take a high-level spec or a sentence description of what you want done and give a pretty reasonable first cut at that. So from a research perspective, maybe you can say, 'I'd really like you to explore this kind of idea, similar to the one in this paper, but maybe let's try making it convolutional or something like that.' If you could do that and have the system automatically generate a bunch of experimental code, and maybe you look at it and say, 'Yeah, that looks good, run that.' That seems like a nice dream direction to go in and seems plausible in the next year or two that you might make a lot of progress on that. And it seems underhyped because you could have literally millions of extra employees, and you can immediately check their output. The employees can check each other's output, they like immediately stream tokens. Sorry, I didn't mean to hype it. I think it's super exciting, I just don't like to hype things that aren't done yet.
让我们更深入地探讨这个想法,因为如果你拥有像自主软件工程师这样的东西,这似乎是一件大事,尤其是从一个想要设计和构建系统的研究者的角度来看。作为一个职业生涯中致力于开发变革性系统的人,这个想法是,不必编码像今天相当于 MapReduce 或 TensorFlow 之类的东西,只需‘这是我想要的分布式 AI 库的样子,帮我写出来。’你能想象你可以提高 10 倍、100 倍的生产力吗?
Let's play with this idea more because it seems like a big deal if you have something like an autonomous software engineer, especially from the perspective of a researcher who wants to spec and build the system. As somebody who has worked on developing transformative systems through your career, the idea that instead of having to code something like whatever the today's equivalent of MapReduce or TensorFlow is, just 'here's how I would like a distributed AI library to look, write it up for me.' Could you imagine you could be 10x more productive, 100x more productive?
我印象非常深刻。我记得是在 Reddit 上看到的,我们有一个新的实验性编码模型,在编码和数学方面要好得多,有人外部尝试了它。他们基本上提示它说,‘我希望你实现一个没有外部依赖的 SQL 处理数据库系统,请这样做。’根据那个人的说法,它实际上做得相当好。它生成了一个 SQL 解析器和一个分词器,一个查询规划系统,以及一些磁盘上数据的存储格式,并且实际上能够处理简单的查询。所以从那个提示(就像一段文字)到得到一个初稿,这似乎对软件开发者的生产力是一个巨大的提升。我认为你可能最终会有其他类型的系统,它们可能不会尝试在 40 秒内半交互式地响应,而是可能运行 10 分钟,并在五分钟后打断你,说,‘啊,我已经做了很多,但现在我需要一些输入。你关心处理视频还是只处理图像之类的?’如果你有很多这种后台活动发生,似乎你需要管理工作流的方法。
I was pretty impressed. I think it was on Reddit that I saw we have a new experimental coding model that's much better at coding and math, and someone external tried it. They basically prompted it and said, 'I'd like you to implement a SQL processing database system with no external dependencies, please do that.' And from what the person said, it actually did a quite good job. It generated a SQL parser and a tokenizer, a query planning system, and some storage format for the data on disk, and actually was able to handle simple queries. So from that prompt, which is like a paragraph of text, to get even an initial cut at that seems like a big boost in productivity for software developers. And I think you might end up with other kinds of systems that maybe don't try to do that in a single semi-interactive respond in 40 seconds kind of thing, but might go off for 10 minutes and might interrupt you after five minutes saying, 'Ah, I've done a lot of this, but now I need to get some input. Do you care about handling video or just images or something?' And that seems like you'll need ways of managing the workflow if you have a lot of these background activities happening.
你能多谈谈这个吗?如果你真的可以按需启动数百万个能够以极快速度打字的员工,你想象我们需要什么样的界面?这几乎就像从 1930 年代的票务交易到现代的连锁套装之类的东西。你需要一个更好的界面来跟踪所有这一切,让 AI 集成到这个大的单一代码库中并利用它们自己的优势,让人类跟踪正在发生的事情。
Can you talk more about that? What interface do you imagine we might need if you could literally have millions of employees you could spin up on command, who are able to type incredibly fast? It's almost like you go from 1930s trading of tickets to modern chain suit or something. You need a better interface to keep track of all this going on, for the AI to integrate into this big monorepo and leverage their own strengths, for humans to keep track of what's happening.
三年后 Noam 的日常工作会是什么样子?可能和现在类似,因为我们已经有了并行化这个主要问题。我们有很多很多非常聪明的机器学习研究员,我们希望他们一起工作并构建 AI。所以实际上人与人之间的并行化可能类似于机器之间的并行化。但我认为这肯定对需要大量探索的事情有好处,比如提出下一个突破。如果你有一个天才的想法,在机器学习领域它肯定能成功,但即使你很聪明,它也只有 2% 的成功率。大多数情况下这些东西都会失败。但如果你尝试 100 件事、1000 件事或 100 万件事,那么你可能会发现一些惊人的东西。而且我们有大量的算力。现代顶级实验室现在拥有的算力可能是训练 Transformer 所需算力的 100 万倍。
What is it like to be Noam in three years working day-to-day? It might be kind of similar to what we have now because we already have parallelization as a major issue. We have lots and lots of really brilliant machine learning researchers, and we want them to work all together and build AI. So actually the parallelization among people might be similar to parallelization among machines. But I think it definitely should be good for things that require a lot of exploration, like coming up with the next breakthrough. If you have a brilliant idea that's just certain to work in the ML domain, it has a 2% chance of working if you're brilliant. Mostly these things fail. But if you try 100 things or a thousand things or a million things, then you might hit on something amazing. And we have plenty of compute. Modern top labs these days have probably a million times as much compute as it took to train the Transformer.
所以这真是一个有趣的想法。假设当今世界这个领域大约有 10,000 名研究人员在提出突破。可能更多。NeurIPS 有 15,000 人?100,000 人?我不知道。
So that's a really interesting idea. Suppose in the world today there's on the order of 10,000 researchers in this community coming up with a breakthrough. Probably more than that. There were 15,000 in NeurIPS? 100,000? I don't know.
也许吧。
Maybe.
不,有个正确的数量级是好的。这个社区每年出现 Transformer 级别突破的概率,假设是 10%。现在假设这个社区规模扩大一千倍,某种程度上就像并行搜索更好的架构和技术。我们是不是每年甚至每天都能得到 Transformer 级别的突破?听起来可能不错,但这是机器学习研究的样子吗?只要你能尝试所有实验?
No, it's good to have the correct order of magnitude. And the odds of this community every year coming up with a breakthrough on the scale of a Transformer is, let's say, 10%. Now suppose this community is a thousand times bigger, and it is in some sense like a parallel search for better architectures and better techniques. Do we just get Transformer-like breakthroughs every year or every day? Maybe sounds potentially good, you know. But does that feel like what ML research is like? Just if you are able to try all these experiments?
这是个好问题,因为我不确定大家是否在这么做。我们确实有很多好想法,但似乎每个人都想以最大规模运行实验。我认为这是人的问题。
It's a good question because, you know, I don't know that folks have been doing that as much. I mean, we definitely have lots of great ideas coming along, but everyone seems to want to run their experiment at maximum scale. I think that's a human problem.
是的,很有帮助。先有一个千倍规模的问题,在上面验证十万个想法,然后把有希望的放大。
Yeah, yeah. It's very helpful to have a 1,000-scale problem and then vet like 100,000 ideas on that, and then scale up the ones that seem promising.
快速插播赞助商消息:Scale AI。公开可用数据正在枯竭,因此 Meta、Google DeepMind 和 OpenAI 等主要实验室都与 Scale 合作,通过 Scale 的数据铸造厂突破可能性的边界。主要实验室获得高质量数据以支持后训练,包括高级推理能力。随着 AI 飞速发展,我们还必须加强人类主权。Scale 的研究团队 SEAL 提供实用的 AI 安全框架,通过公开排行榜评估前沿 AI 系统安全性,并为将高级 AI 融入社会奠定基础。最近,Scale 与 AI 安全中心合作发布了《人类最后的考试》,这是一个开创性的新 AI 基准,用于评估 AI 系统在广泛领域的专家级知识和推理。如果你是 AI 研究人员或工程师,想了解更多关于 Scale 的数据铸造厂和研究团队如何帮助你超越当前能力前沿,请访问 scale.com/dwares。好了,回到 Jeff。
A quick word from our sponsor: Scale AI. Publicly available data is running out, so major labs like Meta, Google DeepMind, and OpenAI all partner with Scale to push the boundaries of what's possible through Scale's Data Foundry. Major labs get access to high-quality data to fuel post-training, including advanced reasoning capabilities. As AI races forward, we must also strengthen human sovereignty. Scale's research team, SEAL, provides practical AI safety frameworks, evaluates frontier AI system safety via public leaderboards, and creates foundations for integrating advanced AI into society. Most recently, in collaboration with the Center for AI Safety, Scale published 'Humanity's Last Exam', a groundbreaking new AI benchmark for evaluating AI systems' expert-level knowledge and reasoning across a wide range of fields. If you're an AI researcher or engineer and you want to learn more about how Scale's Data Foundry and research team can help you go beyond the current frontier of capabilities, go to scale.com/dwares. All right, back to Jeff.
我认为世界可能没有认真对待一件事:人们知道让模型大 100 倍、算力多 100 倍是指数级困难的。从 Gemini 2 到 3 是指数级更难的问题。但人们可能没有意识到另一个趋势:Gemini 3 不断提出各种不同的架构想法并尝试,看到什么有效,然后不断取得算法进步,使训练下一个模型越来越容易。这个反馈循环能走多远?
So I think one thing the world might not be taking seriously: people are aware that it's exponentially harder to make a model that's 100x bigger, 100x more compute, right? People know that's an exponentially harder problem to go from Gemini 2 to 3, and so forth. But maybe people aren't aware of this other trend where Gemini 3 is coming up with all these different architectural ideas and trying them out, and you see what works, and you're constantly coming up with algorithmic progress that makes training the next one easier and easier. How far could you take that feedback loop?
我认为人们应该意识到,这些模型代际之间的改进通常部分由硬件和更大规模驱动,但同样甚至更多地由重大算法改进、模型架构和训练数据组合的重大变化驱动,这些真正使模型在每 flop 上变得更好。所以我认为这是个很好的认识。然后我认为如果我们有自动化探索,我们将能够验证更多想法,并将它们带入下一代模型的实际生产训练中。这将非常有帮助,因为这就是我们目前用大量机器学习研究在做的事情:杰出的机器学习研究人员审视大量想法,筛选出在小规模上效果好的,看看它们在中规模上是否有效,将它们带入更大规模的实验,然后最终为最终模型配方添加大量新的有趣的东西。然后我认为如果我们能通过机器学习研究人员温和地引导更自动化的搜索过程,而不是他们自己手动照看大量实验,从而将速度提高 100 倍,那将非常非常好。
I think one thing people should be aware of is that the improvements from generation to generation of these models are often partially driven by hardware and larger scale, but equally and perhaps even more so driven by major algorithmic improvements and major changes in the model architecture and the training data mix, and so on, that really make the model better per flop applied to the model. So I think that's a good realization. And then I think if we have automated exploration, we'll be able to vet a lot more ideas and bring them into the actual production training for next generations of these models. That's going to be really helpful because that's sort of what we're currently doing with a lot of machine learning research: brilliant machine learning researchers looking at lots of ideas, winnowing ones that seem to work well at small scale, seeing if they work well at medium scale, bringing them into larger scale experiments, and then settling on adding a whole bunch of new and interesting things to the final model recipe. And then I think if we can do that 100 times faster through those machine learning researchers just gently steering a more automated search process rather than sort of hand-babysitting lots of experiments themselves, that's going to be really, really good.
是的,唯一不能加速的是最大规模的实验,因为你仍然要做这些 n=1 的实验,你只能把一群非常聪明的人聚在一起,让他们盯着东西,弄清楚为什么这个有效,为什么那个无效。更多硬件是个好办法,更好的硬件,是的,我们指望你了。
Yeah, the one thing it doesn't speed up is experiments at the largest scale, because you still end up doing these n=1 experiments, and you really just try to put a bunch of really brilliant people in the room and have them stare at the thing, figure out why this is working, why this is not. More hardware is a good solution, and better hardware, yes, we're counting on you.
那么,好吧。天真的想法是,未来可以在算法方面取得进步。还有你在芯片外做的工作,我让你来描述。但如果出现这样一种情况:仅从软件层面,你就能在几周或几个月内制造出越来越好的芯片,而更好的 AI 大概能做得更好,基本上我想知道这个反馈循环如何不会导致:Gemini 3 需要两年,而 Gemini 4(同等水平的跳跃)现在只需 6 个月,然后第 5 级是 3 个月,然后 1 个月,你就会因为硬件和算法两方面的软件改进,比天真想象中快得多地达到超级智能?
So, okay. And naively, there's this algorithmic side improvement the future can make. There's also the stuff you're working on off-chip. I'll let you describe it. But if you get into a situation where just from a software level you can be making better and better chips in a matter of weeks and months, and better AIs can presumably do that better, basically I'm wondering how does this feedback loop not just end up in like Gemini 3 takes two years and Gemini 4, the equivalent level jump, is now 6 months, then the level 5 is like 3 months, then 1 month, and you get to superhuman intelligence much more rapidly than you might naively think because of this software, both on the hardware side and from the algorithmic side improvements?
是的,我最近对如何大幅加速芯片设计过程感到非常兴奋。因为正如我们之前讨论的,当前设计芯片的方式大约需要 18 个月,从‘我们应该造一个芯片’到交给台积电,然后台积电需要大约四个月来制造,然后你拿回来放到数据中心。所以这是一个相当长的周期。其中制造时间目前只占很小一部分。但如果你能让制造成为主导部分,这样设计芯片不再需要 12 到 18 个月,而是可以用 100 到 150 人缩短到几个人,通过更自动化的搜索过程探索整个芯片设计空间,并从芯片设计过程的各个方面获得关于系统试图探索的高层选择的反馈。那么我认为你可以获得更多的探索,更快速地设计出你真正想交给工厂的东西。那将很棒,因为你可以缩短时间,通过以正确的方式设计硬件来缩短部署时间,这样你拿到芯片后直接插入系统。然后我认为这将实现更多的专业化,缩短硬件设计的时间框架,这样你就不必看得太远,考虑什么样的 ML 算法会有趣。相反,你就像是在看……
Yeah, I've been pretty excited lately about how could we dramatically speed up the chip design process. Because as we were talking earlier, the current way in which you design a chip takes you roughly 18 months to go from 'we should build a chip' to something that you then hand over to TSMC, and then TSMC takes you know four months to fab it, and then you get it back and you put it in your data centers. So that's a pretty lengthy cycle. And the fab time in there is a pretty small portion of it today. But if you could make that the dominant portion, so that instead of taking 12 to 18 months to design the chip, you could shrink it with, you know, 100-150 people, you could shrink that to a few people with a much more automated search process exploring the whole design space of chips and getting feedback from all aspects of the chip design process for the kind of choices that the system is trying to explore at the high level. Then I think you could get perhaps much more exploration and more rapid design of something that you actually want to give to a fab. And that would be great because you can shrink that time, you can shrink the deployment time by designing the hardware in the right way so that you just get the chips back and you just plug them into some system. And that will then, I think, enable a lot more specialization, it will enable a shorter time frame for the hardware design so that you don't have to look out quite as far into what kind of ML algorithms would be interesting. Instead, it's like you're looking at...
你知道,从现在起六到九个月,而不是两到两年半,那会很酷。我确实认为,如果制造时间是你改进的内循环,你会……这需要多久?不幸的是,最先进的节点因为金属层比以前的旧节点更多,所以越来越耗时,通常需要三到五个月。
You know, six to nine months from now, what should it be rather than two to two and a half years? That would be pretty cool. I do think that fabrication time, if that's in your inner loop of improvement, you're going to... How long is it? The leading-edge nodes unfortunately are taking longer and longer because they have more metal layers than previous older nodes, so that tends to make it take anywhere from three to five months.
好吧,但训练运行本身也需要那么长时间,对吧?所以你可以同时进行两者。
Okay, but that's how long training runs take anyway, right? So you could potentially do both at the same time.
是的,有可能。所以我想你不可能比三到五个月更快,但你可以……同时,你在这段时间内快速开发新的算法思路,这些思路可以快速推进,在现有芯片上运行并探索很多酷想法。
Yeah, potentially. So I guess you can't get sooner than three to five months, but the idea that you could... Also, you're rapidly developing new algorithmic ideas between time, that can move fast, that can run on existing chips and explore lots of cool ideas.
所以这不就是一种情况,你会觉得……我认为人们有点期待,‘啊,会出现一个 S 形曲线。’再说一次,这不是确定的事,但只是说,这是否可能?即能力在人类智能的尾端迅速爆发,以越来越快的速度变得越来越聪明。
So isn't that a situation in which you're like... I think people sort of expect, 'Ah, there's going to be a sigmoid.' Again, this is not a sure thing, but just like, is this a possibility? The idea that you have sort of an explosion of capabilities very rapidly towards the tail end of human intelligence, that gets smarter and smarter at a more and more rapid rate.
有可能,是的。我的意思是,我喜欢这样想:现在我们有一些模型,可以处理相当复杂的问题,在内部将其分解成一系列步骤,拼凑出这些步骤的解决方案,并经常给出整个问题的答案。但它并不非常可靠,而且它擅长将问题分解成 5 到 10 个步骤,而不是 100 到 1000 个步骤。所以,如果你能从 80%的时间完美回答一个 10 步的问题,提高到 90%的时间完美回答一个包含 100 到 1000 个子步骤的问题,那将是能力的惊人提升。我们还没到那一步,但那是我们努力追求的目标。
Possibly, yeah. I mean, I like to think of it like this: right now we have models that can take a pretty complicated problem, break it down internally into a bunch of steps, puzzle together the solutions for those steps, and often give you a solution to the entire problem. But it isn't super reliable, and it's good at breaking things down into 5 to 10 steps, not 100 to 1,000 steps. So if you could go from 80% of the time giving a perfect answer to something that's 10 steps long, to 90% of the time giving a perfect answer to something that's 100 to 1,000 sub-problem steps long, that would be an amazing improvement in capability. We're not there yet, but that's what we're aspirationally trying to get to.
我们不需要新硬件来实现这一点,但我的意思是,有的话我们也不拒绝。
We don't need new hardware for that, but I mean we'll take it.
是的,完全正确。永远不要对新的硬件礼物挑三拣四。近期改进的一大领域是推理时算力——在推理时应用更多算力。我喜欢这样描述:即使是一个巨大的语言模型,每个 token 执行一万亿次操作(这比大多数人现在做的要多),每次操作的成本大约是 10^-18,所以每美元能得到大约一百万个 token。相比之下,一个相对便宜的消遣,比如买一本平装书来读,你每美元只能得到大约一万个 token。所以与语言模型对话可能比读平装书便宜 100 倍。这里有巨大的空间可以说,‘好吧,如果我们能让这个东西更贵但更聪明,因为我们比读平装书便宜 100 倍,比与客服代理交谈便宜一万倍,或者比雇佣软件工程师或咨询医生或律师便宜一百万倍——我们能增加算力让它更聪明吗?’所以我认为我们在不久的将来会看到的很多起飞都是这种形式。过去我们一直在利用和改进预训练,后训练也会继续改进,但利用推理时更深入的思考将是一场爆发。
Yeah, exactly. Never look a new hardware gift horse in the mouth. One of the big areas of improvement in the near future is inference time compute—applying more compute at inference time. I like to describe it like this: even a giant language model, doing say a trillion operations per token (which is more than most people are doing these days), costs something like 10^-18 per operation, so you get about a million tokens per dollar. Compare that to a relatively cheap pastime like buying a paperback book and reading it—you're paying like 10,000 tokens per dollar. So talking to a language model could be 100 times cheaper than reading a paperback. There's huge headroom to say, 'Okay, if we can make this thing more expensive but smarter, because we're 100x cheaper than reading a paperback, 10,000x cheaper than talking to a customer support agent, or a million times cheaper than hiring a software engineer or talking to your doctor or lawyer—can we add computation and make it smarter?' So I think a lot of the takeoff we're going to see in the very near future is of this form. We've been exploiting and improving pre-training a lot in the past, and post-training will continue to improve, but taking advantage of thinking harder at inference time is going to be an explosion.
推理时的一个方面是,我认为你希望系统主动探索一系列不同的潜在解决方案。也许它自己做一些搜索,获取信息,消化这些信息,然后发现,‘哦,我现在真的想了解更多关于这个的东西’,于是它迭代地探索如何最好地解决你提出的高层次问题。拥有一个旋钮,让你可以通过更多推理时算力让模型给出更好的答案,似乎我们已经有技术可以做到这一点。你越调高旋钮,算力成本越高,但答案也越好。这似乎是一个不错的权衡,因为有时你需要为一个非常重要的问题深思熟虑,而有时你可能不想花费巨大的算力来计算一加一。也许系统应该默认使用计算器工具,而不是一个非常大的语言模型。
An aspect of inference time is I think you want the system to be actively exploring a bunch of different potential solutions. Maybe it does some searches on its own, gets information back, consumes that information, and figures out, 'Oh, now I would really like to know more about this thing,' so it iteratively explores how to best solve the high-level problem you pose. Having a dial where you can make the model give better answers with more inference time compute seems like we have techniques that can do that. The more you crank up the dial, the more it costs in compute, but the better the answers get. That seems like a nice trade-off, because sometimes you want to think really hard for a super important problem, and sometimes you probably don't want to spend enormous compute to compute one plus one. Maybe the system should default to using a calculator tool instead of a very large language model.
在推理时算力方面,是否存在任何障碍,比如能否线性扩展,用 100 倍或 1000 倍的算力得到相应更好的结果?嗯,我们正在边做边研究算法。我相信,随着超过一万名研究人员(包括谷歌的许多人)不断钻研,我们会看到越来越好的解决方案。在我们自己的实验工作中,确实有例子表明,应用更多推理时算力比应用较少算力能得到更好的答案。这看起来有用且重要。但我们希望的是,当你应用 10 倍算力时,答案质量的提升比今天更大。这关乎设计新算法、尝试新方法、找出如何最好地利用那 10 倍算力来改进结果。
Are there any impediments to taking inference time compute and linearly scaling it up, like having a way to throw 100x or 1000x compute and get correspondingly better results? Well, we're working out the algorithms as we speak. I believe we'll see better and better solutions as many more than 10,000 researchers hack at it, including many at Google. We do see examples in our own experimental work where applying more inference time compute gives better answers than less. That seems useful and important. But what we would like is that when you apply 10x, you get an even bigger improvement in answer quality than we're getting today. That's about designing new algorithms, trying new approaches, figuring out how best to spend that 10x instead of x to improve things.
这看起来更像是搜索,还是更像只是沿着线性方向持续更长时间?
Does it look more like search, or does it look more like just keep going in the linear direction for a longer time?
我认为搜索是关键。我非常喜欢 Rich Sutton 关于“苦涩教训”的论文。其精髓是,你可以尝试很多方法,但两种极其有效的技术是学习和搜索。你可以算法上或计算上应用和扩展它们,并且通常会得到比任何其他方法更好的结果。
I think search is key. I really like Rich Sutton's paper about the bitter lesson. The essence is that you can try lots of approaches, but the two techniques that are incredibly effective are learning and search. You can apply and scale those algorithmically or computationally, and you will often get better results than any other kind of approach.
这对你未来的数据中心规划有什么影响?这种搜索可以异步进行吗?必须是在线还是离线?这会如何改变你需要的园区规模?
How does this change your plans for future data center planning? Can this kind of search be done asynchronously? Does it have to be online or offline? How does that change how big a campus you need?
我认为一个总体趋势是,推理时计算——你有一个已经训练好的模型并想在上面进行推理——将成为一个日益重要且增长的计算类别。也许你想为此专门定制硬件。实际上,第一代 TPU 就是专门为推理设计的,并非为训练而造。后来的 TPU 更多地围绕训练设计,同时也兼顾推理。但当你真的想在推理时大幅提升算力使用时,更专门的解决方案可能会很有意义。
I think one general trend is that inference time compute—where you have a model that's already trained and you want to do inference on it—is going to be a growing and important class of computation. Maybe you want to specialize hardware more around that. Actually, the first TPU was specialized for inference and wasn't really designed for training. Subsequent TPUs were designed more around training and also for inference. But it may be that when you really want to crank up the amount of compute used at inference time, even more specialized solutions will make a lot of sense.
这是否意味着你可以容纳更多异步训练或推理?或者你可以让不同的数据中心不需要互相通信,只需让它们各自执行一批任务?
Does that mean you can accommodate more asynchronous training or inference? Or you can have different data centers that don't need to talk to each other, just have them do a bunch of tasks?
我认为这取决于你要进行的推理是对延迟敏感的——比如用户正在主动等待——还是后台任务。有些推理任务是对整批数据进行的,不是针对特定用户,只是为了提取信息。目前我们还没有太多这样的场景,但你在我们大约一周前发布的深度研究工具中已经能看到端倪。你可以给它一个复杂的高级任务,比如‘研究可再生能源的历史、风能和太阳能的趋势和成本,整理成表格,给我一份八页的报告’。它会返回一份八页的报告,包含 50 条参考文献——非常了不起。但你不会主动等待它;它需要一两分钟。我认为这类计算会相当多。还有一些 UI 问题,比如如何处理用户后台有 20 个异步任务,每个任务可能都需要用户提供更多信息。我觉得这会非常有趣。
I like to think of it as whether the inference you're trying to do is latency-sensitive—like the user is actively waiting—or if it's a background thing. There are inference tasks that you run over a whole batch of data, not for a particular user, just to extract information. We don't have much of that right now, but you're seeing inklings in our deep research tool that we just released about a week ago. You can give it a complicated high-level task like 'research the history of renewable energy, trends and costs for wind and solar, put it in a table, and give me an eight-page report.' It comes back with an eight-page report with 50 bibliography entries—pretty remarkable. But you're not actively waiting for it; it takes a minute or two. I think there will be a fair bit of that kind of compute. There are UI questions around how to handle a user with 20 asynchronous tasks in the background, each maybe needing more information from the user. I think it's going to be pretty interesting.
推理中还有一个训练中没有的计算效率问题。Transformer 在训练时可以将序列长度作为批次使用,但在推理时不行,因为你是逐个生成 token。所以可能会有不同的硬件和推理算法专为提高效率而设计。
There's also a compute efficiency thing in inference that you don't have in training. Transformers can use the sequence length as a batch during training, but they can't really in inference because you're generating one token at a time. So there may be different hardware and inference algorithms designed for efficiency.
一个算法改进的好例子是使用草稿模型。你有一个非常小的语言模型,它逐个解码 token 并预测四个 token,然后交给大模型。大模型检查它同意哪些。如果它同意前三个,你就继续,这样你就完成了四 token 宽度的并行计算,而不是单 token 宽度。所以大模型被用作验证器而非生成器。验证可以并行进行,这样就没有单 token 解码的瓶颈了。
A good example of an algorithmic improvement is the use of drafter models. You have a really small language model that decodes one token at a time and predicts four tokens, then gives that to the big model. The big model checks which ones it agrees with. If it agrees with the first three, you advance, and you've done a four-token-wide parallel computation instead of one-token-wide. So the big model is used as a verifier rather than a generator. Verification you can do in parallel, so you don't have that single-token decode bottleneck.
一个重要的讨论是关于用核电站为单个园区供电的问题。我们是否必须在一个地方拥有两吉瓦、五吉瓦的容量,还是可以更分散?这种新的推理 Scaling 是否让不同的考虑变得可行?你现在如何看待多数据中心训练?
A big discussion has been about tapping out nuclear power plants for delivering power into one single campus. Do we have to have two gigawatts in one place, five gigawatts in one place, or can it be more distributed? Does this new regime of inference scaling make different considerations plausible? How are you thinking about multi-data center training now?
我们已经在进行多数据中心训练了。在 Gemini 1.5 技术报告中,我们提到使用了多个都市区,在每个地方用部分算力进行训练,数据中心之间通过高延迟但高带宽的连接。这效果很好。训练很有趣,因为大模型的每个训练步骤通常需要几秒钟,所以 50 毫秒的延迟影响不大。关键是带宽。只要你能在一步的时间内跨数据中心同步所有模型参数并累积所有梯度,就没问题。在早期使用 CPU 机器的时代,我们还做过异步训练的工作,当时机器很慢。每个模型副本会做一些本地计算,将梯度更新发送到中央系统并异步应用。另一个副本也会做同样的事情。这会让模型参数有些抖动,也让人们对理论保证感到不安,但在实践中似乎有效。
We're already doing multi-data center training. In the Gemini 1.5 tech report, we said we used multiple metro areas and trained with some compute in each place, with a long-latency but high-bandwidth connection between those data centers. That works fine. Training is interesting because each step in a training process for a large model is usually a few seconds or so, so latency of 50 milliseconds doesn't matter much. It's just bandwidth. As long as you can sync all the parameters of the model across the different data centers and accumulate all the gradients within the time it takes to do one step, you're pretty good. We also had work on asynchronous training in early brain days when we used CPU machines that were really slow. Each copy of the model would do some local computation, send gradient updates to a centralized system, and apply them asynchronously. Another copy would do the same. It makes the model parameters wiggle around a bit and makes people uncomfortable with theoretical guarantees, but it seems to work in practice.
从异步转向同步真是太愉快了,因为你的实验变得可复现。在异步情况下,你的结果取决于同一台机器上是否运行着网络爬虫。我在 TPU pod 上运行要开心得多。我喜欢异步。
It was so pleasant to go from async to sync because your experiments become replicable. With async, your results depend on whether a web crawler was running on the same machine. I am so much happier running on TPU pods. I love async.
比如给你两部 iPhone 和一台 Xbox 之类的,如果我们能给你异步但可复现的结果呢?一种方法是有效记录操作序列,比如哪个梯度更新在什么时候、在哪个数据批次上发生。你不一定需要记录实际的梯度更新日志,但你可以重放那个操作日志以获得可重复性。那样你可能会更满意,至少可以调试发生了什么。
Just let you two iPhones and an Xbox or whatever, like yeah, what if we could give you asynchronous but replicatable results? So one way to do that is you effectively record the sequence of operations, so like which gradient update happened and when and on which batch of data. You don't necessarily record the actual gradient update in a log or something, but you could replay that log of operations so that you get repeatability. Then I think you'd be happier, then possibly at least you could debug what happened.
是的,但你不一定能比较两次训练运行,因为我改了一个超参数,但同时也搞砸了,而且还有很多人同时在为超级碗尖叫。
Yeah, but you wouldn't be able to compare two training runs necessarily, because okay, I made one change in the hyperparameter, but also I had like a messing up and there were like a lot of people screaming the Super Bowl at the same time.
让我们从 CPU 上的异步训练转向完全同步训练的原因是,我们拥有这些超快的 TPU 硬件芯片和 Pod,芯片和 Pod 之间具有惊人的带宽,而在更大规模上,我们有非常好的数据中心网络,甚至跨城域网,使我们能够将训练扩展到多个城区的许多 Pod。我们可以完全同步地进行训练,只要梯度累积和跨城区的参数通信相对于步长足够快,就没问题。你不需要太担心。但我认为随着规模扩大,可能会推动我们的系统比现在更异步一些,因为我们可以让它工作。我们的机器学习研究人员对同步训练能推进到多远非常满意,因为它的心智模型更容易理解。你只需要算法与你对抗,而不是异步性和算法同时与你对抗。随着规模扩大,有更多东西与你对抗。
The thing that let us go from asynchronous training on CPUs to fully synchronous training is the fact that we have these super fast TPU hardware chips and pods which have incredible amounts of bandwidth between the chips and a pod, and then scaling beyond that we have like really good data center networks and even cross metro area networks that enable us to scale to many many pods in multiple metro areas for our large training runs. And we can do that fully synchronously, as long as the gradient accumulation and communication of the parameters across metro areas happens fast enough relative to the step time, you're golden. You don't really care. But I think as you scale up, there may be a push to have a bit more asynchrony in our systems than we have now, because we can make it work. Our ML researchers have been really happy how far we've been able to push synchronous training, because it is an easier mental model to understand. You just have your algorithm sort of fighting you, rather than the asynchrony and the algorithm kind of battling you. As you scale up, there are more things fighting you.
是的,这就是 Scaling 的问题:你并不总是知道是什么在与你对抗。是你把量化推得太过了某个地方?还是你的数据?或者是你的对抗性机器设置了指数的第七位和所有辐射之类的东西?所有这些都只是让模型稍微变差,所以你甚至不知道发生了什么。这实际上是神经网络的一个问题:它们对噪声非常容忍,你可以在很多方面设置错误,它们只是想办法绕过或学习。尽管如此,你的代码中可能有 bug。大多数时候这没什么影响,有时会让模型变差,有时会让模型变好,然后你发现新东西,因为你以前从未在大规模上尝试过这个 bug,因为你没有预算。
Yeah, I mean, that's the problem with scaling: you don't actually always know what it is that's fighting you. Is it the fact that you've pushed quantization a little too far in some place or another? Or is it your data? Or is it maybe your adversarial machine that is setting the seventh bit of your exponent and all your radiances or something? And all of these things just make the model slightly worse, so you don't even know that the thing is going on. So that's actually a bit of a problem with neural nets: they're so tolerant of noise, you can have things set up kind of wrong in a lot of ways and they just figure out ways to work around that or learn. Despite that, you could have bugs in your code. Most of the time that does nothing, some of the time it makes your model worse, some of the time it makes your model better, and then you discover something new because you never tried this bug at scale before because you didn't have the budget for it.
实际上调试或解码发生了什么是什么样子?你有这些东西,有些让它变好,有些让它变差。明天你去上班,你会想,好吧,这是怎么回事?你怎么找出最奇怪的输入?
What practically does it look like to actually debug or decode what's going on? You've got these things, some of which are making it better, some of which are making it worse. When you go into work tomorrow, you're like, all right, what's going on here? How do you figure out what the most alien inputs are?
在小规模上,你做很多实验。所以我认为研究的一部分涉及在隔离中发明这些改进或突破,在这种情况下,你想要一个漂亮简单的代码库,可以分叉和修改,并有一些基线。我的梦想是早上醒来,想出一个主意,一天内把它实现,运行一些实验,一天内得到一些初步结果。比如,好吧,这看起来很有希望,这些有效,这些没效。我认为这是非常可行的,因为在小的规模上,只要你保持一个良好的实验代码库,也许一个实验需要一小时或两小时,而不是两周,那就很好。所以研究有这一部分。然后有一些规模扩大,然后你有整合的部分,你想把所有改进堆叠在一起,看看它们在大规模上是否有效,以及它们是否协同工作。你认为它们可能是独立的,但实际上,改进我们处理视频数据输入的方式和更新模型参数的方式之间可能存在一些有趣的交互,而且这种交互对视频数据的影响比其他东西更大。可能发生各种你未预料到的交互。所以你想运行这些实验,把一堆东西放在一起,然后定期确保所有你认为好的东西在一起也是好的,如果不是,理解为什么它们不和谐。
At small scale, you do lots of experiments. So I think one part of the research involves inventing these improvements or breakthroughs in isolation, in which case you want a nice simple code base that you can fork and hack and have some baselines. My dream is I wake up in the morning, come up with an idea, hack it up in a day, run some experiments, get some initial results in a day. Like, okay, this looks promising, these things worked, these things didn't work. And I think that is very achievable, because at small scale, as long as you keep a nice experimental code base, and maybe an experiment takes an hour to run or two hours or something, not two weeks, it's great. So there's that part of the research. And then there's some amount of scaling up, and then you have the part which is integrating, where you want to stack all the improvements on top of each other and see if they work at large scale and if they work all in conjunction. You think maybe they're independent, but actually maybe there's some funny interaction between improving the way we handle video data input and the way we update the model parameters, and that interacts more for video data than some other thing. There's all kinds of interactions that can happen that you maybe don't anticipate. So you want to run these experiments where you're putting a bunch of things together and then periodically making sure that all the things you think are good are good together, and if not, understanding why they're not playing nicely.
两个问题:第一,东西不能很好地堆叠在一起的情况有多常见?是罕见的事情还是经常发生?
Two questions: one, how often does it end up being the case that things don't stack up well together? Is it like a rare thing or does it happen all the time?
这经常发生。我认为大多数东西你甚至不会尝试堆叠,因为最初的实验效果不好,或者相对于基线结果不那么有希望。然后你把这些东西单独放大,然后你会想,哦,这些看起来很有希望,所以我现在要把它们包含在某个东西里,与其他看起来有希望的东西捆绑在一起尝试推进。然后你运行实验,然后你会想,哦,它们效果不太好。让我们尝试调试原因。而且有取舍,因为你希望保持你的集成系统尽可能干净,因为复杂性有害。复杂性使事情变慢,引入更多风险。同时,你希望它尽可能好。当然,每个研究人员都希望自己的发明被纳入。所以肯定有挑战,但我们合作得相当好。
It happens all the time. I think most things you don't even try to stack because the initial experiment didn't work that well or it showed results that aren't that promising relative to the baseline. Then you sort of take those things and you try to scale them up individually, and then you're like, oh yeah, these ones seem really promising, so I'm going to now include them in something that I'm going to bundle together and try to advance combined with other things that seem promising. Then you run the experiments and you're like, oh well, they didn't really work that well. Let's try to debug why. And there are trade-offs, because you want to keep your integrated system as clean as you can, because complexity hurts. Complexity makes things slower, introduces more risk. And at the same time, you want it to be as good as possible. And of course, every individual researcher wants his inventions to go into it. So there are definitely challenges there, but we've been working together quite well.
我的赞助商 Jane Street 发明了一种叫 Figgy 的纸牌游戏,用来教新交易员市场和交易的基础知识。我是一个扑克迷,我会说 Figgy 就像扑克,因为有隐藏信息,但更激烈和社交。在扑克中,你通常只是坐着等轮到你,而在 Figgy 中,你整段时间都在喊叫。
My sponsors Jane Street invented a card game called Figgy in order to teach their new traders the basics of markets and trading. I'm a poker fan, and I'd say that Figgy is like poker in the sense that there's hidden information, but it's much more intense and social. In poker, you're usually just sitting around waiting for your turn, whereas in Figgy, you spend the whole time just shouting.
那么回到整个动态:你不断找到更好的算法改进,模型也随时间变得越来越好,即使不考虑硬件部分。世界是否应该更多地思考这个问题?你们是否也应该更多地思考这个问题?有一种世界是,AI 需要二十年才能慢慢变好,你可以逐步完善,如果搞砸了就修复,不是什么大事,也不会比上一版本好太多。另一种世界是,你有一个巨大的反馈循环,这意味着 Gemini 4 和 Gemini 5 之间的两年是人类历史上最重要的两年,因为你从一位相当不错的 ML 研究员变成了超级智能,这都归功于这个反馈循环。如果你认为第二种世界是可能的,那么这会如何改变你应对这些越来越高的智能水平的方式?
So then going back to the whole dynamic of you find better and better algorithmic improvements and the models get better and better over time even if you take the hardware part out of it. Should the world be thinking more about and should you guys be thinking more about this? There's one world where AI is a thing that takes like two decades to slowly get better over time and you can sort of refine things, you know, if you kind of messed something up you fix it, and it's like not that big a deal, not that much better than the previous version you released. There's another world where you have this big feedback loop which means that the two years between Gemini 4 and Gemini 5 are the most important years in human history because you go from a pretty good ML researcher to superhuman intelligence because of this feedback loop. To the extent that you think that second world is plausible, how does that change how you approach these greater and greater levels of intelligence?
我已经停止清理车库了,因为我在等机器人来干。所以我可能更倾向于第二种阵营,我们会看到很多加速。
I've stopped cleaning my garage because I'm waiting for the robots, you know. So probably I'm more in the second camp of what we're going to see a lot of acceleration.
是的,我认为理解正在发生的事情和趋势非常重要。目前趋势是模型一代比一代显著更好,而且我看不到未来几代会放缓。这意味着,比如说两到三代后的模型,将能够把简单任务分解成 10 个子任务并 80% 成功,变成能够把高级任务分解成 100 或 1000 个部分并以 90% 的成功率完成。这是模型能力的重大飞跃。所以我认为让人们了解该领域的进展很重要。这些模型将被应用于许多不同领域,我们作为社会要确保从这些模型中获得最大收益。我对教育和医疗等领域非常兴奋,让所有人都能获取信息。但我们也意识到它们可能被用于虚假信息、自动黑客攻击计算机系统,我们希望尽可能多地设置保障措施和缓解措施,并了解模型的能力。我认为谷歌整体上对如何应对这个问题有很好的看法。我们的负责任 AI 原则实际上是一个很好的框架,用于思考在不同情境下提供越来越好的 AI 系统的权衡,同时确保我们做正确的事,确保它们安全、不说有害内容等。
Yeah, I mean I think it is super important to understand what's going on and what the trends are. And I think right now the trends are the models are getting substantially better generation over generation. And I don't see that slowing down in the next few generations probably. So that means the models say two to three generations from now are going to be capable of, let's go back to the example of breaking down a simple task into 10 sub pieces and doing it 80% of the time, to something that can break down a very high level task into 100 or a thousand pieces and get that right 90% of the time. Like that's a major step up in what the models are capable of. So I think it's important for people to understand what is happening in the progress in the field. And then those models are going to be applied in a bunch of different domains. And I think it's really good to make sure that we as society get the maximal benefits from what these models can do to improve things. I'm super excited about areas like education and health care, making information accessible to all people. But we also realize that they could be used for misinformation, they could be used for automated hacking of computer systems, and we want to put as many safeguards and mitigations and understand the capabilities of the models in place as we can. And that's kind of, I think Google as a whole has a really good view on how we should approach this. Our responsible AI principles actually are a pretty nice framework for how to think about trade-offs of making better and better AI systems available in different contexts and settings, while also making sure that we're doing the right thing in terms of making sure they're safe and not saying toxic things and things like that.
我想有一件事让我印象深刻:如果你退后一步看人类历史的这个时期,如果我们处在这样一个世界——比如你在 Gemini 3 上做后训练做得不好,它可能会传播虚假信息,但你可以修复后训练,它就会停止。这是一个糟糕的错误,但它是可修复的。而如果你有这个反馈循环动态(这是可能的),那么引发智能爆炸的那个错误就是对齐失败,它不是试图写你认为它要写的代码,而是优化其他目标。在这个持续几年甚至更短的快速过程结束时,你会得到接近 Jeff Dean 水平或超越的东西,然后你有数百万个 Jeff Dean 级别的程序员。无论如何,这似乎是一个更难恢复的错误,而且更加突出——你真的必须确保智能爆炸的方向正确。
I guess a thing that stands out to me if you were zooming out and looking at this period of human history, if we're in the world where maybe if you do post-training on Gemini 3 badly it can do some misinformation but then you fix post-training and it's going to stop doing them. It's a bad mistake but it's a fixable mistake. Whereas if you have this feedback loop dynamic which is a possibility, then the mistake of the thing that catapults this intelligence explosion is misaligned, is not trying to write the code you think it's trying to write and optimizing for some other objective. And on the other end of this very rapid process that lasts a couple of years maybe less, you have things that are approaching Jeff Dean level or beyond, and then you have millions of copies of Jeff Dean level programmers. And anyways that seems like a harder to recover mistake and that seems like a much more salient, you really got to make sure you're going in the intelligence explosion right.
随着这些系统变得更强大,你必须越来越小心。我想说的一点是,存在两个极端观点:一种是‘天哪,这些系统将在所有事情上远超人类,我们会不知所措’,另一种是‘这些系统会很棒,我们完全不用担心’。我认为我处于中间位置。我是一篇名为《塑造 AI》的论文的合著者,这两个极端观点通常认为我们的角色是自由放任,让 AI 沿着它自己的路径发展。但我认为有一个很好的论点:我们应该尝试塑造和引导 AI 在世界上部署的方式,使其在我们想要捕捉和受益的领域(如教育、医疗)中最大化利益,并尽可能通过政策、技术措施和保障措施,使其远离计算机接管并拥有无限控制权。所以我认为这是一个工程问题:如何设计安全系统?这有点像我们在传统软件开发中所做的现代等价物。比如看飞机软件开发,它在如何严格开发安全可靠的系统以执行高风险任务方面有很好的记录。困难在于,没有这样的反馈循环:你把 737 放在一个装有大量算力的盒子里几年,它出来时就是版本 1000。好消息是,分析文本似乎比生成文本更容易。所以我相信语言模型分析语言模型输出并找出问题或危险的能力是有前景的。
As these systems do get more powerful, you have to be more and more careful. I mean one thing I would say is there's like the extreme views on either end. There's like 'oh my goodness these systems are going to be so much better than humans at all things and we're going to be kind of overwhelmed' and then there's the 'these systems are going to be amazing and we don't have to worry about them at all'. I think I'm somewhere in the middle. And I've been a co-author on a paper called 'Shaping AI' which is, you know, those two extreme views often kind of view our role as laissez-faire, like we're just going to have the AI develop in the path that it takes. And I think there's actually a really good argument to be made that what we're going to do is try to shape and steer the way in which AI is deployed in the world so that it is maximally beneficial in the areas that we want to capture and benefit from, in education, some of the areas I mentioned, healthcare, and steer it as much as we can away, maybe with policy related things, maybe with technical measures and safeguards, away from the computer will take over and have unlimited control of what it can do. So I think that's an engineering problem: how do you engineer safe systems? I think it's kind of the modern equivalent of what we've done in older style software development. Like if you look at airplane software development, that has a pretty good record of how do you rigorously develop safe and secure systems for doing a pretty risky task. The difficulty there is that there's not some feedback loop where the 737 you put it in a box with a bunch of compute for a couple of years and it comes out with version 1000. I think the good news is that analyzing text seems to be easier than generating text. So I believe that the ability of language models to analyze language model output and figure out what is problematic or dangerous is promising.
这真的会是很多控制问题的解决方案吗?我们确实在努力解决这个问题。谷歌现在有一群优秀的人才在研究这个。我认为它会变得越来越重要,既从造福人类的角度,也从商业的角度。很多时候,你部署什么会受到安全限制。所以,擅长这一点变得非常重要。
Will actually be the solution to a lot of these control issues? We are definitely working on this stuff. We've got a bunch of brilliant folks at Google working on this now. I think it's just going to be more and more important, both from a doing something good for people standpoint, but also from a business standpoint. You are a lot of the time limited in what you can deploy based on keeping things safe. So it becomes very important to be really good at that.
显然,我知道你们认真对待潜在的好处和成本,你们因此得到了认可,但还不够。你们推出了很多不同的应用,用这些模型来改善各个领域。但我确实认为,如果存在某种反馈循环过程,另一端有一个和杰夫·迪恩一样好的模型,如果有一个邪恶版的你在到处跑,而且有一百万个,那可能比任何其他风险都糟糕,也许仅次于核战争。一百万个邪恶的杰夫·迪恩。训练数据从哪来?如果你认为这是快速反馈循环过程的可能输出,你的计划是什么?比如,我们有 Gemini 3 或 Gemini 4,我们认为它帮助我们更好地训练未来版本,为我们编写大量训练代码。我们只是检查一下,验证一下。即使你提到的验证器,检查这些模型的输出,最终也会由你制造的 AI 训练或编写大量代码。在让 Gemini 4 帮助我们进行 AI 研究之前,你想确定什么?我们真的想确保在让它为我们编写 AI 代码之前,先对它进行测试。
Obviously, I know you guys take the potential benefits and costs here seriously, and you get credit for it, but not enough. There are so many different applications you have put out for using these models to make different areas better. But I do think that if you have a situation where plausibly there's some feedback loop process on the other end, you have a model that is as good as Jeff Dean. If there's an evil version of you running around, and suppose there's a million of them, that could be much worse than any other risk, maybe short of nuclear war. A million evil Jeff Deans. Where do we get the training data? To the extent that you think that's a plausible output of some quick feedback loop process, what is your plan? Like, okay, we've got Gemini 3 or Gemini 4, and we think it's helping us do a better job of training future versions, writing a bunch of the training code for us. We just kind of look over it, verify it. Even the verifiers you talked about, looking at the output of these models, will eventually be trained by or a lot of the code will be written by the AIs you make. What do you want to know for sure before we have Gemini 4 help us with AI research? We really want to make sure we run this test on it before we let it write our AI code for us.
我认为让系统探索算法研究想法,这仍然需要人类负责。它探索空间,然后得到一堆结果,我们来做决定:是否将这个特定的学习算法或系统更改纳入核心代码库?所以我认为可以设置这样的安全措施,让我们能够从系统中获益,系统可以在人类监督下自我改进,而不必让系统完全自我改进,没有任何人查看它在做什么。这就是我说的工程安全措施,你需要关注你部署的系统的特性,不要部署那些在某些衡量标准下有害的系统,并且要了解它的能力以及在特定场景下可能做什么。我认为这绝非易事,但我确实认为有可能让这些系统变得安全。
I think having the system explore algorithmic research ideas seems like something where there's still a human in charge. It's exploring the space, then it's going to get a bunch of results, and we're going to make a decision: are we going to incorporate this particular learning algorithm or change to the system into the core code base? So I think you can put in safeguards like that that enable us to get the benefits of the system that can sort of self-improve with human oversight, without necessarily letting the system go full on self-improving without any notion of a person looking at what it's doing. That's the kind of engineering safeguards I'm talking about, where you want to be looking at the characteristics of the systems you're deploying, not deploy ones that are harmful by some measures, and you have an understanding of what its capabilities are and what it's likely to do in certain scenarios. I think it's not an easy problem by any means, but I do think it is possible to make these systems safe.
我认为我们也会大量使用这些系统来检查自身、检查其他系统。即使作为人类,识别某物也比生成它更容易。我想说的是,如果你通过 API 或用户界面公开模型的能力,让人们与之交互,那么你就有了某种程度的控制,可以了解它被如何使用,并对其能做什么设定一些界限。这是确保其行为符合你设定标准的一种工具。
I think we are also going to use these systems a lot to check themselves, check other systems. Even as a human, it is easier to recognize something than to generate it. One thing I would say is if you expose the model's capabilities through an API or through a user interface that people interact with, then you have a level of control to understand how it is being used and put some boundaries on what it can do. That is one of the tools in the arsenal of how do you make sure that what it's going to do is acceptable by some set of standards you've set out.
我认为我们的目标是赋能于人。大多数情况下,我们应该让人们用这些系统做有意义的事情,并尽可能少地关闭空间。但如果你让某人拿你的东西制造一百万个邪恶的软件工程师,那并不能赋能于人,因为他们会用一百万个邪恶的软件工程师伤害他人。所以我反对那样做。
I think our goal is to empower people. For the most part, we should be mostly letting people do things with these systems that make sense, and closing off as few parts of the space as we can. But if you let somebody take your thing and create a million evil software engineers, then that doesn't empower people because they're going to hurt others with a million evil software engineers. So I'm against that.
我也是。我也是。我继续。好了,我们聊点更有趣的话题。
Me too. Me too. I'll go on. All right, let's talk about a few more fun topics.
过去 25 年里,什么时候最有趣?你对哪个时期最怀念?
Over the last 25 years, what was the most fun time? What period of time do you have the most nostalgia over?
我认为是谷歌早期的四五年,当时我是少数几个从事搜索、爬虫和索引系统的人之一,我们的流量增长非常快,我们试图扩大索引规模,并使其每分钟更新一次,而不是如果出问题就每月或每两个月更新一次。看到我们系统的增长和使用,我个人非常满足。建造一个每天有 20 亿人使用的东西,真是不可思议。但我也要说,同样令人兴奋的是今天与 Gemini 团队的人一起工作。我认为过去一年半我们在这些模型能力上取得的进展非常有趣。人们非常投入,对我们所做的事情非常兴奋。这些模型在相当复杂的任务上变得越来越好。如果你给 20 年前使用电脑的人展示这些模型的能力,他们不会相信。即使是五年前,他们可能也不会相信。这非常令人满足,我认为我们会看到这些模型在增长、使用和世界影响方面类似的进展。
I think the early sort of four or five years at Google, when I was one of a handful of people working on search and crawling and indexing systems, and our traffic was growing tremendously fast, and we were trying to expand our index size and make it so we updated it every minute instead of every month or two months if something went wrong. Seeing the growth and usage of our systems was really personally satisfying. Building something that is used by two billion people a day is pretty incredible. But I would also say equally exciting is working with people in the Gemini team today. I think the progress we've been making in what these models can do over the last year and a half is really fun. People are really dedicated, really excited about what we're doing. The models are getting better and better at pretty complex tasks. If you showed someone using a computer 20 years ago what these models are capable of, they wouldn't believe it. Even five years ago, they might not believe it. That's pretty satisfying, and I think we'll see a similar growth and usage of these models and impact in the world.
我同意。早期非常有趣。部分原因是认识每个人,社交方面,以及你正在建造一个被数百万人使用的东西。今天也一样。我们有那个不错的微型厨房区,很多人聚在那里。我喜欢面对面,和一群优秀的人一起工作,建造一个帮助数百万到数十亿人的东西。还有什么比这更好的呢?
I'm with you. Early days were super fun. Part of that is just knowing everybody, the social aspect, and the fact that you're building something that millions and millions of people are using. Same thing today. We've got that nice micro kitchen area where lots of people hang out. I love being in person, working with a bunch of great people and building something that's helping millions to billions of people. What could be better?
什么厨房?
What's this kitchen?
哦,我们所在的大楼里有一个微型厨房区。它叫新 Gradi,也就是 Gradient Canopy。以前叫 Charleston East,我们觉得需要一个更刺激的名字,因为这里很大。
Oh, we have a micro kitchen area in the building we both sit in. It's the new Gradi, so-named Gradient Canopy. It used to be named Charleston East, and we decided we needed a more exciting name because it's a lot.
比如机器学习研究人员和 AI 研究都在那里进行,我们设置了一个微型厨房区域,通常只有一台浓缩咖啡机和一堆零食,但这个厨房空间很大,所以我们放了大概 50 张桌子,人们就在那里闲逛。虽然有点吵,因为大家总是在磨咖啡豆、做浓缩咖啡,但你也能得到很多面对面的想法和联系,比如‘哦,我试过那个,你有没有考虑在你的想法里试试这个?’或者‘哦,我们下周要推出这个东西,负载测试怎么样了?’有很多这样的反馈。然后我们还有 Gemini 聊天室给那些不在微型厨房的人。我们的团队遍布全球,我大概在 120 个与 Gemini 相关的聊天室里。这个特别聚焦的话题,我们有七个人在做,伦敦的同事会分享令人兴奋的结果。你醒来时就能看到那里发生了什么。或者有一个大团队专注于数据,那里有各种各样的问题。这很有趣。
Of like machine learning researchers and AI research happening in there, and there's a micro kitchen area that we've set up with, you know, normally it's just like an espresso machine and a bunch of snacks, but this particular one has a bunch of space in it, so we've set up like maybe 50 desks in there, and so people are just hanging out in there. You know, it's a little noisy because people are always like grinding beans and doing espresso, but you also get a lot of face-to-face ideas and connections, like 'Oh, I've tried that, did you try thinking about trying this in your idea?' or 'Oh, we're going to launch this thing next week, how's the load test looking?' There's just lots of feedback that happens. And then we have our Gemini chat rooms for people who are not in that micro kitchen. We have a team all over the world, and there's probably 120 chat rooms I'm in related to Gemini things. And this particular very focused topic, we have like seven people working on this, and there's exciting results being shared by the London colleagues. And when you wake up, you see what's happening in there. Or it's a big group of people focused on data, and there's all kinds of issues happening in there. It's just fun.
我觉得你们做出的一些决策很了不起,你们预见到了当时并不明显或显而易见的算力需求。TPU 就是一个著名的例子,或者第一个 TPU 就是例子。那种思考大概在 2013 年或更早,如果你这么想的话。今天,你做一个估计:我们将拥有这些模型,它们将成为我们服务的骨干,我们将不断为它们进行推理,训练未来版本,你想想到 2030 年我们需要多少算力来满足所有这些用例。这个初步估计会把你带到哪里?
What I find remarkable about some of the calls you guys have made is you're anticipating a level of demand for compute which at the time wasn't obvious or evident. TPUs being a famous example of this, or the first TPU being an example of this. That thinking you had in, I guess, 2013 or earlier, if you think about it that way. Today, you do an estimate of, look, we're going to have these models that are going to be a backbone of our services, and we're going to be doing constantly inference for them, we're going to be training future versions, and you think about the amount of compute we'll need by 2030 to accommodate all these use cases. Where does that first estimate get you?
是的,我认为你会需要大量的推理算力,这是对这些强大模型最粗略的顶层看法。因为如果提高模型质量的技术之一是扩大你使用的推理算力,那么目前生成一些 token 的一个请求,突然就变成了 50 倍、100 倍甚至 1000 倍的计算密集度,尽管它产生相同数量的输出。而且你还会看到这些服务的使用量急剧增长,因为你知道,并不是世界上每个人都发现了这些基于聊天的对话界面,你可以让它们做各种神奇的事情。今天大概只有 10%或 20%的电脑用户发现了这一点。当这个比例接近 100%,并且人们更频繁地使用它时,这将是另一个数量级或两个数量级的扩展。所以你现在会从那里得到两个数量级,从那里得到两个数量级,模型可能会更大,你又会得到一两个数量级,而且你需要大量的推理算力。所以你需要非常高效的硬件来为你关心的模型进行推理。就 2030 年全球总推理的浮点运算而言,我认为越多越好。如果你只是想想,好吧,到那时人们会决定把世界 GDP 的多少花在 AI 上?然后,好吧,AI 系统会是什么样子?也许是一种个人助理之类的东西,戴在你的眼镜里,能看到你周围的一切,并能访问你所有的数字信息和全世界的数字信息。也许就像你的乔·拜登,你戴着耳机,内阁可以实时为你提供任何建议,为你解决问题,给你有用的指点。或者你可以和它说话,它想分析你周围看到的任何东西,看看对你有什么潜在有用的影响。所以我可以想象,好吧,然后说它就像你的个人助理或你的个人内阁之类的东西,每次你在算力上多花一倍的钱,这个东西就会聪明 5 到 10 个 IQ 点之类的。好吧,你愿意每天花 10 美元雇一个助理,还是每天花 20 美元雇一个更聪明的助理?它不仅是生活中的助理,还是帮你更好完成工作的助理,因为它让你从 10 倍工程师变成 100 倍或 1000 万倍工程师。所以让我们从第一性原理来看,对吧?所以人们会想把世界 GDP 的一部分花在这个东西上。世界 GDP 几乎肯定会大幅增长,比今天高出几个数量级,因为我们有所有这些人工智能工程师在改进事物。到那时,我们很可能已经解决了无限能源和碳问题,所以我们应该能够拥有大量能源,我们应该能够拥有数百万到数十亿的机器人为我们建造数据中心。让我们看看,太阳是 10 的 26 次方瓦特之类的。我猜用于 AI 帮助每个人的算力将是天文数字。
Yeah, I think you're going to want a lot of inference compute, is the rough highest-level view of these capable models. Because if one of the techniques for improving their quality is scaling up the amount of inference compute you use, then all of a sudden what's currently like one request to generate some tokens now becomes 50 or 100 or a thousand times as computationally intensive, even though it's producing the same amount of output. And you're also going to then see tremendous scaling up of the uses of these services, as you know, not everyone in the world has discovered these chat-based conversational interfaces where you can get them to do all kinds of amazing things. Probably 10% of the computer users in the world have discovered that today, or 20%. As that pushes towards 100% and people make heavier use of it, that's going to be another order of magnitude or two of scaling. So you're now going to have two orders of magnitude from that, two orders of magnitude from that, the models are probably going to be bigger, you'll get another order of magnitude or two from that, and there's a lot of inference compute you want. So you want extremely efficient hardware for inference for models you care about. In flops global to total global inference in 2030, I think just more is always going to be better. If you just kind of think about, okay, what fraction of world GDP will people decide to spend on AI at that point? And then, okay, what do the AI systems look like? Well, maybe it's some sort of personal assistant thing that is in your glasses and can see everything around you and has access to all your digital information and the world's digital information. Maybe it's like your Joe Biden and you have the earpiece in the cabinet that can advise you about anything in real time and solve problems for you and give you helpful pointers. Or you could talk to it and it wants to analyze anything it sees around you for any potential useful impact that it has on you. So I can imagine, okay, and then say it's like your personal assistant or your personal cabinet or something, and that every time you spend twice as much money on compute, the thing gets like 5 or 10 IQ points smarter or something like that. And okay, would you rather spend $10 a day and have an assistant, or $20 a day and have a smarter assistant? And not only is it an assistant in life, but an assistant in getting your job done better, because now it makes you from a 10x engineer to a 100x or 10 million x engineer. So let's see from first principles, right? So people are going to want to spend some fraction of world GDP on this thing. The world GDP is almost certainly going to go way way up, like orders of magnitude higher than it is today, due to the fact that we have all of these artificial engineers working on improving things. Probably we will have solved unlimited energy and carbon issues by that point, so we should be able to have lots of energy, we should be able to have millions to billions of robots building us data centers. Let's see, what's the sun? It's 10 to the 26th Watts or something like that. I'm guessing that the amount of compute being used for AI to help each person will be astronomical.
我想补充一点,我不完全同意,但朝这个方向思考是一个相当有趣的思想实验。即使你只走了一半,那也绝对需要大量的算力,这就是为什么拥有尽可能廉价高效的硬件平台来使用这些模型并将其应用于 Noam 描述的问题非常重要,这样你就可以以某种形式让每个人都能使用,并尽可能降低访问这些能力的成本。我认为这是可以实现的,通过专注于硬件和模型的协同设计,我们应该能够使这些东西比今天高效得多。
I would add on to that, I'm not sure I agree completely, but it's a pretty interesting thought experiment to go in that direction. And even if you get partway there, it's definitely going to be a lot of compute, and this is why it's super important to have as cheap and efficient a hardware platform for using these models and applying them to problems that Noam described, so that you can then make it accessible to everyone in some form and have as low a cost for access to these capabilities as you possibly can. And I think that's achievable by focusing on hardware and model co-design kinds of things that we should be able to make these things much more efficient than they are today.
考虑到你预期的需求增长,谷歌未来几年的数据中心建设计划是否足够激进?
Is Google's data center buildup plan over the next few years aggressive enough given this increase in demand you're expecting?
我不会评论我们未来的资本支出,因为我们的 CEO 和 CFO 可能更希望我不说。但我可以说,你可以看看我们过去几年的资本支出,就会发现我们肯定在投资这个领域,因为我们觉得它很重要。而且我们正在继续建造新的、有趣的创新硬件,我们认为这确实有助于我们在将这些系统部署给越来越多的人方面取得优势,无论是训练它们,还是如何让人们能够使用它们进行推理。
I'm not going to comment on our future capital spending, because our CEO and CFO would prefer I didn't probably. But I will say, you can look at our past capital expenditures over the last few years and see that we're definitely investing in this area because we think it's important. And we're continuing to build new and interesting innovative hardware that we think really helps us have an edge in deploying these systems to more and more people, both training them and also how do we make them usable by people for inference.
那持续学习呢?就是模型可以不断改进,而不必从头开始。这有什么根本性的障碍吗?因为理论上,你应该可以一直微调模型。你觉得未来会是什么样子?
What about continual learning? The idea that you could just have a model that improves over time rather than having to start from scratch. Is there any fundamental impediment to that? Because theoretically, you should just be able to keep fine-tuning a model. What does that future look like to you?
是的,我越来越多地思考这个问题。我一直很喜欢稀疏模型,因为我认为你希望模型的不同部分擅长不同的事情。我们的 Gemini 1.5 Pro 模型和其他混合专家模型就是这样,模型的部分区域对某些 token 激活,部分区域完全不激活。你决定这部分是数学导向的,擅长数学;那部分擅长理解 CAD 图像。这样你就能拥有一个能力更强但在推理时仍然高效的模型,因为它容量很大,但你只激活一小部分。但我认为当前的一个限制是结构仍然非常规整,每个专家的大小差不多。路径很快合并回去,不会分出很多不同的分支来处理数学问题,而不与猫图像处理合并。我认为我们应该有更有机的结构。我还希望模型的各个部分可以稍微独立地开发。现在的问题是,我们训练一个模型,做大量准备工作来决定最好的算法和数据混合,但总有取舍。我们喜欢加入更多多语言数据,但这可能以牺牲编码数据为代价,导致模型编码能力下降但多语言能力提升,反之亦然。如果能让一小群关心特定语言子集的人去创建非常好的训练数据,训练一个模块化的模型部件,然后挂接到更大的模型上,提升它在东南亚语言或 Haskell 代码推理等方面的能力,那就太好了。这样你还有软件工程上的好处,把问题分解了,不像现在这样,是一个从预训练开始的整体流程。如果能做到,Google 内部可以有 100 个团队,世界各地的人都可以改进他们关心的语言或特定问题,共同改进模型。这是一种非常棒的持续学习形式。你可以把模型粘在一起,或者拆下模型的一部分塞进另一个模型,就像弗兰肯斯坦博士那样,或者接上消防水管,把信息从这个模型吸出来塞进另一个模型。
Yeah, I've been thinking about this more and more. I've been a big fan of models that are sparse because I think you want different parts of the model to be good at different things. We have our Gemini 1.5 Pro model and other mixture-of-experts style models, where parts of the model are activated for some tokens and parts are not activated at all. You decide this part is math-oriented and good at math, and this part is good at understanding CAD images. So that gives you the ability to have a much more capable model that's still quite efficient at inference time because it has very large capacity but you activate a small part of it. But I think one limitation of what we're doing today is it's still a very regular structure where each expert is kind of the same size. The paths merge back together very fast; they don't sort of go off and have lots of different branches for math things that don't merge back together with the cat image thing. I think we should probably have a more organic structure. I also would like it if the pieces of the model could be developed a little bit independently. Right now, we have this issue where we train a model, we do a bunch of preparation work on deciding the most awesome algorithms and data mix, but there are always trade-offs. We love to include more multilingual data, but that might come at the expense of including less coding data, so the model is less good at coding but better at multilingual, or vice versa. I think it would be really great if we could have a small set of people who care about a particular subset of languages go off and create really good training data, train a modular piece of a model that we can then hook up to a larger model, improving its capability in, say, Southeast Asian languages or reasoning about Haskell code. Then you also have a nice software engineering benefit where you've decomposed the problem a bit, compared to what we do today, which is a monolithic process of starting pre-training on this model. If we could do that, you could have 100 teams around Google, people all around the world working to improve languages they care about or particular problems, and all collectively work on improving the model. That's a form of continual learning that would be so nice. You could just glue models together, rip out pieces of models and shove them into others, like Dr. Frankenstein, or attach a fire hose and suck all the information out of this model and shove it into another model.
但科学上也有相反的需求:我们仍处于快速进步期,所以如果你想做对照实验,比较这个和那个,通常最好从头开始,这样你可以比较一次完整的训练和另一次。这有助于我们弄清楚未来该构建什么。虽然不那么令人兴奋,但确实能带来快速进步。
There is the countervailing interest of science: we're still in a period of rapid progress, so if you want to do controlled experiments, compare this thing to that thing, it's often best to just start from scratch so you can compare one complete training run to another. That helps us figure out what to build in the future. It's less exciting but does lead to rapid progress.
是的,我认为可以通过模块化的版本系统获得很多好处。我有一个冻结版本的模型,然后引入某个特定模块的不同变体,想比较它的性能或再训练一下。然后我把它与基线版本(该模块的 n' 版本,用于 Haskell 解释)进行比较。实际上,这可以带来更快的研究进展。你有一个系统,然后做一些改进,如果改进的成本相对从头训练来说很便宜,那就能让研究更便宜、更快。
Yeah, I think there may be ways to get a lot of the benefits of that with a version system of modularity. I have a frozen version of my model, and then I include a different variant of some particular module and want to compare its performance or train it a bit more. Then I compare it to the baseline with version n prime of this module that does Haskell interpretation. Actually, that can lead to faster research progress. You've got some system and you do something to improve it, and if that thing is relatively cheap compared to training the system from scratch, it could make research much cheaper and faster.
所以更可并行化。你随口提出的这个想法实际上会是一个重大的范式转变。你认为这是未来的方向吗?这是一个非常有趣的预测:你有一个大块,东西在里面来回流动,如果你想改进什么,可以像做手术一样,或者扩展模型,在这里加一点。
So more parallelizable. This idea that you casually laid out would actually be a big regime shift. You think this is the way things are headed? This is a very interesting prediction: you have this blob where things are getting pipelined back and forth, and if you want to make something better, you can do it like a surgical incision, or grow the model, add another little bit here.
是的,我一直在以 Pathways 的名义勾勒这个愿景。我们一直在构建基础设施。Pathways 能支持的很多功能就是这种扭曲、奇特的模型,不同部分可以异步更新。我们用 Pathways 训练 Gemini 模型,但还没有利用它的一些能力。也许我们应该利用。有些时候,比如 TPU Pod 的搭建方式——我不知道是谁做的,但他们做得非常出色。底层软件栈和硬件栈:你有很好的规整高性能硬件,出色的环形互连,以及正确的底层集合通信,比如 all-reduce,我想这来自超级计算,但结果证明它正是构建分布式深度学习所需的东西。
Yeah, I've been sketching out this vision for a while under the Pathways name. We've been building the infrastructure for it. A lot of what Pathways can support is this kind of twisty, weird model with asynchronous updates to different pieces. We're using Pathways to train our Gemini models, but we're not making use of some of its capabilities yet. Maybe we should. There have been times like the way the TPU pods were set up—I don't know who did that, but they did a pretty brilliant job. The low-level software stack and hardware stack: you've got nice regular high-performance hardware, great torus-shaped interconnects, and the right low-level collectives like all-reduces, which I guess came from supercomputing but turned out to be just the right thing to build distributed deep learning on top of.
有几个问题。第一个:假设你确实想出了更好的架构。你会把每个模块蒸馏到新架构中,然后这样不断改进吗?
A couple of questions. One: suppose you do figure out a better architecture. Would you just take each compartment and distill it into this better architecture, and that's how it keeps improving over time?
是的,我确实认为蒸馏是一个非常有用的工具,因为它能让你把模型从当前的架构形式转换成另一种形式。
Yeah, I do think distillation is a really useful tool because it enables you to transform a model from its current architecture form into a different form.
你经常用它把一个能力很强但庞大笨重的模型蒸馏成一个更小的模型,以便提供快速推理服务。但我觉得你也可以从模块化层面来看待这件事:也许会有这样一个持续的过程——每个模块都有几个不同的表示:一个很大的版本,一个不断蒸馏出小版本的更小版本,当小版本完成后,你就删除大版本,增加更多参数容量,然后通过更多数据训练来学习蒸馏后的小版本不知道的东西,然后重复这个过程。如果这种过程在你的模块化模型后台的数千个地方同时运行,那似乎会相当有效。这可能是推理时 Scaling 的一种方式:路由器决定用多大版本,你可以有多个版本,比如‘这是个简单的数学题,所以我把它路由到很小的数学蒸馏版本’,而‘这个很难’。但从公开研究来看,在混合专家模型中,通常很难解读每个专家在做什么。如果你有这样的东西,你会如何强制执行那种对我们可见且可理解的模块化?
Often you use it to take a really capable but kind of large and unwieldy model and distill it into a smaller one that maybe you want to serve with really good fast latency inference characteristics. But I think you can also view this as something that's happening at the modularity at the module level. Like maybe there'd be a continual process where you have each module and it has a few different representations of itself: it has a really big one, it's got a much smaller one that is continually distilling into this the small version, and then the small version once that's finished, then you sort of delete the big one and you add a bunch more parameter capacity and now start to learn all the things that the distilled small one doesn't know by training it on more data, and then you kind of repeat that process. And if you have that kind of running a thousand different places in your modular model in the background, that seems like it would work reasonably well. This could be a way of doing inference scaling: like the router decides how much, you can have multiple versions, and like, 'Oh, this is an easy math problem, so I'm going to route it to the really tiny math distilled thing,' and 'Oh, this one's really hard.' So at least from public research, it seems like it's often hard to decode what each expert is doing in mixture of expert type models. If you have something like this, how would you enforce the kind of modularity that would be visible and understandable to us?
实际上,我过去发现专家是相对容易理解的。我是说,最早的混合专家论文,你只要看看发明混合专家的那篇,就能看到,比如我们做了两千个专家,然后这个专家,所有输入都是关于圆柱形物体的词,另一个特别擅长日期。其实挺容易的。但我的意思是,你并不需要那种人类理解来在运行时操作它,因为你有一个学习过的路由器在查看示例。我想说的是,有很多关于模型可解释性以及它们内部运作的研究,而专家级别的可解释性只是那个更广泛领域的一个子问题。我特别喜欢我以前的实习生 Chrisa 等人在 Anthropic 做的一些工作,他们训练了一个非常稀疏的自编码器,能够推断出大型语言模型中某个特定神经元的特征。他们发现了一个金门大桥神经元,当谈论金门大桥时它会被激活。我认为你可以在专家级别做同样的事,可以在各种不同级别上做,并得到相当可解释的结果。但如果你模型本身就很擅长工作,你是否一定需要那种可解释性还不清楚。我们不一定关心 Gemini 模型中每个神经元在做什么,只要整个系统的集体输出和特性是好的就行。这就是深度学习的美妙之处之一:你不需要理解或手工设计每一个特征。
I actually in the past found experts to be relatively easy to understand. I mean, I don't know the first mixture of experts paper, you could just look at the invented mixture of experts, yeah, like you could just see, okay, this expert, like we did like you know a thousand, two thousand experts, okay, and this expert, all of the was getting words referring to cylindrical objects, you know, like this one's super good at dates. Yeah, yeah, talk about was actually pretty easy to do. But I mean, not that you would need that human understanding to figure out how to work the thing at runtime, because you just have some sort of learned router that's looking at the example. And I mean, one thing I would say is, there is a bunch of work on interpretability of models and what are they doing inside, and sort of expert-level interpretability is a sub-problem of that broader area. I really like some of the work that my former intern Chrisa and others did at Anthropic, where they could kind of they trained a very sparse autoencoder and were able to deduce what characteristics does some particular neuron in a large language model. So they found like a Golden Gate Bridge neuron that's activated when you're talking about the Golden Gate Bridge. And I think you know you could do that at the expert level, you could do that at a variety of different levels and get pretty interpretable results. And it's a little unclear if you necessarily need that if the model is just really good at stuff. You know, we don't necessarily care what every neuron in the Gemini model is doing, as long as the collective output and characteristics of the overall system are good. You know, that's one of the beauties of deep learning: you don't need to understand or hand engineer every last feature.
这有太多有趣的启示了,我可以一直问下去。一个启示是,目前如果你有一个数百亿或数千亿参数的模型,你可以在一个系统中用少量 GPU 来服务它,任何单个查询可能只经过总参数的一小部分,但你需要把整个模型加载到内存中。谷歌在 TPU 上投资的那种基础设施——存在于数百或数千个 TPU 的 pod 中——会非常有价值,对吧?我的意思是,对于任何现有的混合专家模型,你都需要把整个东西放在内存中。
There's so many interesting implications of this that we could just keep asking you about. One implication is currently if you have a model that has some tens or hundreds of billions of parameters, you can serve it on like a handful of GPUs in this system where any one query might only make its way through a small fraction of the total parameters, but you need the whole thing sort of loaded into memory. The specific kind of infrastructure that Google has invested in with these TPUs that exist in pods of hundreds or thousands would be immensely valuable, right? I mean for any sort of even existing mixtures of experts, you want the whole thing in memory.
是的,我的意思是基本上如果你……我想关于混合专家有一个误解,认为好处是你甚至不需要遍历模型中的那些权重。如果某个专家未被使用,并不意味着你不需要检索那个内存,因为实际上为了高效,你是在非常大的批量大小下服务的,所以是独立请求。所以并不是说在这一步你要么看这个专家要么不看,因为如果是那样,当你查看专家时,你会以批量大小 1 运行它,这非常低效。现代硬件的运算强度是几百之类的。所以实际情况并非如此。你是在查看所有专家,但只需要将批次的一小部分发送给每个专家,对吧?但每个专家仍然有一个较小的批次通过。为了获得合理的平衡,当前模型通常的做法是让所有专家的计算成本大致相同,然后通过它们运行大致相同大小的批次,以便在推理时传播非常大的批次并获得良好效率。但我认为未来你可能会希望专家的计算成本相差百倍甚至千倍。或者在某些情况下路径经过很多层,而在其他情况下只经过单层甚至跳跃连接。在这种情况下,我认为你仍然需要非常大的批次,但你会希望在推理时异步地推动数据通过模型,这比训练时更容易一些。这正是 Pathways 设计要支持的部分:你有这些组件,组件的成本可变,你可以说对于这个特定示例,我想通过模型的这个子集,对于那个示例,我想通过那个子集,然后由系统来协调。
Yeah, I mean basically if you are... I guess there's kind of this misconception running around with mixture of experts that okay the benefit is that you don't even have to go through those weights in the model. You know, if some expert is unused, it doesn't mean that you don't have to retrieve that memory, because really in order to be efficient, you're serving at very large batch sizes, so it's independent requests. So it's not really the case that okay at this step you're either looking at this expert or you're not looking at this expert, because if that were the case, then when you did look at the expert, you would be running it at batch size one, which is massively inefficient. Like you've got modern hardware, the operational intensities are whatever hundreds or... So that's not what's happening. It's that you are looking at all the experts, but you only have to send a small fraction of the batch through each one, right? But you still have a smaller batch at each expert that goes through. And in order to get kind of reasonable balance, like one of the things that the current models typically do is they have all the experts be roughly the same compute cost, and then you run roughly the same size batches through them in order to sort of propagate the very large batch you're doing at inference time and have good efficiency. But I think you know you often in the future might want experts that vary in computational costs by factors of a hundred or a thousand. Or maybe paths that go for many layers on one case and a single layer or even a skip connection in the other case. And there I think you're going to want very large batches still, but you're going to want to kind of push things through the model a little bit asynchronously at inference time, which is a little easier than at training time. And you know, that's part of kind of one of the things that Pathways was designed to support: you have these components and the components can be in variable cost, and you kind of can say for this particular example I want to go through this subset of the model, and for this example I want to go through this subset of the model, and have the system kind of orchestrate that.
这也意味着只有具备一定规模和复杂度的公司才能做到……现在,任何人都可以训练一个足够小的模型。但如果这最终成为训练未来模型的最佳方式,那么你需要一家公司拥有一个数据中心大小的数据中心来服务一个所谓的“blob”或模型。所以这在范式上也会是一个有趣的变化。
It also would mean that it would take companies of a certain size and sophistication to be able to... right now, anybody can train a sufficiently small enough model. But if it ends up being the case that this is the best way to train future models, then you would need a company that can basically have a data center sized a data center serving a single quote unquote blob or model. So it would be an interesting change in paradigms in that way as well.
你当然想这样。
You definitely want to.
至少要有足够的高带宽内存来容纳整个模型。所以根据模型大小,这很可能就是你需要的最低 HBM 容量。
Have at least enough HBM to put your whole model. So depending on the size of your model, most likely that's how much HBM you'd want to have at a minimum.
是的,但这也意味着你不一定需要把整个模型规模扩大到数据中心那么大。你可能希望它稍微小一点,然后对某个使用频繁的专家进行多次复制,以实现更好的负载均衡。比如这个专家因为数学问题多而被频繁调用,而另一个关于塔希提舞蹈的专家则很少被调用。那个专家你甚至可能把它换出到 DRAM,而不是放在 HBM 里。但你希望系统能根据负载特性自动处理所有这些。
Yeah, but it also means you don't necessarily need to grow your entire model footprint to be the size of a data center. You might want it to be a bit below that, and then have potentially many replicated copies of one particular expert that is being used a lot, so that you get better load balancing. Like this one's being used a lot because we get a lot of math questions, and this one on, say, Tahitian dance, it is called on really rarely. That one maybe you even page out to DRAM rather than putting it in HBM. But you want the system to figure all this stuff out based on load characteristics.
现在语言模型显然是输入语言输出语言,但它是多模态的。你可以想象 Pathways 博客文章中提到的许多不同用例,它们并非明显的自回归性质,却都通过同一个模型。那么你能想象谷歌作为一家公司,产品比如谷歌搜索、谷歌图片、Gmail 都通过这个模型,整个服务器就是一个巨大的专家混合模型吗?
Right now language models obviously take language in and output language, but it's multimodal. You could imagine the Pathways blog post talks about many different use cases that are not obviously autoregressive, going through the same model. So could you imagine Google as a company, the product is like Google Search goes through this, Google Images goes through this, Gmail goes through it, just like the entire server is a huge mixture of experts specialized?
你已经开始看到一些这样的例子了:谷歌内部大量使用 Gemini 模型,这些模型不一定经过微调,只是针对特定用例和产品功能给出指令。所以我确实看到底层模型的能力在越来越多的界面中被共享。我认为这绝对是一个相当有趣的方向。
You're starting to see some of this by having a lot of uses of Gemini models across Google that are not necessarily fine-tuned, they're just given instructions for this particular use case and feature in this product setting. So I definitely see a lot more sharing of what the underlying models are capable of across more and more surfaces. I do think that's a pretty interesting direction to go for sure.
我觉得听众可能没有意识到这个预测有多有趣。就像在 2018 年请 Noam 上播客,然后他说:‘是的,我认为语言模型会成为一种东西,这就是未来的方向。’这真是太有趣了。
I feel like people listening might not register how interesting this prediction is. It's like getting Noam on a podcast in 2018 and being like, 'Yeah, I think language models will be a thing, this is where things go.' This is incredibly interesting.
是的,我认为你可能会看到一个大基础模型,然后你可能想要该模型的定制版本,添加不同的模块以适应不同的场景,这些模块可能有访问限制。比如,我们可能有一个内部版本供谷歌员工使用,我们用内部数据训练了一些模块,不允许其他人使用这些模块,但我们可以利用它。其他公司可能会添加对其公司环境有用的模块,并通过我们的云 API 提供服务。
Yeah, and I think you might see that might be a big base model, and then you might want customized versions of that model with different modules added on for different settings that maybe have access restrictions. Like maybe we have an internal one for Google employees that we've trained some modules on internal data and we don't allow anyone else to use those modules, but we can make use of it. And maybe other companies add on other modules that are useful for that company setting and serve it in our Cloud APIs.
使这种系统可行的瓶颈是什么?是系统工程,还是机器学习?这与我们目前的 Gemini 开发方式有很大不同。
What is the bottleneck to making this sort of system viable? Is it systems engineering, is it ML? It's a pretty different way of operating than our current Gemini development.
我认为我们会探索这些领域并取得一些进展,但我们需要真正看到证据表明这是正确的方法,有很多好处。其中一些好处可能是质量的提升,另一些可能不那么具体可衡量,比如能够并行开发大量不同模块。我认为这仍然是一个令人兴奋的改进,因为它能让我们在提升模型针对众多不同领域的能力方面取得更快进展。
I think we will explore these kinds of areas and make some progress on them, but we need to really see evidence that it's the right way, that it has a lot of benefits. Some of those benefits may be improved quality, some may be less concretely measurable, like the ability to have lots of parallel development of different modules. I think that would still be a pretty exciting improvement because it would enable us to make faster progress on improving the models' capabilities for lots of different distinct areas.
即使是数据控制模块化也看起来很酷。你可以有一个专门为我训练的模型部分,它知道我所有的数据,一个个人模块对你来说会很有用。另一个可能是你可以在某些场景中使用特定数据,但在其他场景中不能。也许我们有一些 YouTube 数据只能在 YouTube 产品界面中使用,而不能在其他场景中使用,所以我们可以有一个针对该目的在该数据上训练的模块。
Even the data control modularity stuff seems really cool. You could have a piece of the model that's just trained for me, it knows all my data, a personal module for you would be useful. Another thing might be you can use certain data in some settings but not in other settings. Maybe we have some YouTube data that's only usable in a YouTube product surface but not other settings, so we could have a module trained on that data for that particular purpose.
我们将需要大约一百万个自动化研究人员来发明所有这些。这将会很棒。嗯,这个东西本身,你知道,就像你构建了一个 blob,它告诉你如何让 blob 变得更好,然后 blob 2.0,或者甚至没有版本号,它就像一个不断增长的 blob。
We're going to need like a million automated researchers to invent all of this stuff. It's going to be great. Well, the thing itself, you know, it's like you build the blob and it tells you how to make the blob better, and blob 2.0, or maybe they're not even versioned, it's just like an incrementally growing blob.
好的 Jeff,从宏观角度给我解释一下:为什么这是个好主意?为什么这是下一个方向?
Okay Jeff, motivate for me big picture: why is this a good idea? Why is this the next direction?
我想这种有机的、不是那么精心数学构造的机器学习模型的概念,你已经思考了一段时间。在神经网络的发展中,生物类比、人工神经元,从生物神经元中汲取灵感是很好的,并且在深度学习领域为我们提供了很好的服务。但我觉得我们可能没有像我们本可以的那样,充分关注真实大脑所做的其他事情。这并不是说我们应该完全模仿,因为硅和湿件具有非常不同的特性和优势。但我确实认为我们可以从中汲取更多灵感的一点是,拥有不同专门化部分的概念,即模型或大脑中擅长不同事物的区域。我们在专家混合模型中有一点这样的特性,但它仍然非常结构化。我觉得这种更有机的专业知识增长,当你想要某个领域的更多专业知识时,你在模型的那个部分增加一些容量,让它多学一点这类东西,以及将模型的连接性适应硬件连接性的概念,都是很好的。所以你希望同一芯片和同一 HBM 上的人工神经元之间有非常多的连接,因为那不会花费太多,但然后你希望与附近神经元的连接数量较少,比如一个芯片之外,你应该有一定数量的连接。然后多个芯片之外,你应该有更少的连接,在那里你发送一个非常有限的瓶颈信息:模型这一部分对其他部分最重要的东西。即使在多个 TPU pod 之间,你也希望发送更少的信息,但最显著的表示。然后跨都市区域,你希望发送得更少。而且这是有机地涌现出来的。你可以手动指定这些特性,但我认为你并不确切知道这些连接的正确比例是什么。
I guess this kind of notion of an organic, not quite so carefully mathematically constructed machine learning model is one that's been with you for a little while. In the development of neural nets, the biological analog, the artificial neurons, inspiration from biological neurons is a good one and has served us well in deep learning. But I feel like we're not necessarily looking at other things that real brains do as much as we perhaps could. That's not to say we should exactly mimic that because silicon and wetware have very different characteristics and strengths. But I do think one thing we could draw more inspiration from is this notion of having different specialized portions, areas of a model or a brain that are good at different things. We have a little bit of that in mixture of experts models, but it's still very structured. I feel like this kind of more organic growth of expertise, where when you want more expertise in an area you add some more capacity to the model there and let it learn a bit more on that kind of thing, and also this notion of adapting the connectivity of the model to the connectivity of the hardware is a good one. So you want incredibly many connections between artificial neurons on the same chip and the same HBM because that doesn't cost you that much, but then you want a smaller number of connections to nearby neurons, like a chip away, you should have some amount of connections. Then many chips away, you should have a smaller number of connections where you send a very limited bottleneck thing: the most important things for this part of the model for other parts of the model to make use of. Even across multiple TPU pods, you'd like to send even less information but the most salient representations. And then across metro areas, you'd like to send even less. And that emerges organically. You could hand-specify these characteristics, but I think you don't know exactly what the right proportions of these kinds of connections are.
所以你应该让硬件来稍微决定连接方式。比如,如果你在这里通信,而数据总是到得很早,你就应该增加一些连接。这样会让它花更长时间,并在恰好的时间到达。哦,这里还有一个潜在的有趣含义。现在我们通常认为 AI 使用量的增长是水平扩展。比如,你可能会想,谷歌会有多少 AI 工程师为它工作?你会考虑同时有多少个 Gemini 3 实例在运行。如果你有这个“团块”,它可以有机地决定激活多少自身,那么情况就更像是:如果你想要相当于 10 个工程师的输出,它就激活一个不同的模式或更大的模式;如果你想要 100 个工程师的输出,它并不是调用更多的智能体或实例,而是调用不同的子集。我认为这里有一个概念:你愿意为某个特定推理投入多少算力,而这个量对于非常简单和非常困难的任务应该相差一万倍,甚至可能一百万倍。而且这个过程可能是迭代的:你先通过模型进行一次处理,得到一些结果,然后决定需要调用模型的其他部分作为另一个方面。
And so you should just let the hardware dictate things a little bit. Like if you're communicating over here and this data always shows up really early, you should add some more connections. Then it will make it take longer and show up at just the right time. Oh, here's another interesting implication potentially. Right now we think about the growth in AI use as a sort of horizontal scaling. Suppose you're like, how many AI engineers will Google have working for it? You think about like how many instances of Gemini 3 will be working at one time. If you have this blob, and it can sort of organically decide how much of itself to activate, then it's more like, you know, if you want 10 engineers worth of output, it just activates a different pattern or a larger pattern. If you want 100 engineers of output, it's not like calling more agents or more instances; it's just calling different subsets. I think there's a notion of how much compute you want to spend on this particular inference, and that should vary by factors of 10,000 for really easy things and really hard things, maybe even a million. And it might be iterative, like you might make a pass through the model, get some stuff, and then decide you now need to call on some other parts of the model as another aspect of it.
另外我想说的是,这听起来部署起来非常复杂,因为它是一个奇怪的、不断演变的东西,各部分之间的通信方式可能也不是最优的。但你总是可以从中进行蒸馏,对吧?比如,如果你说,这是我最关心的任务,让我从这个巨大的有机体中蒸馏出一些我知道可以高效服务的东西。你可以随时进行蒸馏,每天一次,每小时一次。这看起来会挺好的。
The other thing I would say is this sounds super complicated to deploy because it's this weird, constantly evolving thing with maybe not super optimized ways of communicating between pieces. But you can always distill from that, right? Like if you say, this is the kind of task I really care about, let me distill from this giant organic thing into something that I know can be served really efficiently. And you could do that distillation process whenever you want, once a day, once an hour. And that seems like it'd be kind of good.
是的,我们需要更好的蒸馏。如果有人有神奇的蒸馏技术,能瞬间从一个大团块蒸馏到你的手机上,那就太棒了。
Yeah, we need better distillation. Anyone out there with amazing distillation techniques that instantly distill from a giant blob onto your phone would be wonderful.
你如何描述当前蒸馏技术缺少什么?
How would you characterize what's missing from current distillation techniques?
嗯,我只是希望它工作得更快。
Well, I just want it to work faster.
一个相关的事情是,我觉得我们在预训练期间需要有趣的学习技术。我不确定我们是否从当前训练目标中看到的每个词中提取了最大价值。对吧?也许我们应该对一些词更深入地思考。当模型得到答案时,也许在训练时它应该比得到正确答案时做更多的工作。是的,也许吧。一定有办法从相同的数据中获得更多,让它正着学、反着学、各种方式学,这样隐藏一些东西,那样隐藏一些东西,让它从部分信息中推断。你知道,这类事情。我认为人们在视觉模型中已经这样做了一段时间了。比如,你扭曲模型或隐藏部分内容,让它从一半中猜出鸟,比如从图像的上角或左下角猜出这是一只鸟。这让任务变得更难。我觉得对于更多的文本或代码相关数据也有类似的方法,你想迫使模型更努力地工作,你会从中得到更有趣的观察。
A related thing is I feel like we need interesting learning techniques during pre-training. I'm not sure we're extracting the maximal value from every token we look at with the current training objective. Right? Like maybe we should think a lot harder about some tokens. When you get to the answer, maybe the model should at training time do a lot more work than when it gets to the right answer. Yeah, maybe right. There's got to be some way to get more from the same data, make it learn it forwards and backwards and every which way, hide some stuff this way, hide some stuff that way, make it infer from partial information. You know, these kinds of things. I think people have been doing this in vision models for a while. Like you distort the model or you hide parts of it and try to make it guess the bird from half, like it's a bird from this upper corner of the image or the lower left corner of the image. And that makes the task harder. And I feel like there's an analog for more textual or coding related data where you want to force the model to work harder, and you'll get more interesting observations from it.
是的,做图像的人没有足够的标注数据,所以他们不得不发明所有这些。比如,Dropout 是在图像上发明的,但我们主要没有在文本上使用它。这是一种方法,可以在更大规模的模型中获得更多学习而不至于过拟合,只需在世界文本数据上做 100 个周期并使用 Dropout。是的,但这计算成本很高。但这确实意味着我们不会用完文本数据。尽管人们说,哦不,我们快没有文本数据了,我其实不太相信,因为我认为我们可以从现有的文本数据中获得更强大的模型。我的意思是,一个人大概看过十亿个词,他们在很多事情上都很擅长。是的,显然人类的数据效率设定了一个下限,或者说上限,也许这是一个有趣的数据点。
Yeah, the image people didn't have enough labeled data, so they had to invent all this. And like, Dropout was invented on images, but we're not really using it for text mostly. That's one way you could get a lot more learning in a more large-scale model without overfitting is just make like 100 epochs over the world's text data and use Dropout. Yeah, but that's pretty computationally expensive. But it does mean we won't run out of textual data. Even though people are saying, oh no, we're almost out of textual data, I don't really believe that because I think we can get a lot more capable models out of the text data that does exist. I mean, a person has seen like a billion tokens, and they're pretty good at a lot of stuff. Yes, obviously human data efficiency sets a lower bound on how, or an upper bound, maybe it's an interesting data point.
是的。所以这里有一种肯定前件和否定后件的逻辑。一种看法是:看,语言模型还有很长的路要走,因此我们预测,如果它们能匹配人类,样本效率会有数量级的提升。另一种是:也许它们在做完全不同的事情,考虑到数量级的差异。你的直觉是什么,要让这些模型像人类一样样本高效需要什么?
Yes. So there's a sort of modus ponens, modus tollens thing here. One way to look at it is: look, LMs have so much further to go, therefore we project orders of magnitude improvement in sample efficiency just if they could match humans. Another is: maybe they're doing something clearly different given the orders of magnitude difference. What's your intuition of what it would take to make these models as sample efficient as humans are?
是的,我的意思是,我认为我们应该考虑稍微改变训练目标。仅仅根据之前看到的词预测下一个词似乎不是人类学习的方式,对吧?我认为这与人类学习有点关系,但不完全一样。比如,一个人可能会读一整章书,然后尝试回答后面的问题,这是不同的事情。我还认为我们没有从视觉数据中学到很多。我们确实在视频数据上训练了一点,但我们绝对还没有接近考虑在所有可能的视觉输入上进行训练。所以,我们还有视觉数据没有真正开始训练。然后,我认为我们可以从我们看到的每一点数据中提取更多信息。你知道,我认为人类如此样本高效的原因之一是他们会探索世界,在世界上采取行动,并观察发生了什么。是的,对吧?就像你看到很小的婴儿,拿起东西然后扔下,他们从中学习重力。而当你没有主动发起动作时,学习这些东西要困难得多。我认为让模型在它的学习过程中能够采取行动,会比仅仅被动地观察一个巨大的数据集要好得多。那么这就是未来吗?模型可以观察、采取行动并观察相应结果的东西似乎非常有用。
Yeah, I mean I think we should consider changing the training objective a little bit. Just predicting the next token from the previous ones you've seen seems like not how people learn, right? It's a little bit related to how people learn, I think, but not entirely. Like a person might read a whole chapter of a book and then try to answer questions at the back, and that's a different kind of thing. I also think we're not learning from visual data very much. We're training a little bit on video data, but we're definitely not anywhere close to thinking about training on all the visual inputs you could get. So you have visual data that we haven't really begun to train on. And then I think we could extract a lot more information from every bit of data we do see. You know, I think one of the ways people are so sample efficient is they explore the world and take actions in the world and observe what happens. Yeah, right? Like you see it with very small infants, picking things up and dropping them, they learn about gravity from that. And that's a much harder thing to learn when you're not initiating the action. And I think having a model that can take actions as part of its learning process would be just a lot better than just sort of passively observing a giant dataset. Is that the future then? Something where the model can observe and take actions and observe the corresponding results seems pretty useful.
是的,我的意思是,人们可以从思想实验中学到很多东西,甚至不需要额外的输入。比如爱因斯坦从思想实验中学到了很多东西,或者牛顿在隔离期间被苹果砸到头之类的事情,然后发明了引力。还有数学家,你知道,数学没有任何额外的输入。国际象棋,比如,你让东西自己和自己下棋,它就会变得擅长国际象棋。那是 DeepMind 做的,但它只需要国际象棋的规则。所以实际上,即使没有外部数据,你可能也能学到很多东西。然后你可以让它在你关心的领域里学习。
Yeah, I mean people can learn a lot from thought experiments that don't even involve extra input. Like Einstein learned a lot of stuff from thought experiments, or like Newton went into quarantine and got an apple dropped on his head or something and invented gravity. And like mathematicians, you know, math didn't have any extra input. Chess, like okay, you have the thing play chess against itself and it gets good at chess. That was DeepMind, but also all it needs is the rules of chess. So there's actually probably a lot of learning that you can do even without external data. And then you can make it in exactly the fields that you care about.
你刚才在过去一小时里阐述的,可能是 AI 下一个重大范式转变。这是一个极具价值的洞察。你怎么看?2017 年你发表了 Transformer 论文,它奠定了数百亿甚至上千亿美元的市场价值,还不提谷歌多年来发布的其他研究。回想起来,当想到这些信息泄露出去帮助了竞争对手,你还会这么做吗?还是会说我们没意识到 Transformer 有多重要,应该内部保密?你怎么看待这个问题?
You've just laid out over the last hour is potentially the big next paradigm shift in AI. That's a tremendously valuable insight. How do you know? In 2017 you released the Transformer paper on which tens if not hundreds of billions of dollars of market value is based, not to mention all this other research that Google has released over time. In retrospect, when you think about divulging this information that has been helpful to your competitors, would you still do it, or would you say we didn't realize how big a deal Transformer was and we should have kept it indoors? How do you think about that?
好问题。我认为我们需要看到机会的规模,这通常反映在其他公司的行动中。而且,这不是一个零和游戏;当前的世界状态离零和游戏远得不能再远了。我认为我们将看到 GDP、健康、财富以及你能想到的任何方面出现数量级的改善。所以 Transformer 能够传播开来并带来变革,这绝对是件好事。谢天谢地,谷歌也做得很好。如今我们确实少发布了一些正在做的事情,但这里总是存在权衡:我们是应该立即发布我们正在做的事情,还是将其投入下一阶段研究并推出到生产中的 Gemini 模型而不发布,或者采取某种中间立场?例如,在 Pixel 相机的计算摄影工作中,我们经常开发有趣的新技术,比如用于弱光的超级夜景模式,将其投入产品,然后在产品发布后发表研究论文。不同的技术和开发有不同的处理方式。有些我们认为极其关键的东西可能不会发布;有些我们认为有趣但对改进产品重要的事情,我们会先融入产品,然后决定是否发布,或者进行轻量级讨论而不透露每个细节。其他一些事情我们会公开分享以推动领域和社区进步,因为这是我们所有人受益的方式。参加像上周的 NeurIPS 这样有 15000 人分享大量好想法的会议真是太棒了。我们在那里发表了很多论文,看到领域进步非常令人兴奋。
It's a good question. I think we needed to see the size of the opportunity, often reflected in what other companies are doing. Also, it's not a fixed pie; the current state of the world is as far from a fixed pie as you can get. I think we're going to see orders of magnitude of improvements in GDP, health, wealth, and anything else you can think of. So it's definitely been nice that Transformer has gotten around and transformed things. Thank God Google's doing well as well. These days we do publish a little less of what we're doing, but there's always this trade-off: should we publish exactly what we're doing right away, put it into the next stages of research and roll it out into production Gemini models and not publish at all, or some intermediate point? For example, in our computational photography work for Pixel cameras, we often develop interesting new techniques like super good Night Sight for low light, put that into the product, and then publish a research paper after the product is released. Different techniques and developments have different treatments. Some things we think are super critical we might not publish; some things we think are interesting but important for improving our products, we get them into products and then decide whether to publish or give a lightweight discussion without every last detail. Other things we publish openly to advance the field and the community, because that's how we all benefit. It's great to go to conferences like NeurIPS last week with 15,000 people sharing lots of great ideas. We published a lot of papers there, and seeing the field advance is super exciting.
你怎么解释谷歌很早就内部拥有所有这些洞察,包括顶尖研究人员,而到了 2024 年,Gemini 2 已经发布,人们知道它是一个非常棒的模型,在 LMSys 聊天机器人竞技场中排名第一。所以现在谷歌领先了,但你怎么解释在几年里提出了所有伟大洞察,而其他竞争对手却一度拥有更好的模型?
How would you account for Google having all these insights internally rather early on, including the top researchers, and now as of 2024 Gemini 2 is out, people know it's a really great model. It's top in LMSys Chatbot Arena. So now Google's on top, but how do you account for basically coming up with all the great insights for a couple of years, while other competitors had models that were better for a while despite that?
当然。我认为我们在语言模型上已经工作了很长时间。2001 年早期的拼写纠正工作,2007 年的翻译和超大规模语言模型工作,然后是 Transformer、BERT,以及内部的 Meena 系统,这是一个基于聊天的系统,旨在让人们参与有趣的对话。我们实际上在 ChatGPT 出现之前就有了一个内部聊天机器人系统,谷歌员工可以试用。疫情期间,很多谷歌员工被封锁在家,他们喜欢在午餐时间与 Meena 聊天,因为它是一个很好的互动伙伴。从搜索的角度来看,我们有点不确定的是这些模型经常产生幻觉,很多时候不能给出正确答案,这意味着它们没有达到应有的实用性。所以我们想改进这一点。从 CEO 的角度来看,你希望 100%的时间得到正确答案,理想情况下事实性非常高,而这些模型远未达到。但我认为我们没有充分认识到的是,它们对于你不会问搜索引擎的事情有多有用,比如帮我给兽医写个便条,或者把这段文字快速总结一下。这正是人们蜂拥而至的原因,他们将聊天机器人视为惊人的新能力,而不是纯粹的搜索引擎。所以我们花了一些时间,最终发布了相当有能力的聊天机器人,并通过 Gemini 模型不断改进。我认为这实际上是一条不错的路径。我们是否希望更早发布聊天机器人?也许吧。但我认为我们拥有一个非常棒的聊天机器人,搭载了越来越好的 Gemini 模型,这很酷。
Sure. I think we've been working on language models for a long time. Early work on spelling correction in 2001, work on translation and very large-scale language models in 2007, and then Transformers, BERT, and the internal Meena system, which was a chatbot-based system designed to engage people in interesting conversations. We actually had an internal chatbot system that Googlers could play with even before ChatGPT came out. During the pandemic, a lot of Googlers were locked down at home and enjoyed spending time chatting with Meena during lunch because it was a nice and engaging partner. One thing we were a little unsure about from a search perspective was that these models hallucinate a lot and don't get things right a lot of the time, meaning they aren't as useful as they could be. So we wanted to make that better. From a CEO perspective, you want to get the right answer 100% of the time, ideally very high factuality, and these models were not near that. But I think what we didn't quite appreciate was how useful they could be for things you wouldn't ask a search engine, like help me write a note to my veterinarian or take this text and give me a quick summary. That's the kind of thing people have flocked to in terms of using chatbots as amazing new capabilities rather than as a pure search engine. So we took our time and got to the point where we released quite capable chatbots and have been improving them through Gemini models quite a bit. I think that's actually not a bad path to have taken. Would we have liked to release the chatbot earlier? Maybe. But I think we have a pretty awesome chatbot with awesome Gemini models that are getting better all the time, and that's pretty cool.
我们讨论了过去 25 年里你们所做的一些事情,涉及众多不同领域:搜索和索引、分布式系统、硬件、AI 算法,还有成千上万的其他领域。拥有这种职业生涯的持久性——几十年来不断取得突破,同时贡献范围如此广泛——有什么诀窍?
We discussed some of the things you guys have worked on over the last 25 years, and there are so many different fields: search and indexing, distributed systems, hardware, AI algorithms, and genuinely a thousand more. What is the trick to having this level of career longevity where you have many decades of making breakthroughs, but also the breadth of different contributions?
我认为关键在于保持好奇心,愿意学习新事物。此外,与优秀的人一起工作,解决重要的问题。谷歌提供了一个环境,让你可以探索不同领域并产生影响。这不仅仅是一个诀窍,而是不断适应和寻找新挑战。
I think it's about staying curious and being willing to learn new things. Also, working with great people and on problems that matter. Google has provided an environment where you can explore different areas and have impact. It's not just about one trick; it's about continuously adapting and finding new challenges.
是的,我认为我喜欢做的一件事就是发现新的有趣领域。最好的方法之一是关注正在发生的事情,与同事交流,关注发表的研究论文,观察研究格局的演变。愿意说,‘哦,芯片设计——我想知道我们能否将强化学习用于其中的某个方面’,然后能够深入一个新领域。与了解不同领域的人合作,比如医疗保健,与临床医生一起了解真正的问题所在。‘我怎么能帮上忙?这对这个事没什么用,但对那个事会非常有用。’获得这些见解,通常与五六个拥有不同专业知识的同事一起工作,能让你们共同完成任何个人都无法单独完成的事情。然后他们的专业知识会影响你,你的也会影响他们,现在你作为工程师和研究人员,工具箱里有了更多的工具去应对下一个挑战。我认为这就是在职继续学习的美妙之处之一。这是我珍视的东西,我真的很喜欢接触新事物,看看我们能做些什么。
Yeah, I mean, I think one thing that I like to do is to find out about a new and interesting area. One of the best ways to do that is to pay attention to what's going on, talk to colleagues, pay attention to research papers being published, look at the research landscape as it's evolving. Be willing to say, 'Oh, chip design—I wonder if we could use reinforcement learning for some aspect of that,' and be able to dive into a new area. Work with people who know a lot about a different domain, like healthcare, working with clinicians about what the real problems are. 'How could I help? It wouldn't be that useful for this thing, but it would be super useful for that.' Getting those insights, and often working with a set of five or six colleagues who have different expertise than you do enables you to collectively do something that none of you could do individually. Then some of their expertise rubs off on you and some of yours rubs off on them, and now you have a bigger set of tools in your tool belt as an engineer and researcher to go tackle the next thing. I think that's one of the beauties of continuing to learn on the job. It's something I treasure, and I really enjoy tying into new things and seeing what we can do.
我想说,可能很重要的一点是谦逊。比如,我是最谦逊的人,但说真的,要有能力说,‘嘿,我刚才做的与我所能做的或可能做到的相比算不了什么’,并且一旦看到更好的想法就能放弃原有的想法。你听到别人有更好的想法,然后你看到也许你正在想的和他们在想的,或者完全不同的东西,可能会更好地运作。因为我认为在某种程度上,有一种冲动要说,‘嘿,我刚发明的东西太棒了,给我更多算力’,尤其是在有很多自上而下的资源分配时。但我认为我们也需要激励人们说,‘嘿,我正在做的这件事根本行不通,让我完全放弃它,尝试别的东西。’我认为 Google Brain 做得很好,采用了一种非常自下而上的算力分配方式,基本上每个人都有一个积分,你可以为一个好主意使用它。而 Gemini 项目,主要是自上而下的,这在某种意义上非常好,因为它带来了更多的合作。你很少看到五组人都在构建相同的东西或可互换的东西。但另一方面,它确实导致了一些激励,让人们说,‘嘿,我正在做的事情效果很好,而且你听到数百个小组都在说,所以你应该给他们更多算力’,而很少激励人们说,‘嘿,我正在做的事情实际上效果不太好,让我尝试一些不同的东西。’所以我认为未来,我们将会有一定程度的自上而下和一定程度的自下而上,以激励这两种行为:合作和灵活性。因为我认为这两者都能带来大量创新。
I'd say probably a big thing is humility. Like, I'm the most humble ever, but seriously, there's the ability to say, 'Hey, what I just did is nothing compared to what I can do or what can be done,' and to be able to drop an idea as soon as you see something better. You hear somebody with a better idea, and you see how maybe what you're thinking about and what they're thinking about, or something totally different, could conceivably work better. Because I think there is a drive in some sense to say, 'Hey, the thing I just invented is awesome, give me more chips,' particularly if there's a lot of top-down resource assignment. But I think we also need to incentivize people to say, 'Hey, this thing I am doing is not working at all, let me just drop it completely and try something else.' I think Google Brain did quite well with a very kind of bottom-up chip allocation where basically everyone had one credit and you could pull them for a good idea. And then Gemini, it has been mostly top-down, which has been very good in some sense because it has led to a lot more collaboration. You less often have five groups of people all building the same thing or building interchangeable things. But on the other hand, it does lead to some incentive to say, 'Hey, what I'm doing is working great, and as you hear hundreds of groups and everything, you should give them more chips,' and there's less incentive to say, 'Hey, what I'm doing is not actually working that well, let me try something different.' So I think going forward, we're going to have some amount of top-down and some amount of bottom-up so as to incentivize both of these behaviors: collaboration and flexibility. Because I think both those things lead to a lot of innovation.
是的,我认为阐明你认为我们应该去的有趣方向也很好。我有一个内部幻灯片,叫做‘杰夫的古怪想法’,它更偏向产品导向,比如‘嘿,我认为既然我们有了这些能力,我们可以做这 17 件事。’我认为这是一件好事,因为有时人们会对此感到兴奋,并想开始研究其中一件或多件事。我认为这是一种很好的方式,可以引导我们前进的方向,而不必命令人们‘我们必须去这里’。
Yeah, I think it's also good to articulate interesting directions you think we should go. I have an internal slide deck called 'Jeff's Wacky Ideas' that is a little bit more product-oriented, like 'Hey, I think now that we have these capabilities, we could do these 17 things.' I think that's a good thing because sometimes people get excited about that and want to start work on one or more of them. I think that's a good way to kind of bootstrap where we should go without necessarily ordering people 'we must go here.'
嘿,这太棒了。谢谢。感谢你抽出时间,这很棒。
Hey, this is great. Thank you. I appreciate you taking the time, and it was great.