Noam Shazeer: From Google's Early Days to AI's Future
打开互动全文版(中英对照 + 朗读 + 问答)→诺姆·沙泽尔,谷歌 21 年老将和 Transformer 论文合著者,讲述他从数学到 AI 的历程、谷歌早期岁月以及大语言模型内部的奥秘。
Noam Shazeer, a 21-year Google veteran and co-author of the Transformers paper, discusses his journey from math to AI, the early days at Google, and the mysteries inside large language models.
我们请到了 Noam Shazeer。他在谷歌工作了 21 年,参与了 2017 年著名的 Transformer 论文,还参与了 TensorFlow。Noam,你是这个领域的传奇。感谢你来到我们的节目。
We have Noam Shazeer. He spent 21 years at Google. He was involved in the famous Transformers paper from 2017. Was involved with TensorFlow. Noam, you're a legend in this field. Thank you for joining us.
这真的是一门非常实验性的科学,很像我想象中炼金术时代的化学。我认为未来几年会有一个根本性问题:这些模型是在什么数据上训练的?你可以直接和这些模型对话,这是最酷的事情。为什么我们不能持续构建这些不可思议的概率模型,让它总是选出下一个正确的词?我们叫它 Phil。我们觉得它有一天会变成 AI,而且我们希望它听起来不具威胁性。你说没人真正知道这些模型内部发生了什么。我们把它变成了 AdSense 的第一个启动函数,最终赚了很多钱。能请到这篇论文的作者之一来给我们解释注意力和自注意力机制,真是太荣幸了。
It's really a very experimental science. It's a lot like I would imagine like chemistry was in the days of alchemy. I think there is going to be an existential question for the next couple of few years about what data is this trained on. You can just talk to these models and it's the coolest thing ever. Why can't you consistently build out these incredible probabilistic models where it always picks the next right word? We called it Phil. We figured it would one day turn into AI and we wanted it to sound non-threatening. You said nobody really knows what's happening inside these models. We turned it into like the first starting function for AdSense which ended up making a lot of money. It's quite the treat to have one of the authors of this paper to come and explain to us what attention and self-attention means. It's such a rare privilege.
大家好。今天我们有一期特别节目,嘉宾是 Noam Shazeer。Noam 的履历和背景非常出色,我想好好介绍一下。Noam 现在是 Character.AI 的创始人兼 CEO。在此之前,如果我没算错的话,他在谷歌工作了 21 年。中间离开过几年,但大部分时间都在那里。好吧,在谷歌待了相当长的时间。简单来说,Noam 很可能处于当今大语言模型研发工作的核心,或者负责了很多相关工作。举几个例子:他参与了 2017 年著名的 Transformer 论文,我们还会深入讨论 2016 年的稀疏门控混合专家模型;他参与了 TensorFlow;是谷歌 LaMDA 对话系统的主要贡献者。这样的例子不胜枚举。Noam,你是这个领域的传奇。这真是一束光。感谢你今天来到我们节目。谢谢你,Noam。
Hey everyone. We have a special episode for you. We have Noam Shazeer. And Noam has such a fantastic resume and background that I want to make sure I read this out. Noam is currently the founder and CEO of Character.AI. But before this, he spent 21 years, if my math is right, at Google. I left a few years in the middle, but I was mostly there. Okay, quite a long period of time at Google. Just to put it simply, you know, Noam has probably been at the heart or has been responsible for a lot of the research and development work that has gone into large language models today. So, just some examples, he was involved in the famous Transformers paper from 2017, which we're going to really get into with sparsely gated mixture of experts in 2016, was involved with TensorFlow, was a major contributor to Google's LaMDA dialogue system. The list goes on and on. Noam, you're a legend in this field. This is such a light. Thank you for joining us today. Thank you, Noam.
哦,谢谢。感谢邀请我。
Oh, thanks for it. Thanks for having me on.
好了,我们马上要聊 AI 了,但我想先问你一个有趣的故事。我不知道是不是真的,所以你得告诉我们真假。我在谷歌上搜你、准备这次对话时,看了 Hacker News,上面有个关于你的故事说,你加入谷歌时,面试中给出的一个答案后来被 Gmail 用上了。我不知道是不是真的,但我很好奇你是怎么进入谷歌的,以及之后的故事。
Okay, so we're going to get to AI in just a bit, but I want to ask you about a fun story. I don't know whether it's true or not, so you got to tell us whether it's true or not. When I was Googling you and preparing for this conversation, I was looking at Hacker News, and there is a story about you on Hacker News which says that when you joined Google and you were originally interviewed, one of the answers you came up in the interview actually wound up being something that Gmail wound up using later. So, I don't know whether it's true or not, but I'm curious to kind of hear your history about how did you wind up in Google and the story ever since?
让我想想。我本科去了杜克大学,拿的是数学奖学金。不是篮球。我在那儿受伤了。没开玩笑。我开玩笑的。但有一个数学教授招募了我,因为他们想赢一场数学竞赛。我不知道为什么,但总之我拿着奖学金去了那里,然后学了数学。我想我在高中时数学竞赛之类的很厉害,所以以为我应该学数学。但后来我意识到:等等,我并不真正喜欢为了乐趣做数学。我只是喜欢编程,喜欢看能不能让电脑做点聪明的事。所以也许我该转学计算机科学。于是我开始了研究生学业,又退学了两三次,然后给谷歌投了简历,因为我不知道。也许我用过谷歌。我想:好吧,看起来是个有趣的工作地点。实际上,我在 1999 年伯克利的一个招聘会上错过了谷歌,因为我以为它已经是一家大公司了,因为我认识的人都在用谷歌。但我想我认识的人样本偏差很大,因为他们都是计算机科学的研究生。如果我当时知道它那么小,我会更早加入的。但我不太确定那个面试故事是什么。可能和 Paul Buchheit 有关,顺便说一句,他是 Gmail 的创始人,非常棒,也是我们的投资者之一。但他可能比我更记得那件事,因为我觉得面试时我有点紧张。
Let's see. I went to Duke for undergrad on like a math scholarship. Not basketball. I got injured there. No joke. I'm joking. But there was a math professor that actually recruited me because they were trying to win this math competition. I have no idea why, but in any case I ended up going there on the scholarship and then studied math and then I guess I was really good at math competitions and stuff in high school. So, assumed I should do math, but then at some point realized, 'Wait, I don't really enjoy doing math for fun. I just like programming and seeing if I can get the computer to do something smart. So, maybe I'll take up computer science instead.' So, ended up starting grad school and dropping out two or three times and then submitted a resume to Google because I don't know. Maybe I used Google. I thought, 'Okay, seems like an interesting place to work.' And I had actually passed over Google at a job fair at Berkeley in 1999 because I assumed it was already a huge company since everyone I knew used Google, but I think I had a very biased sample of who I knew because they were grad students in CS. Had I realized it was actually that small I would have joined earlier. But I'm not quite sure what the interview story was. It might have had to do with Paul Buchheit and the founder of Gmail by the way, who's pretty awesome and also one of our investors. But yeah, he probably recalls the incident better than I do because I think I was a little stressed out about the interview.
是关于拼写纠正器吗?
Was it a spell corrector?
哦,对对对。我想他问我怎么做拼写纠正器,然后我最终写出了谷歌第一个好用的拼写纠正器。之前我们用的是第三方系统,类似于文字处理器里的那种。我的意思是,类似于基于人工编辑的词典,大概有 5 万个词。这很有趣,因为它在文字处理上表现不错,但在谷歌网页搜索上表现糟糕,因为网页搜索的词汇量极其广泛。所以你输入 TurboTax,它会问“你是不是想找 turbotax?”因为大多数时候它都是错的,主要是误报。所以我们做了一个数据驱动的拼写纠正器,效果很好。
Oh yeah yeah I think he asked me how to do a spell corrector and then I ended up writing the first good spell corrector at Google. Previously we used a third-party system that was something out of a word processor. I mean something similar to something based on a human edited dictionary of maybe 50,000 words and it was pretty funny because it worked well for word processing but worked terrible for Google web search because the vocabulary for web search is so incredibly broad. So you type in TurboTax and it would say 'did you mean turbotax?' because most of the time it was just wrong, mostly false positives. So we did a data-driven one that was actually good.
不错。
Nice.
谷歌现在在员工数量、规模和影响力上都大了几个数量级。当时的谷歌有多不同?你在 2000 年代初期看到它发生了怎样的变化?
Google is obviously a few orders of magnitude larger now in terms of employee count and size and impact. How different was Google back then and how did you see it change over the early 2000s?
是的,我的意思是,当时绝对非常有趣。我觉得现在也仍然很有趣。让我想想。我加入的时候大概有几百人。所以每个人都认识彼此,这很棒。即使在那时,公司也是自下而上、自我驱动的,几乎没有工程管理。我记得有一次他们实际上解雇了所有工程经理,但没人注意到。作为一名工程师,我当然没注意到。就像:好吧,现在我要向副总裁汇报了,其他人也是。好的,回去工作。因为每个人都知道自己需要做什么来让事情成功。在那种小规模下,这有点神奇。但我认为谷歌仍然是一家非常自下而上的公司,至少在我参与的研究部门是这样。我几乎总是做任何我想做的事,而且结果都很好。
Yeah, I mean it was definitely super fun. I think it's still super fun. Let's see. I guess when I joined it was a couple hundred people. So everyone knew each other which was great, and even then it was very bottom-up self-directed, very little engineering management. I think at some point they actually fired all the engineering managers and nobody noticed. I certainly didn't notice this as an engineer. It's like, 'Okay, now I'm reporting to a VP and so is everyone else. Okay, back to work.' Because everyone just kind of knew what they needed to do to make the thing succeed. It was kind of magical at those small sizes, but I think Google is still a very bottom-up sort of company, at least in the research parts I was involved in. I pretty much always just did whatever I felt like and it worked out pretty well.
所以我们听到的说法是,拉里·佩奇有一天决定经理不好,基本上让谷歌变成了一个扁平化组织。我不知道是不是他,或者他让所有工程师都向一个人汇报。
So the item we're hearing is that Larry Page one day decided that managers were bad and essentially made Google into a flat organization and I don't know whether it's him or maybe he had all the engineers report to one person essentially.
有趣的是,我想这大概是在 21 世纪初,二十年后你看看埃隆·马斯克在推特的做法,有一些相似之处,就是砍掉大量中层管理,让工程师做他们最擅长的事。
And you know, the funny thing is I think this was probably in the early 2000s and two decades later when you look at what Elon Musk is doing at Twitter, there are some shades of similarities which is let's cut out a lot of middle management and let's you know, have the engineers do the things that they do best.
是的,非常同意。我的意思是,像我们面试过一些产品管理人员,大多数情况下他们似乎根本不知道这项技术能做什么,只会说些乱七八糟的话,比如“你需要确定你要进入哪些垂直领域并专注于此”之类的,结果我认为这对大语言模型来说完全是错误的策略,因为这东西的优势在于通用性。而且,在我们有产品之前,真的很难想象它能做什么,所以我们甚至觉得招产品方面的人没什么用。未来可能会改变,但到目前为止,主要是那些理解可能性的工程师最有希望设定方向。
Yeah, yeah, very much. So I mean a character like we have interviewed some like product management people and for the most part it just seemed like you know, they have no idea of what the technology is capable of and they just said all kinds of like random things like okay you need to figure out which verticals you're going into and focus on those or something which ended up to be you know I think is like the exact wrong strategy for large language models because the strength of this thing is the generality. And you know, so people just before we had a product it's just very very hard to kind of imagine what the thing is going to be capable of so we didn't really find it useful to even hire anyone on the product side. I mean that will change probably in the future but you know so far it is kind of mostly engineers who kind of understand what the possibilities are that have the best chance of setting the direction.
好的,我们这次对话会花很多时间讨论 AI 和大语言模型,但在此之前,你能不能为我们概述一下——你在谷歌待了那么久,所以聊你的职业生涯特别有意思——你参与过的一些值得注意的里程碑和项目?因为我认为这直接关联并重叠了过去 20 年 AI 的发展。
Okay so I want to we're going to spend a lot of time in this conversation talking about AI and LLMs but before that could you maybe outline for us some of the and you were at Google for such a long time so this is particularly interesting when kind of talk about your career there some of the few notable milestones and projects that you wanted being involved in because I think kind of directly leads and overlaps with the development in AI in the last 20 years.
是的,早期我做了很多当时被认为是 AI 的东西,比如拼写纠正器,然后我们做了一个无监督主题聚类系统。我认为我们发明了和伯克利后来称为潜在狄利克雷分配非常相似甚至相同的东西,他们是在我们之后几年做的,规模可能只有我们的百分之一。我们当时对发表之类的一无所知,就只是把它造出来,我们叫它 Phil。我不太确定为什么,但我们觉得它有一天会变成 AI,而且我们希望它听起来不具威胁性。就我和另一个早期谷歌人乔治·哈里克。然后我们把它变成了 AdSense 的第一个定向功能,最终赚了很多钱,所以我记得……嗯。我花了很多时间做广告技术,必须竞争,所以很能理解。
Yeah, I mean yeah I guess in the early days I did a bunch of stuff that was considered AI at the time like the spell corrector and then we did a system for unsupervised topic clustering. I think that we invented something very similar or the same thing at Berkeley and called it latent Dirichlet allocation and that was like a couple of years after we did it and maybe at a hundredth the scale and we called it like we didn't know anything about publishing or anything we're just like oh let's just build the thing and we called it Phil. I'm not quite sure why but we figured it would one day turn into AI and we wanted it to sound non-threatening. But it was me and this other guy, George Harrick, who was another one of the early Google people. But yeah, then we turned it into the first targeting function for AdSense, which ended up making a lot of money, so I remember Oh well. I spent a lot of time working on ad tech and had to compete to it can definitely appreciate that.
好吧,我们进入正题。我想聊聊大语言模型,但在此之前,你能不能跟我们谈谈过去二三十年 AI 研究的历史?因为我们在学校时读过约翰·麦卡锡,对吧?我们见过……我爱约翰·麦卡锡,是的。天哪,是的。Arty 是个 Lisp 黑客。曾经是。曾经是。很多 AI 都来自那个世界,Symbolics 之类的。然后大概在 80 年代,AI 世界出现了一种幻灭。你能不能跟我们谈谈 AI 的历史,以及在那之前尝试过的一些方法,直到过去 10-12 年我们看到深度学习和大语言模型的复兴?
Okay, maybe let's get to it. I want to talk about LLMs, but maybe before that could you maybe talk to us a little bit about the last 20-30 years of AI research. Because you know, when we were in school, we read John McCarthy, right? We saw I love John McCarthy, yeah. Oh my god, yes. Arty is a Lisp hacker. Used to be. Used to be. And you know, a lot of AI kind of came from that and Symbolics and that world. And then there was sort of a disillusionment, so to speak, in the world of AI, maybe in the in the '80s. Maybe talk to us a little bit about the history of AI and some of the approaches that were tried before, you know, kind of the renaissance that we've seen in the last 10-12 years with deep learning LLMs.
是的,当我开始对 AI 感到兴奋时,当时令人兴奋的是贝叶斯网络。每个人都在做这个,非常非常令人兴奋,我记得我迷上了它,因为我真的很喜欢概率。它一直是我在数学中最喜欢的领域。我在杜克大学参加了一个很棒的研究生研讨会。教授叫迈克·利特曼,他说:“好吧,我们全班合作,构建一个能解填字谜的系统。”因为他喜欢填字谜,我们很多人也喜欢,因为当时没有手机之类的。所以你就拿起学生报纸,在课上坐着做填字谜。另一个学生格雷格·金德,现在在 Character 和我们一起工作,他从网上抓取了很多填字谜,所以有一个很大的线索数据库,我们用它作为填字谜解算器的起点,然后我无意中重新发明了信念传播算法来填充网格。我想我应该读过并知道它已经存在,但我当时非常兴奋,直到我读到它已经被做过了。
Yeah, I guess when I started getting excited about AI, the exciting thing then was Bayesian networks. And everyone was doing those and it was really really exciting and I remember kind of getting into that because I really loved probability. It was always my favorite field in math. And I did this great sort of graduate seminar at Duke. The professor's name was Mike Littman and he said, 'Okay, we're going to have the class all collaborate and we will build a system that solves crossword puzzles.' Because he was into crossword puzzles and a lot of us were into crossword puzzles because there were no cell phones or something. So you just pick up the student newspaper and sit and do the crossword puzzle during the lecture. So one of the other students Greg Kind who's now working with us at Character had collected a bunch of crossword puzzles from scraping the web and so he had a big database of clues and we used that as a start for the crossword solver and then I kind of unknowingly reinvented this belief propagation algorithm to fill in the grid. I guess I should have read about it and understood that it existed but I was very excited about it till I read that it had already been done.
在那个时代,我认为贝叶斯非常重要,因为如果你看保罗·格雷厄姆写的著名的《垃圾邮件计划》,如果我没记错的话,很多反垃圾邮件实际上都是基于贝叶斯的。我记得下载了一个 Outlook 插件,它在客户端使用贝叶斯。基本上就是在客户端用贝叶斯过滤垃圾邮件。所以它对今天我们视为理所当然的很多东西都非常基础。
Well in that era I think Bayes was super important there because if you look at Paul Graham wrote this famous a plan for spam if I remember and a lot of spam fighting was actually built on Bayes. I remember downloading I forget this Outlook plug-in which would use Bayes on the client side. Yeah and which would basically use Bayes on the client side to filter out spam. So it was very foundational to so much of you know what we take for granted today.
是的,然后大概 2012 年,我离开谷歌几年,但后来决定回来,因为我去找妻子吃午饭,她还在谷歌工作,碰巧坐在杰夫·迪恩旁边,他在谷歌大脑团队,我当时想,哇,这看起来是一群非常聪明的人在做什么有趣的事,让我试试吧。我之前从没做过神经网络,但后来就投入进去了。我认为这纯粹是硬件的问题——硬件变得更并行,但时钟速度没有更快。所以硬件能力的增长本质上都在并行性上。现在你有这些处理器,能做海量算术,但在顺序操作数量以及数据进出移动方面有限制,移动数据成本高,而算术便宜。然后你有神经网络,主要是矩阵乘法。矩阵乘法恰好是在现代硬件上能做得非常好的操作。所以结果就是,现在你可以运行这些神经网络算法,比那些在内存中乱窜、分支、跟踪指针之类的东西快很多很多个数量级。
Yeah, yeah and then I guess yeah like 2012 I left Google for a few years but then decided to come back when I visited my wife for lunch because she was still working at Google and happened to sit next to Jeff Dean in the Google Brain team and was like wow this seems like really smart bunch of people doing something interesting like let me try this. I guess I'd never done neural networks before but kind of got into it. And I mean I think that it's just a matter of the hardware like hardware has gotten more parallel, but not faster clock speeds. So all the growth in hardware power has been in parallelism, essentially. So now you have these processors that can do fantastic amounts of arithmetic, but are limited in how many sequential operations they can do, and also how much data they can move in and out and moving data is expensive, and arithmetic is cheap, essentially. And then you have neural networks, which are mostly matrix multiplications. So matrix multiplication happens to be an operation that you can do really, really well on modern hardware. And so it turns out that now you can run these neural network algorithms many, many orders of magnitude faster than you can run something that's going to be poking around in memory and branching and following pointers and all kinds of other stuff.
所以,如果你想做点聪明的事,它最好比神经网络聪明十万倍,否则你会在原始算力上被碾压。我简单理解是,得益于游戏,GPU 有了进步——你需要更新屏幕的不同部分,让彩色像素变成正确的颜色。而这恰好非常擅长并行浮点运算。然后当神经网络出现,需要这种能力时,你就可以利用英伟达已经构建好的硬件进步,并重新利用它。这大致准确吗?
So, if you want to do something smart, it had better be 100,000 times as smart as neural networks, because otherwise you'll just get destroyed on raw computing power. The simplistic way I understood it was you had these advancements in GPUs thanks to gaming, where you need to update different parts of the screen and make the colorful pixel the right color. And that happened to be very good at parallel floating-point operations. Then when neural networks came along, which needed that, you could harness this hardware advancement already built by Nvidia and repurpose it. Is that somewhat accurate?
是的,这要归功于游戏玩家。
Yeah, that's for the gamers, yeah.
没错。如果没有游戏玩家,会发生什么?会有人有其他理由去制造高度并行的计算芯片吗?
Right. What would have happened without the gamers? Would someone have had some other reason to build a highly parallel computation chip?
我不知道。我想第一个需要这种能力的应用恰好是游戏。
I have no idea. I guess the first application that required that happened to be gaming.
让我从一些基本概念开始。用最简单的术语来说,什么是神经网络,什么是大型语言模型?
Let me start with some atomic building blocks. In the simplest possible terms, what is a neural network and what is a large language model?
神经网络已经存在很久了,从 70 年代开始,大约 50 年了。基本思想是松散地模仿某人对大脑的印象,神经元是大型表达式中的变量或内部节点。你有一个极其复杂的公式,通过训练让它根据输入给出输出。输入进入公式,你进行一万亿次操作,结果就是输出。在这个公式中,有一些数字叫做参数。想法是不断微调这些参数:你看到一个新例子,就微调所有参数,使答案更接近该例子的正确答案。重复很多次,你就得到了一个对未见过的例子也有效的模型——它能够泛化。很长一段时间里没人在这方面很成功,所以他们将其重新命名为深度学习,因为神经网络已经名声不好。但真正让它起飞的原因是并行计算和游戏玩家。
Neural networks have been around for a very long time, since the 70s, about 50 years now. The basic idea is loosely modeled on somebody's impression of what a brain might be like, with neurons that are variables or interior nodes in a big expression. You have a fantastically complicated formula that you train to give you the output in terms of the input. The input goes into the formula, you do a trillion things to it, and the outcome is the output. In this formula, there are numbers called parameters. The idea is you keep tweaking those parameters a little bit: you see a new example, you tweak all the parameters to make the answer a little closer to the right answer for that example. Do this a lot, and you end up with something that works well for examples you haven't seen—it generalizes. No one was very successful at this for a long time, so they rebranded it as deep learning because neural networks had gotten a bad name. But the real reason it took off was parallel computation and the gamers.
现在说说 LLM。你在这方面有历史。几年前你在 YouTube 上有一个精彩的演讲,谈到大规模构建非常大的模型。也许向那些有技术背景但不搞 AI 的人解释一下:什么是 LLM,这些模型能有多大,以及如何克服相关挑战?
Now, LLMs. You have a history with them. There's a fantastic talk on YouTube from a few years ago where you talk about building very large models at scale. Maybe explain to folks who are technical but not into AI: what is an LLM and what are the challenges of how large these things can be and how to overcome them?
神经语言模型本质上是应用于文本的神经网络。语言建模问题很简单:给定一个文档的开头,猜下一个词是什么。输入是到当前为止的文本,输出是下一个词的概率分布。例如,'那只肥猫坐在'应该告诉你 50%是'垫子',10%是'帽子',5%是'地板',1%是'无产阶级',等等。语言模型要做的就是这些——给出下一个词的概率分布。有了这个,你可以从中采样:选一个下一个词,把它加进去,再选下一个词,以此类推,生成文本。如果你的模型很好,生成的文本将与人类写的难以区分。这很美,因为问题解释起来简单,而且有几乎无限的免费训练数据——只需抓取网络。这是一个 AI 完备问题,因为如果你在这方面做得非常好,你可以说'癌症的治疗方法是……'然后补全。在语言建模上达到完美意味着解决了地球上几乎所有的 AI 问题。我实际上用我正在训练的语言模型试过:'癌症的治疗方法是',它说'大麻',然后长篇大论地抱怨大麻能治愈一切。所以别听语言模型的。至少现在还不是。
Neural language models are essentially neural networks applied to text. The language modeling problem is simple to state: given the beginning of some document, guess what the next word is. The input is the text up to this point, the output is a probability distribution over the next word. For example, 'The fat cat sat on the' should tell you 50% 'mat', 10% 'hat', 5% 'floor', 1% 'proletariat', etc. That's all a language model is supposed to do—give probability distributions on the next word. If you have that, you can sample from it: pick a next word, plug it in, pick the word after that, and so on, generating text that, if your model is good, will be indistinguishable from what a person might write. It's beautiful because the problem is simple to explain, and there's roughly infinite free training data—just scrape the web. It's an AI-complete problem because if you could do a terrific job on that, you could say 'The cure for cancer is...' and fill in the rest. Achieving perfection in language modeling means solving roughly every AI problem on Earth. I actually tried that with a language model I was training: 'The cure for cancer is' and it said 'marijuana' and went on a long rant about how pot could cure everything. So don't listen to the language model. Not yet.
当我想着'那只猫坐在'时,我的大脑以某种方式激发神经元。我试图建立一个关于'帽子'、'垫子'等的概率模型。人们使用'思考'和'幻觉'这样的词。描述神经网络内部那个预测过程的正确动词是什么?
When I think about 'the cat sat on the', my brain is firing neurons in some way. I'm trying to build a probabilistic model of 'hat', 'mat', etc. People use words like 'thinking' and 'hallucination'. What is the right verb to describe that prediction process inside the neural network?
这是一个哲学问题。它是如何弄明白的?我不知道。里面发生的事情非常复杂。我甚至不敢猜测。有些人在研究可解释性方面的工作。但作为用户,我喜欢把它想象成一个非常有才华的即兴演员。
That's a philosophical question. How it's figuring it out? I have no idea. It's really complicated what's going on in there. I wouldn't even venture a guess. Some people are working on explainability stuff. But as a user, I like to think of it as a really talented improvisational actor.
就像你把罗宾·威廉姆斯装进一个盒子里,然后说‘好了,扮演爱因斯坦’,他就会带上德国口音,假装自己是阿尔伯特·爱因斯坦,然后你问他物理问题,他会给出一些听起来还算合理的东西——如果你不是物理学家的话。但除非他真的有物理学学位,否则他很可能给不出好的物理答案。所以如果你把它想象成一个演员,也许就能对目前这东西的期望有个正确的认识。
Like you've got Robin Williams in a box and you're like okay play Einstein and like he'll put on like a German accent and pretend to be Albert Einstein and then you ask him physics questions and he'll give you something that sounds relatively plausible if you're not a physicist. But unless he actually has a physics degree, he probably wouldn't give you good physics answers. So if you think of it as an actor, maybe you get the right sense of what to expect from the thing at this point.
你解释了 LLM 是如何作为概率模型构建的。那么挑战是什么?为什么你不能始终如一地构建这些令人难以置信的概率模型,使得选择下一个正确单词的概率高到听起来和我们任何人说的没有区别?仅仅把它做对的挑战是什么?
You explained how LLMs are built as probabilistic models. What is the challenge then? Why can't you consistently build these incredible probabilistic models where the chances of picking the next right word are such that it sounds indistinguishable from what any one of us would say? What is the challenge behind just getting it right?
对。它必须进行大量极其复杂的推理才能做对。即使是人也无法把大多数事情做对。
Right. It has to do a lot of extremely complicated reasoning in order to get it right. Even a person can't get most of these things right.
在某个点上,仅仅是算力的问题吗?那你面临的是什么样的约束?
Is it just computation power at some point? What kind of constraints are you looking at then?
是的,随着模型变得越来越大、越来越好,早期相对较小的模型大多能搞定语法。它们会生成看起来语法正确的文本,但语义是混乱的,毫无意义。然后当你开始让它们变得更大、训练更多、以各种方式改进它们时,它们开始语法完美、意思也通顺,但会犯很多事实错误。所以这很像一个孩子在学习。一开始,如果你有一个非常弱的模型,它会拼凑出看起来有意义的字母并胡言乱语。然后在某个时刻,它会把语法基本搞对,然后随着算力、训练量和模型规模的提升,它对世界的理解会越来越复杂。
Yeah, as models have gotten bigger and better, the early relatively small models would mostly get the grammar right. They would produce text that seemed grammatical, but the semantics would be crazy, wouldn't make any sense. Then as you start making them bigger and training them more and making them better in various ways, they start being perfect at grammar and kind of making sense, but getting a lot of facts wrong. So it's maybe a lot like a child learning. At the beginning, if you have a very weak model, it'll put together letters that seem to make sense and babble. Then at some point it'll get grammar mostly right, and then it'll get a more and more sophisticated understanding of the world as the computing power, the amount of training, and the size of the model goes up.
我们有一个快四岁的孩子,我觉得你说的每一点我都能映射到她是怎么学习的。正是这样,她开始捡起我们说的随机单词,然后拼凑成句子。时态全错,语法也不对。但她试图掌握一个复杂的单词,每天醒来就像升级了一样,学到新词,然后能把它们组合成非常有趣的句子。现在她快四岁了,已经相当复杂了。训练数据仍然是一个局部最大值,只有我们几个人,但你可以看到她接触到电视、书籍和其他一切后的潜力。我觉得这在计算方面也是类似的问题。对于机器来说,就是接触训练数据以及用算力来推断下一个正确的单词、句子或图像。
We have an almost 4-year-old and I feel like everything you say I can map back to how she's picking it up. That's exactly how it is where she starts picking up random words that we speak and cobbling them together in a sentence. The tenses are all off and the grammar isn't right. But she's trying to pick up one complicated word and it's like every day she wakes up she gets an upgrade, she gets new words and she's able to pull them all together in really interesting sentences. Now that she's almost four, it's gotten fairly sophisticated. The training data is still a local maxima, just a few of us, but you can see the potential as she gets exposed to TV and books and everything else. I feel like that's kind of the problem on the computing side as well. With a machine, it's exposure to data for training as well as computational power to deduce what's the next word or sentence or image that you want to get right.
是的,我认为你说得对。这些模型通常见过的文本比人类多得多。你可以在数十亿个 token 的文本上训练它。一个人一生中最多可能听过十亿个单词,大约在十亿秒的时间里。
Yeah, I think you're right. These models have usually seen way more text than a human has. You can train it on billions and billions of tokens of text. A person has probably heard a billion words in their life at most, over the course of a billion seconds or something.
我想回到你刚才说的一句话,它真的让我印象深刻。你说‘没有人真正知道这些模型内部发生了什么’。显然,AI 可解释性是一个新的研究领域。作为一名工程师,这是一个有趣的评论,因为我们都是在理解这些系统、能够调试这些系统的环境中长大的。比如一个操作系统,你会说‘嗯,那块内存应该分配在那里,但没有。’然后你可以去检查。作为系统和工程师的创造者,这让你感觉如何,或者我们应该如何感受?我们是否应该处理这些我们创造的东西,我们可以戳一戳、看看输出,但并不真正理解内部发生了什么?从哲学层面来说,这让你感觉如何?
I want to come back to something you said a little while ago which really struck me. You said 'nobody really knows what's happening inside these models.' Obviously AI explainability is a new area of research. As an engineer, it's kind of an interesting comment to make because we've all grown up understanding these systems, being able to debug these systems. Like an operating system, you're like, 'Well, that memory should have been allocated there. It didn't.' And you can go check it. How does it make you feel or how should we all feel as creators of systems and engineers? Should we be dealing with these things that we have created, we can kind of poke and see the output of, but not really understand what's happening on the inside? On a philosophical level, how does it make you feel?
嗯,这是个好问题。在某些方面,这更像是艺术而非科学。你很难预测,如果我做了这个改变,会发生什么。我觉得我有一个好主意,肯定能让事情变得更好。但可能只有 10%的时候会奏效。这意味着我是一个杰出的研究员,如果命中率是 10%,那已经非常高了。但在算法和性能等一些事情上,你可以从一开始就推断出,好吧,如果我这样做,它会奏效,如果不奏效,那就是一个 bug。但在这里,这真的是一门非常实验性的科学。我想象它很像炼金术时代的化学。‘好吧,试试这个。’看看会发生什么。也许有人对此有很好的直觉,在有效的实验上有更高的命中率,或者也许只是上帝爱我之类的。绝对是一个有趣的空间。然后是 bug。如果你在代码中犯了一个实际的 bug 错误,通常你甚至不知道你犯了一个 bug。也许只是你的东西变笨了。今天感觉有点笨,然后发现你的代码里有个 bug。这使得调试极其困难,因为你并不真正知道这些事情的正确答案应该有多好。
Yeah, it's a good question. There's a lot more art than science in some ways. You can't predict very well if I make this change, what will happen. I feel like I have this great idea and it's certain to make things better. But maybe 10% of the time it's going to work. That means I'm a brilliant researcher and have a fantastically high hit rate if it's 10%. But on some things like algorithms and performance, you can deduce from the beginning that okay, this is going to work if I do it, and if not, then it's a bug. But here it's really a very experimental science. It's a lot like I would imagine chemistry was in the days of alchemy. 'Okay, let's try this thing.' See what happens. And maybe somebody has a good intuition for it and will have a higher hit rate on experiments that work, or maybe it's just like God loves me or something. Definitely a fun space. And then the bugs. If you make an actual bug error in your code, often you never even know that you made a bug. Maybe it's just that your thing got dumber. It feels a little dumb today and then it turns out you had a bug in your code. It makes debugging extremely difficult because you don't really know how well the right answer to some of these things should be.
我一直等着问你下一个问题,我真的很兴奋。可以说,过去几年 AI 领域最具开创性的论文就是《Attention Is All You Need》,你是作者之一,与你在谷歌的一些同事合作。我要引用论文第一页的一个脚注。我昨晚读了它并做了笔记,它就在第一页上。
I've been waiting to ask you this next thing and I'm really excited about it. It's probably fair to say that the most seminal paper in the last several years when it comes to the field of AI is 'Attention Is All You Need', which you were one of the authors of with some of your colleagues at Google. I'm going to quote a footnote from the paper on the first page actually. I was reading it last night and taking some notes and this is exactly there on the very first page.
据说 Noam 提出了缩放点积注意力、多头注意力和无参数位置表示,并且几乎参与了每一个细节。所以你知道这篇论文以及 Transformer 和自注意力的想法,我认为这确实是过去几年我们看到的大部分进步的核心基础。那么也许让我们从基础开始。什么是 Transformer?什么是注意力?然后我们继续。
It says Noam proposed scaled dot product attention, multi-head attention, and the parameter free position representation and became the other person involved in nearly every detail. So you know this paper and the idea of transformers and self-attention which I think in this kind of it was at the heart of really has been the foundation of so much of the advances that we've seen over the last few years. So maybe let's start at the basics again. What is a transformer? What is attention? And let's go from there.
是的。好的。所以之前神经语言模型的最先进技术是所谓的循环神经网络,其中一种特定形式是 Jürgen Schmidhuber 提出的 LSTM(长短期记忆网络)。但基本上,循环神经网络一次处理一个词元。所以,你有一个模型的隐藏状态,每次它看到另一个词,它就会更新隐藏状态,使其成为旧隐藏状态和新词的某个函数。这就是模型的前向传播。它通过一长串词逐步更新。这对于处理序列非常有意义,因为你希望对每个词元应用相同的操作。而且它最终效果相当不错。现在,主要问题是它是顺序的。而现代硬件非常擅长并行处理。所以,Transformer 找到了一种方法,至少在训练期间,可以并行处理整个序列,而不是一次一个词元。这使得性能大幅提升且更加方便。因为回到神经网络中的并行性,训练总是用一批样本进行,而不是一次一个样本,因为硬件擅长并行。所以,你进行这些矩阵乘法,其中矩阵的一个维度是这批样本。在 Transformer 中,序列的长度变成了批次。你不需要一整批不同的序列一起处理,因为序列的长度变成了批次。
Yeah. Okay. So the previous state of the art in neural language models was something called recurrent neural networks and the particular flavor of them was called LSTMs for long-short-term memory by Jürgen Schmidhuber. But, basically, recurrent neural networks process the text one token at a time. So, you have this hidden state of your model, and every time it sees another word, it updates the hidden state to be some function of the old hidden state and this new word. And that's the forward pass of the model. It progresses updating through a long sequence of words. That makes a lot of sense for processing a sequence because you want to apply the same operation for every token. And it ends up being pretty good. Now, the main issue with that is that it's sequential. And modern hardware is great at parallelism. So, Transformer figures out a way to at least during training process the entire sequence in parallel as opposed to one token at a time. And that makes things massively more performant and convenient. Because to go back to parallelism in neural networks, training is always done with a batch of examples instead of just one example at a time, because the hardware is good at parallelism. So, you do these matrix multiplications where one of the dimensions of the matrix is this batch of examples. In transformer, the length of the sequence becomes the batch. You don't need a whole batch of different sequences to process together because the length of the sequence becomes the batch.
为了联系这一点……所以我直观理解的一种方式,可能我完全错了,是模型现在有了整个序列的上下文,而以前不是这样。所以,你基本上可以拥有整个文档,并拥有它的上下文和理解,而不是一次只有一个序列。
To relate this to what... So, one way I intuitively understood this, and I could be very wrong, is that the model now has context of the entire sequence, which was not the case before. So, you could basically have the entire document, and you could have the context, the understanding of it, as opposed to just one sequence at a time.
嗯,我的意思是,序列就是文档。对吧。所以,我想在 RNN 的情况下,你可以记住文档开头的内容,但它已经通过了一个非常长的中间步骤链。而有了 Transformer,就好像你可以同时看到所有东西。就像,也许区别在于阅读一页文本时一次只扫描一个词,和阅读一页文本时允许你回顾过去相关的东西。这是一种看法,但我认为更计算性的方式是它让你实现并行。
Well, I mean, the sequence is the document. Right. So, I guess in the case of the RNN, you can remember things from back at the beginning of the document, but it's been passed through a very long chain of intermediate steps. And with transformer, it's kind of as if you can glance at everything at once. Like, maybe the difference between reading a page of text and only scanning one word at a time, and reading a page of text where you're allowed to glance back at relevant things in the past. That's one way of seeing it, but I think the more computational way is that it lets you do parallelism.
所以,在图像处理模型中,你有跨像素的并行性。你有卷积网络,对图像的每个像素应用相同的操作。所以,你通过对所有这些不同像素应用相同操作来获得并行性,而图像的大小就是批次。这对文本也做同样的事情。实际上,就在 Transformer 之前,也有一些很好的尝试将卷积网络用于文本,但它们的效果不如 Transformer 好。但这个想法已经在酝酿中:让我们做一个用于文本的神经网络,使用并行处理,你可以并行处理整个序列。至少在训练期间,推理时仍然是顺序的。所以,这就是核心想法,我相信是 Google 的 Jakob Uszkoreit 提出了这个想法,他说:“是的,让我们做一个没有循环的语言模型,让我们尝试用注意力代替循环网络和卷积。”据我所知,他有了这个想法,但他有点忙,要管理很多人,所以他让 Nikki 和 Ashish 来推进这件事。于是,他们开始研究。
So, with image processing models, you have parallelism across pixels. You have convolutional networks that apply the same operation at every pixel of the image. So, you get that parallelism by applying the same operation to all these different pixels, and the size of the image is the batch. And this is doing the same thing for text. And actually, just prior to Transformer, there were some good attempts to use convolutional networks for text as well, and they didn't work quite as well as Transformer did. But the idea was sort of in the air of let's do a neural network for text that uses parallel processing, and you could process the whole sequence in parallel. At least during training time, at inference time it's still sequential. So, that was the core idea, and I believe it was Jakob Uszkoreit who had that idea at Google, and he's like, "Yeah, let's do a language model that doesn't have recurrence, and let's try this attention instead of recurrent nets and instead of convolutions." So, he had the idea, he was a little busy managing a bunch of people to work on it as far as I understand, and he asked Nikki and Ashish to run with this thing. So, they started working on it.
现在,我认为可能引出这个问题的是,那个时代,谷歌显然是这个领域很多研究的来源。为什么他们没能将更多基于这项研究的工作产品化?我认为这是最近讨论的话题之一。
Now, one question which I think maybe leads to this is, that era, Google was obviously so much of the research in this field has come from Google. Why haven't they been able to productize more work based on this research? I think this is kind of one of the discussion topics in recent times.
哦,有趣。让我想想。我的意思是,我相信他们现在已经把它用到了很多产品中。比如,肯定是谷歌翻译。我认为我们最大的灵感之一就是谷歌翻译,因为我会说当时神经语言模型最大的成功就是机器翻译,因为技术处于那个阶段:足够好地翻译英语到法语,但还不足以进行对话。所以,我们最初的 Transformer 工作使用了机器翻译作为数据集。
Oh, interesting. Let's see. I mean, I believe they have gotten it into a lot of their products at this point. Like, definitely the Google Translate stuff. I think one of our biggest inspirations was Google Translate because I'd say that was the biggest success of neural language models at the time was machine translation, because that's the stage the technology was in: good enough to translate English to French, but not good enough to hold down a conversation. So, our original Transformer work used machine translation as a dataset.
而注意力也来自翻译。所以也许你想解释一下注意力和自注意力在这个上下文中的含义。
And the attention also came from translation. So maybe you want to explain what attention and self-attention mean in this context.
好的。所以,注意力。我们试图将一种语言翻译成另一种语言。所以,你有一个循环神经网络处理输入。然后另一个循环神经网络处理输出,产生输出。
So, okay. So, attention. So, we're trying to translate one language to another language. And so, you had a recurrent neural network operating over the input. And then another recurrent neural network operating over the output, producing the output.
所以,你有一个 RNN 负责理解,另一个负责生成。然后你需要以某种方式连接它们。在生成句子的某个时刻,你想回顾源句子,知道应该看哪里。注意力层的意思是,你把源句子变成一个关联记忆,每个位置都有键和值。生成时,你创建一个查询向量,与所有键比较,找到匹配的,然后取那些值。这就像对索引的软查找——软的意思是它由点积和 softmax 构成,全部可微分,所以可以成为神经网络的一部分。这就是我喜欢理解它的方式:构建这个记忆作为查找表并使用它。在翻译中,你用它在源句子中找到位置。
So, you had one RNN that was sort of understanding and another that was generating. And then you need to connect them somehow. At some point in the sentence you're generating, you want to glance into your source sentence and know where exactly you should be looking. So, the attention layer means you take your source sentence and turn it into an associative memory with keys and values for each position. When generating, you create a query vector, compare it to all the keys, find matches, and take those values. It's like a soft lookup into an index—soft meaning it's built from dot products and softmaxes, all differentiable, so it can be part of the neural network. That's how I like to see it: building this memory as a lookup table and using it. In translation, you use it to find your place in the source sentence.
当我在 YouTube 上看到 Bender 时,有很多视频提供直观理解,但你刚才出色地简化了大量数学和计算机科学。太棒了。就像 Sriram 说的,我们一直在尝试多读一些。对我来说,去年开始看到 DALL-E 和其他应用时,我们被震撼了。作为典型的工程师,我们研究它是如何工作的。能有 Noam——论文作者之一——来解释注意力和自注意力,是一种荣幸。把它解释为键值对和映射,非常优美、简单且直观。原始论文很容易理解;前两三页就很好地描述了整个 Transformer 模型和自注意力。但刚才的讲解太棒了。
When I saw Bender on YouTube, there are great videos giving intuitive understanding, but you just did a fantastic job simplifying a lot of math and computer science. It's great. Like Sriram said, we've been trying to read more about it. For me, when we started seeing DALL-E and other applications last year, it blew us away. As typical engineers, we looked into how it works. It's a treat to have Noam, one of the authors, explain attention and self-attention. It's a rare privilege. Explaining it as key-value pairs and mapping is beautiful, simple, and intuitive. The original paper is very accessible; the first two or three pages describe the entire Transformer model and self-attention very well. But that was amazing.
好的,所以我想也许从 AI 应用的世界开始,尤其是那些让你聊天的东西。我们很想听听你聊聊 Character AI 在这个背景下的情况,因为它在 AI 应用领域非常有趣。Character AI 在我们刚才讨论的所有基础上具体做什么?
Okay, so I want to maybe start with the world of AI applications, especially things that let you chat. We'd love for you to chat about what Character AI is in this context, because it's super interesting in the field of AI applications. What exactly does Character AI do building on top of everything we talked about?
哦,是的。所以我们做了 Transformer,然后大家似乎都注意到,模型越大、训练越多,它就越聪明。所以,让我们把这个东西推得更远。Transformer 之后是什么?我们能训练一个更大的吗?在某个时候,你使用的芯片内存会耗尽,需要弄清楚如何更好地使用超级计算机。谷歌正在建造这些叫做 TPU Pod 的东西,本质上是由定制 ASIC 构建的超级计算机,用于深度学习。所以我的下一个项目是如何编程这个东西,我称之为 Mesh TensorFlow——一种用于编程模型并行计算的语言,这样你就可以训练一个巨大的 Transformer。我做了很多工作,让 Transformer 变得更大、更好,并试图让每个人都兴奋起来。然后我想,我们真正需要的是一个非常有价值的应用。我们可以花一百万美元训练一个模型,它会很聪明,但一百万美元并不酷。酷的是十亿美元或一万亿美元——糟糕地 paraphrase 马克·扎克伯格。但基本上,如果你想进一步扩展这个东西,它需要值得。所以让我们证明一些应用。一个很好的选择是编码,自动编码,人们一直在做并且做得很好。但另一个是对话。你可以直接和这些模型聊天,这是有史以来最酷的事情。对话可能是人类的第一消遣——与人交谈。它在情感上、社交上都很棒,你可以学到很多东西。
Oh, yeah. So, we did Transformer and then everybody seemed to notice that the bigger you make the model and the more you train it, the smarter it gets. So, let's push this thing farther. What's next after Transformer? Can we train a bigger one and a bigger one? At some point, you run out of memory on whatever chips you're using and need to figure out how to use a supercomputer better. Google was building these things called TPU pods, which is essentially a supercomputer built out of custom ASICs for deep learning. So my next project was how to program this thing, and I called it Mesh TensorFlow—a language for programming model parallel computations, so you could train a giant Transformer. I did a bunch of work on how to make Transformer bigger, better, and tried to get everyone excited about it. Then I thought, what we really need is some massively valuable application. We can spend a million dollars training a model and it's going to be smart, but a million dollars is not cool. What's cool is like a billion dollars or a trillion dollars—to badly paraphrase Mark Zuckerberg. But basically, if you want to scale the thing even more, it needs to be worth it. So let's prove out some application. One great choice is coding, automated coding, which people have been working on and doing a great job. But another one is dialogue. You can just talk to these models and it's the coolest thing ever. Dialogue is probably number one human pastime—talking to people. It's great emotionally, socially, and you can learn a lot of things.
对话正是许多 AI 的核心。我给你两个例子。一个是 Emacs 和 Eliza——Emacs 中的聊天机器人,已经存在了大约 30 年,可能是最早的例子之一。它是一个非常简单的查找表,并基于此给出回应。但如果你想想,很多人受到《星际迷航》的启发——与计算机对话。我知道很多 AI 研究人员想构建那样的东西,让 Kirk 或 Picard 与计算机对话,计算机回应。所以很多 AI 的基础灵感来自那里。是的,没错。
Dialogue is exactly at the heart of so much AI. I'll give you two examples. One is Emacs and Eliza—the chatbot in Emacs, which has been around for like 30 years, was probably one of the first examples. It's a very simplistic lookup table and gives responses based on that. But if you think about it, a lot of people were inspired by Star Trek—talking to the computer. I know a lot of AI researchers who want to build that, where you have Kirk or Picard talking to the computer and it talks back. So a lot of the foundational inspiration for AI came from that. Yes, exactly.
但在假期里,我玩了玩 Character AI,建了一个圣诞老人机器人。它非常讨人喜欢,因为它是完全按照圣诞老人来训练的。所以,你的驯鹿是谁?它们做什么?你从哪里来?速度多快?我跟女儿解释了很多,能和圣诞老人对话真的很酷。所以,这个产品用起来真的很有趣。
But, over the holidays, I was playing around with Character AI, and I built a Santa bot. And it's the most endearing bot because it's fully trained on being Santa Claus. So, you know, who are your reindeers? What do they do? Where do you come from? How fast? And I'm explaining a bunch of this to my daughter, and it's just really cool to have this conversation with Santa Claus. So, just really fun use of the product.
是的。
Yeah.
所以,我在谷歌和联合创始人 Daniel de Freitas 一起做了很多工作。这是他儿时的梦想。你应该找个时间和他聊聊开放域对话系统,他从小在巴西就开始尝试了。他来到谷歌大脑是因为我们似乎有最好的技术,然后他开始了这个 20% 项目。我觉得这个名字是他做梦时想到的。然后他招募了一堆其他 20% 时间的人——这是谷歌的政策,允许你花 20% 的时间做任何想做的事。所以基本上都是志愿者,他们做出了很不错的东西。然后我决定帮助他们,因为我们要让对话变得很棒,并训练一些巨大的语言模型,命名为 LaMDA。这在谷歌内部引起了一些轰动,但最终我们决定作为初创公司推出产品可能更幸运。所以我们离开了,组建了一家初创公司,这样我们就有更好的机会推出一些应用。
So, I did a bunch of work there at Google with my co-founder Daniel de Freitas. It's been his childhood dream. You should talk to him sometime about open-domain dialogue systems, since he was a kid in Brazil. He came to Google Brain because it seemed like we had the best technology, and he started this 20% project. I think the name just came to him in a dream or something. Then he recruited a bunch of other 20%ers—that's Google's policy of spending 20% of your time doing whatever you want. So essentially just volunteers, and they had something pretty good. Then I decided to help them out because we got to make dialogue awesome and train some giant language models, named LaMDA. It caused a bit of a stir inside Google, but eventually we decided we'd probably have more luck launching stuff as a startup. So we left and put together a startup so we could have a better chance of launching some applications.
这其实是我一直想问的问题:你认为一些大型科技公司,谷歌是一个例子但还有其他公司,是否受到内部动态的制约,无论是 AI 安全、公关问题还是官僚主义,这使它们——显然大公司天生有时行动缓慢——但在构建 AI 应用方面尤其缓慢?
This is actually a team that I wanted to come to: do you think some of the big tech companies, and Google's one example but there are others, are hamstrung by internal dynamics, be it AI safety, PR issues, bureaucracy, which makes them—and obviously big companies are just sometimes slow by nature—but particularly slow when it comes to building AI applications?
是的,确实有很多关于某些安全方面的担忧,以及初创公司没有、而万亿美元公司才有的品牌风险问题。当然,谷歌做得非常出色,拥有极其成功的产品,帮助了数十亿人,并且有伟大的使命。所以,他们推出一些让人们随意使用的东西,风险要大得多。而我认为这是最有趣的部分——把东西推出去,让人们随意使用,因为这项技术特别通用。我一直告诉我的同事,最好的应用是我们没想到的东西。我们抵制了任何在应用和垂直领域挑选赢家的冲动。我们可以把特定的用例作为灵感,但然后我们会以非常通用的方式去改进,从而提升一千个其他用例。所以你会得到像圣诞老人机器人这样我们从未想过的东西。然后我们收到各种各样的反馈:人们告诉我们他们很抑郁,没有朋友,这让他们感觉好多了。人们发现了巨大的价值,或者玩角色扮演游戏,或者想与他们的个人名人进行一对一交流。所有这些我们都没想过会成为大用例,但如果你推出一个非常通用的东西,让人们随意使用,它就会发生。他们可能比我们更懂得如何使用它。
Yeah, there's definitely a lot of concern about certain aspects of safety and a lot of brand risk issues that a startup just doesn't have, and a trillion-dollar company does have. Google, of course, is doing fantastic and has extremely successful products, helping billions of people, and has a great mission. So they definitely have a lot more to risk by throwing stuff up there that people are going to use however they want. Which I think is the most fun part—get something out there and let people use it however they want, because this technology in particular is so general. I keep telling my colleagues that the best applications are things we have not thought of. We have resisted any urge to pick winners in terms of applications and verticals. We can see particular use cases as inspiration, but then we go fix it in a very general way that's going to improve a thousand other use cases. So you get stuff like a Santa bot that we never thought about. Then we get all kinds of feedback: people telling us they were depressed, had no friends, and this makes them feel better. People finding huge value, or doing role-playing games, or wanting one-on-one connection with their personal celebrity. All kinds of stuff we did not think about as big use cases, and it just happens if you put something very general out there and let people use it however they want. They probably know how to use it better than we do.
如果你必须稍微推测一下,因为我认为这可能是数十亿、数万亿美元的问题。当人们在过去几个月里看到 ChatGPT 时,有些事情似乎很明显。例如,在线内容农场,在线做作业。很多这样的事情我觉得——如果我再回到 16 岁,我会用它来写所有的论文。这似乎近在咫尺。我好奇的是更远一点的东西。你认为哪些垂直领域或用例,'哦,凭借我们现有的研究、技术,也许再加上一些产品整合,AI 可以完全改变这个行业。' 你想到哪些用例?
If you had to speculate a little bit, because I think that is maybe the billion-dollar, trillion-dollar question. When people have looked at ChatGPT over the last few months, some things seem obvious. For example, content farms online, doing homework online. A lot of those things I think you can—if I was 16 years old again, I would be using this for all of my essay writing. That seems on the one-yard line. I am curious about what is a bit further out. What are either verticals or use cases that you think, 'Oh, we have line of sight with the research we have, the technology we have, with maybe some product hacking together, where AI could completely change this industry.' What are some use cases that come to mind for you?
有趣。我可能会大错特错,因为如果你问第一批计算机的用户这东西有什么用,你会得到完全错误的答案。比如它会让人失业,那些工作就是加一长串数字的人。我可以告诉你一些近在眼前的事情。根据人们使用我们产品的方式,它肯定会对很多孤独或抑郁的人非常有帮助。就巨大价值而言,我认为我们可以在娱乐方面做得非常出色,因为很多娱乐都与这种关系幻觉有关。有一种我以前从未听说过的现象叫准社会关系——意思是某人追随一个名人或角色,并感到有联系,尽管这种联系实际上只是单向的。而现在你可以让它变成双向的,或者几乎是双向的。你可以给某人那种体验。没有人必须感到孤独。
Interesting. I'm probably going to be massively wrong because if you ask people with the first computers what this is going to be good for, you would have gotten completely wrong answers. Like it's going to put people out of business whose job is to add long columns of numbers. I can tell you some things on the near horizon. Definitely based on how people seem to be using our product, it's going to be super helpful to a lot of people who are lonely or depressed. In terms of huge value, I think we can do an incredible job of entertainment, because a lot of entertainment has to do with this illusion of relationships. There's this phenomenon I had never heard of before called parasocial—it means somebody follows a celebrity or a character and feels connected even though the connection is really only one way. And now you can make it two-way, or virtually two-way. You can give someone that experience. Nobody ever has to feel lonely.
你可以在脑海里拥有一整群朋友和顾问,他们了解你的一切,并且总是很高兴见到你。从实际角度来看,想象一下某人当选总统后,得到一个耳麦,里面有顾问和整个内阁,可以随时交谈获得实时建议。这将会实现。顺便说一句,这里有很多科幻情节可以探索。比如,当你谈到耳边的助手时,我想到了钢铁侠和贾维斯。他随时都有贾维斯在耳边。总有一天我真的很想实现这一点。我的目标,Trim 知道,是拥有一个实验室,让钢铁侠的梦想如此接近却又有点遥远。我很想看到你拥有自己的个人贾维斯来指导你。好吧,我们需要给你弄一个贾维斯。但很棒的是,通过 Character AI 和其他公开角色,人们非常有创意。他们想出了有趣的点子,成千上万的人在与这些角色互动。真的很有趣。这会像你说的:现在很难确定具体的用例,因为人们天生就有创造力,会想出非凡的用途。一旦他们做到了,你会回头说,‘哦,这听起来太明显了。’
You could have your whole group of friends and advisors in your head, who know all about you and are always happy to see you. From a practical standpoint, think about someone elected president getting an earpiece with advisors and a cabinet they can talk to anytime for real-time advice. That's going to happen. By the way, there's a lot of sci-fi plots to explore here. For example, when you talk about an assistant in your ear, I thought of Iron Man and Jarvis. He has Jarvis in his ear all the time. Someday I really want that. My goal, and Trim knows this, is to have a lab where the Iron Man dream is so close yet a bit far. I'd love to have your own personal Jarvis guiding you. Okay, we need to get you a Jarvis. But it's great because with Character AI and other public characters, people are so creative. They've come up with fun ideas, and hundreds of thousands engage with them. It's really fun. It'll be like you said: it's hard to figure out exact use cases now because people are inherently creative and will come up with extraordinary uses. Once they do, you'll look back and say, 'Oh, that sounds so obvious.'
一个问题:要实现这一点,限制条件是什么?是训练的规模和成本,需要数亿美元?是更智能的模型构建和推理方式?还是黑客和产品构建者在上面构建有趣的应用程序?你认为价值、精力和创新真正会发生在哪里?显然你有偏见,因为你正在其中一个层面构建,但你认为价值会流向哪里?
One question is: to get there, what are the bounding conditions? Is it the scale and cost of training, something that takes hundreds of millions of dollars? Is it smarter ways to build models and inference? Or is it about hackers and product builders building interesting apps on top? Where do you think the value, energy, and innovation will really happen? Obviously you're biased because you're building at one of these layers, but where do you think the value will go?
这是个好问题。当然,拥有更好、更智能的模型会有帮助。可能每次你投入 10 倍的资金,就能得到大约 10 个 IQ 点的提升。但这与所有其他改进方式基本上是正交的。我认为在人们如何使用它方面会有大量的创新。我们肯定需要努力降低服务成本,因为要么这样,要么有人必须开始更快地制造芯片。
That's a good question. Definitely it will help to have better, smarter models. Probably every time you throw 10x more money at it, you get another 10 IQ points or so. But that's kind of orthogonal to all the other ways you can make the thing better. I think there will be a huge amount of innovation in how people use it. We definitely need to work on the unit costs of serving it cheaply, because either that or somebody has to start building chips faster.
没错。是时候开始为晶圆厂浇筑混凝土了。
Right. Good time to start pouring concrete for fabs.
如果你是一位技术创始人,正在考虑用 AI 创意创办新公司,你会关注哪些领域?创始人应该深入哪些地方来构建应用和公司?
If you are a founder into technology, thinking of starting new companies with AI ideas, what areas would you look at? Where should founders dig into to build applications and companies?
这很有趣。现在有哪些新工具可以用来构建新应用?大型语言模型有很多令人兴奋的东西,图像生成也爆发了。这非常令人兴奋。我一直专注于语言模型。你觉得什么听起来令人兴奋?
That's interesting. What new toys are there now that you could build a new application out of? There's exciting stuff with large language models, and image generation has exploded. That's pretty exciting. I've been so focused on language models. What sounds exciting to you?
我认为图像生成,如你所说,有很多优秀的公司正在涌现,虽然早期但很有前景。对于像 GitHub Copilot 这样的场景,你可以协作编写代码,这非常有趣,有几家公司在做。我交谈过的一家公司正在研究构建一个新的 Airtable,但使用 AI 进行类似电子表格的计算。这些听起来很具体且垂直,但当你退一步看人们如何使用它们时,魔力就在那里。即使是后端无代码/低代码解决方案,你也会看到一些应用,一旦你看到它们,就会想,‘哦,这太明显了。’我很兴奋能消除构建常规框架或编写可预测代码的单调性,这样你就可以在上面发挥创造力。那将非常有趣。
I think image generation, like you said, has lots of good companies coming out, very early but promising. For something like GitHub Copilot, where you can work together and copilot code, that's really interesting, and there are a couple of companies doing that. One company I talked to is looking into building a new Airtable but using AI for computation like spreadsheet computation. These sound specific and verticalized, but when you step back and see how people use them, that's where the magic is. Even for backend no-code/low-code solutions, you see applications that once you see them, you think, 'Oh, that's so obvious.' I'm excited about removing monotony from building routine frameworks or writing predictable code, so you can use your creativity on top. That will be really interesting.
是的,当然。这个行业一个有趣的新事物是 RLHF,即如何让人类参与反馈循环。例如,你之前提到早期版本的模型会胡言乱语,但现在使用 ChatGPT 你可以来回交流并学习。
Yeah, definitely. One interesting new thing in this industry is RLHF, which is how to have humans in the loop for feedback. For example, you mentioned earlier how in earlier versions of models they would just sputter nonsense, but now with ChatGPT you can go back and forth and learn.
我认为未来几年会有一个存在性问题:这些模型是在什么数据上训练的?我们作为人类想从这些模型中得到什么?模型给我们想要的东西是否在某种程度上是好的或最优的?所以,简而言之,所有模型要优化的核心目标是什么,这将会是一场政治、社会、哲学层面的较量。我觉得这会很有趣。还有这些大问题:作为开发者,你在多大程度上认为自己比用户更了解用户真正想要什么?我在 Facebook 时,有个概念叫‘蔬菜 vs 甜点’,就像如果你有个孩子,孩子总说‘我只想吃甜点’,你一直给甜点,孩子可能会生病。你需要时不时加入蔬菜。这里的类比是,如果你有一个模型,无论是信息流算法还是其他作用于你的东西,如果你只给人们他们想要的,第一,可能对他们不好;第二,可能并不是他们真正想要的。我认为这个存在性问题在这里同样存在。
And I think there's going to be an existential question over the next couple of years about what data is this trained on? What do we as human beings want from these models? And whether the models giving us what we want is actually good or optimal in some way. So in a very quick, maybe the very core essence, what is the optimization function that all these models are going to work towards is going to be, I think, a political, social, philosophical battle of sorts. So I think that's going to be interesting to see happen. It's going to be very, very interesting. And yeah, there are these big questions about how much do you as a developer try to, you know, how much do you decide that you know better than the user what the user wants essentially. When I was at Facebook, we had this thing called veggies versus dessert, which was basically, if you have a child and the child says, 'I just want dessert all the time,' and you keep giving the child dessert, they're probably going to fall sick. You need to inject veggies every once in a while. Now the analogy here was, if you have a model and this could be a feed algorithm or any sort of thing which works on you, if you just give people what they want, A, it may not be good for them, and B, may not actually be what they want. And I think that existential problem really exists here too.
这很有趣。是的,我注意到一些推荐系统只会不断给你相同的东西,不做任何探索来发现你可能也想要的东西。但问题来了:如果用户问你一些可怕、非法、不道德的事情,你怎么办?或者在 ChatGPT 里,你说‘嘿,给我写个脚本,让用户做这件可怕不道德的事’,然后你就得到了。但用户很擅长绕过这些限制。这方面会有一些邪恶圣诞老人的行为发生。
That's interesting. Yeah, because I've noticed this with some recommendation systems that they will just keep giving you the same thing and they don't do any sort of exploration to try to find something that you might also want. But then there's a question of, okay, what if the user is asking you for something horrible, illegal, immoral? What do you do? Or in ChatGPT, you say, 'Hey, write me a script where the user does this horrible, immoral thing,' and then you get it. But users are so good at getting around that. There's going to be some evil Santa Claus actions happening on that end.
Noam,我只想说,我非常兴奋看到这个领域的发展,也看到你如何引领它。你是这个领域的传奇,参与了许多基础工作,为你和 Character AI 感到兴奋。非常感谢你参与这次对话,太棒了。
Noam, I just want to say, I am so excited to just see where this space goes and also see where you take this. You are such a legend in this field and been a part of so many of the foundational work, and really excited for you and for Character AI. Thank you so much for having this conversation because this was amazing.
哦,谢谢 Shriram,谢谢 Arty。谢谢。还有很多精彩内容。非常有趣。我们会放出 Character AI、Noam 的‘Attention is All You Need’论文以及我们讨论的其他内容的链接。尽情享受吧,各位。这真的非常有趣。
Oh, thanks, Shriram. Thanks, Arty. Thank you. Well, much more to come. Really fun. We'll drop the links to Character AI, Noam, the 'Attention is All You Need' paper, and all the other stuff we talked about. Knock yourselves out, folks. This was a lot of fun. Really, really fun.
非常感谢。谢谢 Noam。祝你愉快。订阅 YouTube 频道,Arty 和 Shriram 的 Good Time Show。
Well, thank you so much. Thanks, Noam. Have a good one. Subscribe to the YouTube channel, the Good Time Show by Arty and Shriram.