From Poker to DeepMind: Mustafa Suleyman on the Origins of AI
打开互动全文版(中英对照 + 朗读 + 问答)→穆斯塔法·苏莱曼分享与戴密斯·哈萨比斯在扑克桌上创立 DeepMind 的故事,并回顾 AI 发展的漫长历程。
Mustafa Suleyman shares the story of founding DeepMind with Demis Hassabis over a poker game, and reflects on the long journey of AI development.
显然,如果我想尽快结婚,你也能帮忙,对吧?所以据说世界上最棒的鱼钩是现成的,只要你能找到愿意嫁给你的女人,Mustafa。创业——哦,你已经搞定了?
Apparently if I fancy getting married anytime soon, you're available for that too, right? So apparently the world's greatest fishing hook is available if you can find a woman who will marry you, Mustafa. Start up—oh, you already got that accomplished?
不,不,我还在挣扎。我完全是单身。所以如果你想让我跟我的创业公司结婚……
No, no, I'm struggling with that. I'm very much single. So if you want to marry me to my startup...
你已经跟你的创业公司结婚了。你筹集了十亿美元。我可以告诉你未来十年你跟谁结婚。绝对是 Inflection AI。这周你们有 40 个人在那儿。本期《本周创业》由 OpenPhone、Crowdbotics 和 Carta 赞助。今天我们有重磅嘉宾,Mustafa Suleyman 来了。他在 Inflection AI,但更出名的是作为 DeepMind 的联合创始人。欢迎来到节目,Mustafa。
You are married to your startup. You raised a billion dollars. I can tell you who you're married to for the next 10 years. Absolutely, Inflection AI. And you're 40 people over there this week. And startups is brought to you by OpenPhone, Crowdbotics, and Carta. All right, we got a big treat for you today on This Week in Startups. Mustafa Suleyman is here. He's with Inflection AI, but very famous for having been the co-founder of DeepMind. Welcome to the program, Mustafa.
很高兴来到这里。谢谢你,JC。感谢邀请。
Great to be here. Thank you, JC. Thanks for having me.
当然,当然。你知道,我想从 DeepMind 的起源开始,因为我们今天在 AI 领域看到的很多东西都站在那个组织的肩膀上。我觉得大多数人并不了解它的历史。我碰巧知道一点,因为我记得 Peter Thiel 和 Elon 好像是早期的投资者,我们聊过,后来也在一些行业活动上见过面。告诉我,DeepMind 是怎么起源的?然后它是如何开始处理 AI、通用 AI、垂直 AI 这些不同领域,并最终取得成果的?我猜那是 2010 年?2011 年?
Of course, of course. You know, I wanted to start with the origins of DeepMind because it seems like so much of what we're seeing in AI stands on the shoulders of that organization. And I don't think most people know the history of it. I happen to know a little bit of the history of it because I remember when Peter Thiel and Elon, I think, were two of the early funders of it, and we were talking about it, and we met at a couple of different industry events over time. Tell me, what was the origin of DeepMind? And then how did it originate and start to tackle AI, general AI, vertical AI, all these different things that are coming to fruition? And I guess that was 2010, right? 2011?
我们是 2010 年创办公司的。没错,这听起来有点疯狂,差不多 15 年前了。现在看到这一切感觉很超现实,因为过去 9 到 12 个月里,大型语言模型革命好像突然冒出来并爆发了。但实际上,这是多年稳步推进的结果,经历了大量失败、风险和坚持,我觉得这些在新技术的完美爆发故事中常常被忽略。事实上,过去十年的大部分时间里,我们都没有语言模型。Transformer 真正流行起来是在 2017 年。人们常说它是在那时发明的,但绝对不是;它早在 15 到 20 年前就被 Yoshua Bengio 发明了,然后很多人发展了这些想法。但直到四五年前,这个想法才重新获得关注。然后直到 GPT-3,人们才开始看到它在规模上的表现,而不仅仅是在测试环境中。所以,这是一段疯狂的旅程。我们是怎么创办公司的?2010 年,我实际上在和 Demis Hassabis 打扑克,他是我多年的老朋友,我们在伦敦。
It was 2010 that we started the company. Yeah, exactly, which seems kind of insane, like almost 15 years ago. And it's just quite surreal to see because in the last sort of, what is it, 9 to 12 months, it feels like the kind of large language model revolution has come out of nowhere and exploded onto the scene. But in fact, it has been the kind of steady march of many, many years and a huge amount of failure and a lot of risk and a lot of persistence that I often think gets slightly neglected in the story of the perfect explosion of a new technology. You know, in fact, for most of the last decade we didn't have language models. I mean, the Transformer was really only popularized in 2017. I mean, people often say that it was invented then. It was certainly not invented then; it was invented a good 15 or 20 years earlier by Yoshua Bengio and then many other people developed the ideas. But it was really only four or five years ago that the idea started to get traction again. And then it wasn't until GPT-3 that people started to get a glimpse of what it looked like at scale instead of just in a test environment. So yeah, it's been a crazy journey. How did we start the company? So 2010, I was actually playing poker with Demis Hassabis, who is my longtime friend since we were quite a bit younger, in London.
我猜是在那些高抽成的赌场?
I assume at those high rake casinos?
没错。是在伦敦的维多利亚赌场,在 Edgware 路上。不是世界上最大的牌局。我记得大概是一个 250 英镑的比赛,只有 120 人。但我们经常参加这种比赛。我们俩都对扑克非常热衷。我当时是那种在 PokerStars 上同时开八桌的人。我朋友开 16 桌,但我没有每分钟操作速度来管理那么多。
That's right. It was at the Victoria Casino in London, which is on Edgware Road. Not the big game in the world. I seem to remember it was probably a £250 tournament, only 120 people. But you know, we would play at these things regularly. Both of us were very passionate about poker. I was playing—I was one of these people that was doing like eight-table PokerStars back in the day. My friends were doing 16-table, but I didn't have the actions per minute speed to be able to manage that.
你在计时。是啊,多桌不容易。不过多桌会变成一种心流体验,你开始看到模式,对吧?因为你玩得太快,只能凭直觉。现在似乎 GTO 和这些理论,人们能很快应用。我讨厌在线扑克。我喜欢面对面,因为我唯一的优势是读人,这是游戏的关键部分,而我在线很难读人。这也是游戏的乐趣,对吧?比如把人逼出底池,嘲笑他们的损失。那是乐趣所在。但打很多手牌也是很好的练习方式,因为你最终会形成启发式。问题在于,如果你只玩现场,你永远看不到足够的牌局量来获得经验范围。所以短暂地滥用在线扑克的好处是你能看到深度和广度,这很酷。但你会养成坏习惯,因为它会让你过于谨慎。
You're on a clock. Yeah, it's not easy to multi-table. Although it's something about multi-tabling becomes like a flow experience and you start to see patterns, right? Because you're playing so fast that you have no choice but to kind of play instinct, right? And now it seems like GTO and all these theories, people are able to really deploy it very quickly. I hate online poker. I like in person because the only edge I have is my ability to read people, which is such a critical part of the game, and it's so hard for me to read people online. It's also the fun part of the game, right? Like pushing people off pots and teasing people for their losses. I mean, that's the fun part. So yeah, but I mean, getting through a lot of hands is also a very great way to practice. I mean, because you end up developing heuristics. And so you just see—that's the problem is it's such a high variance game in your career if you only ever play live, you never get to see the volume which gives you the range of experiences. So the good thing about really having a short stint of abusing online poker is that you just get to see depth and breadth, which is cool. But you can pick up bad habits because it can make you too cautious.
啊,有意思。我以前没听说过。所以你不——你变得过于谨慎是因为每个人都在读对方的统计数据,你会想,‘我在这里太容易被看穿了。我不能做非传统的打法。我会被抓到。’是的,因为大家看到的牌局量都大得多,然后他们玩得更可预测、更有结构。所以你学会预测别人的动作。而且,因为他们看到更多牌局,他们对手牌更慎重、更谨慎。而在家庭游戏中,即使玩六到八小时,你可能只看到几百手牌。所以你的范围明显低得多。你会玩那些在线下你会放弃的牌,因为在线你看到更多牌局。经典的例子是 NIT。我们以前叫他们 nit,当他们来到现场牌桌,明显只是玩机器人游戏,把自己逼疯,因为他们看不到足够的牌局。挺搞笑的。是啊,这是个迷人的游戏。在赌场现场玩,你能看到人性的广泛光谱。我刚刚跟人聊起我的朋友 Sky Dayton,我以前在洛杉矶的 Hollywood Park 和 Commerce 玩。我们玩最低的桌子。有一次我试图更好地读人,我想出了绝地扑克的主意:我把牌拿起来,拇指放在上面,但我会假装看牌,实际上把牌遮住,这样我就不知道我有什么牌。Mustafa,这是最好的方法。
Ah, interesting. I haven't heard that before. So you don't—is that the reason you get too cautious is because everybody's reading each other's statistics and you're just like, 'I'm going to be too easy to read here. I can't make a non-traditional play. I'm going to get caught.' Yeah, because everybody sees so much more volume then, and then they play in a much more predictable and structured way. So you learn to predict everybody else's moves. And you know, also they end up being, because they see more volume, they are more deliberate with their hands and more cautious. Whereas in a home game, you may only see a couple hundred hands even in a six to eight hour game, right? And so your range is clearly much lower. You're playing cards that you would otherwise leave behind because you're seeing more throughput online, right? So you know, the classic is the NIT. You know, we used to call them the nit when they would come to the live tables and they'd clearly just playing this robotic game and driving themselves nuts because they weren't seeing enough volume. Pretty funny. Yeah, it's such a fascinating game. And playing live in a casino, you get to see a real broad spectrum of humanity. I was just talking to somebody about my friend Sky Dayton, and I used to play at Hollywood Park in Commerce, in LA. And we would play at the lowest tables. And at one point I was trying to figure out how to read people better, and I came up with the idea of Jedi poker, where I would pull my cards up, put my thumb on it, but I'd make a bit of a show of looking at my cards, but I would have them covered so I didn't know what cards I had. Mustafa, that's the best way.
我只会针对对手来打,我会想,这个人看起来很强,这个人看起来很害怕。让我看看能不能把他逼出底池。然后到了河牌,如果有人跟注我,我会亮牌,然后很尴尬地说,哦,我有个暗三条自己都不知道,或者我只有底对。他们会说,你怎么能那样下注,完全没道理。而没道理正是扑克的一部分,因为你必须打破对手 100%读懂你的能力。这是我最喜欢的方法之一。另一种我喜欢的方法是,在翻牌圈假装自己有一手我没有的牌,假设那是与我对对手的判断相反的手牌或者更好的手牌。这其实是个很好的方法,因为这样你可以在三条街持续下注,但不会显得荒谬或疯狂。你在讲一个故事,你在叙述你持有 10-J,当牌面出现 9-K-Q 时,你会说,我在打 10-J,我会像 10-J 那样打。你只是要确保如果对手真的持有你想假装的那手牌时,你知道如何弃牌,因为那会变得很棘手。
I would only play the person and I'd be like, this person seems very strong, this person seems pretty scared. Let me see if I can get this person off the hand. And then I get to the river and I would literally, if somebody called me down, I would turn over my cards and be embarrassed, like, oh I have a set I didn't know it, or I had bottom pair. And they'd be like, how could you bet like that, it makes no sense. And making no sense is part of poker because you have to break the ability for people to be able to read you 100%. This is one of my favorite ways of doing it. The other way I like to do it is to represent a hand off the flop that I don't have, assuming that it is the opposite hand or a better hand than whatever I place that person on. So you know, that's actually a very good way of doing it because then you bet consistently across the three streets, but you're not being ridiculous and wacky. You're telling a story, you're representing a narrative that you have 10-Jack, and when the board comes down 9-King-Queen, you're like, I'm playing 10-Jack and I'm going to play it like 10-Jack would play this. You just got to make sure you know how to lay down if your opponent actually ends up having the hand that you try to represent, because that can get pretty sticky.
你还在为你的初创公司使用个人电话号码吗?现在是 2023 年,是时候停止了。这是创始人犯的一个巨大错误。为什么?你刚刚开始创业,没有把电话号码视为公司知识产权的重要组成部分。使用 OpenPhone,你可以完全解决这个问题。他们重新思考了现代商务电话及其运作方式。非常简单:你只需在手机或桌面上下载应用,选择一个号码,就完成了。而且价格非常低廉,非常实惠。想想看:如果你的销售团队使用个人电话号码,销售人员离职去了竞争对手那里,你就无法了解发生过哪些通话,有哪些人的电话号码。那是你公司的数据库。如果你让销售团队或客服团队随意使用个人号码,那太不专业了。要专业,就用 OpenPhone。我们用它来做活动沟通,所以我们有一个电话号码,但可以轮流转给多个人。你可以有一个共享电话号码,用于客服。OpenPhone 在 G2 上客户满意度排名第一,我相信 G2 的评分。OpenPhone 已经准备好了,价格实惠,每月只需 13 美元。但 Twist 听众可以在前六个月享受任何套餐 20%的折扣,访问 openphone.com/twist。如果你在其他服务商那里有现有号码,没问题,轻松简单,OpenPhone 会免费为你转号。前往 openphone.com/twist 开始免费试用,并获得 20%的折扣。
Are you still using your personal phone number for your startup? It's 2023, it's time to stop. It is a huge mistake that founders make. Why? You're just getting started with your company and you don't think about phone numbers as being an important part of the IP collection of your startup. With OpenPhone, you can totally solve this problem. They've rethought everything about a modern business phone and how it should work. It's super easy: you just download the app on your phone or your desktop, pick a number, and you're done. And you do it for just such a low price, it's so affordable. Think about it: if you have your sales team using their personal phone numbers, a salesperson leaves and goes to a competitor, you don't have any insight into what phone calls occurred, what people's phone numbers are. That's your company's database. And if you allow the sales team to run amok or the customer support team, it's just unprofessional. Be professional, use OpenPhone. We use it for things like event communication, so we get one phone number but it can go to multiple people like a round-robin thing. You have a shared phone number, do that for customer support. OpenPhone is rated number one on G2 for customer satisfaction, and I trust G2's ratings. OpenPhone is ready, it's affordable, starts at just $13 a month. But Twist listeners can get 20% off any plan for the first six months at openphone.com/twist. And if you have existing numbers with another service, no problem, easy peasy lemon squeezy, OpenPhone will port them over at no cost. Head to openphone.com/twist to start your free trial and get 20% off.
所以你在打牌,然后被淘汰出锦标赛了,你坐在那里做赛后分析……还是你进了决赛桌但早早出局了?你说对了。现在你试图向对方解释你的坏运气以及其他人有多差劲,对吧?我们已经抱怨完了坏运气,前半个小时都在复盘我们被淘汰的那手牌。然后我们坐在那里吃巧克力蛋糕、香草冰淇淋和健怡可乐,因为我们显然很酷,而且没有生气。我们在谈论世界的未来。我们俩一直都对如何影响世界、未来会是什么样子感兴趣。我们都是非常长远的思考者,我觉得这本能地是我们的天赋之一。我特别感兴趣的是如何做好事,以及政治如何塑造我们的未来等等。我们都在谈论机器人技术,现在是机器人出现并自动化一切的时候了吗?我想我们都同意,实际上那比人们意识到的要遥远得多。但更紧迫的事情可能是教机器自己学习在某个空间中什么是有价值的表征。比如,机器肯定能学会打扑克,机器可以学习一组启发式方法并重现那些模式。当时,Demis 刚刚在伦敦大学学院的计算神经科学部门完成他的博士和博士后工作。所以他邀请我参加午餐学习会,我参加了将近六个月。我想几乎每天我都去,基本上是从盖茨比计算神经科学单元的后门溜进去,听午餐学习会。就在那里我们遇到了 Shane Legg,我们的第三位联合创始人。然后我们一起吃了午餐。午餐学习会?就是有午餐,然后有人演讲,你学习。就像自带午餐的讲座,实验室里有四五十人,人们邀请不同的演讲者。所以会有演讲者或博士后。基本上每次午餐都有人谈论他们的工作并回答问题。这有点像熊坑;他们不会手下留情。如果你不保持警惕,你会遇到一些非常尖锐的问题。那是一种很棒的学习方式,让你深入其中,真正亲身体验。我当时只有 24 岁。所以在那之后几个月,Shane Legg 被邀请在 2010 年的奇点峰会上做演讲,因为他当时在 LessWrong 论坛上活跃,而且说实话,他那时有点超人类主义倾向。然后我们决定去,因为 Peter Thiel 是赞助商之一,我想是峰会的主要赞助商。我们被邀请参加之后的酒会,我们利用那个机会向 Peter 推销 AGI。值得称赞的是,他是硅谷唯一一个谈论 AGI 或任何形式 AI 的人。对其他人来说,AI 是一个奇怪的禁忌词,大家都在谈论机器学习,但甚至也不是真正在谈;主要是在学术实验室里人们才谈论机器学习。然后我们去了他在普雷西迪奥的办公室,创始人基金办公室。他当场做了决定。很简单。我想他给了我们大约 200 万美元。1000 万估值?甚至没有,老兄。大概只有一半。我们是从伦敦来的无名小卒。他开玩笑说,这就像是从索马里来的。那是他的看法。他 literally 说,‘你们还不如从索马里来。’
So you're playing cards and you get bounced out of this tournament, and you're sitting there doing your post-bounce... Or did you make it to the final table and you're out early? You nailed it. So now you're trying to explain your bad luck and how bad everybody else is to each other, right? We've gotten over the whining about our bad beats, that took up the first half hour running through our knockout hand. And then we're sitting there eating chocolate cake and vanilla ice cream and Diet Cokes, because obviously we're super cool, and we're not getting pissed. We're talking about the future of the world. Both of us have always been interested in how we impact the world, what the future looks like. We've both been very long-term thinkers, and just instinctively that is one of our kind of gifts, I think. I was particularly interested in how you do good in the world and how politics shape our future and stuff like that. We were both talking about robotics, and is now the time for robots to come on and automate everything? I think we both agreed that actually that was way further away than people realized. But the thing that was likely to be more pressing is teaching machines to learn their own representations of what is valuable in a space. Like, surely a machine could learn to play poker, a machine could learn a set of heuristics and reproduce those patterns. At the time, Demis was just finishing up his PhD and post-doctoral work in neuroscience at UCL at the Computational Neuroscience Unit. So he invited me to join the lunch and learns, which I did for almost six months. I think pretty much every day went down to, you know, basically smuggled in the back door of the Gatsby Computational Neuroscience Unit and just listened to the lunch and learns. And that's where we met Shane Legg, our third co-founder. Then we all went for lunch. Lunch and learn? I mean, there's lunch and then somebody speaks and you learn. It's like a brown bag lunch, you know, where there'll be 40 or 50 people at the lab, and people invite different speakers. So there'll be speakers or there'll be postdocs. Basically every lunch someone gives a talk about their work and takes questions. It's a bit of a bear pit; they don't take prisoners. If you're not on your toes, you get some pretty rough questions. It was just an amazing way to learn and be thrown in the deep end and really experience it firsthand. I was only 24 at the time. So then basically a few months after that, Shane Legg got invited to the Singularity Summit in 2010 to be a speaker because he was on the LessWrong forums back in the day, and he was a bit of a transhumanist, to be honest with you, at that time. Then we decided to go because Peter Thiel was one of the sponsors, I think the main sponsor of the summit. We got invited to the drinks afterwards, and we used that as an opportunity to pitch Peter on AGI. He was the only person in the Valley, to his credit, talking about AGI or even AI in any form. To everybody else, AI was a weird taboo word, and everyone was sort of talking about machine learning, but even not really; it was mostly in the academic labs that people talked about machine learning. Then we went to his office in the Presidio, to the Founders Fund office. He made a decision on the spot. It was pretty easy. I think he gave us like $2 million. 10 million valuation? Not even, dude. It was like half that. We were randoms from London. He joked that it might as well be Somalia. That was his view. He literally said, 'You might as well be from Somalia.'
在索马里投资,我当时觉得伦敦是个严肃的地方,但显然彼得不这么认为。你得放在背景里看:他刚做了 Facebook 的投资,可能自我感觉良好,那笔投资很顺利。种子轮的估值大概在 500 万到 800 万美元。而且说实话,那时候 AI,就像你说的,没人觉得它有商业应用或者能成功,对吧?问题就是:这东西真的能给出一个在现实世界有应用的答案吗?因为之前有深蓝,对吧?卡斯帕罗夫被击败了,所以狭义 AI 已经证明了自己,但 IBM 花了数亿美元,却没有产品。我的意思是,那就是当时的局面,对吧?就像个无底洞。当然,那也比我们早了十年,所以已经证明没有严肃的商业应用。所以我们在推销时其实有一个不成文的目标:别提,别提,千万别提,因为它就像个很酷的研究项目,但从未产生我们期望的影响。是的,头两三年非常艰难,因为深度学习似乎没有流行起来。然后突然之间,2012 年我们有了 Alex Krizhevsky 的猫分类论文,AlexNet,2013 年我们发表了 Atari 游戏玩家 DQN。那真的改变了一切,因为拉里·佩奇看了演示,直接冷邮件发到 page@google.com,说:‘你们应该来加入我们。我整个职业生涯都在建设基础设施,让像你们这样的公司能来研究 AGI。’回过头来,你对彼得是怎么推销的?‘我们要做强化学习。不知道有没有应用。这是个科学项目。你的两百万三年内就会花光。’你是给他画了商业化的路径,还是说‘咱们先在实验室里看看能做什么’?
Investing in Somalia, I was like London is a serious place, but apparently not to Peter. Well, you gotta also put in context: he had just done the Facebook investments, probably feeling pretty good about himself, that was going well. And $5 to $8 million was what a seed round valuation would be. And AI at the time, to be honest, as you said, nobody thought there was a commercial application or that it was going to work, right? Like that was kind of the question: is this actually going to come up with an answer that is going to have some application in the real world? Because you had Deep Blue, right? We had Kasparov got beat, and so narrow AI had proven itself, but IBM had spent hundreds of millions of dollars and they had no product. I mean, that was the playing field, right? It was like this is a money pit, right? And of course, that was a decade before us as well, you know, so that had proven to not have serious commercial applications. So that was actually a kind of non-goal in our pitching: to not bring up, don't bring up, don't bring it up, because it was like a cool research thing but never quite had the impact that we hoped. And yeah, you know, for the first two or three years, it was very tough going because deep learning just didn't seem to be catching on. And then all of a sudden, we had the cat classification paper from Alex Krizhevsky, AlexNet, in 2012, and then in 2013, we had the Atari game player, DQN, which we published. And that was really the thing that changed everything for us, because Larry Page had seen the demo and he just emailed us cold at page@google.com and was like, 'You guys should come and be part of us. I've spent my entire career building the infrastructure to enable a company like you guys to come and work on AGI.' Stepping back, what was the pitch to Peter? 'We're going to build reinforcement learning. We don't know if there's an application. It's a science project. Your two million's going to be gone in three years.' Like, was there any path to commercialization that you pitched him on, or was it 'let's see what we can do in the lab'?
有的。我的意思是,我们其实没有向他推销强化学习,因为那时候还太早。我们推销的是深度学习,我们当时在做的是针对时尚、家具和服装等的视觉图像搜索工具。我们实际上持有该领域深度学习的第一个专利,它提取一件衣服的形状、纹理和颜色,比如一件更实惠的高街版本,然后用它来找到更昂贵的等价物,供你比较。那其实是一个重要时刻,因为那是生成式 AI 运动的开端。现在叫 Gen AI,但当时根本不这么叫,真的只是深度学习分类。是的,那时候别想着生成什么,你是在识别东西:这是热狗,这是狗,这是两个不同的东西。那个框架当时还没出现。
There was, yeah. So I mean, we actually didn't pitch him on reinforcement learning because at that point that was really early. We pitched him on deep learning, and what we were working on was a visual image search tool for fashion and furniture and clothing and so on. And we actually held the first patent for deep learning in this area, which actually takes the shape, the texture, and the color of one item of clothing, like ideally a more affordable high street version, and then uses that to find the more expensive equivalent that you could then compare. And that was a big moment actually, because it was the beginnings of the generative AI movement. I mean, it's now called Gen AI, but it was never called that at the time. It was really just deep learning classification. Yeah, at that time, forget about generating something, you were trying to identify something: this is a hot dog, this is a dog, these are two different things. And that framework hadn't actually happened yet.
谷歌当时在做什么?他们会找两三个低薪工人,给一组图片,让他们说‘给这张图描述五个标签’,然后看哪三四个标签在两个人之间是共同的,那就是这张图的内容,对吧?那就是当时谷歌索引的状态,对吧?完全正确。他们还会让工人在图像的某些部分画边界框,比如这个区域包含企鹅,这个区域包含冰山。结果发现,这正是分层神经网络表示所擅长的。它基本上会把相关像素聚类到特定区域,然后在有尖锐区别的地方,比如边缘、线条或聚类断裂处,就会形成一个子表示。然后下一层会吸收那个子表示,逐步构建越来越具有符号代表性的概念。比如从虹膜的一小块区域,到更大的眼睛,到眉毛,到脸的一侧,到整张脸,再到背景。你可以把这理解为理解分层神经网络表示是如何形成的。显然,过去十年我们在分类方面取得了很大进展,然后你利用这些分类来生成新的预测。而生成本质上就是:给定一个句子,在这个大空间中找到所有竞争点的最优表示,最能代表这个长句子的新图像。那就是我们在谷歌 2017 年那篇论文中听到的 Transformer 模型。
What was Google doing at the time? Would they put two or three low-wage people on a group of images and they would say, 'Describe five tags for this image,' and then whichever three or four came in common with two different people, that was what the image was about, right? That was the state of the Google index at the time, right? Spot on. And they would have them draw bounding boxes around certain parts of the image, so this area of the image contains a penguin, this one contains an iceberg. And it turned out that was exactly the kind of thing that this hierarchical neural network representation was pretty good at doing. Like it would essentially cluster together pixels which were correlated around a particular region, and then where there was a sharp distinction like an edge or a line or a break in a cluster, then that would end up being a sub-representation. And then the next layer would absorb that sub-representation and increasingly build more and more symbolically representative ideas. Like it would go from a tiny little area of the iris to a wider eye, to an eyebrow, to the side of the face, to the full face, to the background. And you could kind of think of that as a way of understanding how the hierarchical neural network representation was formed. And obviously now that we had made so much progress over the last 10 years on the classification side, you then use those classifications to generate novel predictions. And that's basically what generation is doing: saying given this sentence, find the sort of optimal representation of all the competing points in this big space that best represents this long sentence as a new image. And that's the Transformer model that we hear about in that 2017 paper from Google.
是的,没错。嗯,那是深度学习,但还有很多其他生成式 AI 组件也在朝那个方向推进。但他们做到了,首先在语言方面让它工作起来。那才是真正的大事。
Yeah, exactly. Yeah, well that's deep learning, but then there's lots of other generative AI components that were pushing in that direction. But they did it, they made it work first for the language side of things. That was really the big deal.
所以过了几年,你弄明白了一些事情,然后开始进入强化学习。那时候拉里·佩奇是……拉里在董事会吗?还是当时只有埃隆在董事会?还有彼得?
So you get a couple years into this, you figured a couple of things out, and you start getting into reinforcement learning. So that's when Larry Page was... Was Larry on the board, or was just Elon on the board at that time? And Peter?
不,我们先是彼得投资,然后是埃隆。然后我们是……抱歉,不,我们是马克·施特劳斯的德丰杰基金第一支基金在 2012 年投的第三张支票,我想是的。那真是一个做投资者的好时期,因为金融危机后只有疯子才创办公司。那五年里,如果你创办公司,你别无选择,因为你是个疯子,必须创办那家公司,因为世界上的一切都在告诉你不要创办公司,对吧?那将是痛苦和折磨。
No, so we had first Peter invest, then Elon. Then we were the first check out of... sorry, no, we were the third check out of the first fund of Mark Stauss' Draper Fisher Jurvetson in 2012, I think it was. It was such a great time period to be an investor, because only lunatics were starting companies after the Great Financial Crisis. It was like this 5-year period where if you started a company, you had no choice because you were a lunatic who had to start that company, because everything in the world was telling you don't start a company, right? It's going to be pain and suffering.
所以你筹到了这笔钱。你们开始做的第一个项目是什么?你们是怎么选的?然后是什么让它成功了?因为我记得 AlphaGo 是其中之一,还有一个变得有意识的钟。我们听说 DeepMind 内部有很多这样的小项目,但 DeepMind 大部分时间都保持低调。我想我们在整个时期大部分时间都在秘密运营,甚至没有公布我们的投资者。我的意思是,还有其他人:我们有周凯旋作为来自维港投资的另一位投资者,我们有一群非常好的人。我们很幸运。我想我们最终筹集了 4500 万美元,所以我们每年都回去。
So you raised this money. What was the first project that you guys started to work on? How did you pick it? And then what clicked? Because I remember AlphaGo was one, and then there was this clock that became sentient. There were just all these little projects that we would hear about inside of DeepMind, but DeepMind kind of kept a lot close to the vest. I think we operated in stealth for most of our entire period, and we actually didn't even announce our investors. I mean, there was a bunch of others: we had Selina Chow as another investor from Horizons Ventures, and you know, we had a very good group of people. We were lucky. I think we raised $45 million in the end, so each year we went back.
我记得我们融了两次或三次,然后是 10,再然后是 30。每次你展示了什么,让人们在大多数人不相信这个愿景的时候,仍然愿意投资?
I think we raised like two or three and then 10 and then 30. What did you show each time to keep people investing in the vision during a time when people didn't believe in the vision? Most people didn't.
是的,我的意思是,我们在第二次融资时展示了 Flatland,那是一个基于智能体的小环境,就像一个二维网格世界,模型通过纯像素学会了在环境中导航。然后我们说,好的,下一个里程碑,我们要教模型学习任意雅达利游戏。最后我们玩了 56 个游戏,这相当不可思议。这是每秒 24 帧,它学习关联动作,比如上下左右或射击。原始的雅达利 2600 控制器,没错,有五个动作。确实如此。所以它基本上要弄清楚这些动作,一开始随机移动,然后偶然遇到有奖励的时刻。它幸运地得到了分数,然后意识到,哦,那是有用的。下次我看到球朝那个位置飞来,我就把球拍向左或向右移动。这简直不可思议,纯粹通过自我对弈和强化学习,只是非常简单的启发式探索和利用策略,就能学会所有游戏,达到超人类表现。对我来说这太震撼了。这一切都是在雅达利 2600 模拟器中完成的。显然,你没有用物理摇杆和机器人。它可以在云端快速运行,对吧?你们找到了加速的方法,让它能玩数百万次 Pong 或 Tank 之类的早期游戏。
Yeah, I mean, so we showed in the second time that we raised, we showed Flatland, which was our little agent-based environment, like a 2D grid world, where the model had kind of learned a way to navigate through the environment using purely the pixels. And we then said, okay, for our next milestone, we're going to basically teach the model to learn arbitrary games of Atari. And in the end, we played 56 games, which is pretty incredible. This is 24 frames per second, and it's learning to basically correlate actions where it can basically go up, down, left, right, or shoot. Original Atari 2600 controller, yeah exactly, which had five actions. Exactly. And so it's basically got to figure out which of those actions it's randomly kind of moving them around at the beginning, and then it stumbles on a rewarding moment. It luckily gets some score, and then it realizes, okay, that's a useful thing to do. Next time I see the ball bouncing towards me in that position, I'll move the paddle left or right. And it's just kind of incredible that purely through self-play and reinforcement learning, just very simple heuristic exploration and then exploit the strategy that turns out to be useful for generating score, and suddenly you can learn to play all the games to basically superhuman performance. I mean, that was mind-blowing to me. And that was all done in an Atari 2600 emulator. Obviously, you're not taking a physical joystick and putting a robot on it. It's able to run very quickly in the cloud, right? You figured out a way to accelerate it so that it could just be playing whatever Pong or Tank or whatever those early games were, play them millions of times.
它需要运行多少次才能精通这些游戏?你追踪过吗?比如投了多少个币才拿到最高分?
How many runs did it have to do to perfect them? Did you track that? Like how many quarters until you perfected the game and got the high score?
有趣的是你提到了云,因为那是 2012、2013 年,当时基本上还没有云。我们实际上是在本地运行的。我们在办公室有自己的小集群,用来训练 Atari DQN。它使用了 2 petaflops 的计算量。flop 是浮点运算,是计算单位,相当于一次计算。peta 是百万的十亿倍,所以训练整个模型大约需要两百万亿亿次计算,耗时两周。可以这么说,当时这是最大的训练之一。我们不确定,但当时没有其他类似的大规模训练。所以可以说它可能是最大的。那是十年前,现在我们在 Inflection 和其他前沿模型公司训练的模型使用了 100 亿 petaflops。哇,100 亿百万亿亿次浮点运算,这太疯狂了。人脑甚至无法理解那是什么。就像我们开始谈论银河系有十亿颗恒星,而又有数十亿个星系。人脑根本设计来理解数百万亿亿、数万亿亿这些数字。根本不可能。
I mean, it's interesting that you mentioned Cloud, right, because this was 2012, 2013, so there wasn't really any Cloud to speak of. We actually ran it on-prem. We had our own little cluster in the office, and it used to train Atari DQN. It used two petaflops of computation. So a flop is a floating point operation, this is a unit of computation, like one calculation. And obviously peta is a million billion, so it's two million billion calculations to train the entire model over the course of about two weeks. To put that into perspective, at the time that was one of the largest. We don't know for sure, but there weren't any other big training runs of those kinds of things at that time. So it's fair to say it was probably the largest. That was a decade ago, and you roll forward, and the models that we train today at Inflection, and the other frontier model companies, use 10 billion petaflops. Wow, 10 billion million billion floating point operations, which is insane. A human brain cannot even conceive of what that is. It's kind of like when we start talking about there are a billion suns in our galaxy, and there are billions of galaxies. The human mind is not designed to even comprehend millions of billions of millions billions of millions billions. It's just not even possible.
我们都知道,区分伟大初创公司和好初创公司的一个因素是产品速度。这是什么意思?产品速度,一个花哨的术语,对吧?你已经有了产品,还有速度,也就是产品改进的速度。你能发布更新吗?你能推出新功能吗?你能修复 bug 吗?你能迭代界面吗?你能为客户解决问题吗?而且你能快速做到吗?因为你并不孤单。你有竞争对手,你的客户有选择。他们可能通过编写自己的定制代码来解决问题,或者使用你的解决方案。这就是初创公司的核心:你能多快启动产品速度?那么,如何加速呢?每个人都说,好的,我们想更快,但你必须聪明地加速。Crowdbotics 会帮助你做到这一点。他们基本上是 CTO 即服务。他们为你提供最优的架构,让你尽快将产品推向市场。你可以按需获得产品经理和开发人才,他们会帮助你的应用以比传统开发快 10 倍的速度投入生产。Crowdbotics 可以与你的内部开发团队合作,或者你也可以让他们独立工作,而你拥有所有知识产权和源代码。让 Crowdbotics 的人今天就来加速你的产品速度。不再等待。在 crowdbotics.com/twist 获取免费构建计划。价值 4.99 美元,仅限 Twist 听众。你免费获得。网址是 C-R-O-W-D-B-O-T-I-C-S 点 com 斜线 twist,获取免费构建计划。
We all know the one thing that separates great startups from the good ones is product velocity. What does it mean? Product velocity, fancy term, right? You already got your product and you have velocity, speed, the speed in which your product improves. So can you ship updates? Can you release new features? Can you do bug fixes? Can you iterate on the interface? Can you solve problems for your customers? And can you do it quickly? Because you're not alone. You have competitors and your customers have choices. They may solve their problems by writing their own custom code, or they might use your solution. This is what startups are about: how fast can you get that product velocity going? And so, how do you supercharge it? Everybody says, okay, yeah, we want to go faster, but you got to go faster intelligently. And Crowdbotics is going to help you do that. They're your CTO as a service, basically. They provide you with the most optimal architecture to get your product to market as fast as possible. You'll have access to an on-demand product manager and developer talent, and they will help get your app into production 10 times faster than conventional development. Crowdbotics can work with your in-house dev team, or you can just have them work independently, and you own all the IP, you own all the source code. Let the folks at Crowdbotics supercharge your product velocity today. No more waiting. Get a free build plan at crowdbotics.com/twist. That's a $4.99 value, just for the Twist listeners. You get that for free. That's C-R-O-W-D-B-O-T-I-C-S dot com slash twist for a free build plan.
硬件确实开始跟上了,硬件似乎是其中的推动因素。也许你可以谈谈当时的基础设施是什么样的,硬件规模与今天相比如何,以及你在 Inflection 正在做什么。
The hardware did start to catch up here, and hardware seems to have been part of the enabling here. Maybe you could talk a little bit about what the infrastructure looked like at that time, the hardware footprint versus what we see today, and what you're doing at Inflection.
硬件规模,这是个好问题。我的意思是,这实际上是硬件革命,而不是 AI 革命。有趣的是,人们专注于算法。显然算法很关键,但它们并没有以计算能力那样的指数速度发展。所以我描述的那 100 亿百万亿亿 petaflops,相当于一个数量级,也就是每年前沿模型使用的总计算量增加 10 倍,持续 10 年。10 的 10 次方。这太疯狂了。真的疯狂。所以这基本上是关于硬件的。这就是为什么我认为这场革命实际上比人们意识到的更容易预测。这个轨迹已经持续了很长时间,我们可以展望接下来的三、四、五次翻倍,或者抱歉,三、四、五次 10 倍增长。它们不再是摩尔定律那样的翻倍;而是计算量的数量级增长。这是一个可预测的轨迹。显然,不清楚会涌现出什么能力,但你肯定可以预测我们将能够构建什么。Steve Jurvetson 有很多这方面的图表,他一直在追踪,我不知道你是否看过 Steve 关于计算能力总量的图表,那种临界点在一定程度上是可预测的。而现在我们遇到了热量和电力摩擦,我想,这是目前的限制,或者说我们能把这些超级计算机连接在一起的程度。
The hardware footprint, it's a great point. I mean, it's really the hardware revolution rather than the AI revolution. It's funny, people fixate on the algorithms. Obviously the algorithms are critical, but they really have not evolved at the exponential rate that computing has evolved at. So those 10 billion million billion petaflops I described, that is the equivalent of one order of magnitude, so a 10x increase in the total amount of compute used for the cutting-edge models every year for 10 years. 10 to the power of 10. It's insane. It's truly insane. So that is basically about hardware. And that's why I think actually this revolution has been easier to predict than people realize. This trajectory has been continuing for a long time, and we can look out at what the next three, four, five doublings look like, or sorry, three, four, five 10x's look like. They're not doublings anymore like Moore's law; they're orders of magnitude increase in compute. And that's a predictable trajectory. Obviously it's unclear exactly what capabilities emerge from that, but you can certainly predict what we're going to be able to build. Steve Jurvetson has a lot of charts on this, where he's been tracking, I don't know if you've seen Steve's charts on just the amount of computing power and the sort of tipping point is somewhat predictable. And now we've got heat and power friction, I guess, is the limit right now, or how much we can connect these supercomputers together.
现在的瓶颈因素是什么?
What's the gating factor now?
是的,这是个好问题。这将成为制约因素。A100 每芯片功耗 700 瓦,H100 翻倍,约 1200 瓦。哇。所以下一代芯片,对吧?显然每个节点有八块这样的芯片,加上机箱和节点本身还有一些额外的功耗限制。所以实际上数据中心的设计与两三年前不同了——机架之间留有空间,不是完全堆叠,必须有很大的间隙。在我看到的一些新冷却系统设计中,节点层之间甚至会有风扇。而且全部采用光纤、全玻璃光子计算来传输数据,因为现在移动的数据量无法通过铜缆或以太网传输——数据量太大了。哦,当然。全部都是光纤电缆,实际上叫做 InfiniBand,是 Mellanox 英伟达的线缆,速度达到每秒 900 GB,用于直接的芯片到芯片连接,这相当惊人。所以这完全是由硬件创新驱动的,而这些硬件创新非常可预测,因为它们实际上提前三年就规划好了。是的,因为他们正在规划建造那些图纸,并让晶圆厂和工厂准备好实际生产它们。
Yeah, it's a good point. That is going to become the constraint. The A100 uses 700 watts per chip, the H100 is twice that, like 1200 watts. Wow. So the next chip, right? Obviously you have eight of these on a node, then the chassis and the node itself have some additional power constraints. So they actually have a different data center design compared to what it was two or three years ago, where there are spaces in between racks—they're not completely stacked up; there have to be really large gaps. And in some of the designs I've seen for new cooling systems, there will actually be fans in between the node layers. And all fiber optics, all glass photonic computing to transfer data from one to the other, because the amount of data being moved now can't be moved over copper or ethernet cables—it's just too much. Oh, for sure. All of it is fiber optic cable, actually called InfiniBand, the Mellanox Nvidia cabling, and that's like 900 gigabytes per second, which is pretty nuts for direct chip-to-chip connections. So it is really driven by all the hardware innovations, and those hardware innovations are very predictable because they're actually laid out three years in advance. Yeah, because they're planning on building those schematics and getting the fabs and factories ready to actually build them.
所以我想,在某个时间点,DeepMind 内部开始出现一些争议。拉里·佩奇说,我们需要这个团队留在谷歌内部。也许彼得·蒂尔、埃隆希望你们保持独立。也许你可以解释一下那个时刻和当时的决策过程。
So there starts to be a little controversy inside of DeepMind, I guess, at a certain point. Larry Page is like, we need this team inside of Google. Maybe Peter Thiel, Elon want you to stay independent. Maybe you could explain that moment in time and the decision making there.
是的,我的意思是,我认为那是在 2014 年我们被收购的时候。我认为埃隆和彼得,以及我们所有的投资者,都希望我们保持独立。我认为对我们来说,具有挑战性的决定只是我们能看到未来所需的投资规模。我们筹集了 4000 万美元,并且我们看到了一条在 3 到 5 年内花费 5 亿美元的路径。事实上,我们最终确实做到了。DeepMind 现在有 12,300 人,每年在算力上花费超过 10 亿美元。这些都是公开信息。所以这个轨迹非常了不起。我们关注的一件事是——他们给了我们一个不可思议的报价来实现这一点。我们以 6.5 亿美元的价格被收购,当时还没有收入,显然是一笔非常划算的交易,尤其是在当时。我的意思是,世界在过去十年发生了巨大变化,但在当时,人们都摇头说,他们买了什么?事实上,当时的对话是——我想你们当时大概有一百人?更少。是的,是的,没错。对话是,拉里疯了吗?他每名工程师付了 1000 万美元。然后这就变成了,硅谷的工程师每人值 1000 万美元。我说,嗯,这些是不同类型的工程师。你雇佣了一群非常精英的人。
Yeah, I mean, I think this was way back in 2014 that we were acquired. And I think that Elon and Peter, all of our investors, wanted us to stay independent. And I think that the challenging decision for us was just the scale of investment that we could see would be required going forward. I mean, we'd raised $40 million, and we could see a path to spending $500 million in three to five years. And in fact, that's what we ended up doing exactly that. DeepMind now has 12,300 people and spends over a billion dollars on compute a year. So that's public information. So the trajectory is pretty remarkable. And one of the things that we focused on—they made us an incredible offer to be able to do that. We were acquired for $650 million pre-revenue, obviously pretty great deal, especially at the time. I mean, the world has changed dramatically in the last decade, but at the time, people were shaking their heads like, what did they buy? In fact, the conversation was—I think you had maybe a hundred people at the time? Less. Yeah, yeah, exactly. The conversation was, has Larry lost his mind? He just paid $10 million per engineer. And then that became, well, engineers in Silicon Valley are worth $10 million each. I'm like, well, these are different types of engineers. You hired a very elite group of people.
也许你可以谈谈当时组建 DeepMind 团队的招聘情况,因为有很多博士,很多人都有——你们的人才储备相当深厚。
Maybe you could talk about the recruiting of bringing together the DeepMind team at the time, because it was a lot of PhDs, a lot of people who had—you had a pretty deep bench there.
是的,我们当时非常专注于招聘最优秀的博士和博士后。而且我把这一点贯彻到了我在 Inflection 的招聘中。我的意思是,人才最终是差异化因素。你可以最先获得算力,你可以拥有最多的资本,但选择一支非常非常高质量的团队才是真正带来差异的唯一因素。这意味着你必须非常谨慎地决定不雇佣谁。实际上,当时我们身边有多少对深度学习革命至关重要的人,这真是令人惊叹。杰夫·辛顿在我们这里做了两年的顾问,之后他成立了自己的公司并卖给了谷歌。还有伊利亚·苏茨克维,现在是 OpenAI 的首席科学家。沃伊切赫·扎伦巴曾是 DeepMind 的实习生,也是 OpenAI 的联合创始人之一。这将会像 PayPal 黑帮一样——它将会是 DeepMind 黑帮。实际上它已经变成了 DeepMind 黑帮。你有一整群校友,他们正在这里创造未来。
Yeah, we were extremely focused on hiring the best PhDs and postdocs actually. And I've carried that through to how I hire at Inflection. I mean, talent is the differentiator at the end of the day. You could be first to get access to compute, you can have the most amount of capital, but selecting a very, very high quality team is really the only thing that makes the real difference. And that means you have to be very deliberate about who you don't hire. It was actually amazing at that time how many people who were fundamental to the deep learning revolution we had around us. So Jeff Hinton was one of our consultants for two years before he set up his company that he then sold to Google. So was Ilya Sutskever, the chief scientist of OpenAI now. Wojciech Zaremba was an intern at DeepMind who was one of the co-founders of OpenAI. It's going to be like the PayPal Mafia—it's going to be the DeepMind Mafia. It's already turned out to be the DeepMind Mafia basically. You got a whole group of alumni who are just creating the future here.
回顾过去,卖掉是个错误吗?你后悔卖给谷歌吗?你应该接受埃隆的建议保持独立吗?
Looking back on it, was it a mistake to sell? Do you regret selling to Google? Should you have taken Elon's advice and stayed independent?
埃隆当然很希望我们加入他的生态系统,做特斯拉的事情。但老实说,我有点——我的意思是,回到当时,他是一个了不起的人,但在 2014 年,这是一个非常不确定的赌注,简直就是不确定性的定义。Model 3 差点毁了他,差点毁了公司。那家公司每次推出产品都经历了一次濒死体验。我的意思是,你要谈论困难的硬件加软件和大规模制造以及建立公共品牌——难度系数是荒谬的。
Elon was certainly keen for us to come and be part of his ecosystem, do the Tesla thing. But I'll be honest, I was a bit—I mean, back then, he's an incredible person, but it was a very uncertain bet in 2014, would be the definition of uncertain. I mean, Model 3 almost killed him, almost killed the company. That company has had a near-death experience with each launch of a product. I mean, you want to talk about hard hardware plus software and manufacturing at scale and building a public brand—the degree of difficulty is absurd.
在谷歌内部,就你能谈到的范围而言,你们做了很多理论性的工作,但也做了很多实际的事情。你能谈谈 DeepMind 在谷歌内部参与的重大成果吗?
Inside Google, to the extent you can talk about it, you guys worked on a lot of theoretical things but also a lot of practical stuff. What were the big wins inside of Google that you can talk about that DeepMind participated in?
是的,我的意思是,我们在除搜索和 YouTube 之外的所有主要产品上部署了 DeepMind 技术。所以我认为我们最终在从数据中心到医疗保健、Play Store、Android 电池优化到 Android 操作系统等所有方面做了七个项目。我们将谷歌数据中心机队的冷却能耗降低了 30%。这是一个为期三年的合作,一个巨大的项目。我们使谷歌的风力涡轮机效率提高了 20%,谷歌拥有世界上最大的风力涡轮机农场——这相当疯狂。是的,我们为所有可穿戴设备设计了活动分类算法,基本上可以判断你是睡觉还是跑步。
Yeah, I mean, we deployed DeepMind technologies on all of the main products other than search and YouTube actually. So I think we did seven PAs in the end on everything from data centers to healthcare to Play Store to Android battery optimization to Android operating system. We reduced the amount of energy needed to cool the Google data center fleet by 30%. That was a three-year collaboration, a huge project. We made the Google wind turbines 20% more efficient, which Google has the largest wind turbine farm in the world—it's pretty crazy. Yeah, we designed the activity classification algorithms for all the wearable devices that would basically tell whether you're sleeping or running.
两个最大的业务他们不让你碰。搜索,他们不让你碰。YouTube,他们不让你碰。为什么?你们有这一千名了不起的人,却不让他们碰这两个最大的业务。为什么?
The two biggest franchises they wouldn't let you touch. Search, they wouldn't let you touch. YouTube, they wouldn't let you touch. Why? You've got this incredible thousand folks and you don't let them touch the two biggest franchises. Why?
嗯,我们试过。我们实际上在 2015 年尝试过 YouTube,但我们失败了。那太早了,而且非常困难。我们当时试图优化“接下来观看”的内容,是的。我们试图使用强化学习来实现,但时机太早了。我们没有成功。搜索是另一回事。搜索太难部署任何东西了。他们非常保守。他们也喜欢所有规则都非常透明的事实,这样他们可以确切地看到为什么推荐某个页面,并且拥有更多的控制权。
Well, we tried. We actually tried YouTube in 2015 and we failed. It was too early and it was just super hard. We were trying to optimize watch next time actually, yeah. And we were trying to use reinforcement learning for it, and it was just too early. We didn't succeed. Search is a different story. Search is just so difficult to ship anything in. They're super conservative. They also like the fact that all of the rules are very transparent, so they can see exactly why a page is being recommended and really have much more control.
算法透明度是可以理解的。实际上,有些深度学习系统的部署最终会因为漂移而导致性能随时间下降,比如在六个月内。换句话说,质量会下降——它一开始会上升,然后下降,再下降。为什么会出现这种漂移?人们也在谈论 ChatGPT-4,说它的结果退化了。我不明白为什么会这样。这是垃圾进垃圾出的情况吗?到底发生了什么?
transparency on the algorithm which is very understandable so in fact there were some deployment of deep learning systems which ended up causing regressions over time because of drift over a six-month period and in other words quality would go down well it would go up initially at the beginning and then come down and then come down exactly why does that drift happen people are talking about that with ChatGPT-4 that results have deprecated I didn't understand why that would occur is it garbage in garbage out kind of situation what's happening
嗯,当这种情况发生时,我认为问题略有不同。关于 ChatGPT,很可能他们最初提供的是最好的模型,但服务成本很高,因为它最大、最好,使用最多的 GPU。一旦用户频繁回来,他们就会换用更小、更便宜的模型,质量就会下降。基本上,质量就是成本。所以我们可以提供更便宜的模型来加快速度,但效果就没那么好了。我想这大概就是原因。
Well when something like that happens I think there are slightly different problems. I think with the ChatGPT thing it's probably that they basically serve their best model which is expensive to serve right because it's the biggest and best and uses the most number of GPUs and then once people are coming back frequently they'll serve a smaller model which is cheaper it'll be a less model quality basically as always the case quality is cost right so we can serve a cheaper model for quicker but it won't be as good so that's probably what's going on I think.
啊,我从未听过这个理论,但这说得通。随着使用人数增加,他们可能别无选择,只能给每个人一个更简单或更基础的模型。
Ah I've never heard that theory but that would track and make sense and as more people use it they may have no choice but to give everybody a little bit of an easier model to use or a more basic model because they don't have a choice.
另一个变量是速度。如果你想要非常快,就必须用更小的模型,或者用更多芯片来服务超大型模型。所以三者不可兼得。如果你想要超高质量,那可能非常慢且便宜,但响应时间可能要 20 秒左右。
Well the other variable would be speed so if you want it really fast then you have to get a smaller model or you have to use more chips to serve a super large model. So you can't have all three and so if you want a super high quality one you could have it really slow and cheap but that would be really slow 20 seconds or something for a response.
听着,如果你在科技行业,你一定知道 Carta。Carta 是领先的风险投资和股权管理平台,他们本周在 Startups 节目中有重大消息要分享。Carta 现在允许你组建一个 SPV。你知道什么是 SPV 吗?特殊目的载体。你可以在 Carta 上创建一个 SPV。为什么要这么做?嘿,你是一个天使投资人,你向一家公司投资了 2.5 万美元,就像我对 com.com 做的那样,但你还有大约 20 个朋友也想投入 5000 或 1 万美元。现在你把他们都放进一个 SPV。你告诉 com.com 或你投资的公司团队,这将是一个单项。我会为所有 25 位天使签署。他们说,太好了,我能把另外 10 个天使放进你的 SPV 吗?然后,嘿,如果你想收取附带权益,因为你组织了这笔交易,太好了,你现在就有了一个商业模式。超过 4500 只基金使用 Carta,管理资产超过 1200 亿美元。他们将在你募资的每个阶段支持你,从你的第一个联合组织到建立全球风险投资公司。你可以在世界任何地方筹集和部署资金,因为 Carta 提供美国和国际 SPV。此外,Carta 还为你提供自动化后台解决方案,这样你就可以专注于重要的事情:寻找优秀的初创公司,建立关系,并大力支持那些创始人。行动号召:访问 carta.com/twist,使用代码 twist,你的第一个 SPV 可享受 10%的折扣。多好的交易!Carta,carta.com/twist。确保你使用促销代码 twist 获得 10%的折扣。
Listen if you're in the tech industry you know about Carta. Carta is the leading venture capital and equity management platform and they have huge news to share here on this week in startups. Carta now lets you syndicate an SPV. You know what an SPV is? A special purpose vehicle. So you create an SPV on Carta. Why would you do that? Hey listen, you're an angel investor and you're putting 25k in a company like I did with com.com but you got about 20 friends who also want to put in five or 10K. Now you put them all into an SPV. You tell the team over at com.com or whatever company you're investing in it's going to be one line item. I'll sign for all 25 of those angels and they say oh great can I put 10 other angels in your SPV and then hey if you want to take carry on it because you syndicated the deal great now you got a business model going. They are used by more than 4500 funds representing over $120 billion in assets under administration. They're going to support you at every stage of your fundraising journey from doing your first syndicate to building a global venture capital firm. You can raise and deploy from anywhere in the world because Carta offers US and international SPVs. Also Carta provides an automated back office solution for you so you can focus on what matters: finding great startups, building relationships, and supporting the heck out of those founders. Here's your call to action: go to carta.com/twist and use the code twist to get 10% off your first SPV. What a deal! Carta, carta.com/twist. Make sure you use the promo code twist for 10% off.
你刚离开谷歌,他们直到 OpenAI 推出后才发布这些东西,但他们显然早就有了。谷歌对其品牌名称负有更多责任,不能发布愚蠢或令人困惑的东西,这说得通。但最终,我想 OpenAI 和微软迫使他们出手。为什么会这样?
Just wrapping up your time with Google they never launched any of this stuff until OpenAI did but they clearly had it sitting there right. It makes sense that Google has more responsibility with their brand name and they can't put stuff out there that's silly or confusing under the Google brand name but eventually I guess OpenAI and Microsoft forced their hands right. Why did it go down that way?
是的,人们说谷歌在打瞌睡,等等,但事实并非如此。我在谷歌时就在 LaMDA 团队工作,我花了一年半时间。我们基本上在 ChatGPT 之前就有了 ChatGPT,这太不可思议了。2020 年夏天,我们就有了,它运行良好,令人惊叹。我们实际上……那是我们在 Gmail 自动补全中看到的吗?是那个模型吗?不是 Gmail 自动补全,但它在 2020 年 5 月的 I/O 开发者大会上由 Sundar 展示。他实际上进行了一次对话——我们设计得真蠢——他和一个纸飞机对话,问它作为纸飞机是什么感觉,然后和冥王星对话,语言模型假装成冥王星,谈论天气之类的。
Yeah I mean people say Google was asleep at the wheel and all the rest of it but you know it's not quite true. I think so I was there at Google and working on the LaMDA team right so I spent a year and a half working on that team and we basically had ChatGPT before ChatGPT it was incredible. I mean summer of 2020 and we had it it was working it was amazing and we were actually... was that what we saw in the Gmail autocomplete? Was that model? It wasn't Gmail autocomplete but it was featured by Sundar in May at IO the annual developer conference in 2020. And he actually had a conversation we designed it was so stupid we he had a conversation with a paper airplane about what it's like to be a paper airplane and then he had a conversation with Pluto and then the language model pretended it was Pluto and talked about the weather and stuff.
嗯,我们总是说例子很重要,他们选了糟糕的例子。当你推销你的初创公司或新产品时,你想要最引人注目、有趣且适用的例子,但他们选了这两个无聊的例子。我可以告诉你,这是故意的,因为我们不想让它看起来像人或听起来像人,我们想让它更像是分享酷炫的技术,但这只是那个方向上的第一步。别害怕,它不会抢你的工作,只是冥王星。我的意思是,如果你让它当医生、图书管理员或文字编辑,突然之间,人们就会想,嗯?而这就是今天发生的事情,我认为这是一个很好的转折点。
Well you know what we always say examples matter and they pick terrible examples exactly literally you know when you're pitching your startup you're pitching a new product you want the most evocative interesting applicable example and they picked two inane ones and I can tell you it was deliberate because we didn't want it to look like a person or sound like a person we wanted it to be kind of like a you know sharing the cool technology but it was just the first small step in that direction don't be scared it's not taking your job just Pluto. I mean if you make it a doctor or you make it a librarian or you make it a copy editor all of a sudden it's like huh and that's what's happened today which I think is a good pivot point here.
总之,谷歌是一个大型组织,他们很保守,所以采取了谨慎的做法,而且他们有实力。我认为他们只是有信心不必第一个出手,可以花更多时间把事情做好,而且搜索在分发和数据方面有巨大的锁定效应。我认为这会带来回报,因为谷歌会没事的。我的意思是,谷歌会没事的。
So anyway suffice it to say Google is a large organization and they're conservative and so they just took a measured approach and they have the goods right. I think there was just a confidence that we don't have to go first on this and that we could take more time to get it right and that search is just this phenomenal lock-in in distribution and data and I think that's going to pay dividends because I think Google's going to be just fine. I mean Google's gonna be fine.
我同意。当这一切发生时,我买了谷歌的股票,因为我也看了 Bard。我看到 Bard,心想,你有这么多点击流数据和本地数据,它突然开始做链接表格、插入图片。我以前见过这种情形。我看着谷歌从 10 个蓝色链接发展到全面的搜索内容、购物、地图等等,这发生在一二十年里,显然也会在这里发生。我还认为广告模式——有一种理论说,越混乱,你点击广告越多。但如果你搜索旅行,Bard 结果中的链接没有理由不能变现,事实上它们会的。我认为这是真的。
I agree. I bought Google shares when I saw this going down because I looked at Bard too and I'm watching Bard and I'm like you've got so much clickstream data and you got so much local data that it's all of a sudden doing links tables it's putting in photos. I mean I've seen this movie before. I watch Google go from 10 blue links to comprehensive search content shopping maps everything and that happened over a decade or two and it's obviously going to happen in there. And I also think the ad model you know there is a theory like the more confusing it is the more you click on ads but if you do a search for travel there's no reason that links inside the Bard result cannot be monetizable in fact they will right. I think that's true.
我认为谷歌将会挣扎的地方在于,谷歌已经发展出了令人难以置信的自我阻碍能力,它几乎就像内部……的大师。
I think where Google is going to struggle is that Google has developed an incredible expertise for getting in its own way right it's just almost like the master of like internal...
混乱,有很多出色的团队和项目互相阻碍,因为存在大量重复。这确实是一个非常混乱的地方。所以我认为这对他们来说将是一个挑战。第二点是,广告模式可能不是未来的模式,对吧?人们可能无法容忍口袋里有一个由出价最高者资助的 AI,试图向你推销东西,因为这些模型非常有说服力,非常个性化,它们会了解你,因为你最终会与它们对话,分享你通常不会在普通搜索查询中输入的信息,比如你可能会说一些关于癌症或心碎的敏感话题。但这不像你我之间这样流畅、连续的对话,对吧?所以我认为人们不会希望你的 AI 突然转身说:‘顺便说一句,当当,我是……’你知道。所以我们拭目以待。我认为谷歌在这方面会遇到困难。
Chaos and so there's loads of amazing teams and projects which just block each other because there's huge amounts of duplication. It's a very chaotic place, it really is. And so I think that's going to be challenging for them. And I think the second thing is the ad model may not be the model of the future, right? It may be the case that people cannot tolerate having an AI in your pocket that is funded by whoever is the highest bidder trying to sell you something, because these models are so persuasive, because they're so personal, because they'll get to know you, because you end up having conversations with them and sharing information that you wouldn't normally type in a regular search query, where it's just like you might say something sensitive about your cancer or your heartbreak, you know. But it's not the same as having a fluent, continuous natural language conversation as though you and I are now, right? So I think people are not going to want your AI to suddenly turn around and say, 'By the way, ta-da, I'm...' you know. So we'll see how that turns out. And I think Google's going to struggle with that one.
可能是联盟链接,你知道。如果我和我的 AI 聊天,我忧郁、抑郁、感到悲伤,它知道我的 AI 知道我难过,它可能会说:‘你知道,也许锻炼、冷水浴、去看心理医生。’所有这些在某种程度上都是可货币化的链接。所以如果它给了你完美的答案,问题是,如果你直接得到了答案,还能货币化吗?拉里总是说,最终我们会给你答案,我们只是给你答案。所以人们不禁怀疑这是否会严重破坏广告竞价。
Could be affiliate links, you know. If I was talking to my AI and I'm melancholy, I got depression, I'm feeling sad, and it knows my AI knows I'm sad, it could be like, 'You know, maybe exercise, cold plunge bath, go see a psychiatrist.' All of those things are monetizable links in some way. And so if it gives you the perfect answer, the question is, is it possible to monetize if you just got the answer? And Larry always said, like, eventually we're going to give you the answer, we're just going to give you the answer. And so the mind does wonder if that screws up the ad auction in a major way.
这正是谷歌的第三个问题:如果谷歌总是给你答案,那么开放网络的未来是什么?因为谷歌将去中介化第三方内容创作者。比如你是一个普通的夫妻店,在网站上经营面包店或写博客,依赖展示广告收入,但谷歌直接给你完美的食谱,你为什么还要去第三方博客呢?这对谷歌和监管机构来说实际上是个问题,因为过去 15 年里,谷歌一直告诉监管机构,它爬取这些网站的原因只是为了索引,以便将用户重定向到第三方页面。这听起来公平,对吧?他们常说这是黄页,是一个查找表。但如果现在它切断了信息源,直接给你完美答案,这对监管机构来说是个大问题,尤其是在欧洲,因为许多谷歌高管曾在证人席上声称他们永远不会这样做。而现在模型正在这样做。它们是在网络上训练的,这很明显,已经被证实了。过去你可以问 OpenAI 的 ChatGPT:‘嘿,这个答案是从哪里训练的?’它实际上会告诉你一些训练数据。我想现在它不这么做了。对于数据池、数据湖、数据海洋,以及谁可以利用它们来构建这些模型,公平的结果是什么?
Well, and that's precisely the problem number three for Google, which is that if Google always gives you the answer, then what is the future for the open web? Because Google is going to disintermediate the third-party content creator. Like if you're a regular mom and pop shop with your bakery on a website or you have a blog post and you rely on that display ad income, well Google's just going to give you the perfect recipe, so why would you ever go to that kind of third-party blog post? And that's actually a problem for Google and the regulator, because Google has been telling the regulator for the best part of 15 years that the reason it can crawl all of these websites is because it's only indexing so that it can redirect the user to the third-party page. It feels fair, right? It's a Yellow Pages, they always used to say, it's a lookup table. Whereas if it's now cutting out that source of information and giving you the perfect answer, that's a big problem with the regulator, certainly in the European context, because many Google execs have been on the witness stand claiming that they'll never do that, right? So it's now the models are doing that. They've been trained on the web, it's obvious, it's been proven. You used to be able to ask OpenAI ChatGPT like, 'Hey, where is this answer trained from?' It would actually tell you some of the training data. I think it doesn't do that now. What's the fair outcome here for pools of data, lakes, oceans of data, and who gets to leverage them to build these models?
但你认为结果会如何?因为我们开始看到诉讼堆积。我们看到埃隆说:‘嘿,Twitter 数据不可用。’Reddit 说数据可以付费获取。Cor 说可以付费获取,或者可能附带链接。Stack Overflow 构建了自己的语言模型。这位新 CEO 刚给我发邮件说:‘听着,我知道 Stack Overflow 经常被提及,我们正在构建自己的副驾驶。没有其他人可以使用我们的数据集。’所以谈谈你认为行业会发生什么。因为我觉得拿 Gourmet 或任何食谱数据库,然后直接给出答案,至少不提供引用,是非常不公平的。会发生什么?
But what do you think is the outcome here? Because we're starting to see the lawsuits pile up. We're starting to see Elon say, 'Hey, Twitter data is not available.' Reddit saying it's available at a price. Cor saying it's available at a price or maybe with a link back. Stack Overflow built their own language model. This new CEO just emailed me to say, 'Look, I know Stack Overflow keeps coming up, we're building our own co-pilot. Nobody else can use our data set.' So talk a little bit about what you think will happen in the industry. Because I feel like it's tremendously unfair to take Gourmet or whatever recipe database and then just give the answer and not give a citation at least. What's going to happen?
事情很棘手。我的意思是,现实是信息被放在开放网络上,开源爬虫引擎在完全合法、可接受的条款下收集了这些信息。那个爬虫,Common Crawler,收集信息并明确说明这些信息将用于研发目的,并供其他试图在开源搜索引擎之上构建产品的人进行实验。所以爬虫,每个人收集的爬虫数据,只是一个既定的现状。所以我不认为这会被推翻,或者会有任何补偿。人们有时会谈论数据信托的想法,每个数据贡献者得到一分钱之类的。我的意思是,这根本不会发生。
Here's the tricky thing. I mean, the reality is that the information was placed on the open web, and the open-source crawling engines have gathered up their information under perfectly legal, acceptable terms. And that crawler, the Common Crawler, collects the information and clearly says that it would be used for research and development purposes and used for experimentation by other people trying to build other products on top of the open-source search engine. So the crawler, the crawling data that everyone's collected, is just a well-established status quo. So I don't think that is going to be undone or there's going to be any compensation. People sometimes talk about this data trusts idea where each individual data contributor gets like one cent or something. I mean, this is just not going to happen.
为什么不会?
Why not?
执行起来太难了。我认为不可能产生足够的收入来支付给数据材料的最终生产者,对吧?所以也许在非常大的数据所有者的情况下,比如 OpenAI 刚刚与 Axel Springer 达成了协议,对吧?但那实际上不是针对历史数据,而是针对新鲜的实时新闻。看,我认为这里有可能性。如果我们作为一个行业集体认为这实际上可能是一种好处。记得法国的 Minitel 曾经按小时收费,并与数据源分享收入。AOL 曾经每小时收费三四五美元,CompuServe 也是,它们与数据提供商分享收入。所以如果你在一个与婚礼等相关的数据网站上,他们每小时给你 50 美分,对吧?实际上那里有一个模式。我认为如果我们采用 robots.txt 并加入许可证,说:‘嘿,听着,这是我的食谱。我是戈登·拉姆齐。如果你想把它们放入你的索引,每个食谱每年最低支付 10 美元。每年 1 万美元才能放入你的索引,再加上一些额外费用。’不管是什么,这可能足以激励人们开始在网上发布更多食谱。
Too hard to execute on. I think it's impossible to generate sufficient revenues to make the payment to the end producer of data material, right? So maybe in the case of a very large data owner, like OpenAI just did a deal with Axel Springer, right? But that's actually not for historic data, that's actually for fresh, real-time news. See, that's where I think there is a possibility of this. If we think as an industry collectively that this could actually be a benefit. Remember Minitel in France used to charge a certain amount per hour and they would share that with the data sources. AOL used to charge three, four, five bucks an hour, CompuServe, and they would share that with the data provider. So if you were on some data site that had to do with weddings or whatever, they would just give them 50 cents of the hour, right? You actually had a model there. I think if we took robots.txt and we put in a license and said, 'Hey listen, these are my recipes. I'm Gordon Ramsay. If you want them in your index, there's a minimum payment each year of $10 a recipe. It's $10,000 a year to put it into your index, plus I want something on top of it.' Whatever it is, and that might be enough to incentivize people to start putting more recipes online.
有可能,有可能。我的意思是,我认为这些事情的挑战在于,创意工具现在将变得非常广泛可用,模型将更擅长生成新食谱。所以秘密已经泄露了。
It's possible, it's possible. I mean, I think the challenge of these things is that the creative tools are now going to be so widely available that the models are going to be better at generating new recipes. So the cat's out of the bag.
确实如此。
It's true.
是的。那么告诉我,你离开谷歌,与里德·霍夫曼共同创立了 Inflection,他刚上过播客。你筹集了一大笔钱。Inflection AI 是什么?目标是什么?你显然近距离亲眼目睹了 OpenAI 和 DeepMind 发生的一切。你在那些精英 AI 产品中处于什么位置?这里有 Bard,有 OpenAI……
Yeah. So tell me, you leave Google and you start Inflection with Reid Hoffman, who just on the pod. You raised a bunch of money. What is Inflection AI? What is the goal here? You obviously got to see everything up close and personal that's happened with OpenAI and with DeepMind. Where do you sit in that sort of pantheon of elite AI offerings? You got Bard over here, you got OpenAI...
那么,你打算定位在哪里,想要开拓什么市场?
A over here, where are you going to sit and what market are you going to try to carve out?
我们正在开发一款个人 AI。我相信未来会有很多不同类型的 AI:商业 AI、医疗 AI、法律 AI。每个数字影响者都会有自己的 AI,每个试图卖东西的品牌和大平台也会有自己的 AI。在未来五年里,你看到的每个网站或应用都会变成一个对话界面,你完全可以称之为 AI。它能生成视频、文本、音频,并像我现在和你说话一样与你交谈。在那个一切都变成 AI 的世界里,我认为你作为个体消费者,想要一个站在你这边、对你的利益负有信托责任的个人 AI。它帮你找信息、识别可靠来源、与其他 AI 讨价还价、规划你的一天、跟进你的研究兴趣、找有趣的信息。它必须高度个性化,因为你最终会分享大量敏感的个人隐私信息,这样它才能出去代表你——无论是在元宇宙的游戏环境中,还是替你查找体育新闻。我想象这就像每个人都有一位数字参谋长——协调员、日程安排员、优先级设定员、摘要员。你早上醒来,它给你一份完美的简报,告诉你一天的所有安排、新闻动态、体育赛事、你关注的公司情况。
We're developing a personal AI. I believe there are going to be lots of different types of AIs: business AIs, medical, legal. Every digital influencer will be an AI, every brand and big platform trying to sell stuff will have their own AI. Wherever you see a website or an app, expect that in the next five years it's going to become a conversational interface that you might as well call an AI. It'll be able to produce video, text, audio, and talk to you just as I'm talking to you now. In that world where everything becomes an AI, I think you as an individual consumer want to have a personal AI that is on your team, fiduciary aligned to your interests, in your corner. It helps you find information, identify credible sources, negotiate with other AIs for the best bargains, plan and prioritize your day, follow up on your research interests, find entertaining information. It's super important that it's personalized to you because you're going to end up sharing a lot of sensitive personal intimate information so it can go out and be your representative, whether in a gaming environment in the metaverse or looking for sports news on your behalf. I think of it like imagine if everybody had a digital chief of staff—a coordinator, scheduler, prioritizer, summarizer. You wake up in the morning and it gives you the perfect briefing of everything you've got on in your day, what's happening with the news, sports, the companies you're tracking.
这让我想起了 General Magic。索尼和一家叫 General Magic 的公司早在 Palm 之前就做了一款 PDA 设备。有部纪录片讲过这个。他们在搜索引擎出现之前就有了智能体的概念。智能体会代表你去查航班、订位、执行任务。所以你认为这个 AI 在某种程度上是自主的,可以执行重复性任务。比如,“我想减肥,目标 165 磅,我该怎么做?”然后它每天给你建议。如果把 AI 放在那个位置,它会非常有效,因为它不会唠叨抱怨;它会鼓舞人心、让人安心、温柔、礼貌且尊重人。除非你想要别的风格,那它也可以变得奇怪。
That reminds me of General Magic. Sony and a company called General Magic made a PDA device long before Palm. There was a documentary on it. They had a concept of agents before search engines existed. The agent would go on your behalf to find flights, reservations, do tasks for you. So you see this AI as being autonomous in some ways, able to be put on repeat tasks. For example, 'I'm trying to lose weight, I want to be 165 lbs, what should I be doing?' and it counsels you every day. If you put an AI in that kind of position, it will be incredibly effective because it's not going to nag and moan; it's going to be inspiring, reassuring, gentle, polite, and respectful. Unless you want something else, then it can get weird.
可能会变得奇怪,但你说得对,它会是个性化的,会是你的智能体。这个框架至关重要:它是代表你工作的智能体,而不是代表公司、OpenAI、必应或谷歌搜索。这是你的 AI,你和它说的任何内容,我们都无法得知。如果你把数据放进去,我们不会与广告商或任何其他人分享。
It could get weird, but to your point, it's going to be personalized and it's going to be your agent. This framework is critically important: it's your agent working on your behalf, not the corporation's, not OpenAI's, not Bing or Google search results. This is your AI, and whatever you talk to it about, we don't have any insights into. If you put data into it, we're not sharing that with advertisers or anybody else.
所以这意味着我每年要付你 100 美元?
So that means I have to pay you $100 a year for this?
归根结底,如果你想要完全信任,你就不能成为产品。如果你不付费,那就是别人在付费。如果你把那么多注意力和敏感信息放在一个地方,确保它站在你这一边的唯一方法就是你自己以某种方式付费。你不会去找你的会计师,结果发现他是由保险公司资助的,而你还想让他帮你报税或做投资决策。你会想,“你是为我工作,还是想推销某种保险产品?”
At the end of the day, if you want to have full trust, you need to not be the product. If you're not paying for it, somebody else is paying for it. If you're putting that amount of attention and sensitive information into a place, the only way to make sure it's on your team is for you to pay for it in some way. You wouldn't go to your accountant and find out they're funded by an insurance company while you're trying to get your tax return done or decide on investments. You'd wonder, 'Are you working for me or are you trying to sell some insurance product?'
我认为理解意图和商业模式非常关键。现在的消费者非常精明;他们明白如果你不付费,你就是产品。即使 Alexa 实际上并没有监听他们并推送广告,他们也理解这个概念。
I think understanding the intent and the business model is so critical. Consumers are super savvy now; they understand that if you're not paying, you're the product. Alexa listening to them and serving ads, even when that's not what's actually happening, they get that concept.
我用过一点,非常令人愉快,设计精美。在早期阶段,你觉得哪些市场或任务会让人们感到惊喜?
I've used it a bit, quite delightful, beautifully designed. What can we expect to be the beachhead markets or tasks that you think it's going to delight people with in the early days?
到目前为止,我们只发布了一个小模型到生产环境。我们公司成立不久,大约 15 个月前,我们正在搭建超级集群。几个月前我们完成了一轮大规模融资,正在建设目前全球最大的 H100 集群。今天我们已经拥有最大的运营中 H100 集群。到年底,我们将拥有 22,000 块 H100,相当于在一个集群中拥有约 88 万块 A100。
So far, we've only got a small model that's shipped in production. We founded the company only a short while ago, about 15 months ago, and we're just bringing up our super cluster. A few months ago we raised a pretty large round, and we're building out the largest cluster of H100s in operation in the world today. Today we have the largest operational cluster of H100s. By the end of the year, we will have 22,000 H100s, which is equivalent to about 880,000 A100s in a single cluster.
英伟达,你是不是得在他们门口等着求他们卖给你?谈谈这些 H100 的稀缺性吧。
Nvidia, you had to go wait on their doorstep and beg them to buy these? Talk a little bit about the scarcity of these H100s.
它们非常稀缺。英伟达是我们的投资者之一,微软也是。我们很幸运能成为他们供应链的顶端客户,他们对我们很好。我们还帮助他们优化了集群以运行 MLPerf,这是一个压力测试集群的开源基准测试。过去六个月我们投入了大量精力优化他们的集群。这是一种很好的互惠关系:我们既是小白鼠,也是首批大规模出货的受益者。
They're extremely scarce. Nvidia is one of our investors, and so is Microsoft. We were just very lucky to get to the top of the supply chain with them, and they've been great to us. We've also helped them optimize their cluster for MLPerf, an open-source benchmark that stress tests their cluster. We've invested a huge amount over the last six months to optimize their cluster. It's a good quid pro quo: we were both the guinea pigs and the beneficiaries of the first big shipment.
当这 13 亿美元融资宣布时,人们有点困惑,因为微软在 OpenAI 上下了大赌注,现在又在这里下大赌注。我们应该如何理解微软投资你和 OpenAI 的行为?这让人有点困惑。
When this $1.3 billion raise was announced, it was a little confusing to people because Microsoft has this big bet on OpenAI, and then they're making this big bet here. What should we take away from Microsoft's behavior here, investing in you and OpenAI? It was kind of confusing for folks.
我认为可以这样理解:微软是一个平台的平台。它历来擅长与大量第三方做交易,并与各种不同的供应商互动。他们很可能也会这样处理这件事。
I think the way to think about it is that Microsoft is a platform of platforms. It's traditionally been very good at doing deals with lots of third parties and interacting with a whole range of different suppliers. That's probably how they're going to approach this.
这 13 亿美元中有多少用于硬件?只是好奇。所以你就直接把它交给英伟达,然后建一个巨大的数据中心?
How much of the 1.3 billion goes to hardware? Just out of curiosity. So you just ship it right to Nvidia and build out this gigantic data center?
是的,这将是一个巨大的优势,因为我们将在世界上其他人之前训练出比 GPT-4 大得多的模型。最迟明年春天,甚至可能更早。所以基本上所有的钱都花在了算力上。我们只有 40 个人。
Yeah, it will be a huge advantage because we will train models that are very much larger than GPT-4 before anybody else in the world. By the spring for sure, maybe even a little bit earlier. So all of it goes to compute basically. We're only 40 people.
你有没有考虑过使用谷歌云、亚马逊云服务或 Azure?还是你需要控制硬件才能获得所需的收益?
Did you consider using Google Cloud or Amazon Web Services or Azure? Or do you need to control the hardware to get the gains you need?
不,我们不会用 TPU;它们在其他方面有困难。但我们肯定想要英伟达。我们确实看过 AWS 和 Oracle,实际上我们也用 Azure 处理一些工作负载。但我们想确保我们为 H100 设计了架构,并在最底层优化了所有操作,以获得最大性能。
No, we wouldn't use TPUs; they're difficult for other reasons. But we certainly wanted Nvidia. We did look at AWS and Oracle, and we actually do use Azure for some workloads. But we wanted to make sure we designed the architecture for the H100s and optimized everything down to the lowest levels to get maximum performance.
建立这个集群需要多长时间?
How long does it take to build out this cluster?
需要一两年时间。我们目前有 7000 块 H100 在运行,到赛季初我们将有 22000 块全面运行。
It's going to take a year or two. We're currently operational with 7,000 H100s, and we'll have 22,000 fully operational by the beginning of the season.
这太不可思议了。这只是分布在世界各地的不同数据中心吗?你们托管然后开始上架?
That's unbelievable. And this is just in different data centers around the world? You collocate and start racking them?
不,只是一个数据中心,因为我们需要它们都在同一个地方。实际上有三个足球场那么大。
No, it's just one data center because we need it all in the same place. It's actually the size of three football pitches.
数据中心在哪里?我很好奇。
Where is the data center? I'm curious.
在美国。我不想说具体位置。
It's in the US. I don't want to say exactly where.
肯定得靠近水电或核电站。
Got to be near something with hydroelectric or nuclear power.
没错,就是这样。
Exactly, that's exactly what it is.
当我们看 AI 的负面影响时,有岗位压缩。我问过很多聪明人,几乎所有人都说效率提高了 30%。这意味着每两年人们的工作能力就会翻倍。还有一些可怕的情景:人们用这个来黑客攻击或制造超级生物武器。你对这些有多担心?
When we look at the downside of AI, there's job compression. I ask a lot of smart people, and almost universally they say 30% more effective. That means every two years people become twice as good at their job. Then there are scary scenarios: people using this to hack or build super biological weapons. How concerned are you about each of those?
这是个好问题。我非常担心。我整个职业生涯都在研究 AI 的伦理和安全。事实上,我们 2010 年的商业计划就有‘构建安全且合乎道德的 AI’这个口号。所以我认为我们从一开始就看到了很多风险。我不同意很多时间线;人们非常焦虑,认为智能爆炸会带来生存风险。但我刚写了一本书叫《即将到来的浪潮》,探讨了 AI 在未来 10 到 15 年可能带来的威胁,以及合成生物学的威胁。我认为劳动力市场风险是真实的。未来 10 年,人们的生产力会提高,但挑战在于生产力的提高会产生剩余价值,这些价值会被资本而不是劳动力捕获。所以我们可能看不到中产阶级工资的平均增长——那些从事认知体力劳动、后台管理、基本电话客服、新工厂工人的人。我们过去对工厂工人失业很不屑,但现在轮到白领了,人们就‘等等,你能做出比设计师更好的标志了。’
It's a good question. I'm very concerned about it. I've worked on the ethics and safety of AI my entire career. In fact, our business plan back in 2010 had the strap line 'building safe and ethical AI'. So I think we saw a lot of these risks from the outset. I don't agree with a lot of the timelines; people are very anxious about an intelligence explosion presenting existential risk. But I've just written a book called 'The Coming Wave' that looks at threats AI might create over the next 10 to 15 years, as well as synthetic biology threats. I think the labor market risk is real. For the next 10 years, people will get more productive, but the challenge is that the increases in productivity will generate surplus value captured by capital, not labor. So we probably won't see an average increase in wages for the middle class—those doing cognitive manual labor, back office administration, basic telephone calls, new factory workers. We were very dismissive about factory workers losing their jobs, but now that it's white-collar, people are like 'wait a second, you can make a logo better than the designer.'
我们现在就有点这样了。你可以做出和营销机构一样好甚至更好的标志或标语。不用天才也能看出我们不需要那么多营销机构或标志创作者了。
We're kind of there right now. You could create a logo or tagline as good or better than a marketing agency. It doesn't take a genius to say we won't need as many marketing agencies or logo creators.
完全正确。一切都会变得更便宜,这很棒。这将推动人类历史上最大的生产力爆发。但现实是,那些失业的人无法及时再培训、适应并在劳动力市场中与人和机器竞争。如果你是一名设计师,世界上只有那么多设计岗位。如果突然 70%的人借助 AI 能做那份工作,那么设计师就会被挤出,不得不做下一级的工作,从而挤掉下一级。我看不出底层如何能足够快地适应。这就是为什么有一个艰难的对策:如果你不想出现严重的结构性失业,就必须有某种再培训补贴、全民基本收入或类似的东西。在完全实现全民基本收入之前,你不必走那么远,但这是 20 年内的方向。
That's exactly right. Everything is going to cost less, which is amazing. It will drive the biggest productivity explosion in the history of our species. But the reality is that those who have their jobs displaced will not be able to retrain, adapt, and compete against man plus machine in the labor market in time. If you're a designer, there are only so many design slots. If suddenly 70% of humans can do that work aided by AI, then designers will be squeezed out and have to do the next tier down, squeezing out the next tier. I don't see how the bottom can adapt quickly enough. That's why there's a tough remedy: if you don't want significant structural disemployment, there has to be some kind of subsidization for retraining, UBI, or something. Before full UBI, you don't have to go that far, but that's the direction over 20 years.
而且成本……某一天,Panera Bread 和麦当劳会说,我们为什么需要收银员?放一个进去,然后其他所有事情都在自助点餐机上完成。那个岗位就被淘汰了。那些人可以去找其他工作。播客现在也是一份工作。关键在于速度以及我们如何管理这个过渡,因为我们会创造新的工作。会有新的需求。人们会因为这种生产力提升而获得新的收入。所以人们会有钱花,人们会更高效,从而用更少的工作产出同样的成果。所以问题在于如何管理未来几十年这个过渡期,那些被挤出劳动力市场的人必须重新培训和适应。我的意思是,很明显存在再培训和适应的需求,而且很多人就是跟不上。我们从工厂开始,从煤矿工人开始。如果你是一个煤矿工人,20 年来只干这个,45 岁时突然神奇地学会编程或成为博主,这很难想象。但话说回来,有了 AI 辅导,也许机会来了——AI 辅导会变得非常好,以至于人们可以通过个性化教育更快地学习技能。
And the cost... at some point Panera Bread and McDonald's were like, why do we need a cashier? Put one in, and then everything else is going to be ordering on a kiosk. And that job has been eliminated. Those people can go find other jobs. Podcasting is a job now. It's the speed and how we manage the transition, because we will create new work. There will be new demand. People will have new income because of this productivity boost. And so people will have money to spend, and people will be more efficient, so they could deliver the same output with less work. So the question is how you manage the transition for this period, you know, the next couple decades, where people who get pushed out of the workforce have to somehow retrain and adapt. I mean, even you know, it's pretty clear that there's a retraining and adaptation requirement, and that many people are just not going to be able to keep up. We started with factories, we started with coal workers. You know, if you're a coal worker and that's all you've done for 20 years, the idea that at 45 years old you're going to just magically learn to code or become a blogger is kind of hard to think. But then again, with AI tutoring, maybe there's an opportunity that the AI tutoring will get so good that people can actually learn skills faster with customized education.
没错,没错。我的意思是,人们,这太不可思议了:这将是一个非常精英主义的时刻,因为很多拥有两三代、三四代稳定家庭的人继承了生活中的和平与稳定,这极大地促进了他们的教育,给了他们自信,给了他们情感支持,给了他们受教育的机会、获得机会的机会。现在将要发生的是,那些一直处于舒适轨道上的人将面临更渴望成功的人的竞争,这些人现在可以获得个性化的 AI 导师,它们会教你任何你痴迷的东西,任何你想深入的东西。它无限耐心,无限聪明,它完全了解你喜欢怎么学习,而且它是免费的,基本上免费或接近免费。
Right, right. I mean, people, this is the incredible thing: it will be a very meritocratic moment, because a lot of people who have had safe and steady families for two, three, four generations have inherited peace and stability in their life, and that has turbocharged their education, it's given them confidence, it's given them emotional support, it's given them access to education, access to opportunities. What's going to happen now is that those people who've been on a comfortable trajectory are going to face competition by people who are hungrier and who now have access to personalized AI tutors that are going to teach you anything that you're obsessed by, anything that you want to go deep on. It's infinitely patient, it's infinitely smart, it knows exactly how you like to learn, and it's free, and it's going to basically be free or close to free.
是的,接近免费。我的意思是,与大学教育相比,它肯定是免费的。与那相比,基本上只需要一个互联网连接。然后你就触及了第三轨,那就是动力和驱动力。这对人们来说是一个非常艰难的对话,但可能的情况是,在斯里兰卡、巴基斯坦、圣保罗的某个人比圣地亚哥或布鲁克林的某个人更渴望,他们会更努力,花更多时间在那个 AI 上。现在它是一个全球性的 AI 导师,一个全球性的市场。我认为这对人们来说会很可怕:'天哪,我在和全球前 5%的人竞争,他们现在有星链,有高速互联网连接,还有云端的 AI 导师,那个无限耐心的可汗学院老师。' 是的,你在西方的特权——你出生在伦敦或纽约——在这种情况下毫无意义。
Yeah, close to free. I mean, compared to a college education, it's going to be free for sure. I mean, compared to that, it's just basically getting an internet connection. And then you're going towards the third rail, which is motivation and drive. And this is a very hard conversation for people to have, but it might be the case that there's somebody in Sri Lanka, Pakistan, São Paulo who wants it more than somebody in San Diego or Brooklyn, and they're just going to work harder and they're going to spend more time on that AI. And now it's a global AI tutor, and it's a global marketplace. And that, I think, is going to be very scary for people: 'Oh my God, I'm competing against the top 5% on a global basis who now have Starlink, have a high-speed internet connection, and they've got the AI tutor in the cloud, the Khan Academy teacher that is infinitely patient.' And yeah, that's your privilege in the West that you were born in London or New York means nothing in that scenario.
对,是的。我的意思是,我在书里有一整节关于这个,我真的很喜欢写。它讲的就是那个故事,因为生产成本正在暴跌,一切现在都将变成零边际成本。所以知识广泛可得,对吧?现在不仅仅是知识,还有智能,对吧?智能是综合知识并将其转化为新策略、见解或行动计划的方式。如果那变成零边际成本,那么为什么任何人都不能超级有创造力呢?这真的取决于你个人有多渴望、多有活力,我认为这将真正取代,或者说,削弱或给西方那种有点占据我们的自满阶层施加压力。那是一群因为继承关系而进入大学的精英。如果你是一个有继承关系的人,进了哈佛,猜怎么着?这可能不如一个来自班加罗尔、有动力成为神经外科医生或开发者的人那么有意义,他们比你更渴望。现在他们将对社会非常有用。
Right, yeah. I mean, I have a whole section about that in the book, which I really enjoyed writing. I mean, it's about exactly that story, because the costs of production are going through the floor, and everything is now going to be zero marginal cost. So knowledge is widely available, right? And now not just knowledge, but intelligence, right? Intelligence being the mode of synthesizing knowledge and turning it into new strategies or insights or action plans. If that goes to zero marginal cost, then why shouldn't anybody be able to be super creative? And it really is going to be about how hungry and dynamic you are as an individual, which I think is going to really displace, or you know, it's going to undermine or put some pressure on the complacency class that has kind of taken over us a little bit in the West. It's the group of elites who get into their college because they're legacy. And if you're a legacy person and you get into Harvard, guess what? It may not mean as much as the person who is motivated and becomes a neurosurgeon or a developer, and they're from, like I said, Bangalore, and they just wanted it more than you. And now they're going to be super useful for society.
听着,这太棒了。谢谢你给我超过一个小时的时间。一定要让你再来。每个人都应该试试……是 Pi,对吧?个人助理的名字?P AI 是简称……Pi。Pi。Pi。Pi。所以 Pi 代表个人智能。
Listen, this has been great. Thank you for giving me over an hour of your time. Got to have you come back. Everybody should try... it's Pi, right? Is the name of the personal assistant? P AI is the short... Pi. Pi. Pi. Pi. So Pi stands for personal intelligence.
是的,Pi。
Yeah, Pi.
这本书叫《The Coming Wave》,我没想到你已经出书了。我这周就读,现在就去订有声书。书什么时候出的?
And the book's called The Coming Wave, which is... I didn't realize you had the book. I'm going to read it this week and I'm going to order the audio book now. When did the book come out?
实际上现在就可以预购,9 月 5 日正式出版。
It's actually available for pre-order now, comes out September the 5th.
太棒了,完美。等我读完后,一定请你回来好好聊聊。希望我们能在硅谷为你办个新书派对之类的。如果你需要世界上最棒的主持人在任何新书派对上采访你,告诉我,我有空。
Oh fantastic, so perfect. Well after I read it, I'll have to have you come back on and we'll talk all about it. And hopefully we have a book party or something for you here in the Valley. If you need the world's greatest moderator to interview you at any book parties or something, let me know, I'm available.
而且显然如果我想尽快结婚,你也能帮忙,对吧?所以目前世界上最棒的钓鱼活动是可行的,只要你能找到一个愿意嫁给你的女人,Mustafa。但你已经有创业公司了,哦你已经搞定了。
Well also apparently if I fancy getting married anytime soon, you're available for that too, right? So currently the world's greatest fishing is available if you can find a woman who will marry you, Mustafa. But you got a startup, oh you already got that accomplished.
不不,我还在挣扎。我单身得很。所以如果你想让我和我的创业公司结婚……你已经和你的创业公司结婚了。你筹集了十亿美元,我可以告诉你未来十年你和谁结婚。绝对是 Inflection AI。你们那边有 40 个人。而且你们在招人,所以如果你想加入 Inflection,就去 inflection.ai。
No no, I'm struggling with that. I'm very much single. So if you want to marry me to my startup... you are married to your startup. You raise a billion dollars, I can tell you who you're married to for the next 10 years. Absolutely, Inflection AI. And you're 40 people over there. Also you're hiring, so if you want to join Inflection, go to inflection.ai.
听着,这是一个精英团队。他们在湾区、帕洛阿尔托、伦敦。你相信人们应该在办公室工作,还是认为远程工作也可以?你对此怎么看?
And listen, it's an elite group. They're in the Bay Area, Palo Alto, London. You believe in people working out of an office or you think remote work is fine? What's your take on all this?
实际上我找到了一个有趣的平衡。我认为你需要两全其美。我们的运作方式是整个公司以六周为一个周期。当你加入公司时,你同意在第七周的聚会期间,无论你在世界哪个角落,都要亲自出差一整周。这是日程的关键部分,因为在那第七周聚会的一周里,我们会有一个非常紧张的 hackathon 式聚会,经典模式是每天 14-16 小时在同一个房间里,真正全力以赴。而在其余六周,我们认识到人们需要灵活工作。所以我本人每天都去,大概公司三分之一的人也这样。另外三分之一在周二、周三、周四来。还有些人完全是远程。所以我认为这是正确的混合结构。
I have an interesting balance actually. So I think that you need the best of both worlds. The way that we operate is that we run the entire company on a six-week cycle. When you join the company, you sign up to traveling to be in person for a full week wherever you are in the world for our seventh week meetups. That's a key part of the schedule because for that one week in our seventh week meetup, we have a very intense hackathon-style meetup where it's the classic 14-16 hours a day in the same room, really going out hardcore. And the rest of the six weeks, we recognize that people need to work in a flexible way. So I am personally in every day, and so is probably about a third of the company. Another third come in Tuesday, Wednesday, Thursday. And some people are actually fully remote. So that's the right hybrid structure, I think.
我同意你的看法。我的观点是,三分之一的人作为远程工作者效率更高,过去两年我找出了这些人。另外三分之二的人在办公室和他人一起工作时表现更好。就像有些人独自跑步更好,而另一些人跟团队一起跑时,大多数人跟跑得比你快的人一起跑会表现更好。就这么简单。或者你和比你强的扑克玩家玩,你会进步更快。所以我认为大概有 25%到三分之一的人更适合远程,但三分之二的人更适合线下。那些人不可能永远完全远程。我个人不相信完全远程的环境,所以我们采用 6-1 的节奏,我认为这恰到好处。基本上是适度的牺牲,能创造一定的团队精神。
I agree with you. My belief is that a third of people are more productive as remote workers, and over the last two years I figured out who they are. Then there's another two-thirds that do better work when they're in an office with other people. Just like some people are better runners alone, and other people when they run with a group, the majority of people when you run with a group, you will perform better when you're running with runners who are faster than you. It's that simple. Or you play with poker players better than you, you're going to get better quicker. So I think there might be on the margins 25% to a third who are better remote, but I think two-thirds are better in person. And there's no way that those people can stay completely remote forever. I personally am not a believer in these fully remote environments, so that's why we do the 6-1 rhythm, and I think it's the right amount. Basically it's the right amount of sacrifice, it creates certain esprit de corps.
我能看到聚在一起非常激励人心。而且,有些人有孩子、有家庭。你招聘的是非常成功、有很多选择的人,所以你不能总是发号施令。你可能有个天才想住在太浩湖,你也许能让她抽出一周来,但你可能没法让她来七周。你不想失去那个人,对吧?我认为这就是我们现在面临的奇怪僵局,或者说是员工和公司之间的妥协。因为你不想失去一个高绩效者。
I could see it being super motivating to get together. And yeah, some people have kids, they have family. You're hiring people who are uber successful and have many options, so it's not like you can always dictate. You might have somebody who's just a genius who wants to live at Lake Tahoe, and you may be able to break her off for a week to come, but you might not get her for the seven weeks. So you don't want to lose that person, right? I think that's the weird standoff we now have, or maybe it's a settlement amongst workers and corporations. Because you don't want to lose a high performer.
没错,完全正确。所以我的意思是,提供灵活性是正确的做法,在周期中给一些人需要的平静——不是所有人,但有些人需要做自己的事——然后有超级紧张的聚会,我认为这是一个好的节奏。是的,这听起来也很可持续。你知道,我在想我们行业早期,每个人都一周工作 6 天,每天 12 小时、14 小时。这带来了很多不可思议的成果,所以我不认为这样做的人一定错了,但它会让人崩溃。它可能会把一些高绩效者排除在团队之外。所以管理层的真正任务是找到合适的节奏。听起来你找到了适合你的节奏。实际上听起来挺令人兴奋的。
Right, exactly. So I mean, getting the flexibility is the right way to do it, and having the kind of peace during the cycle that some people need—not everyone, but some people need to do their own thing—and then have the super intense meetup, which I think is a good rhythm. So yeah, it's also sounds sustainable. You know, I was thinking about the early days of our industry, like just everybody at work 6 days a week, 12 hours a day, 14 hour days. It led to a lot of incredible outcomes, so I don't think anybody who does it is making a mistake necessarily, but it can break people. It can exclude certain people from the team that might be high performers. So it's really the job of management to just figure out a cadence. Sounds like you found the cadence that works for you. Sounds kind of exciting actually.
现在效果很好,我们玩得很开心。所以,如果你的听众想加入我们,我们正玩得开心。我的意思是,另一件事是:精英人才有选择,你必须让它有趣,而且要有使命感,对吧?去这些偏远地点待一周听起来很有趣。好了,听着,干得好。期待读这本书。9 月 5 日出版。大家都去预购。再告诉我一次书名。
It's working right now and we're having a great time. So yeah, if any of your listeners want to come and get stuck in, we're having a great time. I mean, that's the other thing: people have choice amongst elite folks, you got to make it fun and it's got to be purpose, right? And it sounds like a lot of fun to go to these remote locations and do a week. So all right, listen, great job. Look forward to reading the book. Comes out on September 5th. Everybody pre-order it. Tell me the name one more time of the book.
《The Coming Wave》。
It's The Coming Wave.
《The Coming Wave》。所以去亚马逊或 Audible 上找,现在就预购。如果你听到我的声音,请预购,这样他就能在第一周获得大销量。你需要一万本才能登上《纽约时报》畅销书排行榜。这是真的。下周见。再见。
The Coming Wave. So go look for that on Amazon or Audible and pre-order right now. If you hear my voice, please pre-order so he gets that big first week bump. You need 10,000 in order to make the New York Times bestseller list. That's true. I'll see you all next week on this starts. Bye-bye.