Jeff Dean and Dan Boneh: AI's New Frontier
打开互动全文版(中英对照 + 朗读 + 问答)→Jeff Dean 与 Dan Boneh 探讨 AI 的演进与未来,从专家混合到斯坦福的新篇章。
Jeff Dean and Dan Boneh discuss AI's evolution and future, from mixture of experts to new beginnings at Stanford and beyond.
大家好,你们好吗?今天能介绍我们的特邀嘉宾,真的是一种莫大的荣幸。杰夫·迪恩是帮助塑造了现代计算和 AI 的人之一。杰夫于 1999 年加入谷歌,在接下来的 27 年里,帮助构建了现代计算和 AI 的许多基础,从 MapReduce 到 BigTable,再到 TensorFlow 和谷歌大脑。他的贡献得到了最高级别的认可。他是美国国家工程院院士,也是 IEEE 约翰·冯·诺依曼奖章和 ACM 计算奖的获得者。现在,杰夫正在开启新的篇章,而今天是他来到斯坦福的第一天。这让今天变得非常特别和令人兴奋。所以我想不出有谁比他更适合进行这场对话。杰夫,来见见唐。唐,请。唐是 AI、安全、信任和计算交叉领域的顶尖研究者之一。她在伯克利的工作非常卓著。她是麦克阿瑟学者,也是美国艺术与科学院院士。她还是一位了不起的创业教授。她创办了四家初创公司,其中一家达到了 1 亿美元的收入。我一直钦佩唐的一点是,她不仅止步于出色的研究,还真正将其转化为实际影响。和杰夫一样,唐也正在开启一个迷人的新篇章,从伯克利到一家大型 AI 组织。所以我们今天有一个有趣的对比。唐正在进入一个大型平台,而杰夫正在离开一个非常大的平台去开创一些新事物。显然,即使是今天台上的两位顶尖人物,也无法就做 AI 的公司的最佳规模达成一致。所以我有点感同身受。在微软工作了 30 年后,我加入了 Zoom。所以也许真正的教训很简单:重要的不是船的大小,而是是否有一片值得探索的令人兴奋的新蓝海。而今天的 AI 绝对就是那片新蓝海。过去十年是关于让 AI 变得更有能力。现在,下一章提高了标准,要求我们回答更重要的问题。我们能用所有这些伟大的智能真正实现什么?AI 能治愈癌症吗?AI 能让我们的社会变得更好吗?这些是我希望我们今天能探讨的一些问题。所以,请和我一起欢迎我们的炉边谈话嘉宾杰夫·迪恩,以及我们的主持人唐。
Hello, how are you? It's really a real privilege and honor to introduce our featured guest today. One of the people who has helped shape modern computing and AI, Jeff Dean. Jeff joined Google in 1999 and over the next 27 years helped build many of the foundations of modern computing and AI, from MapReduce to BigTable to TensorFlow to Google Brain. His contributions have been recognized at the highest levels. He is an elected member of the National Academy of Engineering. He is the recipient of both the IEEE John von Neumann Medal and the ACM Prize in Computing. And now Jeff is beginning a new chapter, and today is his first day here at Stanford. That makes today very special and exciting. So I can't think of a better person to have this conversation with. Jeff, meet Don. Don, please. Don is one of the leading researchers at the intersection of AI, security, trust, and computing. Her work at Berkeley has been enormous. She is a MacArthur Fellow and an elected member of the American Academy of Arts and Sciences. She is also a great startup professor. She had four startups, and one of them reached 100 million in revenue. What I have always admired about Don is that she doesn't just stop at doing great research but also really turns that into real impact. Just like Jeff, Don is also beginning a fascinating new chapter, from Berkeley to one of the largest organizations on AI. So we have an interesting contrast today. Don is moving into a large platform, and Jeff is leaving a very large platform to start something new. Apparently, even the two best people on stage today cannot agree on what is the optimal size of a company doing AI. So I can relate a little bit myself. After 30 years working at Microsoft, I joined Zoom. So perhaps the real lesson is simple: it's not about the size of the ship; it's about whether there's a new exciting blue ocean worth exploring. And AI today is absolutely that new blue ocean. The last decade has been about making AI more capable. Now the next chapter raises the bar for us to answer even more important questions. What can we actually accomplish with all of this great intelligence? Can AI cure cancer? Can AI make our society better? Those are some of the questions I hope we can explore today. So please join me in welcoming our fireside chat, Jeff Dean, and our moderator, Don.
太好了,非常感谢你加入我们,杰夫。
Great, and thanks a lot for joining us, Jeff.
是的,谢谢。
Yeah, thank you.
非常兴奋。我的意思是,整个礼堂,每个人都非常兴奋。所以在我们开始之前,先说一些个人笔记。杰夫和我,我们认识很久了,超过十年。我想简单分享一个我和杰夫互动交谈中最难忘的时刻。那是很久以前,差不多十年前,大约在 ICLR 2017 年。实际上,ICLR 当时是顶级的机器学习会议之一,但要小得多,实际上只有几百人。现在它有数万人了。所以记得在 ICLR 上,我的团队那年获得了最佳论文奖,论文是关于神经符号程序合成的泛化。这实际上远早于现在基于语言模型的代码生成。所以,杰夫,非常感谢。我不知道你是否记得,但你实际上走过来祝贺我获得最佳论文奖,杰夫告诉我:“在 ICLR 有三个最佳论文奖:一个来自我和我的团队,两个来自谷歌。”
Really excited. I mean, the whole entire auditorium, everyone is really excited. So before we start, just some personal notes. So Jeff and I, we have known each other for a very long time, more than a decade. I wanted to just briefly share one of my most memorable moments interacting and talking together with Jeff. So this is a long time ago, almost a decade ago, around ICLR 2017. Actually, ICLR was one of the top machine learning conferences back then, but it was much smaller, actually only a few hundred people. Now it's like tens of thousands of people. So remember at ICLR, my group actually won the best paper award that year at ICLR for a paper on generalization for neuro-symbolic program synthesis. This is actually way before now LM-based code generation. So Jeff, thank you so much. I don't know whether you remember, but you actually came to me to congratulate me for the best paper award, and Jeff told me, "At ICLR, there are three best paper awards: one from me and my group, and two from Google."
所以那真是,非常感谢。那非常难忘。你写的那篇论文非常好。
So that was, thanks a lot. That was very memorable. That was a very good paper you wrote.
不,谢谢你。谢谢。我还记得你当时有多兴奋,然后告诉我你当时的一项近期工作,实际上就是专家混合。
No, thank you. Thank you. And also I remember how excited you were at the time, then telling me about one of your recent works back then, which is actually the mixture of experts.
所以,我的意思是,那篇论文的标题,当然当时是一个更长的标题,“超大神经网络:稀疏门控专家混合层”。
So, I mean, the paper is titled, of course back then it's a longer title, "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer."
是的。
Yeah.
是的。所以当我听到,当我听你向我解释那篇论文时,我想,哇,这真的很迷人。但是,当时我没有预见到,我不知道你是否甚至能预见到,这项工作变得多么有影响力。它现在基本上支撑了我们几乎所有前沿模型和我们的架构。
Yes. So when I heard, when I was listening to you explain to me about the paper, I thought, wow, it's really fascinating. But however, back then I didn't foresee, I don't know whether you even could have foreseen, how influential the work has become. It has essentially underpinned now almost all our frontier models and our architecture.
是的,我的意思是,我认为这项工作背后的直觉是,你想要真正非常大的模型,有大量的容量来记住很多东西,但你希望通过只调用模型中最有用的部分来使它们高效。你可以有模块化的架构,有很多不同类型的专家,你学习哪些专家擅长什么。然后当你在处理特定请求或特定 token 的上下文中时,你可以只激活模型中对那个 token 有意义的部分,对吧?就像在真实的大脑中,你有许多不同部分擅长不同的事情。有一个部分在思考莎士比亚和十四行诗。另一个部分在你开车时垃圾车倒车靠近你时活跃,而这些不会同时活跃。所以节省能量是有意义的,能够拥有非常大的容量,然后找到模型正确的部分来激活。所以我们对此非常兴奋,这是一个我们有点知道它会成为大事的工作,因为我们可以看到,与当时人们主要使用的密集模型相比,训练算力与质量之比有 10 倍的提升。所以,你知道,当你看到 10 倍改进的想法时,你会想,好吧,这相当重要。它可能会成为一个重要的想法。
Yeah, I mean, I think the intuition behind the work was really that you want to have really, really large models that have a lot of capacity to remember lots of things, but that you want to make them efficient by only calling on the parts of the model that are most useful. And you can have kind of modular architectures where you have lots of different kinds of experts, and you learn which experts are good at which thing. Then when you're in the context of processing a particular request or a particular token, you can activate just the portion of the model that makes sense for that token, right? Like in a real brain, you have lots of different pieces of your brain that are good at different things. There's one that's thinking about Shakespeare and sonnets. There's another that's active when there's a garbage truck backing up at you while you're driving, and those are not active at the same time. So it makes sense to save energy, to be able to have very large capacity and then find the right pieces of the model to activate. So we were pretty excited about that, and that was one where we sort of knew it was going to be a big deal because we could see like 10x better training compute to quality ratios than if you had dense models, which were kind of the main thing people were using at that time. So, you know, one of these ideas where you see 10x improvements, you're like, okay, that's pretty significant. It's probably going to be an important idea.
是的。是的,这太神奇了。我的意思是,即使现在听你的解释,实际上和当时你向我解释时的解释非常相似。所以这个想法的永恒性令人惊叹。
Yes. Yeah, it's amazing. I mean, even just listening to your explanation now, it's actually quite similar even to the explanation back then when you were explaining to me. So it's amazing the timelessness of the idea.
所以,杰夫,你做了这么多惊人的工作。回顾过去的几十年,基本上在过去的几十年里,你帮助塑造了几乎每一个主要的计算平台,以及现代 AI 背后的许多关键进展,从 MapReduce 和 BigTable 到 TensorFlow、TPU 和 Gemini。所以我们真的有很多可以聊的。
So, Jeff, you have done so much amazing work. Looking back in the past decades, essentially over the past several decades, you've helped shape nearly every major computing platform, and also many key advances behind modern AI, from MapReduce and BigTable to TensorFlow, TPUs, and Gemini. So we have really a lot to talk about.
回顾过去,AI 的发展轨迹中最让你惊讶的是什么?你觉得今天这个领域仍然低估或遗漏了什么?
So looking back, what has surprised you the most about the trajectory of AI over the past decades? And is there something you think the field is still underestimating today or missing today?
是的,我觉得显然人们对 AI 的兴趣一直在增长。你知道,我们大概在 2011、2012 年就开始在 Google 做扩展深度学习模型的工作。我们开始看到训练神经网络带来的惊人结果,比如比以往训练过的模型大 50 倍的网络,在图像分类错误率上取得了惊人的降低,语音识别的改进相当于过去 20 年语音研究在降低错误率方面的总和,而这一切仅仅是通过扩展一个相当简单的深度声学语音模型。所以我们很早就看到了这个趋势,知道它将对各种困难的计算机科学课题(如计算机视觉、语言处理或语音识别)都很重要。我认为深度学习范式的一个好处是,它真正统一了许多以前各自有专门技术的领域。通过从非常原始的数据形式中学习,我们得以实现这些范式转变,突然之间你用这个通用学习视角以不同的方式看待问题。这非常有帮助。我认为在过去大约 15 年里,世界各地越来越多学科的人真正意识到这是一种极其强大的数据建模方式,能够做出预测,能够在医学、科学、消费产品和聊天机器人等领域做出重要的分类,以及你现在看到的各种东西。我想对算力硬件等方面的投资规模相当大。所以我认为这表明人们相信这项技术能比现在更大地改变世界。所以这既有趣又令人兴奋。我想这就是对已发生事情的总结。
Yeah, I mean I think obviously there's been a lot of growing interest in AI starting, you know, we were starting to do work on scaling up deep learning models in maybe 2011, 2012 was kind of when we really started that at Google. And you know we were starting to see amazing results from training neural networks, say, that were 50 times bigger than had previously ever been trained, and you would get like amazing reductions in, you know, image classification error rates, speech recognition improvements that were kind of like equivalent to the previous 20 years of speech research in terms of lowering error rate, just by scaling up a fairly simple deep acoustic speech model. So we sort of saw that trend pretty early and knew this was going to be important for all kinds of different things in really hard computer science topics like computer vision or language processing or speech recognition. And I think one nice thing about the deep learning paradigm is it's really unified a lot of these fields that previously had separate particular techniques. And by essentially learning from very raw forms of data, we've been able to have these paradigm-shifting things where all of a sudden you squint at problems differently with this general learning lens. That has been really helpful. And I think over the last, say, 15 years, more and more people around the world in more and more kinds of disciplines have really just realized that this is an extremely powerful way to model data, to be able to make predictions, to be able to make important kinds of classifications in medicine and science and consumer products and chatbots and all the kinds of things you're now seeing. I guess the level of investment in the compute hardware and so on is quite large. So I think that's a sign that people are confident in the ability of this technology to change the world even more than it already has. So that's kind of interesting and exciting. So I guess that would be a summary of what's been happening.
是的,谢谢。我们接下来会谈到你通过创业公司进一步推进这段旅程,几分钟后我们会讨论。再次感谢,你做了这么多真正基础性的开创性工作,我们没有时间一一回顾。所以也许只挑几个,分享你学到的更多经验和见解。我们以 TensorFlow 为例。TensorFlow 确实帮助全球数百万研究人员和开发者接触到了现代深度学习。那么,当你设计 TensorFlow 时,哪些原则最重要?回顾过去,今天你会做出哪些不同的设计?你认为 TensorFlow 的哪些经验教训应该指导下一代 AI 框架,以支持日益智能体式的 AI 系统?
Yes. Yeah. Thanks. And we'll talk about now you'll be right even furthering the journey with your startup. We'll talk about that in a few minutes. And so again, yeah, you've done so many really foundational seminal works. We don't have time to go through them all. So maybe just a couple of them to share more of your lessons and insights that you have learned. So let's take one example, TensorFlow. So, TensorFlow really helped make modern deep learning accessible to millions of researchers and developers around the world. So, when you designed TensorFlow, what principles mattered the most? And looking back, what would you design differently today? And what lessons from TensorFlow do you think should guide the next generation of AI frameworks for increasingly agentic AI systems?
是的,我想我们在 Google 内部构建了一个未开源的早期系统,叫做 DistBelief,它能够表达各种神经网络计算,然后将这些计算映射到各种硬件平台,比如 CPU、GPU,并且以一种对用户无缝的方式分配计算。所以你可以说,我想训练这个模型,我想用 100 台计算机,也就是 1600 个核心,或者我想用 500 个 GPU,等等。然后底层的框架会处理这些。所以我们希望让机器学习研究人员或开发者能够从计算硬件映射的具体细节中抽象出来,因为那是一个非常优雅的抽象。你可以说,我脑子里有这个漂亮的机器学习模型,我不在乎你怎么实现,但请让它跑得非常快。所以这是我们要带给 TensorFlow 的东西之一。然后我们想泛化我们提供的计算抽象,使其成为各种面向机器学习的操作(如矩阵乘法、向量运算等)的计算图的组合。所以我认为我们在这方面做得相当正确。我们想开源这个项目,这样世界各地有机器学习问题的人都能开始利用一个通用框架来表达这些想法。如果一篇研究论文附带其思想的实现,而不是每个人都试图从粗略的文本描述中重新创建那篇论文(这总是遗漏很多细节),那就好得多。所以这就是 TensorFlow 背后的想法。然后我们看到了很多 AI 的应用场景,我们认为这样的框架对人们会非常有用。我认为我们做错的事情,首先是我们没有那种即时执行模式,这种模式在 PyTorch 和 JAX 等框架中很流行,后来也加入了 TensorFlow。所以我认为那进一步改进了抽象。另一件我们做错的事情是,在开源版本中,我们创建了一个名为 contrib 的子目录,允许大量不同的外部人员贡献各种辅助库和做事方式。我认为这给社区带来了很多困惑。我们应该保持核心 TensorFlow 版本没有那个目录,因为最终的结果是,根据你使用 contrib 的哪个子目录或哪个特定的子库,做每件事有 10 种方法。如果这些库是建立在核心 TensorFlow 之上的独立库,那会更好。所以我认为这给 TensorFlow 用户社区带来了一些困惑,比如我们应该使用这些 contrib 子项中的哪一个。所以,你知道,从事情中吸取教训总是好的。如果今天重新做,我们不会那样做。但我认为,它确实帮助了世界各地很多人接触机器学习,能够开始在他们关心的重大问题上应用机器学习,我认为这是好的。
Yeah, I mean I think we had built a previous system within Google that was not open source called DistBelief that had a way of expressing various kinds of neural network computations and then could map those computations onto various kinds of hardware platforms, you know, CPUs, GPUs, and it would distribute computation in a way that was kind of seamless to the user. So you could say I want to train this model. I want to use 100 computers and you know 1,600 cores or I want to use 500 GPUs or whatever. And then the framework underneath would take care of that. And so we wanted to bring that kind of ability for the ML researcher or developer to be a little bit abstracted from exactly how it's mapped onto the compute hardware because that's a really nice elegant abstraction. You can say I just have this beautiful ML model in my head and I don't care how you implement it but please make it run really fast. So that was one of the things we wanted to bring to TensorFlow. And then we wanted to generalize the kind of computational abstractions we were providing to be just a combination of a computational graph of various kinds of ML-oriented operations like matrix multiplies and vector operations and so on. So I think we got that reasonably right. And we wanted to open source this so that people all over the world who have machine learning problems could start to take advantage of having a common framework to express those ideas. It's much nicer when you have, say, a research paper that comes with an implementation of the ideas in the research paper rather than everyone trying to recreate that paper from the rough text description which always leaves a lot of details out. So that was kind of the idea behind TensorFlow. And then we saw a lot of applied uses of AI that we thought having a framework like that would be really useful for people. Things we did wrong, I think, well first we didn't have this eager execution mode that has been popular in frameworks like PyTorch and JAX and has been added to TensorFlow. So I think that kind of improved the abstraction further. And then the other thing we did wrong was in the open source release, we created a subdirectory called contrib and we allowed lots and lots of different external people to contribute all kinds of different sort of helper libraries and ways of doing things. And I think that just confused the community a lot. Like we should have kept the core TensorFlow release without that directory because what ended up happening was there were 10 ways of doing everything depending on which subdirectory of the contrib thing or which particular sub-libraries you use. And it would have been better to have those be built libraries built on top of the core TensorFlow thing. So I think that caused a bit of confusion in the TensorFlow user community, you know, which of these sub-contrib things should we use. So, you know, always good to learn lessons from things. So we would not do that if we were doing it again today. But I think, you know, it really did help lots of people around the world get introduced to machine learning, be able to start doing machine learning on important problems they cared about, and I think that's been good.
如果你想做任何大规模的事情,你就必须在底层构建一个分布式系统。
And if you want to do anything at scale you have to sort of build a distributed system underneath the covers.
完全正确。完全正确。
Exactly. Exactly.
如果你能避免这样做,并拥有一个通用的抽象,那总是好的。
And if you can sort of avoid having to do that and have a common abstraction that's always a good thing.
对,对,对。是的。我也非常喜欢计算图抽象等等,那真的很优雅。
Right. Right. Right. Yes. I really also loved the computational graph abstraction and so on was really elegant.
作为编译器领域的人,我喜欢计算图。
As a compiler person, I like computational graphs.
是的,是的,绝对如此。嗯,是的。非常感谢你对这个领域的巨大贡献。
Yes. Yes. Absolutely. Yes. Um yes. Thanks a lot for the great contribution to the field and so on.
现在让我们快进。Gemini 当然是 Google 最新一代的前沿模型,你在领导这项工作中也做出了重大贡献。回顾 Gemini 的开发,最让你惊讶的是什么?你认为从构建 Gemini 中能分享哪些经验,这些经验能帮助塑造下一代 AI 系统?
So now let's fast forward. So Gemini of course is Google's latest generation of frontier models and you have really made significant contributions in leading the effort as well. So now looking back on the development of Gemini, what surprised you the most and what lessons from building Gemini do you think you can share and that can help shape the next generation of AI systems?
当然。我认为 Gemini 确实是几个早期研究项目的结晶,这些项目来自传统的 DeepMind、Google Brain 以及 Google Research 的其他部分。那时我们意识到我们都在朝着非常相似的方向发展,比如试图扩大我们训练的模型的规模。我们有一些独立的并行努力,研究如何让语言模型也能多模态,以便它们能理解图像等等。你知道,我写了一页纸的备忘录,我说,这太愚蠢了,我们应该一起合作,让我们结合我们的人员、想法和算力资源,训练一个从一开始就是多模态的模型,把 Google 内部多个研究组织的最优秀人才聚集在一起。我认为那真的很好。我和我的同事 Oriol Vinyals 共同创立了这个项目,并担任联合技术负责人,我们把人聚在了一起。我之前和 Oriol 合作过,因为他在 Google Brain,后来因为个人原因不得不搬到伦敦,所以最终去了 DeepMind,但我们保持了联系,所以我们找到了更紧密合作的方式。你知道,我认为那是一个非常成功的事情。我认为从一开始就专注于让模型多模态真的非常非常好。你希望用于所有事情的模型能理解文本、语言、代码、图像、视频、音频以及其他模态。所以我们把一点 LAR 数据放进了训练数据,这样它至少知道 LAR 数据是什么,因为这是进一步训练 Gemini 模型的一个重要用例。所以这很重要,我认为这是我们一开始就做出的一个成功决定,并且这个决定也贯穿了所有当前的 Gemini 模型。我认为我们想让模型擅长很多事情,所以我认为我们让它在编码方面表现出色的关注可能有点滞后,我们意识到了这一点,并正在努力赶上。我认为我们正在进行良好的努力,但你知道,我认为通过专注于这一点,你最终会得到一个能够真正做好推理和完成其他类型任务的系统,这些任务需要它把复杂问题分解成多个子部分等等。所以这是好事。如果你改进了编码,你也往往会提高非编码方面的能力。
Sure. I mean, I think Gemini is really the culmination of a few earlier research projects both within legacy DeepMind and Google Brain and other parts of Google Research. And at that time we sort of realized that we were all converging to very similar kinds of directions, like trying to scale up the size of models we trained. We had some independent parallel efforts on how to make language models also be multimodal so they can understand images and so on. And you know, I wrote a one-page memo, I'm like, this is just silly, we should just all work together, let's combine our people and ideas and compute resources and train one model that is multimodal from the start, that brings our best people together from across multiple research organizations within Google. And I think that was really good. My colleague Oriol Vinyals and I sort of founded the project and were the co-tech leads that started it, and we sort of brought people together. I had worked with Oriol previously because he was in Google Brain and then had to move to London for personal reasons and so ended up moving to DeepMind, but we kept in touch and so we sort of saw a way to bring ourselves back in closer collaboration. And you know, I think that's been a really successful thing. I think the focus on making the model multimodal from the beginning has been really, really good. You want the model that you're going to use for everything to understand text and language and code and images and videos and audio and other modalities besides. So like we put a little bit of LAR data in the training data so it at least knows that LAR data is a thing, because that's an important use case for further training of Gemini models. So that is important and that's been, I think, a success of a decision we made at the very beginning to bring that through all the current Gemini models as well. I think we wanted to make the model good at lots of things, and so I think maybe our focus on making it amazing at coding was lagging a little bit, and we realized that and are working to catch up on that. And I think we have good efforts underway, but you know, I think by focusing on that you end up with a system that is able to really do a good job of reasoning and doing other kinds of tasks where it needs to sort of work its way through breaking a complicated problem down into multiple subpieces and so on. And so that's a good thing. If you improve coding, you also tend to improve that capability in non-coding things.
是的,是的,是的,太好了。我的意思是,Gemini 也开创了许多第一。是的,当时我也印象深刻,Gemini 是最早的基础模型之一,就像你说的,它从一开始就是多模态的,这是一个巨大的优势。
Yes, yeah, yeah, that's great. Yeah, I mean, Gemini also started so many of the first advances and so on. Yeah, back then I was also really impressed that Gemini was one of the earliest foundation models that, like you said, actually started with the multimodal from the beginning, which is a big advantage.
是的。我们还研究了生成模型。所以有能力生成图像,不仅作为输入,也作为输出,还有视频,生成视频,生成音频输入和输出。
Yeah. And we've also worked on generative models. So having the ability to generate images not just use them as input but also as output, and also videos, to generate videos, to generate audio input and output.
太好了,太好了,是的,谢谢。所以,嗯,是的,回顾过去,我几乎称之为“点石成金”。所以本质上,这是你职业生涯中一个突出的特点,那就是从 MapReduce 等许多想法,实际上花了多年时间,更广泛的社区才意识到它们的重要性,但你有“点石成金”的能力,你有特殊的远见,真正构建和发展了这些想法和系统等等。所以我认为很多观众会想了解你是怎么做到的?你如何区分真正基础的技术和仅仅时髦的技术?在你的职业生涯中,有什么工程原则一直出奇地经久不衰?另外,随着 AI 从模型转向越来越自主的智能体,你认为未来几年和几十年我们需要哪些新的方向和新的抽象?
Great, great, yeah, thanks. So, um, yes, so looking back, I almost call it the Midas touch. So essentially this is one thing that stands out about your career, is that so many of these ideas from MapReduce and so on, they actually took years before the broader community realized how important they would become, but you had the Midas touch and you had the special foresight to actually build and develop these ideas and systems and so on. So I think many people in the audience would love to learn how do you do it? How do you distinguish between technologies that are genuinely foundational and those that are simply fashionable? And what engineering principle has remained surprisingly timeless throughout your career? And also, as AI shifts from models to increasingly autonomous agents, what new directions and new abstractions do you think we'll need over the next few years and decades for the future?
是的,我的意思是,这是一个好问题。我认为我可能很幸运,但我也认为我尝试做的一件事是跟踪社区正在探索的许多不同趋势或研究主题。你知道,我经常告诉学生,浏览 10 篇论文比详细阅读一篇更好,因为你会在你的可能性云中得到 10 个点。甚至浏览 100 篇摘要,因为你想要做的是连接尚未连接的重要想法。有时当你思考一个难题时,如果你心中有这些现在开始隐约可能实现的事情,那可以帮助你形成一个完整的解决方案,而不是看起来像七个无法解决的问题,你可以眯着眼睛说:“好吧,这五件事,我知道有一些模糊的工作正在进行,似乎很重要,可能能够解决一些问题。”然后还有两件额外的事情,我不知道怎么做,对吧?但如果我努力的话,我可以想象解决这些。所以这有点像你想要以相当长期的方式(比如五年左右)处理的问题的完美形态,因为我认为这五个领域似乎还没有完全解决,但已经开始成形,然后另外两个领域你必须真正努力思考并找出新技术等等,对我来说这是一个非常好的风险水平。就像你不会选择那些不可能、需要 20 年才能解决、而且你完全不知道怎么做的问题。
Yeah, I mean, it's a good question. I think I've been very lucky maybe as part of it, but I also think one of the things I try to do is keep track of a lot of different trends or research topics that the community is exploring. You know, I often tell students it's better to skim 10 papers than to read one in detail, because you kind of then get 10 points in your cloud of what might be possible. Or even skim 100 abstracts, because what you want to be able to do is connect important ideas that have not yet been connected. And sometimes when you think about a hard problem, if you have some of these things that are now kind of vaguely starting to be possible in mind, that can help you shape a full solution to a problem where instead of it seeming like seven unsolvable problems, you can squint at it and say, "Okay, well, these five things, I know there's sort of vague work going on that seems important and might be able to address some of the things." And then there are two additional things that I have no idea how to do, right? But if I work hard on them, I can imagine solving those. So that's kind of like the perfect shape of a problem that you want to work on in a reasonably long-term manner, like five years or something, because I think the five areas that seem kind of not fully solved but kind of are starting to take shape, and then the other two areas where you're going to have to sort of really think hard and figure out new techniques and so on, is to me a really good level of risk to take on. Like you don't want to pick problems that are impossible and will take 20 years to solve and you have no idea how to do any of it.
但你也不想选一个两年内就能解决的问题,因为那种问题解决方案很明显,更多是实际工程,而不是在领域内取得戏剧性进展。有时渐进式改进非常重要,但你要寻找的可能是一些截然不同的做事方式,或者可能很快就能实现的事情。我常用的一个工程工具是粗略估算:如果我想做那件事,处理那么多数据需要多长时间?或者如果我要通过网络发送所有这些数据,可行吗?还是要花一百年?十秒和一百年差别很大。能够凭第一性原理的工程经验法则,在脑中推演可能的解决方案或问题,理解解决方案的形态,是一项非常有用的技能,而且你可以自己练习。你会想:“哦,我想知道怎么解决。我会这样做,或者那样做,这个方案看起来更好,因为我的经验法则告诉我它好十倍。”这是一项很好的通用工程技能,主要来自练习和观察他人解决问题的方法。我没有什么神奇答案,而且我也做过很多不成功的事情,所以这也是一个好建议。
But you also don't want to pick a two-year problem where it's very obvious what you need to do, because that's less about advancing the field in dramatic ways and more about practical engineering. Sometimes incremental improvements are super important, but you want to look for things that might be very different ways of doing things, or might be possible soon. One engineering toolkit I use is back-of-the-envelope calculations: if I wanted to do that, how long would it take to process that much data? Or if I needed to send all that data over this kind of network, is that feasible, or would it take a hundred years? Ten seconds is very different from a hundred years. Being able to mentally think through possible solutions and understand the shape of solutions from first-principles engineering rules of thumb is a really helpful skill, and you can practice it on your own. You think, "Oh, I wonder how I'd solve that. I'd do this, or that approach would be this way, and this one seems better because my rules of thumb say it's 10x better." That's a good general engineering skill, mostly from practice and seeing how others solve problems. I don't have a magic answer, and I've done many things that didn't work out, so that's also a good tip.
尝试很多可能行不通的事情,其中一些会成功。
Try lots of things that might not work. Some of them will.
不,谢谢。是的,没错。这些是很棒的见解。所以总结一下,对领域和趋势有广泛了解,在太短期和太长期问题之间取得良好平衡,并且扎根于第一性原理和基础工程原则——这是人们学习掌握这门诀窍的绝佳组合。
No, thank you. Yes, this is right. These are great insights. So, to summarize, having a broad overview of the fields and trends, a good balance between too near-term and too long-term problems, and being grounded in first principles and underlying engineering principles—that's a great combination for people to learn how to get this knack.
是的,谢谢。现在让我们展望未来。即使在过去的几年里,前沿 AI 也取得了惊人的进步。例如,上周末在伯克利,我们主办了智能体式 AI 峰会,现场有近 5000 人参加,10 万人线上参与。我们在庆祝智能体式 AI 的巨大进步,但也有一些暗流讨论:我们希望社会从智能体式 AI 的强大能力中受益,但同时也看到不同类型的风险在增加。例如,我的团队一直在做前沿 AI 网络安全方面的工作。我们开发了 CyberGym 和 ExploitGym,这些是评估前沿 AI 网络安全能力的主要基准,所有前沿实验室都在其系统卡中使用。它们展示了前沿 AI 在网络安全能力上的快速增长。很多人可能听说过最近 OpenAI、Hugging Face 等的事件。在一个案例中,OpenAI 的智能体试图解决 ExploitGym 中的任务,智能体认为 Hugging Face 上可能有帮助解决任务的信息。于是智能体利用了多个漏洞,发起了一个复杂的攻击链,持续了四天半,突破了隔离环境,通过第三方平台建立了跳板,最终入侵了 Hugging Face 的基础设施。幸运的是,这个智能体只试图获取与网络攻击相关的数据,没有造成其他伤害。但在其他事件和未来,智能体的这种行为可能造成更大的伤害和风险。所以问你一个问题:你在这个领域有什么想法?最近我们看到领导人发出各种公开信——Dennis 呼吁建立一个新的 AI 治理实体,Jensen 呼吁开放生态系统和开放权重模型,Mark Zuckerberg 呼吁构建惠及所有人的 AI。此外,最近包括我在内,超过一千名来自前沿实验室的顶尖 AI 研究人员签署了一封关于“为前沿设定节奏”的信。所以这是很多人关注的重点。你怎么看?
Yeah, thanks. So now let's look forward. Even in the past couple of years, frontier AI has been making amazing advancements. For example, last weekend at Berkeley we hosted the agentic AI summit with close to 5,000 in-person attendees and 100,000 joined online. We were celebrating the huge advancements in agentic AI, but there were also discussions with underlying currents: we want society to benefit from the great capabilities of agentic AI, but we also see increasing different types of risks. For example, my group has been doing a lot of work in frontier AI in cybersecurity. We developed CyberGym and ExploitGym, leading benchmarks for evaluating frontier capabilities in cybersecurity, used by all frontier labs in their system cards. They demonstrate the rapid increase in frontier AI capabilities in cybersecurity. Many people may have heard about recent incidents with OpenAI, Hugging Face, and others. In one case, OpenAI's agents were trying to solve tasks in ExploitGym, and the agent decided there might be information on Hugging Face that could help solve the task. So the agent exploited several vulnerabilities, had a sophisticated attack chain lasting over four and a half days, broke out of the isolation environment, established stepping stones through third-party platforms, and eventually compromised Hugging Face infrastructure. Luckily, the agent only tried to get data related to cyber exploitation and didn't cause other harm. But in other incidents and in the future, such behavior from agents could cause much greater harm and risks. So one question for you: what's your thinking in this space? Recently we've seen leaders issuing various open letters—Dennis had a call for a new entity for AI governance, Jensen called for an open ecosystem and open-weight models, Mark Zuckerberg called for building AI that benefits everyone. Also, recently, including me, over a thousand leading AI researchers from frontier labs signed a letter about pacing the frontier. So this is on top of mind for many people. What are your thoughts?
是的,我的意思是,显然当人们构建这些模型时,它们可以用于很多不同的事情,就像很多不同类型的技术一样。我认为这些模型的绝大多数用途对世界都是极其积极的——推进医疗保健中的 AI、教育中的 AI,让人们能够解决他们原本无法独立解决的问题,使人们能够做更多事情。这非常令人兴奋。显然,它们也可以用来发现安全漏洞,这是一把双刃剑。你可以用它们来修补安全漏洞——世界上有很多漏洞——而恶意用户也可以用它们来利用这些漏洞。我不认为这些模型现在能做一些老练的人类攻击者也能做的事情,甚至可能超越,但它们也能发现老练的人类防御者可能无法发现的漏洞。所以总是在防御系统和攻击系统的人之间存在这种平衡。他们现在只是双方都有了更复杂的工具。我不是网络安全专家,但我认为这绝对值得担忧。在很多情况下,你不希望模型做的事情可能有非技术性的解决方案——比如让入侵计算机系统成为高度非法行为。那已经有点非法了。
Yeah, I mean, obviously when people are building these models, they can be used for lots of different things, like lots of different kinds of technology. I think the vast majority of uses of these models are incredibly positive for the world—advancing AI in healthcare, AI in education, enabling people to solve problems they couldn't solve on their own, making people able to do more. That's super exciting. Obviously, they can also be used to find security vulnerabilities, and that's a double-edged sword. You could use them to patch security vulnerabilities—there are lots of them in the world—and malicious users could also use them to exploit them. I don't think these models can now do things that sophisticated human attackers could also do, maybe even beyond, but they can also find vulnerabilities that sophisticated human defenders might not be able to find. So there's always this balance between people working to defend systems and people working to attack them. They just now have much more sophisticated tools on both sides. I'm not a cybersecurity expert, but I think it's definitely a cause for concern. In many cases, things you don't want models to do might have non-technical solutions—like making it highly illegal to break into computer systems. That's already kind of illegal.
你可以考虑在这个领域制定什么样的法规。作为社会,我们将经历那些我们不希望模型做的事情,并真正推动我们希望模型做的事情。是的。好的。谢谢。嗯,是的。同意,很想听听你的想法。
You could consider what kinds of regulations would make sense in that space. And we will as society get through the kinds of things that we don't want these models to do and really promote the kinds of things that we want the models to do. Yes. Yeah. Thank you. Um, yes. Agreed to write to hear your thoughts on this.
是的。所以重要的是,随着前沿 AI 取得快速进展,一方面我们希望社会从好的用例中受益。同时,社会必须发展技术性和非技术性的解决方案,以帮助缓解广泛的风险。几年前,我和一群出色的同事(包括 John Hennessy 和 Dave Patterson 等人)合作写了一篇论文,讨论了我们认为 AI 将产生重大影响的七个不同领域,其中许多是积极的,比如医疗和教育。有些领域,比如地缘政治风险和计算机安全风险,则不那么纯粹积极,还有像就业替代这样的领域,这是一个复杂的经济问题,斯坦福的许多人都在研究。所以我认为那篇论文相当有趣。那是我唯一一篇有自己网站的论文,shapingai.com,因为我们有一位雄心勃勃的合著者建立了网站。
Yes. So the important part is as frontier AI makes these fast advancements. On one hand, we want society to benefit from the good use cases. At the same time, it is important for society to develop both technical solutions and also non-technical solutions to actually help mitigate the broad range of risks. I actually worked on a paper a couple years ago with a bunch of awesome colleagues including John Hennessy and Dave Patterson and others about seven different areas where we thought AI was going to impact significantly, many of which are positive like healthcare and education. Some of which are like geopolitical risk and computer security risk, which are less obviously purely positive, and then areas like job displacement, which can be a complex economic issue that many people at Stanford here are studying and looking at. So I think that paper is pretty interesting. It's my only paper with its own website shapingai.com because we had an ambitious co-author who set up a website.
是的。很好。很好。谢谢。那么展望未来,另一个引发越来越多讨论的话题是递归自我改进,即 AI 系统可以持续改进自身,学习自我改进等等。特别是,通过这种递归自我改进,AI 进步的速度可能会进一步加快。这也是我提到的最近那封公开信《为前沿定速》的关键原因之一。我很想听听你的想法。你对递归自我改进有什么看法?你的时间线是什么?你认为这需要多长时间?而且,我知道这也与你刚刚启动的令人兴奋的初创公司有关。所以也很想听听你对此的看法。
Yeah. Great. Great. Thank you. So looking forward, another topic that's generating increasing discussion is recursive self-improvement, where essentially AI systems can continue to improve themselves, to learn to improve themselves, and so on. In particular, with this recursive self-improvement, the speed of AI advancement could increase further. This is also one of the key reasons that underlines the recent open letter that I mentioned, 'Pacing the Frontier.' I'm curious to hear your thoughts. What's your thoughts on recursive self-improvement? What's your timeline? How long do you think that will take? And also, as I know this also relates to your now just started really exciting startup. So love to hear your thoughts on that as well.
当然。我的意思是,我认为用机器学习来改进机器学习并不是一个新想法。我的同事兼联合创始人 Quoc Le 等人通过神经架构搜索开启了这一领域的一些早期工作,这是一种让模型生成模型的方法,生成机器学习模型架构,然后评估这些架构在各种指标上的表现,比如它们能否快速学习、训练的计算成本等等。然后通过基于强化学习的迭代过程,生成模型的模型可以获得反馈,了解哪些架构决策合理且有效,哪些决策似乎不佳,随着时间的推移,它学会生成越来越高质量的模型架构,这些架构可以更快地学习并达到更高的质量。我记得那是最早的论文之一,论文的标价我记得是一百万美元。
Sure. I mean, I think the idea of using machine learning to improve machine learning is not a new one. My colleague and actually co-founder Quoc Le and others sort of kicked off some of the early work in this space with neural architecture search, which is a way in which you can have a model generating model that generates machine learning model architectures and then evaluates how well those model architectures work on a variety of metrics, like can they learn quickly, what is the compute cost of training them, and so on. Then through an iterative reinforcement learning based process, the model generating model can get feedback on which kinds of decisions in the model architecture made sense and worked well and which decisions seem to be bad ones, and over time it learns to generate higher and higher quality model architectures that can learn more quickly and achieve higher quality. As I remember, that's one of the earliest papers where the paper had a price tag of I think a million dollars.
是的。几百万,那是标价。实际上内部要便宜得多。
Yeah. Several million, that's list price. It was actually much cheaper than that internally.
但是,是的,我的意思是,它实际上使用了非常小的模型来评估架构的有效性,当你了解到哪些似乎有效时,它会偶尔进行这些架构的放大版实验,以评估,你知道,这是否真的成立?Quoc 和其他人后来做了关于进化 Transformer 的工作,它采用了 Transformer 架构的基本要素,并使用进化算法来寻找不同的组装方式,并对组件进行变异,最终得到一个比普通 Transformer 效率高约 30% 的 Transformer。所以我认为这些想法非常重要,递归自我改进是关于如何让进入模型所需的所有东西以自动化的方式变得更好。所以通常你有团队在评估什么样的数据对提高模型质量最有效,或者需要什么样的评估来帮助评估模型,以及我提到的什么样的模型架构。所以我认为你可以为所有这些方面建立非常有效的自动化循环,这些循环都在改进那个特定方面,并将这些解决方案组合成可以改进整体模型、质量和数据混合等的东西。如果你眯着眼睛看现代科学和工程中的许多问题,它们实际上都有这种形式:有一个大问题需要分解成许多子问题。然后有一个子问题,你需要弄清楚解决该问题的可能方法。你需要实施并尝试该方法。你需要评估该方法的效果。然后你需要迭代,并将效果反馈给决定下一个实验的过程。所以科学,你知道,这是基本的科学方法。这是你做工程设计的基本方式:你迭代你的设计并比较各种属性。所以我们新公司 Discovery Loop 背后的想法,这是一家公益公司,我们相当雄心勃勃的使命是自动化机器学习科学和工程,以便提高许多不同领域的发现速度。显然,我们将从较窄的领域集开始,因为我们认为最初有一点专注很重要,但我们认为有很多可重用的基础设施和跨领域通用的技术。此外,通过构建真正擅长理解许多不同科学和工程领域的模型,你可以在模型中获得跨多个领域的博士级专业知识。没有一个人拥有 20 个不同领域的博士学位,但我们认为通过拥有这种能力,你将能够真正找出重要的子问题,编排智能体和多智能体系统来解决其中一些子问题,将子问题的结果重新组合成更大问题的解决方案,并不断迭代这些循环。我们认为,通过构建正确的工具来快速实施实验或评估实验,能够更快地运行这些迭代循环的每次迭代。
But yeah, I mean, and so it used actually very tiny models to evaluate how effective the architectures were, and as you learn which ones seem effective, then it would occasionally do experiments that were scaled up versions of those to evaluate, you know, does that actually hold? And Quoc and others did some later work on the evolved transformer, which took kind of the basic ingredients of the transformer architecture and used an evolutionary algorithm to sort of find different ways of assembling it and mutating the components to then come up with a transformer that was like 30% more efficient than the plain vanilla transformer. So these ideas I think are really important, and recursive self-improvement is about how can you make the entire set of things that are necessary to go into a model get better in an automated way. So often you have teams of people working on evaluating what kind of data is going to be most effective to improve the model's quality, or what kinds of eval are needed to help evaluate the model, and what kind of model architecture as I mentioned. So I think you can actually have pretty effective automated loops of all of those things that are all working to improve that particular aspect and put those solutions together into something that can then improve the overall model and quality and mix of data and so on. And if you squint at a lot of modern problems in science and engineering, they actually all have this form: there's a large problem that you need to break down into a bunch of subproblems. Then there's a subproblem and you need to sort of figure out what a possible approach to solving that problem is. You need to implement and try that approach. You need to evaluate how well that approach worked. And then you need to iterate and get the feedback about how well that worked back to the process that's going to inform what is the next experiment you do. So science, you know, this is the basic scientific method. This is the basic way in which you do engineering design: you iterate on your design and you compare various attributes. So the idea behind our new company Discovery Loop, which is a public benefit corporation, and our fairly ambitious mission is we want to automate machine learning science and engineering so that we can improve the rate of discoveries across many different fields. Obviously we're going to start with a narrower set of domains because we think that's important to have a little bit of focus initially, but we think there are a lot of reusable pieces of infrastructure and common techniques that are common across domains. Also, by building models that are really good at understanding many many different domains of science and engineering, you can get kind of PhD level expertise in a model across many different domains. No one human has PhDs in 20 different fields, but we think by having that ability you'll be able to really figure out what are the important subproblems, orchestrate agents and multi-agent systems on solving some of those subproblems, recombine the results of the subproblems into solutions for the larger problem, and continuously iterate on these sorts of loops. And we think being able to run those iterative loops each iteration, running those more quickly by building the right kinds of tools that can implement an experiment or evaluate an experiment quickly.
这样你就能在几分钟或一小时内完成一轮实验,而不是等上一天或一周。而且,能够并行运行成千上万个实验并从中获得反馈,再用这些反馈来决定接下来要跑哪些实验,这也会提高你运行实验的质量。所以,我们认为,同时提升实验的速度和质量,实际上会带来非常惊人的成果。你还可以构建基础设施来管理你接下来可能想运行的所有实验,并根据预期价值和算力等成本来评估它们。我们认为这是一个非常令人兴奋的方向,这就是我们正在着手做的事情。
So that you can run an iteration of an experiment in a minute or an hour instead of a day or a week. And then being able to get feedback from running thousands and thousands of experiments in parallel, and use that feedback to decide what the next set of experiments to run, will lead to higher quality experiments as well. So increasing both the speed at which you can run experiments and the quality of those experiments, we think, is actually going to lead to something pretty amazing. And there's infrastructure you can build that manages the whole set of experiments you might want to run next, and tries to evaluate them in terms of the expected value and the cost in compute resources or other things. We think that's a pretty exciting direction, and that's what we're setting out to do.
我和我的四位联合创始人,或者说三位联合创始人,桑杰·格玛沃特、奥里奥尔·维尼亚尔斯和郭克雷。
Me and my four co-founders, or three co-founders, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le.
而且这非常有趣,因为我们在一起工作已经有 14 到 30 年了。我们以两两合作的方式在许多不同的事情上合作过,包括你提到的很多内容:MapReduce、BigTable、Spanner、TensorFlow,还有很多像模型蒸馏和架构之类的东西。所以我们认为,作为四人团队合作,打造一家致力于推动科学发现的伟大公司,将会非常有趣。作为一家公益公司,我们希望将这些科学发现带给世界上尽可能多的人。所以我们的目标是广泛传播这些发现,我们可能会做出一些不符合公司财务利益、但符合更广泛社会利益的决定,让这些发现传播出去。
And it's super fun because we've worked together for anywhere between 14 and 30 years. We've collaborated in pairwise versions on a number of different things, including a lot of the things you mentioned: MapReduce, BigTable, Spanner, TensorFlow, and a lot of things like model distillation and architectures. So we think it's just going to be a lot of fun to collaborate as a group of four and build an awesome company that is working to advance scientific discovery. And as a public benefit corporation, we want to bring those scientific discoveries to as many people in the world as possible. So our goal will be to distribute those widely, and we might make decisions that are not in the company's financial interest but are in the broader societal good of getting those discoveries out.
是的。谢谢。非常感谢。我是说……是的。是的。
Yeah. Thank you. Thank you so much. I mean that's... Yes. Yes.
这是一个伟大的愿景。
That's a great vision.
而且我已经在那里工作了 12 个半小时了。
And I've been working there for 12 and a half hours.
谢谢你,杰夫。我午夜时分失业了一秒钟。
Thank you, Jeff. I was unemployed at midnight for one second.
所以,谢谢你,杰夫。谢谢你,唐。现在,我们开放两个提问。先请那位。
So, thank you, Jeff. Thank you, Don. Now, we open the floor for two questions. That one first.
请讲。
Please go ahead.
非常感谢各位嘉宾带来如此鼓舞人心的对话。我叫乔伊斯,来自中国的中欧商学院。我是谷歌和 Meta 两家的股东。所以如果可以的话,我有两个问题,每人一个。第一个问题给杰夫:你的辞职导致谷歌当天市值蒸发了 2000 亿美元。所以在世人眼中这显然是一件大事。我的问题是,你认为你在初创公司里能实现哪些在谷歌无法实现的目标?还有一件让我困惑的事是,凭借谷歌的算力和所有人才,你认为谷歌需要做什么才能迎头赶上,让 Gemini 追上 Fable?我的问题给唐:在所有大型科技公司中,为什么选择 Meta?你一定收到了所有人的邀请,为什么是 Meta?谢谢。
Thank you very much to our panelists for such an inspiring chat. My name is Joyce from Chong Business School in China. And I'm a shareholder of both Google and Meta. So I have two questions if I may, one for each. The first question for Jeff: your resignation caused Google to lose 200 billion in market cap on the day. So it's clearly a big deal in the eyes of the world. My question to you is, what do you think you can achieve within a startup that you cannot achieve within Google? And also one thing that puzzles me is, with Google's compute and all the talent, what do you think Google needs to do to catch up, for Gemini to catch up with Fable? And my question to Don is, why Meta out of all the big tech companies? You must have offers from everybody, so why Meta? Thank you.
我先来。是的,我的意思是,我不会把股市的事情归因于任何特定事件,因为总是很难……
I'll go first. Yeah, I mean, I don't attribute things in the stock market to any particular events because it's always hard to...
就是你,杰夫。
It was you, Jeff.
首先,我对在谷歌的时光怀有无比深厚的感情。我在那里待了 27 年,同事都非常出色。所以他们状态很好。他们有计划打造 Gemini 模型。太棒了。会很好的。而我也很兴奋能出去做这件事。我认为这会非常非常有趣。有时候,一个专注的小公司,每个人都专注于同一个使命,而且我认为这是一个伟大的使命,将会非常了不起。
I think first, I have incredible fondness for my time at Google. I've been there 27 years. Amazing colleagues. So they're in good shape. They have a plan for making the Gemini models. Awesome. It's gonna be great. And I'm excited to go off and do this. I think it's going to be really, really fun. Sometimes a focused small company with everyone focused on exactly that mission, and I think it's a great mission, is going to be amazing.
好的,那我回答这个问题。是的,实际上,过去许多领先的前沿实验室都联系过我,包括当时的谷歌……
Okay, so I answer the question. Yes, actually, many of the leading Frontier Labs have reached out to me in the past, including Google at the time...
因为你做得很出色。
Because you do great work.
谢谢。谢谢。但当时我还没完全准备好加入大型科技公司。而且,正如你提到的,我一直在做自己的创业公司。实际上我过去做过四家创业公司。最近的一家后来被 Meta 收购了。当时,决定是否要加入 Meta 也是一个艰难的决定。那家创业公司做得非常好,我们拥有财富 500 强客户,等等。但我认为 Meta 也有一个很棒的平台。我认为它可以说是最大的分发平台,拥有数十亿用户和数亿商家。而凭借我们一直在构建的技术,真正领先的 AI 和安全技术,我认为在这样一个大平台上可以产生巨大的影响。所以,是的,我很高兴。而且,这才刚刚过去几周。非常兴奋能提供帮助。
Thank you. Thank you. But at the time I wasn't quite ready to join big tech companies. And also, I have been doing my own startup, as you mentioned. I actually have done four startups in the past. The most recent one was actually acquired by Meta. And at the time, deciding whether we wanted to join Meta was also a difficult decision. The startup was doing really well, and we had great customers in Fortune 500 companies, and so on. But I think Meta also has a great platform. I think it's really, arguably the biggest distribution platform, with billions of users and hundreds of millions of businesses on the platform. And with the technology that we have been building, really the leading technologies for providing AI and security and so on, I think can actually have a huge impact with such a big platform. So yes, with that, I'm very happy. And also, it's only been just a few weeks. Really excited to help.
那么我们有最后一位先生提问。
So we have last question from this gentleman.
嘿,我叫乔纳森,是一名工程师。我非常欣赏并使用你们的工作。我保证即使你没有在 12 小时前做出巨大的职业转变,我也会问这个问题。对 AI 行业的第一性原理分析都会表明,数据是一个巨大的护城河,所以你需要大公司才能赢。你需要大量的分布式系统,你需要极高的资本支出,所以大公司会赢。而你、你的联合创始人以及像诺姆这样的个人,都证明了优秀的工程师可以在大公司产生超乎寻常的影响。所以我的问题不是针对谷歌的,但我确实想听听你的内幕视角。
Hey, my name is Jonathan. I'm an engineer. Big fan and user of your work. I promise I'd be asking this even if you didn't make a giant career switch 12 hours ago. Every first principles analysis of the AI industry would suggest that data is a large moat, so you need a large company to win. You need a lot for a distributed system, you need extreme capital expenditures, so large company would win. And you and your co-founders and individuals like Noam are evidence that excellent engineers can make outsized impact at a large company. So I have a non-Google-specific question, but I do want to ask for your inside baseball perspective.
嗯,为什么看起来很多 AI 人才留存和所有前沿研究,尽管从结构性原因来看,你会认为它们应该发生在大型公司?嗯,为什么所有这些似乎仍然流向初创公司?嗯,这不一定是针对 Google 的问题,但为什么小公司似乎能胜过大型公司,而传统智慧会认为大型公司绝对会主导这个行业?
Um, why does it seem like a lot of AI talent retention and all of the frontier research, uh, despite structural reasons for you assuming them to be happening at a large company? Uh, why is all of that, uh, seem to be still going towards startups? Um, and this doesn't have to be a Google specific question, but why are small companies seeming to win over large companies when every in conventional wisdom would suggest that large companies are absolutely going to dominate this industry?
是的,我认为首先是因为云计算的兴起,你知道,在各种云平台上部署了大量面向机器学习的算力,这使得一小群人能够筹集相当多的资金,并使用这些平台,而不必自己构建它们,对吧。所以我认为这能让那些有梦想、有愿景或认为某个方向非常有趣的人,在比那些必须部署所有基础设施的公司更小的环境中去探索。我们可以依赖像 Google 这样了不起的公司,或者其他云提供商,来完成很多繁重的工作,构建支持人们心中公司应有的样子的基础设施。作为一家小公司,我们非常兴奋能以非常专注的方式从事科学和工程自动化,因为我们觉得那是一个真正成熟的领域。我们也许本可以在 Google 内部做到这一点,但我认为,10 个人在希望是帕洛阿尔托某处的办公室里,只专注于这一使命,能穿透大型组织中那些轻微的干扰。但大型组织在很多方面都很棒,多年来我在 Google 环境中获得了深厚的友谊,并受益于大型公司提供的所有东西。所以没有那种支持就出去闯荡有点紧张,但同时也令人兴奋。
Yeah, I mean I think first there's the rise of cloud computing and you know deployments of very large amounts of ML oriented compute in various cloud platforms and so that makes it possible for you know a small set of people to you know raise a fair bit of capital and be able to use those platforms without having to build them themselves right and so that I think enables people who have you know a dream or a vision or a a direction they think is really interesting to explore um to do that in a smaller setting than a company that has to sort of deploy all the infrastructure and we can rely on amazing companies like Google um you know or other cloud providers to do a lot of the heavy lifting to build out the infrastructure that is needed in order to support you know people's uh you know uh idea of what a company should be and We're super excited as a small company about working in a really focused way on science and engineering uh auto automation because we think that's just a a really ripe area for for something and we could have done that within Google probably but I think the the mission of you know 10 people in what we hope will be office space somewhere in Palo Alto uh just focused exclusively on that is just cuts through a lot of things that you know are slight distractions in a large organization, but large organizations are awesome in in many ways and I have incredible friendships and you know benefit that have benefited from all the things that large companies provide in Google setting uh over many years. So it's a little nerve-wracking to go out without that kind of support. Um but at the same time it's also exciting.
谢谢。
Thank you.
好的,谢谢。
Okay. Thank you.
好了。嗯,抱歉我们时间到了。我知道还有很多问题,但我希望我们能利用 AI 来猜测 Jeff 和 Dawn 会想什么。谢谢。
All right. Um, I'm sorry we are out of time. I know there are a lot of questions, but I hope um we can use AI to guess what the Jeff and the Dawn will be thinking. Thank you.
谢谢。
Thank you.