开源 AI:效率、创新与外部大脑

Open Source AI: Efficiency, Innovation, and the External Brain

布莱恩·卡坦扎罗 Bryan Catanzaro · Matt Turck 的 MAD 播客 · 2026-07-02 · 约 83 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Bryan Catanzaro 探讨开源 AI 现状、与闭源差距及外部大脑的影响。

Bryan Catanzaro discusses the state of open source AI, the gap with closed source, and the implications of an external brain.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 29)

全文 · Full transcript(中英对照)

引言与开源AI Introduction and open source AI

Host

如果你接受一个事实——我们即将逼近极限——那么这意味着,要获得更多智能,就必须提高效率。如果已经达到极限,我们就无法通过施加更多算力来获得更多智能。我们必须更审慎地利用现有资源。我们制造工具,制造外部器官来帮助解决问题。你看,我们有一个外部胃,我们称之为厨房。现在我们正在创造外部大脑。外部大脑意味着什么?非常深远。实际上没人真正知道。大家好,我是 Matt Turk,欢迎回到 Mad Podcast。开源 AI 又迎来了一个高光时刻,几乎每周都有强大的新模型发布。今天我的嘉宾是 Bryan Catanzaro,他是解读这一切的最佳人选之一。Bryan 领导着 Nvidia 的 NeMo Tron 开源基础模型家族。很多人不知道 Nvidia 在构建前沿 AI 模型上投入了巨大努力,它雇佣了数百名 AI 研究人员,而 NeMo Tron 3 Ultra 在几周前发布后立即成为美国排名第一的开放权重模型。我们从开源 AI 的现状以及中美之间的竞赛开始聊起。然后深入 NeMo Tron,用通俗语言讲解比特训练、混合成员 Transformer 架构、混合专家、多词预测和多教师蒸馏。最后,我们真实地了解现代 AI 研究组织如何运作,如何让众多聪明头脑构建一个模型而不是发表 100 篇论文。请享受与 Bryan Catanzaro 的精彩对话。好的,Bryan,很高兴做这期节目。看起来开源今年表现更好。你们 Nvidia 刚刚发布了 NeMo Tron 3 Ultra,这是一个重要时刻,也是美国最好的开源开放权重模型。就在几天前,GLM 5.2 也发布了,又是一个里程碑。所以,开源 AI 正在加速发展。我觉得这是一个很好的起点。你如何评估当前的情况?闭源和开源之间的差距有多大?

If you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. We build tools, we build external organs that help us solve problems. You know, we have an external stomach, we call it a kitchen. Now we're creating an external brain. What is the implications of an external brain? Pretty profound. Nobody actually really knows. Hi, I'm Matt Turk. Welcome back to the Mad Podcast. Open source AI is having yet another moment with powerful new models arriving almost weekly. And my guest today is one of the very best people to unpack it all. Bryan Catanzaro leads NeMo Tron, Nvidia's family of open foundation models. Now, not everyone realizes Nvidia has a massive effort to build frontier AI models, but it employs hundreds of AI researchers, and NeMo Tron 3 Ultra immediately became the number one US open weights model when it was released just a couple weeks ago. We begin this conversation with the state of open source AI and the race between the US and China. And then we go deep inside NeMo Tron for bit training, hybrid member transformer architecture, mixture of experts, multi-token prediction, and multi-teacher distillation all in plain language. And finally, we get a real look at how a modern AI research organization actually runs, how you get many brilliant minds to build one model instead of 100 papers. Please enjoy this awesome conversation with Bryan Catanzaro. All right, Bryan, excited to do this. It seems that open source is having a better year. So, you guys at Nvidia just released NeMo Tron 3 Ultra, which is an important moment and the best open source open weights model in the US. That was just a few days ago, and then even more recently GLM 5.2 came out, and that was another moment. So, it that things are accelerating in open source AI. It feels like a great place to start. What's your assessment about where we are and how wide the gap between closed source and open source currently is?

Bryan Catanzaro

看到这么多精力投入到 AI 开放技术中,真的很令人兴奋,因为我们知道开放技术让人们能够创新。互联网就是一个很好的例子。我们确实有过封闭的互联网,不知道你是否记得当年的美国在线和 Prodigy。它们很棒。而开放的互联网也同样惊人,对吧?很多不同的公司都因为开放技术而找到了改变工作方式的方法。互联网在零售领域的应用与在医疗或制造业的应用截然不同,但所有这些领域都被互联网彻底改变了。我相信 AI 也是一项变革性技术,而且需要以非常多样化的方式应用。因此,我认为 AI 的开放技术至关重要。看到全球这么多不同的组织持续投资和开发 AI 开放技术,非常令人兴奋。我希望这能继续下去。

Well, it's really exciting to see all of the energy going into open technologies for AI because we know that open technologies make it possible for people to innovate. You know, the internet is such a great example of that. We actually did have closed internets. I don't know if you remember things like America Online and Prodigy back in the day. And they were great. And open internet has also been amazing, right? Like so many different companies have been able to figure out how to transform their work thanks to an open technology. The application of the internet to retail is very different from the application of the internet to healthcare or manufacturing, but all of them have been totally transformed by the internet. AI, I believe, is also a very transformational technology and also a technology that needs to be applied in very diverse ways. And because of that, I believe that open technologies for AI are really fundamental. And it's very exciting to see continued investment and development of open technologies for AI from so many different organizations around the world. And I hope that that continues.

Host

那你觉得开源相比闭源落后多少?过去几年的大趋势是差距在缩小。你认为开源已经接近了,还是闭源模型不断提高门槛?

And what do you sense for how far behind open source is compared to closed source? This has been the big trend of the last few years has been this sort of narrowing gap. Do you think that open source is almost there or the bar keeps getting raised by the closed source models?

Bryan Catanzaro

嗯,我觉得这个问题可能很有诱惑力,因为设定某种竞争很有趣,但我实际上认为整个 AI 社区都在快速前进。比如,看看过去三个月 AI 的进展,无论是闭源还是开源,都令人难以置信。所以,如果你身处一个发展极快的领域,我认为这比不同模型之间可能存在的任何特定差距都更重要,因为最重要的是 AI 作为一个领域如何发展。

Well, I feel like this question, it's maybe a tempting question because it's fun to set up kind of competition, but I actually feel like the whole AI community is moving very fast. And if you look, for example, at the progress in AI, whether it's closed or open, just over the past 3 months, it's been incredible. And so, if you're in a field that's moving really, really fast, I think that's more important than any particular gaps that might exist between different models, because the most important thing is, you know, how is AI developing as a field.

Host

你认为推动开源 AI 持续进步的动力是什么?是社区?是像 Nvidia 这样的大公司支持?还是与中国的全球竞争?是什么推动开源 AI 向前发展?

What do you think the drivers are to continue to progress in open-source AI? Is that the community? Is that big companies like Nvidia being behind it? Is that the global competition with China? What propels open-source AI forward?

Bryan Catanzaro

我认为有很多因素在推动 AI 开放技术向前发展。首先是需求。很多组织希望定制 AI 并深度集成到工作中,这确实需要 AI 的开放技术。所以需求肯定存在。其次,这也是发展技术的最佳方式。几十年来我们看到,开放开发的技术进步更快,因为我们可以相互学习。在我们一生中,AI 的开发和部署是技术领域最激动人心的事情,计算机科学家除了让 AI 变得出色之外,还想做什么呢?如果社区合作是最好的方式,那么这也是推动社区开放开发技术的动力。

You know, I think there's a number of things that are pushing open technologies for AI forward. One is just the demand. There's so many organizations that want to customize AI and want to integrate it deeply into their work in a way that really requires open technologies for AI. So, I think the demand is certainly there. I think also it's just the best way to develop technology. And we've seen this for many decades that technology developed in the open moves quicker because we can all learn from each other. And in an era where we're undergoing the most exciting thing to happen in technology in our lifetimes with the development and the deployment of AI, what else do computer scientists want to work on other than making AI awesome? And if working together as a community is the best way to do that, then that's also a driver that pushes the community towards openly developing technology.

Host

问一个可能有点愤世嫉俗的问题,社区中至少有一部分人想知道,开源作为一个生态系统——不是指 Nvidia,而是泛指——是否部分基于蒸馏闭源模型的能力而进步。在一个我们看到 Anthropic 和 Febo 5 等开始阻止蒸馏的世界里,你认为开源 AI 的进步是否会因此放缓?

To ask maybe a slightly cynical question, there is at least a part of the community that's wondering whether open-source as an ecosystem, not Nvidia, but in general, has been progressing in part based on the ability to distill closed-source models and in a world where we're seeing the Anthropic and Febo 5's of the world starting to discourage distillation, do you think there is a chance that open-source AI progress may slow down in that context? Or as a result?

Bryan Catanzaro

在我看来,毫无疑问,当技术社区决定对我们这个时代最具变革性的技术进行巨额投资时,一定会取得快速进展。而且,这项技术不会被一小群人控制。因为这不是行业运作的方式。我们做出最好的工作,当我们能够以各自的方式思考和应用时,我们的工作才能产生最大的影响。所以,我喜欢闭源 AI API,无论是 Anthropic 还是其他人的。我认为它们很棒。对这些实验室的工作印象深刻。但它们不是世界上唯一的实验室。全球有很多实验室,很多人都有好主意。并不是只有少数实验室垄断了所有好主意。这不是真的。这不是人类运作的方式。这个星球上有很多聪明人。

You know, in my mind there's no question that when the technology community decides to make huge investments in the most transformational technology of our time, that there's going to be rapid progress. And also that that technology is not going to be controlled by a small group of people. Because that's just not the way that the industry works. We do our best work. We have the most impact with our work when we're able to each think about it in our own way and apply it in our own way. So, you know, I love the closed AI APIs, whether from Anthropic or other people. I think they're amazing. Really impressed with the work that those labs are doing. But they're not the only labs in the world. There's lots of labs around the world and lots of people have a good idea. It's not the case that there's only a few labs that have the monopoly on all good ideas. That's just not true. That's not how humanity operates. There's a lot of bright people on this planet.

AI社区与开放性 Community and Openness in AI

Bryan Catanzaro

社区当然非常关心这项技术。它显然具有如此变革性,对这么多事物产生深远影响,所以很多人自然想参与其中。因此,我认为随着时间的推移,我们会看到以社区为导向的 AI 开发和部署方法会继续加强并得到广泛采用,因为这正是人类历史上我们构建事物的方式。

And you know, the community of course cares deeply about this technology. It's obviously so transformational. It has such profound impacts on so many things that of course many people want to be involved in that. And so I think over time we're going to see that community-oriented approaches to developing and deploying AI are going to continue to strengthen and be widely adopted because that's really the history of how we built things as a human species.

Host

你认为这在全球范围内也成立吗?特别是关于中国,有一种看法是,世界上很多人都有好主意,但中国模型的许多进步直接受到闭源模型的启发,或者通过蒸馏产生。这算是种族煽动吗?还是说,作为一位领先的 AI 研究员,你同样对中国涌现的新颖想法印象深刻?

Do you think that is globally true as well? So, you know, in particular with respect to China, this perception that yes, a lot of people have great ideas around the world. However, a lot of progress from Chinese models were directly inspired or perhaps generated through distillation from the closed-source models. Is that just kind of like race bait or from the perspective of a leading AI researcher, you're very impressed by the novel ideas that come out of China as well?

Bryan Catanzaro

也许不寻常的是,我确实在一家中国公司工作了大约两年半。我在百度工作,和吴恩达以及达里奥·阿莫迪一起在硅谷 AI 实验室。我们都为一家中国公司工作,亲眼目睹了百度其他同事的聪明、勤奋、创造力和发明能力。那段经历一直留在我心中。我认为,说其他国家的成就都是靠模仿心态创造出来的,这绝对是错误的。事实并非如此。当然,我们在技术社区中互相学习吗?当然,我们当然互相学习。但我要说,中国 AI 社区对他们所构建的东西如此开放,这对世界来说是一件非常好的事情。我认为这使大量公司能够构建如果没有社区就无法完成的事情。而且我认为这也促进了整个 AI 生态系统的技术进步。所以,我非常感谢中国同事多年来所做的贡献。我也希望鼓励中国以外的世界各地 AI 实验室也保持开放精神。当 OpenAI 不久前发布 GPT OSS 模型时,我非常兴奋,当然谷歌也在 Gemma 上做了出色的工作。看到这些真是令人激动。我们 NVIDIA 也在推进 NeMo Tron。所以,我认为世界其他地方有机会赶上中国,因为我们可以理解作为一个社区共同构建 AI 技术的好处,而我认为中国在这方面坦率地说一直处于领先地位。

You know, perhaps unusually, I actually did work at a Chinese company for about 2 and a half years. I worked at Baidu. I worked in the Silicon Valley AI lab along with Andrew Ng and as well as Dario Amodei. And we all worked for a Chinese company and saw how smart, hard-working, creative, inventive our colleagues were at the rest of Baidu. And you know, that experience has stuck with me. I think it's absolutely false to say that the achievements of some other country are all being created by sort of copycat mentality. It's just not true. Now, do we all learn from each other in the technology community? Of course, of course we learn from each other. But I would say it's been a really good thing for the world that the Chinese AI community has been so open with what they've been building. I think it's enabled a tremendous number of companies to build things that they couldn't have done without community. And I think it's also spurred technological progress throughout the AI ecosystem. So, I'm really grateful for the contributions that our colleagues in China have made over the years. And I would love to encourage a spirit of openness amongst AI labs around the world outside of China as well. I was really excited when OpenAI released the GPT OSS models a while back and then of course Google's been doing great work with Gemma. Absolutely thrilling to see that. And we're pushing NeMo Tron along here at NVIDIA as well. So, I think there's a chance for the rest of the world to catch up to China in the sense that we can understand the benefits of working together as a community to build technologies for AI in a way that I think China has frankly been leading.

开源模型的优势 Advantages of Open-Source Models

Host

对。如今客户使用开源模型的理由是什么?根本优势是什么?

Right. What is the case for a customer to be using open-source models these days? What is the fundamental advantage?

Bryan Catanzaro

每家公司都围绕一个秘密建立。这个秘密不仅涉及他们的知识产权,还涉及他们的平台,即他们如何与问题和客户互动?他们如何思考客户需求的解决方案?AI 的价值总是越紧密地连接这些秘密就越大。因为 AI 关键依赖于数据。所以,输入的数据越有价值,解决方案就越有价值。现在,每家公司考虑如何部署 AI 时,都必须思考:这对我们公司的核心秘密有什么影响?在很多情况下,由于商业秘密、商业模式考量甚至监管要求,有些数据必须依法非常谨慎地处理。当你能够自己思考并实施时,这样做会好得多。考虑 AI 的集成、AI 与客户互动的方式、设置的护栏,每家公司都对其客户有特定的理解,因此也知道客户需要什么。AI 开放技术的惊人之处在于它们允许定制,对吧?所以公司可以仔细考虑,构建对他们真正重要的东西。我在这次对话开始时谈到了互联网,以及互联网的部署在不同行业以非常不同的方式进行。随着我们看到 AI 改变整个经济中工作和娱乐的方式,人们非常渴望这样做。这确实激发了对 AI 开放技术的大量需求。

Every company is built around a secret. This is a secret that has to do with not just their intellectual property, but also their platform, which has to do with how do they interact with problems and customers? How do they think about solutions to what their customers need? And it is always the case that the value of AI is greater when it can be more tightly connected with those secrets. Because AI depends on data critically. So, the more valuable the data that goes in, the more valuable the solution becomes. Now, every company when it's thinking about how to deploy AI has to think through, what are the implications for the core secrets of our company? And there's a lot of circumstances where due to trade secrets or trying to think through the business model or even regulatory requirements that there's data that you really have to treat very carefully by law. And it is much better to do that when you are able to think that through and implement it yourself. Thinking about the integration of AI, the way that AI interacts with customers, the guardrails that are put in place, every company has a specific understanding of its customers and therefore what the customer needs. And the amazing thing about open technologies for AI is that they allow customization, right? So companies can think this through, they can build things that really matter for them. And I started out this conversation talking about the internet and about how the deployment of the internet has been done in very different ways for very different industries. And there's a lot of desire to do that as we see AI change the way that we work and play throughout the entire economy. This is really spurring a lot of demand for open technologies for AI.

Bryan背景与英伟达之路 Bryan's Background and Path to NVIDIA

Host

对。我很想深入探讨一下 Nematron,但在此之前,也许花几分钟聊聊你的故事和背景。你走到今天这一步的路径是什么,包括在百度的经历?

Right. I'd love to go into a bit of a deep dive into Nematron, but before we do that, maybe a few minutes on your story, your background. What was your path to where you are today including the Baidu detour?

Bryan Catanzaro

我 2008 年开始在 NVIDIA 工作。当时我是一名研究生,试图研究用于人工智能的并行计算,我认为 NVIDIA 有机会改变计算机处理 AI 的方式。

So I started work at Nvidia in 2008. At the time I was a graduate student trying to figure out parallel computing for artificial intelligence and I thought Nvidia had a chance of changing the way computers work AI.

Host

这在 2008 年大概是一个孤独的追求吧?

Which was a lonely quest presumably, right, in 2008?

Bryan Catanzaro

哦,那非常混乱。当时人们觉得我疯了。我记得 2008 年去 ICML,我发表了第一篇在 GPU 上训练模型的论文,人们问我为什么在那里。他们说这不是一篇适合 ICML 的好论文,我们这里只做花哨的数学。我说,但我认为计算对 AI 其实非常重要。如果我们能训练更大的模型,拥有更强的学习能力,我们可能能解决更多问题。他们点了点头,说,嗯,我不太确定你为什么在这里。

Oh, it was very chaotic. Back then people thought I was crazy. I remember going to ICML in 2008. I published my first paper training models on the GPU and people asked me why I was there. People said this is not a good paper for ICML. We just do fancy math here. And I was like, well, but I think computing actually matters a lot for AI. If we could train bigger models that had more capacity to learn, we could probably solve more problems. And they kind of nodded their heads and they were like, well, I'm not really sure why you're here.

Host

GPU 不是也用于游戏吗?

Isn't a GPU a thing for gaming as well, presumably?

Bryan Catanzaro

是的,也有这个用途,对吧?我们一直遇到这种想法。实际上,GPU 就是 NVIDIA 定义的东西。我们制造它们。所以,GPU 是我们为了加速世界上最重要的计算而制造的东西,1995 年是图形,而很长时间以来一直是 AI。总之,我开始了在 NVIDIA 的工作。我在研究小组里做一些奇怪的事情,试图为 GPU 上的 AI 制作编译器和库。这导致了第一个 Copperhead 的创建,它是一种嵌入 Python 的语言,可以编译到 GPU,我认为这预示了 TensorFlow 和 PyTorch 中的许多东西。然后这又导致了 cuDNN 的创建,这是 NVIDIA 第一个用于 GPU 上深度学习的产产品。我非常喜欢做那项工作,但我一直想更直接地了解 AI 的应用。

Yeah, there's also that, right? Which we continued to run into that idea. Actually, a GPU is whatever Nvidia says it is. We make them. So, a GPU is a thing that we make in order to accelerate the world's most important computations, which in 1995 was graphics and for a long time now it's been AI. So, anyway, I started at Nvidia. I was in the research group doing strange things about trying to make compilers, libraries for AI on the GPU. That led to the creation of first Copperhead, which was a Python embedded language that compiled to the GPU, which I think foreshadowed a lot of things in TensorFlow and PyTorch. And then that led to the creation of cuDNN, which was Nvidia's first product for deep learning on the GPU. And I really enjoyed working on that, but I was always wanting to see more first-hand about the applications of AI.

百度硅谷AI实验室与Dario合作 Baidu Silicon Valley AI Lab and Working with Dario

Bryan Catanzaro

在英伟达,我主要做 AI 的库和编译器。所以当吴恩达邀请我去百度和他一起建立硅谷 AI 实验室时,我觉得这是个绝佳的机会,因为即使在那个时候,百度在核心业务中应用 AI 就已经非常先进了。所以对我来说,那是一个绝佳的机会。百度硅谷 AI 实验室是一个了不起的地方,充满了才华横溢、工作努力的人。

And at Nvidia, I was mostly working on libraries and compilers for AI. So, when Andrew Ng asked me to go build the Silicon Valley AI lab with him at Baidu, I thought, this is a great opportunity because even back then Baidu was very advanced in its application of AI to its core business. So that was a fantastic opportunity for me. The Baidu Silicon Valley AI Lab was an amazing place, full of brilliant people that were working really hard.

Host

和年轻的 Dario 一起工作是什么感觉?有没有什么迹象表明他会成为今天这样的人?

What was it like working with a young Dario? Were there any signs that he could become who he has become?

Bryan Catanzaro

Dario 从一开始就才华横溢。我记得我面试过他,我是面试小组的成员。当时他一直在做生物信息学,所以没有接触过深度学习或我们现在称之为 AI 的东西。但很明显,他学得极快,而且思考极深。我想我最钦佩 Dario 的是他信念的坚定。我在这个领域工作了很长时间,我也相信 AI 会改变世界,但我不认为我像 Dario 那样完全相信它。也许是因为我博士期间的学术训练充满了谨慎。不知道你是否记得,2005 年时 AI 又老又差。

Dario was brilliant from the beginning. I remember I interviewed him; I was on the panel. At the time, he had been working in bioinformatics, so he hadn't been working on deep learning or the things we call AI these days. But it was very clear that he learned extremely quickly, and also that he thought extremely deeply. I think the thing I admire most about Dario is the strength of his conviction. I've been working in this field for a long time, and I've believed also that AI is going to transform the world, but I don't think that I believed in it as completely as Dario did. Perhaps that was because my academic training during my PhD was full of a lot of caution. I don't know if you remember, but AI was old and bad in 2005.

Host

它永远不会成功。

It will never work.

Bryan Catanzaro

人们用计算机做事情,从 1945 年就开始了,对吧?多年来有太多宏伟的承诺未能兑现,所以我带着很多谨慎进入 AI 领域。事实上,那时我们称之为机器学习,这基本上是一种回避,我们只是不想让人们知道我们在做 AI,因为他们会说:“哦,我们听说过那个。它从来都不成功,对吧?”所以我带着一点这种学术上的谨慎进入 AI,比如,“哦,我们应该有所保留。我不知道现在是不是时候。”而 Dario,他信念的坚定以及他对技术发展时机的理解——这次它真的会成功——以及这对技术应该如何发展、应该建立什么样的机构的影响。我认为他做得非常出色。所以,和他一起工作总是一种愉快的经历。

People did with computers, they started doing it in 1945, right? There had been so many grandiose promises that failed to deliver over the years, so I came to AI with a lot of caution. In fact, back then we used to call it machine learning, which was basically a dodge, like we just didn't want people to know that we were working on AI because then they would be like, 'Oh, we've heard about that. It never works, right?' So I came to AI with a little bit of this academic caution, like, 'Oh, we should hedge a little bit. I don't know if now's the time.' And Dario, his strength of conviction and his understanding of the moment of how the technology was developing, this time it was actually going to work, and then the implications of that on how the technology should be developed, what kind of institutions to build. I think he's done a spectacular job. So, working with him was always a fun experience.

重返英伟达与DLSS Return to Nvidia and DLSS

Host

那么,后来你回到了英伟达,请带我们回顾一下这段历程。

So, then you went back to Nvidia, and walk us through the journey.

Bryan Catanzaro

是的,实际上 10 年前,也就是 2016 年,黄仁勋打电话给我说:“嘿,你愿意回来建立一个应用研究实验室吗?”我觉得那是一个绝佳的机会。我一直热爱英伟达。我喜欢这家公司的运作方式,以及它所持有的信念。英伟达是一家非常独特的公司。它会在很长的时间跨度内坚持执行。我在 CUDA、深度学习技术、光线追踪图形技术、AI 图形技术上都看到了这一点。英伟达一次又一次地不怕投入 5 年或 10 年的研究来改变世界。在一家拥有如此坚定信念和执行力强的公司工作,对我来说是一种理想。我真的很喜欢公司给予研究人员发明未来的支持。所以,我决定回来。我参与的第一个项目后来变成了 DLSS,你的部分观众可能知道它。DLSS 是我们的实时 AI 图形技术,它能让一个小型 GPU 像大型 GPU 一样运行。它的效率大约提高了 10 倍,因为我们不是为每一帧计算每个像素的颜色,而是使用 AI 来推断颜色。如今,当你使用 DLSS 玩游戏时,每 24 个像素中有 23 个是由我们的 AI 模型生成的,玩家们很喜欢它。它已经成为玩游戏的标准方式,因为它反应更快、画面更美。我们的 AI 是在离线状态下用海量数据集训练的,它能够实时渲染出比传统方法更美的图形。我们最近实际上发布了 DLSS 5,它是 DLSS 的完全生成式版本,我对此感到非常兴奋。它代表了 10 年来关于如何让实时图形更美的研究的结合。所以,对我来说,这段旅程的一部分就是实时 AI 图形技术。

Yeah, so 10 years ago actually in 2016, Jensen called me up and said, 'Hey, would you like to come back and build an applied research lab?' And I thought that would be a fantastic opportunity. I've always loved Nvidia. I've loved the way the company works, the convictions the company holds. Nvidia's a very unique company. It follows through over long time periods. I've seen that with CUDA, with our deep learning technologies, with our ray tracing graphics technologies, our AI for graphics. Over and over again Nvidia is not afraid to put in 5 or 10 years worth of research in order to change the world. Working at a company that has that strength of conviction and the ability to follow through is kind of an ideal thing for me. I just really love the support that the company gives its researchers to invent the future. So, I thought I'd come back. The first project that I worked on actually became DLSS, which some of your audience may know about. DLSS is our real-time AI for graphics, and it makes a small GPU run like a big GPU. It's about 10 times more efficient because rather than computing the color of every pixel for every frame, we use AI to infer the color. These days 23 out of every 24 pixels is being generated by our AI model when you're using DLSS to play games, and gamers love it. It's become the standard way of playing games because it's so much more responsive and more beautiful. Our AI we train it offline on huge data sets and it's able to render graphics in real time more beautifully than traditional methods do. We recently actually announced DLSS 5 which is a fully generative version of DLSS and I am so excited about it. It represents a combination of 10 years worth of research on how to make real-time graphics much more beautiful. So that's part of the journey here for me was real-time AI for graphics.

Megatron与语言建模 Megatron and Language Modeling

Bryan Catanzaro

但与此同时,我们也启动了一个语言建模项目,那是在 2017 年,在 Transformer 火起来之前,在语言建模开始席卷世界之前。但我有一种直觉,也许基于我在百度工作时看到的一些东西。我直觉认为,处理文本和理解文本会带来更好的推理,进而让 AI 在各个领域得到更好的应用。于是我们启动了一个名为 Megatron 的项目。Megatron 代表最大最强的 Transformer,这就是我们这么命名的原因。这实际上是一个系统项目,旨在向世界展示如何在英伟达的硬件上训练最大的 Transformer 模型。当时,你的部分观众可能记得也可能不记得,有人声称训练大型 Transformer 模型的唯一方法是在 TPU 上,因为毕竟 Transformer 是谷歌发明的。所以我们看了 Transformer 论文。我们觉得哇,这有惊人的潜力。我们在自己的语言建模任务上试了试,它比我们之前使用的 RNN 好得多。同时,我们立即看到了一个巨大的系统机会,可以共同优化 GPU、网络、所有编译器和软件,使人们能够真正大规模地扩展基于 Transformer 的语言模型。我们认为,这是可以真正产生影响的事情。于是我们启动了 Megatron 项目,它后来帮助整个行业弄清楚了如何训练极大的 LLM,也为今天的 NeMoTron 项目奠定了基础,英伟达在其中训练自己的 LLM 用于自身目的。这就是大致的历史。

But then at the same time we also started a language modeling project and this is back in 2017, you know, before transformers were big and before language modeling started taking over the world. But I just had this intuition, maybe built on some of the things that I had seen while working at Baidu. I just had this intuition that working with text and understanding text was going to lead to better reasoning, which was going to lead to better application of AI in all sorts of domains. And so we started this project called Megatron. Megatron stands for the biggest baddest transformer. That's why we named it that. And it was really a systems project to show the world how to train the largest transformer models on Nvidia's hardware. Back at the time, some of your audience may or may not remember this, but there were claims being made that the only way to train big transformer models was on the TPU because after all the transformer had been invented at Google. So we looked at the transformer paper. We thought wow this has amazing potential. We tried it out on our own language modeling tasks, and it worked so much better than the RNNs that we had been using before. Also, we saw immediately that there was an enormous systems opportunity to co-optimize the GPU, the networking, all of the compilers and software that would enable people to scale Transformer-based language models really dramatically. And we thought, this is something that could really have an impact. So, we started the Megatron project, which then led to helping the whole industry figure out how to train extremely large LLMs, and also led to the foundations of today's NeMoTron project, where NVIDIA trains its own LLMs for its own purposes. So, that's kind of the history.

Host

了不起的历程。好的,那么让我们深入探讨 NeMoTron 的一切。

Great journey. Okay, so let's go into all things NeMoTron.

英伟达为何构建前沿模型 Why NVIDIA builds frontier models

Host

在深入细节之前,有一个显而易见的问题,我相信你已经被问过很多次了,那就是:NVIDIA 为什么一开始就要构建模型,并投入巨大精力打造自己的前沿模型系列?

And before we get into the specifics, there's the obvious question that I'm sure you've been asked many times, which is why does NVIDIA care in the first place to be building models and investing very significant efforts into creating its own family of frontier models?

Bryan Catanzaro

你知道,Nemotron 有两个任务。第一个任务是帮助我们理解如何构建未来的系统。NVIDIA 是一家加速计算公司,这意味着要从第一性原理出发思考世界上最重要的计算挑战,并设计系统——包括大量软件——以便让人们能够发明和部署标准计算无法实现的东西。但要做到这一点,NVIDIA 必须深入理解 AI 的一切工作原理。这就是我们如何为我们的主要产品线共同设计所有系统和软件的方式。所以,Nemotron 的第一个任务是确保 NVIDIA 继续存在,以便我们能够在摩尔定律已死的时代继续提供有意义的加速。如今我们获得的加速来自于专业化。但专业化又来自于理解。所以这就是 Nemotron 的第一个任务:帮助 NVIDIA 理解如何构建其核心产品。NVIDIA 的第二个任务,或者说 Nemotron 的第二个任务,是支持生态系统。NVIDIA 多年来建立的最有价值的东西之一,就是世界各地使用 NVIDIA 技术构建和部署出色 AI 的人们。我们认为,NVIDIA 有必要继续提供开放的 AI 技术来支持这一点。Nemotron 并不想成为唯一的开放 AI 技术。我们热爱所有 AI 技术,原因很简单:每当 AI 得到进一步发展和部署,对我们的业务来说就是机会。所以我们非常明确地致力于发展我们的生态系统,因为这对我们来说是门好生意。但我们并不想成为这个生态系统中唯一的技术提供商。我们乐于看到其他公司也做出贡献。Nemotron 第二个任务最重要的就是确保各种规模的公司都能继续构建和部署自己的 AI。

You know, Nemotron has two jobs. The first job is to help us understand how to build the systems of the future. NVIDIA is an accelerated computing company, and that means thinking through the world's most important computational challenges from first principles, and designing systems, which includes a lot of software, in order to make it possible for people to invent and deploy things that never could have been done with standard computing. But in order to do that, NVIDIA has to deeply understand everything about how AI works. That's how we co-design all of the systems and software for our main product line. So, the first job of Nemotron is to make sure that NVIDIA continues to exist so that we can continue delivering meaningful acceleration in an era where Moore's law has died. And the acceleration that we get these days comes through specialization. But again, specialization comes through understanding. So that's Nemotron's first job: to help NVIDIA understand how to build its core products. NVIDIA's second, or Nemotron's second job, is to support the ecosystem. One of the most valuable things that NVIDIA has built over the years is all of the people around the world who build and deploy amazing AI using NVIDIA's technologies. And we think that it's necessary for open technology for AI to continue to exist from NVIDIA to help support that. Nemotron is not trying to be the only open technology for AI. We love all technology for AI for the very straightforward reason that whenever AI is further developed and further deployed, it's an opportunity for our business. So we are very explicitly trying to develop our ecosystem because that's good business for us. But we're not trying to be the only provider of technologies for this ecosystem. We love seeing other companies contribute as well. The most important thing for Nemotron's second job is just making sure that it continues to be possible for companies of all shapes and sizes to build and deploy their own AI.

摩尔定律已死 Moore's law is dead

Host

顺便问一下,摩尔定律死了?这是官方的说法吗?

By the way, Moore's law is dead? Is that official?

Bryan Catanzaro

它已经死了好几年了。

It's been dead for years.

Host

已经死了好几年了?为什么?

It's been dead for years? Why is that?

Bryan Catanzaro

嗯,你只要看看半导体制造的进展就知道了。摩尔定律最初的表述是经济层面的,对吧?它说的是我们每 24 个月(或某个时间段)能在同一块芯片上以可承受的成本集成两倍数量的晶体管。而如今,这绝对不再是事实了。这种情况可能已经持续了五到十年。现在,我们仍然通过多种方式扩展系统。一种方法是使用更多的硅。晶体管也在继续变小、变高效,尽管速度变慢了,但同时它们也变得更贵了。所以,在摩尔定律还活着的时代,制造未来系统的最佳方式是拿现有系统缩小,可能同时翻倍。但我们已经生活在一个不再能从缩小现有设计中获得经济收益的时代,你必须更聪明地使用系统的每个部分。这是一个加速计算比以往任何时候都更有价值的时代,因为从第一性原理思考问题,并共同设计从晶体管到算法和应用程序的一切,以减少浪费并提供有意义的加速,这比以往任何时候都更有价值。

Well, you just look at the progress in semiconductor manufacturing. The original statement of Moore's law was economic, right? It was about we can afford to put twice as many transistors on the same chip every 24 months, whatever the time period is. And these days that is absolutely not the case. It hasn't been for probably five or ten years. Now, we are still scaling our systems through a number of ways. One is just applying a lot more silicon to it. We are also getting transistors continuing to get smaller and more efficient, although at a slower pace, but they're also getting quite a bit more expensive at the same time. So, in an era where Moore's law was alive, the best way to make the system of the future was to take the system of the present and then just shrink it, and maybe double it at the same time. But in an era where we've been living for a while now, where you don't get economic benefits from taking your existing design and shrinking it, you really have to be more clever about how you use every part of the system. That's an era where accelerated computing is much more valuable than ever because the work of thinking through the problem from first principles and co-designing absolutely everything from transistors to algorithms and applications in order to reduce waste and deliver meaningful acceleration, that's more valuable than ever.

Nemotron发布历史 Nemotron release history

Host

太棒了。回顾一下你刚才说的,我之前提到过,NVIDIA 涉足模型业务在商业上很有道理:第一,它有助于设计更好的芯片;第二,任何对 AI 有利的事情最终都对 NVIDIA 有利,这很合理。Nemotron 的工作是最近才开始的,对吧?我相信可能是在 2023 年。请快速带我们回顾一下关键发布。我记得 2023 年有一个关键发布是 Nemotron 3 8B,还是我漏了一步?

Fantastic. To play back what you were saying, I mentioned earlier, it makes good business sense for NVIDIA to be in the model business because one, it helps design better chips, and two, whatever is good for AI is ultimately good for NVIDIA, which makes a lot of sense. That Nemotron effort is reasonably recent, right? Started in 2023, I believe, maybe. Walk us quickly through the key releases. I believe in 2023, that was Nemotron 3 8B as a key release, or am I missing a step?

Bryan Catanzaro

是的,是的,是的。所以,编号在时间上有点混乱。我几乎感觉我们像是在《指环王》里,就像从古老矿井里挖出一些古代遗物。那是很久以前的事了。最初我们称之为 Nemotron 1 的,实际上是和微软合作的一个项目。我们联合训练了一个 5300 亿参数的模型。我相信那是在 2021 年发布的。所以那是 GPT-3 时代。当时我们称之为 Megatron Turing NLG。Turing 是微软当时对其语言模型工作的称呼。但事后看来,我们把它叫做 Nemotron 1。然后我们陆续又建了几个。我们做到了 Nemotron 3。然后 LLaMA 出现了,我们对此非常兴奋。我们很高兴 Meta 支持开放的 AI 技术空间。于是我们开始将我们的语言模型技术添加到 LLaMA 模型中,这产生了 LLaMA Nemotron 1。那是第一个基于 LLaMA 的推理模型。我们为此感到非常自豪。

Yes. Yes. Yes. So, the numbering is somewhat lost to time. I almost feel like we're in the Lord of the Rings and it's like, there's some ancient relics that we're digging up out of an old mine. This is a long time ago. The original what we originally called Nemotron 1 was actually a project that we did with Microsoft. We jointly trained a 530 billion parameter model. I believe that was released in 2021. And so this is GPT-3 era. And that's what at the time we called it Megatron Turing NLG. Turing was what Microsoft was calling their language model efforts at the time. But that in retrospect we called Nemotron 1. Then along the way we built a few more. We got up to Nemotron 3. And then LLaMA came along and we were really excited about that. We were very happy that Meta was supporting the open AI technology space. And so we started taking our language model technology and adding it to LLaMA models which then resulted in LLaMA Nemotron 1. And that was the first reasoning model built on LLaMA. We were really proud of that.

Host

那是在 2025 年吗?

And that was 2025?

Bryan Catanzaro

可能是 2024 年。我记不清了。大概是那个时候。然后我们继续开发。接着我们发布了 Nemotron 2。我相信是去年。然后我们很快又推出了 Nemotron 3,因为我们需要加入 MoE 支持。Nemotron 2 没有 MoE 支持,这使它与其他模型相比缺乏竞争力,比如 GPT-OSS 20B 因为 MoE 而非常快,所以我们觉得必须加入 MoE,于是就有了 Nemotron 3。现在我们处于一个有点尴尬的状态,因为我们正在开发 Nemotron 4,对吧?但我们已经发布了一个 Nemotron 4,那是在 2024 年,我们发布了一个 340B 的模型,也叫 Nemotron 4。所以我不太确定我们该如何解决这个营销问题。这个营销问题不是我造成的。

Might have been 2024. I believe I can't remember. Somewhere around there. And then we continued to develop that. And then we released a Nemotron 2. I believe it was last year. And then we quickly followed that up with Nemotron 3 because we needed to put MoE support in. Nemotron 2 didn't have MoE support and that made it kind of uncompetitive against other models like GPT-OSS 20B was just like so fast because of MoE and so we were like okay we've got to put the MoE in so that became Nemotron 3. Now we're in a slightly difficult state because we're working on Nemotron 4, right? But we already released a Nemotron 4 which was in 2024 we released a 340B model called Nemotron 4. And so I'm not exactly sure how we're going to solve this marketing problem. I didn't create this marketing problem.

英伟达对NeMo Tron的持续投入 NVIDIA's sustained commitment to NeMo Tron

Bryan Catanzaro

我会尽量说清楚:NeMo Tron 4,或者我们未来发布的任何下一代版本,都与 2024 年的 NeMo Tron 4 不同。但无论如何,我们已经为此努力了很久。我认为,对我们来说,比任何特定版本更重要的,是 NVIDIA 对这些模型开发的持续投入。我们已经做了很长时间。过去一年,我们的模型变得实用得多,这主要反映了两点:第一,整个公司团结起来了。NVIDIA 内部许多不同团队现在都明白这对 NVIDIA 的未来有多重要。因此,有更多的人和更好的想法投入到 NeMo Tron 中。第二,与此同时,我们能够扩大投入其中的算力。显然,拥有良好的计算基础设施对构建 AI 至关重要。我们最近大幅增加了投资,因为我们相信这对公司的未来至关重要。

I'll do my best to make it clear that NeMo Tron 4, or whatever the next generation is, whenever we release it, is different from the 2024 NeMo Tron 4. But in any case, we've been working on this for a long time. I think more important to us than any particular generation is just the sustained commitment that NVIDIA has to developing these models. We've been doing it for a while. I think our models have gotten dramatically more useful in the past year, which is a reflection of two things primarily. One is that the whole company has come together. There are many different teams around NVIDIA that now understand how important this is to NVIDIA's future. So there are dramatically more people and better ideas going into NeMo Tron. And number two, along with that, we've been able to scale the compute resources that go into it. Obviously, it's very important to have good computing infrastructure to build AI. We've recently increased our investment substantially because we believe this is really key to our company's future.

Host

哦,真有意思。

Oh fascinating.

Bryan Catanzaro

但我想继续这个想法。我认为让大家知道我们已经做了很久非常重要。我们正在大幅增加投资,而 NVIDIA 是一家说到做到的公司。你知道,我们在 CUDA 上坚持了十多年,现在在 NeMo Tron 上也是如此。

But I just want to continue the thought. I think it's really important that everybody knows we've been doing this for a long time. We are increasing our investments substantially, and NVIDIA is a company that follows through. You know, we followed through over 10 plus years with CUDA, and we're doing that with NeMo Tron now.

Host

这非常有帮助,因为我认为更广泛的世界才刚刚开始意识到,一个非常实质性的开源前沿 AI 研究工作一直在进行。所以听到这个进展,以及我们马上要讨论的这个模型系列,非常有趣。另一个重要时刻似乎是三个月前,也就是三月,NeMo Tron 联盟的成立。你能简单解释一下那是什么吗?

That's very helpful because I think the broader world is just starting to catch up to the fact that there is a very substantial open source frontier AI research effort that has been happening. So it's very interesting to hear that there's been this progression and now this family of models that we're going to talk about in a second. Another important moment seems to be the creation, just in March, three months ago, of the NeMo Tron coalition. Do you want to explain briefly what that is?

Bryan Catanzaro

NeMo Tron 的存在是为了支持生态系统,我们当时在想,这是一个不同于行业中其他 AI 项目的项目,对吧?因为我们实际上并不试图以任何方式主导。我们只是想支持。我们并不试图控制 AI 如何被整合到这些公司中。我们只是想确保有好的 AI。但我们想,也许如果我们与人们一起开发,那么它对他们会更有用。整合起来会更容易,因为我们会从一开始就考虑他们的需求。NeMo Tron 一直是协作的。我告诉过你,很久以前,我们训练的第一个大模型是与微软合作的,对吧?那是 NVIDIA 和微软研究人员并肩工作的联合努力。我认为那最终对 NVIDIA 和微软都有帮助。我们都从那次经历中学到了很多。所以,因为 NeMo Tron 不是要与其他公司竞争,而是要支持,因为我们无论如何都会公开它,为什么不在构建之前就合作呢?而不是让 NeMo Tron 成为 NVIDIA 独自完成的项目,然后发布到互联网上说,“嘿,你为什么不试试这个?我们认为它可能不错。”我们为什么不通过与感兴趣的合作伙伴在 NeMo Tron 创建之前就合作,并纳入任何形式的反馈、评估、环境、基准测试或其他人想要带来的任何其他技术,来确保它对合作伙伴有益呢?事实证明,整个生态系统中,有很多公司真的希望开放模型成功,因此他们有自身利益,有既得利益,来确保开放技术是优秀的。那么为什么不与他们合作,让他们以任何他们喜欢的方式为改进 NeMo Tron 做出贡献呢?这就是 NeMo Tron 联盟的想法。它不是一个排他性的联盟。我们并不试图成为唯一的模型。与我们合作的所有公司都可以自由地以对他们有意义的方式继续工作,然而这些公司希望与我们合作,因为他们想确保 AI 的开放技术持续快速发展,并且他们有机会影响其发展方式。

So, NeMo Tron exists to help support the ecosystem, and we were thinking, well, this is a different kind of AI project than other projects around the industry, right? Because we're not actually trying to dominate in any way. We're just trying to support. We're not trying to control the way that AI is being integrated into all these companies. We're just trying to make sure there's good AI. But we thought, well, maybe if we worked with people while we develop it, then it's going to be more useful for them. It'll be easier to integrate because we will consider what they need from the beginning. And NeMo Tron has always been collaborative. I was telling you that a long time ago, our first big model that we trained, we did with Microsoft, right? It was a joint effort where NVIDIA and Microsoft researchers worked side by side to build that. That ended up, I think, helping both NVIDIA and Microsoft. I think we both learned a lot from that experience. So, because NeMo Tron is not trying to compete with other companies, but rather support, because we're going to be putting it out there openly anyway, why not collaborate before the thing is built, rather than NeMo Tron being a project that NVIDIA does all on its own and then posts on the internet and says, 'Hey, why don't you try this? We think it might be good.' Why don't we make sure that it's good for the partners that are interested by working with them before NeMo Tron is even created, and incorporating any sort of feedback, evaluations, environments, benchmarks, or any other kinds of technology that other people want to bring. It turns out that the entire ecosystem, there are a lot of companies that really want open models to succeed, and so they have a self-interest, they have their own vested self-interest, to make sure that open technologies are excellent. So why not work with them and let them contribute however they'd like to making NeMo Tron better. So that's the idea of the NeMo Tron coalition. It is not an exclusive coalition. We're not trying to be the only model out there. All the companies that we work with are free to continue doing the work however makes sense to them, and yet these companies want to work with us because they want to make sure that open technologies for AI keep developing quickly and that they have a chance to influence how that happens.

NeMo Tron家族现状 Current state of the NeMo Tron family

Host

太好了。NeMo Tron 系列的当前状态如何?你们有 Nano、Super 和 Ultra。这些模型做什么,它们的用例是什么?

Great. What's the current state of the NeMo Tron family? You got Nano, you got Super, you got Ultra. What do those models do and what are the use cases for them?

Bryan Catanzaro

Nano 是一个总参数量 300 亿、激活参数量 30 亿的模型。Super 是 1200 亿和 120 亿,Ultra 是 5500 亿和 550 亿。它们的设计确实是为了适应小型、中型和大型部署场景。Nano 对于不需要那么多知识或推理的任务可以非常胜任,但显然,对于最强大的模型,你会选择 Ultra。Super 在很多方面是我们最受欢迎的模型,因为它代表了成本和智能之间的良好平衡。所以我们喜欢这种小、中、大的方式来构建系列,因为我们的客户对此反应很好。但从 NVIDIA 的角度来看,人们用 LLM 做的最重要的事情是智能体,对吧?构建智能体式工作流,让智能体代表你工作,日夜为你解决问题。这是一种解决我们必须面对的问题的令人兴奋的方式。让 NeMo Tron 为此目的变得出色是我们的梦想。那是我们的目标。

So Nano is a 30 billion total, 3 billion active parameter model. Super is 120 and 12, and Ultra is 550 and 55. They're designed really to fit small, medium, and large deployment scenarios. Nano can be really capable for things that don't require nearly as much knowledge or reasoning, but obviously for the most capable model you go for Ultra. Super in a lot of ways is our most popular model because it represents a great balance between cost and intelligence. So we like having this small, medium, and large approach to building a family just because our customers seem to respond to that pretty well. But the most important thing from NVIDIA's point of view that people are doing with LLMs is agents, right? Building agentic workflows, having an agent working on your behalf, solving problems for you night and day. Such an exciting way of approaching the problems that we have to solve. And it's our dream to make NeMo Tron amazing for that purpose. That's our goal.

Host

深入探讨一下,从高层次来看,NeMo Tron 系列专注于智能体式推理,特别注重使其高效。这个概括对吗?

To double click on this at a high level, NeMo Tron family is focused on agentic reasoning with a particular focus on making it efficient. Is that the right headline?

Bryan Catanzaro

没错。是的,NeMo Tron 一直采用速度优先的方法来构建模型,因为 NVIDIA 是一家加速计算公司。正如我所说,我们试图从第一性原理思考:“这里的计算问题是什么?”NeMo Tron 3 系列中有很多我们引以为豪的东西。例如,NeMo Tron Ultra 和 Super 是使用 4 位算术进行预训练的。我们在 NVF P4 中预训练了它们。这是一项不平凡的工作,需要发明算法,以便你的模型能够使用如此粗糙的算术收敛到优秀的结果。这需要大量的发明。我们为此感到非常自豪。

That's right. Yeah, NeMo Tron has always been a speed-first approach to building models because NVIDIA is an accelerated computing company. As I was saying, we're trying to think through, 'What is the problem here computationally from first principles?' And NeMo Tron 3 family has a lot of things in it that we're really proud of. For example, NeMo Tron Ultra and Super were pre-trained using 4-bit arithmetic. We pre-trained those in NVF P4. Which is a non-trivial thing to do, to invent the algorithm so that your model can converge to an excellent result using such coarse arithmetic. It required a lot of invention. Really proud of that.

4位与16位格式及效率 4-bit vs 16-bit formats and efficiency

Host

你能解释一下,比如 4-bit 和 16-bit 有什么区别吗?

Do you want to explain maybe for people what 4-bit is versus 16-bit, for example?

Bryan Catanzaro

你知道吗,昨天我在 Hacker News 上看到一篇很棒的帖子,有人让你上传一张图片,然后它会基本上把它海报化,也就是减少颜色数量,以适配不同的数字格式,包括 NVFP4 和 MXFP8 以及其他一些格式。你可以滑动查看它对图片颜色的影响。效果非常显著。4-bit 的位数不多,对吧?只有 16 个值。当然,这些都是所谓的块缩放格式。所以一组数字还会附带一个 8-bit 的缩放因子。具体细节可能相当复杂,所以也许不是那么重要。但我们这样做的原因,首先是在我们的 GPU 上,特别是 Blackwell Ultra,这些格式的吞吐量大幅提高。其次,我们知道这将节省大量能源。思考 AI 计算问题的一种方式是,我们将在极限下运行。无论极限是什么,可能是经济极限,比如我们只有这么多亿美元来购买服务器;也可能是电力极限,我们只能负担得起这么多吉瓦来训练模型。无论极限是什么,我们都会在极限下运行。每个组织都是如此,为什么呢?因为智能的价值如此之高。人们会投资,因为他们知道会得到回报。智能的价值是巨大的。所以如果你接受我们将在极限下运行这个事实,那么这意味着获得更多智能的方法是提高效率。如果我们已经在极限下,就无法通过施加更多力量来获得更多智能。我们必须更明智地利用现有的资源。4-bit 数字格式在移动时成本大幅降低。它们占用更少的内存空间。从内存中移动它们,甚至在芯片上移动它们,消耗的皮焦耳更少。计算时消耗的能量也少得多,这推动了 4-bit 格式的投资。我认为如今 4-bit 格式在部署中已经非常成熟。现在制作一个良好的量化 4-bit 检查点并部署它,可以获得很多推理成本和速度优势,这已经相当直接了。但使用 4-bit 格式进行预训练则更具挑战性,因为你有一个数值求解器在优化权重,它可能非常敏感。所以如果你处理数字不当,你的模型可能会发散,而不是通过预训练得到一个模型,最终只是那次运行发散,这总是令人害怕的。因此,我们能够用 4-bit 预训练 NeMo-Megatron 需要很多创新。我们为此感到非常自豪。

You know, actually there was a fantastic post I saw on Hacker News yesterday where somebody let you upload a picture and then it would basically posterize it, basically reduce the colors to fit different number formats including NVFP4 and MXFP8 and some of the other formats that are out there. And so, you could kind of swipe around and look what it does to the colors of a picture. And it's really quite dramatic. Four bits is not a lot of bits, right? That's only 16 values. Now of course these are all what are called block scaled formats. So groups of numbers also come with an eight-bit scaling factor. The specifics of this can get rather complicated, so maybe they're not quite as important. But the reason why we want to do this is because first of all we have dramatically higher throughput for these formats in our GPUs, specifically on Blackwell Ultra. And secondly, we know that it's going to save an enormous amount of energy. One way to think about the computational problem of AI is that we are going to be running at the limit. Whatever the limit is, it could be an economic limit, like we only have so many billion dollars to buy servers with. It could be a power limit. We only have so many gigawatts that we can afford to train a model with. Whatever the limit is, we're going to be running at that limit. Every organization is, because why? Because the value of intelligence is so high. People are going to invest because they know that they're going to get return. The value of intelligence is enormous. So if you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. Four-bit number formats are dramatically cheaper to move around. They take up less space in memory. They take up less picojoules when you move them from the memory or even on the chip around the chip. Much less energy when you compute on them, and so that's really driving the investment in four-bit formats. I think these days four-bit formats for deployment are very well established. It's pretty straightforward these days to make a good quantized four-bit checkpoint that you can deploy and that gets you a lot of inference cost and speed advantages. But using four-bit formats for pre-training is quite a bit more challenging because you have this numeric solver that's optimizing the weights and it can be quite sensitive. So if you don't treat the numbers right, your model can diverge and instead of actually getting a model done through pre-training, you end up with basically just that run diverged, which is always scary. So it took a lot of invention for us to be able to pre-train NeMo-Megatron in four-bit. We're really proud of that.

混合架构:Transformer与Mamba Hybrid architecture: Transformer and Mamba

Host

好的,太棒了。那么当我们进入稍微更技术性的话题时,NeMo-Megatron 的架构是混合的,对吗?所以它是 Transformer 和 Mamba 状态空间的组合,这是一种稍微更奇特的架构形式。请给我们介绍一下。

Okay, great. All right, so as we get into slightly more technical things, the architecture of NeMo-Megatron is hybrid. Is that right? So it's a combination of transformer and Mamba state space which is slightly more exotic form of architecture. Walk us through that.

Bryan Catanzaro

是的,我们在 2024 年发表了一篇论文,表明通过将状态空间模型与 Transformer 结合,实际上可以得到更智能的模型。我们进行了一次扫描:模型应该有多少是全注意力机制,多少是状态空间模型,才能获得最低的困惑度,也就是你能得到的最佳语言模型。我们发现,实际上你希望它大部分是状态空间模型,带有一点注意力。其背后的直觉是,状态空间模型似乎更擅长那种直观的、印象主义的序列理解,因为它们将整个序列总结到一个恒定空间中。这就是它们的工作原理,对吧?所以它们不是随机查看整个序列,而是在每一步将所有内容总结到一个恒定的缓存或小记事本中。这种约束似乎使它们在一些涉及全局理解的任务上更智能。另一方面,全注意力机制的优势在于它可以挑出非常具体的信息并精确查看。它不会丢失任何东西。没有有损压缩发生。你可以看到全部内容。所以我们发现,将两者结合使用实际上比单独使用任何一种都要好。这与速度优势无关,仅仅是模型更智能了。自从我们发表那篇论文以来,我认为很多其他实验室也发现了这一点。如今很多模型都在使用混合 SSM 方法构建。例如,Qwen 已经这样做了。Kimi 现在使用他们所谓的 Kimi 线性注意力。所以,将某种状态空间模型与全注意力机制结合用于基础架构已经变得相当广泛。现在它也有一些速度优势,因为保存状态空间缓存所需的内存量实际上与序列长度无关,这意味着在训练和推理时,你通常可以在 GPU 上容纳更大的批次,因为内存需求更低。它使 GPU 更饱满、更忙碌,因此也提供了一些非常重要的效率优势。

Yeah, we published a paper in 2024 that showed that you actually get a smarter model by combining state space models with transformers. And we did a sweep: how much of the model should be full attention and how much of it should be a state space model in order to get the lowest perplexity, basically the best language model that you could get. And we found that you actually want it to be mostly a state space model with a little bit of attention. The intuition behind that is that the state-space models seem to be better at this kind of intuitive, impressionistic understanding of a sequence because they're kind of summarizing the entire sequence into a constant space. That's how they work, right? So instead of having the ability to look at the entire sequence randomly, they summarize everything at every step into a constant cache or little scratchpad that they're working on. That constraint seems to actually make them smarter at some tasks that involve like global understanding. On the other hand, the advantage of full attention is that it can pick out very specific bits of information and look at those exactly. It doesn't lose anything. There's no lossy compression going on. You can actually see the whole thing. So we found that using both of these together was actually better than using either one on their own. That is independent of the speed benefit. That is just the model is smarter. Since we published that, I think a lot of other labs have also found this to be true. A lot of models these days are being built with hybrid SSM approaches. For example, Qwen has done that. Kimi is using what they call Kimi linear attention these days. So it's become quite widely adopted to use some sort of state-space model in conjunction with full attention for the base architecture. Now it also has some speed benefits because the amount of memory that you need to hold that state-space cache is actually constant with respect to your sequence length, which then means that generally you can fit much higher batches on the GPU when you're training and doing inference, because the memory requirement is lower. And it keeps the GPU fuller and busier, and therefore provides some pretty important efficiency benefits as well.

混合专家(MoE)架构 Mixture of Experts (MoE) architecture

Host

那么,这些模型也基于 MoE 专家混合架构。我们来介绍一下,也许先提醒大家 MoE 到底是什么。

So, the models are also based on an MoE mixture of expert architecture. We'll go through that and maybe remind people what MoE is in the first place.

Bryan Catanzaro

专家混合是一种稀疏性形式。其想法是,你想在整个互联网上训练一个模型。你想让它记住关于一切历史的绝对一切。但是,当你在回答一个特定问题时,它真的需要思考整个宇宙才能回答那个问题吗?实际上,不需要。这看起来相当稀疏,对吧?我们使用语言模型来探索一个非常小的想法空间,以回答问题或解决问题。我们希望模型能够从整个宇宙中汲取知识。

So, mixture of experts is a form of sparsity. The idea is, you want to train a model on the entire internet. You want it to remember absolutely everything about the history of everything. But, when you're answering a particular question, does it seem reasonable that it needs to actually think about the entire universe in order to answer that question? Actually, no. It seems like it's quite sparse, right? It seems like we're using a language model to explore a very tiny space of ideas in order to answer a question or solve a problem. We want the model to be able to draw from the entire universe.

长上下文窗口与智能体工作流 Mixture of Experts (MoE) Architecture

Bryan Catanzaro

我们希望训练模型,让它尽可能理解一切。但在实际运行时,它并不需要看到所有信息。有很多稀疏性方法试图利用这一特性,但混合专家模型(MoE)是最成功的。它的工作原理是:神经网络有一个学习得到的路由器,对于流经模型每一层的每个 token,它会决定将激活发送给哪些专家子集。路由器会做出选择,决定模型的哪一部分实际与该 token 交互,从而理解它、构建问题表示,然后生成我们要输出的下一个 token。

We want to train it so that it understands everything that it possibly can. But, when it's actually running, it doesn't really need to see all of that information. There's been a variety of approaches to sparsity that try to take advantage of this property, but mixture of experts has been the most successful. And the way that it works is that the neural network has what's called a router that is learned that is going to decide to send activations to a subset of the experts for every token that's flowing through every layer of the model. It's going to be making choices about which fraction of the model is going to actually get to interact with this token as we try to understand it, build up representations of the problem, and then generate the next token that we're going to output.

Host

所以,这有点像我的公司有 550 名员工,但其中 55 人在工程部门。我希望那 55 名专家来参加我的工程会议,而不是公司其他人。

So, it's a little bit like if I have a company with 550 employees, but 55 of them are in engineering. I want the 55 employees who are specialists to come to my meeting about engineering and not the rest of the company.

Bryan Catanzaro

没错。或者你可以把它想象成一个图书馆。如果你去图书馆做研究,你不会读图书馆里所有的书。你的首要任务是弄清楚需要看哪些书才能找到问题的答案。这就是 MoE 背后的思路。MoE 对我们构建的系统有着深远的影响。例如,在 Blackwell 上,英伟达全力投入 MoE。这就是为什么我们构建了 NVL72,它允许最多 72 个 GPU 以极高速率、极低延迟读写彼此的内存。为什么这很重要?因为当你将 token 送入层堆栈时,每一层都有一个路由器将该 token 路由到别处。为什么不把专家分区,不让每个 GPU 上都放所有专家,而是将专家子集分配到每个 GPU,然后在 token 通过网络时动态地在 GPU 之间路由 token?这无法提前预测 token 需要去哪里,因为它非常特定于该 token 和该模型。这就是为什么我们构建了 NVL72,也是为什么 Blackwell 对当今 AI 模型的推理如此出色,因为我们在构建它时深入考虑了混合专家模型。这也说明了 NeMo Tron 的首要任务。如果我们没有致力于理解 AI,我们就无法正确构建 Blackwell。这直接转化为 Blackwell 部署的增加,我们对此感到非常兴奋。

That's right. Yeah, or you can think about it as a library. Like if you go into a library to do research, you don't read all of the books in the library. Your first job is to figure out which books do you need to look at in order to find the answer to your question. So that's kind of the idea behind MOEs. Now, MOEs have fascinating implications for the systems that we build. So with Blackwell, for example, Nvidia went all in on MOEs. That's why we built NVL72, which allows up to 72 of our GPUs to read and write each other's memory at very high speeds, very low latency. And why is that important? It's because as you put a token through the stack of layers, at every layer, you have a router that's routing that token somewhere else. Why don't you partition your experts so that the experts are not sitting every expert on every GPU, but you have a subset of the experts assigned to each GPU, and then you're routing the tokens between the GPUs very dynamically as you push the token through the network. Now, this is impossible to predict in advance where the tokens need to go because it's very specific to that particular token for that particular model. And so that's why we built NVL72, and that's why Blackwell is so amazing for inference for today's AI models, because we thought deeply about mixture of experts when we were building it. And this is speaking to NeMo Tron's first job. If we hadn't been working on understanding AI, we wouldn't have been able to build Blackwell properly. And that has translated directly into increased deployment of Blackwell, which we're very excited about.

Host

你刚才描述的是叫潜在 MoE(latent MOE),还是另一个概念?

Is what you just described called latent MOE, or is that a different concept?

Bryan Catanzaro

潜在 MoE 是我们在 NeMo Triton 3 系列中的一个特定创新,它实际上通过降维投影来减少 MoE 计算期间需要通过 NVLink 传输的通信量。每个 token 产生一个向量,我们的想法是:学习一种压缩该向量的方法,然后将压缩后的数据通过网络发送,最后在另一端解压缩。结果是我们节省了网络带宽,并且在相同的推理成本下获得了四倍数量的专家。你可以把它想象成:我们的图书馆规模变大了四倍,而由于这一创新,我们可以在相同的推理成本下阅读四倍多的书籍。

Latent MOE is a specific innovation that we have in NeMo Triton 3 family and what it does is actually reduces the amount of communication that has to be sent through NVLink during MoE computations by basically down projecting it. So, every token produces a vector and the idea is we're going to take that vector and learn a way to compress it and then send that compressed thing through the network and then we're going to uncompress it at the other end. And as a result, we save on network bandwidth and we also get four times the number of experts for the same inference cost. So, you could think about it as like, our library of books got four times bigger and we get to read four times more books at the same inference cost because of this particular innovation.

Host

MoE 总体上是否正在成为前沿 AI 的默认架构?

Is MoE in general becoming the default architecture for frontier AI?

Bryan Catanzaro

是的,我相信 MoE 长期以来一直是前沿 AI 的默认架构。它们只是推理成本和智能之间的一个非常好的组合。

Yeah, I believe MoEs have been the default in frontier AI for a long time. They're just a really good combination of inference cost and intelligence.

Host

太好了。

Great. Great.

Bryan Catanzaro

但它们也有缺点。它们需要更多内存。如果你的内存非常小,密集模型会更智能。而且它们往往在两种情况下表现最佳:要么你运行批大小为 1,即基本上运行单个作业;要么你运行一个拥有无限查询的大型数据中心。在中间情况下,它们可能有点棘手。

But they have drawbacks as well. They take a lot more memory. If you have a very small amount of memory, a dense model is going to be smarter. And they also tend to work best either if you're running a batch size one, so you're running basically a single job or you're running a huge data center with like infinite queries coming in. In the middle, they can be a little bit tricky.

多令牌预测 Long Context Window and Agentic Workflows

Host

NeMo Triton 3 Ultra 的另一个重要特性是 100 万 token 的上下文,即长上下文窗口。这在整体组合中有多重要?它使模型能够做什么?

Another important characteristic of NeMo Triton 3 Ultra is a 1 million token context, the long context window. How important is that in the overall mix and what does it enable the model to do?

Bryan Catanzaro

上下文长度越长,我们能通过语言模型解决的问题就越有挑战性。这使我们能够将各种信息附加到查询中,比如代码库或指令。长期来看,我希望拥有自己的个人 LLM,能够读取我所有的电子邮件,并帮助我回答相关问题。我们能附加到特定查询的信息越多,模型就越有用。不过,对大量输入数据进行推理的成本会越来越高。这就是为什么上下文长度通常有限制的原因之一。但在 NeMoTron 3 上,我们尽力将其推到极致。我们认为 100 万 token 已经很多了,你可以用它做很多事情。

The longer the context length, the more challenging problems we can solve with a language model. That allows us to do things like append all sorts of information to a query, which could be a code base, it could be instructions. In the long term, I'm hoping that I have my own personal LLM that's able to read all of my emails and help me answer questions about that. The more information that we can attach to a particular query, the more useful the model can be. Now, it can get more and more expensive to reason over large amounts of input data. And so that's one of the reasons why there's usually a limit on how big the context length can be. But with NeMoTron 3, we tried to push it as far as we could go. We think a million tokens is a lot of tokens, and you can do a lot of things with that.

Host

该模型在多步骤、智能体式工作流中是否特别有用?还有一个关于上下文压缩的独立讨论,以确保模型不会在过多的 token 中迷失。你们对此怎么看?

Is the model particularly helpful in multi-step, agentic workflows? And there's this whole separate discussion around the context compaction to make sure that the model doesn't get lost in too many tokens. So, how do you all think about this?

Bryan Catanzaro

百分之百。我的意思是,如果你使用智能体式工作流,压缩是你一直要处理的事情。压缩往往效果不错,因为语言模型很擅长识别最相关的内容并进行总结。压缩时你基本上是在总结上下文。所以,压缩并不是一个坏方法。我认为,能够原生地推理大量数据的模型本质上更有用。所以,我们当然也想在这方面推动边界。

100%. I mean, compaction is a thing if you're using an agentic workflow you deal with all the time. And compaction tends to work pretty well, because language models are pretty good at identifying the most relevant things and summarizing. And you're basically trying to summarize your context when you compact it. So, compaction is not a bad approach. I think having models that can just natively reason about larger amounts of data is just inherently more useful. So, of course, we want to push the boundary on that as well.

Host

好的。你能谈谈多 token 预测吗?这也非常有趣。

Right. Can you talk about the multi-token prediction, which is also very interesting?

Bryan Catanzaro

如果你在低批大小下运行——这在数据中心中是为了获得最高交互性,你希望模型尽可能快地响应,成本更高也可以接受,每个 token 的成本可能更高,但你希望尽快得到结果。或者如果你在本地运行,你可能只运行批大小为 1,因为只有你一个人在使用。

If you're running at a low batch size, which is when you are trying to get the most interactivity if you're in a data center, so you want the model to respond as quickly as possible, and it's okay for it to be more expensive, your cost per token may be higher, but you want the result as quickly as possible. Or if you're running locally, so you might be running a batch size one just because you're the only person using it.

数据来源与合成数据 Multi-Token Prediction

Bryan Catanzaro

事实证明,GPU 有额外的执行能力闲置未用。在这些场景下运行时的主要工作实际上是读取权重。你把 token 推过这些权重,然后再从内存中读取更多权重。但如果你把两个甚至五个 token 推过同一组权重,所需时间基本相同。因为昂贵的不是计算 token 通过权重的数学运算,而是从内存中读取所有权重。所有参数都必须加载进来。因此,多 token 预测的思路就是利用这一点,让模型一次预测多个 token。假设模型预测五个 token,我们知道第一个 token 是正确的,后面四个可能正确也可能不正确。那么在下一次传递时,我们把这四个 token 塞回模型,运行一遍。最后,模型会预测另一组 token,然后我们检查上次额外预测的 token 是否正确。如果正确,我们就直接接受,从而获得大约 4 倍的加速。如果不正确,我们只接受正确的部分,然后继续。这样做的好处是它完全不会降低准确率,因为你用模型进行了双重检查。所有的推测都会在你下一次运行模型时被验证。因此,开启多 token 预测不会降低准确率,但可以带来加速,加速效果取决于预测器的接受率。如果预测器更准确,接受率更高,加速效果也更好。在我们最近的 NeMo Triton 模型中,我们对接受率感到自豪,但一直在努力提升它。这是加速计算的一个很好的例子。在多 token 预测中,速度是模型准确率的函数。模型越准确,推理越快、越便宜,也越准确。这通常不是这样的,但在这里就是这样。这意味着,作为 Nvidia 公司,如果我们想为世界上最重要的计算工作负载提供有意义的加速,这必须是我们思考的重要组成部分。如果推理(2026 年最重要的计算工作负载)能实现 3 倍的成本降低或速度提升,而这又取决于多 token 预测网络的准确率,那么 Nvidia 需要深入理解这一点,因为它会直接影响我们的业务。

It turns out that the GPU has extra execution capabilities that are just lying there unused. The bulk of the work when you're running in these scenarios is actually fetching the weights from memory. And then you push the token past those weights, and then you fetch more weights from memory. But it turns out if you push two tokens or even five tokens through those same weights, it would cost basically the same amount of time. Because the expensive thing is not doing the math to push the token through the weights. The expensive thing is just reading all of those weights from memory. All those parameters, they have to come in. And so the idea with multi-token prediction is to take advantage of this by having the model predict multiple tokens at once. Let's say that the model predicts five tokens. We know the first token is correct. The next four tokens may or may not be correct. So then what we do is on the next pass, we take those four tokens and we stick them into the model, and then run it through. And at the end, we check, the model then predicts another set of tokens, right? Then we check, were the extra tokens we predicted last time correct? If so, then we just accept them, and then we get like a 4x speed up. And if they were incorrect, then we only accept the ones that were correct, and then proceed from there. So the benefit of this is it doesn't degrade accuracy at all because you're using the model to double-check, right? So all the speculation is going to get checked during the next token that you run through the model. So it doesn't degrade your accuracy at all to turn on multi-token prediction, but it can give you a speed up and it's probabilistic depending on the acceptance rate of your predictor. You know, so if your predictor's more accurate, the acceptance rate goes higher, you get a higher speed up. So with our recent NeMo Triton models, we're pretty proud of our acceptance rates, but we're always trying to make them better, always trying to improve that acceptance rate. This is a really good example of accelerated computing. With multi-token prediction, the speed that you get is a function of the accuracy of your model. The more accurate your model is, the faster the inference is, the cheaper the inference is, the more accurate it is. That's not usually how it works, but in this case, that's how it works. And what that implies is that if we're trying as Nvidia as a company to provide meaningful acceleration to the world's most important computational workloads, this has to be an important part of how we think about it. If there's a 3x cost reduction or speed improvement for inference, which is the most important computational workload of 2026, if that's on the table and it depends on the accuracy of the multi-token prediction network, then that's something that Nvidia needs to understand very deeply because it's going to affect our business directly.

Host

太棒了。接下来继续,多教师蒸馏。我们之前简单提过蒸馏。在 NeMo Triton 3 的背景下,这具体是什么意思?

Fascinating. And then to continue on the tour, multi-teacher distillation. We talked about distillation a little bit up front. What does that mean in the context of NeMo Triton 3?

Bryan Catanzaro

在 NeMo Triton 3 Ultra 中,我们使用了称为多领域在线策略蒸馏(multi-domain on-policy distillation)的后训练方法。这意味着我们想要改进模型的许多不同方面。例如,科学理解不同于数学定理证明,不同于编程,也不同于智能体交互。在 NeMo Triton 3 中,我们大概有 10 到 15 个这样的教师模型。思路是,你让这些教师模型在某个特定领域尽可能做到极致,不用担心它们是否全能,只让它们在这个领域变得非常聪明。然后你有一组这样的模型,你想创建一个能学会所有领域的模型。我们使用一种称为 MoPD 的特定强化学习技术来实现,很多实验室现在都在用。好处是,由于教师进行监督,它们可以给学生模型提供非常密集的奖励。基本上每个 token 都受到监督,所以学生可以学得很快,最终在所有方面都几乎和所有教师一样好。一个好处是,这确实有助于团队更好地协作。如果没有这样的技术,假设有 500 人努力改进一个模型,一个团队说“我想让它在这方面更好”,另一个团队说“我想让它在那个方面更好”,就会产生拉锯战:“谁赢?”如果你必须做出选择,比如“我优先考虑这个而不是那个”,那么另一个团队会觉得他们的工作无关紧要。这非常困难。这是 2026 年构建 AI 的挑战之一:你必须想办法让人们协作,尽管最终你只构建一个东西。因此,这项技术对于帮助更多人协作、让 NeMo Triton 更强大起到了关键作用。

So with NeMo Triton 3 Ultra, we did post-training using something called multi-domain on-policy distillation. What that entails is that we have many different aspects of the model we want to improve. For example, science understanding is different from math theorem proving, which is different from coding, which is different from agent harness interactions, right? With NeMo Triton 3, I think we had about 10 or 15 of these teachers. The idea is that you take these teacher models and you push them as far as you can go on some specific domain. You just don't worry about making them good at everything, just make them really, really smart at this one domain. Then you have a collection of these models and you want to create one model that learns to be good at everything. And we do that using a specific reinforcement learning technique that a lot of labs these days use called MoPD. The good thing about this is that because the teachers are supervising, they can give really dense rewards to the student model. Basically, every token is getting supervised, and so the student can learn really quickly and then become almost as good as all of the teachers at all of the things. One benefit of this is that it really helps the team work together better. If you don't have a technique like this and you have, let's say, 500 people working to try to make a model better, and one team's like, 'Well, I'm trying to make it better at this thing.' And then another team's like, 'I'm trying to make it better at that thing.' There can be a tug of war where it's like, 'Well, who wins?' And if you have to make a choice like, 'Oh, I'm going to prioritize this one over that one.' Then you make the other team feel like their work doesn't matter. It's just really hard. It's one of the challenges of building AI in 2026: you have to figure out how to get the people to work together, even though you're only building one thing at the end of the day. So this particular technology has been really instrumental in helping more people work together to make NeMo Triton stronger.

Host

太棒了。所以这既是技术问题,也是人类组织问题。

Fascinating. So it's as much a technology question as a human organization question.

Bryan Catanzaro

完全正确。

Exactly.

Host

好的,太棒了。我们先把这个话题放一放,稍后再回来,因为这是个引人入胜的话题。关于你刚才提到的后训练,你们在 NeMo Triton 中做的一件令人兴奋的事情是发布了训练数据。这包括针对特定强化学习任务的行业数据吗?

Okay. Fantastic. Let's put a pin in this and get back to this in a second because it's a fascinating topic. In terms of the post-training that you just alluded to, one of the exciting things that you all did in the context of NeMo Triton is also to publish the data, the training data. Does that include per industry data for specific reinforcement learning tasks?

Bryan Catanzaro

是的。

Yes.

Host

这就是今天这样对话的美妙之处,你们可以真正谈论这些事情。那么,人们从哪里获得用于后训练强化学习的数据?显然,当今世界的一个关键问题是,LLM 或 AI 系统已经变得非常擅长编程和数学。下一个大问题是,它们能否变得擅长法律、咨询以及各种不同领域。封闭模型的黑箱之一就是人们如何做到这一切,他们从哪里获得数据?在你能谈论的范围内,我非常好奇你们是如何做的。

That's the beauty of a conversation like this today where you guys can actually talk about those things. So, where does one get the data from for post-training reinforcement learning focused efforts? Obviously one of the key questions in the world today is that LLMs or AI systems have become great at coding and great at math. The next big question is can they become great at law and consulting and all sorts of different domains. And part of the black box of the closed models is how people go about doing all of this, where do they get the data from? To the extent that you can talk about all of this, I'd be very curious about how you guys have gone about it.

超越编程与数学的泛化 Data sourcing and synthetic data

Bryan Catanzaro

这个问题不好回答,因为相当复杂,但我会说我们依赖几件事。一是我们会从构建可购买数据集的那些公司购买数据。只要我们有权限重新分发或开放那些数据,我们就会作为 NeMoTron 数据工作的一部分去做。对于 NeMoTron,我们力求最大程度地开放我们发布的数据,因为我们的目标是支持整个生态系统。我们的目标不是成为唯一的模型。当我们听说业内其他模型使用我们的数据集来增强它们的 AI 时,我们非常高兴,因为那意味着我们成功完成了让生态系统蓬勃发展的任务。我们也坚信合成数据生成。我们使用大量算力在自己的系统上运行语言模型,生成合成数据,帮助我们的模型在特定领域更好地解决问题。我们也发布了很多这样的数据。做这件事并不简单。AI 总是垃圾进、垃圾出。你必须非常努力地确保你创建的任何合成数据确实能增加价值,帮助模型更智能地泛化和解决问题。这些就是我们构建数据集的主要方式。

It's not an easy question to answer because it is quite complex, but I would say we rely on a number of things. One is that we do purchase data from companies that are building datasets you can purchase. To the extent that we have the rights to redistribute or open up that data, we do as part of our NeMoTron data effort. With NeMoTron, we are trying to be maximally open with the data we release because our goal is to support the ecosystem. Our goal is not to be the only model out there. We love it when we hear of other models around the industry that are using our datasets to make their AI stronger, because that means we're succeeding at our job to keep the ecosystem thriving and growing. We are also big believers in synthetic data generation. We use an enormous amount of compute running language models on our own systems to create synthetic data that helps our models be better at solving problems in specific domains. We release a lot of that data as well. It's not very straightforward to do this. AI is always garbage in, garbage out. You have to work really hard to make sure that any synthetic data you create is actually adding value, helping the model generalize and solve problems more intelligently. Those are the primary ways we go about building our datasets.

英伟达研究组织与NeMoTron协作 Generalization beyond coding and math

Host

既然我们在讨论不同领域的后训练和强化学习,我很好奇你对泛化方面下一步走向的看法。行业似乎正从编码和数学这些具有可验证奖励的领域,向不同行业推进。你认为这就是趋势吗?AI 行业整体能否像覆盖编码或数学那样高效地覆盖接下来的几个领域?

Since we're talking about post-training and RL in different domains, just curious to get your thoughts on where we go from here in terms of generalization. The industry seems to be marching from coding and math, which are domains with verifiable rewards, to different industries. Do you think that this is where things are going and that the AI industry as a whole is going to be able to cover those next few domains as efficiently as coding or math?

Bryan Catanzaro

编码非常特别,因为它是一项智力密集型活动,创造了巨大的经济价值,这意味着我们有海量的 token 可以学习,还有工具可以验证模型是否真的在解决问题。编码在我们心中永远有特殊地位,AI 也会在这方面持续进步,因为我们与它有这种特殊关系。至于其他领域,我兴奋的是强化学习期间 AI 可以学习的环境将变得显著更多样化。我相信强化学习是一种非常通用的教 AI 解决问题的方式。我们才刚刚开始摸索如何应用它。随着环境变得更加复杂,AI 会更深入地理解它试图解决的问题以及它所能采取行动的后果,然后它就能更好地实际解决这些问题。看看我们今天使用的环境,综合考虑下来仍然相当简单。我认为未来几年它们会变得显著更复杂、更多样化。

Coding is really special because it's a very intellectual exercise that created a lot of economic value, which meant we had an enormous amount of tokens to learn from, as well as tooling that allows us to verify whether our models are actually solving problems. Coding is always going to have a special place in our heart and something that AI is going to continue to get much better at because we have this special relationship with it. With regards to other domains, I think what I'm excited about has to do with significantly more diverse environments for AI to learn in during reinforcement learning. I believe that reinforcement learning is such a general form of teaching an AI how to solve problems. We're just getting started at figuring out how to apply that. As our environments get more sophisticated, the AI learns more understanding of the problems it's trying to solve as well as the implications of the actions it can take, then it becomes much better at actually solving those problems. When I look at the environments we're using today, they're still fairly simple all things considered. I think that's going to become significantly more complex and diverse over the next few years.

研究中的GPU分配 Nvidia's research organization and NeMoTron collaboration

Host

你提到了让 500 人协同工作。我们退一步讲讲。请告诉我们英伟达的研究组织。它是如何架构的?如何运作的?

You mentioned making 500 people work together. Let's take a step back. Tell us about the research organization at Nvidia. How is it structured? How does it work?

Bryan Catanzaro

英伟达不是按照组织架构图来运作的。我们确实有架构图,但它并不是理解我们如何工作的最佳方式。例如,我的团队并不属于英伟达的官方研究团队。我的团队实际上是构建 GPU 的那个组织的一部分。而且我的团队并不是唯一构建 NeMoTron 的团队。公司里大概有 10 个团队在不同部门深度参与了 NeMoTron 的构建:企业软件部门、AI 软件部门,以及实际设计 GPU 的那部分也深度参与了 NeMoTron 的构建。所以有非常多不同的团队必须协同工作。我们总是说使命是老板,而不是组织。但这意味着人们必须自己想办法合作,这很有挑战性,因为人类天生是部落生物,我们很难对不熟悉的人友好,或者信任过去没有成功合作过的同事。实际上,NeMoTron 这个名字就反映了这一点。我们有 NeMo 团队,负责构建 AI 软件;还有 Megatron 团队,主要专注于构建大语言模型的系统研究。我们决定合作,然后开始把我们的项目叫做 NeMoTron,反映了这些团队之间的协作。从那以后,NeMoTron 急剧扩张。有更多的团队参与其中。我们在英伟达内部以这种开放的方式构建它,这一点非常重要。我们邀请公司各地的志愿者来帮助构建英伟达的 AI。我们认为这对公司的未来至关重要。随着这个愿景不断发展,越来越多的人想加入,这太棒了。我们对此非常兴奋。这意味着我们必须想办法组织工作,让每个人都有机会贡献、被倾听,并感觉他们的想法在通往影响力的道路上得到了公平评估。我们有一个正式的流程。我们有一个内部网站,人们可以在上面分享想法,然后这些想法会被分配给 25 位负责人之一,他们负责 NeMoTron 构建的不同部分。他们会与这些想法互动。有些想法会得到进一步发展。有些想法会被推迟到下一轮构建新模型时再考虑。但我们正努力以开放和包容的方式构建 NeMoTron,这样我们才能真正作为一家公司齐心协力来构建它。我认为那些懂得如何协作构建 AI 的组织会成功。那些在控制 AI 归属权上挣扎的组织往往会浪费大量精力。我认为英伟达的成功和 NeMoTron 的成功,直接与我们的协作能力成正比。

Nvidia is not structured according to an org chart. We have one, but it's not actually the best way of understanding how we work. My team, for example, is not part of the official Nvidia research team. My team is actually part of the organization that builds the GPU. And my team is not the only team building NeMoTron. There are probably 10 teams around the company that have significant involvement in building NeMoTron in different parts of the company: in enterprise software, in our AI software division, and the part of Nvidia that actually designs the GPU also significantly is involved in building NeMoTron. So there are so many different teams that have to work together. We always like to say that the mission is the boss rather than the organization. But what that implies is that people have to figure out how to work together, which is challenging in the sense that humans are naturally tribal creatures and it's not natural for us to be friendly with people we don't know very well or trust co-workers that we don't have success working with in the past. Actually, the name NeMoTron reflects that. We had the NeMo team, which was building software for AI, and the Megatron team, which was primarily focused on systems research for building large language models. We decided to work together and then start calling our projects NeMoTron, reflecting the collaboration between these teams. Since then, NeMoTron has dramatically expanded. There are so many more teams that are part of the effort. It's really important that we have structured it in this open way inside of Nvidia. We are inviting volunteers from around the company to come help build Nvidia's AI. We think it's very important to the future of the company. As that vision continues to develop, more and more people want to join, which is fantastic. We're really excited about that. It means we have to figure out how to organize the work so that everybody has a chance to contribute and feel heard and feel like their ideas are fairly evaluated on the path towards impact. We have a formal process for doing that. We have an internal website where people share ideas, and then those ideas are assigned to one of 25 different leads that are over various parts of building NeMoTron. They interact with those ideas. Some of those ideas get further developed. Some of those ideas get deferred until the next time we go around building a new model. But we're trying to build NeMoTron in an open and inclusive way so that we can really come together as a company to build it. I think organizations that figure out how to collaborate to build AI succeed. Organizations that struggle with control over who owns the AI tend to waste a lot of effort. Nvidia's success and NeMoTron's success, I think, is directly proportional to our ability to collaborate.

平衡实用与探索性研究 GPU allocation in research

Host

这是我非常关心的问题。但你之前提到,尽管你在 GPU 领域无可争议的领导者工作,你们组织却没有拥有所有想要的 GPU。那么 GPU 和算力的分配是如何进行的?是基于想法的前景还是早期成功?你们会根据成功与否来分配或撤回 GPU 吗?

It's something that I care deeply about. But you mentioned earlier that despite the fact that you work at the number one undisputed leader in GPUs, your organization doesn't have all the GPUs you would want. So how does the allocation of GPUs and compute happen? Is it based on how promising an idea is or early success? Do you give GPUs, withdraw GPUs based on success?

Bryan Catanzaro

嗯,这是一个非常复杂的问题,也是整个行业在分配算力时面临的难题。在 Neumotron,我们有一个预算,并根据我们认为的项目需求来分配算力。我们有一个层级结构:一系列项目组,每个项目组内有一系列项目。每个项目提出请求,我们有一个两周的周期来审查请求和预算,然后按层级做出决策。话虽如此,我认为我们还能做得更好。在算力分配决策上很难,因为每个研究员都相信,只要再多一千倍的 GPU,他们的想法就能改变世界。他们可能是对的。但我们资源有限,无法为每个想法提供一千倍的 GPU。我们必须在限制内运作。所以这是一个充满挑战的过程。我们尽量纳入更多人的观点,以达成一种共同的理解,不一定是共识。有时某个项目可能觉得它理应得到更多 GPU 却没有。我们希望他们能理解为什么另一个项目被视为更高优先级。这个过程一直在改进,还有更多工作要做,使其更透明、更公平。当然,我的首要任务是获得更多 GPU,这样我们就能资助更多项目。

Well, it's a really complicated question and it's a difficult problem for everyone in the industry to figure out how to allocate their compute. Inside Neumotron, we have a budget, and we allocate compute based on what we think the needs of the project are. We have a hierarchy: a set of programs, and inside each program, a set of projects. Each of them puts forward their requests, and we have a two-week cycle where we review requests and the budget, then make decisions in a hierarchical way. Having said that, I think we can still do better. It's hard when making decisions about compute allocation because every researcher is convinced that their idea could change the world if it just got a thousand times more GPUs. And they might be right. Yet we're running at the limit. We don't have a thousand times more GPUs for every idea. We have to operate within our limits. So it's a challenging process. We try to incorporate as many perspectives as possible to create a shared sense of understanding, maybe not agreement. There may be times when one project feels it deserved more GPUs but didn't get them. We hope they understand why another project was considered a higher priority. This process is always improving. There's always more work to make it more transparent and fair. And of course, my number one priority is to get more GPUs so we can fund more things.

英伟达的登月计划:自下而上与自上而下 Balancing useful and exploratory research

Host

你如何平衡有用研究和伟大的探索性研究?

How do you balance useful research with great exploratory research?

Bryan Catanzaro

我的信念是研究需要自举。这是一个鸡生蛋蛋生鸡的问题。每个研究员都相信,只要他们拥有更多资源,他们的想法就能改变世界。研究员有这种感觉很重要,因为如果没有这种信念,你就不会有动力去做疯狂而新颖的事情。所以你必须相信。但如何将这种信念转化为别人能理解并愿意投资的东西呢?这就是鸡生蛋蛋生鸡的问题:一旦你的想法明显很好且影响巨大,获得资源很容易,但如何在缺乏资源的情况下让它变得明显很好呢?解决鸡生蛋蛋生鸡问题的方法就是自举。这是一种迭代的问题解决方法:先做点小事,获得一些信号表明这是个好主意,然后告诉别人。接着再要一点点资源。如果人们看到实验效果不错,他们会说‘这很有趣,我们应该再多做一点。’然后你就走上了正轨。随着时间的推移,通过快速多次迭代,你可以自举到获得大量资源,并吸引更多人,因为他们看到这个想法将改变世界,并希望参与其中。

My belief is that research needs to be bootstrapped. It's a chicken and egg problem. Every researcher believes that if they just had more resources, their idea would change the world. It's important that researchers feel that way because without that conviction, you wouldn't have the drive to do something crazy and new. So you have to believe. But then how do you translate that belief into something others can understand and are willing to invest in? This is the chicken and egg problem: once your idea is obviously good and impactful, it's easy to get resources, but how do you get it to be obviously good without those resources? The way you solve chicken and egg problems is by bootstrapping. This is an iterative problem-solving approach where you do something small, get some signal that it's a good idea, and tell people about it. Then you ask for a little bit more. If people see that the experiment turned out well, they'll say, 'That's intriguing, we should do a little more.' Then you're on track. Over time, by iterating quickly and many times, you can bootstrap to significant resources and attract more people who see that the idea is going to change the world and want to be part of it.

英伟达的创业文化 Moonshots at Nvidia: bottom-up vs top-down

Host

多年来 Nvidia 的登月项目就是这样启动的吗,无论是在 AI 还是其他领域?所以是自下而上,有人提出好主意,而不是 Jensen 说这就是我们要做的?

Is that how the moonshots at Nvidia got started over the years, whether in AI or otherwise? So it was bottoms-up, somebody coming up with a good idea versus Jensen saying this is what we need to do.

Bryan Catanzaro

嗯,Jensen 也有很多好主意,公司对他的想法反应非常迅速,这也很重要。但 Jensen 一直明确表示,‘这是一家由志愿者组成的公司。我们每个人都是因为选择才在这里。我们可以做别的事情,但我们选择在这里。’因此,我们倾向于以自下而上的方式做决策,尤其是早期研究,因为这像是一种邀请:带来你最好的想法,让我们找出所有最好的想法,然后从那里迈出一步。我们有时会有对公司战略重要的自上而下的想法吗?当然有。NVF P4 预训练就是其中之一。作为领导层,我们决定大力投资 NVF P4 硬件。然后到了发明成功使用它的优化算法的时候。我们告诉团队,‘有一个机会。我们正在进行一项重大投资,如果我们能解决这个问题,它将对公司意义重大。’然后我们让感兴趣的人去研究,结果我们成功了。所以这是自下而上和自上而下的平衡。但即使像 NVF P4 这样有重大战略自上而下成分的项目,也始终带有自举的感觉。实际的技术解决方案非常复杂,有很多活动部件,都来自研究人员本身。我的信念是研究总是来自研究人员。你不能精确地告诉研究如何解决问题,因为那样就不是研究了,而是工程。但在 AI 的世界里,最重要的需要解决的问题都带有研究成分,如果要取得进展,就需要给研究人员创新的自由。

Well, Jensen has lots of good ideas too, and the company is very responsive to his ideas, which is important. But Jensen very explicitly says all the time, 'This is a company of volunteers. Each of us is here because we choose to. We could be doing something else, but we choose to be here.' So we tend to make decisions, especially for early-stage research, in a very bottoms-up way because it's an invitation: bring your best ideas, let's figure out what all our best ideas are, and then take a step from there. Do we sometimes have top-down ideas that are important for company strategy? Of course. NVF P4 pre-training is one of those. We decided as leadership to really invest in NVF P4 hardware. Then it was time to invent optimization algorithms that succeed in using it. We told the team, 'There's an opportunity. We're making a big investment, and if we can figure this out, it will be significant for our company.' Then we let the people interested work on that, and as a result, we succeeded. So it's a balance of bottoms-up and top-down. But it always has this bootstrapping feeling, even with something like NVF P4 where there's a significant strategic top-down component. The actual technical solution, which is very intricate and complex with many moving parts, came from the researchers themselves. My belief is that research always comes from the researchers. You can't tell research exactly how to solve a problem because then it wouldn't be research; it would be engineering. But in a world of AI where the most important problems all have a research component, there needs to be freedom for researchers to innovate if we're going to make progress.

对奇点与智能的看法 Nvidia's entrepreneurial culture

Host

听你这么说,我很惊讶英伟达的文化竟然还这么有创业精神。我在一些大公司工作过,肯定有各种政治斗争,你也提到了部落本能。所以我相信这些都在发生,但考虑到公司成立这么久、取得了如此巨大的成功、内部很多人赚了很多钱,它似乎仍然非常具有创业精神,自下而上驱动,也许还是精英管理。这是正确的理解吗?

Listening to everything you're saying, I'm struck by how entrepreneurial the culture at Nvidia still seems to be. So, like I work at some very large companies. I'm sure there's all sorts of politics and you mentioned the tribal instinct. So, I'm sure all of this is happening, but especially given how long the company's been around, the phenomenal success, the fact that people have been making a lot of money internally, it still seems to be very entrepreneurial, bottoms-up driven, maybe meritocratic. Is that the right takeaway?

Bryan Catanzaro

是的,英伟达非常不寻常的一点是领导层的任期。黄仁勋已经经营公司 33 年了,但他不是一个人。公司里还有很多其他非常资深的高管,他们也在公司待了三十年甚至更久,包括我的老板。这些人记得在很小的英伟达工作的感觉,也知道在很大的英伟达工作的感觉。他们对公司有共同的主人翁意识。英伟达这个地方,我们常说没有人会独自失败。这句话本身就是事实,对吧?你在一个公司工作,这是一个整体。你们一起成功,一起失败。你从事加速计算。加速计算是数千种技术的组合。如果其中任何一个未能实现加速,价值就会被摧毁。如果编译器很烂,芯片再好也没用。归根结底,你卖给那些试图构建 AI 未来的研究人员的是时间和能力。如果他们得不到这些,无论是晶体管、数学单元、编译器、库、网络还是其他任何环节未能达到预期,整个组合就会失败,全部价值都会被摧毁。所以,我们在英伟达的文化中深刻理解这一点,它激励着我们合作的方式。

Yeah, one thing that's very unusual about Nvidia is the tenure of its leadership. Jensen Huang has been running the company for 33 years, but he's not alone. There are a lot of other very senior leaders in the company who have been there for three decades or longer, including my boss. And these people remember what it feels like to work at a very small Nvidia. And they know what it feels like to work at a very large Nvidia. They have a shared sense of ownership for the company. You know, Nvidia is a place we often say no one fails alone. And the point of that, that's just a statement of fact, right? You work at a company, it's one company. You all succeed together, you all fail together. You work in accelerated computing. Accelerated computing is the composition of thousands of technologies. If any of them fail to deliver acceleration, the value is destroyed. It doesn't matter whether the chip is great if the compiler sucks. At the end of the day, the thing that you're selling is time and capability to researchers that are trying to build the future of AI. And if they don't get that, it doesn't matter whether it was the transistor or the math unit or the compiler or the library or the networking or anything else along the way that failed to live up to its expectations, the whole thing in composition fails, the whole value is destroyed. And so, we have a deep understanding of that culturally at Nvidia, and it is something that motivates the way that we work together.

公众对AI的认知 Views on singularity and intelligence

Host

也许在结束对话时,我想把视野拉远,听听你这位深入参与这一切的人对未来走向的看法。谁知道几年后会怎样,但未来一两年也许有些可见性。我读到过你不太相信奇点论。这么说对吗?

Maybe to close the conversation, I'd love to zoom out, get your take from the perspective of somebody who's as deep into all of this as it gets about where things maybe going. So, who knows in a few years, but I don't know in the next year or two, maybe there's some visibility. I read somewhere that you're not really a big singularity kind of a person. Is that fair?

Bryan Catanzaro

对。

True.

Host

为什么?

And why is that?

Bryan Catanzaro

嗯,我认为智能是非常多面的。我经常思考这个问题:如果一家公司要找下一任 CEO,他们会找一个国际数学奥林匹克竞赛冠军吗?可能不会,对吧?尽管那些冠军很厉害,我甚至根本无法参加国际数学奥林匹克竞赛。那些人很棒,他们有着非凡的才智。但那不是经营公司所需的才智。如果我们看看文化中其他非常重要的方面,比如音乐家。成为热门音乐家需要什么样的智能?不要以为那全是运气。不是的。这些人很努力,他们以我作为博士可能不理解的方式非常聪明,对吧?我可能没有那种智能。所以,当我思考智能时,我认为它非常多元且依赖情境。它真的取决于具体情况。不仅仅是原始智力。原始智力有点像引擎的马力,但引擎没有轮子哪儿也去不了,对吧?所以,智能的影响很大程度上取决于智能所处的环境、框架和平台。因此,当我想到这一点时,我认为奇点虽然是一个有吸引力的想法,但它实际上是一个错误的想法,因为它没有真正考虑这些其他因素。所以,我相信人工智能将继续快速发展。它将为世界经济各个领域、从事各种工作的人们解锁重要能力。我对它将带来的机遇感到非常兴奋。我也有些担心我们将如何管理这一过渡。我认为过渡对人类来说通常是困难的。我们通常比较保守。将会发生很多变化。这是对我们思维方式、工作方式、学习方式的深刻改变。最终,我相信人类有能力解决这个问题。我们过去已经做到了。这就是我们。我们建造工具。我们建造外部器官来帮助我们解决问题。我们有一个外部胃,我们称之为厨房。它为我们创造了巨大的价值。没有厨房,我们无法吃某些东西,对吧?现在我们正在创造外部大脑。外部胃对我们这个物种的影响非常深远。它导致了农业,进而导致了有组织的社会、城市的建设方式。那么,外部大脑的影响是什么?非常深远。实际上没有人真正知道。但我相信人类解决问题的能力、学习能力以及以有益方式吸收新技术的能力。我也相信我们星球面临的所有问题都需要更多的智能。每一个问题,无论是不平等、气候变化还是其他我认为非常令人担忧的结构性问题,解决方案都需要发明和智能。对我来说,这意味着未来我们真正能创造的工具只有 AI。因为我们面临的问题都与智能有关。无论解决这些问题的技术方法是什么,解决方案总会被称作 AI。这让我对未来充满希望,但也让我们对挑战保持敬畏,因为我们试图找到与这个新外部大脑共存的新方式。但我相信我们学习和改变的能力,我认为最终这会让我们的生活变得更好。

Well, I think that intelligence is just so incredibly multifaceted. I always think about this question: if a company were to be looking for its next CEO, would it find the next CEO by looking for somebody who won the International Math Olympiad? Probably not, right? Even though it's incredible for people like I could never even compete in any way at the International Math Olympiad. And those people are amazing, right? They have just incredible brilliance. That's not the right kind of brilliance to run a company. If we look, for example, at other aspects of our culture that are really important. For example, musicians. What kind of intelligence does it take to become a hit musician? Don't assume that it's all luck. It's not. These people are working hard, and they're very smart in ways that I might not understand with my PhD, right? I might not have that kind of intelligence. And so, when I think about intelligence, I think it's just so multifaceted and so contextual. It really depends on the situation. It's not just about raw intelligence. Raw intelligence is kind of like the horsepower of an engine, but an engine running without wheels doesn't go anywhere, right? So, the impact of intelligence has a lot to do with the context that the intelligence is put in, the harness, the platform. And so, when I think about that, I think the singularity, although it's an attractive idea, I think that it's really a wrong-headed idea because it doesn't really take into account these other factors. So, I believe that artificial intelligence is going to continue to develop at a rapid pace. It's going to unlock significant capabilities for people in every aspect of our world economy, people doing every kind of work. I'm very excited about the opportunities that it's going to bring. I am also a little bit concerned with how we're going to manage the transition. So, I do think that transitions are hard for humans in general. Like we're conservative generally. And there is going to be a lot of change. This is a profound change in the way that we think, in the way that we work, the way that we learn. Ultimately, I have faith in our ability as humans to figure it out. We've done it in the past. This is who we are. We build tools. We build external organs that help us solve problems. We have an external stomach. We call it a kitchen. It creates enormous value for us. We can eat things that we couldn't eat without a kitchen, right? Now we're creating an external brain. The implications of the external stomach were pretty profound for us as a species. They led to agriculture, which led to organized societies, the way our cities are built. So, we think about what is the implications of an external brain. Pretty profound. Nobody actually really knows. But what I do believe in is the power of humanity to solve problems and to learn and to incorporate new technologies in ways that benefit us. I also believe that the problems we face as a planet all require more intelligence. Every single one of them, whether that's inequality or climate change or any of the other structural problems that I think are very worrisome that we face, the solution to those are going to require invention and intelligence. And what that means for me is that the only kinds of tools that we can really create moving forward are going to be AI. Because the problems that we face are all about intelligence. And regardless of the technological approach to solving those problems, the solutions will always be called AI. And so that makes me hopeful for the future, but also, you know, somewhat respectful of the challenge that it is going to bring to us as we try to figure out how to live in a new way with this new external brain. But I believe in our ability to learn and to change, and I think ultimately this is going to make our lives better.

安全与开源vs闭源 Public perception of AI

Host

你们是否感受到内部正在形成的对 AI 的抵制?这是你们感知到并思考的问题吗?如果是,考虑到你刚才提到的 AI 的种种明显潜力,你认为这是我们行业可能存在的沟通问题吗?

Do you guys feel the AI backlash that seems to be forming internally? Is that something that you all perceive, think about, and if so, do you think it's a communication problem that our industry may have, given what you just said about all the obvious potential of AI?

Bryan Catanzaro

我一直担心公众对技术的看法和互动方式。这很重要。确实,渴望技术进步的社会比不愿改变的社会拥有更多的技术进步。所以我认为思考这个问题很重要。关于 AI 有趣的一点是,我相信当它成为日常生活的一部分时,人们更容易接受它,那时人们不再把它看作 AI,而只是“哦,这是我用的工具”。比如,当你用地图应用导航时,你会在意是不是 AI 在帮你规划路线吗?实际上背后有复杂的 AI,但你不会去想它,对吧?你只是在用工具。所以我觉得人们对 AI 的接受是随着经验而来的。我们使用它的经验越多,就越学会如何高效地使用它,也就越感到舒适。

I'm always worried about the way the public thinks about technology and interacts with it. It matters a lot. It is definitely the case that societies that want technological advancement have more technological advancement than societies that don't want change. So I think it is actually important to think about it. One thing that's interesting about AI is that I believe it tends to be much more accepted when it is part of everyday life, and at that point people stop thinking about it as AI. It's just, oh, this is the tool I use. Like, do you care whether it's AI that's helping you route your car when you ask the map application to help you drive somewhere? There is actually sophisticated AI going into that, but you're not really thinking about that, right? You're just using a tool. So I feel like people's acceptance of AI comes with experience. The more experience we have working with it, the more we learn how to work with it productively. I think the more comfortable we become with it.

结束语 Safety and open vs closed source

Host

好的,Brian。这真是一场精彩的对话。作为最后一个问题,我想确保我们谈到了安全。目前安全的状况如何?开源和闭源在今天的安全讨论中处于什么位置?

Great, Brian. So it's been a fascinating conversation. Maybe as a very last question to make sure we cover it, I want to make sure that we talk about safety. What is the state of safety currently and where does open source and closed source sort of fit in the safety conversation today?

Bryan Catanzaro

安全是现在每个人都在想的事。看到 Fable 的发布以及政府的反应,我认为这是对这些模型安全担忧的结果。它们越来越强大,然后可能被滥用。关于安全,有不同的思考方式和定义方式。我对此可能有一个稍微非正统的观点:我认为开放技术通常更安全,因为阳光更充足。当更多人思考一项技术的安全性、评估它并贡献使其更安全时,这本质上比让一小群人负责所有人的安全更安全。我还认为,对于人工智能,因为它关乎思想,关乎以不同方式探索思想,这种多样性比单一文化更安全。这意味着会有不同的信念。多样性不只是关于容易的事;多样性是关于困难的事。当人们有深刻的意见分歧,他们真的完全不同意对方。让人们能够以多样化的方式探索他们的想法,我认为这比试图创建一个围墙花园(其中某些想法被认为是安全的,某些被认为是不安全的)更安全。这在今天的 AI 环境中是有争议的,我认为这很有趣,因为我们有数百年的传统直接说明了这一点。例如,在美国,我们有关于良心自由和言论自由的法律。这不是因为我们几千年来没有考虑过如果没有这些法律是否会更安全。我们试过。我们实际上尝试过单一文化,规定哪些想法可以安全地谈论,哪些可以安全地相信。我们发现这远不如多元主义安全,在多元主义中,我们官方不对哪些想法安全持立场。我们实际上发现,作为一个社会,支持多样性比自上而下地试图让每个人都安全要安全得多。因此,我相信 AI 的开放技术本质上是构建 AI 最安全的方式。

Safety is on everybody's minds right now. Watching the Fable release and the way the government interacted with that, I think is a consequence of concerns about safety about these models. They get stronger and stronger and then they could be misused. There are different approaches to thinking about safety and trying to define safety. I have maybe a slightly unorthodox opinion about this, which is that I think open technologies are generally safer because there's more sunlight. When more people are thinking about the safety of a technology and evaluating it and then contributing to making it safer, I think that's inherently safer than having a small group of people being in charge of safety for everyone else. I also think with artificial intelligence, because it is really about ideas, it's really about exploring ideas in different ways. That diversity is more safe than monoculture. What that means is that there's going to be different beliefs. Diversity isn't just about the easy stuff; diversity is about the hard stuff. When people have deeply felt disagreements, they really totally disagree with each other. Making it possible for people to explore their ideas in a diverse way, I think it's more safe than trying to create a walled garden where certain ideas are considered safe and certain ideas are considered unsafe. This is controversial in today's AI environment, which I think is interesting because we've had hundreds of years of tradition that speak directly to this. In the United States, for example, we have laws about freedom of conscience and freedom of speech. It's not because we didn't consider for thousands of years whether it would have been safer if we didn't have those. We tried that. We tried actually having a monoculture about which ideas are safe to talk about and which ideas are safe to believe. And we found that to be much less safe than a pluralism where we officially don't take a position about what ideas are safe. We actually found that is much safer as a society to support diversity than it is to try to keep everybody safe top down. And so I believe that open technologies for AI are inherently the safest way of building AI.

Closing remarks Closing remarks

Host

好的,太棒了。用这个有争议的观点来结束对话。Brian,这太棒了。非常感谢。我们很感激你今天抽空和我们交流。

All right, love it. Controversial take here to close the conversation. Brian, it's been fabulous. Thank you so much. We appreciate you spending time with us today.

Bryan Catanzaro

谢谢邀请。

Thanks for inviting me.

Host

嗨,我是 Matt Turk。感谢收听本期 MAD 播客。如果你喜欢这期节目,如果你还没订阅,我们非常感激你能考虑订阅,或者在你观看或收听本期节目的任何平台上留下好评或评论。这真的有助于我们发展播客并邀请到优秀的嘉宾。谢谢,下期再见。

Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

互动版:逐字朗读 + 针对本期提问 →