Building Organizational Superintelligence: Misha on Reflection AI's Vision
打开互动全文版(中英对照 + 朗读 + 问答)→前 DeepMind 研究员 Misha 讨论 Reflection AI 通过编程和企业 AI 构建组织超级智能的使命,并推出代码研究代理作为首个里程碑。
Former DeepMind researcher Misha discusses Reflection AI's mission to build organizational superintelligence through coding and enterprise AI, launching a code research agent as the first milestone.
欢迎回到 Matt 播客。今天的嘉宾是 Misha。Misha 是前 DeepMind 研究员,在交付 Gemini 1.5 后离开谷歌,坚信大语言模型已经跨越了一个关键门槛,并创立了 Reflection AI。这家初创公司最近结束隐身模式,获得了 1.3 亿美元融资,目标是构建超级智能自主系统。本期我们深入探讨了超级智能的含义、构建它的所有模块如今已存在,以及编码和企业 AI 是实现它的最佳路径。我们还聊了他们今天发布的全新产品。最后,我们讨论了在大型 AI 实验室和科技巨头疯狂挖角顶尖 AI 人才的当下,打造超级智能初创公司是什么体验。请享受与 Misha 的这场深刻对话。嘿 Misha,欢迎来到 Mad 播客。感谢你今天亲临现场。
Welcome back to the Matt podcast. Today my guest is Misha. Misha is a former DeepMind researcher who left Google after shipping Gemini 1.5, convinced that LLMs had crossed a critical threshold, and started Reflection AI, a startup that emerged recently out of stealth with $130 million in funding and a goal to build superintelligent autonomous systems. We talked about superintelligence a lot in this episode: what it means, how all building blocks to create it already exist today, and how coding and enterprise AI are the best paths to build it. We also chatted about their brand new product announced today as we publish this podcast. Finally, we covered what it's like building a superintelligence startup at a time when the big AI labs and big tech companies are poaching top AI talent left and right. Please enjoy this deeply insightful conversation with Misha. Hey Misha, welcome to the Mad Podcast. Thanks for being here in person today.
谢谢 Matt,非常感谢你的邀请。
Yeah, thanks Matt. Thanks so much for having me.
你是 Reflection AI 的联合创始人兼 CEO,这家初创公司非常令人兴奋,直到最近还基本处于隐身模式。几周前你们才公开亮相,今天我们要讨论一些激动人心的产品发布。在此之前,请介绍一下公司本身。
You are the co-founder and CEO of Reflection AI, which is a super exciting startup that up until recently was pretty much in stealth. You sort of came out of stealth a few weeks ago and you have some exciting product announcements that we're actually going to discuss today. Before we do that, tell us about the company itself.
我们是一家 AI 研究与产品公司,由我和联合创始人 Giannis 创立。我们两人都是 DeepMind 的长期研究员。团队主要由曾在 DeepMind、OpenAI 和 Anthropic 等公司从事大语言模型训练和强化学习系统构建的人组成。我们聚在一起创建这家公司,共同设计研究和产品来构建超级智能。我认为超级智能听起来可能很抽象,但在初创公司做这件事的特别之处在于,你可以非常具体地定义它对你意味着什么。通常,人们想到超级智能时,可能会想到它在数学、奥赛或竞技编程方面变得非常出色。但这些虽然令人印象深刻,却不太清楚对最终用户有何用处。如果你同时设计产品和研究,就可以对目标方向有更明确的看法。在与许多组织和企业客户交流后,我们意识到,组织级超级智能的形态——真正能帮助组织高效运转的东西——很可能类似于一个预言机。这个预言机深刻理解整个组织,能够像团队中最资深的人那样回答问题,比如首席工程师或掌握公司全局的最高级销售主管。它可能同时处理所有这些事情,而不是拥有不同的功能。所以,我们认为组织级超级智能将是一个深度理解组织的系统。
We're a research and product company in AI, founded by myself and my co-founder Giannis. We are both longtime researchers at DeepMind. The team is largely composed of folks who did large language model training and reinforcement learning, building reinforcement learning systems within companies like DeepMind, OpenAI, and Anthropic. We came together to build this company that co-designs research and product to build superintelligence. I think superintelligence can sound very abstract, but what makes it special doing this at a startup is that you can be very concrete about what that means to you. Generally, when people think about superintelligence, they might think of something that becomes really good at mathematics, like Olympiads or competitive coding. But these things, while impressive, are not clearly useful for an end user. If you co-design a product and the research alongside it, you can be much more opinionated about where you want to go. After talking to many organizations and enterprise customers, we realized that the form factor of an organizational superintelligence—the thing that will really help organizations get a lot done—is probably something like an oracle. An oracle that understands the entire organization extremely deeply, can answer questions at the level of the most senior person on that team, like a principal-level engineer or the senior-most sales leader who has full context on the company. It will probably be superintelligent in the sense that it does all those things at once, as opposed to having distinct functions. So that is what we think an organizational superintelligence will look like: a system that deeply comprehends the organization.
这是一个很好的术语,组织级超级智能,相对于超级智能这个非常抽象的概念。关于超级智能这个概念本身,行业里是否真的就它与 AGI 的区别达成了一致?你知道,我们录制这期节目时,大新闻是扎克伯格用巨额资金挖角顶尖研究人员,从各个组织挖人创建超级智能实验室,然后拯救超级智能作为一家初创公司。人们是否就它的含义达成共识,还是说这只是一个方向性上令人印象深刻,但并非每个人都确切知道其含义的东西?
It's a great term, organizational superintelligence, as opposed to the super abstract concept of superintelligence. Just to say for a second on the concept itself, superintelligence, does the industry actually agree on what that means versus AGI? So, you know, we're recording this. Obviously, the big news is our friend Zuckerberg showering top researchers with extraordinary amounts of money and poaching them from various organizations to create a superintelligence lab, and then this save superintelligence as a startup. Do people agree on what that means, or is that just something that directionally feels very impressive but not everybody exactly knows what it means?
我会说,在这些大型实验室的语境中,超级智能实际上被用作 AGI 曾经的同义词。在《夺宝奇兵》的第一场戏中,他去偷一个雕塑,用一个沙袋替换它,希望一切不变。我有点觉得……
I would say that superintelligence in these large lab contexts is actually being used synonymously with what AGI used to be used for. There's a scene in Indiana Jones and the Raiders of the Lost Ark, the first scene where he goes and tries to steal a sculpture and replaces it with a sandbag, hoping nothing changes. I kind of think one of those...
如果我没记错,那场戏结局并不好。
That scene does not end well if I remember correctly.
结局确实不好,但比喻到此为止,除非我们真正解决安全问题。我认为发生的情况是,根据许多说法,很多人会认为 AGI 尚未实现,但现在有些人可能会认为 AGI 已经实现了。所以,这并非一个尚未实现的二元事件,尽管以前是这样。我认为目标被移动了:‘哦,我所说的 AGI 实际上是超级智能。’所以它和以前的 AGI 一样定义模糊,而且我认为没有人真正就 AGI 的含义达成一致。
It doesn't end well, but the analogy stops here unless we really figure out safety. I think what happened is that by many accounts, many people would argue that AGI hasn't been achieved, but now some people might argue that AGI has been achieved. So it's unclear that this is a binary event that hasn't been achieved yet, even though that was the case before. I think the goalpost just moved: 'Oh, what I meant by AGI was actually superintelligence.' So it's as vaguely defined as AGI was before, and I don't think anyone actually agreed on what we meant by AGI.
好吧。不过听起来确实很酷:2025 年的超级智能就是 2024 年的 AGI。好的。回到正题,我真的很喜欢组织级超级智能这个术语。
Okay. I mean it does sound cool though: 2025 superintelligence is the 2024 AGI. Okay. All right. So coming back, I really like that term of organizational superintelligence.
在你的比喻中,一个资深人员所知道的东西听起来是不是更有吸引力?尤其我假设是一个真正做事的资深人员,对吧?因为通常组织里最资深的人实际上与一线发生的事情脱节。
Does that sound more attractive in your analogy of what a senior person would know? Especially I assume a senior person that actually does the work, right? Because typically the most senior people in an organization are actually disconnected from the reality of what happens in the troops.
是的,我认为更像是那种——每个团队里都有一个靠谱的、经验丰富的人,他们深入一线,什么都懂。先构建这样的系统,然后他们能在脑子里容纳更多上下文,甚至比那些关键团队成员还要出色。我觉得这就像一个组织的预言家。超级智能就是这样的,因为一旦你深刻理解了在任何学科或企业中需要解决的问题,采取行动去解决它就是容易的部分。所以我们一开始非常专注于编码,原因有很多。我们觉得如果你解决了这个问题,它就能解决更普遍的问题。
Yeah, I think it's more like the person who—on every team there's a go-to seasoned experienced person who's very much in the weeds and understands everything. Getting systems first that are like that, and then they can hold a lot more context in their heads and sort of be even better than those critical members of the team. I think that's kind of like an oracle for an organization. That's what a superintelligence looks like, because once you deeply understand the problems that you need to solve in any discipline or enterprise, acting upon them to actually go and solve it is the easy part. And so we're really focused on coding to start, for a number of reasons. We kind of think if you solve this problem, it sort of solves the more general problem.
是的。那些原因是什么?
Yeah. What are those reasons?
一个原因是,我们今天认为编码只是软件工程师的事。但当你构建一个编码模型时,你实际上训练了一个语言模型,它可以通过代码与任何软件交互。这些语言模型与任何软件交互的方式——不仅仅是像 Salesforce 和其他 CRM 及创意工具这样的软件工程软件——大多数交互将通过函数调用、API,也就是通过代码进行。所以今天我们觉得编码模型只是为软件工程师准备的,但事实并非如此。如果你构建了那个系统,它实际上可以与任何软件交互。这就像你构建了一个数字 AI 的手和脚,对吧?就像人形机器人可以用和人类一样的手和脚做任何事情。
One reason is that the way we think of coding today is just for software engineers. But when you build a coding model, you've just trained a language model that can interact with any piece of software through code. The way these language models are going to interact with any piece of software—not just software engineering software like Salesforce and other CRM and creative tools—the majority of those interactions are going to be through function calls, APIs, so through code. So today we think of a coding model as something that is just for software engineers, but that's not true. If you build that system, it can actually interface with any piece of software. It's sort of what you've built is the hands and legs of a digital AI, right? In the same way that a humanoid robot can do anything with the same hands and legs that humans do.
那么编码是大脑,其他一切都成为连接到大脑的腿和手?这个比喻错了吗?
So would coding be the brain and then everything else becomes the legs and the hands that are connected to the brain? Is that the wrong analogy?
是的。你可以把模型看作大脑,然后人们所谓的脚手架或智能体脚手架——那就是可供性,即你能实际做的事情。所有这些软件的可供性,至少对于数字智能来说,将主要通过代码实现。另一种选择是教模型如何拖动鼠标,称为计算机控制。这种情况会发生一些,但我认为这不会是语言模型与软件交互的主要方式。所以如果你解决了编码,你就解决了语言模型应该如何与软件交互的问题。
Yes. It's kind of like you can think of the model as the brain, and then what people call scaffolding or agent scaffolding—that's sort of the affordances, the things that you can actually do. All these affordances for software, at least for digital intelligence, are going to be primarily through code. The other option is teaching a model how to drag a mouse around, called computer control. Some of that will happen, but I just don't think that's going to be the majority of how a language model interacts with software. So if you solve coding, you've just solved how a language model should interact with software.
而且这也是一个更容易处理的问题。这样说公平吗?因为代码更结构化,更接近语法,很适合 LLM 所做的那种工作?
And also it's a more tractable problem. Is that fair because code is more structured and code is closer to grammar, lends itself well to the type of work LLMs do?
不,我认为这非常正确。代码对 LLM 来说几乎是直观的,因为它们最初是在互联网上训练的,所以它们对代码几乎有一种肌肉记忆。这有点像人类进化出了空间地理肌肉记忆,所以我们作为幼儿很快学会如何使用双手。代码对我们来说实际上非常不直观,对吧?这就是为什么 GUI 被发明出来——因为与计算机交互的第一种方式是通过代码,对人类来说非常不直观。所以你们发明了 GUI。对于语言模型来说,情况正好相反。它们在互联网上从未真正见过任何 GUI 数据。那里没有鼠标移动数据。所以它们天生擅长的东西与我们天生擅长的东西正好相反。因此,如果你构建了非常擅长编码的东西,你只是在放大 LM 现有的基础知识。这对语言模型来说是直观的。
No, I think that's very much correct. It's more that code is almost intuitive to LLMs because they were trained on the internet initially, so they kind of have almost a muscle memory for code. It's kind of like humans have evolved to have spatial geospatial muscle memory, so we learn very quickly how to use our hands as toddlers. Code is actually very unintuitive to us, right? That's why the GUI was invented—because the first way to interface with computers was through code, very unintuitive for humans. So you invent the GUI. For language models, it's the opposite. They've never seen any GUI data really on the internet. There's no mouse movement data out there. So the thing that is native to them is the exact opposite of what's native to us. So if you build things that are really good at coding, you're sort of just amplifying the existing base knowledge of the LM. It's intuitive to a language model.
所以这就是为什么是编码。但在编码内部,你可以思考创建一个超级智能编码智能体需要什么。你基本上需要一个能很好生成代码的系统,还需要一个能理解大型代码库及其周围所有知识的系统——比如所有文档中的内容,以及工程师头脑中的那种部落知识。我要说,今天的每个编码工具都专注于代码生成部分——以自主或半自主的方式在你的 IDE 或终端中——但都围绕着代码生成。
So that's why coding. But then within coding, you can think about what it would take to create a superintelligent coding agent. You basically need a system that can generate code really well, and you need a system that comprehends large code bases and all the knowledge around them—like everything that's in all the docs and the stuff that's kind of tribal knowledge that's in engineers' heads. I'd say every coding tool today is focused on the code generation piece—in an autonomous or semi-autonomous way in your IDE or terminal—but it's all around code generation.
结果就是,因为我认为今天的公司并没有真正认真关注理解部分。我们正在构建的东西,我几乎可以称之为——今天也许我们有 AI 工程师达到 L4 工程师的水平,也许明年 L5、L6,一直到 L9。我认为这条路径是清晰的,但如果我们不解决理解部分,我们最终得到的基本上就是患有失忆症的 L9 工程师。
And as a result, because companies today I don't think are really focused in earnest on the understanding piece. What we're building towards I'd call almost like today maybe we have AI engineers that are at the level of an L4 engineer, maybe next year L5, L6, going to L9. I think the path there is clear, and what we're going to get to if we don't solve the comprehension piece is basically L9 engineers with amnesia.
一个 L9 工程师,如果他们来到你的公司但患有失忆症。
An L9 engineer if they came into your company but had amnesia.
他们实际上不会——好消息和坏消息。
They wouldn't be really—good news and bad news.
对,对。这有点像如果你明确告诉我该做什么,我会做得很好,但除此之外,我对你的代码库一无所知。我对你的组织一无所知,我也不会去收集这些信息。我认为现在要解决的问题是构建我们所谓的上下文引擎。就像你的编码智能体可以查询的那个大脑,以便获得它们进行有意义工作所需的所有上下文——不仅仅是初级工程工作,而是工程师花费大部分时间的事情,比如关键基础设施漏洞、处理遗留代码。如果你真的在任何大型组织甚至一个拥有相当大代码库的初创公司里,站在工程师身后观察,你会发现他们只有少数时间在编码。大约 70% 的时间他们在挖掘信息试图理解东西,向其他工程师提问。
Right. Right. It's kind of like if you tell me exactly what to do, I'll go do it really well, but otherwise, I know nothing about your code base. I know nothing about your organization, and I'm not going to gather any of that either. The thing to solve now, I think, is building what we kind of think of as the context engine. Like what's the thing that's the brain that your coding agents can query in order to have all the context that they need to go do meaningful work—not just junior engineering work, but the stuff that engineers spend most of their time on, like critical infrastructure bugs, dealing with legacy code. If you actually hover over the shoulder of an engineer at any large organization or even a startup with a really sizable code base, you'll find that a minority of their time is actually spent coding. Like 70% of their time they're digging through information trying to understand stuff, asking other engineers questions.
而你希望构建一个拥有相同 DNA 的超级智能——它大部分时间都在收集信息,然后一旦有了信息,它就知道如何采取行动。
And you kind of want to build a superintelligence that has that same DNA—that what it's spending most of its time on is collecting information, and then obviously once it has it, it knows how to act on it.
这与将编码和 RAG 结合起来以引入企业上下文的做法有何不同或相似?这是大致方向吗?
How is that different or similar to basically doing a combination of coding and RAG to bring in the context of the enterprise? Is that the general direction?
这有点像,也许几乎就像 RAG 如果有效的话。
It's kind of maybe it's almost like if RAG worked.
所以 RAG 的目的是获取你的编码智能体完成工作所需的所有信息。但这是一个非常原始的方式,而且基本上会失败。它有点……
So the purpose of RAG is to get all the information that your coding agent needs in order to do some work. But it's a very primitive form of doing this and by and large fails. It's kind of...
为什么说它原始?
Why is it primitive?
因为它只是面向搜索的。它只是一个查询。
Because it's just search oriented. It's just a query.
有几个方面,但首先是它很稀疏,因为它从大型代码库中抓取的内容有很多假阴性和假阳性。而且它通常只做一次,对吧?它抓取一次,然后你就只有这些了。对于任何有意义的查询,它都不会给你实际完成任务所需的信息。所以 RAG 智能体相当弱。现在出现了一种新的检索方式,人们称之为智能体式搜索,Claude Code 更接近这种方式。它是一个智能体,使用人类使用的命令行,去查找文件,使用工程师使用的命令搜索关键词并抓取内容,而且它是智能体式地进行的。与一步到位的过程(只是拉入一堆嵌入)不同,它会使用文件搜索工具,然后查看结果,思考,再使用另一个工具,依此类推。
There are a few things about it, but the first thing is that it's what we call sparse, because the stuff it grabs from a large codebase has a lot of false negatives and false positives. And it typically only does it once, right? It'll grab it and then that's all you have. For any meaningful query, it will not have given you the information you need to actually do the task. So RAG agents are pretty weak. What's happening now is a new kind of retrieval that people are calling agentic search, which is more what Claude Code does. It's an agent that uses the same command line that a human does, goes and looks for files, uses the same commands an engineer does to search for keywords and grab stuff, and it does this agentically. As opposed to a one-step process where it just pulls a bunch of embeddings in, it will use a file search tool, then look at what it did, think about it, and use another tool, and so forth.
并且会记住它。是否有内置的记忆概念?
And will remember it. Is there a concept of a memory built in?
它确实会把东西存储在上下文中。
It does store the stuff in a context.
但我的理解方式是,想象你被丢进一个巨大的黑暗丛林,你只有一个小手电筒。这基本上就是今天的智能体式搜索。你在探索,你有这个小手电筒,你必须记住脑子里的一切。这是理解的第一种形式,但也是相当弱的形式。它无法扩展到大型丛林。
But the way I would think about it is that imagine you are dropped into a large dark jungle and all you have is a tiny flashlight. That's basically what agentic search is today. You're exploring, you have this tiny flashlight, and you have to remember everything in your head. That's a first form of comprehension, but it's also a pretty weak form. It doesn't scale to large jungles.
就像如果你后院有一个小丛林,你也许能导航。
Like if you have a tiny jungle in your backyard, then you might be able to navigate it.
一个盆景丛林。
A bonsai jungle.
一个盆景丛林。是的。所以这就是我对现有最先进的智能体式搜索的比喻。而你需要构建的系统是那些能扩大光束孔径的系统,让你看到更多,更聪明地知道去哪里找东西,以及如何记住它们。这些是基础研究问题:记忆、长上下文、推理。但我要说的是,这只是扩大光束孔径,弄清楚如何实际存储它,以及除了代码之外还应该从哪些信息来源获取信息。因为组织的信息存在于他们的聊天记录、项目管理工具中,而且很多都在团队的大脑里。当那位了解某段遗留代码的高级工程师离开时,突然之间,那段代码就成了一个闹鬼的墓地,因为没有人能理解它。
A bonsai jungle. Yeah. So that's how I would liken the existing state-of-the-art agentic search. Whereas the kinds of systems you need to build are ones that expand the aperture of your beam so you're seeing more, are smarter about where they go and look for stuff, and how they remember it. These are fundamental research problems: memory, long context, reasoning. But I would say it's just expanding the aperture of your beam, figuring out how to actually store it, and also what sources of information you should be pulling on other than just your code. Because information for organizations lives in their chats, their project management tools, and a lot of it is in their team's brains. Often when that senior engineer leaves who understands some legacy piece of code, all of a sudden that becomes known as a haunted graveyard because no one else understands it.
那么,如何构建产品来捕捉人们头脑中的知识呢?这就是为什么我认为不能抽象地思考超级智能,因为一半是产品问题。理解问题很深。人们过去常说,当一个问题被标记为 AGI 完备时,意味着这个问题如此之深,以至于如果你解决了它,你就得到了 AGI。我会说这是一个超级智能完备的问题。如果你真的为组织解决了这个编码方面的神谕,你基本上已经构建了拥有超级智能所需的所有能力。
And so, how do you build products that capture that knowledge that's in people's heads? That's why I think you can't think about superintelligence in the abstract, because half of it is a product problem. The comprehension problem is deep. People used to say when a problem was AGI-complete, it meant the problem has so much depth that if you solve it, you get AGI. I would say this is a superintelligence-complete problem. If you really solve this oracle for organizations just for coding, you've basically built all the capabilities you need to have superintelligence.
因为你可以从那里泛化。这是一个流传的想法:编码是通往超级智能的路径。一个变体是智能自动构建智能。但你在这里说的更像是,这是一个如此复杂的问题,如果你解决了它,你就到了。
Because you can generalize from there. It's an idea that made the rounds: coding is a path to superintelligence. A variation is intelligence building intelligence automatically. But what you're saying here is more that it's such a complex problem that if you solve it, you're there.
是的。我认为编码是 ASI 完备的概念在于,你可以使用一个智能编码系统来构建另一个智能编码系统。如果人们喜欢抽象地做研究,那听起来很棒,因为就像我不必考虑实际解决的问题。如果我仅仅构建一个构建智能的智能,它总会以某种方式解决。我不认为这是可行的。我认为构建更好编码智能的编码智能会让算法更高效。也许它如此智能,以至于它可以开始为你做出所有产品决策,比如你应该问用户什么问题,以及你应该构建什么功能。对我来说,这似乎是一个更大、更实际的问题。没有与产品共同设计的研究几乎毫无意义。只有当构建 AGI 或 ASI 的要素未知时,这才有意义。我认为现在它们已知了。
Yeah. I think that the notion of coding being ASI-complete is that you can use an intelligent coding system to build another intelligent coding system. If people like to do research in the abstract, that sounds awesome because it's like I don't have to think about the actual problem being solved. If I just build an intelligence that builds an intelligence, it will figure it out somehow. I don't think that's how it works. I think coding intelligence that builds better coding intelligence will make algorithms more efficient. Maybe it's so intelligent that it can even start making all the product decisions for you, like what questions you should be asking users and what features you should be building. To me, it seems like a much larger, more practical problem. There's almost no research without co-designing product with it. That was only meaningful when the ingredients for how to build AGI or ASI were not known. I think now they are known.
所以你认为我们拥有一切?也许列出一切在组件方面是什么。
So you think we have everything? Maybe list what everything is in terms of components.
是的,我认为我们拥有一切,尽管能力会继续加强。但需要发生的一系列突破,很多已经发生了。还有一些,但我会说主要有四个。第一个是让深度神经网络工作。那是 2012 年的 ImageNet 时刻,当时你可以有一个深度神经网络以人类水平然后超人类水平对图像进行分类。我认为那是第一个突破。第二个突破是强化学习,尽管当时可能不清楚,但现在又清楚了。强化学习告诉我们如何拿一个系统,如果我们有一个奖励,就让它变得超级智能。我们实际上已经多次构建了超级智能。AlphaGo 是超级智能的。AlphaStar,如果你投入更多算力,也会变得超级智能。所以我们已经知道如何构建狭窄的超级智能有一段时间了。
Yeah, I think we have everything, even though the capabilities will continue being tightened. But the series of breakthroughs that need to happen, many of them have happened. There are still some, but I would say there are four of them. The first was just making deep neural networks work. That was the ImageNet moment in 2012, when you can have a deep neural network classifying images at human and then superhuman level. I think that was the first breakthrough. The second breakthrough was reinforcement learning, even though it was not clear maybe then, but now it's clear again. Reinforcement learning told us how to take a system and, if we have a reward, make it superintelligent. We have actually built superintelligence a few times now. AlphaGo was superintelligent. AlphaStar, if you put more compute into it, would have gotten superintelligent. So we've known how to build narrow superintelligence for a while.
所以深度神经网络、强化学习。
So deep neural networks, reinforcement learning.
所以 2012 年,AlphaGo 是 2016 年。好吧。然后是扩展 Transformer,GPT 系列工作。2017、18 年。大概是 18 年,但真正开始发光是在 2020 年代初的 GPT-3。那既是架构创新也是数据创新。它是一种能够消化互联网数据的架构,所以两者兼有。Transformer 不仅仅是架构,它也是数据。然后我想最后一个是 RLHF,这让我们知道了如何以基本方式对齐模型,如何在语言模型之上做基本的强化学习。
So 2012 AlphaGo was 2016 or 2016. Okay. Then scaling up transformers, the GPT series of work. 2017, 18. Kind of 18, but then really I think started shining in the early 2020s with GPT-3. And that was both an architectural innovation and a data innovation. It was an architecture that could consume data on the internet, so it's kind of both. And transformers are not just architecture, it was also data. And then I would say the last ones are RLHF, which was okay, we know how to align models in the basic way, we know how to do basic reinforcement learning on top of language models.
嗯。
Mhm.
然后强化学习又随着推理模型回归了。
And then reinforcement learning coming back again with the reasoning models.
所以如果让我总结,我会说神经网络、深度神经网络、Transformer 和互联网规模的数据,以及各种形式的强化学习。这些是构建超级智能所需的要素。而你描述的强化学习在最后的回归,是否意味着 RLHF 不再必要?那是一个临时解决方案,还是它们是并行能力?
So if I was to summarize, I would say neural networks, deep neural networks, transformers and internet-scale data, and reinforcement learning in its various incarnations. These are the ingredients required to build a superintelligence. And as a return of reinforcement learning at the end that you described, does that obviate the need for RLHF? Was that a temporary solution or are those parallel capabilities?
我认为它们是并行能力,因为它们做不同的事情。RLHF 更多地是为了对齐预训练语言模型。如果你玩过这些未对齐的基础模型,它们真的毫无用处。它们像随机鹦鹉,不遵循指令,熵极高,对吧?能够对齐它们,调整它们使其符合人类消费习惯,我认为这是一个重大突破,那就是 RLHF。而带有推理的强化学习更多的是用强化学习来驱动智能能力,使其在编码、数学或任何有奖励的目标领域表现出色。你可以串联使用它们:推理阶段的强化学习真正扩展了能力,然后 RLHF 将其对齐为人类可消费的形式。
I think they're parallel capabilities because they do different things. RLHF was more designed to align a language model like a pre-trained language model. If you play with one of these base models before they're aligned, they're really useless. They feel like stochastic parrots. They don't follow instructions. They're extremely high entropy, right? The fact that you could just align them, tweak them to align them with something that is human consumable, I think was a pretty big breakthrough and that was RLHF. Whereas RL with reasoning is more RL to drive intelligence capabilities, which is to make it really good at coding, make it really good at math, make it really good at whatever target domain you have rewards for. And you use these in tandem where RL in the reasoning phase really expands on a capability, and then RLHF aligns it to be human consumable.
但它们是同一回事。事实上我会说它们就是同一回事。机制是一样的。
And but they're the same thing. It's in fact I would say it's just the same thing. Like the machinery is the same.
所以我们拥有 AGI、超级智能的所有组件。你的目标是在组织企业环境中创建它。我想我们来谈谈本周的大新闻,那就是你正在推出你的第一个产品。
So we have all the components for AGI, superintelligence. Your goal is to create this in an organizational enterprise context. And I guess let's get into the big news of this week, which is that you're launching your first product.
告诉我们关于它的一切。
Tell us everything about it.
是的,我们正在推出我们的第一个产品,也是通往超级智能道路上的第一个里程碑。它叫 Asimov,就像那位对这个主题有想法的科幻作家。这个 Asimov 是为组织构建的一流代码研究智能体。它不同于编码智能体——那种会为你编写代码的东西,但今天我们觉得那些智能体上下文贫乏。所以如果任务范围明确,现在的编码智能体可以做得很好。但如果你大部分时间都在试图理解工程中的某个棘手问题,为什么某个基础设施错误会发生,或者做工程师大部分时间都在做的那种工作,目前没有工具能真正以那种方式为他们解困。所以我认为 70%的软件工程实际上是代码研究工作,而不是代码苦力工作。我们正在构建一个解决这个问题的产品。工程师大部分时间都在拆解问题,试图理解事情发生的原因,而且常常因为需要问别人问题而遇到瓶颈,那个人要几个小时才能回复,因为他们很忙。我们构建了一个帮助克服这类问题的产品。
Yeah, we're launching our first product and the first milestone on the path to superintelligence. It's called Asimov, like the science fiction writer who had some thoughts on the subject. And this Asimov is the best-in-class code research agent built for organizations. So it's different than a coding agent which is a thing that will go and write code for you, but today we feel those are pretty context-poor. So if the task is well-scoped, coding agents today will do it pretty well for you. If you are spending most your time trying to understand some hairy problem in your engineering, why some infrastructure bug is happening, or really doing what engineers spend most of their time doing, which is this kind of work, there's no tool today that really unblocks them in that way. So I'd say 70% of software engineering is actually a code research job rather than a code monkey job. And we're building a product that addresses that problem. So the majority of time that engineers spend unpacking problems and trying to understand why something's happening, and often times it's bottlenecked because they have to ask someone a question and it takes a few hours for that person to get back because they're busy. We've built a product that helps overcome these kinds of problems.
所以复述一下,就像代码的深度研究。在概念上与代码的深度研究非常相似。是的。它会探索你的代码库和其他知识来源。可能比传统的快速问答产品花费更长的时间,但会返回更好的答案。它的不同之处在于,工程知识不仅仅存在于代码库中。它还存在于各种其他表面领域和其他软件产品中,比如项目管理工具、聊天、文档等等。Asimov 从这些地方提取信息。所以它不仅仅是你的代码库,它聚合了所有这些信息。它有一个我认为至今尚未引入的新概念,我们称之为团队记忆。现在的产品有个体记忆,记住开发者的个人偏好,但它们没有团队层面的组织理解。假设你有一个高级软件工程师,他了解某个微服务 A。你不是那个工程师,现在你需要与他们的微服务交互,而你没有他们拥有的上下文。现在有了 Asimov,它会在聊天中自然捕捉这些信息,而且工程师也可以直接教它。
So just to play it back, like deep research for code. It's very similar in concept to deep research for code. Yeah. A thing that will go explore your codebase and other sources of knowledge. It might take a bit longer than the traditional snappy ask kind of products, but it'll come back with much better answers. Things that make it different are that engineering knowledge does not just live in a codebase. It lives in all sorts of other surface areas and other software products like project management tools, chats, documentation, things like this. And Asimov pulls from those things. So it's not just your codebase, it's aggregating all these sorts of information. It has this new concept that I don't think has been introduced to date that we call teamwide memories. So today products have individual memories which are remembering the personal preferences of the developer, but they don't have a teamwide organizational understanding. Suppose you have a senior software engineer that understands some microservice A. You're not that engineer and you now need to interact with their microservice, and you don't have the context they do. So now with Asimov, it kind of organically captures that information as it happens in chats, but also engineers can just teach it directly.
是的,这太酷了。然后这解决了你之前提到的那个问题,比如如果有人离开,机构记忆也会随之消失。所以你有永久的组织级记忆。
Yeah, that's super cool. And then that addresses the problem you mentioned at some point earlier, which is like if somebody leaves, then the institutional memory goes with them. So you have like permanent organization-wide memory.
完全正确。它正在为你的工程知识构建一个永久的组织记录系统,作为起点。
Exactly. It's kind of building a permanent organizational system of record for your engineering knowledge to start.
是的。关于这一点,有趣的是,一些最兴奋的用户是那些一直在回答别人问题的高级工程师。所以我们看到的是,我们进入一个组织,头三周会有四名高级工程师在填充它的知识。这实际上是我第一次看到工程师对文档感到兴奋。
Yeah. A side note on that is something that's been interesting is that some of the most excited users have been these senior staff-level engineers who are fielding people's questions all the time. And so what we've seen is that we'll go to an organization and the first three weeks there'll be four staff senior staff-level engineers who are just populating its knowledge. It's kind of the first time I've actually ever seen engineers excited about documentation effectively.
这挺独特的。最后一点是关于智能体设计,它使你能够扩大光束孔径,从而能够查看比以前智能体所能查看的更大的代码库。
That's kind of a unique thing. And then the final thing is really around agent design that enables you to increase the aperture of your beam so that it's looking at much larger code bases than agents were able to before.
这仍然是一项正在进行的工作,因为长上下文推理是一个重大的基本问题,我认为还有很多工作要做。但我认为这种新的智能体设计是朝着那个方向迈出的一步。
This is still a work in progress in that long-context reasoning is a big fundamental problem, and I think there's a lot of work to be done there. But I think this new agent design is a step in that direction.
从方向上看,它们为什么以及如何能够做到这一点?
Directionally, why and how are they able to do that?
这并非我们独有。当我观察最先进的智能体是如何设计时,它们被设计成多智能体系统。智能体设计本质上是一个大的推理智能体,这很标准,但它会派遣小的长上下文推理智能体去代码中搜索相关信息的不同部分。所以,这种将大推理智能体(具有较小上下文)与一群小的检索侦察智能体解耦的设计,与迄今为止的做法不同,尽管我相信其他公司也会趋同。你需要根据要解决的问题来设计智能体。如果你想要一个非常快速、能立即回答问题的智能体,那么这可能不是最佳设计。你可能想要一个做 RAG(检索增强生成)的,那很快,或者像 Claude Code 或 Cursor 那样的基本搜索智能体,它使用终端和文件系统,用工程师相同的命令。这比派出许多检索器要快得多。所以我认为我们会看到产品与要解决的问题共同进化,不同的问题会有不同的智能体设计。
This is not unique to us. When I look at how state-of-the-art agents are being designed, they're designed as multi-agent systems. Agent design is effectively a big reasoning agent, and that's pretty standard, but it dispatches small long-context reasoning agents to go search for different parts of relevant chunks of information in the code. So this decoupling of a big reasoning agent with a smaller context and a bunch of little retriever scout agents is a different design than what's happened to date, though I'm sure other companies will converge on it as well. You really want to design your agents for the problems you're trying to solve. If you want a really snappy agent that answers things immediately, then this is probably not the best design. You might want something that does RAG, which is very fast, or a very basic search agent like what Claude Code or Cursor might do, which uses the terminal and file system with the same commands as an engineer. That's a lot snappier than sending out many retrievers. So I think we'll start seeing product co-evolving with the problems you're solving, and there will be different agent designs for different problems.
如果把一个快速的一次性智能体用于小问题,然后试图把它们组合起来,这一定是坏事吗?还是说单个小智能体需要更智能?
Is a snappy one-shot agent necessarily a bad thing if you direct them at small problems and then try to put them together? Or does the individual little agent need to be smarter?
不,我认为这是一个非常有趣的问题。这是现在研究中与以前不同的地方之一:以前只关注模型,但现在很多研究在于智能体设计。所以你的问题是一个开放性问题。你可以有一个混合系统,将一些查询路由到一个智能体系统,其他查询路由到另一个。也可能你找到一个优雅的更简单的多智能体系统,可以同时做这两件事。这是一个开放的研究问题。这非常令人兴奋,而且在某种意义上,它类似于在语言模型之前 DeepMind 和可能 OpenAI 中发生的许多未公开的研究。例如,像 OpenAI 的 Dota 5 或 DeepMind 的 AlphaStar 项目,它们训练了专家级的智能体来玩像《星际争霸》和《Dota》这样的复杂视频游戏。那里有一个大问题,我认为没有多少人意识到:你如何为你的神经网络设计环境来实际执行动作?你能想到的最简单的事情是它像人类一样学习使用键盘和鼠标。结果这行不通。所以当你阅读 AlphaStar 论文时,你会发现他们想出了一种特定的方式来分解动作,以便让神经网络玩《星际争霸》。这实际上就是智能体设计。现在这又回来了,体现在设计脚手架方面。那时它被称为环境设计。但这是同一回事。这些大项目中的很多工作甚至不在于训练模型,而在于弄清楚智能体应该如何与你训练它的环境交互,你从哪里获取数据,你的数据是什么,如何收集数据。
No, I think that's a really interesting question. This is one of the things that's interesting in research now as opposed to before: before it used to be just around models, but now a lot of the research is in your agent design. So what you ask is an open question. You could have a hybrid system that routes some queries to one agentic system and other queries to another. It could be that you figure out an elegant simpler multi-agent system that can do both things. It's an open research question. It's very exciting, and in some sense it parallels a lot of the unspoken research that was happening at DeepMind and probably in OpenAI as well during pre-language models. For example, projects like OpenAI's Dota 5 or DeepMind's AlphaStar project, which trained expert-level agents to play complex video games like StarCraft and Dota. A big question there that I don't think many people appreciated is how do you design the environment for your neural network to actually dispatch actions? The most simple thing you can think of is it learns to use a keyboard and mouse like a human does. That turned out not to work. So when you read the AlphaStar paper, you see that they figured out a particular way of factoring out the actions in order to make neural networks play StarCraft. And what that actually meant is that was agent design. That's sort of now back in terms of designing scaffolding. Back then it used to be called environment design. But it's the same thing. A lot of the project in these big projects was not even on training the models; it was figuring out how the agent should interface with the environment you're training it in, where you get your data from, what your data is, how you collect it.
关于这一点,你提到了多智能体设计等三件事。我认为第一点是它访问哪些数据源。这些智能体访问数据的方式与 RAG(基本上是直接搜索)有什么不同吗?作为一个相关问题,这些数据源是否需要有一个特殊的协议来适应智能体,比如 MCP 风格的基础设施,以便那些到处抓取信息的超级智能智能体能够与它们交互?
So on that point, you mentioned the three things like multi-agent design. I think the first point was what sources of data it accesses. Is there some difference in how those agents access data versus RAG, which is pretty much straight-up search? As a related question, do those data sources need to have a special protocol to lend themselves to agents, like MCP-style infrastructure, so that the super intelligent agents that go around and grab information everywhere can interact with them?
这是一个非常好的问题。有几个有趣的点需要展开。首先我要说的是,根据你要解决的问题,不同类型的搜索基本上都属于搜索范畴。如果你想要非常快,但准确性不太重要,或者只需要大致准确但非常快,那么最弱的搜索形式——RAG——就很棒。然后是我们谈到的更智能体式的搜索,其中智能体使用类似于人类在计算机上可用的工具。这比 RAG 慢,但能给出更好的答案。更慢的是我所说的神经检索:你有一个非常长上下文的模型,然后你让那个模型为你检索东西。你把所有内容都喂给它,如果所有内容放不进一个模型,就使用多个,然后你让那个模型查看你放在其上下文中的内容并检索相关的东西。这叫做神经检索。这将花费最长的时间;它不一定是最好的,但你可以训练它变得非常好。这就是我今天看到的搜索能力谱系。像这样的东西与 MCP 在涉及与不同知识源交互方面的区别在于,MCP 是无状态的;它只是你与另一个软件交互的一种方式。但在这里你需要做的是整理一个知识索引:从软件中获取数据,将其存储在某处,并使其对智能体可搜索。
It's a really good question. There are a couple of interesting things to unpack. The first thing I would say is that, depending on the problem you want to solve, different types of search fall into basically search. If you want really fast but it doesn't really matter how accurate it is, or just needs to be ballpark accurate but really fast, then the weakest form of search, RAG, is great. Then there's the more agentic search we spoke about, where the agent uses tools similar to those available to humans on a computer. That's slower than RAG but gives you better answers. And even slower is what I would call neural retrieval: you have a really long-context model and you ask that model to retrieve stuff for you. You feed it everything, maybe use multiple of them if everything doesn't fit in one, and then you ask that model to look at what you put in its context and retrieve the relevant stuff. That's called neural retrieval. It's going to take the longest; it's not guaranteed to be the best, but you can train it to be really good. So that's the spectrum of search capabilities I see today. The difference between something like this and MCP as it pertains to interacting with different sources of knowledge is that MCP is stateless; it's just a way for you to interact with another piece of software. But what you need to do here is actually collate an index of knowledge: take data from software, store it somewhere, and make it searchable for an agent.
这最终成了很多人的盲点。有人问为什么没人做这个?原因之一是,从商业模式角度看,现有编码工具存在盲点,它们原本是作为面向广泛消费者群体的 SaaS 产品。但像这样的企业不希望它变成 SaaS,所以你必须重建整个业务,才能在企业资源上部署这些东西。大多数公司一直回避这种深度的索引和集成问题,因为这需要他们改变整个市场策略和商业模式。这也是为什么今天在不同编码提供商之间切换很容易,因为它们没有深度集成。你今天可以试试 Cursor,明天试试 Claude Code,再切换到 Windsurf。它们只是语言模型 API 上的薄薄一层,所以从开发者角度看,切换和尝试不同东西很容易。
That ends up being a bit of a blind spot for a lot of people. There's a question of why hasn't this been done? One of the reasons is that from a business model perspective, it falls into a blind spot for existing coding tools, which were meant to be served as SaaS offerings for a broad consumer base. But an enterprise like this will not want this leaving into SaaS, so you have to rebuild your entire business to deploy the stuff on the enterprise's resources. Most companies have been staying away from this problem of indexing and integrations at this level of depth because it requires them to change their entire go-to-market and business model. That's also why it's really easy to switch around between various coding providers today because they don't integrate deeply. You can try Cursor today, you can try Claude Code tomorrow, you can switch to Windsurf. They're pretty thin skins on the language model API, so from a developer perspective, it's easy to switch around and try different things.
有意思。那么这是否意味着 Asimov 需要能够在某种隔离环境、虚拟私有云或本地部署中工作?
Interesting. So is a consequence of that that Asimov needs to be able to work in a sort of air-gapped context, virtual private clouds, on-prem?
我们确实有 SaaS 产品,因为最终如果你要把某样东西部署到本地,它无论如何都得先以 SaaS 形式开始。但对企业的主要好处是,我们目前不完全是本地部署,而是做 VPC,这对很多已经在 AWS、Azure 或 GCP 上拥有云基础设施的大型企业来说已经足够了。这种可部署为 VPC 的概念对组织极其重要。这绝对是关键因素,否则根本无法开始。
We do have a SaaS offering because ultimately if you flip something into on-prem, it has to start off as SaaS anyways. But the primary benefit to enterprises is we're not going fully on-prem today, but we are doing VPC, and that ends up being sufficient for a lot of big enterprises out there who already have their cloud infrastructure on AWS, Azure, or GCP. This notion of it being deployable as VPC is extremely important to an organization. It's definitely a dealbreaker. You can't even start.
那 Asimov 中的强化学习部分呢?它是如何体现的?你们是世界级的强化学习专家。所以整个想法是它通过每次互动不断学习、变得更敏锐吗?
What about the reinforcement learning part in Asimov? How does it manifest? I mean you guys are world-class RL specialists. So is the whole idea that it keeps learning and getting sharper with every interaction?
在语言模型变得有用之前,你必须处于一个先构建最好的语言模型、再想产品的世界。我们当时就在那个世界。Anthropic 花了好几年构建语言模型,然后凭借 Claude 3 起飞。OpenAI 的 GPT-2 和 GPT-3 都不太能产品化,直到 3.5 和 4 才有效产品化。今天我们处于一个不同的世界,语言模型已经相当不错了。所以我们的策略是构建这种多智能体系统。有些部分我们训练模型,因为我们看到了第三方模型的盲点;其他部分目前仍使用第三方模型。随着时间的推移,我们会逐步抽象所有部分。但作为初创公司,我们需要战略性地选择系统哪些部分要自己攻克,因为初创公司的优势是可以更专注于手头的问题。劣势是你必须更战略性地下注。你没有资源一次性训练所有东西,所以只能一步一步来。长期来看,这将是一个通过强化学习进行端到端学习的系统。短期内,我们应用强化学习来修复在部署现有模型时看到的盲点问题。
Before language models were useful, you had to be in a world where you build the best language model and then figure out the product. That's the world we were in. Anthropic spent a few years building the language model and then took off with Claude 3. OpenAI, GPT-2 was not really productizable, GPT-3 neither. It wasn't until 3.5 and 4 that they were able to productize it effectively. We're in a different world today where language models are pretty good. So our strategy has been to build this kind of multi-agent system. Some parts we train models for, we see blind spots from third-party models. Other places the third-party model stays for today. Over time, we're going to abstract all of it. But we're being strategic about which parts of the system we need to go after as a startup, because that's the benefit of being a startup: you can be more focused on the problem at hand. The downside is you have to be more strategic about the bets you take. You don't have the resources to train everything all at once, so you take it one step at a time. In the long term, this is going to be a system that learns end-to-end via reinforcement learning. In the short term, we're applying reinforcement learning to fix problems that we see as blind spots in the existing set of models when we're deploying them.
我喜欢你说话中那种务实的基调,这真的很有意思,而且我敢说有点耳目一新。我看过很多在 AI 中构建美妙事物的不同方式,但务实的语气真的很有趣,而且并不常见。这就是你刚刚推出的产品。也许带我们回顾一下。我提到过你的背景是世界级的强化学习。这一切是怎么发生的?你创办这一切的历程是怎样的?
I love the pragmatic undertone to everything you're saying, which is really interesting and dare I say somewhat refreshing. I look at different ways of building wonderful things in AI, but the pragmatic tone is really interesting and not that widespread. So that's the product you just launched. Maybe take us back a little bit. I alluded to your background as being world-class reinforcement learning. How did that all come about? What was your journey to starting all of this?
小时候,我对物理非常着迷,想成为一名理论物理学家。
As a kid, I got pretty obsessed with physics and wanted to be a theoretical physicist.
就像普通孩子一样。
As regular kids do.
就像普通孩子一样。嗯,如果你是一个被扔到美国偏远地区的俄罗斯犹太孩子,就像我这样。
As regular kids do. Well, if you're a Russian Jewish kid dropped in the middle of nowhere America, which happened to me.
是啊。那我们聊聊这个。所以你出生在俄罗斯,然后小时候移民到了以色列。
Yeah. So let's go into that. So you're born in Russia. Then immigrated to Israel as a kid.
出生在俄罗斯,小时候移民到以色列,然后童年后半段又移民到了美国。
Born in Russia, immigrated to Israel as a kid, and then immigrated to the States for the second half of my childhood.
这是 AI 顶尖人物的共同经历吗?因为我能想到其他人也是出生在俄罗斯,移民到以色列,然后去了加拿大。
Is that a thing that top people in AI do? Because I can think of other people that were born in Russia and immigrated to Israel and then came to actually Canada.
是的,加拿大是一个主要目的地。我认为苏联是一个技术学术文化极其浓厚的国家,苏联解体后很多年轻科学家去了以色列、德国、美国、加拿大。当我们到达美国时,实际上……
Yeah, Canada is a big one. I think the Soviet Union was an extremely technically academic culture, and a lot of the young scientists left when the Soviet Union fell apart to Israel, Germany, America, Canada. When we arrived in the United States, it was actually...
那你在以色列待了多久?
So how long were you in Israel for?
我在那里待了八年。从 1 岁到 9 岁在那里。我出生在圣彼得堡,所以对在那里生活没有记忆,只是去过。
I was there for eight years. So from 1 to 9 was there. I was born in St. Petersburg, so I have no memories of living there, just visiting.
那你提到的‘美国偏远地区’是哪里?什么城市?
So where was 'nowhere America' that you mentioned? What city?
我在华盛顿州的农村。华盛顿州有一条分界线,西边是茂密的森林,东边是沙漠。你能清楚地看到分界线在哪里。这不是一个渐进的过渡。就是树木,一片茂密的森林从某处开始,然后立即过渡到沙漠。我住在沙漠那边。那里有一个国家实验室。我的父母是化学家,他们在那个实验室找到了工作。那个小镇的文化是,它是曼哈顿计划的一个地点,叫做汉福德基地,是浓缩钚的地方。所以它是洛斯阿拉莫斯的姊妹基地。
I was in rural Washington state. Washington state has a hairline where the west side is a lush forest and the east side is a desert. You can see exactly where that starts. It's not a gradual transition. It's just trees, a dense forest starts somewhere, and then it transitions into immediate desert. I lived on the desert side. There's a national lab there. My parents are chemists and they got jobs in this national lab. The culture of that town was it was one of the sites during the Manhattan Project. It was called the Hanford site, where the plutonium was enriched. So it was the sister site to Los Alamos.
这整个氛围很特别。
This is a whole vibe.
确实是一种氛围。
It is a vibe.
你知道,所有东西都围绕那个事件。保龄球馆叫原子保龄球馆,啤酒厂叫原子啤酒厂。街道名字像铀、汞、钚,河边公园——在哥伦比亚河畔——叫莱斯利·格罗夫斯公园,就是那个负责曼哈顿项目的冷酷将军。所以那是个氛围强烈的城镇。最强烈的是,高中吉祥物——城镇叫里奇兰,吉祥物是里奇兰轰炸机。B-52 轰炸机,篮球场上画着蘑菇云,就是这样。那是个氛围强烈的城镇。
You know, it's uh the everything is themed around that event. The bowling alley is the atomic bowling alley. The brewery is the atomic brewery. The streets are like uranium, mercury, plutonium, you know, like uh the park there by the river, it's on the Columbia River, is called Lesley Groves Park. who is like the ruthless uh you know general in charge of that project. So it's an intense town. The actually the most intense thing is that the high school mascot there um the town's called Richland and the mascot it's a Richland bombers. So the B-52 bombers and there are mushroom clouds on the basketball court and that's it. It's a it's an intense town.
所以,好吧,现在一切都说得通了,你知道,这会是种逃离。
So hence okay all makes sense now that um you know this would be the escape.
是啊,我想我并没有把小镇的历史和物理联系起来,尽管确实如此,但更多是我在学习一门新语言。我有很多空闲时间,我父母有他们的讲义书,他们买了费曼系列讲座,那些书就在那儿,我花时间读了读,就迷上了。
Yeah, I guess I was um well, I didn't even think about the town's history as physics, even though that, you know, definitely is the case, but it was more that I was learning to speak new language. Um had a lot of time on my hands and uh my parents had their um lecture books from uh they bought this uh the Feynman kind of series of lectures and uh that was around and I just spent some time reading it and just got into it.
所以那就是通往物理的道路。
So that was the path to physics.
嗯,最后我读了理论物理博士,然后实际上转投了人工智能。我意识到,首先我看到 AlphaGo 问世,简而言之,我意识到我选了一门有趣的科学,但不是我们这个时代的科学。
um ended up going through and doing a PhD in theoretical physics and actually defected to uh artificial intelligence. I realized that first what I saw I I saw AlphaGo come out and the short of it is I realized that I had picked an interesting science but not the science of our time.
差不多就是这样。你知道,我在物理学中学到的所有那些非常有趣、伟大的东西,基本上都是 100 年前、60 到 100 年前完成的。当你是学生时,你不会太考虑时间线。你会想:“哦,这太酷了。太酷了。我想做这个。”但后来,60 年后,这个领域在很多方面已经固化,不像 AI 那样前沿飞速发展。
That was kind of it. You know the all the things I was learning in physics all these kind of very interesting great things were done basically a 100 years ago 60 to 100 years ago. And when you're studying something as a student, you don't really think too much about the timelines. You're like, "Oh, this is the cool. This is so cool. I want to do this." But then, you know, 60 years later, the field has crystallized in many ways and it's not as dynamic as um and and you know, where the frontier is just moving so fast as AI.
所以当我看到 AlphaGo 问世时,对我来说,这看起来就像是,哦,这实际上是这个时代的科学,我需要去做这个。
And so when I saw AlphaGo come out, yeah, to me it just seemed like, oh, this is actually the science of our time and I needed to do that.
顺便提一下,几个月前——或者可能是去年——看到诺贝尔奖全都向 AI 靠拢,非常有趣。你认为 AI 正在吞噬所有其他科学领域吗?
And as a quick segue on that note, like it was super interesting to see the Nobel prizes a few months ago or maybe that was last year at this point like everything converging towards AI. Do you think AI is eating all those uh other scientific fields?
我认为它是在增强,从某种意义上说,AI 对世界的影响已经变得难以忽视。但没有诺贝尔计算机科学奖,而图灵奖已经颁给 AI 好几年了。所以我想诺贝尔委员会觉得需要以某种方式把它塞进去。显然,AlphaFold 等非常有影响力的 AI 突破带来了一个奖项。
I think it's augmenting like it's um and in some sense I think that part of what happened was that AI's impact in the world had clearly become uh hard to ignore. But there's no Nobel Prize for computer science which is a field that you know so yeah the Turing awards have gone to AI for you know a number of years now and so I think that the Nobel committee I'd imagine felt like it needed to somehow shoehorn this uh and so obviously there were like very impactful uh AI breakthroughs with AlphaFold that um resulted in a prize.
但有趣的是,诺贝尔物理学奖颁给了在物理学中影响并不大的东西,但我仍然接受,因为导致这些系统——比如玻尔兹曼机和霍普菲尔德网络——的突破有物理学的味道。杰弗里·辛顿和霍普菲尔德因此获奖。它们非常物理,看起来和物理学家研究的对象一模一样。
Um but what's interesting is that the physics Nobel Prize was given to something that has not really had that much impact in physics but it is um but I still buy it because it's kind of uh there's a physics smell to to the breakthroughs that led to uh you know these systems like called Boltzmann machines and Hopfield networks that uh that Geoffrey Hinton and Hopfield got the prize for. Um they're very physicsy and they're they they look They look like the same exact objects that physicists study.
这非常有趣。我复述一下,你刚才说的部分意思是,诺贝尔奖颁给 AI,与其说是 AI 吞噬一切的结果,不如说是诺贝尔学院的一种政治行为——他们本该有计算机科学奖,但没有,现在看起来有点傻,因为 AI 是世界上发展最快的领域。因此,他们有点在事后补救。
That that's super interesting. Just just to play it back like part of what you were saying is um the Nobel Prize going to AI is less a function of AI um sort of eating everything but almost like a political thing at the Nobel uh academy or whatever it is that they should have had a prize for um computer science and they don't have it and now it looks kind of silly because AI is the fastest moving field in the world. Therefore, they're uh sort of retrofitting it.
完全正确。我并不是在负面意义上说这个。但作为物理学家,当我再看这些对象——玻尔兹曼机、霍普菲尔德网络,AI 发展中的基础对象,尽管今天已不再使用,但其中的一些概念并没有以任何方式渗透到物理学中。所以我觉得,但它们看起来就像物理学家会研究的对象。
Exactly. And I don't and I wouldn't even say that in uh in like I don't mean it in a negative way. I think uh but yeah, I don't think that as a physicist when I look at uh you know again these objects, Boltzmann machines, Hopfield networks, very fundamental objects in the development of AI, even though they're not really even used today, but like some concepts from them are those things haven't permeated physics in any way whatsoever. So I thought that was But but they just look like objects that a physicist would study.
它们在数学上看起来非常非常相似。
They kind of mathematically look very very similar.
物理味。是的。
Physics E. Yeah.
好的。那么你从物理转向 AI,下一步是什么?
Okay. Um All right. So you evolved from physics to AI and then what was the next step?
中间有个过渡期,我创办了一家小型初创公司,通过了 Y Combinator,主要是做机器学习预测库存管理。我真心觉得必须进入 AI 研究前沿,我认为那里会积累很多科学影响力。所以我最终加入加州大学伯克利分校做博士后,在 Pieter Abbeel 实验室工作,那是强化学习和无监督学习研究的顶级实验室之一。无监督学习的主要产出就是大型语言模型和扩散模型。我当时没意识到,但在某种意义上,那一年不是奇迹年,但很特别。实验室里很多人后来创办了有影响力的公司,或者在大型实验室做非常有影响力的研究。举个例子,我最早合作的是 Aravind Srinivas,他现在经营 Perplexity,还有他的联合创始人 Denis Yarats。我们一起写过论文。Jonathan Ho 是扩散模型的发明者之一,他在那里发明了扩散模型,那篇突破性论文就是他的。还有 Jiaming Song,他也是那篇论文的作者,他们创办了 Ideogram 和 Genmo,两家视频和图像生成领域的初创公司。Deepak Pathak 也在那里,他是 Skilled 公司的创始人,一家顶级机器人公司。
Had an interim where I I started a kind of small startup that went through Y Combinator and uh it was basically doing machine learning kind of prediction for um inventory management. really felt like I had to get on the frontier of AI research. I felt that that was going to be where a lot of scientific impact accumulates and so I uh ended up joining UC Berkeley as a postdoc um where I worked in um this lab called the Pieter Abbeel lab which is uh one of these kind of great labs for uh reinforcement learning and what was called unsupervised learning research which is basically you know um large language models are probably the biggest large language models and diffusion models are the biggest kind of outputs of unsupervised learning. I didn't realize that at the time, but that was kind of a in some sense like a not I wouldn't say like miracle year, but it was a special year in that lab. Um given the people who are in it, a large portion of that lab ended up going to starting impactful companies or um you know being uh scientists like doing very impactful work in large labs. Um to give you a sense first people I worked with was Aravind Srinivas who now runs Perplexity and Denis Yarats his co-founder. We worked together on some papers there. Jonathan Ho who uh is one of the inventors of diffusion models was there and invented diffusion you know like the big paper that made them break through there. Um and this guy named Jiaming Song who uh was on that paper and they started you know Ideogram and uh Genmo which are two startups in the kind of video gen and image gen space. Deepak Pathak who is the founder of a company called Skilled which is one of the kind of premier robotics companies was there.
这是哪一年?大概什么时间段?
What year was this? Like what rough time period?
2020 年。
2020.
是啊。现在回想起来,那真是一群了不起的人,比如 DT Grover,他创办了 Inception,一家做代码扩散模型的公司。是啊,就是那样。我刚想到我看到了这个人的推文。
Yeah. It was a I mean, now when I look back at it, it was a pretty incredible group of of people um like a DT Grover who uh started Inception, which is a company that does diffusion models for coding. Yeah, it was just a like Yeah. Another I was just thinking I saw this guy's tweet.
他叫 Kevin Louu,那天在实验室里。他当时还是本科生,最近在 OpenAI 主导了很多小型模型的工作。所以那真是一群了不起的人。当时一点也不明显。大家做的研究很有趣,但我从没预料到会从中诞生出这么多公司。
His name is Kevin Louu, who was in the lab the other day. He was an undergrad then, and most recently was leading a lot of the work for the mini models at OpenAI. So it was just an incredible group of people at that time. It was not obvious at all then. The research people were doing was very interesting, but I would never have predicted that so many companies would have come out of there.
嗯。
Mhm.
下一步就是 DeepMind。我去了 DeepMind。当时我对多伦多很感兴趣。
And the next step after that was DeepMind. I went to DeepMind. At the time I was really interested in Toronto.
是的。
Yeah.
于是我加入了 Vlad Mnih 的研究组,他被广泛认为是深度强化学习领域的开创者。他是深度网络论文的第一作者,那篇论文让神经网络学会了玩 Atari。他的第一批论文基本上定义了深度强化学习这个领域。这些成果在很长一段时间里都是 DeepMind 的成名之作。我加入这个小组研究我们称之为——我们共同组建的团队叫通用智能体团队——目标就是做研究,弄清楚如何构建通用智能体。我觉得当时比现在模糊得多。我们试图解决的大问题就是人们所说的(现在也仍然这么叫)无监督强化学习,也就是如何训练能够自己设定奖励的强化学习系统。如果没有监督下的奖励,就像你可以给一些奖励,但孩子和动物很多学习都是无监督的——他们与环境互动,没有人告诉他们要做什么。所以我们思考如何用这种方式训练强化学习系统。我实际上认为这个话题在语言模型时代又流行起来了,问题是如何做预训练规模的强化学习,如果没有明确的奖励,如何生成大量合成数据。我觉得这现在又是一个非常有趣的问题。但那是我加入研究的大致议程。结果发生的事情是,语言模型开始奏效了,一旦它们开始工作,这彻底改变了我对哪些问题重要、哪些不重要的看法,因为很多我们认为根本性的问题都以这种蛮力方式解决了。所以你不得不重新校准你想解决哪些问题。
So I joined the group of a researcher named Vlad Mnih, who was largely credited with starting the field of deep reinforcement learning. He was the first author of the deep networks paper, which was the paper that got neural networks to play Atari. His first set of papers actually largely defined deep reinforcement learning as a field. Those were basically DeepMind's claim to fame for a very long time. I joined this group to study the problem of what we called—the team we built together was called the General Agents team—and the whole point was to do research to figure out how to build general agents. I think it was much more opaque then than now. The big problem we were trying to solve is what people called, and still do, unsupervised reinforcement learning, which is really how to train reinforcement learning systems that are capable of assigning their own rewards. If you don't have rewards without supervision, like you can give things some rewards, but kids and animals learn a lot in an unsupervised way—they interact with their environments and no one is telling them to. So we were thinking about how to teach reinforcement learning systems in this way. I actually think the subject is coming back in vogue in the age of language models, with the question of how to do pre-training scale reinforcement learning, how to generate a lot of synthetic data if you don't have explicit rewards. I think it's a really interesting question now again. But that was the general agenda of what I joined to study. The short of what ended up happening was that language models started working, and once they started working, that really changed my entire perspective on what problems matter and what didn't, because a lot of the problems we thought were fundamental were solved in this kind of brute force way. So you had to recalibrate on what problems you want to solve.
所以你不得不重新校准你想解决哪些问题。
And so you kind of had to recalibrate on what are the problems you want to solve.
于是我加入了一个当时只有几十人的小项目,这个项目后来变成了 Gemini 1 和 1.5,然后显然还有 Gemini 2 等等。我和我的联合创始人 Giannis 一起加入,他当时领导强化学习团队,也就是 RLHF 团队。我加入了他的团队,主导了为 Gemini 训练奖励模型以及实现算法等工作。那是一段非常激动人心的时期,一个 10 到 20 人的团队基本上就是那里做 RLHF 工作的所有人。
So I joined a small project at the time that was tens of people, and that project became Gemini 1 and 1.5, and then obviously two and so forth. I joined with my co-founder Giannis, who was leading the reinforcement learning team, the RLHF team. I joined his team and led a lot of the work for training reward models for Gemini and implementing the algorithms and so forth. It was a very exciting time when a group of 10 to 20 people were basically all the people doing the RLHF work there.
那么是什么时候决定离开并创办公司的?当时的想法是什么?
And when was the decision to leave and start a company, and what was the thinking?
我们发布了 Gemini 1 和 1.5,然后意识到语言模型已经跨越了实用性的门槛——它们不再是研究对象,而是会非常有用。那是 2024 年初。我们意识到构建超级智能的要素已经具备。我们觉得一切就绪。还有一个问题需要解决:从 RLHF 到让强化学习真正发挥作用,而这基本上在去年通过推理模型实现了。所以我们觉得那会发生。那么问题来了:超级智能用来做什么?我们觉得不能通过远离产品和客户的研究人员来抽象地回答这个问题。你必须从产品愿景和你要解决的问题的角度去定义它。我们不想构建一个在数学奥林匹克竞赛中表现出色的超级智能。这个时代的强化学习与之前的预训练时代的不同之处在于,预训练让模型在所有方面都变得更好。强化学习则更加参差不齐——它让模型在你希望它擅长的方面变得擅长。所以仅仅因为你让模型擅长竞赛编程,这提高了通用编码能力,但并不意味着你会拥有对软件工程代码有用的智能。
Well, we shipped Gemini 1 and 1.5, and we realized that language models crossed this threshold of utility where they're no longer research objects—they're going to be very useful. This was early 2024. We realized that the ingredients were in place to build a superintelligence. We felt everything was there. There was one more piece to solve: going from RLHF to making reinforcement learning work, and that basically happened right over the last year with reasoning models. So we felt that would happen. Then the question was: superintelligence for what? We felt that you can't answer this question in the abstract by being a researcher that's really far away from product and customers. You had to go in and define what that means from a product vision and what problem you're trying to solve. We're not interested in building a superintelligence that will be superintelligence in mathematical olympiads. The difference between this era of reinforcement learning and the previous era of pre-training is that when you did pre-training, you made the models generally better at everything. Reinforcement learning is much more jagged—it makes them good at what you want them to be good at. So just because you made them good at competitive code, that improves general coding capabilities, but that does not mean you'll have a useful intelligence for software engineering code.
我认为一个例子是,Anthropic 在构建面向产品用户而非基准测试的模型方面做得非常好。当我查看学术基准时,Claude 模型始终比其他模型差,常常差得远,但从用户角度来看,它们始终更好。这一定有原因,我认为原因是当你用强化学习训练大语言模型时,它们会变得参差不齐,即在你希望它们擅长的方面变得擅长。虽然有一些泛化能力,但比人们想象的要弱得多。
And I think an example of that is I think Anthropic has done a really good job of building models that are meant for users of their products rather than benchmarks. When I look at academic benchmarks, the Claude models are consistently worse than whatever else is out there, often not even close, and yet from a user perspective they're consistently better. Something has to explain that, and I think the explanation is that when you train large language models with reinforcement learning, they become jagged in the sense that they become good at what you wanted them to be good at. And there are some generalization capabilities, but they're much weaker than people think.
这与你经常听到的说法有些相反,即泛化总是会胜出。我不知道这是否正确或是某种曲解,但就像那句苦涩的教训。那么你所说的并非完全相反,而是解决方案是泛化和专业化的结合。这样说对吗?
Which is a little counter to the narrative that you hear a lot, which is that generalization is always going to win. And I guess I don't know if that's true or a bastardization thereof, but like the rich sudden bitter lesson. And so what you're saying is sort of not the opposite, but that the solution is a combination of generalization and specialization. Is that fair?
嗯,我认为苦涩的教训实际上并没有提到泛化。苦涩的教训是说,我们应该思考和构建的系统是那些能够很好地随搜索和算力扩展的系统。仅此而已。所以他说的是,如果今天的模型是有限的——这实际上是给研究者的教训,但我认为对产品构建者也是如此。
Well, I think that the bitter lesson actually doesn't say anything about generalization. The bitter lesson says that the systems that we should be thinking of and building are ones that scale well with search and compute. That's kind of it. And so what he's saying is if models are limited today—this was actually basically the lesson for researchers, but I think it's for product builders as well.
嗯。
Mhm.
如果你在构建产品时假设这些模型会停留在当前的智能水平,并为此做了大量临时修补,那么在下一次模型迭代中,你做的许多修补可能就没那么必要了。
If you're building your product with the assumption that these models are going to stay at their current intelligence level and you make a bunch of hacks around your product to overcome those things, then in the next iteration of models, a lot of the hacks that you put in place will probably be less necessary.
这正是科研人员看到模型局限性后,用临时修补让它们在特定基准上表现更好的做法。所以同样的教训也适用,但更重要的教训是构建那些善于吸收算力并能随搜索良好扩展的系统。
And that was scientific researchers seeing these limitations of models and plugging them with temporary hacks that would make them better at certain benchmarks. So, it's kind of the same lesson translates there, but the lesson is more to build systems that are good at soaking up compute and scale well with search.
是的,我们会把那个放在节目笔记里。我觉得那可能是 AI 历史上被引用错误最多的博客文章之一。
Yeah, we'll put that in the show notes. I think that's probably one of the most often misquoted blog posts in the history of AI.
是的。我认为关于泛化,真正发生的是,如果你的训练分布涵盖一切,那么测试分布自然落在训练分布内,你就实现了泛化。也许有一个观点没多少人认同,但我确实认为会有一个通用超级智能,但它不会是由一个实验室构建的,而是由多个智能体组成的集合——所有智能的集合将构成一个通用超级智能。因为如果你把预训练看作土壤或基质,你可以从中培育出各种类别的超级智能——医疗超级智能、组织超级智能、科学家和数学家的超级智能——这些都是不同类型的超级智能。其中一些可以合并到一个模型中。但我认为,如果你想象不同的超级智能植物从基质中生长出来,前沿实验室能够捕获其中一些,但也会有新的前沿实验室建立起来,培育其他植物,而这个花园的集合将成为一个通用超级智能,而不是一家公司去培育所有植物。从研究和算力的角度来看,这是可能的,但从产品的角度来看,如果你想构建组织超级智能,你必须与所有这些客户集成,必须有解决方案工程师支持他们,必须有销售人员支持他们。在那种研究者只是训练模型并希望模型成为超级智能的美好图景中,我认为世界不会那样发展,因为如果它不与产品和部署相结合,就毫无意义。
Yeah. Well, I think what's happening with generalization is that if your training distribution is everything, then your test distribution just falls in your training distribution and you have generalization. Maybe one point of view that I don't think many people share, but I do think there will be a general superintelligence, but I think that it won't be one lab that has built it but it'll be the plurality—the collection of all intelligences will be a general superintelligence. Because if you think of pre-training as the soil or the substrate from which you can now grow superintelligence in various categories—medical superintelligence, organizational superintelligence, superintelligence for scientists and math—these are all different types of superintelligence. Some of them you can merge into one model. But I think that there will be, if you think of different superintelligence plants growing from the substrate. Yes, like a frontier lab will be able to capture some of them, but I think that there will be new frontier labs built that grow other plants and that the collection of this garden is going to be a general superintelligence rather than one company going in and growing all the plants. I think from a research and compute perspective that's possible, but from a product perspective, if you want to build organizational superintelligence, you have to go and integrate with all these customers and you have to have solution engineers that support them and you have to have sales people that support them. In this rosy picture of a researcher who just trains models and hopes the model is a superintelligence, I just don't think that the world will play out that way because it's meaningless if it's not coupled to a product and its deployment.
深入探讨一下今天的产品。像自主编码智能体这样的东西现实情况如何?当前最先进水平与未来期望相比有多好?比如在 SWE-bench 上我们是在十几分吗?对于某些任务我们更高吗?目前在基准测试上整体情况如何?
Double clicking on the product today. What is the reality of something like an autonomous coding agent? How good is the state-of-the-art right now versus what hopefully it will be in the future? Like are we in the teens in terms of SWE-bench? Are we higher than that for certain tasks? Where does it all land currently on the benchmarks?
快速补充一点:我认为 SWE-bench 要么——我是说,它正在接近饱和,有趣的是,尽管你在 SWE-bench 上看到 70%之类的数字,那些编码模型很好,但它们并没有解决 70%的工程任务。所以基准测试和现实问题之间存在错位,当你的基准测试不是客户实际使用的东西时,这种情况总是会发生。但就进展而言,我一直——如果说有一件事我对其进展时间表相当激进,我会说事情的发展甚至比我预期的还要快。我认为我们已经从自动补全引擎发展到半自主的东西,再到现在对于初级任务它们可以自主完成——这非常不可思议。在某种意义上,我们可能达到了 L4 级别的初级工程师自主水平,这非常不可思议。
Just a quick side comment: I think that SWE-bench is either—I mean, it's getting to saturation, and it's funny that even though you have numbers like 70% or something on SWE-bench, those coding models are good but they're not solving 70% of engineering tasks. So there's a sort of benchmark-to-real-world problem misalignment, which is always going to be the case when your benchmark is not the actual thing that customers are using it for. But in terms of progress, I've been—if there's one thing I had fairly aggressive timelines on progress in my mind, and I would say things have moved faster even than I expected. I think that we've gone from autocomplete engines to things that are semi-autonomous to now things that for junior tasks they can just do them autonomously—it's pretty incredible. There is in some sense we are probably at an L4 kind of junior junior engineer level of autonomy, which is pretty incredible.
自主意味着 100%成功。不需要代码审查。
And autonomy means 100% success. No need for code review.
你仍然需要像对待 L4 工程师那样进行代码审查,但可靠性达到了这样的程度:代码审查中你会挑出一些小毛病,就像对待普通工程师一样,但他们给你的东西首先已经足够有用,值得你去审查。而不是说“这简直是垃圾,我为什么要花时间审查这段代码?”所以我认为对于初级代码,比如小的 UI 改动和零碎的小事,这类任务很多,所以非常有用。作为一个领域,我们可能达到了 L4 水平。我认为在未来几年内,代码生成能力将继续提升,我们将拥有相当智能的东西。我不知道具体会达到什么水平,因为我认为当你给它们所有规格、所有需求、确切需要构建什么时,它们会非常擅长生成代码。但高级工程师及以上的工作首先是弄清楚需求。他们大约 70%的时间花在那上面——设计东西、弄清楚需求、提前规划,比如如果我构建这个,它会与那些东西冲突吗,从其他团队成员那里获取信息。这才是高级工程师及以上人员真正做的事情,而实现部分通常是直接的部分——你只花大约 20%的时间实际编写代码和实现东西。所以按照目前的进展,我认为一旦你给它们非常具体明确的任务,我们在实现层面将拥有非常能干的智能体。但如果你不解决这种上下文收集或构建上下文核心,我认为它们整体上不会达到高级工程师的水平。
You still need to do code review in the way that you would do with an L4 engineer, but a level of reliability that the code review will—there'll be some nits that you pick off like you do with a normal engineer, but that the thing they gave you is useful enough for you to review it in the first place. As opposed to this is just garbage and why am I spending time reviewing this code? So I think for junior code like little UI changes and small things here and there, of which there are a lot, so it's very useful. We're probably at L4 as a field. I think that over the next couple of years, the code generation ability will continue improving and we'll have things that are quite intelligent. I don't know where exactly I'd place them, because I think they'll be really good at generating code when you give them all the specs, every requirement, exactly what needs to be built. But the whole job of a staff level and above engineer is figuring out the requirements in the first place. So that's kind of 70% of their time is spent on that—designing stuff, figuring out the requirements, planning in advance, like if I build this will it conflict with these things, soliciting information from other teammates. That's really what a staff engineer and above does, and then the implementation part is usually the straightforward part—you spend like 20% of your time actually writing code and implementing things. And so the way things are progressing, I think we'll have very capable agents at the implementation level once you give them something to do that's really concrete and specific. But if you don't solve this kind of contextual gathering or build out the contextual core, I don't think they'll be at that staff level as a whole.
但从你的角度来看,有记忆的 L9,而不是我们之前讨论的有遗忘症的 L9,感觉是一个近期可以解决的问题?我们正在朝着那个方向前进。
But the L9 with memory, not the L9 with amnesia that we were talking about earlier, feels like a tractable near-term problem to solve from your perspective? Like we're well on our way there.
是的。
Yes.
所以当你提到,回到对话开头,你说你从一个编码智能体开始,是不是你描述的很多原则可以横向推广到整个公司?所以机构记忆,我觉得这是一个迷人的概念,你可以插入它,它可以是编码的机构记忆,但接下来它变成营销、产品、人力资源的机构记忆。
So when you said, going back to the beginning of the conversation, that you're starting with a coding agent, is the idea that a lot of those principles that you described can be horizontalized across the company? So the institutional memory, which I find a fascinating concept, that you can plug in that could be the coding institutional memory, but next it becomes the marketing, the product, the HR institutional memory.
是的,没错。从构建的角度来看,这非常相似。你现在在企业客户的本地或 VPC 中,拥有围绕他们代码的集中化知识,并且你已经为他们集中了其他工具的知识,这些工具与下一步紧密相关。对吧?如果你集中来自 Jira 的知识,那既与工程相关,也与产品管理相关。所以我认为,到那时,这仅仅是根据企业需求引入其他工具集成,然后使用户能够按需代表他们行动。所以不仅仅是能够提问,而是让实际的智能体去为他们做事。所以我认为这很快就会到来,对吧?在编码领域。我认为一旦你有了上下文核心,你就可以将其集成到现有的编码产品中。你可以这样想,如果我们把其他编码产品看作是有健忘症的自主软件工程师,你可以为他们填补这个空白。显然你也可以自己构建一个,但我认为这最终变成了一种客户选择的概念,对吧?你希望整体解决方案对客户最好,而作为一家公司,你专注于你认为的、实现超级智能的基本构建块。
Yeah, that's right. It's kind of at that point from a buildout perspective, it's very similar. You know, right, you're now on-prem or in the VPC of an enterprise customer, you have a centralized kind of knowledge around their code and you've already centralized knowledge around other tools for them as well that are immediately adjacent to the next thing. Right? If you're centralizing knowledge from Jira, that's both engineer and product management adjacent, right? So I think at that point it just becomes bringing other tool integrations in based on where you're seeing pull from the enterprise, and then enabling the ability to act on the user's behalf when they want to. So instead of just being able to ask questions, enabling the actual agent to go and do stuff for them. So I think that's going to come sooner than later, right? It's in the coding space. I think once you have the contextual core, you can integrate it into existing coding products. You sort of like, if we think about the other coding products as being these autonomous software engineers with amnesia, you can fill that gap for them. Obviously you can build out one of your own as well, but I think it ends up being kind of a notion of customer choice, right? You want the overall solution to be best for the customer, and you're as a company focused on what you believe is kind of the fundamental building block of enabling superintelligence.
也许放大到结尾,我很好奇关于今天构建 AI 初创公司现实的一些想法。我想从人才角度开始,正如我们几分钟前所说,在这个奇怪的时刻,巨额资金被提供给人才,让他们从一家公司跳到另一家公司。在这种环境下,当你显然非常令人印象深刻但仍然是一家小型初创公司时,如何招募和留住人才?
Maybe zooming out to close, I'm curious on a few thoughts about the reality of building an AI startup today. I guess from a talent perspective to start with, as we were saying a few minutes earlier, like in this weird moment where ridiculous amounts of money are being offered to talent to move from one company to the other. How does one sort of recruit and keep talent in this environment when you are, you know, very impressive obviously but still a small startup?
我认为研究团队的整个组成基本上都在大实验室赚了很多钱。显然包括 Giannis 和我自己。我想记住的是,很多人进入这个领域是因为他们内心是科学家或建设者。所以有财务因素,你当然需要支付足够的报酬,慷慨到让他们不会太在意。但人们真正关心的是发现下一个前沿和下一个突破。在 AI 实验室最激动人心的时刻是当它还不是明显的前沿实验室时。我认为 DeepMind 最激动人心的时刻是构建深度网络和构建 AlphaGo。那些是 DeepMind 的黄金时代。同样,OpenAI 是构建 GPT 系列模型的一、二、三阶段。而对于 Anthropic,我认为是在 Claude 一和二的阶段,现在他们正在收获那些突破性工作的成果。我们倾向于吸引那些有内在驱动力、想成为下一个故事一部分的人,因为他们已经在大实验室,或者他们可以加入大实验室,而且他们总是能够做到。我的意思是,我们会看到 ASI 出现时的情况,但这并不是那么稀缺的机会。当你实际看看有哪些初创公司拥有清晰度、团队和启动新前沿实验室的潜力时,并没有那么多。所以实际上,大实验室的职位比有机会做到这一点的初创公司要多得多。所以我认为人们最终在某种程度上自我选择。我们经常从 OpenAI、Anthropic、Meta、DeepMind 那里赢得候选人。显然,他们在这家公司获得更多的股权,作为所有权百分比。如果他们粗略计算一下,如果他们在这一阶段加入 Anthropic 并获得那个所有权百分比,现在会值多少钱,那绝对是世代相传的。所以我认为,如果你没有一个好的初始团队核心,那会变得非常困难。但如果你有一个非常强大的初始团队,并且人们看到突破的潜力,那么你就成为一个稀缺的选择,因为没有多少地方可以做到这一点。今年晚些时候,我们将推出我认为没有人认为初创公司能做到的事情。我认为我们将在研究方面取得一些突破,每个人都认为你需要一个拥有 10 万 GPU 的巨型实验室才能做到,我认为这将非常有趣和令人惊讶。
I think the entire makeup of the research team was earning a lot of money at big labs basically. Obviously Giannis and myself included. The thing I guess to remember is that a lot of people get into this field because they're scientists at heart or they're kind of builders at heart. So there is the financial element, you certainly need to pay enough to be generous enough where that's not really on top of minds. But people really care about discovering the next frontier and the next breakthrough. The most exciting time to be in an AI lab is when it's before it's obviously the frontier lab. I think the most exciting time at DeepMind was building deep networks and building AlphaGo. Those were the times I think were sort of DeepMind's golden days. Similarly with OpenAI, it was building out the GPT series of models at the one, two, three stage. And I think for Anthropic, it was really in the Claude one and two stage, where now they're kind of reaping the benefits of that breakthrough work. We tend to attract people who have that kind of internal drive of wanting to be part of that next story because they're already at a big lab or they could join a big lab and they'll always be able to. I mean, we'll see when ASI comes around, but it's like that's not really that scarce of an opportunity. When you actually look at which startups are out there that have the clarity and the team and potential of starting a new frontier lab, there are not that many. So there are actually a lot more spots at the big labs than there are at startups that have a shot at this. So I think that people end up in a sense self-selecting. We've won over candidates from OpenAI, Anthropic, Meta, DeepMind regularly. Obviously, they get a lot more equity in this company as a percentage of its ownership. And if they do back-of-the-napkin math of if they had joined Anthropic at this stage and got that percentage ownership, what it would be worth now, it would be absolutely generational. So that tends to be, I think if you don't have a good kernel of the initial team, it becomes very hard. But if you have a very strong initial team, and people see that potential for breakthroughs, then you become a scarce option because there are not that many places where you can do this. And later this year we'll be shipping things that I don't think anyone ever thought a startup could do. I think we're going to be tripping some things on the research side that I think everyone thinks you need to be a giant lab with 100,000 GPUs to do, and I think it'll be quite interesting and surprising.
在产品方面,再次回到产品与研究的讨论,你们是超级深入的博士级世界级 AI 研究员,但在一个你确实想构建产品的背景下,核心团队是否包括引入带来产品的人,或者你们是如何考虑的?
On the product front, again to the discussion about product versus research, you guys are super deep PhD-type world-class AI researchers, but in a context precisely where you want to build product, was it part of the core team to bring in people that would bring product, or how did you think about it?
是的,我们已经招聘了。我想公司大致是半产品半研究。所以我们招聘了一个研究团队,一个产品团队,然后公司的大部分组成可能是三分之二的人来自一些大实验室的研究背景。这些人中,很多人的角色介于研究和产品之间。
Yeah, we've hired out. I guess the company is kind of half product, half research. So we've hired out a research team, we've hired out a product team, and then the majority of the makeup of the company is probably two-thirds of it is people who have research backgrounds at some of the big labs. Of those people, a bunch of them are in this role that is between research and product.
例如,智能体设计研究就介于研究与产品之间,或者评估——你评估模型擅长什么?对我们来说,这通常只是研究的一部分,但它是一个跨职能的端到端的事情。你生成的数据,你为训练模型生成的合成数据,也贯穿所有这些方面。所以从某种意义上说,我认为它吸引了像 Giannis 和我这样的人,我们进入这个领域,只是不想再优化另一个学术基准。我们只想解决实际问题,进行真实的评估。对于这些人来说,这里最终是一个非常好的地方。我认为对于那些更愿意待在一个已知实体中,真正专注于某件具体工作的人,比如训练模型,因为模型太大了,当你进入时,你只会得到你负责的那一小块。
For example, the design of agent design research is very much between research and product, or evaluations—what are you evaluating your models to be good at? That's typically something that's just on research for us, but it's kind of a cross-functional end-to-end thing. The data you're generating, the synthetic data you're generating to train your models, also cuts across all those things. So in a sense, I think it attracts people similar to Giannis and myself who came into this and just wanted to not maximize another academic benchmark. We just wanted to solve real problems and have real evaluations. For those people, this ends up being a really good place. I think for people who would much rather sit in a known entity and really focus on some specific piece of work, like training the model because they're so big that when you enter them, you're given a sliver that you own.
但这仍然非常有趣,因为你有点像成为了一个工匠。所以我认为这实际上是一个非常重要的技能。但如果人们处于那个阶段,那么我认为大型实验室绝对是更好的选择。你已经筹集了一大笔资金——我记得是 1.25 亿或 1.3 亿美元,不管媒体报道的数字是多少。对于像你这样的公司,资本是否同样重要?考虑到当今 AI 的模式,某些类型的公司确实存在一场广为人知的资本竞赛。你这一代创业公司和你这种类型的创业公司需要那么多资本吗?
But that's really interesting nonetheless, because you kind of become a craft person. So I think that's actually a very important skill to pick up. But if that's where people are in their lives, then I think the big lab is definitely a better option for them. And you've raised a bunch already—I think like 125, 130 million, whatever the number in the press. Does capital matter as much for a company like yours? Thinking about modes in AI today, there certainly is a well-publicized capital race for certain types of companies. Does your generation of startup and your type of startup require as much?
资本非常重要。我认为区别在于,你不能以比前沿实验室少 100 倍的资本运营,但当你真正专注时,你可以以少 10 倍——一个数量级——的资本运营。所以我认为资本非常重要,它实际上取决于你何时准备扩大 GPU 数量。基本上就是这样。显然还有人员、数据等等,但这些公司的主要成本是 GPU 支出。所以你在准备进入下一阶段时筹集相应的资本。我认为资本非常重要,但你可以比传统的前沿实验室高效得多。
Capital matters a lot. I think the difference is that you can't operate at 100x less capital than a frontier lab, but you can operate at say 10x—an order of magnitude less capital—when you're really focused. So I think capital matters a lot, and it really is determined by when you are ready to scale up your GPU count. That's basically it. There's obviously headcount and data and so forth, but the primary cost for any of these companies is their GPU expenditures. So you raise capital commensurate to when you're ready to scale to the next stage. I think capital matters a lot, but you can be a lot more efficient than traditionally frontier labs have been.
好的。我们涵盖了很多内容,从过去到超级智能,到阿西莫夫,到你的背景,到你正在构建的东西以及你下一步要构建的东西。我对你提到的未来几个月在研究方面即将发布的公告感到兴奋。非常感谢你来到这里。这太棒了。谢谢。
All right. Well, we covered a bunch from the past to superintelligence to Asimov to your background to what you're building and what you're building next. I'm excited for this announcement as you alluded to that's coming on the research side in the next few months. Thank you so much for being here. This was terrific. Thank you.
是的,谢谢你 Matt。这很有趣。
Yeah, thank you Matt. This was fun.
嗨,我是 Matt Turk。感谢收听本期 MAD 播客。如果你喜欢这期节目,如果你还没有订阅,我们希望你能考虑订阅,或者在你观看或收听本期节目的任何平台上留下好评或评论。这真的有助于我们发展播客并邀请到优秀的嘉宾。谢谢,下期节目再见。
Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks, and see you at the next episode.