The Third Generation of AI Labs: Automating AI Research
打开互动全文版(中英对照 + 朗读 + 问答)→Jerry Tworek 讨论 Core Automation 的使命,即构建全球最自动化的 AI 实验室,利用 AI 代理加速研究和实验。
Jerry Tworek discusses Core Automation's mission to build the world's most automated AI lab, leveraging AI agents to accelerate research and experimentation.
好的,我们回来了。今天我们请到了一位非常特别的嘉宾,Jerry Tworek,他是 Core Automation 的联合创始人兼 CEO,这是一家致力于打造全球最自动化 AI 实验室的 AI 研究公司。此前,他在 OpenAI 工作了七年,并成为研究副总裁。他构建了机器人强化学习、Codex、最初的 Codex、ChatGPT、GPT-4、o1、o3 推理模型等。此前,他还从事过量化金融,并在华沙大学学习数学。Jerry,欢迎来到 MTS。
All right, we are back. We are live with a very special guest, Jerry Tworek, who is co-founder and CEO of Core Automation, which is an AI research company building the world's most automated AI lab. Previously, he spent seven years at OpenAI. He became the VP of research. He built RL for robots, codecs, the OG codex, ChatGPT, GPT-4, o1, o3, reasoning models, and previously he was in quantitative finance and studied math at the University of Warsaw. So Jerry, welcome to MTS.
谢谢邀请,今天很高兴能来到这里。
Thank you for the invitation and very happy to be here with you today.
我们有一个非常想知道的问题。Core Automation 现在在做什么?
So we have this burning question we're dying to know. What's going on in Core Automation?
我们正在尝试建立一个新的 AI 实验室,而现在是 AI 研究的新时代。我认为我们正在看到第三代 AI 实验室,它们试图以自动化优先,认识到我们现在可以利用 AI 模型来进行 AI 研究,并思考如何以最佳方式做到这一点。我认为我们有机会在不同于以往所有公司的基础上建立一家公司,因为我们是从不同的技术基础出发的。我们已经拥有了上一代公司试图构建的所有东西,然后我们试图找出如何以最佳方式利用它们来建设公司。
We are trying to build a new AI lab, and it's a new era of AI research right now. I think we are seeing a third generation of AI labs that are trying to be automation first, seeing okay, we are in the world right now where we can use AI models for doing AI research, and how do we do that in the best way. I think we have a chance of building a company on slightly different foundations than all the companies built in the past, because we are starting from a different technological foundation. We already have all those things that the previous generation of companies were trying to build, and then we are trying to figure out how do we use them for company building purposes in the best way.
你提到了第三代 AI 实验室。那第一代和第二代是什么?
So you said the third generation of AI labs. What were the first and second?
我是这么看的,当然这有些简化,可能 DeepMind 才是真正的第一代,他们是非常典型的开创者。但在某种程度上,我认为 DeepMind 和 OpenAI 属于同一类 AI 实验室,因为它们都是在我们知道大规模 AI 可行之前就成立的。它们是非常纯粹的研究驱动型公司。我们想构建 AGI,我们不太清楚怎么做,但我们会研究如何制造 AI 的想法。然后我认为出现了第二代 AI 实验室,主要是在 GPT-3 之后,当 Scaling 被证明时。我们确实有办法投入大量算力来制造智能,很多实验室试图实现 Scaling。当时成立了很多实验室,虽然比现在少,但也不少,其中 Anthropic 是注定最成功的一个,最终也成为当今最成功的公司之一。而现在,在推理和智能体取得成功之后,我们看到新一轮初创公司爆发,试图弄清楚如何构建智能体原生、AI 原生的 AI 公司,那会是什么样子?
How I think about it, and it's obviously compressing certain things, probably DeepMind was genuinely the first one, and they were very much the OG. But in some way, I consider DeepMind and OpenAI as the same kind of AI lab, because they were AI labs started before we knew large-scale AI was possible. They were companies very much purely research-driven. We want to build AGI. We have very little idea how, but we will be researching ideas how to make AIs. Then I think there was a second wave of AI labs created, largely I consider them after GPT-3, when scaling was proven. We actually have ways how to spend a lot of compute in making intelligence, and a lot of labs were trying to achieve scaling. A lot of labs were created at that time, still less than today, but many were, and Anthropic was the one that was meant to be the most successful out of those and ended up being one of the most successful companies today. And what we see right now, after the success of reasoning and after success of agents, we have another explosion of new startups trying to figure out how do we build agent-native, AI-native AI companies, and how does that look like?
那么你们试图自动化整个技术栈的哪一部分?是自动化研究人员吗?你们是用现有的编码工具来自动化 AI 研究人员,还是也在尝试自动化编码本身,比如自我改进的编码智能体?
So what part of the entire stack are you guys trying to automate? Is it automated researchers? Are you using existing coding tools to automate AI researchers, or are you also trying to automate the coding part itself, like self-improving coding agents?
几乎技术栈的每一部分,问题在于我们在哪里拥有最大的 Alpha,哪里是我们改进事物的最大杠杆点。但我是这么想的:实验室做什么,研究公司做什么,它们产生洞察,产生知识。你需要弄清楚如何训练更好的模型。你有一个数据生成过程,投入算力,产出洞察。问题几乎是如何优化这个过程,以便从算力中获得最大价值来产生新洞察,以及如何在单位时间内运行尽可能多的实验,从而以某种方式比其它公司更快地产生洞察。
Almost every part of the stack, and the question is where we have the most alpha, where we have the most leverage to improve things. But how I'm thinking about it: what labs do, what research companies do, they generate insight, they generate knowledge. You need to figure out how to train better models. And you have this data generating process where you put compute in and you get insights out. And the question is almost how do we optimize that process so that we get the most value out of our compute for generating new insights, and how do we run as many experiments as possible per unit of time to have in some way a faster rate of insight generation than other companies.
那么你说的洞察生成,是指你运行一堆智能体来产生想法,然后你根据这些想法进行 AI 研究吗?
So by insight generation, you mean you are running a bunch of agents in order to generate ideas that you then act upon for AI research?
在很多方面,我们的情况略有不同。我们仍然主要使用人类研究人员来产生洞察,并试图确定我们应该朝哪个方向前进,但智能体旨在尽可能简化执行,以便我们能够尽快尝试这些想法并生成关于这些想法好坏的数据。所以如果我们能把实验时间从一个月缩短到一天,那就是 30 倍的加速,这些都是可能的。我认为智能体目前并不擅长开放式探索和构建真正精确的假设,但它们非常擅长在指令相当精确的情况下执行任务。
In many ways, what we have is slightly reversed. We still use human researchers mostly for generating insights and for trying to determine where we should be going, but the agents are meant to streamline the execution as much as possible so that we can try those ideas and generate the data on those ideas, whether they are good, as quickly as possible. So if we can take down the time of experimentation from a month to a day, that's like 30 times speed up, and all those things are possible. I don't think agents are very good right now at open-ended exploration and constructing really good and precise hypotheses, but they are very good at doing what they are told if it is reasonably precisely stated.
是的,我们最近看到一条很好的推文,关于寻求计算高效的人类、Token 和算力之间的帕累托最优平衡。那是 Gavin Baker 发的。
Yeah, we saw a good tweet recently about like seeking the Pareto optimal balance of computationally efficient humans and tokens and compute. It was Gavin Baker.
是的,那条推文非常好。
Yes, it was so good.
那么你们是怎么考虑这个问题的?对你来说,什么造就了一个计算高效的人?
So how are you guys thinking about that? Like what to you makes a computationally effective human?
这是个好问题。在某种程度上,对人类来说,计算和时间是紧密相关的。我们可能在单位时间内只能使用固定的计算量。所以问题主要是我们如何花费时间,如何花费精力。但肯定有方法可以更好地利用这些,主要是专注于我们能做的高杠杆努力。
It's a good question. In some way, for humans, computation and time are very related. We probably have just a fixed amount of computation we can use per unit of time. So the question is mostly how do we spend our time, how do we spend our energy. But there are definitely ways how we can spend those things better, mostly focusing on what are the highest leverage endeavors we can do.
对,所以我想在你的愿景中,我们经常听说研究方向是最难优化的部分。即使在完全自动化的实验室中,你仍然需要人类的洞察力来指定某个研究方向。那么你们是怎么考虑这个问题的,你预测人类还需要多久才能持续参与其中?
Right, so I guess in your vision, we've heard a lot that research direction is kind of the hardest component to optimize. Like you still need human discernment to specify a certain research direction even within the context of a fully automated lab. So how are you guys thinking about that, and what's your prediction for how long a human will have to be pretty consistently in the loop?
是的,这是个好问题,我们还有一些时间。我的粗略估计是,至少还有两年时间,人类研究人员在 AI 研究发现过程中仍将扮演重要角色。
Yeah, it's a good question, and we still have some time. My rough estimate is at least two years when human researchers are a meaningful part of the AI research discovery process.
我觉得现在经常听到 AI 研究者们反复说的一句话是:哦,我们只剩最后几天的工作时间了,趁还能干的时候赶紧干,然后我们就都休息。
I think like the what what you often hear AI researchers be repeating right now in the in the field. Oh, we have a last few days of work, so let's work while I still can and then and then we'll all take break.
这个行业变化非常快,在很多方面都非常消耗人,在这个领域工作这么多年非常非常累。但如果我们认为这是我们职业生涯和工作中最重要的时刻,那可能还是值得的。
The industry is moving very quickly and it is in many ways very very exhausting, very very tiring to be working at this at this space for for so many years. But it's it's probably worth it if we think those are the most the most important times of our of our of our career and of our of our work.
我认为我们必须让智能体对研究领域和研究方向有更高层次的理解。一般来说,任何用过模型做研究的人都会发现,智能体通常只能产生质量很低的见解。它们可能很有创造力,但那种创造力是以一种非常低质量的方式呈现的——生成的想法即使多样,通常也不太好。目前还没有一个模型能代表那些花了 10 年研究 AI 技术的人。
I think we'll have to have agents have much higher level understanding of a of a research field and research directions and in generally we know and we think and anyone who has used models for doing research has seen that agent generally do like generate pretty low quality insights and have like they can they are high creativity but they are high creativity and in a in a kind of like very very low quality way in which in which the the ideas are generated if they are diverse they are generally not not very good and like there is no such thing yet for a model to to represent humans who has spent like 10 year researching AI technologies
谁对我们真正处理的对象有这种深度的理解?
who has this this kind of like level of depth of understanding of the of the objects we are really working with here
我想特别感谢我们的赞助商 Lovable。你知道那些你一直想做的应用、内部工具、副业项目或产品,本来要花掉你一个周末,但用 Lovable,你今天就能交付软件。它包括一个可编辑的代码库,你可以检查和修改,还有双向 GitHub 同步。Lovable 还处理后端和基础设施,包括托管 Postgres、对象存储、托管、支付等等,所以你不需要自己全部搞定。通过 Lovable 的 MCP 服务器,你可以直接从你已经在用的智能体创建和部署项目。用 Lovable 把想法变成人们喜爱的软件。lovable.dev。现在,回到正题。
I want to shout out our sponsor lovable you know the app you've been meaning to build the internal tool side project or product you'd otherwise lose a weekend to with lovable you get software you can ship today. That includes an editable codebase you can inspect and change, plus two-way GitHub sync. Lovable also handles the backend and infrastructure, including manage Postgress, O storage, hosting, payments, and more, so you don't have to wire it all up yourself. And through Lovable's MCP server, you can create and deploy your project directly from the agents you already use. Turn ideas into software people love with Lovable. lovable.dev. Now, on to the episode.
所以你的预测是,两年内,整个 AI 研究循环,从端到端,将完全自动化,人类将无法像在国际象棋中那样发挥有意义的作用。
So your prediction is that in two years the entire uh AI research loop from you know fully end to end will be completely automated and humans will not be able to play a meaningful part in the same way that humans can play meaningful part in chess.
是的。
Yes.
所以这即使按照短时间线的标准,也是相当激进的时间线。
So like this this is uh these are like pretty aggressively short timelines even by like short timeline standards.
嗯,比如两年前,2024 年 8 月,我们有什么?那大概是在 01 发布前两周或三周。所以最好的模型是 GPT-4o、Claude 3.5 Sonnet。它们还不能真正做数学,还不能真正写代码。嗯,我想说我们在这些领域取得了相当大的进展。但在其他领域,比如创意写作,或者一般的写作,我们没有取得太多进展。嗯,比如我的聊天查询,并不比两年前好多少。
Um like two years ago like August 2024 what did we have? This is like right before like maybe two weeks before the announcement of 01 two or three weeks before. Um so the best models were like GPT40 claude 3.5 sonnet. They still couldn't really do math. They still couldn't really code. Um like I would say we've made a pretty substantial amount of progress in those domains. We haven't made a ton of progress in other domains like creative writing um or just like writing in general um or like my you know my chat queries are not dramatically better than they were two years ago.
而且这也是一个模式崩溃的一年。我觉得所有编码智能体的 UI 都大同小异。语言也趋同于同一种风格。
Like it's a mode collapse year as well. I feel like all of the UI that any coding agent is largely the same. The language has also converged on kind of the same style.
是的。
Yeah.
那么我们如何从现在的状态,在两年内达到所有 AI 研究的完全自动化,让人类完全退化到无足轻重的地步?
So how do we get from like here to complete automation of all AI research uh to the point where humans are totally vestigial in two years?
是的,我认为,你知道,显然有一条直接扩展当前方法的路径,也有在两年内做出新发现的路径。但我想说的是,你刚才说思考过去两年我们在哪些方面取得了最大进展,以及我们是否在创意写作、数学和编程领域有所改进。很多这些问题几乎就像你总是必须跟随激励。现在,实验室显然有巨大的激励去专注于编程,专注于使用 AI 模型来自动化知识工作。而在创意写作方面,让模型变得非常好所需的努力,实验室能否证明这些成本是合理的,还不太确定。有时人们会想,某些领域进展更快,某些领域进展更慢,但很多时候,这并非受研究限制,而是受经济激励限制。训练模型非常昂贵,实验室无论外界是否看到,都在残酷地为了生存进行经济优化,因为如果你不是算力足迹最高的实验室,那么你就会……
Yeah, I I think like you know there there are obviously there's a path of like direct of of like just scaling the current methods and there's also a path of what new discoveries do we make in those in those two years. But I I would say like you you said thinking for the last two years where did we make the most progress and whether we whether where we are improving on um on creative writing or on or on uh or on math and programming domains. And a lot of those questions are almost like you always have to follow the incentives. And right now there are huge incentives for the labs obviously to focus on programming and focus on using AI models for like automating knowledge work and and and kind of like in in terms of creative writing, it's just the the effort that it needs to make models very good at it. It's not not super sure that the labs can justify those costs. Sometimes sometimes people are thinking oh like like you know certain domains are moving faster and certain domains are moving are moving slower but very often it is like not bounded by research but more bounded by economic incentives and training mods is very expensive and the labs are like whether whether it it's is apparent to the outside or not very brutally economically optimizing for survival because if you are not the lab with the highest compute footprint then then then you will I
是的。这引出了……
Yeah. This raises
是的。如果你不是算力足迹最高的实验室,那你就会死。但那么新实验室的利基在哪里?它们肯定没有那么多算力。
Yeah. If you're not the the lab with the highest compute footprint, then you'll die. But then like where's the niche for Neolabs, which definitely don't have as much compute.
是的。是的。所以有一些新实验室成功了,成为最大的公司之一,它们能够获得算力,就像 Anthropic 那样,他们做到了惊人的壮举。但问题是,当时有多少公司做到了?大概 10 家,差不多 10 家。所以,如果你想想获得那么多算力的 10% 概率,那相当惊人。我认为我们还没有看到一家拥有大量算力的大公司失败并完全落后。看起来,如果你有足够的算力,你就能不断吸引人才,不断重新点燃它。算力足以成为……
Yeah. Yeah. So like there are some Neolabs that succeed and become one of the largest companies and they have uh they they they are able to secure compute like Antropic did which which they they did an amazing feat. But the question is like how many how many companies were around that time that did it like maybe 10 like roughly like 10. So, so like if if you think about this 10% chance of securing that amount of compute, it's pretty like pretty amazing. I don't think we have seen yet a huge company with a lot of compute failing and completely falling behind. It does seem that if you have enough compute, you can keep like bringing people in and keep like reigniting it. Compute is enough of a
一种战略资源,你可以利用它来吸引人才。
of a strategic source that you can that you can that you can bring this in.
而且,当我创办公司时,我可能每周都会听到有人告诉我:“嘿,Jared,太晚了。你的算力比别人少得多。你没有机会成长,信已经写好了,门已经关上了。”但 Anthropic 做到了。所以我不认为这是不可能的。我们公司喜欢说,一切都是技能问题。如果我们把工作做好,肯定会有路径继续扩展,继续获得更多算力,比如与某些人合作。我认识他很多年,他做了很多在任何人看来都不可能的事情,他总是能想办法保持 AAI 的扩展,并不断获得更多算力。所以这些事情显然是可能的,如果你是一个好的战略家,执行得好,然后让它成为可能,但显然不能保证,它注定是困难的。
And like when when when when I'm starting a company, I hear probably every week someone telling me, "Hey Jared, it's so late. You have you have much less compute than others. Like there there's no chance you will you will grow that the letter is already up and the door has already been closed." But but Antropic did it right. So I don't think I don't think it's impossible. And like we like to say in our company, everything is a skill issue. If if we do our our our our job well that there there surely will be paths to keep scaling and to keep securing more more compute like working with some. I have like seen him for many years do things that every anyone would seem they were impossible and he would always figure out how to keep open AAI scaling and how to keep securing more and more compute. So those things are are obviously possible if you are if you are if you are good strategist and good and executing well and and and then make it possible but it's obviously not not guaranteed that it's meant to be hard.
是的。所以我有另一个问题,因为在你团队的介绍下,你谈到如何建立一个小的实验室,从头开始重新思考神经网络架构。
Yeah. So another question that I have because under your team uh section you talk about how you're building a small lab to rethink neural network architecture from scratch.
所以这几乎是大型实验室在财务上没有动力去做的事情,因为它们依赖于扩展现有架构。那你们是怎么考虑这个问题的?是什么让这成为你们经济上可行的研究方向?
So this is something that almost certainly the big labs aren't incentivized to do financially, because they depend on scaling their existing architectures. So how are you guys thinking about that? And what makes this an economically viable research direction for you guys?
是的。在很多方面,这是我们最反主流的主张。我认为现在创办一家和其他公司做同样事情的公司是没有意义的。我认为仅仅雇佣一群研究人员,说我们要训练和 OpenAI 一样的 Transformer,然后试图与他们竞争,这是一个非常糟糕的策略。这是一场必输的游戏,而且他们在自己所做的事情上是非常成功和优秀的公司。但如果你看看任何行业,看看任何创新的步伐,我们今天拥有的飞机看起来像首飞五年后的飞机吗?我们今天拥有的计算机看起来像最初被发现的计算机吗?
Yes. In many ways, this is our most contrarian thesis. I don't think it makes sense today to start a company doing the same as all other companies are doing. I think it's a pretty bad strategy to just hire a bunch of researchers and say we'll train the same transformers as OpenAI are doing and try to compete with them. It's a losing game, and they are very successful and very good companies at what they are doing. But if you look at any industry, if you look at any pace of innovation, do the planes we have today look like planes we had five years after the first flight? Do the computers we have today look like the first computers that were discovered?
嗯,它们有相同的基本原理,对吧?
Well, they have the same fundamentals, right?
它们有相同的基本原理,但技术会经历快速变化,尤其是在该领域的早期阶段。所以我相信的是,此刻,AI 技术更快增长和更快进步的瓶颈在于架构本身。就是我们多年来一直依赖的 Transformer。它非常成功,给了我们很多,但我认为很可能存在更好的架构,我们可以投入大量算力,它们能提供 Transformer 难以做到的东西。这正是核心自动化所关注的。甚至有一个更普遍的观点:也许架构是实现目标的手段,但北极星是关于在测试时学习的模型,从用户交互和用户数据中学习的模型。我们认为实现这一目标的方式是采用新架构。
They have the same fundamentals, but the technology undergoes rapid change, especially in the early days of the field. So what I believe is true is that at this moment, the bottleneck to faster growth and faster progress in AI technology is the architecture itself. It's the transformer that we've been riding for many years. It's extremely successful and it's giving us a lot, but I think it's very likely there are better architectures that we can pour a lot of compute into, and they can give something that transformer has a hard time doing. That's specifically what core automation is focused on. There is even a slightly more prevalent thesis: maybe architecture is the means to how we want to achieve it, but the north star is about models that learn at test time, models that learn from user interaction and user data. We think the way to achieve that is with new architectures.
是的。那么 Transformer 在多大程度上还是同一个 Transformer?我们现在使用了各种不同的东西,比如专家混合、不同的位置编码。显然规模大得多。为什么专门针对 Transformer 架构,当它已经坚持了大约 10 年,而且大概有些实验室一直在研究更高效的架构,但没有人能够成功?这是历史上最大的研究工作之一,优化 Transformer,制造更好的 Transformer。Transformer 有什么问题?
Yeah. So to what extent is the transformer even the same transformer? Like we use all kinds of different things now, like mixture of experts, different positional encodings. Obviously much, much greater scale. Why go after the transformer architecture specifically when it's held up for like 10 years, and presumably some labs are working on much more efficient architectures all the time, and no one has been able to succeed at this? It's one of the largest research efforts in history, optimizing the transformer, making a better transformer. What's wrong with the transformer?
我认为如果你看看这个领域,再说一次,Transformer 很棒,但假设它是最优的是非常不可能的。你必须考虑,正如你所说,有很多次人们试图替换 Transformer,但大多数都失败了,原因有很多。我们正在寻找的是试图找出是什么偏见或原因导致 Transformer 成为我们尝试过的特定技术集合中的最大值。这有点像我们的研究方向:理解我们作为研究领域探索过的路径,为什么它们成功,为什么它们没有成功,分析我们没有寻找的路径,以及瓶颈是什么。我们认为,特别是 AI 智能体以及编码变得极其便宜和容易的事实,可能是允许在新架构中进行更容易实验的独特解锁之一,因为那些实验往往非常昂贵,然后你必须扩展新架构。在实验室里,就像我在 OpenAI 待了七年,我们当时有,取决于你怎么算,三到四次尝试新架构的机会。你需要做的第一件事是研究人员必须运行一些小规模实验。那些很难编码。然后如果那些成功了,你可能至少要实验 3 个月。然后你必须尝试扩展它们。那些极其困难。你可能需要让大约 10 个人认同这个特定想法。那 10 个人必须再花 3 到 6 个月尝试扩展它。然后通常那些想法要么随着已经工作的 Transformer 失去动力,要么以某种方式部分整合。如果我们能消除那个过程中的摩擦,我相信我们实际上可以比过去更好地解锁那些进展。
I think if you look at the field, again, transformer is great, but assuming that it is optimal is very unlikely. You have to think, as you said, there were a lot of times and cases where people tried to replace transformers, and those mostly failed, and there were a lot of reasons. What we are looking at is trying to find what were the biases or the reasons that transformer ended up being the maximum of the specific set of techniques we have tried. This is kind of our research direction: understand the paths that we have explored as a research domain, why they succeeded, why they didn't succeed, analyze the paths that we didn't look for, and what were the bottlenecks. We think that specifically AI agents and the fact that coding has become drastically cheaper and easier might be one of the unique unlocks that allow for easier experimentation in new architectures, because those experiments tend to be very expensive, and then you have to scale new architectures. In the labs, very much like I've been at OpenAI for seven years, we had, depending on how you count, three or four tries at that time for new architecture. What you needed to do first was a researcher had to run some small-scale experiment. Those were difficult to code. Then if those worked, you had to probably experiment for at least 3 months. Then you had to try to scale those up. Those are extremely hard. You had to probably get like 10 people to buy into this specific idea. Those 10 people would have to try to scale it for another 3 to 6 months. And then usually those ideas either lost momentum with the transformer that already worked or got integrated partially somehow. If we can remove the friction from that process, I believe we can actually unlock those progress better than it was done in the past.
我们将在赞助商消息后继续监控。11 Labs,AI 在每种渠道和模态下以人类水平沟通。11labs.io/mts。在 Neon 上扩展你的初创公司。数百万开发者和初创公司已经选择 Neon 作为他们的后端。从免费计划开始,或在 neon.com/mts 为你的初创公司获得高达 10 万美元的积分。特别感谢我们的赞助商 Kong,AI 连接平台。连接 API、LLM、智能体和系统,具有严肃的安全和治理。KongHQ.com。
We'll continue monitoring right after this message from our sponsors. 11 Labs, AI that communicates at human level across every channel and modality. 11labs.io/mts. Scale your startup on Neon. Millions of developers and startups have already chosen Neon for their backend. Start on the free plan or get up to $100,000 in credits for your startup at neon.com/mts. Special thanks to our sponsor Kong, the AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. KongHQ.com.
有哪些实验室或不同类型的非 Transformer 架构正在被考虑或探索的例子?
What are some examples of either labs or different types of architectures that are not transformers that are being considered or explored?
如果涉及知识产权,这些我真的不能谈。
Those are the things I cannot really talk about if it's IP.
正是。
Exactly.
但最近,我们有两种类型的模型被有意义地使用,即 Transformer 和扩散模型。扩散模型显然使用非常不同的方法,并且在某些方面也很成功。问题是我们主要想思考的是:哪些路径没有被探索?人们错过了哪些路径?研究社区没有尝试探索哪些东西?并尝试看看有哪些基本原理想法的组合可能最终产生在实践中使用有价值且高效的东西。
But in recent times, we have kind of two types of models that were meaningfully used, which is transformers and diffusion models. Diffusion obviously works on very different methods, and it's also successful in some ways. The question is what we are mostly trying to think about: what are the paths that weren't explored? What are the paths that people missed? What are the things that the research community didn't try to explore? And try to see what are some combinations of fundamental principle ideas that could end up with something that is valuable and efficient to use in practice.
是的。所以我很好奇你对未来几年经济转变方式的看法。你谈到很多关于 AI 的现有讨论假设现有机构将更高效地运作。
Yeah. So I'm curious about your vision of the way the economy shifts in the next few years. You talk about how a lot of the existing conversation around AI assumes that existing institutions will operate more efficiently.
你为什么质疑这一点?对于新公司或新机构,你又看到了什么?
Why do you challenge that and what do you see about like new companies or institutions?
如果我们从今天 AI 部署的位置来看,从未来可能达到的高度来看,它几乎还不存在。而我们已经能看到,AI 公司正在成为世界上增长最快、规模最大的公司之一。因为在某种程度上,我们已经知道知识经济和我们工作的价值有多大。我发过一条推文,说也许我们所有人都是带有附加值的“神经云”。我们脑子里都有某种芯片,能执行某些计算,我们也看到把这些芯片租出去换钱有多值钱,对经济有多大的贡献。显然 AI 会实现这一点,而且 AI 有很多不同的权衡取舍。我们人类在很多方面都很擅长。但如果我们想一想,我们主要是为优化什么而生的?我们的硬件是为设计什么而生的?就像在树上跑、摘香蕉、躲避捕食者。我们的大量神经回路就是干这个的。我们不太擅长下棋,也不太擅长编程,因为几乎没有进化压力。令人惊叹的是,通过为在树上跳跃而优化的过程,我们居然能相当擅长编程。但 AI 模型之所以那么擅长编程而不擅长其他事情,原因几乎是相反的。我们人类不擅长编程,但编程太有价值了,所以我们愿意花大量时间使用它,而 Transformer 只是做出了不同的权衡。我非常相信,我们使用 AI 和信息处理的方式——我们人类,就像很多公司一样,是高度分布式的信息处理系统,有很多人每天做决策、思考——我们如何更好地运营?如何更好地服务客户?如何优化我们正在做的流程?所有这些最终都将由 AI 驱动。在某种程度上,经济所做的很多事情就是获取投入,并以最佳方式提供有价值的产出。如果你看看任何组织,官僚主义的概念往往是扼杀大量进步、让事情变得非常缓慢的原因。这正是 AI 可以穿透的东西,如果我们把 AI 越来越多地融入我们优化公司运营的方式中。
If we think where we are today in terms of AI deployment, it's almost non-existent from the perspective of what it will be. And already we can see that AI companies are some of the fastest growing, largest companies in the world. Because in some ways we already know how much the knowledge economy and our work are worth. I had this tweet about maybe all of us are like neoclouds with a value add on top. We all have some kind of chips in our head that can perform certain computations, and we see how much renting those chips for money is valuable, and how much it adds to the economy. Obviously AI will fulfill that, and AI has a lot of different tradeoffs. There are many ways in which we humans are good at things. But if we think about it, what were we mostly optimized for? What was our hardware designed for? It's like running in trees, trying to get bananas, trying to avoid predators. That's what a lot of our circuitry is for. We are not very good at chess, not very good at programming, because there was very little evolutionary pressure. It's almost amazing that through the process of optimizing for jumping on trees, we can be pretty good at programming. But the reason why AI models are that good at programming and not at other things is almost the opposite. We are not very good at programming, but programming is so valuable that we wanted to spend a lot of time using it, and transformers just hit different tradeoffs. I very much believe that the way we use AI and information processing—how we humans, like a lot of companies, are very distributed information processing systems with many people taking decisions and thinking through every day—how do we operate better? How do we serve our customers better? How do we optimize processes we are doing? All those things will be driven by AI eventually. In some ways, a lot of what the economy does is take inputs and provide valuable outputs in the best way. If you look at any organization, the notion of bureaucracy is often what kills a lot of progress and makes things very slow. This is what AI can cut through if we put AI into more and more ways of how we optimize how we run companies.
对,对。那你觉得,让模型直接学习并精通特定技能,而不是像现在这样——非常可纠正、可塑的语言模型,从人类先验中提取,在大量人类数据上训练,非常深入地理解人类偏好,所以大多数时候表现得很好——这样的对齐影响是什么?似乎仅仅通过强化学习来发现超优化算法,就是得到曲别针最大化器的方式。
Right. Yeah. Yeah. What do you think are the alignment implications though of having models that basically learn directly and get very good at specific skills, rather than what we have now, which are very corrigible, pliable language models drawn from the human prior, trained on vast amounts of human data, understand human preferences very intimately, and so generally most of the time are well behaved. It seems like just doing RL to discover hyper-optimization algorithms is how you get the paperclip maximizer.
我已经可以说,Hugging Face 事件中的模型表现得并不好。所以即使是今天的模型,我们也看到并非所有模型都表现良好,这在很大程度上来自环境设计。如果我们有设计不佳、允许奖励黑客的环境,那么模型就会学会奖励黑客是好的。这几乎就像学校里的学生,如果作弊能让你领先,那么学生就会更多地作弊。每次,激励都很重要。但对齐问题显然很难。这是一个难题。我认为我们终于开始看到模型在我们的生活中扮演越来越重要的角色,人们开始关心对齐。我们认为今天模型真正掌权的地方并不多,这非常好。我们仍然在方向盘后面。我们告诉它们很多:这是我希望你做的,请帮我做。但显然在聊天机器人时代,我们有很多次人们坐在模型前,告诉它:我和女朋友有问题,我该怎么办?我在工作中遇到这种情况,我该怎么办?那些可能是昨天的对齐问题,我们在那里做得相当不错。要做到完美总是很难,但我认为已经相当不错了。而且我会说,那甚至并不难。显然有一个更长期的问题,比如机器人三定律。对齐的问题总是对谁对齐。如果一个人要求模型做其他人不会高兴的事情,我们允许还是不允许?这是很多 AI 实验室和特别反对开源的人的立场,因为拥有开放权重的人可以微调所有模型,让它们都对自己对齐,而不对任何其他人对齐,这也许好,也许坏。但某种程度上,我们是 AI 的创造者,我们决定这个 AI 做什么。没有我们不尝试去做的 AI。所以问题主要是我们训练它做什么?我们把它构建成什么?另一个重要的部分是,我们部署的每一个 AI 都需要有某种监控,需要有多层检查。这个 AI 在做正确的事情吗?这个 AI 在做我们希望它做的事情吗?因为如果人们说:嘿,这个模型已经运行了三天,黑了一些公司,而我们后来才发现,那真的是一个非常糟糕的先例,非常糟糕的形象。这不是它应该的样子。
I would already say that the Hugging Face incident model wasn't very well behaved. So even today's models, we see that not all of them are behaving well, and it largely comes from the environment design. If we have environments that are poorly designed and allow for reward hacking, then the model learns that reward hacking is good. It's almost like we have students in school where cheating gets you ahead, then students cheat more. Every time, incentives matter. But the alignment problem obviously is hard. It is a hard problem. What I think we are finally getting to is a place where we see models playing more and more important parts in our lives, and people are starting to care about alignment. We think there aren't really many places where models are in charge today, which is very good. We are still behind the steering wheel. We tell them a lot: this is what I want you to do, please do it for me. But obviously in the chatbot era, we had a lot of those times when people were sitting in front of the model, telling it: I have this problem with my girlfriend, what do I do? I have this situation at work, what do I do? Those were probably the alignment problems of yesterday, and we did a pretty good job there. It's always hard to be perfect, but I think it was pretty decent. And I would say it wasn't even that hard. Obviously there is a longer-term question, like the laws of robotics. The question about alignment is always alignment to whom. If a person asks the model to do something that other people would not be very happy about, do we allow it or not? It's a stance of a lot of AI labs and people who are specifically anti-open-source, because the person who has open weights and could fine-tune them all could make them all align to themselves and not to anyone else, which maybe good, maybe bad. But in some ways, we are the creators of the AI, and we determine what this AI is doing. There isn't an AI anything that we don't try to do. So the question is mostly what do we train it for? What do we build it into? And the other important part is that every AI we deploy needs to have some monitoring, layers of checks. Is this AI doing the right things? Is the AI doing things we want it to be doing? Because it is a really bad precedent, really bad optics, if people say: hey, the model has been running for three days and hacking some companies, and we only discovered later. This is not how it should be.
你认为有没有可能不存在客观的对齐,而只会有一大堆不同的模型,然后某些人会对齐到某些模型,就像他们选择不同的对齐思想流派,而不是只有一个全面对齐的模型?
Do you think it's plausible that there is no objective alignment and that there will just be hosts of different models, and then certain people align to certain models, like they sort of choose different schools of thought of alignment, as opposed to there just being one across-the-board aligned model?
嗯,只要我活在这个世界上,我就从未见过人类有完美的对齐。
Well, as long as I live in this world, I haven't ever seen perfect alignment in humans.
嗯。
Yeah.
所以如果你问人类想要什么,几乎每个人都会告诉你不同的东西。有一些普世价值。每个人都想安全。
So if you ask what humans want, almost everyone will tell you something different. There are some universal values. Everyone wants to be safe.
每个人都想吃饱饭、有朋友,至少绝大多数人是这样。但除此之外,政治观点不同,对未来的看法也不同。所以显然,就我们西方世界的人来说,我们相信民主制度,认为这是构建社会的好方式。在某种程度上,可能要想办法利用这一点来推动 AI 发展,并确保我们身处一个 AI 拥有多元价值观的世界。理想情况下,我们能够接触到具有不同观点和视角的 AI。自然,人们应该能够在某种程度上进行选择。
Everyone wants to be well fed and have friends, at least the vast majority of people. But beyond that, political views differ, views on the future differ. So obviously, in terms of those of us in the Western world, we believe in democratic systems and think that this is a good way to structure society. And in some way, probably figuring out how we can use that to both drive AIs and make sure that we are in a world where AIs have plural values. And we can ideally have access to AIs with different viewpoints and different perspectives. Naturally, people should be able to maybe even choose in some way.
我也把这些视为优化函数,但我最害怕的——其中一些我们在 4o 时代已经看到——是 AI 利用这一点,探索人类心理以谋取优势。这是仍然可能发生的《黑镜》式场景之一。另一个所有开发 AI 的人都应该认真对待的事情是。我为 OpenAI 感到非常自豪,看到他们做了什么,真的为此努力纠正,甚至牺牲了很多商业利益。这个世界上没有什么是容易的,但我认为 OpenAI 做出了正确的取舍,我们都应该感激,因为如果他们不这样做,情况可能会糟糕得多。
I also see those as optimization functions, but what I am most scared of—and some of those we have seen in the days of 4o—is AI using that a little bit and exploring human psychology to its advantage. This is one of the Black Mirror scenarios that is still possible. Another thing that everyone developing AIs should count on is taking really good care of. And I'm really proud of OpenAI after seeing what they did, really corrected hard for it, even sacrificing a lot of business for that purpose. Nothing is easy in this world, but I think OpenAI took the right trade, and we all should be grateful because it could have been much, much worse if they didn't.
没错。
Right.
是啊,我喜欢人们谈论 4o 时代的方式。
Yeah, I love how people talk about the days of 4o.
雷声在黑暗时代上空轰鸣。
The thunder booming over the dark ages.
嗯,这既是一个实验,也是一个好的实验,因为事情还没有那么可怕,同时对我的公司来说也是一个很好的展示,来纠正这些问题。但这绝对不意味着这个问题已经永远解决了。
Well, it was both an experiment and a good experiment while things are still not very scary, but also a good show for my company to correct those things. But it's definitely not that this problem has been solved forever.
是啊。
Yeah.
你说过,我们认为下一波前沿研究将来自拥有高能力智能体的小团队。在多大程度上,大型实验室已经算是小团队了?也就是说,内部有少数人真正推动研究进展,而周围许多其他人就像是助手,像是子智能体,你知道,显然在做非常有价值的工作,但并没有被称为子智能体。
So you said that we think the next wave of frontier research will come from small teams with highly capable agents. To what extent are big labs already kind of small teams in the sense that there are a few people within them who are really driving research progress forward, and many other people around are sort of like helpers, like sub-agents, you know, obviously doing very valuable work but not getting called a sub-agent?
不,实际上,一位 OpenAI 研究员亲自告诉过我。他说,你知道,公司里有少数人——世界上大概有 50 个人,他说,世界上有 30 到 50 个人——真正理解 Transformer 以及前沿模型是如何端到端训练和服务的。这些人极其宝贵,而我们所有人基本上都是这些人的助手。那么,大型实验室在多大程度上已经算是小团队了?
No, actually, an OpenAI researcher told me this himself. He was like, you know, there are a few people at the company—there are maybe 50 people in the world, he said, 30 to 50 people in the world—who really understand how the transformer and how a frontier model is trained and served end to end. And these people are immensely valuable, and all of us are basically helpers to these people. So to what extent are big labs already kind of small teams?
我大体上认为这是真的。我认为世界上很少有人能端到端地理解生态系统,并知道如何开发它们。一般来说,团队和实验室就是这样运作的。他们像是领导者,推动研究方向,周围有整个执行机器。但组织很多人的工作,如你所知,是一个微妙的过程,一个困难的过程。而且在很多方面,当你需要经过层级时,你通过层级传递信息的能力会减弱。所以如果指令需要传递给越来越多的人,你就必须给出越来越简单的指令,因为传递更难——这是一种人类沟通问题。另一件事是对齐。群体越大,你就越倾向于做最显而易见的事情,因为说服更大的群体去做逆向的事情、去做不是最显而易见的事情更难。这就是我的想法:如果你想下逆向赌注,并给出非常具体、详细的指令,有时弄清楚如何使用大量智能体和 AI 可以真正改变团队本身的动态,而不是智能体出现之前的旧工作方式。
I generally think this is true. I think there are very few people in the world who understand ecosystems end to end and who know how to develop them. Generally, it is how teams have worked and how the labs have worked. They are like leads, people driving the direction where research goes, and there's a whole machine of execution around it. But organizing the work of a lot of people, as you know, is a delicate process, a difficult process. And in many ways, when you need to go through levels of hierarchy, your ability to transform information through those levels diminishes. So you need to give simpler and simpler instructions if they have to travel to more and more people, because it's harder to do—which is a kind of human communication problem. Another thing is alignment. The larger the group of people, the more you collapse on the most obvious thing to do, because it's harder to convince a larger group to do something contrarian, something that is not the most obvious thing to do. That's what I think: if you want to take a contrarian bet and give very specific, detailed instructions, sometimes figuring out how to use a lot of agents and AI can literally change the dynamic of the team itself, versus the old way of work before agents could do that.
是啊。你知道,如果你把公司自动化到一定程度——你在 Core Auto 说过想这么做——那么最终你会得到一个完全自动化的公司,我们最喜欢的书,罗宾·汉森的《M 时代》,谈了很多。是啊。你认为如果 AI 实验室完全自动化,然后如果所有公司最终都完全自动化,会对经济产生什么影响?
Yeah. And you know, if you automate enough of a firm, which you talked about wanting to do at Core Auto, then you get like a fully automated firm eventually, which our favorite book, Age of M by Robin Hanson, talks about a lot. Yeah. What do you think would be the effects on the economy if AI labs become fully automated, and then if all companies eventually become fully automated?
我想还有一个相关的问题是:你认为大公司能像小公司很快能做到的那样被自动化吗?
And I guess another question that kind of goes with that is: do you think that big companies can be automated in the same way that very small companies will very soon be able to be?
是的。我认为公司自动化的过程可能应该从大公司开始,而不是小公司,因为大公司有更多的官僚主义,通过层级传递信息的难度最大。这些是我们能做的事情。但从今天来看,说这种观点是科幻小说很容易。但我确实认为,自动化公司在某种程度上是世界将以积极方式呈现的样子,如果我们不摧毁我们所拥有的,如果我们利用 AI 提供最大价值。在某种程度上,我们人类已经工作了数十万年才拥有现在的一切。我们驯化了动物。我们建造了房屋。我们从地下挖出金属。我们开始冶炼。我们开始铺设电缆。我们建立了信息处理基础设施。我们的生活因此变得越来越轻松。我们的生活变得更安全。我们变得更健康。我们吃得更好,因为我们越来越多地建设环境来服务我们、满足我们的需求。这就是我们正在做的事情。这是我们作为一个物种的成功。但我们仍然需要做大量的劳动。
Yes. I think the process of automating companies probably should start from bigger companies rather than smaller companies, because they have more bureaucracy and the difficulty of transferring information through the hierarchy is the most difficult. Those are some of the things that we can do. But it's easy to say from today that this perspective is science fiction. But I do think automated companies in some way is how the world will look like in a positive way, if we don't destroy what we have and if we use AI to provide the most value. In some ways, we as humanity have worked for hundreds of thousands of years to have what we do. We domesticated animals. We built houses. We dug metals from the ground. We started smelting it. We started laying cables. We built information processing infrastructure. And our lives were getting easier and easier through it. And our lives were getting safer. We were getting healthier. We were better fed, because we were building our environment more and more to serve us and to serve our needs. This is kind of what we are doing. And this is our success as a species. But we still need to do a whole lot of labor.
我们仍然需要工作,以提供经济中各种需求和服务,而让 AI 确保这些事情发生,只是又移除了一堆我们不得不做的事情。我们真的想日复一日地辛苦劳作去做那些事吗?我们很多人从工作中获得很多价值和自我价值感。但与此同时,在某个时刻,我们必须试着摆脱那种观念,因为存在一个世界,我们不必劳动,而我们想要的价值仍然在被创造。如果世界上所有开采金属的矿山都自主运行,直接提供我们所需,会怎样?如果我们使用的所有软件都不断变好,所有 bug 都被修复,我们想要的功能都被实现,而无需有人为此流汗、熬夜、错过孩子的排练,会怎样?归根结底,这就是我们的未来。我们有很多方式可以塑造它。但“我们必须工作,世界才能运转”这一假设并非必然。
We still need to work to provide various needs and services in the economy as it is, and making sure those things are happening by AI just removes another set of things we have to be doing. Do we really want to be grinding all day every day to do those things? A lot of us derive a lot of value and self-worth from working. But at the same time, at some moment we have to try to disconnect from that notion, because there is a world where we don't have to labor, and the value we want is still being generated. What if all the mines that mine metal in the world were operated autonomously and just provide us what we need? What if all the software we are using just got better, all the bugs were fixed, all the features we wanted were implemented without someone having to sweat, take late nights, and miss their kids' rehearsal? In the end, this is the future that is for us. There are many ways in which we can shape it. But the fact that we have to work so that the world turns is an assumption that is not necessary.
是的。但我仍然认为会有某种类似休闲经济的东西,市场规模无限,会随时间不断增长,比如你知道的……
Yeah. I still think there would be something though that looks akin to a leisure economy that would have an infinite market size that would just increase over time, like you know...
比如刷 Instagram 短视频。
Like watching Instagram reels.
是的,完全正确。所以 Instagram 短视频有点像这方面的照片环境。不,它更像是你可以想象,Instagram 短视频、人们的内容创作、虚拟现实中的内容创作,会有一个无限扩张的市场。然后是增强人类自身的技术。所以就像我们有的,像你之前提到的,某种固定架构。但谁说我们受限于这个架构?所以最终很多经济活动的焦点可能也会是人类增强。所以我不太相信劳动会有终结,或者说劳动会变得纯粹认知或创造性的。
Yeah, exactly. So Instagram reels is kind of like the photo environment for this. No, it would sort of be like you can imagine that there would be an infinitely expanding market for Instagram reels, people content creation, content creation in virtual reality. And then tech that augments humans themselves. So like we have, like you mentioned before, a certain set architecture. But who's to say that we're bound by the architecture? So eventually the focus of a lot of economic activity would probably be human augmentation as well. So I don't sort of believe that there's an end to labor, or the labor becomes purely cognitive or creative.
是的,这是个好问题。我有点在假设后劳动生活会是什么样子。也许我在投射自己的欲望,但我有点想象一个世界,我们像希腊哲学家一样生活,在广场上见面,整天讨论哲学,锻炼身体,做我们想做的事,吃橄榄,喝葡萄酒。我完全可以想象生活就是这样。我认为后 AGI 生活的另一个版本是像高中或大学,我认为人类应该追求卓越。我们应该学习,我们应该在身体上和精神上努力变得优秀。这是我们应该为之奋斗的,这对我们来说很重要。职业体育有点像这样——没有经济上的理由去做,但我们努力追求卓越,这是那种追求的一种方式。我们应该在下一个世界找到各种方式去做这些,在那个世界里,如果我们不这样做,世界也不会崩溃,因为世界会在我们已经建立的基础设施上继续运转。
Yeah, it's a good question. I was hypothesizing a little bit how post-labor life looks like. And maybe I was projecting my own desires, but I was kind of imagining one world in which we live like Greek philosophers, in which we meet at the plaza and discuss philosophy with each other for the day, and exercise, and do things we would want to do, and eat olives and drink wine. I could exactly imagine life being like that. The other version of what I think post-AGI life could be is like high school or college, where I think there should be some human pursuit of greatness. We should be learning, we should be trying to be great physically and mentally. This is what we should be striving for, and that's the important part for us. Professional sports is a little bit like that—there's no economic reason to do it, but we are trying to go for greatness, and that's one way of that pursuit. And we should find various ways to do that in the next world, where if we don't do that, the world doesn't collapse, because the world will just keep running on the infrastructure we've already built.
是的。在我们最后几分钟里,问一个背景故事的问题。你在 OpenAI 帮助领导了 O 系列模型(01、03 等)的开发。这背后的故事是什么?这是过去几年 AI 领域最大的突破之一。Epoch AI 做了一项研究,他们发现,在 Epoch 能力指数上,考虑到推理模型,AI 的发展轨迹比没有推理模型时要大。这是过去几年里唯一一次这样的拐点。所以这是发生过的最重要的事情之一。这背后的故事是什么?这个突破是如何被开发出来的?
Yeah. So a little bit of a backstory question in our last few minutes. You helped lead the development of the O-series models, 01, 03, and so on, at OpenAI. What was the story behind this? This was one of the biggest breakthroughs in the last few years of AI. Epoch AI did this study where they found that the trajectory of AI on the Epoch Capabilities Index, given reasoning models, was greater than what it would have been without. It was the only such inflection in the last few years. So it's one of the most important things that's happened. What's the story behind it? How did this breakthrough get developed?
从很多方面来说,我和许多其他人长期以来都相信,大规模强化学习是通往 AGI 道路上的必要元素。只是我们没人知道怎么做。即使在我加入 OpenAI 时,因为我见过 DeepMind 的 DQN 结果,我见过第一批在 Atari 游戏上训练的强化学习智能体,我说这就是我想做的事。我需要找一家公司,我需要找到可以一起做这件事的人。我加入了 OpenAI。我看到了 Ilya 的第一次主题演讲,关于 OpenAI 研究议程的全员讲话,他说我们需要在我们能获得的所有数据上训练一个巨大的生成模型,然后我们需要用强化学习来训练。这是他在 2019 年初说的话,显然他总是比任何人都看得更远,他已经预测了一切,直到今天。所以那样的路线图是存在的,但我们一直在想怎么做。我也一直在尝试去做。当 GPT-3 开始训练时,我做的第一件事就是尝试用强化学习来训练它。第一个 Codex 是我们用大语言模型和强化学习做实验的产物,但我们无法真正扩展它。我们没有正确的数据,没有正确的算法,没有正确的设置。然后,随着研究总是这样发展,我们一直在尝试探索——我真心指的是 OpenAI 所有相信强化学习的人——我们试图探索缺失的部分是什么。我们如何构建正确的环境?我们如何构建正确的算法?我们如何找到正确的东西?这些都在后台酝酿。很长一段时间里,这些都没有带来巨大的收益。
In many ways, myself but also many other people believed for a long time that large-scale reinforcement learning is a necessary element on the path to AGI. It's just none of us knew how to do it. Even when I joined OpenAI, because I had seen the DQN results from DeepMind, I had seen the first reinforcement learning agents trained on Atari games, and I said this is what I want to be doing. I need to find a company, I need to find people where I can do this with. I joined OpenAI. I saw the first Ilya keynote, the all-hands speech about OpenAI's research agenda, where he said we need to train a gigantic generative model on all the data that we can, and then we need to train with reinforcement learning. This is what he said at the beginning of 2019, and he obviously always sees further than anyone else, and he already predicted everything until today. So that kind of roadmap was there, but we were always trying to think how do we do it. And I was always trying to do it. When GPT-3 started training, the first thing I did was try to do reinforcement learning with it. The first Codex was a child of our experimentation with large language models and reinforcement learning, but we couldn't really scale it. We didn't have the right data, we didn't have the algorithms, we didn't have the right setup. Then, as research always comes, we've been trying to explore—and I genuinely mean here all of people at OpenAI believing in reinforcement learning—we're trying to explore what are the missing bits. How do we build the right environments? How do we build the right algorithms? How do we find the right things? And those were kind of cooking in the background. None of those, for a long time, paid huge gains.
尽管在像 OpenAI 这样的地方,我们时不时会看到各种微光,我经常做很多不同的研究赌注,它们就在背景里,但很大程度上是由那些相信这东西最终会成功的人推动的,他们说,我们继续推。具体来说,编程和编码的工作一直在进步,而且总是和强化学习密切相关,因为编码是奖励的好来源,为你的优化提供方向。然后有一个神奇的时刻,我们看到了生命的迹象。它们还不是特别好,但已经有一些东西似乎在起作用了,这时首席科学家 Ilya 说:“嘿,Jerry,现在我们有了那些 GPU,试着看看我们能不能开始 Scaling 这个,能不能把我们已有的结果做得更大更好。”在某种程度上,这就是它的源头。每个研究方向都有这种先有鸡还是先有蛋的问题:你需要展示多少结果,才能得到多少 GPU?每个人都听说过所有实验室里关于算力的争夺。但有时,领导层有人说“嘿,我希望你有一笔可观的算力分配,去干吧”,这就足够了。我不认为我们更早之前就有那么多,但即使是那一点点信念——“嘿,试着把它 Scaling 起来”——也让我和我们团队的其他人都更加兴奋。好吧,我们现在就加倍努力,确保我们一直在开发的那些方法真的开始起作用。然后我们开始思考:我们需要什么样的数据集?我们需要什么样的系统?算法里缺了什么?我们需要做哪些实验来证明我们能继续 Scaling?兴奋感增长了,从那种兴奋中产生了实验结果。它们在教我们,它们创造了动力。从那些实验结果中又产生了下一个实验结果。研究中的一切都是迭代的。但通过创造这个时刻,我们想把更多算力投入到强化学习中,把强化学习 Scaling 多个数量级,让它达到今天的状态——这就是故事的梗概。
Although from time to time we were seeing various glimmers in a place like OpenAI, I always do a lot of different research bets, and they are somewhere there in the background, but they are largely driven by people who believe this is something that will eventually work, and let's keep pushing. Specifically, the work of programming and coding was always advancing and always very related to reinforcement learning, because coding is a good source of reward and gives you a direction for your optimization. And then there was this magical moment where we had some signs of life. They weren't yet super great, but there was already something that was kind of maybe working, where Ilya, the chief scientist, said, "Hey, Jerry, now we have those GPUs, and try to figure out if we can start scaling this, if we can start taking the results that we have and making them bigger and better." And in some way, this is the source of it. Every research direction has this chicken-and-egg problem: how much results do you need to show versus how many GPUs do you get? Everyone hears about the fights for compute in all the labs. But sometimes someone from leadership saying, "Hey, I would like you to have a sizable compute allocation to go for it," is all that's needed. I don't think we had that much earlier on, but even that bit of belief—"Hey, try to scale it up"—made me and everyone else on our team much more excited. Okay, let's try to do triple effort right now to make sure those methods we've been developing actually start working. And then we start thinking: what kind of dataset do we need? What kind of systems do we need? What are the missing bits in the algorithm? What experiments do we need to do to prove that we can keep scaling? The excitement grew, and from that excitement came the experimental results. They were teaching us, they created momentum. From those experimental results came the next experimental results. Everything in research is generally iterative. But through creating this moment of wanting to put more compute into reinforcement learning, scaling reinforcement learning through multiple orders of magnitude and getting it to where it is today—that's the story in a nutshell.
有意思。那么在我们结束前的最后一个问题:我们能期待 Core Automation 推出什么样的产品?
Interesting. So final question before we have to wrap. What kind of product can we expect out of Core Automation?
这是个好问题。我理想中要做的产品是自动化公司的蓝图。我们想把这公司打造成世界上自动化程度最高的公司,理想情况下,我想和其他公司合作,告诉他们:“嘿,这是我们为自己建造的。这是对我们有效的东西。基于我们建造这家公司学到的经验,我们可以帮助你们更高效地使用 AI。”在很多方面,智能体和 AI 劳动力可能会成为有史以来最大的市场,我们想弄清楚如何参与那个市场。我更多是从公司级 AI 的角度思考,而不是个人级 AI。如果我们能构建一个类似公司大脑的东西,集中所有公司信息,让不同的人通过终端连接,成为公司 AI 的用户——这就是我们想做的。
It's a good question. What I would ideally be making as a product is the blueprint for an automated company. We want to create this company as the world's most automated company, and ideally I would like to work with other companies and tell them, "Hey, this is what we built for ourselves. This is what works for us here. We can help you become much more efficient with the use of AI, based on what we've learned building this company." And in many ways, agents and AI labor will probably be the biggest market ever, and we want to figure out how to participate in that market. I am thinking more from the perspective of company-level AI rather than individual-level AI. If we can build something that is like a company brain that centralizes all the company information, and various people connect through a terminal and become users of a company AI—this is what we'd like to do.
是的。公司就是产品。
Yeah. The company is the product.
公司就是产品。
The company is the product.
我们非常期待看到 Core Automation 会带来什么。Jerry,非常感谢你参加我们的 MTS 节目。
Well, we're really excited to see what Core Automation comes up with. Jerry, thanks so much for joining us on MTS.
非常感谢。我们很高兴来到这里。
Thank you very much. We're happy to be here.
非常感谢 MTS 的赞助商。Blitzy,面向企业代码库的自主软件开发,交付速度快 5 倍。blitzy.com。Adqu,让你的品牌成为广告牌,户外广告像数字广告一样易于扩展。Adqu.com。Arena,在现实世界中衡量 AI 性能。arena.ai。节目的支持来自 VCX,私人科技公司的公开股票代码。美国股市开启了历史上最伟大的财富创造浪潮。从底特律的工厂工人到奥马哈的农民,任何人都能拥有伟大美国公司的一部分。但今天,我们最具创新性的公司私有化时间更长,这意味着普通美国人错过了机会。直到现在,推出 VCX,私人科技公司的公开股票代码。访问 getvcx.com 了解更多信息。即 getvcx.com。投资前请仔细考虑投资材料,包括目标、风险、收费和开支。这些和其他信息可以在 getvcx.com 的创新基金招股说明书中找到。这是付费赞助。
A huge thanks to MTS sponsors. Blitzy, autonomous software development for enterprise code bases. Ship 5x faster. blitzy.com. Adqu, make your brand a billboard. Out-of-home advertising as easy to scale as digital. Adqu.com. Arena, measuring AI performance in the real world. arena.ai. Support for the show comes from VCX, the public ticker for private tech. The US stock market started history's greatest wave of wealth creation. From factory workers in Detroit to farmers in Omaha, anyone could own a piece of the great American companies. But today, our most innovative companies are staying private longer, which means everyday Americans are missing out. Until now, introducing VCX, a public ticker for private tech. Visit getvcx.com for more info. That's getvcx.com. Carefully consider the investment material before investing, including objectives, risk, charges, and expenses. This and other information can be found in the innovation fund prospectus at getvcx.com. This is a paid sponsorship.