Defining AGI: The Novel Test and Measuring Intelligence
打开互动全文版(中英对照 + 朗读 + 问答)→关于为 AGI 定义智能的讨论,以写出引人深思的小说为基准,并评估当前如困惑度和基准测试等衡量方法。
A discussion on defining intelligence for AGI, using the ability to write a thought-provoking novel as a benchmark, and evaluating current measures like perplexity and benchmarks.
我认为我们想从智能的定义开始。如果要认真讨论 AGI 或解决 AGI 或走完剩下的路,定义 I 可能是第一步。那么,你如何看待定义智能?
I think we want to start with a definition of intelligence. If we want to talk seriously about AGI or solving AGI or getting the rest of the distance, defining the I is probably the first step. So, how do you think about defining intelligence?
我确实没有很好的答案。我相信很多人都尝试过并失败于给出一个真正具体的智能定义。我记得 2017 年底我做 Reddit AMA 时,有人问我:‘你担心 AGI 吗?我们有能下围棋的机器人,你刚做了一个打扑克的机器人。模型似乎变得相当智能。你认为它们在接下来 10 年内做不到什么?’我说:‘老实说,我现在不太担心 AGI。我可以以超过 90% 的把握说,AI 在接下来 10 年内无法写出一部发人深省的小说。如果 AI 能做到,那我就会非常害怕 AGI。’那是 2017 年。现在还剩大约一年半。我的预测看起来不太乐观。但我觉得那是一个很好的智能例子。我坚持这个观点。我不会移动我的目标。如果我们能有一个 AI 真正写出发人深省的小说,那将是智能的一个伟大标志。
I really don't have a great answer here. I'm sure a lot of people have tried and failed to come up with a really concrete definition of intelligence. I remember when I did a Reddit AMA back at the end of 2017, somebody asked me, 'Are you concerned about AGI? We have bots that play Go and you just made a bot that plays poker. It seems like the models are getting pretty intelligent. What do you think they won't be able to do within the next 10 years?' I said, 'To be honest, I'm not super worried about AGI right now. I can say with greater than 90% confidence that an AI will not be able to write a thought-provoking novel in the next 10 years. If an AI could do that, then I'll be very afraid of AGI.' That was 2017. I have about a year and a half left now. Not looking great for my prediction. But I'd say that's a good example of intelligence. I kind of stick to that. I'm not going to move my goalposts. If we could have an AI that can actually write a thought-provoking novel, that's a great sign of intelligence.
对谁来说发人深省?
Thought-provoking to whom?
有道理。我不会说这是智能的全面定义,但这是我想到的一个例子,说明我们需要达到的目标以及我们尚未达到的地方。
That's fair. I wouldn't say this is an all-encompassing definition of intelligence, but it's one example I've come up with for where we need to go and where we're not quite there yet.
你认为限制步骤是什么?为什么它写不了小说?我们可以做下一个词预测,让它一直生成到书末。这只是一个提示问题吗?
What do you think is the limiting step? Why can't it write a novel? We can do next-token prediction and just let it roll out until the end of the book. Is it just a prompt issue?
2017 年,连拼凑一个连贯的句子都难以想象。我们取得了很大进步,我认为实际上已经很接近了。小说的主要挑战是任务长度。一部小说至少有 10 万个 token。想想人类写一部好的发人深省的小说需要多长时间——他们可能要花一年多。如果你看看像 meter plot 这样的东西,我虽然对它有些意见,但它是一个相当不错的评估。AI 要多久才能做人类需要一年才能完成的事情?我认为大多数事情我们还没到那一步,但我们正在进步。
In 2017, it was unthinkable to even string a coherent sentence together. We've made a lot of progress, and I think we're actually getting pretty close. The main challenge with a novel is the length of the task. A novel is probably at least 100,000 tokens. Think about how long it takes a human to write a good thought-provoking novel—they might spend more than a year. If you look at things like the meter plot, I have my issues with it, but it's a pretty good eval. How long until AIs can do stuff that takes a human a year? I don't think we're quite there yet for most things, but we're making progress.
所以你对 AGI 基准的粗略代理是写小说。有没有什么归一化或控制迭代次数、先前经验、产生小说所需的焦耳?如果思考量相同,但焦耳少 100 倍或训练迭代少 1000 倍,或者完全从头开始没有训练,那不是更令人印象深刻吗?
So your rough proxy for an AGI benchmark is writing a novel. Is there anything normalizing or controlling for number of iterations, prior experience, joules to produce the novel? If it were equal amounts of thoughtfulness but with 100x less joules or 1,000x less training iterations, or tabula rasa with no training at all, wouldn't that be more impressive?
当然,有办法让它更具挑战性或更令人印象深刻。10 年前争论这些都很荒谬。现在细节确实重要。我认为现有模型可能可以通过某种方式搭建框架,输出一部发人深省的小说。这足以达到 AGI 吗?可能不是。我想看到一个相当轻量的框架。但如果你在 2017 年告诉我,到 2026 年我们可以通过搭建 LLM 框架输出一部发人深省的小说,我会说这符合我对 AGI 的定义。
Sure, there are ways to make it more challenging or impressive. This was all absurd to debate 10 years ago. Now the details do matter. I think it's probably possible with existing models to scaffold them in a way that outputs a thought-provoking novel. Is that sufficient for AGI? Probably not. I'd want to see a pretty light scaffold. But if you told me in 2017 that by 2026 we could scaffold LLMs to output a thought-provoking novel, I would have said that's sufficient for AGI under my definition.
很好,你没有频繁移动目标,这经常发生。在现有的智能度量中,比如验证集困惑度(来自 Chris Ray 实验室的每瓦特智能),或者切换到 GSM8K 中井号之后的所有内容,或者一堆基准测试,或者像 Arc Prize 这样的努力。你对这些作为智能的代理有什么看法?
It's good that you're not moving the goalpost frequently, which often happens. Of the available measures of intelligence, like val perplexity (intelligence per watt from Chris Ray's lab), or switching to everything after the hashtag in GSM8K, or the cacophony of benchmarks, or efforts like Arc Prize. What are your thoughts on these as proxies for intelligence?
衡量智能非常困难。这就是为什么我只举了一个例子。我们甚至不知道如何衡量人类的智能。人们谈论不同类型的智能。如果我们没有衡量人类智能的方法,就很难在模型上有一个客观的方法。所有这些都有效,但都有挑战。例如,如果一个模型在数学上微调,它会在 GSM8K 上表现很好,但如果你关心写小说,它就更智能吗?可能不是。这就是 Chollet 的观点:我们必须控制先验经验。这很难做到,尤其是在人类身上。
It's very difficult to measure intelligence. That's why I gave one example. We don't even know how to measure intelligence in humans. People talk about different kinds of intelligence. If we don't have a way to measure it in humans, it's hard to have an objective way in models. All these are valid, but they have challenges. For example, if a model is fine-tuned on math, it'll get good at GSM8K, but is it more intelligent if you care about writing novels? Probably not. That's Chollet's point: we have to control for prior experience. That's hard to do, especially in humans.
你之前有多少经验?有没有根据先前经验进行归一化的做法?所以,我可以进行奖励黑客,但如果不这样做,模型获得相同性能是否更令人印象深刻?
How much prior experience do you have? And is there anything to normalizing by prior experience? So, I can reward hack, but if I don't, is it more impressive that it gets equal performance?
我对此没有一个非常满意的答案,因为我没有一个日常使用的智能定义。我认为评估是多样化的。综合这些评估可能是衡量智能的最佳方式。不同的评估捕捉不同的方面。如果你是一位数学家,你可能关心模型能否证明数学问题。ARC 捕捉了另一种智能,比如适应新环境。这取决于你希望这些模型做什么。
I don't have a super satisfying answer here because I don't have a definition of intelligence that I use day-to-day. I think there's a diversity of evals. Looking at the aggregate of these evals is probably the best way to measure intelligence. Different evals capture different aspects. Maybe if you're a mathematician, you care about whether the model can prove math problems. ARC captures another kind of intelligence, like adapting to novel environments. It depends on what you want these models to do.
我确信在 OpenAI,发布新模型前你们有一套庞大的评估集。可能是静态的、动态的,或者包含人类参与。那具体包括什么?是一套庞大的内部任务集吗?OpenAI 基准?可能不能说。
I'm sure at OpenAI you have some massive eval set before releasing a new model. Maybe it's static, dynamic, includes human in the loop. What does that entail? Is it a sprawling set of internal tasks? OpenAI bench? Probably can't say.
好的。
Okay.
现在考虑量化智能,你最有兴趣的主义是什么?有 Ilyaism——下一个词预测就够了。有 gnomeism——Ilyaism 加上生成器-验证器差距,通过生成可验证问题来获得更多数据。你同意 gnomeism 的这种表述吗?
Now thinking about quantifying intelligence, what are your most interesting isms? There's Ilyaism—next token prediction is enough. There's gnomeism—Ilyaism plus a generator-verifier gap, generating verifiable problems for more data. Do you agree with that phrasing of gnomeism?
我认为这没有涵盖全部。另一个关键部分是推理算力——通过更长时间的思考来提高效率。这是一种预测一系列词元的下一个词预测形式。还有其他扩展推理算力的方式,但我认为这是智能的一个关键要素。
I don't think it captures the whole story. Another key part is inference compute—the ability to be more productive by thinking longer. It's a form of next token prediction where you predict a series of tokens. There are other ways of scaling inference compute, but I consider that a key ingredient to intelligence.
对测试时算力的反驳:如果我有一个人为的例子,比如学习排序,而人类还不知道归并排序,只知道冒泡排序。我有无限多的未排序列表,带有冒泡排序过程和排序后的答案。这是可验证的。我可以生成无限数据。但我永远无法通过测试时算力得到归并排序。为什么?我只会得到一个糟糕的冒泡排序版本。它被训练成精确执行冒泡排序。它如何压缩成不同的算法?
A counter to test-time compute: if I have a contrived example like learning to sort, and mankind doesn't know merge sort yet, only bubble sort. I have infinite traces of unsorted lists with bubble sort procedure and sorted answer. That's verifiable. I can generate infinite data. But I will never test-time compute my way to merge sort. Why not? I'd just get a crappy version of bubble sort. It's trained to do bubble sort exactly. How would it compress to a different algorithm?
如果你只训练它冒泡排序,那么它可能不会做超出这个范围的事情。这就是为什么预训练数据的多样性很重要——以确保充分的探索。
If all you trained it on is bubble sort, then it probably won't do anything beyond that. That's why diversity in pre-training data is important—to ensure sufficient exploration.
如果在词元空间中采样,其中 theta 是程序,它发出对自身的调用,而不是程序归纳输出一个程序,那么仅基于冒泡排序训练的条件下,归并排序轨迹的概率是多少?这就是生成器-验证器差距的用武之地。如果概率非零,就像一百万只猴子打字,其中一只发现了归并排序,很容易验证它更快。这就是人类文明发展的方式——有人发现了火,很容易验证发生了重要的事情。
If sampling in token space, where theta is the program and it emits calls to itself versus program induction outputting a program, what's the probability of the merge sort trace conditioned on only bubble sort training? This is where the generator-verifier gap comes in. If it's non-zero probability, like a million monkeys typing, one discovers merge sort, and it's easy to verify it's faster. That's how human civilization developed—someone figures out fire, it's easy to verify something important happened.
是的,如果对程序空间的搜索是随机的,那效率极低。但如果你有一个先验,并且它进行了巧妙的启发式搜索,学会了搜索程序空间,那么我完全同意你的观点。
Yeah, and if the search over program space is random, it's extremely inefficient. But if you have a prior and it does clever heuristic search, having learned to search program space, then I would completely agree with you.
但即便如此,每年赢得 IMO 的人总是有大约 10 种策略可以学习,或者国际象棋高手知道兵值 1 分,车值 5 分。
But even then, there are ways like the people that win IMO every year always have like 10 strategies that you learn, or the people that are really good at chess know pawn is worth one, rook is worth five.
我不认为 IMO 或国际象棋那么简单。他们尝试这些策略,这是达到至少 25 分的好方法。你有这 10 种策略,做这 10 件事。我同意,要完成剩下的距离会很愚蠢,而且我确定马格努斯·卡尔森会质疑:‘哦,你只需要知道兵值 1 分、车值 5 分就能打败我?’不。但你是在利用别人凝结的知识,然后运行这个算法。也许你的海马体比我的更强,你能做更多的 MCTS,或者它更快,能比我深入,或者你有更好的记忆力。这里有困难的部分,甚至可能涉及 DNA。也许是技能,也许是学习,也许是 DNA。但我的观点是,程序内搜索的启发式仍然严重依赖人类,因为我们用人类程序作为先验来搜索程序空间,而不是随机搜索。
I don't think it's quite that simple for IMO or chess. They try these strategies, it's a good way to get to at least 25. You have these 10 strategies, you do these 10 things. I agree that it would be a pretty silly thing to get the rest of the distance, and I'm sure Magnus Carlsen would take some pause with, 'Oh, all you need to know is that a pawn is worth one and a rook is worth five and then you can beat me.' No. But you're kind of bootstrapping off of someone else's congealed knowledge and then running this algorithm. Maybe your hippocampus is just stronger than mine, you can do a lot more MCTS, or it's faster and you can go deeper than me, or you have a better memory. There's this hard part, maybe it could be even DNA. Maybe it's skill, maybe it's learned, maybe it's DNA. But my point is that the heuristic of in-program search is still heavily informed by humans because we've trained it on human programs as the prior to search the program space versus random.
这很有道理。确实有助于优化那个先验,让它专注于有希望的方向,从而更快地发现有用的事物。
That makes a lot of sense. It definitely helps to sharpen that prior so it focuses on promising directions, you stumble upon useful things faster.
在我们结束这个话题之前,你认为当前的 LLM 设置在多大程度上是生物上合理的,哪些不是,你真的在乎吗?
Before we get off this one, how much do you think the current LLM setup is biologically plausible and what isn't, and do you really care?
你说的生物上合理是什么意思?
What do you mean by biologically plausible?
比如,大脑在做反向传播吗?预训练基本上就是 DNA 吗?人们争论的所有这些映射。
Like, is the brain doing back propagation? Is pre-training basically DNA? All these mappings that people would argue.
我不明白,我不是生物学家或神经学家,所以很难在这里做出断言。我认为深度神经网络方面,据我所知,与大脑的结构有一些相似之处。我不认为与人类大脑的接近程度非常重要,尤其是在这一点上,显然这里有一些有效的东西。我认为它有用,因为也许还有更有用的东西我们尚未发现。对此,我不是合适的人选来评论。
I don't understand, I'm not a biologist or neurologist, so it's a little hard for me to make claims here. I think the deep neural net aspects, for what I understand, have some resemblances to the way brains are structured. I don't think it's super important how close it is to the human brain, especially at this point, clearly there is something working here. I think it is useful in the sense that maybe there are even more useful things we haven't discovered. For that, I'm not the right person to say.
你博士期间的一大创新是,利用测试时算力,用一个可以在 CPU 上运行的相当简单的程序,实际上比许多在大量扑克游戏上训练的策略内 RL 算法更有意义。弗朗索瓦·肖莱说,我们最终会找到运行大脑的算法,它将是一个相当简单的程序,几乎没有先验。你同意还是不同意?
Part of your big innovation in your PhD was that leveraging test-time compute with a quite simple program that could run on a CPU was actually meaningfully better than a lot of the on-policy trained RL algorithms trained on massive amounts of poker games. François Chollet said that we will end up finding the algorithm that is running the brain and it will be a quite simple program with very little priors. Do you agree or disagree?
我同意它可能相当简洁优雅。我认为我们今天使用的许多算法都相当简单优雅。至于它没有太多先验的想法,我不确定是否同意。有可能,但我也认为有不同的解读方式。当很多人听到这个时,他们会想到像 AlphaZero 这样的东西。AlphaGo 取得了巨大成功,在大量人类数据上训练,进行了蒙特卡洛树搜索、自我对弈,并击败了围棋顶尖人类选手。然后 AlphaZero 去掉了人类数据的学习,即先验,只保留规则和算力,结果做得更好。但当在星际争霸这样的游戏中尝试时,效果并不好。
I agree that it would probably be quite simple in elegance. I think a lot of the algorithms we use today are quite simple and elegant. I think the idea that it doesn't have much of a prior, I don't know if I necessarily agree with that. It's possible, but I also think there are different ways to interpret it. When a lot of people hear that, they think about something like AlphaZero. AlphaGo had a big success, trained on a massive amount of human data, did Monte Carlo tree search, self-play, and beat top humans in Go. Then AlphaZero took out the learning from human data, the prior, just rules and compute, and it ended up doing a lot better. But when that was attempted in something like StarCraft, it didn't work out very well.
你认为这与动作空间有关吗?
And you think that has to do with action space?
我认为是的。原则上,如果你用足够大的网络进行足够的强化学习,并且持续足够长的时间,你可以完全不用先验。但事实证明这不太实际。到目前为止,先验对于大型语言模型之类的东西极其重要。有没有可能有一天我们完全摆脱先验,只通过强化学习从头学习一切?我认为理论上可行。但我不认为这很可能,也不认为在可预见的未来这是正确的道路。
I think so. In principle, if you did enough RL with a big enough network and did it for long enough, you could get away with just no prior. But it turns out that's not very practical. So far that prior has been extremely important for things like large language models. Is it possible that one day we move away from the prior completely and have everything learned from scratch with just RL? I think it's in theory plausible. I don't think it's likely and I don't think it's the right path to pursue for the foreseeable future.
你为什么这么说?
Why do you say that?
证据相当有力。对于该领域的许多人来说,很长一段时间里,他们专注于去除人类先验、从头学习。从 2017 年到 2021 年,这曾是强化学习的主导范式。但越来越清楚的是,能够在大型互联网数据上预训练并构建有用的先验非常有效,结果很难反驳。情况可能会改变,但目前证据强烈支持相反的方向。
The evidence is pretty strong. For a lot of people in the field, for a very long time, they were focused on this idea of removing the human prior and learning from scratch. For years, from 2017 through 2021, this was the dominant paradigm in RL. And it just became clearer that being able to pre-train on large-scale internet data and build a useful prior was extremely effective and pretty hard to argue with the results. It's possible that changes, but the evidence is very strong in favor of the other direction right now.
当然,你只是在购买任意水平的技能,但智力的真正衡量标准是技能获取的速度,即技能获取效率。通过提升先验,你基本上越来越接近基准测试作弊。目前尚不清楚构建这个有用的先验是否会损害其在线学习的能力。这不一定是一个权衡。我的意思是,也许有,但如果是这样,我想看到一些证据。
Surely you're just buying arbitrary levels of skill, but the real measure of intelligence is the rate of skill acquisition, the skill acquisition efficiency. By boosting up my prior, you're basically getting closer and closer to benchmark hacking. It's not clear that building up this useful prior is hurting its ability to learn online. It's not like you're having a trade-off there necessarily. I mean, maybe you are, but I would want to see some evidence if that's the case.
这是个公平的观点。我想持续学习的东西,反向传播本身已经证明了一点:无论你在什么数据上训练,随着迭代的改进或继续,神经可塑性会降低。我也经历过这一点。训练得越久,越多的维度变得退化。如果你在训练后期随机采样方向,与训练早期相比,每个方向对损失完全没有影响。
That's a fair point. I guess the continual learning stuff, backprop itself has something kind of proved out: the more you train it, no matter what you're training it on, reduces neuroplasticity as iterations improve or continue. I've experienced this too. The further you train, the more dimensions become degenerate. If you just sample random directions late in training versus early in training, every direction doesn't impact loss at all.
所以这有点像是一个关于神经可塑性降低的论点。但不管怎样,我想确保我们能把剩下的部分讲完。那么关于 AI 的反主流观点。我确定我们之前聊过一点。你今天对 AI 有什么反主流观点吗?
So that's kind of like an argument that it's reduced neuroplasticity. But anyway, I want to make sure we get through the rest of these. So the contrarian view on AI. I'm sure we've talked about this a little bit. Do you have a contrarian view on AI today?
是的,我会说我坚信模型会越来越好。智能会非常快速地持续提升。我确实认为这个领域低估的一件事是推理算力的重要性,尤其是当我们发布更强大的模型时。我的意思是,当我们发布一个新模型时,我们用单个数字在不同的评估上评价它,比如 GPQA,一个数字告诉你这个模型在那个基准上有多聪明。我认为这不再有意义了。这在 GPT-2 时是对的,在 GPT-3 时也有意义,在 GPT-4 时勉强有意义,但即使在那里也有点牵强。我认为自从你可以给一个提示思维链并获得更好的性能以来,这就不再正确了。一旦我们有了推理模型,这显然就不对了。现在我认为我们还在这样做有点傻。但我觉得人们只是期望这样,所以才会这样做。我认为 ARC 实际上在快速摆脱这一点方面做得很好,用推理算力或成本的 x 轴来衡量事物。我认为这是衡量任何基准的正确方式,尤其是推理密集型的。当你开始看像准备框架或负责任扩展政策这样的东西时,这实际上非常重要,因为有一些阈值来决定,比如,如果我们发布一个模型,它的能力是什么,不同的阈值在哪里?当你评估发布的能力并决定它是否危险或超过某个能力水平时,你在那个评估中投入了多少推理?如果你处于智能纯粹是推理函数的阶段,那么你就有任意数量的智能。所以你应该在评估上花无限的钱。这是一个真正的问题,因为你可以发布一个模型并说,我们把它限制在 10 美元的推理上,但如果有人能轻松地把一堆查询组合在一起,花 1000 美元的推理,那么他们实际上就拥有了一个比你发布的更智能的模型。如果他们组合了一百万美元的推理,他们就拥有了一个能力更强的模型。这是否是同一个模型还有争议。显然这在 GPT-2 或 GPT-3 时不是问题,因为你不能把一千个 GPT-3 查询组合起来得到明显更智能的东西。那可能有点成立,但并非真正如此。随着模型变得越来越强大,这越来越成为一个问题。
Yeah, I would say I'm a strong believer that the models are going to keep getting better. Intelligence is going to keep improving pretty rapidly. I do think that one of the things the field is underestimating is the significance of inference compute, especially as we release more capable models. I mean, this idea that when we release a new model, we evaluate it with a single number on different evals, like GPQA, a single number telling you how smart this model is on that benchmark. I don't think this makes sense anymore. It was true with GPT-2, made sense with GPT-3, kind of made sense with GPT-4, but even there it was a bit iffy. I don't think it's been true ever since you give a prompt chain of thought and get better performance. It was clearly not true once we had reasoning models. And now I just think it's kind of silly that we're still doing this. But I think people just expect this, so that's why it's being done. I think ARC has actually been quite good about moving away from this quickly and measuring things with an x-axis of inference compute or cost. And I think that's the right way to measure any benchmark, especially reasoning-heavy ones. This actually matters a lot when you start looking at things like preparedness frameworks or responsible scaling policies, because there are thresholds to decide, like, if we release a model, what are its capabilities, and where do the different thresholds lie? When you're evaluating capability for release and deciding whether it is dangerous or exceeds a level of capability, how much inference do you put into that evaluation? If you're at a point where intelligence is purely a function of inference, then you have arbitrary amounts of intelligence. So you should just spend infinite money on your eval. This is a real problem because you could release a model and say, we cap it at $10 of inference, but if somebody can easily scaffold a bunch of queries together and spend $1,000 of inference, then they effectively have a more intelligent model than what you released. If they scaffold together a million dollars of inference, they have a much more capable model. It's debatable whether it's the same model. Clearly this wasn't an issue with GPT-2 or GPT-3 because you can't scaffold together a thousand GPT-3 queries and get something substantially more intelligent. That was maybe slightly the case but not really. This is increasingly a problem as models become more capable.
这真的很有趣。我几乎觉得当前的趋势是扩展脚手架。Claude Code 就在做这个,注入越来越多的脚手架。随着模型变得更好,你需要的系统提示和脚手架会更少。但我实际上想转向另一个方向。我写了一篇博文,说下一个 Scaling 定律将是训练时递归。虽然你有测试时递归,更多的测试时算力,如果我们试图构建一个图灵机,一个要求是测试时的无界递归。它们在测试时确实有,但在训练时没有。训练时是一次前向传播和教师强制。也许当你让它自由运行时,这会被打破,然后从测试时发现的潜在空间中进行一些蒸馏。从每个参数的角度来看,我们在 ARC-2 上看到的最大进步来自训练时递归,比如 HRM TRM,表明外部精炼循环是主要贡献者。你对未来的训练时递归有什么看法?
That's really interesting. I almost feel like the current trend is scaling scaffolding. Claude Code is doing that, injecting more and more scaffolding. As models get better, you'll get away with less system prompt and less scaffolding. But I actually want to go into another direction. I wrote a blog post about the next scaling law being train time recurrence. While you have test time recurrence, more test time compute, if we're trying to build a Turing machine, one requirement is unbounded recurrence at test time. They do have that at test time, but not at train time. It's one forward pass and teacher forced. Maybe this is broken when you let it go and then have some distillation back from the latent space discovered at test time. We've seen the most progress on ARC-2 from a per-parameter standpoint with train time recurrence, like HRM TRM, showing that outer refinement loop was the main contributor. What are your thoughts on train time recurrence in the future?
嗯,我认为在训练时投入更多算力的想法——预训练和后训练之间有区别。对于后训练,至少有一些算力投入。但听起来你这里更多是在说预训练。我认为这是一个值得追求的好方向。仅仅因为我们今天有能用的东西,并不意味着它是最好的方式,当然也不是唯一的方式。不同的研究方向被探索是有价值的。完全有可能,甚至很可能,我们会想出比今天使用的更有效的东西。
Well, I think the idea of spending more compute during training time—there's a distinction between pre-training and post-training. For post-training, there is at least some compute spent. But it sounds like you're talking more about pre-training here. I think it's a great direction to pursue. Just because we have something that works today doesn't mean it's the best way, and certainly not the only way. It's valuable that different research directions are being pursued. It's entirely possible, if anything likely, that we will come up with something more effective than what we're using today.
你特别看好当前股票中什么会消失?
What particularly are you most bullish will go away in current stock?
我不是做预训练的人,所以我可能不是问这个问题的合适人选。而且,即使我是,我可能也没法说。
I'm not a pre-training person, so I'm probably not the right person to ask. And also, if I were, I probably wouldn't be able to say anyway.
好的。沿着这条线,也许这是合适的过渡:那个研究。我感觉我 2012 到 2014 年在李飞飞的实验室开始。
Okay. Along that line, maybe this is the right segue: that research. I feel like I started in Fei-Fei's lab in 2012 to 2014.
然后现在我回来了,LLMs 吸走了房间里所有的空气,而之前那些分散的不同研究想法已经不再被追求了。所以现在有了这个自动研究员,我有四个不同的 Karpathy 自动研究员在运行,它们正在写我自己的 NeurIPS 论文之类的,我们聊过这个。嗯,你认为学术界的作用是什么,会议的作用,提交论文的重要性,经过审稿人——不管是他们的智能体还是高中生用他们的智能体,甚至更糟。嗯,你对学术界的未来、会议的未来、获得博士学位的重要性有什么看法?
And then now I'm back and the amount of air that LLMs have sucked out of the room versus the diaspora of like different research ideas that were being pursued is just not really happening anymore. And so like and then now with this auto researcher, I have like four different Karpathy auto researchers going on right now that are doing writing my own NeurIPS paper and things like that we talked about. Um what do you think the role of academia is, the role of conferences, the importance of submitting, um going through the reviewers where it's their agent or it's a high schooler with their agent even worse. Um what are your thoughts on the future of the future of academia, the future of conferences, the importance of getting PhDs?
是的,这是个好问题。过去几年我经常被问到这个问题,嗯,我确实有一些想法。我的意思是,这并不像一些研究人员说的那么绝望。我认为学术界实际上还有可行的事情可做。我先谈谈问题,其中一个问题是,很多最令人印象深刻的 AI 能力都是 Scaling(规模扩张)的结果。这在学术界是个问题,因为学术界没有那么多 GPU。这就是现实。我认为情况可能会有所不同。理论物理的例子很好,比如大型强子对撞机。据我所知,那不是私人企业。嗯,所以大学可以获得大量算力,并用它来做学术研究。嗯,如果我负责一所大学,我会花十亿美元建一个大型计算集群,老实说,如果你是一所不在 AI 计算机科学前十名的大学,想快速进入前十甚至成为第一,你就花大笔资金买 GPU,然后去找每个明星教授、每个明星学生说:“嘿,我们每个研究人员的 GPU 比地球上任何其他地方、任何其他学术机构都多。”是的,你会很快吸引很多优秀人才。
Yeah, it's a good question. I've been asked this question a lot over the past couple years and I it's um I do have some thoughts on this. I mean it's not as hopeless as some researchers might say. I think there is actually viable stuff to be done in academia. I'll I'll first talk about the problems, which is one a lot of the most impressive AI capabilities have been a consequence of scale. And that is a problem in academia because there's just not as many GPUs in academia. And that's that's just a reality. I think it could be a bit different. I think the example of theoretical physics is a good example where um there is like the Large Hadron Collider. That's not a private enterprise as far as I know. Um and so universities could get large amounts of compute and use that to do academic research. Um if I was in charge of a university, I would spend a billion dollars to get a large computing cluster and and honestly, if you are a university that's like not in the top 10 for computer science for AI and you want to quickly get into the top 10 or be even be number one, you spend a lot of capital to get GPUs and then you go to like every star professor, every star student and say, "Hey, we have more GPUs per researcher than any other place on Earth, any other academic institution on Earth." Yeah. You're going to get a lot of good talent very quickly.
是的,我发过的最著名的推文是今年 NeurIPS 之后,我们和一群教授吃了晚饭,问他们学校的算力怎么样。我回家后,抓取了 H100 或等效 GPU 的数量,按美元加权,除以 CS 学生人数,然后发了图表。结果除了 MIT 和哈佛,所有人都远低于 1。这很糟糕。是的。
Yeah, my most famous tweet that I've ever posted was after at NeurIPS this year, I just like we had a dinner with a bunch of professors and we just asked them about like what's compute like at your school. And I just went home and I just like grabbed um you know, number of H100 or equivalent um you know, weighted by dollar uh whatever divided by uh number of CS students and then just like posted the the chart and it's like meaningfully below one for everyone except and then I waited a little bit more for do undergrads really need 1 H100? Um and it's basically like besides MIT and Harvard, everyone is below one. It's pretty rough. Yeah.
这很糟糕,我和教职员工聊过,我随口提到我随时可以启动任务的大概 GPU 数量,他们简直要炸了。我觉得很多学术界的人并不理解差距有多大。
It's it's pretty rough and I've talked to faculty and I would casually mention just like you know, roughly how many GPUs like, you know, I have access to whenever I want to spin up a job and they they just their head explodes. Like it's really I think a lot of people in academia do not understand how much how much of a difference it is.
对。嗯,目前学术界和工业界之间。我完全同意。嗯,幸运的是我自己在资助自己,但也要感谢佛罗里达大学,没人知道他们,但他们有整个叫 Gator 什么或 Gator Cloud 的东西。大量的 GPU,还有 UT Austin 也有大量算力。所以,如果你是个博士生,我建议你去那里,老实说这两个之一。斯坦福实际上很糟糕,尽管有几百亿的捐赠基金,我和我们的教授聊过很多。我确实认为学生在决定去哪里时应该问一个重要问题:我能获得多少算力?而且可能不要只问教授,还要问学生他们有多少 GPU 可用,因为那样你会得到更诚实的答案。我相信招聘人员会问你。比如,如果我来 OpenAI 工作,我能得到多少算力?这难道不是选择工作地点的重要因素吗?
Right. Um between academia and industry right now. I completely agree. Um and I'm funding my own uh fortunately, but the co- shout out to University of Florida who no one knows, but they have this whole whole thing called Gator something or Gator Cloud or something like that. It's huge amount of GPUs and uh UT Austin huge amounts of compute. So, if you're a PhD student, I would actually look to go there, one of those two to be honest. And Stanford is actually like really bad despite a tens of billions of dollar endowment um and something I have I've talked to our professors about a lot. I I do think it's an important question for students to be asking uh when they're deciding where to go. Like how many how many how much compute am I going to have access to? And to probably not just ask the professors, but also ask the students about how many GPUs they have access to cuz you'll get a more honest answer that way. I'm sure recruits ask you. Like if if I'm coming to work at OpenAI, like how much compute am I getting? And like is that not an important factor in choosing which where to work?
嗯,我想是的。我的意思是基本上想知道公司有多少算力,以及整体研究情况。嗯,但我确实认为这是一个应该问的重要问题。我不常被问到,因为人们很确信 OpenAI 有大量算力,我认为这是正确的评估。是的。但人们可能应该问这个问题。
Um I I think so. I mean basically want to know how much compute I think the uh company has and for for research as a whole. Uh, but I I do think it's an important question to be asking. I don't get asked that very often because I think people are pretty confident that OpenAI has a ton of compute, which I think is the correct assessment. Yeah. But it is a question that people probably should be asking.
你认为学术界获得算力的解决方案是什么?因为显然,我们不会得到像粒子对撞机那样级别的资金投入数据中心。我的意思是,我们等了斯坦福的 Marlo,他们终于发布了。那是,你知道,几百个 GPU 给整个斯坦福 CS。是的,我在这里没有很好的答案。我的意思是,一个办法是大学可以联合起来,合作做更大规模的实验。嗯,我觉得看到学术界做一个纯粹开源的预训练努力会很酷,比如,嗯,那实际上会……
What do you think the solve is for, um, academia to get the compute? Because clearly the the the, you know, we're not getting particle colliders like, you know, levels of of money into data centers. I mean, we waited for Marlo at Stanford and then they finally got it got it released. It was like, you know, hundreds of GPUs for all of Stanford CS. Yeah, I don't have a great answer here. I mean, one is, um, universities can kind of pull together, uh, work together to, um, do larger scale experiments. Uh, I would I would I think it'd be really cool to see like an academic opens like purely open source pre-training effort, for example, um, that would actually Yeah.
它实际上会有竞争力,嗯,我认为挑战在于,好吧,你怎么做?因为现在 AI 研究的文化是两作者论文,也许五作者论文,如果你要做这样一个大规模项目,你如何在文化上确保每个人的贡献得到适当归属?对。这是一个非常困难的文化调整。嗯。所以这是一个挑战,但我想说清楚,这不是在学术界产生影响的唯一方式。我认为如果你想做大规模实验,是的,这是一个真正的挑战。有些事不需要大量算力也能做。嗯,我认为评估(Evals)是一个我一直指出的例子,老实说,在 OpenAI 这样的地方,我们从第三方评估中获得了很大价值。比如,你们很棒。我猜你们用了一些算力来制作我们的 KGI 3,但可能不是很大。嗯,我的意思是,我想问你,你认为像我们的 KGI 3 这样的东西能在学术机构中制作出来吗?
it's actually be competitive and um, I I think the challenge is like, okay, well, how do you cuz the the culture for AI research right now is very much two author papers, maybe five author papers and if you're going to do a large scale project like this, how do you culturally, um, make sure that everybody's properly attributed for their contribution? Right. It's a very difficult cultural adjustment to make. Mhm. And so that's one challenge, but it's not this isn't the only way to have an impact in academia, I want to be clear. I think this is if you wanted to be able to do large scale experiments like yes, I think this is a real challenge. There are things you could do without large amounts of compute. Um, I think Evals is a consistent one that I point to where honestly, we get a at places like OpenAI, we get a lot of value from third party Evals. Like, you know, you guys are great. And I'm I'm guessing you used some amount of compute to make our KGI 3, but it's probably not a huge amount. Um, it's not like I mean, I I was going to ask you like, do you think that something like our KGI 3 could have been made in an academic institution?
嗯,我的意思是,第一个,是的。第二和第三,不行。第二和第三需要更多的细节关注、资金和规模化,而且没有人靠写维护代码获得博士学位,所以你会想,我就写下一个版本吧。我认为很多人低估了评估的影响。我认为 Dan Hendricks 是一个很好的例子,他通过制作高质量、非常好的评估而名声大噪。
Um, I mean, one, yeah. Two and three, no. Two and three require a lot more um, attention to detail and money and scaling it up and also no one really gets a PhD by writing maintenance code and so like you kind of like I'm just going to write the next version of this. I think a lot of people underestimate the impact of Evals. I think Dan Hendricks is a great example where he really made a name for himself by making really high quality really good Evals.
这是一种非常可行的产生影响力的方式。比如,我们实验室就产出了 LegalBench,还有其他几个类似的。我觉得这很难。说实话,我认为 ARC 是最高质量的基准测试之一,一个博士生甚至一群有经费的博士生都很难做出 ARC 那样的成果。我认为 Karpathy 和 Chollet 投入了非凡的努力、才华、智慧和金钱,这在学术界是做不到的。事后看来,你可能会说,‘哦,我们可以做一堆游戏,确保它们在某种程度上是正交的。’但学术界就是不奖励这种协同努力。
And that is a really viable way to have impact. Like out of our lab came LegalBench, for example, and there were a couple others like that. I think it's just hard. For really, I think ARC is one of the highest quality benchmarks that a single PhD student or even a group of them with funding would still struggle to do what ARC did. I think it was a remarkable amount of effort, talent, intelligence, and money that Karpathy and Chollet put into this that I don't think could have been done in academia. In hindsight, maybe you can say, 'Oh, we can make a bunch of games and make sure they're orthogonal in some way.' But academia just doesn't reward that type of concerted effort.
我认为这是个问题。我同意学术界不奖励这种努力,希望这会改变,或者将会改变。所以,是的,当然学术界也产出了很多优秀的评估,我们 OpenAI 也会关注。我想你还有一个关于会议角色的问题,我觉得看看这会如何演变会很有趣。现实是,很多论文现在是由 AI 写的,很多论文也是由 AI 评审的。我不一定认为这是坏事。我和学术界的人聊过,他们基本上说,‘看,AI 评审的论文,AI 的评审并不那么好,但可能实际上比现在人类的平均评审要好。’
I think this is an issue. I agree that there's an issue where academia does not reward this kind of effort, and I think this hopefully should change or hopefully will change. So, but yeah, I mean certainly there are great evals that come out of academia and evals that we at OpenAI pay attention to. I think there's another question you had about the role of conferences, and I think that's going to be very interesting to see how that evolves. The reality is a lot of papers are currently being written by AI. A lot of papers are currently being reviewed by AI. And I don't necessarily see this as a bad thing. I mean, I've been talking to academics about this, and they basically say, 'Look, yeah, the AI-reviewed paper, the AI reviews are not that great, but they might actually be better than the average human review right now.'
完全同意。我完全同意。这就像大规模的 actor-critic。科学界正在发生的最大的 actor-critic 事情。是的,有相关人员参与其中,确保一切正常。但我同意,以前评审者的方差太大了,所以这很可能是一个净正面效应。
Completely agree. I completely agree. It's like actor-critic in mass. The biggest actor-critic thing happening in science. And like, yeah, there are people in the loop on it that are somewhat making sure everything is copacetic. But I agree that the variance in reviewers was so bad before that it's probably a net positive.
是的,我认为你想把它们与人类配对,但我认为让 AI 评审每篇论文,指出是否有致命缺陷,然后由人类验证,这实际上是个好主意。我认为我们会看到这个趋势。我的意思是,现在,AI 评审,我不知道现在的情况,但六个月前,AI 评审还不太好,但 AI 进步非常快。我认为它们实际上能够,可能最新的模型,我预计会做得很好。我同意。特别是如果你让它做文献综述,它们可以查看参考文献,看看那些论文,我预计它们会做得很好。如果现在不是,我估计到今年年底,它们会做得很好。
Yeah, I mean, I think you want to pair them with humans, but I think it's actually a pretty good idea to have an AI review every paper and just point out if there is some kind of fatal flaw and have that verified by a human. And I think we'll see this trend. I mean, right now, yeah, the AI reviews, I don't even know about right now, but like 6 months ago, the AI reviews were not that good, but the AIs are progressing quite rapidly. And I think they would actually be able to give, probably the latest models, I would expect to actually do a very good job of reviewing a paper. I agree. Especially if you point it to, you know, and they can do a literature review. So they can look at the references and just look at those papers, and I would expect that they would do a pretty good job. If they're not currently, I would say by the end of the year, like, yeah, I think they'd be doing a great job.
现在标准做法是把论文提交给 AI,所有大模型,让它们给你反馈。就像自动评审。你不如用它来挤出论文中的小问题。是的。老实说,我认为用不了多久 AI 就会端到端地写论文了。是的。它们可能已经在写相当大的一部分了,但我认为没有障碍阻止它们首先端到端地写论文,然后做所有实验等等。这需要时间。工作量很大。人类写一篇好的会议论文需要多久?可能至少几个月。但我认为我们会达到那个地步。
It's standard practice now to submit it to your AIs, to all of the big AI models, to just give you back feedback. It's like an automatic review. You might as well do it to squeeze out little issues in the paper. Yep. And to be honest, I don't think it's going to be that much longer until the AIs are just writing the papers end-to-end. Yeah. They're probably already doing pretty decent chunks, but I don't see a barrier to them just first of all, writing the papers end-to-end, but then also doing all the experiments and everything. And it's going to take a while. It's a lot of work. How long does it take a human to write a good conference paper? Probably a couple months at least. But I think we'll get there.
我付了 Overleaf 的积分,所以还有两个问题。我们谈到了地球上这个 OAI 形状的洞,现在当你们停止发布权重和开源很多东西时,有一些回归的迹象。你对 DeepSeek V3 普遍成为事实上的、中国模型普遍成为世界事实上的开源权重有什么看法?我的意思是,这不是我的专业领域,也不是我能决定的。所以我可能没有太多要补充的,但不管怎样,我确实认为开源很有价值。我认为拥有好的开源模型很重要。
I pay for Overleaf credits, so there are two more questions. We talked about this OAI-shaped hole in the planet, and now when you guys stopped releasing weights and open-sourcing a bunch, there was some glimmer of getting back to that. What are your thoughts on generally high level of DeepSeek V3 being the de facto and Chinese models in general being de facto open source weights for the world? I mean, it's not really my area of expertise or my call to make on these sorts of things. So I probably don't have too much to add to this, but for what it's worth, I do think there's a lot of value in open source. I think it is important that we have good open source models out there.
我希望这个洞能被填补,无论是我们还是其他人。我同意。
And I hope that hole is filled, either by us or by somebody else. I agree.
最后一个问题。你在过去十年中经历了巨大的成长,见识了很多。在 AI 之外,过去十年中你经历的最艰难、自我成长最多的时刻是什么?
Last question. You have gone through a tremendous amount of growth over the last decade, and you've seen so much. Outside of AI, what has been the hardest time, most amount of self-growth that you had to go through in the past decade?
是的,好吧,这不在 AI 之外,但我要说 2017 年我参加的 Libratus 比赛是我生命中一个非常关键的时刻。我当时是一个相当不知名的研究生,我很清楚如果我在这个比赛中成功,那将是一件大事和巨大的成就。我也知道这需要大量的努力。我想基本上在 2016 年初,我就很清楚成功的道路是什么,我决定花一年时间非常努力地执行。那一年我投入了很多时间,基本上不停地工作,制作一个非常强大的扑克机器人。让你的整个职业生涯取决于一场扑克游戏,这压力非常大。因为你不知道你的机器人是否好,因为扑克是一个方差极高的游戏。仅仅通过看牌局很难判断它做得好还是不好。我们没有预算雇佣一群人类扑克玩家来对抗它,进行压力测试。
Yeah, well, it wasn't outside of AI, but I would say the Libratus competition that I did in 2017 was a pretty pivotal moment in my life. I was a pretty unknown grad student at the time, and it was very clear to me that if I was successful in this competition, it would be a pretty big deal and a big achievement. And I also knew that it would take a ton of hard work. I think basically in early 2016, it became very clear to me what the path to success was going to be, and I decided to spend a year executing on it super hard. That year I put a lot of hours in, basically working non-stop to make a really strong poker bot. And it was really stressful to have your entire career come down to a poker game. Because you don't know if your bot is any good, since poker is a super high variance game. It's really hard to tell just by looking at the hands whether it's doing a good job or a bad job. We didn't have the budget to hire a bunch of human poker players to play against to try to stress test it.
对。AlphaGo 有。
Right. AlphaGo did.
是的,我们和以前的机器人对战,我们有一种感觉,首先方差非常高。所以我们甚至对胜率没有很好的把握。但我们感觉它对以前的机器人表现不错。问题是真正的挑战是人类非常适应。所以我不知道一旦我们和真正的人类对战,他们能否适应,找到机器人的弱点,发现漏洞。所以那是一个非常紧张的时期,幸运的是最终成功了。
Yeah, we played against previous bots and we had a sense, first of all the variance is super high. So we didn't even have a good sense of the win rate. But we had some sense that it was doing well against the previous bots. The problem is that the real challenge was that the humans were very adaptive. And so I didn't know if once we play this against actual humans, they are able to adapt, the bot's weaknesses, will they find holes. So that was a very stressful period, and fortunately it ended up being successful.
你赢了之后是怎么庆祝的?
How did you celebrate when you won?
我大概花了几周时间才冷静下来,意识到‘哦,真的结束了,我现在可以放松了。’我不知道我是否真的庆祝了。
It took me probably a few weeks to kind of calm down and realize, 'Oh, it's actually over, I can relax now.' I don't know if I actually celebrated.
你去阿鲁巴之类的地方,休息一下?
You go to Aruba or something, get time?
我想我实际上没有庆祝。
I don't think I actually celebrated.
我觉得直到最后一刻我都很难相信我们真的不会突然输掉。但事后有人告诉我,我本可以付出 90%的努力却得到 0%的回报,我觉得这是个很好的人生教训——有时候你确实得拼命努力才能把事情做成。
I think it was really just like I also had trouble believing it until the very end that we actually were not going to lose this suddenly. But somebody did tell me afterwards that I could have put in 90% of the effort and gotten 0% of the reward, and I thought that was a good life lesson that sometimes you just have to work really hard to get something done.
太棒了。非常感谢你,这次访谈非常精彩。
Awesome. Well thank you so much. This was amazing.
嗯,谢谢邀请我。
Yeah. Thanks for having me.
好的,谢谢。酷。
All right. Thanks. Cool.