AI:好得难以置信,差得毫无用处

AI: Too Good to Be True, Too Bad to Be Useful

迪奥戈·阿尔梅达 Diogo Almeida · AI 理事会 · 2026-06-19 · 约 37 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

前 OpenAI 研究员 Diogo 探讨大语言模型过度承诺与交付不足之间的巨大鸿沟,追问为何 AI 在某些任务上表现出色却在其他任务上失灵。

Former OpenAI researcher Diogo explores the massive divide between LLM over-promise and under-delivery, asking why AI excels at some tasks yet fails at others.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

介绍与讲者背景 Introduction and Speaker Background

Diogo

严格来说,我是一家叫 Typeface 的公司的 CEO。这次演讲的绝大部分内容,我会以技术爱好者 Diogo 的身份来讲,而不是 CEO Diogo。演讲的标题是“AI 好得不像真的,差得没啥用”。我考虑过的另一个标题是“AI 到底怎么回事,过度承诺却交付不足?”这其实是一个关于优化的深刻隐喻,但我相信现在 AI 有一个大问题,就是 LLM 的过度承诺和交付不足之间的巨大鸿沟。我很乐意就此辩论。请留到问答环节。这是我最喜欢谈论的话题,如果这里不合适,我会在那个房间和人辩论到天荒地老。

Technically, I am CEO of a company called Typeface. Vast majority of this talk will be as technology lover Diogo, not CEO Diogo. The title of this is AI too good to be true, too bad to be useful. Another topic that I was considering was what's the deal with AI over-promise under-deliver? This is a surprisingly deep metaphor for optimization, but I believe that we have a major problem in AI right now, which is the massive divide between LLM over-promise and LLM under-deliver. Very happy to debate that. Please save it for questions later. This is my favorite topic ever to talk about, and if it doesn't fit, I will be in that room and debate people until the cows come home.

Diogo

好的。

Yeah.

Diogo

我是谁?你们在这儿,可能已经知道了。我在 OpenAI 工作了大约四年半。我是所有重大成果的共同作者。这是我在 Anthropic 接管后的第一次演讲,所以可以说我经历了它的黄金时代,因为现在有点像白银时代,但我是 GPT-4、InstructGPT、ChatGPT 的共同作者,最相关的是 RLHF。我现在不深入讲这个。这将是一个非常技术性的演讲。抱歉。我独特的地方,尤其是从 OpenAI 的角度来看,是 OpenAI 的人很少会批评 ChatGPT。文化上,这其实很奇怪。我并不讨厌 ChatGPT。我基本上每天都用,我只是觉得它造成了很多问题,尽管它是一个现象级的产品。优秀的产品,但它不在我们的 AI 梦想之路上。它没有引领基于 AI 的经济革命,我认为这是早期人们真正想要的。

Who am I? You guys are here, so you probably know that already. I was at OpenAI for about four and a half years. I was co-author to all of its greatest hits. This is my first talk since the Anthropic takeover, so you could say I was there for its golden era because it's kind of like silver right now, but co-author of GPT-4, InstructGPT, ChatGPT, and most relevantly of all, RLHF. I'm not going to get into that right now. This will be a very technical talk. I'm sorry. And what's kind of unique about me, especially from an OpenAI perspective, is it's very rare that people at OpenAI hate on ChatGPT. Culturally, it's actually like this very weird. And I don't hate ChatGPT. I use it basically every day, and I just think that it has caused a lot of problems that I think it's just really worth talking about despite being a phenomenal product. Excellent product, not on the path of It is not our AI dream. And, you know, this is it's not leading to an AI-based economic revolution, which is I think what the early people really wanted.

LLM叙事的鸿沟 The Divide in LLM Narratives

Diogo

我将讨论分解的部分。它们可能有点随机,但对我来说这是最重要的部分。请留在这个部分,试着从中学习。这是我认为每个人都应该知道的事情,即在 LLM 领域,两种叙事之间存在巨大的分歧。你知道,到底怎么回事?我们是像 EA 认为的那样,走在指数级机器神的道路上,还是像经济学家认为的那样,全是炒作,还有奇怪的循环融资等等。我相信每个 AI 从业者都应该知道答案。如果你们愿意,我甚至可以给你们时间大声回答。让我们即兴发挥。但我的问题一直是:经济革命在哪里?如果你看基准测试,我们一直在不停地饱和每个基准。模型实际上一直在变得更好,看起来我们处于指数增长。这还没有元图,但我喜欢这个,因为它显示了 GPQA。模型在 GPQA(谷歌防作弊问答)上超越人类水平,是导致 OpenAI 政变的原因。人们当时想,“天哪,我们是不是解决了所有问题?这太危险了。让我们赶走 Sam。”无论你是否知道,无论你站在 Sam 争论的哪一边,很明显那个模型并不太危险。你知道,那比 01 还差,今天人们根本不在乎。你知道,人们根本不在乎比那更好的模型。然后如果你看另一面,是人们在说 AI 是一个泡沫。基本上没有经济影响,这实际上在技术上也成立。那么,AI 领域到底发生了什么?哦,对了,还有一件事是目标已经真正改变了。即使那些曾经是真正信徒的人,认为这是变革性技术,现在也说,“呃,它会不同。它更像是复兴而不是革命”,他们把目标从自动化大多数有经济价值的工作,转移到要么赚取或拥有 1000 亿美元的收入或利润。这到底是怎么回事?AI 很棒,实际上相信主流叙事,即它只在这么多事情上这么好,实际上对 AI 的潜力来说是非常令人失望的。我稍后会说服你。但我们应该问的问题是,这是每个人都应该知道答案的问题。请想想你的答案。你怎么能有两个故事:它在很多事情上太好了,不像是真的,新闻里到处都是,但基本上对其他所有事情都太差了,没什么用。这些任务,你知道,右边的看起来并不比左边的难。事实上,看起来相反。如果你想理解 AI,你至少需要一个奥卡姆剃刀式的解释,为什么会这样。如果你想构建产品,你应该知道你在哪一边。如果你投资公司,知道你在哪一边。如果你想在公司工作,你是在 vaporware 领域,还是在太好了,不像是真的,这会很棒领域。你知道,我们应该有一个简单的解释。我将过度简化,直接跳到我的答案。实际上,我将快速过一遍整个演讲,因为我想进入问答环节。上次看起来很有趣。看起来超级有趣,所以我就快速讲很多我本来会慢讲的东西。

Going to talk the broken down sections. They're going to be kind of random, but this to me is the most important part. Please stay for this section and try to learn from this one. This is the thing that I think everyone should know, which is there are like such a wide divide in LLM land between like two narratives. You know, like what the hell is going on? Like are we in like the path of the exponential machine God like the EAs think or are we like all hype like the economists think and like there's like weird ass circular financing whatever else. And I believe that everyone in AI should know an answer to this. And if we guys want, I can even give you times to try to answer things out loud. Let's improvise this talk. But my question has always been where is the economic revolution? If you look at benchmarks, we've just been like saturating each benchmark non-stop. The models have actually been getting better and it kind of looks like we're on an exponential. This doesn't have the meta plots yet, but I like this one because it shows GPQA. Models surpassing human-level performance at GPQA, Google-proof question answering, is what caused the OpenAI coup. People were like, "Holy did we just like solve everything? This is too dangerous. Let's kick out Sam." And like whether or not you know, you whether I doesn't matter like which side of that you want to you you stand in that argument on Sam, it's clear that that model was not too dangerous. You know, that was worse than 01, which people today do not give a about. You know, people don't give a about way better models than that. Then if you look at the other flip side, it's people saying like AI is a bubble. There's basically no economic impact and like that is actually also technically true. So, what is actually happening in the field of AI? Oh oh yeah, and like one other thing is that the goal posts have really shifted as well. Like even the people who used to be like the true believers of like this is a transformative technology are now like, "Eh, it's going to be a different. It's going to be more renaissance than revolution, and they've moved the goalposts from automating most economically valuable work to either making or having a revenue or profit of a hundred billion dollars. And like, what the hell? That like AI is awesome, and actually believing the mainstream narrative of it's only so good at so many things is actually extremely underwhelming to the potential of AI. I will I will convince you of this later. So, but the question we should ask, and this is the question that everyone should know an answer to. Please think of your answer. Is how can you have the two stories of it's two being too good to be true at a whole bunch of stuff, everything in the news, but like too bad to be useful for like basically everything else. And these tasks, you know, it doesn't look like the stuff on the right is any harder than the stuff on the left. In fact, it looks like the opposite. And if you want to understand AI, you need at least an Occam's razor explanation of like why is this the case. If you're wanting to like build a product, you should know where you fall. If you're investing a company, know where you fall. If you want to work at a company, like are you in like vaporware land or are you in like too good to be true, this is going to be incredible land. And you know, like we we should have like a simple explanation for all this. I'm going to oversimplify and just jump straight to my answer. Actually, I'm just going to speed run this whole talk because I want to get to Q&A. It looked really fun last time. It looked super fun, so I'm just going to blast through a lot of the stuff I would go slower on.

人在环 vs 机器在环 Human-in-the-Loop vs Machine-in-the-Loop

Diogo

我相信答案很简单:左边所有东西都是人在环辅助,右边的东西是机器在环自动化任务。这就是为什么很多看起来更简单,因为如果你想自动化,你会想自动化那些简单的小工作,只需要可靠地完成。你可以看到无决策客户服务和有决策客户服务之间的关键区别。有很多 AI 聊天机器人做客户服务代理。它们不允许你采取任何行动。它们喜欢给你戴上手铐。哦,我们稍后会回到这个。它们把你扔进文档的迷宫,但它们不允许采取行动,原因很好。模型不够可靠。但它们在辅助方面非常非常擅长。这里超级常见的 FAQ 是编码代理呢?编码代理不是有点像自动化吗?它说计算机语言,你知道,是代码。自动化还有什么?信不信由你,那也是辅助。当任务为辅助优化时,会有某些气味,你可以看到这一点,当你看到模型的输出时。

I believe that the answer to this is simply that all of the stuff in the left were human-in-the-loop assistance and the things on the right were like machine-in-the-loop automation tasks. And that is why a lot of them look simpler because like if you want to automate stuff, you wanted to automate like simple little bits of work that just need to be done reliably. And you can see the critical difference between customer service without decisions and customer service with decisions. There's a lot of AI chatbots out there who make customer service agents. they do not allow you to take any actions. They like to give you the kitty gloves. Oh, we'll get back to this later on. And they throw you into like the labyrinth of documentation, but they are not allowed to take actions for good reasons. The models are not reliable enough. But they are really, really good at the assistant side of things. Super common FAQ here is what about coding agents? Aren't coding agents kind of like automation? It speaks the language of computers, you know, it's code. What more is there to automation? And believe it or not, that is also assistance. There are like certain smells to when a task is optimized for assistance and like you can like kind of see this when you're like seeing outputs of the models.

优化辅助 vs 自动化 Optimizing for Assistance vs Automation

Diogo

尤其是当你得到看似合理、好看但不正确的答案时。我们稍后在讨论优化时会谈到这一点。但最简单的判断标准是:当它针对辅助还是自动化进行优化时,工作流程是否真的涉及人类在环?可以说代码是给人类而非机器使用的语言。人们对此感到惊讶,这让我很惊讶。这些是去年的结果。这是基于较老一代的模型。所以可能不适用,但这基本上是我们目前在 economy 中持续观察到的效应。你看到很多 AI 加速的感知价值,但就实际结果而言,它并没有真正带来回报。这正是你所预期的,如果你优化的是人类偏好,即 RLHF 中的 HF,而不是你真正关心的根本,即逻辑、决策。就像只是正确完成任务,而 LLMs 并不擅长这个。或者至少它们现在没有为此优化。

Like especially when you get plausible, good-looking, but incorrect answers. We'll talk about this later when we talk about optimization. But the simplest razor for when it's optimized for assistance versus automation is: does the workflow literally involve humans in the loop? And debatably code is a language for humans, not for machines. And it's quite surprising to me that people are surprised by this. These were results from last year. This is on an older generation of models. So, it may not apply, but this is basically the effect that we've kept observing in the economy right now. You see a lot of perceived value of AI speed up, but in terms of actual results, it has not really been paying off. And this is exactly what you'd expect if you're optimizing for the human preference, that is the HF in RLHF, instead of actually the root of the thing you care about, which is logic, decision-making. Like just do the task right, which LLMs are just not very good at. Or at least they're not optimized for right now.

学界的斯德哥尔摩综合征 Stockholm Syndrome of the Field

Diogo

结果是该领域的斯德哥尔摩综合征,人们说:“LLMs 更像人类而不是软件。它们不可预测。”随着对精确度的需求上升,AI 的效用下降,这对今天的 LLMs 来说完全正确,但对机器学习来说完全新颖。这不是机器学习的工作方式。机器学习是高度结构化的。它旨在与自动化集成。实际上,随着精确度上升,ML 的效用会飙升。你会信任一个人来平衡 Meta feed flow 的推荐算法吗?那太疯狂了。那太难了,对吧?但我们出于某种原因认为,“不,不。LLMs,它们像 ML 的反面。对它们要小心翼翼。它们只适合效用的感知,而不是真正的效用。”这对我来说有点疯狂。

And the result is a Stockholm syndrome of the field where people are like, "LLMs are more like humans than software. They're unpredictable." As the need for precision goes up, the utility of AI goes down, which is completely correct for today's LLMs, but completely novel for machine learning. This is not how machine learning works. Machine learning is highly structured. It is meant to integrate with automation. And actually, as the precision goes up, the utility of ML skyrockets. Would you trust a human being to balance the recommender algorithm of the Meta feed flow thing? That's insane. That's so hard, right? But we for some reason we think, "No, no. LLMs, they're like the opposite of ML. Kid gloves on them. They are just good for the perception of utility and not true utility." That is kind of nuts to me.

总结:优化辅助 Summary: Optimizing for Assistance

Diogo

完美。本节总结。这是演讲最重要的部分。我演讲的其余部分基本上是一些不太重要细节的闲聊。但今天几乎所有的 LLMs 都是为辅助而优化的。这导致了过度承诺和交付不足之间的巨大鸿沟。我实际上会认为在 RLHF 之前没有 LLM 历史。我是其中的一部分,所以我更倾向于认为 RLHF 部分更好,但我实际上认为那时没有这种鸿沟。就像 GPT-3 没有 ChatGPT 那样的巨大鸿沟,尽管 GPT-3 是一个很棒的模型。如果你想让今天的 LLM 任务奏效,确保它是辅助任务。人们尝试过非常基本的自动化任务。据我所知,我不太出门,但我不相信 drive-thrus 已经自动化了。如果我错了,请有人纠正我。但就像有大量的经济激励这样做。而且你知道,就像在这一节中,我们有一些模糊的简短预示,我们可以作为一个领域做得更好,最终我们会谈到。但我们需要建立我们的技术基础才能达到那里。

Perfect. Summary for this section. This is the most important part of the talk. The rest of my talk is basically a rambling of less important details. But almost all of today's LLMs are optimized for assistance. This causes the massive divide between overpromise and underdeliver. I would actually argue there wasn't LLM history pre-RLHF. I was part of it, so I'm even more biased to think the RLHF part is better, but I actually think that there was not that divide then. Like GPT-3 did not have the massive divide that ChatGPT has, even though GPT-3 was an awesome model. And if you want today's LLM tasks to work, make sure it's on assistance tasks. People have tried on really basic automation tasks. Like to my knowledge, I don't really go out much, but I don't believe drive-thrus are automated yet. Someone correct me if I'm wrong. But like there's lots of economic incentive to do so. And you know, like also we had during this section we had some vague short foreshadowing that we can do better as a field, which eventually we'll get into. But we need to like build up our technical fundamentals to get there.

优化:好得不像真的 Optimization: Too Good to Be True

Diogo

完美。优化。这是我最痴迷的东西,我喜欢谈论优化。这是一个简短的部分,我认为它解释了,你知道,给你技术支撑,说明 LLMs 优化的目标如何可能导致所有这些效应。所以,我在优化中最喜欢的问题是:什么时候一个结果好得令人难以置信?这是一个看起来超级假的结果。我喜欢这个结果,因为它是我见过的最假的结果。你看到 OpenAI 在底部随着参数数量平线。这不是沙袋。这实际上是在高效前沿上,你看到一个小模型完全主导了当时最好的模型。而且这是当时最好的模型,要明确。然后这不是类型安全的。这实际上是 RLHF。那是同一个图。我只是屏蔽了一些点,这样你可以看到差异。但就像这个图有一些直观的东西,这个图没有完全显示,它归结为人们对优化的直觉。

Perfect. Optimization. This is the thing I am nerdliest for, and I love talking about optimization. It's a short section that I think explains, you know, gives you the technical backing of how it could be possible that what LLMs are optimized for causes all of these effects. So, my favorite question in optimization is: when is a result too good to be true? This is a super fake looking result. I love this result cuz it's like the fakest result I've ever seen. You see like OpenAI like flatlining in the bottom with number of parameters. This is not sandbag. This is actually on the efficient frontier and you see like a tiny little us model completely dominating what those like the best model of the time. And this was the best model of the time to be clear. Then this is not type safe. This is actually RLHF. That is the same plot. I just masked out some points so you can see the difference. But like there's something visceral about this plot that this plot doesn't quite show and it comes down to like people's intuitions about optimization.

缩放定律与苦涩教训 Scaling Laws and Bitter Lessons

Diogo

你可以看到我的鼠标。底部的这条线实际上是真正的缩放定律。这是 GPT-2。这是 GPT-3 家族中的一个小版本,而这是完整的 175B GPT-3 在指令遵循任务上。你可以看到,即使是 GPT-2 大小的模型也完全主导了 GPT-3,当时最好的模型。这需要一个解释,人们对该领域的心理模型应该相应更新。这里一个非常有趣的事情是,即使你以极其优化的方式提示 GPT,它甚至没有接近这里最基本的优化版本,这与人们谈论的缩放定律并不相符。这将在后面多次出现。所以,希望这是有信息量的。人们喜欢谈论 Rich Sutton 的苦涩教训。我转述一下,它模糊地说是算力比算法更重要。这是当人们说他们的苦涩教训药丸时,他们就像把计算机扔给它。就像你知道,让我们开始吧。算法不那么重要。我们只想尽可能多地获取算力。这几乎不能解释该领域的任何东西。实际上,有许多结果表明这个简单解释并不——好吧,我们会谈到。我们会谈到。但对我来说,有一个更苦涩的教训可以简单解释,那就是数据比算力重要得多,做正确的任务比数据重要得多。缩放定律奏效以及你得到算力比算法更重要的原因是,在某些领域,数据已经可用,比如在互联网上,或者在 RL 中,更好的是,做正确的任务有点像简单定义的玩具实验,而数据是算力的函数。所以,向 Sutton 致敬,但我不相信那是完全可解释的事情,这将在后面多次出现。我相信这很容易解释 Anthropic 的崛起,而 OpenAI,就他们的对齐是什么、他们的北极星是什么、他们在做什么正确的任务而言,有点像在使其适合聊天机器人、使其适合博士级科学推理东西以及追赶代码之间摇摆。而 Anthropic 显然对此有更清晰的北极星,这就是他们如何追赶的。值得注意的是,如果你看这条曲线,并试图用除做正确任务之外的任何东西来解释它,你基本上就完蛋了。你知道,OpenAI 在数据上的花费比 Anthropic 多。他们的算力比 Anthropic 多得多,我认为你也许可以论证 Anthropic 在算法上更好,但我个人不会那样做。我认为答案很简单,就像你优化什么就得到什么。

You can see my mouse. This line at the bottom is actually the real scaling law. This is GPT-2. This is a small version in the GPT-3 family and this is full on 175B GPT-3 at the task of instruction following. And what you could see is even GPT-2 sized models like completely dominated GPT-3, best model of its time. That warrants an explanation and people's mental models of the field should be updated accordingly. One of the very interesting things here is that even if you prompt GPT in an extremely well optimized way, it doesn't even approach the most basic version of optimization here, which doesn't really square with the scaling laws that people talk about. And this will come up a bunch later. So, hopefully this is informative. People like to talk about Rich Sutton's bitter lesson. I'm paraphrasing, it is vaguely compute matters more than algorithms. This is when people say their bitter lesson pill, they're just like throw computer at it. Like you know, let's go. And algorithms don't matter as much. We just want to like just get as much compute as possible. And this does not explain almost anything in the field. Actually, there's many results that show this simple explanation to not—well, we'll get to it. We'll get to it. But to me, there is a bitterest lesson that it can be simply explained, which is data matters way more than compute, and doing the right task matters way more than data. And the reason why the scaling laws work and you get compute matters more than algorithms is in some domains, the data is already available, like on the internet, or in RL, even better, doing the right task is kind of like a toy experiment that's simply defined, and data is a function of compute. So, kudos to Sutton, but I don't believe that is like a fully explainable thing, and this will come up a bunch. I believe that this explains the rise of Anthropic quite easily, which OpenAI, in terms of what their alignment is, what is their North Star, what right task are they doing, kind of like flops in between like some mixture of making it good for chatbots, making it good for like PhD-level sciency reasoning stuff, and also like playing catch-up on code. And Anthropic clearly has a more clear North Star for this, and that is how they're catching up. And notably, if you look at this curve and you try to explain it with any of the things other than doing the right task, you are kind of cooked. You know, OpenAI spends more on data than Anthropic. They have way more compute than Anthropic, and I think that you could maybe make a case that Anthropic is way better than algorithms, but I wouldn't personally do that. I think the answer is simple, like you just get what you optimize for.

优化与RLHF Optimization and RLHF

Diogo

这就是这一部分的全部教训。哦,没错。你得到的就是你优化的东西。这就像是所有机器学习、AI、LLM 等等的教训。它只是有风格的优化。如果你忘了你在优化东西,那你也就不理解这个领域。而在基于人类反馈的强化学习(RLHF)中,我们大量优化了字符串。酷。

And that's the whole lesson of this part. Oh, hell yeah. You get what you optimize for. This is like the lesson in all of ML, AI, LLMs, whatever. It is just optimization with style. If you forget that you're optimizing stuff, then you are also not understanding the space. And we have optimized a lot for strings in RLHF. Cool.

Diogo

我能大概举手看看谁知道什么是基于人类反馈的强化学习(RLHF)吗?哦,大约一半。嗯。我们会即兴发挥。看看会不会提到。

Can I get a vague show of hands of who knows what RLHF is? Oh, about half. Huh. We'll improvise this. We'll see if it comes up.

Diogo

所以,基于人类反馈的强化学习(RLHF)代表从人类反馈中进行强化学习。它是 ChatGPT 背后的算法。如果你看到像这样的图表,RLHF 就是实现它的原因。据我所知,大约 100% 的实际生产 LLM 都是用 RLHF 或其变体训练的。这就是用来让模型工作的东西。这有很多权衡。我可能在这个演示中讲得太深了,所以抱歉,但希望会有录像。但 RLHF 导致了许多非常奇怪的性质,值得讨论,因为它们非常集中于这种字符串输入字符串输出的范式,这对人类消费很好,但对自动化很糟糕。

So, RLHF stands for reinforcement learning from human feedback. It is the algorithm behind ChatGPT. And if you see plots like this, RLHF is what made that happen. As far as I can tell, roughly 100% of actual production LLMs are trained with RLHF or some variation of it. And this is what is used to just make models work. There's a lot of tradeoffs to that. I might have gone a little bit too deep in this presentation, so sorry, but hopefully there'll be a recording. But RLHF results in a bunch of very odd properties that are worth talking about because they're very centric on this string in string out paradigm that is very good for human consumption but very bad for automation.

Diogo

关于 RLHF,我要讲的第一件事,这有点像趣闻,是 Yann LeCun 的那张著名幻灯片。如果你不知道他是谁,他是现代 AI 的教父之一,绝对是深度学习的教父。他有一张人们喜欢嘲笑他的幻灯片,内容是 LLM 注定要失败。他一直说 LLM 注定要失败。他认为它们是一个糟糕的方向。他提出了一个非常简单的数学论证,即如果你在生成一个 token 序列,这就是 LLM 所做的,如果每个序列都有错误的概率,并且你连续生成很多,整个序列正确的概率会指数下降。这是一个非常有趣的直观论证,但常识表明它并不成立。你知道,就像你让 ChatGPT 写一篇维基百科文章,它会写得非常好。它不会像指数衰减和发散那样。它只会是对或错,就像简短回答一样。所以,这有点奇怪,你知道,你对字符串模型中发生的事情有两种不同的看法,而发生这种情况的原因叫做模式丢弃或模式崩溃。这是 RLHF 无法解决的与魔鬼的交易,因为它确实是一个特性,而不是一个 bug。

The first thing I'm going to talk about with RLHF, and this is kind of like a fun fact, is that there is this famous slide of Yann LeCun. If you don't know who it is, he is like one of the godfathers of modern AI, definitely a godfather of deep learning. And he has this slide that people like to dunk on him about, which is that LLMs are doomed. He keeps saying LLMs are doomed. He thinks that they're a bad direction. And he makes this really simple mathematical argument, which is that if you're making a sequence of tokens, which is what an LLM does, if each sequence is a probability of being incorrect and you make a lot of them in a row, the probability of the whole sequence correct goes down exponentially. And this is like a very interesting intuitive argument, except that common sense shows it doesn't work. You know, like you ask ChatGPT for a Wikipedia article, it's going to Wikipedia article that thing really well. It's not going to like degrade and diverge exponentially. It is just going to be right or wrong, just like a short answer is going to be. So, that's kind of weird, you know, you have like these two different takes on what's going on in string models, and the reason this occurs is something called mode dropping or mode collapse. This is RLHF's unfixable deal with the devil because it really is a feature and not a bug.

Diogo

当你有一个优化曲面时,天哪,我要讲得太技术了,不是吗?讲吧。天哪。当你有一个优化曲面,对于你的底层函数来说太复杂而无法拟合,或者很难拟合时,你最终不得不选择你的误差形状。这往往来自你的优化函数。可能出现的两种粗略误差形状是模式崩溃和模式覆盖。所以,有时你可以在两个模式之间,得到中间的东西,预训练语言模型非常擅长这一点,因为它们的损失函数会这样做。这就是为什么它们如此有创造力,可以写绝对疯狂的东西。另一方面,RLHF 极其模式丢弃,这意味着如果你有两种可能性,它高度倾向于选择安全的选项。这使得事情几乎总是看起来非常合理正确。这应该让你对 AI 的状态有一个很大的疑问。就像,为什么它总是那么合理正确?这并不是所有 LLM 都如此。这是 RLHF 的一个特性。我在做这个时向自己保证不要在这里太书呆子,所以我不会那样做。但这是 RLHF 所做事情的一个非常重要的特性,以及为什么优化中没有免费午餐,因为这是一个极其有价值的特性。如果你尝试模式覆盖,当你在做一个决定时还可以,但如果现在你基于那个决定做另一个决定,并且你这样做 10,000 次,你现在有零机会得到一个合理正确的答案,有 10,000 个 token,Yann LeCun 的答案就对了。所以,这里的 TLDR 是 RLHF 更像一个 GAN。所以,如果你有一个少数类被采样,它可以丢弃那个空间来制作看起来合理的答案,而预训练更像是做一个模糊的图像,而不是从现有训练集中取东西。我还没有声称什么对自动化更好,但我声称这是 RLHF 的一个不可否认的特性,解释了一大堆事情。完美。

When you have an optimization surface, oh man, I'm going to go way too technical, aren't I? Do it. Oh boy. When you have an optimization surface that is way too complicated for your underlying function to fit or it's very hard to fit it, you end up having to choose what the shape of your errors are like. And this tends to occur from your optimization function. And two rough shapes of errors that can occur are mode collapse and mode covering. So, sometimes you can like in between like two modes, you can get stuff in the middle, and pre-trained language models are really good at this because their loss does that. That is why they're so creative and they can write about like absolute crazy stuff. On the flip side, RLHF is extremely mode dropping, which means that if you have like two possibilities, it's highly incentivized to just go with the safe option. And this makes things look very plausibly correct pretty much all the time. And that should give you like a pretty big huh about the state of AI. Like, why is it always so plausibly correct? And this is not true of all LLMs. This is a property of RLHF. I promised myself when making this do not nerd out too much here, so I will not do that. But this is a very important property of the whole of what RLHF is doing and why there is no free lunch in the optimization because this is an extremely valuable property. If you did try to mode cover, that is okay when you're doing one decision, but if now you're doing another decision based on that decision and you do that like 10,000 times, you now have zero chance of a plausibly correct answer with 10,000 tokens and Yann LeCun's answer gets right. So, the TLDR here would be that RLHF works a lot more like a GAN. So, if you have like a minority class that is being sampled, it's okay for it to just drop that space to make plausible-looking answers while pre-training is more like doing a blurry image in between instead of like just taking stuff that is from the existing training set. I'm not making a claim about what is better for automation just yet, but I am making a claim that this is like an undeniable property of RLHF that explains like a whole bunch of things. Perfect.

Diogo

这有点相关于这个题外话:软件工程比得来速更容易吗?也就是说,你知道,我们应该问自己这个问题。任何了解 AI 现状的人,就像有一个元叙事正在进行,软件工程正在被解决。那怎么可能比得来速更容易?我喜欢用得来速作为例子,不是说我非常关心得来速,显然,而是因为失败是如此公开和有趣,就像是对一个行业实际发生的事情的试金石。你知道,真正的 AI 人士试图说的叙事是那些人很烂。你知道,就像他们不知道如何实现它。他们不擅长,如果你付钱给我们,我们会为你实现,据我所知这不是真的,而且真的没有推动多少进展。所以,对我来说类似的问题是,如果我不相信软件工程比得来速更容易,我实际上认为思考的方式是,你会信任一个没有版本控制的编码智能体吗?我本应该在这里放版本控制,但没有版本控制。这有点疯狂,对吧?就像一个没有版本控制的软件编码智能体,那绝对是疯了。但你可以信任人类做到这一点,对吧?或者至少那是在我的时代之前。我的理解是软件工程发生在版本控制之前。有人吗?我不知道。是的,它发生了吗?

And this is a little bit related to this aside of is software engineering easier than drive-thrus? That is, you know, we should be asking ourselves this question. Anyone who knows what's going on in AI right now, like there's a meta-narrative going on that software engineering is getting solved. How could that be easier than drive-thrus? I like drive-thrus as an example, not that I care a lot about drive-thrus, obviously, but because there's the failures are so public and entertaining that like it's kind of like a litmus check of like what's actually happening in an industry. And the you know, the narrative that the really AI people are trying to say is those people suck. You know, like they don't know how to implement it. They're bad at it and if you just pay us, we'll implement it for you, which is as far as I can tell not true and really hasn't moved the needle very much. So, the similar question to me is if I don't believe software engineering is easier than drive-thrus, I actually think that the way to think about it is that would you trust a coding agent without I should have put version control here, but without version control. And that's kind of crazy, right? Like a software coding agent without version control, that's like absolute nuts. But you could trust humans with that, right? Or at least that was back before my era. My understanding is that software engineering happened before version control. Anyone? I don't know. Yeah, yes, did it happen?

Diogo

哦。我当时在谷歌,所以我们做了一些古老的东西。但对我来说,这里的教训是,模型不能被信任来做出你承担风险的决策,因为它们被如此鼓励做出看起来合理的答案,如果你想做出校准的决策,这不是你想要的。如果别人承担风险,那就没问题。这就是为什么你看到 AI 客户支持的兴起。但 AI 客户支持从不采取行动,因为成本总是由用户承担,我们稍后会回到这一点。也许我现在就讲。

Oh. I was at Google, so we did some ancient stuff. But to me the lesson here is that the models just cannot be trusted to make decisions with stakes that you pay because they're so encouraged to make plausible-looking answers, which is not what you want if you want to make calibrated decisions. If others pay, it's all good. That is why you see a rise of AI customer support. But the AI customer support never takes actions because the cost is always to the users, and we'll get back to this later on. Maybe I'll just get into it now.

人们为何讨厌AI Why People Hate AI

Diogo

我认为这就是为什么世界上对 AI 有一种隐隐的恨意。我不觉得人们恨技术。我觉得他们恨的是技术对他们不利。而现在,所有人都有极强的动机永远不把事交给 AI 去信任,但你又不得不用 AI 做事。所有外部化的成本都被转嫁到用户身上,于是他们就想:“算了,这玩意儿太烂了。”我觉得这就是对整件事最现实的看法。

I believe this is why there's this subtle hatred for AI out in the world. I don't think people hate technology. I think they hate when technology is bad for them. And right now, everyone's extremely incentivized to never trust the AI with stuff, but you need to use AI for stuff. And all of the externalized costs are put on the users, which makes them be like, "Fuck it. This sucks." And I think this is just the realistic take of the whole thing.

ChatGPT要完蛋了吗? Is ChatGPT Cooked?

Diogo

好。哦,关于这个再插一句。ChatGPT 会怎么样?他们是不是完蛋了?他们是不是就会停滞在不到十亿用户之类的水平?

Cool. Oh, another aside on this. What's going to happen to ChatGPT? Are they cooked? Are they just going to plateau and be at like sub-billion users or whatever this is?

Diogo

我认为答案是——人们对未来可以有各种乱七八糟的看法。我是个搞优化的人。这件事有一个非常简单的优化视角。ChatGPT 远远没有完蛋。完蛋的是这个世界。OpenAI 还没有——哦,而且就算 OpenAI 不做、他们想当好人,也会有别人来当坏人。这是不可避免的。我不怪 TikTok 让人脑腐,无意冒犯,但脑腐一定会出现。这里面的激励太大了,如果我们作为一个社会不保护自己,它就会发生。而现在我们还处在 ChatGPT 的 MySpace 时代。它很友好,很好玩,老实说,我觉得他们只是搞不清自己在优化什么,所以才会来回反复横跳。很明显,他们的优化团队在这里缺了马眼罩。而我们整个领域甚至还没开始优化。就算是技术人——这跟我的演讲无关,就是优化,我在发牢骚——就算是非常技术的人也会对此感到惊讶。我觉得这对世界会非常糟糕。他们会反驳说,那如果只是脚本小子在调你的提示词,想让你稍微更上瘾一点,比如做性爱机器人什么的呢。那跟一家公司在 ChatGPT 这种能拿到如此丰富反馈的目标上做生产级优化,完全不是一回事。所以这会很糟,做好准备吧。理想情况下政府会介入,但我不太抱希望。但从优化的视角看,这恰恰说明我们连表面都还没刮到,这非常吓人。

I believe that the answer to that — people can have random-ass takes for the future. I'm an optimization guy. There's a very simple optimization take to this. ChatGPT is extremely not cooked. The world is cooked. OpenAI has not yet — oh, and even if OpenAI doesn't do it and they try to be the goodies, someone else will be the baddie here. This is inevitable. I don't blame TikTok for being brain rot, no offense, but brain rot will emerge. There's just a lot of incentive for that, and if we don't protect ourselves as a society, it will happen. And right now we're in ChatGPT's MySpace era. It's friendly, it's fun, and honestly, I think they're just confused about what they're optimizing for, which is why they whiplash back and forth so much. It's kind of obvious that their optimization team lacks horse blinders here. And we as a field have not even begun to optimize. Even technical people — this has nothing to do with my talk, just optimization, I'm just ranting — even very technical people are surprised by this. I think it's going to be really bad for the world. They make the counterargument that, what if you have script kiddies tuning your prompts to try to get slightly better at getting more addictive, if you're making sex bots or whatever else. And that is absolutely not the same as a company doing production-scale optimization on an objective that you get such rich feedback on as ChatGPT does. So this is going to be bad, prepare for it. And ideally governments are involved, but I'm not too hopeful of that. But this has got — the lens of optimization just shows that we have not yet begun to scratch that surface, and that is super spooky.

全力投入人在环 Go All In on Human-in-the-Loop

Diogo

太好了。而这一切的结果——哦,串了——智慧就是,短期内正确的做法就是全力投入。

Perfect. And the results of all of this — oops, string — wisdom is the correct thing to do in the short term is to go all in.

Diogo

其实你应该放弃其他路线。真的,除非你是研究者,否则你应该放弃,因为它不管用。你看,昨天 Thinking Machines 发了一篇帖子,他们全力投入人机带宽。我觉得那是个很好的方向。我不知道怎么说才不显得阴阳怪气,但如果你解决不了自动化问题,那对大多数人来说这就是合理的。那就是你现在该做的事,因为 AI 的每一个用例都是人在回路中。我很喜欢这里的这个总结。LLM 对分层极其抗拒。一般来说,消费者是人,而不是另一层软件,等等等等。这就是这个领域现在的状态。所以如果你押注的是可组合性、分层或者其他类似很酷的东西,我有个坏消息给你。如果你不自己做优化,你就完蛋了,因为优化是在幕后以你无从知晓的方式完成的。

You know, actually you should give up on the other approaches. Truly, unless you're a researcher, you should give up on that because it doesn't work. You see, there's a post from Thinking Machines yesterday, which is they're going all in on human-AI bandwidth. I think that's a very good direction. I don't know how to say this in a way that's not passive-aggressive, but that makes sense for most people if you can't solve the automation problem. That is what you should be doing right now, because every use case of AI is human in the loop. And I love this summary here. LLMs are incredibly resistant to layering. Generally the consumer is a human, not another layer of software, blah blah blah. And that is the state of the field right now. So if you are betting on composability or layering or other cool stuff like that, I have bad news for you. You are cooked if you're not doing the optimization yourself, because the optimization is done behind the scenes in ways that you are not privy to.

理解当今LLM的特性 Understand the Properties of Today's LLMs

Diogo

太好了。这一节的总结是:如果你用今天的 LLM,请理解它们的特性。不要拿有重大利害的事去做决策。

Perfect. Summary for this section is: if you use today's LLMs, please understand their properties. Don't make decisions with stakes.

Diogo

人们试过了。我不知道有谁真正成功过。我见过利害最大的也就是控制一个日历。你们告诉我?你们来啊。我们在这儿来场真正的辩论。但 OpenClaw 正在发生的事让我——我找不到别的词来形容,就是害怕。我对 OpenClaw 正在发生的事非常非常害怕。它似乎正在消亡,但这对世界来说似乎真的很糟。AI 靠常识做不到那个。这行不通,不管人们多狂热,它都不会成功。要有人的回路。这是 OpenClaw 没有遵循的头号教训。我不知道在座有多少是搞商业的,但 FTE 解决不了这个问题。这是其他所有人都在试图解决的共同问题,因为元叙事显然是 AI 很棒,但我所有的内部证据都表明 AI 很烂。那一定是我自己的问题。让我花钱请专家来搞定它。但作为一个跟那些专家合作过、见过香肠是怎么做出来的人,我不相信是那样。

People have tried. I don't know of anyone who's really succeeded. The most I've seen for stakes is having control of a calendar. Do you guys tell me? You guys come on. Let's have a real debate here. But what's happening with OpenClaw makes me — I don't know of another analogy for scared. I'm very, very afraid of what's going on with OpenClaw. It seems to be dying out, but this seems really bad for the world. AI can't do that with common sense. This is not working out, and no matter how crazy people are about it, it's not going to work. Have human in the loop. That is the number one lesson that OpenClaw did not follow. And I don't know how businessy people are here, but FTEs will not solve this problem. This is the common problem that everyone else is trying to solve, because the meta narrative is clearly AI is great, but all of my internal evidence is that AI sucks. It must be a me problem. Let me pay experts to try to do that. But as someone who's worked with those experts and seen how the sausage is made, I don't believe that is the case.

Host

是技能问题。技能问题。

Skill issue. Skill issue.

Diogo

是啊。不过我不觉得这是技能问题。OpenAI 基金里有些公司自动化的程度跟其他所有东西一样低。我大概不该点名,因为我现在不想挑起争端。我会在发布产品之前挑起争端,但不是今天。但这就是现在的叙事,我觉得它说不通。而且我觉得这么误导人其实对世界有害,尽管这么做显然有巨大的经济激励。呼。这对我来说就像心理治疗。

Yeah. Well, I don't think it's a skill issue. There are companies in the OpenAI fund that are automating just as little as everything else. I probably shouldn't name names because I don't want to start fights right now. I will start fights before we release a product, but not today. But that is the narrative right now, and I don't think it makes sense. And I think it's actually bad for the world to be this misleading, even though there's obviously so much economic incentive to do so. Woo. It's like therapy for me.

类型安全的语言模型 Type-Safe Language Models

Diogo

好。类型安全的语言模型。这里是形容词“类型安全”,不是那家公司。类型安全是计算机科学里的一个概念。它讲的是超越字符串,真正整合进编程。我想总体谈谈这些。等到了广告环节就会清楚。我的主张是我们可以做得更好。这是个非常有争议的主张。其实,这大概是人们问得最多的一个问题,当我说——人们会说,这整套叙事说得通,但它们肯定没那么差吧。我告诉你,我第一次得到这个方向时,我当时在 OpenAI,当我想到它时,我就想,这也太明显了。Anthropic 大概领先一年。我们整个领域算是完蛋了。而 OpenAI 在这个方向上完蛋了,可它到现在还没发生。我们仍然没有那种可分层、智能便宜到不用计量的 AI。我大概不该多谈我们做什么,但我相信这是可能的。由此会引出两个问题,而且这还是在技术人层面——所以我想鼓舞人心地说,你可以做得更好。我们可以做得更好。这个领域不会止步于此。

Perfect. Type-safe language models. This is type-safe, the adjective, not the company. Type safety is a concept in computer science. It's about going beyond strings and actually integrating into programming. I want to talk about these in general. It'll be clear when we're in the ad section. My claim is we can do better. This is a very spicy claim. And actually, this is probably the biggest question people ask when I say — people are like, this whole narrative makes sense, but surely they can't be that bad. And I will tell you, when I first got this direction, I was at OpenAI at the time, and when I had it, I was like, this is too obvious. Anthropic is like a year ahead. We are kind of cooked as a field. And OpenAI is cooked in this direction, and it still hasn't happened yet. We still don't have layerable AI that is kind of like the intelligence too cheap to meter. And I probably shouldn't talk too much about what we do, but I believe this is possible. There are two questions that arise from how that can be, and this is still with technologists — so I want to be inspirational and say, you can do better. We can do better. The field doesn't end here.

FLAN的教训:主流实验室也会错 The FLAN Lesson: Mainstream Labs Can Be Wrong

Diogo

如果你今天想用 LLM,短期内你应该放弃这条路,但这是这个领域不可避免的未来,因为显然这就是通往自动化的方式。那么,第一个问题来了:我们怎么能做得更好?这怎么可能?主流实验室是不是太差了,以至于他们根本不知道这一点,只是在瞎搞?我把这个叫做 FLAN 的教训。有一篇论文叫 FLAN,fine-tuned language 什么的,我不太清楚。这是谷歌的论文,有 6000 次引用。我不酸,我发誓。我其实是这项工作的粉丝。他们抄了 OpenAI 的东西,创造了指令微调。然后他们试图抢跑模型。我觉得很酷,因为他们开源了数据。太酷了,我爱数据。那是在我意识到做对任务很重要之前。这其实出奇地难。我们发现最优的量是绝对的零。这就是当时最先进水平的样子。最先进水平可能完全是错的,因为基准测试极具误导性。他们的数据毫无价值。可悲的是,人们还在继续尝试他们的数据集。我时不时还会看到论文说:“我们试了这个数据,发现最优量是零。”主流领域可能是错的。这个领域就是这样运作的。

You should give up on this in the short term if you're trying to use LLMs today, but this is the inevitable future of the field because this is just obviously how you get to automation. So, question one that comes up from the claim of how can we do better is: how can that be possible? Are the mainstream labs so bad that they just don't know this and they're just clowning around? And I call this the FLAN lesson. There was a paper called FLAN, fine-tuned language something, I don't really know. It's a Google paper. It had 6,000 citations. I'm not bitter, I swear. I actually am a fan of this work. They copied the OpenAI thing. They coined instruction tuning. And then they tried to front-run the model. I thought it was cool because they open-sourced their data. That is so cool. I love data. This is before I realized that doing the right task is important. It's actually surprisingly hard. And we found that the optimal amount of this is absolute zero. And that's kind of what the state of the art can be at the time. The state of the art can be straight-up wrong because benchmarks are super misleading. And their data was worthless. And the sad part is that people continue to try their datasets. And I still see papers every now and then that are like, "We tried this data and we found optimal amount was zero." And the mainstream field can be wrong. This is just how the field works.

RLHF并非显而易见 RLHF Was Not Obvious

Diogo

这个问题的另一面是:如果 RLHF 那么明显——现在它极其明显——但当时并不明显。当时 OpenAI 有三种后训练方法。它是所有人最不喜欢的一个。没人喜欢它,因为它是一个对齐项目,而且工作量很大。但它最终坚持下来了,因为我们做对了任务。其他努力都失败了,但这个变得难以置信地成功。席卷全球。商业成功。事后看来很明显,但为什么我们不用 GPT-2 做呢?一个 GPT-2 大小的模型显然完全可以在这个事情上超越 GPT-3,对吧?问题是,做这件事其实超级难。我会说,对于 LLM 来说,这大概发生过一次半。如果你把 LLM 算作从 GPT-3 左右开始,它们已经存在了大概六年半。有 RLHF,我认为这是路上最大的弯道。然后是 RLVR,我会给它一半的功劳,因为人们多少在用,但它仍然是 RLHF,因为 RLVR 其实效果不太好。那是推理相关的东西。对于六年的创新来说,这是非常短的时间。这不是一件容易的事。

And a flip side of this question is: if RLHF was so obvious, which now it's extremely obvious, it wasn't obvious at the time. There were like three post-training methods at OpenAI at the time. It was everyone's least favorite one. No one liked it because it was an alignment project and it was a lot of work. But it ended up sticking because we did the right task. The other efforts kind of failed, but this ended up becoming unbelievably successful. Took over the whole world. Commercial success. And it seems obvious in hindsight, but why didn't we do it with GPT-2? A GPT-2 sized model clearly could completely chance GPT-3 at this thing, right? And the thing is that doing this is actually super duper hard. And I would say that it's happened maybe one and a half times for LLMs. If you count LLMs as starting like GPT-3-ish, they've been around for like what? Six and a half years. And there was RLHF, which I believe is like the biggest curve in the road. And then RLVR, which I'll give like half credit for because people kind of use it, but it's still RLHF because RLVR doesn't really work very well. It's the reasoning stuff. And that's like a very small amount of time for like six years of innovation. And it's not an easy thing to do.

能否兼得所有好处? Can We Have the Best of All Worlds?

Diogo

人们喜欢问的第二个问题是:我们能拥有所有世界的最佳吗?我们能不能得到一个超级擅长辅助的模型?然后微调继续,也许像 ChatGPT 10 会在自动化方面足够好,我们最终解决它,但在此之前我们从辅助中提供大量自动化价值。如果你看旧曲线,你必须外推很远才能希望它泛化到错误的任务。GPT-3 是如何被超越的?它的任务是理解、建模预训练分布。而 ChatGPT 在自动化方面如何被超越?他们优化的是人类偏好。那是不同的。我对此的回应,如果我要再给一个,其实是同一张图。你真的能在面包和罂粟籽上同时最大化吗?你不能两者兼得,对吧?你可以有一个最优混合,但这很难,因为你优化的是两个极其不同的属性。如果它们朝非常不同的方向拉,你最终会粉碎优化空间,得到看起来像尖刺智能或锯齿状前沿的东西。

Second question that people love to ask is: can we have the best of all worlds? Can we just get a model that is super good at assistance? And then just fine-tune to keep on, maybe like ChatGPT 10 is going to be good enough at automation that we finally solve it, but we provide a lot of value of automation from assistance before that. And if you look at the old curve, you'd have to extrapolate so far to hope for it to generalize to the wrong task. How GPT-3 was beat is its task was understanding, modeling the pre-training distribution. And how ChatGPT could be beat at automation is that they're optimizing for human preference. That is different. And my response to this, if I were to give another one, is actually the same image. Can you actually max out on bread and poppy seed or whatever that is? And you just can't have both, right? You can have an optimal mixture of them, but it's just very hard because you get two extremely different properties you're optimizing for. And if they're pulling in super different directions, you end up shattering the optimization space and ending up with what looks like spiky intelligence or the jagged frontier.

试图两者兼得的问题 The Problem with Trying to Get Both

Diogo

我相信这并不好,因为如果你想要真正的可靠性,你需要像九个九的可靠性才能工作。而试图两者兼得就像毒害你的模型。也许在极限情况下,但我认为你无法在这里得到两个世界的最佳。

And this, I believe, is not great because if you want actual reliability, you need like nines of reliability to work. And this is like trying to go for both is like poisoning your models. Maybe in the limit, but I don't think that you can get both the best of both worlds here.

广告:Type Safe的使命 Ad Break: Type Safe's Mission

Diogo

完美。现在是我的广告部分。我时间不多了。我要摘下热情研究员的帽子,谈谈 Type Safe 的帽子。我不能告诉你我们在做什么。Brian Bischof 这么说的。他说,但我们正在做的是放弃所有这些系统,全力投入自动化。我会告诉你更多,但这是一封真实的邮件。这不是 AI。这是一张假图片。但是的,我们只是全力投入自动化。我想谈谈为什么。我们可以谈谈之后会发生什么。但对我来说,我相信 AI 真的是改变世界的东西。像 pre-chain,ChatGPT 发布是什么时候,三四年前,我不太清楚。而现在,我们所有或大多数人拥有的只是 ChatGPT 和现在的 Claude Code。这有点疯狂。我们以为会有这种基础性的新技术,我们在其上构建。而现在我们只是有时打开这个额外的标签页,有点像谷歌。这是对技术的巨大浪费。我认为经济革命仍然可能发生。我会说,让我决定做这个初创公司的是,如果 AI 泡沫破裂,我希望它根本不破裂,所以如果我们设法避免它,我会显得很蠢。但这是我关心的。如果我没有全力投入来避免这种情况,我会为我在促成这件事中所扮演的角色感到无尽的遗憾。我真的认为 RLHF 是路上一个不必要的弯道。我认为它很棒,但我认为人们没有意识到我们本可以走其他弯道。我会再给一个问题,稍后讨论,但我只有 2 分钟。如果你想参与,公司里人们让我说这个。如果你想做真正的自动化,我们有一个注册表单。我们有一个,如果你想跳过等待名单,你应该告诉我们更多关于你在做什么。我们不是所有自动化的完美选择。如果你试图做一个助手人类在循环中的东西,我们是错误的模型。我们想要硬核的、超级无聊的自动化,最终自动化工作。请填写那个。如果你想和我们一起构建,我们有一个招聘页面。我发誓我们非常非常酷。我的联合创始人在这里。

Perfect. Now for my ad component. I have not too much time. I'm going to remove my passionate researcher hat, talk about Type Safe hat. I can't tell you what we're doing. Brian Bischof told me so. He says, but what we're doing is we're giving up on all these systems, going all in on automation. I would tell you more, but this is a real email. This is not AI. This is a fake picture. But yes, we are just going all in on that automation. I want to talk about the why. We can talk about what to happen afterwards. But for me, I believe AI truly is this world-changing thing. Like pre-chain, the fact that ChatGPT released what, three four years ago, I don't really know. And right now, all we have or all what most people have is like ChatGPT and now Claude Code. That's kind of insane. We thought that there would be this foundational new technology that we're building up on. And now we just have this additional tab we sometimes open that's kind of like a Google. This is such a waste of a technology. And I think that economic revolution still can be happening. I will say that what made me decide to do this startup was that if the AI bubble bursts, I'm hoping it doesn't burst at all, so I will look really dumb if we manage to avert it. But this is what I care about. If I did not go all in to avert this, I will feel endless regret for the part I played in making this happen. I really think RLHF was like a curve in the road that was not necessary to happen. I think it's great, but there are other curves in the road that I think people don't realize we could have taken. And I will give like one more question that I'll talk about later on, but I have 2 minutes. If you want to be involved, people are making me say this at the company. If you want to do real automation, we have a sign-up form. We have like a, if you want to skip the waitlist, you should tell us more about what you're doing. We are not a perfect fit for all automation. If you are trying to make like an assistant human in the loopy thing, we are the wrong model for you. We want hardcore super boring ass automation of finally automating jobs. Please fill that out. If you want to build it with us, we have a careers page. I swear we are very very cool. My co-founder is here.

招聘与发布计划 Hiring and Release Plans

Diogo

我们稍后会回答问题,但我们正在招聘各种职位,而且我们增长得非常快。我们准备在几个月内发布这个。其他一切事宜,只需给我们发邮件。

And we will answer questions later on, but we are hiring all sorts of roles, and we are growing super fast. We are getting ready to release this in a couple of months. And for everything else, just email us.

Diogo

哦,我还有一张幻灯片。总结。我的总结是,我们为辅助而优化。这就是对过度承诺、交付不足这个问题的答案。对我来说非常清楚。希望你们也清楚了。我稍后愿意在那个房间里一起测试假设。

Oh, I have another slide. Summary. My summary is we optimize for assistance. That is the answer to the question of overpromise underdeliver. It's super duper clear to me. Hopefully it's clear to you guys. I'm down to test hypotheses together in that room later on.

Diogo

不要让 AI 做有风险的决定。我认为这只会给 AI 带来坏名声。希望大家已经内化了这一点,但也许如果你说得这么清楚,就会变得容易一些。我们应该安装 OpenClaw 来让用户加入任何组织吗?这取决于情况。有没有风险?你是否有安全措施确保它不会造成伤害?如果没有,就不要做。或者如果你想让人嘲笑你,那就去做吧。

Don't have AI make decisions with stakes. I think it just gives AI a bad name. Hopefully everyone has internalized this, but maybe if you say it this clearly it becomes a little bit easier. Should we install OpenClaw to onboard users into whatever organization? Like, it depends. Are there stakes or not? Do you have the security around it to make sure it can't do harm? If not, don't do it. Or if you want people to laugh at you, go for it.

Diogo

如果你对 LLM 感到厌倦,只需知道我们作为一个领域可以做得更好。我知道有些人对真正拥有这一新层技术充满热情。除了基于人类反馈的强化学习(RLHF),还有其他路径,我个人对研究这些路径感到兴奋。可能还有更多,但我对整条路径非常兴奋。

And if you are jaded with LLMs, just know that we as a field can do better. I know that there are some people who are super passionate about actually having this new layer of technology. And there are other paths other than RLHF, which I'm personally excited about working on. There might be more, but I'm super jazzed about this whole path.

Diogo

我还有 11 秒,时间刚刚好。这个问题让我重新——实际上,这个我要超时了。我刚意识到我有背景。

And I have 11 seconds, which is perfectly timed. This was the question that made me re- Actually, I'm going to go over time on this one. I just realized I have like background.

Diogo

当我们发布基于人类反馈的强化学习(RLHF)时,实际上那是在 GPT-3.5 时代左右。我们得到的结果是,模型在遵循指令方面超越了人类。这种字符串输入、字符串输出的东西,我们在这方面超越了人类水平。当时,我们问自己,这是 AGI(通用人工智能)吗?然后整个团队都说,我们不要发布这个,因为有些安全研究员,但我是能力研究员。所以,我说,管他呢,发布吧。结果它非常巨大,你知道,它席卷了全世界,但出于某种原因,它被用于极少的自动化。它基本上只用于文案写作。我相信大约 90% 的用途是文案写作,比如写这些巨大的、推销式的假邮件。这太糟糕了。我不想评判提供价值或其他什么,但那是很难辩护的。

When we released RLHF, actually this was around the GPT-3.5 era. We got the results that the models were superhuman at following instructions. This whole string in string out thing, we were surpassing human level at this. And at the time, we asked ourselves, is this AGI? And then the whole team was like, let's not release this cuz there were safety researchers, but I'm like a capabilities research researcher. So, I said, it, we ball. Like, release it. And that was it was huge, you know, like it took over the whole world, but for some reason it was used for extremely little automation. It was only used basically for copywriting. I believe like 90% of the usage was copywriting, like writing these giant salesy fake emails. And like, that's like so bad for the world. Like, I don't want to judge providing value or whatever else, but like that that it's a hard thing to defend.

Diogo

我过去常问自己,缺少了什么?为什么 AI 平均来说这么聪明,因为你甚至不需要达到人类平均水平。一半的人类比那还笨,但他们仍然能做很多工作。这不是侮辱他们。它应该问我们关于我们的技术。我们应该思考一下。缺少了什么?我认为这不是关于让模型本身更聪明。我认为我们只是做错了任务。

And I used to ask myself, like, what was missing? Like, how can AI be so smart on average cuz like you don't even need to be get to average level human performance. Like, half the humans are dumber than that, but they can still do a lot of work. Not an insult to them. It should like ask us about like our technology. We should think about it. And what was the missing thing? And I don't think it's about making the model smarter per se. I think we're just doing the task wrong.

Diogo

真正让我恍然大悟的是,如果我们有一场基于 AI 的经济革命,那种很多人现在已经放弃的事情,他们只是认为未来就是云代码之类的。有多少调用会是计算机的简单事情?你知道,那种为计算机理解而优化的东西,让 LLM 成为一种新的原语,更像数据库和 API,而不是同事的复制品,我们应该为此优化吗?

And the thing that really clicked for me is if we had an AI-based economic revolution, the kind of thing a lot of people have given up on now, and they're just thinking like that the future is like Claude Codes for days. What percent of calls would be simple stuff to computers? You know, like the kind of thing that's optimized for like computer understanding, having LLMs be a new primitive, more like databases and APIs rather than like facsimiles of co-workers, and should we optimize for that thing?

互动版:逐字朗读 + 针对本期提问 →