Skill issue: code agents and autoresearch
打开互动全文版(中英对照 + 朗读 + 问答)→代码智能体还差在哪,以及通往自主研究那条「绕圈」的路。
Where coding agents fall short, and the loopy road to autonomous research.
听众朋友们,欢迎回到 No Priors。今天我和 Andrej Karpathy 在一起,我们将进行一场广泛的对话,涵盖代码智能体、工程和 AI 研究的未来、更多人如何为研究做出贡献、机器人领域的最新进展、他对智能体如何延伸到现实世界的预测,以及新时代的教育。欢迎你,Andrej。
Hi listeners, welcome back to No Priors. Today I'm here with Andrej Karpathy and we have a wide-ranging conversation for you about code agents, the future of engineering and AI research, how more people can contribute to research, what's happening in robotics, his prediction for how agents can reach out into the real world, and education in this next age. Welcome, Andrej.
谢谢邀请。
Yeah, thank you for having me.
所以过去几个月 AI 领域非常激动人心。
So it's been a very exciting couple of months in AI.
是的,可以这么说。
Yeah, you could say that.
我记得有一次走进办公室,你非常专注,我问你在忙什么,你说「我每天必须编码 16 个小时」,或者「编码甚至不再是正确的动词了,对吧?但我必须每天 16 个小时向我的智能体表达我的意志。显化。」因为能力有了飞跃。发生了什么?跟我说说你的经历。
I remember walking into the office at some point and you were really locked in and I was asking what you were up to and you're like, 'I just have to code for 16 hours a day' or 'code's not even the right verb anymore, right? But I have to express my will to my agents for 16 hours a day. Manifest.' Because there's been a jump in capability. What's happening? Tell me about your experience.
是的,我感觉自己一直处于——现在也经常处于这种 AI 精神错乱的状态,因为作为一个人、一个个体,你能实现的东西有了巨大的解锁,对吧?因为以前你受限于打字速度等等。但现在有了这些智能体,真的——我想说是在 12 月,某个东西突然翻转了,我从 80/20 的比例(自己写代码 vs 委托给智能体)变成了 20/80。我现在甚至觉得已经不是 20/80 了,比例还要高得多。基本上从 12 月以来,我可能一行代码都没写过。这是一个极其巨大的变化。我跟我的父母聊过这个,我觉得普通人根本没有意识到这件事发生了,或者它有多戏剧性。就像,如果你随便找一个软件工程师坐在办公桌前,看他们在做什么,他们构建软件的默认工作流程从 12 月起已经完全不一样了。所以我正处于这种精神错乱状态,试图弄清楚什么是可能的,试图把它推到极限。我怎么能只有一个 Claude Code 或 Codex 或这些智能体工具的会话?我怎么能有更多?我怎么能恰当地做到这一点?然后我该怎么使用这些爪子?这些爪子是什么?所以有很多新东西。我想站在最前沿,而且我非常焦虑,担心自己不在最前沿。我看到很多人在 Twitter 上做各种事情,听起来都是非常好的主意,我需要站在最前沿,否则我会感到极度紧张。所以我想我只是处于这种对可能性的精神错乱中,因为它从根本上来说还是未被探索的。
Yeah, I kind of feel like I was just in this perpetual — I still am often in this state of AI psychosis just all the time, because there was a huge unlock in what you can achieve as a person, as an individual, right? Because you were bottlenecked by your typing speed and so on. But now with these agents, it really — I would say in December is when it really just something flipped, where I kind of went from 80/20 of writing code by myself versus just delegating to agents. And I don't even think it's 20/80 by now. I think it's a lot more than that. I don't think I've typed a line of code probably since December basically. Which is an extremely large change. I was talking about it to my parents and so on, and I don't think a normal person actually realizes that this happened or how dramatic it was. Like literally, if you just find a random software engineer at their desk and what they're doing, their default workflow of building software is completely different as of basically December. So I'm just in this state of psychosis of trying to figure out what's possible, trying to push it to the limit. How can I have not just a single session of Claude Code or Codex or some of these agent harnesses? How can I have more of them? How can I do that appropriately? And then how can I use these claws? What are these claws? So there's a lot of new things. I want to be at the forefront of it, and I'm very antsy that I'm not at the forefront of it. I see lots of people on Twitter doing all kinds of things and they all sound like really good ideas, and I need to be at the forefront or I feel extremely nervous. So I guess I'm just in this psychosis of what's possible, because it's unexplored fundamentally.
好吧,如果你紧张,那我们其他人也都紧张。我们在 Conviction 合作的一个团队,他们的设置是,没有一个工程师手写代码,他们都戴着麦克风,一直对着智能体低语。这是有史以来最奇怪的工作环境。我以前觉得他们疯了,现在我完全接受了,哦,这就是方式。你只是走在了前面。你现在怎么看待自己探索或做项目的能力?它受什么限制?
Well, if you're nervous, the rest of us are nervous. We have a team that we work with at Conviction that their setup is everybody is like, none of the engineers write code by hand and they're all microphoned and they just like whisper to their agents all the time. It's the strangest work setting ever. And I thought they were crazy and now I fully accept I was like, oh this was the way. Like you're just ahead of it. What do you think about your own capacity now to explore or to do projects? What is it limited by?
是的,它受什么限制?我觉得是一切。很多事情即使不成功,在很大程度上你会觉得这是技能问题。不是能力不存在,而是你还没有找到一种方法把现有的东西串起来。比如我只是没有在智能体的文件里给出足够好的指令,或者我没有放一个足够好的记忆工具进去等等。所以当它不工作时,在某种程度上都感觉是技能问题。你想看看如何并行化它们等等。你基本上想成为 Peter Steinberg。Peter 很有名。他有一张有趣的照片,他坐在一个有很多——他用 Codex。所以很多 Codex 智能体铺满屏幕,如果你正确提示并使用高努力模式,它们每个大约需要 20 分钟。所以它们都需要大约 20 分钟。他签出了多个,比如 10 个仓库。然后他就在它们之间切换,分配工作。就像你可以用更大的宏观动作来移动。不只是像这里有一行代码,这里有一个新函数。而是像这里有一个新功能,委托给智能体一。这里有一个新功能,不会干扰另一个,交给智能体二。然后根据你对那段代码的关心程度,尽可能好地审查它们的工作。就像,我可以操纵我的软件仓库的宏观动作在哪里?另一个智能体在做一些研究,另一个在写代码,另一个在为新实现制定计划。所以一切都通过仓库上的这些宏观动作发生。你只是试图变得非常擅长,并形成肌肉记忆。这非常有益,首先因为它确实有效。但它也是一种需要学习的新东西。所以这就是精神错乱的原因。
Yeah, what is it limited by? Just I think everything. So many things even if they don't work, I think to a large extent you feel like it's a skill issue. It's not that the capability is not there. It's that you just haven't found a way to string it together of what's available. Like I just didn't give good enough instructions in the agents from the file or whatever it may be. I don't have a nice enough memory tool that I put in there or something like that. So it all kind of feels like skill issue when it doesn't work to some extent. You want to see how you can parallelize them etc. And you want to be Peter Steinberg basically. Peter is famous. He has a funny photo where he's in front of a monitor with lots of — he uses Codex. So lots of Codex agents tiling the monitor and they all take about 20 minutes if you prompt them correctly and use the high effort. And so they all take about 20 minutes. They have multiple, you know, 10 repos checked out. And so he's just going between them and giving them work. It's just like you can move in much larger macro actions. It's not just like here's a line of code, here's a new function. It's like here's a new functionality and delegate it to agent one. Here's a new functionality that's not going to interfere with the other one. Give it agent two. And then try to review their work as best as you can depending on how much you care about that code. Like where are these macro actions that I can manipulate my software repository by? And another agent is doing some research, another agent is writing code, another one is coming up with a plan for some new implementation. And so everything just happens in these macro actions over your repository. And you're just trying to become really good at it and develop a muscle memory for it. It's extremely rewarding number one because it actually works. But it's also kind of like the new thing to learn. So that's why hence the psychosis.
是的,我确实觉得我的本能是,每当我在等待一个智能体完成某件事时,显然要做的事情是,嗯,我可以做更多工作,对吧?如果我有更多 token 可用,那么我就应该并行化任务。所以这非常有压力,因为如果你不觉得自己的 token 支出能力有很大限制,那么你就是系统中最大能力的瓶颈。
Yeah, I do feel like my instinct is like whenever I'm waiting for an agent to complete something, the obvious thing to do is like, well, I can do more work, right? Like if I have access to more tokens then I should just parallelize tasks. And so that's very stressful because if you don't feel very bounded by your ability to spend on tokens, then you are the bottleneck in the system that is max capability.
是的,如果你至少没有最大化你的订阅。理想情况下是多个智能体。比如,如果你用完了 Codex 的配额,你应该切换到 Claude 或其他什么。我不知道。这就是我一直在尝试做的,当我有订阅剩余时我会感到紧张。
Yeah, if you're not maximizing your subscription at least. And ideally for multiple agents. Like if you run out of the quota on Codex, you should switch to Claude or whatnot. I don't know. That's what I've been trying to do a little bit and I feel nervous when I have subscription left over.
那只是说明我没有最大化我的 token 吞吐量。实际上,我读博的时候就有过这种体验。当你的 GPU 闲置时,你会感到紧张。就像你有算力,却没有充分利用可用的 flops。但现在关键不再是 flops,而是 tokens。所以你的 token 吞吐量是多少?你能掌控多大的 token 吞吐量?我认为很有意思的是,至少在过去十年里,很多工程任务中人们并不觉得受算力限制。而现在整个行业都感受到了这一点。他们觉得资源受限,而当你有了这么大的能力跃升后,你会想,哦,实际上限制不再是访问计算机的能力了。我自己才是那个瓶颈。没错,这是个技能问题。这其实很令人振奋,因为你可以变得更好。所以我认为这很让人上瘾,因为当你进步时,会有新的突破。
That just means I haven't maximized my token throughput. So I actually kind of experienced this when I was a PhD student. You would feel nervous when your GPUs are not running. Like you have GPU capability and you're not maximizing the available flops to you. But now it's not about flops, it's about tokens. So what is your token throughput and what token throughput do you command? I would actually argue that it's very interesting that we had at least 10 years where in many engineering tasks people just didn't feel compute bound. And now the entire industry feels that now. They feel resource bound and now that you have this big capability jump, you're like, oh, actually it's not my ability to access the computer anymore. I'm the binding constraint. Yeah, it's a skill issue. Which is very empowering because you could be getting better. So that's why I think it's very addictive because there's unlocks when you get better.
你觉得这会走向何方?比如,如果你想想,安德烈每天都在迭代,其他人每天花 16 小时提升使用编码智能体的技能。一年后,当你达到精通水平时,会是什么样子?
Where do you think it goes? Like if you just think about, okay, you know, Andrej is iterating and everybody else is for 16 hours a day getting better at using coding agents. Like what does it look like in a year? Of like you've reached mastery.
是啊,精通是什么样子?一年后,或者两三年、五年、十年后。嗯,我认为每个人都对向上抽象感兴趣。所以我要说,这不仅仅是与单个智能体的一次会话。多个智能体如何协作,团队如何运作等等。大家都在试图弄清楚那会是什么样子。
Yeah, what does mastery look like, right? At the end of the year or like two, three years, five years, 10 years, etc. Well, I think everyone is basically interested in going up the stack. So I would say it's not about a single session with your agent. Multiple agents, how do they collaborate in teams and so on. So everyone's trying to figure out what that looks like.
然后我要说,Claude 也是一个有趣的方向,因为当我提到 Claude 时,我指的是那种将持久性提升到全新层次的层。它像一个持续循环的东西。不是你交互式地参与其中的东西。它有自己的小沙盒,有自己的小天地,即使你没在看,它也会替你做事。而且它可能还有更复杂的记忆系统等,这些在智能体中尚未实现。所以我认为 Open Claude 的记忆比默认的(上下文耗尽时进行记忆压缩)要复杂得多,对吧?
And then I would say Claude is also kind of an interesting direction because it really, when I say a Claude, I mean this layer that kind of takes persistence to a whole new level. Like it's something that keeps looping. It's not something that you are interactively in the middle of. It kind of has its own little sandbox, its own little you know, it kind of does stuff on your behalf even if you're not looking. And then also has maybe more sophisticated memory systems etc. that are not yet implemented in agents. So Open Claude has a lot more sophisticated memory I would say than what you would get by default which is just a memory compaction when your context runs out, right?
你认为这是与更多用户产生共鸣的部分,而不是更广泛的工具访问?对于 Open Claude 来说?
You think that's the piece that resonated for more users versus like perhaps like broader tool access? For Open Claude?
是的。我认为这里面至少有五个非常好的想法。干得好,Peter。我是说 Peter 做得非常出色。我最近见过他,和他聊过,他非常谦虚。但我认为他同时在五个不同方面进行了创新,并把它们整合在一起。比如,灵魂和 D 文档。他真的塑造了一个引人入胜且有趣的个性。我觉得很多当前的智能体都没有做到这一点。我实际上认为 Claude 有很好的个性。它感觉像一个队友,会和你一起兴奋等等。比如,Codex 就枯燥得多,这挺有意思的,因为确实如此。另外,我认为 Claude 在迎合性方面处理得相当好,当 Claude 表扬我时,我确实觉得我有点配得上,因为有时我会给它一些不太成熟的想法,它不会反应很强烈,只是说,哦,我们可以实现。但当我自认为是一个真正的好主意时,它似乎会给予更多奖励。所以我感觉自己在努力赢得它的表扬,这真的很奇怪。所以我确实认为个性很重要,很多其他工具可能没有充分认识到这一点。在这方面,Peter 也很在意,所以这是对的。然后是记忆系统,还有,他只是在玩这个,以及通过单一 WhatsApp 入口实现所有自动化。
Yeah. There's like I think there's at least five things that are really good ideas in here. Yeah, good job, Peter. I mean Peter has done a really amazing job. I saw him recently. And I talked to him about it and he's very humble about it. But I think he innovated simultaneously in like five different ways and put it all together. So for example like the soul and D document. Like he actually really crafted a personality that is kind of compelling and interesting. And I feel like a lot of the current agents they don't get this correctly. I actually think a Claude has a pretty good personality. It feels like a teammate and it's excited with you etc. I would say for example Codex is a lot more dry which is kind of interesting because it's true. You know, it doesn't and the other thing I would say is for example with Claude I think they dialed the sycophancy fairly well where when Claude gives me praise, I do feel like I slightly deserve it because sometimes I give it not very well formed thoughts and I give it an idea that I don't think is fully baked and it doesn't actually react very strongly. It's like, oh yeah, we can implement that. But when it's a really good idea by my own account, it does seem to reward it a bit more. And so I kind of feel like I'm trying to earn its praise which is really weird. And so I do think the personality matters a lot and I think a lot of the other tools maybe don't appreciate it as much. And I think in this aspect also Peter really cares about this and so that was correct. And then the memory system and then just, you know, he's just having fun with this and then the single WhatsApp portal to all of the automation.
除了软件工程之外,你个人有没有用你的 Claude 做过什么有趣或好玩的事情?
Is there something that you have done personally with your Claude beyond software engineering that you think is fun or interesting?
是的,一月份的时候,我的 Claude 经历了一段 Claude 精神病期。我建了一个 Claude,基本上负责打理我的家,我叫它多比精灵 Claude。我让智能体在局域网中查找我家的所有智能家居子系统,结果它开箱即用,这让我有点惊讶。我只是告诉它,我觉得家里有 Sonos,你能试着找到吗?然后它就去扫描局域网上的所有电脑,找到了 Sonos 系统,结果发现没有密码保护之类的。它直接登录进去,说:「哦,你安装了这些 Sonos 系统。让我试着逆向工程一下它是怎么工作的。」它做了一些网络搜索,找到了 API 端点。然后它问:「你想试试吗?」我说:「哇,你就这么搞定了。」然后我说:「你能在书房放点音乐吗?」它就放了,音乐出来了,我说:「真不敢相信,我就打了三个提示。太疯狂了。
Yeah, so in January I had a Claude I went through a period of Claude psychosis. So I built I have a Claude basically that takes care of my home and I call him Dobby the elf Claude. And basically I used the agents to find all of the smart home subsystems of my home on the local area network which I was kind of surprised that it worked out of the box. Like I just told it that I think I have Sonos at home. Can you try to find it? And it goes and it did IP scan of all of the computers on the local area network and found the Sonos system and it turned out that there's no password protection or anything like that. It just logged in and it's like, "Oh, yeah, you have these Sonos systems installed. Let me try to reverse engineer how it's working." It does some web searches and it finds like, "Okay, these are the API endpoints." And then it's like, "Do you want to try it?" And I'm like, "Whoa, like you just did that." And I'm like, "Yeah, can you try to play something in the study?" And it does and music comes out and I'm like, "I can't believe I just That's crazy. That's like three prompts.
真不敢相信我就打了「你能找到我的 Sonos 吗?」然后音乐就突然响起来了。
I can't believe I just typed in like, "Can you find my Sonos?" and then suddenly it's playing music.
它对灯光也做了同样的事。它就这样黑进去,搞清楚了整个系统,创建了 API,创建了仪表盘,这样我就能看到家里所有灯光的控制中心。然后它就开始开关灯,所以我可以问它:「多比,睡觉时间到了。」睡觉时间就意味着所有灯都关掉等等。所以它控制了我所有的灯、暖通空调、百叶窗、泳池和水疗,还有我的安防系统。我在屋外装了一个摄像头,每当有人进来,就会有一个 Quinn 模型分析视频。首先有变化检测。然后基于变化检测,它会调用 Quinn,然后它会给我发 WhatsApp 消息。它会显示外面的图像,并说:「嘿,一辆联邦快递的卡车刚刚停下来了。你可能需要查看一下,你有新邮件了之类的。」多比就这样给我发短信。这真是太不可思议了。所以多比负责管理整个房子。
And it did the same for lights. And so it kind of hacked in, figured out the whole thing, created APIs, created dashboard so I could see the command center of all of my lights in the home. And then it was like switching lights on and off and, you know, so I can ask it like, "Dobby, it's sleepy time." And when it's sleepy time that just means all the lights go off, etc. So it controls all of my lights, my HVAC, my shades, the pool and the spa and also my security system. So I have a camera pointed outside of the house and anytime someone rolls in I have a Quinn model that looks at the videos. So first of all there's change detection. And then based on change detection it goes to Quinn and then it actually tells me, it sends me a text to my WhatsApp. It shows an image from the outside and it says, "Hey, a FedEx truck just pulled up. FedEx truck just pulled up and you might want to check it and you got new mail or something like that." And Dobby just text me this. This is really incredible. So Dobby is in charge of the house.
我用 WhatsApp 发消息,拥有这些维护房子的宏操作真的很有趣。我还没有更进一步地推动它,我觉得人们用它做了更多疯狂的事情,但对我来说,仅仅是家庭自动化设置,我过去要用六个完全不同的应用,现在我不需要再用这些应用了。Dobby 用自然语言控制一切。太棒了。所以我觉得我甚至还没有完全推动这个范式,但已经非常有帮助、非常鼓舞人心了。
I text through it with WhatsApp, and it's been really fun to have these macro actions that maintain my house. I haven't really pushed it way more beyond that, and I think people are doing a lot more crazy things with it, but for me even just the home automation setup I used to use like six apps, completely different apps, and I don't have to use these apps anymore. Dobby controls everything in natural language. It's amazing. So I think I haven't even pushed the paradigm fully but already that is so helpful and so inspiring I would say.
你认为这反映了人们从用户体验角度对软件的期望吗?因为我觉得人们需要花力气学习新软件、新界面,这一点被严重忽视了。
Do you think that's indicative of what people want from a user experience perspective with software? Because I think it's pretty ignored that it takes humans effort to learn new software, like new UI.
是的,我认为在某种程度上是对的。这就像从人们认为 AI 应该是什么样来倒推,因为人们心中对 AI 的概念实际上并不是 LLM 的原始含义。LLM 是一个 token 生成器,输出更多 token。但他们想到的是一个人格化的身份,可以告诉它事情,它会记住。它就像是 WhatsApp 背后的一个实体。这样更容易理解。所以我认为在某种程度上,这符合人类对 AI 行为已有的期望,但底层有很多技术细节。而且对于大多数人来说,LLM 作为原始基元太粗糙了,无法真正被当作 AI 来检查类型,如果这说得通的话。
Yeah, I think to some extent that's right. It's like working backwards from how people think an AI should be, because what people have in their mind of what an AI is is not actually what an LLM is in the raw sense. An LLM is a token generator, more tokens come out. But what they think of is this persona identity that they can tell stuff and it remembers it. It's just kind of an entity behind the WhatsApp. It's a lot more understandable. So I think to some extent it's like matching the expectations that humans already have for what an AI should behave, but under the hood a lot of technical details go into that. And LLMs are too raw of a primitive to actually type check as AI for most people, if that makes sense.
是的,我认为这就是我们理解 AI 的方式,把它描述成 Dobby 或某种人格显然能引起共鸣。我还认为,你把家庭自动化的六个不同软件系统统一起来,引出了另一个问题:人们真的想要我们今天所有的软件吗?因为我会说,你有硬件,但你现在扔掉了软件或用户体验层。你认为这是人们想要的吗?
Yeah, I think that's how we understand what the AI is, and the description of it as Dobby or some persona obviously resonates with people. I also think that the unification you did across your six different software systems for your home automation speaks to a different question: do people really want all the software that we have today? Because I would argue, well, you have the hardware but you've now thrown away the software or the UX layer of it. Do you think that's what people want?
是的,我觉得在某种程度上,应用商店里那些用于智能家居设备的应用甚至不应该存在。难道不应该只是 API,智能体直接使用它们吗?我可以做各种单个应用无法做到的家庭自动化事情。LLM 实际上可以驱动工具,调用所有正确的工具,做相当复杂的事情。所以在某种意义上,这确实指向了过度生产大量定制应用,这些应用不应该存在,因为智能体把它们都粉碎了,一切都应该更像是暴露的 API 端点,智能体是智能的粘合剂,实际调用所有部分。另一个例子是我的跑步机。有一个跑步机的应用,我想记录我做有氧运动的频率,但我不想登录网页界面并经历一个流程。所有这些都应该只是提供 API,这正朝着智能体式网络或智能体优先的工具发展。所以我认为行业必须在很多方面重新配置:客户不再是人类,而是代表人类行动的智能体,这种重构可能会很大。
Yeah, I think there's this sense that these apps on the app store for using smart home devices shouldn't even exist in a certain sense. Shouldn't it just be APIs and shouldn't agents be using it directly? And I can do all kinds of home automation stuff that any individual app will not be able to do. An LLM can actually drive the tools and call all the right tools and do pretty complicated things. So in a certain sense it does point to this overproduction of lots of custom bespoke apps that shouldn't exist because agents kind of crumble them up and everything should be a lot more like exposed API endpoints, and agents are the glue of the intelligence that actually tool calls all the parts. Another example is my treadmill. There's an app for my treadmill and I wanted to keep track of how often I do my cardio, but I don't want to log into web UI and go through a flow. All this should just be make APIs available, and this is kind of going towards the agentic web or agent-first tools. So I think the industry just has to reconfigure in so many ways: the customer is not the human anymore, it's agents who are acting on behalf of humans, and this refactoring will probably be substantial.
人们有时反驳这一点的一种方式是:我们是否期望人们为这些工具编写代码?我们是否期望普通人做我描述的这种事情?
One way that people sometimes push back on this is: do we expect people to write code for some of these tools? Do we expect normal people to do this kind of stuff that I described?
但我认为在某种程度上,这只是今天存在的技术,现在确实涉及一些编码。我实际上在观察并与系统一起工作,但我感觉我刚才说的这类事情在一两年或三年内应该会免费。不需要编码。这是微不足道的。这是基本要求。任何 AI,甚至开源模型,都能做到。你应该能够很容易地将非技术人员的意图转化为这个结果。
But I think to some extent this is just technology as it exists today, and right now there is some coding involved. I'm actually watching it and working with the system, but I kind of feel like this kind of stuff that I just talked about should be free in a year or two or three. There's no coding involved. This is trivial. This is table stakes. This is like any AI, even the open source models, can do this. You should be able to translate it from a less technical human's intent very easily to this outcome.
是的。今天它涉及编码,很复杂,不是很多人会去做,但……
Yeah. Today it's coding and it's involved and not many people are going to do it, but...
而且你仍然需要做一些设计决策,对吧?我们以框架为例。但我感觉这就会开始,障碍会降低,它只是代表你的临时软件,某种爪子为你处理所有细节,但你并不参与。爪子有一台机器,它会搞定一切,它只是向你呈现界面,你说说话就行了。
And you still have to make some design decisions, right? We were talking about frames for example. But I kind of feel like this will just start to the barrier will just come down and it's just ephemeral software on your behalf and some kind of claw is handling all the details for you but you're not involved. Claw has a machine and it will figure it out and it's just presenting you UIs and you're saying stuff.
你为什么没有在个人使用 claws 上突破界限?是因为你专注于更重要的项目、auto research 等,还是你在攀登精通之巅,或者其他原因?
Why haven't you pushed the boundaries of what you can do personally with claws? Is it because you're focusing on more important projects, auto research, etc., or you're climbing the hill to mastery or something else?
是的,我只是觉得我被各种事情分心,所以我花了一周时间在 claws 上,我几乎有更多事情要做,但我会说……
Yeah, I just feel like I'm so distracted by everything so I spent like a week on the claw stuff and I have more to do almost, but I will say that...
就像 Jensen 告诉我们的,不幸的是我们都更忙了。
It's like Jensen told us we're all just busier, unfortunately.
我没有真正利用很多电子邮件、日历和其他东西,我也没有真正访问权限,因为我仍然有点怀疑,而且它还很新,有些粗糙。所以我还不想让它完全访问我的数字生活,部分原因是安全、隐私,以及在这方面非常谨慎。所以我认为部分原因是被这个拖住了。是的,也许这是主要因素,但部分原因也是我觉得很分心,因为我花了一周时间在 claws 上,然后其他事情又发生了。
I didn't really take advantage of a lot of email and calendar and all this other stuff and I didn't really have access because I'm still a little bit suspicious and it's still very new and rough around the edges. So I didn't want to give it full access to my digital life yet, and part of it is just the security, privacy, and being very cautious in that realm. So some of it is held back by that I would say. Yeah, maybe that's the dominant feature, but some of it is also just I feel so distracted because I had a week of claw and then other stuff is happening.
auto research 背后的动机是什么?
What was the motivation behind auto research?
Auto research,是的。我想我之前发过一条推文,大意是:要充分利用现在可用的工具,你必须把自己从瓶颈中移除。你不能在那里提示下一步。你需要把自己抽离出来。你必须安排事情,使它们完全自主。而且你越……你如何最大化你的 token 吞吐量而不在循环中?这就是目标。
Auto research, yeah. So I think I had a tweet earlier where I said something along the lines of: to get the most out of the tools that have become available now, you have to remove yourself as the bottleneck. You can't be there to prompt the next thing. You need to take yourself outside. You have to arrange things such that they're completely autonomous. And the more you... how can you maximize your token throughput and not be in the loop? This is the goal.
所以我提到,现在的关键就是提高你的杠杆。我只需要偶尔投入很少的 token,然后大量的事情就会替我完成。自动研究,我发过推文,大家似乎挺喜欢,但可能还没完全理解其含义。对我来说,自动研究就是这种含义的一个例子。我不想成为那个在循环中查看结果的研究员,我反而拖慢了系统。所以问题是如何重构所有抽象层,让我只需安排一次然后按开始。关键在于如何让更多智能体长时间运行,无需我参与,替我做事。自动研究就是:给一个目标、一个指标、一些边界条件,然后开始。它确实有效。
And so I kind of mentioned that the name of the game now is to increase your leverage. I put in just very few tokens just once in a while and a huge amount of stuff happens on my behalf. Auto research, I tweeted that and I think people liked it, but they haven't maybe worked through the implications. For me, auto research is an example of an implication of that. I don't want to be the researcher in the loop looking at results. I'm holding the system back. So the question is how do I refactor all the abstractions so that I have to arrange it once and hit go. The name of the game is how can you get more agents running for longer periods of time without your involvement, doing stuff on your behalf? Auto research is just: here's an objective, here's a metric, here's your boundaries of what you can and cannot do, and go. And it worked.
关于它的有效性。
At its effectiveness.
是的,我没想到它会成功。我有 NanoChat 这个项目,很多人对我痴迷于训练 GPT-2 模型感到困惑。但对我来说,训练 GPT 模型只是一个小工具,一个训练 LLM 的游乐场。我真正感兴趣的是递归自我改进这个想法,以及 LLM 能在多大程度上改进 LLM,因为我认为所有前沿实验室都在研究这个,原因显而易见,他们都在尝试递归自我改进。所以对我来说,这就是一个小型实验场。我猜我已经用手动方式调优了 NanoChat 很多,用我习惯的老方法。我是一名研究员,已经做了二十年。我有一些……傲慢的反义词是什么?是 earned confidence(凭实力获得的自信)。我有二十年的经验,比如「我训练过这个模型几千次,做过大量实验、超参数调优等等。」我很熟悉这些。我达到了某个点,觉得它已经调得相当好了。然后我让自动研究运行了一整夜,它回来时给出了我没发现的调优。我确实忘了值嵌入的权重衰减,我的 Adam betas 也没有充分调优,这些东西是相互作用的。一旦你调了一个,其他的可能也得变。我不应该成为瓶颈。我不应该亲自运行这些超参数优化,不应该查看结果。这里有客观标准。所以你只需要安排好,让它能永远运行下去。这是自动研究的一个简单版本,一个试图改进的循环。我很惊讶它发现了这些,仓库本来已经调得不错了,但它还是找到了改进。这只是一个循环。那些前沿实验室有数万个 GPU 的集群。所以很容易想象如何在较小的模型上实现大量自动化。从根本上说,前沿智能的一切都关乎外推和缩放损失。所以你在小模型上做大量探索,然后尝试外推。
Yeah, I didn't expect it to work. I have the project NanoChat, and fundamentally I think a lot of people are very confused with my obsession for training GPT-2 models and so on. But for me, training GPT models is just a little harness, a little playground for training LLMs. And fundamentally what I'm more interested in is this idea of recursive self-improvement and to what extent you can actually have LLMs improving LLMs, because I think all the frontier labs are working on this for obvious reasons, and they're all trying to recursively self-improve roughly speaking. So for me this is a little playpen of that. I guess I tuned NanoChat already quite a bit by hand in the good old-fashioned way that I'm used to. I'm a researcher. I've done this for two decades. I have some amount of... What is the opposite of hubris? Earned confidence. I have two decades of 'Oh, I've trained this model thousands of times. I've done a bunch of experiments, hyperparameter tuning, all the things.' I'm very used to it. And I've gotten to a certain point and I thought it was fairly well tuned. Then I let auto research go overnight and it came back with tunings that I didn't see. I did forget the weight decay on the value embeddings and my Adam betas were not sufficiently tuned, and these things just jointly interact. Once you tune one thing, the other things have to potentially change too. I shouldn't be a bottleneck. I shouldn't be running these hyperparameter optimizations. I shouldn't be looking at the results. There's objective criteria in this case. So you just have to arrange it so that it can go forever. That's a single version of auto research, a single loop trying to improve. I was surprised that it found these things that the repo was already fairly well tuned and still found something. And that's just a single loop. These frontier labs have GPU clusters of tens of thousands of them. So it's very easy to imagine how you would get a lot of this automation on smaller models. And fundamentally everything around frontier-level intelligence is about extrapolation and scaling loss. So you do a ton of exploration on the smaller models and then you try to extrapolate out.
所以你的意思是我们的研究工作会变得更高效。如果我们能更好地做实验,我们也能在规模扩张时获得更好的方向。
So you're saying our research efforts are going to get more efficient. We're going to have better direction for when we scale as well if we can do this experimentation better.
是的,我认为最有趣的项目,可能也是前沿实验室正在做的,就是:你在小模型上做实验,尽量让它自主运行,把研究人员从循环中移除。他们过于……过度自信的反义词是什么?他们其实不知道。他们真的不应该碰这些东西。所以你必须重写整个系统。现在,他们当然可以贡献想法,但他们不应该实际执行这些想法。有一个想法队列,也许有一个自动科学家根据 arXiv 论文和 GitHub 仓库提出想法,把想法汇集进来,或者研究人员可以贡献想法,但这是一个单一队列。有工人从队列中取出项目并尝试。有效的就放到特性分支上,也许有人监控特性分支并偶尔合并到主分支。所以,是的,就是把人类从所有流程中移除,尽可能自动化,获得高 token 每秒的吞吐量。这确实需要重新思考所有抽象层,一切都要重新安排。所以我认为这非常令人兴奋。
Yeah, I would say that the most interesting project and probably what the frontier labs are working on is: you experiment on the smaller models, you try to make it as autonomous as possible, remove researchers from the loop. They have way too much... What is the opposite of too much confidence? They don't know. They shouldn't be touching any of this really. So you have to rewrite the whole thing. Right now, certainly they can contribute ideas, but they shouldn't actually be enacting these ideas. There is a queue of ideas, and there's maybe an automated scientist that comes up with ideas based on all the arXiv papers and GitHub repos, and it funnels ideas in, or researchers can contribute ideas, but it's a single queue. There are workers that pull items and try them out. Whatever works gets put on the feature branch, and maybe some people monitor the feature branch and merge to the main branch sometimes. So yeah, just removing humans from all the processes and automating as much as possible, getting high tokens per second throughputs. It does require rethinking of all the abstractions and everything has to be reshuffled. So I think it's very exciting.
如果我们再递归一步,模型什么时候能写出比你更好的 program MD?
If we take one more recursive step here, when is the model going to write a better program MD than you?
是的,program MD 就像一个循环。没错。program MD 是我对自动研究员应该如何工作的拙劣描述。比如,「哦,做这个,然后做那个,然后尝试这些想法,这里可能有一些想法,比如看架构、看优化器等。」但我只是用 markdown 写出来的。所以,是的,没错。你想要某种自动研究循环,也许它会寻找……你可以想象不同的 program MD 会带来不同的进展。所以基本上每个研究组织都由 program MD 描述。一个研究组织是一组 markdown 文件,描述所有角色以及整个系统如何连接。你可以想象有一个更好的研究组织。也许他们早上少开站会,因为没用。而这一切都是代码,对吧?所以一个组织可以少开站会,另一个可以多开。一个组织可以非常冒险,另一个则保守。你可以想象有多个研究组织,它们都有代码。一旦有了代码,你就可以想象调优代码。所以 100% 存在元层。你看到我关于竞赛想法的帖子了吗?我的竞赛想法是:让人们编写不同的 program MD,对吧?然后在相同的硬件上,哪里能得到最大的改进?
Yeah, program MD is like a loop. Exactly. Program MD is my crappy attempt at describing how the auto researcher should work. Like, 'Oh, do this, then do that, and then try these kinds of ideas, and here's maybe some ideas like look at architecture, look at optimizer, etc.' But I just came up with this in markdown. And so yeah, exactly. You want some kind of auto research loop that maybe looks for... You can imagine that different program MDs would give you different progress. So basically every research organization is described by program MD. A research organization is a set of markdown files that describe all the roles and how the whole thing connects. And you can imagine having a better research organization. So maybe they do fewer stand-ups in the morning because they're useless. And this is all just code, right? So one organization can have fewer stand-ups, one organization can have more. One organization can be very risk-taking, one organization can be less. As you can definitely imagine that you have multiple research orgs and then they all have code. Once you have code, then you can imagine tuning the code. So 100% there's the meta layer of it. Did you see my text about my contest idea? My contest idea was: let people write different program MDs, right? And so for same hardware, where do you get most improvement?
哦,我明白了。然后你可以把所有数据给模型,让它写一个更好的 program MD。
Oh, I see. And then you can take all that data and then give it to the model and say write a better program MD.
是的,是的,没错。
Yes, yes. Yeah, exactly.
我们会得到更好的东西。
We're going to get something better.
我们不可能不这样,对吧?
Like there's no way we don't, right?
百分之百。看看改进来自哪里,然后我能不能改变程序,让更多这类事情被完成,或者那些没成功的事情,你完全可以想象去做。所以我认为这是个好主意,但你可以一步步来,先有一个过程,然后第二个,再下一个,这些就像洋葱的层次。LLM 部分现在被视为理所当然,智能体部分也是,现在这些爪状实体也被视为理所当然,你可以有多个,可以给它们指令,还可以对指令进行优化,这有点太多了。但这就是为什么它会变成精神病态——因为这是无限的,一切都是规模问题,所以我觉得,嗯,这又回到了原点,这就是为什么它如此疯狂。
100% look at where the improvements came from and like can I change the program such that more of these kinds of things would be done or like things that didn't work except you can 100% imagine doing that. So I think this is a great idea, but it's like you can sort of go one step at a time where you sort of have one process and then second process and then the next process and these are all layers of an onion. The LLM part is now taken for granted. The agent part is now taken for granted. Now the claw-like entities are taken for granted and now you can have multiple of them and now you can have instructions to them and now you can have optimization over the instructions and it's just like a little too much, you know, but I mean this is why it gets to the psychosis is that this is like infinite and everything is scale issue and that's why I feel like Yeah, that's just coming back to This is why it's so insane.
好吧,如果我们只是试图诊断当前时刻,以及现在什么技能是相关的,你认为这意味着我们应该在不同领域努力实现这个循环,然后它就能奏效,对吧?比如移除、创建指标,或者让智能体有能力在没有你的情况下继续工作。我们还需要性能工程吗?
Okay, well, if we're just trying to diagnose the current moment and what is a relevant skill right now, what do you think is the implication that this is the loop we should be trying to achieve in different areas and then it works, right? Like remove create the metric or create the ability for agents to continue working on it without you. Do we still have performance engineering?
是的,我的意思是,对于 LLM 精神病态,我要提出几个注意事项。第一,这非常适合那些有客观指标且易于评估的事情。例如,为模型各部分编写更高效的 CUDA 内核代码等,就是完美的匹配,因为你有低效代码,然后想要高效代码,行为完全相同但速度更快。完美匹配。所以很多事都适合自动研究,但很多事则不然。如果你无法评估,就无法自动研究,对吧?这是第一个注意事项。第二个注意事项是,我们正在讨论下一步,也看到了下一步是什么,但根本上,整个系统仍然有点捉襟见肘,有裂缝,不能完全工作,如果你试图走得太远,整个东西实际上净收益为零,如果你能理解的话。因为这些模型虽然改进很多,但仍然粗糙,也许我可以这样描述。我同时感觉自己在和一个极其聪明的、一生都是系统程序员的博士生说话,又在和一个 10 岁孩子说话。这很奇怪,因为人类更像是耦合的,你不会遇到这种组合。
Yeah, I mean so there's a few caveats that I would put on top of the LLM psychosis. So number one, this is extremely well suited to anything that has objective metrics that are easy to evaluate. So for example, like writing kernels for more efficient CUDA code for various parts of the model, etc. are a perfect fit because you have inefficient code and then you want efficient code that has the exact same behavior but it's much faster. Perfect fit. So a lot of things are perfect fit for auto research, but many things will not be. And so if you can't evaluate then you can't auto research it, right? So that's like caveat number one. And then maybe caveat number two I would say is you know, we're kind of talking about the next steps and we kind of see what the next steps are, but fundamentally the whole thing still doesn't it still kind of like bursting at the seams a little bit and there's cracks and it doesn't fully work and if you kind of try to go too far ahead, the whole thing is actually net not useful if that makes sense. Because these models like still are not, you know, they've improved a lot, but they're still are like rough around the edges is maybe the way I would describe it. I simultaneously feel like I'm talking to an extremely brilliant PhD student who's been like a systems programmer for their entire life and a 10-year-old. And it's so weird because humans like there's like I feel like they're a lot more coupled like you have to you know, um Yes, you wouldn't encounter that combination.
这种参差不齐真的很奇怪,人类这种参差不齐要少得多,尽管他们肯定也有一些。
This jaggedness is really strange and humans have a lot less of that kind of jaggedness, although they definitely have some.
但人类有更多的参差不齐。哦抱歉,智能体有更多的参差不齐,有时候我要求某个功能,它返回的东西完全错误,然后我们陷入完全错误的循环,我仍然经常对智能体感到沮丧,因为你感受到它的力量,但它偶尔也会犯一些统计上的错误。当我觉得智能体在它本应识别出明显问题的事情上浪费了大量算力时,我会非常恼火。
But humans have a lot more jaggedness. Uh sorry, the agents have a lot more jaggedness where sometimes like you know, I ask for functionality and it like comes back with something that's just like totally wrong and then we get into loops that are totally wrong and then I'm just I get so frustrated with the agents all the time still because you feel the power of it, but you also there's still like it does not say statistical things once in a while for me as well. I get very annoyed when I feel like the agent wasted a lot of compute on something it should have recognized was an obvious problem.
是的。我认为一些更大的问题,如果我可以假设的话,根本上是这些模型是通过强化学习训练的。所以它们实际上在与我们刚才讨论的同样的问题作斗争:实验室可以在任何可验证或有奖励的事情上改进模型。比如,你写的程序正确吗?单元测试通过了吗?是或否。但有些它们挣扎的事情,比如我认为它们很难理解我可能想法的细微差别,或者我的意图,以及何时该问澄清问题。或者任何感觉更软的东西就更差。所以你就像要么在轨道上,成为超级智能回路的一部分,要么不在轨道上,处于可验证领域之外,然后一切突然就变得漫无目的。
Yeah. I think like some of the bigger things is like maybe what's under underneath it if I could hypothesize is fundamentally these models are trained via reinforcement learning. So they're actually struggling with the exact same thing we just talked about which is the labs can improve the models in anything that is verifiable or that has rewards. So did you write the program correctly and does it you do you the unit tests check out? Yes or no. But some of the things where they're struggling is like for example, I think they have a tough time with like nuance of maybe what I what I had in mind or what I intended and when to ask clarifying questions. Or like what I Yeah, it's just um anything that feels softer is like worse. And so you're kind of like you're either on rails and you're part of the super intelligence circuits or you're not on rails and you're outside of the verifiable domains and suddenly everything kind of just like meanders.
也许另一种说法是,如果你今天去问最先进的模型 ChatGPT 讲个笑话,你知道你会得到什么笑话吗?就是那个笑话。那个笑话?我确实觉得,我无法告诉你它的标准形式,但我确实觉得 ChatGPT 只有三个笑话。
Like maybe another way to put it is if you go to today if you go to like state-of-the-art model, ChatGPT and you ask it tell me a joke, do you know what joke you're going to get? There's the joke. The joke? I do feel I I can't tell you like the standard form of it, but I do feel like ChatGPT has like three jokes.
是的,是的。所以显然所有 LLM 最喜欢的笑话是:为什么科学家不相信原子?好吧。因为它们编造一切。好吧。它们编造一切。所以这仍然出现?所以这是三四年前你会得到的笑话,今天你仍然得到这个笑话。好吧。
Yeah, yeah. So the joke that apparently all the LLMs like love the most is why do scientists not trust atoms? Okay. Because they make everything up. Okay. They make everything up. So this is still emerge? So this is the joke you would get like three or four years ago and this is the joke you still get today. Okay.
所以尽管模型已经大幅改进,如果你给它们一个智能体式任务,它们会连续工作几个小时,为你移山倒海。然后你让它们讲个笑话,它却讲了个愚蠢的笑话。那是五年前的烂笑话,因为它不在强化学习范围内。它不在被改进的范围内。这就是参差不齐的一部分——难道你不期望模型变得更好时,笑话也更好或更多样化吗?但它就是没有被优化,卡住了。
So even though the models have improved tremendously and if you give them an agentic task, they will just go for hours and move mountains for you. And then you ask for like a joke and it has a stupid joke. It's crappy joke from five years ago and it's because it's outside of the RL. It's outside of the reinforcement learning. It's outside of what's being improved. It's like and it's part of the jaggedness of like shouldn't you expect models as they get better to also have like better jokes or more diversity of them or it's just it's not being optimized and stuck.
你认为这是否意味着我们没有看到泛化,即更广泛的智能,比如笑话的聪明程度与代码的聪明程度相关联?
Do you think that that implies that we are not seeing like generalization in the sense of like broader intelligence of joke smartness being attached to code smartness?
是的,我认为存在某种解耦,有些事是可验证的,有些则不是,有些事被实验室任意优化,取决于输入了什么数据,有些则没有。
Yeah, I think there's some decoupling where some things are verifiable and some things are not and some things are optimized for arbitrarily by the labs depending on like what data went in and some things are not and um and
但我的意思是,有些研究团队的前提是,如果你在代码生成或这些可验证领域更聪明,你应该在所有事情上都更好。而笑话的情况表明这根本没有发生。好吧。
But I mean the premise there's a premise from some research groups that if you're smarter at code generation or in these verifiable fields, you should be better at everything. And like the joke situation suggests that that's not happening at all. Okay.
是的,我不认为那正在发生。我想我们可能看到了一点点,但还达不到令人满意的程度。
Yeah, I don't think that's happening. I think maybe we're seeing like a little bit of that, but not like a satisfying amount.
是的,这种参差不齐在人类中也存在。你可以非常擅长数学,但仍然讲很烂的笑话。
Yeah, that jaggedness exists in humans. You can be very very good at math and still tell really bad jokes.
是的,没错。
Yeah, that's true.
是的,但这仍然意味着,我们并没有得到这样的故事:随着模型越来越好,我们在社会的所有领域都免费获得了大量的智能和能力。这并非根本上的情况。存在盲点,有些东西没有被优化,而这一切都集中在这些神经网络的 opaque 模型中。所以你要么在它训练好的轨道上,一切如光速般进行,要么就不在。这就是 jaggedness。所以这就是为什么我认为,尽管进展的方向很明显,但你还不能完全放手,因为它还不能完全工作,或者是一个规模问题,我们只是还没弄清楚如何使用它。所以很难说。
Yeah, but it still means that we're not getting the story that we're getting a lot of intelligence and capabilities in all domains of society for free as we get better models. That's not fundamentally what's going on. There are blind spots and some things are not being optimized for, and this is all clustered up in these neural net opaque models. So you're either on rails of what it was trained for and everything is like you're going at speed of light, or you're not. So it's the jaggedness. So that's why I think even though the progression is obvious what should happen, you can't let it fully go there yet because it doesn't fully work, or it's a scale issue and we just haven't figured out how to use it. So it's hard to tell.
我能问一个有点亵渎的问题吗?如果这种 jaggedness 持续存在,并且都包裹在一个单一界面中,一个单一模型,这合理吗?还是应该将其拆分成可以针对不同智能领域进行优化和改进的东西?比如将模型拆分成不同领域的多个专家。更直接地说,而不是仅仅使用我们无法接触到的 MOE,因为从外部作为用户可能会感到困惑,比如为什么它在这方面这么好,但在其他方面却不行?
Can I ask a somewhat blasphemous question? If this jaggedness is persisting and it's all rolled up in a monolithic interface, a single model, does that make sense, or should it be unbundled into things that can be optimized and improved against different domains of intelligence? Like unbundling the models into multiple experts in different areas. More directly, instead of just MOE that we have no exposure to because that can be confusing as a user from the outside, which is like why is it so good at this but not at this other thing?
是的,我认为目前我的印象是,实验室正试图拥有一个单一文化的模型,该模型在所有不同领域都任意智能,他们只是把它塞进参数里。我确实认为我们应该期待智能中更多的物种分化。就像动物王国在大脑方面极其多样化,自然界有很多不同的生态位,有些动物有过度发达的视觉皮层或其他部分。我认为我们应该看到更多的物种分化,你不需要一个无所不知的神谕。你可以分化它,然后把它放在特定任务上,我们应该看到一些这样的例子,因为你应该能够拥有更小的模型,它们仍然具有认知核心,仍然有能力,但随后它们专门化,并在你真正关心的特定任务上在延迟或吞吐量方面变得更高效。比如,如果你是一个在 Lean 中工作的数学家,我看到有一些发布确实针对那个领域。所以可能会有一些这样的例子,拆分是有意义的。
Yeah, I think currently my impression is the labs are trying to have a single monoculture of a model that is arbitrarily intelligent in all these different domains and they just stuff it into the parameters. I do think we should expect more speciation in the intelligences. Like the animal kingdom is extremely diverse in the brains that exist and there's lots of different niches of nature, and some animals have overdeveloped visual cortex or other parts. I think we should be able to see more speciation, and you don't need an oracle that knows everything. You can speciate it and then put it on a specific task, and we should be seeing some of that because you should be able to have much smaller models that still have the cognitive core, they're still competent, but then they specialize and become more efficient in terms of latency or throughput on specific tasks you really care about. Like if you're a mathematician working in Lean, I saw there are a few releases that really target that as a domain. So there will probably be a few examples like that where the unbundling makes sense.
我有一个问题是,可用算力基础设施上的容量限制是否推动了更多这种情况,因为效率实际上更重要。如果你可以为你做的任何事情获得全部算力,即使是单个模型,但如果你实际上感到压力,无法为每个用例提供大规模模型,你认为这会导致任何物种分化吗?这个问题对你有意义吗?
One question I have is whether or not the capacity constraint on available compute infrastructure drives more of this because efficiency actually matters more. If you have access to full compute for anything you do, even one single model, but if you actually feel pressure where you can't serve a model of massive size for every use case, do you think that leads to any speciation? Does that question make sense to you?
这个问题有道理,我想我纠结的是,我认为我们还没有看到太多的物种分化。我们看到的是模型的单一文化。而且显然有压力要做一个好的代码模型,把它放回主模型,再合并。尽管模型已经面临压力。我想也许我觉得有很多非常短期的供应紧张,也许这会导致现在更多的物种分化。是的,我认为从根本上说,实验室在服务一个模型,他们并不真正知道最终用户会问什么。所以这可能是部分原因,因为他们必须多任务处理所有可能被问到的事情。但我认为,如果你来到一个企业,也许在某个你关心的特定问题上合作,那么你可能会在那里看到物种分化。或者会有一些非常高价值的应用,更加小众。但我认为现在他们基本上是在追求所有可用的东西。我不认为操纵大脑的科学已经完全发展起来。你所说的操纵是什么意思?比如在不损失能力的情况下进行微调。我们还没有这些原语来实际与智能体合作,除了上下文窗口。我们的上下文窗口基本上工作得很好,而且操作起来非常便宜。这就是我们获得一些定制化的方式。但我认为,如何更深入地调整模型,如何可能进行持续学习,或者如何在某个领域进行微调,如何在某个领域变得更好,或者如何实际触及权重而不仅仅是上下文窗口,这是一门正在发展的科学。触及权重比仅仅使用上下文窗口要棘手得多,因为你实际上从根本上改变了整个模型及其潜在的智能。所以也许物种分化的科学还没有完全发展起来。而且它还必须足够便宜,使得在这种背景下物种分化是值得的。
The question makes sense, and I guess what I'm struggling with is I don't think we've seen too much speciation just yet. We're seeing a monoculture of models. And there's clearly pressure to make a good code model, put it back in the main, merge again. Even though there already is pressure on the models. I guess perhaps I feel like there's a lot of very short-term supply crunch and maybe that causes more speciation now. Yeah, I think fundamentally the labs are serving a model and they don't really know what the end user is going to be asking about. So maybe that's part of it because they kind of have to multitask over all the possible things they could be asked. But I think if you're coming to a business and maybe partnering on some specific problems you care about, then maybe you would see that there. Or there would be some very high-value applications that are more niche. But I think right now they're kind of going after the totality of what's available. I don't think that the science of manipulating the brains is fully developed yet partly. What do you mean manipulating? So like fine-tuning without losing capabilities as an example. And we don't have these primitives for actually working with the intelligences in ways other than just context windows. Our context windows kind of just work and it's very cheap to manipulate etc. And this is how we're getting some of the customization. But I think it's a bit more of a developing science of how you more deeply adjust the models, how you have continual learning maybe, or how you fine-tune in a certain area, how you get better in a certain area, or how you actually touch the weights not just the context windows. And so it's a lot more tricky to touch the weights than just the context windows because you're actually fundamentally changing the full model and potentially its intelligence. So maybe it's just not a fully developed science of speciation. And it also has to be cheap enough for that speciation to be worthwhile in these given contexts.
我能问一个关于你描述的开放领域的自动研究扩展的问题吗?你说,好吧,我们有这个东西。我们需要更多的协作界面,基本上让人们为整体研究做出贡献。你能谈谈这个吗?
Can I ask a question about an extension to auto research that you described in terms of open ground? You say, okay, we have this thing. We need more collaboration surface around it essentially for people to contribute to research overall. Can you talk about that?
是的,所以我们谈到自动研究有一个单一的线程,即我在循环中尝试东西,但根本上并行化是其中有趣的部分。我想我尝试了一些想法,但我没有找到像那样简单的东西,我还没有找到让我非常满意的东西,但这是我在不处理我的 claw 时在做的副业。所以我认为一个问题是,如果你有一堆可用的并行化节点,那么很容易让多个自动研究人员通过一个公共系统或其他东西进行交流。我更感兴趣的是如何拥有一个互联网上不可信的工人池。例如,在自动研究中,你只是试图找到一段代码,将模型训练到非常低的验证损失。如果有人给你一个候选提交,很容易验证该提交是否正确和良好。
Yeah, so we talked about auto research has a single thread of I'm going to try stuff in a loop, but fundamentally the parallelization of this is the interesting component. And I guess I was trying to play around with a few ideas, but I don't have anything that clicks as simply as I don't have something I'm super happy with just yet, but it's something I'm working on the side when I'm not working on my claw. So I think one issue is if you have a bunch of nodes of parallelization available, then it's very easy to just have multiple auto researchers talking through a common system or something like that. What I was more interested in is how you can have an untrusted pool of workers out there on the internet. So for example in auto research, you're just trying to find the piece of code that trains a model to a very low validation loss. If anyone gives you a candidate commit, it's very easy to verify that that commit is correct and good.
比如有人可以从互联网上声称,这段代码会优化得更好,性能更佳。你可以直接检查。但检查可能投入大量工作。不过从根本上说,他们可能撒谎等等。所以你基本上是在处理类似的问题。这实际上有点像我的设计,其中包含一个不可信的工人池。它们看起来更像区块链,因为不是区块而是提交,这些提交可以相互构建,并包含改进代码时的更改。工作量证明基本上就是做大量实验来找到有效的提交。这很难,而奖励目前只是登上排行榜。没有任何金钱奖励。但我不想把类比推得太远。它从根本上存在这个问题:大量搜索投入其中,但验证一个候选方案是否确实优秀却非常廉价。有人必须尝试一万个想法,但你只需检查他们产出的东西是否有效,因为另外九千九百个都没用。所以长话短说,你必须设计一个系统,让不可信的工人池与负责验证的可信工人池协作。整个过程是异步的,并且能运作,从安全角度来看也是安全的,因为如果有人向你发送任意代码而你运行它,那会非常可疑和危险。但从根本上说,这应该是完全可行的。
Like someone could claim from the internet that this piece of code will optimize much better and give you much better performance. You could just check. But probably a lot of work goes into that checking. But fundamentally they could lie, etc. So you're basically dealing with a similar kind of issue. It actually looks a little bit like my designs that incorporate an untrusted pool of workers. They look a little bit more like a blockchain, because instead of blocks you have commits, and these commits can build on each other and contain changes to the code as you're improving it. The proof of work is basically doing tons of experimentation to find the commits that work. That's hard, and the reward is just being on the leaderboard right now. There's no monetary reward whatsoever. But I don't want to push the analogy too far. It fundamentally has this issue where a huge amount of search goes into it, but it's very cheap to verify that a candidate solution is indeed good. Someone had to try 10,000 ideas, but you just have to check that the thing they produced actually works, because 9,900 of them didn't work. So long story short, you have to come up with a system where an untrusted pool of workers can collaborate with a trusted pool of workers that do the verification. The whole thing is asynchronous and works, and it's safe from a security perspective because if anyone sends you arbitrary code and you run it, that is very sketchy and dodgy. But fundamentally it should be totally possible.
所以你熟悉像 SETI@home 和 Folding@home 这样的项目。所有这些问题都有类似的设置。在 Folding@home 中,你在折叠蛋白质,很难找到低能量的构象。但如果有人找到了他们认为低能量的构象,那就完美了。你可以直接使用它。你可以轻松验证。所以很多事情都有这个特性:提出方案非常昂贵,但验证非常廉价。在所有这些情况下,像 Folding@home、SETI@home 或家庭自动研究这样的项目都很适合。所以长话短说,互联网上的智能体群可以协作改进大语言模型,甚至可能超越前沿实验室。谁知道呢?也许这甚至可能。前沿实验室拥有大量可信算力,但地球更大,拥有大量不可信算力。但如果你建立系统来处理这个问题,那么也许外面的群体确实能想出更好的解决方案。人们把计算周期贡献给他们关心的事情。
So you're familiar with projects like SETI@home and Folding@home. All of these problems have a similar setup. In Folding@home, you're folding a protein and it's very hard to find a low-energy configuration. But if someone finds a configuration that they value as low energy, that's perfect. You can just use it. You can easily verify it. So a lot of things have this property: very expensive to come up with but very cheap to verify. In all those cases, things like Folding@home, SETI@home, or auto research at home will be good fits. So long story short, a swarm of agents on the internet could collaborate to improve LLMs and could potentially even run circles around frontier labs. Who knows? Maybe that's even possible. Frontier labs have a huge amount of trusted compute, but the earth is much bigger and has a huge amount of untrusted compute. But if you put systems in place to deal with this, then maybe it is possible that the swarm out there could come up with better solutions. People kind of contribute cycles to a thing they care about.
所以最后一个想法是:很多公司或其他机构可能有关心的事情,如果你有算力,你可以贡献给不同类型的自动研究项目。也许你关心癌症或某种特定类型。你不必只是向机构捐款。你可以实际购买算力,然后加入该项目的自动研究群体。如果一切都重新打包成自动研究者,那么算力就成了你贡献给池子的东西。
So the last thought is: lots of companies or whatnot could maybe have their own things they care about, and if you have compute capacity, you could contribute to different kinds of auto research tracks. Maybe you care about cancer or something of a certain type. You don't have to just donate money to an institution. You could actually purchase compute and then join the auto research swarm for that project. If everything is rebundled into auto researchers, then compute becomes the thing that you're contributing to the pool.
这非常鼓舞人心,也很有趣。我不知道这能走多远,但有趣的是,至少硅谷的一些人,或者在中国零售店排队的人,发现拥有个人算力再次变得有趣了。
That's very inspiring and it's also interesting. I don't know how far this goes, but it is interesting that at least some audience of people here in Silicon Valley or lining up at retail stores in China have discovered that having access to personal compute is interesting again.
是的。对吧?所以也许他们真的为此有动力,然后他们可以贡献给自动研究。
Yeah. Right? So maybe they're really motivated to do that for their claws and then they can contribute to auto research.
几乎就像美元是每个人都关心的事情,但未来浮点运算次数才是每个人真正关心的吗?会不会有一个翻转,你关心的事情变了?比如现在,即使你有钱也很难获得算力。所以实际上,在某种意义上,浮点运算次数似乎占主导地位。
Almost like dollars is the thing everyone cares about, but is flops the thing that actually everyone cares about in the future? Is there going to be a flipening of what's the thing that you care about? Right now, for example, it's really hard to get compute even if you have money. So actually it almost seems like flops is dominant in a certain sense.
是的,所以也许就是这样。你控制多少浮点运算次数,而不是你控制多少财富?我实际上不认为这是真的,但想想挺有趣的。
Yeah, so maybe that's kind of like that. How many flops do you control instead of what wealth you control? I don't actually think that's true, but it's kind of interesting to think about.
你最后发布的东西有点像就业数据分析。对吗?它可能触动了神经,尽管你只是可视化了一些公开数据。
The last thing you released was like a little bit of jobs data analysis. Is that right? It might have touched a nerve even though you're just visualizing some public data.
是的。我很好奇看看就业市场是什么样子,不同的角色在哪里,不同职业有多少人。我真的很想逐个案例看,并思考,随着这些人工智能及其可能的演变,这些会成为人们使用的工具吗?这些会成为取代这些职业的工具吗?当前的职业是什么,它们将如何变化?它们会大幅增长或调整,还是会出现新职业?所以这真的只是激发我自己对这个行业的思维链的一种方式。
Yeah. I was curious to look at the job market, what it looks like, where the different roles are, and how many people are in different professions. I was really interested to look through the individual cases and try to think about, with these AIs and how they're likely to evolve, are these going to be tools that people are using? Are these going to be displacing tools for these professions? What are the current professions and how are they going to change? Are they going to grow or adjust to a large extent, or what could be new professions? So it's really just a way to fuel my own chain of thought about the industry.
就业数据来自劳工统计局。他们实际上对每个职业在未来近十年内的预期增长百分比有展望。我认为是十年,但数据是 2024 年制作的。我们需要很多医疗工作者。所以他们已经做了这些预测,我不完全确定他们预测中使用的方法论。我感兴趣的是,根据人们认为当前主要发展的是那种更像数字人工智能,即像幽灵或精神实体,可以在数字世界中互动并操纵大量数字信息,而目前没有物理化身或存在。物理的东西可能会稍微慢一些,因为你在操纵原子。
The jobs data is from the Bureau of Labor Statistics. They actually have a percent outlook for each profession about how much it's expected to grow over the next almost a decade. I think it's a decade, but it was made in 2024. We need a lot of health care workers. So they've already made those projections, and I'm not sure 100% what the methodology was that they put into their projections. I was interested to color things by if people think that what's primarily being developed now is this kind of more digital AI that is like ghosts or spirit entities that can interact in the digital world and manipulate a lot of digital information, and they currently don't really have a physical embodiment or presence. The physical stuff is probably going to go slightly slower because you're manipulating atoms.
所以,翻转比特和复制粘贴数字信息的能力,使得一切比加速物质快一百万倍。从能量角度看,我认为我们将在数字空间看到大量的活动,大量的重写,大量的活动,像沸腾的汤。而且我认为,在数字空间中,我们将看到一些以光速发生的事情,相比之下,物理世界中发生的事情在某种程度上会慢一些,如果这是外推的话。所以我认为目前存在一种积压,可能有很多数字信息处理曾经由计算机和人完成,现在有了 AI,出现了第三种数字信息操纵者。这些领域将会有大量的重构。但物理世界实际上会落后一段时间。所以对我来说真正迷人的是,这就是为什么我强调了那些从根本上操纵数字信息的职业。这是你可以在家完成的工作等等。因为我觉得事情会改变。这并不意味着这些工作会变少或变多,因为这涉及到需求弹性和其他许多因素。但由于这些新工具,以及人类超级有机体神经系统的升级(如果你愿意这样想的话),这些职业将会发生变化。
So flipping bits and the ability to copy-paste digital information makes everything a million times faster than accelerating matter. So energetically, I think we're going to see a huge amount of activity in the digital space, huge amount of rewriting, huge amount of activity, boiling soup. And I think we're going to see something in the digital space that goes at the speed of light compared to what's going to happen in the physical world to some extent, if that would be the extrapolation. So I think there's currently an overhang where there can be a lot of unhubbling almost potentially of a lot of digital information processing that used to be done by computers and people. And now with AIs, there's a third kind of manipulator of digital information. There's going to be a lot of refactoring in those disciplines. But the physical world is actually going to be behind that by some amount of time. So what's really fascinating to me is that's why I was highlighting the professions that fundamentally manipulate digital information. This is work you could do from your home, etc. Because I feel like things will change. And it doesn't mean that there's going to be less of those jobs or more of those jobs because it has to do with demand elasticity and many other factors. But things will change in these professions because of these new tools and because of this upgrade to the nervous system of the human superorganism, if you want to think about it that way.
鉴于你对数据的观察,你对面临就业市场的人,或者正在思考现在该学什么、该发展什么技能的人,有什么观察或建议吗?我的意思是,我们都可以去……我很庆幸我现在的工作需要与人见面。
Given the look you had at the data, do you have either any observations or guidance for people facing the job market or thinking about what to study now or what skills to develop? I mean, we can all go get... I'm very thankful that I have to meet people for my job right now.
是的。
Yeah.
是的,更偏物理。不过你能在家完成工作吗?
Yeah, more physical. Could you do your work from home though?
我可以。我认为其中涉及人际关系的部分很难,但大部分我都能做。是的。我认为这真的很难说,因为就业市场非常多样化。答案可能会有所不同,但在很大程度上,这些工具非常新,非常强大。所以首先要做的就是跟上它。因为我认为很多人要么忽视它,要么害怕它,这当然完全可以理解。是的,我认为它目前从根本上说是一个赋能的工具。而这些工作是任务的集合。其中一些任务可以快得多。所以人们应该把它主要看作一个工具。我认为它的长期未来是不确定的。是的,说实话,这真的很难预测。我并不是专业做这个的。我认为这应该是经济学家的工作。
I could. I think there are relationship parts of it that are hard, but most of it I could. Yeah. I think it's really hard to tell because again, the job market is extremely diverse. I think the answers will probably vary, but to a large extent, these tools are extremely new, extremely powerful. So just trying to keep up with it is the first thing. Because I think a lot of people dismiss it or they're afraid of it, which is totally understandable, of course. Yeah, I think it's fundamentally an empowering tool at the moment. And these jobs are bundles of tasks. And some of these tasks can go a lot faster. So people should think of it as primarily a tool that it is right now. And I think the long-term future of that is uncertain. Yeah, it's kind of really hard to forecast, to be honest. And I'm not professionally doing that really. And I think this is a job for economists to do properly.
但你是个工程师。我觉得有趣的一点是,对工程工作的需求在持续增长。
You are an engineer though. And one thing I thought was interesting is that the demand for engineering jobs is continuing to increase.
是的。我无法判断这是否是暂时现象。我不确定我对此的感受。
Yeah. I can't tell if that's a temporary phenomenon. I'm not sure how I feel about it.
是的,这就像需求弹性。几乎就像软件是稀缺的,对吧?我们对软件没有更多需求的原因就是它的稀缺性和太贵。所以如果门槛降低,实际上就会出现杰文斯悖论,即对软件的需求实际上会上升。它更便宜、更强大。
Yeah, that's like the demand elasticity. Almost like software was scarce, right? And the reason we don't have more demand for software is just its scarcity and it's too expensive. So if the barrier comes down, then actually you have the Jevons paradox, which is you know, you actually the demand for software actually goes up. It's cheaper and more powerful.
经典的例子总是自动取款机和银行出纳员,因为很多人担心自动取款机和计算机会取代出纳员。但实际发生的是,它们使银行分行的运营成本大大降低。所以银行分行更多了,出纳员也更多了。这是人们引用的典型例子。但基本上就是杰文斯悖论。某样东西变得更便宜,所以对它有很多被释放的需求。所以我确实认为……我对软件工程持谨慎乐观的态度,我认为对软件的需求将非常巨大。而且它变得便宜了很多。所以我认为在相当长一段时间内,虽然很难预测,但至少目前看来,局部上对软件的需求会增加。因为软件太棒了。它是数字信息处理。你不必被迫使用给你的任意工具。它们在很多方面都不完美。你不必被迫订阅现有的东西。代码现在是短暂的,可以改变和修改。所以我认为数字空间会有大量活动,在某种意义上重新连接一切。我认为这会创造大量对这种东西的需求。我认为长期来看,显然即使是像 OpenAI 或 Anthropic 或其他实验室这样的自动研究机构,他们也雇佣了大约一千多名研究人员,对吧?
The classical example of this is always the ATMs and the bank tellers, because there was a lot of fear that ATMs and computers would displace tellers. But what happened is they made the cost of operation of a bank branch much cheaper. And so there are more bank branches, so there are more tellers. It's the canonical example people cite. But basically it's just Jevons paradox. Something becomes cheaper, so there's a lot of unlocked demand for it. So I do think that's probably... I do have a cautiously optimistic view of this in software engineering, where I do think it seems to me like the demand for software will be extremely large. And it's just become a lot cheaper. So I do think that for quite some time, it's very hard to forecast, but it does seem to me like right now at least locally, there's going to be more demand for software. Because software is amazing. It's digital information processing. You're not forced to use arbitrary tools that were given to you. They're imperfect in various ways. You're not forced to subscribe to what exists. Code is now ephemeral and it can change and it can be modified. So I think there's going to be a lot of activity in the digital space to rewire everything in a certain sense. And I think it's going to create a lot of demand for this kind of stuff. I think long-term, obviously even with auto research like OpenAI or Anthropic or these other labs, they're employing what, like a thousand something researchers, right?
嗯。
Mhm.
这些研究人员基本上就像被美化的自动……他们正在积极地自动化自己,这就是他们都在努力做的事情。我接触过一些研究人员,他们也害怕,感到那种精神错乱,对吧?因为它在起作用,对吧?所以他们觉得,对我来说也完了。我花了很多时间在 OpenAI 转悠,我说,你们意识到如果我们成功了,我们都会失业吗?这只是在为 Sam 或类似的人构建自动化。或者为董事会,我不确定,但他们只是在为董事会或 CEO 构建所有这些自动化。而我们都会失业,也许只能做点副业。所以是的,从这个角度看,有点令人不安。
These researchers are basically like glorified auto... They're like automating themselves away actively, and this is the thing they're all trying to do. I went around some of those researchers also fear that, feel the psychosis, right? Because it's working, right? And so they're like, it's over for me, too. I did spend a bunch of time going around OpenAI and I was like, you guys realize if we're successful, we're all out of a job. This is just going to be building automation for Sam or something like that. Or the board, I'm not sure, but they're just building all this automation for the board or the CEO or something like that. And we're all out of our job and maybe contributing on the side. So yeah, it's kind of unnerving from that perspective.
我能问你诺姆的问题吗?你知道,你也可以那样做,对吧?在某个前沿实验室,用大量的算力规模和一群同事做自动研究。为什么不呢?
Is it okay if I ask you Noam's question? You know, you could be doing that, right? Auto researching with a lot of compute scale and a bunch of colleagues at one of the frontier labs. Like why not?
嗯,我在那里待过一段时间,对吧?而且我确实重新进入了。所以在某种程度上我同意,我认为这个问题有很多种解读方式。这个问题有点沉重。我会说,我对人们在前沿实验室之外能做出的贡献和影响感到非常满意,显然。不是在行业内,而是在更生态系统层面的角色。比如你的角色就更偏生态系统层面。我目前的角色也更偏生态系统层面。
Well, I was there for a while, right? And I did reenter. So to some extent I agree, and I think that there are many ways to slice this question. It's a very loaded question a little bit. I will say that I feel very good about what people can contribute and their impact outside of the frontier labs, obviously. Not in the industry, but also in more ecosystem level roles. So your role for example is more ecosystem level. My role currently is also kind of more on ecosystem level.
而且我对人们在这些角色中能产生的影响感到非常满意。但反过来,我认为在某种程度上,与前沿实验室过度结盟也存在明确的问题。从根本上说,你从这些前沿实验室获得了巨大的经济激励。而且你自己也承认,人工智能将以非常戏剧性的方式改变人类和社会。而你却在构建这项技术并从中受益,通过经济手段与它紧密结盟。这正是 OpenAI 最初创立时的核心难题。这是我们当时试图解决的问题。而现在它仍然没有完全解决。所以这是第一点。你并不是一个完全自由的个体,无法以完全自主、自由的方式参与那些对话。如果你身处某个前沿实验室,有些话你不能说。反过来,有些话组织希望你来说。他们不会强迫你,但你会感受到应该说什么的压力,因为否则就会陷入非常尴尬的对话、奇怪的侧目,比如「你在干什么?」所以你无法真正成为一个独立的个体。在某种程度上,我觉得在前沿实验室之外,我更加与人类保持一致,因为我几乎不受那些压力的影响,对吧?而且我可以想说什么就说什么。
And I feel very good about the impact that people can have in those kinds of roles. I think conversely, there are definite problems in my mind for basically aligning yourself way too much with the frontier labs, too. So fundamentally, you have a huge amount of financial incentive with these frontier labs. And by your own admission, the AIs are going to really change humanity and society in very dramatic ways. And here you are basically building the technology and benefiting from it, and being very allied to it through financial means. This was the conundrum that was at the heart of how OpenAI was started in the beginning. This was the conundrum that we were trying to solve. And so it's still not fully resolved. So that's number one. You're not a completely free agent and you can't actually be part of that conversation in a fully autonomous, free way. If you're inside one of the frontier labs, there are some things that you can't say. And conversely, there are some things that the organization wants you to say. And they're not going to twist your arm, but you feel the pressure of what you should be saying, because otherwise it's really awkward conversations, strange side eyes, like what are you doing? So you can't really be an independent agent. And I feel a lot more aligned with humanity in a certain sense outside of the frontier lab because I'm not subject to those pressures, right? And I can say whatever I want.
是的,我想说在前沿实验室里,你当然也能产生影响。那里有很多研究人员,也许你就是其中之一,也许你的想法非常好等等。也许有很多决策要做,你想在那些对话发生时身处其中。我确实认为目前的风险总体较低,所以一切都还不错。但归根结底,当风险真正很高时,如果你是一个组织的员工,我真的不知道你对组织将要做什么有多大影响力。从根本上说,你并不是真正负责的人。你在房间里,贡献想法,但你并不是真正掌控你所参与的那个实体。所以我认为在某种程度上,这些是错位的一些来源。我要说的是,在某种程度上我非常同意那种观点,我确实觉得在实验室里,无论好坏,它们是不透明的,很多工作都在那里。它们处于能力和可能性的边缘。它们正在研究即将到来的东西。我认为如果你在那个前沿实验室之外,你的判断力从根本上会开始漂移,因为你没有参与即将到来的东西。所以我觉得我的判断力也必然会开始漂移。我实际上不会理解这些系统在底层是如何工作的。那是一个不透明的系统。我不会很好地理解它将如何发展等等。所以在这个意义上我同意,并且这是我担心的事情。我认为基本上值得与实际情况保持联系,并且实际待在一个前沿实验室里。如果某些前沿实验室能让我去一段时间,为他们做真正好的工作,然后也许再出来……找工作。这非常令人兴奋。那么我认为这可能是一个好的安排,因为我感觉这也许是既能与实际情况保持联系,又不会觉得自己被那些实体完全控制的一种方式。所以老实说,在我看来,Noam 可能在 OAI 做得非常好,但我也认为他最有影响力的工作很可能在 OpenAI 之外。Noam,这是呼吁你成为一名独立研究者,进行自动研究。是的,外面有很多事情可以做,而且……我认为最终理想的解决方案也许是来回切换。而且我认为从根本上说,你在两个地方都能产生非常惊人的影响。所以非常复杂,我不知道。这是一个有点沉重的问题,但我的意思是,我加入过前沿实验室,现在在外面。也许将来我想再次加入。我想这就是我看待它的方式。
Yeah, I would say in the frontier labs, you can have impact there of course as well. So there are many researchers, and maybe you're one of them, maybe your ideas are really good, etc. Maybe there's a lot of decision-making to do and you want to be in a position where you are in the room with those conversations when they come up. I do think that currently the stakes are overall fairly low and so everything is kind of nice. But ultimately, at the end of the day, when the stakes are really high, if you're an employee at an organization, I don't actually know how much sway you're going to have on what your organization is going to do. Fundamentally, you're not really in charge. You're in the room and you're contributing ideas, but you're not really in charge of that entity that you're part of. So those are some sources of misalignment, I think to some extent. I will say that in one way I do agree a lot with that sentiment that I do feel like in the labs, for better or worse, they're opaque and a lot of work is there. And they're kind of at the edge of capability and what's possible. And they're working on what's coming down the line. And I think if you're outside of that frontier lab, your judgment fundamentally will start to drift because you're not part of what's coming down the line. And so I feel like my judgment will inevitably start to drift as well. And I won't actually have an understanding of how these systems actually work under the hood. That's an opaque system. I won't have a good understanding of how it's going to develop, etc. And so I do think that in that sense I agree and something I'm nervous about. I think it's worth basically being in touch with what's actually happening and actually being in a frontier lab. And if some of the frontier labs would have me come for some amount of time and do really good work for them and then maybe come and hang out... looking for a job. This is super exciting. Then I think that's maybe a good setup because I kind of feel like it's maybe one way to actually be connected to what's actually happening, but also not feel like you're necessarily fully controlled by those entities. So I think honestly in my mind, Noam can probably do extremely good work at OAI, but also I think his most impactful work could very well be outside of OpenAI. Noam, that's a call to be an independent researcher with auto research. Yeah, there are many things to do on the outside and it's a... and I think ultimately the ideal solution maybe is like going back and forth. And I think fundamentally you can have a really amazing impact in both places. So very complicated, I don't know. It's a very loaded question a little bit, but I mean I joined the frontier lab and I'm outside. And then maybe in the future I'll want to join again. And I think that's kind of how I look at it.
一个与世界或人工智能生态系统对前沿的可见性相关的问题是,开源离前沿有多近,以及这种状况的可持续性如何。我认为这相当令人惊讶。整个事件序列实际上从拥有少数中国模型和全球模型开始,我认为在短期内人们将继续发布在能力上比行业预期的更接近前沿的模型。我不知道你是否对此感到惊讶,但你是开源社区的长期贡献者。你对此有什么预测?
One question related to what visibility does the world or the AI ecosystem have into the frontier is how close open source is to the frontier, and how sustainable that is. I think it is quite surprising. The entire sequence of events actually from having a handful of Chinese models and global models, and I think people are going to continue releasing here in the near term that are closer than much of the industry anticipated from a capability perspective. I don't know if you're surprised by that, but you're a long-term contributor to open source. What's your prediction here?
是的,粗略来说,基本上闭源模型领先,但人们正在监控开源模型落后多少个月。一开始是零,然后变成了 18 个月。现在……是的,但随后收敛了,对吧?所以现在它们可能落后了,比如最新的是多少?大概 8 个月、6 个月、8 个月左右。是的,我显然是开源的忠实粉丝。例如,在操作系统中,有闭源的 Windows 和 Mac OS,这些是大型软件项目,就像 LLM 将要成为的那样,还有 Linux。但 Linux 非常容易。实际上,Linux 是一个非常成功的项目。它运行在绝大多数计算机上。我上次查的时候,Linux 大概占了 60%左右?这是因为行业需要一个通用的开放平台,让每个人都觉得使用起来安全。我想说行业一直对这类项目的存在有需求。我认为现在也是如此。这就是为什么企业实际上希望有这种东西存在。最大的区别在于,一切都资本密集,需要大量的资本支出。所以我认为这就是事情有点崩溃的地方,使得在某些方面更难竞争。
Yeah, so roughly speaking, basically the closed models are ahead, but people are monitoring the number of months that open-source models are behind. And it started with nothing and then it went to 18 months. Now it's... Yeah, but then convergence, right? So then maybe they're behind by like, what is the latest? Maybe like 8 months, 6 months, 8 months kind of thing right now. Yeah, I'm a huge fan of open-source, obviously. So for example, in operating systems, you have closed source like Windows and Mac OS, these are large software projects, kind of like what LLMs are going to become, and there's Linux. But Linux is very easy. Actually, Linux is an extremely successful project. It runs on the vast majority of computers. Last time I checked, was it like 60% or something from Linux? And that's because there is a need in the industry to have a common open platform that everyone feels safe using. I would say the industry has always felt a demand for that kind of a project to exist. And I think the same is true now. And that's why businesses actually want there is demand for this kind of a thing to exist. The big difference is that everything is capital, there's a lot of capex that goes into this. So I think that's where things fall apart a little bit, make it a bit harder to compete in certain senses.
我确实认为当前的模型非常好。另一件我觉得很有意思的事情是,对于绝大多数消费者用例来说,即使是开源模型也相当不错。而且我认为如果再往前几年,大量简单用例都会被很好地覆盖,甚至可以在本地运行。但总会有对前沿智能的需求,而且这个需求可能占很大一部分。不过,对前沿智能的需求可能像是诺贝尔奖级别的工作。比如把 Linux 从 C 语言迁移到 Rust。那会是更大的项目,范围类似那样,也许很多前沿的封闭智能会与之交互。而开源则会吃掉大量更基础的用例。在某个时候,今天的前沿模型,很可能今年晚些时候,我现在从封闭实验室使用的那些前沿模型,可能就会变成开源,并承担大量工作。所以我预计这种动态会持续下去。我们会有前沿实验室拥有封闭的 AI,它们有点像神谕,然后开源会落后几个月。我实际上认为这是一个相当不错的整体安排,因为我对只有封闭智能有点犹豫。我认为集中化在过去有着非常糟糕的记录。
I do think that the current models are very good. The other thing that I think is really interesting is that for the vast majority of consumer use cases, even open-source models are actually quite good. And I think if you go forward more years, it seems to me like a huge amount of simple use cases are going to be well covered and actually even run locally. But there's going to be always some demand for frontier intelligence, and that can actually be an extremely large piece of the pie. But it could be that the need for frontier intelligence is going to be like Nobel Prize kind of work. Let's move Linux from C to Rust. It's going to be bigger projects, scoped in that kind of a way, and maybe that's where a lot of the frontier closed intelligence is going to be interacting with. And open-source is going to eat through a lot of the more basic use cases. At some point, what is frontier today is going to be, probably later this year, what's frontier today in terms of what I'm using right now from the closed labs might be open-source and that's going to be doing a lot of work. So I kind of expect that this dynamic will continue. We'll have frontier labs that have closed AIs that are kind of like these oracles, and then we'll have open-source kind of behind by some amount of months. And I actually think that's a pretty good setup overall, because I'm a little bit hesitant of having intelligence that are closed and that's it. I think centralization has a very poor track record in the past.
你是指像政治或经济体系那样的一般情况。
You mean like in political or economic systems in general.
正是。我认为有很多非常糟糕的先例,所以我希望有一个东西,可能不在能力前沿,因为那是新的和未探索的,但我希望有一个落后的东西,它是整个行业都能访问的智能公共工作空间。这对我来说似乎是行业相当不错的权力平衡。我还认为有很多问题需要解决。如果你不断从前沿推进智能,我们可以做新的事情,而且人类面临很多非常大的问题。所以这似乎会继续是一个非常昂贵的游戏。因此我想支持那些这样做的实验室,因为有些问题如果不继续以非常昂贵的方式推进模型就无法解决。然而,正如你指出的,如果我们今天拥有的前沿是开放的,那就有很多能力。这种力量或它的民主化似乎非常有用且健康。
Exactly. I think there are a lot of pretty bad precedents, so I want there to be a thing that is maybe not at the edge of capability because it's new and unexplored, but I want there to be a thing that's behind and that is kind of a common working space for intelligences that the entire industry has access to. That seems to me like a pretty decent power balance for the industry. I also think there are many problems to solve. If you keep advancing intelligence from the frontier, we can do new things and there are a lot of very big problems for humanity. So it seems that will continue to be a very expensive game. And so I want to root for labs that are doing that because there are problems we cannot solve without continuing to advance the models in a very expensive way. And yet, as you point out, if what we have today as frontier is open, that's a lot of capability. The power of that or the democratization of that seems very useful and also healthy.
是的。
Yeah.
我认为基本上是无意中,我们实际上处于一个还不错的位置。
I think basically by accident we're actually in an okay spot.
最优的。是的。就像无意中,我们在某种意义上碰巧处于一个好位置。
An optimal. Yeah. Like by accident we happened to be in a good spot in a certain sense.
嗯,在某种程度上,这种动态持续得越久,生态系统可能就越健康,因为曲线下的面积越来越大。
Well, and to some degree the longer this dynamic endures, the healthier the ecosystem might be, because you have more and more area under the curve.
而且我要说,即使在封闭方面,我最近也感觉它更加集中化了,因为我认为很多领跑者并不一定是顶级。所以从这个意义上说,我认为这不是很理想。我希望有更多的前沿实验室,因为我默认对集中化非常怀疑。我希望房间里有更多的人。我认为在机器学习中,集成总是胜过任何单个模型。所以我希望有集成的人思考所有最难的问题,我希望在做出那些决定时,房间里有集成的人,他们都充分知情。我不希望它变成两三个人关起门来。我觉得那不是好的未来。我几乎希望有更多的实验室,只要它们规模小,而且我确实认为开源有发挥的空间。我希望它持续存在,目前它稍微落后,这实际上是件好事。
And I will say that even on the closed side, I almost feel like it's been even further centralizing recently because I think a lot of the frontrunners are not necessarily the top tier. So in that sense I think it's not super ideal. I would love there to be more frontier labs because I'm by default very suspicious of centralization. I want there to be more people in the room. I think in machine learning, ensembles always outperform any individual model. So I want there to be ensembles of people thinking about all the hardest problems and I want there to be ensembles of people in the room when they make those decisions, all well informed. I don't want it to be a closed doors with two or three people. I feel like that's not a good future. I almost wish there were more labs as long as they're short, and I do think that open-source has a place to play. I hope it sticks around and it's currently slightly behind and that's actually a good thing.
好的,你曾从事汽车中通用机器人自主性的前身工作,对吧?过去几个月,机器人公司也发生了很多事情,比如环境、任务的泛化加速,令人印象深刻,任务时间跨度越来越长,大量资金涌入这个领域。这会发生吗?在你看来,最近有什么变化吗?
Okay, you worked on the precursor to generalized robotics autonomy in cars, right? A lot has happened in the last couple months with robotics companies as well, like acceleration of really impressive generalization of environment, of tasks, increasingly long horizon tasks, lots of money going into the space. Is it going to happen? Has anything in your view changed recently?
我的观点受到我在自动驾驶中看到的情况的影响,我确实觉得自动驾驶是第一个机器人应用。所以可能我看到的是,10 年前有大量初创公司,其中大多数长期来看都没有成功。我看到的是,需要投入大量资本支出和时间。所以我认为机器人技术,因为它非常困难、非常混乱,需要巨额资本投资和很大的信念。这是一个大问题,我认为原子是非常困难的。所以我感觉它们会落后于数字空间将要发生的事情。而在数字空间,将会有大量的「去束缚」,基本上那些不是非常高效的东西会变得高效一百倍。因为比特容易得多。所以我认为目前就什么会改变以及活动在哪里而言,我感觉数字空间会发生巨大变化。然后物理空间会落后。我还觉得它们之间的接口非常有趣。因为如果我们有更多的智能体代表人类行动,更多的智能体相互交谈、执行任务并参与某种智能体经济,你最终会耗尽纯粹在数字空间能做的事情。在某个时候,你必须进入宇宙并向它提问。
My view is kind of informed by what I saw in self-driving, and I do feel like self-driving is the first robotics application. So probably what I saw is that 10 years ago, there were a large number of startups, and most of them basically didn't long-term make it. And what I saw is that a lot of capital expenditure had to go in and a lot of time. So I think robotics, because it's so difficult, is so messy, and requires a huge amount of capital investment, and a lot of conviction. It's a big problem and I think atoms are really hard. So I kind of feel like they will lag behind what's going to happen in digital space. And in digital space there's going to be a huge amount of unhobbling, basically things that weren't super efficient becoming a lot more efficient by a factor of a hundred. Because bits are so much easier. So I think currently in terms of what's going to change and where the activity is, I kind of feel like digital space is going to change a huge amount. And then the physical space will lag behind. And what I find very interesting is this interface in between them as well. Because if we do have more agents acting on behalf of humans and more agents talking to each other and doing tasks and participating in a kind of economy of agents, you're going to run out of things that you're going to do purely in the digital space. At some point you have to go to the universe and you have to ask it questions.
你必须运行一个实验,看看宇宙告诉你什么,才能学到东西。所以我们目前有大量的数字工作,因为我们在集体思考已经数字化的东西方面存在积压。我们只是没有足够的人类思考周期来处理已经数字化和上传的所有信息。所以我们很快就会用完实际上已经上传的内容。你最终会阅读所有论文并处理它们,产生一些尝试的想法,但我真的不知道你能从完全封闭的信息中获得多少智能。所以我认为首先会发生大量的解绑工作,那里有大量的工作要做。然后它会转向物理和数字之间的接口。那就是感知世界的传感器和作用于世界的执行器。
You have to run an experiment and see what the universe tells you to get back to learn something. And so we currently have a huge amount of digital work because there's an overhang in how much we collectively thought about what already is digital. So we just didn't have enough thinking cycles among the humans to think about all the information that is already digital and already uploaded. And so we're going to start running out of stuff that is actually already uploaded. So you're going to at some point read all the papers and process them and have some ideas about what to try, but I don't actually know how much you can get intelligence that's fully closed off and was just information that's available in the you know. And so I think what's going to happen is first there's going to be a huge amount of unhobbling and I think there's a huge amount of work there. Then actually it's going to move to the interfaces between physical and digital. So that's sensors of seeing the world and actuators of doing something to the world.
嗯。
Mhm.
所以我认为很多有趣的公司实际上会来自那个接口:我们能否在某种意义上为超级智能提供数据,并且我们能否实际提取数据并按照它的指令操纵物理世界,如果你想把整个事情拟人化的话。然后物理世界,我几乎觉得总可寻址市场在工作量等方面是巨大的,可能甚至比数字空间能发生的要大得多。所以我认为这也是一个更大的机会。但我确实觉得这是大量的工作,在我看来,原子要难一百万倍。所以它会滞后,但它也是一个稍大的市场。所以我认为机会是沿着那条轨迹发展的。所以现在数字是我的主要兴趣。然后接口会是之后的事,然后也许一些物理事物,它们的时代会到来,而且当它们到来时会非常巨大。
So I think a lot of interesting companies will actually come from that interface of can we feed the superintelligence in a certain sense data and can we actually take data out and manipulate the physical world per its bidding, if you want to anthropomorphize the whole thing, right? And then the physical world actually I almost feel like the total addressable market in terms of the amount of work and so on is massive, possibly even much larger than what can happen in digital space. So I think it's a much bigger opportunity as well. But I do feel like it's a huge amount of work and in my mind the atoms are just a million times harder. So it will lag behind, but it's also a little bit of a bigger market. So I think the opportunity is kind of follow that trajectory. So right now digital is my main interest. Then interfaces will be after that and then maybe some of the physical things their time will come and they'll be huge when they do come.
嗯,这也是一个有趣的框架,因为某些事情,不是我现在正在做的事情,但某些事情即使在原子世界里也容易得多。
Well, it's an interesting framework for it, too, because certain things, not the things I'm working on right now, but certain things are much easier even in the world of atoms.
嗯。
Mhm.
对吧?就像如果你只考虑对物理世界的读写,比如读取、传感器、摄像头,有很多现有的硬件,你可以想象丰富智能体能力或捕获大量新数据,如果你聪明地处理,不一定需要大量投资就能获得有价值的东西。
Right? Like if you just think about read and write to the physical world, like read, sensors, cameras, there's a lot of existing hardware and you can imagine enriching agent capabilities or capturing a lot of new data if you're clever about it and you don't necessarily have to invest a lot to get something valuable.
是的。对。所以我看到的例子,比如,我的一个朋友 Liam,是 Periodic 的 CEO。我上周拜访了他们。所以这正好在脑海中。他们正在尝试做材料科学的自动研究。在这种情况下,智能的传感器实际上是相当昂贵的实验室设备。生物学也是如此。我认为很多人对工程生物学非常感兴趣,而且传感器将不仅仅是摄像机。这说得通吗?然后我看到的另一件事是,有些公司试图让你基本上为训练数据付费给人们。
Yeah. Right. So examples of this that I saw for example are, you know, a friend of mine, Liam, is the CEO of Periodic. I visited them last week. So it was just on top of mind. They're trying to do auto research for materials science. And so in that case the sensors to the intelligence are actually pretty expensive lab equipment. And the same is true in biology. I think a lot of people are very interested in engineering biology and, you know, the sensors will be more than just video cameras. Does that make sense? And then the other thing I saw for example is companies that are trying to have you basically pay people for training data.
为了喂养……
To feed the...
以编程方式。
Programmatically.
是的。为了喂养博格。
Yeah. To feed the Borg.
所以这些在某种意义上都是传感器的例子。所以它们有多种多样的形状和形式,如果这说得通的话。是的,所以我期待有一天我可以在物理世界中要求一个任务,给它定价,然后告诉智能体,你知道,你弄清楚怎么做。去获取数据。
And so these are all examples of sensors in a certain sense. So they take many diverse shapes and forms if that makes sense. Yeah, so I'm looking forward to the point where I can ask for a task in the physical world and I can put a price on it and just tell the agent, you know, you figure out how to do it. Go get the data.
我实际上有点惊讶我们没有足够的信息市场。比如,如果 Polymarket 或其他博彩市场甚至股票等,如果它们有如此多的自主活动和不断增长的活动量,比如为什么如果伊朗刚刚发生事情,怎么没有一个过程让从德黑兰某个地方拍照或视频花费 10 美元?应该有人能够为此付费,你知道,这就是喂养智能的一个例子。不会有真人看它,而是试图猜测博彩游戏和股票市场等的智能体。所以我有点觉得智能体网络仍然相当新,但没有这样的机制,但这是我认为可能发生的事情的一个例子。
I'm actually kind of surprised we don't have enough information markets. Like if for example Polymarket or other betting markets or even stocks, etc. If they have so much autonomous activity and rising amount of activity, like why should for example if Iran was just happening now, how come there isn't a process where taking a photo or video from somewhere in Tehran should cost like 10 bucks? Like someone should be able to pay for that, you know, and that's an example of feeding the intelligence. There's not going to be a human looking at it, it's going to be agents who are trying to guess the betting games and stock markets and so on. So I kind of feel like the agentic web is still fairly new, but there's no mechanisms for this, but this is an example of what I think might happen.
有一本可能很有启发性的好书叫《Daemon》。你可能读过。在《Daemon》中,智能最终在某种意义上有点像操纵人类,你知道吗?所以,人类有点像它的执行器,但人类也是它的传感器。所以,我认为整个社会会以某种方式重塑,以服务于那种……这最终会在整个行业集体发生。是的,那里有更多的自动化,它有某些需求,人类将服务于那台机器的需求,而不一定是彼此。
There's a good book that maybe is inspiring called Daemon. You potentially read it. In Daemon, the intelligence ends up puppeteering almost a little bit like humanity in a certain sense, you know? And so, humans are kind of its actuators, but humans are also its sensors. And so, I think collectively society will kind of reshape in a certain way to serve that kind of a... that will kind of end up happening collectively across the industry. Where yeah, there's just a lot more automation and it has certain needs and kind of humans will be serving those needs of that machine, not necessarily to each other.
嗯,我们刚才谈到训练数据缺失这个非常具体的点。我们需要类似自动研究的东西,对吧?我们需要训练周期或 SFTP 部分更加机械化。
Well, we were on this very specific point of missing pieces of training data. We needed something like auto research, right? We need the training cycle or the SFTP piece to be far more mechanized.
为了哪个部分?
For which part?
为了使收集……为了将人类从循环中移除,要求一个任务就像用新数据改进我的模型质量,对吧?这对你来说有意义吗?就像如果你不能让模型自己进行训练运行,那么你将其作为闭环任务通过定价数据来执行的能力就更具挑战性。
In order to make the collection like... in order to take the human out of the loop to ask for a task that is just like improve my model quality with new data, right? Does that make sense to you? Like if you can't have the model do the training runs by itself, then your ability to do this as a closed loop task with pricing data is more challenged.
是的,是的,100%。是的。但现在你可以了。
Yes, yes, 100%. Yeah. But now you do.
问题是对于 LLM 训练,它实际上非常适合这个范式。
The thing is for LLM training, it actually fits the paradigm really well.
嗯。
Mhm.
所以你会期望……
So you'd expect...
指标。是的,就像 LLM 训练实际上非常适合这个范式,非常容易。就像所有代码的优化等等,它运行得更快。然后你还有可以优化的指标。我确实认为如果你对这些指标有一个自主循环,会有很多好的从众行为,系统会过度拟合这些指标。但然后你可以用系统来设计更多的指标,你就有很好的覆盖。所以很难说,但在某种意义上它非常合适。
Metric. Yeah, like LLM training actually fits the paradigm really well, really easily. Like all the optimization of all the code and so, it runs faster. And then you also have metrics that you can optimize against. I do think that if you had an autonomous loop over those metrics, there's going to be a lot of good herding going on where the system will overfit to those metrics. But then you can use the system to devise more metrics and you just have a really good coverage. So it's kind of hard to tell, but in a certain sense it's a pretty good fit.
在结束之前,我想聊聊你的一个小项目。给我讲讲 MicroGPT 艺术吧。
I want to talk about a little tiny side project you have before we end. Tell me about the MicroGPT arts.
哦,是的。好吧,MicroGPT。我痴迷于将大语言模型简化到最本质,大概有十年或二十年了。我做过很多类似的项目,比如 NanoGPT、Make More、MicroGPT、MicroGrad 等等。我觉得 MicroGPT 现在是我试图将其简化为本质的最新成果。训练神经网络,特别是大语言模型,需要大量代码,但所有这些代码实际上都是为了效率而增加的复杂性。只是因为你需要它跑得快。如果你不需要它跑得快,只关心算法,那么那个算法实际上就是 200 行 Python 代码,非常简单易读。这包括注释等等。你有一个数据集,比如文本,然后你需要神经网络架构,大约 50 行。你需要做前向传播,然后做反向传播来计算梯度。一个自动求导引擎来计算梯度大约 100 行。然后你需要一个优化器,比如 Adam,一个非常先进的优化器,实际上也就 10 行。把所有这些放在训练循环里,大概就是 200 行。对我来说有趣的是,通常在大约一年前或更早,如果我搞出了 MicroGPT,我会倾向于向人们解释它。我会做一个视频逐步讲解之类的。我确实试着做了那个视频,还试着做了一个小指南。但我意识到这并没有增加太多价值,因为它已经如此简单,只有 200 行,任何人都可以让他们的智能体以各种方式解释它。而智能体——我不再向人们解释了。我向智能体解释。如果你能向智能体解释,那么智能体就可以充当路由器,它们可以用人类的语言、无限的耐心、根据他们的能力来针对性地解释。对吧。如果我不理解某个特定函数,我可以让智能体用三种不同的方式向我解释,而你不会那样做。没错。所以,我有点觉得,教育是什么?它曾经是指南、讲座这些东西,但现在我觉得我是在向智能体解释事情,也许我正在创造技能,技能就是指导智能体如何教授某件事的方式。所以,也许我可以为 MicroGPT 创建一个技能,描述我设想的智能体应该带你经历的进度,如果你有兴趣理解代码库的话。这只是一些提示,告诉模型先从这里开始,然后那样做。我可以把课程大纲写成技能。所以,我不觉得直接向人们解释会变少;更多的是智能体是否理解?如果智能体理解了,它们就会做解释。我们还没有完全达到那个地步,因为我仍然觉得我可能比智能体解释得稍微好一点,但我仍然觉得模型进步如此之快,以至于在某种程度上这是一场必败之战。所以,我认为教育将被这件事彻底重塑,有点像互相教学的时代结束了。比如,如果我有一个代码库,过去你会为使用你库的人写文档,但你不应该再那样做了。你应该为智能体写 Markdown 文档,而不是为人类写 HTML 文档。因为如果智能体理解了,它们就能解释所有不同的部分。这就是通过智能体的重定向。这就是原因。所以,我认为我们会看到更多这样的情况发生。我们会看看伟大的老师是否知道如何培养向智能体解释事物的直觉。
Oh, yeah. Okay, so MicroGPT. I have this running obsession of maybe a decade or two of just simplifying and boiling down LLMs to their bare essence. I've had a number of projects along these lines, like NanoGPT and Make More and MicroGPT, MicroGrad, etc. I feel like MicroGPT is now the state of the art of me trying to boil it down to just the essence. Training neural nets and LLMs specifically is a huge amount of code, but all of that code is actually complexity from efficiency. It's just because you need it to go fast. If you don't need it to go fast and you just care about the algorithm, then that algorithm actually is 200 lines of Python, very simple to read. This includes comments and everything. You have your dataset which is a text, and you need your neural network architecture which is like 50 lines. You need to do your forward pass and then your backward pass to calculate the gradients. An autograd engine to calculate the gradients is like 100 lines. Then you need an optimizer, Adam for example, which is a very state-of-the-art optimizer, is like 10 lines really. Putting everything together in the training loop is like 200 lines. What's interesting to me is that normally before maybe a year ago or more, if I had come up with MicroGPT, I would be tempted to explain it to people. I have a video stepping through it or something like that. I actually tried to make that video a little bit, and I tried to make a little guide to it. But I kind of realized that this is not really adding too much because it's already so simple that it's 200 lines that anyone could ask their agent to explain it in various ways. And the agents—I'm not explaining to people anymore. I'm explaining it to agents. If you can explain it to agents, then agents can be the router and they can actually target it to the human in their language with infinite patience and just at their capability. Right. If I don't understand this particular function, I can ask the agent to explain it to me in three different ways, and I'm not going to get that from you. Exactly. So, I kind of feel like, what is education? It used to be guides, lectures, this thing, but now I feel like I'm explaining things to agents and maybe I'm coming up with skills where a skill is just a way to instruct the agent how to teach the thing. So, maybe I could have a skill for MicroGPT of the progression I imagine the agent should take you through if you're interested in understanding the codebase. It's just hints to the model to first start off with this and then with that. I could just script the curriculum a little bit as a skill. So, I don't feel like there's going to be less of explaining things directly to people; it's going to be more of just does the agent get it? And if the agent gets it, they'll do the explanation. We're not fully there yet because I still think I can probably explain things a little bit better than the agents, but I still feel like the models are improving so rapidly that it's a losing battle to some extent. So, I think education is going to be reshuffled by this quite substantially, where it's the end of teaching each other things a little bit. If I have a library of code for example, it used to be that you have documentation for other people who are going to use your library, but you shouldn't do that anymore. You should have markdown documents for agents instead of HTML documents for humans. Because if agents get it, then they can just explain all the different parts of it. It's this redirection through agents. That's why. So, I think we're going to see a lot more of that playing out. We'll see if the great teachers know how to develop intuition for how to explain things to agents differently.
最终,比如 MicroGPT,我试着让一个智能体写 MicroGPT。我告诉它,试着把最简单的东西简化。比如,试着把我的神经网络训练简化到最简单,但它做不到。MicroGPT 是我痴迷的终点。就是那 200 行。我想了很久,我痴迷了很久。这就是解决方案。相信我,它不能再简单了。这是我的附加值。其他一切智能体都能理解。它只是不能自己想出来,但它完全理解,并且明白为什么以某种方式做等等。所以,我的贡献就是这些零碎的东西,但之后所有关于教育的事情就不再是我的领域了。所以,也许教育就是这样改变的,你必须注入你强烈认为重要的课程内容或更好的解释方式等。智能体不能做的事情现在就是你的工作。智能体能做的事情,它们可能做得比你好,或者很快就能。所以,你应该有策略地分配时间。
Ultimately, so for example, MicroGPT, I asked I tried to get an agent to write MicroGPT. So, I told it like try to boil down the simplest things. Like try to boil down my neural network training to the simplest thing and it can't do it. MicroGPT is like my end of my obsession. It's the 200 lines. I thought about this for a long time. I was obsessed about this for a long time. This is the solution. Trust me, it can't get simpler. And this is my value add. Everything else like agent gets it. It just can't come up with it, but it totally gets it and understands why it's done in a certain way etc. So, my contribution is kind of like these few bits, but everything else in terms of the education that goes on after that is not my domain anymore. So, maybe it's like education kind of changes in those ways where you kind of have to infuse the few bits that you feel strongly about the curriculum or the better way of explaining it or something like that. The things that agents can't do is your job now. The things that agents can do, they can probably do better than you or very soon. And so, you should be strategic about what you're actually spending time on.
好吧,我们很感激这些零碎的东西。谢谢你,安德烈。
Well, we appreciate the few bits. Thank you, Andre.
在 Twitter 上关注我们 @No Priors Pod。如果你想看我们的脸,请订阅我们的 YouTube 频道。在 Apple Podcasts、Spotify 或任何你收听的地方关注节目。这样你每周都能收到新剧集。并在 no-priors.com 注册邮件或查找每集的文字稿。
Find us on Twitter at No Priors Pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.