Building the Cutting Edge: Inside OpenAI with Karina Nguyen
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 研究员 Karina Nguyen 探讨团队如何构建产品、AI 进步所需技能,以及她为何从工程转向研究。
Karina Nguyen, AI researcher at OpenAI, discusses how teams build products, the skills needed as AI advances, and why she moved from engineering to research.
你不仅工作在 AI 和大语言模型的最前沿,你实际上正在构建这个前沿。当我第一次来到 Anthropic 时,我想,‘哦不,我真的很喜欢前端工程。’然后我转向研究的原因是,我意识到,‘天哪,Claude 在前端方面越来越好了。Claude 在编程方面越来越好了。我觉得 Claude 可以开发新应用。’你认为未来对产品团队来说,哪些技能会最有价值?创造性思维。你基本上需要产生一堆想法,然后筛选它们,最后构建出最好的产品体验。我认为教模型如何拥有良好的视觉设计审美,或者如何在写作中极具创造性,实际上非常非常困难。你认为人们对模型创建方式最大的误解是什么?当你教模型一些自我认知时,它实际上没有物理身体在物理世界中操作。模型会变得非常困惑。今天我的嘉宾是 Karina Nguyen。Karina 是 OpenAI 的 AI 研究员,她帮助构建了 Canvas、Tasks、01 思维链模型等。在 OpenAI 之前,她在 Anthropic 领导了 Claude 3 模型的后训练和评估工作,构建了具有 100K 上下文窗口的文档上传功能,等等。她还曾在《纽约时报》担任工程师,在 Dropbox 和 Square 担任设计师。很少有人能一窥在 AI 和大语言模型最前沿工作的人如何运作,以及他们如何看待未来的发展方向。在我们的对话中,我们讨论了 OpenAI 的团队如何运作和构建产品,她认为随着 AI 变得更聪明你应该培养哪些技能,模型是如何创建的,为什么合成数据能让模型持续变得更聪明,以及她在意识到大语言模型在编程方面会变得多好之后,为什么从工程转向研究。如果你喜欢这个播客,别忘了在你最喜欢的播客应用或 YouTube 上订阅和关注。这是避免错过未来剧集的最佳方式,也对播客帮助巨大。话不多说,有请 Karolina Noain。
Not only are you working at the cutting edge of AI and LLMs, you're actually building the cutting edge. When I first came to Anthropic, I was like, 'Oh no, I really love front end engineering.' And then the reason why I switched to research is because I realized, 'Oh my god, Claude is getting better at front ends. Claude is getting better at like coding. I think Claude can like develop new apps.' What skills you think will be most valuable going forward for product teams in particular? Creative thinking. And you kind of want to like generate a bunch of ideas and like filter through them and then just build the best product experience. I think it's actually a really, really hard to teach the model how to be aesthetic with really good visual design or like how to be extremely creative in the way they write. What do you think people most misunderstand about how models are created? When you taught the model some of the self-knowledge, you actually don't have a physical body to operate in the physical world. The model would get like extremely confused. Today, my guest is Karina Nguyen. Karina is an AI researcher at OpenAI, where she helped build Canvas, Tasks, the 01 chain of thought model, and more. Prior to OpenAI, she was at Anthropic, where she led work on post-training and evaluation for the Claude 3 models, built a document upload feature with 100K context windows, and so much more. She was also an engineer at New York Times, was a designer at Dropbox and at Square. It's very rare to get a glimpse into how someone working on the bleeding edge of AI and LLMs operates and how they think about where things are heading. In our conversation, we talk about how teams at OpenAI operate and build product, what skills she thinks you should be building as AI gets smarter, how models are created, why synthetic data will allow models to keep getting smarter, and why she moved from engineering to research after realizing how good LLMs are going to be at coding. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. It's the best way to avoid missing future episodes, and it helps the podcast tremendously. With that, I bring you Karolina Noain.
本期节目由 Interpret 赞助。Interpret 将你所有的客户互动——从 Gong 通话到 Zendesk 工单、Twitter 帖子、App Store 评论——统一起来,并使其可用于分析。它受到 Canva、Notion、Loom、Linear、monday.com 和 Strava 等领先产品组织的信任,将客户的声音带入产品开发过程,帮助你更快地构建一流产品。Interpret 的独特之处在于它能够构建和更新针对特定客户的 AI 模型,为你的业务提供最精细、最准确的洞察。将客户洞察与 CRM 或数据仓库中的收入和运营数据连接起来,映射每个客户需求的业务影响,并自信地确定优先级;通过 Interpret 的 AI 助手 Wisdom,让你的整个团队轻松对赢单分析、关键错误检测和识别流失驱动因素等用例采取行动。希望像 Notion、Canva 和 Linear 一样自动化反馈循环并自信地确定路线图优先级?访问 enterpret.com/lenny 与团队联系,并在注册年度计划时获得两个月的免费服务。这是限时优惠。网址是 interpret.com/lenny。
This episode is brought to you by Interpret. Interpret unifies all your customer interactions from Gong calls to Zendesk tickets to Twitter threads to App Store reviews and makes it available for analysis. It's trusted by leading product orgs like Canva, Notion, Loom, Linear, monday.com, and Strava to bring the voice of the customer into the product development process, helping you build best-in-class products faster. What makes Interpret special is its ability to build and update customer-specific AI models that provide the most granular and accurate insights into your business. Connect customer insights to revenue and operational data in your CRM or data warehouse to map the business impact of each customer need and prioritize confidently, and empower your entire team to easily take action on use cases like win-loss analysis, critical bug detection, and identifying drivers of churn with Interpret's AI assistant Wisdom. Looking to automate your feedback loops and prioritize your roadmap with confidence like Notion, Canva, and Linear? Visit enterpret.com/lenny to connect with the team and to get two free months when you sign up for an annual plan. This is a limited-time offer. That's interpret.com/lenny.
本期节目由 Vanta 赞助,我很高兴邀请到 Vanta 的 CEO 兼联合创始人 Christina Cassiopo 参加这个简短的对话。很高兴来到这里。我是这个播客和新闻通讯的忠实粉丝。Vanta 是我们节目的长期赞助商,但对于一些新听众来说,Vanta 是做什么的?它的目标用户是谁?当然。我们在 2018 年创立了 Vanta,专注于创始人,帮助他们开始构建安全计划,并通过 SOC 2 或 ISO 27001 等合规认证获得所有安全工作的认可。如今,我们帮助超过 9000 家公司,包括一些初创公司如 Atlassian、Ramp 和 LangChain,启动和扩展他们的安全计划,并通过自动化合规、集中化 GRC 和加速安全审查来最终建立信任。太棒了。根据我的经验,这些事情需要大量时间和资源,没有人愿意花时间做这个。这非常符合我们的经验,但在公司成立之前和某种程度上期间,我们的想法是通过自动化、AI 和软件,帮助客户高效地与潜在客户和现有客户建立信任。你知道,我们的玩笑是,我们创办这家合规公司,这样你就不用做了。我们感谢你这样做,而且你为听众提供了特别折扣。他们可以在 vanta.com/lenny 获得 1000 美元的 Vanta 折扣。网址是 vanta.com/lenny,可享受 1000 美元优惠。谢谢,Christina。谢谢。
This episode is brought to you by Vanta, and I am very excited to have Christina Cassiopo, CEO and co-founder of Vanta, joining me for this very short conversation. Great to be here. Big fan of the podcast and the newsletter. Vanta is a long-time sponsor of the show, but for some of our newer listeners, what does Vanta do and who is it for? Sure. So, we started Vanta in 2018 focused on founders, helping them start to build out their security programs and get credit for all of that hard security work with compliance certifications like SOC 2 or ISO 2701. Today, we currently help over 9,000 companies, including some startup household names like Atlassian, Ramp, and LangChain start and scale their security programs and ultimately build trust by automating compliance, centralizing GRC, and accelerating security reviews. That is awesome. I know from experience that these things take a lot of time and a lot of resources and nobody wants to spend time doing this. That is a very much our experience but before the company and to some extent during it but the idea is with automation, with AI, with software, we are helping customers build trust with prospects and customers in an efficient way and you know, our joke, we started this compliance company so you don't have to. We appreciate you for doing that and you have a special discount for listeners. They can get $1,000 off Vanta at vanta.com/lenny. That's vanta.com/lenny for $1,000 off Vanta. Thanks for that, Christina. Thank you.
Karina,非常感谢你来到这里。欢迎来到播客。
Karina, thank you so much for being here. Welcome to the podcast.
非常感谢你,Lenny,邀请我。我很高兴能邀请到你,因为你不仅工作在 AI 和大语言模型的最前沿,你实际上正在构建 AI 和大语言模型的前沿。你最近推出了一个功能,基本上是 OpenAI 的第一个智能体功能。我还刚刚做了一项调查。我不知道你是否知道这个。我对我的读者进行了一项调查,问他们每天在工作中使用哪些工具,使用最多的工具中,ChatGPT 排名第一,超过了 Gmail、Slack 和其他任何工具。90%的人说他们经常使用 ChatGPT。这太不可思议了。两年前它还不存在。是的。而且,我们录制这期节目的同一周,OpenAI 宣布了 Stargate,这是一个对 AI 基础设施的 5000 亿美元投资。所以,AI 领域一直在发生很多事情,而你有一个非常独特的视角,了解事情如何运作、发展方向以及工作如何完成。所以我有很多问题想问你。我想谈谈你在 OpenAI 如何运作和工作,你认为事情会如何发展,哪些技能在未来会更重要或更不重要,以及总体上的发展方向。你觉得怎么样?
Thank you so much, Lenny, for inviting me. I'm very excited to have you here because not only are you working at the cutting edge of AI and LLMs, you're actually building the cutting edge of AI and LLMs. You recently launched this feature which is basically the first agent feature of OpenAI. I also just did this survey. I don't know if you know about this. I did this I did a survey of my readers and asked them what tools do you use every day in your work and most use and ChatGPT was number one above Gmail, above Slack, above anything else. 90% of people said that they use ChatGPT regularly. It's It's absurd. It wasn't around 2 years ago. Yeah. Uh also, we're recording this the week that OpenAI announced Stargate which is this half trillion dollar investment in AI infrastructure. So, there's just like a lot happening uh constantly in AI and you have a really unique glimpse into how things are working, where things are going, how thing how work gets done. So, I have a lot of questions for you. I want to talk about how how operate and how you work at OpenAI, where you think things are going, what skills are going to matter more and less in the future, and also just where things are going broadly. So, how does that sound?
听起来很棒。非常感谢。嗯,是的,我非常幸运能在早期加入 Anthropic,在那里学到了很多东西,然后大约 8 个月前我加入了 OpenAI。所以,是的,我很期待多聊聊。
Sounds great. Thank you so much. Um yeah, I was extremely lucky to join early days at Anthropic and kind of learned a lot of things uh there and and I joined Open AI around like 8 months ago. So, yeah, I'm excited to talk more
好的,我肯定会问你这两者之间的区别,但我想先从更技术性的问题开始,直接切入正题。我想谈谈模型训练。人们总是听说模型被训练,这些大模型,需要多少数据、多长时间、多少钱,以及我们如何耗尽数据,这些我都想谈谈。让我先问你这个问题。
Okay, I'm going to definitely ask you about the differences between those, but I want to start more technical and just dive right in. I want to talk about model training. People always hear about models being trained, this these big models, how much data it takes, how long it takes, how much money it toss it takes, how uh how we're running out of data, which I want to talk about. Let me just ask you this question.
你认为人们对模型创建过程最大的误解是什么?
What do you think people most misunderstand about how models are created?
模型训练更像是一门艺术而非科学。在很多方面,我们作为模型训练者非常重视数据质量——这是模型训练中最重要的事情之一:如何确保为你想要创建的特定交互或模型行为提供最高质量的数据?但调试模型的方式实际上与调试软件非常相似。我在 Anthropic 早期学到的一件事,尤其是在 Claude 3 训练期间,是当你教模型自我认知,比如'你在物理世界中没有身体可以操作',但同时我们又有数据教模型函数调用,比如'这是如何设置闹钟',模型会极度困惑:它没有身体,但能否设置闹钟?所以它有时会过度拒绝,说'抱歉,我无法帮助你'。在让模型对用户更有帮助和在其他场景中不造成伤害之间总是存在权衡。关键在于让模型更加鲁棒,能够在各种不同场景中运作。
Model training is more an art than a science. In many ways, we as model trainers think a lot about data quality—it's one of the most important things in model training: how do you ensure the highest quality data for the specific interaction or model behavior you want to create? But the way you debug models is actually very similar to debugging software. One thing I learned early at Anthropic, especially during Claude 3 training, was that when you taught the model self-knowledge like 'You don't have a physical body to operate in the physical world,' but at the same time we had data teaching the model function calls like 'This is how you set an alarm,' the model would get extremely confused about whether it can set an alarm when it doesn't have a body. So it would sometimes over-refuse, saying 'Sorry, I cannot help you.' There's always a trade-off between making the model more helpful for users and not being harmful in other scenarios. It's about making the model more robust across diverse scenarios.
这太有趣了。我从没想过这一点。它训练的大部分数据都假设它是一个人类在描述世界和他们的操作方式,假设有身体并且可以做事,然后模型被告知它没有身体。
That is so funny. I never thought about that. Most of the data it's trained on assumes it's a human describing the world and how they operate, assuming there's a body and you could do things, and then the model is told it doesn't have a body.
好的。既然我们谈到这个话题,我想聊聊数据。我知道你对此有强烈的看法。有一种说法是模型会停止变聪明,因为数据用完了。它们主要是在互联网上训练的,而互联网只有一个,而且它们已经训练过了。还能向它们展示世界的什么?还有合成数据这个趋势。什么是合成数据?你为什么认为它很重要?你觉得它会奏效吗?
Okay. I want to talk a little bit about data while we're on this topic. I know you have strong opinions here. There's this meme that models are going to stop getting smarter because they're running out of data. They're trained largely on the internet and there's only one internet and they've already been trained on it. What more can you show them about the world? And there's this trend of synthetic data. What is synthetic data? Why do you think it's important? Do you think it's going to work?
我认为这里有两个问题。我们可以逐一拆解。人们说我们遇到了数据墙。我认为人们更多考虑的是在整个互联网上训练以预测下一个词的大型预训练模型。但模型在这个过程中实际学习的是如何压缩知识——它学会了建模世界。例如,预测'教我开车'之后的下一个词——只有少数词如'一辆车'匹配。所以模型学习了世界本身。它有时会建模人类行为。当你与非常大的预训练模型对话时,它们极其多样且富有创造力,因为你几乎可以通过它们与任何 Reddit 用户交谈。但现在 L1 系列的新范式是,后训练中的 Scaling 本身并没有遇到瓶颈。这是因为我们从预训练模型的原始数据集转向了通过强化学习在后训练中可以教给模型的无限任务。任何任务,比如如何搜索网络、如何使用电脑、如何写好文章——各种技能。这就是为什么我们说没有数据墙,因为会有无限的任务。而这就是模型变得超级智能的方式。我们实际上在所有基准测试上都趋于饱和。所以我认为瓶颈在于评估。我们没有所有前沿的评估,比如 GPQA,这是一个博士级别的谷歌-proof 问答。智能基准测试已经达到 60-70%以上,这正是博士的水平。所以它们实际上在评估上遇到了瓶颈。
I think there are two questions here. We can unpack one at a time. People say we're hitting the data wall. I think people think more in terms of pre-trained large models trained on the entire internet to predict the next token. But what the model is actually learning during that process is how to compress knowledge—it learns to model the world. For example, predicting the next word after 'teach me how to drive'—only a few words like 'a car' match. So the model learns about the world itself. It models human behavior sometimes. When you talk to very large pre-trained models, they're extremely diverse and creative because you can talk to almost any Reddit user through them. But what's happening now with the new paradigm of L1 series is that scaling in post-training itself is not hitting the wall. That's because we went from raw datasets from pre-trained models to an infinite amount of tasks you can teach the model in the post-training world via reinforcement learning. Any task, like how to search the web, how to use a computer, how to write well—all sorts of skills. That's why we say there's no data wall, because there will be infinite tasks. And that's how the model becomes super intelligent. We are actually getting saturated on all benchmarks. So I think the bottleneck is in evaluations. We don't have all the frontier evals like GPQA, which is a Google-proof question answering at PhD level. Intelligent benchmarks are getting to more than 60-70%, which is what PhDs get. So they're literally hitting the wall in evals.
我想顺着这两个思路继续。首先,关于合成数据的概念。简单的理解是,模型生成未来模型训练所用的数据?你让模型生成所有这些做事的方式、所有这些任务,然后更新的模型在之前模型生成的数据上训练。
I want to follow both those threads. So, the first is on this idea of synthetic data. Is the simple way to understand it that the models are generating the data that future models are trained on? And you ask it to generate all these ways of doing stuff, all these tasks as you described and then the newer models trained on this data that the previous model generated.
有些任务是合成策划的。这实际上是一个研究领域:如何用模型合成构建新任务来学习?有时,当你开发产品时,你会从产品和用户反馈中获得大量数据,你可以用这些数据进行持续学习。有时你仍然想使用人类数据,因为有些任务非常难教——专家才知道某些化学或生物学知识。所以你需要大量利用专家知识。对我来说,合成数据训练更多是为了类似产品结果的快速模型迭代。我们制作 Canvas、任务和新产品功能的方式主要是通过合成训练完成的。
Some tasks are synthetically curated. This is an actual research area: how can you synthetically construct new tasks with models to learn? Sometimes, when you develop products, you get a lot of data from the product and user feedback, and you can use that data for sustained learning. Sometimes you still want to use human data because some tasks are really hard to teach—experts only know certain knowledge about chemicals or biology. So you need to tap into expert knowledge a lot. To me, synthetic data training is more for rapid model iteration for similar product outcomes. The way we made Canvas and tasks and new product features was mostly done by synthetic training.
我们深入聊聊这个。这真的很有趣。我想谈谈评估,但先顺着这个思路。说说这如何帮助你创建了 Canvas。
Let's actually get into that. That's really interesting. I want to talk about Evals but let's follow that thread. So, talk about how this helped you create Canvas.
当我刚来到 OpenAI 时,我有一个想法:让 ChatGPT 改变视觉界面,同时也改变它与人的交互方式,这将会非常酷。从聊天机器人转变为更像一个协作智能体和合作者,是朝着更具智能体式的存在迈出的一步,最终成为创新者。整个团队由应用工程师、设计师、产品、研究人员组成,几乎是白手起家——一群人聚在一起,迅速开始迭代。Canvas 是 OpenAI 最早的项目之一,研究人员和应用工程师从产品开发周期的一开始就一起工作。
When I first came to OpenAI, I had this idea that it would be really cool for ChatGPT to change the visual interface and also change the way it interacts with people. Going from being a chatbot to more of a collaborative agent and collaborator is a step towards more agentic existence that ultimately become innovators. The entire team of applied engineers, designers, product, research kind of formed out of nothing—just a collection of people who got together and rapidly started iterating. Canvas is one of the first projects at OpenAI where researchers and applied engineers started working together from the very beginning of the product development cycle.
我认为我们在过程中学到了很多东西。我一开始就抱着快速迭代模型的心态,这样工程师更容易使用最新模型,同时也能从用户反馈或内部测试中快速学习改进。部署产品后很难预测用户会如何使用。所以,合成训练模型的关键是明确产品功能所需的核心行为。以 Canvas 为例,主要有三个行为。第一,如何触发 Canvas:对于像“帮我写一篇长文”这样的提示,用户意图是迭代长文档,或者“写一段代码”;而对于一般性问题如“给我讲讲总统的事”,用户只是想要答案,就不触发。第二,如何教模型在用户要求时更新文档。我们教模型具备一定的自主性,能够进入文档,选择特定段落并删除或添加内容,比如高亮并重写某些部分。例如,如果用户说“把第二段改得更友好一些”,模型需要找到那个段落并调整语气。所以我们既教如何触发编辑,也教如何进行高质量编辑。对于代码,还有模型是完全重写文档还是进行针对性编辑的问题,这是编辑中的另一个决策边界。最初我们偏向重写,因为觉得质量更高,但后来根据用户反馈逐渐调整。第三,我们教模型如何对任何文档进行评论。我们用一个模型模拟用户对话:比如“帮我写一份关于 XYZ 的文档”,另一个模型生成文档,然后我们注入用户提示如“评论一下,批评我的作品”,再教模型对特定部分进行评论。这一切都通过稳健的评估来衡量进展。这就是我们如何利用合成数据生成 Canvas 的。
And I think there's a lot of things we learned along the way. I came with the mindset that we need rapid model iteration so it's easier for engineers to work with the latest model, and also learn from user feedback or internal dogfooding to improve the model quickly. It's really hard to figure out how people will use a product once deployed. So synthetically training the model involves figuring out the core behaviors you want for a product feature. For Canvas, it came down to three main behaviors. First, how to trigger Canvas for prompts like 'write me a long essay' when the user intends to iterate over long documents, or 'write me a piece of code.' And when not to trigger it for general questions like 'tell me more about the president,' where the user just wants an answer. Second, how to teach the model to update the document when the user asks. We taught the model to have some agency to go into the document, select specific sections, and delete or add content, like highlighting and rewriting certain parts. For example, if the user says 'Change the second paragraph to be friendlier,' the model needs to find that paragraph and adjust the tone. So we teach both how to trigger edits and how to make higher quality edits. For coding, there's also the question of whether the model should completely rewrite the document or make targeted edits. That's another decision boundary within editing. Initially, we biased the model toward rewrites because we thought the quality was higher, but over time we shifted based on user feedback. Third, we taught the model how to make comments on any document. We used one model to simulate a user conversation: 'Write me a document about XYZ,' then another model produced the document, and we injected a user prompt like 'Make some comments, critique my piece of writing.' Then we taught the model to make comments on specific parts. It all came down to measuring progress via robust evals. That's how we used synthetic data generation for Canvas.
好的,这太有趣了。你谈到教模型这个概念,提到用合成数据教模型不同的行为。简单来说,是不是就是通过评估来展示成功的样子?就像这样:'这就是成功完成这个任务的样子',然后模型就学会了'好的,我明白了,这就是我应该做的'?
Okay, this is so interesting. So you talk about this idea of teaching the model and you mentioned how it's using synthetic data to teach the model different behaviors. Is a simple way to think about it basically that you do that by showing it what success looks like using evals? Is that the simple way to think about it? Like here's what doing this successfully would look like and that teaches it, 'Okay, I see. This is what I should do.'
没错,就是这样。你说得对。
Yeah, amazing. Yeah, you got it.
好的,明白了。我想了解一下你日常构建这些东西时的工作状态。是不是就像你坐在那里,跟某个版本的 ChatGPT 对话,然后设计这些评估?
Okay, got it. I want to start unpacking what your day-to-day looks like as you're building these sorts of things. Is it like you sitting there talking to some version of ChatGPT, crafting these evals?
有时候我会这么做。有时候我会和 ChatGPT 一起工作。实际上,我从 Anthropic 学到了很多。人们花大量时间提示模型、调试底层代码,然后你会得到很多新想法,知道如何让模型更好。比如,'哦,这个回答有点奇怪,它为什么会这样?'然后你开始调试,或者找出新方法来教模型以不同的方式回应,或者拥有更好的个性。所以模型个性的塑造方法也是类似的。不过,我觉得我在 OpenAI 的工作已经发生了变化。刚来的时候,我主要做研究方面的工作,写代码、训练模型、写评估,和产品经理、设计师一起工作,教他们如何思考评估。那真是一段很酷的经历。那是一种如何对 AI 功能或 AI 模型进行产品管理的实践。但现在主要是管理和指导。不过下午四点之后我还是会做一些研究代码。但确实,工作内容变了。
Sometimes I do that. Sometimes I do sit with ChatGPT. Actually, I think I learned this so much from Anthropic. It's like people spend so much time just prompting the models and qualifying the low-level bash all the time, and you actually get out a lot of new ideas of how to make the model better. It's like, 'Oh, this response is kind of weird. Why is it doing this?' And you start debugging it or figuring out new methods of how to teach the model to respond differently, or have better personality, let's say. So it's the same thing of how personality is made in the models. They're very similar methods. But yes, I think my time at OpenAI has changed. When I first came, I was mostly doing research-side work. I was writing code, training models, writing evals, working with PMs and designers to teach them how to think about evaluations. That was a really cool experience. It was an adoption of how to do product management of AI features or AI models. But now it's mostly management and mentorship. I still do some research code after 4:00 p.m., though. But yeah, it just kind of changed.
好了,别太强调你是个经理,因为现在大家都在裁经理。谁还需要经理呢?我最近常听到这个。开个玩笑。有趣的是,你花了很多时间教产品团队如何整合评估以及评估的重要性。我听过几次,但自己还没经历过,所以我觉得这是一个重要的线索:编写评估将越来越成为产品团队工作的一部分,尤其是在构建 AI 功能并使其良好运行的时候。你能再多谈谈这具体是什么样的吗?是不是就像坐在那里,拿着一个 Excel 表格,展示输入、输出以及结果有多好?实际工作中具体是怎样的?
All right, don't talk too much about being a manager because everyone's firing their managers. Who needs managers anymore? That's what I hear now. Just kidding. It's interesting that so much of your time was spent on teaching product teams how evals integrate and how important it is. And I've heard this a few times and I haven't personally experienced it yet, so I think it's an important thread to follow is just how writing these evaluations is going to become increasingly an important part of the job of product teams, especially when they're building AI features and working well on them. So, can you just talk a bit more about what that looks like? Is it like sitting there with an Excel spreadsheet, basically showing like here's the input, here's the output, here's how good the result was? Talk about what that actually looks like very practically.
这当然很大程度上取决于你在开发什么,但评估有多种类型。有时我会让产品经理,或者我们新设的模型设计师角色,去查看一些用户反馈。
It certainly depends a lot on what you're developing, but there are various types of evaluations. Sometimes I ask product managers or there's also a new role we have called model designers to go through some of the user feedback.
也许你会想到各种用户对话,在某些情况下应该触发 Canvas。你有真实标签:'这个对话应该触发 Canvas,那个不应该。' 这是一种非常确定性的评估。比如我们推出 Tasks 功能时,如何让模型正确安排日程?这其实很难。但我们构建了确定性评估,比如'如果用户说晚上 7 点,模型就应该说晚上 7 点。' 所以你可以有通过/不通过的评估。有时我让产品经理创建一个 Google 表格,包含不同标签页:当前行为、理想行为以及备注。他们用它来做评估或训练。如果你把表格给模型,它就能自学好的行为。第二种评估是人工评估。你可以有专门的训练师或内部人员。给定一个提示和多个模型的输出,你选择胜率——哪个模型最好,哪个产生最高质量的评论或编辑。随着你开发新模型,它们应该始终胜过之前的模型。所以这取决于你想衡量什么。
Maybe you think of various user conversations that should trigger canvas under certain circumstances. You have ground truth labels: 'With this conversation, it should trigger canvas. With this one, it should not.' That's a very deterministic eval. When we were launching tasks, for example, how do you make correct schedules? It's actually really hard for the model. But we built deterministic evaluations like, 'If the user says 7:00 PM, the model should say 7:00 PM.' So you can have pass/fail evals. Sometimes I ask product managers to create a Google Sheet with different tabs: current behavior, ideal behavior, and notes. They use it for eval or training. If you give the spreadsheet to a model, it can figure out how to teach itself good behavior. The second type of eval is human evaluations. You can have specific trainers or internal people. Given a prompt and various model completions, you choose the win rate—which model is best, which produces the highest quality comment or edit. As you develop new models, they should always win over previous ones. So it depends on what you want to measure.
这太有趣了。我听到的是,产品开发可能从'这是规格 PRD,我们一起构建,然后审查'转变为'AI 为我构建了这个,这是正确的样子。' 我把所有时间都花在定义什么是正确的,以及评估上。你肯定想衡量模型的进展。最稳健的评估是那些提示基线得分最低的,因为如果你训练了一个好模型,它应该在该评估上持续提升。但你也不想在其他智能评估上倒退。这更像是一门艺术而非科学。如果你优化某个行为,你不想损害其他智能领域。这在每个实验室都经常发生。
It's so interesting. Basically what I'm hearing is product development might shift from 'here's a spec PRD, let's build it together, then review it' to 'AI built this for me, and here's what correct looks like.' I'm spending all my time on what correct looks like, and evals essentially. You definitely want to measure progress of your model. The most robust evals are those where prompted baselines get the lowest score, because if you train a good model, it should hill-climb on that eval. But you also don't want to regress on other intelligence evals. It's more of an art than science. If you optimize for one behavior, you don't want to damage other areas of intelligence. That happens all the time in every lab.
提示也是一种原型设计新产品想法的方式。早期在 Anthropic 做文件上传功能时,我就是通过提示模型来做的。当我们推出百键联系人功能时,我在本地浏览器里做了原型。我做了演示,大家非常喜欢,他们想要文件上传的 API。那时我意识到:提示是产品开发和原型设计的新方式,适用于设计师和产品经理。例如,我想要的一个功能是个性化推荐起始提示。所以当你使用 Claude 时,它应该根据你的兴趣推荐起始提示。这完全可以通过提示来实现。另一个功能是为对话生成标题。这是一个很小的微体验,但我很自豪我们实现的方式:我们取了用户最近的五条对话,让模型判断用户的风格,然后对于下一个新对话,生成的标题会匹配那种风格。就是这些小小的微体验。
Prompting is also a way to prototype new product ideas. Early at Anthropic when I worked on the file uploads feature, I was just prompting the model. When we launched hundred key contacts, I prototyped it in the local browser. I did a demo and people loved it, and they wanted an API for file uploads. That's when it clicked to me: prompting is a new way of product development and prototyping for designers and product managers. For example, one feature I wanted was personalized recommended starter prompts. So whenever you come to Claude, it should recommend starter prompts based on your interests. You can literally do that with prompting. Another feature was generating titles for conversations. It's a small micro-experience, but I'm proud of how we did it: we took the five latest conversations from the user, asked the model what the user's style is, and then for the next new conversation, the generated title would match that style. Those little micro-experiences.
你是在 Anthropic 还是 OpenAI 做的这些?
Did you do that at Anthropic or at OpenAI?
在 Anthropic。
At Anthropic.
好的,酷。顺便说一句,我很喜欢 Claude 的文件上传功能。ChatGPT 还没有这个功能,对吗?
Okay, cool. I love the file upload feature that Claude has, by the way. ChatGPT doesn't have that yet, is that right?
我觉得它有,但实现方式很不同。
I think it has, but the implementation is very different.
好的,也许是 PDF 功能,因为我一直用 Claude 的 PDF 功能。那很酷。所以它需要跟上。天哪,你构建的很多功能我每天都在用,很多人也每天都在用。你提到的原型设计观点非常重要。在这个播客里经常提到:AI 最近对产品构建者工作最大的影响就是原型设计。不再是从 PRD 和设计出发,产品经理越来越多地直接展示一个能用的原型,你可以直接体验。
Okay, maybe it's the PDF feature, because I use it all the time with Claude. That's cool. So it needs to get on that. Man, it's wild how many features you built that I use every day and that many people use every day. The prototyping point you made is really important. It's something that comes up a ton on this podcast: how AI has most impacted the job of product builders recently is just prototyping. Instead of going from a PRD and design, PMs more and more just show a prototype of the idea that's working and you can play with it.
好的,我想多花点时间了解你的工作方式。你提到了构建和推出 Tasks 功能。你是这么描述的吗?Tasks?
Okay, I want to spend a little more time on how you operate. So, you talked about building and launching the tasks feature. Is that the way you describe it? Tasks?
是的。
Yeah.
那么,谈谈它是如何产生的,让我们更好地理解你是如何与产品团队合作的,以及 OpenAI 在这方面是如何运作的。你可以分享任何相关信息。
So, talk about how that emerged and let's better understand just how you collaborate with product teams and how OpenAI works in that way. Whatever you can share there.
我认为 Canvas 和 Tasks 属于短期或中期的项目。实际上,Canvas 和 Tasks 的诞生始于一个人做原型并创建规格。这有点像 PRD,一个关于模型行为的规格。我不认为 Tasks 是一个极其突破性的功能。它之所以酷,是因为模型非常通用——它们可以搜索、写科幻故事、搜索股票、每天总结新闻——给人们一些熟悉的东西,比如通知和提醒,非常熟悉。所以为人们创造一种使用熟悉符号的形式,比如 Canvas 就像 Google Docs,但然后你加入魔法动物,它就变得非常强大。
I think Canvas and tasks fall into the bucket of projects that are more short or medium term. Actually, the way Canvas and tasks came about was that it started with one person prototyping and creating a spec. It's kind of like a PRD, a spec of the model's behavior. I don't think tasks is an extremely groundbreaking feature. What makes it cool is that because the models are so general—they can search, write sci-fi stories, search for stocks, summarize news every day—giving something familiar to people, like notifications and reminders, is very familiar. So creating a form factor for people using familiar symbols, like Canvas which is like Google Docs, but then you add magical animal and it becomes very powerful.
但通常的操作方式是,它从一个原型开始。实际上就是一个提示原型,描述你希望模型如何为任务表现。例如,你需要进行一些设计——设计系统、设计思维。比如,如果用户说“提醒我明天早上 8 点去吃午饭”,模型需要从该提示中提取哪些信息来创建提醒?这就是你如何为一项新功能或工具设计规格。Canvas 和 Tasks 都是工具。所以关键在于如何创建工具规格。然后主要是开发 JSON 模式。例如,从这个提示中,模型应该提取用户要求的时间。然后你考虑时间格式,以及你希望模型如何通知你——基本上就是用户是否应向模型提供指令,该指令将在特定时间每天触发。例如,如果你说“每天搜索最新的 AI 新闻”,模型应将其改写为“搜索最新的 AI 新闻”,这个任务将在用户要求的时间触发。然后你设计这个工具规格。
But the way it usually works operationally is that it starts as a prototype. Literally a prompted prototype of how you want the model to behave for tasks. For example, you kind of need to design a bit—design systems, design thinking. Like, okay, if the user says, 'Remind me to go to lunch at 8:00 a.m. tomorrow,' what kind of information does the model need to extract from that prompt to create a reminder? So that's how you design a spec for a new feature, a tool. Canvas and Tasks are all tools. So it's about how you create the tool spec. Then it's mostly developing a JSON schema. For example, from this prompt, the model should extract the time the user requested. Then you think about which format you want the time in, and how you want the model to notify you—basically if the user should give an instruction to the model, and that instruction would fire off every day at that particular time. So for example, if you say, 'Search every day for the latest AI news,' the model should rewrite that into 'Search for the latest AI news,' and this task will get fired at the time the user requested. Then you design this tool spec.
实际上,我不确定——我觉得有时是通过对话。要么有人请我加入他们的团队,他们说‘天哪,我们需要战略,需要支持,需要训练模型’,要么对于 Canvas,主要是我提出了想法,然后在休息期间很快就组建了团队。所以这取决于项目。通常,团队包括产品经理、模型设计师、实际产品设计师、几位研究人员和一群应用工程师。这取决于项目的复杂性。对于 Tasks,从零到一大约花了两个月。对于 Canvas,从零到一花了四五个月。然后你教产品经理如何构建,以及如何长期思考——不仅仅是推出更好的功能,还要思考你希望 Tasks 拥有哪些酷功能。例如,Tasks 更个性化会很好,或者通过语音和移动设备创建任务。所以这就是你获得研究路线图的方式——思考功能未来如何发展。
Actually, I don't know—I feel like sometimes it's through conversations. Either people ask me to join their team and they say, 'Oh my god, we need to be strategic, we need support, we need to train the models,' or sometimes for Canvas, it was mostly that I pitched the idea and it got staffed quite immediately during the break. So it depends on the project. Usually, staffing involves a product manager, a model designer, an actual product designer, a couple of researchers, and a bunch of applied engineers. It depends on the complexity. For Tasks, it took about two months to go from zero to one. For Canvas, it was four or five months from zero to one. Then you teach product managers how to build in balls, and how to think long-term—not just ship a better feature, but think about what cool features you want Tasks to have. For example, it would be nice for Tasks to be more personalized, or to create Tasks via voice and on mobile. So that's how you get a research roadmap—thinking about how the feature will develop in the future.
从零开始,你创建带有评估的数据集,以确保一切顺利。你需要在使用的方法之间进行权衡。我之所以非常喜欢完全依赖合成数据而不是从人类收集数据,是因为它更具可扩展性。它很便宜,成本不到一半——你实际上是从模型中采样,并教授模型核心行为,这些行为会泛化到各种多样化的覆盖范围。当你推出更好的功能时,你从用户那里学到很多,从而可以调整你的合成集,使其匹配用户在产品上的行为分布。这就是你改进的方式。Canvas 从测试版到正式版发布时也是如此。
From zero, you start creating datasets with evals to make sure things go well. You need to have a trade-off between what methods you use. The reason I really love relying purely on synthetic data instead of collecting data from humans is that it's much more scalable. It's cheap, less than half—you literally sample from the model, and you teach the core behaviors in the models, and that will generalize to all sorts of diverse coverage. When you launch a better feature, you learn so much from users that you can shift your synthetic sets to match the distribution of user behavior on the product. That's how you improve. That's what happened with Canvas too, when we launched from beta to GA.
我想帮助人们理解的一件事,我自己也不完全明白,就是如何最简单地区分研究人员、模型设计师和其他参与者的工作?
Something I want to help people understand, and I don't even fully understand, is what's the simplest way to understand the job of a researcher versus a model designer and other folks involved?
我描述的项目主要是面向产品的研究。我团队的另一个部分则更多是关于长期探索性项目,开发新方法并在各种情况下理解它们。所以基本上,开发新方法遵循类似的构建评估的步骤,但更复杂。你需要有一个分布来衡量泛化能力。这更偏科学。例如,对于合成数据,最困难的事情之一是如何使其更多样化。合成数据中的多样性是当前最重要的问题之一。探索注入多样性的通用方法是一项研究探索。其他则涉及开发新能力。你研究新方法,获得初步成功的迹象,然后思考如何使其更通用或有用。这就是长期项目如何变成中期或短期项目。
The projects I described are mostly product-oriented research. Another part of my team is more about longer-term exploratory projects, developing new methods and understanding them under various circumstances. So basically, developing new methods follows a similar recipe of building evals, but more sophisticated. You want to have a distribution to measure generalization. It's more science-y. For example, with synthetic data, one of the hardest things is how to make it more diverse. Diversity in synthetic data is one of the most important questions right now. Exploring ways to inject diversity as a general method that works for all is one research exploration. Others involve developing new capabilities. You work on new methods, get signs of life, then think about how to make it more general or useful. That's how longer-term projects become medium or short-term.
有道理。本质上是在开发让模型更聪明的方法。
That makes sense. Essentially working on developing ways to make the model smarter.
是的。像 o1 这样的新方法是一个重大突破,对吧?它的运作方式——不仅仅是‘这是你的答案’。它实际上会思考,并花时间思考得出答案的过程。
Yeah. New ways like o1 was a big breakthrough, right? The way it operates—it's not just 'here's your answer.' It actually thinks and takes time to think through the process of coming up with an answer.
说到思考未来走向,我想花点时间聊聊这个洞察——你基本上是在构建 AI 的最前沿,处于 AI 发展方向和现状的最尖端。所以我很好奇你的看法:根据你看到的趋势,世界会如何变化,人们的工作方式会如何改变?我知道这个问题很宽泛,但就说未来三年吧。你如何看待世界的变化?人们的工作方式会如何改变?
Speaking of that of thinking about the future where things are going, I want to spend some time on just this insight that basically you are building the cutting edge of AI. Like at the very bleeding edge of where AI is going and where it is. And so I'm very curious to hear just your take on how you think things are going to change in the world and how people work based on where you see things are going and I know this is a broad question, but let's say like in the next 3 years. How do you see the world changing? How do you see people's way of working changing?
在两个实验室工作是非常谦卑的经历。对我来说,刚来 Under Brain 时,我心想:“天哪,我真的很喜欢前端工程。”后来我转向研究,是因为那时我意识到:“天哪,Claude 在前端方面越来越强了,编码能力也越来越好。我觉得 Claude 可以开发应用之类的,甚至能为我的项目开发新功能。”这有点像一种元认知——世界真的在变。当我们首次推出 100K 上下文窗口时,我就在思考形态因素。文件上传很自然,人们很熟悉。但你可以想象,我们可以在 Claude AI 应用中创建无限图表,就像在 100K 上下文中一样。但因为文件上传遵循“形式服从功能”,这种形态让人们可以上传任何东西——书籍、财务报告,然后向模型提问任何任务。我记得企业客户,比如金融客户,对此非常感兴趣。他们发现这实际上是他们工作中非常常见的任务。看到一些冗余任务被这些智能模型自动化,真是疯狂。我们正在进入一个时代,有时我甚至不知道模型给出的答案是否正确,因为我不是那个领域的专家,我甚至不知道如何验证模型的输出。只有专家才知道并能够验证。所以,基本上有几个趋势。第一个趋势是推理和智能的成本急剧下降。我写过一篇博客文章,也许我应该用最新的基准更新它,因为那时大家都在做某个基准,然后很快就饱和了。现在我们需要用另一个前沿评估来写同样的博客。但智能成本在下降,因为它变得更便宜。智能小模型甚至变得比大模型更聪明,这得益于蒸馏研究。比如 Claude 3 Haiku,我在负责它时发现它比 Claude 2 聪明得多,而 Claude 2 要大得多。小模型变得非常智能、快速且便宜。我们正在走向那个世界,这有多重影响。好消息是,人们将更容易获得 AI,这对构建者和开发者非常有利。但这也意味着所有被智能瓶颈限制的工作都将被解锁。比如在医疗领域,我不必去看医生,而是可以问 ChatGPT 或给它一系列症状,让它判断我是感冒、流感还是其他什么。我几乎可以接触到医生,而且已经有相关研究。是的,《纽约时报》有篇报道比较了医生、使用 ChatGPT 的医生和纯 ChatGPT,结果纯 ChatGPT 表现最好,医生反而更差。这很疯狂,对吧?教育方面,如果我年轻时拥有今天的工具,我会学到更多。现在人们几乎可以从这些模型中学到任何东西——新语言、如何构建新应用等等。推出 Canvas 并将其带给人们,让他们能做以前做不到的事情,这很谦卑,也很有魔力。所以教育将产生巨大影响,还有科学研究。我认为任何 AI 研究的梦想都是自动化 AI 研究,这有点可怕。这让我觉得人员管理会保留下来,因为情感智能和创造力本身是最难的事情之一。所以作家们不必太担心,我认为 AI 会减轻很多冗余任务。
It's a very humbling experience to be in both labs, I guess. Like to me when I first came to Under Brain, I was like, "Oh my god, I really love front end engineering." And then like the reason why I switched to like research is because I realized at that time is like, "Oh my god, like Claude is getting better like front end. Like Claude is getting better like coding. I think Claude can like developing your apps or something. And so like it can like develop new features for the thing that I'm working. So it's like it was kind of like this meta realization where it's like oh my god like the world is actually changing and they're like when we first like launched 100K context at that time obviously, you know, I'm thinking about like form factors. It's like yeah like file uploads were like very natural, very familiar to people. But you can imagine we could just like make like infinite charts in the Claude AI app, right? Like as if like it's like a in 100K context. But because like file uploads is like form follows function is like the form factor of the file uploads kind of enable people to just like literally upload anything, the books and like any reports financial and like ask any task to the model. And then I remember it was like you know, enterprise customers like like financial customers were like really interested in that. It's like oh wow like actually they it's actually one of the very common tasks that people do in that setting. It's like kind of crazy to like see how some of the redundant tasks are getting like automated basically by these like smart models. And we're entering them the era where I actually don't know for example sometimes like if a one gives me the correct answer or not because I'm not an expert in that field. And it's like I don't even know how to verify the outputs of the models. It's because like all my experts know that and like they can like verify this. So yes. So basically there are trends that are going on. The first trend is the cost of reasoning and intelligence is drastically going down. I had a blog post about this. Maybe I should update them like latest benchmarks because at that time like everybody was like doing like um so like one benchmark and then you'd like quickly saturate that benchmark. So like now we need to like do the same blog post with with another like frontier eval. But the cost of intelligence is like going down because it's becomes like much cheaper. Smart small models are becoming even smarter than like large models and that's because of like the distillation research. This happened with like Claude 3 Haiku. I was like working like for senior like Claude 3 Haiku and I realized it was much smarter than like Claude 2 which was like way you know bigger like something like that. Um but like the power of like small models become very intelligent and fast and cheap. We are moving towards that world. That has like multiple implications. But the news that like people will have more access to AI and that's really good. Like builders and developers will have much better access to AI. But also it means like all the work that has been like bottlenecked by the intelligence will be kind of like unblocked. So anyone like in I'm thinking about healthcare, right? Like if I have instead of like going to the doctor, I can like ask ChatGPT or give ChatGPT a list of symptoms and ask me like oh which like would I have like a cold, flu, or like something else? Like I can literally get the access to like uh doctor almost and there's like been some like research studies around that. Yeah, there's a New York Times story about that where they compared doctors to doctors using ChatGPT to just ChatGPT and just just ChatGPT was the best Yeah. of them all. Like, doctors made it worse. Yeah, that's crazy, like right? Like, education, I think uh I would have dreamt if like I had the tools that I have today when I was like younger in life would learn so much. But it's like people can now learn almost anything from these models. So, they can learn new language, they can learn how to build new book apps, like I don't know, anything that you want in like I'm so like it's humbling to like have like launch canvas and like bring that thing to the people, enable them to do something else that they couldn't have ever before. I think this is there's something like magical around this experience. Uh so, education has will have massive implications, like I guess like scientific research, right? Like, I I think it's like the dream of like any AI research is like automate AI research. Uh it's kind of scary, I'll say. Um which makes me think that like people management will stay, you know? It's like one of the hardest things to It's like emotional intelligence for the models like creativity in itself is like one of the hardest things. So, writers I I don't think like people should be worried as much. I think it's like I think it alleviate a lot of like redundant tasks uh for people.
太棒了。好的,我肯定想顺着这个话题聊下去。有趣的是,你描述的情况是:你在 Anthropic 做工程师,然后觉得 Claude 会在工程方面变得非常出色,这可能不是长期的职业方向,所以你转向了研究。而 AI 在很长一段时间内仍然需要你来构建它、让它更智能。
This is awesome. Okay, I want to follow this thread for sure. And it's funny that what you described is like you were an engineer at Anthropic and you're like okay, Claude is going to be very good at engineering. This isn't going to be a potentially career long term, so I'm going to move into research. And the AI is going to need me for a long time to build it, to make it smarter.
我想说,Canvas 团队仍然有非常棒的前端工程师,我很喜欢他们,那些真正关心交互设计和交互空间的人。我不认为模型已经达到那个水平。但如果我们能让模型达到前端领域的前 1%,那当然有可能。
I would say we still have like I think Canvas team has still have like a really cool like front engineers that I really like, you know, people who like really care about like interaction design, like interaction space. Like, I don't think like models are there yet. Like, I think if but we can get the models to like this top 1% of like front end or something, um, for sure.
接下来我想沿着这个方向问一个问题,这只是猜测:你认为未来对产品团队来说,哪些技能会最有价值?听众可能会想,这很可怕。我现在应该培养什么技能才能保持领先,避免将来陷入困境?你认为哪些技能会变得越来越重要?
So what I want to move on to next along these lines is just and this is just speculation, but uh, what skills do you think will be most valuable going forward for product teams in particular? So folks are listening and they're like, okay, this is scary. What should I be building now to help me stay ahead and not be in trouble down the road? What skills do you think are going to be most it more and more important to build?
是的,我认为是创造性思维。你需要产生大量想法,然后筛选它们,最终构建出最好的产品体验。倾听。
Yeah, I think like creative thinking. Like you kind of want to like, um, come up with generate a bunch of ideas and like filter through them and then just like build the best product experience. Listening.
你知道,你想构建一些东西,让最通用的模型无法取代你。通常你构建一个东西,让它对特定用户群非常好用。当模型收到用户反馈后,你倾听他们的意见并快速迭代。我不认为我们已经到了那个阶段。有太多想法了。我觉得有大量的想法可以去做,所以我不会担心。事实上,我希望 AI 领域的人能更有创意一些,能跨领域连接点,开发出真正酷的新一代 AI 交互范式。我不认为我们已经解决了这个问题。几年前,我告诉一些人,你应该为未来而构建。现在模型好不好并不重要,但你可以构建产品创意,等到模型真正变好时,它就会非常好用。我觉得这自然发生了。比如 Claude 的 Artifacts。我觉得 Canvas 的早期可以追溯到 2022 年,在 ChatGPT 之前,比如写作 ID。但 Claude 1.3 模型本身还不足以做出非常高质量的编辑,比如编程。我觉得像 Character 这样的初创公司做得非常好,因为他们迭代很快,发明了新的模型训练方式,行动迅速,倾听用户,并且有大规模的分发。这很酷。
You know, you want to build something that the most general model will not replace you. Often you build something and make it really good for a specific set of users. After the model is in user feedback, you listen to them and rapidly iterate. I don't think we are there yet. There are so many ideas. I feel there's an abundance of ideas you can work on, so I wouldn't be worried. In fact, I wish people in AI were a bit more creative and connecting dots across different fields to develop really cool new generations and new paradigms of interaction with AI. I don't think we've cracked this problem at all. A couple of years ago, I was telling some people, you kind of want to build for the future. It doesn't necessarily matter whether the model is good right now, but you can build product ideas such that by the time models are really good, it will work really well. I think it just happened naturally. For example, Claude artifacts. I feel like early days of Canvas was back in 2022, before ChatGPT, like writing ID. But Claude 1.3 model itself was not there to make really high-quality edits, for example, coding. I feel like startups like Character are doing super well because they iterate so fast, invent new ways of training models, move fast, listen to users, and have massive distribution. It's kind of cool.
这真的很有帮助。所以我听到的是,软技能本质上会变得越来越重要和强大。你刚刚谈到了管理、领导他人、创造力、创新见解、倾听。我写了一篇文章(我会附上链接),分析 AI 将如何影响产品管理。我们实际上非常一致。我的感觉是同样的:软技能会变得越来越重要,而被取代的将是硬技能。这很有趣,因为通常人们看重硬技能,比如编程、设计、写作。有趣的是,AI 在这些方面确实很擅长,因为它能获取大量数据,综合起来,然后写作或创造东西,而所有这些模糊的事情,比如影响、说服他人、对齐、倾听、创造力,则不同。我这么说的时候,你有什么想法吗?
That's really helpful actually. So what I'm hearing is that soft skills essentially are going to be more and more important and powerful. You just talked about management, leading people, being creative, coming up with innovative insights, listening. There's a post I wrote that I'll link to where I analyze how AI will impact product management. We're actually very aligned. My sense was the same: soft skills are going to become more important, and the things that will be replaced are hard skills. That's interesting because usually people value hard skills like coding, design, writing really well. And it's interesting that AI is actually really good at that because it takes a bunch of data, synthesizes it, and writes or creates a thing, versus all these fuzzy things around influencing, convincing people, aligning, listening, creativity. Anything along those lines come up as I say that?
我认为教模型如何具有审美感、做非常好的视觉设计,或者如何以极具创意的方式写作,实际上非常非常困难。我仍然认为 ChatGPT 在写作方面很糟糕,那是因为它受限于创造性推理。我认为优先级排序对管理者来说是最重要的事情之一。我觉得 AI 研究的进展受限于研究管理者,因为算力是有限的,你需要把算力分配给你最看好的研究方向。你需要对研究方向有非常高的信念才能投入算力;这更像是一种投资回报的情况。我经常思考,在所有项目中,哪些优先级更高?这是优先级排序,同时在更低的层面,哪些实验现在真正重要,哪些不重要,要能分清主次。所以我认为优先级排序、沟通、管理、人际技能、同理心、理解他人、协作都很重要。我觉得 Canvas 如果不是以人为本,就不会有如此成功的发布。我有机会与李·拜伦(联合创始人)、拉斐尔以及一些最优秀的苹果设计师合作。看到如何创造人与人之间的协作,这真的很酷。这仍然是人性化的东西。
I think it's actually really, really hard to teach the model how to be aesthetic or do really good visual design, or how to be extremely creative in the way they write. I still think ChatGPT kind of sucks at writing, and that's because it's bottlenecked by creative reasoning. I think prioritization is one of the most important things for managers. I feel like AI research progress is bottlenecked by research managers because you have a constrained set of compute and you need to allocate compute to the research paths you feel most committed about. You need to have a really high conviction in the research paths to put the compute; it's more of a return on investment situation. I'm thinking a lot about how across all projects, which are higher priority? It's prioritization, and also at a lower level, which experiments are really important to run right now and which are not, and cut through the line. So I think prioritization, communication, management, people skills, empathy, understanding people, collaboration. I think Canvas wouldn't have been an amazing launch if it wasn't about people. I got a chance to work with people like Lee Byron, who is a co-creator, Raphael, and some of the best Apple designers. It's so cool to see how do you create this collaboration between people. It's something that's still humane.
让我稍微深入一下,因为我想象听众会说,好吧,但一旦我们有了 AGI 或 SGI,它就能做所有这些事情。有一个世界,为什么这些还没完成?我认为很容易就假设所有事情都会被解决。我很好奇关于创造力和倾听的想法,为什么你认为 AI 不擅长这些?除了很难训练它做好之外,有没有什么原因让 AI 和 LLM 特别难以擅长这些?
Let me just follow through a little bit because I imagine people listening are like, okay, but once we have AGI or SGI, it'll do all this. There's a world where why isn't all this done? I think it's easy to just assume all that. I'm curious about this idea of creativity and listening, why you think AI isn't good at it? Other than it's just very hard to train it to do this well. Is there anything there of just why this is especially difficult for AI and LLMs to get good at?
我认为目前有很多原因导致它困难。我认为这仍然是一个不太好的研究领域,也是我的团队正在努力的方向:如何教模型在写作中更有创意?我认为这种让模型更多思考的新范式本身应该会带来更好的写作。但说到创意生成或辨别什么是好的视觉设计和艺术,我觉得模型还没有从人类那里学到足够的例子来很好地辨别。我确实认为这是因为真正擅长这些的人并不多,而且模型无法接触到这些人来学习。所以我想这就是它糟糕的原因。
I think currently it's difficult for many reasons. I think it's still a not-good research area, and it's something my team is working on: how do we teach the model to be more creative in writing? I think this new paradigm of models thinking more should actually lead to better writing in itself. But when it comes down to idea generation or discrimination of what is good visual design and art, I feel like it hasn't learned enough examples from people to discriminate very well. I do think it's because there are not that many people who are actually really good at these things, and it's not accessible for models to learn from these people. So I guess that's why it sucks.
是的,有道理。基本上,像你这样的人还不够多。研究人员把教 AI 做这些事情视为任务,那些拥有极好品味和创造力的人可以教这些。你可以说这将来会实现,但我不……我们不需要继续深入这个话题。让我问你一个具体问题。在我写的那篇文章中,我提出了一个很多人不同意的观点:策略是 AI 工具会越来越擅长并接管的事情。有一种感觉是,策略是人类会继续做得更好的事情,你不能把制定策略、告诉你如何取胜的任务交给 AI。
Yeah, that makes sense. Basically, there's not enough of you yet. Researchers treat teaching it to do these things, people that have incredible taste and creativity that can teach these things. You could argue this will come, but I'm not... We don't need to keep going down that thread. Let me ask you a specific question. In this post I wrote, I made this argument that a lot of people disagreed with: that strategy is something that AI tooling will become increasingly great at and take over. There's this sense that that's the thing that people will continue to be better at and you can't offload to AI basically developing your strategy, telling you what to do to win.
我的观点是,策略不就是获取所有输入、所有可用数据,理解周围的世界,然后制定一个获胜的计划吗?感觉 AI,比如语言模型,在这方面会非常聪明。你怎么看?
My case is isn't strategy just take all the inputs, all the data you have available, understand the world around you and come up with a plan to win? Feels like AI would be like an LM would be incredibly smart at this. What's your take?
我也这么认为。我们会教模型各种工具、能力和推理。比如,对于现在的 Canvas,模型可以汇总所有用户反馈,总结出用户体验中最痛苦的五个缺陷,然后根据自身的构建方式,想办法为自己创建数据集进行训练。我认为这种自我改进离我们并不遥远——模型通过产品开发实现自我提升,就像一个有机体。策略更像是数据分析和连接点。模型可以从一个来源获取用户反馈,结合内部指标仪表盘和其他输入,共同制定计划或提出建议。我认为这是当今 ChatGPT 最常见的用例之一。
I think so, too. I think we would teach the model all sorts of tools, capabilities, and reasoning. For example, with Canvas right now, it would be very cool for the model to aggregate all user feedback, summarize the top five most painful flaws in user experience, and then, knowing how it's built, figure out how to create datasets for itself to train on. I don't think we are far away from that kind of self-improvement—models becoming self-improving through product development, like an organism. Strategy is more about data analysis and connecting the dots. The model can take user feedback from one source, internal dashboards with metrics, and other inputs, then co-create a plan or recommendations. I think this is one of the most common use cases for ChatGPT today.
有道理。本质上,人类一次只能理解有限的信息,查看有限的数据来综合要点。而且正如你所说,现在的上下文窗口非常大。这里有所有信息,那么最重要的事情是什么?
That makes sense. Essentially a human can only comprehend so much information at once and look at so much data at once to synthesize takeaways. And as you said, these context windows are huge now. Here's all the information. What's the most important thing I should do?
是的,科学研究也是如此。理想情况下,模型能够提出新想法,根据先前实验的经验结果迭代实验,并提出新方法。
Yeah, same as in scientific research. Ideally, the model will be able to suggest new ideas, iterate on experiments given empirical results from previous experiments, and come up with new methods.
哦,天哪。好的,为了结束这个话题,你建议人们重点培养和依赖的技能是软技能,比如创造力、管理影响力、协作、寻找模式。这大致是你的想法吗?
Oh, man. Okay, so just to close the loop on this conversation, this part of the thread is the skills you're suggesting people focus on building and leaning into are soft skills like creativity, managing influence, collaboration, looking for patterns. Is that generally where your mind is at?
是的,我一直在思考如何更有效地建立关系。我认为这主要是管理——如何组织研究团队或一般团队,使他们达到最大绩效。这关乎扩展组织或扩展产品研究。低效的管理正在限制人类物种的潜力。
Yeah, I'm thinking a lot about how to make relations more effectively. I think this is mostly management—how to organize research teams or general teams so that they achieve maximum performance. It's about scaling organizations or scaling product research. Inefficient management is limiting the potential of the human species right now.
没错。研究团队内部的管理不善,以及 OpenAI、Anthropic 和其他一些模型。是的,想想都觉得疯狂。天哪。
Right. mismanagement within the research team and OpenAI and Anthropic and some of these other models. Yeah, it's kind of crazy to think about it. Holy moly.
好的,说到 Anthropic 和 OpenAI,你两家都工作过。很少有人在这两家公司都工作过,并了解它们的运作方式。我很好奇你注意到这两者之间的差异,它们如何运作、如何思考、如何处理事情。你能分享些什么?
Okay, so speaking of Anthropic and OpenAI, you've worked at both. Very few people have worked at both companies and have seen how they operate. I'm curious just what you've noticed about the differences between these two, how they operate, how they think, how they approach stuff. What can you share along those lines?
相似之处多于不同之处。当然,在文化细微差别上有些不同。我非常喜欢 Anthropic,在那里有很多朋友,我也喜欢 OpenAI,在那里也有很多朋友。这不是敌对的;实际上,这是一个做同样事情的大社区。我从 Anthropic 学到的是对道德行为、道德工艺和道德训练的真正关怀和匠心。我一直在思考是什么让 Claude 成为 Claude,让 ChatGPT 成为 ChatGPT。我认为这归结于导致模型输出的操作流程。Claude 之所以有更多个性,更像一个图书管理员——一个非常书呆子气的人——是因为它反映了创造它的团队。在角色、个性、模型是否应该跟进问题以及正确的道德行为方面有很多细节。这就是我在 Anthropic 学到的艺术部分。Anthropic 规模小得多;我加入时大约 70 人,离开时 500 人,文化变化很大。我很享受早期创业公司的氛围,人们像家人一样彼此了解。我认为 Anthropic 更擅长聚焦和优先级排序——非常硬核的优先级排序。但我认为 OpenAI 在产品或研究方面更具创新性和冒险精神。在 OpenAI,你有更多的产品自由,几乎可以做任何事情;它更自下而上。人们提出想法并尝试,从而推出更多产品。
It's more similar than different. Obviously, there are some differences in cultural nuances. I really love Anthropic and have many friends there, and I also love OpenAI and still have many friends there. It's not about enemies; it's actually one big community of people doing the same thing. What I learned from Anthropic is a real care and craft towards moral behavior, moral craft, and moral training. I've been thinking about what makes Claude Claude and what makes ChatGPT ChatGPT. I think it comes down to operational processes that lead to the model outputs. The reason Claude has so much more personality and is more like a librarian—someone very nerdy—is because it reflects the creators who made it. There are many details around character, personality, whether the model should follow up on a question, and what the correct ethical behavior is. That's where I learned the art part at Anthropic. Anthropic is much smaller; when I joined it was about 70 people, when I left it was 500, and the culture changed a lot. I enjoyed the early startup vibes where people knew each other like family. I would say Anthropic is much better at focusing and prioritization—very hardcore prioritization. But I think OpenAI is much more innovative and more risk-taking in terms of product or research. At OpenAI, you have more product freedom to do almost anything; it's more bottoms-up. People bubble up ideas and try things, leading to more products launching.
是的,我也是这么想的。感觉 OpenAI 更自下而上,分布式,人们提出想法并尝试。这导致更多产品推出,更多事情被尝试,而另一种风格则是确保我们做的每件事都出色、精良,并对每项投资深思熟虑。
Yeah, that's how I was thinking about it. It feels like OpenAI is more bottoms-up distributed people bubble up ideas, try stuff. There's more and that leads to more products launching I imagine more things just kind of being tried versus more of a let's just make sure everything we do is awesome and great and craft and thinking deeply about every investment.
我从未听过这样的描述。Karina,我们覆盖了太多内容,这能帮助很多人从多种角度思考未来的方向。在进入激动人心的快问快答环节之前,我想知道你是否还有其他觉得值得分享或深入的话题。
I've never heard it described this way. Karina, we've covered so much ground. This is going to help a lot of people with so many ways of thinking about where the future is going. Before we get to our very exciting lightning round, I'm curious if there's anything else that you think might be helpful to share or get into.
我在 Anthropic 早期的一个遗憾是,在 ChatGPT 出现之前,我们有奢侈的时间带着一堆想法几乎每天做原型。我们做了很多酷想法,比如 Slack 里的 Claude,这是最早的工具使用产品之一。Claude 可以在你的工作场所操作——比如当你@Claude 来总结一个讨论串时。如果你和某人有一整段对话,想要一个总结,你可以说'@Claude 总结这个'。迭代模型本身也很有趣。在 Slack 里一直和模型对话创造了一种社交元素,有点像 Discord 里的 Midjourney。人们学到了很多关于提示工程以及如何与 Claude 协作的知识。早期的一个原型是每周一 Claude 会总结整个频道,或者每周五总结一堆频道并发布组织新闻。那是一种很酷的产品形态。我认为产品形态是 AI 中一个非常重要的问题。我们甚至还没弄清楚如何用所有这些模型创造出色的产品体验。范式从同步实时给出答案转变为更异步的智能体在后台工作的范式。但问题在于:智能体应该与你建立信任,而信任是随着时间建立的,就像人与人之间一样。你开始一种协作——你和模型之间的这种协作模式非常重要,因为你们彼此信任,模型从你的偏好中学习,从而变得更加个性化。它会开始预测你接下来想在电脑上执行的操作。这就像我们从个人电脑走向了个人模型。
One of my regrets from early days at Anthropic was that there was a luxury of time pre-ChatGPT to come in with a bunch of ideas and prototype almost every day. We did a lot of cool ideas like Claude in Slack, which was one of the first tool-using products. Claude could operate in your workplace—like when you @Claude to summarize a thread. If you had an entire conversation with someone and wanted a summary, you could say '@Claude summarize this.' It was also fun to iterate on the model itself. Talking to the model in Slack forever created a social element, kind of like Midjourney in Discord. People learned so much about prompting and how to work with Claude. One early prototype was every Monday Claude would summarize the entire channel, or every Friday summarize a bunch of channels and give news about the organization. It was a really cool form factor. I think form factor is a really important question in AI. We haven't even figured out how to create an awesome product experience with all these models. The paradigm shifts from synchronous real-time answer giving to a more asynchronous paradigm of agents working in the background. But then the question is: agents should build trust with you, and trust builds over time, like with humans. You start a collaboration—that collaboration model between you and a model is so important because you both trust and the model learns from your preferences, so it can become more personalized. It will start predicting the next action you want to take on the computer. It's like we went from personal computers to personal models.
为什么这不是一个标配功能?这看起来如此明显,每个大语言模型都应该有一个 Slack 机器人版本。这是我现在可以安装的东西,还是目前还没有?
Why isn't that a thing? That seems like such an obvious feature that every LLM should have as a Slack bot version of themselves. Is that something I can have installed, or is that not a thing right now?
我知道 Slack 里的 Claude 在 2023 年左右被停用了,但那是因为 ChatGPT 之后,焦点转向了消费者或企业用例。当你想讨论新功能时,Slack 里 Claude 的产品形态有些受限。
I know that Claude in Slack was sunsetted in 2023 or something, but that's because after ChatGPT, the focus shifted to consumer or enterprise use cases. The form factor of Claude in Slack was a bit constrained when you want to talk about new features.
我想要那个。我知道 ChatGPT 也有 Slack 机器人。所以我不知道,也许它某天会回来。我愿意为此付费。早期还有其他回忆吗?因为作为 Anthropic 早期的一员,那是一个很特别的地方。还有什么有趣的回忆或故事可以分享吗?
I want that. I know that ChatGPT had Slack bots too. So I don't know, maybe it will come back sometime. I would pay for that. Any other memories from that time of early days? Because that's a really special place to have been as early days Anthropic. Any other memories or stories from that time that might be interesting to share?
我认为第一次让我们感觉'再次点击新闻'的发布是 100K 上下文窗口的发布,模型可以输入整本书并给出总结,或者整个多文件财务报告并回答非常具体的问题。那里有些东西让我觉得'天哪,这真是一个很酷的新能力'——不仅仅是模型能力,更多是来自产品形态本身的能力。我们考虑的其他原型包括 Claude 中的共享工作区,我和 Claude 会有一个共享文档,我们可以一起编辑。我觉得有时候想法会停滞并锁定大约两年,就像这个案例一样。
I think the very first launch when we felt like 'click through news again' was the 100K context launch, where the model could input an entire book and give you a summary, or entire multi-file financial reports and answer very specific questions. There was something in there that made me think, 'Oh my god, this is a really cool new capability'—not just model capability, but more the capability that came from the product form factor itself. Other prototypes we were thinking about included a shared workspace in Claude, where Claude and I would have a shared document that we could both edit. I feel like sometimes ideas take a slog and lock for about two years, just like in this case.
有趣的是,这些里程碑打开了我们对正在发生什么以及走向何方的视野。ChatGPT,我认为是第一个让人感叹'哇,这比我想象的好太多了'。你提到了 100K 上下文窗口,可以上传一本书并提问和总结。我经常在采访嘉宾时使用这个功能,如果他们写了书,我有时没时间读完整本书,就用它来帮我理解最有趣的部分,然后我再深入阅读,澄清一下。然后,我不知道,也许语音是另一个,比如你可以和 ChatGPT 对话。还有其他让你觉得'哇,这比我想象的好太多了'的时刻吗?
It's interesting there are these milestones that kind of open up our view of what is happening and where things are going. ChatGPT, I think, was the first of just like, 'Wow, this is much better than I would have thought.' You talked about 100K context windows where you can upload a book and ask it questions and have it summarized. I actually use that all the time when I have interview guests and they wrote a book. I sometimes don't have time to read the whole book, so I use it to help me understand what the most interesting parts are and then I actually dive into the book, just to be clear. And then I don't know, maybe voice was another one where you could talk to say ChatGPT. Is there any other moments there that you're like, 'Wow, this is much better than I thought it was going to be?'
是的,我认为计算机使用智能体——模型操作桌面——你基本上可以想象一种新的体验,模型可以学习你的浏览方式,并根据这种偏好像你一样浏览。这有点像模拟的人格。它实际上非常类似于这样的想法:'好吧,也许 Sam Altman 没有很多时间。也许我想和他的模拟器、他的模拟对话,问他问题。'这样的模拟环境会非常酷。
Yeah, I think the computer use agents—the model operating the desktop—you can essentially think of a new kind of experience where the model can learn the way you browse and from that preference it can browse just like you. It's kind of like a simulated persona. It's actually very similar to the idea of, 'Okay, maybe Sam Altman doesn't have a lot of time. Maybe I want to talk to his simulator, his simulation, and ask him questions.' Simulated environments like this would be really cool.
这是一个很好的地方来推荐 Lenny 机器人。我有一个这样的机器人。它在我所有的播客和新闻通讯上训练过。它运行在多个模型上。我不知道他们具体用哪个模型,但正是这样。它甚至不是我本人,而是所有上过播客的嘉宾和我写的新闻通讯。你可以直接问它:'我如何增长我的产品?如何制定策略?'它实际上好得惊人。你觉得它反映了你是谁吗?就像它会是什么样子?好吧。最好的部分是你可以和它对话。
That's a great place to plug Lenny bot. I have one of those. It's trained on all of my podcasts and newsletters. And it sits on many models. I don't know which one exactly they use, but it's exactly that. It's not even me; it's all the guests that have been on the podcast and newsletters that I wrote. And you could just ask it, 'How do I grow my product? How do I develop a strategy?' And it's actually shockingly good. Do you feel like it reflects who you are? Like what it would be? Okay. The best part of it is you can talk to it.
它已经建成了。有一个基于这个播客中我的声音训练的 11 Labs 语音版本,效果非常好。有人告诉我他们能跟它聊上几个小时。哇。还有人让它“像在 Lenny 的播客里一样采访我,问问我职业生涯的问题。”它真的和 Lenny 机器人做了一期半小时的播客节目。太有趣了。太不可思议了。未来真疯狂。
It's built. There's an 11 Labs voice version that's trained on my voice from this podcast. And it's actually very good. People have told me they sit there for hours talking to it. Wow. And somebody told it, "Interview me like I am on Lenny's podcast. Ask me questions about my career." And it did a half-hour podcast episode with Lenny bot. That's so fun. It's incredible. Future is wild.
是的,我认为内容转化就像,你知道,我设想有一天,当你在 Canvas 中生成一个科幻故事时,你可以把它转化成音频博客,就像非常自然的内容从一种媒体到另一种媒体的转化。我最早的灵感之一来自《西部世界》的最后一集,多洛雷斯来到一个新的工作空间,开始写故事。她一边写,一个 3D 虚拟现实世界就即时生成。所以我有点想创造那样的东西。
Yeah, I think content transformation is like, you know, I imagine sometime, when you generate a sci-fi story in Canvas, you can transform this into an audio blog, like very natural content transformation from one media to another. One of my earliest inspirations is one of the last episodes of Westworld, where Dolores comes to a new workspace and starts writing a story. And as she writes, a 3D virtual reality world is created on the fly. So I kind of want to create that.
说到媒介,我在想是否该往这个方向走,但很快说一下,OpenAI 的 CPO Kevin Weil。他去年在 OpenAI Friends Summit 做了一个小组讨论,提出了一个非常有趣的观点:聊天是这些工具的一个非常有趣的界面,因为它们变得越来越智能,而聊天作为一种与人交互的范式仍然有效。你可以和爱因斯坦聊天,也可以和一个不太聪明的人聊天,都还是对话。所以这是一种非常灵活的方式,与越来越好的智能体交互。在某个时候它可能不再那么好,你也在谈论添加其他交互方式,但有趣的是,聊天被证明是这一切之上的一个非常强大的层。
Speaking of medium, I was wondering if I should go in this direction or not, but real quick, Kevin Weil, the CPO of OpenAI. He did a panel at the OpenAI Friends Summit last year, and he made this really interesting point that chat is a really interesting interface for these tools because they're just getting smarter and smarter, and chat continues to work as a paradigm to interact with them similar to a human. You could talk to Albert Einstein, you could talk to someone not very smart, and it's all conversation still. So it's a really flexible way to interact with increasingly good intelligence. At some point it'll not be so great, and you're talking about all these ways that you're adding additional ways to interact, but it's interesting chat proved to be a really powerful layer on top of all this stuff.
是的,这真的很酷。我觉得聊天也有社交元素,这非常人性化。就像,你知道,你有时想加入群聊,而和 AI 对话本身就像一种群聊。我觉得这个想法是如何构建这样的功能?我把任务看作一种通用功能,随着模型发展新能力,它会很好地扩展。模型将能够进行更好的搜索,创作更具创意的写作,渲染 React 应用和 HTML 应用,你可以每天有一个新谜题,或者从前一天的故事继续。这扩展得很好。
Yeah, that's really cool. I feel like chat also has a social element, which is very human. It's like, you know, you sometimes want to get into a group chat, and having conversations with AI is kind of like a group chat in itself. I feel like this idea of how do you build features like this? I see tasks as a general kind of feature that will scale very nicely as the models develop new capabilities. The models will be able to do better searches, create more creative writing, render React apps and HTML apps, and you can have every day a new puzzle for you, or continue the story from the previous day. This scales very nicely.
你提到当我们进入这个额外部分时,我们最终深入探讨了智能体使用计算机的想法。我知道这实际上是你们今天要发布的东西,也就是我们录制这一天,等这期节目出来时它已经上线了,叫做 Operator。你能谈谈这个人们将能使用的非常酷的功能吗?
You mentioned something as we were getting into this extra section that we ended up going down is this idea of the agents using a computer. I know this is actually something you were going to launch today, the day we're recording it, which will be out by the time this comes out, called Operator. Can you talk about this very cool feature that people will have access to?
是的,很遗憾我没有参与太多,但我对这个发布感到非常非常兴奋。它基本上是一个智能体,可以在自己的虚拟计算机、自己的虚拟环境中完成任务。你可以做任何字面上的任务,比如“在亚马逊上给我订一本书”。理想情况下,模型要么会跟进问你“你想要哪本书?”,要么非常了解你,开始推荐,比如“这里有五本我可能推荐你买的书”。然后你点击“好的,帮我买”,模型就会进入自己的虚拟小浏览器,完成任务,在亚马逊上买书。如果你给模型提供凭证和信用卡,显然这伴随着很多信任和安全问题,那么它就会像虚拟助手一样为你完成这件事。
Yeah, so I unfortunately did not work on it enough, but I'm really really excited about this launch. It's basically an agent that can complete tasks in its own virtual computer, in its own virtual environment. You can do any literal task like "order me a book on Amazon." Ideally the model will either follow up with you like "which book do you want?" or know you so well that it will start recommending, like "here are some of the five books that I might recommend to you to buy." Then you hit "yeah, help me buy," and the model goes off into its own virtual little browser, completes the task, and buys the book on Amazon. If you give the model credentials and credit card, obviously it comes with a lot of trust and safety, then it will just complete the thing for you as a virtual assistant.
有趣的是,这听起来显然应该发生,就像为什么这还不是现实,这也很令人震惊,我们只是假设这应该存在,就像只要你说一声,AI 就能在电脑上为你做事。这很荒谬。实际上这非常难。
It's interesting how this just sounds like obviously this should happen, like why is this not a thing yet, which is also mind-blowing that we're just assuming this should exist, like just some AI doing things for you on a computer when you just ask it to do. It's absurd. It's actually really hard.
我认为,你知道,他们还在攻克这个难题。我觉得,不知道你听说过一个叫 Tuple 的工具没有,它是一个结对编程产品。不知道你是否喜欢结对编程。Shopify 在播客节目中也提到了同样的数字。是的,这是一个非常酷的产品,你可以随时给任何人打电话,共享屏幕,对方可以访问你的屏幕并开始直接操作你的电脑。它是非常实时的,延迟非常高质量。我也想要同样的东西:我想和我的模型结对编程,模型应该能够和我交谈,或者在 VS Code 中画出我代码中的特定部分,告诉我,教我,我们可以有不同的模式。就像这里,这就是一个现成的产品。我不知道。人们应该去构建它。听起来一个初创公司刚刚从听这个的人中诞生了。
I think, you know, they're still cracking this. I feel like, I don't know if you heard of this tool called Tuple, it's a pair programming product. I don't know if you love pair programming. Shopify is the same number came up on a podcast episode. Yeah, so it's a very cool product where you can just call anyone at any time, share screen, and the other person can have access to the screen and start literally operating your computer. It's very real-time, the latency is very high quality. And I kind of want the same: I want to pair program with my model, and the model should be able to talk to me or draw very specific sections in my code in VS Code, and tell me, teach me, and we can have different modes. It's like right here, this is a product right here for you. I don't know. People should build it. It sounds like a startup just got birthed from someone listening to this.
你提到让智能体像你一样控制电脑并帮忙非常困难。是什么让它这么难,你能简要解释一下吗?
You mentioned that it's very hard to do this agent controlling a computer as you and helping out. What makes it so hard for whatever however much you can explain briefly?
很大程度上是因为目前模型操作的是像素而不是语言或其他。像素对模型来说真的非常非常困难,因为视觉感知。我认为还有很多多模态研究在进行。但相比之下,语言扩展起来容易得多。我的团队正在研究的另一件事是如何非常正确地推导人类意图。就像有时候,模型是否有足够的信息来提出后续问题或完成任务?你肯定不希望智能体跑开 10 分钟,然后带回一个你根本不想要的答案。这实际上会造成更差的用户体验。而这涉及到教模型人际交往技巧。
Much of it is because right now the model operates on pixels instead of language or whatnot. Pixels are actually really really hard for the models because of visual perception. I think there's still a lot of multimodal research going on. But I think language scaled much easier compared to multimodal because of that. Another thing that my team is working on is how do you derive human intent very correctly. It's like sometimes, does the model know enough information to ask a follow-up question or to complete the task? You kind of don't want an agent to go off for 10 minutes and then come back with an answer that you didn't even want. That actually creates a much worse user experience. And this comes with teaching the model people skills.
就像,人们喜欢什么?我们某种程度上会建立用户的心理模型,关心用户,从而提出某些问题。这部分对模型来说很难。这和我们之前谈到的有关,这种软技能、人际交往能力,还不是这些模型的强项。
It's like, what do people like? We kind of create a mental model of the user and care about the user in order to ask certain questions. That part is hard for the models. And that relates to what we talked about earlier where this kind of soft skill, people skills piece is not where these models are strong yet.
好,我要跳过快问快答环节了。我只想问快问快答里的一个问题,一个有趣的问题。是的。那么,当 AI 取代你的工作时,Karina,我很好奇你会做什么。而且它会给你津贴,每月给你津贴。这是你当月的薪水。你想做什么?你想把时间花在什么上?在这个未来的世界里,你会做什么?
Okay. I'm going to skip the lightning round. I want to ask just one question from the lightning round. Something fun. Yes. So when AI replaces your job, Karina, I'm curious what you're going to do. And it gives you a stipend. Gives you a monthly stipend. Here's your salary for the month. What would you want to do? What do you want to spend your time on? What will you be doing in this future world?
我一直在想这个问题,所以我觉得我有很多职业选择。我想我会喜欢当作家,我觉得那会很酷。就写写短篇小说,比如科幻故事、小说。我真的很喜欢艺术史。你知道博物馆里的那些保护者,他们努力保存画作,就只是画作一点点。我觉得做那个会很酷。
I've been thinking about this all the time so I feel like I have a lot of job options. I would love to be a writer, I think. I think that would be super cool. Just write short stories, like sci-fi stories, novels. I really like art history. So you know those conservationists in museums who just try to preserve art paintings, but just paintings a little bit. I think that would be really cool to do.
听起来很美。我不知道。我听到的是,你需要削弱这些模型,让它们不擅长写作,这样你才能继续。不过到那时你不需要为了谋生而做,不需要人们购买。你只是出于乐趣而做。所以即使它们非常擅长写作或艺术保护,那也没关系。
That sounds beautiful. I don't know. What I'm hearing is you need to nerf these models to not get very good at writing so that you can continue. Although at that point you don't need to do it for like you don't need people to buy it. You're just doing it for fun. So it doesn't even matter if they're incredibly good at writing or art conservation.
哦,天哪,这期节目或这次对话真棒。我们生活在一个多么疯狂的时代。Karina,非常感谢你来做客。最后两个问题,如果人们想联系你或跟进任何事情,可以在网上哪里找到你?听众怎样才能帮到你?
Oh man, what an episode or a conversation. What a wild time we're living in. Karina, thank you so much for being here. Two final questions, where can folks find you online if they want to reach out and follow up on anything and how can listeners be useful to you?
你可以在 Twitter 上找到我,账号是 Karina Nguyen。你也可以通过我的网站给我发邮件。我的团队正在招聘,所以我在寻找研究工程师、研究科学家,以及机器学习工程师,比如来自产品工程背景、想学习模型训练的人。我实际上是在为我的团队招聘。我的团队叫做前沿产品研究,我们训练模型,开发新方法,但面向产品导向的成果。
You can find me on Twitter. It's Karina Nguyen. You can also shoot me an email on my website. And my team is hiring, so I'm looking for research engineers, research scientists, as well as machine learning engineers, like people who come from product engineering who want to learn model training. I'm actually hiring for my team. My team is called Frontier Product Research and we train models, we develop new methods but for product-oriented outcomes.
多么棒的工作场所。天哪。人们申请这些非常丰厚的职位的最佳方式是什么?
What a place to work. Holy moly. What's the best way for people to apply for these very lucrative roles?
我想你可以在 Twitter 上给我发私信。或者我还没写职位描述。这就是职位描述。或者你可以申请后训练团队。
I think you can shoot me a DM on Twitter. Or I'm yet to create a job description. This is the job description. Or you can apply into the post-training team.
好的。听着,你会收到大量私信的。希望你有准备。Karina,非常感谢你来做客。这太棒了。
Okay. Listen, you're going to get a flood of DMs. I hope you're prepared. Karina, thank you so much for being here. This was incredible.
非常感谢,Lenny。大家再见。
Thank you so much, Lenny. Bye, everyone.
非常感谢你的收听。如果你觉得本期内容有价值,可以在 Apple Podcasts、Spotify 或你最喜欢的播客应用上订阅本节目。也请考虑给我们评分或留下评论,这真的能帮助其他听众找到这个播客。你可以在 lennyspodcast.com 找到所有过往节目或了解更多关于本节目的信息。下期再见。
Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennyspodcast.com. See you in the next episode.