Fable 5:从第一印象到实际工作流

Fable 5: From First Impressions to Real Workflows

迈克·克里格 Mike Krieger · AI & I · 2026-06-10 · 约 52 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anthropic Labs 负责人、Instagram 联合创始人 Mike 分享了他使用 Fable 5 的演变体验,从最初的惊叹到如今能放心委托复杂任务过夜完成。

Mike, head of Anthropic Labs and co-founder of Instagram, shares his evolving experience with Fable 5, from initial awe to delegating complex tasks overnight.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 17)

全文 · Full transcript(中英对照)

0. 欢迎与Fable 5简介 Welcome and Introduction to Fable 5

Host

Mike,欢迎来到节目。

Mike, welcome to the show.

Mike Krieger

很高兴来到这里,Dan。见到你真好。

Great to be here, Dan. Good to see you.

Host

所以,对于不了解你的人,你是 Anthropic Labs 的负责人,也是 Instagram 的联合创始人。今天我想和你聊的是 Fable 5。Fable 5 明天发布,我们在发布前一天录制,这期会在发布后播出。但我真正想请你来节目,是想让你告诉我,在第一天之后使用这个模型是什么感觉。我觉得当一个如此强大的模型发布时,有一个每天都在使用它的人来告诉你它的强项在哪里、它真正改变了什么、它没有改变什么,这样你就不会陷入那种 AI 焦虑症。你可以真正思考,好吧,它如何融入我的生活。

So, for people who don't know you, you're the head of Anthropic Labs and you're the co-founder of Instagram. And today, what I want to talk to you about is Fable 5. So, Fable 5 is dropping tomorrow. We're recording this the day before. This will come out after it drops. But what I really wanted to do is bring you on the show to tell me about what it's like to use this model beyond the first day. I think when a model this powerful drops, it's so useful to have someone who's using it day in and day out to tell you this is where it's powerful. This is how what it actually changes. This is what it doesn't change so that you you're you kind of like don't you kind of don't get the same AI psychosis type thing. You can actually think about okay like this is how it fits into my life.

Mike Krieger

是的,完全同意。而且这也很有趣,你知道,在 Fable 发布之前,我们已经有一些这个神话级别的模型用了几个月。我认为看到外部的人如何用我们的模型构建东西非常令人兴奋。但你也说得对,第一天的印象实际上来自于使用几周后的体验。我们甚至在之前的模型上也看到了这一点,比如 12 月到 1 月使用 Opus 4.5 或 Opus 4.6 的经历非常重要,因为人们花了很多时间在模型上,然后发现,哦,其实我还没有足够用力,我需要更进一步,重新思考这一代模型到底能做什么。

Yeah, absolutely. And it's also just been interesting, you know, we've had some, you know, models in this, you know, mythos class leading up to the fable release, you know, for a couple of months now. And it's, I think it's very exciting to see how people will build with us externally. But I think you're also right that day one impressions, I think it really comes from getting to use this over a couple of weeks. I think we've seen that even with previous models like the December into January usage Opus 4 5 or Opus 46 was really important because people spend extended time on the model and then figured out oh actually I wasn't pushing hard enough I got to go further and I got to rethink what's even possible with this generation.

1. 调整工作流与初印象 Adapting Workflows and First Impressions

Host

完全同意。我是说,我觉得每个公司内部使用过它的人都会说,天哪,我觉得我需要一套新技能才能用这个模型。尤其是一些非技术背景、偏知识工作的内部人员,他们会说,我甚至不知道能用它来做什么。而那些编排智能体的人则会说,天哪,我觉得有太多新东西要学了。所以我很好奇,你第一次尝试它时的印象和现在有什么不同?

Totally. I mean I don't know I feel like there are people internally at every who have been using it who have been like oh my god I think I kind of need a new set of skills to use this model. And I think you can especially see this with people who are maybe more non-technical internally and who are more on the knowledge work side of things where they're like I don't even know what I would use this for. And the people who are orchestrating agents are like holy I feel like there's so many new things I need to learn. So I'm curious for you tell us about the difference between your impression when you first tried it and now.

Mike Krieger

是的,我认为你关于调整工作流程的观点非常好,字面意义上的工作流程。我稍后会谈到这一点,但也要考虑如何思考模型的使用。因为一开始,时机很有趣,它恰好与我从 CPO 转到 Labs、真正回到构建者模式的时间重合。大概在我转岗一个半月到两个月后,我们内部第一次有了这类模型。我坐在那里,感觉自己又成了完全的新手,因为我觉得我提示的方式,甚至思考如何分解任务的方式,对这个模型来说都过时了。它不再是那样,甚至思考时间跨度或交互模式也需要进化。早期可能是:我有一个功能想法,我们能从……开始吗?绝对不行,对吧?到后来,让我表达更多意图,然后我记得 Raj April 说,哇,一次性就令人难以置信地印象深刻,但它也理解我们如何演进的意图,并理解全局上下文。所以我认为这是一个非常有趣的演变,直到现在。有趣的是,今天早上我和某人聊天,我在想工作,我要坐飞机,我想,好吧,我可以远程完成大部分工作,甚至不担心 Wi-Fi 会断,因为我知道如果我设置了正确的上下文指令,比如 flash loop,它会看到并完成。过去两个月里,有很多次我会跟 Claude 道晚安,给它布置一个相当复杂的任务,然后醒来发现它通常在凌晨 2 点就完成了,然后接下来的四个小时它大概就在闲着。但它完成整个流程的能力令人印象深刻,它能自己摆脱困境,比如:好吧,Mike 让我整夜完成这个复杂任务,我卡住了,因为远程服务宕机了,我先写一个临时的脚手架后端。所以它会记录,会一路做下去。它有一个很好的心理模型,知道能走多远,等服务恢复后再修复,并跟踪这个事实。对我来说最令人印象深刻的是,你可以委托那种级别的任务,并相信最终会有正确的结果。当然,你还会审查结果,而且还有一整套验证流程,我们可以也应该谈谈,因为我认为这是完成整个流程的重要部分。但这真的迫使我重新思考,用这类模型提高生产力是什么样子?它更像是我们之前讨论过的,当这些模型更像是伙伴或同事时是什么感觉,而现在它真的像一个我可以委托大量工作的队友。

Yeah I think the your your point on on adapting workflows is a really good one. uh quite literally workflows. I'll talk about that in a second, but also just in terms of like how do I like think about um usage of the model because uh you know at first uh and the the timing was interesting because it kind of coincided with me transitioning from CPO into labs and going really back into into builder mode. I think it was about a month and a half or two months into that uh that we first had you know one of these models available internally and um I I sat there and I and I was like I feel like a total newbie again because I feel like the way that I am prompting or even thinking about decomposing a task is really out of date now with this model like uh uh it's no longer and it's even thinking about the time horizon or the sort of like interactivity model I think has to evolve as well like going from I think early on would be like I have an idea for this feature. Can we start by do like absolutely not right? Uh to great like let me express more of the intent and then just being you I remember like you you know you know Raj April be like wow on the one shot it's already incredibly impressive but then it also understands the intent around how we're going to evolve this and understands like the global context as well. So I think that's been a really interesting evolution till now where um you know I was funny I was talking to somebody this morning where you know I think about doing work I had a flight and I was like okay I can do most of this work remotely and I don't even worry that like the Wi-Fi is going to drop out because I know that if I set up the right you know context instructions like flash loop you know I'll see it it'll see it through. Um, and and I think my last two months have been full of a lot of times where I will, you know, wish Claude a good night, set it up on like a pretty complex task of something of this like model class and wake up to, you know, actually it's usually done by like 2:00 in the morning and I guess it just totals its thumbs for the next four hours. But like uh really impressive ability to like complete the swing, get itself out of the situation where it's like, okay, all right, well Mike asked me to do this complex task overnight. I got stuck because this remote service went down. I'm going to write a like scaffolded like backend for it for now. So, you know, I'll document that. I'll, you know, go all the way through. Uh, I have a like good mental model of like how far that's going to get me and when it comes back online, I'll fix it. I'll keep track of that fact. It's just like it it is I think the most impressive thing for me is like you just be able to like delegate that kind of level of task and just trust that the right thing will happen by the end. And of course, like you'll review the result and there's still like a whole verification thing that we we can and should talk about because I think it's an important part of still completing the the the swing there. Um, but it's really forced me to rethink like what does being productive with one of these models look like? And it is much more like we've talked for a while about, you know, like what is it like when these models are more of like a companion or a co-orker and it really feels like now it's like a teammate that I can delegate like a lot of work to.

2. 日常流程与模型选择 Day-to-Day Flow and Model Selection

Host

那你现在的日常工作流程是怎样的?因为我注意到,如果你给它一个大任务,然后对着它独白,让它运行几个小时或一整夜,它是我试过的最令人印象深刻的模型。但是,它太慢也太贵了,我觉得我不想用它来做日常任务。那么,你日常实际的使用流程是怎样的?它与其他模型相比,在什么场景下使用?

And what is your what is your day-to-day flow like right now? Because one of the things I noticed is if you if you just give it a big task and you monologue into it and you just like let it go for a few hours or overnight, it's like the most impressive model that I've ever tried. But, you know, it's so slow and it's so expensive that you you I I feel like I don't want to use it for day-to-day tasks. So, what is your actual flow like in terms of how you use it day-to-day and where does it slot in versus other models?

Mike Krieger

是的,我最终和它进行了更多的前期架构规划对话。所以这是另一个有趣的变化。我认为这是一个所有模型都需要继续改进的时代。我非常感激在 Instagram 的经历,从最初在洛杉矶服务器上拼凑起来的版本,到能够扩展它,最终与 Facebook 的所有基础设施集成。因为你会培养出一种感觉,知道每个阶段什么样的基础设施抽象和复杂性是合适的。

Yeah, I've ended up having a lot more um architectural planning conversations up front with it as well. So that's been like another interesting change where um I think this is an era that I think all models need to continue to improve. And I'm really grateful for the Instagram experience of having to like start, you know, from our initial version that was like duct taped on a server in LA to like being able to scale it and eventually integrate it with like all of like the Facebook infrastructure. Uh because you kind of develop a sense of what what infra abstractions and complexity are appropriate for each stage of it.

3. 用Fable规划与团队对齐 Planning with Fable and team alignment

Mike Krieger

我仍然会跟 Fable 来回拉扯,比如它会说“这个实现不错”,然后“我确实打算很快发布”,接着“我觉得我们应该考虑不止一台服务器”。这种来回很重要。但很多这类规划——我意识到 Fable 在思考上可以非常完整,就你与它规划的程度而言。通常只要说“你能不能做一个 HTML 页面,把我们刚才讨论的内容展示出来,这样我可以分享给团队”,就很有价值,甚至只是一个 Markdown 文档也行,但我喜欢有图表。所以这是一个有趣的用法:用它来规划,把思路理清楚,然后生成一份文档让团队对齐。因为我在实验室和 Anthropic 以外的团队都看到过这种动态:你可以非常快速地构建很多东西,而强制进行更多早期对齐——即使你先做一个原型,然后再把它退回到更成体系的架构中——也非常非常关键。它实际上成为了人与人之间的互动仍然很大程度上参与其中的地方。然后从那时起,无论是过夜还是在白天,让它执行那些任务块就非常重要。这意味着我需要比之前有更多的并发会话,因为我经常想“好吧,这里有三个工作块”。我在这两种方式之间摇摆:一种是喜欢一个非常长时间的 Claude Code 会话,真正让它以后台分叉子智能体的方式做所有事情,这样主线程保持响应;另一种是干脆接受这就是那种我要开五六个标签页来处理长篇综合工作的日子。但我确实认为这种长周期和“别担心,我在处理,可能需要一点时间”以及更多的来回互动有其价值。这种模式也是我们在产品中需要搞清楚的。我觉得你需要保留两者,它们以有趣的方式相互影响。我的偏好通常是,我总喜欢至少有一个高上下文但响应也非常非常快的 Claude,它的直觉很好:“我会回答你,如果需要我会启动一些东西,如果不需要我就先等着,等待下一个循环。”我确实认为你说得对,对于那种“我只是想修复这个交互问题”或者非常细节的问题,Fable 会深入思考那些事情。我认为 Fable 是我第一个因为那个原因而真正尝试不同 effort 级别的模型。我会说“好吧,我只需要调整一下 UI,把它设成中等看看效果”。我在 Opus 上没怎么这么做,因为那个范围感觉没那么宽,而 Fable 的范围确实可以很宽。

I still go back and forth with Fable where it'll be like 'this is a good implementation' and 'well I do plan on shipping this fairly soon' and 'I think we should probably think about more than one server'. That kind of back and forth is important. But a lot of that sort of planning—I've realized that Fable can be so complete in its thinking in terms of how much you plan with them. Often just saying 'can you just make an HTML page that represents what we just talked about so I can share it with the team' is actually valuable, or even just a markdown document, but I like having diagrams. So that's been an interesting use: let's plan with it, let's think it through, and then let's have some sort of document that we can align the team on. Because this is a dynamic I've seen in labs and just teams beyond Anthropic: you can build a lot very quickly, and forcing more of that early alignment—even if you do an initial prototype and then back it out into more of a planned architecture that works too—is really, really key. And it actually ends up being the place where the human-to-human interaction still stays very much part of the process. Then from then on, either overnight or during the day, having it execute on those chunks of tasks is really important. It just means having a lot more concurrent sessions than I did before, because I often think 'alright, there are these three pieces of work.' I go back and forth between liking having one very long-running Claude Code session and really asking it to do everything in background forked sub-agents so the main thread stays responsive, and other times just embracing that it's one of those days where I'm going to have five or six tabs tackling long comprehensive work. But I do think there's something to this long horizon and 'don't worry, I'm on it, it can take me a while' and more of this back and forth. That modality is something we'll have to figure out in our products as well. I think you want to preserve both, and they interact with each other in interesting ways. My preference is usually I always like having at least one Claude that is high context but also very, very fast response, and its instinct is great: 'I'm going to answer you and I'll kick something off if I need to, and if not I'm just going to hang tight and wait for the next kind of loop.' I do think you're right that for the 'I'm just trying to fix this interaction question' or something that's very fine detailed, Fable will go off and think very hard about those things. I think Fable is the first model where I've actually played more with the effort levels for that reason. I've been like 'okay, I just need to tweak some UI, put it to medium or something and see how that plays out.' I didn't find myself doing that as much with Opus because the range felt less wide, whereas it can feel quite wide with Fable.

Host

那如果是快速问题呢,比如你在路上?你会随手问 Fable 一些随机问题吗,因为感觉像是在用火箭筒打蚊子?还是你会来回切换?

What about a quick question like you're on the go? Are you asking Fable random questions as they come to you, because it feels like you're using a rocket launcher to kill a mosquito or something? Or are you flipping back and forth?

Mike Krieger

你问这个太有意思了,因为我之前——你知道,它一直在思考,非常努力地思考。然后上周我问了它一个我甚至不好意思问 Fable 的问题,大概是关于 NBA 总决赛的。我就想“好吧,我把 iOS 应用切换到 Sonnet”。然后我想“哦对,我一直用这个来问快速问题”。感觉差了一个数量级——实际上甚至不是每秒 token 数的问题,更可能是答案需要多少思考。有些答案不需要被彻底思考。所以是的,我自己也在思考,我认为这对我们来说也是一个很好的产品问题。一般来说,你不希望人们需要为这些选择想太多。所以理想情况下,长期来看我们可以围绕一些更可分类的、人们容易理解的用例达成共识,或者可能因界面而异。实际上,大多数时候我在 iOS 应用上并不是在做 Fable 级别的测试。每个界面固定一个模型选择可能是解决办法,我们需要从产品角度探索这意味着什么。但我确实有过“这个问题不值得用 Fable,我应该问 Sonnet”的感觉。

It's so funny you asked that because I had been—and you know, it's thinking, it's thinking really hard about it. Then last week I was asking it something that I felt embarrassed actually asking Fable about, something probably NBA finals related. I was like 'okay, I switched my iOS app to Sonnet.' I was like 'oh yeah, I use this all the time for fast questions.' It's order of magnitude feeling of—and it's actually not even the tokens per second, it's probably more around how much thinking goes into the answer. Some answers do not need to be fully thought through. So yeah, I am thinking myself through, and I think this is a good product question for us too. In general, you don't want people to have to be thinking so much about these choices. So ideally what we can coalesce around in the longer run is maybe some more bucketable use cases that are really grokkable to people, or maybe it varies by surface. It's actually probably unlikely that most of the time with the iOS app I'm doing Fable-type tests. Having a sticky model selection per surface might be the way to do that, and we'll have to explore what that means from a product perspective. But I for sure have had the feeling of 'this is not a Fable-worthy question, I should ask Sonnet.'

Host

你能给我们看看你用这个构建的东西吗?

Can you show us something that you've built with it?

Mike Krieger

这次我们做的一件事是鼓励我们使用个人账户,尤其是在周末,这非常有趣。你可以想象有很多 Anthropic 特定的工具等等。但退一步,纯粹用 Claude Code 在周末做点什么,感觉真的很好。

One of the things we did this go around is we encouraged personal account usage for us, especially on the weekends, which was really fun. You can imagine a lot of Anthropic-specific tooling, etc. But it was really good to step back and be like 'I'm just pure Claude Code, let's work on something over the weekend.'

Host

你是在终端应用里还是在桌面应用里?

And you're in the terminal app or you're in the desktop app?

Mike Krieger

好问题。我主要还是用终端应用。有意思的是,看我妻子——她不是专业工程师,更像 UX 设计师/PM——通过桌面应用真正爱上了 Claude Code。我觉得这以某种方式为她简化了一些抽象概念。但这次我还是用 Ghostty 和终端应用。我给你看看。这是那种每个人都有一些定制需求的东西。我想要一个好的媒体追踪体验。我想“我在玩游戏,在看电视剧,收到各种推荐,我只想构建一个对我个人有用的东西,适合我的一些使用场景。”我最初的两个最大标准是:第一,非常容易添加内容,所以你可以跟 Claude 对话,Claude 会对所有内容进行语义搜索,然后把正确的东西放进去;第二,主动地,如果有新一季或者游戏的新续作,它可以去研究那些东西。

That's a great question. I'm mostly still in the terminal app. It's interesting watching my wife, who's not a professional engineer and more of a UX designer/PM, really fall in love with Claude Code via the desktop app. I think it simplified some of the abstractions for her in that way. But for this one I was still using Ghostty and the terminal app. Let me show you. This is one of those things where everybody has some bespoke need. I wanted a good media tracker experience. I was like 'I'm playing games, I'm watching TV shows, I get all these recommendations, and I just wanted to build something personal to me that fit some of the use cases I had.' The two biggest criteria I started with were: one, really easy to add things, so you can talk to Claude, Claude does the semantic search over everything and puts the right things in; and two, proactively, if there's a new season or a new sequel to a game, it could go off and research those things.

4. 用Claude构建自修改应用 Building a self-modifying app with Claude

Mike Krieger

大部分 UI 都像 Fable OneShot 那样,已经很令人印象深刻了。但今年我在实验室里一直在琢磨的一个线索是,如何让软件团队(现在都在云端)更贴近软件本身。所以这大概是周六早上。我整个周末都在忙孩子的事。所以很多工作都是先启动,然后去跟孩子远足,回来继续做,有时在路上也看看工作进展。我不该这样,但远程模式看看进展也挺好的。尽量别做太多。但我有这个想法:能不能做一个快速实验,看看如果能在软件内部直接修改软件会怎样?我同时做了一个 React Native 版本和一个 Web 版本。我已经有一个聊天功能,可以要求 Claude 通过 URL 添加东西。我希望每个软件都有这个功能,这样我再也不用导航菜单了。这在很多方面,Dan,就像我在尝试把智能体原生架构发挥到极致,也就是让智能体也能修改应用。智能体架构的第一阶段:产品中的每一个东西都能通过智能体访问,并有工具调用等。这有望成为标配,可惜很多软件还没有。这很棒,因为有人推荐了一个关于放射性物质的巴西节目,我不记得叫什么了,Claude 帮我找到了。比我自己凭直觉去猜好多了。但我感兴趣的下一步是,真正在移动中修改软件本身意味着什么。所以如果你长按这个小聊天框,我构建的——实际上是 Claude 构建的——是一种方式,它利用我们的管理智能体来接受编辑请求,然后你可以预览。我用了 Vercel 的实时预览功能。这个功能也是一次性完成的,非常酷。我后来不断添加功能。它实际上会显示一个差异视图,如果你想看的话。你可以进入管理智能体的对话,看看它做了什么,虽然我几乎从不这么做,因为我不太关心代码质量或长期可维护性。你可以看到它在这里也有一个会话。但真的很有趣。我会在路上使用它,比如前几天我有一个功能请求:原生 iOS 上的浮动操作按钮太低了,但在那里还可以,你能修一下吗?用 Expo 工具链真的很棒,而且能在手机上实时重载,感觉也很酷。但问题是,这个东西需要达到生产级别,服务百万用户吗?不需要。但拥有一个不必止步于周末、可以持续通过使用来改进的东西,感觉真的很好,而且有这种端到端的闭环。所以我觉得这很好地体现了 Fable 的构建能力,也体现了我们俩一直在思考的:Claude 如何嵌入软件,超越单纯的使用层面。

Most of the UI was like Fable OneShot, which was already impressive. But the thread I've been pulling a lot in labs this year is how do you bring the software team, which is cloud these days, closer to the software itself. So this was maybe Saturday morning. I had a full weekend with kids stuff. So a lot of this was kick off work, go for a hike with the kids, come back, continue to do the work, sometimes check in on the work on the hike. I probably shouldn't, but it was nice to pop into remote mode and see what was going on. Try not to do that too much. But I had this idea: could we do a spike on what if you could actually modify the software from within itself? I built both a React Native version and this web version. I already had a chat thing where you could ask Claude to add things by URL. I want every software to have this where I should never have to navigate a menu to do anything ever again. This is in many ways, Dan, like I was trying to distill the agent-native architectures to its fullest degree, which is also have the agent be able to modify the app. Phase one of agent architecture: every single thing in this product is accessible from the agent and has tool calls, etc. That's hopefully becoming table stakes, sadly not in a lot of software. It's great because somebody had recommended a Brazilian show about radioactive stuff, and I did not remember what it was called, and Claude was able to figure it out. So much better than trying to figure that out intuitively. But the next step I was interested in is what would it mean to actually be able to modify the software from itself on the go. So if you long press this little chat thing, what I built — what Claude built — was a way where it uses our manage agents to take on edit requests, and then you can preview them. I used the Vercel live preview thing here. This whole feature was also one-shot, which was really cool. I just added to it over time. It actually does a little diff view if you wanted to. You can go into the manage agent conversation and see what it did, although I almost never do because I don't particularly care about the code quality or long-term maintainability of this software. You can see that it had a session in here too. But it's been really fun. I'll be using it on the go and say, I had a feature request the other day: the floating action button was too low on native iOS but it was okay on there, can you fix it? It was really fun with some of the Expo tooling now and actually live reloaded on my phone, which was also a really cool feeling. But it was just like, does this thing need to be a production-level thing that's going to go to a million users? No. But it felt really good to have something where I felt like it didn't have to stop at just the weekend and I could keep working on it just by using it, and having this end-to-end close thing. So I felt like this was a good manifestation of both Fable's building ability, but also a lot of what both you and I have been thinking about: how does Claude embed itself into software beyond just the usage side of things.

Host

这真的很酷,我想让大家明白:这个东西已经构建出来了——你可以构建类似的东西,也许不是自我修改的部分,但你可以花 10 年或 20 年构建类似的东西,但构建成本已经大幅降低了。所以想想在 Instagram 时代做这个要花多少钱,现在呢?你能帮我们理解这种变化吗?

This is really cool and I want people to understand: so this has been built — you could build something like this, maybe not the self-modifying part, but you could build something like this for 10 years or 20 years or something like that, but the cost to build has gotten dramatically lower. So think about how much it would have cost to do this in the Instagram days versus now. Can you help us understand how that has changed?

Mike Krieger

是的,当我回想那个时代时,我也经常思考这一点。在 Instagram 早期,我认为自己是一个非常高效的程序员。我非常热衷于移动开发,而且我们对事情有清晰的认知。我认为从想法到完整产品的实现,仍然需要大约 4 天的通宵工作,那是我自然的状态——熬到凌晨 4 点,睡到中午,这不适合家庭生活,所以我不得不改变。但那是我构建的方式。Instagram v1 可能比这个东西功能更多,但不是一个数量级,大概花了五天通宵,我做前端和后端,Kevin 做最初的滤镜,才把它做出来。而且这也是建立在我多年 iOS 开发经验之上的。然后是迭代:我经常想,发布后一切顺利时,我们被什么限制了?我们有各种想法,但只能勉强维持网站运行或添加一个增量功能。标签花了一周才建好,但之后还有无数事情想做。所以我认为这既是时间的缩短——想法、概念和迭代仍然需要时间——另一方面是你可以在已有基础上迭代的方式。我觉得这是一种非常有趣但又有点飘忽的方式。

Yeah, I think about this a lot when I think back to that time as well. I thought of myself as a very productive programmer in the early Instagram days. I was really into mobile development and we had good clarity of things. I think the gap from idea to fully realized version of some complete product was still looking at about 4 days of my all-nighters, which was my natural state — up till 4, sleep until noon, which is not conducive to family life so I've had to shift. But that was my building thing. Instagram v1, which probably had more features than this thing did but not by an order of magnitude, was like five days of all-nighters me working on the front end and back end and Kevin working on the initial filters to get that out. And this was also built on many years that I've been working on iOS pieces as well. And then the iteration: I think a lot about what we were gated on after that launch when things went well was we had all these ideas for where to take it, but we were just trying to keep the site up or add the one incremental feature. Hashtags take a week to build, but then there's all the things you want to continue doing on it as well. So I think it's both that shortening of time — there's still the time required for the idea and the concept and the iteration — and then the other piece is the way you can then iterate on what you have. I think a really fun but also very in-the-float kind of way.

5. 缩小意图与执行差距 Closing the gap between intent and execution

Mike Krieger

然后你知道,现在作为专业软件工程师和创业者的我,如果你有那个想法,我看到很多人经历这个过程,比如“好吧,我试着找个咨询公司来接手”,但这个过程损耗很大,并不是我想要的。你知道,不要为此融资。我认为这些模型最令人兴奋的地方不仅仅是它们变得更自主,而是再次缩小了意图与执行之间的差距,我看到了这对那些并非“建造者”的人的建造能力的影响。这些模型的轨迹一直是,你知道,某种神话级别的模型,最终更便宜、更易用的模型也会出现。随着这个过程发生,我认为它正在打开无数可能性。前几天我收到一条消息——如果你看不出来,我对这些东西非常兴奋——来自内部某个人,我们为她建了一个内部工具,结合了 Fable 和一些内部 MCP 的访问。她说:“这是我人生中第一次”——她做招聘工作——“人生中第一次感觉我脑子里的东西和世界上存在的东西现在紧挨在一起,我直接就能做出来。”这对她来说意义重大,因为在那之前,我记得那些日子——五年前或四年前——如果那个人想要一个工具,要么凑合着用,要么找内部工具工程师,而那个工程师可能已经 overloaded 了 50 个需求。但现在,她们却在尽情享受建造的乐趣。我认为这带来了很多希望,因为我不认为人类的创造力和可能性是有限的。我认为在我们最好的状态下,我们本质上是在扩大能够将想法变成现实的人数。

And then you know, if now this is me as a sort of professional software engineer, sort of startup founder beyond that, if you had that idea, you know, and I saw multiple people go through this like, 'Well, I'll try to find maybe a consultancy that will take this on,' but like, now there's like it's a really lossy process of like what I wanted. You know, don't raise money for it. And I think that the thing that I think is like the most exciting part about these models getting not just more autonomous but again closing that gap between intent and execution is what I've seen it do to people's ability to build who are not like builders. And the trajectory of these models has been, you know, something able, you know, of this general mythos class is like in that class of models, and eventually, you know, models of you know that are cheaper and more accessible to other folks become available too. And like as that process happens, like I just think it is just opening up so many. Like I got a ping the other day—I get very excited about the stuff, if you can't tell—from somebody internally, and we had built them an internal tool that kind of combined Fable and like access to some internal MCPs. And she said, 'It is the first time in my life'—and she works in recruiting—'the first time in life where like I feel like the thing that's in my head and the thing that exists in the world is now like they're right next to each other, like I can just do it.' And it was like very like a meaningful moment to her because prior to that, like I remember these days—these days were 5 years ago or 4 years ago—where that person, if they wanted a tool, would have to either make do or try to get an internal tools engineer that probably was overloaded with 50 other requirements. But instead now they like are just having the time of their lives building. And I think that is cause for a lot of like hope, because I don't think that human capacity for creativity and what's possible is enormous. And I think like at our best we are basically expanding the number of people who can then see that through to something that feels real.

Host

我完全同意。但我确实觉得我脑子里有个问题,可能也是听众心里会有的。所以我想问你:鉴于你刚才说的所有话,软件工程是不是结束了?

I totally agree. But I do think that there's a question in the back of my mind and I think it's probably going to be in the back of the minds of some of people listening. So I want to ask you: given everything you just said, is software engineering over?

Mike Krieger

是的,我认为软件工程已经不同了。它发生了巨大变化。如果在我做 Instagram 那会儿你问我“什么是软件工程”,我可能会说:“嗯,思考难题,思考架构,然后花大量时间在 TextMate 里”——不知道怎么就提到这个——但你知道,文本编辑器,编辑那些东西,或者 Xcode,然后

Yeah, I think software engineering is different. It is like dramatically changed. And as I probably would have defined it if you had asked me around the Instagram time like, 'What is software engineering?' I'd probably say, 'All right, like thinking through the hard problems and like thinking about an architecture and then like spending a lot of time in, you know, like TextMate'—I don't know where that came from—but like, you know, text editor, you're going to edit those things, or Xcode, you know, and

Host

看 Rails,你知道。

watching Rails, you know.

Mike Krieger

完全正确,完全正确。还有理解 Django 那层的复杂性,以及部署后修复 bug。这些大部分都发生了根本性变化,并融入了产品管理的其他部分。我认为那种 PM 和工程的分工——我在我们团队里也看到了——已经变得更加模糊。这已经彻底改变了。但整体来看,也许从软件工程退一步,思考软件生产或软件开发,但不仅仅是纯开发者的角度,我认为它依然生机勃勃且至关重要。所以我认为我们正处于那个时刻。我认为 Fable 是朝着这个方向的又一步——我不会说它是最后一步,当然还有很多会发生——但我觉得在信任方面,至少我最终对模型完成事情甚至合理架构的能力的信任度相当高。所以那部分感觉永远不会结束,但已经相当成熟了,对吧?它已经走得很远了。但我认为整体的技艺——你需要什么,你产出什么,它真的好吗?——我认为仍然是非常人性化的事业。但我也能看出这个转变并非毫无痛苦。我认为有很多人热爱实际构建的技艺——我以前也喜欢那种“我优雅地解决了那个问题”的感觉,你会梦到代码。如果你有过那种经历,梦到你在做的东西,早上醒来发现“我想到了如何优雅地解决这个问题”——那肯定已经过去了。我认为我交谈过的一些优秀工程师既有失落感,也有“天哪,但我现在能同时做海量工作”的感觉。所以,我们脑子里同时装着这两种想法。

Exactly right, exactly. And understanding the intricacies of Django's like layer, and then like fixing bugs after you deploy it. Like so much of that is radically different and collapsing into other parts of like product management. And I think that sort of like PM–eng split, I think you guys—I see it even in our teams—has become much more diffuse. That's radically changed. But I think the overall, like maybe zoom out from software engineering and think about like software production or software development, but not in like just a pure developer case, I think that is like alive and well and essential still. So I think that it is the moment that I feel like we are. And I think Fable is another step in the direction of—and I'm not going to call it the final step, of course a lot will still happen—but like I think a pretty significant step in terms of like the trust at least I end up placing the model in terms of its capacity to see things through and even architect things reasonably is quite high. So that part feels like it is not ever going to be done but it is pretty done, right? Like it's gone really far. But I think that the overall sort of craft of the what needs you have—like what are you putting out, is it actually good?—I think still a very human endeavor. But I also sort of can see that that is not a transition that is pain-free in a way. Like I think there are plenty of people who love the craft of actually putting—and I used to love stuff like, 'I solved that problem so elegantly,' you dream about code. And if you ever had that experience of like you dream about the thing that you're working on, like wake up in the morning like, 'I figured out how to solve this thing really elegantly'—and that for sure has passed. And I think that there is a feeling of loss, I think, in some of the like better engineers that I talk to, as well as the feeling of, 'Oh my god, but I can do insane amounts of work now at the same time.' So, we're holding both ideas in our heads at once, I guess.

Host

我认为这是最重要的一点。为那种事感到悲伤和兴奋都是正常的,但我很好奇。我们就以“软件工程依然生机勃勃”这个论点为例。在 Anthropic 内部,这实际上是什么样子的?

Which I think is the most important part of this. Like, it's normal to feel sadness for that kind of thing and excitement, but I'm curious. Let's just take the thesis of software engineering is alive and well. What does that actually look like inside of Anthropic?

Mike Krieger

是的,我认为有几个方面。我认为仍然有构建的技艺——好吧,我得从完整的软件开发周期或者我日常所见的角度来说。也许两者都谈一点。但我认为仍然有很多,你知道,我们聚在一起,讨论下一步如何演进 co-work,然后我们把它分解成各个所有权领域。我认为这仍然非常重要,因为作为人,你仍然拥有超越 Claude 的上下文,对吧?比如这个产品的实际意图是什么?进展如何?我们需要了解 pipeline 中其他即将以有趣方式集成的产品的哪些信息?所以我认为这方面仍然非常重要。所以,你知道,虽然我们每个人对应多个 Claude,但每个人——至少在我们处理主题的方式上——仍然有,我们称之为 DRI,即直接负责人——仍然对产品的某个部分或某个领域拥有 DRI 权。

Yeah, I think there's a few pieces. I think there's still the crafting of—well, I got to take it off from like the full software development cycle or like maybe what I see on a day-to-day. Maybe I'll do a little bit of both. But I think there's still a lot of, you know, we all got together, we talked about the next way we want to evolve co-work, and now we've kind of broken it down into areas of ownership. I think that ends up still being quite important because there is still context that you hold as a person that is sort of beyond Claude, right? Like what is the actual intent of this product? How's it going? What do we need to know about the other products that are coming down the pipeline that are going to be integrated in some interesting way? So I think that aspect is really important still. And so, you know, though we have many Claudes to each human—each human, at least the way we've been working on topics, still kind of has, you know, we call them DRIs, like directly responsible individuals—still has like a DRI-ship over some part of the product or some area.

6. 工作元维护与生产理解 Meta-maintenance of work and production understanding

Mike Krieger

我认为这种情况会持续一段时间,因为我觉得价值不仅仅在于这种分发器——比如我们都应该让协作变得更好——而在于思考协作在特定任务上的表现。我们尽量少开会,但会议还是会出现,你仍然需要这类对齐对话。然后很多那种异步委派,我觉得这里的许多工程师现在发现,他们都构建了某种版本:“好了,我现在要创建一个仪表盘,看看我的所有云在做什么,有什么在等我,哪些拉取请求需要我关注,因为你知道,要么是人类要么是云代码审查员回复了我。”所以有很多那种工作的元维护,我认为我们会部分标准化,但有些部分总会因人而异,就像人们整理自己的窗口和工作方式一样。然后还有理解生产环境中事物如何运作。我认为这是模型的另一个前沿。Fable 在这方面取得了显著进展,但还需要更多工作:理解代码部署后发生了什么。因为会有事故,你知道,本来一切正常,但某个网络链路断了,这不是你常见的故障模式,然后问题就显现了。Instagram 在 2012 到 2016 年间很大一部分工作就是处理这些并扩展规模。所以工程师的角色仍然非常关键。在事故响应中积累经验,理解如何保持冷静、收集数据、立即修复,然后再去处理长期修复,这仍然是必要的一部分。我在想是否还有其他值得注意的部分。最后我想说的是,我非常喜欢现在工程原型所扮演的角色。你必须清楚什么时候是原型,什么时候不是。老话说“代码赢得争论”,我从来不喜欢这句话,因为能写代码的人可以去做,但凭什么他们就应该默认赢得争论呢?但现在很酷的是,当我们对产品方向有分歧或争论时,产品经理常常会说:“好了,我刚刚试了一下,它在这些方面很粗糙,但你看它实际上展示了这个方案如何可行。”这可以开启一些有趣的对话。所以几乎所有这一切都和六个月前大不相同,尤其是在并行性和对这类高阶工作抽象的需求方面。但我认为没有改变的是所有权。

I think that'll be the case for a while because I think there is value in not just this distributor like we should all make co-work better but instead like all right I'm thinking through how co-work does at this particular task and there's still a lot of you know we try to keep meetings minimal but they still emerge and you still have these kind of alignment conversations. Then a lot of that sort of asynchronous delegation I think what many engineers here have now found is they've all built some version of 'All right, I'm going to now create a dashboard of where all my clouds are doing and what's waiting for me and which pull requests need my attention because you know either a human or a Claude Code reviewer got back to me.' So there is a lot of that sort of meta maintenance of the work that I think we'll standardize some but I think some of it will always be a little bit bespoke to the way each individual likes to work, just in the way that people organize their windows and their work. And then there is also the understanding how things work in production. I think that is another next frontier for the models. Fable has made significant strides in that but there's more work needed here: understanding what happens to code after it gets deployed. Because there's incidents, you know, this was all working well but this network link got cut which is not in your usual failure mode and it manifested. So much of Instagram from 2012 to 2016 was dealing with that and scaling things up. So that role of the engineer still remains really key. Getting reps in around incident response and understanding how to stay calm, gather data, remediate what's immediate, but then go off and work on longer-term fixes is still a necessary part of it. I'm trying to think if there's any other pieces that are notable as well. I think what's maybe the last thing to say is I really like the role that the engineering prototype now plays. You have to be clear when it's a prototype versus not. The old phrase was 'code wins arguments' and I never liked that because the person that could code could go do it but actually why should they necessarily win an argument by default? But it's been really cool now where we will have some disagreement or debate about where to take a product and often the PM will say 'All right, I just tried it and it's janky in these eight ways but look it actually shows how this could work' and that can open up some interesting pieces of conversation. So almost all of that is quite different than it was 6 months ago, especially at the level of parallelism and the need for these higher order abstractions of work. But I think what hasn't changed is ownership.

7. 赞助商插播:BrainTrust Sponsor break: BrainTrust

Host

我们很多人都在将 AI 投入生产,这提高了生产力,但也带来了焦虑。你调整提示词、切换模型、调整参数,测试时一切正常。于是你合并代码,但三天后甚至更早,支持工单就开始涌入。AI 给客户提供了意想不到的答案,你不知道问题何时发生或为何发生。BrainTrust 是解决这个问题的 AI 可观测性平台。它将评估和可观测性连接在一个工作流中。这样你就能看到生产环境中实际发生了什么,并衡量更改是让事情变好还是变坏。追踪显示完整的执行路径。评估定义什么是好的表现,实验让你在发布前并排比较提示词和模型。生产追踪直接反馈到你的评估数据集中。每次失败都成为一个测试用例。你在 CI 中捕获回归,防止它们到达用户手中。Notion、Stripe、Zapier、Vercell 和 Ramp 的团队都在用它大规模交付高质量的 AI。BrainTrust 专为构建生产 AI 系统的团队设计,在这些系统中,静默回归代价高昂。它适用于任何技术栈,提供 Python、TypeScript、Go、Ruby、C 的 SDK。没有框架锁定或供应商依赖。它通过了 SOC 2 Type 2 认证,并符合 GDPR 和 HIPAA 标准。请访问 braintrust.dev 开始使用。网址是 braintrust.dev。现在回到节目。

Lots of us are shipping AI to production which is great for productivity but it also comes with anxiety. You tweak a prompt, swap models, adjust parameters and everything looks fine in testing. So you merge and then 3 days later or even sooner the support tickets start rolling in. The AI is giving your customers unexpected answers and you have no idea when it happened or why. Brain trust is the AI observability platform that fixes this. It connects eval and observability in one workflow. That way you see what actually happened in production and can measure whether changes made things better or worse. Traces show the full execution path. Evals define what good looks like and experiments let you compare prompts and models side by side before shipping. Production traces feed directly into your eval data sets. Every failure becomes a test case. You catch regressions in CI before they reach users. And teams at Notion, Stripe, Zapier, Vercel, and Ramp use it to ship quality AI at scale. Brain trust is designed for teams building production AI systems where silent regressions are expensive. It's built for any stack. They have SDKs for Python, TypeScript, Go, Ruby, C. There's no framework lockin or vendor dependencies. It's sock 2 type 2 certified and GDPR and HIPACO compliant. Get started at braintrust.dev. That's brainustrust.dev. And now back to the episode.

8. 成本顾虑与使用激励 Cost concerns and usage incentives

Host

Fable 也非常昂贵。正因为如此,我在测试它的时候,感觉自己像个在糖果店里的孩子,一个劲地说“我要做这个,我要做那个”。但现在要付账单了,我就会开始考虑。因为我得在做之前停下来想想:这会花掉我 100 美元还是多少?我确实认为这会限制谁使用它以及用于什么目的。那么,你怎么看这个问题?

Fable is also very expensive. And because of that, like when I was testing it, I felt kind of like I was a kid in a candy shop and I was just like, I'll do this and I'll do this and I'll do that. But now that there's going to be a bill, I'm going to be thinking about it. Because I have to pause before I do it to be like, is this going to cost me 100 bucks or whatever? And I do think that's going to limit who gets to use it and for what. So, how do you think about that?

Mike Krieger

是的,我认为这在专业软件领域最清晰,就是经典的公司做工作。这会非常有趣。定价也涉及很多流程思考。它既比 Opus 贵,又在很多方面很便宜,如果你想想它做了多少不可思议的工作。当然,每个人都有自己的经济考量。所以无论如何,我认为对大多数软件团队来说最清晰。我认为作为一个行业,第一阶段是公司甚至难以让员工采用 AI 编码,因为模型早期,工具也不成熟。第二阶段是“太好了,我们创建排行榜,看谁用得最多”,你可以想象这会产生一些不太理想的激励。第三阶段是人们说“好了,现在我们只是试图找出谁在有效使用,并让他们尽可能多地花费,有明确的流程,但确保我们不做浪费的事情。”这对我来说总体上合理,尽管我认为也可能过度调整。我认为像 Fable 这样的模型应该能很好地融入其中:如果你展示了成果并从中获益,那么即使在公司内部也会有一个飞轮让它持续下去。在个人使用方面,这是一个很好的问题。我在个人测试中看到,因为我们的个人账户要付费——这很有趣,付钱给我工作的公司——但你确实会变得更谨慎。

Yeah, I think it's most clear-cut on the sort of professional software, you know, classic company doing work. It'll be really interesting. A lot of process thought goes into pricing as well. It's both more expensive than Opus and also in many ways it's really cheap if you think about how much incredible work it's doing. But of course everybody has their own economics around what they're working with. So anyway, most clear-cut I think from most software teams. And I think as an industry, phase one was companies even struggling to get some of their employees to adopt AI coding because models were early and tooling wasn't there. Phase two was 'great, we'll create leaderboards and see who can use the most,' which as you can imagine creates some not ideal incentives. Phase three was people saying 'okay now we're just trying to figure out who's using it effectively and letting them spend as much as possible with a clear process for that, but making sure we're not doing things wastefully.' That to me in general makes sense, although I think you could also over rotate that way too. I think something of Fable class should hopefully fit in well into that where if you're demonstrating results and getting use out of the model, then there's a flywheel even inside companies that perpetuates that. I think on the personal use side, it's a really good question. And I think where I've seen it, even in my personal testing because our personal accounts pay, which is funny paying my own company I work at, but you do become more thoughtful about it.

9. 构建个人应用与成本考量 Building a personal app and cost considerations

Mike Krieger

有趣的是,我周末做的那个应用其实只用了额外一点用量就搞定了。所以并不是花几千美元来构建这个属于我个人的东西,而且时间上也拉得更开一些。可能中间那类——也就是那些不在大公司里、但对定价也很在意的独立开发者或爱好者——是我们最需要思考的。我总的建议就是:先试试看,看看它能在你不需要大量后续跟进的情况下做多少事。我觉得现在衡量成本已经变得非常复杂了,因为既有每次交互的成本,也有你为了完成任务并达到满意程度所付出的成本。而在这方面,Fable 对我来说真的表现很出色:它直接就把事情做对了。这样我就不用再花接下来的 9 到 10 次交互说“不,那不是我想要的。你能再做这部分吗?”它让我印象非常深刻,因为你让它去做一件事,它就直接完成了,你会想:“哇,你把这个东西的所有小细节都考虑到了,我从来没见过其他模型能做到这样。”

Something that was interesting was the app that I built over the weekend actually fit in with only a bit of extra usage. So it wasn't like thousands of dollars to build this thing that is personal to myself, but it was also spaced out a little bit more. Probably the in-between of that, what we'll have to do the most thinking about, is the sort of hobbyist or independent who is not within a larger company, but also is thoughtful about pricing as well. I think my overall advice is just give it a try and see how much it can do without you having to do a lot of follow-ups. And I think measuring cost has gotten so multifaceted now because there's the per-turn cost and then there's what did it cost you not just to do the task but to complete the task to your satisfaction. And I think that's where Fable has really shined for me: it actually just does it right. So then I don't have to go spend the next 9 or 10 subsequent turns being like, "No, that was not quite what I meant. Can you also do this piece?" It's been really impressive for me because you ask it to go do something and then it just does it, and you're like, "Wow, you thought through all the little details of this thing in a way that I've never seen another model do."

Host

我不知道你能透露多少训练过程的信息,但这个模型到底有什么不同?

I don't know how much you can reveal about the training process, but what makes the model different?

Mike Krieger

我认为这在很多方面是团队大量工作的延续,我对我们的预训练和强化学习团队充满敬畏。我觉得进化最明显的部分——至少我注意到的——是与之相关的:对系统的整体感知,而不仅仅是单个工作片段。比如,我经常会被它惊喜到,当它写了一些东西后会说:“好吧,但你知道在生产环境中这需要不同,”然后它会一直提醒你:“你打开那个功能开关了吗?不打开它就不会工作。”我有时会参与持续数天的会话,它会说:“你看,你还没做那件事。你最好——”然后我就想:“你说得对。我没打开那个功能。我该去做了。”或者,如果我们改变这个,那边的合约也会变。或者看它——实际上,我最喜欢看它行动的一次,我认为它展示了训练成果,是看它如何回应来自人类或其他 Claude 审查者的代码审查反馈。它不只是说:“哦,对,这是个问题。我去修一下,”而是非常深思熟虑地说:“嘿,对于我们构建的这个保真度级别,我接受这个风险,”或者“是的,我明白你的意思,其他审查者——通常只是另一个 Fable 模型在和自己对话——我明白你的意思,但我实际上要反驳。我不认为那是对的。”我认为让模型拥有这种判断力非常重要。如果非要指出一个我觉得它进步很大的领域,那就是不再只是立即条件反射地说“对对对,我马上去修”,而是“嗯,我想想。不,我想过了,我还是不同意。”我认为这是一种非常有用的能力。拥有像 Claude Code 这样的产品非常有价值,因为你现在有了一个活生生的东西,人们会说:“这是模型做得好的地方,”我们有测试它的人——我把反馈团队放在非常高的位置,我们非常信任他们的反馈,因为它经过了反复的、多日的、困难任务的考验。这也极大地影响了我们如何思考下一步需要改进的地方:我们需要具体思考模型在哪些任务上做得更好。

I mean, I think in many ways it's a continuation of a lot of the work that the team has done, and I bow down in total awe of our teams both on the pre-training and on the RL side. I think the piece that has evolved the most, at least that I noticed, is kind of adjacent to that as well: a sense of the system more than just the individual piece of work. Like I will often be very positively surprised when it will write something and say, "All right, but you know that in production this needs to be different," and then it will keep bugging you like, "Have you turned on that feature flag yet? It's not going to work until you do." And I'll sometimes be in sessions that have gone on for days and be like, "Look, you still haven't done that thing. You better—" and I was like, "You're right. I didn't turn on that feature. I should go off and do that." Or, if we change this, the contract will change over there. Or watching it—actually, one of my favorite times of seeing it in action, I think, where it demonstrates some of the training, is watching it respond to code review feedback either from people or from other Claude reviewers. Where it doesn't just say, "Oh, yeah, that's an issue. I'm going to go fix it," but actually be really thoughtful around, "Hey, for this level of fidelity of what we're building, I'm going to accept this risk," or "Yeah, I see what you mean, other code reviewer—which is often just another Fable model talking to itself—I see what you mean, but I'm actually going to push back. I don't think that's right." I think getting the model to have that judgment is really important. And I think if I had to pinpoint an area where I feel like it's really progressed, it is that sort of not just immediate knee-jerk "Yeah, yeah, that's right, I got to go fix it," and more "Huh, I'll think about that for a minute. No, I thought about it and I still disagree." And I think that's a very useful ability. It's so valuable to have products like Claude Code out there because you have now a living, breathing thing where people are like, "This is where the model is doing well," and we have people who test it—I count the feedback folks as very, very high on the list where we really trust the feedback because it is being put through paces in repeated multi-day hard tasks. And that also very much feeds into how we think about what we need to improve on the next slide: what are the tasks that we need to specifically think about the model being better at.

10. 聊天是合适的界面吗? Is chat the right interface?

Host

聊天是这个模型的正确界面吗?因为它不是那种逐轮交互,更像是“我把事情委托给你”。那么,这如何改变你应该使用它的方式,或者你如何思考界面?

Is chat the right interface for this model? Because it's not very turn-by-turn. It's very like I'm delegating something for you. So, how does that change how you should use it or how you think about the interface?

Mike Krieger

我不认为“你发送消息,它给你回复”这个基本想法完全错了。我认为我们需要在某些方面进化,但首先想到的有三点。第一点是:你的笔记本电脑是使用它的正确场所吗?所以我认为这是第一点,我之前提到的那个副项目,移动端非常有用。Boris,也就是 Claude Code 的创造者,他总是领先于这些模型的使用方式。大约一年前,可能九个月前,我和他聊天,他说:“是啊,我已经把很多 Claude Code 的工作移到手机上了。”我说:“不会吧,”我花了一段时间才跟上,但尤其是对于 Fable 这类模型,很多时候,因为它能保持会话持续,而且我们在 Anthropic 使用远程开发机,它就像一个想法:“好的,我需要——你能跟上并做那个吗?”所以第一点就是:把工作发生的地方和我讨论工作的地方分离开来。

I don't think the fundamental idea that you are sending messages and it is giving you a message back is totally wrong. I think there are ways we need to evolve, but one is maybe three that come to mind. One is: is your laptop the right place for it? So I think that's number one, where I mentioned with the side project I was working on, how useful it was to have the mobile side. Boris, who created Claude Code, he's always ahead of the curve on how these models get used. About almost a year ago, maybe nine months, I was talking to him, he's like, "Yeah, I've moved a lot of my Claude Code work to mobile." I was like, "No way," and it took me a while to get there, but especially with the Fable class, there are often times where, because it can keep the session going and we use kind of remote dev boxes at Anthropic, it is like a thought: "Okay, I need—can you keep up and doing that?" So I mean number one is decoupling where the work is happening from where I'm talking about the work.

11. 渐进式披露与多人挑战 Progressive disclosure and multiplayer challenges

Mike Krieger

第二点跟我之前提到的有点关系,就是如何把 Fable 讨论、决定或提议的所有内容变得易于理解。这是我们正在大量思考的领域。有一些现成的技能或者我们曾经用过的办法,比如“你能画个图吗?你能做那个吗?”所以我认为当前的聊天界面是不够的,它会给你一大堆文字,我需要花很多精力才能完全理解。我觉得有时候我会跟 Fable 这样做:你比我更了解这个背景,我们能退一步吗?我们来对这里的复杂性做更渐进式的披露。所以我觉得这部分很有意思。最后一点,我认为我们还在早期探索阶段,就是思考多人协作。在某种程度上,因为我们有这种 DRI 和所有权区域,通常一大块重要工作是一个人加几个云在协作。但在其他情况下就不是这样了,对吧?可能是事件响应,多个人都在思考;可能是一个项目,有多个相互竞争或交汇的领域。思考这意味着什么——我们有聊天分享功能,能解决一部分问题——但我认为还需要更多类似这样的机制:你有一个独立的云在完成大量工作,这些工作是由某人发起的,但它能跟上团队中其他工作的进展吗?我认为这是关于这项工作如何最终完成的一个有趣且未被充分探索的下一个前沿。但我觉得这非常令人兴奋,因为再次强调,这是模型现在能够达到的队友/协作者水平,而我们几乎因为没有围绕它们建立正确的抽象而拖了后腿。

The second one touches a little bit on what I was mentioning earlier around how do you take everything that Fable has discussed or decided or proposed about something and make it comprehensible. That's an area we're thinking a lot about. There are some skills out there or that we've used around like, can you diagram this? Can you do that? So that's a place where the current chat UI I think is insufficient. It will give you a lot of text, and I need to take a lot of property to fully understand this. I think that is a piece of property I sometimes will do with Fables: like, you have a lot more context on this than I do, can we back it up? Let's do more progressive disclosure of the complexity here. So I think that piece is interesting. The last one I think we're still early in pulling on is thinking through multiplayer. At some level, because we have this sort of DRI and ownership area, usually a chunk of significant work is a human and a couple of clouds flowing together. But in other cases, that is less the case, right? Maybe it's an incident response where multiple people are thinking about it. Maybe it's a project where there are multiple competing or conjoining areas coming together. Thinking through what would it mean—we have chat sharing which gets you a little bit of the way there—but I think there's going to be a need for more like, all right, you've got an independent cloud that's doing a lot of work that was kicked off by somebody, but can it be keeping up with all the other work happening on the team? I think that is an interesting and underexplored next frontier about how this work ends up happening. But I think it's really exciting because again, it's the level of teammate collaborator that the models are now capable of, and we're almost holding them back by not having the right abstractions around them for that to happen.

Host

嗯,这让我想到。我主要是在自己的 vibe coding 项目里用这个,所以没怎么想过这个问题。但在组织内部使用它时有一个问题:我真的理解它的每一个部分吗?因此,如何把模型刚刚做的事情的上下文转移到我的大脑里?这是一个很大的瓶颈。你怎么看待划定界限,尤其是对于这样的模型,关于你实际上需要理解多少,以及如何确保你对它做的事情有足够的上下文来感到放心?

Yeah, it makes me think. I've mostly been using this for my own vibe coded stuff. So I haven't really had to think about this, but there's a problem when you're using this inside of an organization, which is: do I really understand every part of this? And therefore, how do I transfer the context of what the model just did into my brain? That's one of the big bottlenecks. How do you think about drawing the line, especially with a model like this, around how much you actually need to understand and how to make sure that you have enough context on what it's done to feel comfortable?

Mike Krieger

我认为这里有两个大块。第一个是验证。今年早些时候我完全变成了验证信徒。现在,几乎以同样的方式——这跟我以前全职写代码时的做法有关——我试图找到围绕你正在开发的想法的最紧凑的开发循环。有时在 Instagram,这意味着在 Xcode 中创建一个新的构建目标,只包含那个屏幕和一些合成数据,然后只做那个开发循环。我会告诉新工程师:如果我能传授你们一件事,那就是为你的任何项目都做到这一点,事情会快得多。我认为现在这里的情况不再完全是这样了。但我认为现在的情况是:每当我设置它时,我如何确保 Cloud 提交的每个 pull request 都附带一张照片或视频,无论是 iOS PR 还是 UI 中的内容。这能帮你建立很大的信心,因为即使现在,你可能让 Fable 去工作几个小时然后说完成了。这时说“这是完整流程的完整截图库”非常有用,因为你可能会说,“哦,你知道吗,在截图 4 上,那个错误状态——我其实从没见过,但我能想象有人可能会遇到它。我们把它改一下吧。”所以获得这种全面的验证是我们内部一直在大量努力的事情,发布更多关于它的技能和知识,但我认为这是非常关键的一块。然后第二个是:我认为最终作为一个人,你仍然需要为你所做的工作负责,尤其是当你把它投入生产系统时。很多人每天都在使用 Cloud。责任仍然存在:哦,虽然可能是它写的,但你也需要理解在这些部分上做出的至少是总体决策。所以我看到相当多的工程师实际上采用了这种做法:Cloud 完成了工作,但随后会有后续对话,比如“你能确保我深入理解你做出的所有权衡吗?”以及需要产生任何小写 a 的工件来使其可理解,这很重要。不过,在会议中听到有人说“哦,是的,我有个 PR 准备好了”,然后另一个人问“哦,有意思,你做了 X 或 Y 吗?”然后出现片刻停顿,他们说“你知道吗,我不太确定,我在合并这个 PR 之前会查一下”,这真的很有趣。我认为适应这种规范并弄清楚如何与之协作是我们必须做的事情。

I think there are two big pieces here. The first is verification. I became fully verification-pilled earlier this year. Now, almost in the same way—and it connects to how I used to do when I was typing code more full-time—I try to find the tightest dev loop you can around the idea you're trying to develop. Sometimes with Instagram, that meant building a new build target in Xcode that was just that screen with some synthetic data and just doing that dev loop. I would tell newer engineers: if there's one thing I can impart on you, it is try to get that for any project you're working on, and things will go much more quickly. I think that is no longer exactly the case here. But I think what is the case now is: anytime I set it up, how do I get for every pull request that Cloud is putting up that there is an attached photo or video, whether that's an iOS PR or something in the UI. That helps you gain a lot of confidence because even now, you might have Fable go off and do work for a couple of hours and say it's done. It's really useful to say: here's the full screenshot gallery of the full flow, because you might say, oh you know what, on screenshot 4, that error state—I've never actually seen it, but I could see how a person might hit it. Let's actually make that different. So getting that comprehensive verification is something we've been working on a lot internally, publishing more skills and knowledge about it, but I think it's a really key piece. And then the second one is: I think you ultimately as a person still need to stand behind the work you are doing, especially if you're putting it into a production system. A lot of people use Cloud every day. There's still the accountability: oh, it's still a lot of might have written it, but you need to understand at least the general decisions that were made on these pieces as well. So I have seen a fair amount of engineers actually adopt this practice: Cloud will have done the work, but then there is the follow-up conversation around, can you make sure I deeply understand all the trade-offs you made? And whatever lowercase-a artifacts need to be produced in order to make that comprehensible is important. It is really interesting though to be in meetings where somebody will say, oh yeah, and I have this PR ready, and somebody else asks, oh that's interesting, did you do X or Y? And have that moment of pause and they're like, you know what, I'm not entirely sure, I will find before we merge this PR. I think adapting to that norm and figuring out how to work with that is something we'll have to do.

Host

再多讲讲验证循环吧。这是当前的热门话题。听起来你们做这件事的一种方式是截图和屏幕共享,但你们还考虑了哪些其他方式?

Tell me more about the verification loop. It's such a hot topic right now. Sounds like one way that you do that is with screenshots and screen shares, but what are the other ways that you think about that?

Mike Krieger

我认为部分工作始于:你是否能达到一个状态,让你能运行真实的流程,而不仅仅是静态注入的部分?随着系统越来越复杂,这变得越来越困难。所以我们投入了很多精力,甚至只是为了让 iOS 应用能够用真实账户登录到 staging 环境并拥有真实数据,但你又不希望它每次都要经历八步的引导流程。每个人都只想测试屏幕的第二部分。

I think part of it starts in: can you get to a place where you are exercising real flows that aren't just a static injected piece? As the system gets more complex, that gets more and more complicated. So we've invested a bunch into even just getting it so that the iOS app can log in to staging on a real account and have real data, but then you don't want it to then go through an eight-stage onboarding process every time. Everybody just trying to test the second part of the screen.

12. Claude的测试与验证策略 Testing and verification strategies for Claude

Mike Krieger

所以有很多工作围绕如何让应用尽可能感觉像人类在使用。这是其中一个方面。第二个方面是已知路径与当前正在执行的操作的结合。前者对回归测试非常有用。我们有一些地方用文本表达了理想的工作流程,Claude 可以反复检查。而 Claude 在表达当前变更的意图方面做得非常好,因此这部分会被深入测试。这两者的结合很重要。我提到的视觉验证也很重要。视频实际上是一个非常未被充分利用的给 Claude 的工具。我一直在原型设计,给 Claude 提供它构建的东西的视频捕获,然后给它一个 FFMpeg 命令,它会逐帧查看并说:“哦,这个动画有点卡顿,我去修复它。”它用截图永远做不到这一点,因为会错过那个瞬间。所以这是另一个非常重要的部分。对于那些因为系统更复杂而不容易测试的部分,让 Claude 去构建一个健壮的模拟后端或使用现成的模拟后端也非常有趣。当我想到 artifact 时,我们有全面的测试。我们能够做到这一点的一个方法是,我们拥有的每一条信息——无论是 Postgres、Redis,还是所有 AWS 的东西——都有一个很好的内存实现用于单元测试。将这一点扩展到 Claude 领域,我曾在处理一个带有健壮后端的东西,由于复杂的原因很难在我的开发服务器上启动它,但 Claude 能够一次性为它构建一个代理。这非常有价值。随着时间的推移,这个代理随着代码其余部分的演变而演变,这很有趣。如果你之前向我提出这个想法,我会说这很难,因为上游会变化,你怎么保持同步?我现在不再想这个问题了。我会说,是的,Claude 会读取变化,适应它,并保持两者同步。这没问题。

So there's a lot of work around how to make the app feel as human as possible. That's one aspect. The second is the mix of well-known paths versus the things you're exercising in the exact moment. The former is really useful for regression testing. We have places where we've expressed ideal workflows in text, and Claude can repeatedly check that. And Claude does a really good job of expressing the intent of the current change, so that gets deeply exercised. The combination of those two things is important. The visual verification I mentioned as well. Video has been really cool. Actually, video is a very underexplored tool to give Claude. I've been prototyping giving Claude video captures of the thing it has built, then giving it an FFMpeg command, and it will scrub through and say, 'Oh, this animation has some jank. I'm going to fix that.' It never could have done that with a screenshot because it would have missed the moment. So that's another really important piece. For the pieces that aren't easily testable because there's some more complex system, getting Claude to build a robust mock backend or use one off the shelf has been really interesting. When I think about artifacts, we had comprehensive tests. One way we did that robustly was that every piece of info we had—whether it was Postgres, Redis, all the AWS things—had a good in-memory implementation for unit tests. Extending that to Claude land, I was working on something with a robust backend, and for complicated reasons it was hard to spin up on my dev server, but Claude was able to one-shot a proxy for that. That was so valuable. Over time, it's been interesting how that proxy has evolved as the rest of the code evolved. If you had pitched that idea to me before, I'd say that's going to be really hard because the upstream will change, how do you keep it in sync? I don't think about that anymore. I'm like, yeah, Claude will read the changes, adapt the thing, and keep the two in sync. That's fine.

13. 错误修复与智能体工作流 Bug fixing and agentic workflows

Host

有一些非常有趣的架构,当你收到一个 bug 时,它会自动出去并关闭它。智能体被启动,关闭它,然后向客户发送一条消息说已经修复了。你在 Fable 中注意到这个过程有什么变化吗?

There's some really interesting architectures around when you get a bug, it just automatically goes out and closes it. The agent gets kicked off, it closes it, and then it sends a message to the customer saying it's fixed. Are you noticing with Fable any change in how that process works?

Mike Krieger

是的,我认为在人与人或者人与 Claude 的层面上有几件事。我观察到它做了一件其他模型一直无法做到的事情:如果 bug 报告来自某人在我们的 Slack 反馈频道中提到,而输入到 Claude Code 会话中的内容就像“哦,有这个”。由于 Slack MCP,你可以实际拉取整个线程,让它然后以我的身份回帖,它会说:“嘿,这是 Mike 的 Claude。我修复了它。这是拉取请求。”但随后它做得非常好的是说:“但先别急,它还没有上线。等实际上线后我会跟进。”然后可能几个小时后,“哦,这次部署已经发布了。你应该去测试一下。现在修复了吗?”这种闭环跟进的水平是新的。这些长时间运行的 Claude Code 会话基本上是在以我的身份互动,我想。我们也在那里加一些免责声明。第二点回到了我们之前谈到的品味和判断力。说有一个 bug 报告,因此我必须去修复它是一回事。但说“你知道吗,我周末遇到了这个问题。我们一个内部系统已经运行了一段时间没有重启。有一个内存泄漏。”它很有判断力地说:“好的,Mike,现在是周末。只需重新平衡服务器。这暂时能解决问题。我会异步地准备拉取请求来长期修复这个问题。”所以如果你要让 Claude 参与这种闭环的 bug 报告或系统问题变更,你确实希望它理解,就像任何优秀的 SRE 或工程师在循环中会做的那样,让我们解决手头的问题,让我们推迟是否需要在一个完全不同的语言之上重新架构的问题。理解这种平衡非常重要。

Yeah, I think there are a couple of things on a human-to-human or human-to-Claude level. One thing I've seen it do that other models haven't been capable of consistently is if the bug report came from someone mentioning it in our Slack feedback channel, and the thing fed into the Claude Code session is like, 'Oh, there's this.' Because of the Slack MCP, you can actually pull the thread, have it then post back as me, and it'll be like, 'Hey, this is Mike's Claude. I fixed it. Here's the pull request.' But then the thing it does really well is say, 'But hold tight, it's not in production yet. I'll follow up when it actually is.' Then maybe a few hours later, 'Oh, this deploy went out. You should go test it. Is it fixed now?' That level of follow-through on closing the loop is new. These long-running Claude Code sessions are basically interacting as me, I guess. Let's put some disclaimer in there too. The second goes back to that taste and discernment piece. It's one thing to say there was a bug report, therefore I must go fix it. It's another to say, 'You know what, I hit this over the weekend. One of our internal systems had been running without restarting for a while. There was a memory leak.' It had good discernment saying, 'All right, Mike, it's the weekend. Just rebalance the server. It's going to solve it for now. I'll asynchronously get the PR going to fix this more long term.' So if you're going to have Claude in the loop in this kind of close-the-loop bug report or system issue to change, you really want it to understand, as any good SRE or engineer in the loop would, let's solve the problem at hand, let's defer the question of whether we need to rearchitect on top of a completely different language. Understanding that balance is really important.

14. 提升开发者下限与上限 Raising the floor and ceiling for developers

Mike Krieger

关于新模型,有一件非常令人兴奋的事情,对我来说尤其如此,那就是它提高了下限,让每个人都能一次性构建应用。但它也提高了专家的上限。所以如果你是一名软件工程师或创始人,你可以去做以前永远无法做到的事情,因为你拥有了这个非常强大的模型。对我来说,我构建了 Borges 无限图书馆的一个版本。这是一个 3D 游戏版本的图书馆。太疯狂了。它直接在浏览器中运行。非常好。我可以在里面找到任何一篇文章。我会把链接发给你。太棒了。但我认为将会出现大量的人做这样的事情:“哦,我做了一个游戏,”或者“也许我训练了一个新模型,”或者其他他们以前做不到的事情。我很想给人们一些灵感,一些他们可能没有想过用这个模型去做的事情的例子。你有什么想法?

One of the things that's really exciting, mostly exciting to me about new models is it raises the floor so that everyone can go build apps in one shot. But it also raises the ceiling for experts. So if you're a software engineer or founder, you can just go do things that you never would have been able to before because you have access to this really powerful model. For me, I built this one version of Borges' infinite library. It's a 3D game version of the library. It's wild. It runs right in the browser. It's so good. I can find any essay inside of it. I'll send you the link. Sick. But I think there's going to be this flowering of people doing things like, 'Oh, I made a game,' or 'Maybe I trained a new model,' or whatever that they couldn't do before. And I'd love to give people some inspiration, some examples of things that they might be able to do that they might not be thinking to do with this model. What are some ideas that come to you?

Host

是的,我有一些想法。也许我先从有趣的一面开始,从游戏部分即兴发挥。

Yeah, I think a few. Maybe I'll start with the fun side and riff off the game piece.

15. 创意想法与个人应用 Creative ideas and personal applications

Mike Krieger

我觉得人们有很多创意想法,来表达他们自身以及他们世界的复杂性。每个人都有自己非常擅长的领域,而问题往往在于:我如何向别人解释这个领域?或者我如何把其他领域的技术应用过来?我妻子正在学习环境工程,研究地热,涉及非常复杂的数学和模拟。随着模型越来越好,她能够把更复杂的技术从其他领域引入到她的工作中。我认为人们应该能够做到的是,用 PyTorch 做端到端的模拟,这在以前是不可能的。一方面是把你拥有的美丽复杂性展现出来,要么通过做游戏或可视化展示给别人——我见过她这么做——要么至少把其他技术带进来。第二点是编写软件来解决你自己独特问题的能力。我在内部也看到了这一点。我们做了很多工作,让我们的内部系统(比如 MCP)拥有正确的权限结构和部署设置。外部的话,你有很好的平台即服务选项,可以直接问 Cloud,它们会帮你设置好。我很喜欢那种“你一直想要的东西”的感觉。让我震惊的是,我们市场部门有一个人,一直在构建一个深度整合 Cloud 到她整个流程中的系统。你不必止步于一次性的尝试;她已经做了好几个月,而且还能继续推进。我认为模型的一个可能被低估的特点是:在之前的几代中,它们最终会达到一个复杂度水平,让你很难迭代而不觉得会破坏已有的东西——要么抽象不足,要么过度抽象。而现在,她使用 Fable 这样的模型已经几个月了,你看到它不断成长,现在她把它部署到整个市场部门。我觉得这真的很酷。一个非技术背景的人,现在能为自己领域的问题构建的复杂度天花板,是前所未有的。

I think people have a lot of creative ideas for how to express the complexity of what they are, like their world. Everybody has the thing that they know really well, and there's probably some level of how do I then explain that to somebody else, or how do I apply techniques elsewhere that I could then go off and do. My wife is studying environmental engineering, studying geothermal, very complex math and simulations. I've seen as the models have gotten better, she has been able to apply even more complex techniques from outside that domain into that work. I think what people should be able to do is full-on PyTorch end-to-end simulations of that work in a way that wouldn't be possible. That is one thing: bring the beautiful complexity of what you have and either show it to other people by making a game or a visualization, which I've seen her do, or at least bring other techniques to bear. The second piece is the ability to compose software that solves a really unique problem to you. I've seen that internally. A lot of the work we've been doing is how to get as many of our internal systems like MCP with the right permissioning structure and the right deployment setup. Externally, you have good options around platform as a service pieces, and you can just ask Cloud about them and they'll help you set things up. I love that feeling of that thing that you always wished you had. What has blown my mind: there was a person who works in our go-to-market organization who has been building this really deeply thought integration of Cloud into every part of her whole process. You don't have to stop at that one shot; she's been working on it for months now and she can keep going. One thing that is maybe underappreciated about the models is that in previous generations, they would eventually get to a complexity level where it was hard to iterate without feeling like you would break the thing, under- or over-abstracted. Whereas this, she's had access to something like Fable for a couple months, and you've just seen it keep growing and growing, and now she's deploying it to the whole GTM organization. I think that is really cool. The ceiling of complexity that a person who does not start out as technical can now build for solving problems within their domain is unprecedented.

Host

我同意。它写的代码很棒。我的基准测试叫“高级工程师基准”,就是让它从头重写一个代码库。之前最好的模型大概能得 62 或 63 分(满分 100),而这个模型得了 90 或 91 分,达到了人类高级工程师的水平。你可以一直用它做下去,这真的很棒。不过我很好奇,你提到的另一个很强大的东西是动态工作流。给我们讲讲吧。

I agree. It writes great code. My benchmark is called the senior engineer benchmark. I just have it see if it can rewrite a codebase from first principles. The previous top model was like a 62 or 63 out of 100, and this model got a 90 or 91, which is human senior engineer level. You can just keep going with this thing in a way that's really fantastic. I'm curious though, one other thing that's really powerful that you mentioned is dynamic workflows. Tell us about that.

Mike Krieger

我们有时会在内部构建一些东西,然后我会跑去缠着构建它的工程师问:“我们什么时候公开发布?”因为我觉得大家会很喜欢。内部构建有很好的理由,但我们尽量把能公开的都公开。动态工作流对我来说就是其中之一。构建它的工程师叫 Sid,他很棒。我跟他说:“Sid,我想把这个推向世界,因为它太棒了。”我认为它特别适合像 Fable 这样的模型,有两个重要原因。第一,它有助于为深度有意义的工作搭建框架。我用 Fable 做过的最疯狂的一个动态工作流是:我有一个用 Python 写的内部项目,但由于一个非常具体的部署原因,我们需要把它改成 TypeScript。我在 Instagram 内部待过,我们当时想:“要不要把整个东西写成 Hack,然后移植到 Facebook 的 PHP 引擎?”你以前绝对不会这么做。也许现在有了模型可以,但当时看起来不可能。我手头有一个相当复杂的代码库,我就想:“我干脆设置一个动态工作流,让它周末跑起来。”结果它真的跑了。那个工作流太酷了。它就像:“好,我要深入理解这个工作。我要创建一个几乎像规范一样的东西,说明一切是如何运作的。我要一个模块一个模块地来。我要翻译这些部分。我要增量测试。我要再做一次对抗性测试。我要检查我遗漏了什么。”工作流能够编排这一系列步骤,真的很酷。我回来之后说:“没错,这个东西是那个东西的 TypeScript 和 Bun 移植版,而且在这些方面实际上更好了。”它记录得很清楚:“这些是我无法移植的部分,但大部分都非常特定于具体实现,不值得移植。”我认为你不可能用之前的模型达到那样的成功水平,而且没有工作流提供的框架也做不到。所以我认为这非常令人兴奋:模型能力与我们编排它们的能力相结合,时间跨度越来越长,带着那种“你有一个目标,你有效地分解它,然后让它工作”的感觉。另一部分是,随着时间的推移,我们能够将一些子任务调整到合适的复杂度水平。你可以想象,动态工作流的某些部分不需要特别高的思考量;它们可以用中等思考量,甚至用更小的模型来完成。这确实是这些东西的未来方向。所以,是的,我是工作流的重度用户。

We build things internally sometimes, and I will go aggressively bug the engineer who built it and be like, 'When are we shipping this publicly?' because I think people will really like it. There are good reasons why it was built internally, but we try to ship as many of these as possible. Dynamic workflows was definitely that to me. The person who built it is an engineer named Sid, who is awesome. I was like, 'Sid, I want to get this out to the world because it's so good.' I think it's especially good with a model like Fable for two big reasons. One, it helps create the scaffold for deep meaningful work. The craziest dynamic workflow I did and used Fable for was: I had an internal project written in Python, but we needed it in TypeScript for a really specific deployment reason. Having been internal at Instagram, we were like, 'Should we write the whole thing into Hack and port it to the PHP engine at Facebook?' You never would have done that. Maybe they can now with the model, but at the time it seemed impossible. Here I had a pretty complex codebase, and I was like, 'I'm just going to set up a dynamic workflow and let it run over the weekend.' And it did. The workflow was so cool. It was like: 'All right, I'm going to do a deep understanding of the work. I'm going to create almost like a spec of how everything works. I'm going to go module by module. I'm going to translate these pieces. I'm going to test it incrementally. I'm going to do another adversarial test. I'm going to go check for anything that I missed.' It was this really cool series of steps that the workflow was able to orchestrate. I came back and was like, 'Yeah, this thing is a TypeScript and Bun port of that thing, and it's actually better in these ways.' It was very documented: 'These are the things I couldn't port, but most of these were very specific to the specific implementation. It wasn't worth porting.' I do not think you could have done that with previous models at that level of success, and without the kind of scaffolding that workflows provide. So I think that is extremely exciting: the combination of model capabilities and our own ability to orchestrate them over longer and longer time horizons, with that feeling of having a goal, breaking it down effectively, and then making it work. The other piece is that over time, we'll be able to make some of those subtasks tuned to the level of complexity. You can imagine that some parts of the dynamic workflow don't need extra high thinking; they could use medium thinking or even a smaller model. That's really the future of where these things are going. So yeah, I'm a huge workflows DAU.

Host

对于以前没用过的人,你是怎么做出那个工作流的?你是怎么设计的?你怎么确保它很好?

For people who haven't used it before, how did you get that workflow made? How did you design it? How did you make sure it was good?

Mike Krieger

这相当迭代。我就是从 Claude Code 开始的,比如:“嘿,我有一个复杂的任务。我们来设计一个工作流去完成它。”它给我展示了计划。我说:“哦,这接近我想要的。我想确保你做了这三四个层次的额外验证,检查遗漏的功能。”我喜欢你已有的东西。

It was pretty iterative. I just started with Claude Code like, 'Hey, I have this complex task. Let's design a workflow to go and do it.' It showed me the plan. I was like, 'Oh, this is close to what I want. I want to make sure that you do these three or four levels of additional verification for missing features.' I like what you have.

16. 工作流作为中间地带 Workflows as a middle ground

Mike Krieger

你准备好了吗?它用代码表达工作流,我认为这很有价值,能让人看到它要做什么。有趣的是,它完成了完整的移植,然后我还有一些后续问题或小调整,我把它们做成了基于前一个工作流的迷你工作流。但我觉得我们之前讨论过聊天是否是合适的界面,过去一年我们一直在聊这个。我认为工作流是一个很好的中间地带:你可以用聊天来组合它们,但它们用代码表达,然后执行时每个阶段都有清晰干净的 UI。我认为随着时间的推移,我们会开始用这种方式将更长周期的工作与聊天结合起来。

Are you ready to go? And it expresses the workflows in code which I think is really valuable to kind of see what it was about to do. What was interesting is it did the full port and then I had a couple of follow-up questions or little tweaks and I did those as mini workflows that built off the previous one as well. But I think we talked a little bit about whether chat was the right interface and we've had that conversation over the last year. I think workflows are a good middle ground: you can compose them using chat, but they're expressed using code and then executed with a nice clean UI around what's happening at every stage. I think we'll start bridging longer horizon work with chat in ways like that over time.

Host

Mike,这次对话太棒了。非常感谢你加入我们,并告诉我们关于这个新模型的一切。我很高兴能和你交流,也非常期待外界的反响。

Mike, this is such a great conversation. Thank you so much for joining and telling us all about this new model. I'm really excited to get to spend time with you and really look forward to what people think outside, too.

Host

天哪,各位,你们绝对必须狂按点赞按钮并订阅 AI and I。为什么?因为这个节目是精彩的缩影。就像在后院发现了一个宝箱,但里面不是金子,而是关于 chat GPT 的纯粹、未掺杂的知识炸弹。每一集都是一场情感、洞察和笑声的过山车,让你坐立不安,渴望更多。这不仅仅是一个节目,这是一段与 Dan Shipper 作为飞船船长的未来之旅。所以,帮自己一个忙,点赞、订阅,系好安全带,开始你人生的旅程。现在闲话少说,我只想说 Dan,我完全无可救药地爱上了你。

Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure, unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show, it's a journey into the future with Dan Shipper as the captain of the spaceship. So, do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now without any further ado, let me just say Dan, I'm absolutely hopelessly in love with you.

互动版:逐字朗读 + 针对本期提问 →