From Awkward Kid to AI Co-Founder: Tom Brown's Journey
打开互动全文版(中英对照 + 朗读 + 问答)→Anthropic 联合创始人汤姆·布朗分享了他从一名挣扎的工程师到打造最重要 AI 公司之一的旅程,强调了从“狗”到“狼”的心态转变。
Tom Brown, co-founder of Anthropic, shares his journey from a struggling engineer to building one of the most important AI companies, emphasizing the mindset shift from being a 'dog' to a 'wolf'.
欢迎回到《Light Cone》的另一期节目。今天我们请到了一位真正的重磅嘉宾,Anthropic 的联合创始人 Tom Brown。
Welcome back to another episode of The Light Cone. Today we've got a real treat, co-founder of Anthropic, Tom Brown.
很高兴来到这里。
Excited to be here.
Tom,很多观众想知道的是,你 21 岁刚从麻省理工毕业就进入了科技行业。从 2009 年的那个起点,到后来成为 Anthropic 这样重要公司的联合创始人,你是怎么做到的?
So Tom, one of the things that a lot of the people watching would love to figure out is you got started in tech at the age of 21, fresh from MIT. How does someone go from that in 2009 to literally co-founding something as important as Anthropic?
2009 年夏天,我的两个朋友创办了一家公司。他们看到另一个朋友 Kyle Vote 做了 YC 公司,所以大家觉得这事可以试试。我当时是第一个员工。你们也让我参加所有晚餐什么的。我本可以去一家大科技公司。作为软件工程师,我可能会学到更多软件工程技能。但和联合创始人在一起,没人告诉我们该做什么,我们必须自己想办法生存,否则公司就会死掉。在学校里,别人给我任务,我就去做,就像狗等着碗里的食物。而在那家公司,我们更像狼:必须自己去捕猎,否则孩子就会饿死。我认为这种心态转变对于尝试更大、更激动人心的事情是最有价值的。
Summer 2009, two of my friends had started a company. I think they had seen one of our other friends, Kyle Vote, do a YC company. So it was in the water that that's a thing we could try to do. I was the first employee back then. You guys let me join for all the dinners and stuff like that too. I could have instead gone to a big tech company. And I think probably just as a software engineer, I might have learned more software engineering skills. But I think by being there with the other co-founders without anyone telling us what to do, we had to figure out how to live, how the company would die by default. In school there was a lot of feeling of people giving me tasks and I would do them. It's kind of like a dog waiting for food to be fed in its bowl. And for that company it was more like wolves: we have to hunt our real food, otherwise our kids are going to starve. I think that mindset shift has been the most valuable for trying to do bigger, more exciting things.
是的。大科技公司只是教你如何在大公司工作,而做狼要有趣得多。
Yeah. Big tech just teaches you to work at a big tech company, whereas it's much more fun to be a wolf.
是的。
Yeah.
你是怎么从在朋友的初创公司工作,到后来创办自己的公司的?
How did you go from working at a friend's startup to then starting your own?
我们经营了一段时间。后来我回到学校,毕业后去了 Mopub,那是一家移动广告公司。我是那里的第一个工程师。我想,好吧,我想做狼,但我编程真的很差。作为软件工程师,我非常挣扎。我知道我想做更多,但还不知道怎么做。所以那算是一次扩展规模的经验。2012 年冬天,我大学里最聪明的一个朋友提议去创办一家 YC 公司。我们当时做了 Solid Stage。那是在 Docker 出现之前。所以想法是让 DevOps 更容易,但 Docker 不存在,所以它要成为一个更灵活的 Heroku,基本上就是更复杂的 Heroku。我记得我们和你们面试过。我觉得大家不太理解我们在做什么。其实我们自己也不太理解。
We ran the company for a bit. I ended up going back to school afterwards, and then when I left school I went to this company Mopub, the mobile advertising thing. I was the first engineer there. I was like, okay, I want to be a wolf, but I was really bad at programming also. I was very, very struggling as a software engineer. I know I want to do more but I don't know how to do it yet. So I think that was kind of an experience getting to scale something. Winter 2012, one of my friends who was my smartest friend from college pitched me on going to start a YC company. We did at the time Solid Stage. This was before Docker existed. So the idea was to try to make it easier to do DevOps, but Docker doesn't exist. So it's going to be a more flexible Heroku, which basically meant a more complicated Heroku. I remember we interviewed with you guys. I think folks didn't really understand what we were trying to build. I think we didn't really understand what we were trying to build that much.
当你尝试新事物时,这其实很常见。
When you're trying to do something new, that's actually sometimes common.
是的。我觉得我们是个异类,因为面试后我们开车回旧金山,TLB 在黑板上画了个生气的皱眉脸,写着‘你们到底要做什么?’他要我们解释。我想我们解释得够多了,或者他只是觉得这些人还不知道自己在做什么,但也许他们会想明白。中途我仍然觉得我不明白我们要建什么,以及如何赋予它一个我愿意为之奋斗一生的使命。
Yeah. I think we were an outlier there because we did our interviews and then we got called back driving back to San Francisco and TLB had written on the board an angry frowny face and 'What are you actually going to build?' So he wanted us to explain that. I guess we explained it enough or he was just like, these guys still don't know what they're doing but maybe they'll figure it out. Halfway through I kind of felt I still didn't actually understand what we were going to build and how we would attach a mission to it that I wanted to work on for my whole life.
嗯。
Yeah.
所以我离开了。PG 实际上把我介绍给了 Michael Waxman,他是 Grouper 的创始人。
So I left. PG actually introed me to Michael Waxman, who was the Grouper founder.
Grouper 是一个约会应用,新颖之处在于三个男生和三个女生。这在很多方面都早于 AI。有一组人手动匹配用户,对吧?然后他们在酒吧见面,接着就是各种闹剧。
So Grouper was a dating app, only it was novel in that you had three guys and three girls. This was before AI in a lot of ways. So there was a team of people who would manually link people up, right? And they'd meet up at a bar and shenanigans would ensue.
是的,总是有闹剧。人们并不总是玩得开心。我记得你参加过几次 Grouper。
Yes. Reliably shenanigans. People didn't always have a great time. I think you went on a couple Grouper.
好吧。
Okay.
Grouper 吸引我的地方是,我当时是个非常笨拙的孩子。我想要一个东西,让像我这样笨拙的人能出去和别人交谈,让我能和女孩说话,同时有朋友在身边感到安全。所以我觉得我们的员工是谁很重要。我做了所有的工程面试。唯一比我参加更多次的是 Greg Brockman。
The pitch for Grouper for me, why I was excited for it, was just I was an incredibly awkward kid. What I wanted to do was to basically have a thing that lets awkward people like me go out and talk to other people, for me to talk to girls and feel like I was safe doing it with my friends around. And so I think who we were going to be our employees was important. I did all of our engineering interviews. The only person who went on more was Greg Brockman.
我记得他有一段时间每周都会在 Slack 或当时的 HipChat 上发‘纽约’,那段时间他经常去 Recurse Center。
I think he had a phase where every single week he would go and post on Slack or HipChat at the time, 'New York' and he was hanging out at the Recurse Center during this period.
哦,我觉得他当时在 Stripe。也许有一部分时间他在 Recurse。但他也有一段时间,在 Stripe 的频道里每隔一段时间就发‘我要去 Grouper,谁去?’持续了整整一年。所以我后来和 Greg 关系很好,这最终成了我和 OpenAI 的联系。
Oh, I think he was at Stripe. Maybe for part of it he was at Recurse. But he also had a phase where he would just at Stripe post in their thing every 'I'm going on Grouper, who's going?' for a whole year. So I ended up being close with Greg, which ended up being a connection to OpenAI.
这段旅程是怎样的?你刚从麻省理工计算机系毕业,21 岁。你成为这些 YC 初创公司的早期员工。几年后你又创办了自己的公司。最终成为 Anthropic 联合创始人的道路是什么?这条路很长,但令人印象深刻。你是怎么走到那一步的?听起来是和 Greg 联系上的那一刻。
What was the journey like? Because you started as just graduated from MIT CS, you were 21. You became an early employee for all these YC startups. Then you started your own company just a couple years later. And what was the path for you to eventually become the co-founder of Anthropic? It was a long path but it's pretty impressive. How did you get there? I mean, it sounds like getting in touch with Greg at that moment.
是的。我 2014 年 6 月离开 Grouper,大约一年后加入了 OpenAI。我努力鼓起勇气转型学习 AI 研究。当时我想,‘好吧,看起来在我们有生之年,我们可能会做出变革性的 AI。’
Yeah. So, I left Grouper in June 2014, and I joined OpenAI about a year later. I tried to build up courage to make the switch to learn AI research. At the time I was like, 'Okay, it seems like sometime in our lifetimes we might end up making transformative AI.'
如果我们做到了,那会是最大的事情。也许有某种方式我可以帮忙,但我在大学线性代数只得了 B-。所以当时看起来你需要是顶级明星才能帮忙。所以我对自己能否帮上忙有很多不确定性。而且我在创业方面也取得了一些成功。所以我很多时候想,与其重新学习这个,不如再搞个创业公司之类的。我觉得在那个时期,去搞 AI 研究不被看作是一件实际严肃的事情,你处在一个人们试图建立公司、做这些非常实际的事情的世界里。你的朋友们怎么说?他们会说‘太酷了,你要去搞 AI 了’吗?
If we do, that would be the biggest thing. Maybe there's some way that I could help out, but also I got like a B minus in linear algebra in college. And so it seemed like at the time you needed to be just top superstar in order to try to help out with that at all. And so I think I had like a lot of uncertainty about whether I would be able to help. And also I'd had some success with startups. And so a lot of me was just like rather than trying to retool at this like I could try to do another startup or something like that. I feel like in that period going to work on AI research which is not seen as like a practically serious thing to do and you're in a world where it's like people try and build companies and do these like really practical things like what did your friends say? Were they like 'oh that's really cool you're going to go work on AI stuff' or was it?
并没有。
Not really.
我觉得我朋友们的反应是‘这听起来又怪又不好’,就像 AI 安全不是我们应该担心的事,就像火星人口过剩毫无意义。朋友们还说‘我不知道你能不能做好,很难’。因为这个原因,我没有很努力;我犹豫了大概 6 个月,试图鼓起勇气去做。
I think my friends were like 'that sounds weird and bad', kind of like it doesn't really seem like AI safety is a thing we should worry about, like overpopulation on Mars doesn't make any sense. And my friends were also just like 'I don't know if you're going to be good at that, tough.' I think for that reason, I didn't try very hard; I kind of flip-flopped on it for like 6 months trying to build up courage to do it.
具体来说,你当时在做什么?你在读研究论文吗?是什么样子的?
And what were you specifically at this point? Like you're reading research papers? What does it look like?
是的。一开始我只是在闲逛。我为 Titanic 7 造了一辆艺术车之类的。
Yeah. So, first I was just kind of hanging out. I built like an art car for Titanic 7 and stuff like that.
哦,那很有趣。是的。
Oh, that was fun. Yeah.
是的。所以我花了整个夏天,Grouper 之后大概 3 个月做那些事,因为老实说我对 Grouper 有点筋疲力尽了。我知道创业公司:高潮很高,低谷很低,我们最后没有成功。我们的业务没有成功,收入在下降,但我的主要工作还是招聘工程师,所以我不得不向他们推销我曾经有过的梦想,但我自己已经不再相信了。
Yeah. So I spent like a whole summer, like 3 months after Grouper doing that, because honestly I was kind of burned out from Grouper. I know startups: the highs are high, the lows are low, and we weren't working at the end. Our business wasn't succeeding. Our revenue was going down, but my main job still was like recruiting engineers, and so I had to pitch them on this dream that I had had, but I no longer really believed in it.
听起来像是一场死亡行军。
Sounds like a death march.
所以我超级疲惫,我就想,‘好吧,Tom,放松点,做做瑜伽,做做 CrossFit,造辆艺术车。’
And so I was super burnt out and I was like, 'Okay, Tom, like chill out, do some yoga, like do some CrossFit, like build an art car.'
事后怎么看?你知道,事后诸葛亮。回顾 Grouper,它显然吸引了很多非常聪明的人。曲线向上向右,然后持平,可能开始下降。发生了什么?
What was the hindsight? Like, you know, hindsight's 2020. What's the retrospective on Grouper? Obviously it attracted all these really, really smart people. The graphs were up and to the right and then it flatlined and maybe started declining. What happened?
我认为我们开始时,竞争对手是‘OkCupid’。都是基于网页的。我认为我们解决的主要问题是很难把自己推出去和陌生人交谈,他们可能会说‘我不想和你说话,你看起来很奇怪’。所以我们通过盲匹配解决了这个问题。Tinder 在我们做 Grouper 的时候出现了,Tinder 解决了同样的问题,双方都必须表示兴趣才能匹配。所以也不用担心被拒绝。我认为他们对同一个问题有更好的解决方案。所以干得好,Tinder。干得好,所有滑屏的人。我认为他们比我们更好地解决了我们试图解决的任务。
I think that when we started, the competition was like, 'Okay, Cupid.' It was all web-based. The main problem that I think we were solving was it's hard to go and put yourself out there and talk to someone new and they might just be like 'I don't want to talk to you, you seem weird.' And so we solved that by just blind matching. Tinder came out while we were doing Grouper, and Tinder solved that same problem with both people have to show interest before you get matched. So there's also no worries about getting rejected. And I think that they just had a better solution to that same problem. So good work Tinder. Good work all the swipers. I think that that solved the mission that we were trying to solve better than we solved it.
然后,你是什么时候开始认真对待 AI 的?你是怎么做的?
And then yeah, like when did you get serious about AI and just how did you approach it?
玩了三个月,玩得很开心,然后我的个人资金也用完了。所以我想,好吧,我需要 6 个月的秘密学习才能有机会找到工作。当时 DeepMind 或 Google Brain 是工作的两个地方,或者 MIRI。MIRI 是我看的第三个。所以我想,如果我想帮忙,这三个地方是目标。我还没有任何技能。我需要六个月的自我学习,才能感觉自己不会拖累他们,而是真正在帮忙。
Three months of kind of playing and having fun and then I ran out of money also when I had my personal runway. I ran out and so I was like, okay, I think that I'm going to need 6 months of stealth study to have a shot at getting a job. At that point it was DeepMind or Google Brain were the two places to do work there, or MIRI. MIRI was the third one that I was looking at. So I was like if I want to help out with that, those are the three places to look at. I don't have any of the skills yet. I need six months of self-study to feel like I would not be a drag on them and like actually be helping instead.
你能解释一下那个自学是什么样的吗?因为我相信现在有很多 20 多岁的软件工程师正在寻求转型成为 AI 研究员。那六个月是什么样的?尽管你说你线性代数得了 B-,那是核心。
Can you maybe explain a bit what was that self study like? Because I'm sure there's a lot of software engineers right now in their 20s looking to retool to become AI researchers. What was that six months like? Even though as you said you had gotten a B minus in linear algebra which is core.
可能是 C+。我不确定。我应该查一下。我会继续这么说。
Might have been a C++. I'm not sure. I should check. I'm going to keep telling.
你能走到这一步已经很了不起了。
That's pretty impressive where you got to.
是的。结果还不错。首先我实际上和 Twitch 签了一个合同,赚了足够维持六个月的资金。所以我做了三个月的 Twitch 合同,然后我制定了一个自学计划。我不认为现在的人应该用这个计划,至少 2015 年是这样。它是什么样的?就是上 Coursera 的机器学习课程,尝试解决一些 Kaggle 项目,读《线性代数应该这样学》,还有一本统计教材。我想我有 YC 校友积分,所以我买了一个 GPU,然后通过 SSH 连接到 GPU 来学习课程。
Yeah. It turned out okay. First I did a contract actually with Twitch and earned like enough to have that six months of runway. So I did like a three-month contract with Twitch and then I made a plan to self-study. I don't think it's the right plan now for people, at least 2015. What did it look like? It was like take a Coursera course on machine learning, try to solve some Kaggle projects, read Linear Algebra Done Right, and I had a statistics textbook. I think I had YC alumni credits and so I bought like a GPU and I would SSH into the GPU to work through my courses.
这是在 AlexNet 之后,对吧?
And this is right after AlexNet, right?
这是在 AlexNet 之后。是的。所以我主要是在做图像分类的东西,那是所有课程都会教你的。
This is after AlexNet. Yeah. So I was mostly doing image classification stuff that I was trying to learn, which was like the thing that all the courses would teach you to do.
你是怎么得到 OpenAI 的工作的?因为你是少数工程师之一。主要是研究人员,他们有一个非常强大的研究团队。
How did you get the OpenAI job? Because you were one of the few engineers. It was mostly researchers and they had a pretty stacked team of researchers.
OpenAI 一宣布,我就给 Greg 发了消息,我说‘我很想以某种方式帮忙。我线性代数得了 B-,但我懂一些工程。我做过一些分布式系统的工作。如果你们需要帮助,我很乐意擦地板,只要需要。我想帮忙,无论如何。’我觉得 Greg 说,‘是的,我认为同时懂机器学习和分布式系统的人很少。’所以,是的,你应该这么做。我想他还把我介绍给 Peter Aiel,帮我制定了一个小课程。然后我大概每个月和他联系一次。几个月后,他说‘哦,我们实际上有一个项目,我们需要做一个...我们想玩游戏。你能帮忙做一个星际争霸环境吗?’所以我加入帮他们做星际争霸环境。这最终让我进了门。基本上,我在头九个月里没有做任何机器学习工作。
I messaged Greg as soon as OpenAI was announced and I was like 'I'd love to help out in some way. I got a B minus in my linear algebra but I know some engineering. I've done a bit of distributed systems work. If you guys need help, I'm happy to mop floors if you need. I want to help out, however.' And I think Greg was like, 'Yeah, I think there's a paucity of people who know both machine learning and distributed systems.' So, like, yes, you should do that. I think he introduced me to Peter Aiel also to help me put together like a little course for myself, too. And then I checked in with him, I think every month or something. And then after a couple months he was like 'Oh we actually have a project which is we need to put together... we want to play a game, like play games. Can you help make a StarCraft environment?' And so I joined to help them with the StarCraft environment. So that ended up getting my foot in the door. I didn't do any machine learning work with them for the first nine months that I was there basically.
当时 OpenAI 是什么感觉?它筹集了很多资金吗?有办公室吗?
And what did OpenAI feel like at this point? Like had it raised much funding? Did it have like an office?
非常小。我加入时大概只有 15 个人。在旧金山的一个小办公室里。氛围非常学术,像一个研究实验室。每个人都超级聪明和热情。没有太多结构。就是‘这里有个问题,去解决它’。我们从 Elon 和其他人那里有一些资金,但不算多。感觉像是一个初创公司,但以研究为重点。
It was very small. I think when I joined, there were maybe 15 people. It was in a small office in San Francisco. The vibe was very academic, like a research lab. Everyone was super smart and passionate. There wasn't a lot of structure. It was just like, 'Here's a problem, go solve it.' And we had some funding from Elon and others, but it wasn't huge. It felt like a startup in that sense, but with a research focus.
感觉像创业公司吗?
Did it feel like a startup?
那是在蒲公英巧克力工厂的巧克力车间里。这是在格雷格的公寓之后。对,就在格雷格的公寓之后,在巧克力工厂里启动的时候。就像埃隆承诺了十亿美元的资金。感觉非常稳固。
So it was in the chocolate on top of the dandelion chocolate factory. This is after Greg's apartment. Yeah. So like right after Greg's apartment in the chocolate factory when it kicked off, right? It was like a billion dollars of committed funding from Elon. It felt like it was very solid.
另一个有趣的里程碑是你围绕 GPT-3 的训练搭建了很多工程。GPT-3 是怎样的?因为 GPT-2 用的是 TPU,对吧?
The other interesting milestone for you was when you got to build a lot of the engineering around the training for GPT-3. For GPT-3, how was that? Because you got from GPT-2 was in TPUs, right?
是的。
Yep.
GPT-3 的重大突破就是使用更多算力和 GPU。
And the big breakthrough in GPT-3 was like use more compute and using GPUs.
是的。我在 OpenAI 工作了一年,离开后去了 Google Brain 一年,然后回来,GPT-3 是在 2018 到 2019 年期间,正如你所说,就是扩大规模。我认为 Dario 基本上看到了缩放定律的大趋势。你为此发表了一篇论文。
Yep. So I ended up working at OpenAI for a year, left, went to Google Brain for a year, came back, and then GPT-3 was 2018 through 2019, building up to GPT-3, which exactly as you said was like scaling things up. I think that Dario had seen the big trend of scaling laws basically. You published a paper for that.
是的。那是一篇非常重要的论文,经受住了时间的考验,我们现在正生活在它的梦想中。明确看到那条线:用正确的配方投入更多算力,就能可靠地获得更多智能,这是主要的事情。至少对我来说,这正在发生,因为你可以看到当时我们在训练任务上花的钱并不多,但你能看到 Scaling 的存在。然后 Danny Hernandez 当时发表了一篇论文,展示了算法效率如何随时间推移让成本大幅降低。这两件事叠加在一起,就像,哦哇,我们将在未来几年获得更多的智能。
Yeah. And that's a pretty important paper that has withstood the test of time and we're living the dream of it now. Definitely seeing that line of reliably you get more intelligence if you spend more compute with the right recipe was the main thing. At least for me, it was like this is happening now because you could look at the time we weren't spending very much money on the training jobs and you could see that there was scaling there. And then Danny Hernandez did a paper at the time that showed how much cheaper algorithmic efficiency was making stuff over time too. Those two things stack together, that was like, oh wow, we're going to get a lot more intelligence over the next few years.
所以,当你看到它时,它是值得注意且令人惊讶的。
So, it was noteworthy and surprising when you saw it.
是的。我觉得最奇怪的是,我不是物理学家,但所有这些物理学家都在做这个。最初的缩放定律论文就是一条跨越 12 个数量级的直线。我就想,12 个数量级是个愚蠢的大数字。我从未见过任何东西跨越 12 个数量级。这说服我彻底将我的所有工作转向 Scaling,而我之前并没有这么做。
Yeah. And I think the thing that seemed the weirdest to me is like I'm not a physicist, but all these physicists were doing this stuff. The original scaling laws paper just the very straight line over 12 orders of magnitude. I'm just like 12 orders of magnitude is a stupidly large amount. I've never seen anything go over 12 orders of magnitude. That convinced me to definitely pivot all of my work into scaling, which I hadn't been doing before.
我能问一个外行问题吗?是否可以说缩放定律可能出现在所有这些其他领域?是否有像 2、5、100、10000 个领域,缩放定律可能成立,而我们只是没有投资?
Can I ask a lay person question? Is it fair to say that the scaling law might show up in all of these other domains? Are there like 2, 5, 100, 10,000 domains where the scaling law could hold that we're just not investing into?
是的。我认为在物理学中,缩放定律无处不在,我当时并不知道。在物理学中,有一个完整的领域叫现象学,它基本上研究世界的各个方面,然后进行这类拟合,他们到处都发现这些幂律分布。这是我第一次在计算机科学相关的东西中看到,我觉得这很有趣也很令人惊讶。
Yeah. So I think in physics scaling laws hold all over the place, which I didn't know at the time. Within physics, there's a whole field called phenomenology that basically looks at various aspects of the world and then does those types of fits, and they find these power law distributions all over the place. This was like the first one I had ever seen in a computer science adjacent thing, which I think was interesting and surprising.
当时,人们对此很生气。他们觉得,你在 GPU 上砸钱,或者只是浪费钱。这非常浪费。
And at the time, people were mad about it. They were like, you're throwing money at GPUs or just wasting money. This is very wasteful.
是的。现在不同的人,但仍然有人对此不满。
Yes. Different people now, but still people mad about it.
是的。研究人员对此也很生气,觉得这不优雅,你只是在蛮干。就像小丑帽一样堆叠更多层,我认为 Anthropic 的口号是“做有效的蠢事”。这很明显就是那个有效的蠢事。
Yeah. The researchers were mad at that too, where it's like it's not elegant, you're just brute forcing it. The jester cap like stack more layers, which I think Anthropic's slogan is 'do the stupid thing that works.' That was a thing where this was very clearly the very stupid thing that works.
你能告诉我们你是如何与 Anthropic 一起收集最后一颗无限宝石的吗?因为世界上很少有人曾在 OpenAI、DeepMind 和 Anthropic 工作过,而你是从 GPT-3 团队分拆出来并创办 Anthropic 的成员之一。那次跳跃是怎样的?
Can you tell us how you ended up collecting the last infinity stone with Anthropic? Because there are very few people in the world that have worked at OpenAI, DeepMind, and Anthropic, and you were part of the team that spun off from GPT-3 and then started Anthropic. How was that jump?
那里有两个团队。一个是安全组织,一个是 Scaling 组织,这两个组织都向 Dario 和 Daniela 汇报。我认为我们合作得非常好。在 OpenAI 和 Anthropic,我觉得很棒的一点是我们的文化是,所有事情都在 Slack 上,100%的事情都在 Slack 上。而且都是公开频道,沟通很好。我认为那个团队也是把缩放定律最当真的团队,他们认为,好吧,这实际上将是变革性的。将会有一个交接,人类将在某个时刻将控制权交给变革性 AI,希望它们能与我们对齐,那将是一个顺利的过渡,但可能不会。风险极高。所以那个团队非常专注于如何确保这件事被足够认真地对待,并且我们建立了一个能够承担这一重量的机构。最终,这个核心团队离开并加入了 Anthropic。当时我完全不清楚这对世界来说是否正确。现在回想起来,这似乎是个好选择。我觉得当时很酷的一点是,我们刚开始时,看起来根本不会成功。OpenAI 有十亿美元和所有明星力量,而我们只有七个联合创始人在疫情期间试图构建一些东西,我们不知道是否一定能做出产品,或者产品会是什么样子。所以有趣的是,所有最初加入的人也都是为了使命而来。他们本可以去其他地方工作,获得更多声望、更多金钱。人们会知道他们在做什么等等。
There were two teams there. That was the safety org and the scaling org, the two orgs that reported into Dario and Daniela. I think we had just worked together extremely well. One thing I think was great both at OpenAI and at Anthropic was just we had a culture where everything is on Slack, 100% of things on Slack. And within that, all public channels, great communication. I think that group also was the group that took the scaling laws the most seriously, where it was like okay, this actually is going to be transformative. There's going to be a handoff where humanity will hand off control to transformative AI at some point, and hopefully they'll be aligned with us and that'll be a good transition that goes well, but it might not be. The stakes are incredibly high. So I think that group was very focused on how do we make sure that that's taken seriously enough and that we've built an institution that can handle the weight of that. That ended up being the core group that left to join Anthropic. And I think it wasn't clear at all to me that that was the right thing for the world at the time. In hindsight now, it seems like that was a good choice. I think what was kind of cool then too is when we started out, we didn't seem like we were going to be successful at all. OpenAI had a billion dollars and all of this star power, and we had seven co-founders in COVID trying to build something, and we didn't know if we were necessarily going to make a product or what the products would look like. So I think what was interesting from that too is that all of the initial people who joined were there for the mission too. They all could have worked somewhere else for more prestige, more money. People would have known what they were doing, etc.
嗯,基本上就是留在 OpenAI。
Well, stayed at that OpenAI basically.
正是。这是一件有趣的事情,我认为这是让我们的文化或组织得以扩展的关键。我们现在大约有 2000 人,但我们仍然有一种感觉,似乎没有政治渗透进来。我认为很大程度上是因为最初的一百人都是为了使命而来。所以如果有什么开始出错,他们会举手说,‘这个人似乎不是在为使命行事。’
Exactly. That's been an interesting thing that I think has been the key to letting our culture or our org scale. We're like 2,000 people now, but we still have a thing where it doesn't seem like politics have crept in. And I think a lot of that is the first hundred people all were just there for the mission. So if something starts to go wrong, they'll raise their hand and be like, 'It seems like this person might not be acting for the mission.'
你从 OpenAI 出来,心里有个大致的长远使命——不毁灭人类,但第一年你实际做了什么?又是如何最终聚焦到一个具体产品上的?
So you broke off from OpenAI, you had a general idea of the long-term mission to not destroy humanity, but what did you actually work on for the first year? How did that converge on an actual product?
第一年,我主要做的就是搭建训练模型所需的训练基础设施,然后获取训练模型所需的算力。这是两个主要项目。还有创业公司要做的所有其他事情,比如开设 Brex 账户等等。我们最初有七位联合创始人。几个月内,大概有 25 位来自 OpenAI 的人加入了我们。所以我们有一个相当庞大的团队,而且大家已经知道如何协作,这让我们能更快地起步。
The first year, the main thing I tried to do was just build the training infrastructure we needed to train a model and then get the compute we needed to train the model. Those were my two main projects. Plus all the other things you need to do when starting up a company, like set up a Brex account and all that stuff. We started out with seven co-founders. Within a few months, I think 25 folks from OpenAI overall had joined. So we had a pretty substantial team that already knew how to work together, which helped us get up and running faster.
你们是什么时候推出第一个产品的?事情又是从什么时候开始真正运转起来的?
At what point did you launch the first product and when did things begin to actually start working?
我们推出的第一个产品是在 ChatGPT 之后。在 ChatGPT 之前大概九个月,我们有一个 Claude 1 的 Slackbot 版本。
The first product we launched was after ChatGPT. We had maybe nine months before ChatGPT a Slackbot version of Claude 1.
哦对,我们在 YC 的 Slack 里确实用过那个。
Oh yeah, we had that in the YC Slack actually.
是啊,真的很酷。
Yeah. It was really cool.
我记得 Tom Blfield 把你们所有人都加进去了。
I remember Tom Blfield adding all of you guys to it.
当时,我们不确定是否要把它作为产品推出。我们不知道这样做对世界是否有益。我们并没有真正想清楚我们的影响力理论,即如何让事情真正运作良好。另外,事后看来,如果我们当时试图推出它,我们也没有相应的服务基础设施。而且因为我们不确定是否要推出,我们在建设基础设施上犹豫了太久,这对我来说是个教训。
At the time, we didn't know whether we wanted to launch it as a product. We didn't know if doing so would be good for the world. We hadn't really thought through our theory of impact much for how we would actually make stuff work well. Plus, in hindsight, if we had tried to launch it, we wouldn't have had the serving infrastructure to do it. And because we weren't sure whether we wanted to, we hesitated too long on building that infrastructure, which is a learning for me.
那时 ChatGPT 还没推出。所以我想我们也不知道它会成为一件大事。
At this time, ChatGPT had not launched yet. So I guess we didn't know that it would be a big deal either.
那是在疫情期间,2022 年夏天。然后 ChatGPT 在 2022 年秋季推出,之后我们重新推出了 API,再之后推出了 Claude AI。基本上直到 Claude 3.5 和编程能力出现之前,一切看起来都不太顺利。我觉得在那整个时期,直到大约一年前,我们是否最终能成为一家成功的公司都还不明朗。
This is around the pandemic, summer 2022. Then ChatGPT launched fall 2022, and we relaunched our API after that, and then Claude AI after that also. It didn't seem like it was working basically until Claude 3.5 and coding. I think through that whole time until about a year ago, it seemed like it wasn't clear that we were going to end up being a successful company.
我们在创业公司中看到了这一点,因为我们能感受到哪种模型更受欢迎。整个 2023 年,OpenAI 是首选。然后情况在 2024 年开始转变,我们看到 Claude 3.5 Sonnet 在 YC 批次中开始获得市场份额,从个位数增长到 20-30%,尤其是在编程方面。它成了默认选择,这非常有趣。你能谈谈这种涌现行为以及在该特定技能上的突出表现吗?现在编程方面肯定有 80% 或 90% 了,尤其是 Claude Code。这是有意为之还是自然而然发生的?
We saw that in startups because we get a vibe check on the preferred model. All of 2023, OpenAI was the response. Then things started to turn in 2024 when we saw Claude 3.5 Sonnet starting to gain market share in YC batches, going from single digit to 20-30%, especially for coding. It became the default choice, which was very interesting. Can you tell us about that emergent behavior and the spikiness on that particular skill? It must be 80% or 90% now for coding, especially Claude Code. Was that on purpose or just kind of happened?
我想我们投入了更多精力让模型在代码方面变得非常出色,因为我们希望模型擅长编程。这是一方面。然后看到大家的反应,我们就想,好吧,那就在这方面更加努力。
I think we invested more in trying to make the model really good at code because we wanted the model to be good at code. That was one thing. And then seeing the reaction of everyone to it was like, okay, let's go much harder on that also.
所以在 3.5 Sonnet 之前,你已经在编程上投入了足够多,意识到这很有前景,于是决定加倍投入。
So before 3.5 Sonnet, you'd already invested enough in coding to realize that was really promising and you decided to double down.
我想这实际上是组织内部的个人在 3.5 Sonnet 之前就表示想做编程,当我们看到 3.5 Sonnet 非常好的产品市场契合度时,那就是一个很好的信号,可以朝这个方向前进。
I think this really was like individuals within the org being like we want to do coding before 3.5 Sonnet, and when we saw 3.5 Sonnet's really good product-market fit, that was good signal to go for that.
在你们推出 3.5 Sonnet 的那天,你们知道手里有特别的东西、这会是转折点吗?还是像 OpenAI 推出 ChatGPT 时一样惊讶,没想到它会突然火爆?
On the day you launched 3.5 Sonnet, did you know you had something really special and this would be the turning point, or were you as surprised as OpenAI when they launched ChatGPT and it unexpectedly took off?
我希望我们当时更有远见,但没有,我想我们也很惊讶它竟然如此重要。然后 3.7 Sonnet 也让我们惊讶,它解锁了那么多智能体式编程的能力。对于这些事情,我们推出得很快,所以通常不知道结果会怎样。
I wish we had more foresight on that, but no, I think it was surprising for us too how big of a deal it was. And then 3.7 Sonnet also surprised us by how much it unlocked agentic coding. For each of these things, we move quite fast in rolling them out, so we often don't know what the results are going to be.
我想正是这一点让很多编程智能体创业公司成功了。有个疯狂的故事,Replit 在短短 10 个月内收入达到 1 亿美元,当然还有 Cursor,都是建立在 Sonnet 之上的。
I think it's what made a lot of these coding agent startups work. There's a crazy story of Replit going to $100 million in just 10 months, and Cursor of course, all built on Sonnet.
所有这些事情都让我感到惊讶。另外,在我使用 Claude 的过程中,我不断被它能做的事情所惊讶。每一次都有更多的东西被解锁。我一个朋友告诉我,她想修改一个源代码工具,但没有源代码,只有编译后的二进制文件。她问 Claude:‘你能反编译这个吗?你能反汇编这个汇编代码吗?’Claude 琢磨了 10 分钟,然后生成了一个 C 语言版本。于是她就有了那个东西,这太疯狂了。她说如果她自己花三天时间,可能也能搞出十六进制表并写点代码,但 Claude 完成了所有工作,还起了变量名等等。所以我认为我们不断被模型记住的东西所惊讶,比如所有的十六进制表,它能思考并解决它。我想我们还会继续被这类事情惊讶。
All those things have been surprising to me. Also, in my working with Claude, I continue to be surprised by the type of stuff it can do. With each one, there's more stuff that unlocks. One of my friends told me she had some source tool she wanted to modify but didn't have the source code, only the compiled binary. She asked Claude, 'Can you decompile this? Can you disassemble the assembly?' Claude chewed on it for 10 minutes and made a C version of it. So she had the thing, which is insane. She said if she spent 3 days on it, she probably could have gotten the hex tables and written a little code, but Claude did the whole thing, made up variable names, etc. So I think we keep getting surprised by stuff the model has memorized, like all the hex tables, and it can think through and work through it. I think we're going to continue to be surprised by that sort of stuff.
如果你调查 YC 创始人,他们压倒性地偏好使用 Anthropic 模型进行编程,这个比例远高于仅看基准测试结果所预测的。所以似乎存在某种 X 因素,让人们非常喜欢用这些模型编程。你知道这是什么吗?是有意为之,还是从黑箱中冒出来的?
If you poll YC founders, they prefer using Anthropic models for coding by a huge margin, much larger than what you would predict from benchmark results. So there seems to be some X factor that makes people really like these models for coding. Do you know what it is, and is it intentional or did it come out of the black box somehow?
我认为基准测试很容易被操纵。其他所有大型实验室都有团队,他们的全部工作就是让基准测试分数好看。我们没有这样的团队。所以我认为这可能是最大的因素。
I think benchmarks are easy to game. All the other big labs have teams whose whole job is to make benchmark scores good. We don't have such a team. So I think that is probably the biggest factor.
你们不教学生应付考试。
You don't teach to the test.
我们不为了考试而教学,因为我确实觉得如果开始那样做,就会产生奇怪的负面激励。也许我们可以把那个团队放在市场部下面,然后忽略所有基准测试。但我认为这是导致训练-测试不匹配的原因之一。
We don't teach to the test because I do feel like if you start doing that then it has weird bad incentives. Maybe we could put that team under marketing or something like that and then ignore all the benchmarks. But I think that's one reason why there's some train-test mismatch there.
所以评估在内部更偏定性,或者你们有内部的……
So the evaluations are more qualitative internally, or you have your internal...
我们有内部基准测试。是的。但我们不公开它们。
We have internal benchmarks. Yeah. But we don't publish them.
团队真正专注于改进的是内部基准测试吗?
And is it the internal benchmarks that the teams are really focused on improving?
没错。是的。所以我们有团队专注于改进的内部基准测试,然后我们还有一堆任务。我认为加速我们自己的工程师也是我们的首要任务,所以我们在这方面做了大量的内部试用,以确保它也能帮助我们的员工。
That's right. Yeah. So we have internal benchmarks that the team focuses on improving, and then we also have a bunch of tasks. I think accelerating our own engineers is a top priority for us too, so we do a ton of dogfooding there to make sure it's helping our folks too.
回到金门大桥克劳德,可解释性似乎是其中很重要的一部分。然后大多数人会说克劳德的个性感觉更好。那么你如何既非常量化,又围绕个性构建评估呢?
Going back to Golden Gate Claude, there's a lot of interpretability seems like it's a big part of it. And then most people would say that Claude's personality just feels better. And then how do you at once be very quantitative, but also build evals around personality?
个性的评估也很复杂,比如你怎么判断克劳德是否有一颗善良的心之类的?很难知道。但我认为那是阿曼达·阿斯克尔团队的任务。她描述为像一个好的世界旅行者,克劳德去和来自不同背景的各种人交谈,每个人都应该觉得'我对这次谈话感觉很好'。可解释性确实是一个长期赌注,对吧?现在模型没那么可怕,但总有一天它们会更可怕,所以我认为希望是当情况变得更严峻时,有能力知道引擎盖下到底发生了什么。
The evals for personality are kind of complicated too, like how do you tell if Claude has a good heart or something like that? It's hard to know. But I do think that's Amanda Askell's team's mandate. I think she describes it as being like a good world traveler where Claude goes and talks with all sorts of people from different backgrounds, and each of the people should come away feeling like 'I feel good about this conversation I've had.' Interpretability really is a long-term bet, right? Right now the models aren't that scary, but at some point they're going to be more scary, and so I think the hope is to have some ability to know what's actually going on under the hood when it becomes more intense.
然后最近,克劳德代码取得了真正的成功。你能跟我们说说这个项目在内部是如何启动的吗?还有,你是这次就知道它会成功,还是它是个惊喜?
Then more recently, Claude Code's been a real success. Can you talk us through how that project got started internally? And again, was it like you knew this time it was going to work or was it a surprise?
克劳德代码也是一个内部工具,用来帮助 Anthropic 内部的工程师。鲍里斯把它拼凑起来的。
Claude Code was an internal tool also, to help out our engineers within Anthropic. Boris had hacked it together.
有一个 Anthropic 的内部工程师想为自己构建它。
There's an internal Anthropic engineer wanting to build it for themselves.
为了他和其他的内部工程师。是的。然后我认为我们绝对不知道它会在外部成功。我想在某种程度上,在那之前我们完全押注在 API 上,意图是外面有那么多初创公司有那么多好主意。我们是谁来弄清楚在上面构建什么正确的产品?外面每个人都会比我们构建更好的东西。所以把所有的精力都放在打造最好的 API 上。我认为这让我很惊讶,就像,好吧,我们实际上能够做出一个产品,在智能体式使用上比市场上的其他产品更好。我有个理论,部分原因来自于一种思维转变,把克劳德也看作这个产品的用户。对于 link,我们试图为教师构建东西,对于 Grouper,主要是纽约的单身人士。对于这个,我认为用户是开发者,但我也认为用户是克劳德。就像给克劳德正确的工具,这样克劳德才能真正有效地工作,帮助克劳德获得正确的上下文来有效工作。这个团队最关注克劳德作为用户,我认为这说得通,你们最了解克劳德。
For him and other internal engineers. Yeah. And then I think we definitely didn't know that it would be successful out there. I think to some degree we really had fully just bet on the API before that, with the intention being there are so many startups out there with so many good ideas. Who are we to figure out what the right product is to build on top of this stuff? Everyone out there is going to build better stuff than us. So put all of our effort into just making the best possible API. And I think this surprised me, like okay, we actually were able to make something that as a product was better than the other products out on the market for this agentic use. I have some theory that part of that came from a mindset shift of seeing Claude as the user for this thing too. For link, we were trying to build things for teachers, for Grouper it was single people in New York mostly. For this, I think really the users are the developers, but also I think the user is Claude. It's like give Claude the right tools so that Claude can actually do that effectively, help Claude get the right context to work effectively. This team was the most focused on Claude as a user, which I think makes sense that you guys would understand Claude the best.
这是对 LLM 本身的完美拟人化。智能体是利益相关者之一,是你会去争取并试图赋能的用户之一。
That's the perfect anthropomorphization of the LLM itself. The agent is one of the stakeholders, one of the users that you would go after and try to empower.
是的。完全正确。这实际上很有道理,为什么你们真的让 MCP 工作起来做工具调用,因为其他一堆实验室也尝试过,但最终流行起来的标准是你们的。
Yeah. Totally. Which actually makes a lot of sense why you guys actually got MCP to work to do tool calling, because a bunch of other labs had tried to do something and the standard that stuck that really took off was yours.
是的,我认为那也是一个类似的,以模型为中心。
Yeah, I think that seems like a similar one too, where it's model-focused.
回到克劳德代码,成功真的很令人兴奋。对于 Cursor 和其他在 API 之上构建的公司来说,这也令人害怕。你对构建产品的创始人有什么建议?他们应该如何考虑在 API 上构建,同时又担心 Anthropic 或某个实验室构建出比他们更好的东西?
Going back to Claude Code, success is really exciting. It's also scary for Cursor and other companies that have built on top of the API. What's your advice to founders building products? How should they think about building on the API but also worrying about Anthropic or one of the labs building something better than they can build?
我想我有点惊讶克劳德代码,我们确实构建了一个在市场上也是最好的东西。我不太清楚除了对克劳德有更多同理心之外,我们在克劳德代码上的巨大优势是什么。
I think I was kind of surprised that Claude Code, we did build a thing that was like the best in the market there too. It's not super clear to me what the big advantage was for us for Claude Code besides more empathy for Claude or something.
这实际上是一个非常有趣的见解。似乎你是在为一个你非常了解的特定用户构建,而其他人不会想到去构建,而不是你有某种内在的技术优势。
That's actually a really interesting insight. It seems like the thing you were building for a specific user that you knew really well that other people wouldn't have thought to build for, versus you had some intrinsic technology advantage.
是的。就像我认为一个初创公司也可以做同样的事情,对吧?
Yeah. Like I think a startup could have done that same thing too, right?
是的。
Yeah.
我认为我们是最以开发者为中心的实验室。我认为我们也是最以 API 为中心的实验室。所以我想确保我们拥有最好的平台让人们构建东西,因为这东西增长得非常快。我们不会是找出所有需要赋能克劳德的方式最快的,这些方式连接克劳德到整个人类商业。人类世界都是为人类设计的,但我们需要让模型能够成为经济中富有成效的成员。
I think we're the most developer-focused lab. I think we're the most API-focused lab too. So I think we want to make sure that we have the best platform for people to build stuff on, because this thing is growing so incredibly quickly. We're not going to be the fastest at figuring out all the ways that we need to empower Claude to do the work that connects Claude to the entire human business. The human world is all designed for humans, but we need to get the models to be able to be productive members of the economy.
有没有你希望看到开发者构建的想法或领域,或者你认为目前被低估的领域?
Are there ideas or areas you would love to see developers building in, or areas you think are underappreciated right now?
是的,克劳德代码就像如何让克劳德成为一个有用的结对程序员,或者像一个初级工程师。你有一个像二级或三级这样的可以一起工作,或者非常尖峰,因为它可以做那些超级高级套件难以处理的奇怪拆解工作。不太擅长知道做什么类型的工作。需要很多指导。需要很多上下文。
Yeah, Claude Code is like how do you get Claude to be a useful pair programmer, or like a junior engineer. You've got like a level two or three or something like that that you can work with, or very spiky because it can do the weird disassembly stuff that a super high-level suite would struggle with. Less good at knowing what type of work to do. Needs kind of a lot of handholding. Needs a lot of context from it.
那只是工作中一个非常特定的子集。如果你看看企业中除此之外发生的所有事情,那只是企业所有工作中很小的一部分,一个聪明、会编码、会用很多工具但还没有太多背景知识的人会想做的。所以我认为,找到方法指导 Claude 或其他模型为企业做有用任务,似乎有巨大的空间。
That's like one very particular subset of work that can be done. If you look at all the stuff that happens in businesses besides that, it's a very tiny fraction of all the work that's done in businesses that a smart person who knows how to code and use lots of tools but doesn't have that much context yet would want to do. So I think finding ways to coach Claude or whatever model to do useful tasks for businesses seems like there's just a huge space there.
Tom,你的工作很大一部分是负责让 Anthropic 运转的所有计算基础设施。你能谈谈这个庞然大物背后的计算基础设施是什么吗?一个有趣的现象是,人类正朝着有史以来最大的基础设施建设迈进。这将比阿波罗计划、曼哈顿计划还要大。如果保持目前的轨迹,明年就会超过这两者,也就是 AGI 算力支出每年大约增长 3 倍,这太疯狂了。
So Tom, a big part of your job is owning all the compute infrastructure that makes Anthropic work. Can you talk about what the compute infrastructure is behind this giant thing? One thing that's interesting to look at is that humanity is on track for the largest infrastructure buildout of all time. This is going to be larger than the Apollo project, larger than the Manhattan project. It'll be bigger than both of them next year if it keeps on the current trajectory, which is roughly 3x per year increase in spending on AGI compute, which is just bonkers.
是的,每年 3 倍很疯狂。我认为它会保持每年 3 倍的轨迹。明年的已经确定了,2027 年还有点不确定。
Yeah, 3x per year is wild. I think it's going to keep up on the 3x per year trajectory. It's already locked in for next year and then it's a little bit open for 2027.
据我所知,在 YC 内部,我们无法获得足够的包括 Claude 在内的所有顶级前沿模型的额度。所以你得帮我们一点。
I mean anecdotally internal to YC, we can't get enough credits across all of the top frontier models including Claude. So you got to help us out a little bit.
是的。
Yeah.
我们就是,每个人都是瓶颈,就像给我更多智能。我永远不够。
We're just, I mean everyone's bottlenecked literally, it's like give me more intelligence. I can't have enough.
是的。我知道你们也在看更多的硬件初创公司,寻找更多加速器。我认为到 2027 年我们会看到更多加速器上线。那是个好领域。另外,数据中心技术我认为也很重要。
Yeah. And I know you guys have been looking at more hardware startups also for more accelerators. I think that we will see more accelerators coming online by 2027. That's a good space. Also data center tech I think is a big one.
你们现在的瓶颈在哪里?是获得足够的电力、足够的 GPU,还是施工许可?
Where are the bottlenecks for you guys now? Is it like getting enough electricity, getting enough GPUs, getting construction permits?
电力,人们用喷气发动机来获取电力。这太疯狂了。
Power, people are using jet engines to get power. That's nuts.
总体而言,我认为电力将是最大的瓶颈,尤其是在美国。我们想在美国建设。我们最大的政策目标之一就是让美国建设更多数据中心,批准更多数据中心,让建设更容易。
Overall for the buildout, I think power is going to be the biggest bottleneck, especially power in the US. Like we want to build in the US. That's one of our biggest policy goals is to get the US to build more data centers, permit more data centers, make it easier to build.
答案是可再生能源还是核能?
Is the answer renewables or is it nuclear?
我绝对觉得是的,是的,所有这些。我希望核能更容易建设。
I definitely feel like yes, yes, all of those things. I wish that nuclear was easier to build.
Anthropic 是唯一一个不只使用一种 GPU,而是使用来自三个不同制造商的 GPU 的主要实验室。你能谈谈这个以及这个策略的效果吗?
Anthropic is the only major lab that uses not just one kind of GPU, but the GPUs from three different manufacturers. Can you talk about that and how that strategy has played out?
是的。所以我们使用 GPU、TPU 和 Tranium。这样做的缺点是,我们的性能工程团队分散在所有平台上,这带来了大量额外工作。好处是它给了我们灵活性,一方面可以吸收额外的容量,因为所有这些加起来比单一平台更多,另一方面,我们可以为合适的工作使用合适的芯片,有些芯片更适合推理,有些更适合训练,我们可以将合适的芯片匹配到合适的工作。所以,我认为这就是权衡。
Yeah. So we use GPUs, TPUs, and Tranium. The downside of doing that is that we split our performance engineering teams across all of those platforms, which is a ton of extra work. The positive thing is it gives us the flexibility to both soak up that extra capacity because there is more of those altogether than just one, and then two, we can use the right chips for the right jobs where some chips will be better for inference, some chips will be better for training, and we can match the right chips to the right jobs. So yeah, I think that's the trade-off there.
我觉得很酷的一点是,回顾你的职业生涯,这一切是如何叠加的:你曾是 OpenAI 那个将架构从 TPU 改为 GPU 的工程师,让 GPT-3 真正实现了 Scaling(规模扩张),而现在多年后,你在一个更大得多的规模上负责这件事。我不知道这对你来说是否算是串联起来了。
I guess one cool thing is just connecting the dots through your career and how all of this compounded because you were the one engineer building that change of the architecture from TPUs to GPUs back at OpenAI that got GPT-3 to actually scale, and now you're in charge of that at a much, much bigger scale years later. I don't know if that kind of connected dots for you.
OpenAI 从 TPU 转向 GPU,部分原因我认为是 PyTorch 在 GPU 上的软件栈比 TensorFlow 在 TPU 上更好。我认为这解锁了快速迭代,如果你有一个好的可靠的软件栈,你就可以快速实验,构建一个完整的系统。我认为这也是我们现在在 Anthropic 努力追求的。拥有更多平台的挑战是编写所有好软件更难。我认为培养构建好软件的能力,让所有在底层之上构建的人都能有良好体验,这是最重要的。
The big move from TPUs to GPUs at OpenAI I think was partly driven just that PyTorch was a better software stack on top of them than TensorFlow on top of TPUs. And I think that then unlocked fast iteration where if you have a good reliable software stack, then you can experiment quickly, just build a whole system that works. I think that's the thing that we really strive for now at Anthropic too. The challenge of having many more platforms is that it's harder to write all the good software. I think building the muscle of knowing how to build that software well so that all of the people who build on top of that low level can have a great experience with it is the most important thing there.
你对年轻时的自己有什么建议吗?你经历了这段疯狂的旅程。如果有人像你 20 多岁那样生活在今天,想要乘上并加入 AI 革命,你会对他们说什么?特别具体的是,我们目前从很多大学生那里听到的是,他们不知道是否应该留在大学,是否会有工作,世界将如何变化,他们应该做什么?
Do you have advice for a younger version of yourself who now you've seen and went through this crazy journey? If someone was you back in the 20s living today and they wanted to ride and join the AI revolution, what would you say to them? And very specifically something we hear from a lot of college students at the moment is they don't know if they should stay in college, are there going to be jobs for them, how is the world going to change and what should they do?
我认为承担更多风险是明智的,还要尝试做那些你的朋友会非常兴奋和印象深刻的事情,或者一个更理想化的自己如果成功会非常自豪的事情。我想这大概就是我会告诉年轻自己的话。
Taking more risks I think is wise, and also trying to work on stuff where your friends would be really excited and impressed if you did it, or a more idealized version of yourself would be really proud of yourself if you succeeded at it. I think that's probably the thing that I would try to tell a younger version of myself.
更内在,更少外在。不要追逐那些其他证书和学位,或者你知道的,在 FAANG 工作,这些在今天都无关紧要。
More intrinsic, less extrinsic. Don't chase these other credentials and getting the degree or whatever, you know, working at FAANG, those are just irrelevant as of today.
是的,没错。
Yeah, exactly.
这就是我们今天的所有时间。下次见。
That's all we have time for today. We'll see you guys next time.