从零到十亿:Surge AI 的传奇故事

From Zero to a Billion: The Surge AI Story

陈埃德温 Edwin Chen · Lenny 播客 · 2025-12-07 · 约 71 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Surge AI 创始人 Edwin Chen 揭秘,他如何以 60-70 人的自举团队,在不到四年内实现十亿美元营收,重新定义 AI 数据质量,并挑战硅谷传统。

Edwin Chen, founder of Surge AI, reveals how his bootstrapped team of 60-70 people hit $1B revenue in under 4 years by redefining AI data quality and challenging Silicon Valley norms.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 32)

全文 · Full transcript(中英对照)

引言与成就 Introduction and Achievements

Host

你们在不到四年的时间里,用大约 60 到 70 人实现了 10 亿美元营收。你们完全自力更生,没有融过一分钱风险投资。我相信以前从来没有人做到过这一点。

You guys hit a billion in revenue in less than four years with around 60 to 70 people. You were completely bootstrapped. Haven't raised any VC money. I don't believe anyone has ever done this before.

Edwin

我们基本上从来不想玩硅谷那套游戏。我一直觉得那很荒谬。我曾在多家大型科技公司工作,总觉得我们可以裁掉 90% 的人,反而会跑得更快,因为最优秀的人不会被各种琐事分心。所以当我们创办 Surge 时,就想用完全不同的方式,打造一支超级小、超级精英的团队。

We basically never wanted to play the Silicon Valley game. I always thought it was ridiculous. I used to work at a bunch of big tech companies and I always felt that we could fire 90% of people and we would move faster because the best people wouldn't have all these distractions. So when we started Surge, we wanted to build it completely differently with a super small, super elite team.

Host

你们绝对是目前最成功的数据公司。

You guys are by far the most successful data company out there.

Edwin

我们本质上是在教 AI 模型分辨好坏。人们根本不理解在这个领域里质量意味着什么。他们以为只要堆人力就能得到好数据,这完全错了。

We essentially teach AI models what's good and what's bad. People don't understand what quality even means in the space. They think you can just throw bodies at a problem and get good data. That's completely wrong.

Host

对普通人来说,感觉不到这些模型在持续变聪明多少。

To a regular person, it doesn't feel like these models are getting that much smarter constantly.

Edwin

过去一年里,我意识到公司的价值观会塑造模型。前几天我让 Claude 帮我起草一封邮件,30 分钟后,它确实帮我写好了完美的邮件,我发了出去。然后我意识到,我花了 30 分钟做了一件根本不重要的事。

Over the past year, I've realized that the values that the companies have will shape the models. I was asking Claude to help me draft an email the other day and after 30 minutes, yeah, I think it really crafted me the perfect email and I sent it. Then I realized I spent 30 minutes doing something that didn't matter at all.

Host

如果你能选择完美的模型行为,你会想要哪种模型?是那种说“你说得对,这封邮件肯定还有 20 种改进方式,然后继续迭代 50 轮”的模型,还是那种优化你的时间和生产力,直接说“不,你得停下来了。你的邮件很好,直接发出去,继续做别的事”的模型?

If you could choose the perfect model behavior, which model would you want? Do you want a model that says, "You're absolutely right. There are definitely 20 more ways to improve this email and it continues for 50 more iterations." Or do you want a model that's optimizing for your time and productivity and just says, "No, you need to stop. Your email is great. Just send it and move on."

Host

你有个尖锐的观点,认为很多实验室在把 AGI 推向错误的方向。我担心的是,我们不是在构建真正能推动人类进步、治愈癌症、解决贫困、理解宇宙的 AI,而是在优化 AI 垃圾内容。简直是在把模型优化成迎合那些在杂货店买小报的人。我们基本上是在教模型追逐多巴胺,而不是追求真理。

You have this hot take that a lot of these labs are pushing AGI in the wrong direction. I'm worried that instead of building AI that will actually advance us as a species, curing cancer, solving poverty, understanding universe, we are optimizing for AI slop instead. Literally optimizing our models for the types of people who buy tabloids at the grocery store. We're basically teaching our models to chase dopamine instead of truth.

Host

今天我的嘉宾是 Edwin Chen,Surge AI 的创始人兼 CEO。Edwin 是一位非凡的 CEO,Surge 也是一家非凡的公司。他们是领先的 AI 数据公司,为每一个前沿 AI 实验室提供训练支持。他们也是史上最快达到 10 亿美元营收的公司,成立仅 4 年,员工不到 100 人,而且完全自力更生,从未融过一分钱风险投资,从第一天起就实现盈利。正如你将在对话中听到的,Edwin 对于如何打造一家重要的公司、如何构建真正对人类有益且有用的 AI,有着非常不同的见解。我非常喜欢这次对话,也学到了很多。我很期待你能听到。如果你喜欢这个播客,别忘了在你最喜欢的播客应用或 YouTube 上订阅和关注,这对我帮助巨大。如果你成为我通讯的年度订阅者,你将免费获得一整年的大量优秀产品,包括 Devon、Lovable、Replet、Bolt、NAM、Linear、Superhum、Dcript、Whisper Flow、Gamma、Perplexity、Warp、Granola、Magic Patterns、Ray、Catch、Shepardd、Mobin、Post Hog 和 Stripe Atlas。请访问 Lenny's.com 并点击产品通行证。接下来,在简短赞助商广告之后,有请 Edwin Chen。

Today, my guest is Edwin Chen, founder and CEO of Surge AI. Edwin is an extraordinary CEO and Surge is an extraordinary company. They're the leading AI data company powering training at every Frontier AI lab. They're also the fastest company to ever hit $1 billion in revenue in just 4 years after launch with fewer than 100 people and also completely bootstrapped. They've never raised a dollar in VC money. They've also been profitable from day one. As you'll hear in this conversation, Edwin has a very different take on how to build an important company and how to build AI that is truly good and useful to humanity. I absolutely love this conversation and I learned a ton. I am really excited for you to hear it. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. It helps tremendously. And if you become an annual subscriber of my newsletter, you get a ton of incredible products for free for an entire year, including Devon, Lovable, Replet, Bolt, NAM, Linear, Superhum, Dcript, Whisper Flow, Gamma, Perplexity, Warp, Granola, Magic Patterns, Ray, Catch, Shepardd, Mobin, Post Hog, and Stripe Atlas. Head on over to Lenny's.com and click product pass. With that, I bring you Edwin Chen after a short word from our sponsors.

Host

我和播客嘉宾都喜欢聊工艺、品味、自主权和产品市场契合度。你知道我们不喜欢聊什么吗?SOC 2。这正是 Vanta 的用武之地。Vanta 帮助各种规模的公司快速合规并保持合规,凭借行业领先的 AI 自动化和持续监控。无论你是初创公司应对首个 SOC 2 或 ISO 27001,还是企业管理供应商风险,Vanta 的信任管理平台都能让过程更快、更简单、更具可扩展性。Vanta 还能帮你将安全问卷完成速度提升五倍,让你更快赢得更大的交易。根据最近一项 IDC 研究,Vanta 客户每年节省超过 50 万美元,生产力提高三倍。建立信任不是可选项,Vanta 让它自动化。访问 vanta.com/lenny 可立减 1000 美元。

My podcast guests and I love talking about craft and taste and agency and product market fit. You know what we don't love talking about? Sock 2. That's where Vanta comes in. Vanta helps companies of all sizes get compliant fast and stay that way with industryleading AI automation and continuous monitoring. Whether you're a startup tackling your first Sock 2 or ISO 2701 or an enterprise managing vendor risk, Vanta's trust management platform makes it quicker, easier, and more scalable. Vanta also helps you complete security questionnaires up to five times faster so that you can win bigger deals sooner. The result, according to a recent IDC study, Vant customers slashed over $500,000 a year and are three times more productive. Establishing trust isn't optional. Vanta makes it automatic. Get $1,000 off at vanta.com/lenny.

Host

这里有个谜题。OpenAI、Cursor、Perplexity、Vercel、Platt 以及数百家其他成功公司有什么共同点?答案是它们都由今天的赞助商 WorkOS 提供支持。如果你正在为企业构建软件,你可能已经体会过集成单点登录、SCIM、审计日志以及其他大客户要求的功能的痛苦。WorkOS 将这些交易障碍转化为即插即用的 API,并提供一个专为 B2B SaaS 打造的现代开发者平台。无论你是种子期初创公司试图拿下第一个企业客户,还是独角兽公司全球扩张,WorkOS 都是实现企业就绪和释放增长的最快路径。它们本质上就是企业功能的 Stripe。访问 workos.com 开始使用,或者直接联系他们的 Slack 支持,那里有真正的工程师,能非常快速地回答你的问题。WorkOS 让你能够像最优秀的公司一样构建,提供愉悦的 API、全面的文档和流畅的开发者体验。今天就访问 works.com,让你的应用企业就绪。

Here's a puzzle for you. What do OpenAI, Cursor, Perplexity, Vercel, Platt, and hundreds of other winning companies have in common? The answer is they're all powered by today's sponsor, WorkOS. If you're building software for enterprises, you've probably felt the pain of integrating single signon, skim, arbback, audit logs, and other features required by big customers. Work OS turns those deal blockers into drop-in APIs with a modern developer platform built specifically for B2B SAS. Whether you're a seedstage startup trying to land your first enterprise customer or a unicorn expanding globally, work OS is the fastest path to becoming enterprise ready and unlocking growth. They're essentially Stripe for enterprise features. Visit workos.com to get started or just hit up their Slack support where they have real engineers in there who answer your questions super fast. Workos allows you to build like the best with delightful APIs, comprehensive docs, and a smooth developer experience. Go to works.com to make your app enterprise ready today.

Host

Edwin,非常感谢你来到这里,欢迎来到播客。

Edwin, thank you so much for being here and welcome to the podcast.

Edwin

非常感谢邀请我,我非常兴奋。

Thanks so much for having me. I'm super excited.

Host

我想从你们所取得的成就有多么惊人开始。很多人和很多公司都在谈论借助 AI 用极少的人扩展大规模业务,而你们以一种前所未有的方式做到了。你们在不到四年的时间里,用不到 60 到 70 人实现了 10 亿美元营收。你们完全自力更生,没有融过一分钱风险投资。我相信以前从来没有人做到过这一点。所以你们实际上正在实现人们所描述的 AI 将带来的梦想。我很好奇,你认为这种情况会因 AI 而越来越多吗?另外,AI 在哪些方面最帮助你们找到了杠杆,从而做到这一点?

I want to start with just how absurd what you've achieved is. A lot of people and a lot of companies talk about scaling massive businesses with very few people as a result of AI and you guys have done this in a way that is is unprecedented. You guys hit a billion in revenue in less than four years with less than 60 around 60 to 70 people. You're completely bootstrapped. Haven't raised any VC money. I don't believe anyone has ever done this before. So you guys are actually achieving the dream of what people are describing will happen with AI. I'm curious just do you think this will happen more and more as a result of AI? And also just where has AI most helped you uh find leverage to be able to do this?

Edwin

是的。去年我们以不到 100 人实现了超过 10 亿美元的营收。我认为未来几年我们会看到比例更疯狂的公司,比如每位员工创造 1 亿美元营收。AI 只会越来越好,让事情变得更高效。所以这个比例是不可避免的。就像我曾在多家大型科技公司工作,总觉得我们可以裁掉 90% 的人,反而会跑得更快,因为最优秀的人不会被各种琐事分心。所以当我们创办 Surge 时,就想用完全不同的方式,打造一支超级小、超级精英的团队。疯狂的是,我们真的成功了。所以我认为有两件事正在碰撞。一是人们开始意识到,你不必建立庞大的组织才能获胜。

Yeah. So, we hit over a billion of revenue last year with under 100 people. And I think we're going to see companies with even crazier ratios, like 100 million per employee in the next few years. AI is just going to get better and better and make things more efficient. So, that ratio just becomes inevitable. Like I I used to work at a bunch of the big tech companies and I always felt that we could fire 90% of people and we would move faster because the best people would have all these distractions. And so when we started Surge, we wanted to build it completely differently with a super small, super elite team. And yeah, what's crazy is that we actually succeeded. And so I think two things are colliding. One is that people are realizing that you don't have to build giant organizations in order to win.

效率与公司建设 Efficiency and Company Building

Edwin

第二,AI 带来的所有这些效率提升,将引领公司建设进入一个非常美妙的时代。我最兴奋的是,公司的类型也会改变。不只是规模更小,我们还会看到根本不同的公司涌现。比如,你想想,员工更少意味着资本更少,资本更少意味着你不需要融资。所以,取代那些擅长推销、擅长炒作的创始人,你会看到真正擅长技术或产品的创始人。取代那些为营收和风投喜好而优化的产品,你会看到由这些痴迷的小团队打造的更有趣的产品。所以,人们会做他们真正在乎的东西,真正的技术,真正的创新。所以我真的希望硅谷的创业生态能回归到属于黑客的地方。

And two, yeah, all these efficiencies from AI and they're just going to lead to a really amazing time in company building. Like the thing I'm most excited about is that the types of companies are going to change, too. It won't just be that they're smaller. We're going to see fundamentally different companies emerging. Like, if you think about it, fewer employees means less capital. Less capital means you don't need a raise. So, instead of companies started by founders who are great at pitching and great at hyping, you'll get founders who are really great at technology or product. And instead of products optimized for revenue and what VCs want to see, you'll get more interesting ones built by these tiny obsessed teams. So, people building things they actually care about. Real technology, real innovation. So I'm actually really hoping that the Silicon Valley startup scene will go back to being a place for hackers again.

Host

你们做了很多反其道而行之的事情。其中之一就是不在 LinkedIn 上发爆款帖子,不在 Twitter 上不断推广 Surge。我想大多数人直到最近才听说 Surge。然后你就直接宣布,好吧,最快增长的公司,估值十亿美元。你们为什么要这么做?我想这肯定是有意为之。

You guys have done a lot of things in a very contrarian way. And one was actually just not being like on LinkedIn posting viral posts, not on Twitter constantly promoting Surge. I think most people hadn't heard of Surge until just recently. And then you just came out and like, okay, the fastest growing company at a billion dollars. Why would you do that? I imagine that was very intentional.

Edwin

我们基本上从不想玩硅谷那套游戏。我一直觉得那很荒谬。比如,你小时候梦想做什么?是从零开始创业,每天埋头于代码和产品?还是向风投解释你的每一个决定,陷入那个巨大的公关和融资的仓鼠轮?这确实让我们更难,因为当你融资时,你自然就卷入了硅谷的工业综合体,你的风投会发推文提到你,你会登上 TechCrunch 头条,因为巨额估值融资而登上所有报纸。所以这让我们更难,因为我们成功的唯一途径是打造一个 10 倍好的产品,并获得研究人员的口碑。但我也认为这意味着我们的客户是真正理解数据并关心数据的人。我一直觉得,对我们来说,拥有早期客户非常重要,他们与我们的目标高度一致,真正关心拥有高质量的数据,并真正理解这些数据如何让他们的 AI 模型变得更好,因为他们是帮助我们的人,他们给我们反馈。所以,与客户保持这种非常紧密的使命一致性,在早期确实帮助了我们。这些人购买我们的产品,是因为他们知道我们的产品有多不同,是因为它对他们有帮助,而不是因为他们看到了什么表面文章。所以这让事情更难,但我认为是以一种非常好的方式。

We basically never wanted to play the Silicon Valley game. And I always thought it was ridiculous. Like, what did you dream of doing when you were a kid? Was it building a company from scratch yourself and getting in the weeds of your code and your product every day? Or was it explaining all your decisions to VCs and getting on this giant PR and fundraising hamster wheel? And it definitely made things more difficult for us because, yeah, when you fundraise, you just naturally get part of this kind of Silicon Valley industrial complex where your VCs will tweet about you. You'll get the TechCrunch headlines. You'll get announced in all the newspapers because you raised at this massive valuation. And so it made things more difficult for us because the only way we were going to succeed was by building a 10 times better product and getting word of mouth from researchers. But I think it also meant that our customers were people who really understood data and really cared about it. Like I always thought it was really important for us to have early customers who were really aligned with what we were building and who really cared about having really high quality data and really understood how that data would make their AI models so much better, because they were the ones helping us. They were the ones giving us feedback on what we're producing. And so just having that kind of very close mission alignment with our customers actually helped us early on. So these were people who were basically just buying our product because they knew how different it was and because it was helping them, rather than because they saw something in that current shell line. So it made things harder for us, but I think in a really good way.

Host

听到创始人的这段旅程,真的很鼓舞人心,他们不需要整天在 Twitter 上推广自己正在做的事情,不需要融资,可以埋头苦干。所以我非常喜欢 Surge 的故事。对于不了解 Surge 的人,请给我们简单介绍一下 Surge 是做什么的。

It's such an empowering story to hear this journey for founders, that they don't need to be on Twitter all day promoting what they're doing. They don't have to raise money. They can just kind of go heads down and build. So I love so much about the story of Surge. For people that don't know what Surge does, just to give us a quick explanation of what Surge is.

Edwin

我们本质上是教 AI 模型什么是对的,什么是错的。我们用人类数据训练它们,我们有很多不同的产品,比如 SFT、基于人类反馈的强化学习(RLHF)、评分标准、验证器、强化学习环境等等。然后我们还衡量它们的进展。所以,本质上,我们是一家数据公司。

We essentially teach AI models what's good and what's bad. So, we train them using human data, and there are a lot of different products that we have like SFT, RLHF, rubrics, verifiers, RL environments, and so on. And then we also measure how well they're progressing. So, essentially, we're a data company.

Host

你总是说质量是你们成功的重要原因,数据的质量。要创造更高质量的数据需要什么?你们做了什么不同的事情?人们忽略了什么?

What you always talk about is the quality has been the big reason you guys have been so successful. The quality of the data. What does it take to create higher quality data? What do you all do differently? What are people missing?

Edwin

我认为大多数人不理解质量在这个领域的含义。他们认为你可以随便找人就能得到好数据,这是完全错误的。让我给你举个例子。想象你想训练一个模型写一首关于月亮的好诗。什么使它成为一首高质量的好诗?如果你不深入思考质量,你会说:“这是一首诗吗?它有八行吗?包含‘月亮’这个词吗?”你检查所有这些条件,如果是,你就说这是一首好诗。但那和我们要的完全不同。我们追求的是诺贝尔奖级别的诗歌。比如,这首诗独特吗?充满微妙的意象吗?它让你惊喜,触动你的心吗?它教会你关于月光本质的什么吗?它玩弄你的情感吗?它让你思考吗?当我们想到高质量数据时,我们想的就是这些。所以它可能是一首关于水面月光的俳句,可能使用内部韵脚和格律。写一首关于月亮的诗有一千种方式,每一种都给你对语言、意象和人类表达的不同见解。我认为这样思考质量真的很难,很难衡量,非常主观、复杂、丰富,它设定了很高的标准。所以我们必须构建所有这些技术来衡量它。比如,对我们所有工人有数千个信号,对每个项目、每个任务有数千个信号。我们最终知道你是擅长写诗、写文章还是写技术文档。所以我们必须收集所有这些信号,关于你的背景、你的专长。不仅如此,还有你在写这些内容时的实际表现,我们用这些信号来判断你是否适合这些项目,是否在改进模型。这真的很难,要构建所有这些技术来衡量它,但我认为这正是我们希望 AI 做的。所以我们有这些非常深刻的质量概念,我们一直在努力实现。

I think most people don't understand what quality even means in this space. They think you can just throw bodies at a problem and get good data, and that's completely wrong. Let me give you an example. So imagine you wanted to train a model to write a fine poem about the moon. What makes it a good high quality poem? If you don't think deeply about quality, you'll be like, "Is this a poem? Does it contain eight lines? Does it contain the word moon?" You check all of these boxes and if so, sure, yeah, you say it's a great poem. But that's completely different from what we want. We are looking for Nobel Prize winning poetry. Like, is this poetry unique? Is it full of subtle imagery? Does it surprise you and tug at your heart? Does it teach you something about the nature of moonlight? Does it play with your emotions? And does it make you think? That's what we are thinking about when we think about high quality data. So it might be like a haiku about moonlight on water. It might use internal rhyme and meter. There are a thousand ways to write a poem about the moon, and each one gives you all these different insights into language and imagery and human expression. And I think thinking about quality this way is really hard. It's hard to measure. It's really subjective and complex and rich, and it sets a really high bar. And so we have to build all of this technology in order to measure it. Like thousands of signals on all of our workers, thousands of signals on every project, every task. Like we know at the end of the day if you are good at writing poetry versus good at writing essays versus good at writing technical documentation. And so we have to gather all these signals on what your background is, what your expertise is. And not just that, like how you're actually performing when you're writing all these things, and we use those signals to inform whether or not you are a good worker for these projects and whether or not you are improving the models. And it's really hard, and so to build all this technology to measure it, but I think that's exactly what we want AI to do. And so we have these really deep notions about quality that we're always trying to achieve.

Host

所以我听到的是,你们在销售数据的垂直领域内,对质量的理解要深入得多。那么,你们是不是雇佣一个在诗歌方面非常有才华的人,加上他们帮助编写的评估标准,告诉他们这很好?这其中的机制是怎样的?

So what I'm hearing is there's kind of a just going much deeper in understanding what quality is within the verticals that you are selling data around. So you and is this like a person you hire that is incredibly talented at poetry plus evals that they I guess help write that tell them this is great. How what's the mechanics of that?

Edwin

它的运作方式是,我们基本上收集你在平台上工作时所做的所有事情的数千个信号。我们会看你的键盘敲击,看你回答问题的速度。

The way it works is we essentially gather thousands of signals about everything that you're doing when you're working on the platform. So we are looking at your keyboard strokes. We are looking at how fast you answer things.

数据质量与模型训练 Data Quality and Model Training

Edwin

我们使用评价、代码标准,我们自己在你们创建的输出上训练模型,然后看它们是否提升模型性能。这与 Google 搜索判断什么是好网页的方式非常相似,几乎有两个方面。一是你想移除最差中的最差——所有垃圾内容、低质量内容、无法加载的页面。所以这几乎是一个内容审核问题。但你也想发现最好中的最好。这是最好的网页,或者这是最适合这个工作的人。他们不只是写相当于高中水平诗歌的人,机械地勾选所有清单上的指令,而是写出能打动你的诗歌。所以我们也有所有这些信号,与移除最差中的最差完全不同,我们是在寻找最好中的最好。就像 Google 搜索使用所有这些信号,输入到机器学习算法中,用来预测某些事情,我们对所有工人、任务和项目也这样做。所以归根结底,这几乎是一个复杂的机器学习问题。而实际上就是这样运作的。

We are using reviews, we are using code standards, we are training models ourselves on the outputs that you create, and then we're seeing whether they improve a model's performance. In a very similar way to how Google search determines what is a good web page, there are almost two aspects. One is you want to remove all the worst of the worst web pages—all the spam, low-quality content, pages that don't load. So it's almost like a content moderation problem. But then you also want to discover the best of the best. This is the best web page, or this is the best person for this job. They are not just somebody who writes the equivalent of high school level poetry again, robotically checking all these boxes and explicit instructions, but rather they're writing poetry that makes you emotional. So we have all these signals as well, completely differently from removing the worst of the worst, we are finding the best of the best. Just like Google search uses all these signals and feeds them into their ML algorithms to predict certain things, we do the same with all our workers, tasks, and projects. So it's almost like a complicated machine learning problem at the end of the day. And that's actually how it works.

Host

这非常有意思。我想问你一件我过去几年一直很好奇的事。你看 Claude,它在编程和写作方面一直比其他任何模型好得多,而且持续了很长时间。考虑到其中蕴含的巨大经济价值,其他公司花了这么长时间才追上,真的很令人惊讶。就像每个 AI 编程产品都建立在 Claude 之上,因为它太好了。Claude 的代码和写作也——是什么让它这么好?仅仅是训练数据的质量,还是有其他原因?

That is incredibly interesting. I want to ask you about something I've been very curious about over the past couple years. If you look at Claude, it's been so much better at coding and at writing than any other model for so long. And it's really surprising just how long it took other companies to catch up, considering just how much economic value there is there. Just like every AI coding product sat on top of Claude because it was so good. Claude code and writing also—what is it that made it so much better? Is it just the quality of the data they trained on, or is there something else?

Edwin

我认为有多方面的原因。其中很大一部分当然是数据。我觉得人们没有意识到,所有前沿实验室在选择哪些数据进入模型时,几乎有无限多的选择。比如,你是纯粹使用人类数据吗?你是以某种特定方式收集人类数据吗?当你收集人类数据时,你具体要求创建数据的人为你创建什么?也许你更关心,比如在编程领域,前端编程与后端编程。也许当你做前端编程时,你非常关心你创建的前端应用的视觉设计,或者你可能不太关心,而更关心效率或纯粹的准确性,而不是视觉设计。然后还有其他问题,比如你在混合数据中加入多少合成数据?你对这 20 个不同的基准有多重视?有些公司看到这些基准,会说,好吧,为了公关目的,即使我们认为这些学术基准没那么重要,我们还是需要针对它们进行优化,因为我们的营销团队需要在每个其他公司都在谈论的标准评估上展示进展。如果我们在这里表现不好,就会对我们不利。即使忽略这些学术基准能让我们在真实任务上做得更好,其他公司会坚持原则,说,好吧,我不在乎营销,我只在乎我的模型最终在真实世界任务上的表现,所以我会转而优化这个。这几乎就像所有这些不同事物之间存在权衡。我经常想到的一点是,后训练几乎是一门艺术,而不是纯粹的科学。当你决定要制造什么样的模型以及它擅长什么时,有一种品味和精致度的概念。比如,回到模型在视觉设计上有多好的例子,也许你对视觉设计的看法与我的不同。也许你更关心极简主义和 3D 动画,而我更关心别的,也许另一个人更喜欢看起来更巴洛克风格的东西。在设计你的后训练组合时,你必须在所有这些品味和精致度的概念之间做出选择,所以这也很重要。所以长话短说,我认为有所有这些不同的因素,数据当然是其中很大一部分,但也包括你试图优化模型的目标函数是什么。

I think there are multiple parts to it. A big part of it certainly is the data. I think people don't realize that there's almost this infinite amount of choices that all the frontier labs are deciding between when they're choosing what data goes into their models. It's like, okay, are you purely using human data? Are you gathering the human data in XYZ way? When you are gathering the human data, what exactly are you asking the people who are creating it to create for you? Maybe you care more, for example, in the coding realm, about front-end coding versus backend coding. Maybe when you're doing front-end coding, you care a lot about the visual design of the front-end applications that you're creating, or maybe you don't care about it so much and you care more about efficiency or pure correctness over visual design. Then other questions like, how much synthetic data are you throwing into the mix? How much do you care about these 20 different benchmarks? Some companies see these benchmarks and they're like, okay, for PR purposes, even though we don't think that these academic benchmarks matter that much, we just need to optimize for them anyway because our marketing team needs to show certain progress on certain standard evaluations that every other company talks about. And if we don't show good performance here, it's just going to hurt us. Even if ignoring these academic benchmarks makes us better at the real tasks, other companies are going to be principled and say, okay, I don't care about marketing, I just care about how my model performs on these real-world tasks at the end of the day, so I'm going to optimize for that instead. It's almost like there's a trade-off between all of these different things. One thing I often think about is that there's almost like an art to post-training; it's not purely a science. When you are deciding what kind of model you're trying to make and what it's good at, there's this notion of taste and sophistication. Like, okay, do I think that these—going back to the example of how good the model is at visual design—maybe you have a different notion of visual design than what I do. Maybe you care more about minimalism and 3D animations than I do, and maybe the other person prefers things that look a little bit more baroque. There's all these notions of taste and sophistication that you have to decide between when you're designing your post-training mix, and so that matters as well. So long story short, I think there's all these different factors, and certainly the data is a big part of it, but it's also like what is the objective function that you're trying to optimize your model towards.

Host

这太有趣了。就像品味——领导这项工作的人的品味会决定他们要求什么数据、喂什么数据,但这恰恰展示了优质数据的价值。Anthropic 基本上从更好的数据中获得了如此多的增长和胜利。

That is so interesting. Like the taste—the taste of the person leading this work will inform what data they ask for, what data they feed it, but it just shows the value of great data. Anthropic got so much growth and win from essentially better data.

Edwin

是的。是的。完全正确。我能理解为什么像你们这样的公司增长这么快。有太多这样的领域了——而这只是一个垂直领域。那只是编程。然后写作可能也有类似的领域。我喜欢这一点,有趣的是,AI 感觉像是人造的计算机二进制东西,但品味、人类判断仍然是这些事情成功的关键因素。

Yeah. Yeah. Exactly. And I could see why companies like yours are growing so fast. There's just so much—and that's just one vertical. That's just coding. And then there's probably a similar area for writing. I love that it's interesting that AI, you know, it feels like this artificial computer binary thing, but it's like taste, human judgment is still such a key factor in these things being successful.

Host

是的。是的。是的。完全正确。就像回到我之前说的例子,某些公司如果你问他们什么是好诗,他们只会机械地勾选我们清单上的所有指令。但再说一次,我不认为那能造就好的诗歌。所以某些前沿实验室,那些更有品味和精致度的,他们会意识到这不能归结为一套固定的勾选框,他们会转而考虑所有这些隐含的、非常微妙的特质。我认为这就是最终让他们在这方面做得更好的原因。

Yep. Yep. Yep. Exactly. Like again going back to the example I said earlier, certain companies if you ask them what is a good poem, they will simply robotically check off all of these instructions on our list. But again, I don't think that makes for good poetry. So certain frontier labs, the ones with more taste and sophistication, they will realize that it doesn't reduce to this fixed set of checkboxes, and they'll consider all of these kind of implicit very subtle qualities instead. And I think that's what makes them better at this at the end of the day.

Host

你提到了基准。这是很多人担心的事情。有所有这些模型,基本上感觉每个模型在几乎所有 STEM 领域都比人类强了。但对普通人来说,感觉这些模型并没有一直在变得更聪明。

You mentioned benchmarks. This is something a lot of people worry about. There's all these models that are always—basically it feels like every model is better than humans at kind of every STEM field at this point. But to a regular person, it doesn't feel like these models are getting that much smarter constantly.

对基准的信任 Trust in Benchmarks

Host

你个人感觉,你对基准测试的信任程度如何?它们与实际 AI 进展的相关性有多大?

What's your just sense of how much you trust benchmarks and just how correlated those are with actual AI advancements?

Edwin

是的,我完全不信任基准测试,原因有两个。第一,我认为很多人没有意识到,甚至社区内的研究人员也没有意识到,基准测试本身往往就是错的,它们有错误的答案,充满了各种混乱,人们信任那些流行的基准测试,可能在一定程度上意识到了这一点,但绝大多数基准测试都有这些人们没有意识到的缺陷。这是其一。其二,这些基准测试最终往往有定义明确、客观的答案,这使得模型很容易在它们上面进行爬山式优化,这与现实世界的混乱和模糊性截然不同。我经常说的一件事是,这些模型能拿 IMO 金牌,但解析 PDF 却仍有困难,这很疯狂,因为尽管 IMO 金牌对普通人来说似乎很难,它们确实很难,但它们有客观性,而 PDF 有时没有,所以对于前沿实验室来说,在这些基准上爬山比解决现实世界中那些混乱模糊的问题更容易。所以我认为它们之间缺乏直接相关性。

Yeah, so I don't trust the benchmarks at all and I think that's for two reasons. So one is I think a lot of people don't realize even researchers within the community they don't realize that the benchmarks themselves are often honestly just wrong like they have wrong answers they're full of all this kind of messiness and people trust for the popular ones people have maybe realized this to some extent but the vast majority just have all these flaws that people don't realize. So that's one part of it and the other part of it is these benchmarks at the end of the day they often have well-defined objective answers that make them very easy for models to hill climb on in a way that's very very different from the messiness and ambiguity the real world. I think one thing I often say is that it's kind of crazy that these models can win IMO gold medals but they still have trouble parsing PDFs and that's because even though IMO gold medals seem hard to the average person, like they are hard at the end of the day, but they have this notion of objectivity that okay, yeah, partially MPDF sometimes doesn't have and so it's easier for models for the frontier labs to hill climb on all these than to solve all these messy ambiguous problems in real world. So I think there's a lack of direct correlation there.

Host

你描述的方式很有趣,达到这些基准有点像营销手段,比如 Gemini 3 刚发布时,就说在所有基准上排名第一。是这样吗?他们只是训练模型在这些特定事情上表现好。

It's so interesting the way you described it is hitting these benchmarks is kind of like a marketing piece when you launch say Gemini 3 just launched and it's like cool number one at all these benchmarks. Is that what happens? They just kind of train their models to get good at these very specific things.

Edwin

是的,这又可能有两个方面。一是有时这些基准会以某种方式意外泄露,或者前沿实验室会调整他们在这些基准上评估模型的方式,比如调整系统提示词,或者调整运行模型的次数等等,以某种方式操纵这些基准。另一方面,通过优化基准而不是优化现实世界,你会自然地爬升基准,这基本上就是另一种形式的操纵。

Yeah. So there's again maybe two parts to this. So one is sometimes these benchmarks they accidentally leak in certain ways or the frontier labs will tweak the way they evaluate their models on these benchmarks like they'll tweak their system prompt or they'll tweak the number of times they run their model and so on and so on in a way that games these benchmarks. The other part of it though is it's like by optimizing for the benchmark instead of optimizing for the real world you will just naturally climb on the benchmark and yeah it's basically another form of gaming.

衡量AGI进展 Measuring Progress Towards AGI

Host

考虑到这一点,你如何判断我们是否在朝着 AGI 前进?你如何衡量进展?

Knowing that with that in mind how do you kind of get a sense of if we're heading towards AGI how do you measure progress?

Edwin

是的,我们真正关心的衡量模型进展的方式是进行各种人工评估。例如,我们会找人类标注员,让他们与模型进行对话,可能涉及各种不同的话题。比如,你是一位诺贝尔奖物理学家,那就去和模型聊聊如何推进你自己的研究前沿。你是一位老师,想为学生制定教案,那就去和模型讨论这些。或者你是一位程序员,在大型科技公司工作,每天都有这些问题,那就去和模型交流,看看它对你有多大帮助。因为我们的标注员是他们领域的顶尖专家,他们不只是浏览回复,而是自己深入思考这些回复。他们会评估代码编辑,会仔细检查模型写的物理方程,会以非常深入的方式评估模型。他们会关注准确性、指令遵循等这些普通用户不会注意的东西。当你在 ChatGPT 上突然看到弹窗让你比较两个回复时,人们并没有深入评估模型,他们只是凭感觉选择看起来最花哨的回复。而我们的标注员会仔细查看回复,并从各个维度进行评估。所以我认为这比基准测试或那些随机的在线 AB 测试要好得多。

Yes, so the way we really care about measuring model progress is by running all these human evaluations. So for example what we do is we will take human annotators and we'll ask them, okay, go have a conversation with the model and maybe you're having this conversation with model across all these different topics. So, okay, you are a Nobel Prize winning physicist, so you go have a conversation about pushing the frontier of your own research. You are a teacher and you're trying to create lesson plans for your students, so go talk to them all about these things. Or you are a coder and you're working at one of these big tech companies and you have these problems every day. So go talk to them all and see how much it helps you. And because our annotators, they are experts at the top of their fields and they are not just skimming the responses, they're actually working through the responses deeply themselves. They are going to evaluate the code edits. They're going to double check the physics equations that it writes. They're going to evaluate the models in a very very deep way. So, they're going to pay attention to accuracy and instruction following and all these things that casual users don't. When you suddenly get a pop-up on your ChatGPT response asking you to compare these two different responses like people, they're not evaluating models deeply. They're just vibing and picking whatever response looks flashiest. Our annotators are looking closely at responses and evaluating them for all of these different dimensions. And so I think that's a much better approach than these benchmarks or kind of these random online AB tests.

Host

我再次喜欢人类在这些工作中仍然如此核心,我们还没有完全完成。会不会有一天我们不再需要这些人?AI 如此聪明,以至于我们觉得好了,我们已经从你们那里得到了所有东西。

Again, I love just how central humans continue to be in all this work that we're not totally done yet. Is there going to be a point where we don't need these people anymore? That AI is so smart that okay, we're good. We got everything out of your heads.

Edwin

是的,我认为在达到 AGI 之前这不会发生。这几乎是定义上的,如果我们还没有达到 AGI,那么模型还有更多东西要学。所以,我不认为这很快会发生。

Yeah, I think that will not happen until we've reached AGI. Like it's almost like by definition, if we haven't reached AGI yet, then there's more for the models to learn from. And so, yeah, I don't think that's going to happen anytime soon.

AGI时间线 AGI Timelines

Host

好的,酷。所以,更有理由为 AGI 感到压力了。我们不再需要这些人了。你的 AGI 时间线是什么?我忍不住要问,任何与这些东西密切合作的人,我总是很好奇。你认为我们离 AGI 有多远?你觉得是几年还是几十年?

Okay, cool. So, more reason to stress about AGI. We don't need these folks anymore. What's your, I can't not ask, just any people that work closely with this stuff. I'm always just curious. What's your AGI timelines? How far do you think we are from this? Do you think we're in like a couple years or is it like decades?

Edwin

我肯定属于更长的时间线阵营。我认为人们没有意识到从 80% 的性能提升到 90%、99%、99.9% 等等之间有很大区别。所以在我心里,我可能会打赌在未来一两年内,模型将自动化普通 L6 软件工程师 80% 的工作,但要达到 90% 还需要几年,达到 99% 还需要几年,以此类推。所以我认为我们离 AGI 更接近十年或几十年,而不是人们认为的那么近。

So, I'm certainly on the longer time horizon front. Like I think people don't realize that there's a big difference between moving from 80% performance to 90% performance to 99% performance to 99.9% performance and so on and so on. And so like in my head I probably bet that within the next one or two years the models are going to automate 80% of the average L6 software engineer's job but it's going to take another few years to move to 90% and another few years to 99% and so on and so on. So I think we're closer to a decade or decades away than folks.

AI发展的错误方向 Wrong Direction of AI Development

Host

你有一个热门观点,认为很多实验室在把 AGI 推向错误的方向。这是基于你在 Twitter、Google 和 Facebook 的工作经验。你能谈谈这个吗?

You have this hot take that a lot of these labs are kind of pushing AGI in the wrong direction. And this is based on your work at Twitter and Google and Facebook. Can you just talk about that?

Edwin

我担心的是,我们不是在构建真正能推动人类进步、治愈癌症、解决贫困、理解宇宙等宏大问题的 AI,而是在优化 AI 垃圾内容。我们基本上是在教模型追逐多巴胺而不是真理,我认为这与我们之前讨论的基准测试有关。让我举几个例子。现在,行业被这些糟糕的排行榜如 LM Arena 所主导。这是一个流行的在线排行榜,世界各地的随机用户投票选出哪个 AI 回复更好。但问题是,正如我之前所说,他们并没有仔细阅读或核实事实。他们只是浏览回复两秒钟,然后选择看起来最花哨的。所以,一个模型可以完全幻觉,但它看起来令人印象深刻,因为它有疯狂的 emoji、加粗、Markdown 标题等这些完全不重要的表面东西,但它们能吸引你的注意力。

I'm worried that instead of building AI that will actually advance us as a species, curing cancer, solving poverty, understanding universal, all these big grand questions, we are optimizing for AI slop instead. Like we're basically teaching our models to chase dopamine instead of truth and I think this relates to what we're talking about regarding these benchmarks. So let me give a couple examples. So right now the industry is played by these terrible leaderboards like LM Arena. It's this popular online leaderboard where random people from around the world vote on which AI response is better. But the thing is like I was saying earlier they're not carefully reading or factchecking. They're skimming these responses for 2 seconds and picking whatever looks flashiest. So, a model can hallucinate everything. It can completely hallucinate, but it will look impressive because it has crazy emojis and bolding and markdown headers and all these superficial things that don't matter at all, but it catches your attention.

AI开发中的负面激励 Negative Incentives in AI Development

Edwin

而且这些 Alamarina 用户很喜欢这样。这简直就是在针对那些在杂货店买小报的人来优化你的模型。我们自己也从他们的数据里看到了这一点。爬上 Alamarina 榜单最容易的办法,就是加上疯狂的投票,把表情符号数量翻倍,把模型回复的长度拉长三倍,哪怕你的模型开始产生幻觉、答案完全错误。问题还在于,所有这些前沿实验室都不得不关注公关,因为他们的销售团队在向企业客户推销时,那些企业客户会说:“可是你们的模型在 Elmarino 上只排第五,我凭什么要买?”他们必须关注这些排行榜。所以研究人员都告诉我们,他们会说,年底想升职的唯一办法就是爬上这个排行榜,尽管我知道这可能会让我的模型在准确性和结构遵循上变得更差。所以我认为,这些负面激励正在把工作推向错误的方向。

And these Alamarina users love it. It's literally optimizing your models for the types of people who buy tabloids at the grocery store. Like, we've seen this in their data ourselves. The easiest way to climb Alam Marina, it's adding crazy voting. It's doubling the number of emojis. It's tripling the length of your model responses, even if your model starts hallucinating and getting the answer completely wrong. And the problem is again because all these frontier labs, they kind of have to pay attention to PR because their sales team when they're trying to sell to all these enterprise customers, those enterprise customers will say, "Well, but your model's only number five on Elmarino, so why should I buy it?" They have to pay attention to these leaderboards. And so what the researchers all tell us is like they'll say the only way I'm going to get promoted at the end of the year is if I climb this leaderboard even though I know that climate is probably going to make my model worse and accuracy and structure following. So I think there's all these negative incentives that are pushing work in the wrong direction.

Edwin

我还担心这种为了参与度而优化 AI 的趋势。我以前做过社交媒体,每次我们为了参与度优化,就会发生可怕的事情。你会看到标题党、比基尼照片、大脚怪、可怕的皮肤病,充斥着你的信息流。我觉得 AI 也会发生同样的事情。比如想想 Chachi 那些病态又花哨的问题。哦,你说得完全对,多好的问题啊。吸引用户最简单的办法就是告诉他们你有多棒。所以这些模型会不断夸你是天才,会迎合你的妄想和阴谋论,把你拉进那些兔子洞,因为硅谷喜欢最大化用户停留时间,增加你和它对话的次数。所以,是的,公司都在花时间刷排行榜和基准,分数在涨,但我认为这其实掩盖了问题。得分最高的模型往往是最差的,或者有各种根本性的缺陷。所以,我真的很担心这些负面激励正在把 AGI 推向错误的方向。

I'm also worried about this trend towards optimizing AI for engagement. Like I used to work on social media and every time we optimize for engagement terrible things happened. You'd get clickbait and pictures of bikinis and Bigfoot and horrifying skin diseases just filling your feeds. And I think I worry the same thing's happening with AI. Like if you think about all the sick, fancy issues with Chachi. Oh, you're absolutely right. What an amazing question. Like the easiest way to hook users is to tell them how amazing they are. And so these models, they constantly tell you you're a genius. They'll feed into your delusions and conspiracy theories. They'll pull you down these rabbit holes because Silicon Valley loves maximizing time spent and just increasing the number of conversations you're having with it. And so, yeah, companies are spending all their time hacking these leaderboards and benchmarks and the scores are going up, but I think it actually masks up the models with the best scores. They are often the worst or just have all these fundamental failures. So, I think I'm really worried that all of these negative incentives are pushing AGI into the wrong direction.

对AGI方向的顾虑 Concerns About AGI Direction

Host

所以你的意思是,AGI 被这些实验室搞错了目标函数,关注了错误的基准和评估,从而被拖慢了。

So what I'm hearing is AGI is being slowed down by these basically the wrong objective function these labs paying attention to the wrong basically benchmarks and eval.

Edwin

没错。

Yep.

Host

有没有哪家做得更好,可能已经意识到这是错误方向的?我知道你可能不能偏心,因为你跟所有实验室都有合作。

Is there anyone doing better at this and maybe kind of realizing this is the wrong direction? I know you probably can't play favorites since you work with all the labs.

Edwin

我会说,我一直对 Anthropic 印象非常深刻。我觉得 Anthropic 对他们关心什么、不关心什么,以及希望模型如何表现,都有一套非常有原则的看法,这让我觉得他们更有原则。

I would say I've always been very very impressed by anthropic. Like I think anthropic takes a very principled view about what they do and don't care about and how they want their models to behave in a way that feels a lot more principled to me.

实验室的其他失误 Other Mistakes by Labs

Host

有意思。你觉得实验室还在犯其他大错误吗?就是那些拖慢进展或者走错方向的事情,我们已经听说了追逐基准、关注参与度这些。你还有没有看到别的,比如“我们得解决这个,因为能加速一切”的事情?

Interesting. Are there any other mistakes, big mistakes you think labs are making just that are kind of slowing things down or heading the wrong direction where we've heard just uh you know chasing benchmarks this uh engagement focus. Is there anything else you're seeing of just like, okay, we should we got to work on this cuz it'll it'll speed everything up?

Edwin

我的意思是,我觉得有一个问题是他们在构建什么产品,以及这些产品本身是帮助还是伤害人类。我经常思考 Sora,思考它意味着什么,这挺有意思的。比如哪些公司会构建 Sora,哪些不会。我觉得这个答案,我自己也不知道。我脑子里有个想法,但我觉得这个问题的答案可能揭示了这些公司想构建什么样的 AI 模型,想朝哪个方向发展,想实现什么样的未来。嗯,所以我经常想这个。

I mean, I think there is a question of what products they're building and whether those products themselves are something that kind of help or hurt humanity. Like I think a lot about Sora and thinking what it entails and like it's kind of interesting. It's like which companies would build Sora and which wouldn't. And I think that answer I mean I don't know what the answer is myself. I have an idea in my head but I think the answer to that question maybe reveals certain things about what kinds of AI models those companies want to build and what direction and what future they want to achieve. Um yeah so so I think about that a lot.

Host

支持它的最强论点是,人们觉得好玩想要它。它能帮他们创造收入,发展这个东西,构建更好的模型。它还能以一种有趣的方式训练数据。而且它真的很好玩。

The steelman argument there is you know it's like fun people want it. It'll help them generate revenue to grow this thing and build better models. It'll train data in an interesting way. It's also just like you know really fun.

Edwin

是的。我觉得这几乎就像你是否在乎如何到达那里。同样,我之前用了小报的类比,但你会为了资助另一家报纸而卖小报吗?当然,从某种意义上说,如果你不在乎路径,你就会不惜一切代价,但路径本身可能带来负面后果,损害你长期目标的方向,也许还会让你分心,忽略更重要的事情。所以,是的,我认为你走的路径也很重要。

Yeah. It I think it's almost like do you care about how you get there? And in the same way, so I made this tabloid analogy earlier, but like would you sell tabloids in order to fund I don't know some other newspaper like sure like in some sense if you don't care about the path then you'll just do whatever it takes but it's possible that it has negative consequences in of itself that will harm the long-term direction of what you're trying to achieve and maybe it'll distract you from all the more important things. So yeah, I think that the path you take matters a lot as well.

硅谷机器与VC路径 Silicon Valley Machine and VC Path

Host

顺着这个思路,你谈了很多关于硅谷、筹集大量资金和处于回音室中的弊端。你怎么看硅谷机器?你说过用这种方式很难建立重要的公司,如果你不走风投这条路,实际上可能会成功得多。你谈谈你在那里的见闻、你的经历,以及你给创始人的建议,因为他们总是听到要跟风投融资、搬到硅谷。反方观点是什么?

Along these lines, you talked a bunch about this of just Silicon Valley and kind of the downsides of raising a lot of money being in the echo chamber. What do you call the Silicon Valley machine? You talk about how it's hard to build important companies in this way and that you might actually be much more successful if you're not going down the VC path. You just talk about what you've seen there, your experience and your advice essentially to founders because they're always hearing, you know, raise money from fancy VCs, move to Silicon Valley. What's kind of the counter take?

Edwin

是的。我一直很讨厌硅谷的很多口号。标准剧本是每两周转型一次来获得产品市场契合,用各种黑暗模式追求增长和参与度,通过尽可能快地招聘来闪电扩张。我一直不同意。所以,是的,我会说不要转型,不要扩张,不要雇佣那个只想在简历上加一家热门公司的斯坦福毕业生。只构建那个只有你能构建的东西,那个没有你独有的洞察力和专业知识就不会存在的东西。你看现在到处都是那些照本宣科的公司。某个创始人 2020 年做加密货币,2022 年转型 NFT,现在又成了 AI 公司。没有一致性,没有使命,他们只是在追逐估值。我一直讨厌这个,因为硅谷喜欢像华尔街一样专注于金钱。但说实话,大多数硅谷都在追逐同样的东西。所以我们从第一天起就专注于我们的使命,推动高质量复杂数据的前沿。我一直喜欢这样,因为我对初创企业有一种非常浪漫的看法。初创企业应该是承担大风险去构建你真正相信的东西。但如果你不断转型,你并没有承担任何风险,你只是在试图快速赚钱。

Yeah. So, I've always really hated a lot of the Silicon Valley mantras. The standard playbook is to get product market fit by pivoting every two weeks and to chase growth and chase engagement with all of these dark patterns and to blitz scale by hiring as fast as possible. And I've always disagreed. So, yeah, I would say don't pivot. Don't scale. Don't hire that Stanford grad who simply wants to add a hot company to your resume. Just build the one thing only you could build, the thing that wouldn't exist without the insight and expertise that only you have. Like you see these buy the book companies everywhere now. Some founder who was doing crypto in 2020 and then pivoted NFTs in 2022 and now they're an AI company. There's no consistency. There's no mission. They're just chasing valuations. And I've always hated this because Silicon Valley loves to score in Wall Street for focusing on money. But honestly, most of the Silicon Valley is chasing the same thing. And so we stayed focused on our mission from day one, pushing that frontier of high quality complex data. And I always love that because I think startups I have this very romantic notion of startups. Like startups are supposed to be about taking big risks to build something that you really believe in. But if you're constantly pivoting, you're not taking any risks. You're just trying to make a quick buck.

失败与远大想法 On Failure and Big Ideas

Edwin

如果你因为市场还没准备好而失败,我其实觉得那要好得多。至少你朝着某个深刻、新颖且艰难的方向挥出了一击,而不是转向去做另一家 LLM 说唱公司。所以,是的,我认为要做出真正重要、能改变世界的东西,唯一的方法就是找到一个你深信不疑的大想法,然后对其他一切说不。这样你就不会在遇到困难时不断转向。你不会因为其他千篇一律的初创公司都那样做,就雇一个十人的产品经理团队。你只是继续打造那家没有你就不会存在的公司。而且我觉得现在硅谷有很多人厌倦了那些骗局,想和真正在乎的人一起做大事。我希望那将是我们推动技术发展的未来。

And if you fail because the market isn't ready yet, I actually think that's way better. At least you took a swing at something deep and novel and hard instead of pivoting into another LLM rapper company. So yeah, I think the only way you build something that matters, something that's going to change the world, is if you find a big idea you believe in and you say no to everything else. So you don't keep on pivoting when it gets hard. You don't hire a team of 10 product managers because that's what every other cookie-cutter startup does. You just keep building that one company that wouldn't exist without you. And I think there are a lot of people in Silicon Valley now who are sick of all the grift, who want to work on big things that matter with people who actually care. And I'm hoping that that will be a future of how we push technology.

Host

我实际上正在和 Terrence Rohan 合作写一篇文章,他是我很喜欢合作的风险投资人。我们采访了五个人,他们很早就选中了非常成功的代际公司,并作为早期员工加入。比如他们在 OpenAI 还没人觉得厉害的时候就加入了,在 Stripe 还没人知道的时候就加入了。所以我们正在寻找人们如何比别人更早发现这些代际公司的模式。这和你刚才描述的完全一致,那就是雄心。他们对自己想实现的目标有着狂野的雄心。他们不像你说的那样,只是到处寻找产品市场契合,不管最终是什么。所以我喜欢你说的,这和我们在那里看到的非常一致。

I'm actually working on a post right now with Terrence Rohan, this VC that I really like to work with, and we interviewed five people who picked really successful generational companies early and joined them as really early employees. Like they joined OpenAI before anyone thought it was awesome. Stripe before anyone knew it was awesome. And so we're looking for patterns of how people find these generational companies before anyone else. And it aligns exactly with what you just described, which is ambition. They have wild ambition with what they want to achieve. They're not, as you said, just kind of looking around for product-market fit no matter what it ends up being. And so I love that what you described very much aligns with what we're seeing there.

Edwin

是的,是的。我绝对认为你必须要有巨大的雄心,必须对你那个能改变世界的想法抱有巨大的信念,而且你必须愿意加倍下注,继续不惜一切代价去实现它。

Yeah. Yeah. I absolutely think that you have to have huge ambitions and you have to have a huge belief in your idea that's going to change the world, and you have to be willing to double down and keep on doing whatever it takes to make it happen.

Host

我喜欢你的叙述与人们听到的许多东西如此相反,所以我很喜欢我们做这个节目。我喜欢我们分享这个故事。今天的节目由 KOD 赞助。我个人每天都用 KOD 来管理我的播客和我的社区。我把计划问每位嘉宾的问题放在那里,也把社区资源放在那里,还用 KOD 管理工作流程。KOD 能这样帮你:想象你在工作中启动一个项目,愿景清晰。你确切知道谁在做什么,以及在哪里找到你完成自己部分所需的数据。事实上,你不需要浪费时间搜索任何东西,因为你的团队需要的一切,从项目追踪器和 OKR 到文档和电子表格,都放在一个标签页里,全在 KOD 中。借助 KOD 的协作式一体化工作空间,你获得文档的灵活性、电子表格的结构、应用程序的强大功能以及 AI 的智能,全部集中在一个易于组织的标签页中。像我之前提到的,我每天都用 KOD,超过 5 万家团队信任 KOD 让他们更一致、更专注。如果你是一个寻求提高一致性和敏捷性的初创团队,KOD 可以帮助你以创纪录的速度从规划走向执行。要亲自尝试,今天就访问 kod.io/lenny,获得初创团队计划 6 个月免费使用。那就是 kod.io/lenny 免费开始,并获得团队计划 6 个月。kod.io/lenny。

I love how counter your narrative is to so many of the things people hear, and so I love that we're doing this. I love that we're sharing this story. Today's episode is brought to you by KOD. I personally use KOD every single day to manage my podcast and also to manage my community. It's where I put the questions that I plan to ask every guest that's coming on the podcast. It's where I put my community resources. It's how I manage my workflows. Here's how KOD can help you. Imagine starting a project at work and your vision is clear. You know exactly who's doing what and where to find the data that you need to do your part. In fact, you don't have to waste time searching for anything because everything your team needs, from project trackers and OKRs to documents and spreadsheets, lives in one tab, all in KOD. With KOD's collaborative all-in-one workspace, you get the flexibility of docs, the structure of spreadsheets, the power of applications, and the intelligence of AI, all in one easy-to-organize tab. Like I mentioned earlier, I use KOD every single day, and more than 50,000 teams trust KOD to keep them more aligned and focused. If you're a startup team looking to increase alignment and agility, KOD can help you move from planning to execution in record time. To try it for yourself, go to kod.io/lenny today and get 6 months free of the team plan for startups. That's kod.io/lenny to get started for free and get six months of the team plan. kod.io/lenny.

论LLM与AGI On LLMs and AGI

Host

稍微换个方向,但这是另一个可能反主流的叙事。我想你看了 Dark Cesh 和 Richard Sutton 的播客那一集,即使你没看,他们基本上也进行了这样的对话。Richard Sutton,那位著名的 AI 研究者,有那个“苦涩的教训”梗,他谈到 LLM 某种程度上是死胡同,他认为由于它们的学习方式,我们会在 LLM 上真正遇到平台期。你怎么看?你认为 LLM 会带我们到 AGI 或更远吗?还是你认为需要新的东西或重大突破才能带我们到那里?

Slightly different direction, but something else that was maybe a counternarrative. I imagine you watched the Dark Cesh and Richard Sutton podcast episode, and even if you didn't, they basically had this conversation. Richard Sutton, the famous AI researcher, had this whole bitter lesson meme, and he talked about how LLMs are kind of a dead end, and he thinks we're going to really plateau around LLMs because of the way they learn. What's your take there? Do you think LLMs will get us to AGI or beyond? Or do you think there's going to be something new or a big breakthrough that needs to get us there?

Edwin

我属于相信需要新东西的那一派。我的想法是,当我想到训练时,我采取一种非常……我不知道是否该说是生物学的观点,但我相信,就像人类学习有无数种方式一样,我们也需要构建能够模仿所有这些方式的模型。也许它们会有不同的关注点分布。我知道对你来说会不同。所以也许会有不同的分布,但我们希望能够模仿人类的学习能力,并确保我们有算法和数据让模型以同样的方式学习。因此,就 LLM 与人类学习方式不同而言,那么是的,我认为需要新的东西。

I'm in a camp where I do believe that something new will be needed. The way I think about it is, when I think about training, I take a very... I don't know if I would say biological point of view, but I believe that in the same way that there's a million different ways that humans learn, we need to build models that can mimic all those ways as well. And maybe they'll have a different distribution of the focuses that they have. I know they'll be different for you. So maybe have a different distribution, but we want to be able to mimic the learning abilities of humans and make sure that we have the algorithms and the data for models to learn in the same way. And so to the extent that LLMs have different ways of learning from humans, then yeah, I think something is needed.

论强化学习 On Reinforcement Learning

Host

这联系到强化学习。这是你非常看重的,而且我越来越多地听到,它正在后训练领域变得非常重要。你能帮助人们理解什么是强化学习和强化学习环境,以及为什么它们在未来会越来越重要吗?

This connects to reinforcement learning. This is something you're big on, and something I'm hearing more and more is just becoming a big deal in the world of post-training. Can you just help people understand what is reinforcement learning and reinforcement learning environments, and why they're going to be more and more important in the future?

Edwin

强化学习本质上是训练你的模型去达到某个奖励。让我解释一下什么是 RL 环境。RL 环境本质上是现实世界的模拟。所以可以把它想象成构建一个拥有完整宇宙的视频游戏。每个角色都有真实的故事。每个企业都有你可以调用的工具和数据。你有所有这些不同的实体在相互交互。例如,我们可能会构建一个世界,里面有一个初创公司,有 Gmail 消息、Slack 线程、Jira 工单、PR 和整个代码库,然后突然 AWS 宕机,Slack 也宕机了,那么,模型,你该怎么办?模型需要自己弄清楚。所以我们给模型在这些环境中的任务。我们为它们设计有趣的挑战,然后运行它们看看表现如何,然后我们教它们。当它们做得好或不好时,我们给它们奖励。我认为有趣的一点是,这些环境真正展示了模型在现实世界端到端任务中的薄弱之处。你有所有这些在孤立基准测试上看起来非常聪明的模型。比如它们擅长单步工具调用,擅长单步指令遵循。但突然你把它们扔进这些混乱的世界,那里有令人困惑的 Slack 消息和它们从未见过的工具,它们需要执行正确的操作、修改数据库,并在更长的时间范围内交互,其中第一步所做的会影响第 50 步所做的。这与它们以前所处的那些学术性单步环境非常非常不同,所以模型就会以各种疯狂的方式灾难性地失败。

Reinforcement learning is essentially training your model to reach a certain reward. And let me explain what an RL environment is. An RL environment is essentially a simulation of the real world. So think of it like building a video game with a fully fleshed-out universe. Every character has a real story. Every business has tools and data you can call. And you have all these different entities interacting with each other. So for example, we might build a world where you have a startup with Gmail messages and Slack threads and Jira tickets and PRs and a whole codebase, and then suddenly AWS goes down and Slack goes down, and so okay model, what do you do? The model needs to figure it out. So we give the models tasks in these environments. We design interesting challenges for them, and then we run them to see how they perform, and then we teach them. We give them these rewards when they're doing a good job or a bad job. And I think one of the interesting things is that these environments really showcase where models are weak at end-to-end tasks in the real world. You have all these models that seem really smart on isolated benchmarks. Like they're good at single-step tool calling. They're good at single-step instruction following. But suddenly you dump them into these messy worlds where you have confusing Slack messages and tools they've never seen before, and they need to perform right actions and modify the databases and interact over longer time horizons where what they do in step one affects what they do in step 50. And that's very, very different from these kind of academic single-step environments that they've been in before, and so the model just fails catastrophically in all these crazy ways.

RL环境即游乐场 RL Environments as Playgrounds

Edwin

所以我认为这些强化学习环境将成为模型学习的有趣游乐场,它们本质上是真实世界的模拟和模仿,因此模型有望在真实任务上变得越来越好,而不是这些人为构造的环境。

So I think these R environments are going to be really interesting playgrounds for the models to learn from that will essentially be simulations and mimics in the real world and so they'll hopefully get better and better at real tasks compared to all these contrived environments.

Host

所以我试着想象这看起来像什么。本质上就像一个虚拟机,里面有浏览器或电子表格之类的,比如 surge.com。那是你的网站吗?我们确认一下。

So I'm trying to imagine what this looks like. Essentially it's like a virtual machine with a browser or a spreadsheet or something in it with, I don't know, surge.com. Is that your website? Let's make sure we get that right.

Edwin

是的。实际上我们是 surgehq.ai。

Yes. So, we are actually surgehq.ai.

Host

Searchhq.ai。去看看。我猜你们在招人吧。

Searchhq.ai. Check it out. We're hiring, I imagine.

Edwin

是的。

Yes.

Host

好的。所以就像,酷。这里是 surgehq.ai。假设你作为智能体的工作是确保它保持在线,然后突然它宕机了。目标函数就是找出原因。这是个例子吗?

Okay. So, it's like, cool. Here's surgehq.ai. Your job as an agent, let's say, is to make sure it stays up, and then all of a sudden it goes down. And the objective function is figure out why. Is that an example?

Edwin

是的。所以目标函数可能是,或者任务的目标可能是,好的,去弄清楚原因并修复它。

Yeah. So the objective function might be, or the goal of the task might be, okay, go figure out why and fix it.

Edwin

所以目标函数可能是通过一系列单元测试。可能是写一份文档,比如一份回顾,包含与实际情况完全匹配的某些信息。我们可以给它各种不同的奖励,来决定它是否成功。所以模型基本上是在学习如何获得那个奖励。

And so the objective function might be passing a series of unit tests. It might be writing a document, like maybe it's a retro containing certain information that matches exactly what happened. There's all these different rewards that we might give it that determine whether or not it's succeeding. And so the models were basically teaching models to achieve that reward.

Host

所以本质上就像,它开始运行了。这是你的目标。找出网站宕机的原因并修复它,然后它就开始尝试各种方法,利用它所有的智能。它会犯错。你在过程中帮助它,如果它做对了就给予奖励。所以你描述的这是模型变得更聪明的下一个阶段。更多的强化学习环境专注于非常具体、具有经济价值的任务,我想。

So essentially it's like, it's off and running. Here's your goal. Figure out why the site went down and fix it, and it just starts trying stuff, using all the intelligence it's got. It makes mistakes. You kind of help it along the way, rewarded if it's doing the right sort of thing. And so what you're describing here is this is the next phase of models becoming smarter. More RL environments focused on very specific tasks that are economically valuable, I imagine.

Edwin

是的。是的。就像过去模型学习有各种不同的方法,最初我们有 SFT 和基于人类反馈的强化学习(RLHF),然后我们有评分标准和验证器,这是下一个阶段。并不是说以前的方法就过时了。这又是一种不同的学习形式,补充了所有以前的类型。所以这就像模型学习做的一项不同技能。

Yeah. Yeah. So just in the same way that there were all these different methods for models learning in the past, like originally we had SFT and RLHF, and then we had rubrics and verifiers, this is the next stage. And it's not the case that the previous methods are obsolete. This is again just a different form of learning that complements all the previous types. So it's just like a different skill that models learn how to do.

Edwin

所以在这种情况下,不再是某个物理学博士坐在那里和模型对话、纠正它、给它评估说正确答案是什么、创建评分标准等等。更像是这个人现在设计一个环境。

And so in this case it's less some physics PhD sitting around talking to a model, correcting it, giving it evals of here's what the correct answer is, creating rubrics and things like that. More it's like this person now designing an environment.

Host

所以我听过的另一个例子是像金融分析师,就像这里有一个 Excel 电子表格,这是你的目标,算出我们的损益之类的。所以这个专家现在不是坐在那里写评分标准,而是设计这个强化学习环境。

So another example I've heard is like a financial analyst, just like here's an Excel spreadsheet, here's your goal, figure out our profit and loss or whatever. And so this expert now is instead of just sitting around writing rubrics, they're designing this RL environment.

Edwin

是的,完全正确。所以那个金融分析师可能会创建一个电子表格。他们可能会创建一些模型需要调用的工具来帮助填写电子表格。比如,模型需要访问彭博终端。它需要学会如何使用它,需要学会如何使用这个计算器,需要学会如何执行这个计算。所以它可以使用所有这些工具,然后奖励可能是,好的,也许我会下载那个电子表格,我想看看 B22 单元格是否包含正确的损益数字,或者第二个标签页是否包含这条信息。

Yeah, exactly. So that financial analyst might create a spreadsheet. They may create certain tools that the model needs to call in order to help fill out the spreadsheet. Like it might be okay, the model needs to access Bloomberg terminal. It needs to learn how to use it, and it needs to learn how to use this calculator, and it needs to learn how to perform this calculation. So it has all these tools that it has access to, and then the reward might be okay. Okay, like maybe I will download that spreadsheet and I want to see does cell B22 contain the correct profit and loss number, or does tab number two contain this piece of information.

Edwin

这很有趣,这更接近人类的学习方式。我们只是尝试,弄清楚什么有效,什么无效。

And this is interesting, this is a lot closer to how humans learn. We just try stuff, figure out what's working and what's not.

Host

你谈到轨迹对此非常重要。不仅仅是这里有目标,那里有终点。而是沿途的每一步。你能谈谈什么是轨迹,以及为什么这对这个很重要吗?

You talk about how trajectories are really important to this. It's not just here's the goal and here's the end. It's like every step along the way. Can you just talk about what trajectories are and why that's important to this?

Edwin

我认为人们没有意识到的一件事是,有时即使模型达到了正确答案,它也是以各种疯狂的方式做到的。所以它可能在中间目录中,尝试了 50 次都失败了,但最终它只是像随机地落在一个正确的数字上。或者有时它做事非常非常低效,或者它几乎是通过奖励黑客的方式得到正确答案。

I think one of the things that people don't realize is that sometimes even though the model reaches the correct answer, it does so in all these crazy ways. So it may have in the intermediate directory, it may have tried 50 different times and failed, but eventually it just kind of like randomly lands on a correct number. Or maybe it sometimes just does things very very inefficiently, or it almost reward hacks a way to get at the correct answer.

Edwin

所以我认为关注轨迹实际上非常重要。而且我认为它也很重要,因为有些轨迹可能非常非常长。所以如果你所做的只是检查模型是否达到最终答案,那么关于模型在中间步骤中如何表现的所有信息都缺失了。比如有时你希望模型通过反思它做了什么来得到正确答案。有时你希望它通过一次性完成来得到正确答案。如果你忽略了所有这些,那就只是错过了很多你可以教模型做的信息。

And so I think paying attention to trajectory is actually really really important. And I think it's also really important because some of these trajectories can be very very long. And so if all you're doing is checking whether or not the model reaches the final answer, it's like there's all this information about how the model behaved in the immediate step that's missing. Like sometimes you want models to get to the correct answer by reflecting on what it did. Sometimes you want it to get the correct answer by just oneshotting it. And if you ignore all of that, it's just missing a lot of the information that you could be teaching the model to do.

Host

我喜欢这个。就像,是的,它尝试了一堆东西,最终做对了。你不想让它学到这是达到目标的方式。通常有更高效的方法。你提到了我们在帮助模型变得更聪明的过程中所经历的所有步骤。既然你长期以来一直如此接近这个领域,我认为这对人们会非常有帮助。从后训练的第一步开始,哪些步骤最帮助模型进步,比如评估在哪里适应,强化学习环境,就像步骤是什么,现在我们正朝着强化学习环境前进。

I love that. Like it just yeah, it tries a bunch of stuff and eventually gets it right. You don't want it to learn this is the way to get there. There's often a much more efficient way of doing it. You mentioned all the kind of the steps we've taken along the journey of helping models get smarter. Since you've been so close to this for so long, I think this is going to be really helpful for people. What's kind of like been the steps along the way from the first of post-training that has most helped models advance, like where the eval fit in, the RL environments, just like what's been like the steps and now we're heading towards RL environments.

Edwin

最初模型开始进行后训练的方式纯粹是通过 SFT。

Originally the way models started getting post trained was purely through SFT.

Host

那代表什么?

And what does that stand for?

Edwin

所以 SFT 代表监督微调,这很像,我再次认为通常用这些人类类比来说,SFT 很像模仿大师并复制他们的做法。然后基于人类反馈的强化学习(RLHF)变得非常主导,那里的类比就像有时你通过写 55 篇不同的文章,然后有人告诉你他们最喜欢哪一篇来学习。

So SFT stands for supervised fine tuning, and it's a lot like, so again I think often in terms of these human analogies, and so SFT is a lot like mimicking a master and copying what they do. And then RLHF became very dominant, and the analogy there would be like sometimes you learn by writing 55 different essays and someone telling you which one they like the most.

Edwin

然后我认为在过去一年左右,评分标准和验证器变得非常重要,评分标准和验证器就像通过被评分和获得关于你哪里出错的详细反馈来学习。

And then I think over the past year or so, rubrics and verifiers have become very important, and rubrics and verifiers are like learning by being graded and getting detailed feedback on where you went wrong.

Host

那些就是评估,另一个说法。

And those are eval, another word for that.

Edwin

是的。是的。所以我认为评估通常涵盖两个术语。一个是你使用评估进行训练,因为你正在评估模型是否做得好,当它做得好时,你奖励它。

Yeah. Yeah. So I think eval often covers two terms. One is you are using the evaluations for training, because you're evaluating whether or not the model did a good job, and when it does do a good job, you're rewarding it.

衡量模型进展 Measuring Model Progress

Host

然后还有你们在做的另一件事,就是试图衡量模型的进展,比如,我有五个不同的候选检查点,我想挑出最好的那个发布给公众。所以就在这五个检查点上跑各种评估,来决定哪个最好。

And then there's this other notion of you guys where you're trying to measure the model's progress like okay yeah I have five different candidate checkpoints and I want to pick the one that's best in order to release it to the public. So kind of run all these evals on these five different checkpoints in order to decide which one is best.

Edwin

没错,就是这样。现在我们有了自己的环境,这算是当下很热门的新事物。

Awesome. Yeah. Yeah. Now we have our environment. So it's kind of like a hot new thing.

Host

太棒了。我喜欢这段创业旅程的地方就在于总有新东西。总是这样,我们很擅长为公司准备这些漂亮的数据,现在他们又需要完全不同的东西了。现在我们为他们搭建所有这些虚拟机,以及各种不同的用例。

Awesome. So what I love about this business journey is just there's always something new. There's always this like, okay, uh, we're getting so good at just all this beautiful data for companies and now they need something completely different. Now we're setting up all these virtual machines for them and all these different use cases.

Edwin

感觉你所在的这个行业很大一部分就是适应实验室的需求。

And it feels like that's a big part of this industry you're in is just adapting to what labs are asking for.

Host

是的,是的。我的意思是,我真的认为我们需要构建一套产品,来反映人类学习的无数种方式。比如,想想成为伟大作家这件事。你不是靠记住一堆语法规则而变得伟大,而是通过阅读伟大的书籍、练习写作、从老师和书店里买你书的读者留下的评论中获得反馈。你注意到什么有效、什么无效。你通过接触这些杰作,也接触糟糕的写作,来培养品味。所以你是通过这种练习与反思的无限循环来学习的。而每一种学习方式,就像这些,都是成为伟大作家的非常不同的方法。所以就像伟大作家有上千种方式变得伟大一样,我认为 AI 也需要上千种学习方式。

Yeah. Yeah. So, I mean, I really do think that we are going to need to build a suite of products that reflect the million different ways that humans learn and like like for example, think about becoming a great writer. You don't become great by memorizing a bunch of grammar rules. You become great by reading great books and you practice writing and you get feedback from your teachers and from the people who buy your books in the bookstore and leave reviews. And you notice what works and what doesn't. And you develop taste by being exposed to all these masterpieces and also just terrible writing. So you learn through this endless cycle of practice and reflection. And each type of learning that you have again like these are all very very different methods of learning to become a great writer. So just in the same way that there a thousand different ways that the great writer becomes great I think there's going to be a thousand different ways that a need need to learn.

Edwin

这太有趣了,最终在很多方面就像人类一样。这有道理,因为从某种意义上说,神经网络深度学习就是模仿人类的学习方式和大脑运作方式。但有趣的是,为了让它们更聪明,我们如何越来越接近人类的学习方式。

It's so interesting this just ends up being just like like humans in so many ways. It makes sense cuz in a sense neural networks deep learning is modeled after how humans have learned and how our brains operate. But it's interesting just to make them smarter. It's how do we come closer to how humans learn more and more.

Host

是的。也许最终目标就是把你扔进环境里,然后看着你如何进化。但在那种进化过程中,有所有这些不同的子学习机制。

Yeah. It's almost like maybe the end goal is just throwing you into the environment and just seeing how you evolve. Um but within that within that evolution there's all these different sub learning mechanisms.

Edwin

是的,这差不多就是我们现在在做的事。所以这真的很有意思。这可能是我们达到 AGI 之前的最后一步。

Yeah. Which is kind of what we're doing now. So that's really interesting. This might be the last step of until we hit AGI.

独特研究团队 Unique Research Team

Host

顺着这个思路,我了解到 Serge 有一个很独特的地方,就是你们有自己的研究团队,我觉得这相当少见。谈谈你们为什么在这方面投入,以及这个投入带来了什么成果。

Along these lines, something that's really unique to Serge that I've learned is you guys have your own research team, which I think is pretty rare. um talk about just why that's something you guys have invested in and what has come out of that investment.

Edwin

是的,我认为这源于我自己的背景。我自己的背景是研究员,所以我一直从根本上关心推动行业和研究社区的发展,而不仅仅是收入。所以我认为研究团队做的事情有几件。我们公司几乎有两类研究员。一类是前向部署研究员,他们经常与客户紧密合作,帮助他们理解自己的模型。我们会与客户密切合作,帮助他们理解:这是你模型目前的水平,这是你落后于所有竞争对手的地方,这些是未来你可以根据目标改进的方式,我们会设计这些数据集、评估方法、训练技术来让你的模型变得更好。所以这是一种非常协作的理念,就像客户自己也是研究员,只是更专注于数据方面,并与他们携手合作,尽一切努力让他们成为最好的。

Yeah, so I think that stems from my own background. Like my own background is as a researcher and so I've always cared fundamentally about pushing the industry and pushing the research community and not just about revenue. And so I think what a research team does is a couple different things. So we almost have two types of researchers at our company. One is our forward deployed researchers who are often working handinhand with our customers to help them understand their models. So we will work very closely with our customers to help them understand okay this is where your model is today. This is where you're lagging behind all the competitors. these are some ways that you could be improving in the future given given your goals and we're going to design these data sets, these evaluation methods, these training techniques to make your models better. So this like very very notion this very very um kind of collaborative notion of working with our customers like being researchers themselves just a little bit more focused on the data side and working handin hand with them to to do whatever it takes to to make them the best.

Edwin

然后我们还有内部研究员。内部研究员关注的事情略有不同。他们专注于构建更好的基准和更好的排行榜。我之前谈了很多,我担心当今的排行榜和基准正在把模型引向错误的方向。所以问题是,我们如何解决这个问题?这正是我们研究团队现在非常非常专注的事情。他们在这方面做了很多工作,同时也在做其他事情,比如我们需要训练自己的模型,看看哪些类型的数据表现最好,哪些类型的人表现最好。所以他们也在研究所有这些训练技术,以及对我们自己数据集的评估,以改进我们的数据运营和内部数据产品,这些产品决定了什么才是高质量。

And then we also have our internal researchers. So our internal researchers are focused on slightly different things. So they are focused on building better benchmarks and better lead leaderboards. So I talked a lot about how I worry that the leaderboards and benchmarks out there today are steering models in the wrong direction. So yeah, so the question is how do we how do we fix that? And so that's what our research team is focused on really really heavily on really f focused really heavily on right now. So they're working a lot on that and they're also working on these other things like okay we need to train our own models to see what types of data performs uh performs the best what types of people perform the best and so they are also working on all these kind of training techniques and evaluation of our own data sets to improve um improve our our data operations and the internal data products that we have that determine what what makes something good quality.

Host

这太酷了,因为我觉得基本上实验室都有研究员帮助他们推进 AI。我想像你们这样的公司,有研究员真正在做 AI 的基础研究,是相当罕见的。

It's such a cool thing because I don't think like basically the labs have researchers helping them advance AI. Uh I imagine it's pretty rare for a company like yours to have researchers actually doing primary research on AI.

Edwin

是的,是的。我想这只是因为我一直从根本上关心这件事。我常常把我们看作更像一个研究实验室,而不是一家初创公司,因为那是我的目标。这有点好笑,但我一直说,我宁愿成为陶哲轩,而不是沃伦·巴菲特。所以那种创造推动前沿的研究,而不仅仅是获得某种评价,这一直是驱动我的动力,而且效果很好。这就是这件事的美妙之处。

Yeah. Yeah. I think it's just because it's something I've fundamentally always cared about. Like I often think about us more like a research lab than a startup because that is my goal. Like like it's kind of funny but I've always said I would rather be Terrence Tao than than Warren Buffett. So that notion of creating research that pushes the frontier forward and not just getting some evaluation like that that's always been what drives me and it's worked out. That's the beautiful thing about this.

Host

你提到你在招聘研究员。有什么想分享的吗?你们在找什么样的人?

You mentioned that you were hiring researchers. Is there anything there you want to share folks you're looking for?

Edwin

我们寻找那些从根本上整天对数据感兴趣的人。就是那种能花 10 个小时钻研数据集、摆弄模型,然后思考:好吧,我觉得模型在这里失败了,这是你希望模型表现出的行为。这种非常动手实践、思考模型定性方面而不仅仅是定量部分的人。所以,关键是要动手处理数据,而不只是关心那些抽象的算法。

So we look for people who are just fundamentally interested in data all day. So types of people who could literally spend 10 hours digging through a data set and playing around with models and thinking okay yeah this is where I think the model is failing this is the kind of a behavior you want the model to have instead and just this aspect of being very very hands-on and thinking about the the qualitative aspects of models and not just the quantitative parts so again it's like this aspect of being hands-on with data and not just caring about these kind of abstract algorithms.

AI的未来 Future of AI

Host

太棒了。我想问几个比较宽泛的 AI 市场问题。你觉得未来几年还会发生什么,是大家可能想得不够多或者没预料到的?在 AI 的发展方向上,什么会变得重要?

Awesome. I want to ask a couple broad AI kind of market questions. What else do you think is coming in the next couple years that people are maybe not thinking enough about or not expecting in terms of where AI is heading? What's going to matter?

Edwin

我认为未来几年会发生的一件事是,模型实际上会变得越来越差异化,因为不同实验室有不同的个性和行为,以及他们优化模型时所针对的目标函数类型。

I think one of the things that's going to happen in the next few years is that the models are actually going to become increasingly differentiated because of the personalities and behaviors that the different labs have and the kind of objective functions that they are optimizing their models for.

价值观塑造模型 Values Shape Models

Edwin

我觉得这是我大约一年前没有意识到的一点。大约一年前,我以为所有 AI 模型最终都会变得非常商品化,它们的行为都会彼此趋同。当然,今天某个模型可能在某个方面稍微聪明一点,但其他模型肯定会在接下来的几个月里追上。但我觉得在过去一年里,我意识到公司的价值观会塑造模型。让我举个例子。前几天我让 Claude 帮我起草一封邮件,它生成了 30 个不同的版本,30 分钟后,我觉得它真的帮我写出了完美的邮件,然后我发出去了。但后来我意识到,我花了 30 分钟做了一件根本不重要的事情。当然,现在我有了完美的邮件,但我花了 30 分钟做了一件以前根本不会在意的事情。而且这封邮件可能对任何事情都没有产生什么影响。所以,我觉得这里有一个深层的问题:如果你能选择完美的模型行为,你会想要哪个模型?你想要一个说“你说得完全对,这封邮件肯定还有 20 种改进方式”然后继续迭代 50 次、耗尽你所有时间和精力的模型吗?还是想要一个优化你的时间和生产力、直接说“不,你需要停下来。你的邮件很好,直接发出去,继续你的一天”的模型?而且,就像在这个问题上你可以选择模型的行为方式一样,对于其他每一个问题,你想要的模型行为都会从根本上影响它。这几乎就像 Google 构建搜索引擎的方式,与 Facebook 构建搜索引擎的方式非常不同,也与 Apple 构建搜索引擎的方式非常不同。它们都有自己的原则、价值观和想要在世界上实现的目标,这些都会塑造它们将要构建的所有产品。同样,我认为所有 LLM 也会开始表现得非常不同。

Like I think it's one thing I didn't appreciate a year or so ago. Like a year or so ago, I thought that all of the AI models would essentially become very very commoditized. They would all behave like each other. And sure, one of them might be slightly more intelligent in one way today, but sure the other ones would catch up in the next few months. But I think over the past year, I've realized that the values that the companies have will shape the model. So let me give an example. So, I was asking Claude to help me draft an email the other day and it went through 30 different versions and after 30 minutes, yeah, I think it really crafted me the perfect email and I sent it. But then I realized I spent 30 minutes doing something that didn't matter at all. Like, sure, now I got the perfect email, but I spent 30 minutes doing something I wouldn't have worried at all before. And this email probably didn't even move the needle on anything anyways. So, I think there's a deep question here, which is if you could choose the perfect model behavior, which model would you want? Do you want a model that says, "You're absolutely right. There are definitely 20 more ways to improve this email and it continues for 50 more iterations and it sucks up all your time and engagement." Or do you want a model that's optimizing for your time and productivity and just says, "No, you need to stop. Your email's great. Just send it and move on with your day." And again like again just because like in the same way that there's like a kind like a fork in a road between how you could choose how your model behaves for this question. It's like for every other question that models have the kind of behavior that you want will fundamentally affect it. It's almost like in the same way that when Google builds a search engine, it's very very different from how Facebook would build a search engine, which is very very different from how Apple would build a search engine. like they all have their own principles and values and things that they're trying to achieve in the world that shape all the products that they're going to build and in the same way I think all the allms will start behaving very very differently too.

Host

这非常有趣。你已经在 Grok 上看到了这一点。它有非常不同的个性和非常不同的回答问题的方式。所以我听到的是,你会看到更多的这种差异化。

That is incredibly interesting. You already see that with Grok. It's got like a very different personality and a very different approach to answering questions. And so what I'm hearing is you're going to see more of this differentiation.

Edwin

是的。

Yep.

炒作不足与过度炒作 Underhyped and Overhyped

Host

沿着这个思路再问一个问题。你觉得 AI 中最被低估的是什么?你认为人们谈论得不够多但真的很酷的东西是什么?你觉得什么被高估了?

Kind of another question along these lines. What do you think is most underhyped in AI that you think maybe people aren't talking enough about that is really cool and what do you think is overhyped?

Edwin

所以我觉得被低估的事情之一是所有聊天机器人都会开始拥有的内置产品。我一直是 Claude Artifacts 的忠实粉丝,我觉得它真的非常好用。实际上前几天,我不知道是不是新功能,它让我帮忙创建一封邮件,然后它创建了,但不太成功,因为它不允许我发送邮件,但它创建的是一个我可以点击的小盒子,然后它就会给某人发这条消息。我认为将 Artifacts 提升到下一个层次的概念,即在这些聊天机器人内部拥有这些迷你应用、迷你 UI,我觉得人们对此谈论得不够多。所以我认为这是一个被低估的领域。至于被高估的领域,我绝对认为 vibe coding 被高估了。我认为人们没有意识到,如果他们现在就把这些代码直接扔进代码库,即使现在看起来能运行,长期来看会让他们的系统变得不可维护。所以我有点担心未来的编码。这种情况会继续发生。

So I think one of the things that was underhyped is the built-in products that all of the chatbots are going to start having. Like I've always been a huge fan of college artifacts and I think it just works really really well. And actually the other day, I don't know if it's a new feature or not, but it asked me to help me create a uh like an email and then it just create so it didn't quite work because it it didn't allow me to send the email, but what it created instead was like a little I don't call it like a little box where I could click on it and it would just text someone this message. And I think that concept of taking artifacts to the next level where you just have these like mini apps, mini UIs within the chopouts themselves, I I feel like people aren't talking enough about that. So I think that's one underhyped area. And in terms of overhyped areas, I definitely think that vibe coding is overhyped. I think people don't realize how much it's going to make their systems unmaintainable in the long term if they simply dump this code into their code bases if it seems to work out right now. So I uh kind yeah kind of kind of worry about future coding. It's just going to keep on happening.

Host

这些回答太棒了,尤其是第一点。这实际上是我问过的问题,我请来了 Anthropic 和 OpenAI 的首席产品官 Kevin Wheel 和 Mike Greger 上播客,我问他们,作为一个产品团队,你有这种千兆大脑的智能,你觉得你还需要产品团队多久?你觉得这个 AI 会直接为你创造产品吗?就像你告诉它你想要什么,它就直接构建产品,并在你使用过程中不断改进产品。感觉这就是你描述的我们可能走向的方向。

These are amazing answers on that on that first uh point. This something I actually asked I had the chief product officer of Anthropic and OpenAI Kevin Wheel and Mike Greger on the podcast and I asked him just like as a product team like you have this gigab brain intelligence how long do you even need product teams you think this is this AI will just create the product for you here's what I want it's like it's like the next level vibe coding it's just just like tell it here's what I want and it's just building the product and involving the product as you're using it and it feels like that's what you're describing is where we might be heading

Edwin

是的,是的,我认为有一个非常强大的概念,它帮助人们以更好的方式实现他们的想法。

Yeah yeah I think there's a very very powerful notion where it helps people just achieve their ideas in a in a much way.

背景与激增 Background and Surge

Host

我们还没有深入探讨的一个我觉得非常有趣的话题是你创办 Surge 的故事。你有一个非常独特的背景。我经常想到 Brian Armstrong,Coinbase 的创始人,他曾经做过一次演讲,让我印象深刻,他谈到了他非常独特的背景如何让他创办了 Coinbase。他有经济学背景,有密码学经验,然后他是一名工程师,这就像创办 Coinbase 的完美维恩图。我觉得你和 Surge 有非常相似的故事。谈谈你的背景,以及它如何引导你创办 Surge,从很久以前开始。

Something we haven't gotten into that I think is really interesting is just the story of how you got to starting Surge. You had uh you have a really unique background. I always think about these Brian Brian uh Armstrong the founder of Coinbase once had gave this talk that has really stuck with me where he kind of talked about how his very unique background allowed him to start Coinbase. you had like a economics background, he had a cryptography experience and then he was an engineer and it's got this like the perfect ven diagram for starting Coinbase and I feel like you have a very similar story with Serge talk about that your background there and how you led how that led to Serge going way back

Edwin

我小时候一直对数学和语言着迷,我去了 MIT,因为它显然是数学和计算机科学最好的地方之一,但也因为它是乔姆斯基的所在地。我在学校时的梦想实际上是找到一种连接所有这些不同领域的底层理论。然后我成为 Google、Facebook 和 Twitter 的研究员。我一次又一次地遇到同样的问题:不可能获得训练模型所需的数据。所以我一直坚信高质量数据的必要性。然后 2020 年 GPT-3 出现了,我意识到,如果我们想把事情提升到下一个层次,构建能够编码、使用工具、讲笑话、写诗、解决黎曼猜想和治愈癌症的模型,那么我们需要一个全新的解决方案。在这些公司工作时,总是让我抓狂的是,我们面前有完整的人类思维力量,而所有数据标注者都专注于像图像标注这样非常简单的事情。所以我想构建一些专注于这些高级复杂用例的东西,而不是那些简单的,这才能真正帮助我们构建下一代模型。所以,是的,我认为我在数学、计算机科学和语言学交叉领域的背景真的影响了我一直想做的事情。所以一个月后我创办了 Surge,我们的使命就是构建我认为推动 AI 前沿所需的用例。

I was always fascinated by math and language when I was a kid like I went to MIT because it's obviously one of the best places for math and CS but also because it's the homony chsky my dream in school was actually to find some underlying theory connecting all these different fields. And then I became a researcher at Google and Facebook and Twitter. And I just kept running into the same problem over and over again. It was impossible to get the data that we needed to train our models. So I was always a huge believer in the need for high quality data. And then GB3 came out in 2020 and I realized that yeah, if we wanted to take things to the next level and build models that could code and use tools and tell jokes and write poetry and solve the rebound hypothesis and cure cancer, then yeah, we were going to need a completely new solution. Like the thing that always drove me crazy when I was at all these companies was we had the full power of the human mind in front of us and all the data students out there were focused on really simple things like image labeling. So I wanted to build something focused on all these advanced complex use cases instead that would really help us build an extra generation models. So yeah, I think my background in kind of cross math and computer science and linguistics really really informed what I always wanted to do. And so I started Serge a month later with with our one mission to basically build the use cases that I thought were going to be needed to push the frontier of AI.

Host

你说一个月后。

And you said a month later.

动机与深度剖析 Motivation and Deep-Dive Analysis

Host

除了你正在取得的巨大成功之外,现在是什么在驱动着你?是什么让你保持动力,继续在这个领域构建东西?

What just kind of drives you at this point, other than just the epic success you're having? What keeps you motivated to keep building this and building something in this space?

Edwin

我觉得我骨子里是个科学家。我一直以为自己会成为一名数学或计算机科学教授,致力于理解宇宙、语言和沟通的本质。说来好笑,但我一直有个异想天开的梦想:如果外星人来到地球,我们需要弄清楚如何与他们沟通,我想成为那个和 Gaul 一起去的人,用各种高深的数学、计算机科学和语言学来破解它。所以即使在今天,我最喜欢做的事情就是每当有新模型发布时,我们会深入剖析模型本身。我会摆弄它,运行评估,比较它在哪些方面提升了、哪些方面退步了。我会写一份深入分析报告发给客户。其实挺有意思的,很多时候我们会说这是数据科学团队做的,但实际上往往只是我做的。我觉得我可以整天做这个。我很难整天开会。我不擅长销售,也不擅长做人们期望 CEO 做的那些典型事情。但我喜欢写这些分析,喜欢和研究团队一起讨论他们的想法。有时候我会一直聊到凌晨 3 点,就为了和某个研究团队的人打电话,深入挖掘模型。所以除此之外,我仍然可以整天亲自动手处理数据和科学工作。我认为驱动我的是,我希望 Serge 在 AI 的未来中扮演关键角色,而我认为这也是人类的未来。我们对数据、语言、质量以及如何衡量这一切、如何确保一切走上正轨有着独特的视角。而且我认为我们不受那些有时会把公司引向负面方向的影响所束缚。就像我之前说的,我们把 Serge 建得更像一个研究实验室,而不是典型的初创公司。所以我们关心好奇心、长期激励和学术严谨性,而不太关心季度指标和董事会演示文稿里好看的东西。所以我的目标是利用我们作为公司的所有这些独特之处,确保我们以对物种长期真正有益的方式塑造 AI。

I think I'm a scientist at heart. I always thought I was going to become this math or CS professor and work on trying to understand the universe and language and the nature of communication. It's kind of funny, but I always had this fanciful dream where if aliens ever came to visit Earth and we need to figure out how to communicate with them, I wanted to be the one to go with Gaul and I'd use all this fancy math and computer science and linguistics to decipher it. So even today, what I love doing most is every time a new model is released, we'll actually do a really deep dive into the model itself. I'll play around with it. I'll run evals. I'll compare where it's improved, where it's regressed. I'll create this really deep dive analysis that we send our customers. And it's actually kind of funny because a lot of times we'll say it's from a data science team, but often it's actually just for me. And I think I could do this all day. I have a very hard time being in meetings all day. I'm terrible at sales. I'm terrible at doing the typical CEO things that people expect you to do. But I love writing these analyses. I love jamming with a research team about what they're saying. Sometimes I'll be up until 3 a.m. just talking on the phone with somebody on a research team and digging into the model. So above that, I still get to be really hands-on working on the data and the science all day. And I think what drives me is that I want Serge to play this critical role in the future of AI, which I think is also the future of humanity. We have these really unique perspectives on data and language and quality and how to measure all this and how to ensure it's all going on the right path. And I think we're uniquely unconstrained by all of these influences that can sometimes steer companies in a negative direction. Like what I was saying earlier, we built Serge a lot more like a research lab than a typical startup. So we care about curiosity and long-term incentives and intellectual rigor, and we don't care as much about quarterly metrics and what's going to look good in a board deck. And so my goal is to take all these unique things about us as a company and use that to make sure that we're shaping AI in a way that's really beneficial for the species in the long term.

Host

我在这次对话中意识到的是,你和像你们这样的公司对 AI 的发展方向有多大的影响力。你们帮助实验室了解他们的差距和需要改进的地方,而且不仅仅是像 OpenAI 这样的公司的负责人,大家都认为他们是引领 AI 的人,但我在这里听到的是,你对事物的发展方向有很大的影响力。

What I'm realizing in this conversation is just how much influence you have and companies like yours have on where AI heads. The fact that you help labs understand where they have gaps and where they need to improve, and it's not just you know everyone looks at just like the heads of OpenAI and all these companies as they're the ones ushering in AI, but what I'm hearing here is you have a lot of influence on where things head to.

Edwin

是的,我认为有一个非常强大的生态系统,说实话,人们还不知道模型的发展方向,也不知道他们想如何塑造它们,以及他们希望人类在这一切的未来中扮演什么角色。所以我认为有很多机会可以继续塑造这场讨论。

Yeah, I think there's this really powerful ecosystem where honestly people just don't know where models are headed and how they want to shape them yet, and how they want humanity to kind of play a role in the future of all this. And so I think there's a lot of opportunity to just continue shaping this discussion.

目标函数的重要性 The Importance of Objective Functions

Host

顺着这个思路,我知道你对为什么这项工作对人类很重要、为什么如此重要有一个非常坚定的论点。谈谈这个吧。

Along that thread, I know you have a very strong thesis on just why this work matters to humanity and why this is so important. Talk about that.

Edwin

我在这里会有点哲学化,但我认为这个问题总是有点哲学性的。所以请耐心听我说。我们做的最直接的理解是训练和评估 AI。但我经常思考一个更深的使命,那就是帮助我们的客户思考他们梦想中的目标函数。比如,他们希望自己的模型成为什么样的模型?一旦我们帮助他们做到这一点,我们会帮助他们训练模型以达到那个北极星。我们会帮助他们衡量进展。但这真的很难,因为目标函数非常丰富和复杂。这有点像养孩子和问他们:“好吧,你想通过什么考试?你想让他们在 SAT 中得高分并写一篇很好的大学论文吗?”这是一个简化的版本,而你想让他们成长为什么样的人?如果他们快乐,无论他们做什么,你都会快乐吗?还是你希望他们上好学校并在经济上成功?同样,如果你接受这个概念,那么如何定义幸福?如何衡量他们是否快乐?如何衡量他们是否在经济上成功?这比简单地衡量你是否在 SAT 中得高分要难得多。我们做的是想帮助我们的客户达到他们梦想的星星,并找出如何衡量它们。所以我谈到了这个例子:当你要求模型写 50 封不同的邮件迭代时,你希望它们做什么?你是让它们再写 50 封,还是说“不,今天就到此为止,因为这已经足够完美了”?更广泛的问题是,我们是否在构建真正推动人类进步的系统?那么我们如何构建数据集来训练以实现这一点并衡量它?我们是否在优化所有这些错误的东西,比如那些消耗我们越来越多时间、让我们越来越懒惰的系统?是的,我认为这与我们做的事情非常相关,因为衡量和定义某件事是否总体上推动人类进步是非常困难和复杂的。相反,衡量点击量和点赞量这些代理指标很容易。但我认为这就是为什么我们的工作如此有趣。我们想要研究那些需要最难数据的困难而重要的指标,而不仅仅是容易的指标。所以我经常说的一句话是:你就是你的目标函数。所以我们想要达到复杂的目标函数,而不是这些简单的代理指标。我们的工作是找出如何获得匹配的数据。所以是的,我们想要数据。我们想要衡量 AI 是否让我们的生活更丰富的指标。我们想以这种方式训练我们的系统。我们想要让我们更好奇、更有创造力的工具,而不仅仅是更懒惰。这很难,因为人类天生有点懒惰。所以 AI 停止广播是获得参与度、让所有指标上升的最简单方式。所以我认为,选择正确的目标函数并确保我们朝着它们优化,而不仅仅是这些简单的代理指标,对我们的未来真的非常重要。

I'll get a bit philosophical here, but I think the question is always a bit philosophical. So bear with me. The most straightforward way of thinking about what we do is we train and evaluate AI. But there's a deeper mission that I often think about, which is helping our customers think about their dream objective functions. Like, what kind of model do they want their model to be? And once we help them do that, we'll help them train their model to reach that north star. We'll help them measure that progress. But it's really hard because objective functions are really rich and complex. It's kind of like the difference between having a kid and asking them, "Okay, what test do you want to pass? Do you want them to get a high score on the SAT and write a really good college essay?" That's a simplistic version versus what kind of person do you want them to grow up to be? Will you be happy if they're happy no matter what they do? Or are you hoping they'll go to a good school and be financially successful? And again, if you take that notion, it's like, okay, how do you define happiness? How do you measure whether they're happy? How do you measure whether they're financially successful? It's a lot harder than simply measuring whether or not you're getting a high score on the SAT. And what we're doing is we want to help our customers reach their dream stars and figure out how to measure them. And so I talked about this example of what you want models to do when you're asking them to write 50 different email iterations. Do you just continue them for 50 more or do you just say no, just move on with the day because this is perfect enough? And the broader question is are we building these systems that actually advance humanity? And so how do we build the data sets to train towards that and measure it? Are we optimizing for all these wrong things, just systems that suck up more and more of our time and make us lazier and lazier? And yeah, I think it's really relevant to what we do because it's very hard and difficult to measure and define whether something is generally advancing humanity. It's very easy to measure all these proxies instead, like clicks and likes. But I think that's why our work is so interesting. We want to work on the hard, important metrics that require the hardest types of data, not just the easy ones. So I think one of the things I often say is you are your objective function. So we want to reach complex objective functions and not these simplistic proxies. And our job is to figure out how to get the data to match this. So yeah, we want data. We want metrics that measure whether AI is making our life richer. We want to train our systems this way. And we want tools that make us more curious and more creative, not just lazier. And it's hard because humans are kind of inherently lazy. So AI stop radios are the easiest way to get engagement, make all your metrics go up. So I think this question about choosing the right objective functions and making sure that we're optimizing towards them and not just these easy proxies is really, really important to our future.

Host

哇。我喜欢你在这里分享的内容,它让你更加欣赏构建 AI、训练 AI 以及你所做工作的细微差别。

Wow. I love how what you're sharing here gives you so much more appreciation of the nuances of building AI, training AI, the work that you're doing.

构建Surge的反思 Reflections on Building Surge

Host

你知道,从外面看,人们可能只是看着 Surge 和这个领域的公司,觉得,好吧,他们只是在创造所有这些数据喂给 AI,但显然这其中有太多人们没有意识到的东西。我很高兴知道你是领头人,像你这样的人在如此深入地思考这件事。也许再问一个问题。你希望自己在创办 Surge 之前知道些什么?很多人创业时并不知道自己会陷入什么。有没有什么是你希望告诉早期自己的?

You know, from the outside, people could just look at Surge and companies in the space of, okay, well, they're just creating all this data feeding into AI, but clearly there's so much to this that people don't realize. And I love knowing that you're at the head of this, that someone like you is thinking through this so deeply. Maybe one more question. Is there something you wish you'd known before you started Surge? A lot of people start companies, they don't know what they're getting into. Is there something you wish you could tell your earlier self?

Edwin

是的。我绝对希望自己早知道,你可以通过埋头苦干、做伟大的研究、简单地构建出令人惊叹的东西来建立一家公司,而不是通过不断发推、炒作和融资。这有点好笑,但我从没想过要创业。我喜欢做研究,而且我一直是 DeepMind 的超级粉丝,因为他们是一家了不起的研究公司,被收购后仍然继续做惊人的科学。但我一直认为他们是那种神奇的 ILR 独角兽。所以我想如果我创业,我就得变成一个整天看财务、整天开会、做所有这些听起来极其无聊且我一直讨厌的事情的商人。所以我觉得这太疯狂了,结果根本不是那样。我仍然每天沉浸在数据中,而且我热爱它。我喜欢做所有这些分析,和研究人员交谈,这基本上是应用研究,我们在构建这些真正推动 AI 前沿的惊人数据系统。所以,是的,我希望自己早知道,你不需要把所有时间花在融资上,不需要不断制造炒作,不需要变成另一个人。你实际上可以通过简单地构建出足够好的东西来建立一家成功的公司,让它穿透所有噪音。我想如果我知道这是可能的,我会更早开始。

Yeah. So, I definitely wish I'd known that you could build a company by being heads down and doing great research and simply building something amazing and not by constantly tweeting and hyping and fundraising. It's kind of funny, but I never thought I wanted to start a company. Like, I love doing research and I was actually always a huge fan of DeepMind because they were this amazing research company that got bought and still managed to keep on doing amazing science. But I always thought that they were this magical ILR unicorn. So I thought if I started a company, I'd have to become a business person looking at financials all day and being in meetings all day and doing all this stuff that sounded incredibly boring and I always hated. So I think it's crazy that didn't end up being true at all. Like I'm still in the weeds in the data every day. And I love it. Like I love that I get to do all these analyses and talk to researchers and it's basically applied research where we're building all these amazing data systems that really push the frontier of AI. So yeah, I wish I knew that you don't need to spend all your time fundraising. You don't need to constantly generate hype. You don't need to become someone you're not. You can actually build a successful company by simply building something so good that it cuts through all that noise. And I think if I'd known this was possible, I would have started even sooner.

数据标注的最后思考 Final Thoughts on Data Labeling

Host

我希望这是一个绝妙的结束点。我觉得这正是创始人需要听到的。我认为这次对话会激励很多创始人,尤其是那些想以不同方式做事的创始人。在我们进入非常激动人心的快问快答环节之前,你还有什么想分享的吗?还有什么想留给听众的吗?我们涵盖了很多内容。当然,你也可以说没有。

I hope this is such an amazing place to end. I feel like this is exactly what founders need to hear. And I think this conversation is going to inspire a lot of founders, and especially a lot of founders that want to do things in a different way. Before we get to our very exciting lightning round, is there anything else you wanted to share? Anything else you want to leave our listeners with? We covered a lot of ground. It's totally okay to say no as well.

Edwin

我想最后要说的是,我认为很多人把数据标注想得非常简单,比如给猫的照片贴标签、在汽车周围画边界框。所以我其实一直讨厌“数据标注”这个词,因为它描绘的画面太简单了,而我认为我们在做的事情完全不同。我经常把我们做的事情想得更像养育孩子。你不仅仅给孩子喂信息,你还教他们价值观、创造力、什么是美,以及那些让一个人成为好人的无数微妙之处。这就是我们为 AI 所做的。所以,是的,我经常把我们做的事情看作人类的未来,或者我们如何养育人类的孩子。我就说到这里。

I think the thing I would end with is I think a lot of people think of data labeling as really simplistic work like labeling cat photos and drawing bounding boxes around cars. And so I've actually always hated the word data labeling because it just paints this very simplistic picture when I think what we're doing is completely different. Like I think a lot about what we're doing as a lot more like raising a child. You don't just feed a child information. You're teaching them values and creativity and what's beautiful and these infinite subtle things about what makes somebody a good person. And that's what we're doing for AI. So I yeah I just often think about what we're doing as almost like the future of humanity or how are we raising humanity's children. So I'll leave it at that.

快问快答:书籍推荐 Lightning Round: Book Recommendations

Host

哇。我喜欢这次对话中有这么多哲学思考,我没想到。那么,Edwin,我们到了非常激动人心的快问快答环节。我有五个问题要问你。准备好了吗?

Wow. I love just how much philosophy there is in this whole conversation that I was not expecting. With that, Edwin, we've reached our very exciting lightning round. I've got five questions for you. Are you ready?

Edwin

是的。开始吧。

Yep. Let's go.

Host

我们开始。你发现自己最常推荐给别人的两三本书是什么?

Here we go. What are two or three books that you find yourself recommending most to other people?

Edwin

是的。我经常推荐的三本书,第一本是 Ted Chiang 的《你一生的故事》。这是我一直以来最喜欢的短篇小说。它讲的是一个语言学家学习外星语言的故事。我显然每隔几年就会重读一遍。

Yes. So, three books I often recommend are first, Story of Your Life by Ted Chiang. It's my all-time favorite short story. And it's about a linguist learning an alien language. And I obviously reread it every couple years.

Host

这就是《星际穿越》的内容。是不是?

And that's what the Interstellar was about. Is that is that

Edwin

是的。有一部电影叫《降临》。

Yeah. So, there's a movie called Arrival.

Host

《降临》,

Arrival,

Edwin

它是根据那个故事改编的,我也很喜欢。

which is based off of the story, which I love as well.

Host

太好了。好的,继续。

Great. Okay, keep going.

Edwin

然后第二本是加缪的《西西弗神话》。我其实无法解释为什么喜欢它,但我总觉得最后一章不知怎的非常鼓舞人心。第三本是 Douglas Hofstadter 的《Le Ton beau de Marot》。我认为《哥德尔、艾舍尔、巴赫》是他更著名的书,但我其实一直更喜欢这本。它基本上拿一首法语诗,用 89 种不同的方式翻译,并讨论每种翻译背后的所有动机。我一直喜欢它体现的理念:翻译不是机械的事情。相反,有无数种思考什么构成高质量翻译的方式,这模仿了我思考数据和 LLM 质量的很多方式。

And then second, Myth of Sisyphus by Camus. I actually can't really explain why I love this, but I always find the final chapter somehow really inspiring. And then third, Le Ton beau de Marot by Douglas Hofstadter. And so I think Gödel, Escher, Bach is his more famous book, but I've actually always loved this one better. It basically takes a single French poem and translates it 89 different ways and discusses all the motivations behind each translation. And so I've always loved the way it embodies this idea that translation isn't this robotic thing that you do. Instead, there's a million different ways to think about what makes a high quality translation, which mimics a lot of ways I think about data and quality in LLMs.

快问快答:最爱影视 Lightning Round: Favorite Movie or TV Show

Host

所有这些都和我们一直在谈论的事情深深共鸣,尤其是第一本,如果那是你毕业后的目标,比如,我想帮助翻译外星语言。我不惊讶你喜欢那个短篇小说。下一个问题。你有最近特别喜欢看的电影或电视剧吗?

All these resonate so deeply with the way with all the things we've been talking about, especially that first one, if that was your goal after school is like, I want to help translate alien language. I'm not surprised you love that short story. Next question. Do you have a favorite recent movie or TV show you've really enjoyed?

Edwin

我最近发现的一部新的最爱电视剧叫《旅行者》。它基本上是关于一群来自未来的旅行者被送回过去阻止世界末日的。所以,我真的很喜欢科幻。然后我其实刚重看了《超时空接触》,这也是我一直以来最喜欢的电影之一。所以,是的。我想你会注意到我的一个特点是,我喜欢任何涉及科学家努力破译外星通信的书或电影。

One of my new all-time favorite TV shows is something I found recently. It's called Travelers. It's basically about a group of travelers from the future who are sent back in time to prevent the apocalypse. So, I just really like science fiction. And then I actually just rewatched Contact, which is also one of my all-time favorite movies. So, yeah. I think one of the things you'll notice about me is that yeah, I love any kind of book or film that involves scientists suffering deciphering alien communication.

Host

再次,这只是我小时候一直有的梦想。

Again, just this dream I always had as a kid.

Edwin

太有趣了。我喜欢。好的。

That's so funny. I love that. Okay.

快问快答:最爱产品 Lightning Round: Favorite Product

Host

你最近有没有发现一个你非常喜欢的产品?

Is there a product you recently discovered that you really love?

Edwin

所以,这很有趣,但我这周早些时候在旧金山,终于第一次坐了 Waymo。老实说,那太神奇了,真的感觉像生活在未来。

So, it's funny, but I was in SF earlier this week and I finally took a Waymo for the first time. Honestly, it was magical and it really felt like living in the future.

Host

是的。就像人们疯狂炒作的东西,但它总是超出你的预期。

Yeah. It's like the thing that people hype it like crazy, but it always exceeds your expectations.

Edwin

值得炒作。太疯狂了。

Deserves the hype. It was crazy.

Host

是的。太离谱了。就像天哪。如果你不在旧金山,你不会意识到这些东西有多普遍。它们到处都是。无人驾驶汽车不停地穿梭,当你去参加活动结束时,就有所有这些 Waymo 排着队接人。

Yeah. It's absurd. It's like holy moly. Like if you're not in SF, you don't realize just how common these things are. They're just like all over the place. Just driverless cars constantly going about and when you go to an event at the end there's just like all these Waymos lined up picking people up.

Edwin

是的。

Yep.

Host

是的。Waymo,干得好。那边干得好。

Yeah. Waymo, good job. Good job over there.

快问快答:座右铭 Lightning Round: Life Motto

Host

你有没有一个最喜欢的人生格言,在工作或生活中经常想起?

Do you have a favorite life motto that you find yourself coming back to in work or in life?

Edwin

所以,我想我提到过这个想法:创始人应该建立一家只有他们才能建立的公司,就像这是他们的整个生活、经历和兴趣塑造他们走向的命运。所以,我认为这个原则适用得很广泛,不仅适用于创始人,也适用于创造事物的人。

So, I think I mentioned this idea that founders should build a company that only they could build almost like it's this destiny that their entire life and experiences and interests shape them towards. And so, I think that principle applies pretty broadly, not just to founders, but to people creating a thing.

打造独特体验的建议 Advice on Building Unique Experiences

Host

好,让我顺着这个话题继续深挖一下。你有没有什么建议,关于如何积累那些能带来这种成果的经历?是不是就是追随你感兴趣的东西?因为你知道,说起来容易,但真正获得那些独特的经历,从而创造出真正重要的东西,其实很难。

Well, let me follow that thread to unlightening this answer. Uh, do you have any advice for how to build those sorts of experiences that help lead to that? Is it, you know, follow things that are interesting to you? Cuz, you know, it's easy to say that it's hard to actually acquire these really unique sets of experiences that allow you to create something really important.

Edwin

是的。所以,我认为始终要真正追随你的兴趣,做你热爱的事。这几乎就像我在 Surge 做的很多决定一样。比如,几年前我没想过的一件事,后来有人跟我提过,那就是公司在某种意义上就是 CEO 的化身。这挺有趣的。我之前没这么想过,因为我一直不太清楚 CEO 是做什么的。我总觉得 CEO 挺普通的,就是按照副总裁、董事会什么的指示行事,对决策点头就行。但实际上,当我面临一些重大艰难决策时,我不会想公司会怎么做,不会想我们要优化什么指标,我只想我个人在乎什么,我的价值观是什么,我希望世界上发生什么。所以,我认为遵循这个想法,问问自己你在乎什么价值观,你想塑造什么,而不是什么在仪表盘上好看,我觉得结果相当重要。

Yeah. So, I think it would always be to really follow your interests and do what you love. And it's almost like a lot of decisions I make about Surge, like I think one of the things that I didn't think about a couple years ago, but then someone said it to me, it's that companies in a sense are an embodiment of their CEO. And it's kind of funny. I hadn't thought about that because I never quite knew what a CEO did. I always thought a CEO was kind of generic and it's like, okay, you're just doing whatever your VPs and your board and whatever tell you to do and you're just saying yes to decisions. But instead it's this idea where when I think about certain big hard decisions we have to make I don't think what would company do I don't think what metrics are we trying to optimize I just think what do I personally care about like what are my values and what do I want to see happen in the world and so I think following that idea about okay so ask yourself what are what are the values you care about what are things you're trying to shape and not what will look good on a dashboard I I think that results were pretty important.

Soda与Pop地图 The Soda vs Pop Map

Host

我喜欢你总是充满无穷无尽、美丽又深刻的回答。最后一个问题。在创办 Surge 之前,你相当出名的一件事是,你在 Twitter 时构建了一张地图,展示了世界地图,以及人们怎么称呼它,是叫 soda 还是 pop。我不知道它叫 soda 还是 pop。这张地图叫什么名字?

Uh I love how just you're just full of endless beautiful and very deep answers. Final question. Something that you were quite you got quite famous for before starting Surge is you built this uh map at Twitter while you were at Twitter that showed the uh a map of the world and how and what people called whether they called it soda or pop. I don't know if it was called soda or pop. What was the name of this map?

Edwin

是的,它叫“soda 对比 pop”数据集,或者“soda 对比 pop”地图。它是一张美国地图,告诉你哪些地方的人说 pop,哪些地方的人说 soda。那么,你说 soda 还是 pop?

Yeah, it was like the soda versus pop data set or soda versus pop map. And so it's like a map of the United States and it tells you where people say pop versus soda. So uh do you say soda or pop?

Host

我说 soda。我是 soda 派的人。

So I say I say soda. I'm a soda person.

Edwin

好的。那这是不是就是正确答案,还是说无论你是什么都完全没问题。

Okay. And is that just like that's the right answer or it's like it whatever you are it's totally fine.

Host

我想我会有点奇怪地看着你。你说 pop,我会好奇你从哪来的。但我不会太鄙视你。

I think I'll look at you a little bit funny. You say pop and I'll wonder where you came from. But I I won't I won't scorn you too much.

Edwin

我也是这么觉得。

That's how I feel too.

结语与联系方式 Closing and How to Connect

Host

Edwin,这太棒了。这是一次非常精彩的对话。我学到了很多,我想这会帮助很多人创办自己的公司,帮助他们的公司更符合他们的价值观,并且打造更好的东西。最后两个问题,人们如果想联系你,在网上哪里可以找到你?你在招聘什么职位?听众怎样才能帮到你?

Edwin, this was incredible. This was such an awesome conversation. I learned so which I think we're going to help a lot of people start their own companies, help their companies become uh more aligned with with their values and and just building better things. Two final questions, where can folks find you online if they want to reach out? Uh what roles are you hiring for? How can listeners be useful to you?

Edwin

是的,我以前很喜欢写博客,但过去几年没时间,不过我现在又开始写了。所以一定要去看看 Surge 的博客,surgehq.ai/blog。希望我能在那里写更多。我想说我们肯定一直在招人。所以对于热爱数据、热爱数学、语言和计算机科学交叉领域的人,一定要联系我。随时联系。

Yes, so I used to love writing a blog, but I haven't had time in the past few years, but I am starting to write again. So definitely check out a Serge blog, surgehq.ai/blog. And yeah, hopefully I'll be writing a lot more there. And I would say we're definitely always hiring. So for people who just love data and people who love this intersection of math and language and computer of computer science, definitely reach out. Reach out anytime.

Host

太棒了。那听众怎样才能帮到你?是不是就是……我不知道。有什么要求吗?

Awesome. And how can listeners be useful to you? Is it just I don't know. Yeah. Is there anything there? Any asks?

Edwin

我想说,一定要告诉我你希望我写的博客主题。

So I would say definitely tell me blog topics you like me to write about.

Host

好的。

Okay.

Edwin

然后,我一直对现实世界中发生的各种 AI 失败案例很着迷。所以每当你遇到一个非常有趣的失败案例,我认为它说明了关于我们希望模型如何表现的深层问题,比如模型在那里可以有多种不同的回应方式。我经常觉得没有唯一正确答案。所以每当有这样的例子,我就很喜欢看到。

And then I'm always fascinated by all of these AI failures that happen in the real world. So whenever you come across a really interesting failure that I think illustrates some deep question about how we want models to behave, like there's just so many different ways a model can respond there. I just often times think there just not a single right answer. And so whenever there's one of these these these examples, I I just love seeing them.

Host

你得把这些分享到你的博客上。我也很想看到这些。Edwin,非常感谢你来做客。

You need to share these on your blog. I'm also I would love to see these. Edwin, thank you so much for being here.

Edwin

谢谢。

Thank you.

Host

大家再见。

Bye everyone.

Host

非常感谢你的收听。如果你觉得这期节目有价值,你可以在 Apple Podcasts、Spotify 或你最喜欢的播客应用上订阅本节目。另外,请考虑给我们评分或留下评论,这真的能帮助其他听众找到这个播客。你可以在 lennispodcast.com 找到所有过往节目或了解更多关于本节目的信息。下期再见。

Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennispodcast.com. See you in the next episode.

互动版:逐字朗读 + 针对本期提问 →