Simulating 8 Billion Lives: AI for Wicked Problems
打开互动全文版(中英对照 + 朗读 + 问答)→AI 模拟 80 亿人能帮助解决气候变化等棘手问题吗?
Can AI simulation of 8 billion people help solve climate change and other wicked problems?
我们能否创造一个容纳 80 亿人生活在地球上的模拟?我认为这非常有趣,而这正是我们的愿景。一旦达到那种状态,你能为社会解答的问题类型也会随之改变。对我来说,这些问题包括:我们能否帮助解决气候变化?如果你把气候变化看作一个问题的空间,这就是我们社会科学家常说的“棘手问题”——许多行动者带着相互竞争的激励,试图做出一个非常复杂的协调决策。模拟能帮助我们解决这个问题吗?
Can we create a simulation of 8 billion people living on earth? I think that's quite interesting and that really is the vision and once you get to that kind of state the kind of questions that you can help answer for the society also start to change from my perspective and for me it's questions like can we help solve climate change. If you look at climate change as a problem space, this is what we like social scientists would often call the wicked problems problem where you have many actors with competing incentives for trying to make a very complex decision a coordinating coordination decision. Can simulation help us solve that?
在进入今天的正题之前,我有一小段话想对听众说。谢谢你们。如果不是你们选择点击并收听我们的内容,我们就不可能为你们带来这些你们显然想要的 AI 工程、科学和娱乐内容。几乎每天都有赞助商找上门来。但幸运的是,有足够多的你们真正订阅了我们,让我们能够在没有广告的情况下维持这一切,而我们希望保持这种状态。但我只想请大家帮一个忙。你们能做的最有力、完全免费的一件事,就是点击那个订阅按钮。这是我对你们唯一的请求。这对我以及我的团队意义重大,他们每周都非常努力地为大家带来 Inspace 的内容。如果你们订阅了,我保证我们会永不停止地努力,让节目变得更好。现在,让我们开始吧。今天,我们请到了 June 来到播客。很高兴能开始这一期。这是一家非常令人兴奋的公司。我想先问你一个问题:跟我们讲讲你的人生故事吧。你是怎么走到今天的?
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it. Today, we have June in the podcast. Excited to kick this one off. Very exciting company. I want to kick off and ask you the question, you know, talk us through the story of your life. How have you gotten here?
对,当然。真的很高兴能来到这里。我的人生故事:我出生在韩国,在那里生活了大约 11 年。然后我们全家搬到了波士顿,那是我 11 岁的时候。我的父母都是医生,他们当时在做博士后研究。我父亲是外科医生,他在波士顿儿童医院做学术休假。所以我在那里长大。其实我离科技并不近。我更像是一个音乐、艺术、绘画那种类型的人。我其实是在高中稍晚一些时候才开始画画的,但那是我过去常做的事。离开韩国后,我大部分时间在东海岸长大。我在新罕布什尔州住了好几年,然后去宾夕法尼亚州上大学。在大学里,我更多地接触了科技领域。我最初是接受艺术家训练的,我确实认为那会是我的职业道路。所以那不是爱好,而是真的想靠这个谋生。后来我逐渐对这个想法产生了浓厚兴趣:最伟大的艺术家往往创造自己的媒介,而我们今天拥有的最佳媒介其实是计算。所以我决定深入探索,事情一件接一件,显然我们可以深入聊,但我逐渐对研究产生了兴趣,然后我就走到了今天。
Right. Yeah, for sure. So really excited to be here. A story of my life. So I was born in Korea. And I lived there for a good 11 years or so of my life. And then my family moved to Boston. So we moved when I was 11. And my parents were doctors. So they were basically going through their post-doctoral studies. My dad was a surgeon. So he was doing his sabbatical years actually at the Boston Children's Hospital. So I grew up there. Not too close to tech actually. I was very much like, you know, music, artsy, painting, like that kind of guy. I actually got into painting a little bit later in high school. But that's what I used to do. And then I grew up mostly in the east coast after Korea. So I lived a good number of years in New Hampshire and then I went to college in Pennsylvania. And I got into more of this tech scene in college. So I was originally trained to be an artist. I actually thought that would be my actual professional career. So it wasn't a hobby, it was actually like, hey, let's make a living out of this. And then gradually I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was actually in computation. So I decided to go deeper into that, and one thing led to another, and obviously we can go deeper into this, but I decided that research was something that gradually I got interested in, and here I am.
显然,你在研究方面投入了很多。你有一篇 2023 年最佳论文,就是生成式智能体论文,通常被称为“小维尔”论文。
So there's obviously a lot that you packed into the research component. You had one of the best papers of 2023 which was the generative agents paper commonly known as the smallville paper.
是的。
Yeah.
你可以随意回顾你提到的其他内容,但大多数人显然是从这篇论文认识你的。你有没有关于多少人读过它的统计数据?Archive 会给你一些数据,对吧?
Feel free to call back to anything else that you mentioned, but most people would have heard of you from this obviously. Do you have any statistics of how many people have like read it? Archive gives you something, right? Some stats.
是的,这是个好问题。有多少人读过?我其实不确定。我知道我们确实在跟踪引用次数,我知道它涨得很快,但读者数量,Google
Yeah, it's a good question. How many people have read it? I'm actually not sure. I know this, that we do keep track of the number of citations which I know is going up quite fast, but the readership the Google
Google Scholar 显示有 72,000 次引用。
Google Scholar has 72,000.
它引起了更大的轰动,实际上是一篇相当有影响力的论文。它被引用了很多次。
It made a bigger hit and it was actually a pretty instrumental paper. It was like one that got cited so many times.
它经常被提及,当人们问年度最佳论文,或者你最近读过的最佳论文时,就是这一篇。
It is frequently like when people ask what is the best paper of the year like best paper you've read recently it's that's this one.
我觉得记忆组件被低估了,你知道,一个非常好的早期记忆系统,但确实是最大的论文之一。
I thought the memory component was pretty underrated, you know, like very good early memory system, but yeah, one of the biggest papers, you know.
是的,是的。那么也许我可以谈谈这篇论文是如何成型的。
Yeah. Yeah. Yeah. So maybe I can talk a little bit about how this particular paper came together.
当我开始做研究时,那是 2020 年,我在斯坦福开始攻读博士学位。那一年我们即将迎来 GPT-3.5,GPT-3 即将可用。我们已经有 GPT-2,你能感觉到市场上正在出现一类新的模型,团队对此非常感兴趣。普遍的共识是:这个模型真的能对任何事有用吗?很奇怪,这些模型并没有被训练来做任何特定任务,但我们决定赌一把。于是斯坦福的一大群学者,实际上由我的联合创始人之一 Percy Liang 领导,聚集在一起
So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year when we were about to get GPT-3.5, GPT-3 to be available. So we already had GPT-2 and you could sense that there's this new class of models that was just becoming available in the market, and the team got very intrigued, and the general consensus was, well, is this model actually going to be useful for anything? It's really strange that these models are not trained to do any particular task, but we decided to take a bet. So a large group of scholars at Stanford, and it was actually led by one of my co-founders, Percy Liang, came together
他创造了“基础模型”这个词
who coined foundation models
他创造了“基础模型”这个词,我们写了那篇论文,这个词就出自那里,叫《基础模型的机会与风险》。在这个过程中,我开始深入思考的是:这是一个在我们的生态系统中根本全新的模型。它之所以新,是因为它同样没有被训练来做任何特定的事情,但它的前提是它可以做任何事、一切事。如果用生物学类比,它就像一个干细胞。我对这个想法非常感兴趣:如果我们真正思考这项技术会催生哪些杀手级应用,那会是什么?我的许多同事用它来做更简单的分类、简单的生成。有趣的是这些模型能做到这些,但从交互的角度来看,并不那么有趣。我们几十年前就知道怎么做了。我们最终得出的结论是:这些模型实际上是在来自网络的非常广泛的数据上训练的,对吧?所以这些是人类行为数据。是社交媒体、维基百科,所有这些数据。所以如果你从正确的角度去戳它,你就能看到人类行为浮现出来。那实际上非常逼真,我们以前从未见过。
who coined the term foundation model, we wrote this paper where that term came from, called 'Opportunities and Risks of Foundation Models'. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn't again trained to do anything in particular, but it was its premise was it could do anything and everything. It was like a stem cell if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simpler classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We've known how to do that for many decades. And what we came down to was these models are actually trained on this very broad data from the web, right? So these are human behavioral data. It's social media, Wikipedia, all these kind of data. So if you poke at the right angle, then you could see human behavior that would just pop out. That's actually quite realistic and we've never seen that before.
我们决定做的这个练习,是我们这群同事——我、迈克尔·伯恩斯坦、PC·杨,他后来成了我在 Simile 的联合创始人——坐下来玩一个我们称之为“时间机器游戏”的游戏。想象我们坐上时间机器,快进 10 年,然后回头看:哪一项应用会是最重要、最有趣、最鼓舞人心的?我们想,如果我们能直接重建我们所生活的世界呢?我的意思是,很难有比这更宏大的野心了。比如,让我们创造一个世界。这就是我们的起点。最初,我们有一篇论文,是生成式智能体论文的前身,叫《社会模拟》(Social Simulacra)。
The exercise that we decided to do, and this is something that this particular group of colleagues—myself, Michael Bernstein, PC Young, who ended up becoming my co-founder at Simile—we sat down and played this game we call the time machine game. Imagine we were to get on a time machine, fast forward 10 years, and look back. What would have been the single application that would have mattered, that would be the most interesting and inspiring? Well, we thought, what if we could just recreate the world that we live in? I mean, it's really hard to get more ambitious than that. Like, let's just create a world. And that's where we started. Initially, we had this paper that was a precursor to the generative agents paper, called Social Simulacra.
在你继续之前,在时间机器练习中,还有其他最雄心勃勃的候选项目吗?可能会是什么?接下来是什么,你知道,第二名或第三名是什么?
Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What could have been? What were the next, you know, what was number two or number three?
好的。我们考虑过一个非常接近的第二名,它基本上演变成了更多这类自动化工具,但尤其是关于真正个性化的智能体、真正为你做事的愿景。
Okay. So there is a close second that we were considering, which basically ended up becoming more of these automation tools, but especially the vision around really personalized agents that actually do things for you.
那也在发生。
That's also happening.
它也在发生,但对我们来说这很有趣,对吧?我们决定采用模拟这个想法的原因是——我的意思是,我是个超级科幻迷,而创造模拟这个想法我个人真的非常着迷。我喜欢这个想法。看到一个像这样的游戏小镇,看着这些智能体生活在其中,真的很酷。但与此同时,我的赌注是,如果你想用这项技术创造一个真正出色的个人助理,你首先需要的其实是一个出色的用户模型。比如,我告诉模型:嘿,你能帮我去买晚餐吗?它点了夏威夷披萨,而我不喜欢披萨上有菠萝。那它就完全失败了。它不犯这个错误的唯一方法,就是对我有深刻的理解。我这里举了一个非常简单、愚蠢的例子,但你可以想象这种对人的核心理解是多么重要。比如,如果我们有家人或最亲密的朋友,他们对我们是怎样的人有一个良好的心智模型。那是我们社交联系的基础。所以我们的赌注也是,围绕模拟、创造准确的人类表征的技术,应该先于那些将自动化我们所生活的世界的更复杂的智能体。这就是我们的赌注。但那是一个非常接近的第二名,我仍然对它非常着迷。我认为有很多有趣的工作正在进行。不过,我在这里的热辣观点是,我认为我们还没有看到一个真正的个人助理,在真正符合那条工作线的雄心的方式上真正有用。我认为有一些早期应用显然很有趣,即使你现在和 ChatGPT 或 Claude 交谈,它们显然也对我们了解很多。所以我认为它做的很多生成都更加定制化了,但我认为那个领域的野心非常大,而且我认为我们还没有完全具备所有正确的要素。
It's also happening, but it was sort of interesting for us, right? The reason why we decided to go with the idea of simulation—I mean, I was a huge science fiction nerd, and this idea of creating simulation I was personally really just fascinated by. I love the idea. It's really cool to see like a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you actually need first is an amazing model of your users. So for instance, I told the model, hey, can you go buy late dinner for me? And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, for instance, if we have our family or closest friend, they have a good mental model of who we are. That's the basis of our social connection. So our bet also was that this technology around simulation, creating accurate representations of people, ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I'm still very much fascinated by it. I think there's a lot of interesting work that's going around. My hot take actually here, though, is I don't think we've actually seen a true personal assistant that's actually useful in ways that actually meets the ambition of that particular line of work. I think there are early applications that are obviously interesting, and if you talk to even ChatGPT nowadays or Claude, they obviously know a lot about us. So a lot of the generation it's doing, I do think it's much more tailored, but I think the ambition is quite large in that field, and I don't think we quite have all the right ingredients just yet.
那么像 OpenClaw 这样的客户端个人智能体,你希望看到它们拥有哪些目前没有的东西?
So like OpenClaw, all these client personal agents—what do you want to see from them that they don't currently have?
我确实认为它在慢慢接近,但我通常希望它们对人能有更深刻的理解。现在你看看这些模型——我的意思是,OpenClaw 基本上利用的就是一个 Markdown 文件,我认为这相当聪明,对吧?所以如果你看看生成式智能体的论文,这实际上和我们当时的直觉是一样的。最初我们为生成式智能体创建记忆架构时——那是在 2022 年左右,所以我们甚至还没有智能体式架构或“智能体”这个词的概念——但我们与今天一些工作共享的直觉是,我们最初想:我们是否要把记忆做成知识图谱?我们是否要训练一个定制模型?诸如此类。而我们决定的是:不,不,忘掉这一切。这些语言模型实际上非常擅长建模文本、理解和推理文本,所以把所有东西放在一个 Markdown 文件或文本文件里,就完成了。我觉得我们能这样做很有趣,而且这样做有很多优点。但也有局限性。那就是你检索和理解极其庞大的数据的方式。这需要很多工作。所以我认为这项技术正在变得更好。不过,我也确实认为,有些事情你无法仅仅通过提示模型来塑造。所以在某种程度上,你确实需要触及模型本身的参数。所以我认为这类工作确实需要发生,而且显然它正在发生。问题是我们能把它推进多远?我们如何获取数据,以及你如何创建一个生态系统,让人们持续向这个模型提供数据?让它了解他们。
I do think it's slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now you look at the models—I mean, OpenClaw, what it's basically leveraging is basically a markdown file, and I think it's quite clever, right? So if you look at the generative agents paper, this actually was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents—and this is like back in 2022, so we didn't really quite have the idea of even agentic architecture or the term agent—but the intuition that we shared with some of the work that's coming out today was we initially thought, well, do we want to make the memory into, let's say, a knowledge graph? Do we want to train a bespoke model? All these kind of things. And what we decided to do was no, no, just forget about all this. These language models are actually quite good at modeling text and understanding and reasoning about text, so just put everything in a markdown file or text file, you're done. I thought that was quite interesting that we could do that, and there's a lot of strength in doing that. But also there is limitation. It's the way you retrieve and make sense of data that's extremely large. It takes a lot of work. So I think that technology is getting better. I also do, however, think there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there's these kind of work that I do think does need to happen, and obviously it is happening. The question is how far can we take it? How do we source data, and how do you also create an ecosystem where the people are continuously feeding data to this model? So it's learning about them.
为什么需要在模型内部做这件事,背后的直觉是什么?
What's the intuition between why you need to do it in the model?
关于何时训练甚至后训练一个模型,而不是仅仅提示一个模型,我背后的直觉是:如果模型必须学习它运作所在的世界的底层物理,那么它就必须学习新的社会物理。在它不需要训练的地方,是它已经拥有物理的地方。我们信任那个物理。它已经有了基础统计,但它只是在试图对环境做出反应。那么我认为你可以仅仅通过提示来让它产生行动。我认为模型还没有——至少是公开的模型还没有——学到人类完整的社会物理映射。而这实际上是 Simile 的核心论点之一,对吧?之所以如此,一个核心原因是,如果你看看模型训练所用的数据,这些模型是在网络数据上训练的,比如网络上可用的任何东西。这些是非常有趣的数据集,但它们从根本上说是自我暴露的态度数据,夹杂着一些行为数据。它还没有学到人们真正深层的行为本质。不仅仅是人们说他们在网上做什么,而是他们在现实生活中实际做什么。而这实际上是我认为我们尚未捕捉到的人类黑暗知识之一。正是这些数据也需要被纳入模型创建中。
My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train is where it already has the physics. We trust the physics. It already has the base statistics, but it's just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don't think the model has yet—at least the models that are out in the open—has yet learned the complete mapping of social physics of humanity. And this actually is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like whatever was available on the web. And these are really interesting datasets, but they are fundamentally the self-exposed attitudinal data with some behavioral data sprinkled around here and there. And it has yet to learn really deep behavioral nature of people. Not just what people say they do online, but what they actually do in real life. And this is actually one of the sort of what I would consider to be the dark knowledge of humanity that we haven't quite captured. And it's these kind of data that would also need to get factored into the model creation.
你称之为行为基础模型。
You call it behavior foundation model.
这里有个不错的总结,但除此之外,你需要什么类型的数据?你在模型层面做了什么改变?你实际上是如何建模行为基础模型的?
There's a good oneliner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about actually modeling a behavior foundation model?
我们把数据分为三类。第一类是访谈数据。例如,定性数据,丰富的定性数据很有趣。它不是行为数据,但我们确实会问人们:“嘿,给我讲讲你的人生故事。”
We think about data in three buckets. One bucket is interview data. For instance, it's quite interesting—qualitative, rich qualitative data is interesting. It's not behavioral, but we would literally ask people, 'Hey, tell me the story of your life.'
是啊,我们在这里做的就是这件事。
Yeah, that's what we're doing here.
没错。你们在采访开头问的那个问题,也正是我们会问的问题。显然,我们会让参与者比我说得更深入一些。也许我可以借此多讲一些我的人生故事。但这些数据之所以有趣,是因为通过了解关于人的这些长尾信息,你实际上能获得很多关于这个模型的质感——就像把这个人当作一个模型。所以,即使了解他们的童年记忆、创伤、初恋,这些信息在难以预测的方面都相当有启发性。这是第一类。
Exactly. The question you all asked at the beginning of this interview is literally the question we also ask. Obviously, we ask our participants to go a little bit deeper than how far I went. Maybe I can actually give more of my life story in light of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you actually get a lot of texture around this model—like this person as a model. So even understanding their childhood memory, or even their trauma, their first love—these kinds of things are quite informative in ways that are really hard to predict. So that's one.
然后是我认为的行为数据的两部分。一部分是观察性行为数据。这些可能是交易数据,或者通过抓取网页获得的数据。你可以想象为什么这些数据集会很有趣——它们提供了人们行为的基础统计。
Then there are two tranches of what I would consider behavioral data. One behavioral data is observational. These might be transaction data, or data you can get by scraping the web. You can imagine why these datasets would be interesting—they give you the base statistics of people's behavior.
但还有最后一类数据,我个人认为可能是最重要的:描述人们行为原因和机制的数据。其中一部分来自访谈数据,即定性数据,因为人们会谈论他们为什么做出某些决定。但真正能看到行为最大方面的地方是随机对照试验,比如 RCT。想象一下,你基本上有相同的设置,但有几个不同的变量试图调整。你能从中得到真实的人类行为吗?想象你必须做一个特定选择,比如是否喝咖啡。你喝咖啡的那天和没喝的那天——你的行为会改变吗?这就是描述原因或机制的数据集。这在建模人类时非常重要。
But then there's the last category of data, which I personally think is perhaps the most important: the data that describes the cause and mechanism—the whys—of people. Some of this is covered by the interview data, the qualitative data, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized control trials, like RCTs. Imagine you basically have the same setup but you have a few different variables that you're trying to tweak. Can you actually get realistic human behavior out of it? Imagine you had to make a particular choice, like whether to drink coffee or not. The day you drank coffee versus the day you didn't—does your behavior change? That's a dataset that describes a cause or mechanism. This is quite important in actually modeling people.
这之所以重要,是因为很多时候人们来找我们——或者不仅仅是找我们,而是人们对模拟感兴趣的原因——并不是因为他们想预测未来。如果你想在股市中获胜,预测未来很有趣。但大多数决策者想知道的是我们如何塑造未来。听到你的销售额将在两个季度后暴跌,对你并没有帮助。他们只会说:“哇,真糟糕。”他们想知道的是:“那么,我们现在需要做什么来避免那个未来?”这就是因果机制。
The reason why this is important is often times when people come to us—or not just to us, but the reason why people are interested in simulation—isn't because they want to predict the future. If you're trying to win against the stock market, predicting the future is interesting. But most decision makers want to know how we can shape the future. It doesn't really help you to hear that your sales are going to tank in two quarters. They'll just say, 'Wow, that sucks.' What they want to know is, 'Well, what do we need to do now to avoid that future?' That's causal mechanism.
而且这类数据也很难获得,因为世界是我们的真实基准,但它只发生一次。所以在一个非常受控的设置中,除了一个变量外其他一切都相同,这种数据集几乎很少出现。所以这就是为什么这类数据集既难以获得,又对建模人类行为非常重要。
And this is also very hard data to come by, because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of dataset almost rarely happens. So this is the reason why this dataset is both hard to come by but also quite important if you're trying to model human behavior.
所以我认为行为数据是最难获取的数据集。外面有什么?甚至可能有什么?因为你不会知道我生活的很多细节。我甚至没有关于自己的数据,比如我想分析自己的健康或习惯,但我并没有记录所有事情。那么你怎么能有那些数据呢?
So the behavior I think is the hardest dataset to acquire. What is out there? What is possible even? Because you're not going to know a lot of details about my life. I don't even have data for myself on, like, I want to analyze my own health or habits, and I just don't log everything. So how can you have that data?
所以我们实际上进行了很多随机对照试验。
So we actually run a lot of randomized control trials.
是啊。但你把人们放在实验室里,观察他们睡觉还是什么?
Yeah. But you put people in the lab, you watch them sleep or what?
我们确实非常重视同意流程。所以人们知道我们邀请他们成为这个社区的成员,既分享数据,也让他们以不同形式被代表。但我们带很多人到实验室或虚拟实验室,在那里我们设计实验,向他们提出真实的行为决策。在这些实验设置中,态度与行为之间的区别通常在于你的决策是否有真实的利害关系。这最终使它成为行为数据。
So we do actually care a lot about the consent process. So people know that we invite them to be a member of this community to both share data and also have themselves represented in different forms. But we bring a lot of people to the lab or virtual lab where we design experiments that pose real behavioral decisions to them. And often in these kinds of experimental setups, what makes the difference between attitudinal versus behavioral is if the stake in your decision is real. That's ultimately what makes it behavioral.
在这些设置中,我们受到社会科学、心理学等领域同事的启发。他们进行研究时使用的技术是,想象有一个在线商店,你邀请人们来,他们在实验中购买的任何东西,他们真的会收到该商品。这些就是让利害关系变得真实的事情。所以我们进行了很多这样的实验,我们也与公司合作。而且现在我们也有客户非常乐意让我们一窥他们用户表现出的行为类型,这样我们就能更深入地了解人们在这些不同平台上的行为。
So in these kinds of setups, we are inspired by our colleagues in social sciences, psychology, and so forth. When they run studies, the techniques they utilize are, imagine there's an online store that you invite people to come by, and whatever they purchase in this experiment, they actually get that item delivered. These are the kinds of things that make the stakes real. So we run a lot of these experiments, and we also partner with firms. And right now we also have customers who are quite excited to give us a glimpse of the kinds of behaviors their users exhibit, so that we can get a little bit deeper understanding of how people behave on these different platforms.
我认为在客户方面,他们有很多关于用户的数据——谁买了,他们有行为数据。你能给我们举个例子,说明有人来找你是为了什么?他们想解决什么问题?在这个过程中,你会为他们定制模型吗?你有现成的东西吗?那是什么样的?
I think on the customer side, they have a lot of data about their users—who has bought, they have the action data. Can you kind of walk us through an example of what does someone come to you for? What questions would they want solved? In the process, do you customize a model for them? Do you have something off the shelf? What does that look like?
如今,当人们使用我们的模型时,通常是为了更好地了解他们感兴趣的人群。所以通常在合作开始时,我们会聚在一起,了解他们希望我们建模的人群。例如,如果你是一家面向全美销售的消费品公司,那可能相当直接——你想建模美国总体人口。但同时,如果他们想进入某个垂直领域或市场,想象他们想更好地了解,比如说,住在加利福尼亚的 20 多岁和 30 多岁的人——那是一个更具体的人群。
Today, when people leverage our models, it's often to better understand the population of their interest. So usually at the start of the relationship, we basically come together and hear about what population they want us to model. So it might be that if you're a CPG company selling to all of the US, it might be fairly straightforward—you want to model the general population of the US. But at the same time, if there's a vertical or a market that they're trying to go into, imagine they want to better understand, let's say, people in their 20s and 30s living in California—that's a much more specific population.
所以我们了解这些人群,然后我们在征得同意并提供激励的情况下招募这些人。我们基本上收集他们的一些数据,并创建这些人的模型。然后我们的产品允许你基本上查询他们。它可以输入一个过滤器,即你想交谈的人群的描述,就像我刚才提到的那个,以及一个环境。环境可以就是调查问题、行为实验或 A/B 测试。通常核心用例首先是概念测试之类的事情。
So we hear about these populations, and we go recruit these people with consent and with incentives. We basically collect some of their data and create a model of these people. Then what our product allows you to do is basically query them. It can take as input a filter that is a description of the population you want to talk to, just like the one I just mentioned, and an environment. The environment can literally be survey questions, behavioral experiments, or A/B testing. Often times the core use cases are things like concept testing to start with.
但你也知道,人们有时会想做焦点小组,或者我们服务的另一个有趣用例,就是为上市公司模拟财报电话会议。这些是我们经常起步的用例。
But also, you know, people sometimes want to do focus groups, or one of the sort of fun use cases that we also serve is actually modeling things like the earnings call for public companies. So these are the use cases that we often start with.
概念测试。这是个既定术语吗?我从来没听说过概念测试。
Concept testing. Is that an established term? I've never heard of concept testing.
是的。它基本上涉及他们有不同的信息、不同的产品、不同的想法。
Yeah. So it basically has to do with they have, let's say, different messaging, different products, different ideas.
这就像一种营销活动。嗯,好的,明白了,明白了。
It's like a marketing exercise. Yeah. Okay. Got it. Got it.
政治。
Politics.
我们确实与盖洛普有战略合作伙伴关系,当然盖洛普深入政策领域等等。目前我们还没有深入涉足政治领域。
We do have a strategic partnership with Gallup, and of course Gallup is deep into policy space and so forth. Right now we have not worked deeply with politics like that area just yet.
不过,我很好奇是否有需求,或者他们是否真的有不同需求,这些需求在根本上与你们现有用户或人群不兼容。
However, I'm curious if there is demand or if they really would have different needs that somehow fundamentally don't mix with your existing users or people.
我认为肯定有需求。是的。
I think there's certainly demand. Yeah.
但我们非常关注这项技术如何被采用,以及我们最终会对社会产生的影响。我确实认为政治是一个公司必须特别谨慎对待其运作和影响方式的领域。所以这也是我们希望在服务政治等市场之前,确保我们形成足够的护栏和对如何利用这项技术的视角。
But we are very much mindful of how this technology gets adopted and the societal impact that we'll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guard rail and perspective on how to leverage this technology before we go on to serve markets like the politics.
我给大家举个例子。我最喜欢的剧之一是《白宫风云》。我不知道大家有没有看过。其中一个关键情节是总统患有多发性硬化症,但他们还没有——他们需要想办法如何公开。所以他们用一个假州长进行民意调查,让人们回应,然后他们试图根据民意调查的结果来决定如何应对,比如我们该怎么处理。
I'll give people an example. One of my favorite shows is the West Wing. I don't know if people have watched it. One of the key storylines is like the president has multiple sclerosis, but they haven't—they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond on the poll, and they try to make decisions based on the results of that poll on like how well they'll be received, like where, how should we play this.
我就想,嗯,你知道,我觉得这种反事实的情况,如果我能信任模拟的话,我肯定会用它来做这个。是的。
And I'm like, well, you know, I think those kind of counterfactual things I would actually use a simulation for this if I could trust it, for sure. Yeah.
在那部剧里,结果如何?
In that show, how did it go?
在那部剧里,基本上就是已成定局。他们说:“我们知道这很糟,只是不知道有多糟。”然后民意调查回来了,结果是“真的很糟”。然后他们还是照做了。
In that show, it basically was like a foregone conclusion. They were like, "We know it's bad. We just don't know how bad." And then the poll came back. It was like, "It's really bad." And then they just did it anyway.
部分原因是这是剧,对吧?所以你在最大化戏剧性。
Part of it is it's a show, right? So you're maximizing drama.
能有多糟?哦,太糟了。
How bad could it be? Oh, it's horrible.
是的。在某种程度上,我认为这是作为你们客户的一部分技巧或挑战,那就是如果我大致知道并能直觉到效果会怎样,我还需要你吗?我需要什么样的效果敏感度才能做决定,对吧?比如,如果我的支持率是 50%,
Yeah. And to some extent, I think that is part of the trick or the challenge with being a customer of yours, which is that if I roughly know and can intuit what the effect is going to be, do I need you? What sensitivity of effect do I need in order to make a decision, right? So for example, if my approval rating is 50%,
然后我有一条负面新闻出来,支持率降到 30%。
and I have this negative piece news item comes out and it drops to 30.
嗯。
Yeah.
如果降到 20%,如果降到 40%,我在乎吗?不。我知道它会降。这是负面的。那我什么时候在乎模拟呢?
If it drops to 20, if it drops to 40, do I care? No. I know it drops. It's negative. So when do I care about simulations?
你做了明显糟糕、不受欢迎的事,人们不喜欢你。是的。我的意思是,这是模拟。
You do something that's clearly bad, that's not popular, and people don't like you. Yeah. I mean, it's a simulation.
嗯,所以有几件事。首先,显然,有些用例是每天,比如开发者、设计师、政策制定者、营销人员——他们每天都会创造资产、创造新产品。结果发现,很多决定事后看来是显而易见的。是的,这当然很糟。但我们仍然进行这些研究,因为理解幅度和理解某事的严重程度实际上相当困难。即使我们觉得,当然,这说得通。我的意思是,这就是我们犯这么多错误的原因。就像每次有人上网说了什么引发巨大反弹的话,你看着那说,真是个白痴。然而,这很难。这是第一点。
Well, so there are a couple of things. One, obviously, is there are use cases where every day, for instance, developers, designers, policy makers, marketers—every single day they create assets, they create new products. And it turns out that many of the decisions in hindsight are sort of obvious. Yes, of course this is bad. But we still run those studies because understanding the magnitude and understanding how acute something is is actually quite difficult. Even if we feel like, of course, this makes sense. I mean, this is the reason why we make so many mistakes. Like every time somebody goes online and says something that has huge backlash, you look at that and like, what an idiot. However, it's tough. That's one.
这里还有另一个方面,这再次说明了为什么模拟实际上不同于预测。在模拟中,在理想情况下,模拟试图展示的是每一步,或者我们需要采取的每一步以达到某个结果。对吧?所以在最先进的模拟中,有时我们建议的下一步可能相当反直觉。我有时给出的类比,我把它放在更现实的例子里,但你知道,正如我提到的,我是科幻小说的超级粉丝,我不知道有多少观众读过《基地》系列之类的。
There's also another aspect here, which is again, this is the reason why simulation is actually different from prediction. In simulation, in the ideal case scenario, what simulation is trying to show is each step of the way, or each step that we need to take to get to a certain outcome. Right? So in the most advanced simulations, sometimes the next step that we're suggesting might actually be quite counterintuitive. The analogy that I sometimes give, and I ground it in a more realistic example, but you know, as I mentioned, I'm a huge fan of science fiction, and I don't know how many of the audience members have read things like the Foundation series.
我们已经多次提到心理史学。
We've mentioned psychohistory a number of times.
好的,太棒了。所以我可能真的在对合适的人说话。如果你读过《基地》系列,第一幕就是有一群科学家发现我们的银河帝国将要崩溃,我们将有三万年的动荡。他们基本上运行心理史学,这个模拟器试图教他们,好吧,我们如何把动荡缩短到一千年?他们计划好了,而计划的第一步是把那些说“好吧,这要来了”的科学家流放到银河系中某个随机的地方。
Okay, fantastic. So I might actually be talking to the right crew. If you read Foundation series, literally the first act is there's a group of scientists who have found out that our Galactic Empire is going to collapse and we're going to have 30,000 years of unrest. And they basically run psychohistory, the simulator that tries to teach them, okay, how can we keep this unrest to 1,000 years? And they plan this out, and the first step of that plan is to get the scientists who say, "Okay, this is coming," exiled into this random place in this, you know, galaxy.
端点星。
Terminus.
完全正确。这太反直觉了。就像,你把那群对银河帝国崩溃发出警告的科学家送到一个偏远地方,这是多么奇怪的举动。这怎么可能是正确的第一步?嗯,结果在这个特定的模拟中,那确实是正确的举动。就是这类事情,对吧?而这类推理之所以可能,是因为你展示了阶跃函数或导致特定结果的每一步。所以模拟在最高形式上真正让你做的是,你给它的不是一个问题或疑问,比如人们对调查会怎么回答。那不是我们做的。我们告诉它的是,在《基地》的背景下,我们有一个目标。我们想把动荡控制在一千年。我们现在需要走什么路径才能达到那个特定的未来?这就是模拟让你做到的。
Exactly. And that's so counterintuitive. Like, what a strange move that you literally sent the group of scientists who was raising voice around the potential collapse of Galactic Empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation that actually was the move. It's these kind of things, right? And the reason why these kind of reasoning is possible is because you're showing the step function or each step that results in a particular outcome. So really what simulation allows you to do in its highest form is you give it not a problem or question like what would people answer to the survey. That's not what we do. What we tell it is here is a goal that we have in the context of Foundation. We want to keep the unrest to a thousand years. What is the path that we need to take now to get to that particular future? And that's what simulation allows you to do.
现在把它转化到真实市场。想象你是一家汽车公司,你即将发布一款电动车,你试图理解,嗯,我们如何营销电动车以确保股价上涨?但如果答案是,你可以用 XYZ 方式营销你的电动车,但那可能会改变人们对非电动车的看法,实际上让你的整体销量下降。这不太直观。尤其是如果你试图优化的只是电动车销量,而你只追踪这一点,那么那可能实际上导致一个完全错误的解决方案,或者至少是一个与你预期不同的解决方案,无论对错。
Now translating that into real market. Imagine you're an automobile company and you're about to release an EV, and you're trying to understand, well, how do we market EV to make sure that our stock price goes up? But what if the answer comes down that well, you can market your EV in XYZ way, but that might change people's perception around the cars that are not EV and actually make your overall sales go down. Not very intuitive. Especially if all you're trying to optimize is EV sale, and that's the only thing you're tracking, then that might actually result in a completely wrong solution, or at least a different solution than what you would have expected, whether it's right or wrong.
是的,这就是模拟的力量。
Yeah, that's the power of simulation.
给听众说明一下,我们之前和 Shopify 的 ML Parin 聊过类似的话题,他们在做 Sim Jim。我不知道他有没有跟你提过。非常相似。
For listeners, we covered a similar topic with ML Parin from Shopify where they are working on Sim Jim. I don't know if he ever talked to you about it. It's very similar.
目标是提高转化率,但整个路径非常不寻常。
The goal is to increase conversion, but the journey is very unusual.
路径不寻常。
Journey is unusual.
对。他实际上是在购物轨迹上寻找干预点,这跟你说的类似。不是关于态度层面的——这是你的说法。
Yeah. He's actually trying to look for interventions on a shopping trajectory, which is similar to what you're saying. It's not about the attitudinal—that's your word for it.
是关于行为。
It's about behavior.
是关于行为,这正是区别所在,对吧?不是关于短期方向,而是如何影响多轮交互。
It's about behavior, and that's exactly the difference, right? It's not about the near-term direction, but more about how do you affect multiple turns of interactions.
对。你开头也说过一句很好的话。不是关于人们想知道结果,而是关于他们如何改变它。改变到达那里的方式。但我想回到我们怎么知道这是有根据的?比如你怎么运行评估?你怎么测试模拟能实现?基本上,如果我用你最喜欢的 LLM,比如 Opus、GPT-5,让某个智能体来规划这些事情。
Right. You had a good quote at the start about this as well. It's not about people wanting to know the outcome. It's about how they can change it. Change the way to get there. But I want to take it back to how do we know this is grounded? Like how do you run evals? How do you test that simulations come through? Basically, if I was to do the same thing that you described with say your favorite LLM, Opus, GPT-5, have some agent to map out these things.
如果我给它同样的目标、同样的目的,做一个像样的系统,我们得到的答案会有多大不同?你说你需要改变模型权重。你有自己的解决方案,但我们离这有多远,你怎么检查它是否有根据?
How different are the answers we would get if I give it the same goal, the same objective, make a decent system. You're saying that you need to change the model weights. You have your own solution to this, but how far off are we and how do you check if it's grounded?
你网站上有一些有趣的内容,实际上指出了你如何运行真正的评估,但如果你能带我们了解那方面,你知道,我认为这是人们的一大担忧。他们会说 LLM 会幻觉,你只是在层层幻觉,对吧?
You have some interesting stuff on your site that actually points to how you run real evals, but if you could take us through that side, you know, I think that's one of the big concerns that people have. They're like LLM hallucinate, you're just hallucinating layer after layer, right?
我们做这件事的方式,实际上是我们继生成式智能体论文之后写的论文,它真正成为了基础,至少对 simile 以及模拟和合成面板领域来说。是的,这篇论文叫《生成千人的模拟》。我们在这篇论文中做的是,我们实际上把一千名从美国代表性抽样的人带到了虚拟实验室,我们基本上花了两个小时收集相当广泛的数据。在这项特定研究中,我们非常关注访谈数据,其脚本来自一个叫“美国之声项目”的项目,然后我们还会搭配大量行为数据等等,凡是我们能在两小时内收集到的。然后我们实际上会让这些人离开几周,在那段时间里,我会用这些数据创建他们的数字孪生,两周后我会让人类参与者回来,完成一系列调查、实验、行为研究。所以我们实际上有这份清单,基本上包括行为经济游戏之类的东西。我们会实际运行大五人格测试、综合社会调查。我们还会继续运行发表在 PNAS 上的随机对照试验,然后让他们的数字孪生预测源个体在这些研究和调查中会如何表现。正是在这里,我们基本上能以人们复制自己的准确度的 85% 来复制人们的行为和态度。所以这实际上是第一篇真正给出验证结果的论文,证明我们能够准确地建模个体。
The way we do this, and this is actually the paper that we worked on after the generative agents paper, that really became the foundation, at least for simile and also the field of simulation and synthetic panels. Yeah, this is the paper called 'Generating Simulations of a Thousand People.' Here's what we've done for this paper. We actually brought a thousand people that are representatively sampled from the US to a virtual lab, and what we basically did was we spent two hours collecting fairly wide-ranging data. In this particular study, we focus a lot on this interview data, whose script was taken from this project called American Voices Project, and then we would also pair it with a lot of behavioral data and so forth, whatever we can collect within two hours. And then we would actually send these people away for a couple of weeks, and during that time I would use this data to create their digital twins, and I would bring the human participants back after two weeks and have them complete a battery of surveys, experiments, behavior studies. So we actually have the list here, which basically included things like the behavioral economic games. We would run literally like big five personality test, general social survey. We would also go ahead and run the randomized control trials that were published on PNAS, and we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we basically could replicate people's behaviors and attitudes 85% as accurately as people would replicate their own. So that actually was the first really paper that gave this validated results that we can actually model individuals in an accurate way.
我们现在最终发现的,当然,在 AI 领域——所以这篇论文在 2024 年底发表——在 AI 领域,一年半、两年,那就是一辈子。
And what we ended up finding now, of course, in AI space—so this paper came out at the end of 2024—in AI space, a year and a half, two years, that's a lifetime.
对。我只是,对于没看 YouTube 的听众,我想说标题数字是 85% 的准确率,这比你展示的所有其他方法都有很大改进。
Yeah. I just, for listeners who are not seeing the YouTube, I just want to say like the headline figure is 85% accuracy, which is a big improvement over all the other methods that you showed.
但对我们来说特别引人注目的部分,尤其是当我们进一步改进这项技术时,是生成式智能体模型。像 ChatGPT、即将推出的 Claude 这样的生成式 AI 模型,确实给了你正确的基础。然而,它们没有考虑的是人们真实的态度和行为方面,尤其是在你关心的群体中。所以这些模型今天真正擅长的是,它们基本上试图成为超级理性的客观机器,对吧?所以你去 Major Scale 这样的地方获取数据,你与专业程序员、科学家交谈,创建一个在推理方面惊人的模型。这就是它们所做的。Similar 实际上并不关心这些。我们在这里讨论的模型,我们试图创建的是和我一样笨的模型,对吧?所以如果我犯那些错误,模型必须犯同样的错误。
But the part that was actually particularly striking to us, especially as we improved this technology even further, was the generative agents model. Generative AI models like ChatGPT, Claude that's coming out, it does give you the right foundation. However, what they do not consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about. So what these models are really, really good at today is they're trying to basically become the super rational objective machines, right? So you go get their data from places like Major Scale, you talk to professional programmers, scientists to create a model that's amazing at reasoning. That's what they do. Similar actually doesn't care about any of this. The models that we're talking about here, what we're trying to create are models that are as dumb as I am, right? So if I make those mistakes, the model has to make the same kind of mistake.
哦,那很难。
Oh, that's very hard.
那很难。
That's very hard.
你在解决更多的 VX 悖论。
You're solving more of a VX paradox.
正是如此。这实际上是完全不同的数据和训练目标。这也是我们实际上看到前沿模型与该领域创建的模拟在人类行为预测性能上存在相当大差异的地方,在某些情况下,前沿模型的性能会一路下降到 20%、30%。特别是如果你进入更小众的群体,涉及我们客户真正关心的主题。在更普通的群体上,可能在 50% 到 60% 左右。所以它不是很稳健。你不会想根据这类发现来做决定。如果你能把它提高到 85%,那最终就是人们非常兴奋的地方。
That's exactly. And this is actually completely different kind of data and training objective. This is also where we actually see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models and simulations being created in the space, where in some cases the model performance of frontier models go all the way down to 20, 30%. Especially if you go into that more niche population on topics that our customers would actually care about. On more gem pop, it might be around 50 to 60%. So it's not very robust. Like you wouldn't want to make your decision off of these kinds of findings. If you can bring that up to 85%, that is ultimately what people end up getting very excited about.
对。我们想继续沿着论文路线走吗?
Yeah. Do we want to keep going on the paper routes?
对,当然。所以最后一篇是有点意思的。这篇论文是我们千智能体论文的后续论文,基本上想法是现在我们能否进一步增强模型,并实际上基于大量随机对照试验进行后训练。所以这是一篇有趣的论文。在很多方面,数据总是建模中最有趣的部分。我们在这里得到的数据是,有一个叫开放科学基金会的平台。所以一些听众可能对此熟悉。特别是在过去 5 年左右的社会科学中,一直存在对研究可复制性的担忧,对吧?所以科学家们承认这是一个危机,我们重新运行研究,但并没有看到相同的发现。这很棘手。而这种情况经常发生的原因基本上是生存偏差,发表的论文通常需要在我们的实验中保持所谓的 p 值小于 0.05。这基本上表明我们看到的結果是假阳性的概率只有 5%。
Yeah, for sure. So the last one was sort of an interesting one. So this paper was the follow-up paper that we had to the thousand agents paper, where basically the idea was now can we augment the models even further and actually post-train a model based on a lot of randomized control trials. So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this platform called Open Science Foundation. So some of the audience might be familiar with this. There has been, especially in the social sciences over the past 5 years or so, this concern around replicability of studies, right? So it's a bit of a crisis the scientists acknowledged, where we rerun the study and we don't actually see the same finding. It's tough. And the reason why that was often the case was there's basically the survival bias, where the papers that get published often need to maintain what we call the p-value of less than 0.05 in the experiments that we ran. That basically suggests that there's only a 5% chance that the results that we saw is a false positive.
但棘手的地方在于那些没有发表的论文,而且我们发表的任何东西仍有 5% 的概率完全是随机生成的。比如有 5% 的概率这个效应不是真的,只是由于抽样偏差碰巧显得是真的。正因为如此,科学家们开始做的是预先注册他们的研究。在运行实验之前,他们会去这个平台说:这是数据,这是我们要收集的人群,这是假设。他们只会说:这是我们的假设,这是我们相信的,而且你不能事后修改这些假设。这实际上给了我们更多的科学统计信心,让你最终看到的任何效应都是真实的。所以这最终创造了一个非常有趣的平台,这个平台上现在有成千上万个真实世界的实验和假设,而且很多质量确实很高,比如专业设计的行为研究和随机对照试验。所以我们从这个平台获取了数据和研究,基本上用它们来证明一个观点。显然这个特定的模型不是我们商业化的东西,因为这显然是开放科学的一部分,但这个特定的数据集帮助我们证明了一个观点:通过收集大量设计良好的随机对照试验,我们可以显著提高模型预测人类行为的能力。这就是这篇论文的内容。
But the tricky part was all the papers that were not published, and there's still a 5% chance that whatever we publish is actually totally randomly generated. Like there's a 5% chance that this effect is not real but it just happened to be real because of the sampling bias. So because of that, what scientists started to do was they started to pre-register their studies. So before running an experiment, they would go to this platform and say, here is the data, here is the population that we're collecting, and here is the hypothesis. And they would just say, here is our hypothesis, this is what we believe, and you cannot retroactively change those hypotheses. This is what actually gives us more scientific statistical confidence that whatever effect you ended up seeing is actually true. So that ended up creating this really interesting platform where there's one platform that now contains tens of thousands of real-world experiments and hypotheses, and a lot of these are actually really high quality, like professionally designed behavior studies and randomized control trials. So we actually got the data and the studies from this platform and basically used that to make a point. And obviously this particular model is not something that we're serving commercially, because this obviously was a part of the open science, but this particular dataset helped us make a point that by collecting a lot of these randomized control trials that are really well designed, we can make significant improvements in models' capability to predict human behaviors. So that's what this paper was about.
这些是在个体层面做的吗?比如我需要为每个个体、每家公司调整模型吗?是基础模型有变化,然后做一些轻微的后训练吗?你能分享些什么吗?
Is this stuff done on an individual level? Like do I need to tune the model per individual per company? Is there foundation model changes and then some slight post-training? Anything you can share there?
所以这个特定的模型实际上是在我们拥有的个体层面的数据上训练的,但我们对两种方式都做了实验,这实际上也是我们在 Sim 2 最终做的。我们总是训练两个不同的模型。一个是我们所谓的人群层面模型,另一个是我们所谓的个体层面模型。两者实际上都接受非常相似的输入,即子人群或个体的描述以及一个刺激。在这项特定的工作中,我们也做了同样的事情。我们报告的结果更偏向于个体,因为我们确实认为这在很多方面是更困难的任务,但这就是我们所做的。
So this particular model actually was trained on data we actually had at the level of individuals, but we experimented with both, and this is actually what we end up doing at Sim 2. We always train two distinct models. One is what we call the population-level model. The other is what we call the individual-level model. And both actually take very similar input, which is the description of a subpopulation or individual and a stimulus. In this particular work, we've done the same here. The results that we are reporting are much more geared towards individuals because we do actually think that is a harder task in many ways, but that's what we have done.
你有没有见过关于人类能解决而模型不能解决的问题?比如现在,你知道,我住的地方离洗车店步行 5 分钟,开车 10 分钟。我应该走路还是开车?
Have you seen anything on the questions that humans can solve that models can't solve? So like currently, you know, I live 5 minutes walk away from a car wash. It's a 10-minute drive. Should I walk or drive?
模型会说,哦,走到洗车店,但你知道你没有车。
The model will say, oh, walk to the car wash, and you know you don't have your car.
在模拟中,像这样的事情是问题吗?你会认为这对人类来说很简单,但如果模型说你应该走到洗车店,这里有什么问题吗?
Is anything like this a problem in simulation? You would assume it's very simple for a human to think about, but if the model is saying you should walk to the car wash, anything here?
这与其说是我们能解决什么,不如说是人们会犯哪些模型会忽略的偏见或错误。比如,想象一下,当我在斯坦福而不是在斯坦福时,我住在帕洛阿尔托,所以从校园步行大约 40 分钟。你问模型:‘好吧,我们回家。我能做什么?’它很可能会叫一辆 Uber,或者给我公交时间。但在很长一段时间里,我实际上非常喜欢走回去。我这样做的原因不是为了效率。它确实帮助我思考。我喜欢每天走半小时、40 分钟左右,在那里我可以思考想法、研究,只是沉浸在自己的思绪中。这是一种非常人类的活动。除非模型见过这种情况并真正理解这种活动的重要性,否则它确实会错过这些特征。所以我认为这基本上是我们试图建模的。从根本上说,人类的东西可能不是最有效的事情,可能不是正确的事情,但却是让我们成为自己的东西。
It's less about what can we solve, but I think it's more about what biases or mistakes do people make that models miss. Like for instance, imagine that you are, you know, when I was instead of Stanford, I lived in Palo Alto, so it's about, I would say, 40-minute walk from the campus. You ask the model, 'Okay, let's go home. What can I do?' It would likely call an Uber or, you know, give me the bus time. But for the longest time, I actually really liked walking back. And the reason why I wanted to do that was not for efficiency. It actually really helped me think. And I like to walk for, you know, half an hour, 40 minutes or so a day, where I just get to think about ideas, research, just get lost in my thoughts. That's a very human activity. Unless the model has seen that and actually understands the importance of that activity, it would actually miss these kind of features. So that actually I think is fundamentally what we're trying to model. Like what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.
我很好奇是否有你真正想要的数据集,能实质性地帮助你。其中一个版本可能很有趣。作为数据集,哪个对你来说更有价值?所有的 LinkedIn、所有的 Twitter、所有的 Facebook。
I'm curious if there are some datasets that you really want that would materially help you. One version of this may be interesting. Which is more valuable to you to acquire as a dataset? All of LinkedIn, all of Twitter, all of Facebook.
呃,你知道,说实话,有点难排名。部分原因是,你知道,有这样一个产品说法:没有反馈是错误的,因为它教会你一些关于用户的东西。不管是什么样的反馈。
Uh, you know, to be honest, it's a little bit hard to rank. In part because, you know, there's this product saying where no feedback is wrong because it teaches you something about your users. Doesn't matter what kind of feedback.
我觉得有点像那样。
I think it's a little bit like that.
所以就是越大越好。
So just whatever is bigger.
那不同的领域呢?比如所有的亚马逊数据怎么样?
What about a different domain? Say it was what about all of Amazon data?
购物数据,对吧?
Shopping data, right?
购物数据。
Shopping data.
所以亚马逊数据很有趣,因为它非常行为化。虽然人们在社交媒体上的行为,你可以眯着眼说那也是行为,但交易数据总是有趣的。它也是最常见的。然而,如果我们只看纯粹的社交媒体,如果我必须真正选择,Facebook 可能很有趣,因为我确实认为它是最接近人们默认版本的东西。因为你去 LinkedIn,那是一个非常专业的环境,所以人们会有所防备,对吧?那仍然有趣,因为那是真实的人类态度和行为,但这不是你的基础状态。你去 Twitter,Twitter 上的人有他们自己疯狂的人设。或者取决于你是谁,比如我的 Twitter 个人资料和人设最初非常学术。嘿,我在这里分享我的研究。现在我分享与 Simile 相关的东西。但 Facebook 是那种更私密的空间,人们只是和朋友联系。在那种方式下,我确实认为它更能展示一个人是谁。所以如果我必须选择,我可能会选 Facebook。
So Amazon data is interesting in that it's very much behavioral. Although like what people do on social media, you could sort of squint and say that is also behavioral, but the transaction data is always interesting. It is also most commonly available. However, if we were to look at purely social media, like if I had to really pick, Facebook likely is interesting because I actually do think it is most sort of a default version of people. Because you go to LinkedIn, it's very much a professional environment, so people put up their guards up, right? And that still is interesting because that is true human attitude and behavior, but it is not your base state. You go to Twitter, and Twitter people have their own crazy personas. Or depending on who you are, like my Twitter profile and persona is very much initially, I was very much an academic. Hey, I'm here to share my studies. Now I share things that's related to Simile. But Facebook is one of those more private spaces where people just connect with their friends. In that way, I actually do think it shows you a little bit more about who that person is. So if I had to pick, I likely pick Facebook.
是的。而且你感兴趣的是整个人、他们的背景和哲学。我想说这只是注入变体和偏见的方式,是不是太临床或太机器学习导向了?我猜广泛的问题是,这是否比随机的组合爆炸版本更好?所以我们有一个链接到 10 亿人设的论文,他们基本上没有做你们做的任何基础工作。是的,
Yeah. And you're interested in like the whole person and their background and philosophy. I guess is it too clinical or too machine learning oriented to just say this is just ways to inject variants and biases? The broad question I guess is like is this any better than a randomized like combinatorial explosion version? So we have a link to the 10-centent uh billion persona paper where they basically did not do any of the groundwork that you are doing. Yeah,
是的,
Yeah,
他们只是做了一个交叉矩阵,这里是世界上所有的职业,这里是世界上所有可能的背景。对它们做点积,就这样。这就是你为 10 亿人准备的提示。
they just sort of did like a cross matrix of here's all the professions in the world. Here's all the people possible backgrounds in the world. Do a dot product across all of them and that's it. That's your prompt for a billion people.
是的,这会有点用。
Yeah, this will do something.
我不知道它能否做到你所做的,但它能让你走一段路,走一部分路。
I don't know if it'll do what you do, but it gets you some way, some percent of the way there.
所以,这其实是一篇有趣的论文。这篇论文出来时,我欣赏的是它的规模,显然你确实逐渐想要能够模拟真正庞大的社会和互动。所以规模绝对令人钦佩。它严重依赖于训练模型时所用的已知统计数据。所以,如果你相信这些统计是正确的,这其实不失为一个好方法。但这里的论点,也是我们在市场上看到的,就是如果这行得通,那我们实际上已经解决了模拟问题。
So, this actually was an interesting paper. What I admired about this paper when it came out was the scale, and obviously you do gradually want to be able to simulate really large societies and interactions. So the scale is definitely admirable. It is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is actually not a bad way to go about this. But the thesis here, and this is something that we also have seen in the market, is that if this works, then we actually have solved simulation.
对。
Right.
因为我调查,比如,好吧,美国人口中有 5% 从事建筑,另外 5% 从事医疗,等等,对吧,然后你继续往下数,然后你再看另一边,5% 的人具有大五人格中的神经质之类的,就这样。
Because I survey, like, okay, 5% of the US population is in construction, the other 5% is in medicine, whatever, right, and then you just keep going down the list, and then you do the other side, 5% has, like, you know, the Big Five personality of like neurotic or whatever, that's it.
就这样。所以如果你相信我们利用的底层数据集和平台拥有所有正确的统计信息,那么这实际上就解决了问题。那时你只是在检索已经嵌入模型参数中的知识。但不幸的是,这不是我们所看到的。关于人们的知识非常详细且小众,如果你只看一个例子,可能觉得很平凡,但当你把它们放在一起时,实际上相当丰富。你确实需要做大量定制数据收集来更好地理解人们。而且,你知道,我认为这也是这份工作有趣的地方,你想深入了解人们,而深入了解他们的过程实际上需要对细节非常关注,你需要关注并尊重人们所过的日常生活。
That's it. So if you believe that the underlying dataset and the platform that we're leveraging has all the right statistics, then this actually will have solved it. You're at that point merely retrieving the knowledge that is already embedded in the model parameters. That's not unfortunately what we see. There's such detailed and also niche knowledge about people that, if you just take one example, it might feel very mundane, but it's actually quite rich when you put together. You actually do need to do a lot of bespoke data collection to better understand people. And this is also, you know, I think what makes this particular job fun, which is you want to deeply understand people, and the process of deeply understanding them actually requires a lot of attention to the details, and you do need to pay attention to and pay respect to the daily lives that people lead.
我想谈谈 Scaling(规模扩张)模拟。那么,我们不能模拟什么?我们能模拟什么?Scaling 如何影响这一点?模型有多大?如果我们从 80 亿参数,比如几亿,到 1000 亿参数,万亿呢?在某个规模、某个训练量下,我们是否会得到有趣的涌现?你发现了什么不寻常的东西,以及从中有什么收获?
I want to talk about scaling simulations. So what can't we simulate? What can we simulate? And how does scaling affect this? So how big are the models? What if we go from, you know, 8B, like a couple hundred million, like 100 billion parameters, trillion? Do we get any interesting emergence at a certain scale, at a certain amount of training? Do you uncover anything unusual, and any learnings from that?
我们在 Simile 看到的是,我们确实后训练我们自己的模型。我们实际看到的是模拟中缩放定律的早期迹象。你摄入的关于人类的数据和算力越多,你实际上开始在模拟和预测人们方面获得模型性能的可预测和可预测的收益。
What we are seeing is, at Simile, we do post-train our own model. The thing that we're actually seeing is the early glimpse of scaling laws in simulations. The more data about humans and more compute you ingest, you actually start to get predictive and predictable gains of the model performance in simulating and predicting people.
我们需要一个缩放定律。
We need a scaling law.
你知道,它 Scaling 得很好。每当你发现它,那是一件美好的事。我们开始看到它的迹象,这相当令人兴奋。
You know, it's scaling well. Whenever you find it, it's a beautiful thing. And we're starting to see the glimpse of it, which is quite exciting.
但如果你谈论模拟作为一个整体的雄心,那不仅仅是构建一个模型。它是关于构建一个模型,然后创建智能体,这些智能体成为更大生态系统中的个体。所以你基本上是在创建这种多智能体模拟。将来,你希望这些多智能体模拟也生活在一个非常丰富的环境中,对吧?我们真正想要达到的是,嘿,我们能否真正创造——让我们再做一次时间机器游戏——5 年、10 年后,我们能否创造一个地球上 80 亿人的模拟?我认为那相当有趣,那确实是愿景。一旦你达到那种状态,你能帮助社会回答的问题类型也开始改变。从我的角度来看,答案从根本上讲是关于涌现,即社会和大型群体的涌现行为。所以,例如,让我兴奋的问题——也许这有点,你知道,我有我学术的一面——对我来说,问题是像,我们能帮助解决气候变化吗?如果你把气候变化看作一个问题空间,这就是我们,像社会科学家常说的“棘手问题”,即有许多具有竞争激励的参与者试图做出非常复杂的决策,而协调这种协调决策在现实生活中很难真正解决,这也是我们无法解决它的原因。模拟能帮助我们解决这个问题吗?另一个是,我们能否真正理解民主崩溃的信号?或者我们能否理解,或揭示货币体系的起源故事?这些是我们从未真正有好的回答方式的社会问题。如果我们能创造我们社会的模拟,你必须相信这些是我们能解决的问题。所以这确实是这个领域的雄心。而且,你知道,我也认为,是的,我的意思是,我认为那里有一个诺贝尔奖可以赢得,这并不令人惊讶。而且我认为我们可以产生一些惊人的社会影响,帮助人们做出更好的决策。
But if you talk about the ambition of simulation as a whole, it's not merely about building a model. It's about building a model, then creating the agents that become the individuals in a much larger ecosystem. So you're basically creating this multi-agent simulation. Down the line, you want these multi-agent simulations to also live in a very rich environment, right? What we are really trying to get to at that point is, hey, can we actually create—let's do a time machine game again—5 years, 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that's quite interesting, and that really is the vision. And once you get to that kind of state, the kind of questions that you can help answer for the society also start to change. From my perspective, the answers are fundamentally about emergence, the emerging behavior of society and large groups of people. So, for instance, the kind of questions that I get excited by—and maybe this is a bit, you know, I have my academic side of me—for me, it's questions like, can we help solve climate change? If you look at climate change as a problem space, this is what we, like social scientists would often call the wicked problems problem, where you have many actors with competing incentives who are trying to make a very complex decision, and coordinating that coordination decision is very difficult to really solve in real life, which is also the reason why we couldn't solve it. Can simulation help us solve that? Another one is, can we actually understand the signals for collapsing democracy? Or can we understand, or can we uncover, the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the kind of problems that we can solve. So that's really the ambition of this field. And you know, I also think, yes, I mean, I think there's a Nobel Prize to be won there, which wouldn't be surprising. And I think there's some amazing societal impact that we can have to help people make better decisions.
诺贝尔经济学奖。
Nobel Prize in economics.
经济学。
In economics.
我明白了。我明白了。我们期待你有一天写出那篇论文。
I see. I see. We're rooting for you to write that paper one of these days.
但是,你知道,当我进入模拟领域时,深深启发我的学者之一实际上是托马斯·谢林。
But, you know, one of the scholars that I was deeply inspired by when I was coming into the space of simulation actually is the scholar named Thomas Schelling.
谢林,是的。
Schelling, yeah.
他工作的典型例子是,他是基于智能体建模的创造者之一。那是在 1970 年代和 80 年代。那是非常早的时期,但这确实是模拟的最早范例之一。而那个时代的典型模型——当然许多这些模拟都试图解决他们时代最相关的社会问题——它被称为隔离模型。所以种族隔离是我们关心的一个大话题。他们所做的是,他们实际上创建了一个网格世界,其中有红点和蓝点,这些点在当年就像是智能体,它们有一个简单的规则来支配它们的行为:如果你的邻居中有一定百分比是不同颜色的,并且超过某个阈值,那么你就随机移动到一个新位置。这篇论文或这个基于智能体的模型的一个惊人发现是,在很长一段时间里,人们认为社会中的隔离是由明确和公开的种族主义引起的。但如果你看这个模型,人们倾向于与同色人种居住——这种偏好可能非常微小,但非常小的差异实际上导致社会随着时间的推移完全隔离。这对很多人来说非常反直觉。而实际上,这项特定的工作最终影响了住房政策。例如,混合收入住房就受到了这类工作的启发。而托马斯·谢林最终因奠定模拟的早期基础而获得诺贝尔奖。
So the canonical example of the work that he's done was he was one of the creators of agent-based modeling. So this was like in the 1970s and '80s. It's very early days, but this was truly one of the first exemplars of simulations. And one of the canonical models from that time—and of course many of these simulations are trying to tackle the societal problems that are most relevant for their era—it was called the model of segregation. So racial segregation was a big topic that we cared about. And what they've done was they actually created this grid world where they had red dots and blue dots, and these dots were, back in the day, like they were the agents, and they had a simple rule that governed their behavior: if a certain percentage of your neighbors are of a different color, and if that goes above a certain threshold, then you move to a new location at random. One of the striking findings of this paper or this agent-based model was, for the longest time, people thought the segregation within society was caused by explicit and overt racism. But if you look at this model, people's preference towards living with people of the same color—that preference can be very minute, but the very small difference actually causes the society to segregate completely over time. This was very counterintuitive for a lot of people. And this actually, this particular work ended up informing housing policies. Mixed-income housing, for instance, got really inspired by this kind of work. And Thomas Schelling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations.
从更科学的角度看,我确实看到的机会是智能体模型。在 1980 年代、90 年代以及某种程度上 2000 年代初,它们曾产生过影响,但现在有点被社区遗忘了,因为你可以想象,红点和蓝点并不是对人类的丰富描述。但随着生成式 AI 的出现,尤其是生成式智能体,我们有机会创建足够高保真的智能体模型,帮助我们做出真正复杂的决策。这就是我看到的机遇。
The opportunity that I do see here in the more scientific terms is agent-based models. For the longest time, they had impact in the 1980s, '90s, and to some extent early 2000s, but they have now sort of gotten forgotten by the community a little bit because, as you can imagine, red dots and blue dots is not really a rich description of people. But with the emergence of things like generative AI, and in particular generative agents, we do have an opportunity to create these kinds of agent-based models that are high fidelity enough to help us make really complex decisions. And that's the opportunity that I see.
如果自然可行,那么是的,这类工作会带来更低的价格。
If naturally works, then yes, and that is the kind of work that will result in a lower price.
是的。
Yeah.
顺便说一句,我在新加坡长大。新加坡 80% 的人住在公共住房里,而公共住房正是出于这个原因实行种族配额,这真的很有意思。
For what it's worth, you know, I grew up in Singapore. 80% of Singapore is in public housing, and public housing has enforced racial quotas for exactly that reason, which is really interesting.
好的。那么,我们谈 Scaling,我们谈所有这些可能的智能体应用。
Okay. So, we talk about scaling, we talk about all these sort of agent possible applications.
我担心成本。
I'm scared about the cost.
即使我们只限于美国,而不是 80 亿人。是的。但建模这么多数以亿计的人要花多少钱?
If you even let's just keep it to the US, not 8 billion people. Yeah. But how much does it cost to model so many hundreds of millions of people?
如今很多时候,显然我们不会在行业和模拟技术的这个阶段从那种规模开始,但即使建模数千或数万人,我们也能为用户提供极其丰富和有意义的洞察。今天我们每周都在收集数万人的数据,而且我们实际上有样本库合作伙伴,能让我们触达全球数千万人。这就是我们今天所做的。
Often times today, obviously we don't start at that scale at this stage of the industry and simulation technology, but we can actually get to our users extremely rich and meaningful insights even by modeling thousands or tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands of people, and we actually have panel partnerships that get us to tens of millions of people globally. So that's what we do today.
顺便提一句,一旦你为一项研究收集了一个人,你能在后续所有研究中重复使用同一个人吗?
And just as a side note, once you've collected one person for one study, can you reuse that same person for all the subsequent studies?
完全正确。好的。这个模型和这些智能体的美妙之处在于它们是领域无关的。你真正想理解的是这些人的根本本质,他们的社会物理。
That's exactly right. Okay. The beauty of this model and these agents is the fact that they are domain agnostic. What you're really trying to understand is what is the fundamental nature of these people, what's their social physics.
显然,人的很多方面会随时间变化,比如过去一周你去过几次 CVS——显然那会变——但也有很多特质是已知永远不会变的,比如你的风险承受能力不会随时间真正改变,它非常稳定。所以我们试图捕捉的就是这类东西。但我们目前运营的规模是数百或数万到数十万,在我们部署的许多核心用例中,这个人口规模绰绰有余。真正在这一点上,你关心的不是人数,而是你是否覆盖了感兴趣的正确子群体。这也是人们想要更大样本的原因。不是因为他们真的想要更强的统计保证,而是他们能否真正筛选出任何他们感兴趣的人群。然而,你也可以想象,10 年后,如果我们真的相信算力会扩展,我们将有更多的算力可用,我们对模拟的雄心也会相应扩展。我的意思是,我们绝对有理由创建整个数据中心规模的模拟,或者我的直觉是,我认为在未来几年内,我们将开始创建实际上成本与训练一个基础模型相当的模拟。但也许它对社会如此有价值,以至于这是不言而喻的。我的意思是,即使今天,我们也在训练一堆新的基础模型,只是为了说我们训练了它们,我们花费了数千万。但如果我们能创建一个社会层面的模拟,真正解决气候变化,我今天就会运行它。我现在就会筹钱去运行它。
And obviously there are a lot of things about people that do change over time, like even things like how many times have you been to CVS the past week—obviously that will change—but there are so many traits about people that are also known to never change, like your risk tolerance doesn't really change over time; it's very consistent. So it's these kinds of things that we're trying to capture. But the scale we are operating at right now is hundreds or tens of thousands to hundreds of thousands, and in many of the core use cases that we are deployed in, this is more than enough population to cover those. Really at that point, what you care about is less the number of people but more do you have the right subpopulation of interest covered. And this is also the reason why people want a larger sample. It's not because they actually want stronger statistical guarantees; it's more that can they actually filter down to any population of their interest. However, you can also imagine in 10 years, if we truly believe that compute is going to scale, that we'll have much more availability for compute, and our ambition for simulation is also going to scale accordingly. I mean, there's definitely a reason for us to create an entire data center worth of simulations, or in my hunch here, I do think in the next some number of years we will start creating simulations that will actually cost as much as training a foundation model. But perhaps it's going to be so valuable to society that it would be a no-brainer. I mean, right now even today, we are training a bunch of new foundation models just so we can say we trained on them, and we spend tens of millions. But if we can create a simulation at the level of society that would actually solve climate change, I would run that today. I would raise the money right now just to run that.
太棒了。我想后续问题是,如果你让模拟彼此对话,它是否会复合增长?
Amazing. I guess the follow-up question is, does it also compound if you let the simulations talk to each other?
还是说他们今天已经这样做了?他们没有,对吧?据我所知,这取决于你试图运行什么样的模拟。在多智能体模拟设置中,智能体确实会彼此对话,对吧?
Or do they already do that today? They don't, right? As far as I understand, it depends on what kind of simulation you're trying to run. In the multi-agent simulation setup, the agents do talk to each other, right?
这正是 Smallville,对吧?
Which is exactly Smallville, right?
没错。
That's right.
但很多时候,比如在电子商务中,你只是一个人,所以没有对话的必要。但总有层次,对吧?比如你根据周围人购买和谈论的东西来决定买什么,对吧?
But a lot of times, for example, in e-commerce, you're just by yourself, so there's no point talking. But there's always levels, right? Like you decide what you will buy based on what other people around you buy and talk about, right?
这取决于情况。再说一次,我是从成本角度出发的。我会想,天哪。如果存在某种组合效应,数千人彼此对话,那么那 100 万倍的成本可能会很高。
It depends. Again, I'm coming at this from a cost point of view. I'm like, oh my god. Like if there's some combinatorial thing of thousands of people talking to thousands of people, then that 1 million X's might cost.
我对成本方面有非常不同的看法。比如,在现实中运行这些研究实际上要昂贵得多,对吧?运行任何这样的研究,你需要有人来做。你需要招募人员。这非常昂贵,有时实际上不可行。
I have a very different view of the cost side. Like running these studies in reality is actually a lot more expensive, right? Running any study like this, you've got to have people do it. You've got to sign people up. It's very expensive and sometimes not feasible to actually run the study.
但结果或你做出的决策对他们来说非常昂贵,对吧?所以花 X 百万在某件事上,而你知道整个过程成本是 1 亿,那不妨一试。那里有很多价值。这是一个小成本,但某种程度上我对成本方面感到兴奋。显然,当你部署技术时,你通常希望以能替代现有预算或基本上提高效率的方式部署,那是最好的部署方式。然而,你捕捉技术长期价值的方式实际上是论证:不,实际上是上行空间,通过使用模拟做出更好的决策,你节省或赚取了数亿甚至数十亿美元,这是一个可以提出的论点。
But the outcome or the decisions you make are very expensive on them, right? So spend X million on something that you know the overall process cost 100 million, might as well. There's a lot of value to be had there. It's a small cost, but I'm excited on the cost side actually to some extent. And obviously when you deploy technology, you often want to deploy in a way where you can replace existing budget or you can basically make things more efficient, and that is the best way to deploy. However, the way you capture the long-term value of the technology actually is making an argument that no, it's actually the upside that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars, and that's a case to be made.
随机跑题问题。如果你在做大量推理,大量多智能体工作,你是否到了训练一个非常稀疏的模型有意义的阶段,你期望进行数百万美元的运行?你是从模型架构的角度还是推理效率的角度考虑这个问题,还是你仍处于“它有效,它有效”的研究阶段?
Random tangent question. So if you're doing a lot of inference, a lot of multi-agent stuff, are you at the point where it makes sense to train a model that's very sparse, you're expecting to do multi-million dollar runs? Are you thinking about this in model architecture standpoint or inference efficiency, or you're still at the research phase of it works, it works?
我们还没有完全到那一步。效率我们确实考虑得相当多。我的意思是,这项技术现在部署在世界一些最大的企业公司中,我们确实处理大量试图同化世界人口的查询。所以效率是一个持续的问题。显然,我们不想过早过度优化。所以我不说这是当前最高优先级,但这绝对是我们相当仔细考虑的事情。
We're not super there yet. Efficiency we actually do think quite a bit about. I mean, this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process significant number of queries that are trying to assimilate the populations in the world. So efficiency is a consistent thing. Obviously we don't want to over-optimize too early. So I wouldn't say this is the highest priority right now, but this is definitely something that we think pretty carefully about.
是的。
Yeah.
还有其他案例研究吗?你提到了 CVS、盖洛普、德勤和 Wealthfront。
Are there other case studies? So you talked about CVS, talked about Gallup, Deote, Wellfront.
Wealthfront 是个有趣的案例。因为他们想做的事情之一,他们是首批想要真正进行产品测试的客户之一,这种测试超越了仅仅询问人们想法的范畴,比如行为实验等等。所以实际上我们必须处理多模态输入,包括图像,但也可以想象这些智能体在 Figma 原型或网站中穿行。我们的智能体还能做的一些事情是,你可以给它一个域名或网站 URL,它实际上会去使用一段时间。正是这类事情,Wealthfront 是最早的客户之一,并且对这个可能性非常兴奋。
Wealthfront is an interesting one. Because one of the things they're trying to do, they were one of the first customers that wanted to actually do product testing that goes beyond just asking people what they think about, let's say, behavior experiments and so forth. So there really what we had to do was reason about multimodal input. So images, but also you can also imagine like these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is you can be given a domain like or like a website URL and actually go use it for a while. It's these kind of things and Wealthfront was one of the first customers and was very excited about this possibility.
那么,人们有没有问过,是否有我们未覆盖的需求,比如 UI 测试,对吧?我想尝试新功能,想发布新功能,测试 UI,模拟人们会怎么做。你今天看到有哪些有趣的需求?
Well, have people been asking like is there any demand that we have not covered like UI testing, right? I want to try a new I want to ship a new feature, test the UI, simulate how people will do it. Any any interesting things that you're seeing demand for today?
很多需求确实来自那些历史上使用人工样本组的地方。我们现在基本上可以用智能体和合成人群来替代,但这显然不是在取代人工样本组。在很多方面,Simile 构建的模拟是有根据的。所以我的思考方式是,我们试图大规模地代表人类,从这个角度看,用例是我们预期的,但让我惊讶的是部署的规模。事实证明,在这些组织和群体中,人们每天要做很多决策,我们希望能够说我们听取了人们的意见,我们咨询了我们的用户,但实际上很少如此,因为接触人们并真正问他们很多问题是困难的,既昂贵又耗时,但最重要的是人们根本没有时间。如果我必须为某个特定供应商回答一千个调查问题,即使我想做,我也永远不会做,事实就是如此。模拟能做的就是确保人们的声音始终出现在为他们做决策的房间里。所以这个特定产品发布的所有利益相关者,理想情况下他们都被咨询了,这正是这项技术真正试图实现的。
A lot of the demand does come from basically like the places where people have historically used human panels. We can basically now replace with agents and the synthetic populations and this is obviously not replacing human panel. In many ways the simulation that Simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale and in that way the use cases are what we would expect but it's the scale of deployment that surprises me. It turns out there are so many decisions that people make every day in these organizations groups and we want to be able to say we listen to people we have consulted our users but in reality that is rarely the case because getting to people and actually asking them many questions it's difficult it's both costly time consuming but most importantly people are just not available. If I had to answer a thousand survey questions for this one particular vendor even if I wanted to do that like I would never do it and that's very much the case. What simulation can do is ensure that the voices of people is always represented in rooms where the decisions for them is made right. So all the stakeholders of this particular product launch ideally they're consulted that's what this technology really is trying to enable.
在我看来,这意味着它更偏向于消费者导向,对吧?就像任何拥有足够广泛客户群的东西,你会从你所代表的多样性中受益。对于不熟悉这个市场的人来说,有哪些粗略的统计数据?市场规模是多少,我相信你有一些大概的数字?显然市场规模是个模糊的问题。
In my mind that means it's skewed towards more consumer focus right? Like anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics just for people who are not familiar with this market in general? What's the market size that I'm sure you have some like rough numbers? Obviously market size is like a vague question.
是的。
Yeah.
但人们大概花多少钱?
But like how much do people spend?
市场研究是一个千亿美元的行业。
So market research is a hundred billion dollar industry.
嗯。
Yeah.
但关于模拟,模拟不是市场研究的工具。模拟是人类决策的工具。所以这里关于 TAM 是什么的问题实际上相当棘手,对吧?因为很容易说,市场研究的 TAM 大约是 1000 亿。那么这是 TAM 吗?其实不是,对吧?因为在很多方面,你试图为所有人类决策提供信息。你基本上试图为每一个关于人类、为人类做出的决策提供信息。那 TAM 是什么?真的很不清楚。而且说实话,你知道,我有科学背景,有研究背景。所以我进入这个领域时并没有计算“哦,人类决策的 TAM 是什么”,但我只能假设,如果我们能为每一个关于人类、为人类做出的决策提供信息,那一定很大。
But the thing about simulation is simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what is a TAM here is actually quite tricky, right? Because it's easy to say, well, market research TAM is roughly 100 billion. So is that a TAM? And not really, right? Because in many ways, you're trying to inform all human decision-making. You're trying to basically inform every decision that is made about human for humans. What is a TAM for that? It's really unclear. And I'll be honest, you know, I have a scientific background. I have a research background. So I didn't come into the field actually calculating oh what is the TAM for human decision-making but I just had to assume well if we can inform every decision that is made about human for human that has to be big.
有价值的东西。
Something valuable.
正是如此。
Exactly.
我的意思是,在某种程度上,你知道你现在是独角兽创始人,作为 CEO 你必须关心。但我确实认为,是的,我们走进这些董事会,你报价数百万美元的合同,你必须说,好吧,这是你在人力上的花费,这是我们为你节省的,而且 85% 相似。
I mean to some extent you know you are a unicorn founder now and you you have to care as a CEO. But like I I do think like yeah we go into these boardrooms with people that you're quoting millions of dollars of contracts for like you have to say well well here's what you spend on humans and here's what we save you and it's 85% similar.
当然,价值主张是我们深切关心的,比如我们真正为用户和决策者提供的价值是什么,但这也是作为创始人,我认为估值只讲述了故事非常肤浅的一面,我通常尽量不去想太多估值,因为那也不是激励团队的东西,当然也不是。你知道,再说一次,研究者的有趣之处在于,我们满足于生活在学术界,薪水接近——我的意思是,我们薪水还行,但不算多,作为学术界的研究者,但真正驱动我们的是影响力,是我们能为社会中的个人提供的价值,最终驱动我们的是影响力。我们提供的模拟是否对人们的决策产生真正的影响,以推动社会进步?如果答案是肯定的,那么是的。我的意思是,那一定是伟大的生意,我们从数字中看到了这一点,我们确实深切关心那个上行故事,但那是更高层次的东西。
And certainly the value case is something that we care deeply about like what is the value that we actually provide to the users and the decision makers but this is also where like you know as a founder I think valuation only tells one very superficial aspect of the story and I try not to think too much about valuation in general because that's not what also motivates the team or certainly doesn't. You know, again, the interesting thing about researchers is we are happy living in academia getting paid next to I mean we get paid okay I mean we don't get paid that much I mean as a researcher if you're in academia but it's the impact and it's the value that we can provide to the individuals in the society that really drives us and in that way ultimately what drives us is the impact. Does the simulation we provide have a real impact in people's decision-making in ways that progresses our society forward? If the answer is yes, then yes. I mean that has to be great business and we see that in numbers and we do care deeply about that upside story but that's the higher bit.
你有任何时间线预测吗?我们谈到了你提出的模拟的缩放定律。好吧,也许有一天我们可以模拟如何解决气候变化。我们现在在哪里?如果那不是最终状态,最终状态是什么?进展看起来是什么样的?
Do you have any timeline predictions? So we talked about scaling laws of simulations you brought up. Okay, maybe one day we can simulate how to solve climate change. Where are we now? If that's not the end state, what is an end state? And what does progress look like?
你知道,所以我有时告诉人们,模拟作为一个行业,感觉很像 GPT-3.5、GPT-4 在 AGI 传奇中的位置,基本上是我们现在拥有的技术足够强大,能够在我们正在攻克的垂直领域产生真正的影响,同时还有很多进展尚未到来,我认为这就是现状。所以在我看来,我确实认为在数据和算法方面都会继续有突破。未来几年也会出现更激进的 Scaling。但我认为这大致就是我们所在的位置。
You know, so what I sometimes tell people is simulation as an industry. It feels a lot like where GPT-3.5, GPT-4 was for the AGI saga which basically is we have now technology that is powerful enough to do real damage on the verticals that we are tackling at the same time there's a lot of progress that is yet to come and that's I think where this is. So the way I see it I do think there will continue to be breakthroughs both in data and obviously in algorithm. And there will be much more aggressive scaling that will also happen over the next few years. But I think that's roughly sort of where we are.
我想这就是大致的话题范围了。还有什么我们应该问你的,或者你希望人们更多地问你关于 Simile 的事情吗?
I think that was about the rough set of topics. Anything else that we should have asked you or you wish people asked you more about about Simile?
你知道,我认为对我来说,模拟真正迷人的地方在于,它是一项非常有影响力的技术,但实际上也是一项非常有趣的技术,无论是对人类社会、我们的哲学意味着什么,以及我有时解读模拟的方式,回到我的背景,实际上正如我之前提到的,我的职业生涯始于画家。那是一种职业追求。我实际上画人物画。所以我最初是在现实主义工作室接受训练,那是我花了很多年做的事情。
You know, I think what's for me what's actually quite fascinating about simulation is it is very impactful technology but actually is also very interesting technology both in terms of like what it means for human society, our philosophy, and the way I sometimes interpret simulation is so going back to my background I actually as I mentioned earlier I started my career as a painter. It was a professional pursuit. And I actually did painting for figures. So I got my training originally in sort of the realism studios and that's what I spent a lot of my years doing.
模拟很像绘画,对吧?最好的画作能让你对想表现的对象有深刻的理解。它总不是完美的再现,没有一幅画是完美的,总有些细微的差异和出入。但它所做的是试图突出对象最重要的方面。
Simulation is a lot like painting, right? And the best paintings teach you something deep about the subject that you're trying to represent. And it is always not a perfect representation. No painting is perfect. There's always some small differences and discrepancy. But what it does is it tries to highlight the thing that matters the most about the subject.
本质的精髓。
The essential essence.
是的。他提到了你的一些作品。正好展示一下。对,这些就是一些作品。这其实是我还是研究员时维护的个人网站。我想很多人会说,像毕加索,任何后现代的东西都非常注重本质。
Yes. He's brought up some of your work. Just nice to put it up. Yeah. So these are some of the works. So this is actually from my personal website that I maintain when I was still a researcher. I think a lot of people will say like, you know, like a Picasso, like anything postmodern is like very much focused on the essence.
是的。
Yes.
对。嗯,是的。但我不知道其中有没有哪一幅能唤起你想讲述的故事。
Right. Um Yeah. But I don't know if any any one of these evokes something that you like to tell the story of.
不,这就是那种事情,你知道,每一幅画、素描,无论是什么,都在试图把你对主题的深刻感受浮现在表面。你知道,当我还是画家和艺术家时,我真正深切的关注点其实是人类生活中更平凡的一面。这实际上出现在我做过的一些作品中,比如我对一个乡村小镇做了完整的研究,基本上到处给人拍照,他们并没有做什么特别的事,只是过着日常生活。我觉得那是最有趣的。我是有这种视角的人,世界围绕着一个分形形状展开,你有两个选择去理解这个分形。要么向外走,尽可能探索,以理解分形的更广阔形状;要么向内走,因为外部形状与内部形状相似。理解平凡的一面,模拟在这方面有很多相似之处,对吧?你试图理解人们最平凡的方面,当它们放在一起时,会教你关于那个个体和社会的深刻道理。所以我认为这就是模拟的有趣之处,就像 AGI 帮助我们更好地理解或真正批判性地思考人性和人类智能一样。模拟实际上是一种理解人类社会和集体生活的练习。所以我觉得这特别有趣。
No, it's it's one of those things where, you know, each of these paintings, drawings, whatever may be, it is trying to surface something about the subject that you feel deeply about onto the surface. You know, when I was a painter and artist, the topic that I cared really deeply about actually was the more mundane aspect of human lives. This actually shows up in some of the work that I've done where I did this entire sort of study of a rural town where I basically went around and took photos of people not really doing anything special, but just living their everyday lives. I thought that was the most interesting thing. I'm somebody who has this perspective where the world is oriented around this fractal shape and you have two choices to understand the fractal shape. You either go outward and try to explore as much as you can to understand the broader shape of the fractal, or you go inward because the outward resembles the inward shapes. And understanding the mundane aspect of it was very much that simulation has a lot of this, right? You're trying to understand even the most mundane aspect of people, when put together teaches you something really deep about that individual and the society. So I think that's what's interesting about simulation, sort of the same way that AGI helped us better understand or really think critically about humanity and human intelligence. Simulation is really an exercise of understanding more about human society and our collective lives. So that I find to be particularly interesting.
是的。现在你让我想起一些最好的传记作者、纪录片制作人甚至摄影师,他们给你拍照。但在给你拍照之前,我必须花时间,必须像跟着你一个星期来了解你,你知道,有些艺术家会这样做。你的作品中有部分,有一本非常著名的书叫《工作》。我不知道你是否之前被提到过。
Yeah. Now you're reminding me that some of the best biographers, documentarians and even photographers, they're taking a photo of you. But before I take a photo of you, I must spend, I must like follow you for a week just to understand you, you know, which some artists do. Part of your work, there's a very famous book called Working. I don't know if you've been referred to it before.
是的。它非常非常有名,你知道,甚至到了有维基百科页面的程度,这种对人们生活的深入理解和采访,看似平凡,却以非常引人入胜的方式讲述。是的,也是 1970 年代。
Yeah. It's very, very famous, you know, to the point of having a Wikipedia page about this kind of really in-depth understanding and interview of people about their lives, which seems mundane but is told in a very compelling way. Yeah. 1970s as well.
好的。那是一个了不起的十年。
Okay. It was an amazing decade.
实际上在结束问题之前,你说你开始了类似的问题,对吧?如果我们现在做,10 年后,我们能模拟什么?如果你取得了重大进展,你会模拟什么?有没有我们提出的问题之外的问题?你认为最有影响力的是什么?你会展望 10 年后吗?
Actually before closing question, you said that you started similar question, right? If we do that now 10 years down, what can we simulate? What would you simulate if you've made significant progress? Are there any questions outside of the ones that we brought up? Anything that you think is most impactful? Anything that you would go vision 10 years out?
在很多方面,正如我提到的,我是一个非常注重影响力的人。所以真正激励我的是,我想在 10 年后问,作为社会,我们必须问的最重要的社会问题是什么。我很想解决这个问题。比如,我们需要全民基本收入(UBI)吗?这可能是一个有趣的问题。
In many ways, as I mentioned, I am somebody who is very much impact driven. So the what would actually inspire me is I would want to ask 10 years later what would actually be the most important societal question that we as a society have to ask. I would love to tackle that. Like for instance, do we need UBI? That could be an interesting one.
哦,有人做过吗?
Oo, has anyone done that?
嗯,我是说,你知道,我们正在考虑。
Well, I mean, you know, we're thinking about it.
我能获得访问权限吗?我们能直接
Can I get access? Can we just
OpenAI?这就像现在的琐事,比如开放,或者我认为 Sam Altman 实际上资助了在非洲的一项研究,答案是否定的。
OpenAI? This is like just trivia now, like opening or I think Sam Altman actually funded a study on this in Africa and the answer was no.
答案是否定的。但这是关于实施的问题吗?
The answer was no. But what was it something about the implementation?
是的。
Yeah.
但这就是问题所在。
But but this is the thing.
你看,当 Sam Altman 资助这项特定研究时,他花了 1400 万美元,相当多。但这就是问题所在。这就是你想运行模拟的原因。你在这项研究上花 5 年、4000 万美元,得到一个发现。但如果你能立即运行模拟很多很多次,那就是价值所在。
See when Sam Altman funded this particular study, he spent $14 million, quite a bit. But this is the thing. This is the reason why you want to run a simulation. You spend 5 years, $40 million on this one study and have one finding. But if you can run simulation many many times instantly, then that's the value.
我觉得那个你本可以在模拟中完成。就像如果你能做住房研究,你就能做 UBI 研究。我是说,拜托。
I feel like that one you could have done in a simulation. Like if you can do the housing study, you can do the UBI one. Like I mean come on.
我认为有时候人们会花钱是因为他们想验证你的想法,对吧?就像有时候你只是想知道它实际上是否正确?你必须测试它。
I think sometimes people will spend the money because they want to verify what you think, right? Like sometimes you just want to is it actually right? Like you got to test it.
好的。结束问题。我们现在处于模拟中的概率有多大?
Okay. Closing question. What are the chances we are in a simulation right now?
这是一个有趣的问题,在某个时刻我只是回答是的,我们肯定在模拟中。但我确实觉得,无论我们是否在模拟中,我不认为这会让我们的体验变得不那么真实,我认为这基本上就是我所相信的。也许我们生活在模拟中,也许不是,但
It's a fun question and I at some point I just answer yeah we're definitely in a simulation. But what I do feel however is whether we are in a simulation or not, I don't think that makes our experience any less real, and I think that's fundamentally like what I believe in. Maybe we live in a simulation maybe not but
对我们来说是真实的。是的。
It's real to us. Yeah.
是的。对我来说,我真的不在乎。
Yeah. For me I don't really care.
是的。除非你死了,然后在更高的层级醒来。
Yeah. Unless you die and you wake up in like the level higher.
那会很有趣。
That would be interesting.
我觉得你不会在乎。你知道,一旦你死了,然后你发现你就像
I feel like you wouldn't care. You know, once you die, then you find out you like
我死的时候再担心。
I worry about it when I die.
我想另一件事,好吧,我喜欢这个问题的数学答案,即你在模拟中的可能性数量远远超过你不在模拟中的可能性数量。是的。除了最简单的答案,即让你成为一个文明在计算上非常昂贵。嗯,好的,你非常慷慨地付出了时间。恭喜你所有的成功。你知道,我在你的 smallville 论文之后遇到你,不知道你能建立如此庞大的公司,然后现在你就像,嗯,这是一个千亿美元的市场,但这只是我们开始的地方。所以这非常令人兴奋。千亿美元的市场不是终点,那只是如果你想得太小的一部分。
I think the other thing that I Okay, so I like the mathematical answer to this, which is like the sheer number of possibilities that you are in a simulation far outweigh the sheer number of possibilities that you're not. Yes. Except for the simplest answer which is it is computationally very expensive to have you be a civilization. Um okay great you've been very generous with your time. Congrats on all your success. You know I met you just after your smallville paper and had no idea that you could build like such an enormous company and then now you're like well it's a hundred billion dollar market but that's just where we're starting. So this is very exciting. $100 billion market was not the TM that was only part if you're thinking too small.
嗯,我确实相信,我最后的说明可能又是,我喜欢科幻。你看科幻中任何先进的文明,都有两个孪生支柱技术,一个是某种形式的 AGI,另一个是模拟。所以我认为这里的市场相当大。
Well, I do believe that I made you my final note here might be again I love science fiction. You look at any advanced civilization in science fiction, there's two twin pillar technology, one's AGI in some form and the other is simulation. So I think the market's pretty big here.
是的。跟我们说说公司吧。你们刚刚筹集了很多资金。
Yeah. Tell us about the company. You guys just raised a lot.
你们一半是研究实验室,一半是公司。我猜你们在招人。你们在哪里?
You're half a research lab, half a company. I guess you're hiring. Where are you based?
是的,我们的总部在 Mission Rock,离我们现在所在的地方不远。我们在旧金山,但也在沿海地区。我们的团队——我想说总部在旧金山,很多技术人才也在旧金山。我们其实在纽约刚开了一个较小的办公室。作为一家公司,我们很有趣,因为今天显然有 AI 新实验室和 AI 产品公司,而我们的公司两者兼是。这家公司由四位联合创始人创立:我、Michael Bernstein、Pong Laney、Ellen、Michael Percy,我们都是研究者。当然,Michael 是 2013 年 ImageNet 启动 AI 革命的合著者之一,在以人为本的 AI 方面发挥了重要作用,创造了“基础模型”这个词,显然是当今伟大的 AI 研究者之一。Laney 是我的商业搭档,她为所有增长最快的 AI 原生公司提供从 C 到 AMV 的支持。但我们的公司有这样的基因:我们正在创造的技术愿景在不断发展。我们招的人基本上都是我的实验室伙伴。我们现在大约有 60 人,其中 15% 到近 20% 的人来自我的实验室。这其实很有趣,因为很多人后来去了 AAI、Google Gemini 等地方。我们已经好几年没有真正聚在一起工作了。但现在他们回来了,真正在构建这个我觉得非常令人兴奋的愿景,而且这种兴奋是共享的。所以在我们公司,有一群研究者试图做一些没人正在做的事情,我们认为这可能是最有影响力的。但与此同时,这也是今天就能产生影响的技术。所以我们有一群出色的工程师、产品人员和设计师,他们和我们一起,试图想象如何帮助人们理解模拟能做什么,并据此做出现实世界的决策,同时将其部署到当今一些最大的客户那里。这感觉很独特。
Yeah, so we're based in Mission Rock, not too far away from where we are right now. So we're in SF. But we're also by Coastal. So we have our team—I would say our headquarters is in SF, and we have a lot of our technical talent in SF. We do have a smaller office that just opened up actually in New York. We are as a company an interesting one in that today obviously there are AI neolabs and then there are AI product companies. Similarly, truly is both. So this is a company that was founded by four co-founders: myself, Michael Bernstein, Pong Laney, Ellen, Michael Percy, and I are all researchers. So of course Michael was one of the co-authors of the ImageNet kickstarter, the AI revolution back in 2013, has been instrumental in human-centered AI, coined the term foundation model, and obviously one of the great AI researchers today. And Laney is my business counterpart, where she lends all the fastest-growing AI-native companies from their C to AMV. But we have this DNA at the company where the vision of the technology that we're creating is continuously developing. We are getting people who were basically my labmates. We are right now about 60 or so people. 15%, almost 20% of the company population actually are just my labmates from my person's lab. And it's actually quite fun because many of them then had gone on to open AAI, Google Gemini, and these places. And so it's been a few years since we really got together and had a chance to work together. But now they're coming back and really building out this vision that I find to be quite exciting, and that excitement is shared. So there is that motion at simile where we are a group of researchers trying to do something that no one is working on, that we find to be the most impactful potentially. But at the same time, this is again technology that can make impact today. So we have an amazing group of engineers, product people, and designers who are sitting here with us, basically trying to imagine what it looks like to help people understand what simulation can do and make real-world decisions with this, having both and then deploying it to some of the largest customers in the world today. It feels quite unique.
是的,这非常吸引人。其中一部分是行动号召,比如你们在招什么人?你已经做了一部分,就是你有一个非常有才华的团队。你们在招什么人?
Yeah, it's very compelling. One part of it was this is the call to action like who are you hiring? You've done part of it which is you've got a very talented group. Who are you hiring?
比如什么职位?
Like what roles?
说实话,目前我们在招一小部分人。我们总是很高兴能引进出色的研究人才。
So honestly at this point we are hiring a small section. We are always excited to bring on amazing research talent.
所以如果你有兴趣和我们的实验室伙伴一起工作,我们总是欢迎出色的研究者。但我们也招聘出色的工程师,其中一些是我最尊敬的。很多人其实来自我们有个人关系的地方。所以很多成员来自 Figma、Notion、Harvey 等公司。但更广泛地说,也来自我们团队非常欣赏的公司。所以产品端和基础设施端的工程师,我们都在寻找。
So if you're interested in working with our labmates, we're always welcoming of amazing researchers. But also we hire amazing engineers, some of whom I respect the most. Many of them actually come from places where we have personal connections with. So many of the members are from Figma, Notion, Harvey, and so forth. But also more broadly from the companies that we as a team heavily admired. So engineers both on the product side and infra side, we're all looking for those hires.
好吧,很多人。我觉得你说得很有道理。所以,谢谢,我们模拟中见。
Well, lots of people. I think you make a really good case. So, thanks, and we'll see you in the simulation.
太棒了。到时候见。
Amazing. See you all there.