Simulating Human Societies with AI: From Smallville to Real-World Applications
打开互动全文版(中英对照 + 朗读 + 问答)→Simile 创始人 June 探讨了生成式智能体在小镇实验中模拟涌现社会行为的过程,以及利用模拟指导社会的愿景。
June, founder of Simile, discusses how generative agents in Smallville simulated emergent social behaviors, and the vision for using simulations to guide society.
我是一个深受科幻小说启发的人。当你阅读那些描述技术成熟度足够高的社会的科幻作品时,你总会看到两大支柱:某种形式的 AGI(通用人工智能)和某种形式的模拟,它们真正帮助引导社会。我今天确实看到了一个机会,可以率先尝试构建这种模拟。即使在 5 年前,我也不会这么说。但这是我们多年来深入研究后逐渐建立起来的信念。
I am somebody who is quite inspired by science fiction. When you read science fiction that covers societies that have progressed far enough in technological maturity, you always see two pillars: some version of AGI and some version of simulations that really help guide the society. I do see an opportunity today to really take the first crack at building the simulation. I would not have said that even 5 years ago. But that is a conviction that we have built up over the years as we are going deep into this research.
今天,我们很高兴邀请到 Simile 的创始人兼首席执行官 Joon。Simile 正在建立一个应用 AI 实验室,模拟人类行为和社会。我非常兴奋能请你来讨论你们正在构建的东西。
Today, we're delighted to have Joon, founder and CEO of Simile. Simile is building an applied AI lab simulating human behavior and societies. And I'm very excited to have you here to discuss what you're building.
我也很兴奋。谢谢邀请。
Same here. Thanks for having me.
好的,带我回到 2023 年 4 月,加利福尼亚州斯坦福。具体来说,是斯坦福的 Smallville。那是什么?
Okay, take me back to April 2023, Stanford, California. Specifically, Smallville, Stanford, California. What was that?
Smallville 是我们在斯坦福运行的一个项目。我们的想法是,我们观察到大型语言模型现在可以编码大量人类行为,这些行为嵌入在来自网络和社交媒体等的训练数据中。如果你从正确的角度去探究,你实际上可以从这些模型中获得很多微观行为。所以,给定一个非常具体的场景演示或描述,比如“人物 X 会做什么?”,它实际上会生成非常有趣的行为。我们发现这非常有趣,并且认为这是我们一直在等待的、用于创建真正复杂的智能体行为的要素。因此,Smallville 实际上是一个实验,我们决定尽可能推动这一点,看看由这些智能体创建的社会会是什么样子。所以,我们基本上创建了生成式智能体,它们与具有记忆、规划和反思能力的生成式 AI 模型配对,从而创造出这些智能体生活在这个小镇中的生活体验。Smallville 基本上是一个有 25 个智能体居住的游戏小镇。每个智能体都有一个人物描述,但它们实际上会早上起床、做日常事务、去上班、像人一样建立关系,并且会出现涌现现象,比如举办派对等等。这就是我们进行的实验。
So, Smallville was a project that we were running at Stanford where the idea was that we made this observation that large language models can now encode a lot of human behavior that is embedded in its training data from the web and social media and so forth. If you sort of probe it from the right angle, you can actually get a lot of micro behaviors out of these models. So, given a very specific demonstration or description of a situation, what would person X do? And it would actually generate really interesting behaviors. We found that to be so interesting and we found that to be the ingredient that we have been waiting for for creating really complex agentic behaviors. So, Smallville actually was an experiment where we decided that if we push this as far as possible, what would a society that is created by these agents look like? So, we basically created generative agents that is paired with generative AI model with memory, planning, and reflection to basically create this lived experience of agents living in this small town. So, Smallville was basically a game town of 25 agents living in it. Individual agents had a description of persona, but they would actually wake up in the morning, do their routines, go to work, actually have relationships sort of like people would, and they would actually have emergent phenomena like having parties and so forth. So, that was the experiment that we ran.
实验中最令人惊讶的事情是什么?
What was the most surprising thing to come out of the experiment?
其中一个令人惊讶的事情是——实验,模拟本身实际上发生在情人节的前一天。所以,你实际上会看到这些智能体,其中一个智能体在想:‘嗯,我经营一家咖啡馆。’她是一位咖啡馆老板,名叫 Isabella。她想了想:‘如果我能办一个情人节派对,邀请很多朋友和顾客,那就太好了。’所以,你实际上会看到她在情人节前一天四处走动,为派对收集材料,告诉她的顾客:‘嘿,我们要办这个派对,请来吧。’到了情人节那天,你实际上会看到这个涌现出来的派对真的形成了,所有这些智能体都来到了咖啡馆。
So, one of the surprising things was—the experiment, the simulation itself actually takes place the day before a Valentine's Day. So, you actually see these agents, one of the agents actually thinking, 'Well, I run a cafe.' So, she's a cafe owner, her name is Isabella. She goes and thinks, 'It'd be great if I can do a Valentine's Day party where we invite a lot of friends, customers.' So, you actually see her on the day before Valentine's Day going around, actually gathering materials for the party, actually telling her customers, 'Hey, we're going to have this party. Please come.' And on the day of Valentine's, you actually see this emergent party that actually gets formed with all these agents coming to the cafe.
有人没被邀请吗?
Did anyone not get invited?
嗯,有些人确实收到了邀请,但他们忘了。这是确实发生的一件事。有些智能体没有明确收到邀请,但我们有一个收到邀请的智能体 Claus,他决定约他的暗恋对象出去约会。所以,他实际上会带约会对象来,他们会在咖啡馆里办派对。所以,相当超现实。
Well, some of the people did get an invitation, but they forgot. That's one thing that did happen. Some of the agents did not explicitly get invited, but we had one agent who got the invite, Claus, who decided to ask his crush out on a date. So, he would actually bring in the date, they would actually have a party at this cafe. So, quite surreal.
那么,你最初是怎么想到要构建 Smallville 的呢?你是在研究人类心理学和社会行为,还是说这来自于技术本身?
So, how did you end up building Smallville in the first place? Like, were you studying kind of human psychology and social behavior, or was this coming from the technology out?
我的团队一直对模拟感到兴奋,我们很早就看到了模拟失败的愿景。所以,我作为斯坦福研究员的职业生涯实际上始于 2020 年。那一年 GPT-3 即将问世。它还没有完全准备好,但即将发布。我们开始看到它的第一批演示。在我第一年,我们与许多斯坦福研究员一起写了一篇论文,题为《基础模型的机会与风险》,由我的一位联合创始人 Percy Liang 领导,他现在是斯坦福基础模型中心的负责人。当我们写那篇论文时,我真正关注的部分是:嗯,这是一类我们过去从未见过的新模型,这些模型可以以我们过去没有的方式实现高度泛化。我开始思考:如果我们能想象我们可以用这些模型创建什么样的交互,那会是什么?当时我的许多同事对这些智能体或模型能够进行分类或简单生成感到惊讶。这确实令人难以置信,因为这些模型并没有真正被教导去做这些事。但让我惊讶的不是这些模型能做到这一点,因为从交互的角度来看,我们早就知道如何做到这一点。有趣的部分是:这些模型实际上可以包含人类行为。这意味着什么?如果我们尽可能推动这一点。所以,我来自的研究传统包括我们所说的社会计算。人机交互中的社会计算实际上与这个想法有关:我们如何构建一个更好的技术平台,以实现社交互动和协作?构建社交平台最困难的挑战之一不一定是测试系统的 UI/UX,而是当你有几十人、几百万人、甚至几十亿人时,所有这些人是如何聚集在一起产生既有好也有坏的涌现现象的,以及我们如何为规模设计?到目前为止,我们还没有一个工具可以让我们测试这一点。今天我们测试它的唯一方法就是实地测试。你发布原型,看看会发生什么,有时这实际上会付出真正的代价。显然,它在人力和时间方面成本很高。但与此同时,如果你有一个糟糕的设计,想象一下社交媒体上的一个信息流更有可能传播某种负面情绪,那么这显然是我们想要避免的。但现在这只能在实地测试中检验。所以,我们想看看我们是否真的可以创建一个模拟,让你测试这一点。因此,在 2022 年,也就是生成式智能体出现的前一年,我们写了一篇名为《社会拟像》的论文,这实际上是我们最终写出的智能体论文的前身。核心论点是:想象你正在构建一个 subreddit。
So, my particular team has been excited about simulations, and we saw the vision of simulation failure early on. So, my career as a researcher at Stanford really started back in 2020. That was the year when GPT-3 was about to come out. It wasn't quite there yet, but it was just about to come out. We started to get its first demos. And my first year, we wrote this paper called 'Opportunities and Risks of Foundation Models' alongside many of the Stanford researchers, and was led by one of my co-founders, Percy Liang, who is now the head of the Center for Foundation Models at Stanford. And when we were writing that, the part that I was really focused on was, well, here's a new class of models that we have not seen in the past, that these models can be very generalizable in ways we didn't quite have in the past. And I got into thinking, well, if we can imagine the kind of interaction we can create with these models, what would that be? And many of my colleagues back then were surprised that these agents or these models can do classification or simple generation. And that was really incredible to see because these models didn't really know or weren't really taught to do that. But the part that was surprising to me wasn't that these models can do that because from interaction perspective, we've known how to do this for a long time. The interesting part was, well, these models can actually include human behavior. What does that mean? If we were to push this as far as possible. So, part of the tradition I come from research included what we call social computing. And social computing within human-computer interaction really has to do with this idea of how can we build a better technological platform that would enable social interactions and collaboration. One of the most difficult challenges of building a social platform is not necessarily testing the UI/UX of the system, but it's more about when you have tens of people, millions of people, and down the line billions of people, how do all these people come together to create the emergent phenomena that's both good and bad, and how can we design for scale? And so far, we didn't really have a tool that would enable us to test for that. The only way we test it today is you basically field test it. You release your prototype, see what happens, and sometimes it actually comes at a real cost. Obviously, it's high cost in terms of human hours and the time it takes. But at the same time, if you have a bad design, imagine you have a feed on social media that is more likely to propagate certain emotion that is negative, then obviously that is something that we want to avoid. But this now gets tested in the field. So, we wanted to see whether we can actually create a simulation that would actually let you test for this. So, 2022, this was actually a year before generative agents, we worked on a paper called 'Social Simulacra', which actually really was the precursor to the agent paper that we ended up writing. The core thesis was, imagine you're building a subreddit.
你是一个子版块的设计师,想看看人们会在子版块里做什么,这对即使是有经验的设计师来说也是一项出奇困难的任务。我们基本决定:“嘿,我们有这个模型,看起来挺独特。我们用这个模型来创建整个子版块的模拟吧。”所以,你定义目标,定义管理策略,然后用成千上万个——当时我们不叫它们智能体,而是叫它们角色——但用成千上万个角色填充它。这基本上是 Mob Book 的 22 版本,有趣的是它真的回来了。当我们看到这个时,我们确实从中得到了很多非常重要的见解。哪些是好的行为?我们模拟了一个社区,其核心理念是人们互相讨论匹兹堡的观光景点。突然之间,你开始看到这些角色真的合作讨论:“嘿,XYZ 地方太棒了。你想一起去旅行吗?”并在模拟的子版块中实时规划这些旅行。所以,我们就这样兴奋起来了。因此,我们很早就看到了愿景、兴奋点和潜在应用。但之后我们必须做的工作是展示如何超越简单的角色,创建能够随时间思考的复杂智能体,因为我们想模拟社会的纵向方面。然后还要验证这些模拟在实践中是否准确。
You're a designer on a subreddit, you want to see what people might do in the subreddit, which is surprisingly hard task even for practiced designers. And we basically decided, "Hey, we have this model, seems unique. Let's use this model to create simulations of the entire subreddit." So, you define the goal, you define the moderation strategies, and you populate it with thousands of what back then we didn't call them agents, but we called them personas, but populate it with thousands of personas. This is basically 22 version of Mob Book, which is quite interesting that it actually came back. And when we saw that, we actually got a lot of really important insights out of this. What are the good behaviors? We actually simulated a community where the entire idea was for people to discuss with each other the sightseeing of places to sightsee in Pittsburgh. And all of a sudden, you start to see these personas actually collaborate to actually discuss, "Hey, XYZ places are amazing. Do you want to actually go to a trip together?" And actually plan those trips live in the simulated subreddit. So, that's how we got excited. So, we saw the vision and the excitement and the potential applications very early on. But then the work that we had to do was then demonstrating how can we go beyond simple personas to create complex agents that actually can think over time because we want to simulate the longitudinal aspect of our society. And then actually validating that these simulations are actually accurate in practice.
有没有一个模型演化的节点,让你觉得:“好了,我们做到了。模型已经足够好,能够忠实再现人类社会了?”
Was there a point of model evolution at which you felt like, "Okay, we're there. The models are good enough for us to actually have a faithful representation of human society?"
所以,GPT-3 刚出来时,Social Simulacra 就是用 GPT-3 构建的,非常粗糙。它没有做任何指令微调,也不遵循你的指令。所以,仅仅为了让它听你的话并做你想做的事,你不得不在提示上耍一些奇怪的花招。但你能看到它的潜力。模型确实编码了大量人类行为,你能看到轨迹。当我们发表生成式智能体论文时,虽然还不是 ChatGPT,但我们有了指令微调。所以,我们可以构建更复杂的智能体,能够推理自己的记忆。这在做 Social Simulacra 时是不可能的。自那以后,模型当然也改进了。所以,今天我们的基础模型已经达到了一个点,我们可以真正想象构建这类应用了。但我觉得这里有趣的是,今天如果你看很多大语言模型公司,无论是 OpenAI、Anthropic 还是许多新成立的实验室,他们创建的模型,我认为他们的北极星是类似于构建超级智能机器。这些机器应该是理性的,应该非常擅长解决有客观答案的技术问题。
So, GPT-3 when it came out and Social Simulacra was built with GPT-3 and was very junky. It didn't do any instruction tuning. It did not follow your instructions. So, just to have it to listen to you and do what you wanted to do, you had to do some weird tricks with prompting and so forth. But you could actually see the promise. The model actually had encoded a lot of human behavior and you could actually see the trajectory. And when we had the generative agents paper, wasn't quite ChatGPT, but we now had instruction tuning. So, we could actually build much more complex agents that can reason about its memory. That wasn't really possible when we did Social Simulacra. And since then of course the models have improved. So, where we are today is the models at its foundation level have reached a point where we can actually imagine building these kind of applications. Now the part that I do think however that's quite interesting here. Today, if you look at many of the large language model companies, whether it's OpenAI, Anthropic, and many of the new labs that are getting formed, the models they are creating are models that I would consider to be their North Star to be something that is similar to let's build a super intelligent machines. These machines are meant to be rational. And these machines are supposed to be really amazing at technical problems that have an objective answer.
那么,也许那甚至不是对真实人类社会的最佳模拟。
So, maybe that's not even the best simulation of true human society then.
结果发现人是非理性的,对吧?我们有很多主观的价值观、偏好和品味。所以,你实际上开始看到随着模型规模增大,它在预测和模拟人类行为方面的能力出现分歧。所以,我们目前这种建模范式在真正模拟人类的能力上已经有些停滞了。所以,它处于一个不错的起步基础水平,但要让它真正出色,我们确实需要下一个前沿,更侧重于模拟人们的多样性。
Turns out people are irrational. Right? We have a lot of subjective values, preferences, and taste. So, you actually start to see divergence in model size going up and the performance in its ability to predict and simulate human behavior. So, we have sort of plateaued with current modeling paradigm, our ability to really simulate humans. So, it is sort of at the starting good foundational level, but to make it really amazing, we do need the next frontier that is more geared towards actually modeling people's diversity.
非常有趣。你是在什么时候意识到你在 Smallville 所做的工作可以变成一家公司的?
Very interesting. At what point did you realize that what you did with Smallville could become a company?
对。所以,这个应用的潜力在早期通过 Social Simulacra 等模拟就给了我很大启发。但随着时间的推移,我意识到研究和公司有着非常不同的功能。如果你想做广度研究,研究是一个绝佳的工具。你在实验室里,周围是一群非常聪明的人,每个研究人员都有自己的一个小课题。他们去探索,其中一些课题会开花结果,成为惊人的研究成果。但我们不一定以完成工作而闻名,我们通常不是把研究影响带到现实世界的人。公司是深度研究的机器。你对某个领域有信念,找到一座你想攀登的山。这个工具让你汇集资源和一群优秀的人,毫不犹豫地追求一个单一的愿景。我们是在生成式智能体之后大约半年获得这种信念的。在最初的生成式智能体论文之后,我们收到了大量兴趣,最初来自社会科学家,他们想在我们的平台上运行实验和所有随机对照试验。很快之后,许多《财富》500 强公司看到了这个演示,他们的董事会成员和 CEO 有时会来斯坦福参加 Soda,他们开始问:我们运行所有这些调查和实验,有很多关于市场的研究问题我们今天无法回答。我们能在模拟中运行吗?这让我非常感兴趣,因为它展示了一条清晰的研究影响现实世界的路径,而我们并不总是有这样的机会。所以那时我们决定要验证模拟是准确的。于是我们出去创建了 1000 名美国人口的模拟。我们证明了使用我们的架构和模型,我们可以预测人们的行为,准确率达到人们自我复制的 85%。当我们看到这个时,我们认为,好吧,这是一个我们可以放心提供给用户作为模拟他们重要决策的平台。所以那时,联合创始人——我、Percy 以及 Michael Bernstein(他是斯坦福的研究员和我的导师),他们俩实际上都是我的导师。我们三个人已经一起工作了五年,现在在 Similac 已经六年了。但就是在那个时候,我们聚在一起进行了最初的对话:这能成为一家公司吗?
Right. So, again, the promise of the application was something that I was very much inspired by early on with simulation with Social Simulacra and so forth. But, the part that I realized over time is research and a company have very different functions. Research is an amazing vehicle if you want to basically do breadth research. You are in a lab surrounded by really smart set of people, and each of the researchers owns a small piece of thesis. And they go explore some of those thesis blossom into amazing research product. But, we're not necessarily known for finishing our job. We're not usually the one to bring that research impact to the real world. Company is a machine for depth research. You have a conviction on an area. You find a hill that you want to climb. This is the vehicle that let you put together resources and an amazing group of people to go after a singular vision without hesitation. And we got that conviction, I would say about half year after Generative Agents. After the original Generative Agents paper, we got so much inbound interest, initially from actually social scientists who wanted to run their experiments and all the RCTs on our platform. Then very soon after, many of the Fortune 500 companies who saw this demo and their board members and CEOs who sometimes visit Stanford for Soda and they start asking well we go run all of these surveys and experiments and there's so many research questions about the market that we cannot answer today. Can we run that in simulation? That sort of really intrigued me because that showed a clear line towards a real world impact for research which is not always the case that we have that kind of opportunity. So that is when we decided we actually want to validate the simulations are accurate. So we went out and actually created simulations of 1,000 people of the US population. We demonstrated that using our architecture and the models we can actually predict people's behaviors 85% as accurately as people replicate their own. When we saw that we thought okay, this is something that we feel comfortable providing to our users as a platform for simulating their really important decisions. So that's when the co-founders, myself, Percy as well as Michael Bernstein was a researcher and my advisor at Stanford. Both of them were actually my advisors. So the three of us have been working together for five years and now at this point at Similac six years. But that's when we got together to have the initial conversation of can this be a company?
明白了。太棒了。也许给我讲讲今天一个客户从开始到结束的互动过程。比如,典型的客户是谁,来自哪个部门,他们来找你,问什么,你提供什么产品或服务给他们?
Got it. Amazing. Maybe walk me through a customer engagement end-to-end today. Like who's a canonical customer and in which department and they come to you, what are they asking you and what product or service do you deliver to them?
好。也许我可以举一个具体的例子。CVS 在过去大约半年里一直与 Similac 合作。他们是一个很棒的合作伙伴。
Right. So maybe an example that I can give to make this concrete. So CVS has been partnering with Similac for the past I would say nearly half a year. And they've been an amazing partner.
我们最初与 CVS 的主要买家取得联系的方式是,那位负责人是负责人类洞察的高级副总裁。起源故事是,他读了我的论文,该论文验证了智能体模拟,并认为:“我们必须把这项技术带到 CVS,因为今天我们受限于能够实地测试的问题数量。我们也受限于人类社会的物理现实。做调查和实验是一回事,但如果你最终想模拟整个市场,并真正绘制出你向领导层建议的决策的所有二阶影响,那完全是另一回事。”所以,他一直在寻找解决方案,而他的表弟恰好认识我,并告诉我们的买家 Shri,这篇论文的作者实际上正在寻找创业机会。这就是我们建立联系的方式。
The way we initially got in touch with our main buyer at CVS, the lead is a senior VP who leads human insights. The origin story there was he basically read my paper that validated the agent simulations and thought, "We have to bring this to CVS because today we are bottlenecked by the number of questions we can field test. And we're also bottlenecked by truly the physics of human society. It's one thing to ask surveys and experiments. Totally different thing if down the line you actually want to simulate the entire market and actually map out all the second order impact of the decisions you suggest to your leadership." So, he's been looking around for the solution and his cousin happened to know me and basically told our buyer Shri that the authors of the paper are actually looking to start something. So, that's how we got connected.
在这种特定的合作中,通常的情况是,我们的客户今天非常习惯于与民调公司或样本库公司合作。他们会去问这些公司:“XYZ 是我们想更好了解的人群。我们能针对这些主题进行研究吗?”对于 Simily 来说,初始阶段看起来非常相似。所以,我们的买家来找我们,告诉我们:“我们想更好地了解 XYZ 人群。”然后 Simily 会通过我们与供应商的合作——例如,我们现在与盖洛普(一家民调和样本库公司)建立了战略合作伙伴关系——去接触真实的人类。因此,这些模拟基于真实数据,但我们会联系那些人,收集我们认为关于该个体既高效又具有泛化能力的数据。想象一下,你有 15 分钟时间,你能在那段时间里回答或询问这些人的神奇问题是什么?我们收集这些数据,利用这些数据创建这些人的智能体或模拟,这些模拟基本上可以用来回答远超原始领域的大量问题。我们加载整个平台,它基本上是一个 SaaS 产品。我们的客户可以来询问任何关于他们感兴趣的人群的问题。
In this particular engagement, usually the way this goes is our customers are very much used to working with polling companies or panel companies today. And there they go and basically ask these companies, "XYZ are the populations that we're interested in better understanding. Can we go run a research study of these topics?" That initial stage looks very similar for Simily. So, our buyers come and they tell us, "We want to better understand XYZ population." Then Simily goes out and we have through our partnership with vendors, we have a strategic partnership now with Gallup for instance, who is a polling and panel company, where we go out, work with our vendors to actually reach out to real humans. So, these simulations are grounded in real data, but reach out to those people, collect data that we believe are efficient and generalizable about that person. So, imagine you have 15 minutes, what are the magical questions you can answer or you can ask these people during that time? We collect that data, use that data to create agents or simulations of these people that can basically be used to answer a large number of questions that goes way beyond the original domain. We load that entire platform and it's basically a SaaS product. Our customers come and they can basically ask any questions about the group of people of their interest.
真有趣。这让我想起了自动驾驶汽车,你知道,你从道路上收集一堆数据,然后能够用模拟来增强它。这是一个类似的概念,还是与你在这里所做的有很大不同?
So interesting. It reminds me of in autonomous vehicles, you know, you go and collect a bunch of data from the road and then you're able to augment it with simulation. Is this a similar concept or are there big differences to what you're doing here?
这是一个类似的概念,因为当然,对于自动驾驶汽车,你想创建一个基于真实世界物理的模型。但你想创建一个能够泛化到训练数据之外的模型。它需要在两个不同地点、不同天气条件下都能泛化。非常相似的概念,我们想要做的是接触真实的人,并理解这些人的一些基本特征,这些特征是我们无法编码到模型中的。
It is a similar concept in the sense that of course with the self-driving vehicles, you want to create a model that is based on real-world physics. But you want to create a model that is generalizable beyond your training data. It needs to be generalizable in two different locations with different weather conditions. Very similar concept where what we want to create is we want to reach out to real people and for these people want to understand something fundamental about these people in a way that we cannot code into the model.
我本以为大型语言模型能够很好地代表整个世界,你几乎可以缩小范围。你可以告诉 Claude,你是一个 34 岁的女性,生活在双海岸大都市区。它应该能够有一个忠实的表现。所以我实际上很惊讶你会去找盖洛普。也许你能解释一下为什么你必须要出去收集任何真实世界的数据?
I would have thought that the large language models would be such a good representation of the whole world that you could almost narrow it down. You could tell Claude, you are a 34-year-old woman living in a bi-coastal metropolitan area. And it would be able to have a faithful representation. So I'm actually surprised that you go out to Gallup. Maybe can you just explain why you have to go out and collect any real-world data at all?
是的。这里的一个大问题是关于“说做差距”。人们说的事情和他们实际做的事情之间存在差距。这个差距是真实存在的。而许多大型语言模型是在态度数据上训练的。从根本上说,那是人们在网上说过的话。这确实覆盖了其训练数据的很大一部分。所以 Simily 的模拟平台所做的一件事就是弥合这个差距。因此,我们最终收集的很多数据本质上是行为数据。它还包括像“告诉我你的人生故事”这样的问题所得到的数据。事实证明,如果我们理解了这个人的人生故事,从中得到的数据就是我们所说的关于这个人的长尾信息。它不是关于你在特定时刻做了什么。它不是像“你对政治的看法”这样非常宽泛的问题。而是关于你在哪里长大,你在生活中必须做出的一些艰难决定。这些数据的有趣之处在于,它是建立态度和行为之间转换层的绝佳方式。所以,我们结合了这些数据集,但根本上,这就是我们想要弥合的差距。
Yeah. One of the big questions here is the question around say-do gap. There are things that people say and then there are things that people actually do. And the gap there is real. And a lot of the large language models are trained on attitudinal data. Fundamentally, it is the things that people have said online. That does cover a large quantity of its training data. So one of the things that Simily's simulation platform does is actually closing that gap. So a lot of the data that we end up collecting by nature are behavioral. It also includes data that actually goes into literally questions like just tell me the story of your life. Turns out, if we understand the person's story of your life, the kind of data you get from it is what we consider to be the long tail information about this person. It's not about what you've done in this particular moment. It's not about very broad questions like what's your view on politics. It's about where you grew up, what were some of the difficult decisions you have to make in life. And what's interesting about this data is it's an amazing way to build a translational layer between attitudes and behavior. So, we combine these kind of data sets, but fundamentally, that's the gap that we want to close.
你们有什么样的行为数据?
What sort of behavioral data do you have?
Simily 确实进行了很多实验。所以,我们训练的那种模型,例如,我们有一个庞大的随机对照试验(RCT)库。这些是在社会科学背景下进行的随机对照试验,围绕定价研究展开。所以,我们正在训练的一个模型基本上可以说是人类行为的基础模型。我们拥有来自 RCT 的所有行为信号。我们能否将这些信号编码到模型中,使得最终结果是一个能够预测任何 RCT 结果的模型?同时,我们一直与客户进行的、令我们非常兴奋的对话之一是,客户进来看到这个潜力,他们的想法是,比如 CVS 有 9000 万客户。我们如何利用这类数据来创建更好的模拟?所以,也有关于如何以负责任和道德的方式利用客户内部现有数据,然后使用这些数据来创建 Simily 模型的增强版本的讨论。当然,这将更加微调,针对这些客户的人群。但这就是我们将要利用的那种数据。
So, Simily does run a lot of experiments. So, the kind of models that we have trained, for instance, we have a huge repo of RCTs. So, randomized control trials that were run in social scientific context, that were run around pricing studies. So, one of the models that we are training is basically the foundational model of human behavior in quite a literal sense. We have all the behavioral signal from RCTs. Can we actually encode that into the model so that the end outcome is a model that can basically predict the results of any RCTs? At the same time, one of the conversations that we keep on having with our customers that we're very excited by is our customers then come in, see that potential, and their mind goes to, well, we have 90 million customers, let's say, for a CVS. How can we leverage this kind of data to create better simulations? So, there's also conversation around how can we, in a responsible and ethical way, leverage existing data that is also in-house for our customers, then use that to create an augmented version of Simily's model. So, that, of course, is going to be more fine-tuned, specific to the population of these customers. But that's the kind of data that we would be leveraging.
我明白了。你们通常是通过语音进行这些访谈吗?是填写的调查问卷吗?形式是什么?
I see. And are you doing these interviews typically by voice? Is it a survey that is filled out and what's the modality?
所以范围很广。简单的答案是两者都有。如果你想获得关于人们的长尾信息,访谈非常棒。所以在我们 2024 年进行的原始研究中,我们确实会问问题,比如“告诉我你的人生故事”。现在我们做的方式是训练我们自己的模型。所以这是一个强化学习循环。但基本上想象一下,这里的目标函数是如何用最少的时间获得关于这个人的最大可见度。这就是我们做的事情之一。
So it's a huge breadth. The quick answer here is it's both. Interviews are fantastic if you want to get the long-tail information about people. So we actually do in the original study that I conducted back in 2024, we literally ask questions, tell me the story of your life. Now the way we do it is we are training our own model. So it's a reinforcement learning loop. But basically imagine the objective function here is how can you spend the minimum amount of time to get the maximum amount of visibility about this person. So that is one of the things that we do.
所以基本上是在训练一个访谈者,它不询问事实信息或关于某个特定平台的经历,而是询问人们的生活故事,这些故事可以用来训练我们自己的智能体模型。而对于更事实性或更离散的选择题、调查问卷等,这些也非常高效。它们在时间和数据上都很高效,因为人们可以在短时间内填写很多问题。所以我们确实会利用这些。例如,如果你只是想更广泛地了解人们对某些话题、某些政策等的看法。
So basically training an interviewer that is not really asking for factual information or an experience about a particular platform but just what are the life story that people have that can be used to train our own model for these agents. And then for the more factual or sort of more discrete choices, choice questions, surveys, and so forth. These are also very efficient. These are the time and data efficient because people can fill out many of the questions in short period of time. So for those we actually do leverage them. And for instance, if you want to just have a broader understanding of people's viewpoints on certain topics, certain policies, and things like that.
你把自己描述为一个应用型 AI 实验室。你如何考虑在哪些地方自己构建模型,哪些地方依赖其他现有模型?
You describe yourself as an applied AI lab. How do you think about where you want to build your own models versus where you want to rely on other existing models?
在构建我们自己的模型方面,我们的核心论点是,有一个很棒的模型需要构建,它能够真正编码人们价值观、偏好和品味的多样性,这是单纯的理性模型无法做到的。所以一种表述方式是,我们正在构建的模型,可以想象今天的模型类似于智能单元的 CPU。它是一个在极其理性的数据上训练的单一模型,擅长解决非常复杂的客观问题。而我们的模型更接近于智能单元的 GPU。这里的想法是,我们实际上不需要一个在某些方面超人的模型。事实上,我们想要一个尽可能像人的模型。但我们希望确保这些模型作为个体子单元,能够代表不同子人群的真实观点。所以,当我们看到这个差距时,我们就去开发自己的模型。但同时,我们也利用前沿模型来协调研究。前沿模型在制定研究计划方面非常出色。这就是这些模型被利用的地方。
So in terms of building our own model, we have the core thesis here is there is an amazing model to be built that really encodes the diversity of people's values, preferences, and taste in ways that simply a rational model cannot do. So one way I actually pose this, we're sort of building say imagine the current today's model are akin to the CPU of intelligence unit. It's a single model trained on amazingly rational data that is amazing at solving very complex objective questions. Similarly small model is much more akin to developing something that is close to closer to the GPU of the intelligence unit. Where the idea here is we don't actually need a model that is superhuman at Similarly. In fact, we want model that's as human as possible. But we want to make sure that these models at the sort of individual sub units can represent the real viewpoints of different sub populations. So, where we see that gap, that's when we go develop our own model. But at the same time we do leverage frontier models for instance as a way to coordinate the research. Frontier models are amazing at coming up with a research plan. So, that's where those models actually do get leveraged.
非常有趣。人们通常来找你问的问题是关于新产品发布、如何营销公司、定价,还是所有这些?
Very interesting. Are people typically coming to you with questions around new product launches, how they should be marketing their companies, pricing, all of the above?
所以,以上所有情况都有。不过我们的旅程通常从非常具体的用例和他们试图解决的问题开始。概念测试是一个大项,也非常直接。他们有一个新概念、新产品想法或新市场信息想要测试,想听听用户对 XYZ 的看法。这是他们快速测试这些想法的一种方式。然后他们很快看到的愿景是:现在我们每个月测试 5 到 10 个不同的想法,但如果我们能瞬间测试数千个不同子人群中的数千个不同想法,那会是什么样子?这就是他们最初看到的愿景。然后我们真正进入细节:模拟从这里走向何方?他们很快开始问,这能否用于产品测试,但不仅仅是提交一张图片,而是让这些智能体去体验这个产品 10 分钟,然后告诉我们他们经历了什么、看到了什么。所以你基本上增加了时间维度。然后你进入多智能体模拟之类的事情。我们的一些客户非常常规地要求我们模拟他们的财报电话会议。这实际上是一个起初让我惊讶的用例,但也是一个出奇常见的需求。因为 CEO 和董事会成员总是需要考虑:“嘿,我们该如何设计财报电话会议?观众会如何反应?”所以这也是我们做的事情,而且这非常是一个多智能体模拟。
So, it is all of the above. Our journey usually does however start with a very concrete use cases and problems they are trying to solve. Concept testing is a big one. It's also very straightforward one. So, they have a new concept, new product idea, new market message they want to test. And they want to hear from their users what they would think about XYZ. This is one way for them to quickly test those ideas. And then the promise they quickly see is well, right now we're very much in the practice of testing five to 10 different ideas a month. But what does it look like for us to test instantly thousands of different ideas across thousands of different sub populations? That's the initial vision they see. Then we really get into the nitty-gritty details of well, where does simulation go from here? They then pretty soon start asking well, can this be used to do product testing, but not just simply submitting like an image, but imagine basically asking these agents go experience this product for 10 minutes and tell us about what you experienced, what you saw. So, you're basically adding temporal dimension. Then you go into things like multi-agent simulation. Some of our customers very routinely actually ask us to simulate their earnings call. This is actually a use case that both surprised me at first, but this is also surprisingly a common ask. Because of course the CEOs and board members always need to think about, "Hey, how are we going to design our earnings call? How would the audience react?" So, that is something that we also do and this is very much a multi-agent simulation.
你知道,一旦你有了一个模拟的客户群体,似乎有太多潜在的用例可以测试,对吧?我很好奇在模拟中进行研究和测试的价值,相比之下,比如你有一个新产品概念想要测试,为什么不去跑一千个 Facebook 广告,实际获得点击率呢?难道真实世界的数据不是比模拟数据更有用吗?模拟数据是关于人们可能如何行为的,然后你再用自己的模型进行校正。
You know, it seems like there's so many use cases that could potentially be tested once you have like a simulated almost customer population, right? I'm curious the value of research and testing in sim versus just like, let's say you have a new product concept that you want to test. Why not just go run a thousand Facebook ads and like you actually get the click-through rates on this stuff? Isn't that real-world data almost more useful than the simulated data on how people might behave that you then correct for with your own models?
这是一个很好的问题。我认为在某种程度上,答案首先与规模有关,然后随着发展,真正的新能力来自于你可以模拟交互。规模问题其实很直接:是的,你完全可以运行 Facebook 广告和 Facebook 测试。但你在模拟中可以运行的实验是大规模的实际行为模拟。你可以引入任意数量的用户,甚至不必受 Facebook 上可用人群数量的限制。而且它也更具代表性,因为只有某些群体的人会真正响应在线实验。但同样,我们正在创建的模型,其关键承诺之一就是代表性。我们做了艰苦的工作,实际获取具有代表性的群体,然后收集能够恰当代表他们的数据。所以,规模和代表性是许多用户不容易获得的。这实际上是我们经常听到的常见需求或痛点。许多人的问题不在于问这些人什么问题,而在于首先,我们如何接触到我们感兴趣的人群?这是一个巨大的底线。然后随着发展,你可以真正开始想象——这也是我们一些最具前瞻性的客户现在正在进入的领域——你决策的所有下游影响是什么?不仅仅是关于你是否有这个特定产品,你会喜欢还是不喜欢?你会为此付费还是不付费?我们想要回答的不仅仅是这些初始问题,而是想要理解:假设你是一家汽车公司,你在某个市场推出了一款电动汽车。也许电动汽车做得非常好。所以我们可以帮你围绕电动汽车的营销和产品进行概念测试。但这会对非电动汽车的认知产生什么影响?它会改变市场认知吗?那么对产品线的其他部分意味着什么?你如何以更基于证据的方式平衡这些决策的二阶影响?今天,没有办法测试这个。
So, it's a great question and I think to some extent here the answer has to do with initially scale and then down the line truly the new capability that comes because you can simulate the interactions. The scale question here is actually quite straightforward where yes, you can absolutely run Facebook ads and Facebook testing. But the kind of experiments that you can run in simulation is actual behavior simulation at scale. All right, so you can basically pull in any number of users, doesn't even have to be bounded by the number of population that's available on Facebook. And it's also much more representative because only certain groups of people will actually respond to the online experiments. But similarly, the model that we are creating, one of the key promises is that it is representative. We do the hard work of actually gaining the representative set of people and then collecting the data that they actually represent them properly. So, the scale representativeness is something that many of our users do not have easy access to. This is actually one of the common ask also that we do get or common sort of pain points that we have heard. Where the question that many of these people have isn't about like what questions do we ask these people, but it's about in the first place, how can we get to the population that we're excited to talk to? That's a huge bottom line. Then down the line, you can actually really start to imagine and this is something that our customers and some of the most forward-looking customers are now going into, which is what are all the downstream implication of the decisions you make? It's not just about whether imagine you have this particular product, will do you like it or do you not like it? Would you pay for this, not pay for this? It's not necessarily the just that initial questions that we want to answer and finish, but we want to understand, imagine you're a car company, you launched an electric vehicle in this market. Maybe the electric vehicle does really, really well. So, we can help you do concept testing around marketing and the product around the electric vehicle. But what does that do to the perception of, let's say, non-electric vehicle? Does it change the market perception? Then what does it mean for the rest of the product line? And how do you balance those kind of second order impact of your decision in a way that is more evidence-based? Today, there's no way to test for this.
我想了解你如何看待模型在模拟真实人类行为方面的预测能力。我猜你们在这方面有很多评估。你们的北极星指标是什么?你们做得怎么样?你认为理论极限是什么?
I'd love to understand how you think about how predictive your model is in actually simulating real human behavior. I imagine you have a lot of evals on this. I guess what is your North Star metric? How do you guys do on that? And what do you think is the theoretical limit?
这是个好问题。先从理论极限说起。确实存在极限,因为人类行为有真正的随机性。如果你问我同一个问题,我的回答会略有不同。所以人类行为确实有这种随机性。然而,即使在今天,我们在预测人类方面还有很多性能提升空间。我们的测量是在群体层面进行的。对于更定量的问题,我们测量回答的分布。具体来说,我们使用总变差距离,它衡量真实分布与模拟分布之间的接近程度。这是我们所有客户用例都使用的指标。我们设定了一个阈值,认为足够用于决策。比如,总变差距离小于 0.15,我们认为这是做出决策的有力证据。这就是我们针对这类更定量、更偏向问答的用例想要达到的北极星状态。这也涵盖了随机对照试验,这是客户许多核心用例。现在还有一个有趣的问题:多智能体交互呢?我们模拟的所有下游影响呢?这些评估是什么样的?
It's a great question. So, theoretical limit, let me just start from there. Certainly it does exist in the sense that humans are genuinely random. There's a lot of randomness: if you ask me the same question, I'll actually answer the question slightly differently. So, there is certainly that degree of randomness in human behavior. However, there's a lot of gains in performance that we can have even today in the way we're predicting people. So, the measurement that we do is at the level of population. We measure the distribution of responses if it is more quantitative. So, we actually measure total variation distance, which basically shows how close are the distributions of the ground truth versus the simulated information. And that is a metric that we run across all the use cases that our customers have. And we have certain threshold that we believe is good enough for decision-making. So, TVD of less than 0.15, we believe is actually quite strong evidence for making decisions. So, that is a North Star state that we want to hit for this class of use cases that are more quantitative, that's more question and answers. This also does cover RCTs, which is many of the core use cases our customers have. Now, there's actually a really interesting question to ask around: well, what about multi-agent relation? What about all the downstream implications that we're going to be simulating? What does the evaluation of those look like?
对,那会不会出现级联错误?比如这个智能体有 85% 的准确率,然后它告诉另一个智能体一些信息,随着多智能体交互,错误会累积吗?
Yeah, and then do daisy chain errors as you kind of, if this one is 85% accurate and then this agent is telling another agent something, do you accumulate errors as you go towards multi-agent?
没错。我们的核心论点之一是,模拟大致分为两类:一类是收敛型模拟,另一类是发散型模拟。有时两者并存,关键取决于你的研究问题。对于收敛型问题,少量误差并不重要。当然,误差不能大到完全脱离现实,但即使误差随时间累积,只要收敛的拉力足够强,你仍然能理解最终结果。一个很好的例子是模拟人际网络。这样的网络总会形成一个中心节点,网络科学家称之为无标度网络。这也是谷歌 PageRank 的核心观察:无论网络如何形成,总有一些网页获得指数级更多的链接。这是人类行为中非常基本的特性,我们在模拟网络中也能看到。只要以一定阈值精度复现人类行为,这种收敛就会发生。而另一些问题通常是发散的,比如经典问题“第一次世界大战是否不可避免?”有时很难通过多次模拟得到完全相同的结果。想象一下模拟一场选举:每次都是同一个人获胜吗?每个决策都有很多影响,所以结果会发散。对于这类问题,核心评估围绕置信度展开。比如,运行模拟 100 次,有多少次结果是 X?我们如何利用这一点,通过自助重采样来计算模拟的置信区间?这些都是我们提出的问题。模拟的一大优势还在于展示发散时的多种可能结果,让人们观察、理解导致这些结果的因果机制,并为这些未来做好准备。这就是发散模拟的一些含义。
Exactly. And one of the core thesis here is we basically see two categories of simulations. One simulation is what I would consider to be simulations that converge. The other categories of simulations are the simulations that diverge. And sometimes they actually coexist and it's really about what research question you have. Questions that converge don't actually matter if you have a little bit of error. Now, the error cannot be obviously so dramatic that it is completely detached from reality, but you actually are okay even if the errors do compound over time because the pull towards the convergence is strong enough that you'll actually understand where everything would fall. A good example here actually is if you simulate a network of people. Then that network will always have a hub that gets formed. This is what network scientists would call the scale-free network for instance. This is actually what powered Google, too. The core observation of PageRank was that it doesn't matter how these networks actually get formulated, you actually see some web pages that get exponentially more links attached to them. This is a very fundamental behavior in humans that we also see in simulated networks. And that convergence always happens as long as you're replicating human behavior with certain threshold accuracy. Now, there are other questions that generally do diverge. It's like classical questions like 'Was World War I inevitable or was it not?' And there it is sometimes difficult to run the same simulation over time and get the same exact outcome. Imagine you're running a simulation of an election. Will the same person win the election every time? There are a lot of implications of every single decision that does happen. So, it does diverge. There, the core evaluation is around confidence. So, you imagine you run the simulation 100 times, how many of those times do the results come out to be X? And how can we actually use that to basically create a bootstrap resampling to calculate the confidence around the simulations? Those are some of the questions that we do ask. And a huge part of this also of the power of simulation is then to show when it diverges, to show the diversity of possible outcomes, so that people can actually look, understand the cause or mechanism of how we got to those outcomes, and prepare for those futures. So, those are some of the implications of divergence in simulations.
有没有什么数学描述来解释为什么某些事情会收敛或发散?比如我想,如果有一个平均函数,可能会收敛;如果是二元结果分裂,可能会发散。
Are there any mathematical descriptions of like why something would converge or diverge? Like I'm imagining if you have an average function, maybe you converge, and then if it's like a binary split of outcomes, then you might diverge.
是的,你的直觉很接近。从技术上讲,这也是一个研究课题。Similarly 这家公司正是在深入这个研究课题。我认为模拟领域就像推断统计学的早期阶段。推断统计学的科学家们经过大量研究和探索,才确定 P 值小于 0.05 是科学上足够有力的证据。Similarly 正在为模拟领域设定类似的阈值和标准。所以你的直觉是对的,关于如何建立稳健的数学方程来预测何时发生什么,这确实是模拟领域的研究前沿。
Yeah. So, the intuition, I think, is close. And technically, this is also a research topic. Similarly is a company where we do go deep into this research topic. In the sense that I see simulation as a field as akin to developing the day one of inferential statistics. You know, inferential statistics scientists actually had to do a lot of this question and research over time to decide that P less than 0.05 is actually evidence that is strong enough for science. Similarly, is working on setting the same kind of threshold and standards for the rest of the field. So, those are the intuition. I think that's exactly the right intuition in terms of actually how to make a robust mathematical equation around when's going to like what's going to happen when. It is a real research frontier for simulations.
谢谢你提到这一点。我很好奇,似乎很多财富 500 强公司都在找你们。我想知道是否存在一些尚未开发的企业用例,可能解决社会的重大谜题。比如,我想到经济学、央行决策。我个人认为,宏观领域没人真正懂什么,很多问题源于人类心理。所以对我来说,宏观经济学本质上就是大规模模拟人类行为。甚至风险投资用例中,我们内部也经常争论:价值是否归属于这家公司?你可以模拟 AI 栈的所有不同层次,几乎就能找出持久性和价值在哪里积累。
Thank you for bringing that up. I'm curious, it seems like a lot of Fortune 500 companies are coming to you. I'm wondering whether there are non-existing corporate use cases that might solve great mysteries of our society. For example, I'm wondering about economics, central bank decisions. Oftentimes, I personally believe that in macro, nobody knows nothing. And a lot of the issues come from human psychology. So to me, macroeconomics is a function of simulating human behavior at scale. I'm thinking even in the venture capital use case, we often debate internally: does value accrue to this company or not? You could run a simulation of all the different layers of the AI stack and almost figure out where durability and value accrues.
如果你有一种完美的人类行为模拟器,你能做的远不止服务五家《财富》500 强企业。你同意吗?如果同意,你们在服务政府吗?
If you had a kind of perfect simulator of human behavior, there's so much more you could do than serving the five Fortune 500. Do you agree with that? And if so, are you serving governments?
是的,这很有趣。当我们还在这个领域做研究时,我让我的导师 Michael 和 Percy 感到兴奋的方式是告诉他们:‘看,如果我们做对了,这里能拿诺贝尔奖。’我真心相信这一点。经典的经济学模拟,比如基于智能体的模型,曾经开创了我们对诸如‘隔离是如何发生的?隔离的原因和机制是什么?’等问题的理解,这并不奇怪。像 Thomas Schelling 这样的学者构建了极其简单和基础的基于智能体的模型,但它们揭示了人类宏观行为的深层含义。他后来获得了诺贝尔奖。我在这里看到了同样的机会,但以一种增强的方式。那时,基于智能体的模型非常确定。在 30 年前的隔离模拟中,单个智能体只是一个红点或蓝点。每次迭代,它们会环顾四周,看看有多少邻居是相同颜色,如果阈值低于某个水平,它们就会移动。仅此而已。但现在,我们可以创建真正复制个体全部丰富性的智能体,并运行同样的模拟。我们能提出的问题超越了商业用例。例如,在宏观经济学中,经济学家问我诸如‘银行挤兑何时发生?’或者关于气候变化的问题。解决气候变化的一个核心障碍是许多国家的集体行动问题。我们能模拟这个吗?或者民主即将崩溃的信号是什么?我们能理解货币体系的起源故事吗?这些应该是这个领域的北极星。想象一下这在实践中会是什么样子很有趣。这些将涉及大规模模拟,许多智能体相互交互。我看到了一个未来,今天模拟运行得很快,但如果有一次模拟花费一亿美元,需要好几个月才能运行呢?当我们运行它时,它解决了我们社会的一个基本问题。这确实是这个领域一个令人兴奋的可能性。
Yeah. So, it's interesting. When we were still researching in this area, the way I got my advisors, Michael and Percy, excited about this was I told them, 'Look, we do this right, there's a Nobel Prize to be won there.' And I truly believe that. It's also not surprising that classical economics simulations, like agent-based models, pioneered our understanding of topics like 'How does segregation happen? What are the cause and mechanism for segregation?' Scholars like Thomas Schelling built agent-based models that were extremely simple and rudimentary, but they showed something deep about human macro behaviors. He went on to win a Nobel Prize. I see the same opportunity here, but in an augmented way. Back then, agent-based models were very deterministic. In a simulation of segregation from 30 years ago, an individual agent was simply a red dot or blue dot. Every iteration, they would look around, see how many neighbors are the same color, and if that threshold goes below a certain level, they would move. That was it. But now, we can create real agents that replicate the full richness of individuals and run the same kind of simulations. The questions we can ask go beyond commercial use cases. For instance, in macroeconomics, economists asked me things like, 'When does a bank run happen?' Or questions about climate change. One core blocker to solving climate change is the collective action problem of many nations. Can we simulate that? Or what are the signals of a democracy about to collapse? Can we understand the origin story of the monetary system? These are the kind of simulations that ought to be the North Star of this field. It's interesting to imagine what that would look like in practice. These would involve very large-scale simulations with many agents interacting. I see a future where today a simulation is quick and fast to run, but what about a simulation that costs a hundred million dollars to run once and takes many months? When we run it, it solves one of the fundamental questions of our society. That is a genuinely exciting possibility for this field.
我同意。我甚至认为政治,例如,可能会永远改变。今天每个人都有议程,声称某些政策变化会如何影响事物。那么,我们为什么不直接运行模拟呢?
I agree. I'm even thinking politics, for example, could be forever changed. Today everyone has an agenda of how they say some policy change will impact things. Well, why don't we just run the simulation?
并理解所有下游影响。
And understand all the downstream implications.
是的。
Yeah.
不仅仅是今年会发生什么,而是未来五到十年意味着什么?
And not just what's going to happen this year, but what does it mean in the next five to ten years?
完全正确。太迷人了。我本想最后问你,什么让你对未来感到兴奋?是我们刚才谈到的,还是别的什么?
Exactly. Fascinating. I was going to close by asking you what makes you excited about the future? Is it what we just talked about or is it something else?
我是一个深受科幻小说启发的人。当你阅读那些描绘技术成熟度足够高的社会的科幻小说时,你总会看到两个支柱:某种版本的 AGI 和某种真正帮助指导社会的模拟。我今天看到了真正率先构建模拟的机会。即使在 5 年前,我也不会这么说。但这是我们在深入研究这个领域多年后建立起来的信念。令人兴奋的是,今天有一个清晰的用例可以服务我们的用户。但还有很多尚未到来,我认为这些将逐步构建起一个类似于人类社会 CERN 的模拟器。我的联合创始人 Percy 有时会说:看看最伟大的科学创新,它们往往始于一次惊人的测量。哈勃望远镜真正改变了我们理解宇宙的轨迹。模拟可以成为人类社会的哈勃望远镜。所以让我兴奋的是,虽然有很多关注点放在自然科学上,但模拟如何真正解锁我们对人性和社会科学的理解,以及我们如何实际用它让社会变得更好?这很令人兴奋。
I am somebody who is quite inspired by science fiction. When you read science fiction that covers societies that have progressed far enough in technological maturity, you always see two pillars. You have some version of AGI and you have some version of simulations that really help guide the society. I see an opportunity today to really take the first crack at building the simulation. I would not have said that even 5 years ago. But that is a conviction we have built up over the years as we go deep into this research. What's exciting is there's a clear use case today that can serve our users. But then there's a lot yet to come that I think will build up to actually building a simulator akin to the CERN of human society. One of the things my co-founder Percy sometimes says is: look at the greatest scientific innovations, they often start from an amazing measurement. The Hubble telescope really changed the trajectory of how we understand the universe. Simulation can be that for human society. So what excites me is there's a lot of focus on natural sciences, but how can simulation really unlock our understanding of humanity and social sciences, and how can we actually use it to make our society a better place? That's exciting.
完全同意。我记得读到过有人兴奋地提到,有一个微小但令人惊叹的可能性,即我们所知的经济学领域实际上可能通过模拟得到解决。我会把这个扩展到不仅仅是经济学,而是所有涉及人类行为和社会科学的东西,最终就是我们周围的一切。
Totally. I remember reading somebody was excited about the small but breathtaking chance that the field of economics as we know it may actually become solved by simulation. I'd extend that not just to economics but to everything that deals with human behavior and social sciences, which ultimately is everything around us.
确实如此。
Truly.
是的。太棒了。非常感谢你今天参加节目,分享了 Smallville 的故事以及你现在在 Simile 的工作。我真的很享受这次对话。
Yeah. Wonderful. Thank you so much for joining today and sharing the story of both Smallville and what you're now up to at Simile. I really enjoyed the conversation.
我也是。谢谢你邀请我。
Same here. Thank you for having me.
嗯。
Mhm.