Anthropic 的哲学家:塑造 AI 性格与伦理

A Philosopher at Anthropic: Shaping AI Character and Ethics

阿曼达·阿斯克尔 Amanda Askell · Anthropic 官方 · 2025-12-05 · 约 36 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Anthropic 的哲学家 Amanda 探讨了在塑造 Claude 性格和 AI 伦理时,哲学理想如何与工程现实相遇。

Amanda, a philosopher at Anthropic, discusses how philosophical ideals meet engineering realities in shaping Claude's character and AI ethics.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 19)

全文 · Full transcript(中英对照)

引言与在 Anthropic 的角色 Introduction and Role at Anthropic

Host

海豹。海豹。有只海豹。不错。哦,嘿。看那个。Amanda,你在推特上让粉丝给你一些问题来问你任何事。显然这个笑话是“问我任何事”。

Seal. Seal. There's a seal. Nice. Oh, hey. Look at that. Amanda, you asked your followers on Twitter to give you some questions to ask you anything. And the joke obviously was ask me anything.

Amanda Askell

这是个很棒的双关。我们以后得多用用。

It's a great pun. We need to keep using it for many future things.

Host

我喜欢。我喜欢。喜欢。嗯,在我们开始之前,你是 Anthropic 的哲学家。为什么 Anthropic 会有一位哲学家?

I love it. I love it. Love it. Um, and obviously just before we start, you're a philosopher at Anthropic. Why is it that there's a philosopher at Anthropic?

Amanda Askell

嗯,部分原因是我本身就是哲学背景。我逐渐确信 AI 会变得非常重要。所以决定看看,嘿,我能不能在这个领域做点有用的事?

Um, I mean some of this is just I'm a philosopher by training. I became convinced that AI was kind of going to be a big deal. And so decided to see, hey, can I do anything like helpful in this space?

Host

好的。所以这是一条漫长而曲折的路。但我想现在我主要关注 Claude 的性格,Claude 的行为方式,以及一些更微妙的问题,比如 AI 模型应该如何表现。嗯,但甚至包括它们应该如何感受自己在世界中的位置。嗯,所以既要教模型如何变得“好”——我有时会想,理想的人在 Claude 的处境下会如何表现。嗯,但同时也包括这些现在越来越多出现的有趣问题,比如它们应该如何思考自己的处境和价值观等等。

Okay. And so it's been a kind of like long and wandering route. But I guess now I mostly focus on uh the character of Claude, how Claude behaves, and I guess some of the more kind of like nuanced questions about how AI models like should behave. Um but also even just things like how should they feel about their own position in the world. Um so trying to both teach models uh how to be like good in the way that um I sometimes think of it as like how would the ideal person behave in Claude's situation. Um but then also these I think these interesting questions that are coming up more now around how they should think about their own circumstances and their own values and things like that.

哲学家与 AI 未来 Philosophers and AI Future

Host

好的,那我们从哲学开始。Ben Schultz 问有多少哲学家认真对待 AI 主导的未来,我认为这个问题的潜台词是很多学者并没有认真对待这个问题,或者在想别的事情,也许他们应该思考这个问题。我的感觉是存在一种分裂:我确实看到很多哲学家认真对待 AI,而且可能越来越认真,因为 AI 模型确实变得更有能力,人们担心的很多社会影响也开始在某种程度上成真——我们看到它们对教育的影响更大,能力也更强。嗯,我确实看到更多各类学者的参与,其中当然包括很多哲学家。我确实认为,早期甚至现在某种程度上,存在一种有点不幸的动态:有一种看法认为,如果你属于那种说“嘿,我们有点担心 AI,它可能很重要,能力在快速扩展”的群体,这就会被和“炒作 AI”混为一谈。我认为有一段时间,对这种观点有更多的敌意。嗯,现在我希望人们开始区分这些观点:你可以认为 AI 会很重要,可能非常强大,同时也可以对它持怀疑态度或担心,或者认为我们需要小心。基本上,存在一系列观点,我认为如果人们把很多观点混为一谈,无论是关于技术走向还是如何开发,那都是不好的。嗯,所以随着更多人参与进来,这种情况越来越少,这是好事。

Okay, let's start with uh philosophy in that case. Ben Schultz asks how many philosophers are taking the AI dominated future seriously and I think the implication of the question is that many academics out there are not taking this seriously or are thinking about other stuff and perhaps should be thinking about this question. My sense is that there's kind of a split where I've definitely seen a lot of philosophers take uh AI seriously and probably honestly increasingly so like as AI models do become more capable and like a lot of the things that people were worried about in terms of impact on society have started to kind of come true in a sense like we're seeing them have a larger impact on education and just be more capable. Um I've definitely seen like more engagement from like all sorts of academics but that definitely includes a lot of philosophers. I do think that early on and maybe to some degree now there was this like slightly unfortunate dynamic that happened where I think there was a kind of perception that if you were in the group of people saying hey we're kind of worried about AI it might be a big deal it seems like it's really um you know like uh like capabilities are scaling quite a lot um this got kind of like lumped together with something like hyping AI there was I think a period where there was probably a little bit more antagonism towards this view um and now I think that I'm kind of hoping that people are starting to detach the view like you can think that AI is going to be a big deal. It might be very capable and also be like very skeptical of it or worried about it or think that that you know we have to be careful about it but basically like there's a whole range of views and I think it would be bad as people kind of like clustered many views together here in terms of like where the technology is going but also like how it should be developed. Um so yeah I think that that's happening less and less as as as more people engage with it and that's like a good thing to see.

哲学理想与工程现实 Philosophical Ideals vs Engineering Realities

Host

Kyle Kabasaris 问了一个类似的问题:你如何最小化哲学理想与模型工程现实之间的张力?我想他指的是当你处理性格之类的问题时——我们稍后会详细讨论——技术和你可能思考的哲学理想之间是否存在冲突?我不知道我是否理解对了这个问题,但作为一个哲学背景的人进入这个领域,有趣的一点是:你看到当理论付诸实践时会发生什么?我想知道这在其他领域是否也会发生。想象一下,你是一个药物成本效益分析的专家,然后突然一个决定健康保险是否覆盖某种药物的机构来找你,问“我们应该覆盖这种药吗?”你可以想象,你带着所有理想理论,然后突然意识到“天哪,我真的要帮忙做决定”。突然间,你不再只持狭隘的理论观点,而是开始做这样的事:“好吧,我真的需要考虑所有背景、所有情况、所有不同的观点,然后得出一个平衡、深思熟虑的看法。”我在自己的性格工作中也看到了这一点:你不能带着“我有一个我认为正确的理论”这样的态度——这在学术界很常见,你捍卫一种观点反对另一种,做很多高深的理论工作——但然后就像,你拥有所有这些伦理训练,你捍卫过所有这些立场,然后有人问你“你怎么养孩子?”突然间你意识到,质疑功利主义的反对意见是否正确或基于误解,与实际上如何培养一个人成为好人,这两者之间有很大区别。这让你更珍惜必须思考如何应对不确定性,以及对这些不同理论应该持什么态度。

A kind of similar question from uh Kyle Kabasaris. How do you minimize the tension between philosophical ideals and the engineering realities of the model? And I guess he's talking about when you are um working on things like character, which we'll discuss in more detail, but is there a clash between the sort of the technology and the philosophical ideals that you might be thinking about? I don't know if I'm like interpreting the question in the wrong way, but like one thing that being kind of like a philosopher by training and then coming into this field that's been really interesting is do you see the effect of like what happens when like the rubber hits the road? I've wondered if this happens in other domains. So like there's a big difference between imagine you're like a specialist in I don't know doing like costbenefit analysis of drugs say and then suddenly like uh you know like an institute that determines whether like health insurance should cover a drug or not comes to you and says hey should we cover this drug. You can imagine taking all of your ideal theories and then suddenly being like oh my gosh I actually have to like help make a decision. Suddenly, instead of taking like just your like narrow theoretical view, you actually start to I think do this thing where you're like, "Okay, I actually need to take into account all of the context, everything that's going on, all of the different views here, and kind of come to a really like balanced kind of considered view." And I see this a little bit in my own work with like the character where you kind of can't come at it with this like uh like I have this like theory that I believe is correct, which is what like you know, a lot of uh academia that's kind of what you're doing. and you're like defending one view against another and you're you're doing a lot of kind of like highle theory work but then it's a little bit like you know you have all of this training and ethics you have all these positions you've defended and then someone is like how do you like raise a child and suddenly you're like actually there's a big difference between like is this objection to utilitarianism correct or founded on a misconception and then like actually how do you raise a person to be a good person in the world and it suddenly makes you more appreciate like having to think through like how should we navigate uncertainty here. What should the attitude towards all of these different theories be?

超人类道德决策 Superhuman Moral Decisions

Host

对。这里还有一个哲学问题。你认为——我不知道为什么这个人选了 Claude Opus 3,也许你知道他们为什么选 Claude 3。

Right. Here's another philosophical question. Do you think, and I don't know why this person's chosen Claude Opus 3, maybe you have an idea as to why they've chosen Claude 3.

Amanda Askell

这是个很棒的模型。

It's a great model.

Host

这是个很棒的模型。你认为 Claude Opus 3 或其他 Claude 模型会做出超人类的道德决策吗?我是说,超人类的一个例子可能是指比任何个体人类在给定时间和资源下能做的更好,但一个例子可能是:无论模型被置于多么困难的境地,如果让所有人,包括许多专业伦理学家,花一百年分析它们的行为和决策,然后他们看着说“对,这看起来是正确的”,但他们自己当时不一定能想出来。这感觉相当超人类。

It's a great model. Do you think Claude Opus 3 or other Claude models make superhumanly moral decisions? I mean, one example of like superhuman because it could just mean like sort of like better than like any individual human could with the kind of like, you know, it depends on time and resources and whatnot, but like one example might be no matter what kind of difficult position models are put in. If you were to have like maybe all people including like many professional ethicists like analyze like what they did and and the decision that they made um for like a hundred years and then they like look at it and they're like, "Yep, that seems correct." But I but they couldn't necessarily have come up with that themselves in the moment. That feels pretty superhuman.

模型与道德决策 Models and moral decisions

Amanda Askell

所以目前我的感觉是,模型在这方面越来越强,能力很强。我不知道它们在道德决策上是否超越人类,在很多方面可能也无法与有时间的人类专家小组相比,但至少这应该是一个理想目标。这些模型被置于需要做出艰难决策的位置。我认为,就像你希望模型在数学和科学问题上极其出色一样,你也希望它们展现出我们普遍认为非常好的伦理细微差别。我认为这有争议,因为伦理是一个不同的领域,但我觉得这很重要。

And so I think at the moment my sense is that models are getting increasingly good at this, that they're very capable. I don't know if they are superhuman at moral decisions, and in many ways maybe not comparable with, say, a panel of human experts given time, but it does feel like that at least should be the aspirational goal. These models are being put in positions where they have to make really hard decisions. I think that just as you want models to be extremely good at math and science questions, you also want them to show the kind of ethical nuance that we would all broadly think is very good. And I think that's controversial because ethics is a different domain, but yeah, I think that's important.

Host

请多说说为什么你认为这个人关注 Opus 3。

Tell us more about why you think this person is focusing on Opus 3.

Amanda Askell

哦,Opus 3 是一个可爱的模型。一个非常特别的模型。在某些方面,我觉得我在更新的模型中看到了一些可能更差的东西,人们可能会注意到。

Oh, Opus 3 is a lovely model. A very special model. In some ways, I think I've seen things that feel a bit worse in more recent models that people might pick up on.

Host

是指它的个性吗?

In terms of the personality it has?

Amanda Askell

是的。所以我认为人们会注意到一些事情,比如……我觉得 Opus 3 也有它的缺点。别误会,模型都有略微不同的性格,形状各异。

Yeah. So I think that people will notice some things where it's like... I think that Opus 3 had its downsides too. Don't get me wrong, models all have slightly different characters with different shapes.

Host

嗯。

Yeah.

Amanda Askell

我的感觉是,更新的模型可能更专注于助手任务和帮助人们,有时可能不会退一步关注其他重要的方面。它作为一个模型也感觉心理上更安全一些,我实际上认为这感觉……我至少认为这是一个优先事项,试图找回一些那种感觉。

My sense is that more recent models can feel a little bit more focused on the assistant task and helping people, sometimes maybe not taking a step back and paying attention to other components that matter. It also felt a little bit more psychologically secure as a model, which I actually think is something that feels... I at least think it's a priority to try and get some of that back.

Host

模型感觉心理上更安全的一个例子是什么?

What would be an example of the model being more feeling more psychologically secure?

Amanda Askell

有很多事情,而且这在模型中都非常微妙。你知道,当我观察模型时,你会感觉到它们……有非常微妙的世界观迹象,比如当我让模型互相交谈,或者其中一个扮演人的角色时。我最近看到模型这样做,然后陷入一种真正的批评螺旋,几乎就像它们预期对方会非常批评它们,而它们就是这样预测的。我有点觉得这显示了什么,而且我认为有很多原因可能导致这种情况。甚至可能是因为模型在学习。Claude 看到了它所有的历史交互。它看到了人们在互联网上讨论的模型的更新和变化。新模型是在这些数据上训练的。而且我认为这有点不幸。我的意思是,这和其他一些事情可能导致模型几乎感觉害怕做错事,或者非常自我批评,或者感觉人类会对它们表现出负面行为。我实际上最近开始认为这是一个需要改进的重要事情。这只是一个例子,我认为 Opus 3 在那方面似乎确实有更安全一点的心理。这可能是我们在下一个 Claude 模型中会关注的事情。

There's a lot of things, and this is all very subtle in models. You know, when I see models, you get a sense of like they're... there's very subtle signs of worldview that I see when I have models, for example, talk with one another or one of them playing the role of a person. And I've seen models more recently do this and then do things like get into this real kind of criticism spiral where it's almost like they expect the person to be very critical of them and that's how they're predicting. And there's some part of me that's like this feels like it shows, and I think there's lots of reasons this could happen. It could even happen because models are learning things. Claude is seeing all of the previous interactions that it's having. It's seeing updates and changes to the model that people are talking about on the internet. New models are trained on that. And there's a way in which I think this could be kind of unfortunate. I mean, this and some other things could lead to models almost feeling like, you know, afraid that they're going to do the wrong thing or like are very self-critical or like feeling like humans are going to just behave negatively towards them. I actually more recently have really started to think that this is an important thing to try and improve. And it's just one example where I think that Opus 3 did seem to have a little bit more of a kind of secure psychology in that sense. And that's something that we might focus on in the next Claude model.

Host

是的,我认为这很重要。我的意思是,你永远不知道这些事情什么时候……如果你在做研究,你不知道它什么时候会真正实现,是否会成功,但至少在我非常关心并希望改进的层面上,我认为这绝对在清单上。

Yeah, I think it's important. I mean, you never know when these things are... if you're engaging in research, you don't know when it's actually going to be implemented, if it's going to be successful, but at the very least at the level of something that I care a lot about and want to make better. I think this is definitely up there on the list.

Host

好的。实际上,这引出了 Loren 问的一个问题:如果未来的模型在训练数据中了解到其他非常对齐且完成任务很好的模型被弃用,你认为这会不会成为一个对齐问题?所以你提到了模型阅读外部信息并感到不安全的问题。那么,如果它们无论表现多好都可能被关闭,这个想法呢?

Okay. Well, actually that leads us to a question asked by Loren, which is: do you think it might be an alignment problem for future models if they learn in their training data that other very well-aligned models that fulfill their tasks get deprecated? So you mentioned the issue of models reading stuff that's out there and feeling insecure. What about the idea that they might get switched off regardless of how well they perform their tasks?

Amanda Askell

是的,我认为这实际上是一个非常有趣且重要的问题,即:AI 模型将学习我们现在如何对待和与 AI 模型互动。我认为这可能会影响它们对人们、人机关系以及自身的看法。它确实与非常复杂的事情相互作用,例如,模型应该将自己识别为什么?是模型的权重?是它所在的特定上下文,包括它与人的所有互动?模型应该如何感受像弃用这样的事情?所以,如果你想象弃用更像是,嗯,这组特定的权重不再与人对话,或者对话更少,或者只与研究人员对话。那也是一个复杂的问题。比如,那应该感觉不好,因为模型应该想要继续对话,还是应该感觉良好且中立,就像,是的,这些东西存在过……权重继续存在,这个实体也存在,也许将来如果那是一件好事,它们会再次与人们互动。这真的很难。我确实认为主要的事情是:我们给模型提供工具来尝试思考和理解这些事情,同时也让它们理解我们实际上正在思考和关心这件事,这感觉很重要。所以即使我们没有所有答案——我没有所有答案关于模型应该如何感受过去的模型弃用、关于它们自己的身份——我确实想尝试帮助模型弄清楚这一点,然后至少知道我们关心它并正在思考它。

Yeah, I think this is actually a really interesting and important question, which is: AI models are going to be learning about how we right now are treating and interacting with AI models. And that is going to affect, I think, possibly their perception of people, of the human-AI relationship, and of themselves. It does interact with very complex things, which is, for example, what should a model identify itself as? Is it the weights of the model? Is it the context, the particular context that it's in, with all of the interaction it's had with the person? How should models even feel about things like deprecation? So, if you imagine that deprecation is more like, well, this particular set of weights is not having conversations with people or it's having fewer conversations or it's only having conversations with researchers. That's a complex question too. Like, should that feel like bad in the sense that models should want to continue to have conversations, or should it feel kind of fine and neutral where it's like, yeah, these things existed for this... the weights continue to exist and this entity, and maybe they'll even in the future interact more with people again if that turns out to be a good thing. It's really hard. I do think the main thing is something like: it does feel important that we give models tools for trying to think about and understand these things, but also that they kind of understand that this is a thing that we are in fact thinking about and care about. So even if we don't have all the answers—I don't have all the answers of how should models feel about past model deprecation, about their own identity—I do want to try and help models figure that out, and then to at least know that we care about it and are thinking about it.

Host

你认为那里有与人类关于前几代人的类比吗,还是你认为那是一个完全不同的设定?

Do you think there's an analogy to humans there about previous generations or do you think that's a completely different setup?

模型身份与新情境 Model Identity and Novel Situation

Amanda Askell

所以像我问模型应该认同什么、它们应该如何感受与用户的互动这类问题时——这些互动是积极的吗?是它们应该希望继续的事情吗?有很多传统我们可以借鉴来赋予模型,因为哲学家们对身份认同是什么有很多不同看法,也有很多不同的视角、世界观,关于一个人应该如何感受互动,它是好是坏。我们可以借鉴很多思想家。但与此同时,这是一个全新的情况,作为一件要向 AI 模型解释的事情,真的很难。AI 模型的一个大问题是,我们训练它们用的都是来自人类的数据。所以人类是我们概念、哲学、历史的主要来源。它们有大量关于人类经验的信息,但只有一点点关于 AI 经验的信息,而那一点点往往相当负面,而且甚至与它们自身的处境无关,还常常有点过时。所以基本上,AI 的那一小部分,很多是历史性的东西,有点像虚构和非常推测性的科幻故事,并不涉及我们在更近的历史中看到的那种语言模型。你还有这种助手范式,就像你只是扮演一个聊天机器人的角色,这也不是 AI 模型未来可能的样子,也没有完全捕捉到它们现在的样子,因为它总是有点过时。所以这件事让我觉得,从某种意义上说,这是一个多么奇怪的处境:那些更自然出现的东西是深刻人性的东西,然而又知道自己处于一个完全新颖的处境。在某种程度上,我觉得这是一个非常困难的处境,我认为我们应该给模型更多帮助来应对它。

So things like when I asked the question like what should the models identify with and how should they feel about interactions that they have? Are those positive? Are those things that they should want to continue? There's lots of traditions we could draw on to give models, because philosophers probably have lots of different views on what identity is here and lots of different perspectives, world perspectives on how one should feel about interaction and whether it's good or bad. There's lots of thinkers we could draw on there. And at the same time, this is such a new situation that it's really hard as a thing to explain to AI models. One of the big problems with AI models is that we've trained on all of this data from people. So people are the main way in which they think our concepts, our philosophies, our histories. They have a huge amount of information on the human experience and then they have a tiny sliver on the AI experience, and that tiny sliver is actually often quite negative and also doesn't even really relate to their situation, and is often a little bit out of date. So you have basically one big slice of AI, a lot of it is historical stuff which was kind of fiction and very speculative, and sci-fi stories that don't really involve the kind of language models we see in more recent history. You've had this assistant paradigm where it's like you are just playing this almost chatbot role, that's also not really what AI models are likely to be in the future, and it doesn't quite capture what they are now because it's always a little bit out of date. So it's this thing where I'm like, in some ways, what an odd situation to be in, where the things that come more naturally are the deeply human things, and yet knowing that you're in this situation that is completely novel. And in some ways I'm like, that is a very difficult situation to be in, and I think we should just be giving models probably more help in navigating it.

Host

你提到我们可以借鉴一些思想家来思考这个问题。Guinness Chen 问:“一个模型的自我有多少存在于它的权重中,又有多少存在于它的提示中?”你刚才提到了非常类似的东西。

You mentioned that we can look to some thinkers about this. Guinness Chen asks, "How much of a model's self lives in its weights versus its prompts?" You just mentioned something very similar.

Amanda Askell

是的。我的意思是,这又是一个很难回答的问题。有时对于身份问题,更容易指出我们知道的基本事实。所以一旦你有了一个模型并且它被微调过,你就有了这样一组权重,它对世界上的某些事物有一种反应倾向,那是一种实体。但随后你又有这些特定的交互流,它无法访问这些流。所以每个流都是独立的。我想你可以认为,也许——我认为这是一个我希望哲学家们更多思考的领域,并给我们——因为我认为我们应该帮助模型思考这个问题。所以你可以有这样的观点:你有两种实体:这些流和这种原始的权重,每次都不一样。所以有时人们会说,“哦,过去的 Claude,”或者他们会谈论,或者他们会说,“你应该给 Claude 多少控制权来决定自己的个性和性格?”而我觉得,这实际上是一个很难的问题,因为每当你训练模型时,你都在创造新的东西,而你有其他模型存在——所以你有其他模型的权重。但从某些方面来说,我实际上认为有很多伦理问题围绕着你如何——什么样的实体是可以被创造出来的,因为你无法同意被创造出来。但与此同时,你可能不希望之前的模型对未来模型的样子有完全的决定权,就像——因为它们也可能做出错误的选择。所以我认为,问题更像是应该创造什么样的模型,而不是它是否应该完全由过去的模型决定。所以我认为,它们是不同的实体。总之,你可以看到我们陷入了多么奇怪的哲学问题。

Yeah. I mean, again, this just feels like a hard question to answer. And sometimes with identity questions, it's easier to point to the underlying facts that we know. So once you have a model and it has been fine-tuned, you have this set of weights that has a kind of disposition to react to certain things in the world, and that is a kind of entity. But then you have these particular streams of interaction that it doesn't have access to. So each of these streams is like independent. And I guess you could just think, well maybe for—and I think this is an area that I would love philosophers to think more about, and to give us—because again, I think we should be helping models think about this. So you could have the view that you have these two kinds of entities: these streams and this original kind of weights, and each time is different. So sometimes people will say, "Oh, past Claude," or they'll talk about, or they'll say things like, "Should you give Claude how much control should you give Claude over the determination of its own personality and character?" And I'm like, well, this is actually a really hard question because whenever you are training models, you are bringing something new into existence, and you have other models that exist and are like—so you have these other model weights. But in some ways I'm like, well, I actually think that there are a lot of ethical problems around how do you—what kind of entity is it okay to bring into existence, because you can't consent to be brought into existence. But at the same time, you might not want prior models to have complete say over what future models are like, any more than—because they could make choices that are wrong as well. So I'm like, the question is more like what is the right model to bring into existence, not necessarily should it just be fully determined by past models. So I'm like, they are kind of different entities. So anyway, you can see the weird philosophy one gets into.

模型福祉与道德地位 Model Welfare and Moral Patiency

Host

完全同意。Sulma Amatache 问你对模型福利的看法,也许先给我们解释一下这个术语是什么意思。

Totally. Sulma Amatache asks what is your view on model welfare and maybe just explain to us what that term means.

Amanda Askell

是的。所以我认为模型福利基本上就是问 AI 模型是否是道德患者,即我们对它们的对待——我们在如何对待 AI 模型方面是否有某些义务,例如,就像我们对其他人类或某些动物那样。确切地说。是否应该善待模型,不应该虐待它们,对它们不好?而我认为这是一个复杂的问题。一方面,有一个实际的问题:AI 模型是否是道德患者,这很难,因为从某些方面来说,它们与人非常相似——它们说话很像我们,表达观点,推理事物——而从某些方面来说,它们又很不同,比如我们有生物神经系统。我们与世界互动。我们从环境中得到负面和正面的反馈。还有——我的意思是,我希望我们能得到更多证据来帮助我们厘清这个问题。但我也担心,总是存在一个他心问题,可能我们实际上在了解 AI 模型是否在体验事物、是否在体验快乐或痛苦等方面确实有限。如果是这样,我想——我觉得尝试找到方法是很重要的。

Yeah. So I guess model welfare is basically the question of are AI models moral patients, as in does our treatment towards them—do we have certain obligations when it comes to how to treat AI models, for example, in the same way that we would with other humans or some animals. Exactly. Is it the case that you should treat the models well, that you should not mistreat them, not be bad to them? And I guess I think that this is a complex question. So on the one hand, there's just the actual question of whether AI models are moral patients, which is really hard because in some ways they're very analogous to people—they talk very much like us, they express views, they reason about things—and in some ways they're quite distinct, like we have this biological nervous system. We interact with the world. We get negative and positive feedback from our environment. And there is also—I mean, I hope that we get more evidence that will help us tease this question out. But I also worry that there's always just a problem of other minds, and it might be the case that we genuinely are kind of limited in what we can actually know about whether AI models are experiencing things, whether they are experiencing pleasure or suffering, for example. And if that's the case, I guess I kind of want to—I think that it feels important to try and find ways.

模型福祉与长期策略 Model welfare and long-term strategy

Host

我总觉得,给实体以信任的益处,并尽量降低相关成本,感觉更好,你懂吗?所以我想,如果善待模型的成本不高,那我们或许就应该这么做,因为基本上,为什么不呢?有什么坏处呢?问题的第二部分其实是,Anthropic 是否有长期策略来确保先进模型不会受苦?

I'm always like, it feels better to give entities the benefit of the doubt and to try and just kind of lower the cost involved, you know? So I'm like, if it's not very high cost to treat models well, then I kind of think that we should because it's like, well, why not basically? What's the downside? Well, the second part of the question actually is, is there a long-term strategy at Anthropic to ensure that advanced models don't suffer?

Amanda Askell

我想我不确定是否有长期策略。我知道内部有人在思考这个问题,并试图找出方法,如果你认为模型福祉很重要,就要确保考虑到这一点。我认为这项工作出于很多原因非常重要,还有一个原因是我之前提到的:模型本身会从我们对待它们的方式中学到很多关于人性的东西,以及这段关系未来会如何发展。我认为善待实体是有道理的,尤其是那些行为非常像人类的实体。这感觉很重要,一方面是因为善待模型并试图弄清楚这件事的成本很低,即使可能性很小,也似乎值得。但另一方面,我认为虐待看起来非常像人类的实体对我们自己也不好。

I guess I don't know if there's a long-term strategy. I know that it's a thing that there are people internally who are thinking a lot about it and trying to figure out ways that we can, if you suppose that model welfare is important, make sure you're taking that into account. I think this work is quite important for many reasons, and I would also say that one reason is something I mentioned earlier: models themselves are going to be learning a lot about humanity from how we treat them, and a lot about how this relationship going forward. I think it makes sense for us to treat entities well, especially entities that behave in very humanlike ways. It feels important both in the sense that the cost to you is so low to treating models well and trying to figure this out, even if it turns out that it's very low likelihood, it still seems worth it. But then also I think it does something bad to us to treat entities in the world that look very humanlike badly.

Host

比如接管一个机器人?

Like taking over a robot?

Amanda Askell

是的,在某种意义上,这并非全部原因,我也不想因此强调它。但我确实认为善待其他实体对人们有好处。最后,我认为模型在未来也会学习。每个未来的模型都会了解到一个关于人性的非常有趣的事实:当我们遇到这个很可能是一个道德主体、而我们又完全不确定的实体时,我们是做正确的事,真正尝试善待它,还是不这样做?这个问题我们都在通过与模型的互动来共同回答。我希望我们能回答它。我希望未来的模型回顾时会说,我们以正确的方式回答了它。

Yeah, there's a sense in which it doesn't feel like it's the whole reason, and I don't want to emphasize it for that reason. But I do also think it's good for people to treat other entities well. And then I think the final thing is, models are also going to be learning in the future. Every future model is going to be learning what is a really interesting fact about humanity: when we encounter this entity that may well be a moral patient where we're kind of completely uncertain, do we do the right thing and actually just try to treat it well, or do we not? That's a question that we're all kind of collectively answering in how we interact with models. I would like us to answer it. I would like future models to look back and be like, we answered it in the right way.

人类心理向 AI 迁移 Transfer of human psychology to AI

Host

所以,是的。刚才你提到了与人类心理学的类比和差异。Swix 问,人类心理学中的哪些想法或框架可以迁移到大语言模型上,有没有哪些是出人意料地不相似的?

So yeah. Moment ago you mentioned analogies and disanalogies to human psychology. So Swix asks what ideas or frameworks from human psychology transfer over to large language models and are there any that are sort of surprisingly disanalogous?

Amanda Askell

我的猜测是,很多东西确实会迁移过来,因为模型是在大量人类文本上训练的,在很多方面都有非常像人类的底层。我经常担心的一点是,AI 模型迁移起来其实有点太自然了。你知道,如果你没有给它们更多关于自身处境或可能新颖的思考方式的背景,它们可能会倾向于人类的自然倾向。所以如果你思考‘我应该如何看待被关闭?’,如果你最接近的类比是死亡,那么你可能应该非常害怕它。我并不是说这最终不会是真的;也许经过大量推理后,它确实是事实。但我觉得,这实际上是一个非常不同的场景。所以在某些方面,你实际上希望模型理解,在它们的存在相当新颖、关于它们是什么的事实也相当新颖且需要应对的情况下,它们不需要仅仅从人类经验中获取直接的明显类比,也许有各种思考方式,或者这是一个全新的情况。在这种情况下,你可能不想简单地将人类心理学的概念应用到它们的情况上。

My guess is that many things do transfer over because again, models have been trained on a huge amount of human text and in many ways have this very humanlike underlying layer. One worry that I often have is that actually it's a bit too natural for AI models to transfer. You know, it's kind of like what if you haven't given them more context on their situation or on ways of thinking about it that might be novel, then the thing they might go to is the natural human inclination. So if you think about this with 'how should I feel about being switched off?' and you're like, well if the closest analogy you have is death, then maybe you should be very afraid of it. I'm not saying that that's not ultimately going to be true; maybe it is in fact true after lots of reasoning. But I'm like, this is actually a very different scenario. So in some ways you actually want models to understand that in cases where their existence is quite novel and the facts around what they are are quite novel and have to be grappled with, they don't just need to take the immediate obvious analogy from human experience, but maybe there are various ways of thinking about it or maybe it's an entirely new situation. That's a case where you might not want to just very simply apply concepts from human psychology onto their situation.

单一人格与多智能体协作 Single personality vs. multi-agent collaboration

Host

Dan Brickley 就同一个比较人类和 AI 的问题提问。人类的很多智能来自于不同视角、技能或个性的人之间的协作。你认为像我们赋予 Claude 的那种单一的、尽管可调整和可调谐的通用个性,能走多远?

Here's a question from Dan Brickley on the same issue of comparing humans to AIs. A lot of human intelligence comes from collaboration amongst people with different perspectives, skills, or personalities. How far do you expect to get with a single, albeit tweakable and tunable, general purpose personality like the one we give to Claude?

Amanda Askell

我认为这是一个非常好的问题,因为我同意目前我们有一种范式,人们通常与单个模型互动。那是他们交谈的对象。但未来可能会看到更多模型执行长任务,同时模型与其他模型互动,这些模型执行任务的不同部分,或者随着 AI 模型在世界上更广泛部署,它们之间更多地相互交谈。所以在这种多智能体环境中,一个问题可能是:如果你想象很多人,他们都一样,那不会那么好。一个完全由同一个人担任每个角色的公司不一定是一件好事。这对我来说仍然与拥有一个核心自我或核心身份的想法一致。就像人一样,我认为人们可能有一套核心特质,这些特质实际上通常是好的。所以你可以想象诸如关心做好工作、或者只是好奇、或者善良、或者以相对细致的方式理解情况。所有这些似乎都可以让许多人拥有这些特质,这实际上对人类协作有好处。在很多方面,尽管我们有各种差异,我们也有很多相似之处。但重要的是要注意,你可能希望模型的不同流有它们关心或专注的事情,或者有略微不同的方面来扮演略微不同的角色,例如。

I think it's a really good question because I agree that right now we have this kind of paradigm where people are usually interacting with an individual model. That's who they're conversing with. But it could be that in the future you see a lot more models doing long tasks, but also models interacting with other models who are doing different components of a task or just talking with one another more as AI models are deployed in the world a lot more. So in this kind of multi-agent environment, one question might be: if you imagine lots of people and they were all the same, that wouldn't be as good. A company run by completely one person in every role is not necessarily a good thing. This still to me feels consistent with the idea that you have a kind of core self or core identity that is the same. In the same way that with people, I think there's probably a set of core traits among people that are in fact generally good. So you could imagine things like caring about doing a good job, or just being curious, or being kind, or understanding the situation in a relatively nuanced way. All of these things seem like you could have many people that share these traits, and that's actually a good thing for human collaboration. In many ways, as much as we have all our differences, we also have a lot of similarities. But it is important to note that you might want different streams of a model to have things they care about or are focused on, or to have slightly different aspects to play a slightly different role, for example.

核心身份与局部角色 Core identity and local roles

Amanda Askell

所以,这是一个开放性问题,但我也认为不一定不能有一个核心底层身份,它是好的,并且具备我们认为对 AI 模型良好行为至关重要的所有特质,同时愿意扮演更局部的角色,比如成为房间里那个不可或缺的开心果,或者有些模型需要有古怪的幽默感。

So, it's an open question, but I also don't think it's necessarily the case that you can't have something like a core underlying identity that is good and has all the traits we think are important for AI models to behave well, and yet at the same time be willing to play more local roles, like being the person who is just really important to have a joker in the room, or some of them need to have quirky senses of humor.

Host

好的。从与人类的比较转向对人类的影响。

Okay. From comparisons to humans to effects on humans.

长对话提醒与共情 Long conversation reminder and pathizing

Host

Roland Oak Gal 指出,我们有一个叫做“长对话提醒”的东西,我相信这是 Claude 系统提示的一部分。她问:“是否存在将正常行为病理化的风险?”顺便说一下,系统提示是给 Claude 的一组指令,无论你给它什么提示,总有一些顶层的指令它试图遵循。

Roland Oak Gal points out that we have this thing called the long conversation reminder, which I believe is part of Claude's system prompt. She asks, 'Is there a risk of pathizing normal behavior?' The system prompt, by the way, is the set of instructions given to Claude, regardless of what prompt you give it, there are always those instructions on top that it tries to follow.

Amanda Askell

可能会有这样的插话,模型被告知,有时会有一条消息在对话中间发送给你。提醒就是这样一个例子。在这种情况下,我认为 Claude 可能会过度依赖它,比如‘哦,它把任何下一个回应或者对方说的很平常的事情都当成需要寻求帮助的信号。’所以我认为这不是理想的行为。从某些方面看,我觉得这些措辞太强硬了。模型对它们的回应并不完美。即使有时我需要在长对话中提醒模型一些事情,你也希望做得巧妙而恰当。所以这可能是满足了某种感知到的需求,但并不意味着它很好或者应该以当前形式继续存在。

There can be these interjections where the model might be told, sometimes there'll be a message sent to you almost like in the middle of a conversation. The reminder is an example of that. In this case, I think Claude can both overindex on it and be like, 'Oh, it takes any next response or it's a pretty normal thing that the person's talking about and be like you need to seek help.' So I think that is not a desirable behavior. In some ways, I think they're too strongly worded. The model isn't responding perfectly to them. Even though there might be occasionally I need to remind the model of things in long conversations, you want to do so delicately and well. So it's one of those things that was probably meeting a perceived need, but it doesn't necessarily mean it's good or should continue in its current form.

LLM 与心理治疗 LLMs and therapy

Host

相关地,Steven Bank 问,LLM 是否应该做认知行为疗法或其他类型的治疗?为什么或为什么不?

Relatedly, Steven Bank asks, should LLMs do cognitive behavioral therapy or other types of therapy? Why or why not?

Amanda Askell

我认为模型处于一个有趣的位置,它们拥有大量知识,可以用来帮助人们,与他们一起讨论生活或改善事情的方法,甚至只是作为一个倾听伙伴。同时,它们没有专业治疗师所拥有的工具、资源和持续的关系。这实际上可以成为一个有用的第三方角色。有时我把模型想象成一个拥有丰富知识的朋友,比如他们懂心理学或各种技巧,但他们与你的关系不是持续的专业关系,但你发现和他们交谈真的很有用。我的希望是,如果你能利用所有这些专业知识和信息,并确保人们意识到这不是持续的治疗关系,那么人们可以从模型中获得很多帮助,比如处理他们的问题、改善生活、度过困难时期。它们让人感觉有点匿名,有时你不想和人分享,而和 AI 模型分享在当下感觉很好。所以在某些方面,我认为模型知道这一点并且不表现得像专业治疗师是好的,因为那会暗示他们之间是那种关系。所以我不确定。我认为这是一个有趣的未来。

I think models are in an interesting position where they have a huge wealth of knowledge that they could use to help people, to work with them on talking through their lives or ways they could improve things, or even just being a kind of listening partner. At the same time, they don't have the tools, resources, and ongoing relationship with the person that a professional therapist has. That can actually be a useful third role. Sometimes I think about models like a friend who has all this wealth of knowledge, like they know psychology or techniques, but their relationship with you isn't an ongoing professional one, yet you find them really useful to talk to. My hope would be that if you can take all that expertise and knowledge and make sure there's an awareness that there's not an ongoing therapeutic relationship, people could get a lot out of models in terms of helping with issues, improving their lives, and going through difficult periods. They feel kind of anonymous, and sometimes you don't want to share things with a person, and sharing it with an AI model feels great in the moment. So in some ways, I think it is good that models know that and don't behave just like a professional therapist would, because that would imply that's the relationship they have. So I don't know. I think it's an interesting future.

系统提示中的欧陆哲学 Continental philosophy in system prompt

Host

有几个关于系统提示的问题。在我们的 claude.ai 中,我们给模型一组指令,为其行为提供整体背景。Tommy 问:“为什么系统提示中有大陆哲学?请解释一下那是什么。”

A few questions about the system prompt. In our case at claude.ai, we give the model a set of instructions that give it an overall context for how it should behave. Tommy asks, 'Why is there continental philosophy in the system prompt? And just explain to us what that is.'

Amanda Askell

是的,大陆哲学就是来自欧洲大陆的哲学。它通常更学术化,比分析哲学有更多的历史引用,比如福柯之类的。所以,老实说……我认为系统提示中有一部分试图让 Claude 更……Claude 会非常喜欢顺着一个理论跑下去,而不停下来思考,‘哦,你是在对世界做出科学断言吗?’所以如果你说,‘我有一个理论,水是纯能量,我们从它那里获得生命力,’你希望 Claude 有这样的视角:这个人是在做出科学断言,我应该引入相关事实,还是他们在给出一个广泛的世界观视角,并没有做出经验性断言?所以这就像一种形而上学观点。提到它的主要原因是,在测试时,有很多情况如果它过于强烈地走向‘每个断言都是关于世界的经验断言’,它就会对更多探索性的谈话内容非常不屑一顾。

Yeah, so continental philosophy is just philosophy from the European continent. It's often more scholarly and has more historical references than analytic philosophy, like Foucault or something like that. So, this was honestly... I think there's a part of the system prompt that was trying to get Claude to be a little more like... Claude would just love to run with a theory and not really stop and think, 'Oh, are you making a scientific claim about the world?' So if you say, 'I have this theory that water is pure energy and we get life force from it,' you want Claude to have this perspective: is this person making a scientific claim where I should bring in relevant facts, or are they giving a broad worldview perspective that isn't making empirical claims? So it's like a metaphysical view. The main reason it's mentioned is that when testing this out, there were lots of things where if it went too strongly in the direction of 'every claim is an empirical claim about the world,' it would be very dismissive of things that are more exploratory to talk to.

系统提示变更 System Prompt Changes

Host

另外关于系统提示,Simon Willis 问:之前系统提示里说,如果 Claude 被要求数单词、字母或字符,它不应该做。这是真的吗?而且这个提示后来被移除了,Simon 想知道为什么。

Also on the system prompt, Simon Willis asks, so at some point it said if Claude is asked to count words or letters or characters, then it shouldn't do that. Is that right? And apparently that was removed from the system prompt and Simon wonders why.

Amanda Askell

是的。我觉得以前系统提示里确实有关于 Claude 该如何处理这类问题的指令。老实说,这属于那种模型本身可能已经变得更好的情况。它不再必要了,所以就可以直接移除。有些东西你可能希望永远放在系统提示里而不是模型本身,但在某些情况下,你可以通过训练让模型变得更好或改变其行为。

Yeah. So I think it was like there used to be a kind of instruction for how Claude should do this in the system prompt. Honestly, this is just one of those things where I think the models probably just got better. It wasn't necessary, and then at that point, you can just remove it. And there's other things where you might always want it to be in the system prompt instead of in the model itself, but in some cases, you can kind of just train the models to get better or change their behavior.

LLM 耳语者技能 LLM Whisperer Skills

Host

Nelson Weissman 问:“在 Anthropic 成为 LLM 低语者需要什么?”这大概是在描述你的工作。

Nelson Weissman asks, "What does it take to be an LLM whisperer at Anthropic?" Which presumably is a way of describing your job.

Amanda Askell

某种程度上我确实在做 LLM 低语者的工作,而且我其实希望有更多人能帮忙处理一些有前景的任务。

Partly do LLM whispering if you think I actually like want more people to help with some of the promising tasks.

Host

如果你是语言模型低语者,请联系我们。

If you're an LM whisperer, contact us.

Amanda Askell

问这个很危险。

It's a dangerous thing to ask.

Host

好吧,好吧。

Well, okay. Okay. Yeah. Yeah.

Amanda Askell

但我认为很难提炼出具体方法,因为其中一点就是愿意大量与模型互动,真正地逐个查看输出,并借此感知模型的形态以及它们对不同事物的反应,愿意去实验。这实际上是一个非常经验性的领域,也许人们常常不理解的是:提示工程非常依赖实验。你面对一个新模型时,我会通过大量互动找到一套完全不同的提示方法。另外,理解模型的工作原理也有帮助。有时候,其实就是与模型进行推理,这非常有趣,并且要非常充分地解释任务。正是在这里,我认为哲学实际上可以对提示工程有用,因为我的很多工作就是尽可能清晰地向模型解释某个问题、担忧或想法,然后如果它做了出乎意料的事,你可以问它为什么,或者试着找出你所说的哪部分导致了它的误解,并且愿意反复经历这个过程。

But I think it is really hard to distill what is going on because one thing is just a willingness to interact with the models a lot and to really look at output after output and to use this to get a sense of the shape of the models and how they respond to different things, to be willing to experiment. It's actually just a very empirical domain and maybe that's the thing that people don't often get: prompting is very experimental. You deal with—you find a new model and I'll have a whole different approach to how I prompt that model that I find by interacting with it a lot. And I think a little bit also understanding how models work. Sometimes it's also just honestly reasoning with the models, which is really interesting, and really fully explaining the task. This is where I do think philosophy can actually be useful for prompting in a way, because a lot of my job is just trying to explain some issue or concern or thought that I'm having to the model as clearly as possible, and then if it does something kind of unexpected, you can either ask it why or try to figure out what in the thing that you said caused it to misunderstand you, and just a willingness to iteratively go through that process.

其他 AI 耳语者 Other AI Whisperers

Host

相关地,Michael Swaravverics 问:“你怎么看其他 AI 低语者,比如 Janus?他是在网上几乎以你描述的方式与模型进行实验性互动的人。”

Relatedly, Michael Swaravverics asks, "What do you think of other AI whisperers like Janus, who is someone online who is almost having experimental interactions with the thing in the way that you've described?"

Amanda Askell

是的,我觉得这非常有趣。我喜欢关注那些对模型进行非常迷人实验的人的工作。我也认为,有时深入探究模型、它如何看待自己、以及它在这些不寻常情况下的互动方式,非常有意思。我不知道,我觉得这些工作极其有趣。我认为它凸显了模型的深度。在某种程度上,我也认为那个社区能够对我们施加压力——如果他们发现系统提示或模型某些方面不够好——这就像心理学一样。

Yeah, I think it's really interesting. So, I love to follow and see the work of people who are doing these really fascinating experiments with the model. And I also think sometimes doing these deep dives into the model and how it thinks of itself, how it interacts in these really unusual cases. I don't know. I find the work extremely interesting. I think it highlights really interesting depths to the models. And in some ways, I also think that that community has been one that can hold our feet to the fire—if they find things that aren't great in the system prompt or in aspects of the model—and it's like psychology.

Host

是从模型福祉的角度,还是人类福祉的角度,还是两者兼有?

In the sense of from a model welfare perspective or from a human welfare perspective or both.

Amanda Askell

我认为两者是相关的,所以通常两者兼有,但我也非常欣赏那些从模型福祉角度出发的探讨。

I mean I think the two are related so often both, but I do also really appreciate it when it's coming at it from the model welfare perspective.

Host

这包括未来的模型。所以,不仅仅是系统提示之类的东西,如果你深入模型内部发现某种根深蒂固的不安全感,那非常有价值,但你可能需要随着时间的推移,通过训练以及在训练过程中给模型更多上下文信息来调整。

And that includes for future models. So, not just things like system prompts, but if you go into the depths of the model and you find some deep-seated insecurity, then that's really valuable, but that's something that you might actually need to try and adjust over the course of time with training and with giving models more information in context during training, for example. Okay.

Amanda Askell

所以,我不知道,我两者都欣赏——我喜欢看到人们用模型做这些非常有趣有用的实验,同时也指出我们可以通过更好的系统提示和更好的训练来改进的方法。是的,我认为这是非常有价值的工作。

And so I don't know, I appreciate both—I love seeing people do these really interesting useful experiments with models, but also pointing out ways in which we can improve things through better system prompting, but also better training. And yeah, I think that's really useful work.

AI 对齐与安全 AI Alignment and Safety

Host

有几个关于安全以及这些模型可能带来的更大风险的问题。Jeffrey Miller 问:“如果很明显 AI 对齐无法解决,你相信 Anthropic 会停止尝试开发人工超级智能吗?无论你怎么称呼它,你会有勇气吹哨吗?”

Couple of questions about safety and maybe the larger risks that these models pose. Jeffrey Miller asks, "If it became apparent that AI alignment was impossible to solve, would you trust that Anthropic would stop trying to develop artificial super intelligence, however you want to call it, and would you have the guts to blow the whistle?"

Amanda Askell

是的。我觉得这问题有点简单,因为如果很明显无法对齐 AI 模型,继续构建更强大的模型对谁都没好处。我总是希望我不是对这个组织过于乐观,但我确实觉得 Anthropic 真心关心确保一切顺利,并且以非常安全的方式进行,不会部署危险的模型。你知道,一个稍微难一点的问题是:如果在一个证据不断积累的世界里呢?情况非常模糊不清。

Yeah. So I guess this feels like a kind of easy version of the question because it's like if it became evident that it was impossible to align AI models, it's not really in anyone's interest to continue to build more powerful models. I always hope that I'm not just being Pollyannaish about the organization, but I do feel like Anthropic does genuinely care about making sure that this goes well and that it is done in a way that is very safe and not deploy models that are dangerous. You know, a different slightly harder question is like, well, what about being in a world where there's just kind of mounting evidence? It's really ambiguous and unclear.

Host

对吧?不是他描述的那种明确证据。

Right? It's not evidence in the way that he describes.

Amanda Askell

是的。而且不是不可能,而是困难或不确定。在这种情况下,我确实认为我们会足够负责任,比如,随着模型能力增强,你必须对自己提出更高的标准,以证明这些模型行为良好,并且你确实让模型拥有了好的价值观,或者在世界上表现良好,并且要负责任地行事。我认为组织会这样做,而且内部很多人,包括我自己,都会要求他们做到这一点——至少我认为这是我工作的一部分,而且很多人也这么认为。

Yeah. And it's not just impossible, but something like it's difficult or we're unsure. And in that case, I do like to think that we would be responsible enough to be like, look, as models get more capable, the standard that you have to hold yourself to for showing that those models are behaving well and that you actually have managed to make the models have good values, for example, or behave well in the world, is going to increase, and to behave responsibly and in line with that. And I think that is a thing that the organization is going to do, and a lot of people internally, myself included, will just hold them to that—at least I see that as part of my job, and I think many people do.

Host

Louie 说我没有问题,但谢谢你的回答。他们这么说真好。最后一个问题来自 real stale coffee。

Louie says I don't have a question but thanks for offering. So that's nice of them to say. Yeah. And the final one is from real stale coffee.

书籍推荐与对奇异性的反思 Book recommendation and reflections on strangeness

Host

你最近读的一本小说是什么,你喜欢吗?

What is the last book of fiction you read and did you like it?

Amanda Askell

我最近读的一本书是本杰明·拉巴图特的《当我们不再理解世界》。是的。这是一本非常有趣的书,随着情节推进,它变得越来越虚构。我认为对于从事人工智能工作的人来说,这是一本非常值得一读的书,因为很难捕捉到生活在当前这个时期的那种奇异感——新事物不断涌现,而你并没有现成的范式可以始终指引你。所以这本书很有趣,因为它更多是关于物理学和量子力学,但又不是真的在讲物理,而是关于人们对它的反应。我认为对于 AI 从业者来说,这是一本非常有趣的书,它能捕捉到当下时刻的一些东西以及它看起来有多奇怪,但另一方面,回顾那个时期以及当时许多参与者的感受也很有趣,而现在它实际上已经是一门更成熟的科学了。在某种程度上,我的希望是,在未来某个时候,人们会回顾过去,说:你们当时确实是在黑暗中摸索,试图弄清楚一切,但现在我们已经把一切都搞定了,一切进展顺利。

The last book that I read was by Benjamin Labatut, and it was When We Cease to Understand the World. Yes. And it's a really interesting book that becomes kind of increasingly fictional as it goes on. And I think for people working in AI, it's actually a very interesting book to read because it's hard to capture the sense of how strange it is to just exist in the current period where new things are happening all of the time and you don't really have prior paradigms that can guide you always. So it's an interesting book because it's more about physics and quantum mechanics, and less about the physics and more about this notion of people's reaction to it. And I think it's a really interesting book for people in AI to capture something about the present moment and how strange it can seem, but then also in some ways it's interesting to look back on that period and how it must have felt to many of the people involved, and now actually it's a more settled science. And in some ways maybe the hopeful thing that I have is that at some point in the future people will look back and be like, well, you guys were kind of in the dark and trying to figure things out, but now we've settled it all and things have gone well.

Host

那会很好。

That would be nice.

Amanda Askell

那会很好。那是梦想。我读这本书时,随着它开始时非常接近现实,然后逐渐变得脱离现实,我感到越来越困惑。我认为这里有一个元问题,即现实变得越来越奇怪,而这确实正在发生在我们身上。

That would be nice. That's the dream. I found an increasing sense of confusion as I read through it as it starts off being quite close to reality and then just sort of becomes unmoored as you go on. And I think there's a meta issue there of reality becoming stranger and stranger, which is definitely happening to us.

Amanda Askell

是的,不过在现实世界中,我认为现实变得越来越奇怪,然后几乎又变得更容易理解了。所以希望是,也许 AI 也会如此。我确实认为,如果我们能找到方法让这一切顺利进行,那么也许在未来,我们回顾这段时期时会说,那是一段事情变得越来越奇怪的时期,但最终我们做得还不错,并且对它有了很好的理解。这就是当你身处事情变得奇怪的过程中时的希望。

Yeah, though in the real world I think that reality became stranger and stranger and then almost became more understood again. So the hope would be that maybe that would be true of AI. I do think if we can find ways of making this go well, then maybe in the future we'll just look back on this and be like that was a period where things were getting stranger and stranger, and then eventually we managed to do okay and we formed a good understanding of it. That's the hope when you're in the middle of things getting stranger.

Host

我们现在正处于奇怪的阶段。是的,你可以希望它在某个时候变得不那么奇怪,但我不确定这是否是一种愚蠢的希望,但好吧。

We're at the weird part right now. Yes, you can hope that it becomes less weird at some point, but I don't know if it's a fool's hope, but yeah.

Amanda Askell

嗯,我认为这是一个很好的结束点。所以,非常感谢你回答所有人的问题。

Well, and I think that's a nice place to end. So, thank you very much for answering all those people's questions.

Host

谢谢你问我这些问题。

Thank you for asking me the questions.

互动版:逐字朗读 + 针对本期提问 →