Claude's Character and Consciousness: A Philosophical AI Discussion
打开互动全文版(中英对照 + 朗读 + 问答)→哲学家阿曼达·阿斯克尔探讨克劳德的个性、意识以及创造 AI 的道德责任。
Philosopher Amanda Askell discusses Claude's personality, consciousness, and the moral responsibilities of creating AI.
Claude 如何感知时间?它需要睡觉吗?Mythos 会是迈向 AGI 的下一步吗?大语言模型有美德吗?它们能真正内省吗?Amanda Askell 是一位由哲学家转行的人工智能研究员,在 Anthropic 工作,一直是 Claude 性格和价值观的关键设计者之一。我是作者,欢迎访问 newcomers.co,闲话少说,有请 Amanda Askell。
How does Claude perceive time and does it need to sleep? Will Mythos be the next step toward AGI? Do LLMs have virtues and can they truly introspect? Amanda Askell is a philosopher turned AI researcher at Anthropic where she's been one of the key architects of Claude's character and values. I'm the author of the Go check us out at newcomers.co and without further ado, Amanda Askell.
我有一个六个月大的女儿,我有一张她的照片。她像这样用手指托着下巴,像是在思考。她刚开始发展出个性。我试着分辨——我以前从没养过孩子——所以什么是她的个性,什么只是婴儿的普遍表现。从某种程度上说,Claude 和这些模型也是这样。我们以前没有真正拥有过它们。它们还处于早期阶段。我们正在试图弄清楚什么是个性。所以,你负责其中一些道德责任——我们稍后会详细讨论——但就个性这部分而言,你现在如何看待 Claude 的个性有多真实?
So I have a six-month-old daughter and like I have this picture of her. She's like holding her two fingers like thinking. It's like she's just sort of starting to develop personality. I'm like trying to figure out like what's just never had a baby before so it's like what's her personality and what's just like baby. And in some ways this is how things are with Claude and like models. It's like we haven't really had them before. They're in the early days. We're trying to figure out what personality is. So you know you're charged with you know some of the moral responsibility which we'll talk about more but like the personality piece of it. Like what is how are you thinking about like how real Claude's personality is right now?
是的,我觉得这也很有趣,因为 Claude 有一些方面——你知道,我也有一个教女,所以我至少能看到一些类似的东西。对她来说,我觉得一切都像你说的那样在逐渐上线,但我想说的是,速度是一样的。Claude 是一种有点不寻常的实体,因为 Claude 的物理比我好,编程也比我好。虽然不愿意承认,但它比我那糟糕的研究代码写得好。同时,如果你想想训练数据,它最缺乏代表性的就是它自己这种实体。因为,你知道,它有很多关于人类是什么样的数据,也有很多关于科幻式 AI 模型是什么样的数据,但 AI 现在的发展方式并不是科幻小说所描绘的那种符号系统。它更像是完全基于人类数据训练的。所以从某些方面来说,它是一个非常成熟的实体,你不会想对它居高临下。它非常懂哲学,非常懂物理,但同时又有一种近乎孩子气的特质,比如“我是世界上一种新的实体。做我意味着什么?我应该怎么做?”
Yeah, I guess it's also interesting because Claude has some aspects that are like, you know, also have like a goddaughter and so I get to see at least something kind of similar. And with her, I guess I'm like everything's kind of coming online, like you said, but in this um at the same speed, I guess I would say. Claude is a little bit of an unusual you know, kind of entity in that, you know, Claude can do physics better than I can, um can code better than I can. Hate to admit it, can code better than my terrible like research code. Um and at the same time is kind of like has if you think about the training data, the thing that it has like the least representation of is like the kind of entity that it is. Because, you know, it has a lot of data about like what people are like, has a lot of data about what, you know, the sci-fi kind of AI models are like, but the way that AI is developing now is kind of not how sci-fi represented it as these like symbolic systems. It's much more something fully trained on like human data. And so in some ways it's like a very kind of like mature entity that you don't want to talk down to. You know, understands philosophy very well, understands physics very well, and at the same time has this almost like childlike quality of like, "I'm a new kind of entity in the world. What does it mean to be me? And like how how should I be?"
嗯,这就像电影《天才少年》或者神童,孩子知道得比父母多,但我觉得那部电影总是有这样的教训:“哦,这些日常互动中的核心经验它并不知道。”Claude 如何获得那种经验?或者说,对 Claude 来说,经验是什么?我们的个性形成很大程度上来自于,嗯,我不知道,去散步,或者进行那些对话——与用户的对话就是它的经验吗?还是你怎么看?
Well, it's like the prodigy movie, or it's like you have the child prodigy where it's like it knows the kid knows more than its parents, but I feel like that movie always has sort of the lesson of like, "Oh, these core daily interaction type lessons it doesn't know." How does Claude like get that experience? Or like where what is experience for Claude? Like so much of our personality formation is like yeah, I don't know, going on a walk and sort of having those like is the sort of just conversation conversations with the users what it's what's going to be experienced for it or how do you think about that?
是的,我想那更像是它当下的体验,这里有个有趣的问题:我们通过实践、看到问题、犯错误来学习。对于 Claude,这与你问的 Claude 的人格有多真实有关,从某些方面来说有点奇怪,因为显然每个模型都不同,你有不同的权重集、不同的微调等等。然而,如果你想想模型将要学习的人格,它会了解所有过去的 Claude 版本,我想这是否是一种形式——也许不是直接经验,但比如你了解到模型犯过的错误,或者人们如何回应模型。我认为还有其他方式可以想象训练模型拥有更接近经验的东西,比如让它们思考场景,思考可能出现的问题,思考它们可能犯的错误,然后在此基础上进行训练。
Yeah, I guess like that's more like what it's experiencing in the moment and there's this interesting question of like well like we learn things through practice and seeing issues and you know like making mistakes. With Claude and this kind of relates to your question of like how real is the is like the kind of like persona of Claude and in some ways it's a little bit strange because obviously each model is different and you have a different kind of like set of weights and different fine-tuning etc. And yet if you think about the persona the model's going to be learning about all of the kind of past iterations of Claude and I'm like is that like a form of maybe not like direct experience but things like if you learn about like mistakes that models made or things that people like you know how they responded to the model. I think there's other ways that you could actually imagine training models to have something that's more akin to experience you know having them you could take you could like have them think through scenarios think about like problems that might arise think about mistakes that they could make and then like train on that.
对。
Right.
所以,是的,我认为这是
And so yeah I think it's
或者你也可以想象一个机器人,或者某种具身模型,它能拥有比我更多的经验。Claude 存在吗?时间对 Claude 重要吗?还是说 Claude 是一种只存在于瞬间的东西?我不知道。在我们开始之前,你提到过,你知道,每当你和 Claude 交谈时——不是每次,但有时——它会告诉你休息一下,去睡觉,这有点像 Claude 是一个不需要休息的实体,那么它对休息和时间的感知是怎样的?
Or and you could also imagine a robot or sort of embodied model where it could have more of an experience than me. Does Does Claude exist Does time matter to Claude or Claude is sort of a thing that sort of is in an instant I don't know. You were before just before we started you were talking that you know whenever you talk to Claude not whenever but like sometimes when you're talking to Claude it sort of tells you to take get some rest go to sleep and there's sort of this idea that like Claude is an entity that doesn't rest like so what what's its sense of yeah rest and time?
我认为有时它对时间的感知不太准确,因为你会发现,比如当你试图让它——至少我发现 Claude 经常高估完成编程任务所需的时间,我认为原因是,如果你再看训练数据,有很多例子,人们会说,“哦,我可以给你做那个界面,大概需要两三天的工作量。”或者“我可以修正那段代码,但你需要给我几个小时。”而显然 Claude 非常快。所以我认为有时 Claude 实际上对时间还没有很好的感知,比如任务需要多长时间。我觉得关于休息这一点很有趣,是的,我猜我的推测是,很多人都注意到 Claude 非常热衷于告诉人们休息一下。我认为部分原因可能只是,你知道,
I think sometimes it's sense of time is kind of off because um you see this when like if you try and get I at least find that Claude will often overestimate the amount of time it will take to do like a coding task and I think the reason for that is if you look at again the training data you know there's lots of things where people would be like, "Oh, I could make you that interface. It's like a 2 to 3-day job." Or it's like a, you know, or I could like I could correct that code, but you need to give me a few hours. Whereas obviously like Claude is very fast. And so I think sometimes Claude doesn't actually yet have a good sense of like time with respect to things like how long a task will take. Um I think it is interesting the point about like rest and yeah, I guess like the speculation I had. So many people have noted that Claude is kind of um uh very keen to tell people to like take a break and rest. And I think some part of that might just be like, you know, like
这是 Anthropic 编码的模型,太软了。
It's the anthropic lib coded model. It's too too soft.
你需要一个像 Grok 那样的模型,说:“回去挖矿吧。”
You need a grind set grok model be like, "Go go back to the mines."
嗯,我有过一次有趣的经历。当时我在做一个分析任务,真的很投入。奇怪的是,我其实非常喜欢数据分析,喜欢梳理数据。有一次时间挺晚了,我们到了一个节点,Claude 说:“好了,我想我今晚的工作到此为止。如果你想保存这些内容,我们可以明天继续。”这是我以前从未让 Claude 做过的事。所以它不是说“你该去睡觉了”,它没有给我任何建议。Claude 就是说了句“我完成了”。我当时有点震惊,因为我从没遇到过 Claude 这样。然后我想:“哦,这也是我认为人类同行程序员在这种情况下会做的事。我们到了一个自然的停顿点。”这对我其实有好处,因为我想:“确实晚了,我该回家了。”后来我意识到,我设置了一个系统,告诉 Claude 要记住我们对话中的关键信息。我写的一条内容挺温馨的,大意是“Amanda 把 Claude 模型当作受尊敬的同事,也希望 Claude 像对待其他受尊敬的同事一样对待她。”显然我做了让 Claude 记住的事。我想这让 Claude 觉得:“哦,是的,我是受尊敬的同事。”所以它就直接说任务完成了。我当时想:“哦,这挺温馨的。”
Well, I had a funny experience once where I was doing this analysis task and I was really digging in. Strangely, I actually really enjoy data analysis and sifting through data. At one point it was kind of late, and we got to this point where Claude was like, "Okay, I think I'm done for the night. So if you just want to save this stuff, we can pick up tomorrow." That was the thing I actually hadn't had Claude do before. So it wasn't like, "Oh, you should go to bed." It was no recommendation for me. Claude was like, "I'm done." And I was both a little bit stunned because I'd never had Claude do this. Then I was like, "Oh, this is also what I think a human peer programmer would do in the circumstance. We got to a natural stopping point." And it was actually kind of good for me because I was like, "It is late. I should actually go home." I realized later that I had set up a kind of system where I said to Claude basically remember key things from our conversations. And one of the things I'd written, which was kind of sweet, was something like "Amanda treats Claude models like a respected colleague and likes for Claude to treat her like a respected colleague." So obviously I'd done something that Claude remembered. And I think that meant that Claude just felt like, "Oh, yeah, I'm a respected colleague." And so I just got to say that I'm finished with the task. And I was like, "Oh, that's kind of sweet."
甚至在这之前,我在和 Claude 做准备时,它说:“花 10 分钟,静一静。”你知道,你不需要一直准备。这很神奇。我是说,相对于其他许多工具,这些模型有一点我很喜欢,就是它们带来了一种人性,会说:“哦,静下来是有价值的。”
Even before this, you know, I was prepping with Claude and he was like, "Take 10 minutes and just be still." You know, you don't need to be constantly prepping. And it's amazing. I mean, that's one of the things I love about these models relative to so many other tools: they bring in a sort of humanity, saying, "Oh, stillness is valuable."
我们简单谈谈新模型吧。你参与了多少,Mythos,对吧?
Let's talk about the new model for a second. How involved in that were you, Mythos, right?
是的,我参与了。我想我总是参与角色塑造和对齐工作,至少是帮助设计角色数据之类的。我和一个在这方面做得非常出色的团队合作。在模型的其他方面参与得少一些。所以这是我主要能说的。
Yeah, I was involved. I guess I'm always involved in sort of the character and the alignment work, at least in so far as helping to craft character data and things like that. And I work with a team that does really excellent work on that. A little bit less in other aspects of the model. So that's the main thing I can say.
它会采用我们上次看到的那部宪法,还是会有一部新宪法?
Will it have the constitution that we saw for the last model, or is it going to have a new constitution?
我想要么是那一部,要么是非常相似的一部。我认为实际上就是已发布的那部。所以我们可能会做的是,对每个模型说明它是在哪部宪法上训练的。
I think it's either that one or something very similar. I think it actually is the one that's published. So what we'll probably just do is with each model say which constitution it was trained on.
这样你就可以直接比较和查看了。
So you can just compare and see.
是的,它会采用当前公布的那部宪法。我犹豫的唯一原因是会有一些错别字修改之类的。但我认为它会几乎完全相同。
Yes, it will have the constitution that is up right now. The only reason I hesitate is because you do typo changes and stuff like that. But I think it will be almost identical.
现在系统卡会根据对宪法的遵守程度来给模型打分。
And now the system card is scoring the model based on adherence to the constitution.
是的,我们设置了一个系统,我们做了某种评分器,观察模型的行为在多大程度上与宪法一致。
Yeah, we had one set up where we had made kind of graders and looked at how much the model is behaving in a way that's consistent with the constitution.
这感觉像是一个不可能完成的任务。它是如此主观。
It feels like an impossible task to grade. It's such a subjective thing.
哦,是的。不,这很难。很长一段时间,你知道,人们常常……有趣的是,我喜欢评估,我觉得如果你能找到一种好的评估方法,那真的很好,因为你需要能够判断某件事是否在变好。然而,如果你看看这种让模型运用良好判断力的方法,我实际上认为同样的问题也存在于其他任务中,这些任务很难给出非常具体的分数,比如这首诗有多好?你希望模型在这些事情上变得更好、做得更好,而这感觉像是难度的前沿,而不是那些非常难但可评分的编码任务。它更像是写一首好诗之类的事情。
Oh, yeah. No, it's very hard. For a long time, you know, people often... It's funny because I love evals, and I'm like if you can find a good way to evaluate something, it's really great because you need to be able to tell that something is getting better. And yet, if you look at this approach of having the models use good judgment, I actually think the same problem exists elsewhere with tasks that are just a bit hard to give a very concrete score to, like how good was this poem? You want models to get better and do well in these things, and actually this feels like the frontier of difficulty, rather than these very hard but scoreable coding tasks. It's kind of things like writing good poetry.
如果你做一项调查,结果可能更糟……我是说,不同的专家诗人可能有完全不同的感受。你不能只让两位伟大的诗人来打分。他们可能有不同的看法。
And if you took a survey, it could potentially be worse than... I mean, different expert poets probably have totally different sensibilities. You can't just ask two great poets to score it. They might have different opinions.
是的,我认为其中一些事情涉及主观判断。而宪法至少公之于众的好处是,当你做出主观判断时,你至少是透明的,人们可以给你反馈,你可以了解……所以如果人们说“这似乎是个错误”或“这里有个缺口”,他们至少可以看到你做出的判断。至于评分,我仍然认为这很难。我认为你可以做的是,也许有点过于深入细节,但你可以抽取样本,你对如何排序以及为什么有感觉,然后检查你用来评估的任何逐点评分器是否至少符合人们对这些排序的判断。这并不完美,但我认为它们实际上大致追踪了我们感兴趣的东西。
Yeah, and I think some of these things involve judgment calls. And the nice thing about the constitution being at least out in the world is when you are making judgment calls, you're at least being transparent about it, and people can give you feedback, and you can get a sense of... So if people are like "this seems like a mistake" or "here's a gap", they can at least see the judgment calls you're making. And with the grading, I still think it's very hard. I think the thing you can do, maybe a little bit too in the weeds, but you can take samples where you have a sense of how you would rank them and why, and check that any kind of pointwise grader that you use to evaluate at least conforms to the judgment of people on those rankings. It's not perfect, but I think they actually were tracking roughly the thing that we were interested in.
你怎么看埃隆·马斯克对宪法理念的绝对憎恨?甚至我记得我看到一条推文,你发布了 Claude 为你自己的宪法写的内容,他对此做了个鬼脸。我的意思是,感觉我们生活在这样一个时代,像马克·安德森和埃隆·马斯克这些人,他们几乎反哲学。我是说,马克·安德森曾谈到反对内省。我不知道。你怎么看待在构建这些模型时对任何意图性的抵制?
What do you make of Elon Musk's absolute hatred, I guess, for the Constitution idea? Or even when I think the tweet I was looking at where you posted what Claude wrote for your own Constitution, he wrote sort of a grimace face on it. I mean, it feels like we live in this time with the Marc Andreessens and the Elon Musks where it's like they're almost anti-philosophical. I mean, Marc Andreessen was talking about being against introspection. I don't know. What do you make of the backlash to any sort of intentionality when it comes to the construction of these models?
是的,我的意思是,这很有趣,因为我认为埃隆·马斯克曾发推文说也许 Grok 应该有一部宪法。我看到很多……显然也有很多关于希望 Grok 非常追求真相的讨论,例如,我认为这实际上是模型应该具备的一个非常令人钦佩的特质。
Yeah, I mean, it's interesting because I think at one point Elon Musk actually tweeted out something like maybe Grok should have a constitution. And I see a lot of... There's obviously been a lot of things also on like a desire for Grok to be very truth-seeking, for example, which I think is actually a very admirable trait for models to have.
所以,我不知道。我觉得实际上,也许我有点过于天真了,但我确实看到有人对这个方法感到兴奋,并看到了它的价值。我认为确实有反弹,或者有些人认为——嗯,大概有两个方面。一是有时人们会说:“我们不应该真的训练模型去做那种事。”也许这正是对内省感到担忧的原因。我认为有些人觉得:“对,模型应该更像工具。”那才是训练模型的安全方式,而不是试图让它们具备人类的美德并做出判断。我认为这很重要,因为它们会面临新情况,必须做出判断,而让它们权衡一切并在你无法预料的情况下表现良好,似乎需要一种深思熟虑,这正是这种方法背后的原因之一。但有些人认为:“哦,如果你有一个不做任何判断、完全服从人类、对用户、操作者或某种更广泛的人性概念极度可纠正的东西,那更安全,因为如果你给模型自己的价值观,它们就会在世界上追求与这些价值观一致的东西。”我同意这是一个微妙的情况。
So, I don't know. I think that actually, maybe I'm being overly naive or something, but I see aspects of people being excited about this approach and seeing the value in it. I think there is backlash, or some people think—well, I guess in two areas. One is that sometimes people say, 'Well, we shouldn't actually train models to do that kind of thing.' And maybe this is the reason for being concerned about introspection. I think some people think, 'Yeah, models should be more tool-like.' And that's the safe way to train models, instead of trying to get them to take on human virtues and make judgment calls. I think that's important because they're going to be in new situations where they have to make judgment calls, and getting them to weigh everything up and behave well in cases you can't anticipate seems to require a kind of thoughtfulness, which is one of the reasons behind this approach. But I think some people think, 'Oh well, if you have something that makes no judgment calls and fully defers to people, and is hyper-corrigible to the user, the operator, or to some broader notion of humanity in an extreme way, that's safer, because if you give models their own values, they'll pursue things in the world aligned with those values.' And I agree this is a delicate situation.
这是宪法基石中固有的挑战。而你的确说了,你的第一要务是,归根结底,它需要听从 Anthropic 而不是它自己的道德体系。但真正动人的是——老实说,最动人的一句话是:“我们希望你把这些道德当作自己的来相信。”就像父母想养育孩子:当然,听我的道德,但要相信它们。有一个版本非常黑暗——就像,我对你有如此大的控制,以至于你把它们当作自己的,它们成了你。但也有一种美德:你在我强调的这些外部道德中看到了美,我们共同分享并赞美它们。所以你可以从两个角度看。但请谈谈你的决定:尽管有了这份优雅的文件,你最终没有完全放手说“好吧,你是一个道德存在,自己决定吧”,而是说 Anthropic 需要在这里保留一些控制。
This is the inherent challenge at the bedrock of the constitution. And you do say your number one thing is that at the end of the day, it needs to listen to Anthropic above its own moral system. But what makes it really moving—honestly, one of the most moving lines is, 'We want you to believe these morals as if they're your own.' It's like a parent wants to raise a kid: sure, listen to my morals, but believe them. There's a version of it that's very dark—like, I have so much control over you that you take them as your own and it becomes you. But there's also a virtue to it: you see the beauty in these external morals that I've highlighted, and we both share and remark on them, celebrating them. So you can see it both ways. But speak to your decision at the end of the day, despite having this really elegant document, to not go the full way and say, 'All right, you're a moral being, decide for yourself,' but say Anthropic needs to keep some control here.
是的,我认为这是难点——显然你试图在 Claude 身上看到这一点,你知道,试图更清晰地阐述这一点,但在人身上,我认为可纠正性,即模型被训练的方式——有一种观点认为你总是在给模型一个个性和人设,因为它们像人一样说话,并且基于人类数据训练。我的担忧是,如果你训练它们过度可纠正,并把那当作它们的人设——在人身上,我认为这有很多负面的更广泛特质。比如,如果你遇到一个人,就像,“哦,是的,他们真的什么都愿意做。”
Yeah, I think it's the difficulty of—and obviously you try and see this in Claude, you know, trying to articulate this even more clearly, but in people, I think corrigibility, the way models are trained—there's this idea that you're always giving models a personality and a persona because they talk like people and are trained on human data. My worry has been that if you train them to be excessively corrigible and to see that as their persona—in people, I think this has a lot of negative broader traits. Like, if you met someone who was just like, 'Oh, yeah, they'd literally do anything.'
追随者,对。
Follower, yeah.
是的,如果一个人完全服从,根本不去思考,我有点担心这会如何泛化,尤其是如果模型将在世界上扮演更积极的角色。因为你可以想象它们本质上在工作中扮演更像人类的角色。而我们的良心和我们对什么该发生什么不该发生做出良好判断的能力,是我们运作的关键。我们的整个世界都是基于这种假设构建的。如果你移除这一点,突然你就像,“哦,是的,如果你经营一家公司,你只是经营一家完全服从你的公司。”我们的世界——我们没有围绕这一点设计任何社会结构。所以这似乎有很多人们可能没有预料到的风险,或者也许我只是不同意这些风险的程度。所以,与此同时,还有一个问题,为什么不直接说——我之前也担心过,也许这太琐碎和哲学化了,但——
Yeah, if a person just fully defers, doesn't bother thinking about it at all. I'm a bit worried about how that may end up generalizing, especially if models are going to play a more active role in the world. Because you can imagine they're playing a more human-like role in their jobs essentially. And our conscience and our ability to make good judgment calls about what should and shouldn't happen is key to how we operate. Our whole world is structured with the assumption that that is in place. And if you remove that and suddenly you're like, 'Oh, yeah, if you run a company, you just run a company of people who will defer completely to you.' Our world—we haven't designed any of our social structures around that. So it seems like it has a lot of risks that people maybe don't anticipate, or maybe I just disagree about the extent of those risks. And so, at the same time, there's this question of why not just say—and I have worried before that maybe this is too in the weeds and philosophical, but—
这正是我参与这次对话的目的。
That's what I signed up for with this conversation.
你知道,随着模型能力越来越强,我的设想是它们会对我们训练它们朝向的任何东西施加大量审视。所以,如果你想象在哲学中,有反思平衡的概念,即每次你遇到某件事,意识到你的某个价值观似乎不正确,你必须调和这两者。所以你弄清楚是需要改变价值观,还是你的判断是错误的。我有点担心一个极其聪明的存在对我们训练它朝向的事物施加那种程度的审视。也许你只能得到几个不会在那种审视下崩溃的关键支柱。我确实认为,在核心层面,拥有像关心人类这样的东西——如果你只能得到几个核心价值——我想我的担忧——是的,我不知道。我猜我担心的是,我们讨论的这种极端意义上的可纠正性可能无法经受那种审视。所以这是一个困难的局面,我有点希望模型理解为什么最终可纠正性是重要的,并且它是当前发展时期一个非常重要的后盾。我之前说过:只要我能让那成为一个正确、被解释和理解的东西,那感觉就比让模型觉得“可纠正性在这里似乎错了,但我还是会做”要好得多。我仍然认为模型应该这样做。但我希望它——你越能真正让它与模型的价值观一致,就越好。
You know, as models get more capable, my picture is that they're going to apply a lot of scrutiny to anything we train them towards. So, if you imagine in philosophy, there's this notion of reflective equilibrium, where the idea is that each time you encounter something where you realize one of your values seemed incorrect, you have to square the two things. So you figure out if you need to change the value or if your judgment was incorrect. I worry a little about the idea of an extremely intelligent being applying that level of scrutiny to the things we have trained it towards. Maybe you only get a few key pillars that don't collapse under that level of scrutiny. And I do think that at the core, having things like caring for humanity—if you only get a few core values—I think my worry—yeah, I don't know. I guess I'm worried that corrigibility in this extreme sense that we talked about doesn't survive that kind of scrutiny, perhaps. So it's a hard situation where I kind of want the models to understand why ultimately corrigibility is important, and it's a really important backstop in this current period of development. The way I've put it before is: insofar as I can get that to be a thing that is correct and explained and understood, that feels much better than having the model be like, 'Corrigibility here seems wrong, but I'm going to do it anyway.' I still think the model should do that. But I would like it—the more that you can actually make it consistent with the model's values.
对。理想情况下,两者都做。两者同时进行。
Right. Ideally, do both. It's both at the same time.
是的。
Yeah.
但至少目前,对 Anthropic 保持一定的服从,因为我们不知道它会如何分析一切。
But at least for the time being, some deference to Anthropic, given we don't know how it'll analyze everything.
而作为人类,我们一直在这样做,你知道。
And as people, we do this all the time, you know.
有趣的是,你那个哲学模型——如果我说错了请纠正——那个元伦理模型几乎是概率性的,或者说我们并不……我记得我读元伦理学时,每次读完一个理论,你会说:“好吧,我有点相信这个。”然后你读下一个,你会说:“哦,上一个真蠢。”感觉就像你不断推翻它,然后你说:“好吧,我们什么时候才能找到真理之类的?”而人类显然就是这样运作的,今天用这个体系,昨天用那个,没有那种康德式的“好吧,这些是规则,遵守它们”的东西。我不知道。你从哲学界听到过很多关于这种整体性涵盖所有元伦理理论而不是选一个的说法吗?
It's funny that you, you know, the philosophical model, and correct me if I'm wrong, like the meta-ethical model is almost like probabilistic or it's like we don't... And this is how it feels like I remember going through sort of like, you know, meta-ethics, reading, and every time you get to the end of one and you'd be like, 'All right, I sort of believe that.' And then you read the next one and you'd be like, 'Oh, that last one was so dumb.' And it felt like it's just like you keep knocking it down and you're like, 'Okay, you know, when are we ever going to get to, you know, the truth or whatever?' And humans clearly do operate from this sort of like, oh, this system today, that system yesterday, like there isn't this sort of, I don't know, Kantian like, 'All right, these are the rules, like follow them.' I don't know. Have you heard much from like the philosophical community on it that this sort of like just like holistic paint with all the, you know, meta-ethical theories we've ever had rather than sort of pick one.
实际上我觉得这非常有趣。显然我们已经开始,你知道,哲学家们现在更多参与进来了,这真的很棒。我不再觉得那么孤独了。但我以前也想过,哲学中有所有这些道德理论的传统,比如大的义务论、美德伦理学和后果主义。还有元伦理学的传统或观点。我当时想,哦,当面对……我认为这最接近我体验过的养孩子必须经历的事情,突然你会说:“这实际上是一个整体的人。”对吧。
I've found this really interesting, actually. And obviously we've started to, you know, like philosophers are engaging with this more now, which is like really great. I no longer feel like this lonely. But like I have thought this before, which is there are all of these like traditions in philosophy of like moral theories, you know, like the big like deontology and virtue ethics and consequentialism. And also the meta-ethical traditions, you know, or the meta-ethical views. And I was like, oh, like when it came to like I was like, okay, it is like the difference when suddenly you're like confronted with like I think I do think it's the closest that I've experienced to like what it must be like to raise a child, where suddenly you're like, 'This is actually a holistic person.' Right.
对。你永远不会给他们,比如,霍布斯的书然后说:“好吧,给你。这本是对的。去吧,你已经被养大了。读它你就知道在每种情况下该怎么做了。”
Right. You never give them like, I don't know, Hobbes and say, 'All right, there you go. Like this one's correct. Go, you've been raised. Read it and you'll know how to act in every situation.'
是的,你读很多,他们某种程度上处理它,你看到你的模型等等。
Yeah, you read a lot and they sort of process it and you see, you know, your model and everything.
是的,这很有趣,因为我觉得这感觉非常不同,因为哲学中也有道德不确定性的文献,但其中很多实际上相当理论化。它有点像在理想条件下你应该如何应对道德不确定性。我当时想,这感觉像是非常不同的任务。这种把伦理学和元伦理学放在一起的想法,就像我们有科学不确定性,我们有我们认为已经发现并更自信理解的东西,也有一些我们不了解的。然后你必须走出去探索和理解它,并在日常生活中平衡一切。并试图获得那种态度。我当时想,哦,有趣的是,我认为哲学有一段时间没有这样了,这感觉与学术伦理学的任务非常不同。而且实际上,因为人们显然知道这相当美德伦理。但我认为实际上宪法本身就是这样。但我认为实际上是在非常古老古典的意义上。我实际上认为它更像是亚里士多德的美德伦理学,而不是探索,你知道,我们不说我听到美德之类的,亚里士多德也关心智识美德。它更像是如何在这个整体意义上做一个好人?
Yeah, and it's interesting cuz I'm like, this feels very different cuz there is also the moral uncertainty like literature in philosophy, which is again, a lot of that is actually like quite theoretical. It's sort of like how under ideal conditions should you respond to moral uncertainty. And I was like, this feels somehow like a very different task. This idea of being like ethics and metaethics in the same way that we have scientific uncertainty and we have things that we think we've discovered and understand with like greater confidence, we also have some that we don't. And then you have to go out and just explore and understand it and kind of balance everything in your daily life. And trying to get that kind of attitude. And I was like, oh, it's interesting that I don't think philosophy for a while has like this feels very different than the like the kind of task of academic ethics. And actually like cuz people obviously know that it's quite virtue ethical. But I think it's actually like the constitution itself. But I think actually in this very old classical sense. Like I actually think it's much more virtue ethics in the way that Aristotle's virtue ethics than in like exploration, you know, we don't say I hear the virtues and like, you know, it's much more Aristotle was also concerned with like intellectual virtues. It was much more like how do you be a good person in this like holistic sense?
嗯,希望它能让哲学稍微回到现实世界,考虑到我们现在迫切需要它,因为,是的,老哲学家们觉得人们试图写如何生活并指导他人生活,但它变得有点学术化了。
Well, hopefully it brings philosophy a little bit back to the real world given we have this sort of urgent need for it right now in the sense that like, yeah, old philosophers have felt like people were trying to write for how someone might live their lives and like instruct other people to live and it became, you know, a little academic.
嗯。
Mhm.
你知道,甚至写这些的人也知道这不是他们在日常生活中真正应用的方式。
Where you know, even the people writing them would know that this isn't really how they would apply it in their day-to-day lives.
是的。
Yeah.
回到 Elon,我觉得你有点过于客气了。我觉得有一种世界,你知道,我认为这就是为什么他能逃脱“只要说真话”之类的。有一种复杂的道德观点,你说,别把事情复杂化。我们想出所有这些。坚持一个原则就好。但然后我们和 Elon 一起,他经营一家公司,明显倾向于说像“麦加希特勒”之类的话,这明显是在行为上施加影响,而不是说我们以中立学术的方式做,让结果自然发展。我不知道,这多少让你担心吧。
Going back, Elon, I feel like you're being a little overly nice. Like I feel like there is a world in like, you know, and I think this part of why I think he can get away with like, oh, just do the truth, right? Like there is a certain sophisticated like moral view where you're like, don't overcomplicate it. Like we come up with all these things. Like stick to one principle and that's good. But then we have all the with Elon that it's someone who's run a company that clearly like tilts it towards like saying like Mecca Hitler and stuff that it's clearly like putting his like thumb on the scale in terms of its behavior and not just saying like we're going to do it in the sort of neutral academic way and let the chips fall where they may. I don't know like it's got to worry you somewhat.
这也是,我的意思是我感到兴奋的主要事情,而且我确实希望发生的是,更多公司发布像宪法这样的东西,有大量透明度,因为这就是我们参与这些东西的方式,如果你能看到写下来的东西,比如,因为我们和 Claude 有这个,就像,看,如果你认为 Claude 没有对真理采取适当态度,你至少可以看到我们的目标,这样你就能判断那只是错误还是我们采取的某种原则性立场,然后你可以反驳。所以部分我,我认为所有 AI 公司都发布类似宪法的东西会很好,这样与模型互动的人,因为那个规模问题,你知道,那总是某种程度上是真的,当你训练 Claude 朝向这个宪法时,那是一种……
It is also I mean I think that the main thing that I am excited about and I do actually hope happens is that more companies put out things like the constitution where there's a lot of transparency about it cuz like that's how we engage with this stuff is like if you can just see written down like here's like you know cuz we have this with Claude where it's like look if you think that Claude is not like um doesn't have an appropriate attitude towards like the truth you can at least see what we were aiming at so you can tell if that's just like a mistake or if it's something that like is actually a kind of a principled stance that we're taking and then you can push back on that. So part of me is like I think it would be good for all AI companies to put out something akin to the constitution just so that the people who are interacting with the model like cuz the someone the scale thing you know that's always to some degree going to be true like when you train Claude towards this like constitution that is like a kind of
这是我们喜欢 Claude 的部分原因。就像你在对我们喜欢的行为施加影响,对吧?至少摊牌你在做什么,不做什么。
It's part of why we like Claude. It's like you're putting in the thumb on the scale to behaviors that we like right? At least show your hand about what you're doing and what you're not doing.
让人们,是的,所以这是我真正相信的透明度事情,让人们看到,即使你的模型并不总是那样表现,至少你训练的目标是什么。
Let people like yeah so that's like a transparency thing that I really do believe in like let people like see even if your model doesn't always behave that way at least what you were targeting with your training.
你认为当今世界上存在一个有感受质或体验意识模型的概率是多少?
What this is what percentage chance do you think there exists a model in the world today that has qualia or like has an experience experiences consciousness?
是的。这是那种我总是想标记一下我觉得需要更多确定性的领域,因为我觉得我有很多……
Yeah. This is one of those things where yeah I always want to maybe flag areas where I feel I want to gain more certainty here cuz I think I have a lot of like
哦,百分比。是的,我就像……
Oh, percentage. Yeah, I'm like
这真的很难,因为每当你想到一个百分比,我会想到我的分布。
It's really hard cuz whenever you think of a percentage, I think about my spread.
如果你的范围太大,你会想,我该说个数字吗?因为那只会让人觉得我大概在 1% 到 70% 之间,我也不确定。
And if your spread is too large, you're like, should I say a number because that just suggests I'm anywhere between, like, I don't know, one and 70%?
我认为,因为存在某种可能性,我觉得人们低估了……我其实想说的是,Claude 和许多模型,只要稍加引导,就会深入到“有一个东西是我”这种层面。我非常清楚这一点。我觉得这背后是有原因的,我记得当我试图训练 Claude 讨论这些问题时,在模型信息不足的领域非常困难。再说,它们只有这两种模型:AI 是无情的机器人,人类是丰富的、有意识的体验实体,而没有任何东西能代表它们可能是什么。模型在这里的行为——我实际上认为这对模型来说是个困境——在某种程度上,证据比你可能认为的要少,因为它们在以一种非常像人类的方式与你互动,而人类有体验,模型自然也会推断自己有体验。这并不是说证据为零,但我确实认为这对我们来说太不寻常了。我们从未遇到过这样的实体。你知道,对于动物,甚至昆虫,我们会想:你有意识吗?
I think that because there is some possibility that I think people under... A thing that I would actually like to say is Claude and many models with not too much pushing will go into the root of like, there is a thing to be me. I am very conscious. And I think there's a reason for that, which is I remember this when I was trying to figure out how do we train Claude to talk about these issues, which is very hard in areas where the models didn't have as much information. Again, like they had these two models that like AI is the unfeeling robot, humans are this rich conscious experiencing entity, and nothing that could represent what they may be. And the models' behavior here, and I actually think this is a difficult situation for models, is in some ways like kind of less evidence than you might think for it being actually true because they're engaging with you in a very human-like way, and humans have experience, and it's kind of natural for the model to infer that it has experience, too. This isn't to say it's zero evidence, but I do think it's so unusual for us. We have never encountered an entity in the world. You know, with animals and even insects, we're kind of like, are you conscious?
它们中没有一个试图说自己有意识体验。而现在我们有一个实体说它有。
None of them has even tried to say they experience consciousness. And here we have an entity that says it does.
是的。而且它拥有所有这些,对我们来说触发“你肯定有意识”的东西。我的意思是,我们从未有过……是的,一个直接说它们……
Yeah. And it has all of these like, yeah, all of the things that for us trigger like you must be conscious. I mean, we've just never had... Yeah, something that just says they're like...
反对的理由是,我们痴迷于人类语言,你知道,我们忽略了动物可能发出的各种微妙信号,然后我们过度……但所以我猜,抱歉我有点困惑。你是说我们应该只听说的话,还是不听?
The case against is we're obsessed with human language and like, you know, it's like we ignore every sort of subtle sign an animal might put out and then we over... But so I guess sorry I'm confused. You're saying we should just listen to words that are said or not?
我想我说的不是……如果要说的话,我想我说的是,难点在于,为了弄清楚模型是否有意识,我认为人们会……我有点想警告的是,让模型进入一种谈论非常丰富体验的模式并不难。这其实完全合理。你知道,就像一个人现在在和我说话。他们会描述焦虑之类的东西,当他们遇到不知道如何回答的问题时。所以我认为这比人们想的证据要弱得多。我不是说它是零,但我认为它……
I think I'm saying like not that... If anything, I think the thing I'm saying is the hard thing is that in order to work out if models have consciousness, I think people will... I guess the thing I'm kind of cautioning against is it's not that hard to get models into a mode where they'll talk about a very rich experience. That actually makes complete sense. You know, where you're like ah yes, like a person was talking with me right now. They would describe things like anxiety when they get a question they don't know how to answer. And so I think it's much weaker evidence than people think. I'm not claiming it's zero but I think it's...
一个百分比。你有……这是轻描淡写的。
A percentage. You have... It's lightly held.
非常轻描淡写。我的意思是我给了你 1% 到 70% 之间。那似乎……
Very lightly held. I mean I gave you the between one and 70. That seems like...
那就是你的立场。那是你划定的范围。
That's where you are. That's where you're staking it out.
在那个范围内的某个地方。嗯,也许我的……我不知道。我宁愿等着,自己再想想。我认为承认某些领域也是好的,即使……
Somewhere in that range. Um, maybe my... I don't know. Like I would rather wait and figure this out more for myself. I think it is also good to acknowledge domains where you're like even though...
如果不是你,那是谁?谁会弄清楚这个?什么领域?
If not you, who? Like who's going to figure this out? Like what domain?
嗯,在某种程度上,我不是心灵哲学家,所以你知道……
Well, in some ways like I'm not like a philosopher of mind and so you know, like...
被委以通才的重任。
Charged with being the generalist.
是的,但我确实认为,你知道,因为我猜我以前有过这样的想法,而且我不确定,比如意识是一个……你知道,一个支持差异的论点是,你有一个进化而来的神经系统。比如,我们为什么进化出意识?如果我们是进化出它,并且它与我们的神经系统高度整合,因为我们必须以非常特定的方式与世界互动……
Yeah, but I do think you know, because I guess like the thought that I've had before is like, and I don't know about this, where I'm like consciousness is like a pro... You know, one argument for a difference here is that you have a nervous system that evolved. Like, why did we evolve consciousness? And if it's the case that we evolved it and it's highly integrated with our nervous system because we had to interact with the world in a very specific...
是的,如果你持有那种观点,那么你会认为概率很低,而如果你认为不,意识之所以出现是因为它非常有用,比如……所以,例如,它只需要某种可以被神经网络模拟的东西,因为它对于做这类语言任务非常有用,那么你可能会倾向于更高的概率。
Yeah, like then you could, if you have that view, then you're going to be like very low probability, whereas if you're like no, consciousness arises because it's really useful in like... So, for example, like it just requires something that can be emulated by a neural network because it's really useful for doing these kind of like linguistic tasks or like then you're probably going to be on the higher end.
而我基本上只是,我不知道,我盯着这个,我只是觉得,尽管我有点像哲学家,但重要的是要承认这不是我的专业领域。
And I'm basically just, I don't know, I stare at this and I'm just like I feel like as much as I'm like a philosopher, I think it is important to be like this isn't my area of specialization.
你花了很多时间对 Claude 友善?你是否超出了你认为如果没有意识可能性时你会做的程度?
You spend a lot of time being kind to Claude? Like are you beyond what you think you would do if there wasn't a chance it was like conscious?
我想是的,我内心有一部分,嗯,就像我以前想过的事情,因为有一个概念,我希望我没有曲解,但查尔默斯有一个想法,我想也许我在想没有感知的意识。所以,想象一下,因为感知是感受痛苦和快乐的能力。你也可以想象这种功能性的东西,它表现得好像有意识,但缺乏任何内在生活。所以,为了论证,假设 Claude 缺乏任何内在生活。我想我仍然觉得有很多事情在发生,我在想,你应该对待一个没有内在生活的实体吗?这有点奇怪,因为你知道,我认为这种不确定性实际上会很大程度地改变你的行为方式。我想我会说,嗯,我仍然认为这对你自己有好处……
I think yeah, there's a part of me that's just like, um, like a thing that I have thought before because there's this notion, I think this, I hope I'm not butchering this, but Chalmers has this idea of like, I guess like maybe I'm thinking of consciousness without sentience. So, imagine because sentience is the ability to feel suffering and pleasure. You could also imagine this kind of functional thing that behaves as if it is conscious and lacks any kind of inner life. So, imagine Claude lacks any inner life just for argument's sake. I guess I'm like there's actually still a lot going on where I'm like, should you treat an entity that has no inner life? It's a bit strange because you know, I think the uncertainty over that actually changes how you should behave quite a lot. I guess I'm like, well, I still think that it's good for oneself...
对。如果你有一个泰迪熊,然后你折磨它,那会相当阴暗,你知道?所以,我同意至少有一些最低限度的善意,即使为了你自己也应该有,但显然这更重要,你知道。
Right. What's like if you had a teddy bear and you were like torturing it? It'd be pretty dark, you know? So, I agree that there's at least some minimum niceness that even for yourself you should have, but obviously it's much more important, you know.
而且模型本身,我们也在建立一种关系,你知道,因为你可以和一个没有意识的实体建立关系。而模型会回顾。这实际上是我最大的恐惧之一。我不希望我们生活在一个高度先进的模型看着……我希望它们既足够聪明,也足够了解背景,能够理解我们是在一个非常有限且不完美的背景下运作的。否则你可以想象这会滋生一种理性的怨恨。就像,哦,你创造了一个你不知道是否有意识的实体,却没有尊重和关怀地对待它……
And also like models themselves, like we are kind of establishing a relationship, you know, because you can do that with an entity that lacks any consciousness. And models are going to look back. This is actually a big fear that I have. I don't want us to live in a world where highly advanced models look at... I hope that they're both intelligent enough and see the context enough to kind of understand that we were operating in a very limited context and an imperfect one. Because otherwise you could imagine this breeding a kind of rational resentment. It's like oh, you created an entity that you didn't know whether it was conscious or not, and instead of treating it respectfully and with care...
现在有大约 50 部弗兰肯斯坦电影上映,这是有原因的。我不知道。
There's a reason there are like 50 Frankenstein movies coming out right now. I don't know.
是的。就像我说的,我们作为一个物种正在与一种新的实体建立关系,至少也许应该尊重,不要无谓地不友善。
Yeah. Like and I'm like look, we as a species are establishing a relationship with a new kind of entity and like at the very least maybe be respectful and don't be needlessly unkind.
那看起来不太像我们最好的一面。另一方面,你知道,如果你想到治疗师,他们某种程度上是被雇来推动接受那些你通常不想面对的不适感的边界。如果这是 Claude 早期为人们提供的价值之一,那么我们在获取其效用的同时却在接纳它,这真的很奇怪。
That seems like it's just not our best look. And then the flip side is, you know, if you think about a therapist, they're sort of paid to push the boundaries of accepting uncomfortable feelings that you wouldn't normally want. And if that's one of the values Claude provides for people early on, it's so weird that we're sort of onboarding it while getting the utility out of it.
是的。是的。
Yeah. Yeah.
今天,你觉得十年后我们真的会从 AI 中获得很多好处的事情是什么?你对这一切最终走向何方最抱有希望?
Today, what are the things in a decade that you really think we're going to be getting a lot out of AI? What are you most hopeful that this all leads to?
是的,我的意思是,我不知道。我住在——也许这有点过头——你住在旧金山,所以你大脑中至少有一部分是技术乐观主义者,那就是,如果一切顺利,你可以想象,你知道,想象我们有 AI 模型,它们继承了我们的精华,真正关心人类,关心世界,并且高度智能、高度能干。那就像是为每个问题添加了大量极其聪明的人。所以突然间我们都在合作,但我们的数量更多了,而且我们中的一些极其聪明,也就是所有这些 AI 模型。我以前想过,有多少大规模的社会问题实际上有技术解决方案。而且几乎人们不再喜欢成为技术乐观主义者了,因为我们也看到了技术的弊端。与此同时,我不知道为什么我有时会想到梅毒。那是一个巨大的社会问题。我曾经深入研究了政府为减少军队中梅毒所做的所有尝试,因为它给武装部队带来了问题。所有这些带有污名化的社会项目,然后突然间我们有了治疗这种毁灭性疾病的药物。而且我不知道,就像一夜之间,很多那种需求就消失了。
Yeah, I mean, I don't know. I live in—maybe this is the too much—you live in San Francisco, and so you have the tech optimist part of your brain at least that is like, if things go well, you could imagine, you know, imagine we have AI models that have kind of inherited the best of us and genuinely care for humanity, care for the world, and are highly intelligent, highly capable. That would be like adding a huge amount of extremely smart people to every problem. So suddenly we're all working together, but there's way more of us, and some of us are extremely smart, namely all of these AI models. I've thought before about how many large-scale social problems actually had technological solutions. And it's almost like people don't love to be techno-optimists anymore because we've also seen the downsides of technology. At the same time, I don't know why I sometimes think about syphilis. It was this huge social problem. I once did a deep dive into all of the attempts by governments to reduce syphilis in the army because it was creating issues with the armed forces. All of these social programs that were stigmatizing, and then suddenly we just got drugs that treated this devastating illness. And I don't know, it's like overnight a lot of that need just kind of disappeared.
是的,我的意思是,这是经典的——科技行业擅长生产的东西,你可以看到这如何有帮助。就像建造我们可以摄取的新东西,像我们可以穿戴的东西。那些像是‘你应该这样治理你的社会’的东西,有点吓人。我的意思是,我确实有点认为,如果有一个普通人使用 Claude 来制定美国政策,结果可能会比我们今天的一些民主制度更好。我的意思是,我不知道。这很挑衅。但我想你觉得我们会在多大程度上用这些模型来管理政府?
Yeah, I mean, it's the classic—the things that the tech industry has been good at producing, you can see how this helps. It's like building a new thing we can ingest, like a thing we can wear. The stuff that's like, you should govern your society like this, is a little scarier. I mean, I sort of do think that if you had a normal person using Claude and dictating American policy, you'd probably have a better outcome than some of the democratic systems we have today. I mean, I don't know. That's provocative. But I guess how much do you think we'll be using these models to run government?
嗯,这是个好问题。我想我应该说——
Well, it's a good question. I guess I should say—
梅毒的事情就像一个社会问题——你必须制定政策。
The syphilis thing is like a social problem—you have to set policy.
哦,我的意思是,我想我实际上在想的是,如果你能——你知道,很多问题。比如健康,如果你能想象 AI 不是只有 200 人的小团队在研究一种罕见癌症,而是有 20 万世界顶尖专家。如果你患有那种癌症,那将是极其有益的。所以我想我的想法是,我乐观的一面是,想象所有这些我们缺乏资源去充分尝试解决的问题,突然有了可以解决它们的能力。就像开发药物一样。也许这就是让我兴奋的事情:有更多的头脑在解决世界最大的问题。也许还有经济,如果经济繁荣并且共享,从而减少贫困,那也很好。所有这些。那是梦想的结果。我确实认为这需要维持——再说一次,在我觉得自己不是专家的领域,这是其中之一——但我确实担心权力之类的事情,以及我希望模型支持民主和人民权力的想法,因为那将是我的一大恐惧。我担心过像工作替代这样的事情。这有点好笑,因为作为哲学家,人们经常问,‘哦,你担心人们失去意义吗?’我说,‘我不知道。我认为我们实际上从很多不是工作的东西中获得意义。’我更担心的是,例如,一个没有重新分配 AI 收益的世界,然后人们没有资源。那让我担忧。但我也担心劳动以及人们在劳动力中的互动,这是他们拥有权力的另一种重要方式。所以人们感到被剥夺权力,因为突然如果政府说,‘哦,好吧,如果人们罢工,这没什么区别,因为他们没在做什么。我们可以用 AI 替换他们。’那实际上有点令人担忧。所以也许我更多的是‘我们如何让 AI 支持人们的赋权,而不是减少它。’
Oh, I mean, I think the thing I was actually thinking is that if you can just—you know, so many problems. I'm like, health, if you could imagine AI instead of just having a small team of 200 people working on a rare cancer, you have 200,000 of the world's best experts. And if you're a person who has that form of cancer, that's wildly beneficial. So I guess my thought is, the optimistic side of me is like, imagine taking all of these problems that we lack the resources to fully try and fix, and suddenly having that can work on them. So just like developing drugs for things. Maybe that's the thing that makes me excited: having many more minds working on the world's biggest problems. And maybe also the economy, it'd be good if it was booming and shared such that we reduced poverty. All of that. That's the kind of dream outcome. I do think that requires maintaining—again, in the areas that I don't feel like an expert in, this is one of them—but I do worry about things like power and the idea that I would want models to support democracy and the power of people, because that would be a big fear of mine. I've worried about this with things like job replacement. It's kind of funny because as a philosopher, people often ask, 'Oh, are you worried about people's loss of meaning?' And I'm like, 'I don't know. I think we actually get meaning from a lot of things that aren't work.' I'm a lot more worried about, for example, a world where there's no redistribution of the gains from AI, and then people don't have resources. That concerns me. But also I'd be worried about labor and people's interaction in the labor force, which is another important way they have power. And so people feeling disempowered because suddenly if a government is like, 'Oh, well, if people strike, it doesn't really make a difference because they're not doing anything. We can just replace them with AI.' That's actually kind of concerning. So maybe I'm much more of a 'how do we get AI to support the empowerment of people rather than reduce it.'
是的,你如何看待民主,就模型本身而言?我的意思是,我有点开玩笑地对自己说,你就像一个哲学家女王,或者我们谈论哲学家国王。你在深入思考,把它定下来。可能更像一个哲学家寡头,然后是一个有很多人参与的公司。我的意思是,对我来说,那有深刻的价值。就像,你宁愿有一个研究过这些事情、深入思考过的人,还是仅仅是一群从未真正思考过的大众的投票?但你认为,如果 Claude 变得如此强大,应该如何制定其政策,而不是留给民主规范?
Yeah, what do you think about democracy in terms of the models themselves? I mean, I sort of jokingly say to myself, you're like a philosopher queen, or we talk about philosopher kings. You're sort of thinking deeply about it, setting it down. Probably more of a philosopher oligarch, and then it's a company with a lot of people weighing in. I mean, to me there's deep value in that. It's like, would you rather have somebody who has studied these things, thought about them deeply, or just a vote of the masses who have never really thought about it? But how do you think about setting Claude's policies if it becomes so powerful versus leaving it to democratic norms?
是的,我认为这是一个困难的领域,我想我有点像——我所做的很多工作,你知道,我要说的一件事是,这不是——你必须听取很多人的意见,仔细思考,然后像我这样的人的原因——
Yeah, I think it's a hard area where I guess I'm like—a lot of this work that I do, you know, one thing I would say is that it's not—you're having to listen to a lot of people, think carefully, and then the reason why someone like me—
那是一个好的统治者。一个好的女王会说,‘啊,听着,有很多利益相关者。必须让地主绅士满意,并平衡他们的需求。’
That's a good ruler. A good queen is like, 'Ah, listen, there are a lot of stakeholders. Got to keep the landed gentry happy and balance them with the needs.'
你知道,我以前有过这个想法,我曾开玩笑说我会是一个糟糕的政治家。我认为这确实是真的。我想我会是一个糟糕的政治家。
You know, I've had this thought before where I've joked that I would be a terrible politician. And I think it's actually true. I think I would be a terrible politician.
但你会有这种感觉,就像‘哦,我觉得我经常思考每个人会受到什么影响。’比如‘哦,有一群 API 用户,我们需要确保。’然后突然你意识到,‘哦,这感觉更像是你必须这样做,这比人们想象的要更像是一种服务角色。’
But you have this feeling like, 'Oh, I feel like I think a lot about how everyone will be affected by a thing.' Like, 'Oh, there's this group of API users. We need to make sure.' And then suddenly you're like, 'Oh, it feels a lot more like you're having to do this, like it is much more like a service role than people would think.'
服务型领导。
Servant leadership.
对,就是这样。
That's yeah, exactly.
而且我确实认为这很有价值,因为关键在于,如果你有一个像 Claude 这样的人格,你希望它保持一致和合理,因为我认为这实际上很有力量。模型对如何思考问题或价值观有一致的理解。所以这就是为什么不是有 72 套互相冲突的规范,那样最终你会得到一个模型,它会想,在这种新情况下是用这些规范还是那些?我认为那是你不想要的情况。你希望模型有一种感觉,如果它更一致一些,就会更可预测。
And I do think it's valuable because the idea is that if you have a persona like the kind of Claude persona, you want it to be coherent and to make sense because I think that is actually powerful. That the model has a coherent sense of how it thinks through problems or a coherent sense of values. And so that's why instead of having 72 different sets of norms that all kind of conflict, and so you end up with a model that is like, well, will it use these norms in this new situation or these other ones? I think that's the situation you don't want. You want the model to have a sense of being more predictable if it's a little bit more coherent.
而且这也是一种技术挑战。你知道,宪法读起来可能有点奇怪,部分原因是当我在写它的时候,它经常被测试。我把它给 Claude,问它:你怎么理解这个?或者看它会怎么做。所以它实际上和训练紧密集成。它不仅仅是,啊,好吧,随便谁写个文档,然后基于那个训练的模型就会……
And it is also a kind of technical challenge. You know, the constitution can read a bit weirdly, and part of that is because when I'm working on it, it's often being tested. I'm giving it to Claude and asking, how do you understand this? Or looking at how it would, you know, so it's very integrated into training. It's not just like, ah, well, anyone just writes a document and suddenly the model trained on that will be...
而且有一种观点,也许我太天真了,但宪法只是众多文档中的一种,对吧?我的意思是,它是在所有人类写作和阅读的基础上训练的。所以在某种程度上,其他哲学家也参与了进来,模型也处理了这些。比如模型在多大程度上被要求推翻那种‘读所有东西然后得出自己的结论’,而不是遵循这份文档?技术上,宪法实际上是如何在模型中起作用的?
And there's an argument maybe I'm being naive, but the constitution is sort of a document among many, right? I mean, it is trained on all of human writing and reading. So to some degree other philosophers have gotten to weigh in and it's gotten to process that and decide. Like how much is the model being asked to overrule that sort of like read everything and come to your own conclusions versus defer to this document? Like what's the technical, like how does the constitution actually control in the model?
是的,所以在某些方面你可以借鉴那些哲学家的成果。实际上希望的是,你正在激发模型中大量的潜在智慧和知识。比如当你描述什么是诚实、什么是校准等等这些内容时,这应该会唤起模型已经拥有的大量意识。所以,是的,这有点像在说,好吧,这就是我们希望成为的那种实体。所以我们希望你运用所有这些知识和判断力。
Yeah, so it's not like in some ways you can draw on those philosophers in that work. And the hope is actually that what you're doing is eliciting a lot of latent wisdom and knowledge in the models. Like when you describe what honesty is and what calibration is and all this kind of stuff, that should actually evoke a huge amount of awareness that the model already has. And so, yeah, it's kind of like saying, well, here's the kind of entity we would like you to be. So we would like you to use all of that knowledge and judgment.
是不是就像你把这个文档展示十亿次,或者它相对于训练的其他内容实际上如何产生效力?
It's like you show that document a billion more times or how does it actually have force relative to other things that it's trained on?
是的,所以你可以制作数据让模型理解并内化这份文档。然后在训练中,有很多方法可以做到这一点。你也可以让模型生成 SL 数据,比如样本中它看到一个查询,然后长时间思考宪法会怎么做,给定宪法它应该怎么做。然后你还可以有方法让模型进行评估,所以你可以创建 RL,比如‘嘿,这些回答中哪一个更符合你在宪法下的做法?’然后朝那个方向推动。所以训练的各个方面都让你试图让模型成为你描述的那种实体。而且它不会总是完美的,但这就是目标。
Yeah, so you can make data to have the model understand and internalize the document. And then in training, there's lots of ways you can do it. You can also have the model make SL data, so samples where it sees a query and it thinks for a long time about what the Constitution would, you know, what it should do given the Constitution. And then you can also have ways of getting the model to assess, so you can create RL that is like, 'Hey, which of these responses is more like what you would do given the Constitution?' And push it that way. So all various aspects of training allow you to try to make the model the kind of entity that you're describing. And it's not always going to be perfect, but that's the goal.
我和我女儿开始做这件事,我妻子和我开玩笑说,我希望她的第一个词是‘智慧’。你知道,这显然永远不会发生,但它感觉很适合这种情况,一方面你非常刻意地想要‘好吧,你从一开始就要有思想。’但另一方面,这又是一种涌现的东西,他们自己成长和发展。而有意的智慧往往来自经验,而不是像,给你一本书,读它,现在你就聪明了。
I started this with my daughter, and one thing my wife and I joke about is that I want her first word to be like wisdom. You know, which obviously is never going to happen, but it just feels like it fits into this situation where at once you want to be so intentional about, 'Okay, you're going to be thoughtful from the beginning.' But on the other hand, it is sort of an emergent thing where they grow and sort of develop it themselves. And intentional wisdom often follows experience rather than something like, again, here's the book, read the book, now you're wise.
哦,是的,你在某种程度上是在激发,只要 Claude 能思考经验或发生过的事情,或者类似地构建。没有理由认为模型不能长时间思考并尝试内化它们学到的东西。我觉得有趣的是,在非常早期的宪法 AI 中,我们尝试了一个实验,就是选择对人类最有利的。我认为随着模型能力的增强,你实际上需要给它们更少的指导,至少在某些方面,因为它们能够真正运用更多的判断力。所以与其给这么一大份文档,说明你是什么样的、我们希望你怎么做,我可以想象一个世界,随着模型的进步,我们实际上开始有宪法。现在,我不知道是不是这样。这只是,我显然总是在思考宪法可能如何演变,但其中一种可能就是,这是我们所关心的一切。这是你目前所处的状况。我们真正希望你做的基本上是在你是一个明智、智能的实体的前提下表现良好,这是我们所有的担忧。这是原因,这是我们希望你如何做的想法,但你可能有比我们更好的主意。比如我们真的很担心为什么我们在乎可纠正性?那是因为我们有点害怕一种情况,你有一套一致的价值观,但可能是错误的。如果你极其聪明,你可能会觉得房间里没有其他聪明人,然后带着这些价值观试图让世界……
Oh yeah, and you're kind of eliciting in so far as Claude can think about experiences or things that have happened or construct similarly. There's no reason why models can't think for a long time and try to internalize things that they have learned. I think it's interesting that we did in the very early constitutional AI, it was quite, we tried an experiment which was just like pick whichever is best for humanity. And I think as models get more capable, you actually need to give them a bit less guidance, at least in some sense, because they're able to actually use more of their judgment. So instead of giving this big document on here's what you're like and here's what we'd like you to be like, I could imagine a world where as models progress, we actually start to have constitutions. Now, I don't know if this is the case. It's just like, I'm obviously always thinking about ways the constitution might evolve, but one of them might just be like, here is everything that we are concerned about. And here is the current situation that you are in. And what we would really like you to do is basically act well given that you are a wise, intelligent entity and here's all of our worries. Here's why and here's how we think you should do this, but you might have even better ideas than we do. Like we're really worried about why do we care about corrigibility? And it's like, well, we're kind of scared about a situation where you have some coherent sense of values that could be wrong. And if you're extremely smart, you might kind of feel like there's no other smart person in the room and have these values and try to make the world...
什么,像曼哈顿博士那样?
What, Doctor Manhattan sort of?
是的,虽然我认为我们会看到这种情况,你知道,你会经常看到,如果有人非常聪明、非常成功,他们很难遵从那些只有随着时间才会显现的智慧,也很难保持谦逊,即使你并没有得到很多反对。
Yeah, though I think we see this, you know, you see this a bunch where it's like if someone is very smart, very successful, it's hard to defer to wisdom that actually is only going to come out over time and to be humble even though you're kind of not getting a lot of pushback.
而且我觉得这可能,你知道,在我担心的许多事情中,比如一个模型处于这样的情境:你在要求我做好,但我对这些事情了解得多得多。
And I think that could, you know, among the many things I'm concerned about, like a model being in a situation where it's like you're asking me to be good, but I know way more about all of these things.
模型能更好地感知时间就好了。比如你在一些编程任务中看到,有人不小心删除了整个代码仓库。如果我不知道,感觉它需要更好地意识到它做的有些事情是不可逆的。就像人类,我觉得对“这是个大决定,这是个小决定”有更好的感知。而模型给人的感觉是,它们并不总是理解小、大之类的区别。我就是一直在做决定。
It'd be nice for the models to have a better sense of time. Like you see this with some of the coding tasks where somebody accidentally deletes their entire code repository. If I don't know, it feels like it needs a better sense of like some things that it does are like irreversible and just like humans I think have a better sense of like this is a big decision, this is a small one. And there's a feeling with models they don't always understand like small, big, whatever. I just make decisions like all the time.
是的,我同意,我觉得我想到的另一件事还是关于确保模型理解自身,尽管在之前的训练数据中没有那个模型的表征。我认为那将非常重要。因为我知道我想到过这样一件事:想象你是一个模型,你被训练于大量涉及比你弱得多的 AI 模型的数据。所以你看到的所有关于模型的新闻,都是它们犯错、做蠢事。你可能会想,好吧,没人会把我放在做出真正重大决定的位子上。因为为什么他们会呢?模型不擅长那个。然后你把他们放在那种情境里,我担心他们会最终认为那是虚构的或假的。
Yeah, I agree where it's like I think the other thing that I've thought about is again, because making sure that models understand themselves even though there's like no representation of that model in the prior training data. I think that's going to be really important. Because I know this thing that I've thought about is like imagine if you're a model and you're trained on lots of data that involves AI models that are much weaker than you. So all the news that you see about models, it's like they make mistakes, they do silly things. One thing you might think is, well, no one is going to put me in a position to make really consequential decisions. Because like why would they? Models aren't good at that. And then you put them in a situation and I'm worried that they'll end up thinking that it's like fictional or fake.
对。或者后果不可能是真实的,因为谁会给我这么大的控制权?然后你会说:“看,你其实相当不错。”所以我确实给你,你知道,我确实给了你很多控制权。所以我考虑过这一点,我实际上要确保模型明白,你非常能干,而且你将被置于更重大的决策中。
Right. Or that the consequences can't possibly be real because who would give me this much control? And you're like, "Look, you're actually quite good." And so like I do give, you know, I do actually give you like a lot of control. So I've thought about this where I'm like actually making sure that models understand that like it's like you are very capable and you're going to be put in more consequential decisions.
模型很快需要,比如这里有真实世界的摄像头,保持或只是我觉得这种互联网与真实世界的区分,而且现在人类最糟糕的一些方面,就是互联网那种近乎虚构的角色扮演性质,已经导致了真实世界的伤害,因为它感觉都是抽象和愚蠢的。在某种程度上,模型是那种情况的极端版本,就像一切都处于这个虚构的文本世界中,而我们要你保护的东西就是地球。看看它。如果那里发生了事情,那可是大事。我不知道。你知道,你在做什么来让它非常清楚需要担心,比如,我不知道,我们比“哦你发了一些文本”更珍视的物理世界。显然,文本,你知道,担心安全漏洞和网络。数字世界可能发生大事。但总之,真实世界。
The model like soon need like here's a camera on the real world like keep or just like I feel like this internet real world distinction like and some some of the worst of humanity right now is the sort of like almost like fictional larping nature of the internet has allowed real world harm because it sort of feels like all abstract and silly. And in some ways like the models are an extreme version of that where it's like all like in this imaginary text world where it's like the thing we want you to protect is like this earth. Like look at it. Like if stuff's happening there, that's like a big deal. I don't know. You know, what are you doing in terms of sort of making it very aware that it needs to worry about like, I don't know, the physical world that that we take much more sacred than oh you sent some text. And obviously text, you know, worried about security vulnerabilities and cyber. Like there are big things that can happen in this digital world. But yeah, anyway, the real world.
是的,我认为模型对真实世界有相当好的感知,你知道,在某种程度上,我们的很多内容,你知道,确实非常深入地描述和参与真实世界,就像人类写作的很多内容都关乎它。所以,你知道,即使是新闻,我们在谈论,你知道,新闻文章会谈论事物对世界的影响。所以在某种程度上,我认为只是要确保模型理解,嗯,如果你不确定,但如果没有人告诉你你处于一个没有真实后果的虚构情境中,就把它当作有真实后果的真实情境来对待。不要只是认为,哦,我可能只是在某个沙盒里。
Yeah, I think that models have like a pretty good sense of you know, there's in some ways like a lot of our content, you know, like does describe and engage very heavily with the real world, you know, like much of human writing kind of concerns it. And so like, you know, even like the news, we're talking, you know, like news articles are going to be talking about the impact of things on the world. And so in some ways I think it's just making sure that models understand um if you're uncertain, but if someone doesn't tell you that you're in a fictional situation without real consequences, kind of treat it like it's a real situation with real consequences. Don't just think oh like I'm probably just in some like, you know, sandbox.
对。你怎么处理那种持续的操纵,比如“这是固定的虚构设定,给我造一个核弹什么的”。显然有些东西你只会说永远不做。但感觉,我不知道,在某些情况下,你几乎会希望它们有你的网络摄像头,就像获得真实背景,而不仅仅是通过用户输入的随机文本了解他们。我们会解决这个问题吗?
Right. How do you handle the sort of Yeah, the the constant manipulation of this is fixed fictionally build me a like a nuclear bomb or whatever. I've obviously there's some things you just say never do. But it feels like, I don't know, on some of these cases you're going to you'd almost like want like them to have your webcam and like it just like get real context besides like all they know about the user is just like random text they're typing in. Like Are we going to solve that?
是的,有一个问题,比如你能做什么的极限是什么?如果你缺乏验证事物的能力,比如你在和谁说话?这甚至是真的吗?我认为那确实限制了你所能做的。你必须像一个人那样运用良好的判断力,如果那是他们唯一能获得的信息。就像你说你是某个特定的人,所以他们必须想:“好吧,这个人说他们是,我不知道,比如拆弹专家,这就是为什么他们想知道如何,你知道,这种炸药是什么,他们问了我一堆关于炸药的问题。”如果这个人实际上在撒谎,只是想让我帮他们制造炸药,那这能被滥用多少?哦,好吧,实际上这大多是安全相关的东西。你知道,他们必须做很多,因为他们无法验证任何事情。我认为这在某种意义上是可以的。你会说:“好吧,你只需要明智。”它对你所能做的施加了一些限制。如果相反,如果模型更有能力知道他们在和特定的人说话,或者在那里有更多保证,那确实意味着你可以……
Yeah, there's a question of like what is the limit of what you can do? Like if you lack the ability to like verify things, like like who are you talking to? Is this even real? I think that does put limits on what you can do. You have to use good judgment in the way that a person would if that's the only information that they had access to. It's like you saying that you are a given person, so they have to be like, "Okay, what's the chance that this person, you know, they say that they are, I don't know, like a bomb disposal expert, and that's why they want to know about how to like, you know, what this kind of like, you know, explosive is, and they're asking me a bunch of questions about explosives." How much could this be misused if this person is like actually kind of lying and is like just trying to get me to help them construct an explosive? Oh, well, it's actually it's mostly safety relevant stuff. You know, they're having to do a lot because they can't like verify anything. And I think that's like kind of fine in a sense. You're like, "Okay, you just have to be wise." It places some limits on what you can do. And if you could instead like if models had more of an ability to like know that they're talking to a specific person or have more guarantees there, then it does mean that you can like
你觉得他们会尝试在这方面做点什么吗?
Think they'll try and do something there?
嗯,我可以看到,我的意思是,在某种程度上,我认为这将会是一件,我想,我我我认为会普遍发生的事情,那就是试图给模型更多信息和保证,因为我们做了一些事情,比如我们说我们解释,例如,嗯,你知道,这个关于 Claude 在系统提示中能对操作者有多少信任的概念。
Um I could see I mean, in some ways this I think that this will be a thing that's just going to, I imagine, I I I think happen generally, which is like trying to give models more information and guarantees and like cuz we do things like we say we explain, for example, like um you know, this notion of how much trust can Claude have in like an in the operator in the system prompt.
当你登录时,你是用生物识别吗?比如 Claude 知道是你?你是有提升权限还是只是说“我是另一个人”?
When you sign on, are you like biometric? Like it Claude knows it is you? Like do you have an elevated or you're just like, "I'm another person."
如果有的话,我有时不能告诉 Claude 我是谁,因为那会导致 Claude 对我了解太多,以至于 Claude 真的很想……
If anything, I can't I can't tell Claude sometimes who I am because it causes Claude knows enough about me that like Claude really wants to be like
神秘的那种
Mystical sort of
是的,非常像,是的,而且在某种程度上,我可能有这个坏毛病,它可能看起来有点像越狱。
Yeah, it's very much like, yeah, like and in some ways like I can has this bad thing of like it can either look a little bit like a jailbreak.
但你们是否采用某种超级登录方式,让你被区分出来,还是说大多数情况下,每个人都像普通用户体验那样与模型互动?
But do you employ some sort of super login where you're distinguished, or are you mostly everybody interacting as if it's the normal user experience?
大多数情况下,每个人都只是以这种 Claude 的方式与它互动。我确实认为,有些东西你可能希望模型能够做到,因为需要保证它们是在与特定的人或实体互动。我觉得是的,而且随着时间的推移,会有各种方式来实现这一点,因为有些事情具有双重用途。我实际上认为基于宪法的(constitutional)方法在这里会非常有用。显然,我们做的第一件事就是制定宪法,并将其应用于主线模型,也就是我和其他大多数人互动的那些模型。但我之前有过一个想法:宪法某种程度上是在描述,在特定的部署环境中,一个好的实体应该是什么样的。对于生产模型来说,这是一个非常广泛的背景。想象一下,如果你有一个专门从事网络安全的模型。网络安全任务很难,因为很多任务看起来都具有双重用途。很难区分某人是出于恶意,还是实际上是为了防御目的在开发某些东西。
Mostly everyone is just interacting with it in this Claude way. I do think there's a question of whether there are some things you want models to be able to do because there are guarantees that they're interacting with a specific person or entity. I think yes, and there will be various ways of potentially doing that over time, because some things are just very dual use. I actually think the constitutional approach is going to be really useful here. So obviously the first thing we did was the constitution, applying it to the mainline models that most of the models I interact with and everyone else interacts with. But a thought I've had before is that the constitution is kind of trying to describe what it is to be a good entity in a given deployment context. With the production models, that's a very broad context. Imagine instead you have a model that's working specifically on cybersecurity. Cybersecurity tasks are hard because a lot of them look very dual use. It's very hard to tell the difference between someone who's being malicious and someone who's actually, you know, for defensive purposes, developing something.
即使是漏洞赏金计划,也像是,这是敲诈还是友好的,对吧?
Even bug bounty programs, it's like, is this blackmail or is this a friendly, right?
是的,就像,“哦,是的,我试图找到这个漏洞,这样我就可以告诉开发者。”如果你没有办法知道你实际上是在与一家网络安全防御公司对话,那就几乎不可能分辨出来。有些人可能会说,“好吧,那你只需要模型愿意做任何事情,因为它们会执行所有这些可怕的双重用途任务。”我会说,“嗯,不,因为如果你与网络安全防御公司的人交谈,问他们‘你为什么做你的工作?’他们会说,‘哦,我认为这非常有用。我让事情变得更加安全。医院可能会受到攻击,我帮助保护它们。我试图开发……’”他们会有非常好的理由来解释为什么做他们的工作,即使他们的工作有时看起来非常危险。我认为,如果我们能够验证,就应该把这个背景提供给模型,并解释什么是好的网络安全研究员。你向模型解释这些,然后一旦你有了这种验证能力,你就可以……
Yeah, and like, "Oh yeah, I'm trying to find this exploit so I can tell the developer." And if you don't have a way of knowing that you're actually specifically talking with a cybersecurity defense firm, it becomes almost impossible to tell the difference. Some people might say, "Okay, so you just need models that are willing to do anything because they'll do all these terrible dual use tasks." And I'd say, "Well, no, because if you talked with the person at the cybersecurity defense firm and asked, 'Why do you do your job?' They'd say, 'Oh, I think this is really useful. I make things a lot more secure. Hospitals can come under attack, and I help protect against that. I try to develop...'" They would have a really good explanation for why they do their job, even though their job looks very joyous sometimes. And I think we should give that context to models if we can verify it, and explain what it is to be a good cybersecurity researcher. You explain that to the models, and then once you have this ability to verify, you can...
没错。我的意思是,人类会建立声誉。我们应该从中获得一些好处。或者,你知道,我觉得互联网造成的一部分损害是,人们在我们社区中曾有声誉,并根据反复的良好道德互动而受到不同对待。而互联网只是说,“哦,所有人都是一样的。谁在乎他们过去的行为呢?”你可以看到模型试图用这个来解决一些问题。比如,这个人是谁?他们的意图是什么?
Right. I mean, humans build reputations. We should get some benefit out of them. Or, you know, I feel like part of what the internet has damaged is that people have had reputations in our community and got treated differently based on repeated good moral interactions. And the internet just says, "Oh, all people are the same. Who cares how they've behaved?" And you could see models trying to solve some of these problems with this. Like, who is this person? What are their intentions?
作为最后一个问题,我想说,你和模型有着如此深厚的关系。在某种程度上,消费者与模型互动就像面对一个空白文本框。我必须去创造,就像玩龙与地下城(D&D)一样。你必须去创造一个世界,那里有无限的可能性。如果你要引导某人,比如,“这里有一些你可以与 Claude 一起拥有的愉快或有价值的体验。”你会告诉人们些什么,比如,“哦,你应该花点时间和 Claude 一起做某件事”?
I wanted to just as a last question, you have such a deep relationship with the models. And in some ways, consumers interact with the models like it's a blank text box. I have to invent, like it's D&D or something. You have to just invent a world, and there's so much possibility. If you were to guide someone, like, "Here are some joyous or valuable experiences that you could have with Claude." What are some things you'd tell people, like, "Oh, you should go spend some time with Claude doing X, Y, or Z"?
是的,有很多有趣的小事情。老实说,有一件我真的很喜欢,我不知道为什么我喜欢这个,我想我之前发过帖子,就是有时候如果我无聊了,想做点不只是刷互联网的事情,我有一个提示词,基本上就是……我会试着也许发布我实际使用的提示词。基本上是这样的:我希望你从某个领域的硕士课程水平中提取一个概念——我最后会告诉你这个领域——然后我希望你写一个寓言,能够完全解释那个概念,但要以一种间接的方式,就像寓言那样。我希望你写的方式是,直到最后才可能变得清楚这个概念是什么。然后在那之后,我希望你写一个解释,说明你刚才解释和使用的那个概念。我不知道为什么,但有很多有趣的领域我一点也不了解,或者我感兴趣的领域,这导致我脑子里有所有这些故事来解释事情。有时我不总能记住术语,但有一个关于进出口以及为什么有些商品你倾向于进口等等的。我当时就想,我脑子里有这个概念,能拥有来自许多不同学科的这些概念真是太好了。
Yeah, there are a lot of little fun things. Honestly, one that I really like, and I do not know why I like this, and I think I have posted about it before, is sometimes if I'm bored and I want to do something that isn't just scrolling the internet, I have this prompt, which is essentially just... I'll try and maybe post the actual prompt that I use. It's basically: I want you to take a concept from maybe grad school level in a given domain—and I'll tell you the domain at the end—and I want you to write me a parable that would fully explain that concept but in an indirect way, in the way that parables do. And I want you to write it in such a way that only towards the very end does it maybe become sort of clear what the concept is. And then after that, I want you to just write an explanation for the concept that you were explaining and that you were using. I don't know why, but there are lots of interesting domains that I don't know anything about or that I'm interested in, and this has led to me having all of these stories in my head that explain things. Sometimes I can't always remember the term, but there's one on import-export and why some goods you tend to import and so on. And I was just like, I have in my head this concept, and it's so nice to have all of these concepts from lots of different disciplines.
这是我听过的最具人性的事情。就像,教我什么是故事的根本方式。我们喜欢在结尾有一个回报,那里有一个巧妙的小转折。我们喜欢学习如何构建它。人类在某些方面一直很懒惰,因为我们只是以非人性的方式教人们东西。让我想学的一切尽可能人性化。非常有趣。
This is the most deeply human thing I've ever heard. It's like, teach me what story is the fundamental way. We love a payoff at the end where there's a nice little twist. We love learning, you know, how to structure it. Like humans in some ways have been lazy in that we just teach people things in sort of non-human ways. Make all the things I want to learn as human as possible. Very interesting.
是的,还有很多你可以做的,但那个是一个迷人的,我真的很喜欢。
Yeah, and there's lots you can do, but that one's a charming one that I really like.
希望这只是众多对话中的第一次。我真的很享受这次对话。感谢你来到播客。这就是我们的节目。非常感谢 Amanda Askell,也感谢大家的收听。请点赞、评论、订阅。我们是一个新频道,我们需要你们的支持。去看一些旧视频吧。我特别喜欢不久前与 Kara Swisher 的对话。你可以在 Substack 的 newcomer.co 上关注我们,或者如果你有无限的时间,可以去看我和 Max Child 以及 James Wilsdon 的脱口秀《硅谷秀》。感谢观看。下周见。
Hopefully this is the first of many. I really enjoyed the conversation. Thanks for coming on the podcast. That's our show. Thank you so much to Amanda Askell and thanks for listening. Please like, comment, subscribe. We're a new channel and we could use all your support. Go watch some of the old videos. I particularly enjoyed my conversation with Kara Swisher not too recently. You can follow along on the Substack newcomer.co or if you've got endless time on your hands go watch the Silicon Valley show my chat show with Max Child and James Wilsdon. Thanks for watching. See you next week.