Maximizing Luck Surface Area: Neil Nando on Building a Career in AI
打开互动全文版(中英对照 + 朗读 + 问答)→尼尔·南多分享如何在正确的时间出现在正确的地点,并对机会说“是”,从而在 26 岁时领导谷歌 DeepMind 团队。
Neil Nando shares how being in the right place at the right time and saying yes to opportunities helped him lead a team at Google DeepMind at age 26.
我学到的最重要一课是,你只管去做就行。这在一定程度上是我所说的最大化你的运气表面积。你要尽可能多地创造好机会降临的可能性。你要做一个有时会说“好”的人,这样人们才会把机会带给你。我最终意外地领导了 DeepMind 团队。我加入 DeepMind 时,本打算做一名独立研究员。没想到,我加入几个月后,负责人决定辞职,之后几个月里,我逐渐接替了他的位置。我当时不知道能否胜任。现在看来,一切还算顺利。对我来说,这既说明了拥有运气表面积的重要性——要让自己处于能遇到这类机会的境地,也说明你应该对事情说“好”,即使它们看起来可怕,你也不确定能否做好,只要风险不大就行。
One of the most important lessons I've learned is that you can just do things. Part of this is what I think of as maximizing your luck surface area. You want to have as many opportunities as possible for good opportunities to come your way. You want to be someone who sometimes says yes so people bring things to you. I ended up running the DeepMind team by accident. I joined DeepMind expecting to be an individual researcher. Unexpectedly, the lead decided to step down a few months after I joined, and in the months since, I ended up stepping into their place. I did not know if I was going to be good at this. I think it's gone reasonably well. To me, this is both an example of the importance of having luck surface area, being in a situation where opportunities like that can arise, but also you should just say yes to things even if they seem scary, you're not confident they'll go well, so long as the downside is pretty low.
AI 正在重塑世界。这是当下最重要的事情之一。我简直无法想象,自己会不想站在这个领域的前沿,帮助它变得更好。今天,我有幸请到了传奇人物 Neel Nanda,他领导着 Google DeepMind 的机制可解释性团队。机制可解释性是一个研究项目,旨在弄清楚 AI 模型为什么做它们所做的事,它们是如何做到的,以及为什么选择做一件事而不是另一件事。直到几年前,我们对这些还知之甚少。但 Neel,你是少数几个帮助这个研究项目从四年前一个很小的努力,发展到如今相当有规模、有数百人贡献答案的人之一。非常感谢你再次来到播客。
AI is reshaping the world. This is one of the most important things happening. I just can't really imagine wanting a career where I'm not at the forefront of this, helping it go better. Today I have the pleasure of speaking with the legendary Neel Nanda, who runs the mechanistic interpretability team at Google DeepMind. Mechanistic interpretability is a research project to try to figure out why AI models do what they do, how they're doing it, and why they're choosing to do one thing over another. Until a couple of years ago, we had fairly little insight into that. But Neel, you've been one of a handful of people who've helped grow this entire research project from a quite small effort four years ago to now something quite meaningful, with hundreds of people contributing to answering those kinds of questions. So thanks so much for coming back on the podcast.
是的,我很高兴来到这里。
Yeah, I'm really excited to be here.
人们可能不会立刻意识到你只有 26 岁。尽管你只参与了机制可解释性研究大约四年,但你已经为该领域贡献了一些基础性成果。你在 Anthropic 工作过,现在领导着 Google DeepMind 的一个团队。你发表了许多论文,获得了大量引用。过去四年里,你还指导了大约 50 人,其中 7 人现在在前沿 AI 公司从事重要工作,还有几人担任重要的政府或监管职务。所以可以说,过去四年你非常忙碌且高产。你是如何在职业生涯早期就取得如此多成就的?
One thing people might not immediately realize about you is that you are just 26. And despite having only been involved in mechanistic interpretability research for four years or so, you've contributed some foundational results to the field. You've worked at Anthropic and now lead a team at Google DeepMind. You've published many papers and gotten plenty of citations. You've also mentored about 50 people over the last four years, and seven of them are now doing important work at frontier AI companies, and several are in important government or regulatory roles. So it's fair to say you've had a busy and very productive last four years. How have you managed to accomplish so much so early in your career?
嗯,我很想说是自己聪明绝顶。我觉得自己做得还行,但主要是运气——这类问题其实应该更常得到这个答案。稍微展开一点:我认为这是天时地利人和,以及善于抓住机会、为自己创造合适机会的结合。我进入机制可解释性领域恰逢其时。当时这个领域很令人兴奋,但规模很小。如果你进入一个快速发展的领域,比如创业公司或研究领域,你可以极快地成为最有经验的人之一,而这实际上与你有多优秀关系不大。我们稍后可能会讨论一些关于如何为自己创造合适机会的话题。另一件大事是,我真的觉得管理和指导他人很有趣。很多研究人员很不情愿做这件事,但我似乎很擅长。通过让 10 到 20 个其他研究人员变得更好,我可以产生更大的影响。当你做得足够多,效果就会累积。最重要的是,我没有浪费五年时间读博士。
Well, I'd love to say this is just because I'm incredibly smart and brilliant. I think I'm decent at what I do, but mostly luck, which really should be a much more common answer to this kind of question. Maybe fleshing that out slightly: I think it's a combination of being in the right place at the right time and being good at taking opportunities as they arise and making the right opportunities for myself. I got into mechanistic interpretability at the right time. There was a lot of excitement but it was tiny. If you get into a fast-growing thing like a startup or a research field, you can become one of the most experienced people extremely fast in ways that are not actually that related to how good you are. We'll discuss later maybe some of the things around being good at making the right opportunities for yourself. The other big thing is I just find managing and supervising people really fun. Lots of researchers grudgingly do this, but I seem pretty good at it. I can have so much more impact by making 10 to 20 other researchers better. When you do this enough, it adds up. And the most important thing was not wasting five years of my life in a PhD.
是的,我忘了提你其实只有本科学位。所以你可能省了好几年。你提到了冷邮件。我想你收到不少。什么样的冷邮件能真正引起你的注意并让你回复?
Yeah, I forgot to mention that you actually only have an undergraduate degree. So you've saved many years there potentially. You mentioned cold emails. I imagine you get a fair few. What kinds of cold emails actually get your attention and get you to engage?
我给冷邮件的主要建议是:假设读邮件的人很忙,并且会在不确定的地方停止阅读。因此,你需要尽可能简洁,并优先呈现关键信息。理想情况下,加粗关键词或短语。就我个人而言,收到邮件时,我通常会先看一眼。如果我能在一分钟内回复,回复的可能性就很大。如果需要更多精力,我的门槛就很高。所以问问自己:我想要的东西是否能在几分钟内完成,比如一个具体问题?如果是,就把它放在最显眼的位置。否则,更多是关于吸引注意力。一些好的做法:一个非常有效的方法是展示能力。一种方式是有令人印象深刻的成就或资历。我知道人们有时不愿意自夸,但老实说,自夸很有帮助。我希望人们直接说,‘以下是我最令人印象深刻的地方,方便你优先考虑。’显然这很嘈杂且烦人,但鉴于我收到的邮件多于我能回复的时间,我需要快速排序。一个既能展示自己又不显得讨厌的技巧是说:‘我知道你肯定收到很多这样的邮件,所以为了帮你优先处理,这里是一些关于我的关键信息,等等。’我热烈鼓励任何给我发邮件的人这样做。我也热烈鼓励任何考虑给我发邮件的人,不如去给别人发邮件,因为一个常见错误是联系某个领域最突出的人,比如团队负责人或知名教授。但他们越突出,收到的邮件就越多,越忙,回复你的可能性就越小。对于人们可能想要的东西——技术问题、指导、建议——一个更初级的人完全有能力给出建议,尤其是如果你想要像指导这样耗时的事情。
The main advice I have for cold emails is: assume the person reading this is busy and will stop reading at an uncertain point. Therefore, you need to be as concise as possible and prioritize key information. Ideally, bold key words or phrases. Personally, when I get an email, I'll typically look at it a bit. If I can respond within a minute, decent chance I will. If it requires more effort, I have a pretty high bar. So ask yourself: is the thing I want something that could be done in a minute, like a specific question? If so, make that extremely prominent. Otherwise, it's more about catching someone's attention. A few things that are good: one thing that can be very impactful is signaling competence. One way is having done impressive things or credentials. I know people are sometimes unwilling to boast, but honestly it's really helpful if they boast. I want people to just say, 'Here are the most impressive things about me so you can prioritize.' Obviously this is noisy and annoying, but given that I get more emails than I have time to respond to, I need to prioritize fast. One trick for talking about how you're impressive without seeming like a jerk is to say, 'I'm sure you must get many of these emails, so to help you prioritize, here's some key info about me, blah blah blah.' I enthusiastically encourage anyone sending me an email to do that. I also enthusiastically encourage anyone considering sending me an email to send someone else an email instead, because a common mistake is reaching out to the most salient person in an area, like the lead of a team or a prominent professor. But the more prominent they are, the more emails they get, the busier they are, and the less likely they are to respond. For many things people might want—technical questions, mentorship, advice—a more junior person is pretty well placed to give that advice, especially if you want something time-consuming like mentorship.
我认为初级人员通常能够给领域新人提供相当好的指导。比如,加入我团队不到六个月的人,更可能有空闲时间,但仍能给出很多有用的建议。一个推论是,你应该多给论文的第一作者发邮件,而不是给作者列表末尾的知名学者,因为他们更可能有时间。关于简洁性,有个技巧:有时我会让人写一个详细文档,然后给我一段摘要,再附上文档链接,这样我就不会看到一封 3000 字的邮件而懒得读。但如果我感兴趣,就可以点进去看。我也很喜欢用要点列表——容易浏览,清楚表达需求。我还要说,我对任何发邮件寻求指导的人的回复是,我指导人的主要方式是通过目前开放的数学项目申请。我真该在网站上显眼位置写上这个,我已经对太多人说过同样的话了。我也不认为资历是展示能力的唯一方式。通常,做一些有趣的事——比如为有意义的开源库做贡献,或以某种方式帮助完成一篇论文——比属于某个实验室更让我感兴趣。其次,你可以通过提出好问题来展示能力。如果你询问研究者的工作,可以深入很多层面。如果我看到一个问题,心想‘哦,这不是我常被问到的,但确实很合理’,或者有人说‘我想这样扩展这篇论文,是个好主意吗?’而它确实是个好主意,我有时间的话很乐意和他们聊聊。
I think junior people are often very able to give quite good mentorship and quite useful mentorship to someone new to the field. For example, someone who joined my team in the last six months is a lot more likely to have free time but still has a lot of useful advice to give. One consequence of this is you should try to email first authors of papers a lot more than the fancy academic at the end of the author list because they're just more likely to have time. One trick on conciseness: sometimes I've had people write a doc with a bunch of detail and then give me a one-paragraph blurb and link to the doc if I want to read more. That way, I don't look at it and think, 'This is a 3,000-word email, I'm not going to read this.' But if I'm interested, I can click through. I'm also a big fan of bullet points—just be easy to skim and clear about what you want. I will also say that my response to anyone who emails me asking for mentorship is that the main way I mentor people is through the maths program applications currently open. I should really write that prominently on my website somewhere; I've said that reply to too many people. I also don't think credentials are the only way you can signal impressiveness. Often, having done something interesting—like contributing to a meaningful open-source library or helping out with a paper in a certain way—is more interesting to me than being part of lab X. Secondly, you can signal competence by asking good questions. There's a lot of depth you can go into if you're asking about a researcher's work. If I look at something and think, 'Oh, that's not a question I normally get asked, but it's actually very reasonable,' or if someone says, 'I want to extend the paper like this. Is that a good idea?' and it's a good idea, I'm down to see them if I have time.
是的,我觉得我没你那么厉害,收到的陌生邮件也没你多,但我几乎完全赞同这些。你认为初级人员在实践中能在多大程度上利用 LLM 或 AI 来扩大规模,更接近知识前沿或研究能力,比四年前这些工具还没那么有用时更快?
Yeah, I think I'm not nearly as big a deal as you, and I don't think I get as many cold emails as you do, but I can endorse almost all of that. To what extent do you think junior people can in practice use LLMs or use AI to scale up and get closer to the frontier of knowledge or research ability more quickly today than they could four years ago when these tools weren't really that useful?
哦,太多了。我认为现在如果你想进入一个领域而不使用 LLM,那你就错了。这不是说盲目地用 LLM 做所有事,而是把它们当作工具——了解它们的优势和劣势,以及它们能帮上忙的地方。我认为这在过去六到十二个月里变化很大。我以前在日常生活中不怎么用 LLM,但几个月前,我开始努力成为 LLM 的高级用户。现在我会在对话中随口说,‘啊,你有没有考虑过用 LLM 这样解决你的问题?’那么,人们应该怎么看待这个呢?我认为 LLM 非常擅长的一件事是降低进入一个领域的门槛。它们在实现领域内的专家级表现上很差,但在初级水平上相当不错。它们不一定完全可靠,但初级人员也不可靠。这是个错误的衡量标准。有几件事我觉得人们不一定做对了。有些对于擅长使用 LLM 的人来说可能显而易见,但请耐心听。首先,系统提示非常有用。你可以给模型非常详细的指令,说明你希望它帮助完成的任务类型,这会让它做得更好。我很喜欢一些主要提供商的项目功能,你可以在里面放一堆不同的提示和可能有用的上下文文档,然后在那里对话。关于提示的几个技巧:如果你觉得自己不擅长写提示,就让 LLM 帮你写提示。我个人觉得很有用的一点是,我发现用语音听写时更容易思考和头脑风暴。LLM 非常擅长处理我杂乱的口述,把它当作连贯的输出。所以我会对着它漫无边际地说‘任务是什么?我的标准是什么?我不想要的失败模式是什么?’然后它就会帮我写提示。如果出了问题,你可以给 LLM 详细反馈你不喜欢的地方,让它根据这些反馈和它做错的具体例子重写提示,然后下次复制进去。
Oh, so much. I think if you're trying to get into a field nowadays and you're not using LLMs, you're making a mistake. This doesn't mean use LLMs blindly for everything, but understand them as a tool—their strengths and weaknesses, and where they can help. I think this has actually changed quite a lot over the last six to twelve months. I used to not really use LLMs much in my day-to-day life, and then a few months ago, I started a quest to become an LLM power user. Now I'll randomly work into conversation, 'Ah, have you considered using an LLM like this for your problem?' So, how should people think about this? I think one of the things that LLMs are actually very good at is lowering barriers to entry to a field. They're quite bad at achieving expert performance in a domain, but they're pretty good at junior-level performance. They aren't necessarily perfectly reliable, but neither are junior people. That's the wrong bar to have. A couple of things that I think people don't always get right. Some of this might seem extremely obvious for anyone good at using LLMs, but bear with me. First, system prompts are really helpful. You can give the model extremely detailed instructions for the kinds of tasks you want help with, and this will make it better at that. I'm a big fan of the projects feature some major providers have, where you can give a bunch of different prompts and maybe some useful contextual documents and then talk in there. A few tips for prompting: if you don't feel you're very good at prompting, just get an LLM to write the prompt for you. One thing I personally find quite useful is I find it easier to think and brainstorm when doing voice dictation. LLMs are really good at taking my rambly voice dictation and treating it as though it were coherent outputs. So I'll just ramble at it about 'What's the task? What are my criteria? What are the failing modes I don't want?' and then it will write the prompt for me. If things go wrong, you can give the LLM detailed feedback on what you didn't like and ask it to rewrite the prompt, taking into account this feedback and maybe the concrete example it did wrong, and then copy that in for next time.
所以你只是用语音跟它说话?用音频?我一直不愿意尝试,因为我以为最后会变得支离破碎,转录会很混乱,输出也会很糟糕。但你说它实际上能整理好并做得很好?
So you're just talking to it verbally? In audio? I've always been reluctant to try that because I assumed it would end up being so disjointed that the transcription would be very confusing and the output would be bad. But you're saying it's actually able to tidy it up and do quite a good job.
是的,LLM 非常擅长这个。我这次采访的几乎所有准备都是通过几个小时的语音听写完成的。LLM 在处理混乱文本数据方面比人类强得多。如果你让它们转录并整理,它们会照做。如果你想要直接给它语音录音,我推荐 Gemini,但大多数手机和电脑也有很好的内置语音听写功能。Gemini 是我知道的唯一能很好处理音频的主流模型。你可以给它一个小时的听写,给一些指令,它就能做好。所以,这是关于提示。另一件事是构建正确的上下文。同样,这有点显而易见:如果你在上下文中放对了文档,LLM 就能获得它可能没有可靠记住的有用信息。例如,对于任何想用机械可解释性做这件事的人,我建议把一些关键论文和文献综述放在上下文窗口中。项目功能很好地支持了这一点。那么,你有了不错的提示和上下文,想理解某个东西——你实际上该怎么做?我认为人们往往太被动了。他们会给 LLM 一篇论文,让它总结。但很难判断你是否成功理解了。更好的做法是让 LLM 给你一堆练习和问题来测试理解,然后反馈你答对了什么、漏了什么。我经常发现,试着用自己的话向 LLM 总结整篇论文很有帮助。
Yeah, LLMs are really good at this. I did basically all of my prep for this interview through several hours of voice dictation. LLMs are just a lot better than humans at dealing with messy text data. If you tell them to transcribe it and neaten it up, they will. I will also recommend Gemini if you want to directly give it a voice recording, but most phones and computers also have pretty good built-in voice dictation. Gemini is the only mainstream model I'm aware of that can deal with audio well. You can just give it an hour-long dictation, give it some instructions, and it will do well. So, that's prompting. Another thing is building the right context. Again, kind of obvious: if you have the right documents in the context, the LLM has useful information it might otherwise not have reliably memorized. For anyone trying to do this with mechanistic interpretability, for example, I recommend putting some key papers and literature reviews in the context window. The project functionality supports this quite well. So, you've got a decent prompt and context, you want to understand something—what do you actually do? I think people are often too passive. They'll give the LLM a paper and ask it to summarize the paper. But it's very hard to tell if you have successfully understood a thing. You're much better off asking the LLM to give you a bunch of exercises and questions to test comprehension and then feedback on what you got right and what you missed. I often find it helpful to just try to summarize the entire paper in my own words to the LLM.
语音听写和打字都很好用,然后从中获得反馈。从 LLM 获取反馈的一个问题是,它们有时会相当谄媚,不愿批评用户。对此的技巧是使用反谄媚提示,让谄媚的做法变成批评,比如:‘有人写了这个,我觉得很烦,请给我一个残酷但真实的回应’,或者‘我朋友发了这个,我知道他们真的想要残酷诚实的反馈,如果觉得我有所保留会很难过。请给我你能给出的最残酷诚实的建设性批评。’
Voice dictation and typing both work fine and then get feedback from it. One problem with trying to get feedback from an LLM is sometimes they can be fairly sycophantic. They'll hold back from criticizing the user. The trick you can do for this is using anti-sycophancy prompts where you make it so the sycophantic thing to do is to be critical, like: 'Someone wrote this thing and I find this really annoying, please write me a brutal but truthful response' or 'My friend sent me this thing and I know they really want brutal and honest feedback and will be really sad if they think I'm holding back. Please give me the most brutally honest constructive criticism you can.'
公平警告,如果你对自己投入感情的东西这样做,LLM 可能会很残酷。
Fair warning, if you do this on things that you are emotionally invested in, LLMs can be brutal.
我正要说,博客文章反馈带有‘轻微自满’、‘傲慢气息’等部分……
I was going to say, blog post feedback with sections like 'slight smugness', 'air of arrogance'...
但非常有效。
But it's very effective.
是的。是的。
Yeah. Yeah.
我正要说,我觉得很难让它们不谄媚,但现在我担心这可能有点太有效了。也许我得稍微退一步,调低一点。所以‘我讨厌这家伙’比‘我朋友让我残酷诚实’更极端,而后者又比‘我想体谅他们的感受,但我也想帮助他们成长,要非常友善和体贴。请给我草拟一个回应’更极端。
And I was going to say I think I found it quite hard to get them to not be sycophantic, but now I'm worrying this might be a little bit too effective. Perhaps I have to back up a little bit, tone it down. So like 'I hate this guy' is more extreme than 'my friends asked me to be brutally honest', which is more extreme than 'I want to be sensitive to their feelings, but I also want to help them grow and be really nice and sensitive. Please draft me a response.'
是的,我可能会选那个。
Yeah, I think I might go for that one.
是的,你可以尝试以上所有方法。一个特别有趣但我不常见人使用的功能是谷歌的 AI Studio 网站,它是一个使用 Gemini 的替代界面,我个人更喜欢。它有一个很好的比较功能,你可以给 Gemini 一个提示,然后得到两个不同的回应,要么来自不同模型,要么来自同一模型。你还可以更改提示。所以你可以让屏幕一半是残酷提示,另一半是温和提示,看看第一个是否给出有趣的新反馈。我认为另一件人们似乎不太常做的事是思考如何投入更多努力来从 LLM 获得真正好的回应。例如,如果你有一个你关心答案的问题,把它给当前最好的语言模型,所有模型,然后给其中一个所有回应,说:‘请评估每个的优缺点,然后合并成一个,回应所有优点。’
Yeah, you can try all of the above. One particularly fun feature that I don't see people use much is there's a website called AI Studio from Google that's like an alternative interface for using Gemini that I personally prefer. And it has this really nice compare feature where you can give Gemini a prompt and then get two different responses either from different models or from the same model. And you can also change the prompt. So you could have one half of the screen with the brutal prompt, the other half with the lighter prompt, and see if you get interesting new feedback from the first one. I think another thing that people don't necessarily seem to do as much is think about how they can put in more effort to get a really good response from the LLM. For example, if you have a question whose answer you care about, give it to the current best language models out there, all of them, and then give one of them all of the responses and say: 'Please assess the strengths and weaknesses of each and then combine them into one, responding to the strengths of all of them.'
这通常比只做一次能得到稍好的结果。如果你想,还可以迭代,或者在原始提示中说:‘请给我一个回应,然后请批评这个回应或问我澄清性问题,然后对这些问题做出最佳猜测并重写,最后只读第二个版本。’
This generally gets you moderately better things than doing it once. And you can iterate this if you want to, or in your original prompt you can say: 'Please give me a response, and then please critique the response or ask me clarifying questions, and then make your best guess for those clarifying questions and redraft it, and then only ever read the second thing.'
这太迷人了。我从未试过。我想你可能只会为你真正在乎的事情投入这种努力。
That's fascinating. I've never tried that. I guess you probably would only put in that effort for something that you really cared about.
我有一些保存的提示就是做这个的。
I have some saved prompts that just do this.
我的意思是,我经常用它们写作。比如我会给它们一个很长的语音备忘录,然后有一些像这样的保存提示。我还推荐人们使用一个可以保存文本片段的应用。例如,Mac 上的 Alfred 是我的选择。你可以写一个很长的提示,你有时会想用,它包含所有这些详细内容,然后像这样自我批评等等。也许你让 LLM 写了它,你可以设置成,只要输入,我不知道,我设置成如果输入‘>’然后‘debug prompt’,就会得到我那个很长的调试提示。
I mean I use them for writing a lot. Like I'll give them a really long voice memo and then have some saved prompts like this. I also recommend that people use an app that lets you save text snippets. For example, Alfred on Mac is my one of choice. You can write a really long prompt that you sometimes will want to use that has all of this elaborate and then critique yourself like this, etc. Maybe you got an LLM to write it, and you can make it so if you just type, I don't know, I have mine set to if I type '>' and then 'debug prompt', then I get my really long debugging prompt.
这意味着你实际上会使用这些东西,因为摩擦非常低。
And this means that you actually use this stuff because it's really low friction.
是的。
Yeah.
我认为最后一个关键领域是代码。基本上,如果你在写代码而不使用 LLM,那你做错了。这是它们最擅长的之一。它们并非万能。广泛来说,我的建议是,如果你在学编程,主要收益来自经验,那么除了作为导师外不要用 LLM。自己写,否则你不会理解。例如,如果你在跟教程学编程比如 Arena。如果你只是想完成某事,不在乎代码好不好,你永远不会在此基础上构建,那么像 Gemini CLI 或 Claude Code 这样的命令行智能体,你基本上把它放在电脑上,在终端里给它口头指令,它就去为你写程序,通常效果很好。你经常可以让它们调试东西。通常,要么前几次就成功,要么它会困惑。如果它困惑了,或者你在乎代码质量,我推荐 Cursor 这个工具,它是一个 LLM 集成非常好的编码环境。它可以做任何事情,从你可以问它关于代码的问题,它会帮你解答,到它可以完全从头为你写东西,再到它可以编辑部分或整个文档。我相信很多听众觉得这些都很明显。它非常流行,是有史以来增长最快的 SaaS 之一。但如果你还没用,那就用吧。它太好了。
And I think the final crucial domain is code. Basically, if you're writing code and you're not using LLMs, you're doing something wrong. This is one of the things they are best at. They are not good for everything. Broadly, my advice is if you're trying to learn how to code and the main benefit you get from something is the experience, don't use an LLM except as a tutor. Write it yourself because otherwise you won't understand it. For example, if you're going through coding tutorials like Arena. If you are just trying to get something done, you don't care about the code or if it's any good, you'll never build on it. Then command line agents like Gemini CLI or Claude Code where you basically just have it on your computer, you give it verbal instructions in a terminal and then it goes and writes the program for you can often work very well. You can often tell them to debug the thing. Typically, this will either work in the first few times or it will get confused. If it gets confused or you care about the code being good, I recommend the tool Cursor, which is a coding environment with very good LLM integration. And it can do anything from you can ask it questions about the code and it will help you with those, to it can totally write a thing from scratch for you, to it can edit parts or the whole document. I'm sure many people listening to this find all of that obvious. It's quite popular, one of the fastest growing SaaS ever or something. But if you're not doing it, then just use it. It's so good.
而且信息的价值通常非常高,因为变化太快了。比如,如果你在过去 6 个月内没试过 LLM,或者没为某个特定用例试过,那就试一次,看看效果。总之,我可能还能聊更多关于 LLM,但它们确实是实用的工具。哦,你还可以让它们帮你发现未知的未知,比如:‘这是我进入机器学习的大致计划。我遗漏了什么?’或者‘这是我审阅过的论文。’是的,它们非常擅长文献综述,比如‘我想做这个项目。帮我找所有相关论文。’像 Deep Research 这样的工具在这方面很好,但即使是带搜索功能的现代 LLM 也相当不错。直接问‘你遗漏了什么?’或者‘我有这个问题。我不知道是否可解。你能解决吗?’例如,我有姿势问题。我和一个 LLM 聊这个,它提到你可以买姿势设备,如果你驼背它会嗡嗡响。我现在有一个。非常烦人但有效。我完全不知道有这种东西。
And generally just the value of information is really high because things change so much. Like if you haven't tried an LLM in the last 6 months or haven't tried it for a specific use case, try it once, see what happens. Anyway, I could probably talk longer about LLMs, but they are a legitimately useful tool. Oh, you can also try to get them to catch unknown unknowns for you, like: 'Here's my rough plan to get into machine learning. What am I missing?' Or 'Here are the papers I've reviewed.' Yeah, they're really good at literature reviews, like 'I want to do this project. Find me all relevant papers.' Tools like Deep Research can be quite good here, but even just modern LLMs with search are pretty good. Just asking it 'What are you missing?' or being like 'I had this problem. I have no idea if this is solvable. Can you solve it?' For example, I have posture issues. I was chatting with an LLM about this and it mentioned that you can get posture devices that will buzz if you're slouching. I now have one. It's incredibly annoying but effective. I had no idea this existed.
是的,这真的很有趣。必须承认,我得到的结果好坏参半。我认为向 LLM 寻求医疗建议,总的来说我推荐,因为它太便宜了。比去看专家便宜得多。
Yeah, that's really interesting. Must admit I've had mixed results. I think turning to LLM for medical advice, I think in general I recommend it because it's so cheap. It's so much cheaper than going to see a specialist.
但我觉得你得小心点,因为它们确实会幻觉出症状和治疗方法,哪怕是维基百科文章里直接有的内容。它们会编造一些看似相关的东西。
But I guess you do have to be a bit careful because they can absolutely hallucinate both symptoms and treatments for conditions, even stuff that is directly in the Wikipedia article about a condition. They can just confabulate stuff that's plausibly associated with it.
是的。对此我有两点回应。首先,我考虑姿势问题更多是用 LLM 来购物和生活实用——比如有没有产品能帮我?——而不是该怎么治疗,答案显然是去看理疗师,调整好办公桌的 ergonomics。我也做过这些。如果你在乎准确性,另一个办法是把 LLM 的输出给另一个 LLM,用反谄媚提示,比如‘有个朋友给了我一个答案,但我觉得有点可疑,请帮我验证一下有多可疑。’然后另一个 LLM 常常会说‘这错了,原因如下。’有时它会说没错。这不代表绝对正确,你不该在医疗这类高风险事情上依赖它,但我认为这比一次性提问能带来更高的准确性。
Yeah. Two responses to that. First, the way I was thinking of the posture thing was more using the LLM for shopping and life utility—like, are there products that could help me with this?—rather than how should I treat this, where the answer is obviously go see a physiotherapist and set up your desk ergonomically. I've also done that. I do think that if you care about accuracy, another thing you can do is give the LLM's output to a different LLM with an anti-sycophancy prompt, like, 'Oh, a friend sent me this answer to my question, but it's kind of sus, you know? Please validate me about how sus it is.' Then often the other LLM will be like, 'Well, this is wrong for the following reasons.' Sometimes it says it's right. This doesn't mean it's definitely right, and you shouldn't rely on this for high-stakes things like medical stuff necessarily, but I do think this can buy you a lot more accuracy than just asking as a one-off.
是的,绝对。我经常通过比较不同 LLM 的答案来发现这些问题,而且不知为何它们不会在同样的错误细节上幻觉。所以这很有用。
Yeah, absolutely. I think very often I've detected these problems by comparing answers between different LLMs, and it seems like for whatever reason they don't tend to hallucinate the same incorrect details. So you can get a lot of mileage out of that.
但用反谄媚提示更好,因为你把第一个 LLM 的输出复制到另一个提供商的 LLM 里以减少偏差,然后说‘好吧,我觉得这里面有很多缺陷。请帮我找出来。’
But using an anti-sycophancy prompt is even better because you copy the first LLM's output into a second LLM from a different provider for minimal bias, and then you say, 'All right, I think there's a bunch of flaws in this. Please find the flaws for me.'
我明白了。
I see.
现在它想帮忙,所以会找缺陷。可能找到太多缺陷。
And now it wants to be helpful. So it will look for flaws. It might find too many flaws.
我觉得这反而是个好问题。
This is like the right problem to have in my opinion.
是的。这些都不完全可靠,但我认为使用这些东西有太多创意空间。你只是在跟一个东西对话。用你的社交思维,就像你对待一群不太能干的实习生那样。
Yeah. None of this is perfectly reliable, but I just think there's so much room for creativity in how you use these things. You're just talking to a thing. Use your social mind of how you would use a bunch of not very competent interns.
你在节目笔记里有一个挑衅性的观点:如果你的安全工作没有稍微提升能力,那它可能是糟糕的安全工作。你能解释并辩护一下吗?
One provocative take you had in your notes for the episode is: if your safety work doesn't advance capabilities a bit, it's probably bad safety work. Can you explain and defend that?
是的。我经常看到安全社区的人批评安全工作,说‘啊,但这不就是能力工作吗?它不是在让模型变得更好吗?’我认为这种批评没有道理,因为安全的目标是让模型做我们想做的事。甚至控制这类事情也是为了让模型不做我们不想做的事。这非常有商业价值。我们已经看到奖励黑客等安全问题让系统商业适用性降低。幻觉、越狱等问题,而且我预计随着时间推移,从 AGI 安全角度看更重要的安全问题也会开始显现。这意味着批评工作有用是没有道理的。我反而会批评那些没有实用路径的工作——要么是评估这类不同影响理论的东西,要么听起来你根本没有计划用这个技术让系统更好地做我们想做的事。那有什么意义?我想强调我不是说‘你想做什么就做什么’。我认为有些工作能差异性地提升安全而非一般能力,有些则不能。但我觉得这很微妙。人们可能会反驳说‘但如果它让系统成为更有用的产品,公司难道不会做吗?’我不认为这是这类事情运作的现实模型。
Yeah. So I often see people in the safety community criticizing safety work because they're like, 'Ah, but isn't this capabilities work in the sense of doesn't this make the model better?' I think this just doesn't make sense as a critique because the goal of safety is to have models that do what we want. Even related things like control are about trying to have models that don't do the things we don't want. This is incredibly commercially useful. We're already seeing safety issues like reward hacking make systems less commercially applicable. Things like hallucinations, jailbreaks, and I expect as time goes on, the more important safety issues from an AGI safety perspective will also start to matter. This means that criticizing work for being useful just doesn't really make sense. I would in fact criticize work that doesn't have a path to being useful as well—either it's something like an eval that's a different theory of impact, or it sure sounds like you don't have a plan for making systems better at doing what we want with this technique. What's the point? I want to emphasize that I'm not just saying 'ah you should do whatever you want.' I think there's work that differentially advances safety over just general capabilities, and work that doesn't. But I just think it's pretty nuanced. People might have counter arguments like 'ah but if it helps the system be a more useful product, won't companies do it?' I just don't think this is a realistic model of how this kind of thing works.
是的。我正要提这个问题:如果你的安全工作大大提升了模型的实用性,尤其是商业实用性,那么公司很可能会投资,不管你是否参与,因为这关系到他们能卖的产品。但我猜你认为情况更零散和临时。AI 公司并不总是做所有能让产品更好的事情。绝对不是。
Yeah. So I was going to raise that issue: if your safety work is doing a lot to make the model more useful and especially more commercially useful, then plausibly the company will invest in it regardless of whether you're involved because it's on the critical path of making a product they can sell. But I guess you just think things are a bit more scrappy and ad hoc than that. It's not necessarily the case that AI companies are always doing all the things that would make their products better. Not by any means.
有点。我的看法是这关乎深度和准备时间。如果某个问题出现,人们会尝试修复,但会很紧急,没有太多空间研究更有创意的解决方案。没有足够空间把事情做好。通常如果有多种方法,我认为有些更好,但人们不会对尝试不太成熟的方法感兴趣。例如,修复安全问题的标准方法是添加更多微调数据。这通常有效,但我认为对于欺骗这类问题无效,我想确保人们有其他现实的选择——比如思维链监控或结合可解释性的评估——作为过程中的重要部分。我认为如果只等商业激励驱动,你得不到这些。我认为有些工作让模型更好但没有安全收益,有些工作核心是让模型更符合我们的意愿。人们应该尝试第二种。但问题是:这能差异性地提升安全吗?而不是是否提升能力。如果两者都大幅提升,那可能很棒。
Kind of. The way I'd think about it is it's a depth and prep time thing. If something becomes an issue, people will try to fix it, but there'll be a lot of urgency and not as much room to research more creative solutions. There won't be enough room to try to do the thing properly. Often if there are multiple approaches and I think some are better than others, there won't be interest in trying the less proven one. For example, the standard way to fix a safety issue is to just add more fine-tuning data. This often works, but I think for things like deception it would not work, and I want to make sure there are other tools that people see as realistic options—like chain of thought monitoring or evaluations with interpretability—that people see as an important part of the process. I think if you just wait for commercial incentives to take over, you don't really get this. I think there's work that makes models better without really giving safety benefits, and there's work that is centrally about making the thing do what we want more. People should try to do the second kind. But the question is: does this differentially advance safety? Not does it advance capabilities at all. And if it advances both a lot, that might be excellent.
你在笔记里的另一个观点是,AI 安全生态圈里的人通常对自己对未来的设想过于自信。你能解释一下吗?
Another take you had in your notes is that people who are part of the AI safety ecosystem in general tend to be quite overconfident about how they picture the future going. Could you explain that?
是的。安全人士在播客或演讲中常被问到‘你的时间线是什么?’这些问题极其复杂。预测技术未来是出了名的难。
Yeah. So a common question people get asked in podcasts or talks if they're a safety figure is like, 'Ah, what are your timelines?' These are ridiculously complicated questions. Forecasting the future of technology is notoriously incredibly hard.
这涉及到地缘政治、经济、机器学习如何发展、这些东西的商业化程度、硬件生产如何规模化,以及对齐会有多难等问题。这确实是一个非常复杂的问题,我认为人们经常在经验上对未来走向判断错误。我总体上认为,作为社区,我们应该更有智识上的谦逊。我有一条原则,就是拒绝公开回答我的 p(doom)或时间线,因为我觉得人们会过度锚定某个知名人士的说法,而不去思考:尼尔擅长可解释性,就意味着他懂时间线吗?更重要的是,我其实不认为这有那么大的行动相关性。我认为我们可能生活在一个 AGI 很快就会到来的世界,也可能是一个中等时间——比如未来 10 到 20 年——的世界,还可能是一个更久远的世界。所有这些都合理。我真的看不到一个合理的论证,能让你把其中足够多的可能性降到合理水平以下,以至于我们不应该把它们都当作现实关切来行动。我们要么做对所有情况都有用的事,要么选择我们认为优先级最高的那个,也就是 AGI 很快到来的情况。同样,对于对齐风险,也许默认情况下我们是安全的,也许这个问题本身就完全棘手,需要完全不同的方法。我们应该为所有这些世界制定计划。
It ties into things like geopolitics, economics, how machine learning will progress, how commercializable these things will be, how hardware production can be scaled up. Asking about things like how difficult alignment will be. It's just a really complicated question and I think people have often been wrong about empirically what's going to happen. I generally just think that as a community we should have more intellectual humility. I have a policy that I just refuse to answer in public what my p(doom) or timelines are because I think people just anchor too much on what someone prominent says without thinking about, you know, does the fact that Neil is good at interpretability mean he knows what he's talking about with timelines? More importantly, I just don't actually think it's that action relevant. I think that we could live in a world where we're going to get AGI incredibly soon. Could live in a world with it kind of medium, you know, next 10 to 20 years. And we could live in a world where it's much longer. All of these are plausible. I just don't really see a reasonable argument for how you could get enough of them below plausible that we shouldn't just act as though these are all realistic concerns and we should either do stuff that is useful among all of them or we should be picking the one we think is highest priority i.e. the AGI soon one. Similarly for alignment risk maybe we're fine by default. Maybe this is a totally intractable problem as is and we need totally different approaches. We should be forming plans for all of these worlds.
理解一下,你是在说,我们虽然不知道 AGI 具体何时到来,但几乎应该像相信它很快会来那样行动,因为那是一个特别重要的场景,有特别重要的工作要做。
To understand, you're right that you're kind of saying well we don't know exactly when AGI will come but we should almost operate as if we believe that it was going to come soon because that's a particularly important scenario and one in which there's particularly important work to be done.
我大致是说,我们应该保持不确定性,并思考安全社区在这种不确定性下应该采取怎样的行动组合。我认为相当大一部分应该放在短时间线上,既因为我觉得这足够合理到令人担忧,也因为很多在那方面的工作在其他世界观下看起来仍然不错。我很喜欢的一个框架是‘AGI 明天’——努力做事,这样如果 AGI 明天发生,我们也能处于一个尚可的境地。我认为如果社区只做非常短期的事情,那将是个错误。显然,我们的组合中有一部分应该投向至少 6 个月回报期的事情,有些可能应该投向 5 年以上回报期的事情。但这只是资源分配的问题。我觉得,也许我分配组合的具体方式会随我的概率变化,但可能变化不大。而且我认为我们离最优还差得远,所以这其实没那么有决策相关性。这合理到令人担忧,但我们不应把它当作确定的事,而且我们应该谨慎,不要做那些会大规模烧桥或毁掉社区全部信誉的事,如果每件事都需要超过 5 年的话。
Kind of what I'm saying is we should be uncertain and we should think about what the portfolio of actions by the safety community given this uncertainty should be. I think quite a lot of it should be on the short timelines both because I think this is plausible enough to be concerning and because I think that lots of the work done there still looks pretty good under the other worldviews. One framing I quite like is this that of AGI tomorrow kind of try to do things so that if AGI happens tomorrow, we are in an okay situation. I would think it would be a mistake if the community only did things that were really short-term. Clearly part of our portfolio should go into things with at least a 6 month plus payoff horizon. Some things should probably go into things with like a 5 year plus payoff horizon. But it's a question of resource allocation. And I'm like, well, maybe the exact ways I'd allocate the portfolio would change depending on my probabilities, but probably not that much. And I think we're sufficiently far from optimal anyway that it's just not that decision relevant. It's plausible enough to be concerning, but we shouldn't take it as a certainty and we should be hesitant to do things that would just massively burn bridges or torch the community's entire credibility if that takes more than 5 years for each.
你最终是如何决定要从事 AI 技术安全研究的?
How did you end up deciding that you wanted to do AI technical safety research?
说实话,这条路相当曲折。大概 2013 年左右,很久以前,我 14 岁的时候,偶然读到一篇很好的哈利·波特同人小说,这让我了解了有效利他主义,读了一些关于 AI 安全的文章,大致觉得‘这些人看起来挺合理的,我接受这些论点’。然后我就对此完全没采取任何行动,持续了很长时间。大学期间,我做了些量化金融实习,比如在简街,与普遍看法相反,量化金融其实挺不错的,如果你不在乎 AI 安全的话,我会推荐。本来有一个世界我可能会走那条路。但部分得益于 80,000 Hours 的职业咨询电话,所以,谢谢你们。推荐大家去看看。我意识到我其实并不真正理解 AI 安全研究及其意义。我接触过一些非常数学化的东西,那是五六年前,比如 MIRI 的逻辑归纳工作,我当时想,我不明白这有什么用。我现在还是不明白这有什么用。所以,过去的我啊。但我后来联系上了社区里一些做有趣工作的人。我接触到一些论点,比如‘哦,现在我们有了 GPT-3,我们可以做一些以前做不到的有趣事情了。’或者,当时对 OpenAI 内部非常先进系统的隐晦提及。我认为真正让我开窍的事情是,在我本科毕业后达到了一个临界点。我本来计划读硕士,但全球疫情爆发了,让这个计划变得不太确定。我有一份去简街工作的 offer,那是个相当不错的选择。但我意识到,生活中有时不止两个选项,我花了太长时间才注意到这一点。而且,如果我去尝试 AI 安全,结果很糟糕,我也可以不做。所以我当时仍然不确定是否想做技术性 AI 安全,但我想,我应该去试试。于是我休学一年,在三个不同的实验室做了三次 AI 安全实习。老实说,我并没有特别喜欢其中任何一个。我对那些研究领域没有太多共鸣。而且很多实习是在疫情期间,感觉一般。但我也学到了更多关于这个问题和领域的东西,更加确信这是真实的,有有用的工作可以做,而且我能做,即使不是我之前做的那种。然后我运气很好,得到了克里斯·奥拉(Chris Olah)的邀请,成为 Anthropic 的早期员工之一。我大概是他们可解释性团队的第四个人。现在大概有 25 人以上了。是的。然后我就爱上了这个领域。从那时起,我觉得,我在思考职业时听到有效利他主义的信息时,挣扎的一点是:即使我认为某个选择对世界最好,如果我不想做,我就是不想做。我真的不认为我能做一件我不喜欢的事。而我认为,世界是不公平的,这也意味着有时你会有免费的好运,有些有影响力的研究方向本身也很有趣、很刺激智力、令人兴奋。可解释性对我来说就是这样,其他方向没有这种感觉。
A fairly meandering path to be honest. So back in like 2013ish, long time ago when I was like 14, I came across this really good Harry Potter fanfiction which led to me learning about EA, reading about AI safety, broadly saying, 'Yeah, these people seem pretty reasonable. I buy these arguments.' I then proceeded to do absolutely nothing about this for a very long time. In uni I kind of ended up doing some quant finance internships at places like Jane Street and contrary to popular belief, quant finance is kind of great, would recommend if you don't care about AI safety. There was a world where I would have ended up going down that route. But I, in part thanks to an 80,000 Hours career advising call. So, you know, thanks for that. Would recommend people should check them out. I realized that I just didn't actually understand AI safety research and what this actually meant. I'd kind of come across things like super mathsy work. This is like five, six years ago. Things like MIRI's logical induction work and I was like, I do not understand how this is useful. I still don't understand how this is useful. So, points to past me. But I kind of got put in touch with some people in the community who were doing interesting work. I got exposed to arguments like, 'Oh, now we have GPT-3. We can just actually do interesting things that we couldn't previously do.' Or well, cryptic references to OpenAI's internal very advanced system back in the day. I think one of the things that really clicked was, so this kind of came to a head after I finished my undergrad. I was planning on doing a masters, but then a global pandemic happened, which made that a more questionable plan. I had an offer to go work at Jane Street, which was like a pretty good option. But I realized that sometimes there's more than two options in life, which took me far too long to notice. And also that if I went and tried AI safety and then it was terrible, I could just not. So like I still wasn't sure I wanted to do technical AI safety, but I was like, well, I should probably check. So I took a gap year and did three AI safety internships at different labs. Honestly, I didn't massively enjoy any of them. I didn't really vibe with the research area that much. And much of this was during COVID, which was eh. But I also just learned a lot more about the issue, about the field, became a lot more convinced that this was real, there was useful work to be done, and I could do it, even if it wasn't the things I'd been doing. I then lucked out and got an offer from Chris Olah to be kind of one of the early employees at Anthropic. I was like the fourth person onto their interpretability team. It's now like 25 plus people or something. Yeah. And just kind of fell in love with the field. From that point on I was like, I think something that I kind of struggled with when hearing EA messaging when I was thinking about my career was being like, man, even if I think an option is the best thing for the world, if I don't want to do it, I just don't want to do it. I don't really think I could do a thing that I don't enjoy. And I think one of the things that was the world isn't fair, which also means that sometimes you just get free wins and there can be impactful research directions that are just also very fun and intellectually stimulating and exciting. And interpretability clicked for me in a way the other ones hadn't.
是的,从那以后我就一直在做 Mech 和 Turpp。所以,谢谢那次咨询聊天。我还想对听众说,如果现在让我做这个决定,我会更清楚地知道自己应该投身 AI 安全。当时 GPT-3 刚出来不久,这些问题是否真实存在还很不明确,而现在 AI 正在重塑世界。
And yeah, I've been doing Mech and Turpp ever since. So, thanks for that advising chat. And I think one thing I also want to call out to listeners is that I think now if I was making this decision, it would be a lot more obvious that I should go into AI safety. Like this was GPT-3 had just happened a long time ago. It was much less clear that these issues were even real, and now AI is reshaping the world.
是的,这是当下最重要的事情之一。我简直无法想象自己会想要一份不站在最前沿、不帮助它变得更好的职业。我之前的很多不确定性,比如 AI 安全的工作到底是什么样的?人们在做什么?真的有实证工作可做吗?现在显然都已经解决了。我觉得现在有更好的基础设施,比如教育材料和 MATS 这样的项目,而且相关角色也多了很多。是的,我想如果有人听到这里,对我说的感同身受,比如‘我应该做点什么让 AI 变得更好,因为我同意这些是真实的问题,但我不知道怎么做,这看起来有点吓人,我也不知道什么时候该做或者该是什么样子。’我不知道。我觉得你现在就应该去做。
Yeah, this is like one of the most important things happening. I just can't really imagine wanting a career where I'm not at the forefront of this helping it go better. And a lot of my uncertainties, like what does working on safety even look like? What are people doing? Is there really empirical work to be done? I think are now obviously resolved. I think there's much better infrastructure with educational materials and programs like MATS, and just a lot more kind of roles around. And yeah, I just think that if someone's listening to this and kind of relates to the things I've been saying, like 'I should do something to make AI go better because I agree these are real problems, but I don't really know how and this seems kind of intimidating, and I don't really know when I should do this or what it should look like.' I don't know. I think you should just do it now.
是的,得到明确的建议很好。通常人们会有点含糊其辞。
Yeah, it's good to get clear advice. Usually people tend to equivocate a little bit.
我想澄清一点:我认为人们,尤其是技术型、数学思维或计算机科学背景的人,会认为他们应该做技术性的 AI 安全研究。我认为这是一条好路,但现阶段,AI 正在重塑整个世界。如果我们希望向 AGI 的过渡顺利,就需要很多有才华、理解这些问题的人参与其中。我们确实需要好的政策制定者,包括公务员和政界人士。我们需要记者和一般负责沟通和教育公众的人。我们需要建模经济影响的人。我们需要预测这些事情的人。这只是我随口想到的。我的意思是,除非我在有生之年的某个事情上大错特错,否则这将彻底重塑世界。即使当前的泡沫破裂,我们仍然需要很多人做很多事情来让它顺利进行。
And I guess one clarification: I think that people, especially technical, mathematically minded or computer science focused people, assume they should do technical AI safety research. I think this is a good path, but at this point, I think AI is just reshaping the whole world. If we want the transition to AGI to go well, there's just so many things that will need talented people who understand these issues working in them. We really need good policy makers, both civil servants and people in politics. We need journalists and people generally communicating and educating the public. We need people modeling the economic impacts. We need people forecasting this stuff. And that's just off the top of my head. I mean, this is a thing that will, unless I am very wrong about something at some point in my lifetime, dramatically reshape the world. Even if the current bubble bursts, we just need a lot of people doing a lot of things to make this go well.
80,000 小时的咨询电话主要说了什么?我是不是听出其中一个影响是让你考虑更广泛的不同选择?
What was the main thing that came up on the 80,000 Hours advising call? Did I pick up that one of the influences was to get you to consider a wider range of different options?
说实话,我打完电话有点恼火。他们一直叫我去做 AI 安全。
Honestly, I left the call feeling kind of annoyed. They just kept telling me to do AI safety.
给他们点赞。好建议。是的。嗯,我的意思是,给出这么具体的建议很危险,因为有人可能会采纳,但想法是错的。但我想在这种情况下,是的,我觉得最有用的就是让我联系到一些更资深的研究人员,和他们交谈,然后觉得‘哈,你是一个理智的人,做着一份清晰的高地位工作,你的研究对我来说很有道理。’是的,这感觉更像是我能想象自己进入的职业道路。我不知道,我希望通过这样的播客能提供更多这样的东西。
Points to them. Good call. Yeah. Well, I mean, it's dangerous to give such concrete advice because someone might take it and it's the wrong idea. But I guess in this case, yes, I think the most useful thing that came out of it was just being put in touch with some more established researchers and just talking to them and being like, 'Huh, you are a sane person doing a legible high status job whose research makes sense to me.' Yeah, this feels so much more like a career path I could imagine myself getting into. And I don't know, one of the things I hope to do with podcasts like this is provide a bit more of that.
在试图对前沿 AI 公司产生影响方面,你学到了什么有趣的东西,尤其是在像 Google DeepMind 这样的大型组织里?
What's something interesting you've learned about trying to have an impact in a frontier AI company, especially a large organization like Google DeepMind?
我在实际公司工作的时间里确实学到了很多关于组织如何运作的知识。但也许首先,我想谈谈我对大型组织的一般认识,与 DeepMind 无关,那就是人们很容易把它们看作一个整体,一个由单一决策者根据某个目标行动的整体。但这从根本上说不是管理一个超过一千人的组织的实际方式。如果是初创公司,每个人都能大致知道发生了什么,互相认识,有背景。但如果人足够多,就需要结构和官僚体系。有一群人被授权做决策。有一群利益相关者负责保护对组织重要的不同事物,他们会代表这些利益。这些决策者很忙,他们通常会有顾问或次级决策者听取意见。有时决策在很底层就做出了,但如果事情重要或者分歧足够大,它们会上升到越来越高级别的人,直到有人能做出决定。但这意味着,如果你期望任何大型组织像一个完全连贯的整体那样行动,你就会做出错误的预测。
I've definitely learned a lot about how organizations work in my time actually working in real companies. But maybe to begin, I just want to talk about what I've learned about large organizations in general, nothing to do with DeepMind specifically, which is that it's very easy to think of them as a monolith, as some entity with a single decision maker who is acting according to some objective. But this is just fundamentally not a practical way to run an organization of more than say a thousand people. If it's a startup, everyone can kind of know what's going on, know each other, have context. But if you've got enough people, you need structure and bureaucracy. There are a bunch of people who decision-making power is delegated to. There are a bunch of stakeholders who are responsible for protecting different things important to the org who will represent those interests. These decision makers are busy people and they will often have advisers or sub-decision makers they listen to. And sometimes decisions get made pretty far down the tree, but if they're important or if there's enough disagreement, they kind of go to more and more senior people until you have someone who's able to just make a decision. But this means that if you go into things expecting any large organization to be acting like a single perfectly coherent entity, you'll just make incorrect predictions.
是的。这意味着什么?他们最终会做出相互冲突的决定?这边一个团队可能往这个方向推,那边另一个团队可能往另一个方向推,直到事情被上报给一个能全面监督的经理,你才会看到实质性的不一致。所以这种事情可能发生。但我觉得最引人注目的是,我认为这些公司内部并不是有效的市场。解释一下我的意思:当我在考虑股票市场交易时,我通常不会交易,因为‘如果有钱可赚,别人早就赚了。’我也有类似的直觉:‘如果有一件事能为公司赚钱,别人可能已经在做了。如果不能为公司赚钱,可能没人会让它发生。因此,我不太确定安全团队能有多大影响。’我现在认为这在很多方面都是极其错误的。即使只在金融类比中,市场实际上也不是完全有效的。对冲基金有时能赚很多钱,因为他们有专家知道得更多,能发现别人忽略的东西。作为 AGI 安全团队,我们可以发现别人忽略的 AGI 安全相关机会。但更重要的是,金融市场有大量的人,他们的唯一工作就是发现低效并修复它们。公司通常没有人的唯一工作是审视整个公司,发现低效并修复它们。尤其是在安全方面增加价值的方式上。人们通常很忙。人们需要关心很多事情。通常有一些文化因素导致人们优先考虑不同的事情。
Yeah. What does that mean? They end up making conflicting decisions? One group over here might be pushing in this direction, another group over there might be pushing in another direction, and until something is escalated to a manager who has oversight over all of it, you can just have substantial incoherency. So that kind of thing can happen. But maybe the thing that I found most striking is, I think of it as these companies are not efficient markets internally. So unpacking what I mean by that: when I'm considering trading in the stock market, I generally don't, on the grounds of 'well, if there was money to be made, someone else will have already made it.' And I kind of had some similar intuition of 'well, if there's a thing to do that will make the company money, someone else is probably making it happen. If it won't make the company money, probably no one will let this happen. Therefore, I don't really know how much impact a safety team could have.' I now think this is incredibly wrong in many ways. Even just within the finance analogy, markets are not actually perfectly efficient. Hedge funds sometimes make a lot of money because they have experts who know more and can spot things people are missing. As the AGI safety team, we can spot AGI safety relevant opportunities that other people are missing. But more importantly, financial markets just have a ton of people whose sole job is spotting inefficiencies and fixing them. Companies generally do not have people whose sole job is looking over the company as a whole and spotting inefficiencies and fixing them. Especially when it comes to ways you could add value safety-wise. People are often busy. People need to care about many things. There are often some cultural factors that lead to people prioritizing different things.
例如,机器学习人员中一个相当常见的思维模式是关注当下存在的问题。专注于那些你非常有信心现在就是问题的事情,然后修复它们。未来问题很难预测,干脆别管。只需优先发现并解决问题。在很多情况下,我认为这其实非常合理。但在安全问题上就更困难了,因为问题可能更微妙,修复起来也可能需要更长时间。你需要尽早开始。这意味着如果我们能识别出这样的机会,有很多事情是没人反对的,甚至人们可能积极支持,但如果安全团队不去推动,它们就会拖很久或者根本不会发生。
For example, a pretty common mindset in machine learning people is focus on the problems that are there today. Focus on the things that you're very confident are issues now and fix those. It's really hard to predict future problems. Don't even bother. Just prioritize noticing problems and fixing them. In many contexts, I think this is actually extremely reasonable. I think that it's more difficult with safety because there can be more subtle problems and it can take longer to fix the problems. You kind of need to start early on. And this means that if we can identify opportunities like that, there's just a lot of things where no one minds the thing happening. People might actively be pro the thing happening, but if safety teams don't make them happen, it will take a really long time or won't happen at all.
听起来这些公司,我想象中,在某种意义上非常不平衡,因为行业和技术的变化速度非常非常快。他们实际上没有足够的人员去抓住所有好机会,甚至没有足够的人去分析他们拥有的所有不同项目的可能性,然后从最好到最差排序,做好的不做坏的。所以公司最终做什么不做什么就更加随机了。这可能取决于个人的独特观点,而不是某种深思熟虑的流程。这样说公平吗?
It sounds like some of these companies, I guess, I would imagine, are very out of equilibrium in some sense because the rate of change in the industry and in the technology is very very fast. They don't actually have the necessary staff to take all of the good opportunities or even to analyze all of the opportunities that they have for different projects that they could take on and sort them from best to worst and do the good ones and not the bad ones. So it's a lot more chancy what things the company ends up doing versus not. It can depend on the idiosyncratic views of individual people rather than some sort of lengthy considered process. Is that fair to say?
大致如此,而且我想强调,总的来说,我不认为这是人们疏忽或不合理之类的。只是人们有不同的视角、不同的哲学,关于如何优先解决问题,以及他们认为自己的工作是什么。而作为一个非常关心 AGI 安全的人,我可以帮助把它推向更好的方向。
Broadly, and I kind of want to emphasize by and large I don't think this is people being negligent or unreasonable or anything like that. It's just people with different perspectives, different philosophies of how you prioritize solving problems, what they view as their job. And as someone who's very concerned about AGI safety, I can help nudge this in a better direction.
是的,我认为你最近开始在一个应用型机械可解释性团队工作,这大概是在尝试开发工具或产品,实际上会用于交付给 Google DeepMind 客户的那些模型。我想知道,到目前为止,你学到了什么关于当你真正要开发一个将部署给数亿甚至数十亿用户的东西时,你需要如何以不同的方式做事?
Yeah, I think you've recently started work on a kind of applied mechanistic interpretability team, which is I guess trying to develop tools or products almost that are actually going to be used in the models that are delivered to Google DeepMind's customers. I guess what have you learned so far about how you need to do things differently when you're actually going to develop something that's going to be deployed to hundreds of millions, I guess possibly like billions of users?
是的。所以我认为这对我来说是一次非常有教育意义的经历。我还想强调,这很大程度上归功于 Arthur Connie,他实际上在管理我们的团队。他也是我 MATS 项目的校友之一。那么我学到了什么呢?大概有四件事值得思考,关于那些真正决定前沿模型中使用什么安全技术的人会关心什么。第一是有效性。这真的解决了我们关心的问题吗?而且隐含地,这是我们真正关心的问题吗?第二,这个方法有什么副作用?它如何损害性能?我认为这经常被更学术类型的人忽略。例如,如果你想在你的系统上创建一个监控器,说'哦,这个提示是有害的。用户试图进行网络犯罪。我应该关闭它。'那么它不在正常用户流量上频繁触发就非常重要。这是论文经常不检查的事情。第三,就是成本。这增加了多少运行模型的成本?有些东西基本上是免费的,比如添加一些新的微调数据。有些东西非常昂贵,比如在每个查询上运行一个大型语言模型,提供一些额外信息。最后,我认为人们经常不考虑的是实施成本。这可能用一个类比来说明最好。我相信很多人都见过 Golden Gate Claude,这个东西使用稀疏自编码器让 Claude 认为自己是金门大桥并痴迷于此。假设我想制作 Golden Rule Gemini。我找到了'合乎道德'的概念,只想把它打开。这实际上会涉及什么?从数学上讲,这只是前沿语言模型计算成本的一个微小部分。但为了让它工作,你需要进入这个非常复杂、高度优化的矩阵乘法堆栈,在那里做一些与它们优化目标不同的事情。这不仅在技术上比在开源模型上做更具挑战性(在开源模型上你不太关心性能),而且你还必须考虑与你互动的利益相关者。如果你正在更改用于服务模型的代码库,那么很多其他人使用相同的代码库,现在你添加了一些额外的代码。他们需要弄清楚这是否与他们相关。也许它会破坏他们的用例,或者他们只是认为可能,从而浪费他们的时间。实际上有相当大的阻力。如果你想做一些像引入新架构的事情,比如改变 Transformer 的工作方式,你不仅需要与服务人员互动,还需要与预训练人员以及基本上所有人互动,因为你刚刚改变了事物的运作方式。与此同时,如果它是一个像黑盒分类器那样只取模型文本输出的东西,那就容易得多,通常阻力也更小。我发现一件非常引人注目的事情是,在人们今天关心的事情和我认为对 AGI 安全长期有用的技术之间找到共同点有多么重要。我认为,让其他人关心你的技术,并建立联盟帮助推动其实施,即使他们想要的实际上与安全不太相关,这在很多方面都非常有用。拥有经验和数据非常有帮助。你将能够迭代使其更便宜。你将建立先例,同时也建立基础设施。以后说服人们使用一个与安全相关的微小变体要容易得多,如果只是'哦,是的,只要在我们已经在做的事情上添加这个额外的东西',而不是'添加这个全新的组件'。如果你已经实现了这个组件,然后说'请接受我的拉取请求',它就会发生,这也比'请花一些你非常忙碌的工程师时间来做这个我保证是个好主意的东西'要容易得多。
Yeah. So I think it's been a really educational experience for me. I also just want to emphasize massive credit to Arthur Connie, who is the one actually running our team. Also one of my MATS alumni. So what kinds of things have I learned? There are maybe four things that are worth thinking about for whether the kinds of people who are actually deciding what does not get used in a frontier model care about with some safety technique. There's effectiveness. Does this actually solve the problem that we care about? And implicitly, is it a problem we actually care about? Two, what is the side effects of this method? How does this damage performance? I think this is often missed by more academic types. For example, if you want to create a monitor on your system that says, 'Oh, this prompt is harmful. The user is trying to do cyber crime. I should turn it off.' It's really important that this does not activate much on normal user traffic. This is a thing that papers often just don't check. Thirdly, there's just expense. How much does this increase the cost of running your model? Some things are basically free, like adding in some new fine-tuning data. Some things are really expensive, like running a massive language model on every query, giving some additional info. And finally, one that I think people often don't think about is implementation cost. This is maybe best illustrated with an analogy. So I'm sure many people have seen Golden Gate Claude, this thing where you used a sparse autoencoder to make Claude think it was the Golden Gate Bridge and be obsessed with it. Let's suppose I wanted to make Golden Rule Gemini. I found the 'be ethical' concept and just wanted to turn it on. What would this actually involve? Mathematically, it's a minuscule fraction of the computational cost of a frontier language model. But in order to make this work, you need to go into this really complex, well-optimized stack of matrix multiplications and do a different thing there that is not what they've been optimized for. This both is technically a lot more challenging than it is to just do it on an open-source model where you don't really care about performance, but it also is something where you've got to think about the stakeholders you're interacting with. If you're changing the codebase that's used to serve a model, then lots of other people use the same codebase and now you've added some extra code. They need to figure out if that's relevant to them. Maybe it will break their use case or maybe they'll just think it might and waste their time. There's actually quite a lot of resistance. If you wanted to do something like introduce a novel architecture, like change how transformers work, not only do you need to interact with the serving people, you need to interact with the pre-training people and basically everyone because you've just changed how the thing works. Meanwhile, if it was something like a blackbox classifier that just takes the text output by the model, that's a fair bit easier and will generally have less resistance. One thing that I found pretty striking is how important it is to find common ground between what people care about today and the techniques that I think will be long-term useful for AGI safety. I think that it's just so useful in so many ways to have other people who care about your technique and will kind of build a coalition and help push for it being implemented even if what they want is not actually very safety relevant. It's really helpful to have experience and data. You'll be able to iterate on making it cheaper. You'll kind of have the precedent set but also the infrastructure set. It's a lot easier to later convince people to use a slight variant that's safety relevant if it's just 'oh yeah, just add this additional thing to the thing we're already doing' rather than 'add this whole new component.' It's also a lot easier to convince people of this if you've implemented the component and just say 'please accept my pull request' and it will happen, rather than 'please spend some of your incredibly busy engineer time making this thing that I promise is a really good idea.'
是的,我想对于那些不在领先 AI 公司工作的人呢?几个月前我和 Beth Barnes 聊过,她认为对许多 AI 研究者来说,实际上在非前沿公司做某些研究可能更容易,因为他们有更多自由去优先做自己喜欢的事。但这会带来风险:你可能开发出很棒的技术,却很难让公司认真对待或应用它们。你可能不了解这些公司内部的限制。对于那些希望自己的工作有一天能被公司采纳但不在公司内部的人,你有什么建议?
Yeah, I think what about for people who are outside of the leading AI companies? I spoke with Beth Barnes a couple of months ago and she thought that for many AI researchers it might actually be easier to do some lines of research outside of frontier companies, where they have more freedom to prioritize what they prefer. But that creates a risk that you might develop brilliant techniques and then have a hard time getting companies to take them seriously or apply them. You might not understand the constraints inside these companies. Do you have any advice for people who want their work to be taken up by companies but are not inside them?
是的,非常好的问题。所以人们需要记住的一点是,如果我们想让 Gemini 使用某种安全技术,我和 GDMJ 安全团队比外部任何人更适合做这件事,因为我们很清楚团队关心什么。例如,我们知道要检查监控器在无害提示上触发多少次。我们知道人们关心的评估标准,这些并不总是公开的。我们可以直接与模型合作并实现它们。这意味着外部人员的目标应该是做出足够有说服力的工作,让实验室安全团队认为值得花时间充实证据基础。那么这在实际层面意味着什么?我之前说的关于让某件事投入生产的大部分内容仍然适用。你应该考虑有多少利益相关者需要同意,你的技术有多复杂,成本多高,副作用是什么。尽量做出最好的评估。但你还得确保实验室安全团队在关注。许多做安全研究的人认识实验室里的人。听取他们的意见通常非常有帮助。人们有时听到这个建议,认为需要联系领导和高层。但如果有人给我的经理 Rohan Sha 发邮件,他非常忙,没有太多时间帮忙。而如果有人联系过去一年加入我团队的人,他们可能更有能力帮忙,并且了解足够多的背景,如果项目有前景,可以上报给我这样的人。像 Redwood 和 Mesa 这样的外部组织做得很好的一点是,他们在整个研究过程中与实验室的人大量交流,这意味着他们有很好的模型。双方可以有对话,这比研究做完后才试图引起注意更可能成功。虽然说起来容易做起来难。
Yeah, really good question. So one thing that's quite useful for people to bear in mind is that if we want to convince Gemini to use some safety technique, me and the GDMJ safety team are much better placed to do this than anyone outside because we have a good model of what the team cares about. For example, we know to check how much the monitor will trigger on harmless prompts. We know the evaluations people care about which aren't always public. We can work directly with the models and implement them. This means the goal of people outside should be to do sufficiently convincing work that lab safety teams think it is a good enough use of their time to try to flesh out the evidence base. So what does this mean on a practical level? A lot of what I said about getting a thing used in production still applies. You should think about the number of stakeholders who'd need to agree, how complex your technique is, how expensive, what the side effects are. Try to produce the best evaluations you can. But you also want to make sure lab safety teams are paying attention. Many people doing safety research know someone at a lab. Getting their takes can often be really helpful. People sometimes hear this advice and think they need to reach out to leads and senior people. But if someone emails my manager Rohan Sha, he's very busy and doesn't have much time to help. While if someone messages someone who joined my team in the past year, they might be more able to help and also know enough to get useful context and escalate it to someone like me if it's promising. One thing that external orgs like Redwood and Mesa do very well is they talk a bunch to people at labs throughout the research process, which means they have good models. There can be dialogue both ways, and this is much more likely to go well than just trying to get people to pay attention after you've done the research. Though easier said than done.
像 Google DeepMind 这样的公司面临的一个有趣挑战是,随着我们快速进入这个由 AI 主导的未来,涉及众多产品、商业竞争和地缘政治竞争,所有公司都被迫在加速的时间表上就内部治理和安全措施做出可能非常重大的决定,可能没有足够的时间来分析所有问题。你认为公司内部的人如何帮助引导这些过程走向积极的结果?这是否是人们应该考虑作为产生积极影响的重要方式,还是说这种心态不对?
An interesting challenge that companies like Google DeepMind face is that as we're barreling forward into this AI-dominated future with many different products, a commercial race, a geopolitical race, all companies have been forced to make monumental decisions about internal governance and safeguards, on an accelerated timescale with potentially not enough time to analyze all questions. How do you think someone inside the company can help steer those processes towards positive outcomes? Is that something people should have in mind as a significant way to have a positive impact, or is that the wrong mindset?
是的,这是个好问题。我认为如果你是一个关键决策者信任并听取意见的人,你可以产生很大的影响。这不一定表现为玩弄政治。我认为被视为一个中立可信的技术顾问有很大的潜在影响力,这个人真正懂行,理解安全,但也不意识形态化。安全社区有时有问题,认为必须总是把每件事都夸大其词地说成危险,因为这是人们会听的唯一方式。我非常关心 DeepMind 安全团队要指出哪些是危言耸听,哪些是真正令人担忧的,因为只有这样,当真正令人担忧的事情发生时,人们才会听。我认为人们常常忽略这一点:你可以只是一个受人尊敬的技术顾问。这并不意味着不擅长理解官僚机构如何运作并驾驭它们,而且确实需要花很多时间不做研究。但它更像是识别关键决策者,为自己建立聪明、能干、深思熟虑的声誉,成为一个被征求意见的人。比如我的经理 Rohin Shaw 就很棒,我认为他通过这种方式对帮助 DeepMind 走向更安全的方向非常有影响力。
Yeah, that's a great question. I think there is a lot of impact you can have if you are someone that key decision makers trust and listen to. This doesn't necessarily look like playing politics. I think there's a lot of potential impact in being seen as a neutral trusted technical adviser, someone who really knows their stuff, understands safety, but also isn't ideological. The safety community sometimes has issues with thinking it must always hype everything up as dangerous because that's the only way people will listen. I care quite a lot about the DeepMind safety team calling out things that are scaremongering for scaremongering and things that are actually somewhat concerning as actually somewhat concerning, because that's the only way you get people to listen when something actually concerning is happening. I think people often miss this: you can just be a well-respected technical adviser. This doesn't mean not being good at understanding how bureaucracies work and navigating them, and it does involve a lot of time spent not doing research. But it looks much more like identifying the key decision makers, building a reputation for yourself as smart and competent and thoughtful, and being someone who gets called on for advice. So like my manager Rohin Shaw is fantastic and I think very influential at helping DeepMind go in safer directions largely via this kind of approach.
是的。今年早些时候我和 Beth Barnes 聊过,我认为她对人们进入这些大型 AI 公司并引导他们做出更好决策的能力持有点悲观的态度。部分原因可能是她更偏技术,她和许多有机器学习技术背景的人聊过,这些人进入这些角色时希望影响政治或管理层面的决策。但她认为,要影响这些事情,你必须专门掌握与人沟通、围绕你想要的结果建立共识的技能。
Yeah. When I spoke with Beth Barnes earlier in the year, I think she had a slightly pessimistic take about the ability for people to go into these very large AI companies and steer them towards better decisions. I think in part because she's more of a technical person and she's spoken with many people with machine learning technical backgrounds who went into roles hoping to influence political or management decisions. But she thinks to influence those things, you have to specialize in the skill set of speaking to people, building consensus around the outcomes you want.
那些专注于做好机器学习研究的人,首先他们大部分时间都花在研究上,而且他们可能并不具备那种内部或组织政治方面的优势,而这种优势能让你处于最佳位置去改变项目所做的战略或安全决策。这种怀疑的看法你认同吗?
And people who are focused on doing machine learning research well, firstly, they just most of their time is spent doing the research, and they potentially don't necessarily have all the strengths in the kind of internal or organizational politicking that would put you in the best position to change the kinds of strategic or safety decisions that the project is making. Does that skeptical take resonate with you?
是的,这很复杂。我认为这确实有一定道理。有时我看到人们只是说,‘好吧,我关心安全,所以我应该去前沿 AI 公司工作,这样事情就会变好。’我会想,‘这通常比什么都不做要好,但我认为你可以通过制定计划做得更好。’一种计划就是我刚才概述的那种,要么非常擅长在组织内周旋,要么成为广受尊敬的技术专家。但 DeepMind 内部对安全产生影响的绝大多数人并不是这样做的。相反,更像是少数资深人员与组织其他部分对接。而我的团队中的人能产生的影响,很大程度上是通过做好研究来支持那些推动更安全变革的人。因为优秀的工程师和研究人员可以做到:证明某项技术有效、发明新技术、降低成本、为技术的有效性建立良好的证据基础并表明它没有大的副作用或成本,或者直接实现该技术。这样决策者就更容易说‘好,我们接受你已经为我们写好的代码’。这在某种意义上比成为在组织内周旋的人更耗时。而且我认为,如果组织中有受人尊敬的人——比如我非常尊重 DeepMind 的一些安全领导,如 Rohin Shah 和 Anca Dragan——那么努力赋能这些人会很棒。我也认为赋能一个团队可以非常有影响力。我们最近在 AG80 团队中成立了一个小型工程团队,其工作就是加速每个人的研究。你可能会想,‘啊,这是一家成熟的技术公司,肯定没有什么可以改进的了。’但有很多方法可以改进一个具有相当特殊用例的小团队的研究迭代速度。我认为这个团队已经非常有影响力了。负责该团队的 Vikrant 实际上让我打个广告:人们可能没有意识到的是,如果你是一名优秀的工程师,一个在复杂技术栈中工作快速且胜任的人,并且对加速研究感到兴奋——尤其是如果你在 Google 工作——即使你对研究或机器学习一无所知,也能产生相当大的影响。如果你符合这个描述,请务必给我发邮件。
Yeah. So it's complicated. I think there's definitely some truth to that. Sometimes I see people who just say, 'Well, I care about safety, so I should go work at a frontier AI company, and this will make things good.' I'm like, 'This is probably better than nothing in general, but I think you can do much better by having a plan.' One type of plan is the kind of thing I just outlined of either be very good at navigating an org or be a widely respected technical expert. But the vast majority of people having an impact on safety within DeepMind are not doing that. Instead, it's more like there are a few senior people who kind of interface with the rest of the org. And a lot of the impact that, say, the people on my team can have is by just doing good research to support those people who are pushing for safer changes. Because good engineers and good researchers can do things like show that a technique works, invent a new technique, reduce costs, build a good evidence base for the effectiveness of the technique and that it doesn't have big side effects or costs, or just implement the technique. So it's a lot easier for decision makers to say yes, we will accept the code you have already written for us. And this is a lot more time consuming in some senses than being the person who navigates the org. And I think that if there are people that people respect in the org, and like I have a lot of respect for some of DeepMind safety leadership like Rohin Shah and Anca Dragan, just trying to empower those people can be great. I also think that trying to empower a team can be very impactful. We've recently on the AG80 team started a small engineering team whose job is just trying to accelerate everyone's research. And you might think, 'Ah, this is like a mature tech company. Surely there's nothing that could be improved.' But there's just a lot of ways you can improve the research iteration speed of a small team with fairly idiosyncratic use cases. And I think this team has been super impactful. Vikrant who runs that team actually asked me to give a plug that something people may not realize is if they are a good engineer, someone who's fast and competent at working within deep complex tech stacks and feel excited at the idea of accelerating research, especially if you work at Google, you can have a pretty big impact even if you don't really know anything about research or ML. And you should totally shoot me an email if that describes you.
我认为所有领先 AI 公司内部的决策者面临的一个困难是,真的极难知道当前需要什么样的保障措施,或者明年、未来五年需要什么样的保障措施,而哪些措施实际上不会对安全产生实质性影响,可能只是浪费时间和投资,并使你处于商业劣势。听起来你在说,作为公司内部的技术专家,你可以做一件非常有用的事情:提供准确的信息,告知决策者哪些是真正正在显现的威胁,以及哪些保障措施实际上有助于化解这些威胁。这比仅仅不断地说‘我们应该总是做更安全的事情’(因为那是你的整体世界观)更有可能建立善意并真正影响决策。
I think one thing that must be difficult for decision makers inside all of the leading AI companies is that it's just legitimately extremely hard to know what sort of safeguards are required at this point or are going to be required in the next year or are going to be required in the next 5 years versus which ones are just not actually going to move the needle on safety and might just be kind of a waste of time and investment and put you at a commercial disadvantage. And it sounds like you're saying something incredibly useful that you can do as a technically savvy person inside the companies is to provide accurate information to inform decision makers about what are the real threats that are panning out and what sort of safeguards would actually help to diffuse them. And that is much more likely to build goodwill and actually influence the decision than just like constantly saying we should always be doing the safer thing because that's kind of your overall worldview.
是的,我认为这是一个很好的总结。我也认为‘总是做更安全的事情’实际上是一种极其简化的观点。像前沿语言模型这样极其复杂的系统,可以变得更安全的方式空间非常大。我认为产生影响的部分技巧就是找到那个共同点。需要明确的是,我并不是说人们只应该做与当今模型直接相关的安全研究。相反,我的意思是它最终需要相关。我要对你刚才说的补充一点:让一项技术在生产模型中使用,比安全团队决定研究它要困难得多。我们有相当大的自主权,可以用我们最好的判断来决定研究什么。如果我们认为某件事在 6 到 12 个月内会变得重要,人们会看到紧迫性,那么我们现在就可以开始准备。而且我认为,总的来说,有句名言说,实现政策变化的方法之一就是等待危机。我强烈希望当安全问题出现时,我们有好的解决方案准备就绪,因为我个人认为许多安全问题会有某种警告信号和初步前兆,在测试或类似情况中出现。如果你因为问题阻碍了发布而紧急尝试修复,你所能产生的解决方案往往不如你提前准备的那样深入和有效。我们产生影响的途径之一就是这种准备。
Yeah. I think it's a pretty good summary. I also think the 'just do the safer thing' is actually an incredibly oversimplistic perspective. The space of ways that an incredibly complex system like a frontier language model can be safer is just really big. And I think part of the skill of having an impact is finding that common ground. And to be clear, I'm not just saying people should only do safety research that is directly relevant to today's models. Rather, what I'm saying is it will need to eventually be relevant. One caveat I may add to what you just said is it's a lot harder to get a technique used in a production model than it is for the safety team to just decide to work on it. We have quite a lot of autonomy to use our best judgment for what to work on. And if we think that something's going to be a big deal in 6 or 12 months and people will see the urgency then, we can start to prepare that right now. And I think that in general, there's this famous saying that one of the ways to get policy change to happen is to wait until a crisis. And I would strongly prefer that when safety issues arise, we have good solutions ready to go because I personally think that many safety issues will have kind of warning signs and tentative precursors that crop up in testing or things like that. And the thing you can produce to fix it if you're really urgently trying to do it because it is blocking a launch is often kind of less deep and effective than what you can do if you prepare right. And one of our paths to impact is this kind of preparation.
你说的危机是什么意思?你是在想象未来的某个时候,产品正在做一些公司不希望的事情,人们争相寻找解决方案,使其能够以更安全的方式继续提供给用户,而你希望提前准备好可靠的技术,这些技术实际上能解决问题。把这些技术打包好,可能在未来的某个时候,当实际使用它们的需求更大时,随时可以投入使用。
What do you mean by crisis? You're sort of talking about we're imagining sometime in future when the product is doing things that the company doesn't want and people are scrambling to find solutions that will allow it to continue to be served to users in a safer way, and you kind of want to prepare techniques ahead of time that are sound, that actually would address the problem. Kind of have those packaged and potentially ready to go at some future time when there's much more demand for actually using them.
是的。所以在某种意义上,我认为如果发生这种情况,安全团队就没有做好自己的工作。我希望的是,这些未来可能成为危机的事情能够提前被标记出来,让人们通过他们认为真实的评估注意到,这样我们就能在任何实际事件发生之前修复它们。我非常喜欢这些风险管理框架,比如 DeepMind 的前沿安全框架、Anthropic 的负责任扩展政策,因为我将它们视为按我们自己的方式创造可控危机。例如,Gemini 2.5 Pro 触发了早期预警信号,因为它能够帮助人们进行攻击性网络能力方面的活动。
Yeah. So in some sense I think that if that happens, the safety team has failed to do their job right. What I want is for these things that would be future crises to kind of be flagged ahead of time in a way that people notice on evaluations they think are real, such that we can fix them before any actual incidents happen. I'm a really big fan of these risk management frameworks like DeepMind's frontier safety framework, Anthropic's responsible scaling policy, because I kind of think of them as creating managed crises on our own terms. So, for example, Gemini 2.5 Pro triggered the early warning sign for being able to help people do things with offensive cyber capabilities.
攻击性网络能力。
Offensive cyber capabilities.
攻击性网络能力。没错。
Offensive cyber capabilities. Exactly.
我认为我们做这些测试并注意到这一点非常好,因为这样就有了一个阈值,我们说过届时会采取缓解措施。所以现在有很多精力和兴趣来提前开发好的缓解措施,准备在我们触发‘如果缓解不当可能会很危险’的情况时使用。我们现在已经在我们的前沿安全框架中加入了欺骗性对齐。随着时间的推移,这些风险的证据基础越来越强,就更容易将目前更理论性的风险纳入其中。我的意思是,我们已经看到了像对齐造假这样的现象。如果我们能提前识别出风险是什么,并就好的评估方法以及达到阈值时需要什么样的缓解措施达成一致,那么这就能产生类似危机的效果,而无需承受危机的所有负面影响。
And I think that it's very good that we do these tests and notice this because there is then a threshold where we've said we'll have mitigations in place. So there's now a lot of energy and interest in developing good mitigations in advance ready to go by the time we trigger the 'this could actually be dangerous if it wasn't mitigated well'. And we now have deceptive alignment in our frontier safety framework. And as time goes on and the evidence base for these risks starts to get stronger and stronger, it becomes much easier to put the kind of currently more theoretical kinds of risks in there. I mean, we're already seeing things like alignment faking. And if we can identify ahead of time what the risks will be and agree on good evaluations for them and on what kind of mitigations are needed when that threshold is reached, then that creates the same kind of effect as a crisis without all of the bad effect of having a crisis.
你认为对于一个相当注重安全且机器学习能力很强的人来说,他们去像 Google DeepMind 这样已经很注重安全的公司工作,是否可能产生更大的影响?或者,他们去一个安全文化较弱、安全投入较少的公司,在那里他们可能是少数将安全作为重要个人优先事项的人之一,这样是否可能产生更大的影响?人们是否应该认真考虑去那些不太注重安全的实验室?
Do you think that for someone who is reasonably safety focused and has strong machine learning chops, are they likely to have more impact by going and working at a company that is already quite safety focused like Google DeepMind, or perhaps could they have even more impact by going and working at a company where there's maybe less of a safety culture and less safety investment, and they might be one of relatively few people who have that as a significant personal priority? Should people seriously consider going to the less safety focused labs?
是的。所以我认为这介于两者之间。我非常关心每个制造前沿 AI 系统的实验室都拥有优秀的安全团队,无论该实验室的领导理念如何。但我认为,在那些更难推动这类事情且兴趣较低的实验室里,某些类型的人可以产生很大影响,但许多人基本上不会。我在这样的实验室里产生影响的方式是,我认为那些只想做出更好模型的人和那些想做出更安全模型的人之间有很多共同点。世界不一定充满权衡。例如,我希望任何实验室的人都能从事诸如监控模型的思维链以减少奖励黑客之类的工作。在我看来,这相当具有商业适用性,而且我认为这类事情很自然地会流向更安全相关的事情,同时与当前的安全也有一定关联,因为监控模型做你不喜欢的事情的思维链是非常通用的。如果你想实施更精细的控制方案,这是一个很好的基础。
Yeah. So I think it's somewhere in the middle. I think that I care a lot about every lab making frontier AI systems having an excellent safety team, whatever the leadership philosophy of that lab is. But I think that in the kinds of labs where it's harder to get this kind of thing through and there's less interest, I think there are certain kinds of people who can have a big impact, but many people will largely not. The kinds of ways I'd approach having impact in a lab like that is I think there's just a lot of common ground between people who just want to make better models and people who want to make safer models. The world is not necessarily just full of trade-offs. For example, I would love people in any lab to be working on things like monitoring the chain of thought of a model to reduce reward hacking. This is pretty commercially applicable in my opinion, but also I think it's the kind of thing that very naturally flows into much more safety relevant things, while also being somewhat relevant to safety today because monitoring the chain of thought for the model doing things you don't like is very general. If you wanted to implement some more elaborate control scheme, that's an excellent basis.
听众如何判断自己是否适合这样做?
How could a listener tell if this might be a fit for them?
我认为擅长这件事的人是那种有在组织中有效工作和周旋经验的人。有主动性,能够发现并为自己创造机会的人。能够与一群在相当重要的事情上与自己意见相左的人一起工作而不会感到不适的人。理想情况下,不会觉得每天上班都咬牙切齿。只是,你知道,这是一份工作。这些人,我不同意他们的观点。我可以把分歧放在一边。我认为这是一种更健康的态度,不太可能导致倦怠,也会让你在外交上更有效。最后,是那些善于独立思考的人。我个人觉得很难保持与周围人截然不同的观点。我可以,但需要主动努力。有些人非常擅长这个,有些人比我差得多。如果你处于,我不知道,我的水平或更差,可能你不应该这样做。
I think the kind of person who would be good at this is someone who has experience working in and navigating organizations effectively. Someone who has agency, who's able to notice and make opportunities for themselves. Someone who's comfortable at the idea of working with a bunch of people who they disagree with on maybe quite important things. Where ideally it doesn't feel like you're grinding your teeth every day going to work. It's just, you know, this is a job. These are people. I disagree with them. I can leave that aside. I think that's a much healthier attitude, less likely to lead to burnout and will also just make you more effective at diplomacy. And finally, people who are good at thinking independently. I personally find it a bit hard to maintain a very different perspective from the people around me. Like I can, but it requires active effort. Some people are very good at this, some people are much worse than me at this. If you are, I don't know, my level or worse, probably you shouldn't do this.
对于那些有兴趣在 AI 领域发展并希望被这些前沿 AI 公司录用的人,你有什么建议?你能提供一些通用建议吗?
What advice do you have for people who are interested in pursuing a career in AI and getting hired by one of these frontier AI companies? Is there any general advice that you can offer?
我认为人们真正关心的两个关键技能是研究能力,尤其是经过验证的研究经验,以及工程技能。部分只是软件工程技能,部分则是机器学习工程技能。机器学习现在不那么重要了,因为现在不是每个人都在训练自己的模型,而是大规模的训练运行,人们负责不同的部分。这里还有几点补充。我最初进入这个领域时,对工程技能的含义感到相当困惑。人们最初接触的往往是像 LeetCode 这样的东西,这些在编程面试中问到的简单谜题,你甚至可能不用写代码,只需解释解决方案。在我看来,这实际上并没有那么有用,而且是一种非常特定的工程技能。在科技公司工作非常重要的是一种我认为是深度经验的工程技能。它是在一个庞大复杂的代码库中工作的能力,这个代码库有数百人同时在开发,你依赖无数其他人的代码,你无法真正一次性记住整个代码库,你需要知道什么时候深入挖掘找出问题所在,什么时候进行抽象。如今,你需要知道如何最好地利用 LLM 来帮助你导航。我想区分这一点,因为我认为如果不直接在这种科技公司工作,这种技能很难获得。它不是必需的。如果你是一个足够出色的研究员,你可以弥补,但它非常有价值。还有机器学习工程,特别是能够编写高效的机器学习代码、训练模型,这项技能意味着要让它高效,同时还要在它不可避免地以令人困惑的方式出错 100 次时进行调试,与普通编程不同,你不会得到漂亮的错误跟踪告诉你哪里出错了。相反,它只是表现不佳。我确实认为这可能正在发生变化,因为我认为特别是在过去 6 到 12 个月里,编码智能体已经变得非常出色,我们稍后可能会讨论这一点。但我认为这些更高级的工程技能可能更难自动化,并且在更长时间内仍然有用。在研究方面,我推荐的一种思考方式是,论文是一种便携式的凭证。无论你在哪里写了论文,或者你是在做非常著名的博士研究,还是一个随机的独立研究者,这都不重要。
I think the two key skills that people really care about are ability to do research and especially proven experience doing research, and engineering skill. Partially just software engineering skill and partially machine learning engineering skill. Machine learning is a bit less important nowadays given that it's less everyone training their own models and more like massive training runs and people having different pieces. A few additional thoughts here. I was quite confused about what engineering skill meant when I first got into this field. People's first exposure is often things like LeetCode, these simple puzzles you get asked in coding interviews where maybe you don't even write code, you just explain the solution. This is in my opinion not actually that useful and a very specific kind of engineering skill. Something that matters a lot for working at a tech company is what I think of as deep experienced engineering skill. It's the ability to work in a large complex codebase where hundreds of people have been working in the same codebase and you're depending on countless other people's code and you can't realistically keep the whole thing in your head at once and you need to know when to go diving deep to figure out what's going wrong and you need to know when to abstract something. Nowadays you need to know how best to use LLMs to help you navigate this. I want to distinguish this because I think that this is just a much harder skill to gain without just directly working in that kind of tech company. It's not essential. You can make up for it if you're just a fantastic enough researcher, but it is pretty valuable. There's also ML engineering, in particular, things like being able to write efficient ML code, train things, where really that skill means get it to be efficient, but also debug it when it inevitably goes wrong 100 times in really confusing ways where unlike normal programming, you don't get these nice error traces telling you what went wrong. Instead, it just performs bad. I do think this might be changing somewhat because I think especially over the last 6 to 12 months, coding agents have gotten really good in a way we might discuss a bit later. But I think that these kind of more senior engineering skills are likely to be much harder to automate and remain useful substantially longer. On the front of research and research experience, one way I recommend thinking about it is a paper is kind of a portable credential. It doesn't matter where you wrote the paper or whether you were doing a really prestigious PhD or a random independent researcher.
如果你做出了好的工作,并且人们有足够的时间意识到这是好的工作,他们就会在意。人们常常依赖启发式判断。比如,如果你在知名实验室读博,那会让人们更愿意花时间关注你。但通常博士学位并不那么重要,只要你有良好的研究记录。如果你做的是与前沿语言模型相关的研究,人们会更在意。另外,让实验室里有人了解你的工作并认为它很酷,这对你非常有利。比如联系你欣赏其工作的实验室成员,告诉他们你的研究,或者在会议上与他们见面,或者在你申请公开职位时联系他们——这通常没什么用,因为人们会收到大量邮件,但我认为设定预期非常有用。
If you have done good work and people have enough time to realize it is good work, they care. People often rely on heuristics. For instance, if you've done a PhD at a prestigious lab, that makes people more likely to spend time paying attention. But often PhDs don't really matter that much as long as you have a good research track record. People will care more if you do research relevant to frontier language models. It can also be very in your interest to have people at the lab who know about your work and think it is cool research. Things like reaching out to people at the lab whose work you admire, telling them about your work, or trying to meet them at conferences, or reaching out when you apply to an open job ad—often that doesn't work because people get lots of emails, but an expectation I think is extremely useful.
所以我觉得你关于如何在机器学习研究中取得成功以及如何被 AI 公司雇用的回答中,一个共同的主题就是成为一个高效、有思想的研究者。我想没人会否认你在过去四年里发表了大量令人印象深刻的研究。你现在在 Google DeepMind 领导一个八人团队,过去四年里还指导了大约 50 名初级研究员。关于什么造就了一个好的研究者,以及他们倾向于采用哪些实践来取得大量成果,你学到了什么?
So I guess a common theme in some of your answers about how to be successful in ML research and how to get hired by an AI company is just to be a productive, thoughtful researcher. I think no one would deny that you've published an impressive amount of research over the last four years. You're also leading a team of eight people now at Google DeepMind, and you've supervised about 50 junior researchers over the last four years. What have you learned about what makes a good researcher and what practices they tend to adopt to get a lot done?
是的。我认为人们常常有些误解,不清楚成为一个好研究者需要哪些具体技能和心态。我们肯定会进一步探讨这一点。但在高层次上,你需要擅长编码,这样才能完成任务和做实验,尽管通常只要擅长在交互式环境(如 Python 笔记本)中进行非常粗糙的小规模编码就足够了,而不需要构建复杂的基础设施或使用复杂的库。所以关键技能是能够快速粗糙地完成工作并做对。与此相关的是快速迭代的能力。那些能够快速行动、快速迭代、看到新结果并做出决定的人,他们的产出和成功率差异非常显著。这引出了另一个相关点:优先级排序。研究是一个极其开放的空间,比许多其他工作更需要你在不确定性下行动。这既是心理层面的问题——有些人觉得这比其他人难得多——也是知道何时深入问题、何时退一步的问题。人们常常一开始就过于偏向其中一个方向。另一个关键点,尤其是在机器学习中,是实证科学的心态和怀疑精神。你试图导航一个巨大的决策树,你会得到实验证据,需要解读它。你需要理解这些证据告诉你什么可能是真的、什么可能不是,你的证据是否有弱点和缺陷,你应该更努力推进还是不能更新,以及何时证据足够好可以继续前进。能够回顾你过去几周的研究,对你可能非常坚持的假设进行红队测试——思考它们可能如何有缺陷,以及你可以做哪些实验来检验它们的稳健性。最后一个关键技能我认为很复杂且常被误解。我称之为研究品味,基本上意味着对研究项目中正确的高层决策有良好的直觉。这可能包括:我该选择什么问题?因为有些项目是好主意,有些是坏主意。如果你选了好项目,你会很开心。部分在于知道项目内哪些方向最有前景、最有趣。部分更底层,比如你能否设计一个好的实验来检验你关心的假设,这个实验既要可行、易于实现(因为你大致知道什么样的实验可行、什么样的难),又要真正触及你关心的东西。
Yeah. So I think people often somewhat misunderstand the exact skills and mindset that go into being a good researcher. I'm sure we'll unpack this more. But at a high level, you need to be decent at coding just so you can get things done and do experiments, though often it's sufficient to only be good at very hacky small-scale coding that you're doing in an interactive thing like a Python notebook, rather than building complex infrastructure or working within complex libraries. So the key skill is being able to be hacky and get things right. Related to this is the skill of just iterating fast. It's very striking how different the productive output and success rate is between people who can just do things fast and iterate fast and see the new results they have and decide. That gets me onto a related point of prioritization. Research is an incredibly open-ended space, and more so than in many other lines of work, you need to be good at acting under uncertainty. This is both a psychological thing—some people find this a lot harder than others—but also knowing when to dive deeper into a problem, when to zoom out. People often start off too far in one of those two directions. Another key thing, especially in machine learning, is this empirical scientific mindset and this notion of skepticism. You're trying to navigate this incredibly large tree of possible decisions you could make, and you will get experimental evidence and need to interpret it. You need to understand what this tells you about what could and could not be true, whether there are weaknesses and flaws in your evidence, and you should actually push harder and you can't really update it versus when it's good enough and you can move on. Being able to look back on your last weeks of research and take the hypotheses you're maybe quite attached to and try to red team them—think about how they could be flawed and what experiments you could do to test their robustness. The final key skill I think is complex and often misunderstood. It's what I call research taste, which basically means having good intuitions for what the right high-level decisions are going to be in a research project. This can mean what problems do I choose? Because some projects are a good idea, some are a bad idea. If you choose a good project, you're happy. Some of this is knowing what directions within a project are most promising and most interesting. Some of it is more low-level, like can you design a good experiment to test the hypothesis you care about that is both tractable and easy to implement because you kind of know what that looks like and what hard things look like, but also actually gets at the thing you care about.
你认为人们对研究品味有什么误解?
What do you think people misunderstand about research taste?
我认为人们误解的一点是,他们以为研究品味只关乎选择项目,但实际上它适用于多个层面。我有一篇博客文章,希望可以放在描述里,试图解释这个术语的真正含义。其中一个特别混乱的地方是,这些技能有不同的反馈循环。对于编码,你可以在几分钟或几小时内知道代码是否有效。对于可解释性的概念理解,你通常可以在几小时内弄清楚——你可以读论文,或者请教更有经验的人。然后像优先级排序这样的事情,你常常要过很久才知道它是不是个好主意。而研究品味,比如设计实验——你可能在几天到几周内得到反馈。项目内的方向是否正确?几周甚至几个月。项目本身好不好?几个月。我认为很有启发性的想法是,把自己想象成一个神经网络。你试图训练你的直觉,使其善于对研究问题做出预测,因为并没有一门科学能告诉你这个研究问题是否会成功。最终,它确实归结为直觉。尽管直觉通常基于逻辑、论据和证据,你希望获得高质量的数据和大量的数据。你应该预期,在早期,你会更快地学会简单的技能,而研究品味则慢得多。所以我给刚进入研究领域的人的建议是,不要担心研究品味。它非常痛苦。先学习其他技能,然后研究品味会更容易掌握。理想情况下,这表现为找一个导师,利用他们的研究品味。但如果你没那么幸运,这可以表现为做一些相当没有野心的项目,对论文进行渐进式改进,或者尝试一些随机想法,你并不真正关心它们是不是好主意。你只是想尝试做点什么,如果结果不好,你愿意放弃。
I think one thing people misunderstand is they think that it's only about choosing the project when actually it applies on many levels. I have a blog post we can hopefully put in the description trying to unpack what this term actually means. One of the particularly messy things is that these skills have different feedback loops. For coding, you can kind of tell if your code worked in minutes or hours. For conceptual understanding of interpretability, you can often figure out within hours—you can read papers, maybe ask someone more experienced. Then there are things like prioritization where you often don't know if it was a good idea for quite a long time. And there's research taste where things like designing an experiment—maybe you can get feedback within days to weeks. Was this a good direction within a project? Weeks maybe months. Was this a good project? Months. I think it's quite instructive to think of this like you're a neural network. You're trying to train your intuitions to be good at making predictions about research questions, because there's not really a science of will this research question work out or not. Ultimately, it does come down to intuition. Though often the intuition is grounded in logic and arguments and evidence, and you want to get high quality data and lots of data. You should expect that early on you are going to learn the easy skills much faster and research taste much slower. So my advice to people getting into research is just don't worry about research taste. It's a massive pain. Learn the other skills first and then it will be much easier to learn research taste. This ideally looks like finding a mentor and using them for research taste. But if you're not lucky enough to have that, this can look like just doing pretty unambitious projects, incremental improvements to papers or random ideas where you don't actually care if they're good ideas or not. You just want to try doing something, you're willing to give it up if it doesn't turn out very well.
你在笔记中提到研究过程有三个不同阶段,听起来你认为人们可能会对自己处于哪个阶段感到困惑,并可能对他们每天所做的工作抱有错误的态度。你能解释一下这三个阶段是什么,以及每个阶段需要什么样的不同心态吗?
You mentioned in your notes that there are three different stages to the research process, and it sounds like you think people can get a little bit confused about which one they're in and potentially have the wrong attitude to the work that they're doing on a given day. Could you explain what the three different stages are and what the different mentality you need to bring to each?
是的。我把它们称为探索、理解和提炼。探索是指从有一个研究问题到深入理解问题以及哪些事情可能或不可能为真的过程。我认为人们常常没有意识到这是一个阶段,尤其是在可解释性中,它往往占据项目的一半以上。你要找出该问的正确问题,找出你可能有的误解。然后一旦你对问题有了某种假设——比如你从认为某个行为可能有趣,比如为什么这个模型会自我保存?探索会达到你心想‘嗯,我有一个假设:它感到困惑’的程度。然后理解是指从假设出发,试图证明一些结论,或者至少提供足够的证据来说服自己。探索是非常开放的。那里的北极星是‘我想获取信息’,这通常不需要非常定向。而当你理解时,你通常要有一个具体的假设,并制定一个计划,看哪些实验能为你提供支持或反对的证据。以自我保存为例,你可能会想,‘啊,如果模型感到困惑,那么改变提示的这部分可能会让它不那么困惑。’然后提炼是指从你相当确信这件事为真,到将其传达给世界,使其更加严谨,并变成那些没有深入项目细节的人也能信服的东西。
Yeah. So, I call them explore, understand, and distill. Explore is when you go from having a research problem to a deep understanding of the problem and what kinds of things may or may not be true. I think people often don't realize this is a stage, but especially in interpretability, it's often more than half of a project. You figure out the right questions to ask, figure out the misconceptions you might have had. Then once you have some hypothesis for what could be true about your problem—like you go from thinking that there would be something interesting about some behavior, like why does this model self-preserve? Exploration would be getting to the point where you're like, 'Hmm, I have the hypothesis that it's confused.' Then understanding is when you go from a hypothesis to trying to prove some conclusions or at least provide enough evidence to convince yourself. Exploration is very open-ended. The north star to have in mind there is 'I want to gain information,' which often doesn't need to be very directed. While when you're understanding, you often want to have a specific hypothesis in mind and form a plan for what experiments would get you evidence for and against this. In the case of self-preservation, you might be like, 'Ah, well, if the model's confused, then changing this part of the prompt might make it less confused.' And then distillation is when you go from being pretty convinced this thing is true to communicating it to the world, making it more rigorous, and turning it into something that someone who hasn't been deep in the guts of the project might be convinced by.
那这又是如何让人陷入困境的呢?
And how does that trip people up?
在实践中,人们似乎常常没有真正意识到探索是一个阶段。如果他们不知道要回答的假设,或者头脑中没有关于项目走向的清晰计划或影响理论,他们就会觉得自己在失败。这意味着我认为人们常常卡住或感到压力。还有一种错误是人们陷入死胡同,因为他们认为自己找到了正确的东西去研究,但实际上他们对问题了解不够。探索并不意味着只是瞎搞。你可以理性地思考单位时间内获得的信息量。你可以使用某些工具,比如阅读模型的思维链,这可能会给你非常丰富的数据。有些工具可能不会给你丰富的数据,比如测量某种技术的准确率——那只是一个数字。我认为这是为了获得表面积。你要做那些非常开放、有很多不同东西可以观察和思考的技术。你要最大化迭代速度。如果有一堆廉价但不那么严谨的技术,你应该快速做这些,而不是做一个高投入的事情。那是针对理解的。你不需要严谨,因为目标是形成想法和假设。通常,选择一个微假设,然后尝试测试它,再在一周内放大视角是合理的。一个实际的做法是,我建议人们每天或每两天检查一次:‘我学到了什么吗?’如果答案是否定的,考虑改变。然后探索需要不同的心态。也许有人很擅长瞎搞,但他们不擅长真正进入科学模式去做好的实验。我认为人们常犯的一个错误是他们太努力去做很多实验或非常全面的实验,而不是思考什么样的关键实验能真正证明事情有效。然后理想情况下你也做其他实验,但没错,你必须同时考虑质量和数量。然后是提炼:我认为人们犯的错误是,你在一个研究项目中花了数周甚至数月的时间,一切都对你很清楚。这对其他人——那些真正重要的人,因为他们会读你的工作——并不清楚。你需要站在别人的角度,提供所有相关的背景。我对此的实际建议是向人们寻求反馈。把你的想法告诉他们。把草稿给他们看。看看什么让他们困惑。看看他们不理解什么,并优先解决这些问题。另一个非常有帮助的事情是,非常明确你论文的要点是什么。人们常常认为他们需要讲述项目中所有事情的故事,但实际上人们只会真正记住论文中的几个想法。我所说的叙事:就是几个想法,然后为什么你支持它们——比如为什么这些想法是真的,为什么你应该关心这个?你开始论文时只解释想法和非常简短的证据。然后你详细解释想法。然后你详细解释证据。但所有这些都集中在这个目标上:让人们记住并相信这些想法,这个叙事。我有各种关于如何做研究的博客文章,包括一篇关于如何写论文的,我希望会放在描述里。
In practice, people often don't seem to really get that exploration is a stage. They feel like they're failing if they don't know the hypothesis they're trying to answer or they don't have a clear plan or theory of impact in their head for where this project is going. This means that I think people often get stuck or feel stressed. There's also the mistake where people get stuck in rabbit holes because they think they found the right thing to look at, but actually they just don't know enough about the problem. And exploration doesn't mean just messing around. You can think rationally about the information gained per unit time. You can use certain kinds of tools, like reading the model's chain of thought, that might give you pretty rich data. There are certain tools that are probably not going to give you rich data, like measuring the accuracy of some technique—that's just a number. I think of this as gaining surface area. You want to do techniques that are very open-ended and have loads of different things to look at and think about. You want to maximize your iteration speed. If there are a bunch of cheap techniques that aren't super rigorous, you should do those fast rather than one really high-effort thing. That's for understanding. And you don't need to be rigorous because the goal is to form ideas and a hypothesis. Often it makes sense to pick a micro hypothesis and then try to test it and then zoom out within a week. One practical implementation of all this is I recommend people check in once a day or every two days: 'Did I learn anything?' If the answer is no, consider changing. And then exploration requires a different mindset. Maybe someone's really good at messing around, but they're bad at really getting into the scientific mode of making great experiments. One mistake I think people often make is they try too hard to have lots of experiments or really comprehensive experiments rather than thinking about what would be the kind of killer experiment that just really establishes that the thing works. And then ideally you also do other experiments, but yeah, you got to think about both quality and quantity. And then distillation: I think people make the mistake that you've spent weeks to months in the weeds of a research project. Everything's clear to you. This is not clear to anyone else—the people who actually matter because they'll be reading your work. You need to meet people where they are and provide them all the relevant context. My practical advice for this is to get feedback from people. Run your ideas by them. Run a draft by them. See what confuses them. See what they don't understand and prioritize fixing that. Another thing which helps a lot here is being very deliberate about what the point of your paper is. People often think they need to tell a story about everything they did in a project, but actually people will only ever really remember a few ideas from a paper. What I think of as the narrative: it's a few ideas and then why you support them—like why are these ideas true and why should you care about this? You begin the paper by just explaining the ideas and very briefly your evidence. You then explain the ideas in detail. Then you explain your evidence in detail. But all of this is just centrally focused on this goal of having people remember and believe the ideas, this narrative. And I have various blog posts on how to do research, including one on how to write papers, which I hope will be in the description.
是的,听起来你在沟通阶段确实需要提炼。我的意思是,如果人们从你的论文中记住一句话,那么从某种意义上说你已经很幸运了。但这意味着你真的必须非常努力地思考你想传达给他们的那句话是什么。然后我想如果有人记住三句话,你知道,一个叙事的开始,那会是什么?因为你在让论文走红方面相当成功。你会花很多时间精心设计推文,或者精心设计你真正想要传播的论文核心点的提炼吗?
Yeah, I guess it sounds like you really have to distill things down at a communication stage. I mean, if people remember a single sentence from your paper, then in a sense you're quite lucky. But that means you do really have to think very hard about what is the sentence that you're trying to get through to them. And then I guess if someone remembers like three sentences, you know, the beginnings of a narrative, what might that be? Because you've had quite a bit of success getting your papers to go viral. Do you spend a lot of time crafting the tweet or crafting what is the distillation of the core point from the paper that you actually want to spread in the wild?
是的,相当多的时间。这也是我试图让我的 MATS 学者关注的事情。我可能想区分炒作和沟通。我认为关键的是人们理解论文中的想法以及为什么他们应该关心。这看起来可能和试图制造炒作一样,但实际上它是基于某些东西的。
Yeah, quite a lot of time. And this is something I try to get my MATS scholars to focus on. I maybe want to distinguish between hype and communication. I think the crucial thing is that people understand the ideas in the paper and why they should care. And this can kind of look the same as trying to build hype, but it's actually based on something.
一个常见的玩笑是,你应该在摘要、引言、图表、标题和其他所有内容上花费相等的时间。因为如果你用内容的长度乘以阅读它的人数,它们大致相等。
A common joke is you should spend an equal amount of time on the abstract, introduction, figures, title, and everything else. Because if you take the length of a thing times the number of people who will read it, they are about equal.
是的。
Yeah.
而且我还没有找到最好的方法,把项目的 20%时间花在把标题做好上。
And I have not yet figured out the best way to spend 20% of a project making the title good.
我们在节目中几乎接近这个比例了,因为对于 YouTube 或 Twitter 来说,提炼出为什么有人会想听某一集非常重要。我觉得我们还没到 20%,但在上升。
We are almost approaching that on the show because it's so important for YouTube or Twitter to distill why anyone would want to listen to a given episode. I think we are not at 20% yet, but it's going up.
我从你的过程中受到启发。至少,你应该列一个短名单,四处征求意见。不要只凭自己的直觉。
I've been inspired by your process. At the very least, you should make a short list, shop them around for feedback. Don't just go with your own intuition.
Twitter 是人们在机器学习领域交流研究的主要方式之一,无论好坏。现在也有对 BlueSky 的兴趣。如果你想写一篇关于论文的好推文串,关键要记住的是,95%的人不会读第一条推文之后的内容。所以第一条推文需要包含你的整个叙述。
Twitter is one of the main ways people communicate research in ML for better or for worse. Nowadays there is also interest in BlueSky. The key thing to have in mind if you want to write a good tweet thread about your paper is that 95% of people do not read beyond tweet one. So tweet one needs to contain your entire narrative.
这迫使你简洁。这很好。
This forces you to be concise. That is good.
这是 280 个字符的推文,不是新的整条线程推文。不要发很长的推文,因为人们需要点击展开,他们不会看的。
This is a 280-character tweet, not the new whole thread tweets. Don't do the really long tweets because then people have to click on them and they won't.
是的。
Yeah.
另一件事是有一张引人注目的图,以一种人们不需要太多背景就能理解的方式传达有趣的内容。这通常值得投入大量精力,而且通常也应该是你论文中的第一张图。再说一次,这不仅仅是炒作。我认为这实际上是一种非常有效的沟通形式。如果更多人看到你的工作并正确理解,那就太好了。
The other thing is have a single eye-catching figure that conveys something interesting in a way that people can understand without needing too much context. This is often worth a lot of effort and it should also typically be the first figure in your paper. Again, this isn't just about hype. I think this is actually a very effective form of communication. If more people see your work and get a correct understanding, that's great.
另一件要记住的事是,看到你推文的人只是在他们的信息流中看到它,夹杂在一堆其他内容中。他们甚至不一定知道你是机器学习研究员。所以你需要把它放在所有可能事物的思维空间中,使用他们可能认识的关键词,但如果你想让人理解,就不要用术语。有一句动机句,比如要解决什么问题。好的推文串是一门艺术,大多数人似乎并不具备。这很可悲。
The other thing to bear in mind is that people who see your tweet will just see it in their feed amongst a bunch of other stuff. They won't necessarily even know that you're an ML researcher. So you need to situate it in the mind space of all possible things, use kind of keywords they might recognize, but don't use jargon if you want people who don't know the jargon to understand it. Have a sentence of motivation, like what's the problem being solved. There is an art form to good tweet threads that most people do not seem to have. It's very sad.
所以你需要解释论文的全部结果,将其定位在领域中,传达你有必要的专业知识,并且在大约 280 个字符内不使用任何术语。
So you need to explain the full results of the paper, situate it in the field, communicate that you have the necessary expertise, and do it without any jargon in about 280 characters.
你不需要传达你的专业知识。
You don't need to communicate your expertise.
好的。
Okay.
我觉得直接说‘论文’就可以了。
I think it's fine to just be like 'paper'.
是的,而且永远不要把论文链接放在第一条推文里,也永远不要只截图摘要和作者名单及标题,因为那真的很无聊,没人会读。
Yeah, and never put a link to the paper in the first tweet and never just screenshot the abstract and list of author names and title because that's really boring and no one will read it.
人们会感受到一种强烈的压力,我自己有时也有这种感觉,就是你想在多大程度上炒作一个结果有多令人兴奋,或者说一个播客有多好,而不是坦诚地谈论内容和它的局限性。你自己如何权衡这种取舍?
There is a very strong pressure that people can feel, and I feel this myself sometimes, about how much you want to hype how exciting a result is or say a podcast is, versus be frank about the content and its limitations. How do you strike that trade-off yourself?
我不确定我自己总是能很好地把握这个平衡,但大致上,我认为在你的工作中突出地讨论局限性非常重要。理想情况下,如果有一个关键限制,它应该出现在摘要中,因为很多人只会读摘要。在我看来,如果人们对你的工作有错误的正面印象,那是不好的。
I'm not sure I always do an amazing job straddling this balance myself, but roughly I just think it is very important to talk about the limitations of your work prominently in the work. Ideally, if there's a key limitation, it should be in the abstract because many people will only read the abstract. In my opinion, it is bad if people have a mistakenly positive impression of your work.
我认为人们可能经常忽略的一个激励因素是,经验丰富的研究人员很擅长判断一篇论文是否只是充满炒作而没有实质内容,并且通常会降低对作者的看法。这些人往往可能是将来雇佣你、决定是否收你为博士生或与你合作的人。这是一个真实的代价。我也认为保持正直是好的,而且这些是统一的。
An incentive that I think people might often not track is that experienced researchers are pretty good at telling when a paper is just full of hype and has no substance and will generally lower their opinion of the authors. Often these are the people who might be hiring you in future or deciding whether to take you on as a PhD student or potentially collaborating with you. This is a real cost. I also think that having integrity is nice and it's great that these are aligned.
是的。
Yeah.
所以我认为你绝对必须讨论局限性,否则人们会想‘这是谁写的?’
So I think definitely you have to discuss the limitations, otherwise people will be like 'what wrote this?'
是的,但如果你把它们埋在最后,人们可能仍然会那样想,因为他们不会读到结尾。
Yeah, but also if you bury them at the end, people may often still think that because they don't read to the end.
而且你可能会打动很多人,获得很多炒作。我不会撒谎说这没用,但我认为细致的沟通通常是有相当激励的。
And you might also impress a lot of people and get a lot of hype. I'm not going to lie and say this is useless, but I think the nuanced communication is often fairly incentivized.
在过去四年里指导了大约 50 人,你学到了研究导师可以增加价值的方式,尤其是当你有这么多人,无法在任何人身上投入很多时间时?你如何用每周只有几个小时甚至更少的时间来帮助他们很多?
In supervising about 50 or so people over the last four years, what have you learned about the ways a research supervisor can add value, especially with so many people that you can't put many hours into any individual? How can you help them a great deal with only a few hours a week or even less?
我为我的数学学者们采用的系统似乎效果不错,我每六个月左右再招八个人。我把他们分成两人一组,每组做一个项目,我每周和他们进行一次大约一个半小时的检查。我不太擅长让人继续前进,所以我现在有大约 16 个人,这有点问题,但八个人我完全可以在一个晚上搞定,这很好。那么我实际上做了什么,为什么每周一个半小时既足够又比没有重要?
The system I've converged on for my math scholars, which seems to work pretty well, is I take on another eight every sixish months. I put them into pairs, each pair is doing a project, and I have a kind of one and a half hour check-in with them once a week. I'm not very good at getting people to move on, so I currently have like 16 and it's kind of an issue, but eight I can totally do in like one evening a week, it's great. So what do I actually do and why is one and a half hours a week both sufficient and more important than nothing?
这回到了核心观点:研究有不同的技能,这些技能获得的难度不同。另一部分是它们使用的时间成本不同。大部分时间会花在写代码上,而且你在这方面有很好的反馈循环。我不能为他们写代码,我没有时间。实际上,我经常甚至不读我的学者们写的代码,只是希望它没有太多错误。但还有一些技能,比如优先级和研究品味,我以某种方式发展出来了,而且我可以以一种对我来说时间成本很低但对他们来说很难复制的方式使用它们。所以我的理想是,我尝试选择那些能够相当自主地工作,并且在这些技能上自己也做得不错的人。
This gets back to the core idea that research has different skills that have different difficulty to gain. Another part is that they have different time costs to use. The majority of your time will be spent writing code, and you get very good feedback loops on this. I can't write the code for them, I don't have the time. In practice, I often don't even read the code my scholars are writing and I just hope it's not full of bugs. But then there are these skills like prioritization and research taste that I somehow have developed and I can use in a way that's very cheap on my time but would be hard for them to replicate. So my ideal is that I try to select for people who can do fairly well autonomously and do decently on these skills themselves.
你自己没有博士学位。我想你那样可能节省了很多年的学习时间。
You don't have a PhD yourself. I guess you potentially saved many many years of study that way.
我想作为一个局外人,我很难想象现在 AI 行业的最佳选择是离开四年、五年甚至六年去写论文、读一个完整的博士,至少如果你能在行业里找到某种机器学习研究的工作的话。这种直觉对吗——现在不是花那么多时间直接推进职业生涯的时候?
I guess as an outsider, it's hard for me to imagine that the best move in the AI industry right now would be to go away and spend four years or five or six years possibly writing a thesis, doing a full PhD, at least if you could get a job in the industry in some sort of ML research some other way. Is that the right intuition that this is no moment to be taking that much time away from advancing your career directly?
这很复杂。我认为人们常犯的一个错误是认为你必须完成博士学位。这完全是错的。你应该把博士看作一个学习和获得技能的环境,如果有更好的机会出现,通常很容易休学一年。有时候你应该直接退学。如果你有一个选择,比如在一家严肃机构做研究工作,而且你预期它比你的博士更好,那读博的目的就达到了。你已经完成了,提前完成,离开就好。比如,我最近雇了 Josh Engles,他当时在读可解释性方向的博士,已经做了一些非常出色的工作,我已经确信他很优秀,并说服他提前几年退学加入我的团队。我知道我很偏颇,但我认为这是正确的决定。
It's complicated. So I think one common mistake people make is they assume you need to finish a PhD. This is completely false. You should just view PhDs as an environment to learn and gain skills, and if a better opportunity comes along, it's often quite easy to take a year's leave of absence. Sometimes you should just drop out. If you've got an option that's, I don't know, having a job doing research at a serious organization that you expect to go better than your PhD, that's kind of the point of doing a PhD. You're done. You're done early. Leave. Like, I recently hired Josh Engles, who was doing a PhD in interpretability, had done some really fantastic work, and I was already convinced he was excellent and managed to convince him to drop out a few years early to come join my team. I know I'm very biased, but I think this is the right call.
那么除此之外,人们到底该不该读博呢?我认为进入研究领域最有用的方式之一是导师指导。你可以从博士导师或实验室其他人那里获得指导。如果你加入行业研究团队,你也可以从同事和经理那里获得指导。但不同地方在你能学到什么以及经历的价值上差异很大。人们常常没有意识到的是,你将与之共事的人——无论是行业里的经理还是博士导师——作为管理者的技能极其重要,而且这与成为一名优秀研究者的技能截然不同。通常最著名的博士导师是最差的,因为他们非常忙,没有时间给你,或者因为优秀的研究者并不一定能转化为耐心的导师,不关心培养人,不为他们腾出时间。行业角色也是如此。我喜欢认为我创造了一个环境,让团队里的人可以学习,有一定的研究自主权,并成长为研究者,因为至少,如果团队里的人都是优秀的研究者,我的生活会轻松很多。激励是完全一致的。但有些团队自主权少得多,他们可能只希望你做工程师之类的工作,或者你可以做研究但只能在一个非常具体的议程上。我确实认为工程是一项真正的技能。在博士阶段很难学到,在行业里容易学到,但做研究也是如此。
So aside from that, should people do them at all? I think one of the most useful ways to get into research is mentorship. You can get mentorship from PhD supervisors or other people in the lab. You can get it from colleagues and your manager if you join an industry research team. But different places vary a lot in what you'll learn and how valuable an experience it will be. Something people often don't realize is that the skill as a manager of the person you'll be working with, either as your manager in industry or your PhD supervisor, is incredibly important and is a very different skill from being a good researcher. Often the most famous PhD supervisors are the worst because they are really busy and they don't have time for you, or because being a good researcher just doesn't translate into being a patient supervisor, caring about nurturing people, making time for them. The same applies for industry roles. I like to think I make an environment where people on my team can learn, have some research autonomy, and grow as researchers, because if nothing else, my life's a lot easier if people on my team are great researchers. The incentives are very aligned. But there are some teams where you get much less autonomy and they might want you to just be an engineer or something, or you might be able to do research but on a very specific agenda. I do think engineering is a real skill set. It's often hard to learn in a PhD and easy to learn in industry, but so is doing research.
我要特别指出的一点是,我认为我们仍然需要能够设定新研究议程、提出创造性新想法的人,而博士通常是更好的环境,因为你可以做任何你想做的事,没有人能真正阻止你,而且导师通常对此没问题。行业里的经理则不然。我尊敬的一些读过博士并推荐读博的人强调这是关键原因之一:我们确实需要更多能够引领新研究议程的人,而独自被扔进荒野一段时间来摸索这一点可能是学习它的最佳方式之一。你可以这样想:在学术界,你通常有很高的自主权,但没有太多支持。有些人茁壮成长,有些人表现糟糕。众所周知,公共健康负担比性病还严重。我们可以在描述里链接那篇博客文章。
One thing I will particularly call out is I think we still need people who can set new research agendas, come up with creative new ideas, and often PhDs are better environments for this because you can just do whatever you want and no one can really stop you, and often supervisors are fine with this. Managers in industry less so. Some of the people I respect who've done PhDs and recommend it emphasize this as one of the crucial reasons: we really want more people who can lead new research agendas, and being thrown into the wilderness on your own for a while to figure this out can be one of the best ways to learn this. You can think of this as in academia, you often have high autonomy but not much support. Some people thrive, some people do badly. Famously, the public health burden is worse than STDs. We can link to that blog post in the description.
如果你没听过那个说法,去谷歌一下。
Google that if you haven't heard that claim.
我认识很多读博期间非常痛苦的人。我也认识一些热爱读博、根本不想离开或非常庆幸自己完成的人。在行业里,你通常自主权较低,但团队之间差异很大。不过你可以参与更大的项目,资源更充足。你通常会学到更多工程知识,也常常学到与前沿模型研究更相关的技能。但团队之间的差异通常起主导作用。
I know a lot of people who've been very miserable doing PhDs. I also know people who've loved doing PhDs and just did not want to leave or are really glad they finished. In industry, you often have lower autonomy, but it varies a lot between teams. But you can be part of larger projects. You're better resourced. You'll often learn a lot more about engineering, and you'll often learn skills that are more relevant to frontier model research. But the variance between teams typically dominates.
一个具体建议:如果你在考虑加入某个团队或学术实验室,和那个导师的学生或团队里的其他人聊一聊,在一个你认为他们可能会诚实的坦诚环境中,问很多关于那个导师的问题。他们花多少时间?和这个人合作有什么好处?如果可以的话,他们会改变什么?如果这个人名声很好但学生说坏话,那就是个陷阱,你不应该去。我发现初级人员试图进入角色和招聘角色的人之间存在如此大的信息不对称,这让我很沮丧。我尽量做到:如果我想雇某人,但我不确定这是否是他们最好的选择,我就直接告诉他们。但确实存在很多信息不对称。如果你收集到正确的信息,你可以做出更好的决定。有时候你的上级,尤其是某些博士导师,非常糟糕,把自己的利益置于你的利益之上。我听说过有的导师会尽可能阻止学生毕业,以便获得最大的研究劳动力,我认为这很糟糕,我强烈反对这样的导师。
One concrete tip: if you're considering working in some team or academic lab, talk to the students of that supervisor or other people on that team in a candid context where you think they'll probably be honest with you, and ask a lot of questions about what that supervisor is like. How much time do they spend? What benefit do they get from working with this person? What things would they change about them if they could? If the person has a great reputation but the student says bad things, it's a trap and you shouldn't go there. I find it quite frustrating that there's such a big information asymmetry between junior people trying to get into roles and the people hiring for the roles. I try to make a point of if I'm trying to hire someone and I'm not sure this is the best option for them, just telling them this. But yeah, there is a lot of information asymmetry. Expect that if you gather the right kinds of information, you can make better decisions. Sometimes the person in authority over you, especially certain PhD supervisors, is kind of terrible and prioritizes their own interests over yours. I've heard of supervisors who will try to stop their students graduating for as long as possible so they get the maximum research labor, and I just think this is awful and massively jot supervisors like that.
所以我们在这一集里从你这里得到了很多职业和生活建议。这类事情多年来在节目中变得不那么重要了。我想一个原因是,我越来越担心对一个人有帮助的建议可能对另一个人有害,因为人们差异很大,他们可能需要听到不同的东西。生活和工作建议在多大程度上具有普适性,我对此有点怀疑和紧张。你有什么想法吗?我想你经常处于需要给各种人建议的位置,尤其是比你资浅的人,如果你试图帮助他们。你有多担心你建议的东西可能不好?
So we're getting a bunch of career and life advice from you throughout this episode. That sort of thing has become a bit less of a focus for the show over the years. I think one reason is I've become a bit more worried that advice that's helpful for one person can be potentially harmful for someone else, because people are just so different and they might well need to hear different things. How much life and career advice generalizes is something I'm a little skeptical and nervous about. Do you have any thoughts on that? I guess you're in a position where you often have to give people all kinds of different advice, people who are more junior than you if you're trying to help. How much do you worry that the things you suggest could be bad?
也许我们把这一点分成在这种场合给出的建议和在一对一环境中给出的建议。
Maybe let's divide this into advice given in a context like this and advice given in a one-on-one setting.
所以,在这样一个我只是对很多人讲话的场合,我同意你刚才说的一切。人们应该对我说的每句话都持保留态度。我试图控制这一点以及我选择说什么,但我可能搞砸了,而且我很容易不自觉地假设别人和我想法一样。不过,有几个原因让我觉得自己能说些有用的话。我非常清楚,我成功的一个重要原因是运气好,我在对的时间出现在对的地方。所以我并不试图专门就这类事情给出建议,除了元层面的东西,比如好的机会有哪些特征。有些方面比如如何做研究,我确实有相当多的经验。我确信我仍然会偏向某一类人,而可解释性是一个特定的领域。所以人们应该对我说的持保留态度。我也非常同意不同的人需要不同的建议。有一篇很棒的博客文章讲的是“相反建议的法则”。我尝试以这样的形式给出建议:有些人走极端 A,有些人走极端 B,正确做法在中间。我给出我认为正确的样子,试图稍微缓解这个问题。但人们确实应该意识到,他们通常会有一种偏见,认为自己需要更多做已经在过度做的事情。
So, in a context like this where I'm just speaking to however many people, I agree with everything you just said. People should take everything I say with a mountain of salt. I'm trying to control for this and what I choose to say, but I probably mess up and I definitely find it easy to slip into modeling people as though they think like I do. And yeah, but a few reasons that I think I can say somewhat useful things. I'm definitely very conscious that a big part of the reason I've been successful is I got lucky. I was in the right place at the right time. So I am not trying to give advice about that kind of thing specifically, other than meta-level things like here are the traits of good opportunities to look for. There are things like how to do research where I actually have a decent sample size at this point. I'm sure that I still kind of filter for a certain kind of person and interpretability is a certain kind of field. So people should take what I say with a grain of salt. I also very strongly agree that different people need different advice. There's this fantastic blog post on the law of equal but opposite advice. I try to give advice in the form: some people do extreme A, some people do extreme B. The correct thing to do is somewhere in the middle. Here's what I think the correct thing looks like to try to mitigate this a bit. But yeah, people should definitely be aware and also be aware that they will typically have a bias towards thinking they need to do more of the thing they are already doing too much of.
对,对。但我不确定,我觉得有时候你能看出来。还有一些事情比如如何写冷邮件,我对给出建议感觉不错,因为这更多是关于我如何看待事物,尽管有些部分比如把所有的亮点放在前面,可能效果一般。因人而异。
Right. Right. But I don't know. I think you can sometimes tell. And then there are things like how to write cold emails where I feel pretty good about giving the advice because it's more about how I perceive things, though definitely some parts of that like put all the ways you're impressive up front are like eh. Your mileage may vary.
好了,我们得结束了。我们的时间到了。所以,我想你在这一集里分享了很多不同方面的建议。有什么你觉得比较通用的建议,你特别希望人们从这一集里记住的?
All right. We've got to wrap up. We've reached the end of our booking here. So, I guess yeah, you've had a lot of different advice to share across many different fronts in this episode. What's something you would, I guess earlier we were talking about the importance of distilling things down and having the one sentence that people remember. What's a piece of advice that you think is pretty general that you would really like people to recall from this episode?
是的。我认为我学到的最重要的一课是:你可以直接去做事。这听起来可能有点老套,但我认为这里面有很多不显而易见的东西。首先,做事本身是一项技能。我是个完美主义者,经常不想做事,或者觉得‘啊,这有风险,可能会出错’。我打破这一点的方法是挑战自己一个月每天写一篇博客文章。就这样我遇到了我过去四年的伴侣。哦,太好了。这也帮助我产出了很多公开成果,推动了机械可解释性领域的发展。另一部分是最大化你的‘运气表面积’。你想要尽可能多地创造好机会降临的可能性。你要认识人,要成为一个有时说‘是’的人,这样别人会把事情带给你。你要参与很多事情。你也要愿意做一些有点奇怪或不寻常的事情。比如,我 YouTube 频道上最受欢迎的一个视频有大约 3 万次观看,是我花了 3 小时通读一篇著名的机械可解释性论文《Transformer 电路的数学框架》,并给出评论。我没有剪辑,直接上传到 YouTube。人们很喜欢。最后举个例子,我最终意外地管理了 DeepMind 团队。我加入 DeepMind 时以为自己会是个独立研究员。但没想到,我加入几个月后,负责人决定离职,我就在那个月接替了他的位置。我当时不知道能否做好,但结果还不错。对我来说,这既体现了拥有‘运气表面积’的重要性——处于能产生这类机会的环境中,也说明你应该对事情说‘是’,即使它们看起来有点吓人,你不确定能否成功,只要风险很低。最坏的情况,我领导团队不力,然后辞职,他们另选他人。现在看来一切顺利。
Yeah. I think one of the most important lessons I've learned is that you can just do things. And this may sound kind of trite, but I think there are a bunch of non-obvious things I've had to learn here. I think the first one is that doing things is a skill. I'm a perfectionist. I often don't want to do things or I'm like, 'Ah, it seems risky. This could go wrong.' And the way I broke this myself was I challenged myself to write a blog post a day for a month. And that's how I met my partner of the past four years. Oh, wonderful. And also helped me produce a bunch of the public output that's helped build the field of mechanistic interpretability. Another part of this is what I think of as kind of maximizing your luck surface area. You want to just have as many opportunities as possible for good opportunities to come your way. You want to know people. You want to be someone who sometimes says yes so people bring things to you. You want to just get involved in a bunch of things. You also want to be willing to do things that are kind of a bit weird or unusual. I don't know. One of the most popular videos on my YouTube channel with like 30,000 views was I read through one of the famous mechanistic interpretability papers, 'A Mathematical Framework for Transformer Circuits' for 3 hours and just gave takes. I did no editing, put it on YouTube. People were into this. And yeah, maybe as a final example, I kind of ended up running the DeepMind team by accident. I joined DeepMind expecting to be an individual researcher. Then unexpectedly, the lead decided to step down a few months after I joined and in the month since I kind of ended up stepping into their place. I did not know if I was going to be good at this. I think it's gone reasonably well. And to me, this is both an example of the importance of having luck surface area, being in a situation where opportunities like that can arise, but also you should just say yes to things even if they seem kind of scary, you're not confident they'll go well, so long as the downside is pretty low. And like worst case, I just didn't do a good job leading the team. I stepped down. We had to pick someone else. It seems to have gone pretty well.
今天的嘉宾是 Neel Nanda。非常感谢你来到 80,000 Hours 播客,Neel。
My guest today has been Neel Nanda. Thanks so much for coming on the 80,000 Hours podcast, Neel.
非常感谢你的邀请。
Thanks a lot for having me.