How to Have a Career in AI Safety Research
打开互动全文版(中英对照 + 朗读 + 问答)→DeepMind 研究科学家 Yan Leike 讨论了他如何使机器学习稳健且有益的工作,以及如何为人工智能安全领域的职业生涯做准备。
Yan Leike, a research scientist at DeepMind, discusses his work on making machine learning robust and beneficial, and how to prepare for a career in AI safety.
听众朋友们好,这里是 80,000 Hours 播客,一档关于世界上最紧迫的问题以及如何利用你的职业生涯来解决这些问题的节目。我是 Rob Wittman,80,000 Hours 的研究总监。今天的节目是关于如何成为一名机器学习研究员,专注于确保 AI 系统完全按照我们的意图行事。这期节目部分是对第 3 集与 OpenAI 机器学习研究员 Dario Amodei 博士的访谈的回应。如果你还没听过那期节目,我建议你先听一下,因为它解释了更广泛的问题,会让这次访谈更容易理解。如果你对机器学习或人工智能没有太大兴趣,可以跳过这期,因为我们的目标是深入探讨一些更具体的问题。快速通知:Effective Altruism Global 是致力于尽可能做好事的人的主要会议。下一次 EAG 活动将于 6 月第二个周末在旧金山举行,80,000 Hours 团队的大部分成员都会参加。活动的目标是提高人们的知识、技能和社交网络,使他们能够产生更大的社会影响。如果你喜欢这个节目,你很可能会喜欢 EA Global。我看到很多人参加后受益匪浅,无论是结识的人还是发现的机会,所以如果你试图通过职业生涯做很多好事,你应该考虑参加。组织者正在挑选能够从活动中获得最大收益的人,所以如果你已经相当熟悉有效利他主义的关键思想,并希望掌握更复杂的问题或从社区其他人那里获得帮助,他们更有可能接受你的申请。如果你对有效利他主义还比较陌生,通常最好先参加一个社区主办的 EA Gx 活动。今年他们将在澳大利亚、欧洲和美国东海岸举办。你可以在 eaglobal.org 申请,如果在 3 月 18 日(本周日)之前申请,你可以通过早鸟票省下一笔钱。话不多说,有请 Jan Leike 博士。
Hi listeners, this is the 80,000 Hours podcast, the show about the world's most pressing problems and how you can use your career to solve them. I'm Rob Wittman, director of research at 80,000 Hours. Today's episode is about how to have a career as a machine learning researcher focused on ensuring AI systems do exactly what we intend them to do. It builds on and is in part a response to episode 3 with Dr. Dario Amodei, a machine learning researcher at OpenAI. If you haven't listened to that episode yet, I'd recommend doing that first, as it explains the broader issue and will make this interview make a whole lot more sense. If you don't have that much interest in machine learning or artificial intelligence in general, you should feel free to skip this episode, as our goal was to dive into a bunch of more specific questions. One quick announcement: Effective Altruism Global is the main conference of people involved in the discipline of doing as much good as possible. The next EAG event is in San Francisco on the second weekend of June, and most of the 80,000 Hours team will be there. The goal of the event is to increase people's knowledge, skills, and network to enable them to have a greater social impact. If you like this show, you're very likely to enjoy EA Global as well. I've seen many people go and benefit a great deal from the people they met and the opportunities they found out about, so you should definitely think about going if you're trying to do a lot of good with your career. The organizers are looking to choose people who can get the most out of the event, so they're more likely to accept your application if you're already fairly familiar with the key ideas of effective altruism and are looking to master more complex questions or get help from other people in the community. If you are fairly new to effective altruism, it's usually best to first join a community-hosted EA Gx event. This year they're going to be in Australia, Europe, and the US East Coast. You can apply at eaglobal.org, and if you do so before the 18th of March, which is this Sunday, you can save money by getting early bird tickets. Without further ado, I bring you Dr. Jan Leike.
今天我和 Jan Leike 对话。Jan 是伦敦 DeepMind 的研究科学家,也是牛津大学人类未来研究所的研究员。他的研究旨在使机器学习稳健且有益,因此他研究的问题包括如何设计或学习一个好的目标函数,以及如何使机器学习更加稳健。感谢你来到播客,Jan。
Today I'm speaking with Jan Leike. Jan is a research scientist at DeepMind in London and a research associate at the Future of Humanity Institute at the University of Oxford. His research aims to make machine learning robust and beneficial, so he works on questions like how can we design or learn a good objective function, and how can we make machine learning more robust. Thanks for coming on the podcast, Jan.
嘿 Robin,很高兴来到这里。
Hey Robin, great to be here.
我们稍后会谈到人们如何为从事类似你的工作做准备,但首先,你最近开始在 DeepMind 工作。你在帮助他们做什么?
We'll get to how people can prepare themselves to do work that's similar to yours, but first, you recently started working at DeepMind. What are you helping them with?
我是 DeepMind 技术 AI 安全团队的一员。所以我基本上在研究关于让 AGI 安全的技术问题。
So I'm part of the technical AI safety team at DeepMind. So I'm basically looking at technical questions regarding making AGI safe.
80,000 Hours 的核心是让人们致力于世界上最重要、被忽视且可解决的问题。那么你为什么认为你正在做的事情是人类面临的最紧迫的问题之一?
80,000 Hours is all about getting people working on the most important, neglected, and solvable problems in the world. So why do you think what you're working on is one of the most pressing problems that humanity faces?
AI 有潜力成为一种强大的技术,我们可以用它来对世界产生许多积极影响。特别是 AI 和机器学习最近正在经历快速改进的时期,我预计它们将继续如此。如果是这样,我们可以用它来在许多问题上取得进展,比如全球贫困、动物痛苦等全球性问题。但任何新技术都有风险,我们应该事先了解这些风险,以便我们能够驾驭这个领域并明智地使用技术。这就是 AI 安全的意义所在,我正在研究这些问题的技术方面。
AI has the potential to be a powerful technology that we can use to make a lot of positive impact in the world. AI and machine learning in particular have recently been undergoing a period of rapid improvement, and I expect that they will continue to do so. If this is the case, then we can use it to make progress on lots of problems, like global problems such as global poverty, animal suffering, and others. But with any new technology, there are risks that we should understand beforehand so that we can navigate the space and use the technology wisely. This is what AI safety is about, and I'm working on the technical side of these problems.
如果我们不提前准备,未来可能会面临哪些问题?
What are some of the problems that we might face in future if we don't prepare for them ahead of time?
人们喜欢描述的经典场景是,你正在构建一个非常强大的人工智能,你给它一个可能容易指定但实际上不是你关心的目标函数,然后这个 AI 最终非常努力地优化,你得到了你指定的东西,但不是你想要的。这是我研究的一个重点:我们如何将好的目标函数放入机器中?
The classical kind of scenario that people like to describe is that you're building a very powerful artificial intelligence and you give it some objective function that is maybe easy to specify but not actually what you care about, and then this AI just ends up optimizing really hard and you get what you specified but not what you wanted. This is one of the things my research focuses on: how can we get a good objective function into your machine?
人工智能和 AI 安全的风险最近在媒体上被广泛讨论。人们对这个领域有哪些常见的误解?
Risks from artificial intelligence and AI safety have been spoken about a lot in the media recently. What are some common misconceptions that people have about the field?
你指的是公开辩论之类的事情?
So you're referring to the public debate and things like that?
是的,例如。
Yeah, for example.
我认为公开辩论中存在很多无益的两极分化。一方面,有人说我们应该真正关注近期问题,比如自动驾驶汽车和失业;另一方面,有些人非常关心遥远的未来,并且对此非常危言耸听。我认为两极分化这个领域是无益的。我们真正需要的是经过深思熟虑、信息充分的辩论,特别是当我们考虑构建 AGI 和非常强大的东西时。作为社会,我们需要做出很多决定:我们如何处理这个问题?它有什么影响?为了真正有效地做到这一点,我认为我们还必须让公众更好地了解 AI 的真实情况。
I think there's a lot of unhelpful polarization going on in the public debate. On the one hand, there are people saying we should really focus on near-term issues like self-driving cars and unemployment, and on the other side, there are people who really care about the far future and tend to be very alarmist about it. I think polarizing the space is kind of unhelpful. What we really need is well-reflected, informed debates about it, especially when we think about building AGI and something really powerful. There are lots of decisions that we have to make as a society: how do we deal with that? What kind of implications does that have? In order to really do that productively, I think we also have to give the public a better understanding of what's really going on in AI.
在 DeepMind 内部,你们怎么称呼这个问题?我的意思是,我称之为 AI 安全,但专家们是否使用其他术语?
What do you call the problem within DeepMind? I mean, I refer to the issue as AI safety, but is there another term that experts use?
有很多相关的术语,比如对齐问题或对齐,还有 AI 战略和政策之类。我认为 AI 安全是我们不幸沿用下来的一个术语。我不认为这是最好的术语,因为它隐含了一种含义,即人工智能研究在其他方面是不安全的,但事实并非如此,不过它现在已经是一个既定术语了。
There are a bunch of terms that are related, like the alignment problem or alignment, things like AI strategy and policy. I think AI safety is a term that we're unfortunately stuck with. I don't think it's the greatest term because it has this implicit connotation that somehow artificial intelligence research is unsafe otherwise, which is not true, but it's kind of an established term at this point.
你和你的同事在 DeepMind 研究的具体问题有哪些细节?
What are some of the details of the specific issues that you and your colleagues are researching at DeepMind?
我们最近与 OpenAI 合作发表了一篇论文,题为《基于人类偏好的深度强化学习》,其中我们训练了一个神经网络来学习一个奖励函数,供智能体最大化。使用我们的方法,你基本上可以学习并教会智能体任何你想到的目标函数。在我们的案例中,我们教一个小机器人做后空翻,如果你必须手动指定什么函数能让它看起来好看,这真的很难。在我们的案例中,你只需要看一堆智能体行为的视频片段,然后根据它们看起来有多好来排序。
We just recently released a paper together with OpenAI called 'Deep Reinforcement Learning from Human Preferences', where essentially we train a neural network to learn a reward function for an agent to maximize. Using our approach, you can basically learn and teach the agent any objective function that you have in mind. In our case, we taught a small robot to do a backflip, which is really hard to do if you had to hand-specify what function would make it look good. In our case, all you need to do is look at a bunch of video clips of the agent's behavior and rank them based on how good they look.
它看起来很像一个后空翻,而且这对人类来说比自己做后空翻要容易得多。那么,这个过程有望如何提供帮助呢?
Much it looks like a backflip, and this is a humanist kind of like a lot easier to do than say doing a backflip yourself. So what's the hope about how this process will help?
所以短期内,这只是用来解决以前难以解决的新问题。我不是在说像后空翻这样的事情,而是那些实际有用的事情。但从长远来看,这个项目的愿景是,当我们真正构建 AGI 时,目标函数会是什么?这里的想法是,这是朝着学习人类价值观迈出的一小步,或者你希望购买的家用机器人做什么,以一种你不需要成为强化学习专家,但可以通过其他对人类来说非常容易的形式提供反馈的方式。所以你只需上传和下载。我会说这更像是“我想要这个”或“我不想要那个”,就像在验光师那里测试你的眼睛一样,对吧?
So in the short term, this is just useful to solve new problems that were kind of difficult to solve before. I'm not talking about things like the backflip, but things that are actually useful. But in the longer term, the kind of vision for this project is that we're thinking about when we actually build AGI, what would be the objective function? And the idea here is that this is kind of a small step into the direction of learning what humans value, or what you would want a household robot that you buy to do, in a way that you don't need to be an expert in reinforcement learning, but you can just give feedback in other forms that are really easy for humans. So you just upload it and download it. I'll say this is more like what I want or less like what I want than something else. It's like being at the optometrist when they're testing your eyes, right?
我的意思是,我认为最终我们希望以人类想要的方式获取反馈。现在我们做的是,有两个视频片段,你大致说左边更好、右边更好,还是它们差不多?但我认为如果有更好的反馈方式就好了。也许你看一个视频,然后说,哦这部分看起来很好,这部分看起来不太好,但总体上你对视频中发生的事情没有强烈的看法。
I mean, I think ultimately we want to take feedback in the way that humans want to give it. And right now what we do is we have two video clips and you kind of say, is the left better, right better, or are they kind of the same? But I think it would be great if we have better ways of giving feedback. Maybe you watch a video and you say, oh this part looks really good, this part doesn't look very good, but overall you don't have strong views about what happens in the video.
几个月前我采访了 OpenAI 的 Dario Amodei,他提到了这个后空翻面条以及你如何训练它。所以这是 DeepMind 和 OpenAI 之间的合作,对吗?
I interviewed Dario Amodei at OpenAI a few months ago, and he mentioned this backflipping noodle and how you'd go about training it. So this is a collaboration between DeepMind and OpenAI, right?
太好了。是的,这是你们合作的众多项目之一吗?我们刚刚开始与他们合作技术 AI 安全,目前我们正在进行更多的后续合作工作。我希望未来能找到更多可以合作的项目。总的来说,我希望看到 OpenAI 和 DeepMind 更紧密地合作。
That's great. Yeah, is this one of many projects that you guys work on together? So we just started the collaboration on technical AI safety with them, and we're currently doing more follow-up work that we collaborate on. And my hope is that we can find more projects in the future that we can collaborate on. Overall, I would like to see OpenAI and DeepMind working more closely together.
那么回到这个后空翻面条的强化学习过程,有哪些方式即使这个训练系统也可能失败,机器学习算法可能做出我们意想不到的事情?
So coming back to this reinforcement learning process with the backflipping noodle, what are some ways that even that training system might still fail and the machine learning algorithm might do things that we didn't intend?
是的,我认为这是一个非常好的问题,这也是我非常感兴趣的事情,因为我想知道所有可能失败的方式,并提前考虑它们。在那个项目中,我们注意到的一件事是,如果我们不提供在线反馈,即系统学习时的交互式反馈,你可能会遇到退化解决方案。基本上,奖励预测器,即影响奖励函数的组件,你停止给它反馈,然后随着智能体继续学习,状态分布发生变化。然后奖励预测器必须在分布偏移下表现良好,所以它看到以前没有见过的新输入,然后开始错误预测奖励。我们实际看到的是,在某些情况下,智能体只是学习到非常奇怪、不是你本意的东西,但根据奖励预测器它们看起来不错,因为奖励预测器不知道它在新的分布中在做什么。
Yeah, I think that's a very good question, and this is the kind of thing I'm very interested in because I want to know all the ways in which things can fail and think about them ahead of time. In the case of that project, one thing that we noticed was that if we don't give feedback online, meaning interactively while the system is learning, you can run into degenerate solutions. Basically, the reward predictor, the component that influences the reward function, you stop giving it feedback, and then as the agent continues to learn, the state distribution changes. Then the reward predictor has to do well under distributional shift, so it's seeing new inputs it hasn't exactly seen before, and then it starts mispredicting the reward. What we actually saw was that there were some cases where the agent just learns really weird things that you didn't intend, but they look good according to the reward predictor because the reward predictor doesn't know what it's doing in the new distribution.
如果我能试着用连我自己都能理解的语言来解释:基本上,你让人们向机器学习算法提供反馈,然后该算法会尝试预测人类在其他情况下会说什么。对吗?
If I can try to put that into a language that even I would understand: basically you have people offering feedback to a machine learning algorithm that is then going to try to predict what humans would say in other cases. Is that right?
是的。发生的事情是,如果人类停止向该过程提供信息,然后它试图评分的面条动作超出了它熟悉的分布和它有人类评分经验的分布,那么它就会开始给出无意义的答案,因为它不再有任何基础,没有相关的人类答案可以借鉴。所以答案开始崩溃,你基本上会得到后空翻风格的随机变化。
Yep. And what happens is if humans stop providing information to that process, and then the kinds of noodle movements that it's trying to score move outside the distribution of what it's familiar with and what it has experience with humans rating, then it will just start giving nonsense answers because it no longer has any basis, it doesn't have relevant human answers to draw on. So you just start getting the answers to start breaking down, and you'll get random changes basically in the backflipping style.
是的,所以这是神经网络实际上非常不擅长的事情。它们没有良好的置信区间;它们不擅长指定自己的不确定性。所以当你把它们扔到一个不是它们训练过的新问题时,它们不会退缩并说“哦,我不知道该怎么做”。它们仍然给出非常自信的答案。在我们的案例中,你仍然对发生的任何事情分配奖励,你甚至没有意识到你不应该这样做。
Yeah, so this is something that neural networks are actually really bad at. They don't have good confidence intervals; they're not good at specifying their own uncertainty. So when you throw them into a new problem that isn't what they trained on, it's not like they back off and say, 'Oh, I don't know what to do here.' They just still give very confident answers. And in our case, you're still assigning rewards to whatever happens, and you don't even realize that you shouldn't be doing this.
我们在 Dario 的播客上描述了这个后空翻面条的训练过程,但也许只是让我们回忆一下。那么这里的新颖见解是什么?我只是喜欢“后空翻面条”这个词。
We described the training process for this backflipping noodle on the Dario podcast, but maybe just refresh our memories. So what's the novel insight here? I just love the term 'backflipping noodle'.
好的,是的。所以这是一个三部分的过程。一部分是人类观看视频片段并对它们进行排名。第二部分是我们所说的奖励预测器,它基本上学习人类如何对不同的行为进行排名,并将其转化为分配给该特定行为的奖励。第三部分只是一个常规的强化学习算法,它试图最大化奖励,因此基本上试图最大化奖励预测器认为人类想要的东西。
Okay, yeah. So it's a three-part process. One part is a human looking at video clips and ranking them. The second part is what we call the reward predictor, which basically learns how the human would rank different behavior and turns that into the reward that is assigned to that particular behavior. And the third part is just a regular reinforcement learning algorithm that tries to maximize reward, and thus basically tries to maximize what the reward predictor thinks the human wants.
所以担忧是,如果我们试图在现实生活场景中使用这种方法,而人类没有被要求足够频繁地给出他们对机器人所做不同事情的评分,或者如果情况发生变化以至于人类会给出不同的评分,那么机器人会非常自信地做出完全不是我们想要的事情,因为它根本不理解自己所处的新情况。
So the concern would be that if we tried using this approach in a real-life situation and humans weren't called on to give actual scores of how they rate different things that the robot was doing often enough, or if the situation changed such that humans would give different scores, then the robot would very confidently do things that were not at all what we would like, because it just doesn't understand that the new situation it has found itself in.
是的,这正是我想要更好理解的问题。比如,我们到底需要多少反馈?我们能否随着时间的推移减少反馈,这基本上就是我们在那项工作中所做的?但这也取决于环境如何变化,以及你如何知道何时应该要求更多反馈,或者你应该对行为的哪些部分要求反馈?
Yeah, and this is exactly the sort of questions that I would love to understand better. Like, how much feedback do we exactly need? Can we give less feedback over time, which is kind of what we did in that work? But it also depends on how the environment changes, and how can you know when you should ask for more feedback, or what kind of parts of your behavior should you ask for feedback for?
有没有办法让它意识到它现在正在评分的行为与以前见过的行为完全不同?所以我的意思是,我想你可以插入某种异常检测机制,对吧?异常检测是机器学习已经思考了相当长时间的一类问题。我们还没有尝试过。我认为这可能是一个值得尝试的好方法。我不太清楚异常检测目前的效果如何。
Isn't there a way of just getting it to realize that it's now scoring behaviors that are quite different from what it's seen before? So I mean, I guess you could insert some kind of anomaly detection mechanism into this, right? And anomaly detection is a problem class that machine learning has thought about for quite a while. We haven't tried doing that. I think that might be a good thing to try. I'm not quite clear on how well anomaly detection works right now.
嗯,它确实能扩展。那么你认为这是一个重大进步,还是只是我们在找到如何让机器学习在重要应用中真正安全之前必须尝试的众多事情之一?
Well, it scales. So do you see this as a big step forward, or is it just one of many things we're going to have to try before we figure out how to really make machine learning safe for important applications?
我认为这是朝着一个以前很少被探索的方向迈出的一小步。所以我认为一旦我们做了更多的后续工作,其他人开始在此基础上构建,它会非常有用。总的来说,还有很多其他问题我们也想思考和考虑,比如我们如何安全地探索,如何在最大化奖励的同时隐含地识别副作用,以鼓励你的智能体不造成不必要的副作用,这又意味着什么?更广泛地说,还有很多其他我认为非常重要的问题,但目前很少有人真正深入思考它们。其中一些问题正变得越来越流行,机器学习安全就是其中之一。
I see this as a small step into a direction that wasn't explored very much before. So I think it'll be very useful once we do more follow-up work and other people start building on this. Overall, there are so many other problems we also want to think about and consider, like how can we explore safely, how can we maximize reward while implicitly recognizing side effects to encourage your agent not to cause unnecessary side effects, and what does that even mean? More broadly, there are lots of other questions that I think are really important, but right now very few people are thinking very hard about them. Some of them are becoming more popular, so machine learning security is one of them.
有一些非常直观的例子:你拿一张图片,对它进行非常微小的扰动,以至于你几乎看不到或者甚至用肉眼根本看不到,但一个原本自信的分类器现在却彻底改变了它给出的分类。这不仅适用于图像分类,更一般地说,如果你有一个深度神经网络,你只需要对输入进行微小的扰动,就能以你无法预料的方式改变输出。你需要有一个底层神经网络过程的副本才能找出如何创建这些略有不同却被分类为完全不同类别的图像吗?
There are very visceral examples where you take an image and perturb it very minimally, so you can hardly see it or can't even see it with your own eyes, but a regular confident classifier that was trying to classify the image now just vastly changes the classification it gives. That doesn't only apply to image classification, but more generally, if you have a deep neural network, you only need to perturb the input minimally, but you change the output in ways you wouldn't anticipate. Do you have to have a copy of the underlying neural net process to figure out how to create these images that are slightly different but get classified as a completely different kind?
这是其中一件引人注目的事情:你甚至不需要那样做。相反,你可以做的是在相同的数据集上从头训练一个神经网络,可能使用稍微不同的架构,然后攻击你刚刚训练好的自己的神经网络。你从中得到的输入扰动也会转移到其他模型上。这就是所谓的黑盒攻击,这些黑盒攻击实际上出奇地有效。即使你看向深度强化学习,人们用深度学习训练强化学习智能体,你甚至可以用一个全新的算法或完全不同的算法来训练一个智能体。比如说你训练了 DQN 来执行某个任务,而我试图攻击你的 DQN。我刚刚训练了另一个算法 FEC,并攻击它,然后输入仍然会转移到你的问题设置中。我认为这是我们真正应该想办法解决的问题。这有点令人惊讶;对我来说非常惊讶。
So this is one of the striking things about this: you don't even need that. Instead, what you could do is train a neural network from scratch on the same dataset, possibly using a slightly different architecture, and then attack your own neural network that you just trained. The kind of input perturbation you get out of that then also transfers to other models. This is what is called black-box attacks, and these black-box attacks actually work surprisingly well. Even if you look at deep reinforcement learning, where people trained reinforcement learning agents with deep learning, you can even train an agent with an entirely new algorithm or with entirely different algorithms. So say you trained DQN to perform some task, and I'm trying to attack your DQN. I just trained another algorithm, FEC, and attack that, and the input still transfers to your problem setting. I think that's something we should really figure out how to fix. It's a bit surprising; it was very surprising to me.
我最近看到一篇论文声称,很难对自动驾驶汽车使用这种扭曲图像攻击,因为它们通常只从一个特定的角度或特定的缩放级别起作用,而且由于汽车移动得很快,它们会从许多不同位置得到物体外观的平均值。你知道这方面的情况吗?
I saw a paper recently claiming that it would be very difficult to use these kinds of distorted image attacks against self-driving cars because they usually only work from one particular angle or one particular level of zoom, and because cars are moving quite quickly, they get much more of an average of what something looks like from a lot of different positions. Do you know anything about that?
我认为最近有一些研究表明,你不需要对停车标志做太多改动就能骗过分类器。我的意思是,考虑到我之前谈到的所有其他事情,在这一点上并不那么令人惊讶。所以目前机器学习安全的底线是:它非常容易受到攻击,而且有很多不同的攻击方式,我们还没有真正好的防御策略。这有点让人想起互联网的早期,那时每个人都开放了所有端口之类的,攻击软件很容易,而我们只是慢慢变得更好。
I think there was some recent research that showed that you don't have to do much to the stop sign to actually fool a classifier. I mean, given all the other things I talked about that came previously, at this point it's not so surprising. So the bottom line story of machine learning security at the moment is that it's really easy to attack, and there are lots of different ways you could attack, and we don't really have very good defense strategies yet. It's kind of reminiscent of the early days of the internet, where everyone had all their ports open or something, and it was easy to attack software, and we just slowly got better at that.
你可以采取哪些方法让深度强化学习更加鲁棒?
What are some approaches you could take to make deep reinforcement learning more robust?
我认为在这个领域有很多有趣的问题值得探索,因为强化学习实际上就是关于通用智能体与环境交互,你可以想到很多不同的鲁棒性问题出现在这个背景下。例如,对于探索:我如何以安全的方式探索我最初不太了解的环境,这样我就不会做出任何不可逆的决定,或者在这个过程中不会发生坏事?或者例如我之前提到的副作用:我们如何确保在最大化奖励的同时,也尽量不必要地干扰环境?还有机器学习安全问题,比如有人攻击深度强化学习智能体并试图让它做某些事情,你如何防御?另一件事是深度强化学习算法以不稳定著称。例如,你在十个不同的随机种子上训练你的深度强化学习算法,最终得到的性能方差可能相当大。我认为让它更稳定会很好,这样你就能有更可靠的性能,更频繁地收敛到同样好的性能水平。所以不是试图为你所有的随机种子获得最高的性能,而是想要一个算法,其最终结果变化不大,并且可靠地执行。这些都是重要且有趣的问题。我们可以做前沿研究,而且在某种程度上,我认为研究人工智能安全是令人兴奋的,因为它还不是一个非常成熟的领域。有很多不同的低垂果实,正好成熟可以采摘,如果你现在在这个领域,你可能是采摘它们的人之一。一切都等着被拿下。
I think there are lots of interesting questions to explore in this space because reinforcement learning is really just about general-purpose agents interacting with the environment, and there you can think of lots of different robustness questions that come up in this context. For example, for exploration: how can I explore my environment that I initially don't really know much about in a safe way, so I don't make any irreversible decisions or nothing bad happens while I do that? Or for example, side effects that I mentioned earlier: how can we make sure that while you're maximizing reward, you also try not to disturb your environment unnecessarily? There are also machine learning security questions like somebody attacks a deep reinforcement learning agent and tries to get it to do certain things, and how can you defend against that? Another thing is that deep reinforcement learning algorithms are known for being notoriously unstable. For example, you train your deep RL algorithm on ten different random seeds, and the variance in the performance you get out of that at the end can be quite large. I think it would be good to just make it more stable so that you can have more reliable performance, more frequent convergence on the same good level of performance. So instead of just trying to get the highest level performance for all the random seeds you take, you want an algorithm where the final outcome doesn't change that much and just reliably performs. These are all important and interesting questions. We can do cutting-edge research, and in a way, I think working on AI safety is exciting because it's not a very well-established field yet. There are lots of different low-hanging fruit that are just really ripe to pick, and if you're in this field now, you could be one of the people who picks them. It's all up for grabs.
DeepMind 的其他人在研究哪些有趣的问题,你能谈谈吗?
Are there any other interesting problems that other people at DeepMind are working on that you're able to talk about?
DeepMind 有很多有趣的项目在进行,在 DeepMind 工作的好处之一就是你能够看到它们逐步展开。去年有很多关于 AlphaGo 的事情,其中大部分在我去那里工作之前就发生了。但现在很多人对星际争霸非常兴奋,我们最近为此发布了一个研究环境,这样其他人也可以在上面工作。是的,我认为还有很多其他非常令人兴奋的事情在进行。
There are lots of interesting projects going on at DeepMind, and one of the perks of working at DeepMind is that you kind of get to see them as they unfold. So last year there was a lot of stuff going on with AlphaGo, most of that happened before I was even working there. But right now a lot of people are really excited about StarCraft, and we recently released a research environment for that so that other people can also work on it. And yeah, I think there are lots of other really exciting things going on.
你的日常工作是什么样的?
What is your work like on a day-to-day basis?
很多不同的事情。我花很多时间阅读 arXiv 上的论文。
Lots of different things. I spend a lot of time just reading arXiv papers.
你开会、和其他研究人员讨论研究。有时候你只是坐下来思考问题,找出接下来该做什么,做计划,和研究工程师讨论他们在做什么,实现方面的事情等等。我想你的同事都超级聪明,对吧?
You sit in meetings, talk to other researchers about the research. Sometimes you just sit down and think about problems, figure out what would be the next good things to work on, plan, talk to research engineers about what they're working on, implementation side, things like that. I guess your colleagues are insanely smart, right?
是的,这真的很令人兴奋。DeepMind 拥有一大批各个领域的世界级专家,你可以向他们请教问题。
Yeah, this is really exciting. DeepMind has a whole bunch of the world experts on various topics, and you get to ask them questions.
这是这份工作的福利,还是说有些世界上最聪明的人和你竞争,会有点威胁?
Is that a perk of the job, or is it a bit threatening having some of the world's greatest minds competing with you?
两者兼有。我的意思是,他们不完全是在和我竞争;我们是一起工作的。但和有真才实学的人进行良好的对话,并且知道你在最前沿的研究项目的最佳机构,这很令人愉快。
It's kind of both. I mean, they're not exactly competing with me; we're working together. But it's enjoyable to have really smart people to have good conversations with, and know that you're at the best organization at the forefront of the research project.
那么 DeepMind 是一个有趣的工作场所,还是说非常安静、非常学术?
So is DeepMind a fun place to work, or is it just extremely quiet and very studious?
不,我们有一个非常大的开放式共享办公空间,自然会有很多对话。在我的办公桌旁,有两位著名教授就坐在我对面,你会随机开始和他们互动。人们在茶水间和咖啡馆相遇。每周五都有派对,有饮料、披萨等等。
No, we have this really big open office shared space, and there are lots of conversations that can naturally arise. At my desk, I have two famous professors sitting right across from me, and you just randomly start interacting with them. People meet at micro kitchens and cafes. And every Friday there's a party with drinks and pizza and everything.
有很多人和你一起研究这些安全话题,还是只有少数几个人?
Do you have a lot of other people working with you on these safety topics, or is it just a handful?
目前我们还没有那么多人。我认为对于想要建立职业生涯的人来说,专注于这类问题是一个很好的机会。我们与 OpenAI 和未来人类研究所合作。有很多合作在进行,但我真的希望有更多人研究这些问题。
Right now we don't have that many people yet. I think there's a great opportunity for someone who wants to build their career to focus on these kinds of problems. We collaborate with OpenAI and with the Future of Humanity Institute. There's lots of collaboration going on, but I really wish there would be more people working on these problems.
你的职业生涯中走过了怎样的道路,最终来到了 DeepMind?听起来你是德国人,对吧?
What path in your career did you take to end up working at DeepMind? It sounds like you're German, right?
是的,没错。我在德国弗莱堡读了本科,专业是数学和计算机科学。然后我完成了计算机科学硕士,之后去了澳大利亚攻读机器学习博士。
Yes, that's correct. I did my undergraduate degree in Germany in Freiburg, in math and computer science. Then I finished a master's in computer science, and then I went to Australia to do a PhD in machine learning.
你和 Marcus Hutter 合作过吗?我猜他是你的导师。
Did you work with Marcus Hutter? I imagine he was your supervisor.
是的,没错。
Yeah, that's right.
你在堪培拉过得愉快吗?
Did you have a good time in Canberra?
堪培拉很适合让你专注于生产力。我会说这是一种非常礼貌的说法。
Canberra's pretty good to make you focused on your productivity. I would say that's a very polite way of putting it.
博士毕业后你做了什么?然后去了未来人类研究所吗?
What did you do after your PhD? Did you then go to the Future of Humanity Institute?
是的,没错。我在那里做了大约六个月的博士后,然后加入了 DeepMind。
Yes, that's right. I was a postdoc there for about six months before I joined DeepMind.
你认为职业生涯中哪些决定是正确的?
Which decisions in your career do you think you made the right call on?
我认为在那个时候攻读机器学习博士是一个非常正确的决定。我非常喜欢和我的导师 Marcus Hutter 一起工作,他教了我很多东西。我当时去 FHI 的原因是我认为自己想更专注于理论研究,因为那是我博士的重点。但现在我改变了想法,我认为在这个领域实证研究更有价值,而 DeepMind 是我做这件事的最佳场所。
I think getting into a machine learning PhD at the time that I did was a very good decision. I really liked working with my supervisor, Marcus Hutter, and he taught me a lot. The reason for me to go to FHI when I did was that I thought I wanted to focus more on theoretical research, because that was the focus of my PhD. But now I've changed my mind, and I think that empirical research is more valuable in this space, and DeepMind is the best place for me to do that.
是什么让你改变了想法?
What changed your mind about that?
很多事情。目前在这个领域,实证工作似乎有些未被充分探索,有很多事情可以做,而且用这种方法似乎更容易。我认为人们应该做理论,用理论方法解决安全问题,但对我来说这似乎更难。
A number of things. It seems right now the empirical work is kind of underexplored in this space, and there are lots of things to do, and it seems easier to approach it that way. I think people should do theory and figure out safety problems using theoretical methods, but it just seems harder to me.
你有没有考虑过做政策工作,比如你在 FHI 的前同事 Miles Brundage?还是说技术研究显然更适合你?
Did you ever consider doing policy work, like your previous colleague at FHI, Miles Brundage? Or was technical research clearly a better fit for you?
我认为政策研究中有很多非常重要且有趣的问题需要解决,比如关于自主武器、自主黑客等等。人们确实应该思考这些问题。但就我而言,我在技术方面有比较优势,因为我的背景非常技术性,我对政策和国际关系真的不太了解。
I think there are lots of really important and interesting questions to be solved in policy research, things about autonomous weapons, autonomous hacking, and all of this. People should really be thinking about them. But in my case, I have a comparative advantage in working on the technical side because my background is very technical, and I don't really know anything about policy and international relations.
在 DeepMind 找工作的申请流程是怎样的?
What was the application process for getting a job at DeepMind like?
我的情况有点不典型。我和 DeepMind 的三位创始人都进行了面试,因为我申请的时候,技术安全团队非常小,他们正在组建中。通常面试过程包括 DeepMind 测验,他们会问你很多关于计算机科学、数学、统计学和机器学习的不同问题。
In my case, it's somewhat atypical. I had interviews with all three of the DeepMind founders because at the time I applied, the technical safety team was very small, so they were just building it up. Usually part of the interview process is the DeepMind quiz, where they ask you lots of different questions about computer science, math, statistics, and machine learning.
几个月前我在播客中采访了 Dario Amodei。他是你认识的人,而且你在 OpenAI 与他合作。他说的有什么你不同意或者有不同看法的吗?
A couple of months ago I spoke with Dario Amodei on the podcast. He's someone you know and you're collaborating with at OpenAI. Was there anything he said that you disagree with or have a different perspective on?
哦,是的,我几乎同意所有事情。有一件事我会以不同的方式回答。当你问如何判断自己是否适合做研究时,Dario 的回答是,你应该拿一些最近的论文,实现模型,快速完成,看看你做得有多快,是否觉得有趣,以及是否能复现结果。我会更强调研究的其他部分。判断你是否适合做研究的良好指标是,研究对你来说感觉有趣且容易,它自然而然地发生。你在空闲时间会思考研究问题,那些你不知道答案、也许没人知道答案的问题。你倾向于沉迷于谜题和问题,而且你能够清晰简洁地向还不理解的人解释你的想法,比如新颖的想法。所以这结合了产生想法、阅读文献、执行项目和向他人展示。在实现模型方面,在 DeepMind 我们有研究工程师,他们的工作是与研究人员合作进行实现。他们非常擅长实现模型和调整模型,这解放了研究人员的时间,让他们思考更高层次和概念性的问题。如果你从事理论研究,那么擅长实现显然不那么重要。
Oh yeah, I agreed with almost all of the things. There was one thing where I would have answered the question differently. When you asked how to figure out whether you're a good fit for research, Dario answered that you should just take some recent papers, implement the models, do that very quickly, and see how quickly you can do it, whether it's fun to you, and whether you can replicate the results. I would have answered that question in a way that puts more emphasis on other parts of research. Good indicators for whether you're a good fit for research are that research feels fun and easy to you, and it just kind of happens naturally. You end up thinking about research questions during your downtime, the kind of questions you don't know the answer to, and maybe nobody knows the answer to. You tend to get obsessed with puzzles and problems, and also that you can explain your thoughts, like novel thoughts, to people who don't understand them yet, clearly and concisely. So it's a combination of having ideas, reading literature, executing a project, and presenting it to other people. In terms of implementing models, at DeepMind we have research engineers whose job is to work with researchers on the implementation side. They tend to be very good at implementing models and tweaking them, which frees researchers' time to think about more high-level and conceptual questions. If you're working on theoretical research, being really good at implementation is obviously less important.
让我们转向个人职业选择的问题,以及听众如何自己为解决这个问题做出贡献。我们都认为,成为一名 AI 安全研究员是人们可能做的最有影响力的事情之一,但这也可能非常非常困难。那么,允许某人做出有用贡献的数学能力下限是什么?
Let's move on to the issue of personal career choice and how listeners can potentially make a contribution to solving this problem themselves. So we both think that being an AI safety researcher is one of the highest impact things that people can potentially do, but it's also potentially very, very difficult work. What's the lower bound of math ability that would allow someone to make a useful contribution?
所以我认为这取决于你采用哪种方法,对吧?如果你在做理论工作,那么你真的需要非常非常擅长数学。如果你在做更多的实证工作,那么你应该理解基础数学领域,比如线性代数和统计学等等。比如什么是中心极限定理,什么是特征向量——解决这类入门级的东西。但在实证方面,你通常不需要非常擅长证明定理之类的东西。
So I think that depends on what kind of approach you're taking, right? If you're doing theory work, then you should really be really, really good at math. If you're doing more empirical work, then you should understand the fundamental math fields like linear algebra and statistics, and so on. Things like what is the central limit theorem, what is an eigenvector—solving entry-level stuff like that. But on the empirical side, you usually don't have to be really good at proving theorems and stuff like that.
拥有哪种思维方式最重要?
What kind of thinking is most important to have?
除了参与其中,真正有用的思维方式是批判性思维。比如说你在读一篇研究论文,问问自己:这篇论文好在哪里?哪些地方可以做得更好?你如何扩展它?然后你可能就能指出某个研究输出的弱点。这在你自己写研究时也很有用,对吧?因为你应该知道自己研究的弱点在哪里,以及如何改进和扩展它。所以如果你有更多时间做这件事,就知道该关注什么。在某种程度上,做研究有点像训练生成对抗网络:你需要一个好的判别器来判断你做的事情好不好,然后你可以训练你的生成器来生成那些东西。另一种思维方式当然是分析性思维。你需要对数字有良好的直觉,对数学和算法有直觉——比如什么东西计算成本高——而且你必须能够编程。作为研究员,你不一定最终要写很多代码,但你必须能够做到并理解如何做。我认为总的来说,非常重要的一点是,你应该能够自如地在一个你不甚了解的领域中导航,因为研究必然处于人类知识的前沿,涉及我们不理解的东西。所以你必须能够应对未知。这与你的本科学位所培养的技能形成对比,本科阶段你基本上是在学习我们理解得很好的东西,更多的是需要记住并快速理解它们,而不是处理未知。
Beyond engaging in the kind of thinking that is really useful is critical thinking. So say you're reading a research paper, ask yourself: what is good about this paper? What should be done better? How could you extend it? And then you might be able to point to the weaknesses of a particular research output. This is also useful when you're writing your own research, right? Because you should know where your own research has weaknesses and how you could do better, and how to extend it. So if you have more time to work on it, what to focus on. In a way, doing research is kind of like training a GAN: you have to have a good discriminator to figure out whether what you're doing is good, and then you can train your generator to generate that stuff. Another type of thinking is, of course, analytical thinking. You need to have good intuitions about numbers, you have to have intuitions about math and algorithms—like what kind of things take how expensive to compute—and you have to be able to code. As a researcher, you don't necessarily end up coding so much, but you have to be able to do it and understand how to do it. I think overall, something that's really important is that you should be comfortable with navigating a space that you don't really understand very well, because research is necessarily on the frontier of human knowledge and things that we don't understand. So you have to be comfortable with dealing with the unknown. This is in contrast to the skill set that your undergraduate degree selects for, where you're basically learning about things that we understand well, and it's more about being able to remember them and understand them quickly, rather than dealing with the unknown.
如果有人已经熟悉机器学习,已经接受过一些训练,人们可以尝试哪些事情来看它是否适合他们?
If someone's already familiar with machine learning, if they already have some training in it, what kinds of things can people try to see if it's a good fit for them?
所以如果你已经知道如何在机器学习领域做研究,那么你就拥有了在安全领域做研究所需的所有技能。而且,这真的不是两件不同的事情。不存在 AI 和 AI 安全之分;它们都是 AI 问题,或者都是机器学习问题。所以你会用同样的工具来处理它们,用同样的方式思考它们。这真的是同样的问题。
So if you already know how to do research in machine learning, then you should have all the skills that you need to do research in safety. And really, these are not two different things. There's not like AI and AI safety; they're both AI questions, or they're both machine learning questions. So you'll approach them with the same tools, you'll think about them in the same way. It's really the same problems.
如果人们对机器学习了解不多,他们如何判断这是否是一条适合他们继续走下去的合理道路?
And if people don't know much about machine learning, how can they figure out if this is a sensible path for them to get further down?
所以这涉及多个方面,对吧?有这个问题:你对机器学习有多兴奋,你对从事这方面工作有多兴奋?你是想更多地从事研究方面,还是更多地从事实现方面,比如研究工程?你如何判断自己是否具备这些技能?网上有很多资源,你可以通过教程并实现各种深度学习模型。如果你想从事实现方面,这是一件事。我认为如果你想从事研究方面,通常最好读个博士或获得某种同等经验,以便真正提升研究技能。你怎么知道自己是否适合?你应该在大学导师的指导下做一个研究项目,或者实习,看看进展如何,做这件事有多开心。
So there's various aspects to this, right? There's the question of how excited you are about machine learning, how excited you are about working on that. Do you want to work more on the research side or more on the implementation side, like research engineering? And how can you figure out whether you have each of these skills? There are a lot of resources online where you can go through tutorials and implement various deep learning models. That's one thing if you want to work on the implementation side. I think if you want to work on the research side, it's usually good to get a PhD or some kind of equivalent experience in order to really scale up on the research skills. How do you know if you're a good fit for that? You should just work on a research project with a supervisor at university or an internship, and see how well that goes, how much fun you have doing that.
一会儿再谈博士,但对于未来考虑从事 AI 安全的人来说,理想的本科专业是什么?
Talk about the PhD in a minute, but what would be the ideal undergraduate degree for someone who's thinking about working on AI safety in the future?
我认为完美的本科专业是计算机科学和数学。如果你有另一个定量学科的本科学位,比如物理,那也可以。但通常你应该了解所有不同的材料:微积分、编程、算法、机器学习、深度学习、强化学习、统计学等等。一般来说,优先选择较难的课程而不是较容易的课程是好的。如果你必须在数学课程或应用课程之间做选择,我会推荐数学课程,即使它看起来与你实际想做的事情不太相关。我认为通常早点开始做研究是个好主意,并尝试在完成硕士学位之前至少发表一篇论文,因为这会让你在申请博士时处于更有利的位置。这算是你能够做研究的证明,人们可以据此评估你的水平。找一个不一定是最有名的导师也很好,因为如果他们真的很有名,通常没有太多时间给你。但找一个非常擅长指导的人,这样你可以从他们那里学到很多,得到很多反馈。到那时,你也更容易发现自己是否理解。
I think the perfect undergraduate degree is computer science and mathematics. If you have an undergraduate degree in another quantitative subject like physics, that's also fine. But usually you should know all the different material: calculus, coding, algorithms, machine learning, deep learning, reinforcement learning, statistics, and so on. Generally, it's good to prioritize harder courses over easier ones. If you have to choose between a math course or an applied course, I would recommend the math course even if it doesn't seem that related to what you actually want to do. I think it's generally a good idea to start doing research early and try to publish at least one paper before you finish your master's degree, because that would put you in a much better position when you apply for PhDs. That's kind of a proof that you're able to do research, and people can evaluate how good you'll be based on that. It's also good to find a supervisor who is not necessarily the most famous supervisor, because if they're really famous, they usually don't end up having much time for you. But somebody who's just really good at supervising, so you can learn a lot from them and get a lot of feedback. And at that point, you also have an easier time finding out whether you get it.
说到这个,DeepMind 并不是那种任何人都可以走进去见到所有人的开放办公室。如果有人完成了本科学位,并且对这个领域感兴趣,有没有什么办法可以让他们见到在这个领域工作的人,可能讨论这个话题?
Speaking of which, DeepMind doesn't exactly have open offices that anyone can just walk into and meet everyone. If someone has finished an undergraduate degree and they have an interest in this whole area, is there any way that they can go about meeting people who are working in the field, potentially to discuss the topic?
是的,所以我建议去参加顶级的机器学习会议,比如 ICML、NeurIPS、ICLR 等等。理想情况下,你可以看看发表的论文,在线查看它们,浏览并找到你感兴趣的论文。然后仔细阅读它们是有意义的。
Yeah, so I would recommend going to one of the top machine learning conferences, like ICML, NeurIPS, ICLR, and so on. Ideally, you can look at the papers that are coming out, you can look at them online, and you can go through them and find the papers that you find interesting. And then it makes sense to read them in detail.
所以听起来你非常鼓励人们攻读机器学习博士学位。有没有其他值得提及的替代方案?
So it sounds like you're pretty enthusiastic about people doing machine learning PhDs. Are there any alternatives doing a machine learning PhD that are worth mentioning?
是的,通常招聘研究人员的要求是博士学位,但也有例外。还有其他途径能给你类似的经历,比如 Google Brain 的驻留项目就是一个例子。也有人读博中途退出去工业界实验室工作。对于某些职位,比如我们说的工程岗,博士学位并非严格必需,你只需要编程能力强、理解机器学习并能跟进研究。我之所以建议有条件的人攻读机器学习博士,是因为目前这个领域人才最紧缺。在 DeepMind 的技术安全团队,我们很想招更多人。我们需要机器学习博士或同等经验的人,让他们从事安全研究。问题在于,具备所需背景并且对 AI 安全充满热情的人不够多,所以我们无法招到更多人。
Yeah, so usually the requirement for someone to hire us as a researcher, and there's some exceptions for that. And there are other routes that give you similar kind of experience. The Google Brain residency is one example of that. There's also people who start their PhD and then they kind of stop midway through and go work in one of the industry labs. For some positions like we said, engineering a PhD is not strictly required, and you basically have to be really good at coding and understand machine learning and follow the research. And the reason why I would recommend people to get a machine learning PhD if they're in a position to do so is that this is kind of where we're currently the most talent constrained. So at DeepMind, for the technical area safety team, we love to hire more people. We have like a machine learning PhD or equivalent experience, and just get them to work on safety. And the problem is that there's not enough people who have that required background and also excited about working on AI safety, so that we could hire more.
你认为人们在多大程度上应该先走传统的机器学习路径,和像你这样的 DeepMind 的人一起工作,而不是自己或在 DeepMind 之外的志同道合群体中尝试做 AI 安全研究?
To what extent do you think people should go down traditional machine learning paths and work with people like you at DeepMind first, versus trying to do AI safety research on their own or in groups of like-minded people outside of an institution like DeepMind?
是的,我认为这很重要,因为如果人们过早地尝试成为独立的安全研究者,往往不会很成功。我认为通过更成熟的路径来获得必要的研究技能是个好主意。所以从某种意义上说,机器学习博士并不是我们真正关心的,而是一系列相关技能的集合,这些技能对于成为研究者非常有用。
Yeah, so I think this is quite important because I think if people try early on to be independent researchers in AI safety, they tend not to be very successful. And I think going through one of these more established paths to getting the required research skills is a really good idea. So in a way, a machine learning PhD is not actually the thing we care about, but it's more a package of a number of related skills that are really useful if you want to be a researcher.
人们在决定何时从最主流的机器学习研究问题转向你特别感兴趣的安全话题时,会面临什么权衡吗?
People face any trade-offs in deciding when to switch from the most mainstream machine learning research questions to working on the kind of safety topics that you're particularly interested in?
是的,这可能会发生。当你对自己掌握的技能有信心时,就应该发生转变。比如,如果你在读机器学习博士,可以在博士后期;或者博士毕业后去其他地方。我通常看到有些人过于强调尽快进入 AI 安全研究,而忽视了技能积累。有时人们非常担心从事能力研究,认为那不是安全研究。但我认为这些担忧被过分强调了,人们不应该太担心,而应该专注于自己的职业资本。
Yes, that could happen. That should happen basically when you feel confident that you have acquired the skills. So that could be for example in the later years of your PhD if you're doing a machine learning PhD, or it could be after you finished your PhD and you go somewhere else after that. I think usually I see people who overemphasize going into AI safety research as quickly as they can over skill building. Or sometimes people are really concerned about working on capability research where that is meant to be something that is not safety research. But I think these concerns are overemphasized and people shouldn't worry too much about them, and rather really just focus on their career capital.
在准备好申请 DeepMind 之前,人们可以在哪里工作?有没有可能有人读了博士但还没准备好去 DeepMind 工作,他们可以采取哪些中间步骤来弥补差距?
Where might someone work before they're ready to apply to DeepMind? Is it a possibility that someone could do a PhD but not yet be ready to work at DeepMind, and are there any intermediate steps that they can take to bridge the gap?
是的,我认为有很多有趣的事情可以做。有机器学习实习。我认为一个做实习的好地方是蒙特利尔的 Mila。还有各种机器学习初创公司可以去工作。你也可以在工业界做博士后。我认为选择很多。
Yeah, I think there's lots of interesting things to do. There's machine learning internships. So I think one good place to do an internship is at Mila in Montreal. There's various machine learning startups where you might work. You might do a postdoc in industry. I think there's a long list of options.
如果一个人学了机器学习但找不到 AI 安全研究的工作,或者决定不再做这个,他们有什么选择来过上好的生活或者为世界做很多好事?
If someone goes up in machine learning but then they can't find a job in AI safety research, or they decide that's no longer what they want to do, what are the kind of options do they have for either having a good life or doing a lot of good for the world?
是的,我认为现在如果你有机器学习博士学位,你会非常抢手,很多人想给你钱让你和他们一起做很酷的机器学习项目。所以有很多不同的工业实验室可以加入。特别是如果你想在更应用的场景中用机器学习做好事,有很多项目可以加入,很多公司。比如 DeepMind 就有用机器学习做健康领域的事情。但总的来说,我认为现在如果你有这个学位,找到一份高薪舒适的工作不会有问题。
Yeah, I think right now if you have a machine learning PhD, you're really hot and there's lots of people who just want to throw money at you to be around them and make cool machine learning things happen. So there's lots of different industry labs that you could join. In particular, if you want to use machine learning in a more applied setting to do good in the world, there's lots of projects that you could join, lots of companies. There's stuff going on at DeepMind for using machine learning for health, for example. But yeah, I think right now if you have that degree, you wouldn't have problems to find a really well-paying comfortable job.
稍微跳出框框想一想,有没有其他领域的选择,比如从政,或者其他方式让有机器学习经验的人能够帮助处理 AI 安全问题?
Thinking a bit outside the box for a minute, are there any options in other areas like going into politics or any other ways that people with machine learning experience might be able to help to deal with AI safety issues?
是的,在政策和政府领域有很多关于 AI 的问题需要处理。让这些机构中真正理解机器学习技术细节的人参与进来,将有助于进行非常知情的讨论。所以我认为有技术背景的人也应该认真考虑这一点,因为那里有很大的影响力潜力。我知道你不久前在播客中采访过 my language,他写了一本关于如何做政策的优秀指南。我们会在节目笔记中放一个链接,这样人们可以了解 AI 安全的政治、政策和战略方面。
Yeah, so there's a lot of questions around AI that come up in the policy and government space that I think we really need to deal with. And having people in these institutions who really understand the technical details of what's going on in machine learning is going to be really helpful to have really informed discussions about this. So I think this is something that people with a technical background should also really seriously consider, because I think there's a lot of great potential for impact there. And I know you had my language on the podcast just a while ago, and he's written this excellent guide on how to do a policy. We'll stack up a link to that in the notes on the episode so people can find out about the politics and policy and then strategy side of AI safety.
人们可能采取的另一种方法是努力赚钱再捐赠。如果他们有能力赚很多钱,他们可以这样做,然后捐给他们认为有助于让 AI 更安全的人、组织或项目。你认为这是一个明智的方法吗?
Another approach that people might take is trying to earn to give. So if they have an ability to make quite a lot of money, they might go do that and then try to donate it to people or organizations or projects that they think will help to make artificial intelligence more safe. Do you think that that's a sensible approach for people to take?
目前,我认为技术安全领域的人才缺口远大于资金缺口,所以我不会太担心把自己的钱投入这个领域。我们真正需要的是有人来研究这些问题。
So at the moment, I would say that the talent gap in technical safety is just so much larger than the funding gap that I wouldn't really worry about putting my own money into this area. It's really like we need people to work on these problems.
除了做机器学习研究,还有其他技术方法来解决 AI 安全吗?
Are there any other technical approaches to AI safety other than doing machine learning research?
是的,还有其他方法,最著名的是机器智能研究所提出的高度可靠智能体设计议程。对于这类工作,你的背景可能是逻辑学博士之类的。我的印象是,机器学习博士的人才缺口更为严重,而且我预计几年后从事技术 AI 安全研究的人员中,很多都会是机器学习背景的。
Yeah, there's some other approaches, most notably the agenda for highly reliable agent design that came out of the Machine Intelligence Research Institute. And for that kind of work, your background would probably be a PhD in logic or something like that. Yeah, my impression is that the talent gap for machine learning PhDs is much more severe, and I expect that in terms of the number of technical AI safety researchers that will be working on this in a few years' time, I think a lot of them will be machine learning.
你选择的职业道路最大的缺点是有很多工作要做,我想责任重大,因为你正在处理一个非常重要的问题。
The biggest downside of the career path that you've taken is you have a lot of work to do, I guess a lot of responsibility because it's such an important issue that you're working on.
是的,现在有很多事情需要做,但做的人不多,所以我必须阻止自己试图做所有事情,然后惨败,因为实在太多了。我想我们得好好排个优先级。
Yeah, right now there are all these things that need to be done and there are not many people to do them, so I have to prevent myself from trying to do all of them and then failing miserably because it's really too much. We have a lot of prioritization to do, I guess.
有些人担心 AI 安全,也有兴趣做贡献,但他们不愿参与,因为他们觉得‘我不是世界上最聪明的数学家之一,我真的能有所作为吗?’你觉得这种想法是错的吗?
Some people are worried about AI safety and interested in contributing, but they are reluctant to get involved because they think, 'I'm not one of the smartest mathematicians in the world, so can I really make a difference?' Do you think that's misguided?
是的,我认为解决这些问题真的不需要是数学天才。更重要的是要有研究思维,能批判性思考,提出新想法,并坚持解决难题,即使不清楚何时能找到解决方案。你需要适应不确定性和无解的问题。
Yeah, I think you really don't need to be a math genius to work on these problems. I think it's much more important that you have a research mindset, can think critically, come up with new ideas, and stick with difficult problems even when it's not clear when you'll find a solution. You need to be comfortable working with uncertainty and unanswerable questions.
假设有人正在考虑攻读机器学习博士学位。你对他们准备或读博期间有什么一般性建议?
Suppose someone is thinking of doing a PhD in machine learning. What general advice do you have for them in preparing for that or what they should do while doing their PhD?
在决定是否读博以及读什么时,有很多因素需要考虑。阅读关于如何读博的一般性建议很有帮助,这些建议在网上和书籍中都有。目前,攻读机器学习博士学位非常困难,因为许多资深人士已经离开学术界去了工业界,所以学术界培训学生的教授很少。此外,由于机器学习现在非常热门,很多人试图进入这个领域,使得博士申请竞争极其激烈,尤其是在顶尖学校。另一方面,有很多可用的工具,比如在线课程(例如 Coursera 上的深度学习课程)、深度强化学习课程、教程和开源框架(如 TensorFlow)。大多数机器学习论文都在 arXiv 上,可以免费阅读。然而,这些东西并不能教授研究技能。为此,读博非常有用。但我通常不愿意建议别人读博,除非他们非常确定想在这个领域工作,因为这是一个很大的承诺。
While figuring out whether to do a PhD and what to do, there are many factors to consider. It's helpful to read general advice on how to do a PhD, which is available online and in books. Currently, pursuing a PhD in machine learning is very difficult because many senior people have left for industry, so there are few professors in academia to train students. Also, because ML is very exciting right now, many people are trying to enter the field, making PhD applications extremely competitive, especially at top places. On the other hand, there are many tools available, like online courses (e.g., deep learning on Coursera), deep reinforcement learning courses, tutorials, and open-source frameworks like TensorFlow. Most ML publications are on arXiv, so you can read them for free. However, these things don't teach research skills. For that, doing a PhD is very useful. But I'm often reluctant to tell people to do a PhD unless they're really sure they want to work in the area, because it's a big commitment.
机器学习博士通常需要多长时间?
How long does a machine learning PhD take?
在美国,通常需要五到六年。在欧洲,时间较短,大约三到四年,但申请时需要硕士学位。所以这是一个更大的承诺。如果中途决定不继续,可以拿硕士学位离开吗?在美国,我想大约两年后可以。也有一些人开始读博,然后中途被顶级工业实验室聘用,所以你可能不会缺少选择。如果你在机器学习博士期间表现良好,我认为你不会缺少机会。
In the US, they can take five or six years. In Europe, they tend to be shorter, around three or four years, but they require a master's degree when you apply. So it is a bigger commitment. If you decide you don't want to continue halfway through, can you leave with a master's? In the US, I think after two years or so. There are also cases where people start a PhD and then get hired by a top industry lab halfway through, so you're probably not short of options. If you do well in a machine learning PhD, I don't think you will have a lack of options.
如果有人本科不是数学或计算机科学,他们如何转向攻读机器学习博士?
What if someone did their undergraduate degree in something other than math or computer science? How can they pivot to get into a machine learning PhD?
这取决于你的背景和职业阶段。如果你有物理学博士学位,也许你应该学习机器学习,尝试实习并发表论文。如果你处于本科或硕士阶段,也许攻读机器学习硕士是合适的。这真的取决于你做过什么以及你已经有多少研究经验。
That depends on your background and where you are in your career. If you have a PhD in physics, maybe you should read up on machine learning, try to do an internship, and get a paper published. If you're at a bachelor's or master's level, maybe a master's in machine learning is the right thing. It really depends on what you've done and how much research experience you already have.
听起来 DeepMind 需要所有能得到的帮助,而且你很想招聘有资格做这类研究的人。你有什么最后想说的来激励人们加入你们吗?
It sounds like you could use all the help you can get at DeepMind, and you'd be really interested in hiring people qualified to do this kind of research. Is there any last thing you'd like to say to inspire people to come and join you?
是的,当然。我认为有很多绝佳的机会可以在世界上做好事,特别是 AI 安全是一个有前景的高影响力领域,因为做的人很少。在 DeepMind,我们真的很努力地招聘更多人来从事这项工作,而且一直很困难。所以如果你擅长机器学习,我非常希望你能来和我们一起工作。
Yeah, of course. I think there are all these amazing opportunities to do good in the world, and in particular, AI safety is a promising area for high-impact work because so few people are doing it. At DeepMind, we're really trying hard to hire more people to do that, and it's been quite difficult. So I would really love it if you're good at machine learning, you should come work with us.
今天的嘉宾是 Jan Leike。感谢你参加 AI 播客。
My guest today has been Jan Leike. Thanks for coming on the AI Podcast.
非常感谢你的邀请。
Thanks so much for having me.
如果你喜欢这期节目,记得考虑在 ea-global.org 申请参加旧金山的有效利他主义全球大会。感谢收听,下周见。
If you enjoyed that episode, remember to consider applying for Effective Altruism Global San Francisco at ea-global.org. Thanks for joining. Talk to you next week.