Nando de Freitas on AI as a Tool for All Humanity
打开互动全文版(中英对照 + 朗读 + 问答)→Nando de Freitas 探讨 AI 作为通用工具的潜力、他在 DeepMind 的工作,以及通过非洲 Indaba 和拉丁美洲 Khipu 等倡议推动 AI 民主化的努力。
Nando de Freitas discusses AI's potential as a universal tool, his work at DeepMind, and his efforts to democratize AI through initiatives like Indaba in Africa and Khipu in Latin America.
今天的嘉宾是 Nando de Freitas。Nando 是世界领先的人工智能研究者之一。他出生于津巴布韦,在约翰内斯堡的威特沃特斯兰德大学获得学士和硕士学位,剑桥大学博士,伯克利博士后,曾任不列颠哥伦比亚大学温哥华分校教授、牛津大学教授,创立了 Dark Blue Labs(后被 Google/DeepMind 收购),现任 DeepMind 研究总监。Nando 在顶级 AI 会议上多次获得最佳论文奖。他近期的作品包括 Gato(首个多模态、多任务、多具身通用智能体)、AlphaCode(竞赛级代码生成)、学习沟通、学习学习,以及通过观看 YouTube 玩硬探索游戏。Nando 在非洲和南美洲作为深度学习与人工智能的教育者和社区建设者非常活跃。Nando,非常高兴你能来。欢迎来到节目。
Our guest today is Nando de Freitas. Nando is one of the world's leading artificial intelligence researchers. Born in Zimbabwe, bachelor's and master's from the University of the Witwatersrand in Johannesburg, PhD from Cambridge, postdoc at Berkeley, professor at the University of British Columbia, Vancouver, professor at Oxford, founder of Dark Blue Labs, which was acquired by Google/DeepMind, where he's a research director today. Nando has won a number of best paper awards at the top AI conferences. Some of his recent works include Gato, the first multimodal, multitask, multiemodiment generalist agent, competition-level code generation with AlphaCode, learning to communicate, learning to learn, and playing hard exploration games by watching YouTube. Nando is very active as an educator and community builder for deep learning and artificial intelligence in Africa and South America. Nando, so great to have you here with us. Welcome to the show.
能来这里太激动了。我是这个播客的超级粉丝,能有机会来我感到非常荣幸。谢谢。
It's so exciting to be here. I'm a huge, huge fan of this podcast, and so that I feel privileged to have an opportunity to be here. Thank you.
能邀请到你,我深感荣幸。现在,Nando,在深入今天的对话之前,我要感谢我们的播客赞助商 Index Ventures 和 Weights & Biases。Index Ventures 是一家风险投资公司,投资于从种子轮到 IPO 各个阶段的杰出企业家。在旧金山、纽约和伦敦设有办事处,支持包括 AI、SaaS、金融科技、安全、游戏和消费等多个领域的创始人。就我个人而言,Index 是 Covariant 的投资者,我强烈推荐与他们合作。Weights & Biases 是一个 ML Ops 平台,通过实验跟踪、模型和数据集版本管理以及模型管理,帮助你更快地训练更好的模型。OpenAI、Nvidia 以及几乎所有发布大型模型的实验室都在使用它。事实上,我在伯克利和 Covariant 的许多同事都是 Weights & Biases 的重度用户。现在,Nando,AI 已经成为主流。世界似乎已经接受了它,尤其是随着 ChatGPT 等最近的进展,它已成为世界的一部分。过去并非如此。这与我们早期在机器学习社区的日子相比是一个巨大的变化。你认为这对地球上的大多数人意味着什么?
I'm so honored to have you on. Now, Nando, before diving into today's conversation, I'd like to thank our podcast sponsors, Index Ventures and Weights & Biases. Index Ventures is a venture capital firm that invests in exceptional entrepreneurs across all stages, from seed to IPO. With offices in SF, New York City, and London, the firm backs founders across a variety of verticals, including AI, SaaS, fintech, security, gaming, and consumer. On a personal note, Index is an investor in Covariant, and I couldn't recommend any higher working with them. Weights & Biases is an ML Ops platform that helps you train better models faster with experiment tracking, model and dataset versioning, and model management. They're used by OpenAI, Nvidia, and almost every lab releasing a large model. In fact, many, if not all of my Berkeley and colleagues at Covariant are big users of Weights & Biases. Now, Nando, AI has become mainstream. It's something the world seems to have, especially with recent advances in ChatGPT and so forth, accepted as something that's really part of what the world is. Didn't used to be that way. It's a big change from our early days in in the machine learning community. What do you think this will mean for most of the population on this planet?
这很难想象。我的意思是,首先,自从我们二十多年前认识以来,变化太大了,我们曾一起参加 NeurIPS 等等,现在街上每个人都在谈论 AI。所以,它已经成为普遍意识的一部分。那么 AI 会为我们所有人做什么呢?我仍然最喜欢把它看作一种工具。或者一系列工具,让我们能做更多事情,就像显微镜和望远镜扩展了我们的视野,个人电脑扩展了我们的能力。我希望 AI 能扩展我们能做的事情,无论是创造新知识,还是解决我们面临的一些最棘手的问题,比如环境问题、能源问题。最重要的是,我希望它能造福全人类,而不仅仅是某个特定群体,而是真正为人类服务。我认为这是一个自然的方向,最终会非常重要。从长远来看,在未来的几个世纪或几千年里,它将对我们的生存至关重要。所以,我认为它可以成为一个奇妙的工具,只要我们明智、有洞察力且富有同情心地使用它。
It's hard to imagine. It's I mean, first of all, it's it's been such a big change since um I you know, since I met you a happy couple of decades ago, uh and we used to go to NeurIPS together and so on and and indeed now everyone on the street is talking about AI. So, it it has become part of I guess the universal conscious. And um what will AI do for all of us? I you know, I like still my favorite way to think about it is to think of it as a tool. And um and I hope or a family of tools that will allow us to do more, just like microscopes and telescopes extended the range of all personal computers for that matter extended the range of things we can do. Um my hope is that AI will extend the things that we can do, whether it's to create new knowledge, whether it's to solve some of the hardest problems that we face, environmental problems, energy problems. Um and above all, um I would love to see it um being for the benefit of all humans, not just one particular group, but really for humankind. I think it's a natural direction um that will be ultimately very important, I think. And so in the very long range of um many centuries or millennia to the future, will be essential for our survival. So, I think it's um could be a wonderful tool um provided that we we use it wisely, insightfully, and with compassion.
现在,你在一个主要机构——可以说是当今 AI 研究的领先机构——DeepMind,对吧?总部在伦敦。同时,你谈到希望它能惠及整个世界。嗯,你如何确保它朝着这个方向发展?
Now, you're at one of the main people would say the leading institution for AI research today, DeepMind, right? Based in London. At the same time, you talk about something that you hope will benefit the whole world. Um yeah, how how how you're going to make sure it plays out that way?
嗯,我所做的一件事是,我想我们每个人都可以做一点小事,我认为这很重要。嗯,诚然,我不知道如何做到,但我能做的就是做一点,我个人一直参与志愿服务。所以,我喜欢教学。我一直在做这件事,因为这是我唯一能做的事。我帮助创建了 Indaba,为前两届 Indaba 筹集资金等等。我甚至帮助检查编程练习,简化它们,并运行实验室。顺便说一下,Indaba 是我们在非洲举办的一个会议。它是一个聚集学生、初创公司等的活动,旨在促进人工智能和技术在非洲的应用,增加非洲大陆对 AI 的接触。这也启发了世界各地的其他一些努力。我还参与的一个是拉丁美洲的 Khipu。我也参加了第一届 Khipu,几周后,我将参加三年来的第一次会议。抱歉,是三年。我已经很久没有出差参加会议了。我真的很期待,因为与拉丁美洲或非洲的学生在一起非常令人振奋。去那里非常重要。除了内容,除了谈论卷积网络或任何当前流行的 AI 话题之外,当学生看到你时,有时我们会忘记我们在 AI 中的身份,比如,你知道,Peter Bill,你是最伟大的明星之一。我一直很钦佩你的工作,等等。数百万人也钦佩你。所以,如果你来到拉丁美洲,亲自参加 Khipu,而不仅仅是做一个演讲——我希望将来,这对任何 AI 人士都是如此。我认为你在这个节目中采访过的大多数人,只要他们出现在那里,学生们看到他们可以和你交谈,他们就会意识到:“嘿,这个人并没有我想象的那么聪明。我和这个人一样聪明,我可以和他们讨论研究。”这给了人们一种赋权感,让他们意识到自己也能做到。我可以走得很远。我知道这一点,因为我就是非洲那些孩子中的一个,我记得和帝国理工学院的 Longyear 教授交谈过,他是控制领域的人。作为一个学生,见到帝国理工学院的教授,对我来说就像见到达赖喇嘛一样,那是一个不可思议的转变时刻,因为那时我意识到我可以成为对话的一部分。我怎么强调都不为过,但你能给人们的第一件事就是让他们相信自己能做到,并且可以参与其中。
Um so, one of the things that I've done, um you know, I guess each of us can do a little thing and I think it's important. Um you know, admittedly, I would not know how to do it, but all I can do is do a little bit and um and I've personally been engaged in volunteering. So, um I love teaching. So, I've continued doing that cuz that's the one thing I can do. And so, I helped with the creation of the Indaba, you know, raising funding and so on for the first two Indabas. I even helped going over the coding exercises and simplifying them and so on with the labs to run. And have run the labs of the first Indabas. And also um the Indaba, by the way, that's it's uh one of these meetings that um we offer um in Africa. It it's a gathering of people, students, startups, and it's um aimed at uh promoting um artificial intelligence and the use of technology in Africa, increasing access uh for the African continent um to AI. Um and that also inspired the creation of several other efforts um across the world. One that I'm also involved with is Khipu in Latin America. So, I also attended the first Khipu, and in a few weeks, I will be going to my first meeting in in 3 weeks. It's a Sorry, 3 years. It's been a long time since I've traveled to conference. Um and I'm really looking forward to it cuz it's being with the students in Latin America or in Africa is extremely energizing. And it is so important to go there. Beyond the content, beyond talking about conv nets or um uh whatever the the the flavor of the day is in AI. Um when students see you, and sometimes we forget um who we are in AI, like, you know, Peter Bill, you're one of the greatest stars. I've always admired your work, and so on. And millions of people admire you, as well. So, if you come to Latin America, if you come to Khipu in person, and not just give a um I hope hopefully when in the future, and this is true of anyone in AI. I think most of the people you've interviewed in this show, just by being there, and the students seeing that they can have a conversation with you, and they actually realize that, "Hey, this person is um not as clever as I thought. I you know, I'm as just as clever as this person and I I can discuss research with them and that gives people a sense of empowerment um that allows them to realize that um I I can also achieve that. I can go far. And it And I know this because I was one of those kids in Africa and I remember talking to um Professor Longyear, who is this control guy from Imperial College. And for me meeting um as a student a professor from Imperial College was it was like, I don't know, meeting the Dalai Lama or um you know, it it was an incredible transformative moment because that's when I realized I could be part of that dialogue. And I couldn't overemphasize that, but it's the first thing you can give people is just the ability to believe that they can do it and that they can be part of it.
这就是我做的部分工作,我鼓励所有观看这个播客的人考虑参加深度学习或 Indaba,以及东南亚和世界各地的会议。对我们来说,尤其是 AI 领域的高级研究人员,这样做非常重要。这能让那里的社区成长,而且与那里的人和初创公司交流也很重要。你希望培育一个完整的生态系统。同样重要的是,这个生态系统开始成长,并且由那里的人们来领导和指引它。我很喜欢 Indaba 的发展,它催生了非洲各地的许多本地化会议,即 IndabaX 会议。每年参加这些会议时,能与这么多年轻、有抱负的人交流,那是一年中最美好的时光。
And so that's a little bit that I do and I encourage everyone watching this podcast to consider attending Deep Learning or Indaba or some of the meetings in Southeast Asia and throughout the world. It's really important for us, especially senior researchers in AI, to do that. And that just allows for the community to grow there and it's also important to talk to the people there and the startups. There's a whole ecosystem that you want to grow. And it's also important that that ecosystem starts growing and that it's the people there that started leading it and directing it. I've loved how the Indaba grew and it led to many localized meetings, the IndabaX meetings, all throughout Africa. It's the most wonderful time of the year when you go to one of these meetings and you're able to engage with so many young, ambitious people.
你提到这个很有意思,因为我现在想到了几件事。第一,你显然在里面偷偷给我塞了个邀请。我很乐意将来接受这个邀请,尽管今年我自己可能还不会旅行。但我很期待。另一件让我想到的是,当你谈到那种能量和兴奋时,虽然我自己还没去过非洲或南美的活动,但 Black in AI 研讨会把很多研究人员带到了 NeurIPS,对吧?那个研讨会充满了巨大的能量。我想象这和你说的有点相似,那里的能量绝对令人惊叹。
It's interesting you bring it up because a couple of things come to my mind now though. One, you clearly snuck an invitation in there for me. I'd love to take you up on that someday even though this year I probably won't travel yet myself. But I look forward to it. The other thing that comes to my mind is when you talk about the energy, the excitement, while I haven't been to one of the events in Africa or South America yet myself, the Black in AI workshops bring a lot of the researchers to NeurIPS, right? And that workshop has tremendous energy. I have to imagine it's a bit similar to what you're talking about and it's absolutely amazing the energy that's there.
确实如此。这太棒了。你知道,我们谈论这个领域由 GPU 或 Transformer 带来的变革。但我认为人的变革同样惊人。我真的很喜欢我们社区的发展方向。我认为我们还有很多工作要做,但我们在多样性和包容性方面取得的进展,我认为是这个领域发生的最重要的事情之一。我只希望我们都能继续朝着这个方向前进,确保我们的工具、研究等能够充分利用每个人在多样性方面所能提供的。并确保我们正在构建的是一个包容性的技术,一个包容性的社区。我认为这对于我们长期想要实现的目标至关重要。
Very much so. It's wonderful. It's one of... You know, we talk about the transformations that we've had in the field brought in by the GPUs or brought in by the transformers. But I think the people transformation has been equally amazing. I really love where our community is heading. I think there's a lot more work that we all have to do, but the work we're moving in terms of diversity and inclusion, I think it's been one of the most important things that has happened to the field and I just hope we can all continue to move in that direction towards making sure that our tools, our research, and so on capitalizes on what everyone can offer on diversity. And to ensure that it's an inclusive technology that we're building and it's an inclusive community. I think that's essential for what we want to achieve in the long run.
现在,如果我能顺着这个说下去,如果我看看 AI 当前的趋势,其中一个趋势是构建非常大的模型。而构建这些模型非常昂贵。所以,即使我们可能教育人们等等,教育可能还不够。我们实际上需要以某种方式给人们相当大的资源,才能做这类事情。或者也许有其他方式。显然,大模型不是唯一发生的事情,但它是很多兴奋点所在。所以我很好奇你怎么看,这个趋势如何与你刚才所说的一切相互作用。
Now, if I can riff off of that, if I look at the current trends in AI, one of the trends is building very large models. And building those models is very expensive. So, even though we might educate people and so forth, educating might not be enough. We need to actually give people quite large resources in one way or another to be able to do those kinds of things. Or maybe there's other ways. Obviously, the large models are not the only thing happening, but it is where a lot of the excitement is. So, I'm curious how you think about that, how that trend interacts with everything you just said.
是的。我认为你问到了点子上。大模型不是一切。我认为大模型是一个非常令人兴奋的研究方向。我的意思是,我们都知道额外的算力让我们能够构建更大的模型,而更大的模型确实往往具有小模型没有的特性。它让我们能够做一些用小模型无法做的研究。所以,世界上正在进行的大规模实证工作非常重要,并且正在让我们取得巨大飞跃。但与此同时,外部也可以发生很多创新,这些创新可以是基础性的工作。我们一直看到这种情况发生。我的意思是,我们都在谈论注意力机制,好像它是公司的东西,但我记得有一个学生提出了这个想法。我记得在牛津面试那个学生时,他提出了注意力机制的想法。然后它去了蒙特利尔,与蒙特利尔的研究人员进一步开发了这个想法。Kyunghyun Cho 和 Yoshua,最终,你知道,经过多次迭代,这项技术才完全成熟……它催生了我们现在在 GPT-3 中使用的 Transformer。但它始于一个学生提出这个想法。还有更多这类想法。最近我们看到了结构化状态空间模型,我们也看到了很多关于微分方程的工作。我认为人们在基础工作方面可以做很多事情。还有另一种我认为非常重要的工作,那就是我们如何使用这些工具?因为即使我们有了大语言模型,最终我认为我们会学会如何更高效地服务它们,并允许人们以某种方式微调它们。如果人们能够高效地微调它们,并通过不同组织提供的各种云服务获得足够的算力,那么我认为就有可能解决下一个大挑战,那就是我们用这些工具做什么。我们如何让这些工具在我们的社区中发挥作用?因为当你去——我每年经常去阿根廷,我在那里有家人,我也去南非。当你到了那些国家,你会遇到一个不同的现实,与你在伯克利遇到的或我在伦敦遇到的非常不同。
Yeah. I think what you are saying to the question. The large models is not everything. I think the large models is a very exciting direction of research. I mean, we all know that extra compute allows us to build bigger models, and bigger models do tend to have properties that smaller models don't have. It allows us to do a certain kind of research that we can't do with small models. So, the large-scale empirical work that's going on in the world is very important and is advancing us in big leaps. But at the same time, there's a lot of innovation that can happen outside and that innovation can be in fundamental work. And we see this happening all the time. I mean, we all talk about attention as attention is this corporate thing, but I do remember there was a student that came up with this thing. I remember interviewing the student in Oxford that that is the attention idea. And then it went on to Montreal and then developed the idea further with researchers in Montreal. Kyunghyun Cho and Yoshua and eventually, you know, it took many more iterations until the technique was fully... it gave birth to the transformers that we use now in GPT-3. But it starts with a student coming up with this idea. And there's a lot more of these types of ideas. Recently we see the structured state space models and we also see a lot of work with differential equations. There's a lot of things in terms of fundamental work that I think people can do. There is another type of work that I think is very important, which is how do we use the tools? Because even if we do have the large language models and eventually I think we are going to learn how to serve them more efficiently and allow people to fine-tune them one way or the other. And if people can fine-tune them efficiently and have access to enough compute through the various clouds that different organizations offer, then I think it does become possible to then address the next big challenge, which is what we do with these tools. How do we make these tools useful in our communities? Because when you go, I often go every year to Argentina, I have family there, and I go to South Africa as well. And when you're transported to those countries, you encounter a different reality, which is very different than what you encounter in Berkeley or what I encounter here in London.
所以 Nando,我喜欢你所说的,世界不同地区有不同的需求、挑战,因此也有不同的机会来构建 AI 应用,甚至是 AI 公司。你能举一些例子吗?
So Nando, I like how you're saying that different parts of the world have different needs, challenges, and hence different opportunities to build AI applications, AI companies for that matter. Can you give some examples?
我可以。这可能是为什么去旅行、做志愿者和参加世界各地的会议很有用的另一个原因。举个例子,当我参加德班的一个节日时,我了解到一个手机应用,它使用大多数人实际上都能用到的廉价手机。其中一个文本应用被用来提供更好的医疗保健。一个我觉得特别深刻的例子是,妈妈们经常尝试给医生发短信问问题,比如,为什么我的孩子不能吃东西?我给我的孩子喝水。我的孩子呕吐,所以无法保持水分。已经一天了。我该怎么办?有时只需要一条简单的信息和一些非常简单的建议,比如烧开水。这就能帮助那个人,因为世界上很多人死于脱水。
I can. And this is probably another reason why it's useful to travel and volunteer and attend meetings all over the world. When I went to the festival in Durban, to give you an example, I learned about this phone app using cheap phones that most people actually have access to. And one of these text apps is used to provide better healthcare. One particular example that I thought was very poignant was where often moms will try to text a doctor to ask a question, for example, why is my child hasn't been able to eat? I give my child water. My child vomits, so it's not able to stay hydrated. It's been one day or so. What do I do? And sometimes all it takes is a simple message and some very simple piece of advice like boil water. And that is able to help that person because a lot of people around the world are dying of dehydration.
你知道,一些非常简单的疾病,我们在英国或美国甚至都不会想到。所以第二天,我们给初创公司提出的问题是:因为我们没有足够的医生,我们如何尝试自动化系统,以便能够为需要帮助的人提供帮助?我认为借助语言模型,我们可以做得更好一些。这是关于增强人们的能力。所以我对 AI 的一个希望是,它能增强每个人,让他们在一定程度上更有能力提供急救或处理某些医疗状况。这是一个很好的问题例子,我在这里从未遇到过,所以不会去想。但当你去那里时,你会意识到这是一个大问题,通过一些创新,你实际上可以拯救许多生命。当然,还有很多其他问题我们在这里不会想到:不同的经济体系、安全挑战、经济问题、农业等等。只有通过成为那个世界的一部分,去那里,与人交谈,你才能了解到所有这些有趣的研究问题,并能够做出贡献,希望看到我们技术的真正好处,帮助地球上大多数人。
You know, very simple diseases that we don't even think about in the UK or the USA. And so the next day our question for the startup was: because we don't have enough doctors, how do we try to automate the system so that we can provide help to people who are in need? And this is something I think with language models we can do a bit better. It's about enhancing what people can do. So one of the things I hope for AI is that it enhances every person to become, to some extent, better capable at providing first aid or dealing with some medical conditions. That was a great example of a problem I don't think about because I never encountered it here. But when you go there, you realize it's a big problem, and with some innovation you could actually save many lives. Of course, there are many other problems we don't think about here: different economic systems, security challenges, economic problems, agriculture, and so on. It's only by being part of that world, by going there, by talking to people, that one learns about all these interesting research problems and is able to contribute and hopefully see the true benefits of our technology, helping most of the people on the planet.
Nando,我真的很喜欢你花那么多时间确保你触及的世界比我们大多数人都多,并且清楚地说明我们如何也能做出贡献。与你刚才谈论的一切相比,我觉得问下一个问题有点平淡,但对我们这样的人来说,我们看到的大量工作是编码。编码有时非常耗时且令人沮丧,因为大多数人清楚自己想要什么,但计算机就是不能完全正确地做到。你最近是 DeepMind 的 AlphaCode 的主要贡献者之一。你能多谈谈 AlphaCode 吗?
I really like how much time you spend on making sure you reach much more of the world than most of us end up spending time on, and then making it clear how we can also contribute, Nando. I feel a little pedestrian asking the next question in contrast to everything you just talked about, but for people like us, a lot of the work we see happen is coding. And coding can be very time-consuming and frustrating at times, because most people have a clear idea of what they want, but somehow the computer won't do it exactly right. You recently were one of the big contributors to AlphaCode at DeepMind. Can you say a bit more about AlphaCode?
首先,这是一次美妙的经历。作为项目的一部分,它是由一个大团队完成的,目睹人们的能力真是太好了。AlphaCode:每年很多学校,伯克利在这方面很擅长,参加国际或国内的编程竞赛。你会得到一个用英文描述的问题,例如,你有某些商品,你想以最佳方式分配它们来解决问题。来自所有大学的最聪明的学生组成团队,迅速编写代码来解决问题。从 AI 研究的角度来看,这似乎是一个非常有趣的挑战,因为它易于衡量且兴趣广泛。我们如何从那个描述自动生成解决问题的代码?AlphaCode 就是一次尝试。它利用了 Transformer,这并不意外。它们接受描述并生成代码。它还使用了一些其他技巧,比如采样多个解决方案,以及巧妙的思路来缩小样本范围,以便我们能够竞争并获得合理的成功率。它的表现确实非常出色。它并不完美。我自己看一段代码时,有一个 for 循环在做一堆事情,然后你意识到为什么这个 for 循环在这里?这只是机器创建的多余代码,没有被使用。有点像研究生在考试中什么都写;有些是解决方案,有些是他们正在思考的东西。它并不完美,但我喜欢有一篇博客说,感觉就像房间里有一只狗在说英语,但每个人都在指责这只狗语法不好。那篇博客最好地描述了我的感受。对我来说,它能做到这一点太神奇了,因为 20 年前我在 UBC 参与过这些编程竞赛的学生活动。看到机器能做到这一点,真是非常了不起。我认为有一天这些机器会从排名前 20%的表现进化到成为顶尖的程序员。除了这个挑战之外,它还导致许多人,不仅仅是 DeepMind,还有 OpenAI 等,提出了改进软件工程师和程序员工作流程的工具。我觉得这提醒了我:我发现自己现在打字时,输入一个'P',它就能完整写出我在调试时想到的整个 print 语句。感觉就像它在读我的心思。它变得非常真实。我和我的学生交谈,他们几乎一直开着这些代码补全插件。这并不意味着它总是自动补全正确,但很多时候它是正确的,如果不正确,他们可以不接受。显然,这极大地加快了他们在研究项目中的编码速度。
First of all, it was a wonderful experience. Being part of it is a project done by a large team, and it was wonderful to witness what people were capable of doing. AlphaCode: every year many schools, Berkeley is very good at this, compete internationally or nationally in coding competitions. You're given a problem in English, for example, you have certain goods and you want to distribute them in the best way to solve that problem. The smartest students from all universities get together in teams and quickly try to hack code that solves the problem. From a research perspective in AI, it seemed like a very interesting challenge because it's easy to measure and there's wide interest. How could we go automatically from that description to code that solves the problem? AlphaCode was an attempt at this. It capitalizes on using transformers, no surprise there. They take that description and generate code. It uses a few other tricks like sampling many solutions and clever ideas to narrow down the samples so that we can compete and get a reasonable success rate. It was truly remarkable how well it did. It wasn't perfect. I myself was looking at a piece of code and there's a for loop doing a bunch of things, and then you realize why is this for loop here? It's just extra code the machine created that is not used. It's a bit like grad students who throw everything in the exam; some of it will be the solution, some of it is things they were thinking about. It's not perfect, but I love that there was a blog that said it felt like there was a dog in the room speaking English, but everyone was pointing at the dog for not having proper grammar. That blog best described how I felt about it. For me, it was amazing that it could do that because I was involved with students in these coding competitions at UBC 20 years ago. To see a machine doing that is quite remarkable. I think one day these machines will evolve from having a performance of being in the top 20% to actually being some of the top coders. Out of it, besides that challenge, it has also led to many people, not just DeepMind but also OpenAI and others, coming up with tools to improve the workflow of software engineers and coders. I find that a reminder: I mean, I find myself nowadays like I type a 'P' and it completely writes the whole print statement I had thought about when debugging. It feels like it's reading my mind. It's becoming very real. I talk with my students and they have these code completion add-ons turned on pretty much all the time. It doesn't mean it always auto-completes correctly, but very often it does, and if it doesn't, they can always not accept it. Apparently, it speeds up their work tremendously as they're coding for their research projects.
确实如此,尤其是如果你不是每天都编码,只是每周偶尔花几个小时,那么你总是对语法有点生疏,但这个东西正好为你解决了这个问题。它能提高你的生产力,真是令人惊叹。现在,我认为关于智能体栈和辅助代码,真正有趣的一点是通用性,对吧?编码是一件非常通用的事情。相比之下,像学习下棋或围棋这样更具体的事情。那些突破非常有趣,因为方法最终在不同游戏之间共享,但智能体并没有共享。
It does, especially if you're not coding every day and you just dive in every week maybe a couple of hours, then you're always a bit rusty on the syntax, but this thing just solves that problem for you. It's quite amazing how much it can improve your productivity. Now, I think one of the things that's really interesting about the agent stack and help code is the generality, right? Coding is such a general thing. Compare that to something more specific like learning to play chess or Go. Those breakthroughs were very interesting because the methods were shared across the different games in the end, but the agents were not shared.
那是一个单独训练的智能体,是的,用同样的方法,但本质上是靠自己的神经网络单独训练来完成任务的,下一个游戏又是另一个单独的智能体和神经网络。而编程则要通用得多,事实上,你的很多工作都在推动通用性的前沿,从这些在单一任务上过度训练的专家智能体——它们也许能在这一件事上与人类竞争得很好,但这并不意味着它们真正理解其他任何东西。所以我认为你最近发表的 Gato 工作是推动通用性达到全新水平的一个很好的例子。你能多谈谈你在那项工作中做了什么吗?还有是什么启发了你?是什么让你认为这是可能的?因为对我来说,它效果这么好、一个智能体能做这么多事情,这很令人惊讶。
It was a separate agent separately trained, yes, with the same method, but separately trained by its own neural net essentially to get the job done and a separate agent separate neural net for the next game and so forth. And coding is much more general and in fact a bunch of your work is really pushing the frontier of generality, going from these specialist agents that are hyper trained on this one thing and sure maybe can compete with humans on this one thing really well, but then it doesn't mean they really understand anything else. And so I think the Gato work that you recently published is a great example of pushing the generality to a whole other level. Can you say a bit more about what you did in that work? And also maybe what inspired you? What made you think that this would be possible? Because to me it was surprising how well it worked and how it was such a general agent doing so many things with one agent.
是的,我认为对许多研究者来说,一直以来的梦想之一就是能够设计一个单一的智能系统,能够完成各种任务。我认为最引人遐想的是,它能做任何人类能做的事情,也许还能做其他事情,而且可能以与人类略有不同或完全不同的方式完成,但我们一直追求这种通用性。事实上,人们开始把 AI 称为 AGI,就是为了强调通用性的重要性。而且,如你所知,我是参加 CIFAR 会议的人之一,深度学习就是从那里开始的。在那个社区里,在加拿大高等研究院,你当时也参加了一些会议,还有吴恩达和其他优秀的人。我们总是与神经科学有联系。所以,一旦你想到神经科学,你就会意识到,这块组织是非常通用的。这块脑组织既可以用于感知,也可以用于行动。甚至你会看到这些植入物,你可以把带针的东西植入嘴里,再配上摄像头。如果你切断了视觉,随着时间的推移,通过这个针和适应深度处理,你实际上开始能够不用看就判断深度,有趣的是,视觉皮层仍然被招募来做这件事,尽管感觉是触觉的。所以这是触觉,不再是视觉。当然,神经科学中有很多研究,这是一个非常古老的想法,可以追溯到 70 年代,它指出可能只有一个通用算法可以做所有事情。那么这就引出了一个问题:当我们有了这些新的架构和进展,我们现在是否能够重新审视这个问题?我们该如何去做?我很幸运,在 DeepMind 期间遇到了一位名叫 Crete 的同事,他对同样的事情充满热情,并且非常擅长执行计划和想法。我也很幸运有一个才华横溢的年轻团队,他们非常渴望实现这一目标。所以我们花了三年时间研究这个,很多时候我们放弃了发表论文等,这通常是研究者的动力,而只是专注于引入越来越多的数据集,构建越来越多的基础设施。这个过程有起有落,因为有时你会想,我们是不是应该加快速度,或者只是有一个想法。但这就是这个项目的由来。我们的目标就是看看我们能否提出一个单一的模型,一个单一的大型神经网络,能够接受所有类型的输入,无论是面部、对话、语音、扭矩还是预感知等,并且能够与你交谈或控制机械臂。最终,你看到了结果。我们取得了一些成功。我认为这只是第一步。我也不认为这是唯一的工作。我很感激你那样想。我认为这是最早的工作之一。但你自己提出了决策 Transformer,这在精神上是一个非常相似的想法。我认为当时还有其他想法,比如 Sergey Levine 等。但正如我所说,我认为机器学习领域多年来一直有这样的驱动力。事实上,我记得在 CVPR 上与吴恩达讨论过这个问题,他对他的反向传播模型也非常感兴趣。那正是在他创办 Google Brain 之前。我相信,正是对这个梦想的信念在某种程度上推动了这一切。我认为我们还有很长的路要走,但这个梦想仍然存在。我很高兴人们现在看到这可能是可能的,并且有越来越多的人对此感兴趣。当然,仍然存在许多问题、许多挑战、许多我们尚未展示的东西。但这是我感到兴奋的事情,也是我想继续研究的事情,我认为包括你在内的许多其他人现在也在研究它。顺便说一句,你最近的论文,我也必须提一下,你也把它放在这里供所有研究者阅读,你最新的论文是利用 YouTube 视频生成视频。我认为这很出色,我认为这种想法对于构建更通用的智能体非常重要。
Yeah, I mean it's I think for many researchers it's always been one of the dreams is to be able to design one single intelligence system that is capable of doing, and I think what captures the imagination is that can do anything that a human can do, for example. And perhaps it can do other things and maybe it does it slightly different or completely different than how a human would do it, but we've always aimed for this generality. In fact, you know, people started calling AI AGI because to emphasize the importance of generality. And I mean as you know, I was one of the people that attended the CIFAR meetings where the sort of deep learning came from right from the beginning. And I think within that community, within the Canadian Institute for Advanced Research where, you know, I think you were attending a few of those meetings and Andrew Ng and all these wonderful people you had in your meetings. There was always a connection with neuroscience. And so, the moment you think about neuroscience, you actually realize, well, this piece of tissue is sort of very general. This piece of brain tissue, it can be used for perception or it can be used for action. And even you see these implants where you can implant something with pins in your mouth and with a camera. If you cut your vision, with time through this pin and adapt depth processing, you actually start being able to tell depth without seeing, and funny enough is the visual cortex that gets recruited to do this still, even though the sensation is tactile. So this is haptic, no longer visual. And of course there's many works in neuroscience, it's a very old idea going back to the 70s, kind of pinpoint this that there's perhaps just one universal algorithm that can do it all. And so then it does beg the question: when we have these new architectures and advances, are we now at the point where we can revisit that question and how would we go about doing it? And I was fortunate enough to have encountered at my time at DeepMind a colleague called Crete, who is very passionate about the same things and who is incredible at executing plans and executing ideas. And it's also fortunate to have a brilliant team of young people who are very keen to make this happen. And so we spent three years working on this, and at many points we sort of forewent going for publications and so on, which is often a drive that we researchers have, and just focused on bringing in more and more data sets, building more and more infrastructure. And there's ups and downs in that process because sometimes you wonder, are we should we not just speed this up and just here's a thought. But so that's how the project came to be. And we just aimed to see whether we could come up with a single model, a single big neural network that would be able to take all types of input, whether it's facial, whether it's dialogue, whether it's speech, whether it's torques, or pre-perception and so on, and be able to either talk to you or control a robot arm. And so eventually, well, you've seen the results. We have had some success with it. I think I do see that as a first step. I also don't see that as the only work. I'm very grateful for you thinking of it that way. I think it was one of the first. But you yourself came up with decision transformers, which is a very similar idea in spirit. And I think there were other ideas at the time, Sergey Levine and so on. But as I said, I think there's been a drive in machine learning for many years. And in fact, I remember discussing this at CVPR with Andrew Ng, who was also very interested in his backprop models. And in fact, that was at the time just before he went and started Google Brain. And it was believing in this dream that kind of drove that, I think, to some extent. And I think we still have a long way to go, but that dream is still alive. And I'm happy that people now see that it may be possible and that there's more and more people interested in working on it. And of course there remain many problems, many challenges, many things we haven't shown. But yeah, it's definitely something I'm excited about and that I've wanted to continue working on, and I think many others including yourself are working on it now. And by the way, your recent paper I have to mention, you're putting it out here as well for all researchers to go and read, your latest paper generating videos using YouTube videos. I think it's brilliant and it is this kind of idea that I think will be very important to build more general agents.
是的,这很有趣,它们都非常相关。我认为在你写的 Gato 论文中,它确实与决策 Transformer 密切相关,但我会说有一个根本性的不同,那就是你可以将一切标记化的洞见。人们过去认为多模态学习有这种模态和那种模态,各有各的。而你意识到,如果你正确地——当然标记化需要做对——但如果你正确地将所有内容转化为一个标记序列,一个单一的架构就可以处理所有内容并泛化,我认为这非常美妙。
Yeah, it's interesting, it's all very related. I think in the Gato paper that you wrote, yes, it ties to decision transformer quite closely, but I would say there's something fundamentally different which is the insight that you can tokenize everything and that people used to think of multimodal learning as there's this mode and that mode and it's all its own thing. And you realize that if you properly, of course a tokenization needs to be done right, but if you properly turn it all into a sequence of tokens, a single architecture can just process it all and generalize, which is really beautiful I think.
是的,这也是基于一个古老的想法。我第一次尝试做类似的事情是用一种叫 PAQ 的方法。我想很多人可能不知道 PAQ,但它是一种为文本压缩而发明的方法,并且在很长一段时间内是最先进的。PAQ 会把图像、文本等各种东西转换成序列,基本上是比特序列,直到——你知道,如果你想高效,你就要用到比特,并尝试编码。
Yeah, it's also based on an old idea. I first tried to do something like that using a method called PAQ. So I think a lot of people might not know PAQ, but it was a method invented for text compression and it was for a very long time state of the art. And so PAQ would take images and text and all sorts of things and would transform it to sequences and basically sequences of bits until, you know, if you want it to be efficient you kind of go to bits and you try to make the code.
这个架构本质上是利用上下文——也就是比特串的上下文——来门控一个神经网络并预测下一个比特。它由许多神经元组成,每个神经元都试图预测下一个比特。这是一个非常有趣的架构。所以我对它非常着迷,与此同时,Ilya 还在多伦多作为学生研究他的 LSTM 3 RNN。我记得我当时在尝试说服 Ilya 从 RNN 转向 BRCU,但我很高兴他没有听我的,因为他的工作实际上非常可行。但想法已经存在了:用一个单一模型来处理图像等等。事实上,当时我可能用的是 RBM 或稀疏编码模型。然后我取隐藏单元序列,用 PHU 和某种关系……当时我有一个学生叫 Byron Knoll,他现在在 Google,继续研究这个和类似的想法。他是一位出色的程序员,能够在比特级别进行所有操作。所以,这个想法——所有东西都应该只是一个单一序列——已经在文献中出现了。我记得我和 Scott 讨论过这个,他非常兴奋,并大力推动这个想法,最终在 DeepMind 取得了成功。实际上,Sergey Levine 和他的团队也提出了一个非常相似的模型。我想他们也是把所有东西都 token 化,然后使用序列 Transformer。所以这非常相关。我认为我们可能做得不同的是,我们投入了大量时间。我们可能开始得更早一些,并确保引入了尽可能多的数据集和任务。这是一个巨大的挑战,因为如果你想让一个智能体在环境中做 600 件事,那就意味着在训练期间需要运行 600 个环境。这在评估和数据处理等方面是一项巨大的工程。领导这项工作的 Sergio 和 Gomez 以及 DeepMind 团队,我认为他们做了出色的工程工作。Gabe 也是。是的,有很多工程师非常努力地工作,使基础设施成为可能,从而能够进行这些实验。
And then it is essentially context that was so kind engineered features using the context of the string of bits to gate a neural network and predict the next bit. And then it's an architecture consists of many neurons and every neuron is just trying to predict the next bit. It's a very interesting architecture. So I became quite fascinated by it at the same time that Ilya was working on his LSTM 3 RNNs as a student in Toronto. And so I remember doing working in this and I'm trying to convince Ilya that he should switch from his RNNs to a BRCU, but I'm glad he didn't because that actually was quite workable. But it was but the idea was already there where we were using a single model to deal with images, to deal with And and in fact at the time I think I was using an an RBM or or sparse coding model. And then and then I was taking a sequence of the the hidden units and then I was using doing PHU with a a and by by relation say I was a student of mine at the time, Byron Knoll, who is now at Google and has continued working on this and similar ideas. And he was a brilliant coder and he was able to hack all the stuff in at the bit level. Bit operations. Um but yeah, so that idea was all the idea of using everything should be just a single sequence was already um there in the literature. And so um yes, Scott I remember I discussed this with him. He was very excited about it and he definitely pushed for it and he was a champion of this idea DeepMind and eventually succeeded with it. And I think actually Sergey Sergey Levine and his team they also came up with a very similar model. Well, I think they were also tokenizing everything too just use a sequence transformer. So it's very closely related to that. I think what we did differently perhaps is we actually devoted a lot of time to it. I think we might have started a bit earlier and so working on it and we just made sure that we brought in as many data sets and as many tasks as possible. And and that's a big challenge because once you have if you want an agent to do 600 things in environments that means you need to be running 600 environments during your training. And that's a massive endeavor in in evaluation and data processing and so on. So um yeah, one of the the people that led this Sergio and Gomez and in the DeepMind team I think they did wonderful engineering work. Um Gabe as well. Yeah, there were lots of engineers really working very hard to to make that infrastructure possible to be able to ground those experiments.
谈谈通用智能体。我们讨论过的 Gato 本质上是提前从自身经验中学习,然后之后可以做事。我们还做了一些工作,你不需要提前完成所有学习。你可以去看视频,检索视频。这在我看来非常聪明。这基本上符合大多数人类的行为。如果我要修理家里的东西,我会找一个 YouTube 视频,展示如何修理这个排水管之类的。我不能做非常复杂的事情,但简单的事情,我去看个视频,突然就能做了,尽管我以前从未做过。但你们用 AI 智能体做到了这一点。你能再多说说智能体获得了哪些能力吗?
Talk about generalist agents. Gato we talked about is essentially learning ahead of time from its own experiences and then can do things later. We also did some work where you don't have to do all the learning ahead of time. You get to go watch videos, retrieve videos. Which to me seems brilliant. That essentially matches what most humans do. If I have to fix something in the house, I will find a YouTube video that shows me how I'm supposed to, you know, fix this drain or something. I can't do very complicated things, but the very simple things that I go watch a video and all of a sudden I can do it even though I've never done it before. But you've done this with AI agents. Can you say a bit more about what kind of capabilities did the agents acquire?
那个早期的工作,实际上是从我看我侄子玩《我的世界》开始的。他仅仅通过观看就变得非常擅长。我当时想,这就是我们应该学习游戏的方式。我们不应该在环境中做强化学习。网络上有那么多内容,我们应该从中学习。现在很多人都在这样做,我认为最近人们在《我的世界》上做得非常好。但当时我想,让我们尝试一下。让我们利用所有存在的东西。当然,这里有一个领域问题:YouTube 上的各种 Atari 版本与你的模拟器有多大不同。但如果你碰巧有一个模拟器,并且有很多视频,那么你可以利用它们。你可以做一些变换。会有小的领域差距,但你可以处理它,然后在模拟环境中尝试并学习玩游戏。至少对于 Atari,我们能够通过观看视频来玩游戏,并且比从头开始使用更昂贵的 RL 智能体玩得好得多。我认为我们当时的那篇论文展示了可能性的味道。在研究中,我经常喜欢找那些人们还没有真正努力尝试过的问题,只是为了展示某件事是可能的。因为我认为这就像激励学生一样:一旦他们相信某件事可以做到,他们就会去做。这项研究有点像这样。如果你能展示某件事是可能的,那么很快就会有其他论文和产品等远远超越它。我仍然认为我们在利用这个想法方面还有很长的路要走。正如你所说,互联网上还有所有那些“如何做任何事”的视频,我们可以利用这些视频。不仅仅是玩游戏,我认为最终,随着我们能够构建更好的世界视频模型——比如使用扩散技术、掩码 Transformer 等——我们很快就会看到这个想法的不同实现,人们真正能够从 YouTube 上的教程视频中学习,并将其迁移到实际任务中。我认为最明显的应用是如何让机器人递送东西。因为正如你所知,你在这方面比任何人都更有经验。收集机器人数据非常困难。我们经常看到人们通过远程操作机器人的手的视频,但当你真正尝试这些远程操作时,它并不真的有效。做好远程操作非常困难。当然,你还受到人类远程操作机器人时体力的限制。实验室里还会发生另一件事:你通过远程操作收集数据,然后你的夹爪坏了,或者你必须升级它,因为公司出了新的夹爪。任何这些机器动力学或外观上的变化都意味着你必须从头开始收集数据。所以,收集数据非常耗时。而且,仅仅因为必须重新收集而丢弃数据也非常浪费。
And that early work where we um so it you know, it actually started by me watching my nephew playing a Minecraft. And and and it became very good at it just by watching. And I was like, but that's how we should be learning games. We shouldn't be doing reinforcement learning in an environment. There's so much content in the web that we should just learn from it. And now it's many people do that and it's kind of I think people have done it very well with Minecraft recently. Um but at the time I thought, you know, let's try to do that. Let's try to capitalize on everything that exists out out there. And of course there is a question then of domain, you know, how different are all the versions of Atari in YouTube from the simulator that you have. But if you do happen to have a simulator and you have many videos, then you can take advantage. You can do some transformations. There'll be a small domain gap. Um but you can deal with that and then be able to sort of try things in the simulated environment and learn to play games. So at least for Atari we were able to go and just watch videos of what's going on and then be able to um play those games and be able to essentially, you know, play them much better than if we were using um much more expensive RL agents from scratch. Um I think that paper that we did back then, that was sort of a taste of what could be possible. Um, I I often in research I like to um find problems that you know, that people haven't really tried them really hard and and to just to show that something is possible. And because I think that's you know, just like when you inspire students and you go there and you just they once they believe it can be done, then they they will go ahead and and do it. I think what this research is a bit like that. If if you can show that it's possible to do something, um, then soon after there'll be a chain of other papers and products and so on that will far improve upon that. Um, and I still believe that we have a long way to go in terms of taking advantage of that idea. Um, there's still all of that video as you said, all how to do anything on the uh on the internet and we could take uh advantage of that video. Um, not just to play games, but I think ultimately as we're able to um, especially now that we're building much better video models of the world, um, you know, with diffusion techniques of masking transformers and so on. Um, I think we uh we will soon um, see different realizations of that idea where people really can sort of learn um, from uh how-to videos in YouTube and they're able to transfer it uh to uh um I think the the most obvious thing is how to get robots to deliver things. Because it's as as you know, uh because you have more experience in this than anyone, I think. Uh collecting robot data is very hard. Um we often see these videos Yes. of people teleoperating their hands of the robots and but then when you actually go and try these teleoperation stuff, it doesn't really work. It it's very hard to do good teleoperation. And of course, you're limited by how many humans um have the stamina to actually teleoperate uh robots. And and then there's this other thing that happens in the lab, which is you're teleoperating to collect data, and then, you know, your grip of breaks down, or it or you have to upgrade it because the company has a new gripper. And um and any of those changes in dynamics uh of the machine or in the appearance and so on, just mean that you have to collect data again from scratch. And so, it's very time-consuming to collect data. It's also very wasteful to just have to collect again and throw data away.
而且我认为,即使我们把全世界所有实验室的数据都汇集起来,可能还是不够。所以最终,你不得不去 YouTube 上看人们怎么做事情,然后想出创新的方法把知识迁移过来,让机器人能直接观看——就像我们人类看到别人做某事然后自己去做一样。我认为这种能力在未来几年内会越来越常见。我觉得这对机器人学来说可能是变革性的。我很好奇,如果机器人能直接观看人类做某事然后自己去做,这会对机器人学产生多大的变革?根据你的经验,是什么阻碍了我们实现这一点?是算法问题?还是硬件问题?
And I think even if we were to pull all the data from all the labs in the world, perhaps it's still not enough. So eventually, you do have to go to YouTube, see how people do things, and then be able to come up with innovative ways of transferring that knowledge so that the robots can actually just watch, just like we humans see others doing something and then we do it. I think that ability is something I expect we're going to see a lot more of within the next years. I think it could be transformational for robotics. I'd be curious to hear how transformational it will be in robotics if the robots can just watch a human doing something and then are able to go and do it. What prevents us from doing that based on your experience? Is it algorithmic? Is it hardware?
这是两者的结合。但我认为,如果我们能拥有一个足够强大的神经网络,让它观看一个人做某事,然后知道机器人应该如何去做,那将打开很多机会。可以是观看同一个场景中的人实时操作,这相对容易一些;也可以是从网上检索相关场景的视频,就像我们人类做的那样。我们通常不会有一个视频是别人在我们自己家里修理同样的东西,而是另一个房子里做非常类似的事情的视频。所以,是的,我认为这将是变革性的。我自己思考机器人学时,它本质上就是重复运动。机器人已经解决了很多问题很长时间了:制造汽车、制造电子产品等等。但世界上只有两三百万台机器人。然后这些机器人的机会就用完了,因为世界其他地方没有那么重复。再想想我现在正在做的事情:仓储以及类似难度的事情、农业、回收。我认为这是当前一代可能实现的事情。但绝对,我认为你所说的是那之后的下一个阶段。然后你可以想象:一旦你能做到那样,从一段视频中学习如何作为机器人做某事,我认为你就可以进入家庭,变得非常有用。有一个问题,正如你刚才提到的,一切都是为人类建造的。机器人的夹爪足够通用,能完成我们在家里想让它们做的所有事情吗?可能不行。但也许这会强烈激励一些非常优秀的机械工程师深入钻研,制造出坚固的类人手掌,从而能完成类似的事情。所以,是的,这对我来说有点奇怪的问题,因为感觉这已经是我 10 年或 20 年研究的目标了。想到也许我们能在未来几年内实现这一点,甚至更早,感觉有点奇怪。在某些方面,解决这个问题比解决机器人学更容易,因为你还得处理硬件问题。而硬件问题并不容易。还有经济因素也参与其中。制造机器并不便宜。我确实认为这很重要。我仍然认为机器人学是人工智能和智能技术研究中最重要的美德之一。因为对于危险活动,当某项活动对人类来说太危险时,你需要机器人。这些问题尤其出现在人类必须应对的灾难中,比如我们最近目睹的地震。在这些场景中,有机器能帮忙会非常有用。我也认为,最终,对于太空探索,依赖机器变得非常重要,因为人类去火星会很艰难,但我们已经可以送机器人去那里。所以我认为机器人是第一步。但如果有一天我们真的成为一个星际物种,那么机器人将扮演关键角色。它们将是使我们能够实现这一目标的赋能技术。
It's a combination. But I think if we can actually have a neural network that is sufficiently capable to watch a person do something and then know how a robot should do it, it would really open up many opportunities. It could be watching a person live in the same context, which would be a little easier, or it could be retrieving videos online from related contexts, like we as humans do. We usually don't have a video where somebody is fixing the same thing in our own house; it's a video in another house doing something very similar. So yeah, I think it'd be transformative. When I think about robotics myself, it's essentially repeated motion. Robots have solved problems for a very long time: building cars, building electronics, and so forth. But that's only two or three million robots in the world. Then you run out of opportunities for these robots because the rest of the world is not so repetitive. And then think about what I'm doing now: warehousing and related things in terms of difficulty, farming, recycling. I think that's the current generation of things that are possible. But absolutely, I think what you're talking about is the next thing after that. And then you can think: once you can do that, learn from one video how to do something as a robot, I think you can go into houses and be super helpful. There's a question, as you alluded to now, that everything is built for humans. Are the robot grippers general enough to do all the things we want them to do in a house today? Probably not. But maybe it'll give a very strong incentive for some really great mechanical engineers to dive in and build some robust human-size hands that could do similar things. So yeah, it's kind of a weird question for me because it feels like it's been the goal of my research for 10 or 20 years. And it's kind of weird to think that maybe we could actually do this in the next, who knows, handful of years, maybe even sooner. In some ways, solving that is easier than solving robotics because you also have to deal with the hardware question. And the hardware question isn't easy. There are economic factors that also factor into the equation. Building machines is not cheap. I do think it's important. I still think robotics is one of the most important virtues of research in AI and intelligent technology. Because for dangerous activities, where an activity would be too dangerous for a human, you need robots. These arise especially when problems arise that humans have to deal with, like earthquakes, as we've actually witnessed recently. It would be really useful to have machines that could help in those scenarios. I also think ultimately, for space exploration, it becomes really important to rely on machines because it would be very harsh for humans to go to Mars, but we can already send robots there. So I think robots are the first step. But if we do one day end up being an interplanetary species, then robots will play an essential key role in that. They'll be the enabling technology that will allow us to achieve that.
完全同意。从机器人切换到一个稍微不同的话题,正如你之前在这个对话中提到的,研究中特别有趣的是推动那些之前没有太多生命迹象的东西,然后突然你意识到这里可能有生命迹象。让我们展示一些新东西。有一项工作让我印象深刻,并影响了我很多工作,那就是你的“学习如何学习”的研究。你有一篇论文标题很巧妙,叫《通过梯度下降学习如何学习,而梯度下降本身也是学来的》。据我理解,核心思想是:为什么我们还要手动编写学习程序?难道学习程序本身不能也被学习吗?这在某种程度上是一个递归过程。这篇论文是几年前的了。我很好奇:也许你可以快速回顾一下你在论文中做了什么,同时也给出你今天对“学习如何学习”的看法。
Absolutely, agreed. Switching from robots to a slightly different topic, as you've alluded to in this conversation in the past, it's particularly interesting in research to push things that haven't really shown much sign of life before, and all of a sudden you realize that there can be some sign of life here. Let's show something new. A work that really stuck with me and influenced a lot of my work is your learning to learn work. You had this paper cleverly titled 'Learning to Learn by Gradient Descent by Gradient Descent.' As I understand it, the big idea is: why should we write the learning program still? Can't the learning program also be learned? It's a recursive process in some ways. This paper is from a couple years ago. I'm really curious: maybe you can quickly recap what you did in the paper, but also give your perspective on what you think about learning to learn today.
就我的经历而言,开始于我的合作者 Misha Denil 向我推荐了这篇论文。当时还有 Sutskever 等人做的类似工作。我第一次看到这个想法时觉得它非常棒,值得用我们现代的神经网络重新审视。我对此很着迷,因为我认为,尽可能多地利用我们拥有的数据。但有时我也觉得纯粹出于科学好奇心很有趣:事物是如何产生的?事物是如何涌现的?特别是,我有点哲学思考。我把进化看作一个学习过程,它创造了这些能够学习的生物机器。所以它是一个学习过程,一个动态适应过程,导致了能够学习微积分和代数等惊人事物的动态适应过程的产生。因此,对我来说,我们需要赋予神经网络这种能力是合理的。于是我们开始研究。我们尝试了很多东西,一开始有点沮丧,因为我们进展不大。然后 Martin,当时非常年轻的合作者 Martin Andrychowicz,他后来实际上和你一起工作并做出了出色的成果,他坚持了这个想法,最终让它成功了。
In terms of my experience with it, it started when I was pointed to this paper by one of my collaborators, Misha Denil. There was this paper by Sutskever doing this with, I think, several others. And it looked like a wonderful idea when I first saw it, that it was worth revisiting and doing it with our modern neural networks. I was fascinated by it because I think, as much as possible, I always believe in capitalizing on the data that we have. But sometimes I also find it interesting just for scientific curiosity: how do things arise? How do things emerge? In particular, I was interested in philosophizing a bit. I was thinking of evolution as a learning process that then creates these biological machines that can learn. So it's a learning process, a dynamical adaptive process that has led to the generation of dynamically adaptive processes that are capable of learning amazing things like calculus and algebra. And so to me, it made sense that we needed to endow our neural networks with that capability. So we started working on it. We tried quite a few things, and it was a bit frustrating in the beginning because we weren't getting very far. And then Martin, who was a very young collaborator at the time, Martin Andrychowicz, who actually went to work with you afterwards and did wonderful work with you, he sort of persevered with the idea and eventually got it to work.
嗯,还有,你知道,在另外几个人的帮助下,你和我……然后我变得非常兴奋,因为我觉得,这里有一个机会,让神经网络去学习其他神经网络应该是什么样子,算法应该是什么样子。就像我们仍然在手工设计学习算法。不幸的是,现在还是 Adam。很多年大多数人都用它,所以我们没能设计出比它好得多的东西。但是,有可能我们可以设计出更好的算法,让一个神经网络生成新的神经网络来解决另一个问题。你知道,这催生了很多关于学习如何学习的 brilliant 想法。那个领域出现了大量工作。很多工作实际上来自伯克利,如果你和你在南方的合作者……最近,我确实看到这一点在大 Transformer 和语言模型中体现出来。这些大模型的一个涌现能力让我非常惊讶,尤其是在少样本学习中:当你预训练你的模型时,我觉得你学到了那个初始的东西。然后当你给出一些上下文,特别是用少样本提示,给它几个输入-输出例子,然后你给出输入,它给出输出,我认为在提示过程中为 Transformer 创建上下文,本质上就是……你的模型学到的是你能够提示模型,让它从少量例子中学习。所以,我认为这就像少样本学习如何学习,规模惊人且具有通用性。所以,我现在看到这个想法实际上正在大量发生,这是大语言模型真正了不起的事情之一,就是看到这种涌现能力。但是,如果你仔细剖析并做数学计算,它和我们之前做的事情,比如一次性模仿学习和机器人技术,并没有太大不同。我们当时没有大模型,但我认为想法非常相似。当然,当时限制更多。现在,只是一个单一的大模型在做这件事。
Um and also, you know, with the help of few others, so you and so And um And then I became really excited about it because I I thought, you know, there's an opportunity here for uh neural networks to learn what the other neural networks should be like, what that what the algorithm should be. It's like we still hand engineer the learning algorithm. Um Unfortunately, it's still Adam. But many many years that most people use, so we haven't been able to engineer anything much better than that. Um Um but um there was the possibility that we could engineer better algorithms, that we could engineer have a neural network generate new neural networks that would solve another problem. Um And you know, that led to some lots of brilliant ideas on learning to learn. There there's an explosion of works in that area. Um a lot of those works actually came from Berkeley. If you from your collaborators in South And um And more recently, um I actually do see that manifesting itself with the big transformers and the language models. So this one of the emerging capabilities of these big models that has really amazed me is that I still see them um as especially in the few shot is like you you when you pre-train your model, I feel like you've learned that initial thing. And then when you give say some context, especially if you're prompting it with uh a few shot prompting, giving it a few sort of input example solution, input example solution, and then you're giving input example, and you're giving a solution, I see creating that context for the transformer and during prompting as essentially basically um um Yeah, with your model, what you learned was the ability for you to be able to prompt the model so that it's it's learning from few examples. So, I I see that as like few shot learning to learn um at amazing scale and generality. So, I I I I think now I see that the idea is actually happening a lot and that's been one of the sort of really remarkable things of the big language models is that is to see this emergent capability. But, if you start dissecting and doing the math carefully um it's not too far from the things that we were doing before with or you know, one shot imitation learning and robotics. Uh we didn't quite have the the the big models or something, but I think the ideas were quite similar. Um of course, there was a lot more constrained than what is. Now, it's just like one single big model doing it.
在语言方面,当前很多工作都集中在非常大的模型上,对吧?当然,它们取得了一些惊人的成功,拥有前所未有的能力。我的意思是,现在它们受到如此多的关注是有充分理由的。但是,还有另一条工作线,你实际上做了一些早期工作,这仍然让我着迷,那就是 AI 系统在必须共同解决问题的背景下学习交流,从而发明了交流的概念,因为这有助于它们更有效。也许有点像人类和动物学会交流的方式,因为这有助于它们更好地生存,完成更有趣的事情。嗯,最近这方面不太活跃,但我很好奇。我知道我有点让你为难,因为你可能没有准备这个问题,但是 Nando,你对将这些事情与当前的大语言模型结合起来有什么看法?或者也许几年后这些东西会取代今天的大语言模型,因为从根本上看,它更类似于人类如何获得语言。
In language, a lot of the current work focuses on the very large models, right? And of course, they had some amazing successes, unprecedented capabilities. I mean, there's a good reason there's so much focus on them right now. But, there's another line of work which you have actually done some of the early work which still fascinates me which is where the AI systems learn to communicate in the context of having to solve problems together and and hence invent the notion of communication because it helps them be more effective. Maybe a bit the same way humans and animals have learned to communicate um because it helps them survive better, helps them get more interesting things done. Um and it hasn't been as active recently, but I'm I'm curious. I know I'm putting a bit on the spot here because you probably haven't prepared for this question, but what what are your thoughts Nando on combining some of these things with the current large language models or maybe even these things possibly in a few years supplanting today's large language models because somehow it's fundamentally more similar to how humans came to language.
是的,在这个背景下重新审视这个想法绝对值得。我认为那会是一个好项目。嗯,你知道,这又回到了这个问题:你是使用所有数据,还是尝试从更科学的角度理解语言是如何产生的?我认为这仍然是一个开放问题。我们不知道我们是如何走到这一步的。所以,如果你务实,你必须意识到我们拥有全人类多年来在互联网上记录的所有人类语言。所以你可以利用这些来训练大模型,为什么还要费心去做其他类型的研究呢?这可能和我们研究历史、做许多其他事情的原因相同。嗯,因为这是一个未知。这是一个我们不知道答案的事情,我认为我们很多人都想知道语言是如何产生的。所以,那个最初的项目是由我当时的学生 Jacob Buster(现在是牛津的教授)和 Jan Assael(现在也在 DeepMind)领导的。我们尝试使用多智能体强化学习,创建了一些环境,其中智能体解决问题的唯一方法是相互发送有意义的信息。他们还设置了一些瓶颈,比如在 softmax 之前进行采样等,以强制模型在信道有噪声时使用离散通信,因为我们想知道它们是否会想出离散的符号或离散的表示来更有效地通信以完成任务。我们还试图看看它们是否有可能最终以组合的方式组合这些符号,就像我们用语言组合一样。事实证明这非常困难。我认为已经取得了一些进展,但我们还有很长的路要走。我的意思是,我喜欢的一个例子是,在自然界中,有一种猴子,我忘了它的学名,在东非,它们有特定的警报叫声,用来告诉同伴附近有老鹰、蛇或美洲豹。如果这些动物来了,它们会发出不同的声音来指示附近是什么动物。当然,猴子的反应会不同:如果是老鹰,它们会跳上树;如果是地面捕食者,它们会做其他事情。所以,很容易构建一个强化学习系统,让智能体学会这样做。对我来说,很难弄清楚的是下一步,当一只猴子最终学会操纵那个符号,或者开始想“我想要所有这些食物”的时候。
Yeah, it's definitely worth revisiting that idea in this context. I think that that that would be a good project. Um So you know, it comes back to this question of do you use all your data or do you look at try to understand take a more scientific view and try to understand how did language emerge? And which is still I think well, maybe it's still an open question. We don't know how it is that we got to this point. Um so if you're pragmatic, you have to realize that we have all this human language produced by the whole human race through the years on the internet recorded mostly. Um so you could take advantage of that to train big models and why should we bother to go into this other type of research? It's probably the same reason why we do history, why we do many other things. Um because it's it's an unknown. It's something for which we don't know the answer and and and I think it's I think many of us would love to know how language comes to be. And so that initial project that was led by my student then Jacob Buster, the professor now in Oxford, and Jan Assael was also now at DeepMind. Um trying to use multi-agent reinforcement learning and sort of created environments where the only way the agents could possibly uh solve the problems was if they were to send information to each other that was meaningful. And they also put some of these bottlenecks like sampling before the softmax and so on to force um the models to use discrete communication if then if the channels were noisy cuz we wanted to know whether they would come up with um discrete um symbols or discrete uh representations to communicate more efficiently um to be able to solve the task. Um We also tried to see whether it would be possible for them to learn to sort of eventually combine these symbols in uh uh compositional ways in just in just in the same way that we compose with language. Um and that turns out to be very hard. And I think there was there's been some progress, but there's still we still have a long way to go. I mean, one example that I love is um like in nature you have mon- you know, like there's there's monkey in um I forget the name of that I at some point I knew the scientific name of this monkey. So, the East Africa and but they essentially have different um they have certain alarm keys and and they want to tell if there's nearby it's uh um an eagle or a snake or some um I can't remember. I think it's a jaguar there, too. If one of these animals is coming, then you they will produce a different sound that is indicative of what animal is nearby. And of course, the response of the monkeys will be different if it's an eagle or you know, if it's an animal predator on the ground. If it's an eagle they can just jump up tree. Um but if they're on the canopy on top then you know, they become Sorry, if it's an eagle it's usually on top of the tree cuz they don't want to you know, they they could be vulnerable from the air. Um so they will do something appropriate to uh safeguard themselves from the predator. Um and so it's it's easy to build an RL system where the agents will learn to do that. Um the thing that was very hard to figure out for me is the next step when eventually one monkey sort of learns to manipulate that symbol. Or starts thinking you know, I I would like to get all this food.
但还有另一只我不喜欢的猴子,我不想让它得到我的食物。所以也许我会告诉这只猴子它是只鹰,尽管它确实是只鹰。这样我就能得到所有食物。不幸的是,这是一种相当欺骗性的开启智能的方式,但缺乏更好的例子,它能够推理这些机会,达到那种高层次的抽象。当这些机会,仅仅是离散符号的声音,超越了直接的物质世界时,我们就能开始操纵它们,就像我们人类使用词语,或者想出这些水平的 8 字形,我们称之为无穷大。当我们操纵它们时,我们能够创造关于宇宙中从未见过的事物的知识,最终用一些 fancy 望远镜看到它们。科学,尤其是宇宙学等,很大程度上源于我们操纵这些抽象符号。我们可以做出远远超出我们能看到或经历的任何事物的预测。而如何创造新符号、新知识,我认为仍然是一个大问题。即使在 AI 中,这可能对语言模型变得非常相关,不仅仅是复述人类的行为或组合网络上的文本,而是语言模型可能坐下来喝杯咖啡,开始思考宇宙的奥秘,开始创造新符号、新抽象,从而进行推理,对宇宙做出预测,并提出可测试的方法来验证这些预测。所以我认为我们不可避免地会重新审视学习语言的问题,这当然与学习交流有很大关系。
But there is this other monkey that I don't like and I don't want this other monkey to get my food. And so maybe I will just tell this monkey that it's an eagle even though it's an eagle. And then that way I can get all the food. That was a rather deceitful way to start intelligence, unfortunately, but in lack of a better example, it's being able to then sort of reason about those chances to go to that sort of high level of abstraction. It's when those chances, just the sounds of the discrete symbols, sort of transcend the immediate material. And then we can start manipulating them, just like we humans use words or we come up with these figure eights that are horizontal; we call them infinities. And when we manipulate them, we're able to create knowledge about things we've never seen in the universe, and eventually we see those things with some fancy telescopes. So much of science has come, especially cosmology and so on, as a result of us manipulating these abstract symbols. We can make predictions that are so far beyond anything that we can see or will ever experience. And how you could create new symbols, create new knowledge, I think is still a big open question. Even in AI, perhaps it's going to become very relevant for language models, which is not just to paraphrase what humans do or sort of combine what text there is on the web, but actually language models that maybe sit down with coffee and start thinking about the mysteries of the universe and start creating new symbols, new abstractions that allow them to then do some inferences and be able to make some predictions about the universe and be able to also come up with testable ways to verify those predictions. So I think invariably we will revisit this question of learning language, which of course has a lot to do with learning to communicate.
这很有趣,因为你现在谈到的是能够实际运行实验的智能体,对吧?不仅仅是输出文本,而是可能按下网站上的某个按钮,或者物理地嵌入机器人中,在现实世界中尝试事物,观察结果,从中提出新假设,并希望比以往学得更快,对事物运作有更深的理解。这真的很迷人。我很好奇我们能否暂时退一步,Nando。当你思考未来,比如 5 到 10 年的人工智能时,你最兴奋的是什么?
It's very interesting because you're now talking about agents that could actually run experiments, right? Don't just output text, but would maybe press some buttons somewhere on a website or maybe physically embodied in a robot, try things out in the real world, see what happens, and from that make new hypotheses, and learn faster, hopefully, than they could otherwise and have a deeper understanding of how things work. That's really fascinating. I'm curious if we can just zoom out for a moment, Nando. When you think about the next, let's say, 5 to 10 years in artificial intelligence, what are the things that you are most excited about?
嗯,我认为很多我们今天讨论过的事情。我想象我们今天讨论的一切。我认为现在的主流讨论当然是关于大型语言模型。当然,那里出现了重要的问题:如何让语言模型真正事实准确?所以有真理的问题。有关于如何让这些技术真正安全的问题?如何让它们对人类真正有用?然后还有工程问题,确保我们用更少的能量训练它们。如何让它们更便宜?更容易访问等等。当然,所有这些都伴随着安全性和保障性。还有如何将它们情境化,因为我们讨论了交流与语言。有不同类型的交流。有时,比如你在这里采访我,我知道如果我只开始告诉你我今天午餐吃了什么,然后花剩下的时间在这上面,那是没有意义的;我不认为观众会欣赏。所以情境在很大程度上决定了你谈论的内容。情境有多种形式:可能取决于你在哪里,你和谁说话。我认为这是我们需要取得进展的地方。我确实认为我们会看到很多进展,因为有很多驱动力会让我们朝那个方向前进。此外,我对科学问题、我们讨论过的思想实验感兴趣。特别是我想到我们在脑海中做的实验,这些思想实验就像爱因斯坦曾经谈论的:想象有人以光速旅行等等。我希望有一天语言模型能做到这一点。因为物理学家做的很多实验,比如宇宙学家等,斯蒂芬·霍金擅长的那种,都是你用纸笔做的实验。所以只要你有一种使用工具、外化知识、外部存储、检索知识的方法,并且然后你可以分组事物并创建新的抽象来讨论问题,我希望我们会看到很多科学进步。我认为科学进步会——我的意思是,你能想象如果机器开始告诉我们,‘不,伙计们,你们对如何建造聚变反应堆的理解全错了。让我告诉你们,我花了最后几秒钟思考。我推导出了相当于过去 200 年数百万人类物理学家的成果,以下是我关于如何建造一个小型核反应堆的结论。我创建了一个模拟;你去测试,验证它是否正确。这是你需要验证它的方法;有一些测试。’我的意思是,那将是美妙的。我们开始在生物学中看到一点。生物学已经发生了一场革命。显然,我为 AlphaFold 感到非常自豪。随着 AlphaFold,我们看到许多生物学家对神经网络和机器学习产生了浓厚兴趣。所以我们看到了很多关于这些纳米机器和蛋白质等的工作。我们还没有完全达到,但我希望神经网络最终能够思考问题、推理许多复杂变量,并能够像我们一样通过推理、抽象来推导新知识,也许还能建议应该进行哪些实验、创造哪些机器。是的,这就是我希望这将带我们达到的。因为如果它解决了我们的能源问题,并带来安全的聚变反应堆为整个世界提供能源,那将解决饥荒等问题和经济问题。
Well, I think a lot of the things we've talked about today. I imagine everything we talked about today. I think the mainstream discussion now is of course on big language models. And of course important questions arise there: how do you make the language models actually be factual and accurate? So there's the question of truth. There are questions about how do you make these technologies actually safe? How do you make them truly useful for humans? And then there are engineering questions that will make sure that we train them with a lot less energy. How do we make them also cheaper? Easier to access and so on. And of course all that comes with security and safety. And also how do we contextualize them, because we were talking about communication and language. There are different types of communication. Sometimes there is communication where, for example, you're interviewing me here. I know that it's pointless to just start telling you about what I ate for lunch today and spend the rest of the hour on that; I don't think the viewers would appreciate it. And so the context to a large extent determines what you talk about. Context comes in many ways: it could be where you are, who you're talking to. I think that's something we need to make progress on. And I do think we're going to see a lot of progress because there are a lot of driving forces that will make us go in that direction. Also, I'm interested in the sort of scientific questions, the thought experiments that we talked about. Particularly I'm thinking of the experiments that we do in our minds, these thought experiments like Einstein used to talk about: imagine someone traveling at the speed of light and so on. This is something that I would hope one day the language models could do. Because a lot of the experiments that physicists do, for example cosmologists and so on, the sort of thing Stephen Hawking was good at, these are experiments you do with pen and paper. So provided you have a way of using tools, of externalizing knowledge, of storing externally, retrieving this knowledge, and provided that then you can sort of group things and create new abstractions to be able to talk about problems, I would hope that we would see a lot of scientific progress. And I think that scientific progress would—I mean, can you imagine if the machine starts telling us, 'No, guys, you had this all wrong with how to build fusion reactors. Let me tell you, I've spent the last few seconds thinking about it. I derived the equivalent of millions of human physicists for the last 200 years, and here are my conclusions on how you could build a small nuclear reactor. And I've created a simulation; you go and test that, verify that it's correct. This is what you need to verify it; there are some tests.' I mean, it would be wonderful. And we're starting to see that a bit in biology. It's been a revolution in biology. Obviously I'm extremely proud of AlphaFold. And with AlphaFold, we've seen a lot of biologists getting very interested in neural networks and machine learning. So we see a lot of work on these nanomachines and proteins and so on. And we're not quite there, but I would hope the neural networks eventually, with the ability to think about problems and reason about many complex variables and be able to derive new knowledge just like we do by reasoning, by abstracting, and maybe by advising on what experiments should be conducted, what machines could be created. Yeah, that's what I would hope this will take us. Because if it takes us to solving our energy problems and getting safe fusion reactors to bring energy to the entire world, that would solve the problems of famine and so on and economic problems.
我们开始看到比过去更好的气候模型,如果模型能够推理并帮助我们做出更好的预测,知道我们可以采取哪些干预措施来保护环境,那将很有趣。甚至经济和政治系统也非常复杂,即使是善意的政策制定者也难以理解。所以如果机器能帮助人们提出更好的经济政策、环境政策等,那将非常棒。我认为这值得成为主流。至少这是我的梦想。我很想看看这会如何发展。当然,这在很大程度上是增强科学家,但我也想增强那些从事体力劳动的人。我的父母去工作;小时候,我和父亲一起在建筑工地工作——他是一名电工——那是艰苦的工作。对于地球上的大多数人来说,你无法像我一样探索想法。我很幸运能在这样一个美好的环境中工作,周围都是非常聪明的人,他们喜欢向我解释事情并一起头脑风暴。但对许多人来说,你去工作是为了养家糊口,赚钱,照顾家人。有时工作很辛苦,你很累,你弄伤了脚趾但还得去上班。这就是现实。所以对于许多繁琐的任务,帮助人们很重要,但这需要艰苦的工作。这就是机器人技术的用武之地——不一定是取代劳动力,而是增强或让人们更容易完成工作的机器,让他们能够专注于更有创造性的努力。就业市场将会演变,但重要的是,如果我们让所有特权工作都演变,我们也要确保不那么特权的工作也演变,这样他们就能拥有更有意义的生活,也许能花更多时间与家人在一起。我知道这听起来很反乌托邦,还有经济学、动机、心理因素等问题,但这些都是我们需要不断努力的事情。
We're starting to see better climate models than we had in the past, and it would be interesting if the models could reason and help us make better predictions about what interventions we could do to protect our environment. Even economic and political systems are very complex for well-intentioned policymakers to comprehend. So if machines could assist people to come up with better economic policies, environmental policies, and so on, that would be wonderful. I think it deserves the mainstream. At least that's my dream. I'd love to see how this goes. And of course, a lot of this is augmenting scientists, but I also would like to augment people who are doing labor. My parents went to work; as a kid, I worked with my dad in construction—he was an electrician—and it's hard work. For most people on this planet, you don't get to explore ideas like I do. I'm so fortunate to work in this wonderful environment full of incredibly smart people who love to explain things and brainstorm. But for many, you go to work to bring home the bacon, make money, look after your family. Sometimes it's hard work, you're tired, you break a toe but still have to go to work. That's the reality. So for many tedious tasks, it would be important to help people, but that involves hard work. That's where robotics comes in—not necessarily replacement of labor, but machines that augment or make it easier for people to do their jobs, allowing them to focus on more creative endeavors. The job market will evolve, but it's important that if we evolve all the privileged jobs, we also ensure jobs for the less privileged evolve so they can have more meaningful lives, perhaps spend more time with their families. I understand this is dystopian, and there are questions of economics, motivation, psychological factors, but these are things we need to work on.
这真的很美,Nando。我想问你最后一个问题。我知道你的经历:剑桥博士,伯克利博士后,UBC 教授,然后是牛津,你自己的创业公司,最后到了 DeepMind/Google。但在此之前,是什么最初让你对人工智能感到兴奋?
This is really beautiful, Nando. I'd like to ask you one last question. I know your path: PhD at Cambridge, postdoc at Berkeley, professor at UBC, then Oxford, your own startup, ending up at DeepMind/Google. But before then, what was the original thing that got you excited about artificial intelligence?
我只是很幸运。我认为生活中很多事情都不是你做出的决定,而是你发现自己处于某种环境中,然后尽力而为。我不能说我提前计划得很好。也许我计划读一个学位,知道它至少会持续四年。但通常你会遇到挑战或机会,必须确定该做什么。我很幸运,在本科三年级时,有一个关于神经网络的课题。我的控制学教授,威特沃特斯兰德大学的 MacLeod 教授提到了它。我想,‘这些是什么?’于是我去图书馆。起初我想,‘哦不,我不想做这个,这会是像 TCP/IP 协议之类的东西。’但完全不一样。你打开书,里面有一张大脑的图片。我被迷住了。图书馆里有几本书,我读得停不下来。然后我不得不去用 MATLAB 实现反向传播。那是大约 30 年前的事了。本科时接触到了正确的东西。我很幸运有这个机会。
I was just lucky. I think a lot of things in life are not decisions you make, but you find yourself in a context and make the best of it. I couldn't say I planned things well ahead. Maybe I planned to do a degree, knowing it would last at least four years. But often you encounter a challenge or opportunity and have to identify what to do. I was lucky that in my third year of undergrad, there was a project on neural networks. My control professor, Professor MacLeod at the University of the Witwatersrand, mentioned it. I thought, 'What are these?' So I went to the library. At first I thought, 'Oh no, I don't want to do this, it's going to be like TCP/IP protocols.' But it was completely different. You open the book and there's a picture of a brain. I got hooked. There were a few books in the library, and I couldn't stop reading. Then I had to go to MATLAB and implement backprop. That was about 30 years ago. Right exposure during undergrad. I was lucky to have that opportunity.
那么,我感到非常幸运你能抽出时间来参加播客。真的很享受这次对话。非常感谢你抽出时间。
Well, then I feel really fortunate you had the time to come on the podcast. Really enjoyed the conversation. Thanks so much for making the time.
非常感谢你,Peter。我喜欢你的播客,多年来一直很享受收听,尤其是在疫情期间。它总是一个好伴侣。你做得非常棒。它真的把我们做的事情带给了那么多人。
Thank you so much, Peter. I love your podcast, and it's been a pleasure to listen to it through the years, especially during the pandemic. It was always a good companion. It's a wonderful job you do. It really brings what we do to so many people.
谢谢你,Nando。真的很感激。
Thank you, Nando. Really appreciate it.