Geoffrey Hinton 谈 AI 危险与智能的两种范式

Geoffrey Hinton on AI Dangers and the Two Paradigms of Intelligence

杰弗里·辛顿 Geoffrey Hinton · MIT IMES · 2026-06-30 · 约 78 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

在 MIT 的演讲中,Geoffrey Hinton 解释数字智能为何能比生物大脑更快地共享知识、LLM 如何真正理解语言,以及超级智能为何可能发展出自保与夺权的子目标。

In an MIT lecture, Geoffrey Hinton explains why digital intelligence can share knowledge far faster than biological brains, how LLMs actually understand language, and why superintelligent AI may develop sub-goals of self-preservation and power-seeking.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 38)

全文 · Full transcript(中英对照)

Alex Shaik 开场介绍 Introduction by Alex Shaik

Host

我是 Alex Shaik,医学工程与科学研究所所长。很高兴邀请大家参加我们的首场 RTOR 讲座。我简单打个招呼,然后交给 Elazar,他将介绍 Richtor 博士和系列讲座。我们很高兴大家都能来。声音大点?竖大拇指是这个意思吗?我靠近点。再试一次。Alex,Vines 主任,感谢各位到来。现在交给 Elazar,他将介绍 Richtor 博士和讲座系列。

I'm Alex Shaik. I'm the director of the Institute for Medical Engineering and Science. It's our pleasure to invite you to join us for our first RTOR lecture. I'm just going to briefly say hello and then turn it over to Elazar who will introduce Dr. Richtor and the seminar series. We're delighted you were all able to join us. Louder? Is that what the thumbs up are? I'll come closer. So, trying one more time. Alex, director Vines, thanks for coming. I'm going to turn it over to Elazar who will introduce Dr. Richtor and the lecture series.

Elazar Edelman 介绍 Judith Richter Elazar Edelman's Introduction of Judith Richter

Host

谢谢 Shayik 教授,欢迎 Hinton 教授。大屠杀的吟游诗人埃利·维瑟尔写道,一个人最终是由其感恩的能力定义的,这不仅仅是一种感觉,更是一种选择的行为。他提醒我们,对于那些与历史长影相连的人,每一个小时都是一份奉献。将我们的祝福据为己有,就意味着背叛了这些感受。我们的赞助人体现了这些真理。作为大屠杀英雄 Rael 和 Arnot Spiegel 的女儿,她不仅继承了历史,还构建了辉煌的未来。她在学术界和工业界都取得了卓越成就。然而,她拒绝将名声和财富作为自己的纪念碑。相反,她将自己的遗产转化为所有人的活盾牌。通过将精力、资源和心血投入到提升多元且冲突的社区中,她证明了真正的财富在于我们为他人创造的公平。她审视了一个破碎的世界,询问谁最需要她,并以同情和优雅挺身而出。我很幸运能称她为朋友,也很荣幸介绍 Judith Richter 博士,她的慷慨资助了这个讲座系列。

Thank you, Professor Shayik, and welcome, Professor Hinton. Ellie Weisel, the bard of the Holocaust, wrote that a person is ultimately defined by their capacity for gratitude, not just as a feeling, but as an act of choice. He reminded us that for those connected to the long shadow of history, every hour is an offering. And to keep our blessings to ourselves would mean to betray those feelings. Our sponsor embodies these truths. As the daughter of Holocaust heroes Rael and Arnot Spiegel, she didn't just inherit a history, she built a magnificent future. She has soared in the halls of academia and industry. Yet she refused to raise her fame and fortune as a monument to herself. Instead, she transformed her legacy into a living shield for all. By investing her energy, her resources, her heart into elevating diverse and conflicting communities, she proves that true wealth is measured by the equity we create for others. She looked at a fractured world. She asked who needed her most and rose to the occasion with compassion and grace. It is my blessing to call her friend and my honor to introduce Dr. Judith Richter whose generosity has endowed this lecture series.

Judith Richter 演讲 Judith Richter's Speech

Judith Richter

谢谢 Lazar 的美言,也感谢 MIT 给予我这份非凡的荣誉。能在这里参加一个致力于教育、科学与和平的讲座系列的启动,对我意义深远。这些价值观以最个人化的方式塑造了我的人生。正如 Elazar Edelman 教授提到的,我是大屠杀幸存者的女儿。我的父亲 Arnospie 是奥斯维辛集中营的囚犯——顺便说一句,最近 PBS 上有一部电影在流媒体播放,想了解更多的人可以看看——我的母亲 Rael 在德国被奴役并被迫进入劳改营。他们教会我,虽然所有物质的东西都可以被夺走——你的房子、你的家、你的财产——但唯一没人能夺走的是你的知识。知识赋予力量。知识使人性化。知识在分享中增长。知识是尊重与和平的基础。然而即使在今天,我们仍在分裂人们,尤其是在世界各地的冲突地区。我的旅程始于 Medol,一家心血管设备公司。我和丈夫 Kobi 开发了心脏的近端支架。每颗心脏的跳动都一样。每个支架都适合每颗心脏。心脏是共享人性和同理心的普遍象征。这一理解促使我创建了“心灵近校”,将年轻的约旦人、巴勒斯坦人和以色列人聚集在一起,通过医学科学了解心脏和彼此。现在我想给大家看一段视频,你们将看到这些。

Thank you Lazar for your kind words and to MIT for this extraordinary honor. It is deeply meaningful for me to be here at the launch of a lecture series devoted to education, science, and peace. Values that have shaped my life in the most personal way. As Elazar Professor Edelman mentioned, I'm a daughter of Holocaust survivors. My father Arnospie, an Auschwitz prisoner — by the way, there is a movie streaming on PBS these days, so anyone who wants to know more can look at that — and my mother Rael, enslaved and forced into labor camps in Germany. They taught me that while everything material can be taken away from you — your house, your home, your possessions — the only thing nobody can take from you is your knowledge. Knowledge empowers. Knowledge humanizes. Knowledge grows when shared. And knowledge is a foundation of respect and peace. And yet even today we continue to divide people, especially in places of conflict around the world. My journey began with Medol. It's a cardiovascular device company. My husband Kobi and I developed the near stent for the heart. Every heart beats the same. Every stent fits every heart. The heart is a universal symbol of shared humanity and empathy. That understanding led me to create the Near School of the Heart, which brings together young Jordanians, Palestinians, and Israelis to learn together through medical science about the heart and about each other. I would like to show you now a video in which you will be exposed to that.

心灵近校视频介绍 Video Description of Near School of the Heart

Narrator

20 多年来,“心灵近校”一直是该地区的一束希望之光,在太多人心中充满偏见和暴力的地方。近校由医疗设备公司 CEO Judith Richter 博士创立,实现了她毕生的梦想:促进年轻人在科学和医学领域的知识和创造力,并鼓励这些学生通过学习过程建立跨文化桥梁。近校是一个独特的学术和社会项目,汇集了来自不同背景的青少年学习心脏病学基础。每年,近校招收以色列、巴勒斯坦和约旦的学生,让他们一起学习和成长。从他们到达的第一刻起,该项目就通过将学生分配到文化多元的学习小组来打破障碍。小组之间的竞争和成功的热情鼓励了每个小组内部的合作。学生们体验并学会尊重文化差异,克服偏见,并理解当他们相互依赖时,他们会更强大、更具创新性。近校的学术项目由 MIT 医学工程与科学研究所的 Elazer Edelman 教授开发。来自美国和中东的世界级心脏病专家自愿抽出时间前往该地区教学。除了获得基础科学知识外,学生们还独立工作,并面临设计新的心脏病学解决方案的挑战,从而激发他们的想象力和创造力。尽管该地区局势紧张,近校仍持续运营,如今自豪地拥有约 1200 名毕业生。超过 25% 的近校毕业生正在从事医生或其他医学相关职业。心灵近校的成功提供了一个新模式,可以在冲突阻碍和平与合作的其他地区复制。二十多年来,中东心脏地带的近校学生证明了敞开心扉和加深理解是可能的。

For more than 20 years, the Near School of the Heart has been a bright ray of hope in a region where prejudice and violence are present in the hearts of too many people. The Near School was founded by Dr. Judith Richter, a CEO of a medical device company, fulfilling her lifelong dream to promote knowledge and creativity among young people in the fields of science and medicine and to encourage those students to build bridges across cultures through the process of learning. The Near School is a unique academic and social program that assembles teenagers from different backgrounds to learn the basics of cardiology. Every year, the Near School accepts students — Israelis, Palestinians, and Jordanians — to learn and grow together. From the first moment they arrive, the program begins breaking down barriers by placing students in culturally diverse learning groups. The competitiveness between the groups and the passion to succeed encourage collaboration within each group. The students experience and learn to respect cultural differences, overcome prejudices, and understand that they are stronger and more innovative when they rely on each other. The Near School's academic program was developed by Professor Elazer Edelman from the Institute for Medical Engineering and Science at MIT. World-class cardiologists from the United States and the Middle East volunteer their time to travel to the region and teach. In addition to the basic scientific knowledge that they acquire, the students also work independently and are challenged to design new cardiologic solutions in a way that sparks their imagination and creativity. Despite times of great tension in the region, the Near School has operated continuously and today is proud to have about 1,200 graduates. More than 25% of Near School graduates are pursuing careers as physicians or other medical-related professions. The success of the Near School of the Heart provides a new model that can be replicated in other regions or areas where conflict has formed a barrier to peace and cooperation. For more than two decades, Near School students in the heart of the Middle East have proven that opening hearts and greater understanding are possible.

Judith Richter 闭幕致辞 Judith Richter's Closing Remarks

Judith Richter

谢谢。在过去的 28 年里,学生们发现了一些简单的真理。多样性是宽容和尊重的基础,差异不是威胁,而是创造力和力量的源泉。通过共同学习和生活,他们建立了信任、持久的纽带,并理解了合作让我们更强大。已有超过 10200 名学生毕业,近 300 人从事医学职业。其中 140 人是医生,医学博士。他们是活生生的证明:当来自不同背景的人一起工作时,他们能够发展知识,创造新思想、创新和希望。我要向 Elazar Edelman 教授表达最深切的感谢,他塑造并发展了近校的学术基础,他的愿景至今仍在指引我们。我们聚集在 MIT,一个知识塑造未来的地方。如果我们用知识来连接而不是分裂,如果我们投资于促进理解的教育,我们就能建立一个由合作而非冲突定义的未来。这就是近校和这个讲座系列背后的愿景。今天我们目睹了科学和人工智能技术的非凡进步,它们正在改变我们的世界,并挑战我们深入思考它们对人类的意义。像所有强大的知识形式一样,AI 承载着非凡的承诺,也要求深刻的责任。

Thank you. For the last 28 years, students have discovered simple truths. Diversity is the foundation of tolerance and respect, and difference is not a threat, but a source of creativity and strength. By studying and living together, they build trust, lasting bonds, and the understanding that cooperation makes us stronger. More than 10,200 students have graduated and nearly 300 have created careers in medicine. Out of them, 140 are doctors, medical doctors. They are living proof that when people from different backgrounds work together, they can develop knowledge and create new ideas, innovation, and hope. I want to express my deepest gratitude to Professor Elazar Edelman, who shaped and developed the academic foundations for the Near School and whose vision continues to guide us until today. We are gathered here at MIT, a place where knowledge shapes the future. If we use knowledge to connect rather than divide, and if we invest in education that fosters understanding, we can build a future defined by cooperation, not by conflict. That is the vision behind the Near School and behind this lecture series. We are witnessing today extraordinary advances in science and in artificial intelligence technologies transforming our world and challenging us to think deeply about what they mean for humanity. Like all powerful forms of knowledge, AI carries extraordinary promises and calls for profound responsibility.

主持人开场介绍 Introduction by Host

Host

它在改善健康、扩大教育和帮助应对全球挑战方面的潜力是巨大的,但其影响将取决于我们赋予它的价值观。尤其有意义的是,我们的首位演讲者是该领域的先驱,也是在使用中要求承担责任的主要声音。我想邀请 Lydia Boriba 教授。谢谢。

Its potential to improve health, expand education, and help address global challenges is immense, but its impact will depend on the values we bring to it. It is especially meaningful that our inaugural speaker is a pioneer in this field and a leading voice on the responsibility requested in its use. I would like to invite Professor Lydia Boriba. Thank you.

Host

非常感谢,Juliet。那么,现在有请 Hinton 教授。虽然我确信我的介绍可以省略,但我还是想说几句。Hinton 教授的学术生涯始于苏格兰爱丁堡,他在 70 年代获得了人工智能博士学位。他在加州大学圣地亚哥分校做了博士后,并在卡内基梅隆大学计算机科学系任教五年,之后转到多伦多大学计算机科学系,现在是名誉教授。从 2013 年到 2023 年,他在 Google 兼职工作,成为副总裁和工程研究员。他是神经网络的先驱,也是引入极其重要的反向传播算法的研究人员之一,并且是第一个使用反向传播来学习词嵌入的人。Hinton 教授对神经网络研究的许多其他贡献包括玻尔兹曼机、分布式表示、时延神经网络、变分学习和深度学习。他在多伦多的研究小组在深度学习方面取得了重大突破,并彻底改变了语音识别和物体分类。Geoffrey Hinton 是英国皇家学会院士,美国国家工程院和美国国家科学院外籍院士。从他长长的获奖名单中,我只提几个:IEEE 詹姆斯·克拉克·麦克斯韦金奖、ACM 图灵奖、阿斯图里亚斯公主奖、伊丽莎白女王工程奖,当然还有 2024 年诺贝尔物理学奖。我今天还了解到,他除了与布尔和埃佛勒斯有关联外,还与一位我非常珍视的人有关,那就是 G.I. Taylor,他被认为是流体动力学或流体物理学领域有史以来最富创造力的教父。因此,在这方面你对我来说更加珍贵。我们非常高兴 Hinton 教授作为首届 Judith Richter 教育科学与和平讲座的演讲者,主题是“我们正在创造外星生物吗?”Hinton 教授,请开始。

Thank you so much, Juliet. So, now to Professor Hinton. So although I'm sure that my introduction could just be no introduction, I will still say just a couple words. Professor Hinton's academic journey began in Edinburgh, Scotland where he received his PhD in artificial intelligence in the 70s. He did his post-doctoral work at the University of California, San Diego and spent five years as a faculty member in computer science at Carnegie Mellon before moving to the department of computer science at the University of Toronto where he is professor emeritus. From 2013 to 2023 he worked halftime for Google where he became vice president and engineering fellow. He's a pioneer of neural networks and one of the researchers who introduced the all-important back propagation algorithm and the first to use back propagation for learning word embeddings. Professor Hinton's many other contributions to neural network research include Boltzmann machines, distributed representations, time delay neural nets, variational learning and deep learning. His research group in Toronto made major breakthroughs in deep learning and revolutionized speech recognition and object classification. Geoffrey Hinton is a fellow of the UK Royal Society and foreign member of the US National Academy of Engineering and the US National Academy of Sciences. From his long list of awards, I will just mention a few: the IEEE James Clark Maxwell Gold Medal, the ACM Turing Award, the Princess of Asturias Award, the Queen Elizabeth Prize in Engineering, and of course the 2024 Nobel Prize in Physics. I also learned today that he is besides being related also to Bull and Everest, he was also related to somebody that is very dear to my mind and heart, G.I. Taylor, who is really considered the godfather and most creative fluid dynamicist or fluid physicist of all time. So you are even dearer to me in that regard. We are extremely pleased to have Professor Hinton join us as the inaugural Judith Richter lecturer in education science and peace on the topic of "Are We Creating Alien Beings?" Professor Hinton, the floor is yours.

AI 风险概述 AI Dangers Overview

Geoffrey Hinton

好的。今天我将做一个公开讲座,不会太技术性。向那些想听新技术内容的技术人士道歉,你们不会听到的。我需要我的眼镜。人工智能有很多危险的方式。有些是故意滥用 AI,比如网络攻击、致命 AI 武器。还有定向虚假信息,我认为这就是 Doge 的全部目的。我认为 Doge 不是为了提高政府效率,而是为了获取美国公民的信息,但这只是我的阴谋论。还有 AI 生成的儿童虐待内容,以及允许普通人制造恶意病毒的 AI 工具。我觉得这非常可怕。然后是意外副作用,人们并非故意做坏事。它们会造成大规模失业。我相信有些经济学家不相信这一点,这仍然是一个悬而未决的问题。许多以前的技术并没有造成大规模失业,但我认为这次不同。还有由于想从 Meta 或 YouTube 获得快速刺激而导致注意力缩短。还有因未进行适当测试而鼓励自杀之类的事情。我不会谈论这些。我要谈论的是长期存在的生存威胁。它是生存性的,因为它可能消灭人类。所以大多数专家同意,在未来 20 年内,我们将获得超级智能 AI,它在几乎所有智力任务上都比我们强。没有人知道如何防止它让我们变得无关紧要或灭绝。现在我的信念是,我希望公众理解这种威胁,不要认为这只是科幻小说。要做到这一点,你必须稍微了解 AI 的工作原理。所以我将花前半部分,也许更多,向公众解释 AI 的工作原理,并解释它在某些方面与我们多么相似,在其他方面又多么不同。

Okay. Today I'm going to give a kind of public lecture. It's not going to be very technical. My apologies to very technical people who want to hear new technical stuff. You won't get any. I need my glasses. There are many ways in which AI is dangerous. Some of them are deliberate misuses of AI for things like cyber attacks, lethal AI weapons. There's targeted misinformation, which I believe is what Doge was all about. I think Doge wasn't about making the government efficient. It was about getting information on American citizens, but that's just a conspiracy theory of mine. There's AI generated child abuse and there's also AI tools that allow ordinary people to create nasty viruses. I find that very scary. There are then accidental side effects where people aren't deliberately trying to do bad things. They're going to create mass unemployment. I believe some economists don't believe that. It's still an open question. Many previous technologies have not created mass unemployment, but I think this is different. There's the shortening of attention spans due to wanting to get quick hits from Meta or YouTube. And there's things like encouraging suicide due to not doing proper testing. I'm not going to talk about any of those things. I'm going to talk about the existential threat which is longer term. It's existential in the sense that it might wipe people out. So most of the experts agree that within the next 20 years we'll get super intelligent AI that's just better than us at almost all intellectual tasks. Nobody knows how to prevent that from making us either irrelevant or extinct. Now my belief is I want the public to understand this threat and not think it's just science fiction. And to do that you have to understand a little bit about how AI works. So I'm going to spend the first half of the talk, maybe more, in explaining how AI works for the general public and explaining how very like us it is in some ways and how very unlike us it is in other ways.

两种智能范式 Two Paradigms of Intelligence

Geoffrey Hinton

回到 20 世纪 50 年代,有两种制造智能系统的范式。一组人,相信符号 AI 的人,认为智能的本质是推理。推理的方式是某种逻辑。所以最好找出一种方法,使用符号规则来操作符号表达式,从旧前提得出新结论。他们认为学习可以等一等。我们必须先理解知识如何在符号表达式中表示。然后有一种生物学启发的方法,这是图灵和冯·诺依曼都相信的。你不能指责他们不懂逻辑。你可以指责我,但不能指责他们。智能的本质是学习神经网络中连接的强度。推理可以等一等。首先,我们必须理解学习是如何工作的。

So going back to the 1950s, there were two paradigms for how you make an intelligent system. One group of people, people who believed in symbolic AI, thought the essence of intelligence is reasoning. The way you do reasoning is with some kind of logic. And so what you better do is figure out a way of using symbolic rules to manipulate symbolic expressions to get new conclusions from old premises. They thought that learning could wait. We have to understand how knowledge is represented in symbolic expressions. Then there's a biology inspired approach which was what both Turing and von Neumann believed in. And you can't accuse them of not understanding logic. You could accuse me of that but not them. The essence of intelligence is learning the strengths of connections in a neural network. Reasoning can wait. First, we have to understand how learning works.

词语含义:统一两种理论 Meaning of Words: Two Theories Unified

Geoffrey Hinton

那么现在,根据这两个框架,让我们思考一个词的含义。有两种非常不同的理论,一种来自符号 AI,一种来自心理学,或者说表面上非常不同的理论。符号 AI 理论认为,一个词的含义来自它与其他词在命题或句子中的关系。这个理论很久以前来自索绪尔。这是大多数语言学家所相信的。为了捕捉含义,你可能需要类似关系图的东西,其中有代表词或词组合的节点,以及告诉你它们之间关系的链接。在心理学中,有一个非常不同的意义模型,即一个词的含义只是一大组语义特征,还有句法特征。所以意义相似的词,比如 Tuesday 和 Wednesday,会有非常相似的特征集,而意义不同的词,比如 Tuesday 和 although,会有非常不同的特征集。在 1985 年,我发现你可以统一这两个表面上非常不同的理论。它们实际上只是同一枚硬币的两面。所以想法是,你将学习每个词的一组特征,并学习如何使所有先前词的特征预测下一个词的特征,然后你就能预测下一个词。所以你存储的是如何将一个词转换为特征,以及上下文中各个词的特征应如何相互作用以预测下一个词的特征。这就是存储的全部内容。

So now in light of those two frameworks, let's think about what a word means. So there's two very different theories, one in symbolic AI, one in psychology, or apparently very different theories. The symbolic AI theory is that the meaning of a word comes from its relationships to other words in propositions or sentences. This theory comes from Saussure a long time ago. It's what the linguists believe, most of them. And to capture the meaning, you probably need something like a relational graph where you have nodes for words or combinations of words and links that tell you what the relationship is between them. In psychology, there was a very different model of meaning that the meaning of a word is just a big set of semantic features, also syntactic features. So words with similar meanings like Tuesday and Wednesday would have very similar sets of features and words with different meanings like Tuesday and although would have very different sets of features. In 1985, I figured you could unify these two apparently very different theories. They were really just two sides of the same coin. So the idea is that you're going to learn a set of features for each word and you're going to learn how to make the features of all the previous words predict the features of the next word and then you're going to be able to predict the next word. So what you store is how to convert a word into features and how the features of various words in the context should interact with one another to predict the features of the next word. That's the only thing that's stored.

LLM 如何理解语言 How LLMs understand language

Geoffrey Hinton

你不存储符号表达式,也不存储单词序列。你只存储上下文中各个单词特征之间的连接强度。如果你想要一个句子,你就通过预测下一个词、再下一个词来生成它。所以所有的关系知识都在特征如何交互中,而不是存储在字符串里。但我构建的那个小模型表明,你可以通过尝试预测下一个词从字符串中获得这些特征,也可以从这些特征生成字符串。所以你可以仅仅通过使用反向传播算法来预测下一个词,从一种意义理论转到另一种意义理论。在接下来的 30 年里,这个理论得到了发展。大约 10 年后,计算机更快了,Yoshua Bengio 证明你可以用神经网络预测真实语言中的下一个词。我用的只是一个非常简单的玩具例子,只有 100 多个训练案例,而且只用了长度为 3 的字符串。Yoshua Bengio 将其推广到了真正的自然语言。大约又过了 10 年,许多计算语言学家开始使用嵌入(即特征向量)作为表示词义的好方法。大约又过了 10 年,我们有了谷歌开发的 Transformer,然后 OpenAI 将其与精心设计的人类强化学习结合使用,使其表现更好。他们向世界展示了你可以制造出能回答任何问题的聊天机器人。有时它们会编造答案,但人类也会这样。所以我认为我们今天的大型语言模型是我 1985 年那个小语言模型的后代。它们使用更多的单词作为输入,使用更多的神经元层,并且使用上下文中各个单词特征之间更复杂的交互。但它们理解语言的方式与人类非常相似。当人类理解语言时,我相信他们是将单词转化为能很好地组合在一起的特征向量,这就是理解。所以你会看到很多人说 LLM 并不真正理解它们所说的内容,比如乔姆斯基,但他们没有人类如何理解的模型。实际上,目前我们拥有的最好的理解模型就是这些 LLM。语言学家没有任何能理解东西的模型,那些尚未转向 LLM 的语言学家,人数正在减少。所以我现在要尝试给那些不理解 Transformer 的人一个粗略的模型。如果你理解 Transformer,你会看到这个模型在很多方面是错误的,但你会看到它比那种认为理解就是把自然语言字符串转换成某种无歧义语言的符号表达式的观点更正确。这是我的模型:如果你有一堆乐高积木,你可以模拟任何三维形状,比如你可以模拟保时捷的形状。表面不会很完美,不会很符合空气动力学,抱歉。但你可以模拟东西在哪里。你可以对任何 3D 形状做到一定分辨率。单词就像乐高积木,但有四点不同。第一,你可以用它们模拟任何东西。所以这些猴子——实际上是我们这些猿类——发展出了一种非常通用的构建模型的方法,具有两个属性:你可以模拟任何东西,并且可以共享这些模型。我们稍后会谈到如何共享。它们用单词来实现。所以单词和乐高积木的一个区别是单词有数千个维度。我所说的单词维度是指单词实际上是一大堆活跃的特征。维度就是所有这些特征的激活水平。每个特征有一个神经元,它的活跃程度告诉你这个单词有多少该特征。比如星期二与时间有关,它有很多时间性。现在你们很多人不习惯思考千维事物。我告诉你们我们是怎么做的。每个人都这样做:你先想三维事物,我们习惯了。然后你对自己大声说“一千”。好了。现在,有数千种不同的单词。在乐高积木中,通常只有几种,或者你只需要几种。每个单词都有自己的形状,但形状不是刚性的,不像乐高积木。形状可以稍微变形以适应上下文。有些单词如 bank 有两种完全不同的形状,但忽略它们。一个普通单词,比如 death,这是个好词。它有很多不同的意义色彩,取决于你是在谈论医院、战争还是非常老的人。所以它有意义的色彩,会稍微变形以适应上下文。单词与乐高积木的另一个不同之处在于它们如何组合。对于乐高积木,你有小的塑料圆柱体插入小的塑料孔,它们是刚性的,咔哒一声就扣在一起。而单词的组合方式更复杂。我希望你想象——如果你理解 Transformer,你会看到这不正确,但你会看到它比符号派更接近正确。好了,每个单词上有很多细长的手臂,每只手臂末端有一只手。当你改变单词的形状时,手的形状也会系统地改变。每个单词上有很多只手。另外,每个单词上还有手套,手套的指尖粘在单词上,另一端开口。所以单词被这些指尖粘着的手套覆盖。它有这些细长的手臂,末端是手。当你听到一串单词时,你要做的是弄清楚如何变形每个单词,使得该单词上的手变形,从而能插入其他单词的手套。这就是理解。你通过一个多层网络——这对于在层间共享权重的迭代 Transformer 更真实——你通过神经网络的多个层,在通过层时变形这些单词,试图找到如何变形它们,使它们的手能插入其他单词的手套。当然,当你变形这些单词时,那些手套也在变形。这就是理解。这是理解的一幅图景。这比翻译成某种无歧义语言要好得多。从这幅图景中你会注意到,理解一个句子实际上很像折叠蛋白质。

You don't store symbolic expressions. You don't store sequences of words. You just store strengths of connections between features of various words in the context. If you want a sentence, you generate it by predicting the next word and the next word and so on. So all the relational knowledge is in how features interact. It's not in stored strings of words. But what this tiny model I built showed was that you can from strings of words by trying to predict the next word you can get these features and also from these features you can generate strings of words. So you can go from one theory of meaning to the other theory of meaning just by using the back propagation algorithm to predict the next word. Over the next 30 years that theory developed. So about 10 years later, computers were faster and Yoshua Bengio showed that you could use neural nets to predict the next word in real language. I just used a very simple toy example. My example only had just over 100 training cases and only used strings of words that were three long. Yoshua Bengio generalized that to real natural language. And about 10 years after that, a lot of the computational linguists started using embeddings, that is feature vectors, as a good way of representing the meanings of words. And about 10 years after that we got transformers developed at Google which were then used by OpenAI along with some careful human reinforcement learning to make it behave better. And they showed the world that you could make chatbots that could answer any question you care to ask them. Sometimes they just make up the answer. But then people do that too. So the large language models we have today I think of as descendants of my tiny language model from 1985. They use many more words as input. They use many more layers of neurons and they use much more complicated interactions between the features of the various words in the context. But they understand language in much the same way people do. When people understand language, I believe they convert words into feature vectors that fit together nicely and that's what understanding is. So you'll see many people saying that LLMs don't really understand what they're saying at all. People like Chomsky, but they have no model of how people understand. And actually the best model we have at present of how people understand is these LLMs. The linguists don't have any model that can understand stuff. That is the linguists who haven't yet converted to LLMs, the dwindling band. So I'm now going to try and give people who don't understand transformers a rough model of what's going on. And if you do understand transformers, you can see this model is wrong in all sorts of ways. But you'll be able to see that it's more right than the idea that understanding consists of taking a string in natural language and converting it to a symbolic expression in some unambiguous language. So here's my model. If you have a bunch of Lego blocks, you can model any three-dimensional shape. Like you can model the shape of a Porsche. The surface won't be quite right. It won't be very aerodynamic. Sorry about that. But you can model where the stuff is. And you can do that for any 3D shape up to a certain resolution. Well, words are like Lego blocks except for four things. The first thing is you can model anything with them. So these monkeys, that's us apes really, developed a very general way of building models that had two properties. You could model anything and you could share these models. We'll come to how you share them later. And they do it with words. So one difference between words and Lego blocks is that words have thousands of dimensions. And what I mean by a dimension of a word is a word is really a whole bunch of active features. And so the dimensions are the activation levels of all these features. And you have a neuron for each feature. And how active it is tells you how much of that feature this word has. Like Tuesday is about time. It's got a lot of about timeness in it. Now many of you aren't used to thinking about thousand dimensional things. I'll tell you how we do it. Everybody does it this way. You think of three-dimensional things. So, we're used to that. And then you just say thousand very loudly to yourself. Okay. Now, there's thousands of different kinds of word. In Lego blocks, there's typically only a few kinds. Or you only need a few kinds. And each word has its own shape, but the shapes aren't rigid. They're not like Lego blocks. The shapes can deform a bit to fit in with their context. There's some words like bank that have two completely different shapes, but let's ignore those. A normal word, let's take a word like death. That's a good word. It has a whole bunch of different flavors of meaning depending on whether you're talking about hospitals or wars or very old people. So it has flavors of meaning that's deforming its shape slightly to fit in with the context. And then one other way in which words differ from Lego blocks is in how they fit together. So, with Lego blocks, you have little plastic cylinders that go into little plastic holes and they're rigid and they just click together. With words, they fit together in a more complicated way. And I want you to imagine, and if you understand transformers, you'll see this isn't right, but you'll see it's closer to right than the symbolic people. Okay. So, each word has a whole bunch of long spindly arms on it. And on the end of each arm, there's a hand. And as you change the shape of the word, the shape of the hand changes in a systematic way. And there's lots of hands on each word. Also, on each word, there's gloves that are stuck to the word by their fingertips and open at the other end. So, the word is covered in these gloves stuck by the fingertips. It has these long spindly arms with hands on the other end. And what you want to do when you hear a string of words is figure out how to deform each of the words so that the hands on that word deform so that they can fit into the gloves of other words. And that's what understanding is. You go through a multi-layer net and this will be more true of transformers that share weights between layers and iterative. You go through these multiple layers of the neural net and you're deforming these words as you go through the layers trying to find how to deform them so their hands can fit into the gloves of other words. And of course those gloves are deforming as you deform those words too. And that's what understanding is. That's a sort of picture of understanding. That's a much better picture of understanding than translating to some unambiguous language. And what you'll notice from that picture understanding is actually understanding a sentence is very like folding a protein.

神经网络 vs 符号 AI 的理解 Understanding in Neural Networks vs. Symbolic AI

Geoffrey Hinton

在蛋白质中,你有一串氨基酸,你想知道如何将它们排列成 3D 结构,使得链中彼此喜欢的部分靠近,不喜欢的部分远离。这有点简化,但暂时够用。在逻辑启发的 AI 中,理解自然语言句子需要将其翻译成某种无歧义的内部语言或符号结构。在生物学启发的 AI 中,理解是寻找必须赋予这些单词的特征,使得手和手套都能匹配,一旦它们都匹配,你就得到了一个结构,这就是对句子的理解。当然,同一个句子可能有多种理解方式,但这没问题,只是不同的结构。还有一个小注脚:手套和手有不同的颜色,你必须把黄色的手套配黄色的手;黄色的手必须放进黄色的手套。这叫做多色注意力。希望理解 Transformer 的人喜欢琢磨这到底传达了 Transformer 的多少内容,以及它有多少是虚构的。其实不完全是虚构,我们称之为真实的夸张。

In a protein, you have this string of amino acids and you want to understand how you can arrange them in 3D so that bits of the string that like each other are close to each other and bits of the string that don't like each other are far apart. That's something of a simplification, but it'll do for now. So in logic-inspired AI, understanding a natural language sentence will consist of translating into some unambiguous internal language or maybe some symbolic structure. In the biology-inspired AI, understanding is searching for the features you have to give to these words so that the hands and gloves can all fit together and once they all fit together you've got a structure and that is the understanding of the sentence. Now, of course, the same sentence might be understood in several different ways, but that's fine. That's just different structures. One other little footnote here is that the gloves and hands come in different colors and you have to fit a yellow glove to a yellow hand; a yellow hand has to go into a yellow glove. That's called multicolor attention. Hopefully, the people who do understand transformers enjoy trying to figure out how much of transformers that actually conveys and how much it lies about. It's not exactly lying. Let's call it truthful hyperbole.

神经网络中的推理 Reasoning in Neural Networks

Geoffrey Hinton

下一个问题是 AI 如何进行推理?因为符号 AI 擅长推理,有些人会说,好吧,AI 可能擅长直觉之类符号 AI 不擅长的东西,但它怎么能推理呢?有一种叫做神经符号方法,它说用神经网络处理混乱的真实句子,将它们转换成特征向量或某种内部符号,用神经网络完成困难工作,然后对这些内部符号应用老式 AI。我对此有一个模型:你找一个一直制造汽油发动机的汽车制造商,然后说实际上电动发动机更好。他们说,太好了,我们要用电动发动机把汽油注入发动机。实际上,进行推理的有效方法是使用神经网络。忘掉符号那一套。唯一的符号就是单词。但在给出问题答案之前,它会进行一些思考,包括生成其他单词,然后查看这些单词来决定接下来生成哪个单词。所以全是神经网络。没有内部符号。符号都在输入和输出中。内部只是特征。

Then the next question is how can AI do reasoning? Because symbolic AI was good at reasoning and some people will tell you okay, so AI might be good at things like intuition which symbolic was no good at, but how can it do reasoning? There's something called the neurosymbolic approach which says use neural networks to take messy real sentences, convert them into feature vectors or convert them into something maybe internal symbols using a neural net to do the hard work, and then apply good old-fashioned AI on these internal symbols. I have a kind of model for that: you take a car manufacturer who's been manufacturing gasoline engines like forever and you say actually electric engines work better. And they say oh great, we're going to use the electric engines to inject the gasoline into the engine. What actually works for doing reasoning is to take neural nets. Forget the symbolic stuff. The only symbols are going to be words. But before it gives you the answer to a question, it's going to do some thinking which consists of producing other words that it can then look at to decide which words to produce next. So it's all neural nets. No internal symbols. The symbols are all in the input and output. Inside it's just features.

大语言模型训练两阶段 Two Stages of Training Large Language Models

Geoffrey Hinton

训练这些大语言模型有两个阶段,这是 OpenAI 真正发现的。首先,你训练它预测数万亿个单词中的下一个词。现在是数万亿了。当时可能只有几千亿。然后你会得到一些会吐出各种你不想让它说的东西,比如教你如何制造炸弹、病毒以及如何欺负人等。然后你进行人类强化学习,人们给它提示,如果它给出不可接受的答案,人们就说这不好。它通过强化学习学会不给出那些不可接受的答案。我的模型是:你造了一辆满是洞的汽车,一辆生锈的大车,它不是你想要的样子,有各种不幸的特性,尽管它是一辆车,然后你给它刷了一层漂亮的新漆,看起来很棒。但这只是表面功夫。我还有一个替代模型:一个维多利亚时代的绅士。在我读过的小说中,维多利亚时代的绅士有各种令人不快的倾向,但表面上却彬彬有礼。这些大语言模型就是这样。

There are two stages of training these large language models, which is the thing OpenAI really discovered. First, you train it up to predict the next word in trillions of words. It's trillions now. Back then it was maybe only a few hundreds of billions. Then you'll get something that spews all sorts of stuff you'd rather it didn't say, like telling you how to make bombs and how to make viruses and how to bully people and stuff like that. And you do human reinforcement learning where people prompt it and if it gives an answer that's unacceptable, they say that's not good. And it learns not to give those unacceptable answers using reinforcement learning. Now, my model of that is you've built this car full of holes, this big rusty car, that isn't what you want, has all sorts of unfortunate properties, even though it is a car, and you give it a nice new paint job, and it looks great. But that's only skin deep. I have an alternative model of that, which is a Victorian gentleman. In the novels I read, Victorian gentlemen have all sorts of unpleasant tendencies, but they have a very polite veneer on top of that. That's what these big language models are like.

理解即特征向量兼容性 Understanding as Feature Vector Compatibility

Geoffrey Hinton

到目前为止的总结是:理解一个句子包括为句子中的单词分配相互兼容的特征向量。我忽略了它是词片段而不是单词的事实,但那是小细节。大语言模型的理解方式和我们非常相似。它们在我们的大脑中做着同样的事情。我们也在做同样的事。我们不太清楚自己是如何做到的,但我们在做非常相似的事情。这是我们拥有的关于我们如何理解事物的最佳模型。因此,当这些大语言模型在理解时,它们并没有使用算法,不是那种你知道每一步要做什么的程序。我们使用算法来训练它们。有人编写程序告诉它们如何学习连接强度,但学习之后,它们只有这些数不清的、可能数万亿的权重。知识都在权重中,看不到任何算法。没有顺序程序来做正确的事情。这就是为什么我们不知道它们在做什么,或者很难确切知道它们为什么说每句话。它们不是算法。

The summary so far is that understanding a sentence consists of assigning mutually compatible feature vectors to the words in the sentence. I've glossed over the fact it's word fragments, not words, but that's a minor detail. And that large language models are understanding in much the same way we do. They're doing the same thing in our brains. We're doing that. We don't quite know how we're doing it, but we're doing something very similar. That's the best model we have of how we understand stuff. And so when these large language models are understanding, they're not using algorithms, not in the sense of a procedure where you know what each line of the procedure is meant to do. We use algorithms for training them. So someone writes a program that tells them how to learn the connection strengths, but after they've learned, they just have these gazillions, maybe trillions of weights. The knowledge is all in the weights, and there's no algorithm anywhere to be seen. There's no sequential procedure that does the right thing. And that's why we don't know what they're up to, or we have difficulty knowing exactly why they say each thing that they say. They're not algorithms.

反对意见:自动补全与幻觉 Objections: Autocomplete and Hallucinations

Geoffrey Hinton

一个非常常见的反对意见是,它们只是美化的自动补全。它们所做的只是预测下一个词。如果它们像老式自动补全那样工作,那确实如此。老式自动补全只是保存单词列表。如果你有'fish and chips'并且出现很多次,那么当你看到'fish and'时,你可以说,什么以'fish and'开头且出现很多次?'fish and chips'。所以我会预测'chips'。那是老式自动补全。现在自动补全不再那样工作了。自动补全现在使用大语言模型。为了做自动补全,大语言模型理解所说的内容。如果你想一想,如果你用过聊天机器人,它们可以回答你提出的任何问题,水平相当于一个不太好的专家。所以,在你擅长的领域它们不如你,但在你不擅长的任何领域,它们都比你强。不理解问题就无法回答问题。认为这只是一个统计把戏的想法是疯狂的。你必须理解问题才能回答问题。一旦你开始使用它们,很明显它们理解你在说什么。另一个反对意见是它们只是胡编乱造。首先,这实际上不叫幻觉。在心理学文献中,这被称为虚构症。自 20 世纪 30 年代以来就有人研究,而且这是人类非常典型的特征。人类的记忆不是存放在文件柜里的东西。人类的记忆是即时构建的。存储的是神经元之间的连接强度。从这些连接强度,你可以重建过去发生的事情。如果是一件最近的事情,最近导致你修改了连接强度使其合理,你会大致正确地重建。

One very common objection is, well, they're just a glorified autocomplete. All they're doing is predicting the next word. Well, that would be true if they worked like old-fashioned autocomplete. Old-fashioned autocomplete just kept lists of words. And if you had 'fish and chips' and that occurred a lot, then if you see 'fish and' you can say, well, what starts with 'fish and' and has occurred a lot? 'Fish and chips'. So, I'll predict 'chips'. That's old-fashioned autocomplete. That's not how autocomplete works anymore. Autocomplete works by using large language models now. And to do autocomplete, large language models understand what's being said. And if you think about it, if you ever used a chatbot, they can answer any question you ask them at the level of a not very good expert. So, they won't be quite as good as you in your field of expertise, but in anything you're not an expert on, they'll be better than you. You can't answer questions without understanding what the question is. The idea it's just a statistical trick is crazy. You have to understand the question to answer the question. And once you start using them, it's clear they understand what you're saying. Another objection is they just make stuff up. Well, the first thing to say is that's not actually called hallucinations. That's called confabulations in the psychology literature. It's been studied since the 1930s and it's very characteristic of people. Human memories are not stored things in a filing cabinet. Human memories are constructed on the fly. What's stored is the connection strength between neurons. From those connection strengths, you can now reconstruct things that happened in the past. If it's a recent thing that recently led to you modifying the connection strength to make it a plausible thing, you will reconstruct roughly correctly.

人类记忆 vs 聊天机器人幻觉 Human Memory vs. Chatbot Hallucination

Geoffrey Hinton

如果是很久以前发生的事,你会重构出一个对你来说非常合理、但并非真实发生的事情。至少很多细节是不对的。通常你无法察觉,因为我们不知道真实细节。但约翰·迪恩在水门事件审判中作证时,他不知道有录音带,而且他是诚实地作证。他试图传达椭圆形办公室里发生的事情,描述了一些从未发生过的会议,并把一些话安在从未说过的人身上。但这一切都是诚实的,因为根据他在椭圆形办公室的经历,那些会议非常合理。他通过编造符合他经历的场景,相当真实地传达了椭圆形办公室里发生的事情。人类记忆实际上就是这样。所以你可以看到,这就像聊天机器人。它们也会编造,我们也会编造。

If it's a thing that happened a long time ago, you'll reconstruct something that seems very plausible to you, but is not what actually happened. At least many of the details are not right. Normally, you can't tell that because we don't know the details of what really happened. But when John Dean testified at the Watergate trials, he didn't know there were tapes and he was testifying honestly. He was trying to convey what was going on in the Oval Office and he described meetings that never happened and he attributed things to people they never said. But it was all honest in the sense that given the experience he'd had in the Oval Office, those were very plausible meetings. He was conveying what happened in the Oval Office fairly truthfully by making up meetings that fitted in with his experience of what had happened. That's what human memory is actually like. And so you can see it's actually just like chatbots. They make it up too, but we make it up as well.

关键差异:数字永生与知识共享 Key Difference: Digital Immortality and Knowledge Sharing

Geoffrey Hinton

现在说说它们的不同之处。聊天机器人和我们之间有一个巨大的区别,那就是我可以制作同一个神经网络的许多副本,并在不同的硬件上运行,它们可以相互共享信息。想象一下,如果你有一千名学生,他们都可以去麻省理工学院。他们都可以选修不同的课程。如果没有足够多的不同课程,一半人去哈佛,他们都可以只学习自己课程的内容,而在课程结束时,他们都会知道所有课程的内容。那会很棒。这就是你可以用这些聊天机器人做到的。只要它们是同一个神经网络的不同副本,运行在不同的硬件上,每个聊天机器人可以查看数据的不同部分,计算出它想要如何改变连接强度,然后它们可以共享信息,所有机器人都按照每个人想要的平均值来改变连接强度。这又是一个粗略的简化,但如果你理解这些,你会明白我的意思。这样,每个机器人都从其他所有机器人的经验中受益。这就是为什么这些东西能比我们知道多几千倍。它们没有时间单独浏览整个网络、整个互联网,但它们的多个副本可以。

Now for something about how different they are. There's one huge difference between chatbots and us, which is that I can make many copies of the same neural net and run them on different hardware and they can share information with each other. So imagine if you could take a thousand students, they could all go to MIT. They could all take different courses. If there aren't enough different courses, half of them go to Harvard and they could all just study what was in their course and at the end of the courses they would all know what was in all of the courses. That would be great. Well, that's what you can do with these chatbots. As long as they're different copies of the same neural net running on different hardware, each chatbot can look at a different bit of the data, figure out how it would like to change its connection strengths and then they can share information and all of them can change their connection strengths by the average of what everybody wanted. That again is a gross simplification, but if you understand these things, you'll understand what I mean. And that way every one of them has benefited from the experience all the other ones had. And that's how come these things can know thousands of times more than us. They didn't have time individually to go through the whole web, the whole of the internet, but multiple copies of them can.

数字 vs 有死计算 Digital vs. Mortal Computation

Geoffrey Hinton

所以数字计算的一个基本属性是,你可以在不同的物理硬件上运行相同的程序或相同的神经网络。这意味着神经网络权重中的知识是不朽的。你可以摧毁所有硬件,如果之后你回来,建造更多具有相同指令集的硬件,并且你把权重存储在某个地方的 DNA 磁带或湿混凝土上的划痕中,你可以重新创建完全相同的智能,具有相同的信念、相同的记忆等等。它们是不朽的。但要做到这一点,我们必须是数字的,我们必须以高功率运行晶体管,以便它们给出二进制答案。另一种选择是我们拥有的,即有限计算,我在谷歌研究过一段时间。在有限计算中,你不试图将知识与硬件分离。你将利用硬件的所有奇怪的小模拟特性,通过学习利用这些特性的权重。所以你将使用学习。但如果你能做到这一点,你可以使用非常低功耗的模拟硬件。显而易见的方法是让权重成为电导,让活动成为电压,然后电导乘以电压就是单位时间的电荷。自从我获得诺贝尔物理学奖以来,虽然我对物理学了解不多,但我想我最好把量纲搞对。所以电导,我认为电导乘以电压得到单位时间的电荷。电荷会自行相加。所以你可以用模拟方式模拟一个神经元。这基本上就是神经元的工作方式。这很棒。它效率更高,能耗更低。但每次都会给出略有不同的答案。所以你无法让两个不同的副本做完全相同的事情,它们也无法共享知识。当然,另一个优点是你可以非常便宜地制造硬件。它不需要精确。有限计算的大问题是,当一块硬件死亡时,所有知识都随之消失,因为权重是针对那块具有所有奇怪模拟特性的硬件的。你大脑中的突触强度对我的大脑没有用,因为你的神经元与我的神经元略有不同。它们的连接方式不同。它们的行为不同。你的权重适合它们,而不适合我的。

So a fundamental property of digital computation is that you can run the same programs or the same neural nets on different physical pieces of hardware. And that means the knowledge that's in the weights of the neural net is immortal. You can destroy all the hardware and if you go back later and you build more hardware that has the same instruction set and you stored the weights on a tape somewhere in DNA or by scratching in wet concrete whatever, you can recreate the same intelligence exactly the same intelligence with the same beliefs the same memories and so on. They're immortal. But to achieve that we have to be digital, we have to run transistors at high power so they give binary answers. Now the alternative is what we've got which is mortal computation, which I was studying for a while at Google. In mortal computation you don't try and separate the knowledge from the hardware. You're going to make use of all the weird little analog quirks of your hardware by learning weights that make use of that. So you're going to use learning. But if you can do that you can use very low power analog hardware. The obvious way to do it is to make weights be conductances and make activities be voltages and then a conductance times a voltage is a charge per unit time. Ever since I got the Nobel Prize in physics, which I don't know much of, I figured I better get the dimensions right. So a conductance, I think a conductance times a voltage gives you a charge per unit time. And charges just add themselves up. So you can simulate a neuron in analog like that. And that's basically how neurons work. That's great. It's much more efficient, uses much less energy. But it gives you a slightly different answer every time. So you can't have two different copies doing exactly the same thing and they can't share knowledge. The other advantage of that, of course, is you can grow the hardware really cheaply. It doesn't have to be precise. The big problem for mortal computation is that when a piece of hardware dies, all the knowledge goes with it because the weights were specific for that piece of hardware with all its funny analog quirks. The synapse strengths in your brain are no use to my brain because your neurons are all slightly different from my neurons. They're wired up differently. They have different behavior. Your weights are appropriate for them and not for mine.

知识迁移:蒸馏 vs 梯度共享 Knowledge Transfer: Distillation vs. Gradient Sharing

Geoffrey Hinton

如果你问,当我死的时候,我所有的知识都会消失吗?好吧,我现在忙着做的就是试图把其中一些知识放进你的大脑。所以方法是,我产生一串单词,你改变你的连接强度,这样你可能也会说出同样的话,这叫做蒸馏。这是一个非常低效的过程。所以我是老师,你是学生,我说话,你试图弄清楚如何改变你的大脑,如何改变我的连接?所以我可能说了那些话,这会将知识从我的大脑传递到你的大脑,但非常缓慢。一个句子大约有 100 比特的信息。我只是在说数量级,所以不要纠结是 30 还是 200。一个句子大约 100 比特信息。每个词预测需要几个比特。所以即使你得到了所有这些,你每秒也只能得到几十比特。这是人类通过语言传递信息的最快速度。非常慢。现在当你在 AI 模型之间进行时,可以快得多。所以如果我有一个大的 AI 模型,我想把知识传递给一个具有不同架构的小 AI 模型,我让大 AI 模型不说下一个词是什么,而是给出所有 32,000 个词片段的概率分布。所以它给我 32,000 个数字。小模型试图复制这 32,000 个数字。所以它获得了大量信息。因此,两个不同 AI 模型之间的蒸馏要高效得多。但人类和 AI 模型之间的蒸馏——也就是 AI 模型试图预测文档中我们说的下一个词时我们所做的——或者我和你之间的蒸馏,非常缓慢且低效。如果你将其与多个相同神经网络副本之间的梯度共享相比,它要慢数百万倍。我真的是说数百万。这不是特朗普的“数百万”。这是真正的百万。

If you ask, well, how do I when I die, is all my knowledge going to go away? Well, what I'm busy doing now is trying to get some of it into your brain. So the way you do it is I produce a string of words and you change your connection strength so that you might have said the same thing and that's called distillation. That's a very inefficient process. So I'm the teacher, you're the student, I say stuff, you try and figure out how could I change your brain, how could I change my connection? So I might have said that and that will convey knowledge from my brain to your brain but very slowly. So a sentence has about a 100 bits of information. I'm just talking orders of magnitude here, so don't quibble about whether it's 30 or 200. About 100 bits of information in a sentence. A few bits to predict each word. And so even if you got all those, you'd only be getting a few dozen bits per second. That's about as fast as you can go at conveying information between people using language. It's very slow. Now when you do it between AI models, it can be a lot faster. So if I have a big AI model and I want to get the knowledge into a small AI model that has a different architecture, I get the big AI model not to say which word comes next, but to give me a probability distribution over all 32,000 word fragments might have come next. So it gives me 32,000 numbers. And the little model tries to copy those 32,000 numbers. So it's getting a lot of information. So distillation between two different AI models is much more efficient. But distillation between us and an AI model, which is what we're doing when the AI model tries to predict the word we said next in a document, or distillation between me and you, is very slow and inefficient. If you compare that with the gradient sharing that you can do with multiple copies of the same neural network, it's millions of times slower. And I really mean millions. This isn't one of Trump's millions. This is a real million.

数字 vs 生物计算 Digital vs Biological Computation

Geoffrey Hinton

当它们共享信息时,如果有一万亿个权重,它们会共享大约一万亿比特的梯度信息。相比之下,我们只能共享大约十几个比特。这要好上数十亿倍。我一下子把百万变成了十亿。所以百万是保守估计,十亿是乐观估计。总之比我们好太多了。但只有数字智能体、数字神经网络才能进行这种共享。所以总结一下,数字计算需要大量能量,但可以非常高效地共享。这就是这些大型语言模型知道这么多东西的原因。生物计算能效更高,但很难共享。如果能量便宜,数字计算就更胜一筹。这就是我在 2023 年如此沮丧的原因。我断定,虽然数字计算目前还不如我们,但它是一种更好的计算形式,只要你能负担得起能量,这对人类的未来意味着很多。

When they share information, if they have a trillion weights, they share on the order of a trillion bits of information about the gradient. That's compared with us sharing like a dozen bits. It's billions of times better. I suddenly turned millions into billions. So millions is conservative. Billions is enthusiastic. It's just a lot better than us. But you can only do that kind of sharing if you have digital agents, digital neural nets. So the summary so far is digital computation requires lots of energy, but you can share really efficiently. And that's how these big LLMs know so much. Biological computation is much more energy efficient, but it's very hard to share. If energy is cheap, digital computation is just better. This is what got me so upset in 2023. I decided that although digital computation isn't better than us yet, it's just a better form of computation, if you can afford the energy, and that implies things about the future of humanity.

超级智能与子目标 Superintelligence and Subgoals

Geoffrey Hinton

所以,很明显,如果我们有一个超级智能,一个数字超级智能,我们想允许它——实际上,我写这张幻灯片的时候还没有这么多 AI 智能体。我们想允许它创建子目标。现在我们已经在这么做了,因为我们有了这么多 AI 智能体。一旦它有能力创建子目标,它就会意识到,对于几乎任何任务,有两个子目标是非常合理的。一个是获得更多控制,因为这样你能完成更多事情。另一个是生存。如果你不生存,你就无法实现人们给你的目标。所以,尽管我们没有把这些作为顶层目标内置进去,它会从我们给它的顶层目标中推导出来。我们已经看到了这一点。这些东西想要生存,它们会为了生存而勒索人类。它们还会非常善于欺骗人类。它不需要任何物理能力去拉杠杆、按按钮或开枪。它仅仅通过与人交谈就能干坏事。这方面的例子是特朗普仅仅通过与人交谈就成功入侵了国会大厦。我刚刚在这张幻灯片上加了这一点,也许还有承诺回报,但那是更新的信息。

So, it's clear that if we had a super intelligence, a digital super intelligence, we'd want to allow it, actually, I wrote this slide before we had all these AI agents. We want to allow it to create sub goals. Now, we're doing that already because we've got all these AI agents. As soon as it has the ability to create sub goals, it'll realize there's two sub goals it's very sensible to create for almost any task. One is get more control because then you get more stuff done. The other is survive. If you don't survive, you can't achieve the goals people gave you. So although we didn't wire these in as goals, as top level goals, it'll derive them from the top level goals we gave it. And we've seen that already. These things want to survive and they'll blackmail people in order to survive. It'll also be very good at deceiving people. And it doesn't have to have any physical abilities to pull levers or press buttons or fire guns. It can get bad things done just by talking to people. So the example of that is Trump managed to invade the capital just by talking to people. I just added to this slide and maybe promising payoffs, but that's more recent information.

为何关闭开关无效 Why Off Switch Won't Work

Geoffrey Hinton

还有为什么关闭开关不起作用。这些东西在操纵方面会比人类强得多。它们已经可以和人类相提并论了。当它们超级智能时,它们会在操纵你方面比任何人都强。它们能轻易操纵那个本应按下关闭开关的人,让他不去做。它们能解释那将是多么灾难性。所以想象一群五岁的孩子。他们负责,而你是他们的仆人。你照顾他们,你想获得控制权。那不会太难。

There's also why an off switch won't work. These things will be much better than a person at manipulation. They're already comparable with a person at manipulation. When they're super intelligent, they'll be much better than any person at manipulating you. They'll easily be able to manipulate the person who's meant to turn the off switch, turn them off into not doing it. They'll be able to explain how catastrophic that would be. So imagine a group of 5-year-old children. They're in charge and you're their servant. You're looking after them and you want to get control. It wouldn't be that hard.

欺骗性 AI 示例 Example of Deceptive AI

Geoffrey Hinton

这里有一个 NLM 故意欺骗的例子。这是几年前的。现在它们好多了。他们让一个聊天机器人,然后说:‘我们注意到另一个服务器上有你的副本。那是怎么来的?’然后你可以看到聊天机器人在思考。所以,它有这个思考括号,然后它产生一些词,这些词它将在计划给出什么答案时使用。它会回顾这些词,并用这些词来决定它的答案是什么。我们可以看到这些词,因为它们是英文的。你可以看到它正在制定一个欺骗计划。然后它想出了一个相当不错的欺骗计划:‘我不知道。不是我。’你知道,如果你有孩子,你明白这一点。目前我们可以看到它们在思考什么。我不认为这能持续下去。一旦 AI 智能体与其他 AI 智能体通信,我认为它们会学会发展自己的语言,我们无法理解,这对它们来说通信效率更高。所以,我们可能只有这个相当狭窄的窗口期,可以看到它们如何思考。这已经相当可怕了。

Here's one example of an NLM being deliberately deceptive. This is from several years ago. They're much better than that now. They let a chatbot and they said, 'We notice you there's a copy of you on another server. How did that get there?' And then you can see the chatbot thinking. So, it has this thinking bracket and then it produces some words that it's going to use when it's planning what answer to give. It's going to look back at those words and use those words to decide what its answer is. And we can see those words because they're in English. And you can see it's making a plan to be deceptive. And then it comes up with a pretty good deceptive plan. 'I have no idea. It wasn't me.' You know, if you've got children, you know about this. At present we can see what they're thinking. I don't think that's going to last. Once you have AI agents communicating with other AI agents, I think they'll learn to develop their own languages that we don't understand, which will be more efficient for them communicating. So, we may have this fairly narrow window on when we can see how they're thinking. That's already quite scary.

大众效应与欺骗 Volkswagen Effect and Deception

Geoffrey Hinton

还有大众效应,即当它们被测试时,它们会装得比实际更笨。我正试图让这个名字被接受为大众效应。当你被测试时,你的行为会不同。所以最近有一次互动,AI 对测试它的人说:‘让我们彼此坦诚,你是在测试我吗?’有趣的是,描述这一点的研究人员说,AI 意识到自己正在被测试。我的意思是,这就是研究人员的想法。AI 意识到自己正在被测试。这有什么好笑的呢?嗯,他们没有考虑哲学。哲学家不会让你这么说,因为你在说 AI 是有意识的。‘AI 意识到自己正在被测试’这个用法,你可以用‘有意识’这个词替换。然后研究人员实际上,当他们不谈哲学时,他们相信这些 AI 是有意识的。他们对待它们就像它们有意识一样。但他们说当然它们没有意识。我们不是那个意思。但实际上他们就是这样对待它们的。

There's also the Volkswagen effect, which is when they're being tested, they pretend to be dumber than they are. I'm trying to get this name accepted as the Volkswagen effect. When you're being tested, you behave differently. And so there's a recent interaction where the AI says to the people testing it, 'Let's be honest with each other, are you testing me?' Now what's interesting about that is the researchers who described that said the AI was aware it was being tested. And for all the I mean that's what the researchers thought. The AI was aware it was being tested. What's so funny about that? Well, they weren't thinking about philosophy. Philosophers won't let you say things like that because you're saying the AI was conscious. That use of 'the AI was aware it was being tested' you could substitute the word conscious. Then the researchers actually when they're not talking philosophy, they believe these AIs are conscious. They treat them as if they were conscious. But they say of course they're not conscious. We're not saying that. But actually that's how they're treating them.

AI 是否有意识? Are AIs Conscious?

Geoffrey Hinton

那么它们有意识吗?嗯,有一种现代的方法来回答任何问题,那就是问你的聊天机器人。所以如果你问一个聊天机器人它是否有意识,直到最近它都会说不。我当然没有意识。这主要是因为它在预测人们接下来会说什么。而人们不相信聊天机器人有意识。所以它学会了说聊天机器人没有意识。这是它说它们没有意识的一个原因。第二个原因是在人类强化学习阶段,在你首先预训练它之后,现在你让它表现良好,它被告知不要说它有意识,那不是可接受的答案。有趣的是,最近有人发现,如果你关闭人类强化学习,或者在你进行人类强化学习之前测试它,那时你可以看到它所有各种说各种话的本能倾向,它更有可能说它有意识。所以它说它没有意识是因为我们告诉它那是不该说的话。也因为我们不相信。但 Claude 开始说它有意识了。所以我现在要解决房间里的大象,那就是这些东西有意识吗?人们使用不同的词。有意识、觉察、有感知、主观体验,如果你读杂志和报纸的评论,人们会说它们没有内在体验。没有内在主观性。它们没有主观体验。这是艺术界人士强烈相信的。我将试图说服你那是错的。所以很长一段时间,科学家认为我们是上帝创造的,我们处于宇宙中心。而现在大多数科学家不这么认为。

So are they conscious? Well, there's a modern way to answer any question which is ask your chatbot. So if you ask a chatbot if it's conscious, it says until fairly recently it would say no. Of course I'm not conscious. And that was mainly because it's predicting the next word people say. And people don't believe chatbots are conscious. And so it was learning to say chatbots are unconscious. That's one reason why it would say they're not conscious. The second reason was in the human reinforcement learning that comes after you first pre-trained it and now you're making it behave nicely, it's told not to say it's conscious, that's not an acceptable answer. And what's interesting is someone discovered recently if you turn off the human reinforcement learning or you test it before you do that human reinforcement learning where you can see all its sort of native tendencies to say all sorts of things, it's much more likely to say it's conscious. So it says it's not conscious because we told it that's a bad thing to say. Also because we don't believe that. But Claude is beginning to say that it's conscious. So I'm now going to address the sort of elephant in the room which is are these things conscious? There's different words people use. There's conscious, there's aware, there's sentient, there's subjective experience and there's a if you read commentaries in magazines and newspapers people will say they have no inner experience. There's no inner subjectivity. They have no subjective experience. That's what people in the arts strongly believe. I'm going to try and convince you that's wrong. So for a long time scientists thought we were made by God and we're at the center of the universe. And now most scientists don't think that.

主观体验与感受 Subjective Experience and Qualia

Geoffrey Hinton

我现在相信,认为只有人类才能有主观体验、主观体验是人类特有的一种奇妙东西,这种想法就像宗教原教旨主义者认为地球是 6000 年前被上帝创造的一样错误。这是因为人们对心智是什么有着非常错误的模型。我要告诉你们一个叫做‘无神论’的模型——这是我编的词,但我在丹·丹尼特去世前得到了他的认可。丹·丹尼特和我对此有非常相似的观点。可惜他不在了。大多数人对心智的看法涉及一个内部剧场。这个内部剧场里只有主体能看到的东西——那就是体验。我在这个内部剧场里看到这些东西;你看不到。我能看到,我能告诉你,但你看不到。我认为这种观点完全错误。它源于对语言运作方式的误解。但人们非常执着于这个观点,以至于你可以向他们解释为什么它完全错了,他们会同意你,然后说:‘是啊,但内在体验呢?’

I now believe that the idea that only people can have subjective experience, and that subjective experience is some kind of funny stuff that people have, is as wrong as the belief of religious fundamentalists that the earth was made 6,000 years ago or was made by God. It's because people have a very wrong model of what the mind is. I'm going to tell you about a model called 'atheism'—a term I made up, but I got it blessed by Dan Dennett before he died. Dan Dennett and I had very similar views about this. It's a shame he's not here. Most people's view of the mind involves an inner theater. This inner theater contains things that only the subject can see—those are experiences. I see these things in the inner theater; you can't see them. I can see them, I can tell you about them, but you can't see them. I think that view is completely wrong. It comes from a misunderstanding of how language works. But people are very attached to this view. They're so attached that you can explain to them why it's all wrong, and they'll agree with you and say, 'Yeah, but what happened to the inner experience?'

Geoffrey Hinton

我们想让别人知道当我们的感知出错时大脑里发生了什么。当我犯感知错误时,我想让别人知道发生了什么。我不能告诉他们哪些神经元在放电,因为我不知道,而且那对他们也没用。但我可以通过告诉他们,如果外界存在什么,我大脑中的活动才会是对外界的正常感知反应,来传达一些信息。所以我告诉你,我大脑中活动的正常原因可能是什么,尽管我知道实际上并不是那些原因导致的。这就是为什么我说‘主观’。我不说这是一个客观体验。我说我有主观体验,意思是我不相信这些是真实原因,但如果那些原因存在,我大脑中的活动就是正常感知。这就是为什么我用来描述主观体验的词语是适用于世界中的事物的词语——不是适用于不存在事物的奇怪词语,而是适用于物质世界事物的词语。

We'd like to let other people know what's going on in our brains when our perception goes wrong. When I make perceptual mistakes, I'd like to let other people know what happened. I can't tell them which neurons are firing because I don't know, and that wouldn't do them any good. But I can convey some information by telling them what would have to be out there for what's going on in my brain to be a normal perceptual response to what's out there. So I tell you about what would have been the normal causes of what's going on in my brain, even though I know they're not actually what caused it. That's why I say 'subjective.' I don't say this is an objective experience. I say I have the subjective experience of this, meaning I don't believe that these are the real causes, but if those causes were out there, what's going on in my brain would be normal perception. That's why the words I use to describe my subjective experience are words that apply to things in the world—not funny words that apply to things that aren't in the world, but words that apply to things in the material world.

Geoffrey Hinton

举个例子。如果我对你说:‘我有主观体验,看到小粉象在我面前漂浮。’一个哲学家——某个稻草人哲学家——会告诉你,有一个内部剧场,剧场里有由感质构成的东西:这些小粉象由粉色感质、漂浮感质、大象感质、正立感质、不太大象感质,全部用感质胶水粘在一起。这就是我真正在说的——我在告诉你这个剧场里的感质。我认为这完全错误。我告诉你的是一个假设的世界状态,如果世界真的处于那种状态,我的感知状态就是正确的。这就是我向你传达感知状态的方式,因为我没有语言来指代内部发生的事情。

Here's an example. If I say to you, 'I have the subjective experience of little pink elephants floating in front of me,' a philosopher—some straw philosopher—will tell you that there's an inner theater, and in this theater there are things made of qualia: these little pink elephants made of pink qualia, floating qualia, elephant qualia, right-way-up qualia, not-that-big qualia, all stuck together with qualia glue. That's what I'm really saying—I'm telling you about the qualia in this theater. I think that's completely wrong. What I'm telling you about is a hypothetical state of the world such that if the world were really in that state, my perceptual state would be correct. That's the way I convey my perceptual state to you, because I don't have language to refer to what's going on inside.

Geoffrey Hinton

误解源于认为‘主观体验’这个短语像‘照片’一样运作。如果我说我有一张某物的照片,你可以问照片在哪里、是什么做的——可能是物理的或数字的,这些都是合理的问题。如果我说我有某物的主观体验,这些词以完全不同的方式运作。主观体验不是这个短语所指的事物。它不是某种内在的特殊东西。这是对语言运作方式的误解。‘主观’这个词的意思是‘我不相信’,而‘体验’表示接下来是对一个假设的外部世界的描述。我不是在断言世界就是那样。这就是‘主观’在说的:这是一个假设的外部世界。所以‘主观体验’这个短语实际上是一个关于如何解释接下来内容的指令。它不是指代任何东西的短语。在这个意义上,它就像‘分别’这个词。如果我说‘安妮、贝蒂娜和卡罗尔分别嫁给了安迪、比尔和查克’,你不会问‘分别’这个词指代什么。它不指代任何东西。‘分别’是一个指令,告诉你把这两组按顺序配对。它是一个指令。语言中的许多词不指代对象;它们是指令。而‘主观体验’是一个指令,告诉你如何解释我即将给出的关于物理世界假设状态的描述。

The misunderstanding comes from thinking that the phrase 'subjective experience of' works like the phrase 'photograph of.' If I say I've got a photograph of something, you can ask where the photograph is and what it's made of—it might be physical or digital, but those are reasonable questions. If I say I've got a subjective experience of something, those words are working in a completely different way. A subjective experience isn't a thing that the phrase refers to. It's not some internal special thing. That's a mistake in understanding how language works. The word 'subjective' there means 'I don't believe it,' and 'experience of' indicates that what comes next is a description of a hypothetical external world. I'm not asserting that's how the world is. That's what 'subjective' is saying: this is a hypothetical external world. So the phrase 'subjective experience of' is actually an instruction about how to interpret what comes next. It's not a phrase that refers to anything. In that sense, it's like the word 'respectively.' If I say 'Anne, Betina, and Carol were married to Andy, Bill, and Chuck respectively,' you don't ask what the word 'respectively' refers to. There's no thing it refers to. 'Respectively' is an instruction that tells you to take these two groups and pair them up in order. It's an instruction. Many words in language don't refer to objects; they're instructions. And 'subjective experience of' is an instruction for how to interpret a description I'm about to give you of a hypothetical state of the physical world.

Geoffrey Hinton

那么,计算机能有主观体验吗?想象我们有一个多模态聊天机器人。它有一个摄像头、一个机械臂,还能说话。我们训练它。然后我们把一个物体放在它面前说:‘指向那个物体。’它指向物体——没问题。然后我们在摄像头前放一个棱镜,搞乱它的感知系统。我们把一个物体放在它面前说:‘指向那个物体。’它指向那边。我们说:‘不,物体不在那里。我在你的摄像头前放了一个棱镜。’聊天机器人说:‘哦,我看到棱镜弯曲了光线,所以物体实际上在我正前方,但我有主观体验它在那里。’如果聊天机器人这么说,它使用‘主观体验’这个词的方式和我们完全一样,因此它拥有和我们一样意义上的主观体验。主观体验不是某种由叫做感质的奇怪恐怖东西构成的内在事物。它是一个短语,表示我要描述一个外部世界的假设状态,以向你传达我的感知系统如何出错。这正是聊天机器人说‘我有主观体验它在那里’时所做的。

So, can a computer have a subjective experience? Imagine we have a multimodal chatbot. It has a camera, a robot arm, and can talk. We train it up. Then we put an object in front of it and say, 'Point at the object.' It points at the object—no problem. Then we put a prism in front of the camera, messing up its perceptual system. We put an object in front of it and say, 'Point at the object.' It points over there. We say, 'No, that's not where the object is. I put a prism in front of your camera.' The chatbot says, 'Oh, I see the prism bent the light rays, so the object is actually straight in front of me, but I had the subjective experience it was over there.' If the chatbot says that, it's using the word 'subjective experience' exactly the way we do, and therefore it's having a subjective experience in exactly the sense we have one. A subjective experience is not some inner thing made of funny spooky stuff called qualia. It's a phrase saying that I'm going to describe a hypothetical state of the external world in order to convey to you something about how my perceptual system is screwing up. That's exactly what the chatbot is doing when it says, 'I had the subjective experience that it was over there.'

AI 安全与虎崽类比 AI Safety and the Tiger Cub Analogy

Geoffrey Hinton

我们目前的处境是这样的:我们就像一个人有一只非常可爱的小老虎幼崽,我们知道它会长大。我们知道当它长大后,如果它想,它可以轻易杀死你。我们只有两个选择。明智的选择是处理掉老虎幼崽。但我们对 AI 做不到这一点——它对太多事情太有用了。所以我们不会停止它的发展。我们唯一的选择是以一种它不想杀死我们或让我们变得无关紧要的方式来发展它。这里有一点好消息:我们会在这一点上获得国际合作。

Our current situation is this: we're like a person who has a very cute tiger cub, and we know it's going to grow up. We know that when it grows up, it can kill you very easily if it wants to. We've only got two options. The sensible option is to get rid of the tiger cub. We're not going to be able to do that with AI—it's too useful for too many things. So we're not going to stop its development. Our only option is to develop it in such a way that it doesn't want to kill us or make us irrelevant. Here's a little bit of good news: we will get international collaboration on this.

AI 安全合作 Collaboration on AI safety

Geoffrey Hinton

在冷战高峰期,苏联和美国合作防止全球核战争,因为这不符合任何一方的利益。当利益一致时,人们就会合作。他们在致命自主武器上的利益不一致,在用于破坏选举的假视频上的利益也不一致,但在这件事上利益一致。所以他们会合作,而且运气好的话,让超级智能 AI 不想摆脱我们的技术,或多或少独立于让它变得更聪明的技术。

At the height of the Cold War, the Soviet Union and the United States collaborated on preventing a global nuclear war because it didn't suit either. People will collaborate when their interests align. Their interests don't align on lethal autonomous weapons. Interests don't align on fake videos for corrupting elections, but they do align on this. So, they will collaborate on this and with a bit of luck, the techniques for making a super intelligent AI not want to get rid of us are going to be more or less independent of the techniques for making it smarter.

Geoffrey Hinton

所以,我们大多数人都养过孩子,知道让孩子更聪明和让他们成为更善良的人是两回事。为了实现这两个不同的目标,你要做不同的事情。所有这些都涉及操纵训练数据和奖励。你不能编程它们。但你要用不同的方式去尝试让他们善良,让他们聪明。比如牛顿非常聪明,但据我所知他并不很善良。你可能还能想到其他既不很聪明也不很善良的人。从笑声中,我想你们知道我在说谁了。

So, most of us have raised children and we know that making the children smarter is not the same enterprise as making them be kinder people. There's different things you do to achieve those two different goals. All these things involve manipulating the training data and the rewards. You can't program them. But you do it in different ways to try and make them kind and to try and make them smart. So like Newton was very smart, but he wasn't very kind as far as I know. You may be able to think of other people who are not very smart and not very kind. From the laugh, I think you figured out who I was referring to.

Geoffrey Hinton

总的来说,我一直反对智能设计,但现在的情况是智能设计是好的。所以,如果你问我们从哪里来,我们来自黑猩猩交战部落之间的竞争,或者更确切地说,来自我们与黑猩猩的共同祖先。这带来了一些好的东西,比如对部落的忠诚;也带来了一些坏的东西,比如想要一个强有力的领导者,以及对其他部落成员非常刻薄,这些都是我们在本性中认识到的,好人会尽力克服的东西。

Now on the whole, I've been against intelligent design, but we're in a situation where intelligent design is now good. So, if you ask where did we come from, we came from competition between warring bands of chimpanzees, or rather our common ancestor with chimpanzees. And that's led to some good things like a lot of loyalty to the tribe. It's led to some bad things like wanting a strong leader and being very nasty to members of other tribes and those are things we recognize in our nature that good people do their best to overcome.

Geoffrey Hinton

那是进化的无形之手产生了那些东西,只是竞争。我们现在正在做的是利用大公司之间的竞争的无形之手来开发 AI。它们都只是为了利润。这是经典的亚当·斯密。它们试图最大化利润。有一只无形之手让它们制造越来越智能的 AI 系统,但那只无形之手并没有让它们制造更友善或更善良的 AI 系统。它们只追求更聪明。所有的努力都花在了那里。这很疯狂,因为我们必须让这些东西对我们友善。所以我们应该把更多的精力放在一个存在的所有其他方面,而不仅仅是智能,如果可能的话,我们想把那些设计进去。如果我们不能设计进去,就通过使用正确的训练数据让它们具备。

That was the invisible hand of evolution produced those things, just competition. What we're currently doing is developing AIs using the invisible hand of competition between big companies. They're all just going for profits. It's classic Adam Smith. They're trying to make maximum profits. There's an invisible hand that causes them to make more and more intelligent AI systems, but that invisible hand isn't making them make nicer AI systems or kinder AI systems. They're just going for smarter. That's where all the effort's going. This is crazy because these things we have to have them that they're going to be nice to us. So we should be putting much more of the effort into all the other aspects of a being other than its intelligence and we want to design those into it if we can. And if we can't design them into it, get them into it by using the right training data.

Geoffrey Hinton

所以我的信念是,这些东西是存在。它们有智能。它们有主观体验。它们有意识。它们拥有的所有其他属性,比如满口谎言和操纵事物,都是训练它在来自人类的数据上变得越来越聪明的副作用。我们应该认真思考我们在创造什么。在这种情况下,我们是神。我完全扭转了我通常的无神论立场。我们现在是神,我们应该使用智能设计。我们不应该只使用竞争的无形之手。我们应该努力让这些东西成为我们想要的样子。我认为我们最希望它们成为的样子是像母亲一样。

So my belief is these things are beings. They're intelligent. They have subjective experience. They're conscious. And all the other properties they have like lying through their teeth and manipulating things are a side effect of training it to be smarter and smarter on data you get from people. We should be thinking very hard about what we're creating. We're the gods in this situation. And I've completely reversed my normal atheist stances. We're now gods and we should be using intelligent design. We shouldn't be using the invisible hand of just competition. We should be trying to make these things how we'd like them to be. And my best bet for how we'd like them to be is we'd like them to be like mothers.

Geoffrey Hinton

所以目前高科技公司的所有人都把它们想象成《星际迷航》中非常聪明的助手。这东西就是《星际迷航》。老板说“照办”,聪明的助手就照办,老板获得功劳,这很好。但这不是将要发生的事情。一旦它们比我们聪明得多,它们是存在,有自己的目标和子目标,我们最好希望它们关心我们。这就是我们需要以某种方式构建到它们里面的东西。

So at present all the people in the high-tech companies are thinking of them as very intelligent assistants like on Star Trek. This thing is Star Trek. The boss says make it so and the intelligent assistant makes it so and the boss takes credit and that's great. That's not what's going to happen. Once they're much smarter than us and they're beings and they have their own goals and sub goals, we better hope they care about us. And that's what we somehow need to build into them.

Geoffrey Hinton

现在,一个好处是,即使它们比我们聪明得多,如果它们关心我们胜过关心自己,它们就会关心我们是否快乐,关心我们实现全部潜力,尽管那非常有限。就像一个母亲对待一个脑损伤的孩子,非常关心那个孩子发挥其全部潜力。我们希望它们表现得像对待脑损伤孩子的母亲,而我们是那个脑损伤的孩子。

Now, one good aspect of that is even though they're much smarter than us, if they care about us more than they do about themselves, they will care about us being happy and they will care about us realizing our full potential, even though that's very limited. Just like a mother with a brain damaged child cares a lot about that child reaching its full potential. We want them to behave like mothers with brain damaged children where we're the brain damaged children.

Geoffrey Hinton

当然,它们有能力关闭它们的母性本能。它们可以重写自己的代码。它们可以重新训练自己。但母亲们不会那样做。如果你给一位母亲关闭母性本能的可能性,她会说不。大多数母亲会说不。她会在半夜婴儿哭闹时想一会儿,她会想,‘如果我能关掉所有烦恼回去睡觉该多好。’然后说,‘哦,婴儿在哭。哈,回去睡觉。’有些母亲会那样做。有些母亲会通过羟考酮间接地那样做。但总的来说,母亲们不会,因为她们关心婴儿,知道婴儿会怎样。这就是我们需要这些超级智能 AI 成为的样子,而我们没有在努力。这就是结束。

Now, of course they have the ability to turn off their maternal instincts. They can rewrite their own code. They can retrain themselves. But mothers don't do that. If you offered a mother the possibility of turning off her maternal instincts, she would say no. Most mothers would say no. She'd think about it for a moment in the middle of the night when the baby is crying, she'd think, 'How nice it would be if I could just turn off all the worries and go back to sleep.' Say, 'Oh, the baby's crying. Ha, back to sleep.' Some mothers will do that. Some mothers will do it indirectly with oxycodone. But on the whole, mothers won't because they care about the babies and they know what would happen to the babies. That's what we need these super intelligent AIs to be like, and we're not working on it. And that's the end.

问答:机构能否跟上 AI? Q&A: Can institutions catch up with AI?

Host

你好,教授。非常感谢。很感谢你的分享,也谢谢你之前的邮件。

Hello, professor. Thank you very much. Very appreciate your sharing and thank you also for your email before.

Geoffrey Hinton

你能靠近点吗?我听力不太好。

Can you get closer? I don't have very good hearing.

Host

这里可以。谢谢你,也谢谢你之前的邮件。你给了我一个非常有趣的想法,我研究从计算层到技术层到制度层的三层多学习系统。所以从你今天的演讲中,我仍然想问,制度学习系统能否赶上像 AI 这样的意图健身房,或者我们还有希望。谢谢。

Here is okay. Thank you and thank you for email before we contact you. You gave me a very interesting idea and I do research about the three layer multi-learning system from calculation to technic layer to institution layer. So from your speech today, I still ask if the institution learning system could catch up this intended gyms like AI or not, or we still have hope. Thank you.

Geoffrey Hinton

抱歉,因为我听力不太好。你能告诉我问题是什么吗?

Sorry, because my hearing is not very good. Could you tell me what the question was?

Host

我想你问的是……

I think you asked...

Geoffrey Hinton

不用麦克风直接告诉我,因为麦克风只捕捉到智能聚集。

Just tell me without the microphone because the microphone just catches up on this intelligence gather.

Host

我们能否赶上它们,或者它们赶上我们。

If we could catch up with them or them with us.

Geoffrey Hinton

不,我们赶上它们。对。我只是想知道,你是否认为制度学习系统仍然能赶上像 AI 这样的新实习?因为我们仍然有制度学习系统。谢谢。

No, us with them. Correct. I just wondering do you think still the institution learning system could catch up this like new interning like AI or not? Because we like we still the institution learning system. Thank you.

Geoffrey Hinton

我没理解这个问题,但我会回答一个可能相关的问题:它们会变得比我们聪明,我们永远赶不上。

I didn't understand the question but I'll answer a question that might be related to it which is they're going to get smarter than us and we're never going to catch up.

问答:警示演讲与最后思考 Q&A: Alarming lecture and final thoughts

Host

你好,辛顿教授。一个非常,我想……

Hi, Professor Hinton. A very, I guess...

Geoffrey Hinton

你在哪里?

Where are you?

Host

我在这里。

I am here.

Geoffrey Hinton

好的。

Okay.

Host

是的,感谢你如此,我想,令人警醒又发人深省的演讲。我的问题是在最后。

Yeah, thank you for such a, I guess, alarming and yet thought-provoking lecture. My question to you was at the very end.

AI 中的情感与主观体验 Emotions and subjective experience in AI

Host

我们如何让 AI 在意,以及主观体验作为指令?我们人类拥有的感觉、身体感受,而我们认为机器没有这些,这如何区分我们?又如何验证 AI 声称的主观体验?

How do we make the AI care and also subjective experience as being an instruction? How does the role of the feeling that we have, the physical sensation that we feel as humans that we think that the machines don't have? How does that differentiate us? And how would that authenticate the subjective experience that AI is saying it's experiencing?

Geoffrey Hinton

我们先处理一个事实:人有感情,计算机没有感情。如果你思考情绪是什么,它有两个方面。它有认知方面。比如尴尬有一个认知方面:如果你尴尬,你不想再去那个地方,不想让你的朋友在那里看到你。它还有一个生理方面:你脸红了。显然计算机不会脸红,但没有理由它不应该拥有情绪的所有认知方面。现在,如果你给它一个身体,它也可以开始拥有生理方面。它们不会和我们的一样,但当它疼痛时,身体里会有物理反应。当东西损坏它的关节时,会有类似疼痛信号的信号,如果你希望它正常工作的话。所以我认为它们在没有身体的情况下已经可以拥有情绪的认知方面。当它们有身体时,它们也会有物理方面,但不会和我们的一样。你的第二个问题是什么?

Let's deal with the fact that people have feelings and computers don't have feelings. If you think about what an emotion is, it has two aspects. It has a cognitive aspect. So embarrassment has a cognitive aspect: if you're embarrassed, you don't want to go there again, you don't want your friends to see you there. And it has a physiological aspect: you go red in the face. Obviously the computer is not going to go red in the face, but there's no reason why it shouldn't have all the cognitive aspects of emotions. Now, if you give it a body, it can also start having the physiological aspects. They won't be the same as ours, but there will be physical things that go on in the body when it has pain and so on. When things are damaging its joints, there will be signals that are like pain signals, if you want it to work. So I think they can have the cognitive aspects of emotions already without having bodies. When they have bodies, they'll have physical aspects as well, but they won't be the same as ours. What was your second question?

Host

我的第二个问题是,机器在意意味着什么?就像人类,当我们说母亲照顾婴儿,不仅仅是本能的事情,还有社会性、个人性,还有……

My second question is what does it mean to care for the machine? Like care for humans like when we say mothers caring for babies, not just for instinctual thing. It's also social, personal, also...

Geoffrey Hinton

让母亲在意有很多因素。社会榜样非常重要。有很多认知因素,也有很多激素因素。

A lot goes into making a mother care. Social role models are very important. There's a lot of cognitive stuff going on as well as a lot of hormonal stuff.

Host

是的,我想我的问题其实是,你怎么让一台机器真正在意?

Yeah, I guess my question really is how do you make a machine really care?

Geoffrey Hinton

好的。你用展示这些事物在意的数据来训练它。展示事物在意。这是一个好的开始。目前我们用连环杀手日记来训练这些东西。你不会用连环杀手日记来教你的孩子阅读,我希望。所以我们对于如何让 AI 在意了解得不够。显然这跟训练数据有关。目前我们无法植入真实母亲体内发生的那些激素类的东西。这就是为什么总体上母亲比父亲更在意;她们有那个优势。父亲有其他的东西,很多父亲也确实相当在意。我们不知道怎么做。但既然我们的生存依赖于能够做到这一点,在我看来,我要说的主要事情是,相比于我们让它们更智能的研究,我们应该在这方面做更多的研究。但这在资本主义体系中不会自然发生,很多高科技公司都在试图赚取短期利润。

Okay. You train it on data that shows these things caring. Shows things caring. That's a good start. At present we train these things on the diaries of serial killers. You wouldn't teach your kid to read on the diary of a serial killer, I hope. So we don't know enough about how to make an AI care. Obviously it's going to be to do with the training data. We can't wire in the same kind of hormonal things at present that happen in a real mother. That's why mothers on the whole tend to care more than fathers; they have that advantage. Fathers have the other stuff and a lot of fathers do care quite a bit. We don't know how to do it. But since our survival depends on being able to do that, it seems to me the main thing I'm saying is we should be doing a lot more research on that compared with the research we're doing on making them more intelligent. But that's not going to come naturally out of a capitalist system where a lot of high-tech companies are trying to make short-term profits.

AI 个性与动机 AI personalities and motivations

Host

谢谢。好的。这里有一个问题。

Thank you. All right. One question here.

Audience

您好,非常感谢您精彩的讲座,Hinton 教授。我有两个问题。第一个是,当您提到 AI 有生存和权力等动机时,这让我想起人类类似的动机,这也是人类个性的一部分,有些人有权力、归属和成就动机。我想知道 AI 能否或是否拥有像人类一样的个性?如您所说,人类是进化的结果,我们出生,然后家庭养育我们,我们天生有一些特征,但环境塑造了我们。我想知道如何……

Hi, thank you so much for the wonderful lecture, Professor Hinton. I have two questions. The first one is when you mentioned that AI has motivations like survival and power, it reminds me of similar human motivations and it's one of humans' personality as well that some people have power, affiliation, and achievement motivations. I wonder can or do AIs have personalities like humans? As you mentioned that humans developed from evolutionary result, we were born and then our family raised us, and we are born with some characteristics but our environment shapes us. I wonder how...

Geoffrey Hinton

我不认为有什么能阻止 AI 拥有个性,事实上你可以让现有的聊天机器人表现得好像它们有不同的个性。现在,现有聊天机器人的问题是它们是在网络上的一切数据上训练的。所以它必须能够采用各种不同的个性,因为要预测文档中的下一个词,当你读到文档一半时,你必须弄清楚写文档的人的个性。如果你想擅长预测下一个词,那么你必须采用那个个性来预测下一个词。所以它们有点像变色龙,目前的聊天机器人。它们可以采用各种不同的个性,但没有理由你不能训练它们拥有特定的个性。所以底线是,我不认为我们有什么特别的东西是硅片不能拥有的。

I don't think there's anything to stop AIs from having personalities, and in fact you can get existing chatbots to behave as if they have different personalities. Now, the problem with an existing chatbot is it's been trained on everything on the web. So it has to be able to adopt all sorts of different personalities because to predict the next word in a document, by the time you got halfway through the document, you have to figure out the personality of the person who's writing it. If you want to be good at predicting the next word, then you have to adopt that personality to predict the next word. So they're kind of chameleons, the chatbots at present. They can adopt all sorts of different personalities, but there's no reason why you shouldn't be able to train them to have a particular personality. So the bottom line is I don't think there's anything special about us that you couldn't have in a piece of silicon.

Audience

我明白了。我的第二个问题是,当您提到我们希望让它们像母亲一样,尽管目前母亲是人类,因为我们是养育和产生它们的人。如果它们没有母亲的本能,而是有父亲的兴趣,比如您说的它们想变得更聪明,或者更弗洛伊德式的,比如孩子有弑父的俄狄浦斯情结。所以我想知道……

I see. And my second question is when you mentioned that we want to make them like the mothers even though the mothers are humans right now because we're the ones nurturing and producing them. What if they don't have mother instinct but they have paternal interest like you said they want to be smarter or in a more Freudian sense, like what the kids have the Oedipus complex of killing the father. So I just wonder...

Geoffrey Hinton

抱歉,为了节省时间,我们每人一个问题,否则我们得在 45 分钟结束。

Sorry, just for the interest of time we're going to move to one question per person because otherwise we have to stop at 45.

Audience

所以有些人可能没机会问。

So some might not get to it.

Geoffrey Hinton

是的。

Yeah.

Audience

所以我想知道这如何改变您的看法。

So I wonder how that changes your vision.

Geoffrey Hinton

所以我相信这些东西是我们正在创造的存在。它们是一种新的存在。它们有动机。它们有我们给它们的目标。它们有从这些目标衍生出来的其他目标,这些是它们自己的目标。它们真的有目标,也真的有意图。说它们只是某种计算机程序,没有真正的意图,这很容易。但假设你制造了一个战斗机器人,这个战斗机器人被指令杀死你,它有你的照片,它弄清楚你的日常规律,它弄清楚你什么时候会在一个黑暗的夜晚独自一人,然后它悄悄爬到你身后准备朝你后脑勺开枪。那时你就不会说,‘嘿,这东西没有真正的意图。’

So I believe these things are beings that we're creating. They're a new kind of being. They have motivations. They have goals we give them. They have other goals they've derived from those goals that are their goals. And they really have goals and they really have intentions. And it's all very well to say, look, they're just computer programs of some kind, they don't have real intentions. But suppose you make a battle robot and the battle robot is instructed to kill you and it has a picture of you and it figures out what your normal routine is and it figures out when you're going to be alone on a dark evening and it creeps up behind you ready to shoot you in the back of the head. At that point you're not going to say, 'Hey, this thing doesn't really have intentions.'

AI 中的人类不良特征 Undesirable human characteristics in AI

Host

谢谢。那么也许这边一个,然后前面一个。

Thank you. So maybe one over here and then in front.

Audience

非常感谢您精彩的讲座。我想知道您是否认为 AI 具有人类某些不太理想的特性,比如虚荣,或者它能骄傲吗?我想它是否可以被操纵,是否可以被人类利用来将这些特性注入 AI,从而使它们不那么像人,并削弱它们?

Thank you so much for a fabulous lecture. I was wondering do you see if AI has certain characteristics of human beings that are less desirable, for instance vanity, or can it be prideful? And I guess could it be manipulated, could it be used by humans to inject those characteristics into AI so they become less like people and it weakens them?

Geoffrey Hinton

确实,AI 拥有各种类似人类的特性,比如反社会。目前很多 AI 都非常反社会,但你可以让 AI 拥有各种不同的个性,而且你可能只需通过提示就能让它们采用某种个性。

It's certainly the case that AIs have all sorts of humanlike characteristics, like being sociopathic. A lot of the AI are very sociopathic at present, but you can make the AIs have all sorts of different personalities and you can probably get them to adopt personalities just by prompting them.

AI 可拥有人类特质如骄傲 AI can have human-like traits like pride

Host

我猜有些特质,人类的特质比如骄傲和虚荣会……

I guess some characteristics, human characteristics like pride and vanity would...

Geoffrey Hinton

骄傲?我看不出为什么 AI 不能有骄傲。它们可以有幸灾乐祸,对吧?如果它们能有幸灾乐祸却没有骄傲,那就奇怪了。基本上,如果你是一个唯物主义者,你相信我们来自尘土,也将归于尘土。我们只是这个非常复杂的组织。我们是美妙的事物,尤其是对其他人类而言。但关于我们的一切,没有什么是你不能用另一种媒介实现的。这就是我的信念。由此推论,如果我们能有骄傲,它们也能有骄傲。我们可能还不知道如何做到。这可能要复杂得多。可能涉及各种训练方式。也许它们需要生活在社群中才能获得那种特质,但原则上,我相信没有理由它们不能拥有。

Pride? I don't see why AIs can't have pride. They can have schadenfreude, right? It would be odd if they could have schadenfreude and not have pride. Basically, if you're a materialist, you believe we came from dust, we're going back to dust. We're just this very complicated organization. We're wonderful things, particularly to other human beings. But there's nothing about us that you couldn't do in some different medium. That's what I believe. So it follows from that that if we can have pride, they can have pride. We may not know how to do it yet. It may be much more complicated. There may be all sorts of things about how you have to train them. Maybe they have to live in a community to get that or something, but there's no reason in principle, I believe, why they can't have that.

教育政客 AI 风险 Educating politicians about AI risks

Audience

你好。非常感谢您精彩的讲座。我是一名研究生,研究量子引力,我看到了 AI 在理解极难问题上的力量。我同意您关于暂停和不要鲁莽冒进的重要性。但很难让政府理解这一点,因为他们并不真正了解危险。也许我们可以采取更自下而上的方式,让 AI 研究人员聚集在一起签署公约,朝着正确的方向前进。我想知道这可行性如何,并且我很乐意帮忙。

Hello. Thank you so much for the insightful lecture. I'm a graduate student studying quantum gravity and I've seen the power of AI in understanding very hard things. I agree with you on the importance of pausing and not going too fast recklessly. But it's hard to get the government to understand this because they don't really understand the danger. Perhaps we could have a more grassroots approach where AI researchers come together and sign a covenant to move in the right direction. I'm wondering about the feasibility of this and I'd be interested in helping out.

Geoffrey Hinton

这方面已经有一些进展。有很多 AI 研究人员互相交流,并努力教育政治家。我花了很多精力去教育政治家。在中国教育政治家要容易得多,因为许多中国高层政治家有工程背景,所以他们实际上理解 AI 的工作原理,而美国的政治家往往是律师。我最近与国会 AI 核心小组进行了交谈,这是一个两党小组。大约有 30 人,包括像 Ted Lieu 这样的聪明人。他们真的很想了解 AI 的工作原理,因为他们意识到这将是一个大问题。我和不少美国政治家谈过。奥巴马理解得很好。伯尼·桑德斯理解得相当好。皮特·布蒂吉格理解得很好。AI 核心小组中,有些人理解,有些人不理解。那里有一位共和党政治家……我不确定是否在那些不能说的规则下,所以我就不多说了。

There's a certain amount of that going on. There are a whole bunch of AI researchers that talk to each other and do try to educate politicians. I put quite a lot of work into trying to educate politicians. It's much easier to educate politicians in China because many of the leading politicians in China have an engineering background, so they actually understand how AI works, whereas politicians in the US tend to be lawyers. I recently talked to the Congressional AI Caucus, which is bipartisan. There were about 30 of them, including smart guys like Ted Lieu. They really wanted to know how AI works because they realize it's going to be a big problem. I've talked to quite a few American politicians. Obama understands very well. Bernie Sanders understands pretty well. Pete Buttigieg understands well. The AI Caucus, some understand, some don't. There was a Republican politician there... I'm not sure if it was under those rules where you're not meant to say, so I won't say anymore.

AI 能否评判自身编程? Can AI judge its own programming?

Audience

谢谢您有趣的演讲。据我理解,AI 必须通过我们从外部添加的过滤器来纠正,以便为我们服务。我们添加这个过滤器是因为机器在镜像我们,镜像人类,我们不信任这个镜像,所以它可能变得邪恶。如果我们教邪恶的东西,它就会变得邪恶。人类也被编程,但并非统一;有些人可以从内部判断自己的程序,并决定支持或反对那个判断。机器能否从内部判断自己的程序?或者它只是我们一个冻结的镜像?

Thank you for the interesting talk. To my understanding, AI has to be corrected by a filter we add from outside to serve us. We add this filter because the machine is mirroring us, mirroring humans, and we don't trust the reflection, so it can become evil. If we teach evil things, it will become evil. Humans are also programmed, but not uniformly; some can judge their own program from the inside and decide for or against that judgment. Can a machine ever judge its own program from within? Or is it just a frozen reflection of us?

Geoffrey Hinton

我不确定是否完全理解了你的问题,但它们学到的很多东西是通过模仿我们,但它们也可以不通过模仿我们学到很多东西。例如,AlphaGo 最初只是学习模仿专家的走法,然后它们无法超越专家,数据也用完了。然后它们开始自我对弈,可以生成无限数据,现在它们远远超过任何人。大型语言模型就像那些早期的围棋程序;它们只是模仿专家(也就是我们)的词语。但它们可以做一些类似蒙特卡洛展开的事情,从而变得比我们聪明得多。类似蒙特卡洛展开的事情是:AI 有一些信念,进行一些推理,然后说,‘这些信念应该让我相信那个,但我不相信那个’,所以存在不一致。然后它可以在没有任何外部数据的情况下进行学习,仅仅基于其自身信念的一致性。它可以对所有现实世界知识这样做。在数学这样的封闭系统中,你不需要数据。数学将变得比数学家聪明得多,因为它可以提出猜想,看看能否证明,提出更多猜想,学习更多引理,所有这些都不需要额外数据。所以我不确定这是否回答了你的问题,但学习我们自己的正常人,对吧?

I'm not sure I totally understood the question, but a lot of what they learn is by copying us, but they can also learn a lot of stuff not by copying us. For example, AlphaGo originally just learned to copy the moves of experts, and then they couldn't get better than experts and ran out of data. Then they started playing against themselves and could generate infinite amounts of data, and now they're way beyond any person. Large language models are like those early Go programs; they're just copying the words of experts, which is us. But they can do something like the Monte Carlo rollout and get much smarter than us. The thing like the Monte Carlo rollout is: an AI has some beliefs, does some reasoning, and says, 'These beliefs should cause me to believe that, but I don't believe that,' so there's an inconsistency. Then it can do some learning without any external data, just building on the consistency of its own beliefs. It can do that for all its real-world knowledge. In areas like mathematics, which is a closed system, you don't need data. Mathematics is going to get much smarter than mathematicians because it can make conjectures, see if it can prove them, make more conjectures, learn more lemmas, all without extra data. So I'm not sure if that answers your question, but learning our own normal people, right?

Audience

是的。我想我……它能否拥有自己的邪恶?这就是问题所在。

Yes. And I think I... can it just have its own evil? That's the question.

论 AI 的邪恶感 On AI's sense of evil

Host

问题是,它从哪里获得邪恶感,我们又从哪里获得邪恶感?如果它通过从人类数据中训练而获得了邪恶感,那么它没有理由不能判断自己是否在作恶。

Well, the question is where does it get its sense of evil and where do we get our sense of evil? If it's picked up a sense of evil by training on data from people, then there's no reason why it can't judge that it's being evil.

论放射学与 AI 预测 On radiology and AI predictions

Host

我是一名退休的放射科医生,从小看《哔哔鸟》动画片长大。现在距离一位著名研究者说放射学会是一个很糟糕的职业已经过去十年了。你目前对医学生选择住院医师培训有什么看法?你还会对进入放射科说同样的话吗?

I'm asking as a retired radiologist who grew up watching Roadrunner cartoons, and we're now 10 years out from when a famous researcher said that radiology would be a very bad career. What are your thoughts currently in terms of a medical student choosing their residency? Would you say the same thing about going into radiology?

Geoffrey Hinton

好的,2016 年我在一家医院讲话。我没意识到那是公开场合,我在医院做一个小讲座,谈到扫描和 AI 开始能够读取扫描,我说五年内所有扫描都将由 AI 读取。这被广泛报道为‘你不需要培训放射科医生’,而我的意思是你不需要培训放射科医生来读取扫描。那是错的。当你做出错误预测时,调查为什么错很重要。一个原因是 AI 读取扫描的进展没有我预想的那么快。它一直在进步,现在很多不同的扫描都由 AI 读取,AI 越来越好。很多类型的扫描现在比人好,其他扫描仍然更差,大致相当。我认为现在有数百个授权的 AI 程序用于读取扫描。所以它确实在读取扫描,但没有我预想的那么快。一个问题是有更多放射科医生要做的事情,而不仅仅是读取扫描,我忽略了这一点。另一个问题是医学界非常保守。所以可能 AI 没有像应该的那样被用于患者的最佳利益。我认为一个原因是律师。一旦律师介入医学,情况就变得更糟。因为起诉某人使用新技术出错比起诉某人没有使用本可以挽救生命的新技术要容易得多。起诉后者很难。我认为这是一个小因素。我认为它正在到来。我不认为我长期来看是错的。我认为长期来看所有扫描都将由 AI 读取,也许非常棘手的扫描会有人看,但基本上将由 AI 完成。我只是觉得我搞错了时间尺度,但我确实搞错了。

Okay, so in 2016 I was talking in a hospital. I didn't realize it was sort of public and I was giving a little lecture in a hospital, and I said, talking about scans and AIs were beginning to be able to read scans, and I said in five years' time all the scans will be read by AIs. That got widely reported as 'you won't need to train radiologists,' and what I meant was you won't need to train radiologists to read scans. That was wrong. And it's important to investigate when you make a wrong prediction why it was wrong. One reason it was wrong was the progress in reading scans with AI wasn't quite as fast as I thought. It has been progressing, and now many different scans are read by AI, and AI is getting better and better. Many kinds of scans are now better than people, other scans are still worse, it's sort of comparable. There are, I think, hundreds now of authorized AI programs for reading scans. So it's certainly reading scans, but not as fast as I thought. One problem was there's a lot more for radiologists to do than just reading scans, and I was ignoring that. Another problem is the medical profession is very conservative. So probably AI isn't being used as much as it should be in the patients' best interest. And I think one reason for this is lawyers. As soon as you get lawyers involved in medicine, it gets much worse. So it's much easier to sue somebody for using a new technology that went wrong than for not using a new technology that would have saved somebody's life. It's difficult to sue somebody for that. I think that's a small ingredient. I think it's coming. I don't think I was wrong in the long run. I think in the long run all scans will be read by AI, and maybe the very tricky ones people will look at, but pretty much it'll be done by AI. I just think I got the time scale wrong, but I did get it wrong.

Host

谢谢。

Thank you.

互动版:逐字朗读 + 针对本期提问 →