The Godmother of AI on Vision, Learning, and Human-Centered AI
打开互动全文版(中英对照 + 朗读 + 问答)→李飞飞博士探讨视觉的进化、人脑的学习方式,以及保持人工智能以人为本的重要性。
Dr. Fei-Fei Li discusses the evolution of vision, how human brains learn, and the importance of keeping AI human-centered.
我认为人类最不擅长的一件事,就是老一辈总是哀叹年轻一代,好像年轻人什么都不懂。他们很粗鲁。他们忘了过去。但纵观人类历史的长河,总体而言,我们是在朝着更好的方向前进。我不否认那些暴行,不否认那些挫折,不否认这些。但从根本上说,我对人类是乐观的。我看着孩子们,他们充满好奇。当然,他们被这项技术极大地娱乐了,但他们也开始使用它。我担心的是老师和一些家长,因为我认为我们当今的社会,尤其是硅谷,并没有为他们服务。我们把他们忘了。大家好。为了庆祝我的新书《协议》的发布,我很高兴地宣布,我很快将举办三场现场活动。第一场在纽约市无线电城音乐厅,9 月 17 日。第二场在洛杉矶杜比剧院,10 月 8 日。第三场在旧金山共济会礼堂,10 月 28 日。在每场活动中,我都会讨论书中的话题,而我最喜欢的部分是直接回答你们——观众的问题。要购票,你可以访问 hubermanlab.com/events,并使用代码 protocols 获得提前访问权。再次强调,那是 hubermanlab.com/events,使用代码 protocols 获得提前购票权。欢迎收听 Hubberman Lab 播客,我们在这里讨论科学以及基于科学的日常工具。我是 Andrew Huberman,斯坦福大学医学院神经生物学和眼科学教授。我今天的嘉宾是 Dr. Fay Lee,她是斯坦福大学的计算机科学家和教授,也是人工智能和计算机视觉的先驱和杰出人物之一。如大家所知,每天有数百万人使用 AI 聊天机器人查找信息。当然,许多人担心 AI,担心它的走向,以及它可能如何取代某些人类工作或以某种方式降低我们的生活体验。今天我们从一个神经科学的角度讨论智能到底是什么,以及 AI 可以并且正在被用于善途的方式,即真正增强学习、健康,并丰富而非削弱人类体验。我们首先讨论人类大脑在各个年龄段如何学习新信息,大脑在该过程中遵循什么规则,以及 AI 因为基于互联网内容,如何既类似于人类大脑的学习又有所不足。我们还讨论了 AI 和机器人在医学中的激动人心的应用。需要明确的是,Fei-Fei 承认并回应了关于 AI 的许多合理担忧。但作为斯坦福大学以人为本人工智能研究所的主任,她的目标是确保人类和整个人类在 AI 的未来走向中得到代表。正如你很快会听到的,Dr. Fei-Fei 是一位杰出的科学家和教育家。她被称为“AI 教母”,不仅因为她引领了 AI 技术,还因为她坚持 AI 的伦理和善意用途应始终处于 AI 和机器人的核心。所以无论你是年轻人还是长者,今天的对话都将为你提供信息和力量,让你以真正有益于你并丰富你生活的方式理解和使用 AI。在我们开始之前,我想强调,这个播客独立于我在斯坦福的教学和研究角色。但它是我希望并努力向公众提供零成本的科学和科学相关工具信息的一部分。与此主题一致,今天的节目确实包含赞助商。现在,开始我与 Dr. Fei-Fei 的讨论。Dr. Fei-Fei,欢迎。
I think the biggest thing humanity never learns is the older generation lamenting about the future generation as if the future generation doesn't know anything. They're rude. They're forgetting the past. But if you look at arc of history, of humanity, by and large, we advance for the better. Now, I'm not denying the atrocities. I'm not denying the setbacks. I'm not denying this. But fundamentally I'm a optimist in humanity. I look at kids, they're curious. Of course, they get massively entertained by this technology, but they also are starting to use it. What I worry about are teachers and some parents because I think our society today and especially Silicon Valley are not doing them a service. We're forgetting about them. Hey everyone. To celebrate the launch of my new book entitled Protocols, I'm pleased to share that I'll be hosting three live events very soon. The first live event is in New York City at Radio City Music Hall on September 17th. The second event is in Los Angeles at the Dolby Theater on October 8th. And the third live event is in San Francisco at the Masonic on October 28th. At each of these events, I'll be discussing topics from the book and my favorite part, taking questions directly from you, the audience. To get tickets, you can go to hubermanlab.com/events and use the code protocols to get early access. Again, that's hubermanlab.com/events and use the code protocols to get early access to tickets. Welcome to the Hubberman Lab podcast where we discuss science and science-based tools for everyday life. I'm Andrew Huberman and I'm a professor of neurobiology and opthalmology at Stanford School of Medicine. My guest today is Dr. Fay Lee, a computer scientist and professor at Stanford and one of the pioneers and luminaries of artificial intelligence and computer vision. As you all know, millions of people use AI chat bots to look up information every single day. And of course, many people are concerned about AI, where it's going, and how it might replace certain human jobs or degrade our experience of life in one way or another. Today we discuss from a neuroscience perspective what intelligence really is and the ways that AI can and is being used for good meaning to truly enhance learning health and to enrich rather than diminish the human experience. We start off by talking about how human brains of all ages learn new information. What rules the brain follows in that process and how AI because it is based on the content of the internet both resembles and falls short of what human brains can learn. and we discuss exciting uses of AI and robotics in medicine. To be clear, FFE acknowledges and addresses the many valid concerns about AI. But as the director of the Stanford Institute for Human- Centered Artificial Intelligence, her goal is to make sure that humans and humanity at large are represented in where AI goes next. As you'll soon hear, Dr. Fa Lee is an extraordinary scientist and educator. She has been called the godmother of AI for her ushering in of AI technologies, but also for her insistence that the ethics and benevolent uses of AI stay central to AI and robotics. So whether you are young or old, today's conversation will inform and empower you to understand and use AI in ways that truly benefit you and enrich your life. Before we begin, I'd like to emphasize that this podcast is separate from my teaching and research roles at Stanford. It is however part of my desire and effort to bring zero cost to consumer information about science and science related tools to the general public. In keeping with that theme, today's episode does include sponsors. And now for my discussion with Dr. Fay Lee. Dr. Fay Lee, welcome.
谢谢。我很高兴来到这里,Andrew。
Thank you. I'm excited to be here, Andrew.
是的,这一天终于来了。你是 AI 领域的杰出人物,但我也认为你是一位神经科学家和计算机科学家,我们通过视觉科学有着共同的路径。所以我想从视觉开始。视觉、看见和光,对于 AI 及其未来走向有什么特别之处?因为我认为对大多数人来说,这些听起来像是非常不相关的主题,但实际上那正是一切的起点。
Yeah, this is a long time coming. And yes, you are a luminary in this AI field, but I also consider you a neuroscientist and computer scientist, and we share a common path through vision science. And so I'd like to start in vision. What is so special about vision and seeing and light as it pertains to AI and where it's all going? Because I think for most people those probably sound like very divorced themes but actually that's where it all starts.
是的。我认为视觉是智能的基石,几乎以两种平行的方式。一种是进化教给我们的。你知道视觉与动物智能和人类智能的进化关系。另一种是计算机视觉与 AI 的关系。所以我将分别探讨。我常说,5.4 亿年前,动物第一次看到了光。这些是简单的海洋动物,三叶虫及其近亲。在那之前,几乎没有感知。大约同一时期,触觉和触感也开始在动物身体中出现,但没有听觉,没有嗅觉,但绝对没有神经系统。但第一批感光细胞产生了一种进化力量,推动了动物的进化,因为感知外部世界改变了你的自我认知,改变了你与外部世界的关系。简单来说,如果你能看到食物,那会改变你的生活,对吧?从进化的角度看,你成为别人的食物,同时你也在积极寻找食物。你在积极寻找配偶,等等。所以正是因为感知和知觉,进化在动物物种形成方面以惊人的速度加速。化石研究告诉我们,在动物第一次看到光之后的 1000 万年,就是我们所说的进化大爆炸或动物物种形成的寒武纪大爆发。快进到现在,我认为视觉不仅在动物的早期进化中发挥了巨大作用,而且在高级智能及其出现方式中也发挥了巨大作用。你和我都是视觉研究的学生和科学家。据估计,人类大脑中一半的皮层活动涉及视觉功能。儿童在发育过程中先有视觉,后有语言。所以视觉至今在动物智能的进化和人类日常生活中都扮演着核心角色。与此同时,视觉作为一门学科或 AI 的一个领域,确实在几个方面对我们所看到的现代 AI 时刻起到了关键作用。首先是算法,神经网络算法。神经网络算法最早是计算机科学家在 20 世纪 50 年代初开始涉足的。
Yeah. I see vision as a cornerstone of intelligence in almost two parallel way. One is what evolution has taught us. You know what's the evolution of vision and animal intelligence and human intelligence. The other one is computer vision and AI what that relationship is. So I'll go into each evolution. I always say that 540 million years ago animals saw the first light. These are simple sea ocean animals, trilobytes and and the the cousins. And before that there was very little sensing. Uh around that same time tactile and haptics was starting also to emerge in animal bodies but but there was no hearing. There's no you know smelling there's no but there's absolutely no nervous system. But the first photoreceptive cells created a evolutionary force that propelled animals to evolve because sensing the external world changes your self-perception changes the way your relationship with the external world. To put it simply, if you seek you can see food, it changes your your life, right? from a evolution point of view and you become someone else's food and also you're actively seeking food. You're actively seeking mates and and and all that. So really because of sensing and perception evolution took a incredibly accelerated pace in terms of uh animal speciation. Fossil studies have told us that 10 million years after the first uh light for animals was what we call the the big ban of evolution or Cambrian explosion of animal speciation. And fast forward I think vision has always played a huge role in not only in the early evolution of animals but as well as um advanced intelligence and how that emerged. You and I are both vision student and and scientists. It is estimated half of the cortical AC activities in human brain is involved in visual function. Children were first visual before they were verbal in development. So vision really to this day plays a central role in both the evolution of animal intelligence as well as in the daily life of human human life. Now in parallel, vision as a uh as a discipline or as a area of uh artificial intelligence was really played a pivotal role in what we see as this modern AI moment in a couple of ways. First of all is the the uh algorithms the neuronet network algorithms. Neural network algorithms were first computer scientists start dabbling that in the early 1950s.
安德鲁,你可能还记得,20 世纪 50 年代初神经科学领域发生的事,神经科学家如休伯尔和维塞尔开始记录哺乳动物大脑中的视觉细胞,并逐渐意识到神经细胞存在层级结构,它们层层堆叠,在这些层级间传递神经信息。从视网膜收集光线,一直到识别出你面前有一个形状。我们在哺乳动物大脑中看到的这种神经架构,也是神经网络算法灵感的来源之一。如今,神经网络算法运行在数千亿甚至数万亿的参数上,其复杂性已远超我们在哺乳动物大脑或视觉通路中记录到的情况,但它们的起源非常接近,大约在半个多世纪前。这是视觉对 AI 贡献的一个方面。
And Andrew, you might remember what's happening on the neuroscience side in the early 1950s is that neuroscientists like Hubel and Wiesel were starting to record visual cells in the mammalian brain and starting to realize there is a hierarchical structure of nervous cells that stack against each other and pass neuroinformation across these hierarchies. And it goes from collecting light from the retina all the way to recognizing there is a shape in front of you. And that very neuroarchitecture that we see in the mammalian brain is also part of the inspiration of neural network algorithms. Now today's neural network algorithms run on hundreds of billions and even trillions of parameters. It has the complexity that departs from what we recorded in the mammalian brain or the visual pathway, but the origin is very close to each other about half a century ago, a little more than half a century ago. That's one aspect of vision's contribution to AI.
视觉对 AI 的贡献还有另一个同样关键的方面,那就是通过大数据。这更接近我自己的工作。世纪之交的 AI 是一个机器学习领域。许多不同的实验室、不同的研究科学家都在尝试不同的算法,不仅仅是神经网络,还有其他方法,比如贝叶斯方法、支持向量机方法等术语。这些方法是什么并不重要,但这是一个探索阶段,我们试图让这些算法发挥作用,以便让机器能够阅读或看见。我们一群计算机视觉科学家在这些算法上挣扎,而我是 2006 年普林斯顿大学的一名年轻教师,第一年任教,我和我的学生们在审视这些算法,以及这些算法学习时输入的数据是多么少。于是我转向认知神经科学文献,即视觉文献,开始研究人类能学习多少、能看见多少,这些数字令人难以置信。人类到 6 岁时能学习数万种不同的物体类别,而且接触视觉世界的量也是巨大的,对吧?婴儿从出生那一刻起大部分时间都能看见,所以他们被大数据淹没。因此我们推测,数据的缺乏是 AI 进展缓慢的一个巨大原因。于是我们与其他只专注于算法的人分道扬镳,说我们需要数据,我们需要用数据来驱动这些算法。长话短说,我们领导了 ImageNet 项目,收集了人工智能领域第一个互联网规模的大型数据集,但实际上是透过视觉领域,因为 ImageNet 收集了 1500 万张图像。ImageNet 的目标是驱动机器识别日常物体,比如麦克风、杯子、椅子。这项工作与神经网络算法以及 GPU 计算的进步相融合。到 2012 年,这项工作,即现代 AI 三要素的融合,成为现代 AI 的决定性时刻。
There is another aspect of vision's contribution to AI that is also pivotal, which is through big data. That comes closer to my own work. AI around the century was a field of machine learning. A lot of different labs, different research scientists were trying out different algorithms, and it's not just neural networks. There are other methods, jargon words like Bayesian methods, support vector machine methods. It doesn't matter what these methods are, but it's an explorative phase that we're trying to get these algorithms to work so that we can empower the machine to read or to see. A group of us computer vision scientists were struggling with these algorithms, and I was a very young faculty, first year faculty 2006 at Princeton, and my students and I are looking at these algorithms and how little data were fed into these algorithms to learn. So I turned to cognitive neuroscience literature, namely vision literature, and started to study how much humans learn, how much humans can see, and the numbers were incredible. Humans by age six can learn tens of thousands of different object categories, and the exposure to visual world is also massive. Right? Babies can see most of the time the moment they're born. So they're inundated with this big data. So we conjectured that the lack of data was a huge part of the reason for the lack of progress in AI. So we took a departure from everybody else who were really focusing only on algorithms and said that we need data. We need data to drive these algorithms. So long story short, we led this ImageNet project that collected the first ever internet-scale large dataset for the field of artificial intelligence, but really through the field of vision because ImageNet is a collection of 15 million images. And the goal of ImageNet was to drive machines to recognize everyday objects, you know, microphones, cups, chairs. And that work converged with the advances in neural network algorithms as well as in GPU computing. And by 2012, that work, the convergence of the three elements of modern AI, became the defining moment of what modern AI is.
我记得大约在 2012 年,在冷泉港举办的隔年一次的视觉课程上,似乎有一场辩论,讨论计算机能否像人类一样学会识别特定人脸。现在我想大多数人会说计算机实际上比人类做得更好,即使有那些在这方面异常出色的超级识别者。
I recall somewhere around 2012, it seems there was this debate at this vision course at Cold Spring Harbor that was held every other summer, like could a computer learn to recognize specific faces as well as humans. Now I think most people would say computers are actually much better at it than humans are, even though you have these super recognizer people who are exceptional at this.
你能告诉我们,这项技术是如何从基本上会把你和表亲,甚至某个长得有点像你的人搞混的状态,发展到如今精确无比的地步?我们是怎么走到这一步的?
Could you tell us how is it that this technology went from a state basically where it would confuse you and maybe a cousin or even someone that looks somewhat like you, to the point where now it is exquisitely precise? How do we get here?
我想一定要深入探讨这项技术的融合。我认为在 21 世纪的第二个十年,就像你说的 2012 年左右,巨大的融合是 GPU 计算的能力,它基本上加速或并行化了计算,让你可以在算法中运行更多的浮点运算,对吧?你需要那种速度。然后,经过几十年的研究,神经网络算法本身也变得更加成熟。你知道,正如我们所说,从 20 世纪 50 年代开始,人们开始创造这些非常简单的算法,它们的行为类似于神经元,但简单得多。神经元,如你所知,非常复杂,但这里的想法是,你有一个节点单元,它接受一些输入并输出另一个输入,在内部它只是一个函数,一个非常简单的函数。所以你把它们堆叠起来。这就是神经网络。但到了 2010 年左右,这些算法的成熟度已经达到了一个非常好的水平。但同样重要的是,对大数据的认识。互联网无疑推动了这一点,它使数据更容易获得。但那个顿悟时刻是‘哇,大数据需要成为这个等式的一部分。我们需要用大数据来驱动这些算法学习这些模式。’所以这三者的融合真正引发了 AI 的革命。
I want to definitely double triple click on the convergence of this technology. I think around the second decade of the 21st century, so like you said around 2012, the huge convergence was the capability of GPU computing, which basically accelerated or parallelized computing so that you can have more flops going through algorithms, right? You need that speed. Then you also have, after many decades of research, neural network algorithms themselves are getting more mature. You know, starting as we said in the 1950s, people started to create these very simple algorithms that behave similarly to neurons but much simpler. Neurons, as you know, are very complex, but here the idea is that you have one unit of node that takes some input and outputs another input, and within it, it's just a function, a very simple function. So you stack them together. That's what neural networks are. But by the time it's around 2010ish, the maturity of these algorithms has gotten to a level that it's becoming really good. But also, last but not least, the recognition of big data. The internet definitely fueled that. It made data more available. But the reckoning moment of 'wow, big data needs to be part of that equation. We need to use big data to drive these algorithms to learn these patterns.' So this convergence of these three things really set off the revolution of AI.
具体的时刻也值得一提,因为你提到了人脸识别。这就是 ImageNet 挑战赛。我的实验室从 2010 年开始,在我们收集了这个庞大的数据集之后,当时 GPU 还不成熟,我们向研究社区连续多年发起公开挑战,邀请人们解决这个名为物体识别的重大计算机视觉问题。任务非常简单。我们有一个包含一千个不同物体类别的数据集,这个数据集超过一百万张图像。我们称之为测试数据集,算法的任务是:我给你看一张图片,你必须说出里面的主要物体,如果你猜对了,就得一分;如果猜错了,就不得分。所以那个 ImageNet 挑战赛,几年后我们通过斯坦福一位非常聪明的研究生来基准测试人类表现,错误率大约是 4%。所以随机猜测的概率是千分之一。
The specific moment is also worth mentioning because you mentioned face recognition. This is the ImageNet challenge. My lab put forward that starting 2010, after we collected this humongous dataset, we at that point GPU was not yet mature, and we put out a public challenge for the research community for multiple years in a row and invited people to solve this major computer vision problem called object recognition. The task was very easy. We have a dataset of a thousand different categories of objects, and this dataset is more than a million images large. It's what we call the testing dataset, and the task for the algorithm is: I'll show you a picture, you have to name the main objects inside, and if you guess right, you get a point. If you guess wrong, you don't get a point. So that ImageNet challenge, we later a couple of years later benchmarked human performance by a very smart graduate student at Stanford, and that was roughly 4% error rate. So random chance will be one over a thousand.
对吧?所以对人类来说,4% 的错误率并不算太差。
Right? So 4% for humans is not that bad.
最初几年,机器不如人类。转折点是 2012 年,神经网络、ImageNet 数据集和 GPU 的融合。即使在那一年,尽管错误率被降低了,哦,顺便说一下,人类表现的错误率是 4%。抱歉,我需要纠正一下。错误率被降到了百分之十几。它还没有达到人类表现的水平。
The first few years machines were not as good as humans. The turning point was 2012, the convergence of neural network, ImageNet dataset, and GPU. Even that year, even though the error rate was cut, oh by the way, the human performance error rate was 4%. Sorry, I need to correct that. The error rate was cut down to the teens. It wasn't where human performance was.
所以这是看图像,并从一千个标签中分配一个。
So this is looking at images and assigning one out of a thousand labels.
明白了。
Got it.
是的。
Yeah.
但 2012 年是如此重要,因为这一年,神经网络算法使错误率比之前的算法大幅下降。
But 2012 was so momentous that year because the error rate from previous algorithms dropped a lot by this neural network algorithm.
我们知道在研究界,当如此剧烈的事情发生时,就意味着一个转折点。但还记得,从 2012 年到 2016 年,又花了三年时间,算法才在命名一千个物体上超越人类。
And we know in the research community when something this drastic happens it means an inflection point. But it still took another three years, I remember by 2012 to 2016, for the algorithm to beat humans in naming a thousand objects.
我能问一下,这位非常聪明的研究生那 4% 的错误来自哪里?是他们不认识这些物体,还是因为时间压力下的识别?比如他们被快速输入图像,偶尔会做出错误的判断。
Could I ask you where this 4% error is coming from in this very smart graduate student? Is it that they don't recognize the objects or it's a recognition against time pressure? Like they have to be fed images fast enough that occasionally they do an incorrect assignment.
我认为时间压力不是主要问题,尽管对研究生来说,他们也不想一直做这个。但你知道,人脑的记忆有限,无论是长期还是短期。所以记住一千个物体类别的模式,即使有些类别你很熟悉,也不是那么容易。
I don't think the time pressure was the main issue, even though for a graduate student to do this, I don't think they want to do this forever. But I think the human brain, as you know, has limited memory, whether it's long-term or short-term. So retaining the patterns of a thousand object classes, even if some classes you're familiar with, is not that easy.
所以存在混淆,而且,比如不同种类的狗非常相似。
So there is the confusion, and also, for example, different species of dogs get really close.
嗯,那是个挑战。
Mhm. And that's a challenge.
我想快速休息一下,感谢我们的赞助商 Lingo。Lingo 是一款日常可穿戴设备,全天候追踪你的血糖。血糖驱动许多关键过程,支持能量、身体成分和长期健康。当血糖持续飙升和暴跌时,我们就开始看到代谢功能障碍。随着时间的推移,这甚至可能发展为糖尿病前期。目前,美国约有 1.15 亿成年人患有糖尿病前期。大多数人不知道,而且男性患病比例高于女性。通常,糖尿病前期早期没有明显症状,所以人们往往不会去检查。但事实是,代谢健康每天都在塑造你身体的功能,无论你是否感觉到。使用 Lingo 追踪血糖可以帮助你了解食物、活动和压力如何全天影响你的血糖。我个人使用过 Lingo,它是我改善代谢健康的宝贵工具。如果你想尝试 Lingo,美国和英国的 Huberman Lab 听众可以在四周计划上节省 10%。只需访问 hellolingo.com/huberman 了解更多信息。条款和条件适用。再次强调,那是 hellolingo.com/huberman。今天的节目也由 Wealthfront 赞助。在当今市场不断变化和新闻混乱的金融环境中,很容易对如何储蓄和投资感到不确定。Wealthfront 是帮助你控制资金同时管理风险的解决方案。近十年来,我一直信任 Wealthfront 来应对这种波动。通过 Wealthfront 现金账户,我可以从项目银行获得 3.3% 的年收益率(APY)。我知道我的钱在增长,直到我准备好花掉或投资。我喜欢 Wealthfront 的一个特点是,我可以全天候即时免费提款到符合条件的账户。这意味着我可以随时移动资金,无需等待。当我准备好从储蓄转向投资时,Wealthfront 让我无缝地将资金转入他们专家构建的投资组合之一。限时优惠,Wealthfront 为 Huberman Lab 听众提供基础利率上 75% 的独家 APY 提升,为期 3 个月,这意味着你可以在高达 15 万美元的存款上获得最高 4.05% 的可变 APY。已有超过 100 万人信任 Wealthfront 来储蓄更多、赚取更多,并自信地建立长期财富。如果你想尝试 Wealthfront,可以访问 wealthfront.com/huberman 获得提升优惠,并立即开始赚取 4.05% 的可变 APY。那是 wealthfront.com/huberman 开始。这是 Wealthfront 的付费推荐。客户体验会有所不同。Wealthfront 经纪公司不是银行。基础 APY 截至 2026 年 1 月 30 日,可能发生变化。更多信息请参阅节目描述。
I'd like to take a quick break and acknowledge our sponsor, Lingo. Lingo is an everyday wearable that tracks your glucose 24/7. Glucose drives a lot of key processes that support energy, body composition, and long-term health. When glucose is constantly spiking and crashing, that's where we can start to see metabolic dysfunction. And over time, that can even progress to pre-diabetes. Right now, about 115 million adults in the US have pre-diabetes. Most don't know it, and a higher percentage of men have it than women do. Often, there aren't clear symptoms of pre-diabetes early on, so people don't tend to look into it. But the fact is that metabolic health is shaping how your body functions every day, whether you feel it or not. Tracking your glucose with Lingo can help you see how food, activity, and stress impact your glucose throughout the day. I personally have used Lingo and it's been an invaluable tool for improving my metabolic health. If you would like to try Lingo, Huberman Lab listeners in the US and UK can save 10% on a four-week plan. Just visit hellolingo.com/huberman for more information. Terms and conditions apply. Again, that's hellolingo.com/huberman. Today's episode is also brought to us by Wealthfront. In today's financial landscape of constant market shifts and chaotic news, it's easy to feel uncertain about how to save and invest your money. Wealthfront is the solution that helps you take control of your money while managing risk. For nearly a decade, I've trusted Wealthfront to navigate this volatility. With the Wealthfront cash account, I can earn 3.3% annual percentage yield or APY on my cash from program banks. And I know my money is growing until I'm ready to spend it or invest it. One of the features I love about Wealthfront is that I have access to instant no fee withdrawals to eligible accounts 24/7. That means I can move my money where I need it without waiting. And when I'm ready to transition from saving to investing, Wealthfront lets me seamlessly transfer my funds into one of their expert-built portfolios. For a limited time, Wealthfront is offering the Huberman Lab audience an exclusive 75% APY boost over the base rate for 3 months, meaning you can get up to 4.05% variable APY on up to $150,000 in deposits. Over 1 million people already trust Wealthfront to save more, earn more, and build long-term wealth with confidence. If you'd like to try Wealthfront, you can go to wealthfront.com/huberman to receive the boost offer and start earning 4.05% variable APY today. That's wealthfront.com/huberman to get started. This is a paid testimonial of Wealthfront. Client experiences will vary. Wealthfront brokerage is not a bank. The base APY is as of January 30th, 2026 and subject to change. For more information, please see the episode description.
我能理解在视觉领域这样做的理由。但类似的事情是否在听觉和声音方面探索过?我的意思是,作为人类,我们非常擅长识别语音语调、情感语气等。但如果要我区分,即使 15 种不同的声音频率,作为非音乐家,我可以告诉你,对我来说非常困难。
I can see the rationale for doing this in the vision domain. But has a similar thing been explored with hearing with sounds? I mean, it's, you know, as humans, we are amazing at recognizing speech inflection, emotional tone, things like that. But if I had to discriminate, you know, even 15 different sound frequencies, I can tell you as a non-musician, it would be very difficult for me.
绝对。我认为你看到的是闸门打开了,AI 的每个子领域,无论是语音识别、声音识别、自然语言处理(这不仅仅是识别)、视觉,所有领域在技术上都得到了真正的推动。我们在斯坦福有同事正在研究鲸鱼声音,对吧,鲸歌,现在使用机器学习和 AI。语音识别是另一个在 AI 革命早期表现优异的领域。当然,技术继续进步。到 2016 年、2017 年左右 Transformer 论文发表时,它很快显示出比早期的 ImageNet AlexNet 算法更强大。不是计算机视觉领域取得了下一个重大进展,而是自然语言处理领域。因为配方没有改变,现在我们有了更强大的神经网络算法,称为 Transformer,但互联网上有更多的数据,至少是更容易获得的文本形式的数据,而且我们现在有更强大的 GPU。所以像 OpenAI 和 Google 这样的公司迅速围绕这一非常重要的技术集结。从 2017 年到 2022 年,仍然花了大约 5 年时间才达到自然语言处理中的 ChatGPT 时刻。但那又是向前迈出的一步。
Absolutely. I think that what you see is the floodgate got open and every sub area of AI, whether it's speech recognition, sound recognition, natural language processing, which is more than recognition, vision, all areas got really a boost in terms of the technology. We have colleagues at Stanford who are studying whale sounds, right, whale songs, using machine learning and AI now. And speech recognition is another area that did so well in the early days of this AI revolution. And of course, the technology continues to advance. By the time the Transformer paper was published around 2016, 2017, it quickly showed that it is even more powerful than the early ImageNet AlexNet algorithm. It was not the field of computer vision that made the next big progress. It's the field of natural language processing. So because the recipe hasn't changed, now we have an even more powerful neural network algorithm called Transformer, but we have even more data on the internet, at least more readily available data on the internet in the form of texts, and now we have more powerful GPUs. So companies like OpenAI and Google quickly rallied beyond this very important technology. And it still took about 5 years from 2017 to 2022 to get to the ChatGPT moment in natural language. But that's yet another step forward.
所以我认为对于不是计算机科学家或神经科学家的人来说,自然的人类体验也许会引起共鸣。也许我可以从这个角度来构建我的问题。当一个孩子学到有一种东西叫“小猫咪”时,他们会说“哦,猫”。然后他们通常会省略“小”字。他们可能会说“小猫咪”,然后学会“猫”。
So I think for people who are not computer scientists nor neuroscientists, the natural human experience will perhaps resonate with them. And maybe I can just frame my question through that lens. So when a child learns that there's something called a kitty cat, they go, "Oh, cat." Then they usually drop the kitty part. They may say kitty and then they learn cat.
如果他们与猫有足够的互动,他们就会意识到猫是什么。即使他们从侧面、从后面看到它,最终如果他们看到一条看起来有点像猫的尾巴,而且你知道,它在一些书后面,你会说“那是什么?”他们很可能会说“猫”。即使他们也见过狐狸和其他有尾巴的动物,仅仅基于他们的经验,他们也在做出概率判断。这本质上就是 AI 能做的。这本质上就是机器学习能做的。
And if they have enough interactions with a cat, they'll realize what a cat is. Even if they see it from the side, from the back, and eventually if they see a tail that looks a little bit like a cat and it's, you know, behind some books, you say, "What is that?" They're very likely to say cat. Even if they've also seen foxes and other animals with tails, just based on their experience, they're making a probability judgment. And that's essentially what AI can do. That's essentially what machine learning can do.
嗯。但在我看来,从计算器到我们现在拥有的 AI 的进步过程中,必须有一个关键时刻,才能看到尾巴的图像并做出合理的假设,即如果它在室内,它很可能是猫,因为狐狸通常不在室内。诸如此类。
Mhm. But it seems to me that there's a key moment that had to happen in the progression of, you know, from calculators to the AI we have now, to be able to see an image of a tail and make the reasonable assumption that it's most likely a cat if it's indoors or something like that because foxes generally aren't indoors. This sort of thing.
那么,机器学习和 AI 是在什么时候获得了这种上下文学习的能力,并能得出某物最可能的归属?因为展示苹果、香蕉和橙子是一回事,它们都是水果。好吧,你能区分它们。你也能把它们和汽车、卡车等区分开。但物体恒常性这一点,如果某物在移动,你只能看到部分图像。
So, at what point did machine learning and AI gain the ability to do kind of contextual learning and come up with the most likely assignment of what something is? Because it's one thing to show apples and bananas and oranges, they're all fruit. Okay, you could distinguish them. You could distinguish those from cars and trucks, etc. But this object constancy piece that if something is moving, you're only getting a partial image.
这不是大多数人对智能的理解,但它是让我们的大脑以及其他动物的大脑,尤其是我们的大脑如此非凡的一部分原因,也是为什么我们认为自己可能是地球上最聪明的物种,如果不是最聪明的,那也肯定是最擅长技术开发的。
This isn't what most people think of in terms of intelligence, but it's part of what makes our brains and the brains of other animals, but especially our brains so remarkable and why we consider ourselves probably the smartest species on earth and if not the smartest and certainly the best at technology development.
是的。
Yeah.
那么 AI 是什么时候实现这一点的,又是如何被编入这些计算机,让它们能够做到的呢?
So when did AI achieve this and how was that scripted into these computers to allow them to do that?
那么让我们就这个问题来说,你描述得非常好,这个问题就是看到猫尾巴的一瞥就能认出是猫,或者有高概率是猫的迹象。有趣的是,几代机器学习的计算机科学家都尝试过这个问题。所以在今天机器能可靠地做到之前,有各种不同的算法,你知道,你可以想象一种常识性的思考方式,哦,也许我们应该识别所有的家具,知道这是室内,所以不太可能是狐狸。所以虽然有类似的规则被内置到前几代算法中,也有这样的规则,比如我们不要猜测它是猫,而是只猜测 10 种潜在动物中的一种,猫是其中之一。这限制了搜索或猜测的空间,这会有所帮助。所以尝试了很多想法。那么它变得可靠得多的时刻是当前这个时代,当这些算法(比如 Gemini 或 GPT)学到的大量数据真正在机器中创造了能力,学到了如此多的知识和模式,以至于当呈现一张猫尾巴从书架后面伸出来的新照片时,那个模式激活了所谓的学习权重或学习参数,使机器对物体的评估或猜测更接近它所见过的东西,那很可能是猫尾巴或只是尾巴,因为数据太多了。
So let's just take the problem very you you have described it so well this problem of seeing a glimpse of a cat tail and being able to recognize cat right or or or a sign of high likelihood there is a cat. The interesting thing is Andrew generations of machine learning computer scientists have tried this problem. So before today that machines can reliably do it there were different algorithm you know you can imagine a common sense way of thinking about this is oh maybe we should recognize all the furniture to know it's a indoor so it's unlikely to be a fox. So though there are rules like that that it was built into uh previous generations of algorithms, there are also rules like well let's only instead of guess it's a cat, let's only guess one out of the 10 potential animals, you know, cat being one of them. That limits the the the the search or guess uh space and that would help. So many ideas were tried. So when was the moment it became much more reliable is this current era when the huge data that these algorithms have learned let's take Gemini or GPD uh have learned really created the capability in the machines uh uh learned space so much knowledge so much pattern that when presented with this more or less maybe a newish photo of a cat's tail sticking outside of a bookshelf. That pattern activated the learned what we call learned weights or learned parameters that put put the machine's um assessment or or guess of this this object closer to what it has seen which is likely to be a cattail or or just tail because there's just so much data.
明白了。
Got it.
这就是我认为作为神经科学家的我们与人类大脑不同的地方,因为那个学习你说的“猫咪”的孩子没有机会下载互联网上的猫图片。他们可能只见过三只猫,最多十只,但他们却能通过不同的学习途径识别出那是猫尾巴而不是狐狸尾巴。这些是我们尚未完全解开的谜团。但我确实想指出,今天用海量数据学习的 AI 算法与人类进化方式之间的差异。
This is where Andrew as neuroscientists I think we depart from human brain because that child who learns about what you say kitty cat would not have the chance to download the internet of images of cat. They likely have seen three cats, 10 cats at most, but yet they're able to identify that tail as a cattail instead of a fox tail through a different kind of learning pathway. These are the mysteries we haven't fully solved. But I I do want to point out that departure between today's AI algorithm that is learned with the humongous amount of data versus how uh humans have evolved.
如果我们继续从简单的物体识别上升到我们所说的高级大脑功能,比如更接近大多数人听到“智能”这个词时想到的,哦,那一定是某种更高级的东西,创造力、想象力。让我们先看一个中间步骤,然后再看一个更远的步骤。继续以猫为例,如果一台计算机或一个孩子通过尾巴学会了识别猫,不管怎样,他们看过猫移动,那么对那个大脑、那个孩子或那台计算机来说,那是一个全新的世界,因为他们现在知道猫通常是朝头的方向移动,而不是尾巴。这些是简单的学习规则,对吧?它可能追老鼠,但也可能躲狗。也许是这样,也许不是,等等。所以似乎所谓的“智能”的下一层是给移动方向、不移动的方向、以及该物体可能与之互动的其他物体分配可能性。这对人们来说听起来很基础,但这就是大脑学习的方式,也是机器学习的方式。那么,下一个大的转折点是什么时候?比如给计算机 AI 一张猫的图片,说“帮我让这只猫动起来,让它像猫一样移动”,而不给它任何关于如何移动四肢的具体指令。但我想那是一个相当快速但非常重要的转变,在我们称之为 AI 的整个事情中,因为那就是大脑所做的。
If we continue to um ascend the kind of hierarchy from simple object recognition to what you and I would call higher order brain functions like moving more towards what most people they hear the word intelligence and they just think oh it must be some higher order thing creativity imagination. Let's go to um a middle step and then a and then a much further step out. So staying with the cat example, if a computer or a child learns to recognize a cat through the tail, the whole thing, whatever, and they've seen a cat move, it's a very new world at that point for that brain, that child or that computer because now they know that the cat generally moves in the direction of its head, not its tail. These are simple simple learning rules, right? It might go after mice, but it might run from dogs. Maybe yes, maybe no, and on and on. And so it seems that the next layer up in terms of quote unquote intelligence is to assign likelihoods of direction to move, directions not to move, other objects that that object is likely to interact with. This all sounds very basic to people, but like this is how brains learn and this is how machines learn. So when was the next sort of big inflection in terms of like giving a computer AI um a picture of a cat and saying um uh animate this cat for me, make it move like a cat without giving it any specific instructions about how to move its limbs etc. But I would imagine that was a pretty quick but a but a remarkably important transformation in this whole thing that we call AI because that's what a brain does.
是的。你问这个问题真的很有趣,你表达得也很美。我从未想过以这种方式向公众阐述,但那个时刻是在视频成为训练数据的一部分时到来的。所以你看,我又要回到训练数据了。大约在 2023 年,ChatGPT 出现后不久,多个研究团队开始将视频纳入训练数据。当然,我不会深入细节,比如算法,有一些小的变化和调整。所以,记得 2024 年 1 月,Sora 发布了,人们看到视频可以生成,正如你所说的。人们可以输入“一只猫跑向一只老鼠”,然后几秒钟的片段就会被生成,猫会以合理的方式移动腿跑向老鼠。那时仍然有错误,即使到今天也不完美,但情况已经好多了,但那打开了视频生成的大门,正如你描述的那样。那么那里发生了什么?实际上并不像你想象的那么革命性,因为底线仍然是数据。作为科学家,我可以告诉你,有各种各样的算法调整、变化和改进等等。但总的来说,如果你退一步看,它仍然是这个伟大的神经网络时代的一部分,对吧?但发生的是,我们现在能够以某种方式处理视频数据,再次通过一些巧妙的工程将其标记化,随便你怎么称呼。现在我们可以生成这些短视频片段,这些片段由帧组成,看起来像合理的猫的运动。
Yeah. So it's it's really funny you asked this and you put it beautifully. I never thought it to put it in this way for a uh public audience but that moment came when video become part of the training data. So see again I'm going back to the training data. So around 2023 very shortly after uh chatbt moment multiple research teams start to put video into the training data. Of course, I'm not going to get into the nuance stuff, the algorithm. There's a little bit of uh changes and variations. So, remember tw January 2024, Sora was released and that's where people see a video can be generated literally what you just said. People can then type and say a cat running towards a mouse and then a a a few second clip would be generated and there would be a cat moving its leg in a plausible way running towards the mouse. At that time there were still mistakes still to even today it's not perfect but things gotten have gotten a lot better but that opened the floodgate of video generation as you described it. So what happened there? What happened there is actually not as revolutionary as you might think because the bottom line is it's still data. As a scientist, I can tell you there are all kinds of algorithm tweaks and changes and improvements and and all that. But overall, if you zoom out, it's still part of this great neuronet network era, right? But what happened is that we're now able to um process video data in a way again some clever engineering tokenize it whatever you call it. And now we can generate these short clips of videos which is frames put together that look like plausible cat movement.
你可能会问,算法是否知道猫腿的肌肉结构,以便当算法展示猫以合理的方式移动爪子时,你知道,按顺序。我会说算法不知道,但它拥有的是那么多视频,尤其是互联网上的猫视频,那么多猫的视频,所以它学会了应该是什么样子。所以在某种程度上,人类也是这样做的。我们大多数没有受过教育的人不会知道猫的肌肉如何运动。我仍然不知道。你知道,我们医学院的同事可能知道,但我们只是习惯了看到猫这样移动,所以我们对猫如何移动有一个合理的想法。所以那是相似的。那就是 AI 的相似之处。它是统计。它是大量的数据,向你展示了什么是猫运动的合理生成。
Now you might ask, does the algorithm know the muscle structure of a cat's legs so that when the algorithm shows that the cat is moving in a plausible way with the paws, you know, in a sequence? I would say the algorithm doesn't, but what it does have is so many videos, especially cat on the internet, so many videos of cat, so it learned what it should look like. So in a way, humans do that. Most of us without education would not know how muscles move in cats. I still don't know. You know, our colleagues in medical school might know, but we have just got so used to seeing cats moving this way that we have a plausible idea of how cats move. So that is similar. That's how similar AI is. It's the statistics. It's the large amount of data that showed you what is the plausible generation of cat movements.
是的。所以当人们几乎肯定听说过大脑是一个预测机器,一个学习机器时,这正是你所指的。
Yeah. So when people have heard almost certainly that the brain is a prediction machine, it's a learning machine, this is exactly what you're referring to.
是的。
Yeah.
让我们去一个非常遥远的大脑功能方面,我们知道人类存在这些,那就是思想和创造力。现在可能有思想和创造力的规则。它们比视觉系统的例子更难确定。比如如果它是尾巴而且在室内,那很可能是猫。这类事情,但它们在那里。规则在那里。
Let's go to a really far out there aspect of brain function that we know exists in humans, which is thoughts and creativity. Now there are probably rules for thoughts and creativity. They're a little bit harder to tack down than examples from the visual system. Like if it's a tail and it's indoors, it's likely a cat. This kind of thing, but they're there. The rules are there.
如果你用苹果作为例子,我们可以从低层次看到苹果,到中层次看到苹果总是掉落,而不是飞走。在最高层次,支配苹果运动的方程是什么?
If you use apple as an example, we could have gone from low-level seeing an apple to mid-level seeing apple always drop, not fly off. At highest level, what is the equation that governs the apple's movement?
对,所以那是上升到更高阶、更还原论的分析。你怎么看这个想法:虽然 AI 确实是智能的,它能做大脑能做的事情,也许甚至肯定能做单个人类大脑不能做的事情,我们知道这一点是因为它在国际象棋上击败了人类等等。现在的想法,据我理解,是 AI 在互联网上训练,图像、讨论、视频、歌曲,但那不是人类认知的全部,对吧?所以,是否有 AI 的方面,无论是 chat 还是 Claude 还是甚至最强大的未发布的机器学习和 AI 工具,还没有访问人类大脑功能的特征,因为它们从未被上传到互联网,至少不是 AI 可以提取的方式?所以,例如,你可以放一首交响乐进去,它遵循某些音乐、数学和声音的规则,这有道理,但你整天有想法,我整天有想法,这些想法不太能与语言契合,以至于我可以直接在互联网上打出来。请听我说完。我知道这是一个长问题,但我觉得这是你完全有能力回答的一件事,而且自从我在犹他州见到你以来,我已经等了一年半才问你。在艺术世界里,我们有所谓抽象的东西,对吧?偶尔有人会画出一幅画或素描,它看起来不像任何具体的东西。这在音乐中也会发生,你只是感觉到某种东西,比如有一个基本的规则或与之相关的情感。就像他们触及了大脑功能的某个方面,但你说不出它是什么。我觉得这是对 AI 或对我来说复杂的事情,难以理解 AI 如何能做到,因为你可以把那件艺术品放进 AI 里,说,你知道,这揭示了人类经验的什么基本特征,而它只能访问互联网上的内容。那么,你怎么能用 AI 捕捉复杂的情感和体验星座?这对我来说似乎是差距。而且我确信我们会用 AI 达到那里,但我没有看到从神经科学到 AI 的任何直接方式。就像我们可以通过视觉运动、悲伤、快乐来逐步推进。你可以提取很多东西,但很难达到这些无法言说、书写或绘制的更高阶抽象表征。如果我只是说,给我你童年家乡怀旧的例子。你可以写下来,但那只是文字。不是,我不能以第一人称理解你的经历。
Right, so that's ascending to like a higher order more reductionist analysis. What do you think about the idea that while AI is indeed intelligent, it can do things that brains can do, maybe even well certainly things that individual human brains can't do, we know this by virtue of beating humans at chess and this sort of thing. The idea right now, as I understand it, is that AI is trained on the internet, images, discussions, videos, songs, but that's not all of human cognition, right? So, are there aspects of AI that are whether or not it's chat or it's Claude or even the most powerful not yet released machine learning and AI tools that don't have access to features of human brain function yet because they've never been uploaded to the internet, at least not in a way that the AI can pull out? So, for instance, you know, you could put a symphony there and it follows certain rules of music and mathematics and sound like that that makes sense, but you have thoughts all day long and I have thoughts all day long that don't quite mesh with language in a way that I can just type them out on the internet. Stay with me here. I know this is a long question, but I feel like this is the one thing you are perfectly poised to answer, and I've been waiting to ask you this for a year and a half since I saw you in Utah. In the world of art, we have this thing called abstraction, right? And occasionally somebody will come up with a painting or a drawing that doesn't look like anything specific. This happens in music, too, where you just feel something like there's like a fundamental rule or an emotion associated with it. Like they've tapped into some aspect of brain function, but you can't say what it is. I feel like this is the sort of thing that is complicated for AI or for me to understand how AI could do because you can put that piece of art into AI and say, you know, what fundamental feature of human experience does this reveal and it only has access to what's on the internet. So, how can you capture a complex constellation of feelings and experience with AI? That seems to be the gap for me. And I'm sure we'll get there with AI, but I'm not seeing from neuroscience to AI in any kind of direct way. The same way we could ratchet through visual motion, sadness, happiness. You could pull out a lot of things, but it's hard to get to these higher order abstract representations that can't be spoken or written down or drawn. If I just say, give me your example of whatever nostalgia for your childhood home. You could write about it, but those are just words. It's not I can't understand your experience at a first person level.
完全同意。Andrew,我知道你对这个问题投入了很多思考,我认为这是一个非常重要的问题。让我们一步一步来。首先,TLDR 简短回答是,我同意你的观点,我们必须非常小心地认识到 AI 能做什么,可能做什么,而不是推测 100 年或什么的。我认识到你刚才说的这些是极其微妙、个性化、难以描述甚至未被捕捉的人类认知行为,因为它们未被捕捉,所以它们没有被上传到互联网,我们今天的 AI 没有办法做到这一点。所以当你称互联网为 AI 数据的来源时,让我们非常清楚什么是互联网。互联网不是随机的东西。互联网是人类行为在多模态形式中的最大集合。让我们进一步分解。互联网上有世界人口在上面打字,已经很多年了,到现在已经几十年了。那种打字是一种感知机制,捕捉了一切,从青少年的闲聊到深奥的科学文章,这些都被数字化并上传了。对吧?所以捕捉人类语言是互联网非常擅长的。然后互联网捕捉图像。怎么捕捉?因为我们现在有数码相机。这在智能手机和数码相机中非常普遍。所以人类喜欢拍照,从你家里的猫到自拍,到美丽的 BBC 捕捉的照片。那些也被上传到我们的数字领域。除此之外,还有视频。视频现在有声音,有动作,也被上传到我们的数字领域。除此之外,还有音乐。我们甚至不进入版权的法律讨论,但让我们把它放在一边。我只是在谈论数据的形式。演讲、歌唱、音乐和管弦乐也被上传到数字领域。所以现在我们创造了这个巨大的人类知识库,以文字形式的人类知识,以视频形式的人类行为,以声音形式的人类表达甚至自然的任何东西,现在 AI 在这些数据上训练。这就是为什么它如此强大。这就是为什么特别是在文字方面,AI 能识别模式,能综合模式,因为这么多已经在那里了。但你刚才谈到的事情,比如说毕加索对表达那位年轻女子肖像的特定方式有极其深刻的想法,那个想法从未被捕捉。事实上,作为神经科学家,如果我问你那个想法来自哪个大脑区域,你不知道,对吧?是布罗卡区?是 V1?是运动区?是前额叶?我们不知道。
Totally. Andrew, I know you put a lot of thoughts into this question and I think it's a very important question. Let's peel this one step at a time. First of all, TLDR short answer is I agree with you that we do have to be very careful recognizing what AI can do, is likely to do, not conjecturing over a 100 years or whatever. I recognize what you just said are these extremely nuanced personalized hard to characterize or not even captured human cognitive behaviors and because they were not captured they then were not uploaded on the internet and we don't have today's AI doesn't have a way to do that. So when you call internet which is the source of AI's data, let's be very clear what is internet. Internet is not some random thing. Internet is the biggest collection of human behavior in multimodal forms. Let's break it down further. Internet has the world's population typing on it for many many at this point multiple decades. That typing is a sensing mechanism that captured everything from teenager chit chat all the way to deep scientific articles which got digitized and get uploaded. Right? So that capturing human language is what internet is super good at. Then internet captures images. How? Because we now have digital cameras. That's so prevalent in smartphones and digital cameras. So that humans love taking photos from, you know, the cat in your house to selfies to beautiful, you know, BBC captured photos. Those also got uploaded in our digital sphere. On top of that, there's videos. Videos now has sound, has movements that also got uploaded to our digital sphere. On top of that, there's music. We're not even getting into the legal discussion of copyrights, but let's just table that aside. I'm just talking about the forms of data. The speeches and singing and music and orchestra that also got uploaded into the digital sphere. So now we have created this humongous library of human knowledge in words, human behavior in videos, human expressions or even nature's whatever in sound and now AI gets trained on that. That is why it's so powerful. This is why especially in the words front that AI can recognize patterns, can synthesize patterns because so much of this is already there. But the thing that you just talked about that when let's say Picasso had that incredibly profound thought about that particular way of expressing that portrait of the young woman, that thought has never been captured. In fact, as neuroscientists, if I ask you which brain area did that thought come from, you don't know, right? Is it Broca? Is it V1? Is it motor? Is it prefrontal? We don't know.
也许它无处不在,因为那种想法是如此个性化、如此特别。你可以称之为创造力,可以称之为情感,或者随便你怎么叫。你可以叫它猫 231,随便什么名字。那个想法没有被捕捉到,因此不在互联网上,因此 AI 没有见过它。所以这就是人类仍然独特的地方。但我们也需要肯定 AI,因为 AI 学到了很多东西,它能以高度创造性的方式组合信息。你还记得第 37 手吗?
Maybe it's diffused everywhere because that thought is so personalized, so special. You can call it creativity, you can call it emotion, you can call it whatever you want. You can call it cat 231 whatever name you can give it. That thought is not captured, therefore it's not on the internet, therefore AI has not seen it. So that is where humans still remain so unique. But we also need to give credit to AI because AI has learned so many things, it can combine information in highly creative ways. Did you remember move 37?
这是 AlphaGo。没错。
This is AlphaGo. Right.
对。
Right.
第 37 手象征着 AI 的创造力。我认为这既是对的,但也可能被断章取义,因为那是在 AlphaGo 对阵李世石的时候,我记得是五局中的第三局,AlphaGo 作为一个计算机算法,下出了一步人类围棋大师从未想过的一手。那确实是不可思议的一手,对吧?因为人类集体,这些大师们,从未想过。但如果你深入看 AI 在那里做了什么,那是因为首先,围棋是一个高度数学化的游戏。它有非常明确的数学目标,非常明确的数学规则。所以当 AI 拥有更大的算力和记忆棋步的方式时,它能做到人类大脑通常做不到的事情。那能叫创造力吗?我认为是,但我们必须认识到那是一种特殊的创造力。
Move 37 has symbolized AI's creativity. I think it's both true but can be taken out of context because that was a game when AlphaGo was plain Lee Sedol, and in I think it's the third game out of the five games that AlphaGo as a computer algorithm made a move that the human masters of Go never thought about. And that is an incredible move, right? Because really, humans collectively, these are the masters, never thought about it. But if you really go deep into what AI did there, it was because first of all, Go is a highly mathematical game. It has very clear mathematical objectives, very clear mathematical rules in terms of moves. So when AI has bigger compute and ways to retain how many moves it can remember, it was able to do things that human brains don't typically do. So is that called creativity? I think it is, but we do have to recognize that's a special kind of creativity.
我和我们这个时代一位杰出的数学家聊过,我问他关于数学未解问题以及 AI 如何能为此做出贡献,他非常乐观。他说今天的数学有很多问题。尽管它们很难,即使作为菲尔兹奖得主,我可能也忘了数学中有已知的方法可以解决这些问题,因为我有一个人脑。我不记得过去几百年所有的数学解法,即使我是菲尔兹奖得主。所以 AI 可以帮助我们解决这些问题。但作为数学家,他也告诉我,他说,我不知道 AI 是否能解决所有数学问题,因为其中一些问题需要尚未被发明的解法,那会把创造力推向一个全新的水平。这就是我们应该好奇的地方:会是人类的创造力,还是 AI 通过迭代改进达到人类没有的创造力,还是两者结合的创造力?我目前的猜想是混合的,人类与 AI 并肩工作将帮助我们解决这些尚未被发明解法的难题。
I was talking to an incredible mathematician of our time, and I was asking him about the unsolved problems of mathematics and how AI can contribute to that, and he was very positive. He said there are many problems in today's mathematics. As hard as they are, even as a field medalist, I probably have forgotten there are known methods in math that can solve these problems because I have a human brain. I don't remember all of math's solutions in the past hundreds of years, even if I were a field medalist. So AI can help us to solve these problems. But as a mathematician, he was also telling me, he said, I don't know if AI can solve all of math problems because some of these math problems require solutions that have not been invented, that will push creativity to a whole different level. And this is where we should be curious: is it going to be human creativity, or AI would go through its iterations of improvement and get to a point of creativity that humans don't have, or is it a combined creativity? My current conjecture is hybrid, that humans working alongside AI would help us to solve these problems whose solutions have yet to be invented.
然后你所说的,特别是你提到情感,那更加个性化。这不一定关乎逻辑,不一定是演绎推理。也许,安德鲁,你看着这个杯子说这是个好杯子。但如果它在我心中唤起一种情感,一个童年时刻,一个灰色杯子可能意味着只有我和我最好的朋友分享的东西呢?那是我大脑中完全无法触及的信息,从未上传到互联网,无论今天的 AI 多么强大,都无法访问。所以我对这个杯子的反应,以及因为那段记忆我可能会用它做什么,可能完全不同。你可以称之为创造力,称之为表达,称之为讲故事,或者用很多方式称呼它。但那是 AI 无法触及的地方。
And then what you said, especially you touched on emotion, is even more personalized. This is not necessarily logic, this is not necessarily deductive reasoning. This is maybe, Andrew, you look at this cup and say it's a great cup. What if it evoked an emotion in me, a childhood moment that a gray cup might mean something that only me and my best friend share? That is a completely inaccessible piece of information in my brain that is never uploaded on the internet, and no matter how mighty AI is today, cannot access that. So my reaction to this cup and potentially what I would do with it because of that piece of memory can be completely different. You can call it creativity, you can call it expression, you can call it storytelling, you can call it in many ways. But that's where AI cannot access.
我觉得在不太遥远的未来某个时候,计算机将能以非侵入性的方式访问我们的大脑活动。
I feel like at some point in the not too distant future, computers will have access to our brain activity in non-invasive ways.
所以你知道,我甚至可以想象在 5 到 10 年后,我现在头上戴着什么东西,你看不见。
So you know, like I might even imagine in 5 or 10 years, I'm wearing something on my head right now, you can't see it.
那是一个非常非常细的电极,听起来像那样,不管怎样,就像一些电极,就在我头骨外面,不打扰我,感知我大脑内部的活动,也许还感知我的心率、自主神经活动、我的警觉程度,并将其与我说的话和做的事进行比较。这一切完全触手可及,而且会实现。你我都知道这一点,而且这可能已经开始让人们感到害怕,但让我们保持善意,好吗?有这样的世界,我拥有的一台电脑,我不担心数据泄露或类似的事情,我们可以管理那个问题,它感知我的所有这些方面,并捕捉到这样一个事实:是的,我说的话可能很重要,但我内在状态和大脑活动中有一些方面是我自己都没有意识到的。
It's a very, very fine hairet, makes it sound like it, whatever, like some electrodes that are just there on the outside of my skull, not bothering me, sensing my activity inside the brain, maybe also sensing my heart rate, autonomic activity, how alert I am, and comparing that to what I'm saying and what I'm doing. This is all totally within reach and it's going to happen. You and I both know this, and it's probably already starting to scare people, but let's keep it benevolent, right? There's this world where a computer that I own, and I'm not worried about data getting out or anything like that, we can manage that problem, is sensing all these aspects of me and is picking up on the fact that yes, what I say might be important, but there are aspects of my internal state and brain activity that I'm not even aware of.
是的。
Yeah.
我可以决定与我这个方面合作,说,让我们想出一幅我从未见过的有趣画面,但它来自我一些重要的经历,基于任何东西,它可能会向我揭示这一点,因为它能访问我大脑活动的无意识特征。我认为这在不太遥远的未来很可能发生。也许如果人们在自身经验的范围内思考它,比如这不会立即传到互联网上,也不会被用来对付他们,你实际上是在了解自己。
And I can decide to collaborate with this aspect of me and say, let's come up with a really interesting picture that I've never seen before, but comes from some experience of mine that's important, based on whatever, and it could reveal that to me because it has access to my unconscious features of my brain activity. I think this is very likely to happen in the not too distant future. And perhaps if people thought about it within the bubble of their own experience, like this isn't immediately going to the internet or it's not going to be used against them, you're actually learning about yourself.
当然。
Of course.
我觉得大多数人对自己的情况有内在的兴趣,也对其他人有兴趣,谢天谢地。但我觉得,这太神奇了。比如我很想知道为什么我在某些方面会出错,过得不好,或者为什么有些天我过得最好,或者我的想法从哪里来。我可以详细说明哪些状态,但我不知道怎么做,除非好吧,一杯咖啡好,一杯半好一点,两杯太多。如果我现在,你想想我们处理这件事有多原始,这有点疯狂。这太疯狂了。每个人都有不同的方法,我们都试图把它弄对,然后当你弄对的时候你已经老了,然后你必须更新它。而且我们现在可能根本没有充分利用我们的生物学和大脑。
And I feel most people have an inherent interest in what's going on for them, also with other people, thank goodness. But they're, I think, like amazing. Like I would love to know why I trip up in certain ways and don't have the best day, or why some days I have the best day, or where ideas come from in me. What states I could, you know, kind of elaborate on, but I'm not going to know how to do that except okay, one cup of coffee good, one and a half a little better, two is too much. If I like right now, if you think about how primitively we go about this, it's kind of crazy. It's crazy. And everyone has a different method and we all try and get this right, and then you've aged enough by the time you get it right that then you have to update it. And like we're probably not getting the most out of our biology and our brains at all right now.
不,我们没有。这就是为什么我一直说,当人们谈论 AI 时,这让我困扰。有些人让它听起来像是在取代人类。但我们真正做的,你所描述的,是关于增强和扩展人类。对吧。这甚至不必像科幻小说那样,用智能电极读取你的脑电波。仅仅让 AI 学习你的写作模式,就已经能帮助你成为一个更好的沟通者,一个更有效的沟通者,一个更高效的沟通者,这是我们在今天的 AI 中可以释放的一种赋能能力。我认为最重要的事情之一,安德鲁,作为神经科学家和教员,我们知道能动性对人类如此重要。你知道,这归结为每个个体层面的动机、能动性和尊严。
No, we're not. And this is why I keep saying, this is why it bothers me when people talk about AI. Some people make it sound like it's replacing humanity. But what we really, what you describe is about enhancing and augmenting humanity. Right. This is where it doesn't even have to go as sci-fi as a smart hairet accessing your brain waves. Just AI learning your patterns of writing can already help you to be a better communicator, a more effective communicator, a more efficient communicator, and that is an empowering capability that we could unleash in today's AI. I think one of the most important things, Andrew, that as a neuroscientist and also faculty we know is agency is so important for humanity. You know that boils down to motivation, agency, and dignity at every individual level.
我认为我们需要认识到,我们必须把 AI 视为一种帮助我们保持自主性的工具。它不应剥夺我们的自主性,而当今 AI 领域的领导者也不应试图说这项技术会剥夺人们的自主性。
And I think we need to recognize that we need to think about AI as a tool that helps us in our agency. It should not take away our agency, and people who lead in today's AI should not try to talk like that this work will take away agency from people.
是的。我认为那些非常熟悉技术的人,无论是计算机、生物学,还是任何技术,比如汽车,他们往往会成为这方面的专家,以至于我们忘记了这对普通人来说可能很可怕,而且围绕它的语言表达至关重要。
Yeah. I think people who are very familiar with the technology, whether it's computers or biology or any technology, cars for that matter, they become such nerds of that thing that we forget that it can be scary to people and that the languaging around it is essential.
我记得在 90 年代初,我相信你也记得,当时基因检测被视为一种东西,比如你想做吗?你想抽血检查吗?因为天哪,你可能会看到一些让你害怕的东西。而现在关于自愿做核磁共振之类的讨论也在发生。这些都不是人们必须做的。
And I remember a time in the early 90s, I'm sure you remember this too, when genetic testing was viewed as this thing like, would you want to have it? Would you want to do a blood test? Because oh my goodness, you might see something that could really scare you. And that discussion is happening now around self-elected MRIs and things like that. None of which people have to do.
但我持的观点是信息越多越好。但我逐渐明白,不是每个人都这么想。有些人不想知道。
But I come from the stance like more information is better. But I've come to understand that not everyone feels that way. Some people don't want to know.
是的。但他们应该有选择权。与此同时,我们应该有足够的公共教育和沟通,让人们了解利弊,但不要剥夺他们的选择权,也不要说,既然你不懂这个,让我替你做决定什么对你好。那是不好的。现在关于 AI 的言论变得非常扭曲,因为懂行的人往往居高临下地对公众说话。无论动机是积极的还是消极的,都有一种论调:你们不知道这是什么,我来告诉你们,我会让你们开心或安全,然后我来替你们做决定。这些都不健康,也没有帮助。
Yeah. But they should have the choice. In the meantime, we should have enough public education and communication to let people know the pros and cons, but not to deny them the choice, and also not to take away, you know, and say, well, since you don't understand this, let me decide for you what's good. That is not good. And the rhetoric around AI right now is getting really skewed because people who know what this is tend to talk down at the public. Whether the motivation is a positive one or negative one, there's a rhetoric of you guys don't know what this is and I will tell you, I will make you happy or safe whatever it is, and I will decide for you. These are not healthy and not helpful.
是的,我同意。我认为,你知道,我们开始这个播客的原因之一就是展示那些真正怀有善意的科学家和医生,他们无意把问题简单化,但确实希望人们理解事物,而很多人会觉得健康信息是最需要理解的重要内容之一。
Yeah, I agree. And I think, you know, one of the reasons for starting this podcast was to showcase the scientists and physicians who really have a benevolence about them and they have no interest in dumbing things down, but they do have an interest in people understanding things, and many people would feel that health information is among the more important things to understand.
绝对如此。
Absolutely.
嗯,谢天谢地,你打破了刚才描述的那种典型形象。还有少数其他人,但你确实在最高层面上这样做,真正鼓励人们思考 AI 带来的协作、存在的自主性,以及是否使用它等等。
Well, thankfully you're breaking the mold of the phenotype you just described. And there are a few others, but you've really been doing this at the highest levels, really encouraging people to think about the collaboration that is AI, the agency that exists, and whether to use it or not to use it, and so forth.
关于自主性,我确实认为对每个人来说,无论你是学生、教师、医生还是政策制定者,重要的是了解这一点,不一定要学会编程。我认为那不一定必要,取决于你的工作。所以比如,如果你是艺术家、教师或医生,你不一定需要编程,但要了解这项技术是什么,了解如何自己使用它来增强自己的能力、学习、工作或表达。通过学习,人们会感到更有掌控力。通过学习,你就不那么害怕尝试。通过学习,你保留了自主性和尊严,因为归根结底,无论技术或医学多么先进,作为人类,我们想要的是那种帮助我们生活得更好、保持尊严、让社区更美好的善意。
One of the agency, I do think it's important for individual humans, whether you're a student, a teacher, doctor, or policy maker, is to learn about this, not necessarily learn about how to code. I don't think that's necessary, depending on your job. So for example, if you're an artist or a teacher or doctor, you don't necessarily need to code, but learn about what this technology is, learn about how you can use it yourself to empower yourself, your learning, your work, or your expression. By learning, one feels more in control. By learning, you're less scared of trying. And by learning, you retain that agency and that dignity, because at the end of the day, no matter how advanced technology or medicine is, as humans, we want that benevolence that helps us to live better, keep our dignity, and make our community better.
我想快速休息一下,感谢我们的赞助商 AG1。我很高兴地宣布,AG1 刚刚推出了他们最新的配方 AG1 Pro。AG1 Pro 采用了有临床支持的 AG1 配方,包含维生素、矿物质、益生菌和适应原,并添加了三种重要的新成分:一水肌酸、HMB 钙和锌肌肽。每份含有 5 克一水肌酸,支持肌肉力量和表现,以及大脑健康。HMB 钙支持肌肉恢复并减少肌肉分解,锌肌肽支持并改善肠道内壁。这三种成分都有令人信服的科学依据,因此我很高兴看到它们被添加到现有的 AG1 配方中。你们大多数人都知道,我每天服用 AG1 已经将近 14 年了。我在知道播客是什么之前就开始服用了。这是一个很棒的产品,现在新的 AG1 Pro 配方让它变得更好。如果你想尝试 AG1 Pro,可以访问 drinkag1.com/huberman 获取特别优惠。AG1 会在你的首次订阅时赠送一瓶免费的 Omega-3 辅酶 Q10。再次强调,访问 drinkag1.com/huberman 即可在首次订阅 AG1 时获得免费的一瓶 Omega-3 辅酶 Q10。今天的节目也由 Element 赞助。Element 是一种电解质饮料,包含你所需的一切,没有你不需要的。这意味着电解质钠、镁和钾,比例正确,但不含糖。适当的水分补充对大脑和身体功能至关重要。即使是轻微脱水也会降低你的认知和身体表现。同样重要的是摄入足够的电解质。电解质钠、镁和钾对你身体所有细胞的功能至关重要,尤其是你的神经元或神经细胞。饮用 Element 可以非常轻松地确保你获得充足的水分和电解质。我的日子通常开始得很快,意味着我必须立刻投入工作或锻炼。所以,为了确保我早上醒来时水分充足且电解质充足,我会喝 16 到 32 盎司的水,并溶解一包 Element。我在进行任何体育锻炼时也会喝溶解在水中的 Element,尤其是在炎热的日子里,我出汗很多,流失水分和电解质。Element 有很多美味的口味。事实上,我全都喜欢。我喜欢西瓜、覆盆子、柑橘,而且我非常喜欢柠檬水口味。所以,如果你想尝试 Element,可以访问 drinkelement.com/huberman 领取免费样品包,任何购买均可。再次强调,访问 drinkelement.com/huberman 领取免费样品包。
I'd like to take a quick break and acknowledge our sponsor, AG1. I'm excited to share that AG1 has just launched their newest formulation, AG1 Pro. AG1 Pro takes the clinically backed AG1 formula, which is a blend of vitamins, minerals, probiotics, and adaptogens, and adds three important new ingredients: creatine monohydrate, calcium HMB, and zinc carnosine. Each serving has 5 grams of creatine monohydrate to support muscle strength and performance, as well as brain health. Calcium HMB to support muscle recovery and reduce muscle breakdown, and zinc carnosine to support and improve the lining of your gut. All three of these ingredients have compelling science to support them, and therefore, I love seeing them added to the existing AG1 formula. As most of you know, I've been taking AG1 every day for nearly 14 years now. I started taking it long before I even knew what a podcast was. It's a great product, and it's now made even better with the new AG1 Pro formula. If you would like to try AG1 Pro, you can go to drinkag1.com/huberman to get a special offer. AG1 is giving away a free bottle of Omega-3 co-enzyme Q10 with your first subscription. Again, go to drinkag1.com/huberman to get a free bottle of omega-3 co-enzyme Q10 with your first AG1 subscription. Today's episode is also brought to us by Element. Element is an electrolyte drink that has everything you need and nothing you don't. That means the electrolytes, sodium, magnesium, and potassium, all in the correct ratios, but no sugar. Proper hydration is critical for brain and body function. Even a slight degree of dehydration can diminish your cognitive and physical performance. It's also important that you get adequate electrolytes. The electrolytes, sodium, magnesium, and potassium are vital for the functioning of all cells in your body, especially your neurons or your nerve cells. Drinking Element makes it very easy to ensure that you're getting adequate hydration and adequate electrolytes. My days tend to start really fast, meaning I have to jump right into work or right into exercise. So, to make sure that I'm hydrated and I have sufficient electrolytes when I first wake up in the morning, I drink 16 to 32 ounces of water with an Element packet dissolved in it. I also drink Element dissolved in water during any kind of physical exercise that I'm doing, especially on hot days when I'm sweating a lot and losing water and electrolytes. Element has a bunch of great tasting flavors. In fact, I love them all. I love the watermelon, the raspberry, the citrus, and I really love the lemonade flavor. So, if you'd like to try Element, you can go to drinkelement.com/huberman to claim a free Element sample pack with any purchase. Again, that's drinkelement.com/huberman to claim a free sample pack.
技术可以成为连接器而非分离器,我认为这必须成为讨论的核心。
The idea that technologies can be connectors as opposed to separators, I think, has to sit at the center of the discussion.
是的。我们都知道他们是谁,有好几位。但这个领域的知名人物,你知道,他们也处于一个发展过程中,正在学习如何面对公众,而且这个过程发生得非常快。比如,你知道,显微镜对着他们,摄像机对着他们,所以每一个微小的缺陷都被放大了。
Yes. And we all know who they are, there are several of them. But the big names in this field, you know, they are also in a developmental process where they're learning how to be public facing, and it happens very fast. Like, you know, the microscope is on them and the cameras are on them, and so every subtle dysfunction is magnified.
所以我愿意相信他们会足够快地成熟起来,意识到这一点,而且我认为他们确实如此,其中一些人确实如此,公众需要听到正确、真实的信息,但要以一种让他们理解的方式呈现。这就是医学和学术界那种不为人知的秘密,你要打破这种模式。我愿意相信我在打破这种模式,因为不分享事物的运作方式确实有一种力量。
So I like to think that they will mature quickly enough to realize that, and I think they are, that some are, that the public needs to hear the correct, the true message, but in a way that makes them understand. That's the kind of dirty secret of medicine and academia that you break this mold. I like to think I break this mold, is that there's a power in not sharing how things work.
但归根结底,这对谁都没有好处。就像你揭开面纱,让人们进来,人们会感到更安全。
But it doesn't serve anybody well at the end of the day. Like you pull back the veil and let people in and people feel safer.
是的。不分享确实有一种力量。但说“相信我,我会告诉你”也有一种力量,而作为教育者,我们两者都不会做。我们不会去课堂上说“相信我,2 加 2 等于 4”。我们实际上会说“这是如何分解并学习它的,这样下次你自己就能做到”。对吧。
Yeah. There's a power in not sharing. There's also a power to say just trust me, I will tell you, and neither as educators, that is, we don't go to our lectures and say just trust me, you know, 2 plus 2 equals 4. We actually say here's how you break it down and learn about it so next time you can do it yourself. Right.
我还认为,尤其是你,你的播客作为公共传播和知识教育的一部分非常重要。我还认为我们需要听到不同背景的声音,对吧?因为有很多学者、技术专家、建设者、思想家一直在处理 AI、使用 AI、深入思考如何用 AI 赋能人们,这些声音非常重要。
I also think that especially you, your podcast is so important as part of public communication, education of knowledge. I also think that we need to hear voices of different backgrounds, right? So because there are plenty of scholars, technologists, builders, thinkers out there who have been dealing with AI, using AI, thinking hard about how to use AI to empower people, and these voices are so important.
嗯,当然,除了你之外,我也会考虑邀请其他人来主持,但既然你在这里,我接下来要谈一个我认为大多数人都会同意的事情,如果它存在的话会非常棒,而且它已经开始发生了,那就是利用 AI 来增强健康发现、疾病治疗等等。
Well, certainly I'll take names of people to host in addition to you, but since you're here, I'm going to go next to something that I think most everybody would agree would be a wonderful thing if it existed and it's already starting to happen, which is the use of AI to augment health discovery, treatment of disease, and so on.
所以,用之前的 AlphaGo 例子,人们肯定还记得猫的例子,那些只是遵循某些规则。AlphaGo 是一套非常复杂的规则,但如果你学会了,它是一套受限的规则。
So, using the AlphaGo example from before, and people surely still remember the cat example, those just follow certain rules. AlphaGo is a very complicated set of rules, but if you learn them, there's a constrained set of rules.
对于猫来说,它似乎不受约束,像无限的可能性,但它足够受限,机器和人类都能学得很好。
With the cat, it seems unconstrained, like infinite possibilities, but it's constrained enough that machines and humans can learn it really well.
当你开始进入医学领域,医学有医学的规则,科学有科学的规则。你有一个问题,提出假设,检验假设,尝试排除假设,等等,就像科学方法一样。在医学中,每个领域都有自己的方法。我们观察,观察疾病,观察谁康复了,我们有病例报告,我们做随机对照试验。所以有规则,而且互联网知道这些规则。因此,LLM 可以很好地用于挖掘健康信息,因为有受限的规则。
When you start getting into medicine, there are rules of medicine. There are rules of science. You have a question, you pose a hypothesis, you test the hypothesis, you try and rule out your hypothesis, and so on, like the scientific method. And in medicine, every field has its methods. We observe, we observe disease, we observe who recovers, we have a case report, we do a randomized control trial. So there are rules and the internet knows these rules. So LLMs can be used to mine health information very well because there are constrained rules.
但我认为你我都知道,因为我也认为你是一位生物学家,生物学的规则仍在向我们揭示。这并不是说皮肤科医生、神经外科医生和肿瘤科医生不知道自己在做什么,但他们是在自己学到的受限规则内行事。即使他们不断学习和更新这些规则,现在似乎每个月都会出现一个违反规则的发现。比如我学到动作电位是单一的,它们看起来总是一样的,要么放电要么不放电。
But I think you and I both know, because I also consider you a biologist, that the rules of biology are still revealing themselves to us. Which is not to say that the dermatologists, neurosurgeons, and oncologists don't know what they're doing, but they're doing what they're doing within a constrained set of rules that they learned. And even if they continue to learn and update them, it's every month it seems now that a discovery comes out that violates the rule. Like I learned that action potentials are unitary. They always look the same. You either fire or not.
但就在 12 年前,有一篇论文表明动作电位的形状可以变化很大。它发表在《自然》杂志上。嗯。
But there was a paper not but 12 years ago that showed that the shape of an action potential can vary quite a lot. It was published in Nature. Mhm.
每个人都看到了,但没人想处理它。这太过分了。它改变了规则。
Everyone saw it and then no one wanted to deal with it. It's just too much. It changes the rule.
神经元应该是分级或全或无的。而“全或无”在每一本教科书里都有。所以现在如果我拿一堆神经活动,给它这个规则,哦,你知道,动作电位可以大,也可以小,在同一个神经元里。这完全混淆了我们对神经科学的一切理解。
Neurons are supposed to be either graded or all or one. And the all, I mean it's in every single textbook. So now if I take a bunch of neural activity and I give it the rule, oh well, you know, action potentials can be big, they can be small in the same neuron. It completely confuses everything we understand about neuroscience.
而我们对大脑的理解就完全崩溃到零了。是的。但如果你给 AI 一个规则,说这个信号可能有一百种不同的形状,那么 AI 可能比最最优秀的研究生做得更多,我敢说,斯坦福,或者公平地说,麻省理工或加州理工。我不认为人类能做到,而 AI 能在这个问题的时长内做到,诚然这个问题有点长。所以,我想听听你的想法,医疗保健中的人类、公众和 AI 如何协作来帮助解决疾病,并理想地提出新的发现规则,以便我们最终能在真正改变人类进程的层面上理解我们的生物学。
And it just, our understanding of the brain just breaks down to zero. Yeah. But if you gave AI the rule that it could be, you know, a hundred different shapes of this signal, well, AI could probably do a lot more than even the very, very best graduate student at, dare I say, Stanford or, to be fair, MIT or Caltech. I don't think it can do it, and it can do it like in the duration of this question, which admittedly is a bit long. So, I'd like to get your thoughts on how is it that humans in health care, the general public, and AI can collaborate to help solve disease and ideally come up with new rules for discovery so that we can finally understand our biology at a level that can really change the course of humanity for the better.
是的。不,安德鲁,这可能是,也许你触及了 AI 最激动人心的用途之一,那就是科学发现。在生物医学领域,你知道,科学发现直接与人类健康和疾病相关。我认为我们已经准备好彻底重写科学发现的方式,因为几个世纪以来,我不知道多久了,它依赖于聪明的人类保留从其他聪明人类那里学到的知识,并以我们自身肌肉的速度做事,我想,你知道。当然,很可能有超级对撞机之类的,但总的来说,做科学发现的方式,人脑或科学家的大脑是这个过程唯一的核心角色。现在我们有了一个新工具,它的大脑可以保留海量信息,可以帮助我们综合知识,可以以你我都无法做到的方式跨学科。例如,我们恰好都在视觉神经科学 AI 领域。我对嗅觉一无所知,比如我甚至不知道我们同事知道的大多数这些词怎么拼写,对吧?所以对我们的大脑来说太难了,但现在我们有了一个可以打破的工具。所以我认为我们需要改变,我们需要使用这个工具。我们绝对,我刚才在想,150 年前,或者我不知道确切时间,电改变了我们生活中的一切,对吧?我确信那是一个我们思考变化、机遇、可怕时刻的时刻。我认为我们必须认识到,科学发现是 AI 和健康领域最激动人心的机会之一,对吧?信息如何被综合,信息如何不仅呈现给临床医生,也呈现给患者,以及患者如何从诊断到治疗参与这个过程,我们现在能做的太多了。
Yeah. No, Andrew, this is probably, perhaps you touch one of the most exciting usage of AI, which is scientific discovery. And in the case of biomedicine, you know, scientific discovery directly connects to human health and diseases. I think we're ready for a complete rewriting of how scientific discovery can be done, because for ages, I don't even know how long, it relies on smart humans retaining what they have learned from other smart humans and doing things at the speed of our own muscles, I guess, you know. Most likely, of course, there's like super colliders and all that, but by and large, the ways of doing scientific discovery, human brain or scientists' brain are the only central character in this process. Now we have a new tool whose brain can retain humongous amount of information, can help us synthesize knowledge, can go across disciplines in ways that you and I cannot go. So for example, we happen to be both in the vision neuroscience AI domain. I know nothing about, you know, olfactory, like I don't even know how to spell most of probably these words that our colleagues know, right? So it's so hard for our brain, but now we have a tool that can break open. So I think that we need to change, we need to use this tool. We absolutely, I was just thinking 150 or I don't know exactly when years ago, electricity changed everything in our life, right? I'm sure that's a moment we were thinking about how the changes, the opportunities, the scary moment. I think we have to come to reckon that scientific discovery is one of the most exciting opportunity for AI and for health, right? How information can be synthesized, how information can be presented not only to clinicians but also to patients, and how patients can participate in that process from diagnosis to treatment, there is just so much we can do now.
是的。
Yeah.
我的意思是,我不会说 AI 比所有医生都好,但几个月前,AI 帮我区分了眩晕和低血压,而搞错的人中有一位是研究前庭系统的耳鼻喉科医生。
I mean AI I won't say AI is better than all doctors but AI was able to disambiguate vertigo from low blood pressure for me a few months back and one of the people who got it wrong is a ENT who works on the vestibular system.
你提供了什么信息?
What information did you provide?
只是一两天的主观感受。
Just my subjective experience over a day or two.
好的,不错。
Okay, good.
嗯,结果是我服用了医生开的一种药,出现了轻微但不良的反应。走路时感觉整个世界都在下坠,然后有点旋转,我觉得天哪,这感觉像眩晕,但我记得头晕和头昏眼花是不同的。所以我开始研究这个,然后果然,是血压问题。它把我的血压降得太低了。
Um, turns out it was a medication that a doctor had prescribed me that I had a like a mild but adverse event and it's a weird thing to step and feel like the whole world's dropping down and then kind of spinning and I thought my goodness like feels like vertigo but I remember dizzy and lightheaded or different. So I started like looking into that and then and um sure enough it was a it was a blood pressure issue. It brought my blood pressure down too low.
但我咨询了一些我们认识的聪明医生,嗯,他们都不在斯坦福。我说的是实话。
And but I consult we know some smart doctors um none of these were at Stanford. I will say that this is the truth.
我们应该保持理智上的诚实。
We should just be intellectually honest.
但这太了不起了。当我把它反馈给他们时,他们说:“这太不可思议了。”你知道,如果你不是和我通电话、在我的诊所里,我本可以做些额外检查,但公平地说,这次是零成本。花了一个早上就知道,如果我喝了电解质,虽然我本以为会过量,但两小时后我就没事了。当然,这里有可能存在安慰剂效应,但两小时后我确实好了。
But it's just remarkable. And when I ran it back to them, they were like, "That's really incredible." You know, had you not been on the phone with me and in my clinic, I would have been able to do some additional testing to be fair. But this was zero cost. It took a morning to know if I drank some electrolytes at what I would have thought would be excessive level that by two hours later, I would be fine. Now, of course, there's the possibility of a placebo effect here, but two hours later, I was fine.
所以,这对病人来说也是一种安慰。这并不是说不要去看医生,但这太不可思议了。我是说,这现在已经存在了。
And so, it's also very consoling to the patient to have this. And so, it's not to say don't go to a doctor, but it's incredible. I mean, this exists now.
医生可以使用这种工具。顺便说一句,我有一个非常有趣的例子。你知道我们不得不重新安排这次对话,因为我父亲当时正在斯坦福接受手术,由一位出色的外科医生主刀。但手术是由机器人完成的,达芬奇机器人系统,因为这是一次肝脏手术,而那位出色的外科医生在操控机器人。所以这是一次深度的人机协作。手术后,我问那位外科医生:“你能想象吗,如果你做过一百万例,这对人类外科医生来说是不可能的,但让我们收集所有人类外科医生对这种肝脏手术的数据,我们能否训练一个自动 AI 来做这个?”答案并不明确。所以我们深入探讨了一下,因为肝脏是一个非常复杂的器官。它血管极其丰富,有很多血管,而且每个人的肝脏都非常不同。所以考虑到每年接受肝脏手术的患者数量,即使你汇总全球的肝脏手术数据,也可能没有足够的数据来训练这些算法。这说明了一个非常重要的事实:AI 从模式中学习。当模式不丰富时,我们就必须小心。我们必须知道如何使用 AI 或如何不使用 AI。你知道,在这种情况下,人类与机器人协作远比一个训练不足的机器人独自做手术要好。但同样的问题可能也适用于外科医生,因为一个外科医生能接受多少手术训练呢。所以这些是人类和 AI 完全可以协作的机会,可能会产生最好的结果。目前,未来还有待观察。我们能否创建一个肝脏的人工模拟,从而训练无限的可能性?这些都是摆在我们面前的极其开放的科学可能性。但也有一些情况,比如你的情况,眩晕与低血压的区分可能已经被报道了很多次,数据库中有足够的数据让 AI 学会了。所以我们现在可以利用这一点,为那些无法立即获得医生帮助的人提供服务。
Doctor can use this tooling. By the way, I have a very interesting example. You know that we have to reschedule this our conversation because my father was going through a surgery right at Stanford with an incredible surgeon. But the surgery was done by a robot, the Davinci robot system because it was a liver surgery and the surgeon, incredible surgeon was driving the robot. So it was a deep human machine collaboration. After the surgery, I asked the surgeon, I said, "Do you imagine if say you've done a million, which is impossible for a surgeon, but human surgeon, but let's collect all of human surgeons for this liver, this type of liver surgery data. Can we possibly train a automatic AI to do this?" The answer was not clear. So we went a little bit down the rabbit hole because liver is a very complicated organ. It's extremely vascular. It has a lot of vessels and everybody's liver is very different. So given the reality of how many patients undergo liver surgery per year, even if you aggregate the world's liver patient surgeries, you might not have enough data to train these algorithm. So this speaks of a very important fact that AI learns from patterns. When the patterns are not abundant, then we have to be careful. We have to know how to use AI or how not to use AI. You know in this case that having a human collaborating with the robot is way better than a underlearned robot doing the surgery by itself. But the same issue might be true for surgeons because how many surgeries a surgeon can get trained on. So these are opportunities that humans and AI can totally collaborate with and might reveal the best result. Right now the future remains to be seen. Can we create a artificial simulation of a liver that we can now train infinite possibility? These are all incredibly open scientific possibilities that is waiting ahead of us. But then there are situations like your situation where the vertigo versus low blood pressure probably have been reported so many times that in the database there's enough of that that AI has learned that. So we can then now take advantage of that for people who don't have immediate access to doctors.
太棒了。你父亲的手术顺利吗?
Amazing. Is your father's surgery went okay?
很顺利。实际上,由于机器人手术的腹腔镜能力,失血量比典型手术少了 10 倍。
It did. It actually lost 10x less blood than a typical surgery thanks to the laparoscopic capability of a robot surgery.
我想谈谈一些我们认为人类独有的特征,这些特征可能存在也可能不存在。你来告诉我。这些是真正的问题,不是带有倾向性的问题。然后我还想了解一下 AI 是如何构建的,才能让这些事情发生。比如,直觉。我们都喜欢把直觉想象成某种神秘的东西,它当然很强大,但这是我们拥有的、没人能夺走的、无法被模仿的东西。但我也可以把直觉分解为:嗯,这是我长期的经验。它是一个数据集,加上一些身体和大脑的感觉,以及一些预测线索,比如上次我有这种感觉时发生了这件事。前两次我有那种感觉时事情并没有那样发展,所以我要走这条路。我的意思是,你可以把这些规则分配给计算机。但我们更深层的自我还有其他方面,如果我可以这样称呼它们的话。比如,我们不知道直觉在身体中的映射位置,你可以做成像实验,但你不可能同时收集所有的神经元、激素和一切。这些东西并没有一个真正的位置,甚至没有一个网络可以指向,比如创造力、直觉、预感,那种你确实感觉到某事即将发生但尚未发生的想法。AI 能获得什么样的规则来赋予它这些能力呢?在这里,我想在能量的背景下讨论它,如果你愿意的话。所以无论这个东西是什么,它就像线粒体驱动细胞更倾向于一件事而不是另一件事,就像恐惧或快乐一样,对吧?我们刚才在谈论能量。但在 AI 系统内部,我不是计算机科学家,在 AI 系统和 GPU 内部,我们能否通过特定的学习规则分配更多的能量流?所以也许有一天我们可以告诉它,基于你所知道的关于我姐姐的一切,我很爱她,你对我们的姐弟关系将如何演变有什么直觉?你觉得我们今年生日做些什么会很好,和以前不同?而且它只能访问互联网,它能否真正变得像心智或身心一样,产生一种什么可能真正有价值的感知?还是它只需要越来越多的提示,就像它会一直问我问题,所以实际上是我在做工作。
I'd like to talk a little bit about some features that we think are uniquely human that may or may not be. You'll tell me. These are genuine questions, not loaded questions. And then I'd also like to get educated on how AI is structured to allow these things to happen. For instance, intuition. We all like to think of intuition as this like mystical very like it certainly is powerful, but this thing that like we own that no one can take from us that can't be mimicked kind of thing. But I could also break intuition down to be well, it's my experience over time. It's a data set coupled to some bodily and brain sensations and some prediction cues like the last time I felt this this happened. The last two times I felt that things didn't work out that way so I'm going to go this ways. I mean that you could assign these rules to a computer. But there are other aspects of our deeper self if I can refer to them that way. Like we don't know where intuition is mapped in the body could do an imaging experiment but you're not going to collect all the neurons and hormones and everything simultaneously. who don't really have like a location or even a network to to point to like things like creativity, intuition, premonition, the idea that you know you really sense something is coming on but it hasn't happened yet. What sorts of rules can AI get that could give it these sorts of capabilities? And here I'm want to talk about it in the context if you will of energy. So whatever this thing is, it's like mitochondria driving cells more around one thing versus another, the same way fear or happiness would, right? We were just talking about energy. But within AI systems, and I'm not a computer scientist, within AI systems and GPUs, can we actually allocate more energetic flow through particular learning rules? So we could tell maybe someday you know based on everything you know about my sister who I love you know what is your intuition about how our brother sister relationship will evolve over time and what is your sense about what would be great for us to do perhaps for our birthdays this year that's different than before giving and it only has access to the internet can it actually become sort of mindlike or mindbody like and come up with a sort of sense of what might actually be worthwhile or does it just need more and more prompts like it's just going to keep asking me questions so I'm actually doing the work.
安德鲁,这个问题真有趣。嗯,为了讨论起见,我想把直觉和创造力分开,也许我们稍后再回到合并的话题。
Such a interesting question Andrew. So um I do want to separate intuition from creativity for the sake of argument here and maybe we'll come back to merging.
那我们聊聊这个直觉的问题,关于“考虑到我兄弟姐妹的爱,会发生什么”,这真的是直觉吗?今天你使用 AI 聊天机器人时,你会提示“我是斯坦福教授,神经科学家”,提供这些信息,这已经被称为上下文。我不知道你是否称之为直觉,但因为你提供了那条信息,AI 给你的答案就已经不同了。如果我输入“我是 14 岁少年,喜欢赛车”,即使问同样的问题,它也会给出定制化的答案。这是数学上的,我不会称之为能量,这只是一个数学事实,关于这些算法如何利用上下文来定制输出,这在计算机科学中并不深奥。那是一种相当浅层的直觉,因为你已经能用语言描述它,或者你可以说“我上传一张图片”,那也是可表达的,AI 能理解。你刚才说的更深层的直觉,就像你甚至不知道它们来自哪里,对吧?是因为我闻到了什么?是荷尔蒙?是情绪混合?还是我的早餐?那种直觉,AI 会怎么处理?我认为那是无法获取的。目前还没有感官设备能捕捉那些数据并输入给 AI,甚至无法输入给人类。比如,作为夫妻,有时你们会摩擦不断。
So let's talk about this intuition of given my sibling love what's going to happen right is it really intuition so today when you go to a AI chatbot you're going to prompt you know I'm a Stanford professor and a um a um neuroscientist um give me this information that is already called context I don't know if you call it intuition but because you gave that piece of information. The AI's answer for you is already going to be different if I type that I'm a 14 year old teenager, you know, loving race cars. Even if we ask the same question, it'll have customized answer. That is a mathematical I wouldn't call it energy. I want to be that is just a mathematical uh fact of how these um these algorithms takes these context and tailor the the the outputs and it's called context. It's not that deep in the in computer science. That's one type of intuition that is fairly shallow because you already are able to use language to describe it or you can say I'll upload an image that that also is is already expressable and then AI gets it. The deeper intuition you just said is like you don't even know where they come from, right? Like is it because I smell something? Is it hormones? Is it you know the the mixture of mood? Is it my breakfast? That intuition, what would AI do with it? That is what I would say is inaccessible. There's no sensory apparatus yet that can glean that data and feed it to not only AI cannot even feed it to, you know, for example, sometimes as a couple you might have moment that you're just rubbing each other in the wrong way.
从来没有。不,我开玩笑的。当然,有。
Never. No, I'm just kidding. Yeah, of course.
如果你们非常熟悉,你多少能感觉到,但说不清楚。也许你只是悄悄离开,让那个人独处。所以这意味着,无论那个人有什么直觉,他们甚至无法用语言或手势表达出来,作为信息传递给另一个人。所以当你无法获取那些高度个性化的直觉时,无论是另一个人还是机器都无法处理,因为没有途径。除非我们戴上脑电波收集器或皮肤电导传感器,也许到那时它们才变得可获取。所以我们必须认识到,我想说的是,关键在于数据是否可获取,无论是通过语言、图片、成像还是脑电波,它必须是可获取的信息。如果可获取,并且我们收集了足够多,就可以训练机器;如果机器训练得好,它就能私下使用,别管隐私泄露,机器可能就能利用它。安德鲁,我想做的不是让它听起来神秘,而是尝试给出一个科学过程,描述如果发生,会如何发生?
If you're really familiar with each other, you kind kind of can sense it, but you can't quite tell. Maybe you just leave quietly, leave that person alone. So that means whatever that intuition that person has, they could not even express it in words or or a gesture to give it to another person to use as a piece of information. So when you cannot even access that neither a human a different human nor a machine can can do anything about it because there's no access to that highly individualized intuition. There's no technology that can do that till you say we put brainwave collectors or you know skin conductance sensors. I mean by the time we do those maybe they become accessible. So we have to recognize. So so what I'm trying to say here is it's not what's not very deep is is the data accessible you know either through language or through picture or through imaging or through brain waves whatever it is it needs to be an accessible piece of information. If it's accessible then if we have collected enough of that you can train machines with or if a machine is well trained it can like you said in a private way forget about privacy uh uh uh breach but in a private way the machine can probably you use it. What I'm trying to do, Andrew, here is not to make it sound mystical, but try to give it a scientific process to describe if it were to happen, how would that happen?
是的。因为我听到的是,基于大数据集和规则的模式识别能让我们走得很远。我们之前谈到医生失败的地方,机器人或机器可能做得更好,或者它们协作比单独任何一方都更好。你知道,作为神经科学家,你在职业生涯中花了很多时间观察细胞。令人惊叹的是,电生理学家几十年来发展出的直觉。我不是生理学家,但我学会了根据那些没有写在论文里的特征来识别细胞。比如,如果某个东西有更直的边缘,有一定的形状和圆度,我现在就能告诉你,那是视网膜中的瞬时 OFF 型细胞。最终,我们开发了基因标记来证明这在每个案例中都是正确的。但你也看到一些不符合规则的。机器可以学习这些,计算机可以学习。结合所有论文的信息,我们现在有了相当完整的视网膜零件清单。
Yeah. Because um pattern recognition based on big data sets and rules get us a long way is what I'm hearing. And we earlier we were talking about where doctors fail and robots and machines perhaps do better or they collaborate to do better than either one alone. You know I as a neuroscientist you spend a lot of time looking at cells at some point in your career. And it's amazing how like the electrophysiologists for decades if not longer you develop an intuition. I'm not really a physiologist, but I learned to recognize cells based on like kind of these things that were not written up in any papers. But like if there was kind of a like a like a straighter edge along this thing and it had a certain shape and roundness, like I tell you right now, that's a transient offpha cell in the retina. Eventually, we we developed genetic labels to reveal that that was true in every case. But then you also saw some that didn't fit the rule. Machines can learn that, computers can learn that. And with all that information from all those papers, now we have a pretty good parts list of the retina.
酷,那行得通。然后你可以应用规则,比如它们这样放电,那样放电。好的,我接受所有这些。我想我试图用直觉表达的是,也许我没给最好的例子,比如人类的一些内部状态,很难想象机器能重现,但也许它们可以,比如动机。机器、机器人会有动机吗?我们有动机的规则。比如当我非常有动力去做某事时,我们称之为紧迫感,一种紧迫状态。我可能会更快行动。更少的激活能量。你说“走吧”,我站起来更快。机器可以朝某个方向更快移动。但你能说“嘿,我希望你寻找这个,但要有更高的紧迫感”吗?还是它们只是受限于数学规则?
Cool. That works. And then you can apply rules like they fire this way, they fire that way. Okay, I'm good with all of that. What I think I was trying to get to with intuition, and I probably didn't give the best example, is like what are some internal states of humans that are really hard to imagine machines could recapitulate, but perhaps they can like motivation. Do machines, do robots get motivated? We have rules of motivation. Like when I'm really motivated to do something, we call that urgency, a state of urgency. And I might move faster to do it. Less activation energy. You say, "Let's go." I stand up a little bit faster. Machines could like go quicker in a certain direction. But can you say, "Hey, I want you to seek this out, but with a heightened level of urgency, or are they just constrained by the mathematical rules they can work with?"
所以你可以把这一点构建到数学中。某些东西,无论你称之为动机,还是在机器学习领域我们称之为目标函数,你可以把某些东西构建到数学中。比如现在你使用 GPT,它有不同模式,比如“深入思考”模式或“快速回答”模式。如果你不知道这是如何工作的,你会觉得“哦,这很有趣”,一个更有紧迫感,给我更快的答案,另一个需要更深入搜索,对吧?并且需要更长时间给出答案。所以,作为人类,如果你过度拟人化,你可能会称之为紧迫感或动机。但事实是,这只是算法的不同目标。你可以说,思考更快的有时间限制或 token 限制,思考更慢的可以激活模型的不同部分,需要更长时间。所以这在数学上变得非常枯燥,并不深奥。但对人类来说,你可以称之为动机或紧迫感。但让我们更深入,因为你问的比那更深,对吧?有些认知状态,无论是动机、紧迫感、恐惧还是爱,对人类来说都很难获取和表达。机器今天有这些吗?没有。让我们明确一点。我们倾向于想象机器有感觉,但它们没有那些数据,没有那些数学目标函数。所以当机器说“你今天生病了,我很抱歉”时,这和你朋友对你说完全不同。因为机器说这句话是因为它通过模式学习到,当有人告诉它“我病了”时,你应该说“你病了,我很抱歉”,而不是“你病了,我很高兴”,因为那些数据存在。而你的朋友听到后,他们真心希望你健康。他们爱你。他们不想看到你受苦。他们有那种共情的感觉,“哦,哇。”
So you could build this in the mathematics. So certain things whether you call it motivation or in machine learning world we call them objective functions you can build certain things into math for example now you go to say JBT it has different mode like think deeper mode or or like give me a quick answer mode if you don't know how this works you're like oh this is interesting one has more urgency that gives me a quicker answer the other one has to go deeper into the search, right? And and take longer to give me the answer. So, as a human, if you anthropomorph anthrop morph morphalize it too much, you might call it urg urgency or motivation. But the truth is this is just a different kind of um objective for the uh algorithm. You can say, well, the the one that think quicker has a time limit or token limit. the one that thinks slower can activate a different part of the model that would take longer. So it become actually mathematically very dry and not that deep. But for a human you can call that motivation or urgency. But let's go deeper because you're asking something deeper than than that, right? Is that there are cognitive states that humans you truly just whether it's motivation or urgency or fear or love that is very hard to access and express. And do machines have it today? No. Let's make it very clear. we tend to imagine that the machines feel or or they're not they don't have that data they don't have that mathematical objective function so they can say when the machine says I'm sorry you're so sick today it's very different from how your friend says it to you because the machine said that because it has learned through pattern when someone tells it I'm sick you should say I'm sorry you're sick instead of I'm so glad you're sick because that data exists. Whereas your friend who hears that, they genuinely want your well-being. They love you. They want they don't want to see you suffer. They have that empathetic feel of, "Oh, wow."
如果你在痛苦中,我经历过痛苦。所以这不是镜像神经元,但至少是对痛苦含义的记忆。机器没有这些。所以我们需要确保区分这一点。很多驱动人类、触发人类的东西在今天的机器中并不存在。我们的运作方式与今天的 AI 根本不同,我们必须认识到这一点、尊重这一点,这就是公共沟通如此重要的原因。我们不能让公众对此感到困惑。
If you're in pain, I've experienced pain. So it's not mirror neurons, but it's at least a memory of what pain means. The machine doesn't have any of that. So we do need to make sure we differentiate that. So a lot of what drives human, what ticks human, what triggers human doesn't exist in today's machine. We operate fundamentally different from today's AI and we have to recognize that, respect that, and this is where public communication is so important. We cannot confuse the public about this.
我想快速休息一下,感谢我们的赞助商之一 David。David 制作与众不同的蛋白棒。他们最新的 Bronze Bar 含有 20 克蛋白质、仅 150 卡路里和 0 克糖。我必须说,这些是我吃过的最好吃的蛋白棒,多年来我尝试过很多蛋白棒。这些新的 David 蛋白棒有棉花糖基底,覆盖着巧克力涂层,绝对美味。我当然也吃常规的天然食品。我吃肉、鸡肉、鱼、鸡蛋、水果、蔬菜等。但我也坚持每天吃一两根 David 蛋白棒作为零食,这样很容易达到每磅体重一克蛋白质的目标。这让我摄入所需的蛋白质而不会摄入过多卡路里。我喜欢所有 David Bronze Bar 口味,包括曲奇面团、焦糖巧克力、双重巧克力、花生酱巧克力。它们实际上尝起来都像糖果棒。再说一次,它们很棒。但同样,它们没有糖,只有 150 卡路里和 20 克蛋白质。如果你想尝试 David,可以访问 davidprotein.com/huberman。现在,David 提供优惠:买四盒,第五盒免费。你也可以在亚马逊或 Target、沃尔玛、Kroger 等商店找到 David。再次提醒,要获得第五盒免费,请访问 davidprotein.com/huberman。
I'd like to take a quick break to acknowledge one of our sponsors, David. David makes protein bars unlike any other. Their newest bar, the Bronze Bar, has 20 grams of protein, only 150 calories, and zero grams of sugar. I have to say, these are the best tasting protein bars I've ever had, and I've tried a lot of protein bars over the years. These new David bars have a marshmallow base, and they're covered in chocolate coating, and they're absolutely incredible. I of course eat regular whole foods. I eat meat, chicken, fish, eggs, fruits, vegetables, etc. But I also make it a point to eat one or two David bars per day as a snack, which makes it easy to hit my protein goal of one gram of protein per pound of body weight. And that allows me to take in the protein I need without consuming excess calories. I love all the David Bronze Bar flavors, including cookie dough, caramel chocolate, double chocolate, peanut butter chocolate. They all actually taste like candy bars. Again, they're amazing. But again, they have no sugar and they have 20 grams of protein with just 150 calories. If you'd like to try David, you can go to davidprotein.com/huberman. Right now, David is offering a deal where if you buy four cartons, you get the fifth carton for free. You can also find David on Amazon or in stores such as Target, Walmart, and Kroger. Again, to get the fifth carton for free, go to davidprotein.com/huberman.
我觉得人们会假设 AI 聊天机器人内部有情感、有人或其他什么,因为我们太以语言为导向了。它在跟我们说话。它在给我写东西。我们现在比 30 年前做得更多。
I feel like people assume there's an emotion, a person or whatever inside of the AI chatbot because we're so language oriented. It's talking to us. It's writing things to me. And we do that more now than we did 30 years ago.
是的。
Yeah.
当然,我们已经非常习惯于接收相当贫乏的语言交流。短信不像长篇散文。语言已经改变。交流方式已经改变。更贫乏而不是更丰富。
Certainly, we've gotten very accustomed to receiving communications in fairly deprived language. Texts are not like extensive prose. Language has changed. Modes of communication have changed. More deprived as opposed to more enriched.
是的。
Yeah.
但很快,我猜面孔会开始进入画面。
But at some point soon, I'm guessing faces are going to start to enter the picture.
无意双关。我们离那有多远?比如如果你或我给对方发短信说“哦,下周这个时候在校园见,喝咖啡。”多久之后,那条短信会变成你对我说话的照片或视频图像?
No pun intended. How far off are we from? Like if you or I were to text the other person, 'Oh, see you on campus for coffee next week at this time.' How soon is it that that text is going to be actually a photo or video like image of you just talking to me telling me that?
我的意思是,这在现在是微不足道的。技术已经存在。
I mean, this would be trivial to do nowadays. The technology is there.
嗯。
Mhm.
但我们现在必须稍微退一步,思考社会参数、法律影响。我的意思是,人类能够用工具做很多事情,但我们不会做所有事情。例如,今天任何汽车制造商都可以说每个星期五刹车不工作。这是一个微不足道的技术。汽车电脑里有一个时钟,它每个星期五就关闭刹车。但我们不这样做,因为它对人类社会的深远不良影响。这就是规则、法律、社会规范、道德介入的地方,我认为这是我们退出纯技术讨论 AI 并需要进入 AI 社会讨论的地方。
But we have to now zoom out a little bit and think about the social parameters, the legal implications. I mean, humans are capable of doing a lot of things with our tools, but we don't do all of them. For example, today any car manufacturer can say every Friday the brake doesn't work. This is a trivial technology. There's a clock in the car's computer and it just turns off the brake every Friday. But we don't do that because it has deeply bad implications to our human society. That's where rules come in, laws come in, social norms come in, morality comes in and I think this is where we exit the pure technical discussion of AI and need to enter the social discussion of AI.
嗯。好吧,让我们这样做,因为我知道生物学家或技术专家的一件事是他们喜欢快速前进,因为它令人兴奋。这是下一个前沿,对吧?我记得很久以前我有一个朋友,他研究病毒和将这些病毒放入的方法,这些不是传染病病毒。这些是用于在动物中表达基因作为实验工具的病毒载体。但后来有机会将狂犬病病毒,一种改良的狂犬病病毒,放入果蝇中。
Mhm. Well, let's do that because one thing that I know about biologists or technologists is they like to go fast because it's exciting. It's the next edge, right? I remember long ago I had a friend he was studying viruses and ways of putting these weren't infectious disease viruses. These were viral vectors for getting genes expressed as experimental tools in animals. But there came the opportunity to actually put the rabies virus, a modified rabies virus into Drosophila, into fruit flies.
天哪。
Oh my god.
现在,在我看来,如果你绝对确定,100% 确定那是一个非功能性的狂犬病病毒版本,那很好,因为你可以把其他货物放在里面,做各种重要的实验,信不信由你,关于疾病等等。但如果有一只果蝇不知怎么逃出来,你得到了真正的狂犬病病毒。它有可能与另一只交配,然后它们最终找到其他果蝇。我不知道这是否会是显性或隐性情况,但现在你有携带狂犬病的果蝇,而且那些东西移动得非常快。所以,你不做那个实验是有原因的。但对他们来说,思考这件事很兴奋,然后他们被拒绝了,对吧?有充分的理由。我很感激,对吧?去任何生物学系,你都会看到一些果蝇飞来飞去。顺便说一句,它们喜欢醋。你知道,所以它们会飞到你的沙拉上。但这里的重点是技术专家喜欢快速前进。他们喜欢感知事物的下一个前沿。那么,在政府、公众、技术专家之间,现在我只是把生物学和医学排除在外。这场对话如何能够以某种方式发生,既能让每个群体都足够满意,又不会阻碍我们?因为我们现在据说也在进行 AI 竞赛。所以这需要更快,而不是更慢。你怎么看?
Now, that's fine and good in my opinion if you are absolutely certain, 100% certainty that that is a nonfunctional version of the rabies virus because you can put other cargo in there and do all sorts of important experiments on, believe it or not, disease and things like that. But if there's just one fruit fly that somehow escapes and you get the actual rabies virus. There's the potential it mates with another and then they eventually find the others. I don't know if this would be a dominant or recessive situation, but now you have fruit flies with rabies and those things move really fast. So, there's a reason why you don't do that experiment. But it was exciting for them to think about and then they got denied, right? For good reason. I was grateful, right? Go to any biology department, you're going to see some fruit flies flying around. They love vinegar, by the way. You know, so they're coming to your salad. But the point here is that technologists love to go fast. They love sensing that next edge of things. So how is it that between government, the general public, technologists, and now I'm just leaving out biology here and medicine. How is it that that conversation can occur in a way that's going to satisfy each of those groups enough, not hold us back? Because we're also supposedly in an AI race right now. So that warrants going faster, not slower. How do you think about this?
我的意思是,Andrew,这就是为什么我八年前从 Google 回到斯坦福,并启动了以人为本 AI 研究所。这些是我们必须面对的深刻社会问题。在 2018 年,还没有 ChatGPT。但作为 AI 科学家,我知道这只会加速。这就是为什么我去找我的同事和大学领导,说让我们建立一个框架,但这不仅仅是我或斯坦福的框架。整个社会在各个方面都需要觉醒到社会影响,正如我们在人类历史上所做的那样,无论是汽车、飞机还是生物技术,它是多维的,涉及多方利益相关者。例如,你们作为生物学家,不会偷偷溜进实验室,试图把狂犬病放入果蝇中,因为这是职业规范和伦理训练。
I mean, Andrew, this is why I returned from Google eight years ago back to Stanford and started the Human-Centered AI Institute. These are profound societal questions we had to face. And back in 2018, there was no ChatGPT. But as an AI scientist, I knew that this is only going to accelerate. This is why I went to my colleagues and university leadership and said let's put a framework but it's not just my framework or Stanford's framework. The entire society in every way need to wake up to the social implication as we have done this in human history whether it was cars or airplanes or biotech is that it's multi-dimensional with multistakeholders right there is the professional norm for example you guys as biologists don't sneak into the lab and try to put rabies into drosophilas or fruit flies because that's a professional norm and your ethical training.
有行业规则,比如 IRB。如今大学校园里的每一项人体实验都受 IRB 监管框架约束,所以我们可以参考这个。还有法律和监管法规,取决于应用对象是人还是农作物,等等。所以 AI 也必须经历同样的过程,对吧?我们需要有职业规范,需要教育。计算机科学家没有接受过伦理和社会研究的训练。你知道,他们现在开始补课了。我的意思是,这就是为什么包括斯坦福在内的许多大学正在拼命地把这部分课程纳入我们的教育。所以那些是规范和教育的部分。但我们也应该与政府合作。不同类型的政府和社会有不同的规范、传统和遗产。并审视监管措施应该应用在哪些地方。AI 例如跨入生物学领域,FDA。我认为这是一个非常重要的领域,要研究 AI 应该如何被用来帮助,同时也要设置护栏防止伤害,这样我们才能避免伤害。我不希望看到的是,一个人或少数几个人从工业界出来,告诉所有人该怎么做。我认为那会很危险,因为市场力量不同于社会规范,文化和遗产不同于教育和伦理。这些是需要共同解决的多方利益相关者问题。
There are industry rules, for example IRBs. Every human subject experiment today on university campuses is subject to the IRB regulatory framework, so we can look at this. And then there are laws and regulatory laws depending on if it's applied to humans versus crops, or you know. So AI has to go through the same. Right? We need to have our professional norms. We need to have education. Computer scientists are not educated in ethics and societal studies. You know, they're starting to. I mean, this is why a number of universities including Stanford are feverishly putting that part of curriculum into our education now. So those are the norms and education. But we also should work with the government. And different kinds of governments and societies have different kinds of norms and traditions and heritage. And look at where the regulatory measures should apply. AI, for example, crossing biology, FDA. I think that's a very important area to look at how AI should be used to help but also guardrail to harm, so that we can avoid harm. What I would not like to see is one person or a few people coming from industry and telling everybody what to do. I think that would be dangerous because market forces are different from societal norms, and culture and heritage are different from education and ethics. And these are multistakeholder problems to solve together.
我喜欢这个回答,而且这个问题非常非常及时。我们对话的这个方面肯定会随着时间推移而扩展,但你一针见血了。我想听听你对人类大脑如何被机器塑造,以及机器如何被我们对人脑的理解所塑造的看法。那么,先问第一个问题。很多人,父母和孩子们都在想,哦,我的孩子现在什么都学不到了。他们只会用聊天机器人查所有东西。但如果你回顾学习的历史,关于计算器、电脑、打字机等等,也有过类似的争论。然而,一个有趣的问题是,我们脑袋里的这个硬件进化出来是为了处理世界上的物理事物——光、声音、气味等等。然后它前面有一块很酷的部分,前额叶皮层,它能学习学习规则,并能更新这些学习规则。所以如果说有什么的话,我们被赋予了一台学习如何学习的机器,并且能更新学习。所以孩子们就是这样调整和使用 LLM 的。所以,我作为伴随着个人电脑长大的一代——当然我在帕洛阿尔托长大——就像这里有 Pong,那里有 Apple 2e,我觉得,哦酷,大脑可以围绕技术成熟,与技术协作,我认为我的生活因此大大丰富了。但我认为智能手机,也许还有摄像头与智能手机的结合,正如乔纳森·海特等人指出的那样,造成了一种局面,大多数人喜欢这些技术是因为它们的便捷和舒适。
I love that answer, and it's something that's very, very timely right now. This aspect of our conversation is surely going to expand over time, but you bullseyed it. I'd like to get your thoughts on how the human brain is being shaped by machines and how machines are being shaped by our understanding of the human brain. So, first question first. Many people, parents and kids are thinking, oh, my kid is never going to learn anything now. They're just going to look everything up on a chatbot. But if you look back in the history of learning, similar arguments were made about calculators, and computers, and the typewriter, and on and on. However, it is an interesting question that this hardware that we have in our heads evolved to process physical things in the world—light, sound, smells, etc. And then it got this really cool piece up front, the prefrontal cortex, that can learn learning rules and can update those learning rules. So if anything, we were gifted with a learning-to-learn machine and updating learning. So that's how kids can adjust and use LLMs. So I, as a generation that grew up with the personal computer—granted I grew up in Palo Alto—it was like here's Pong and there's the Apple 2e, and I think, oh cool, the brain can mature around technology, collaborate with technology in a way that I think my life has been greatly enriched by it. But I think the smartphone, and perhaps the camera-smartphone combination as people like Jonathan Haidt have pointed out, have created a situation where most people love these technologies for the ease and convenience.
但我们现在都多了一点意识,或者说多了很多意识,我们也在放弃一些东西。
But we're all a little bit more aware now, or a lot more aware, that we're giving up something too.
是的。
Yeah.
而且它们是人们,尤其是年轻人,可能会掉进去的陷阱。
And that they're traps that people, in particular young people, can fall down.
是的。那么在你看来,如果确实存在三种情况,对年轻大脑如何被现有 AI 丰富、不受影响或受到伤害,非常乐观和非常悲观的观点是什么?我们就先停留在现有情况。
Yeah. So what is the very optimistic and very pessimistic view in your mind, if three flavors actually exist, of how young brains can be enriched, unaffected, or can be harmed by AI as it exists now? Let's just kind of stay with what we've got.
好问题,安德鲁。答案几乎从我们之前的对话中自然得出,因为你用了“动机”这个词,而我用了“能动性”。最糟糕的结果是,我们年轻一代的能动性和人类层面的学习和生活动机被工具夺走。所以,末日滚动、被动观看短视频,这些都不利于能动性,人类能动性。学习从根本上尊重你所说的硬件,需要时间,需要努力,有时还需要一些痛苦。我们的大脑就是这样。不管晶体管如何移动,我们的神经元以某种方式移动,我们的化学物质、荷尔蒙以某种方式移动。所以对年轻一代来说,无论社会将如何不同,工作将如何不同,我们的身体需要经历一个深度发展阶段,学习必须发生。而这种学习的能动性、学习的动机,不能被任何人夺走,不应该被人类夺走,也不应该被机器夺走。那将是我的担忧,即如果 AI 使用不当,能动性和动机被夺走,那么我们留下的是一代又一代人,他们的大脑没有得到适当发展。另一种危险是,以能动性和动机的名义,我们的学生被剥夺了工具,因为我们担心你作弊,或者担心你只是从 ChatGPT 得到了答案。那也非常糟糕,因为有了适当的能动性、适当的动机、适当的使用工具的方式,我们可以借助 AI 比以往任何时候都学得更深入。我刚才想到,我曾经是医学预科生。天哪,有机化学很难,你知道。我记得试图学习分子、它们的取向,但助教时间太短,或者和其他课冲突,我的教授只有固定的办公时间。学习那个真的很挣扎。对吧?如果今天我有一个人工智能伙伴,我会问很多关于有机化学的问题,因为我知道我卡在哪里,对吧?我有学习的动机。我只需要指导。那将是我学习的一个非常强大的工具。所以我们不应该剥夺学生的这些。所以两件事都让我担心:要么剥夺工具,要么夺走能动性和动机。当然,另一面是好的:让我们找到一种方法,保持我们孩子和学生的动机和能动性。让我们找到一种方法,给他们访问权限和正确使用这些工具的方式。那么这一代、下一代以及许多后代将比我们聪明得多,因为他们被赋予了超能力。
Great question, Andrew. And the answer almost falls out of our previous conversations because you used the word motivation and I was using the word agency. The absolute bad outcome is that our young generation, their agency and human-level motivation of learning and living, is taken away by tools. So doom scrolling, passive watching of shorts, all this are not helping agency, human agency. Learning fundamentally respecting the hardware you're talking about takes time, takes effort, sometimes takes some pain. That is just how our brain is. It doesn't matter how transistors move, our neurons move in certain ways, our chemistry, our hormones move in certain ways. So for young generation, no matter how the society will be different, jobs will be different, our human body needs to go through a deeply developmental phase where learning needs to happen. And that agency of learning, that motivation of learning, cannot be taken away by anybody, should not be taken away by humans, nor should it be taken away by machines. That would be my concern, which is that if AI is not used right, the agency and motivation is taken away, then we are left with generations or generations to come who have not properly developed the brain. The other kind of danger is in the name of agency and motivation, the tools are denied to our students because we're worried you cheat or worry you only got your answer from ChatGPT. That is very bad as well, because with the proper agency, proper motivation, proper ways of using this tool, we can go a lot deeper with AI than we have ever learned. I was just thinking about I was a premed student for a while. Man, organic chemistry was hard, you know. I remembered trying to learn the molecules, their orientations, but the TA hours are too short or it overlaps with my other class, and my professors only have certain number of office hours. It was just a struggle to learn that. Right? If today I were to have an AI companion, I would ask so many questions about organic chemistry because I know where I'm stuck, right? I have the motivation to learn. I just need guidance. That would be such a powerful tool for me to learn. So that we should not deny students from. So both things worry me: either denying the tool or taking away agency and motivation. Of course, the flip side is great: let's find a way to keep our children and students' motivation and agency. Let's find a way to give them the access and the right way of using these tools. Then this generation, this coming generation, and many generations to come will be way smarter than us because they are superpowered.
我喜欢这个回答。我对神经可塑性和年轻一代也很有信心。是的,甚至我们自己的——我知道我们老了,但是——
I love that answer. I have great faith in neuroplasticity and the younger generations too. Yeah, even our own—I know we're old but—
不,让我们给自己一些肯定。可塑性在整个生命周期中都存在。
Not so—let's give ourselves some credit. Plasticity does exist throughout the lifespan.
甚至我们自己的神经可塑性,对吧?我觉得 AI 是我学习的一个很好的工具。
Even our own neuroplasticity, right? Like I find AI a great tool for my learning.
我的意思是,对我来说,它是对它能做什么的一个非凡发现。
I mean, for me it's been a remarkable discovery of what it can do.
但我倾向于从消费者的角度来处理它,如果我对某件事一无所知;如果我对所问的事情有一些知识储备,则从创造者的角度来处理。
But I tend to approach it from the position of consumer if I know nothing about something, and from the position of creator if I have some knowledge set inside of whatever it is I'm asking.
嗯,实际上我还有另一件事,因为我的斯坦福本科生去年教了我一些东西,我意识到在 ChatGPT 之前,有时我会偷懒。如果我有问题,我会问我旁边我认为聪明的人。
Well, I actually have another thing because my Stanford undergrad taught me something last year, and I realized before ChatGPT sometimes I got lazy. If I have a question, I ask the person I think is smart next to me.
现在我意识到我不应该问懒惰的问题,因为在占用别人时间之前,先获取信息要容易得多,而问太懒惰的问题就是在浪费别人的时间。AI 正在迫使我不要过于懒惰。
Now I realize I should not ask lazy questions because it's so much easier to get information before you spend somebody else's time to ask something that's too lazy. And AI is forcing me not to be too lazy.
提示词的精确性对于从 AI 获取最佳信息有多重要?
How essential is the specificity of the prompt to getting the best information out of AI?
提示词非常重要。
Prompting is very important.
那是一项技能,对吧?
And that's a skill, right?
那是一项技能。这就是为什么公共教育如此重要,教育如此重要。我希望看到我们的 K12 学校教授提示词。这里有个小测验:人类最优秀的提示者是谁?
That is a skill. This is why public education is so important. This is why education is so important. I would love to see our schools K12 teaching prompting. Here's a quiz. Who is humanity's best prompter?
我肯定考不及格。
I'm going to flunk this quiz.
苏格拉底,如果他活着的话,因为那就是提示的方法。对吧?想想看,苏格拉底的方法是什么?就是通过提问来提示和寻求真理。我们应该回到那种方式,教孩子们这样做。
Socrates, if he were alive, because that is the method of prompting. Right? Think about it. What is Socrates' method? It's prompting and seeking truth by asking questions. And we should go back and teach kids that.
而且在讨论的时候散散步。
And taking a walk while you have those discussions.
是的。
Yes.
这其实可能是一个很好的过渡,来谈谈具身 AI 的概念。你知道,把说话的面孔和听到的词语联系起来,完全是另一回事。我的一位童年好友,我希望你很快能见到他,因为你们俩都会从对话中受益匪浅,我只想做个旁观者。埃迪·张博士,神经外科主任,生物工程师,他研究言语和语言。他和其他人已经弄清楚了神经活动转化为控制喉部和咽部的过程。他基本上让患有闭锁综合征的人能够重新说话。
Which actually is a good transition perhaps to this notion of embodied AI. You know, it's a world apart to attach a face speaking to hearing words. My good childhood friend, who I hope you'll meet soon because you both would benefit from the conversation so much and I just want to be a fly on the wall. Dr. Eddie Chang, chair of neurosurgery, bioengineer, and he studies speech and language. He and others have figured out the transformation of neural activity to control of the larynx and pharynx. And he's brought people essentially out of locked-in syndrome so they can speak.
哇。
Wow.
十年来第一次,他有一个不幸瘫痪的病人,可以通过电脑说话。他还有其他许多这样的例子。但令人难以置信的是,当他开始把 iPad 放在这个人旁边,特别是那位坐轮椅的女性,他们有她婚礼上的视频。所以他们知道她的声音,知道她的情感模式,也知道她身体的一些动作。现在她通过 iPad 说话,旁边是她僵硬的真实面孔。
For the first time in 10 years, he has this patient who was sadly paralyzed and he could speak through a computer. He has others, many examples of these in fact. But the incredible thing is when he started putting an iPad next to this person, one woman in particular who's wheelchair bound, they had a video of her at her wedding. So they knew her voice. They knew her emotive patterns. They knew a bit about how she moved her body as well. And she now speaks through an iPad next to her frozen real face.
但她可以与世界互动,世界也可以与她互动,这种深度完全不同于仅仅通过麦克风。就像斯蒂芬·霍金那样。
But she can interact with the world and it can interact with her in a completely different level of depth than if it were just a microphone. The sort of Stephen Hawking thing.
而且它通过机器学习不断更新。
And it's constantly being updated through machine learning.
什么,而且现在还注意到她说话的对象以及他们的反应。我的意思是,这就是具身。是的。
What, and now also paying attention to the people she's speaking to and their responses. I mean this is embodiment. Yes.
诚然,它是在二维平面上,但这比单纯的机器人声音,甚至仅仅是准确的声音,都是指数级的飞跃。
It's on a 2D flat screen, admittedly, but this is like an exponential leap over just robot sound or even accurate sound alone.
这不仅仅是人的具身,也是具身 AI 进入机器人领域。对吧?正如我一直说的,AI 的下一个前沿是超越语言的,因为人类最初是在语言之前发展的。进化花了五亿年没有语言交流,而在正确的版本中,世界会因为机器人帮助人类而变得更好。
It's not just embodiment of people, it's also embodiment, embodied AI goes into robotics. Right? The next frontier of AI, as I have been saying, is beyond language because again humans develop first preverbally. Evolution took 500 million years without verbal communication, and the world would, in the right version, be a lot better place with robots helping humans.
你能给我一些例子吗?我喜欢这个想法,但再次,我意识到我可能太深入技术兔子洞了,可能吓到一些人。所以,机器人,我们有自动驾驶汽车。实际上,Waymo 总是为我和我的小狗停下来。我那漂亮的小狗才六个月大。当他想过马路时,你怎么能不停下来?很多人不会停,他们早上差点撞到我们。Waymo 非常尊重人。
Could you give me some examples? I love this idea, but again, I realize I'm probably a little too deep into the technology rabbit hole and it's probably scaring some people. So, robots, we've got self-driving cars. Actually, the Waymo always stops for me and my puppy. My beautiful little six-month old puppy. How could you not stop when he wants to cross the street? A lot of people won't stop. They'll almost run us over in the morning. The Waymo is very respectful.
Waymo 必须学习规则,对吧?
The Waymo has to learn the rules, right?
完全正确。完全正确。所以,那里有一种善意,而这种善意并不总是存在于人类中,但你认为这会首先出现在哪里?如果我们把时间拉长到 12 个月后,它会是什么样子?
Exactly. Exactly. So, there's benevolence there that doesn't always exist in humans, but where do you think this is going to show up first? And what's it going to look like if we zoom out 12 months from now?
12 个月对机器人技术来说有点太快了。两年,三年。
12 months is a little bit too fast for robotics. Two years. Three years.
我想说如果我们把时间拉长到 30 年。
I would say if we zoom out 30 years.
30 年。好的。我不是说那是机器人第一次上街。我们已经有自动驾驶汽车了。我只是说,特别是涉及硬件的技术需要更长时间才能实现。但我希望在你有生之年和我有生之年,能看到机器人成为我们社会的一部分,帮助我们。例如,我是一个独生子女,照顾两位年事已高且病重的父母,而且他们也不会说英语。我做的事情多得惊人,对吧?所以我希望得到帮助。这不会减少家庭责任,不会减少爱,不会减少必要的沟通。但体力劳动,某些部分我很想得到帮助。我们住在加利福尼亚州。我们都经历的一件事是什么?
30 years. Okay. I'm not saying that's the first time robots hit the street. We already have robotic cars. I'm just saying it takes longer for especially hardware-involved technology to manifest. But I would say hopefully in you and my lifetime, I would love to see robots being part of our society, helping us. For example, I'm a single grown-up child taking care of two very advanced aged and very sick parents and they happen not to speak English either. The amount of work I do is incredible, right? So I would love to have help. It doesn't take away family's responsibility. It doesn't take away love. It doesn't take away the necessary communication. But the physical labor would really, certain part I would love to get help. We live in the state of California. What is the one thing we all experience?
交通。
Traffic.
高税收。加州某些地方没有交通问题,但有野火。
High taxes. Certain part of California doesn't have traffic but wildfires.
哦,哇。是的。
Oh wow. Yes.
对吧?谁在扑灭这些野火?让人类冒着生命危险去救援自然灾害并不是一个好主意。对吧?所以我的家人和父母碰巧有足够的财力。但我只是在想,独居老人。他们怎么去买菜?怎么去拿药?现在,我们开始看到一些配送服务,但如果他们想去散步或去公园呢?所以,有太多事情了。哦,顺便说一句,你在医学院。我们没有过剩的护理人员,我们缺少护理人员。我们的护士非常疲惫,过度工作。过去一个月我一直在医院陪我父亲,看着护士们做的大量工作。我们知道在一个班次里,护士要走好几英里取东西、拿药。有太多这样的例子。你能想象机器人帮忙吗?对吧?所以我们的社会可以通过很多方式构建,并从帮助中受益。
Right. Who is fighting these wildfires? Putting humans in danger to rescue natural disaster is not a great idea. Right? So my family and my parents happen to have enough means. But I was just thinking, elderly living alone. How do they go get groceries? How do they go get medicine? Now, there might be some shipping we are starting to see, but what if they want to go for a walk or want to go to a park? So, there are just so many things that, oh, by the way, you're in the school of medicine. We don't have an excess of caretakers. We have a shortage of caretakers. Our nurses are deeply fatigued and overworked. I was literally in the hospital with my dad for the past month and just watching the amount of work nurses do. We know that on a given shift nurses walk miles to fetch things, get medicine. There's just so many. Can you imagine robots helping, right? Like so there are just so many ways that our society can be structured and can benefit from help.
哦,我喜欢这些例子。根据你描述的,我想到了很多。你知道,交通协管员。是的。你可以想象,通过视频,一个因为年龄或疾病而困在家里的人可以导航到商店,从货架上取东西。不一定非得是完全断开连接,他们只是编程然后它回来。那也可能是一个选择。
Oh, I love these examples. So many spring to mind based on what you described. You know, crossing guards. Yeah. You imagine with video that somebody who's homebound because of age or illness could navigate to the store and pick things off the shelf. It doesn't have to be so disconnected that they just program and it comes back. That could be an option, too.
是的,我认为我们必须修正我们对这幅图景的看法,因为我认为关于机器人和计算机,有两件事让人们害怕。一是它们的物理硬度,对吧?所以我们与它们共享空间的方式,与我们与其他事物共享空间的方式非常不同。当然,我不是在想,哦,你会和机器人拥抱,虽然有些人可能会这么想。那不是我的想法。但我在想,好吧,如果我有一个机器人,可以叠衣服、吸尘、浇花、喂我的鱼。虽然我喜欢自己喂鱼,我真的很享受,我喜欢看它们吃东西。
Yeah, I think that we have to revise our notions of what this picture looks like because I think there are a couple things about robots and computers that scare people. One is their physical hardness, right? And so the way we share space with them is very different than the way we share space with other things. Of course, I'm not thinking, oh, like you cuddle with a robot, although some people might think that. That's not my mindset. But I am thinking like, okay, if I had a robot that could fold clothes, vacuum, water the plants, and feed my fish. Although I like to feed my fish myself. I really enjoy it. I love seeing them eat.
我喜欢触觉上的接触,你知道,就是真正地触摸它们。它们会从我手里吃东西。
I love being tactile, you know, in literally in touch with them. They'll eat from my hand.
你的小狗喜欢你的鱼吗?
Does your puppy like your fish?
嗯,他喜欢。他有自己的鱼缸。我刚给他买了一些热带鱼。对,就在他的小……
Uh, he does. He has his own fish tank. I just got him some tropical fish. Yeah. Right in front of his little...
他在照顾它们。
He's taking care of them.
嗯,他只是看着它们。我觉得他还没有能力照顾它们。不幸的是,他的前额叶皮层不够发达。而且他是一只斗牛犬混种,虽然很善良,但不是最聪明的品种。
Well, he looks at them. He's not equipped to take care of them yet, I don't think. Unfortunately, there's not enough prefrontal cortex in him. So, and he's a kind, but he's a bulldog mut. They're not the smartest breed.
它们只有很少的学习规则,但它们非常善良。
They only have a few learning rules, but they're very kind.
嗯,但如果你想要一只能照顾鱼缸的狗,你可能需要像西部高地梗之类的。那不一样,前额叶皮层更发达。
Um, but if you want a dog that can take care of a fish tank, you probably need like a West Highland Terrier or something like that. Different, more prefrontal cortex.
带他去。
Take him.
但这里的想法是……
But the idea here is...
如果一个机器人做一件事,另一个机器人做另一件事,感觉我生活中到处都是硬件。我觉得人们大概就是这样感觉的。
If one robot is doing one thing and another robot is doing another, it feels like a lot of hardware in my life. And I think that's kind of how people feel.
但你可以想象一个多形态机器人。你知道大白吗?
But you could imagine a multimorphic robot. Do you know Baymax?
我不知道。
I don't.
迪士尼的机器人,大概 10 年前、15 年前吧。你可以谷歌一下图片。这是一个白色的医疗机器人。一个医疗保健机器人,不是很蓬松,而是非常海绵感,感觉像一个大气球。
Disney's robot probably 10 years ago, 15 years ago. It's a Google the image. This is the white medicine robot. A healthcare robot that is very not fluffy. It's very spongy like it's feels like a big balloon.
嗯。对。所以,你可能会喜欢那个。对,更多曲线。
Mhm. Yeah. So, you might like that. Yeah. More contours.
对。对。
Yeah. Yeah.
而且,嗯,同一个机器人能做更多任务。感觉这个世界我适应起来比满世界都是机器人的想法要快得多。
And um more multitasking from the same robot. Feels like a world that I could adjust to more quickly than the idea of my world filled with robots.
是的。安德鲁,我觉得当我们想象未来、讨论我们如何想象未来时,我不断回到“自主权”这个词。人类应该有自主权来决定我们如何想象这个未来。不能只是一家公司或某个投资者决定世界应该充满金属机器人,对吧?我们的社会应该集体地、主动地去想象。我在这套 AI 话语中担心的一点是,公众被置于被动反应的位置。
Yes. Again, Andrew, I think as we imagine the future and we talk about how we imagine the future, I keep coming back to the word agency. Humanity should have the agency to decide how we imagine this. It cannot just be a company or, I don't know, an investor decide that the world should be filled with metal-like robots, right? Like our society should be collectively proactively imagining and one thing I worry in this AI rhetoric is that the public is put in a position of being reactive.
当感觉有些人只是在决定……
When it feels some people are just deciding...
而多方利益相关者并没有参与共同设计未来。
And the multistakeholders are not participating in this designing the future together.
就像你父亲手术的例子,把机器人和医生结合起来,对吧?如果我们把脆弱的问题和明显能改善情况的机器人结合起来,情况就会朝着正确的方向改变。所以,我想到几个例子,比如,我想大多数人都会同意,如果孩子能自己走路上学和回家,那很好,但你会担心安全。但如果一个机器人真的能很好地守护你的孩子,甚至能报警或实际保护你的孩子,那就太棒了。给他们更多在世界中的自主权。
Like with your example of your father's surgery to cross the robot with the physician, right? If we cross a problem where there's a vulnerability with a robot that clearly makes things better, the picture changes in the right direction. So, I'm thinking of a few examples off the top of my head like um I think most people would agree that if their kids could walk themselves to school and home, it would be great, but you worry about safety. But if a robot was really a good guardian of your kid to the point where they could alert the authorities or maybe even physically protect your child, that would be awesome. Give them more agency in the world.
你想想,嗯,一些更阴暗但不幸真实存在的网络掠夺行为。父母只能监督孩子的行为到一定程度。孩子们能意识到的也有限。但你可以想象,嗯,一个化身在你身边,真正为你辩护,能发现危险并让掠夺者远离。
You think about um some of the darker but nonetheless unfortunately real predatory behavior online. Parents can only oversee their kids' behavior so much. Kids are only aware of so much that's happening. But you could imagine um kind of an avatar in there with you that's really advocating for you that can spot things and keep predators at bay.
就是这样。这是个很棒的创业点子。
Here you go. That's a great startup idea.
那会很酷。
Like that would be cool.
但我觉得画面中缺少的是,对我来说。我记得见过一个了不起的人。我知道有些人说他有点难相处,但那个了不起的人在帕洛阿尔托市中心散步,那时我是博士后,还是孩子时在帕洛阿尔托体育世界工作,那就是史蒂夫·乔布斯。不穿鞋,看起来像个嬉皮士。是的,他在工作中对人喊叫,你知道,现在 HR 可能不会对他太友好,但他明白我们称之为计算机的东西需要有圆润的边缘。
But here's what's missing I think from the picture for me. I remember seeing this incredible guy. I know people some say he was kind of prickly but this incredible guy walking around downtown Palo Alto when I was a postdoc and when I was a kid growing up working at the Palo Alto Sport and World and that was Steve Jobs. No shoes, kind of look like a hippie. Yes, he shouted at people at work, and you know, probably HR wouldn't look too kindly upon him nowadays, but he understood that these things we call computers needed to have rounded edges.
是的。
Yes.
它们需要无缝地放进我们的口袋。它们需要在首页放鲍勃·迪伦之类的,这样能软化人与技术的关系。有些人会说,好吧,这太过分了。这是特洛伊木马。但我不这么认为。真正理解人性的人才能让这些显然是善意的机器人与人类合作发生,因为正如你指出的,并且完全尊重那些构建 AI 的技术人员和做惊人科学的科学家,无论是他们呈现的方式还是他们能分享的内容,都有一种生硬感,这真的是一个隔阂。
They needed to fit kind of seamlessly in our pocket. They needed to have Bob Dylan on the landing page or whatever so that it softened the relationship to technology. Some people would say, well, it went too far. It was a Trojan horse. But I don't think so. Somebody who really understands human nature to allow these like what are clearly going to be benevolent collaborations between robots and humans to happen because as you've pointed out and with total respect to the technologists that have built AI and the scientists that do amazing science, there's a hardness to either the way they're being presented or what they're capable of sharing that is a real separator.
是的。我不是治疗师,但如果我能拥抱他们,我会说:“听着,各位,你们是房间里最聪明的人,先生们女士们。公平地说,你们是房间里最聪明的人。”但人们不喜欢你们,因为他们不理解你们,也许你们需要一个合作者来帮助你们分享愿景,以一种不会让媒体有机可乘的方式,因为媒体在制造这种鸿沟上是有罪的,就像这些技术人员,他们要来对付我们。我认为那完全是媒体的把戏。那只是为了赚钱。现在有很多事情在发生。所以谁是史蒂夫·乔布斯或斯泰西什么的,我的意思是可能是男人也可能是女人。一个真正理解人类的人。
Yes. And I'm not a therapist, but if I could like wrap my arms around him, I'd be like, "Listen, guys, you're the smartest people in the room, guys and gals. To be fair, you're the smartest people in the room." But people don't like you because they don't understand you and they're maybe you need a collaborator to help you share your vision in a way that isn't going to allow the press because the media is guilty of building this chasm because it's like these technologists, they're coming for us. I think that's a total trick of media, too. That's just to put money in their pocket. Like there's a lot going on right now. So who's the Steve Jobs or the Stacy whoever it's I mean could be a man could be a woman. Someone who really understands human beings.
他们中有很多,我们中有很多,你知道,斯坦福成立了以人为本 AI 研究院。嗯,有你。有你。
Many of them there are many of us you know I mean Stanford started Human-Centered AI Institute. Well, there's you. There's you.
好吧。但有很多。
Okay. But there are many.
对。
Yeah.
有很多企业家在做 AI 药物发现、AI 医疗、AI 老龄化、AI 心理健康的惊人创业。这些人关心 AI,对吧?有很多设计师和产品经理在努力……我确实觉得扩音器太集中在那些拍胸脯、以某种特定方式谈论技术的人身上。所以,你知道,即使是这个播客,我希望也在产生积极影响,就是把人的角度、圆润的人的角度、人的视角、人的未来带入这些对话。我不感到绝望,安德鲁。我是教育者。我是建设者。我是技术专家。我看到身边很多人,包括我整个创业公司。
There are plenty of entrepreneurs who are doing incredible startups on AI for drug discovery, AI for health care, AI for aging, AI for mental health. These people care about AI, right? There are many designers and product managers who are trying to... I do think the megaphone is too much focused on people pumping their chest and talking about tech in a certain particular way. So, you know, even this podcast is making a positive difference, I hope, is to put that human angle, the rounded human angle, human perspective, human future into these conversations. I don't feel despair, Andrew. I'm an educator. I'm a builder. I'm a technologist. I see many people around me, including my entire startup.
这些才华横溢的年轻技术专家本可以加入任何他们想去的初创公司或企业,但他们来到 World Labs,是因为他们想赋能他人。所以我看到了很多人,但我觉得还不够。你说得对,我认为当前的公众讨论并不平衡。有太多极端言论,要么是极端的末日论和缺乏安全感,简直把人吓坏;要么是极端的乌托邦主义,仿佛技术不会出错。那是不真诚的。人们会说,好吧,你是乐观派,你当然这么说。所以我认为我们应该回到中间地带,讨论这项技术是什么、如何使用它,以及我们如何共同拥有引导未来的主动权。
These brilliant young technologists could join any startup or company they want, but they come to World Labs because they want to empower people. So I see many people, but I don't think there's enough. You're right, I don't think the public discourse is balanced right now. There is too much extreme rhetoric, either in terms of extreme doomerism and lack of safety, like it's just freaking people out, or extreme utopian as if technology can do no wrong. And that's disingenuous. People would say, well, okay, you're the optimist, of course you say that. So I think we should come to the middle and talk about what this technology is, how to use it, how we can collectively have that agency to guide the future.
是的。我的一位播客同事向我指出了一点,这本该显而易见但并非如此,而且这显然是你身上体现的众多特质之一,那就是人们其实并不想听关于机器的故事。但人们喜欢听到某个人治好了他们狗的癌症,或者他们的孩子出现了疯狂的症状。他们毫无头绪。医生也毫无头绪,而他们指尖的 AI 解决了问题。
Yeah. One thing that was pointed out to me by one of my podcast colleagues, which should have been obvious but wasn't, and clearly this is something that you embody among many other things, is people don't really want to hear stories about machines. But people love hearing that some person cured their dog's cancer or their child that was experiencing crazy symptoms. They had no clue. The doctors had no clue, and their fingertips AI solved the problem.
这些才是真正需要放大的故事,因为我觉得我们能与之产生共鸣,而且它们是美丽的故事。它们是不可思议的故事,但它们得到的关注远不及其他那些东西。
These are the stories that really need amplification because I think that we can relate to them, and they're beautiful stories. They're incredible stories, but they're not getting nearly as much attention as the other stuff.
是的。那是一个挑战。
Yeah. And that's a challenge.
是的。
Yeah.
你知道,传统媒体并不真正关心事物的长期发展轨迹。它们处于 12 到 24 小时的周期中。但也许还有其他一些人,他们真正试图谈论 AI 的善意用途,那些我们应该知道的合作?
You know, traditional media doesn't really care about the long arc of things. They are on a 12 to 24 hour cycle. But other names perhaps of like people who are really trying to talk about the benevolent use of AI, these collaborations that we should be aware of?
Stephan 在 HI 的通讯、我们的网站、我们的研讨会上,我们推广了很多这样的工作。
Stephan for HI's newsletter, our website, our seminars, we promote a lot of those work.
我很想了解更多关于你的初创公司的信息,因为你不会随意选择项目。那么这个项目是什么?目标是什么?
I would love to learn more about your startup, because you don't pick projects haphazardly. So what is the project? What's the goal?
我的初创公司,与另外几位联合创始人共同创立,名为 World Labs。我们在 2024 年初共同创立了它。对我来说,这真的是我毕生的事业。你知道,我们都来自视觉领域,并认识到语言之外还有更多。智能才是真正激励我深入思考 AI 前沿下一章的动力。我们认识到,解锁空间智能和物理智能确实是下一章。当然,这并不排除语言。语言技术是不可思议的。正是在这里,我们可以投入更多时间来构建模型,或最终构建产品,以帮助解锁空间智能的能力,比如生成对创作者、机器人训练、建筑设计或实现交互式环境非常有用的 3D、4D 世界。无论你谈论的是医疗应用、教育应用、机器人应用还是工业应用,这些能力都超越了语言本身。因此,World Labs 就是基于这一前提创立的。我们仍然是一家年轻的公司。我们非常专注于模型,正在构建这个基础模型,我们最初由许多博士组成,但现在我们开始构建产品。所以这还只是开始。这非常令人兴奋,作为一名技术专家,我内心深处觉得自己是一个建设者。你知道,也许还因为我是一名移民,所以卷起袖子,和那些聪明绝顶的年轻一代一起从零开始构建东西,真是太令人兴奋了。
So my startup, co-founded with a couple of other co-founders, is called World Labs. We co-founded it at the beginning of 2024. It really is, for me, a kind of my life's work. You know, we both come from vision and the recognition that there's more beyond language. Intelligence is what really motivated me to think hard about what's the next chapter of AI frontier. And we recognize that unlocking spatial and physical intelligence is really the next chapter. It's not excluding languages, of course. The language technology is incredible. It's where we can devote more time to build models or eventually build products that can help unlock capabilities in spatial intelligence, like generating 3D, 4D worlds that are deeply useful for creators, for robot training, for architecture design, or to enable those interactive environments. Whether you're talking about healthcare usage, education usage, robotics usage, or industry usage, these capabilities go beyond language per se. And so World Labs was founded based on that premise. We are still a young company. We're very much a model-focused company where we're building this foundation model, and we started with a lot of PhDs, but now we're starting to build products. So it's still the beginning. It's very exciting, and as a technologist, I feel deep in my heart I'm a builder. You know, it's maybe because I'm also an immigrant, so that rolling your sleeves up and just get in with the young generation that's so incredibly smart and just build something from scratch is just so exciting.
我记得大约 15、20 年前,有汽车四处行驶拍摄图像。现在仍然在行驶拍摄图像。但我想肯定也有空中视角,但你可以想象那种能飞过神经元的小型无人机,几乎能观察一切,或者无人机收集挪威峡湾每个角落的信息。这是否已经用于绘制三维世界的地图?
I recall a time, not but 15, 20 years ago, when there were cars driving around taking images. Still still driving around taking images. But I imagine that there are certainly aerial views as well, but you could imagine little tiny drones like the type that could fly through a neuron and just kind of look at everything, or drones picking up information about every nook and cranny of the fjords in Norway. Has that been done to sort of map the three-dimensional world?
首先,我们不要把无人机进入人们的家和财产说得那么可怕。我认为捕捉世界图像的能力确实在飞速发展,对吧?比如我们的手机是不可思议的传感器。它们不是无人机,但人们拍很多照片。当然,我们的相机技术也在进步。World Labs 所做的不仅仅是捕捉真实世界的图像。我们让人们能够想象他们脑海中的景象。只要你能输入一个句子,或展示一张图片或草图,表达你的想象,我们就试图将其转化为世界和环境。为什么这有用?因为娱乐产业会使用它,设计产业会使用它,机器人产业非常需要它来训练环境等等。因此,捕捉真实世界和捕捉你想象世界的结合,就是新的前沿。
First of all, let's not make it sound scary that drones are getting into people's homes and properties. I think that the ability to capture imageries of the world is really rapidly advanced, right? Like our cell phones are incredible sensors. They're not drones, but people take a lot of photos. And of course, our camera technology has improved. What World Labs is doing is not just taking real world images. It's we allow people to imagine what's in their mind's eye. As long as you could type a sentence or show a picture or a sketch of what you imagine, we try to turn that into worlds and environments. Why is it useful? Because entertainment industry would use it, design industry would use it, robotics industry very much would use it for training environments and all that. So the combination of capturing what's in the real world as well as capturing what's in your imagined world is the new frontier.
如果你不介意,我想再花几分钟谈谈从想象到实际的过程。因为这里是洛杉矶,我想到很多人写剧本,然后试图让他们的电影拍成。但理论上,有了 AI,你可以把剧本交给 AI,它理论上就能制作电影,对吧?从文字到图片再到视频。当然,你可能需要在需要帮助的地方稍作编辑。这已经实现了吗?有没有一部成功的电影从头到尾完全使用 AI 制作?
If you don't mind, I'd like to just take a couple of more minutes and talk about this moving from imagination to something. Because this is Los Angeles, it occurred to me that a lot of people write scripts and then they try to get their movie made. But with AI, in theory, you could take a script and give it to AI and it could make the movie in theory, right? Going from words to pictures to video. And you could maybe edit it a little bit here and there where it needed help, of course. Has that been done? Has a successful movie been made start to finish using AI?
所以,这是一个非常微妙的话题。这也是我们触及人们对 AI 和创造力的担忧的地方,如果不小心,听起来可能像我们在剥夺讲故事者和创作者的工作。对吧?所以让我们把工作讨论和技术讨论稍微分开一点,尽管它们是纠缠在一起的。技术已经足够先进,将剧本转化为镜头、视频镜头已经变得非常好。我们已经看到短片,甚至接近长片的电影由 AI 工具组装而成。我们确实看到了,而且有许多公司,美国公司、亚洲公司,在创造技术。但真正深刻人性化且重要的是讲故事和故事创作的每一个部分。背后有独特情感、故事、技巧、世界观、镜头运动、角色塑造的人类。很多这些都是好莱坞和小说作家的核心。
So, this is a very nuanced topic. This is where we also get into people's weariness of AI and creativity, if not careful it might sound like we're taking away from storytellers and creators' jobs. Right? So let's separate this job conversation from the technology conversation a little bit, even though they're entangled. Technology has advanced enough that taking scripts and generating shots, video shots, is getting really good. We have seen short movies, even almost feature-length films being assembled by AI tools. We have, and there are many companies, US companies, Asian companies, creating technology. But what remains deeply human and that is important is every part of storytelling and story creation. There are humans behind it with their unique emotion, story, technique, how they see the world, how they move the cameras, how they characterize characters. A lot of that is what Hollywood and novel writers is about.
那么,如何用现代工具满足人类对讲故事的需求和渴望,这其实是一个挑战,因为好莱坞非常担心 AI 会取代讲故事的人、演员和编剧,他们的工作正受到影响。我认为确实如此,但影响的方式是什么?我们对此做了什么?谁在以一种建设性的方式努力?你知道,这本身不是我的行业,但我希望看到这方面更细致的工作,以及更细致的公众讨论。
So how do we meet the human need and human desire of storytelling with modern tools is actually a challenge because there is a fear very much coming from Hollywood that AI is taking over and storytellers and actors and screenwriters the jobs are being impacted and I think it is but how is it being impacted, what are we doing about it? Who is working in a constructive way? You know, this is not my industry per se, but I would love to see much more nuanced work in this and also nuanced public discussion about that.
但我确实认为,就像医疗保健一样,我们刚才谈到 AI 如何能迅速改变和颠覆传统的医疗方式。我认为 AI 绝对在改变我们讲故事的方式。说到故事,我有一个联合创始人叫 Ben,我和 Ben 见过本·阿弗莱克。所以我开玩笑说 Ben 见 Ben,他也在非常前卫地思考用 AI 工具拍电影。所以,在这个时刻,让技术专家和讲故事的人或电影制作人之间进行对话至关重要。
But I do think just like healthcare, we were talking about how AI can rapidly change and disrupt the old ways of doing healthcare. I think AI is absolutely changing the way we're doing storytelling. So one story speaking of which I have a co-founder whose name is Ben and Ben and I met with Ben Affleck. So I was joking Ben meeting Ben who is also thinking very avant-garde about using AI tools for filmmaking right. So having conversations between technologists and storytellers or movie makers at this moment is critical.
是的。我觉得在每个技术的例子里,都有一个交叉点,当真正懂行的人拥抱一项技术,然后它就腾飞了,比如斯皮尔伯格之类的,或者这些可能不是最好的例子,但就像史蒂夫·乔布斯和沃兹尼亚克的交叉,一个设计师型、对技术好奇的人,和一个真正的计算机科学家(请原谅我提到乔布斯电影)。对吧?这种融合,这些合作真的很关键,所以你需要一个内行和一个外行来做这件事,因为你要理解两种文化,以及如何把行业和人们包括进来。
Yeah. I feel like in every example of technology, there's some crossover point that when somebody who's truly an insider embraces a technology and then it just kind of takes off like you know uh Steven Spielberg or something like that or these probably aren't the best examples but like the Steve Jobs Wozniak crossover kind of a designer technology curious guy and a real forgive me to the Jobs film but a real computer scientist. Right? That merge these collaborations are really key so you need an insider and an outsider to do it right because you have to understand both cultures and how to include the industry the people.
所以我真的希望,因为 World Labs 也和视觉特效行业合作,对我来说,让我们的客户和用户感到被赋能非常重要。技术不应该夺走他们的工作,技术应该让他们的工作更好,为他们的创造力提供超能力。这就是我对这项技术的看法,也是我希望与用户和客户合作的方式。
So I really hope right cuz World Labs works with VFX industry as well it's so important for me that our customers and users feel empowered it's not that technology should be taking their jobs away technology should be making their jobs better superpowering their creativity And that's how I see this technology and that's how I would like to work with the users and customers.
想想真是疯狂,你知道,我小时候在帕洛阿尔托的加利福尼亚大道上,有一家店叫 Keeble and Shuchat,那只是一家照相馆和相机店。是的。你走进去,冲洗胶卷,柜台后面有那些人,他们告诉你你可以租长焦镜头之类的。嗯,这些都不存在了。或者一切都数字化了,你知道,但仍然有相机店。所以,行业可以演变,它们不总是被消灭。
It's wild to think that, you know, when I was a kid on California Avenue in Palo Alto, there was this store, Keeble and Shuchat, and it was just a photograph store and camera store. Yes. You go in there, you got your film developed, and there were all these guys behind the counter, and they tell you all you could rent a long-distance lens and this kind of thing. Um, none of that exists anymore. Or everything went digital, you know, but there are still camera stores. So, industries can morph. They don't always get obliterated.
是的,它在演变。人们也会重新学习技能、提升技能。你知道,我们和很多创作者合作,他们使用 AI 工具,因为他们看到了技术的发展方向,他们想重新学习和提升自己。所以,我认为变革的时刻既是机遇也是失去的时刻。我们需要真正深思熟虑。
Yeah, it morphs. People also get reskilled, upskilled. You know, we are working with a lot of creators who are using AI tools because they see where technology is going and they want to reskill and upskill themselves. So, I think moments of change is moment of both opportunity and loss. We need to really be thoughtful about that.
我的最后一个问题是关于年轻一代的。他们对 AI 有什么感觉?因为有一种……
My last question is about the young generation. How do they feel about AI? Because there is this
你说的年轻是指多大?
How young are you talking about?
我说的是 7 到 20 岁的孩子。
I'm talking about kids between the age of uh seven and 20.
好吧,那正好是我的孩子。
Okay, that's literally my kids.
是的。所以,我可能问这个问题是有原因的。你知道,他们对此感觉如何?他们兴奋吗?因为有一种现象,比如电脑出现后,你的书法老师开始紧张,因为人们不只是打字了。现在他们都用指尖写字,没人会写字了,我们写……这些故事已经流传很久了,说如果我们不像拥抱未来那样拥抱过去,我们就会溶解在自己神经元的池塘里。我愿意相信两者兼顾才是重要的。但孩子们感觉如何?他们怎么想?
Yeah. So, I might have asked that question for a reason. You know, how do they feel about it? Are they excited by it? Because there is this phenomenon where like computers come along and you know your handwriting teacher is getting nervous that people aren't just typing. Now they're all writing with their fingertips and no one's going to know how to write and we wrote for there's these stories have been around for a long time about how we're just going to dissolve into a puddle of our own neurons if we don't embrace the past as much as the future. And I like to think some of both is what's important. But how do the kids feel? What do they think?
这其实是我作为教育者和技术专家的一个特别关注的项目。我每到一处,都尽量和学生、家长、老师交谈,因为我认为他们是最被遗忘的人群。我们的政策制定者、技术专家和投资者,他们不谈论老师、家长和学生。他们都有自己的观点,也都有孩子,但他们不谈论这个。我对孩子总是抱有希望。也许因为我是教育者,因为我认为人类最大的问题之一就是老一代总是哀叹下一代,好像下一代什么都不懂。他们粗鲁,他们忘记过去。但如果你看人类历史的轨迹,总的来说,我们在进步。我不是否认暴行,不是否认挫折,不是否认这些,但你知道,人类,我根本上是个乐观主义者,对吧?所以这就是我的出发点。如果你是个彻底的悲观主义者,也许我们已经站在错误的立足点上了。但你看孩子们,他们好奇,这就是为什么他们是孩子,他们好奇,他们当然会被这项技术极大地娱乐,但他们也开始使用它。我担心的是我们的老师和一些家长,因为我认为我们当今社会,尤其是硅谷,没有为他们服务。我们忘记了他们,我们教训他们,我们斥责他们,我们看不起他们。他们是我们社会中最重要的人。我们应该和他们交谈,我们应该提升他们,我们应该支持他们,我们应该为他们提供资源。K12 或 K16 的老师,他们承担着我们社会最重要、最关键的负担。
This is actually my pet project as an educator and technologist. Everywhere I go, I try to talk to students, parents, and teachers because I think that is the most forgotten population, our policy makers and our technologist and our investors. They don't talk about teachers, parents, and students. They all have opinions and they all have kids, but they don't talk about it. I always have hope for kids. Maybe because I'm an educator because I think the biggest thing humanity never learns is the older generation lamenting about the future generation as if the future generation doesn't know anything. They're rude. They're forgetting the past. But if you look at arc of history of humanity, by and large, we advance for the better. Now, I'm not denying the atrocities. I'm not denying the setbacks. I'm not denying this but you know humanity there fundamentally I'm a optimist in humanity right so so that's where I come from so if you're a total pessimist maybe we're already on the wrong footing but I look at kids they're curious that's why they're kids they're curious they of course they get massively entertained by this technology but they also are starting to use it what I worry about are our teachers and some parents because I think our society today and especially Silicon Valley are not doing them a service. We're forgetting about them. We are lecturing them. We are berating them. We are looking down at them. They are the most important people in our society. We should be talking to them. We should be uplifting them. We should be supporting them. We should be providing resources to them. K12 teacher or K16 teachers, they share the most important critical burden of our society.
我给你讲个真实的故事。2022 年 11 月,ChatGPT 发布了。显然,我在技术方面是内行,但我做的第一件事是给我孩子所在小学的校长发邮件,说:“我想来给您的学生和老师们做一次客座讲座。”这不是因为我有多特别,而是因为我想让他们实时了解正在发生的事情。因为没有人,硅谷没有人,没有投资者,没有数十亿美元的投资公司,也没有数万亿市值的公司。当 ChatGPT 发布时,首先想到的是我们附近的老师怎么样?没有人那样想。但我们需要和老师交谈,我们需要给老师展示。当然,他们会问孩子作弊怎么办。没关系,他们问那些问题。让我们展示给他们看,让我们和他们合作,赋予他们力量,让他们想出应对的方法。他们也很聪明,他们渴望改变,他们只是被遗忘了。所以,我对孩子抱有希望,但为了不是盲目的希望,我认为我们都应该记住我们的老师,帮助我们的老师和家长,这样我们才能帮助我们的孩子。
I'll tell you a real story. November 2022, ChatGPT came out. Obviously, I'm an insider in terms of technology, but the first thing I did was emailing the principal of the elementary school my kid was in and said, "I would like to come and guest lecture for your students and teachers." It's not because I'm so special. It's because I want them in real time to know what's happening. Because nobody, nobody in Silicon Valley, no investors, multi-billion dollar investment firms or multi-million dollar, multi-trillion dollar companies. When ChatGPT came out, the first thing is what about our teachers in the neighborhood? Nobody think like that. But we need to be talking to teachers. We need to show teachers. Of course, they're going to ask the question about what if kids cheat. It's okay. They ask those questions. Let's just show them. Let's work with them and empower them to come up with ways to deal with that. They are smart, too. They are eager to change. They're just forgotten. So, I have hope for kids, but in order not to have a blind hope. I think we should all remember our teachers and help our teachers and parents so that we can help our kids.
我非常喜欢这个回答,我知道很多听众也有同感。
I absolutely love that answer and I know that sentiment is shared by many many people listening.
上帝保佑老师们,他们需要帮助、支持和信息,因为现在他们打开的不是你的播客,而是大多数播客时,他们只是害怕。他们非常害怕,听到这些末日论。他们听到末日预言者,或者他们说:‘哦,别担心,这是乌托邦。’这两种信息都无法帮助我们的老师,如果老师得不到帮助,我们的孩子就得不到帮助。
God bless the teachers and they need help, support and information because now they turn on not yours but most podcasts they're just scared. They're so scared they hear these doomerism. They hear the doomsayers or they say, 'Oh, don't worry. It's utopian.' Neither of these messages can help our teachers and if they're not helped, our kids are not helped.
完全同意。
Couldn't agree more.
是的,完全同意。
Yeah, couldn't agree more.
李飞飞,非常感谢你在百忙之中抽出时间。我很高兴听到你父亲没事。这也是你日程的一部分,照顾父母、孩子,以及其他所有事情,同时来教育我们这件事,它不仅重要,而且是我们所处位置和未来方向的一个关键楔子。基于你今天分享的内容,我更加持有谨慎的乐观态度。也感谢你一路教我们更多神经科学知识,因为这些机器受大脑启发,大脑也受这些机器影响,这就是我们生活的世界。我非常乐观,很大程度上是因为你存在于这个世界,感谢你抽出时间来分享。我知道很多人非常感激。所以,谢谢你。
Fei-Fei, thank you so much for taking the time out of your incredibly busy schedule. I'm so glad to hear your father's okay. And that is also part of your schedule, taking care of your parents, kids, and all the rest to come educate us on this thing that's not just important, it's a major wedge of where we're at and where we're headed. And I share great optimism with caution even more so on the basis of what you shared today. And also thank you for teaching us more neuroscience as we went along because these machines are informed by the brain and the brain is informed by these machines and this is the world we're living in and I have great optimism in no small part thanks to the fact that you exist in this world and thank you for taking the time to come here to share. I know many people are very grateful. So thank you.
谢谢。安德鲁和我非常感激这次对话。这是一个文明时刻。
Thank you. Andrew and I really appreciated this conversation. It's a civilizational moment.
感谢您今天收听我与李飞飞博士的讨论。想了解更多关于她的工作,请查看节目说明中的链接。如果您从本播客中学习或喜欢本播客,请订阅我们的 YouTube 频道。这是支持我们的绝佳零成本方式。此外,请在 Spotify 和 Apple 上点击关注按钮来关注本播客。在 Spotify 和 Apple 上,您可以给我们留下最高五星的评价。您现在还可以在 Spotify 和 Apple 上给我们留言。也请查看今天节目开头和整个过程中提到的赞助商。这是支持本播客的最佳方式。如果您对我有疑问,或对播客、嘉宾或希望我在 Huberman Lab 播客中考虑的主题有评论,请在 YouTube 的评论区留言。我会阅读所有评论。对于那些还没听说的人,我有一本新书即将出版。这是我的第一本书,名为《Protocols: An Operating Manual for the Human Body》。这本书我写了五年多,基于 30 多年的研究和经验。它涵盖了从睡眠到锻炼、压力控制、专注力和动机相关的各种方案。当然,我为书中包含的方案提供了科学依据。这本书现在可以在 protocolsbook.com 上预售。那里有各种供应商的链接,您可以选择最喜欢的一个。再次强调,这本书叫《Protocols: An Operating Manual for the Human Body》。如果您还没有在社交媒体上关注我,我在所有社交媒体平台上都是 Huberman Lab。包括 Instagram、X、Threads、Facebook 和 LinkedIn。在这些平台上,我讨论科学和科学相关工具,其中一些与 Huberman Lab 播客的内容重叠,但大部分与 Huberman Lab 播客上的信息不同。再次强调,在所有社交媒体平台上都是 Huberman Lab。如果您还没有订阅我们的神经网络通讯,神经网络通讯是一份零成本的月度通讯,包含播客摘要以及我们称之为方案的一到三页 PDF,涵盖从如何优化睡眠、如何优化多巴胺、刻意冷暴露等一切内容。我们有一个基础健身方案,涵盖心血管训练和抗阻训练。所有这些都完全免费。您只需访问 hubermanlab.com,点击右上角的菜单标签,向下滚动到通讯,然后输入您的电子邮件。我要强调,我们不会与任何人分享您的电子邮件。再次感谢您收听今天与李飞飞博士的讨论。最后但同样重要的是,感谢您对科学的兴趣。
Thank you for joining me for today's discussion with Dr. Fei-Fei Li. To learn more about her work, please see the links in the show note caption. If you're learning from and or enjoying this podcast, please subscribe to our YouTube channel. That's a terrific zero-cost way to support us. In addition, please follow the podcast by clicking the follow button on both Spotify and Apple. And on both Spotify and Apple, you can leave us up to a five-star review. And you can now leave us comments at both Spotify and Apple. Please also check out the sponsors mentioned at the beginning and throughout today's episode. That's the best way to support this podcast. If you have questions for me or comments about the podcast or guests or topics that you'd like me to consider for the Huberman Lab podcast, please put those in the comment section on YouTube. I do read all the comments. For those of you that haven't heard, I have a new book coming out. It's my very first book. It's entitled Protocols: An Operating Manual for the Human Body. This is a book that I've been working on for more than five years and that's based on more than 30 years of research and experience. And it covers protocols for everything from sleep to exercise to stress control protocols related to focus and motivation. And of course, I provide the scientific substantiation for the protocols that are included. The book is now available by pre-sale at protocolsbook.com. There you can find links to various vendors. You can pick the one that you like best. Again, the book is called Protocols, an operating manual for the human body. And if you're not already following me on social media, I am Huberman Lab on all social media platforms. So that's Instagram, X, Threads, Facebook, and LinkedIn. And on all those platforms, I discuss science and science related tools, some of which overlaps with the content of the Huberman Lab podcast, but much of which is distinct from the information on the Huberman Lab podcast. Again, it's Huberman Lab on all social media platforms. And if you haven't already subscribed to our neural network newsletter, the neural network newsletter is a zero-cost monthly newsletter that includes podcast summaries as well as what we call protocols in the form of one to three-page PDFs that cover everything from how to optimize your sleep, how to optimize dopamine, deliberate cold exposure. We have a foundational fitness protocol that covers cardiovascular training and resistance training. All of that is available completely zero cost. You simply go to hubermanlab.com, go to the menu tab in the top right corner, scroll down to newsletter, and enter your email. And I should emphasize that we do not share your email with anybody. Thank you once again for joining me for today's discussion with Dr. Fei-Fei Li. And last, but certainly not least, thank you for your interest in science.