AI Hype Hurts Young Innovators
打开互动全文版(中英对照 + 朗读 + 问答)→一位著名计算机科学家认为,围绕 AI 和 AGI 的危言耸听和过度乐观叙事使年轻人士气低落,分散了他们对机器学习与系统构建真正机会的注意力。
A prominent computer scientist argues that alarmist and exuberant narratives around AI and AGI demoralize young people, distracting from real opportunities in machine learning and systems building.
迈克尔·乔丹教授,非常荣幸能邀请您来到 MLST,尤其是《自然》杂志不久前称您是最有影响力的计算机科学家。
Professor Michael Jordan, it's such an honor to have you on MLST, especially given that nature said that you were the most influential computer scientist a little while back.
有趣的是,我受训成为统计学家和认知科学家,但我接受这个称号。
It's funny because I was trained as a statistician and a cognitive scientist, but I'll take it.
太棒了。那么,迈克尔,您刚刚发表了一篇论文,题为《人工智能的集体主义经济学视角》。请给我们一个简要介绍。
Amazing stuff. Well, Michael, you've just published a paper called a collectivist economic perspective on AI. Give us the elevator pitch.
我从来不是搞 AI 的人。所以,在某种程度上,我很容易介入并审视那些自称 AI 研究者的人,然后问:你们在做什么?你们的重点是什么?你们的目标是什么?可悲的是,我认为他们往往没有非常明确的目标。
I was never an AI person. So, in some ways, it's easy for me to come in and look at people who are self-professed AI researchers and sort of say, what are you doing? What is your point? What's your goal? I think sadly they often don't have a very clear goal.
《自然》杂志说你是最有影响力的计算机科学家。它存在于现实世界中。这是一个抽象概念,但它是真实的东西。就像 F=MA。它是一组预测。所以,如果我像在某个坐标系中写下 F=MA 一样写下游戏,我现在就能预测会发生什么。我认为我们不需要这种对智能和理解的拟人化,这既不必要也不恰当,而且对许多问题来说是一种干扰。为什么要说它理解?我认为这是科幻小说,科幻小说对社会很重要,但以目前被推广的程度和那些声音,它确实伤害了 25 岁和 20 岁的年轻人。你知道,这些年轻人数量庞大,他们对科技感到兴奋,想要建造东西来帮助他们的家庭和国家——实际上更多是家庭而非国家。他们看到了真正的机会,但领导者却告诉他们:我们玩够了,开发了一堆算法。我们做到了,我们只对纯粹的、理解智能感兴趣,尽管他们并不理解智能。他们构建了梯度下降算法,现在你们不能做这个,因为它很危险。它有很大概率会毁灭人类,或者超级智能很快就会到来,所以没什么可做的了。这发生在你们有生之年。这太令人沮丧了。太令人沮丧了,我认为这是最困扰我的事情。我的意思是,第二个困扰我的问题是,那里没有经济思维。所以,当前这一代只是……没有太多思考,没有太多智力上的东西。只是,是的,可以构建它。可以从任何地方窃取数据,因为互联网允许这样做,并且不向数据原始创造者返还任何价值。可以在此基础上运行贪心下降,但需要大量资金。但现在可以从那些没有深入思考的人那里获得资金。我不认为构建你不理解的系统是坏事。但我认为这种与现实脱节的程度在人类历史上是不寻常的。
Nature said that you are the most influential computer scientist. It exists in the real world. This is an abstraction, but it's a real thing. It's like F=MA. It's a set of predictions. So if I write down a game just like I wrote down F=MA in some coordinate system, I can now predict what'll happen. I don't think we need to see this anthropomorphizing of intelligence and understanding all that is not necessary, not appropriate and is a distraction for many many problems. Why say it understands? I think it's science fiction and I think science fiction is important for society but it's also at the level it's being promoted and those kind of voices it's really hurting 25 and 20 year olds. You know these young folks of whom there are huge numbers are excited about technology and they want to build things that help their family and help their country actually more their family than their country honestly. And they see real opportunities in doing that and they're kind of being told by the leaders, well, we had our fun and we developed a bunch of algorithms. We did it and we were just interested in the pure, you know, understand intelligence even though they didn't understand intelligence. They built gradient descent algorithms and now you guys, you can't do this because it's dangerous. It's going to wipe out humanity with a high probability or superintelligence will arrive soon so there's nothing left to do. That's in your lifetime. That is so demoralizing. So demoralizing and that thing I think that bothers me the most. I mean the second part that bothers me is there's no economic thinking going on there. So the current generation is just way too you know there's not much thought going on not much intellectual stuff. It's just yeah it's possible to build it. It's possible to steal the data from wherever you want to because that's what the internet allowed to happen and not return any value to the person who originated the data. It's possible to run greedy descent on that, but you need huge amounts of money. But it's now possible to get it from people who aren't thinking very deeply. I don't think it's bad to build systems you don't understand. But I think this level of detachment from reality is unusual for human history.
顺便问一下,您对 AGI 这个词怎么看?
And what do you think about the term AGI, by the way?
AGI 对我来说只是一个公关术语。有些人认为它很有趣,因为你必须有这些伟大的抱负。我认为它只是扭曲。我认为它让年轻人困惑。正如我今天会稍微谈到的那样,我认为那些经常在播客和其他场合出现的所谓思想领袖最令人担忧的一点是他们的危言耸听或过度乐观的语气。我认为 20 岁和 25 岁的年轻人看到这些,会想:我是要乐观还是危言耸听?只有这两个选择。我希望我们即将进行的这次对话能让年轻人明白,还有其他的方式来看待生活和科技。
AGI to me is just a bit of a PR term. And some people think it's fun because you have to have these great aspirations. I think it's just distortion. I think it confuses young people. And as I will talk about today a little bit, I think that one of the things I find most alarming about the so-called thought leaders that one will see often on podcasts and other venues is the alarmist tone or the exuberant tone. And I think 20 and 25 year olds are watching that and saying am I going to be exuberant or I'm going to be alarmist? Those are the two choices. And I hope that this conversation we're about to have is one that makes it clear to young people that there is other ways to approach life and in technology.
我从未真正将自己视为 AI 研究者。我没读过 AI 方面的书。这个术语是 50 年代创造的,约翰·麦卡锡等人创造它时有特定的目标。他们也有特定的方法,比如逻辑推理等,但这些并没有真正成功。与此同时,在 60、70、80 年代,出现了机器学习。实际的方法,如决策树、最近邻、逻辑回归和隐马尔可夫模型,是在其他文献中发展起来的,主要是统计学、运筹学等,这带来了工业上的成功故事。因此,供应链、商业和交通系统都使用了大量的机器学习,直到今天仍然如此。它们使用了基于梯度的方法,事实上,云是为了处理亚马逊的机器学习工作负载而开发的。这就是我成长的传统。我试图思考大规模系统构建,这些系统也能服务于多人。AI 这个流行词大约在五年前回归,因为开始使用的数据是语言数据,所以现在的系统不仅仅是预测供应链、商业、价格等,它输出人类流畅的语言,人们说:‘天哪,我们解决了古老的 AI 问题。’事实上,从某种意义上说,如果你狭义地定义 AI 问题,比如图灵测试。但机器学习这一传统一直在持续,到那时已经融合了来自不同领域的人,并且确实在工业界产生了影响,现在仍然如此。但 AI 这个流行词因为损失函数而回归,现在在我看来,它对研究路径产生了扭曲效应,影响了我们对研究方向的思考,也影响了我们对商业模式和技术发展方向的思考。AI 还不够,他们还必须创造这个被大肆炒作的流行词 AGI,我们将大量讨论经济学作为智能的来源,一种社会智能,当它与机器学习风格的智能结合在一起时,你现在可以谈论规模,不仅仅是计算机的数量和数据量,还有人类的数量。对我来说,至关重要的是,人类作为这些新兴系统中的生产者和消费者的角色应该得到尊重、放大和深思。
I've never actually thought of myself as an AI researcher. I didn't read an AI book. The term was coined in the 50s and John McCarthy and others had particular goals in mind for coining it. And they had particular methods in mind like logical inference and so on that didn't really quite pan out. In the meantime in the 60s and 70s, 80s something arose called machine learning. The actual methods like decision trees and nearest neighbor and logistic regression and hidden Markov models were developed in other literatures mostly statistics operations research and so on and that led to industrial success stories. So supply chains and commerce and transportation systems all used and still to this day use vast amounts of machine learning. They used gradient-based methods and the cloud was developed to handle machine learning workloads at Amazon in fact. And so that's the tradition I came up in. I was trying to think about systems building at scale that would also serve multiple people. The AI buzzword returned I think maybe five or so years ago because the data that started to be used was language data and so the box now is not just making predictions about supply chains or commerce or prices or whatever it spits out a human fluent language and people said, 'Oh my god, we've solved the old AI problem.' In fact, in some ways if you define the AI problem narrowly like the Turing test. But there was this ongoing tradition of machine learning and by that time had incorporated people from all different kinds of fields and it was really having an impact industry still is. But the AI buzzword returned because of loss and now to my view it's been a distortionary effect on the path of research on how we think about where research should go but also on the path of how do we think about business models and how do we think about where technology is going. And AI wasn't enough they had to create this big hyped up buzzword AGI which we will talk a lot about economics as a source of intelligence, a social intelligence, and when it's put together with machine learning style intelligence you can now talk about at scale, not just numbers of computers and amount of data, but numbers of humans. And that's critically important to me that the role of humans as producers and consumers in these emerging systems should be respected, amplified, and thought about.
大脑就是一台计算机,如果我们模仿它,借鉴它的特性,并行化并增强它,它就能做出伟大的事情。但也就到此为止了。并不是说社会有一个目标,我们要努力实现这个或那个。它只会为我们解决问题,然后我们就快乐了。我离开硅谷的部分原因就是人们总是那样说话,我厌倦了。那里缺乏深入的、长期的思考。它变成了一场竞赛,一场金钱竞赛。
The brain is a computer, and if we mimic that and take aspects of it and parallelize it and make it more powerful, it'll just do great things. It kind of stops there. It's not that there's a goal in society that we're going to try to do this or that. It'll just solve problems for us and then we'll be happy. I got away from Silicon Valley partly because that's just the way that people talk and I got tired of it. There's not a lot of intellectual, deeper long-term thought going on. It became a rat race and a money race.
我的观点源于一个悠久的传统,即人们用社会科学的视角看待智能。我们是社会性动物。我们的智能很大程度上来自于我们汇集观点和思想,以及我们有文化来保留它们。此外,社会为我们的智能提供了背景。在一个背景下聪明的行动在另一个背景下就不聪明了。这一切都非常短暂且依赖情境。我们需要社会科学的理念来理解这意味着什么。我说社会科学时,包括经济学、博弈论。背景是,别人可能想利用我,或者与我合作,而我不确定。所以我必须试探、发出信号,并建立机制来有效互动。经济学以数学方式研究这些,这吸引了我,因为我倾向于数学思维。
My perspective comes from a long tradition of people having social science perspectives on intelligence. We are social animals. A lot of our intelligence comes from the fact that we aggregate opinions and thoughts, and we have cultures that retain them. Moreover, society provides a context for our intelligence. A smart action in one context is not smart in another. It's all very fleeting and contextual. Social science ideas are needed to appreciate what that means. When I say social science, I include economics, game theory. The context is that somebody else out there is trying to take advantage of me or collaborate with me, and I don't really know. So I have to put out feelers, do signals, and create mechanisms where we can interact effectively. Economics studies that in a mathematical way, which attracts me because I am mathematically inclined.
我不是 AI 的批评者。我想把它做对、做得更好,理解在这个世界上拥有智能意味着什么,以及安全、有趣,并思考长期问题。对我来说,你必须从形式或数学层面做到这一点。仅仅构建东西并发布出去是不够的。我说集体主义时,是指这项技术大多基于数十亿人的输入。所以已经有一个集体在输入,它也旨在服务数十亿人。所以它服务的是一个集体。有一个潜在的巨大网络。经济学至关重要。我想写下可操作的数学思想。
I'm not a critic of AI. I want to make it right and make it better, and understand what it means to be intelligent in this world, and safe, and interesting, and think about long-term issues. To me, you have to do that formally or mathematically at some level. It's not enough just to build things and put them out there. When I say collectivist, I mean that most of this technology is based on inputs from billions of people. So there's already a collective putting input in, and it's meant to serve billions. So there's a collective it's serving. There's a big network that's latent. Economics is critical. I want to write down actionable mathematical ideas.
这很有意思。20 世纪 70 年代,德雷福斯提出了第一步谬误,这与麦科德效应有关。我们创造出惊人的东西,就以为离无所不能只差一步。这些系统令人难以置信——它们能生成优美的文本、解决问题、编程。但奇怪的是,它们实际上并没有帮我们太多?我们本以为会带来革命。
This is interesting. In the 1970s, Dreyfus came up with the first step fallacy, and it's related to the McCord effect. We create something amazing and think we're only one step away from being able to do anything. These systems are incredible—they produce beautiful text, solve problems, do programming. But isn't it weird that they don't actually help us that much? We thought it would revolutionize.
这一点也不奇怪。模型还是旧的 AI 模型:我们只是构建一个智能的东西,它只升级了一点点。它会成为一个更好的搜索引擎。这没问题。我确实认为搜索引擎是人类的一大进步。但现在它变得不仅仅是搜索。它就像一个坐在你肩膀上的秘书,对你耳语。这是一个愚蠢的商业模式。我不认为很多人会想要那样。他们会关掉它。他们想自己思考。也许一天结束时需要个总结,但他们不想一直这样,与这个实体互动。这不是一个很好的商业模式。与此同时,我们有庞大的医疗系统、交通系统和金融系统,它们都基于数十亿智能体之间的数据流。这些系统已经成熟,可以用更经济的方式思考。智能体是什么?它们想从中得到什么?潜在的协作和竞争是什么?市场在几千年前就出现了,我们学到了一些原则,但我们可以改进它们。思考在几十亿年前就出现了,但我们并不完美。不仅在思考方面,而且在狭隘地追求自己的议程、伤害他人方面,即使我们不想这样。人类是美妙的——我们珍视生命、创造力、情感、爱——但人类也会做坏事。技术应该能够帮助你。你需要思考系统,思考生态系统。这些并不是真正的系统;它们是大统计盒子,做输入输出。这不是系统思维的方式。有一个更低层次的系统,计算机系统,但我想超越它。我想说:这属于什么生态系统?它与谁互动,以什么速率,什么质量,创造什么价值?我说价值时,通常指金钱。我希望这个东西能创造就业。我不希望它只是回答问题、为我们做事。我希望它创造工作和创造力的机会。
It's not weird at all. The model is the old AI model: let's just build something intelligent, and it only got upgraded a little bit. It's going to be a better search engine. That's fine. I do think the search engine was major progress for humanity. But now it became more than search. It's like a secretary sitting on your shoulder, whispering things to you. It's a dumb business model. I don't think many people will want that. They'll turn it off. They want to think for themselves. Maybe at the end of the day a summary, but they don't want this all the time, interacting with this entity. It's not a very good business model. In the meantime, we have huge healthcare systems, transportation systems, and finance systems that are all based on data flows among many billions of agents. They are ripe for thinking in a more economic way. What are the agents? What are they trying to get out of it? What kind of cooperation and competition is latent? Markets arose thousands of years ago, and we learned some principles, but we can improve them. Thinking arose billions of years ago, but we're not perfect. Not just in terms of thinking, but also in terms of narrowly following our own agenda and hurting other people even though we don't want to. Humans are wonderful—we prize human life, creativity, emotion, love—but humans also do bad things. Technology should be able to aid you. You need to think about the system, the ecosystem. These are not really systems; they're big statistical boxes that do inputs and outputs. That's not a systems way of thinking. There's a lower-level system, the computer system, but I want to be above that. I want to say: what ecosystem does this belong to? Who is it interacting with, at what rate, what quality, what values are being created? When I say value, I often mean money. I want jobs out of this thing. I don't want it just to answer and do things for us. I want it to create opportunities for work and creativity.
里奇·萨顿在设计与进化方面被频繁引用。我看了大卫·多伊奇关于解释的演讲。他说物理学家追求有原则的低层解释,但有时高层解释也很好,也许经济学就是其中之一。对于像伊利亚·苏茨克弗这样的人,他们谈论人类价值函数,说我们只需将 LLM 变成多智能体系统,就能免费获得所有经济方面的东西,你怎么看?
Rich Sutton is quoted a lot regarding design versus evolve. I watched a talk by David Deutsch about explanations. He said physicists go for principled low-level explanations, but sometimes high-level courses are really good, maybe economics is one. What do you say to folks like Ilya Sutskever who talk about human value functions and say we just turn LLMs into multi-agent systems and get all the economic stuff for free?
这不是一个好的工程思维方式。如果你是四五十年代的化学工程师,说我们就把很多东西混在一起让它工作,你可以做到,但你会得到很多爆炸和经济上不可行的东西。你会伤害很多人。我认为这些人很多没有考虑到已经受到伤害的人。Facebook 等已经伤害了很多年轻人。很多青少年有心理健康问题。计算机科学家根本没有谈论过这个。
It's just not a good way to think about engineering. If you were a chemical engineer in the 40s and 50s saying we're just going to throw a lot of stuff together and make it work, you could do it, but you'd get a lot of explosions and economically nonviable things. You'd hurt a lot of people. I think a lot of these people are not thinking about all the people that are being hurt already. Facebook and so on have damaged a lot of young people. A lot of teenagers are having mental health problems. This is not something that has been talked about by computer scientists at all.
现在我们又在谈论另一层次的替代。工作可能会消失,但这很残酷。当然,像以往一样,它会创造新的工作。我只是不喜欢那样说。所以你得说,退一步想想。你的目的是什么?你是想创造一种新的市场,让人们能够进来,让他们的才华得到重视和欣赏,让人们可以发布他们可能需要的东西的竞标,合作可以出现,生产者-消费者的关系可以被探索、理解和开发?这可以是计算和人类的混合。我认为最终我们都会融合,但在这个过程中,用所有这些隐喻来做如此颠覆性的事情,不是好的社会科学或好的数学。这只是隐喻。是的,你可以构建它,因为上一代人创造了这些惊人的东西来收集数据,我们可以做梯度下降和临时架构,是的,这有效。这很惊人,但不要给做这件事的人太多功劳。那是二三十年前的人做的。所以当前这一代只是太——没有太多思考,没有太多智力内容。只是,是的,可以构建它。可以从任何地方窃取数据,因为互联网允许这样做,而不给数据原始者任何价值回报。可以对此做贪婪下降,但需要大量资金,现在可以从那些没有深入思考的人那里获得。所以我可能看起来比我想要的更悲观。我的意思是,建设者有很多好处,但每一个以前的工程发展时代——电气、化学、机械——都有一些建设者,但他们有很多概念和很多思想家。事实上,所有这些工程学科都有类似麦克斯韦方程或牛顿方程的东西来帮助他们。这里,只是非常聪明且会编程的人,然后有很多直觉。我似乎从未看到任何感觉深刻智力的东西。感觉像科幻小说。
And now we're talking about yet another level of displacement. Jobs may go away, but that's tough. It'll create new ones, of course, like always. I just don't like to talk that way. So you have to say, step back a moment. What is your point? Are you trying to create a new kind of market where people could come in and have their talents valued and appreciated, where bids could be put out for things that people might need, and collaborations can emerge, and there could be producer-consumer relationships being explored and understood and developed? This could all be a mix of computation and humans. I think eventually we'll all kind of merge, but along the way, doing something so disruptive with all of these metaphors is not good social science or good mathematics. It's just metaphors. And yes, you can build it because the previous generation of people created these amazing things that collect data, and we can do gradient descent on it and ad hoc architectures, and yes, that works. It's amazing, but let's not give so much credit to the people that did that. It's the people 20 or 30 years ago who did that. So the current generation is just way too—there's not much thought going on, not much intellectual stuff. It's just, yeah, it's possible to build it. It's possible to steal the data from wherever you want because that's what the internet allowed, and not return any value to the person who originated the data. It's possible to run greedy descent on that, but you need huge amounts of money, and it's now possible to get it from people who aren't thinking very deeply. So I may seem more dark than I want to. I mean, there's a lot of good in builders, but every previous era of engineering development—electrical, chemical, mechanical—had some builders, but they had a lot of concepts and a lot of thinkers. In fact, all of those engineering disciplines had something like Maxwell's equations or Newton's equations to help them. Here, it's just people who are very smart and can code, and then have lots of intuitions. It seems to me I don't ever see anything that feels deeply intellectual. It feels like science fiction.
嗯,我想另一件没有帮助的事情是这些系统像汤一样。甚至有一个领域叫做机械可解释性,试图深入汤中,几乎就像他们在寻找 UFO。他们试图找到这些做推理或其他事情的原则性电路。我想你可以愤世嫉俗地说,这不像工程师建桥。
Well, I suppose another thing that doesn't help is that these systems are like soup. And there's even a field called mechanistic interpretability that tries to dig into the soup, and it's almost like they're searching for UFOs. They're trying to find these principled circuits that do reasoning or whatever. And I guess you could say cynically that it's not like when engineers build a bridge.
嗯,我比那稍微不那么消极。我不认为构建你不理解的系统是坏事。但你必须围绕它放置一些东西。周围的东西就像 AI 安全这样的流行词。这是一个流行词。好吧。你真正需要的——我是说,一个人——你无法向我解释你为什么选择这个 Airbnb 而不是另一个。你今天做的所有选择对我来说都是无法解释的。它们来自你的大脑。我不需要知道你选择的所有原因和细节。我需要知道的是你有点可预测,如果我给你某些选项,你可能会选这个而不是那个,因此我可以制定自己的计划,我们可以开始互动。所以这是经济学的一部分。经济学的思维方式说,我不理解所有那些外部实体,但我可以用一些经验法则或定量预测来互动,不受伤害甚至从中获得价值。所以不,我认为没有必要理解所有细节。现在,输入输出行为你通常需要比我们现在更好地理解。例如,如果我在银行被拒绝贷款,而银行使用了基于过去数据的大型 AI 程序,我想知道为什么。而为什么并不意味着你查看内部并给我看一些电路。没有人会想要那样。他们会想要,嗯,根据我们在这个大网络中使用的嵌入,有大约 50 个人和你非常相似。在那 50 个像你的人中,有些人得到了贷款,有些人没有。来,让我给你看看那些人是什么样的。你开始看到,哦,我明白了他们在这方面与我不同。这对我来说是可操作的。我现在可以改变事情。所以你必须围绕这个预测系统构建系统。例如,那是一个最近邻系统。那个系统会提供人们可能认为更像解释的东西。所以不仅仅是试图进入某物的内部。再次,化学工程——当然有热力学,很多东西被理解,但很多现象很长时间都不被理解。你混合一堆东西,某些波产生,某些事情发生,你利用它并继续前进。但你对输入输出行为和约束有所了解。我认为当前一代神经网络将继续——它们有很好的缩放行为。它们将继续存在。但它们真的必须被视为更大生态系统的一部分。然后你问,神经网络在这种情况下能做什么?它缺少什么?如果我有多个呢?它们如何彼此互动以及与我们互动?需要什么透明度才能使整体互动有效,无论我是否理解所有细节?
Well, I'm a little less negative than that. I don't think it's bad to build systems you don't understand. But then you've got to put things around it. And the things that are around are like buzzwords like AI safety. It's a buzzword. Okay. What you really need—I mean, a human—you can't explain to me why you picked this Airbnb over another one. All the choices you've made today are inexplicable to me. They come out of your brain. And I don't need to know all the whys and wherefores of your choices. What I need to know is that you're somewhat predictable and that if I make certain options available to you, you're likely to take this one versus that one, and therefore I can make my own plans and we can start to interact. So that's part of economics. The economic style of thinking says I don't understand all these other entities out there, but there are certain rules of thumb that I can use or quantitative predictions I can put in place that allow me to interact and not get hurt and even get value out of it. So no, I don't think it's necessary to understand all the details. Now, the input-output behavior you often have to understand better than we can now. For example, if I'm denied a loan at a bank and the bank used this big AI program based on past data, I want to know why. And why doesn't mean that you look in the internals and show me some circuit. No one's going to want that. They're going to want, well, there were like 50 people that are pretty much like you according to the embedding we're using in this big network. And of those 50 people like you, some got the loan, some didn't. Here, let me show you what those people are like. You start to see, oh, I see that they differ from me in this way. That's actionable to me. I could now change things. So you have to build systems around this predictive system. That's a nearest neighbor system, for example. And that system will supply what people might consider more like an explanation. So it's not just trying to go into the internals of something. Again, chemical engineering—there's certainly thermodynamics and lots of things are understood, but lots of phenomena were not understood for a long time. You mix up a bunch of stuff, and certain waves are created and certain things happen, and you exploit that and move on. But you understand something about input-output behavior and constraints. I think the current generation of neural nets will continue—they have very nice scaling behavior. They'll continue to be there. But they really have to be thought of as part of a bigger ecosystem. And then you ask, what can the neural net do in this context? What's it missing? What if I have multiple of them? How do they engage with each other and with us? What transparency is needed for the overall interaction to be effective, whether or not I understand all the details?
出于某种原因,如果我错了请纠正我,我有一种直觉,行为主义是坏的——仅仅因为没有机械理解而只看输出,有一个著名的例子,母鸡不知道它的脖子会被折断。实际上,一个例子是 AlphaFold。所以我上周在谷歌采访了 John Jumper,你对那 2 亿个预测蛋白质做了一些分析,你发现它们非常好,但缺少了一些东西,但你可以使它们更稳健。
For some reason, and correct me if I'm wrong, I have an intuition that behaviorism is bad—that just by not having any mechanistic understanding and only looking at the outputs, there's the famous example of the hen that didn't know its neck was going to be broken. And one example of this actually is AlphaFold. So I interviewed John Jumper last week at Google, and you did some analysis on those 200 million predicted proteins, and you found they were very good, but there was something missing, but you could robustify them.
你可以使它们更稳健。没错。我认为这是一个很好的例子。所以你知道,我非常钦佩 AlphaFold。我不认为它像 LLM。我认为它是有针对性的。它是针对一组特定问题,并且做得非常好。
You could robustify them. That's correct. And I think that's a good example. So you know, I'm a big admirer of AlphaFold. I don't think it's like an LLM. I think it's targeted. It was for a particular set of problems and it does it very well.
我们实证发现的问题是,当你问某些类型的问题时,特别是我们做过一个研究,看蛋白质中的量子涨落是否与磷酸化相关,即蛋白质在细胞中是否活跃。你可能会认为这些涨落导致链段悬垂或产生坏蛋白质,进化不会利用它们。但结果发现很多似乎被磷酸化了,意味着它们在细胞中具有反应性。这提示了一个假设检验:磷酸化的是/否与量子涨落的是/否之间是否存在关联?这是一个小的 2x2 列联表,你对其做统计检验。问题是,如果你只使用已知晶体结构的蛋白质数据,你没有足够的数据来高功效地检验这个假设,因此即使看起来有关联,你也不能拒绝无关联的原假设。另一方面,如果你使用 AlphaFold 的 2 亿个蛋白质,你可以高功效地检验假设并拒绝原假设。但我们发现,那个 2x2 列联表统计量的置信区间非常窄,而且远离真实值,即金标准值。我们在一个又一个领域都发现了这一点。为什么会这样?原因是训练集中可能没有很多量子涨落蛋白质的例子,因为过去研究不多,而且很难结晶。例子不多意味着 AlphaFold 很可能不会给出很好的答案,但它不会告诉你它没有给出误差条,而且它不会专门针对你问的问题给出误差条。这正是我想要误差条的地方,而它在构建和设计时并不知道那个问题。
The issue we found empirically was that when you ask certain kinds of questions, in particular we did one where we were looking whether quantum fluctuations in a protein were associated with phosphorylation, meaning the protein was active or not in the cell. You might think that these fluctuations, which lead to strands hanging off or kind of bad proteins, evolution wouldn't use them. But it turned out that a lot of them seem to be phosphorylated, meaning they are reactive in the cell. That suggests a hypothesis test: is there an association between yes/no phosphorylation and yes/no quantum fluctuation? So that's a little 2x2 table, and you do a statistical test on that. The problem is that if you just use known protein data that has a crystal structure known, you don't have enough data to test that hypothesis with high power, and so you can't reject the null hypothesis that there's no association even though there looks like there is. If on the other hand you use 200 million proteins out of AlphaFold, you can test the hypothesis with high power and you reject the null hypothesis. But what we found is that the confidence interval on that statistic of that 2x2 table was extremely narrow and way far from the truth, the true value of the gold standard. And we found this in domain after domain. So why is that? What's happening is that there are probably not many examples in the training set of proteins with quantum fluctuation because it hasn't been studied much in the past and it's hard to crystallize. Not many examples means it's quite possible AlphaFold won't give a great answer, but it won't tell you that it doesn't give you error bars, and it doesn't specifically on the question you're asking. That's where I want the error bars, and it didn't know about that question when it was built and designed.
好的。那么现在我有一个很好的统计问题。如果我在那 2 亿个数据中加入一点真实数据呢?我能移动误差条使其保持较窄吗?这样我既有高功效,又能覆盖真实值。
Okay. All right. So now I have a good statistical question. What if I add a little bit of ground truth data to the 200 million? Can I shift the error bar so it stays somewhat narrow? So I have high power, but it covers the truth.
答案是肯定的,有一种方法论。我们开发了一种叫做预测驱动推断的方法,正是做这个的。它会像经典统计设置一样覆盖真实值,但使用的是这种高度有偏的架构。现在它整体上不再有偏。事实上,它的整体准确性很高。但对于我问的问题,它可能非常有偏。这在科学中会经常发生,因为科学家很少只对重复研究过去感兴趣。他们感兴趣的是知识边缘的全新事物。而这正是这些基础模型表现最差、偏差最大的地方。因此,任何基础模型都需要具备收集一点真实数据,并用类似程序合并,然后给出更可信答案的能力。这都不是科幻。这是可以做到的,也是真正需要做的,我相信 AlphaFold 的人会认同。他们不会觉得奇怪或惊讶。但很多其他人在谈论偏差之类的问题,他们要么不担心,说数据足够多偏差就会消失,要么只是批评架构和输出,但没有提出能帮助我们前进的科学方法。这就是我们目前的状况。
And the answer is yes, there's a methodology. We've developed something called prediction powered inference that does exactly that. And so it'll cover the truth just like in a classical statistical setting, but it's using this rather highly biased architecture. And now it's not biased overall. In fact, its accuracy is high overall. But for the question I'm asking, it might be very biased. And that's going to happen a lot in science because scientists are rarely interested in just studying the past over again. They're interested in brand new things on the edge of knowledge. And that's where specifically these foundation models will be most poor and most highly biased. So there needs to be around any foundation model the ability to maybe collect a bit of ground truth data to merge it in with some procedure like this and then to give out a more trustable answer. That's all not science fiction. That's what can be done and what really needs to be done, and I'm sure the AlphaFold people are on board with that. They would not find that weird or surprising. But a lot of other people out there talk about bias and all that, and they either don't worry about it, they say it'll go away if we have enough data, or they just critique the architectures and critique the outputs but they have no scientific method in mind that'll help us go forward. So that's kind of the state we're in.
我稍微挑战了一下约翰关于 AlphaFold 理解的程度,他基本上对“理解”这个词很反感。
I challenged John a little bit about the extent to which AlphaFold understands, and he was basically allergic to the word understands.
我们并不试图告诉你一切。我们不是整个细胞的模型。这些机器让我们预测。它们让我们控制。我们必须自己推导理解。对吧?我们现在可以在人工产物上进行实验。我们可以查看 2 亿个预测结构,而不仅仅是 20 万个实验结构,以帮助我们理解,但它并不为我们执行理解的行为。它执行的是预测和可能的控制行为。
We are not trying to tell you everything. We are not a model of the entire cell. These machines let us predict. They let us control. We have to derive our own understanding at this moment. Right? We can experiment now on the artifact. We can look at the 200 million predicted structures, not just the 200,000 experimental structures, in order to help us understand, but it doesn't do the act of understanding for us. It does the act of predict and maybe control.
为什么 AlphaFold 应该理解?
Why should AlphaFold understand?
嗯,理解意味着什么?我是说,他向我勾勒了一下。他有点说这是一个奇怪的外星人工制品,不是创造出来的,而是精炼出来的。有这个循环路径。你可以多次输入。你可以在中途破坏它,网络只是迭代地先解决复杂部分,然后精炼、精炼、精炼。就像,我们能将其解释为一个理解过程吗?
Well, what would it mean to understand? I mean, he was sketching it out to me. He kind of said that this is a weird alien artifact and it's not like it's created. It's refined. There's this recycle pathway. You can put the thing through multiple times. You can kind of corrupt it halfway through and the network is just iteratively solving the complex bit first and then refining, refining, refining. And like, could we interpret that as an understanding process?
我认为我们不需要这样看。我认为这种将智能和理解拟人化的做法既不必要,也不恰当,而且对很多问题来说是一种干扰。为什么要说它理解?我的一些背景来自于亲眼目睹 20、30 年前在工业环境中部署机器学习算法。我第一次去西海岸时,大约 2000 年访问了亚马逊。他们使用大量数据进行供应链建模,用的是当时的神经网络,也就是随机森林,而且效果很好。他们能对某些船只在印度洋是否延误等做出非常出色的预测,因此某些零件不会及时到达,整个供应链每天处理数十亿产品并发送给数亿人。所以没有任何人类能理解那个大黑箱里发生了什么。但这并不必要。事实上你可以问:那个整体系统理解运输和物流吗?答案是,谁在乎。它执行了一个非常重要的优化和预测过程,使得可以围绕它构建工程系统。它降低了不确定性。它使得某种库存和规划成为可能,这就是你需要的。你不在乎它是否必须被冠以“理解”或“智能”这样的词。那是给媒体看的。这就是我对很多推出 AGI 和 AI 术语的人的问题。媒体趋之若鹜,他们知道尽管我们不清楚理解和智能意味着什么,而且我们自己的研究意识到我们不在乎也不需要它们。我们想构建好的系统。
I don't think we need to see. I think this anthropomorphizing of intelligence and understanding all that is not necessary, not appropriate, and is a distraction for many many problems. Why say it understands? Some of my heritage comes from seeing in real life in industrial settings machine learning algorithms being rolled out 20, 30 years ago. When I first went to the west coast, I visited Amazon in around 2000. They were using huge amounts of data to do supply chain modeling using the neural networks of the day. It was random forests and it was really working. They could make really fantastic predictions of whether certain ships would be delayed in the Indian Ocean or whatever, and so certain parts wouldn't arrive in time, and the overall supply chain takes billions of products and sends it to 100 millions of people per day. And so there's no way that any human can understand what's happening in that big black box. But it's not necessary. And in fact you can ask: does that overall system understand transport and logistics? And the answer is who cares. It does a very important optimization and prediction process that allows an engineering system to be built around it. It brings down uncertainty. It makes it possible to do kind of stockpiling and planning, and that's what you ask for. You don't care whether it has to have a word like understand or intelligence applied to it. That's for the media. That's kind of my problem with a lot of these people rolling out AGI and AI terminology. The media laps it up, and they know that even though we don't have a clue what understanding and intelligence mean, and we in our own research realize we don't care or need it. We want to build good systems.
我们无法本质化它。像弗朗斯甚至大卫·克拉科夫这样的人,把智能说成是适应、综合,当然还有颗粒表征。但万一还有一步呢?我们不要拟人化。假设理解不是终点,而是通往那里的路径。我们知道在现实世界中我们是集体智能,就像盲人摸象。我们都走自己的路,过自己的生活,对同一个整体有不同的视角。那么,如果更好的理解形式只是能够用积木从你的角度重构事物,而不是试图本质化它呢?
We can't essentialize it. Folks like France or even David Krakow talk about intelligence as adaptation, synthesis, of course, grain representations. But what if there is a bit of a step? Let's not anthropomorphize it. Let's say that understanding is not about the endpoint; it's about the path which led us there. We know that in the real world we're a collective intelligence, and there's the blind men and the elephant. We all take our own paths and lives, and we have different perspectives on the same whole. So what if a better form of understanding is just being able to reconstruct the thing from your perspective using building blocks, rather than trying to essentialize it?
这听起来都很好。只是这不是我们做研究的人会用的语言。我们当然会稍微考虑这些术语,但我们会试图把它变成某种均衡或优化问题。这里有可用的信息、数据、算力和错误率。我们试图围绕它建立一点结构。总有那么一个创意时刻。我记得小时候对跳高感兴趣。你会走向横杆,用各种方式跳过去。当时奥运选手用的是滚式。然后有个叫迪克·福斯伯里的家伙说:‘不,如果我背越式跳,我能做得更好。’没人想到过那样做。他一做,大家都跟着做,横杆提高了大约半米。是什么过程导致了那个?是理解过程吗?那只是一点‘试试不同的东西’,加上尝试和测试的能力。大量的工业规划就是‘试试看什么有效’。那些叫做 A/B 测试,一直在做。我对此没有异议。它不是基于理解,但它导致了优化的系统,能做出人们以前没想过的事情。所以是理解与那种方法的结合。但仅仅是理解?我过去是认知科学家,对神经科学感兴趣。人们应该对这些感兴趣。它们很迷人,但它们不是思考如何构建在世界上工作的系统的前沿,也不是构建下一代系统的前沿。很多人一直在说我们必须把逻辑或符号放回去,因为这来自我们之前对人类行为的看法。可能人类能够进行一些逻辑推理,并且可能有一些符号,无论它们是内置于某个复杂的网络还是以某种方式具体化。我的直觉和你的一样好。但真正的目标是:我骨子里是个工程师,一个数学倾向的工程师。我想说,你想要实现什么?你是想取代老师?还是想让医生更好?你想做什么?这个问题的抽象和切入点是什么?然后你如何从中抽身,以一种优雅且能激励他人的通用方式来做?
That all sounds great. It's just not the language that those of us who do research would use. We would think in those terms a little bit, of course, but we would try to turn it into some kind of an equilibrium or optimization problem. Here's the information that's available, here's the data, here's the power, and the error rates. We try to put a little bit of structure around it of that form. There's always this creative moment. I remember when I was a kid interested in high jumping. You would go up to the bar and jump over it in various ways. The barrel roll was the technique the Olympians were using. Then this guy Dick Fosbury came along and said, 'No, if I go backwards I can do better.' No one had thought about doing that. As soon as he did it, everybody did it, and the bar went up by like half a meter. What process led to that? Was it an understanding process? It was just a little bit of 'let's try something different' mixed in with the ability to try it out and do tests. A huge amount of industrial planning is 'try it out and see what works.' Those are called AB tests, done all the time. I've got nothing against that. It's not based on understanding, but it's led to optimized systems that can do things people hadn't thought about before. So a blend of that with understanding. But just understanding? I was a cognitive scientist, interested in neuroscience. One should be interested in those things. They're fascinating, but they aren't the leading edge of thinking how to build systems that work in the world, and they're not the leading edge of trying to build even the next generation systems. A lot of people keep saying we've got to put logic back in or symbols because that came from our previous view of what humans are doing. Probably humans are capable of some logical reasoning and probably have some symbols, whether they're built into some complicated network or reified somehow. My intuition is as good as yours. But really the goal is: I tend to be an engineer at heart, a mathematically inclined engineer. I want to say, what are you trying to achieve? Are you trying to displace teachers? Are you trying to make doctors better? What are you trying to do? What would be the abstractions and the points of entry into that problem? Then how can you pull back from that and do it in some general way that's elegant and will inspire others?
看到不同学科的科学家从多学科角度攻克这个问题,真是太有趣了。例如,物理学家在非常低的层次上工作,谈论粒子系统的动力学等等。我真正着迷的是:你从经济学角度入手,传统上经济学被这种能动性视角主导,你谈论均衡和激励等等。这如何融入其中?你如何将一个非常复杂的系统几乎分解成这种新的思维框架?
It's so interesting seeing different scientists from a multi-disciplinary perspective attack this problem. Physicists, for example, they work very low level and talk about the dynamics of particle systems and whatnot. What I'm really fascinated in: you come at it from an economics perspective, which is traditionally dominated by this agential lens, and you talk about equilibria and incentives and so on. How does that come into it? How would you take a very complex system and almost decompose it into this new frame of thinking?
这还没有做得足够多,让我有大量好例子,但我们一直在看一些规模适中的例子。例如,我们研究了一点药物发现和监管。假设我是一家制药公司:我测试各种蛋白质,把它们扔进动物体内,也许还有几个人,看看什么有效。我有一些所谓的理解,我知道背后的进化生物学,这指导了我。但到了某个时候,必须有人在现实世界中真正测试,监管机构必须介入说:‘是的,那可以上市,或者不行。’所以现在你有了一个由科学家和制药公司——不止一家,而是很多家——以及蛋白质组成的复杂网络。现在你必须思考这个系统如何运作。希望监管机构试图在整个系统上使假阳性率和假阴性率都低。这就是目标。所以这是一个统计问题。但等等,经典的统计问题会从某个来源收集独立同分布数据。不,数据来自自利的制药公司。他们的动机是什么?钱,也许还有帮助。他们想帮助人和赚钱。所有这些对你作为监管机构来说是隐藏的。所以经济思维开始起作用。它对我隐藏,但并非任意。我可以通过各种方式探查。这变得非常经济。经济学家思考如何定价。如果我有很多人乘坐我的航班——一千人刚到达,想从这里去伦敦——每个人都有一个不同的价格点,那个价格点会根据他们有多急切而实时变化。这不仅仅是因为他们有很多钱;而是因为他们有需求,而我不知道那些需求是什么。所以我做的是设置各种服务和价格,涵盖各种可能性,这样总体上我可能赚到足够的钱,每个人都会有点满意,服务也会继续。这是知道一些事情和承认你不知道其他事情的结合,但把它放在一个能够处理这种不对称和激励混合的系统中。
It's not been done really enough for me to have tons of great examples, but we've been looking at modestly scaled examples. For example, we looked at a little bit of drug discovery and the regulation. So I'm a pharmaceutical company: I test out all kinds of proteins, I throw them in animals and maybe a few humans to see what's working. I have some understanding, quote unquote, and I know the evolutionary biology behind it, and that guides me. But at some point someone's got to really test this out in the real world, and regulatory agencies have to come in and say, 'Yeah, that goes to market or it doesn't.' So now you've got a tangled web of scientists and pharmaceutical companies—not just one but many—and proteins. Now you've got to think about how that system is behaving. Hopefully the regulatory agency is trying to overall over the entire system have the number of false positives be low and false negatives be low. That's the goal. So it's a statistical problem. But wait, a classical statistical problem would just gather IID data from some source. No, the data is coming from self-interested pharmaceutical companies. What's their motivation? Money and maybe to help. They want to help people and money. All that is hidden from you as a regulatory agency. So the economic mindset comes into play. It's hidden from me, but it's not arbitrary. I can probe in various ways. That becomes very economic. Economists think about how you set prices. If I have a lot of people coming on my airline—a thousand people who just arrived wanting to go from here to London—every one of them has a different price point, and that price point will shift in the moment based on how eager they are. It's not just because they have a lot of money; it's because they have needs, and I don't know what those are. So what I do is set various services and various prices that bracket the possibilities, so that overall it's likely I'll make enough money, everybody will be kind of happy, and the service will go forward. It's a blend of knowing a few things and admitting that you don't know other things, but putting it in a system that actually can work with that kind of mix of asymmetries and incentives.
所以,这里的激励是:你有一个特定的服务和价格。如果你选那个,你很可能能上飞机,你很可能得到你需要的那些好东西等等。那并不会强迫你做什么,而是激励你。在制药领域,如果我能激励他们主要提交那些他们做过一些测试、或者他们相信是相当不错的药物,而不是随便扔给我任意的药物,那么也许整个系统就能达到你想要的错误率。因为如果你不这样做,那么如果一种药有十亿人会用,不管它是否真的有效,你都能赚钱。所以你需要让它上市。怎么上市?你就把它扔给监管机构,可能有个假阳性。他们得到了假阳性,就让它上市了。你赚了一大笔钱。如果这种情况足够多,激励就全错了,整个系统就无法控制第一类和第二类错误。这些是我们实际研究过的例子。但我希望你能理解,当这一切开始在社会中真正展开时,不会是有几个大语言模型,每个人都像用搜索引擎一样咨询它们。那不是模型。而是会有本地数据,就像我告诉你的关于预测能力推理那样。每个人都必须审查他们收到的信息。还会有本地数据,因为我花了一些成本收集了它,我不想随便送人。谢天谢地,终于 Anthropic 在付钱给人了。那一定是未来。所以我的数据会有一些竞争价值,不会随便送出去。那么现在,如果你开始与很多想从互动中获得价值的人互动,你就必须谈论激励。他们发送数据的激励是什么?不仅是发送数据,还要发送正确的数据,发送真实的数据,不要对抗。我无法想象这一切在社会中全面展开,以及我们一生中的所有决策,如果没有一个深刻的微观经济学视角伴随数据的梯度下降。
So, the incentives there are that you have a certain service and price. If you pick that one, you're likely to be able to get on the airplane, you're likely to have the goodies you need or whatever. And that then doesn't make you do something; it incentivizes you. In the pharmaceutical world, if I could get them to be incentivized to mostly send in drugs they've done some testing on or they have some belief it's a pretty good one and not just throw arbitrary ones at me, then maybe the overall system will actually have the error rate you want it to. Because if you don't do that, then if it's a drug that a billion people will use, you're going to make money whether it really works or not. So you need it to get to market. How do you get it to market? You just throw it at the regulatory agency and there maybe is a false positive. They get a false positive and put it on the market. You make a ton of money. If there's enough of that, the incentives are all wrong and the overall system will not control type one and type two errors. Those are examples we've actually worked on. But I just hope you can appreciate when all this stuff starts to really roll out in society, it's not going to be that there are a few big LLMs and everyone consults them like a search engine. That's just not the model. It's going to be that there's local data, like I told you about with prediction power inference. Everyone has to vet what's coming at them. There's also going to be local data because I collected it with some expense and I want to just give it away. Thank god finally Anthropic is paying people for money. That has got to be the future. So I'm going to have some competitive value in my data and not just give it out. And so now if you start interacting with lots and lots of people that want to get some value out of the interactions, you have to talk about the incentives. What's the incentive for them to send the data? Not only send the data, but send correct data, send truthful data, don't be adversarial. I cannot imagine a fully-fledged version of all this rolling out in society and all of our decision-making throughout our lives without a deeply microeconomic perspective accompanying the gradient descent on data.
你提到了这个三层模型。有一个例子是,你可能有一些消费者,他们有他们的数据,然后有谷歌,谷歌在使用数据,消费者获得服务,然后谷歌可能把数据卖到别处。这有点像传统模型。我们从这个开始吧。
You spoke about this three-layer model. There was an example where you might have consumers and they might have their data, and you've got Google, and then Google is using the data, the consumers are getting a service, and then Google might sell the data over here. That's kind of like a traditional model. Let's start with that.
好的。所以这些真的有点像玻尔原子之类的东西。我们作为科学家,试图找出一个最小模型,能展示我们想研究的一些行为。那么让我们考虑一个数据市场,因为数据不仅仅是用来分析构建大语言模型的东西。它也是可以买卖的,有价值,而且还有隐私问题。所以我们搭建一个小型最小模型来研究它。我们做过的一个叫做三层数据市场。它在现实世界中存在。这是一个抽象,但它是真实的东西。你有一个或多个用户进入一些平台。平台提供服务,比如支付服务,当我使用它时,他们从我这里获取数据,了解我做了什么样的购买,并用这些数据来改进他们的服务。这是一个很好的小循环。问题是,他们很少能从那项服务中赚到足够的钱。他们抽取一小部分,商家不喜欢给他们,所以他们必须做其他事情来维持业务。通常很长一段时间以来,大概 20 年了,他们一直在把数据卖给第三方数据买家。这些人不是试图破坏人们隐私的坏人。他们试图做市场研究,了解什么会有效,人们真正在做什么。这是行为研究,对他们有价值。他们为此付费。谷歌不需要这个,因为他们创造了这个人工广告市场,我们可以多谈谈,那种东西超级强化了所有这些胡闹。但其他公司比如万事达卡就必须卖他们的数据。所以现在它是一个三层的东西,一旦引入了第三层,均衡就必须改变,因为发送数据的用户失去了一些东西。他们失去了一点隐私。一个我完全不了解的第三方正在获取关于我的数据。我不能就这么接受。但我又不能走开。所以系统现在有了压力。在一个有效的经济系统中,会发生的是,你不会只是等待监管机构进来说不允许这样做。你会做的是,平台会说,‘我们会为你提供可调节的差分隐私级别,需要一些成本’,或者‘我们公司,我是谷歌,我会提供 0.3 级别,而另一家公司说我会提供 0.7 级别。’所以用户看到后说,‘啊,0.7,那更好。我真的很在乎我的隐私,所以我去那里。’那家公司然后会开始获得更多数据,他们的服务会变得更好。你就有了一个很好的小反馈循环。但现在数据买家会看那个人的数据。0.7 意味着数据中添加了更多噪声。对数据买家来说价值更低。数据买家会说,‘我会花更少的钱。我给你更少的钱。我会给谷歌更多的钱。’所以现在你可以看到这里有相互冲突的倾向。激励是一致的,但对每个人来说并不是最优的。所以现在的数学不仅仅是一个优化问题。数学是一个均衡问题。但这是一个涉及统计断言、数据以及你能用这些数据预测多少等等的均衡问题。所以你用误差线和统计预测来量化它。你把所有这些放在一个大的数学系统中,你可以找到作为各种系统参数的函数的均衡。例如,监管机构是否可以要求一个最低隐私级别?或者是否存在某种异质隐私预算?等等。
Okay. So those are really kind of Bohr atom kind of things. We're being scientists there. We're trying to say what's a minimal model that exhibits some of the behavior we want to study here. So let's think about a data market because data is not just something you analyze to build a big LM. It's also something you'd sell and buy and has value, and also there are privacy concerns about data. So let's put a little minimal model together where we could study that. One we've done is called a three-layer data market. It exists in the real world. This is an abstraction, but it's a real thing. You've got a user or multiple users coming into some platforms. The platforms provide a service, like a payment service, and as I use that, they get data from me, they learn about what kind of purchases I've made, and they use that data to make their service better. That's a nice little loop there. The problem is that rarely do they make enough money off of that service. They take a small cut that the merchants don't like to give them, so they have to do other things to stay in business. Typically now for a long time, probably 20 years, they've been selling their data to third-party data buyers. These are not evil people trying to ruin people's privacy. They're trying to do market research, learning what would work and what people are really doing. This is behavioral studies, and there is value to them. They pay for it. Google doesn't need this because they created this artificial advertising market, which we could talk more about, that kind of superpowered all this nonsense. But other companies like Mastercard would have to sell their data. So now it's a three-layer thing, and as soon as that third layer was introduced, the equilibrium has to shift because the user who's sending their data in just lost something. They lost a little bit of privacy. Some third party that I don't know anything about is getting data about me. I can't just accept that. But I can't walk away. So there's a stress on the system now. In an effective economic system, what would happen is that you wouldn't just wait for the regulator to come in and say no this can't be done. What you would do is that the platforms would say, 'We'll offer you a tunable level of differential privacy for some cost,' or 'We'll just say that our company, I'm Google, I'll offer you level 0.3, and some other company says I'll offer you level 0.7.' So the user looks at that and says, 'Ah, 0.7, that's better. I really care about my privacy, so I'll go there.' That company then will start to get more data, and their service will get even better. And you got a nice little feedback loop there. But now the data buyers will look at the data from that person. At 0.7 means more noise has been added to the data. It's less valuable to the data buyer. The data buyer will say, 'I'll spend less. I'll give you less money for that. I'll give more money to Google.' So now you can see there are conflicting tendencies here. The incentives are aligned, but they're not optimal for everybody. So now the mathematics is not just an optimization problem. The mathematics is an equilibrium problem. But it's an equilibrium problem that involves statistical assertions, data, and how much you can predict with this data, and so on. So you quantify that with error bars and statistical predictions. You put that all together in a big mathematical system, and you can find the equilibria as a function of various system parameters. For example, is there a minimal level of privacy the regulators could require or not? Or is there some heterogeneous privacy budget? Etc.
你可以加入各种要素,然后画一个小图,看看均衡如何移动,均衡中三个参与者的总效用加起来就是社会福利。你可以问,这个均衡的社会福利有多高,跟那个比怎么样。另一个监管者可能会说,我更喜欢这个,因为它的整体社会福利更高,法律可以按那个水平来制定。所以,即使这是一个玩具模型,它包含了我非常感兴趣的要素:预测模型、数据市场、金钱激励,以及一个已经实际运作但人们没有很好思考的系统,就像药物发现领域一样。但如果从经济学的角度来看,你可以让系统变得更好。
You can put in various things and now you do a little plot of how the equilibria moved and the equilibria have overall utilities for all the three players summed up. That's the social welfare. You can ask how high is the social welfare at that equilibrium versus this one versus this one. And another regulator could look at that and say, well I prefer this one because it's overall higher social welfare and laws could be made at that level. So, even though this is a toy model, it has the ingredients that I'm very interested in: predictive models, data markets, money incentives, and a real system that is already kind of working, but people aren't thinking about it very well, just like in the drug discovery domain. But if you take an economics point of view, you can make the system better.
所以把它建模成一个动力系统,我们可以模拟,然后得到这些模式……
不,在那种情况下我们不需要模拟;你实际上可以写出方程并计算均衡。这是一个斯塔克尔伯格博弈,你可以找到均衡。但在其他情况下,你会模拟。关键是,我的很多机器学习同事对不动点算法、寻找均衡以及参数变化如何影响均衡知之甚少。那是经济学的东西。机器学习的人非常擅长优化,但这不是一个优化问题。数学的其他分支中有很多算法可以找到帕累托前沿,并统计地作为市场规模和人口规模的函数。在这个时代,这两者几乎从未交汇,这有点令人惊讶。经济学家从未有大量数据来指导他们的市场设计,所以他们只是写下一堆方程,做出理性假设,然后数学地找到均衡。而机器学习的人从未考虑过均衡;他们只是有大量数据,并用它来做显而易见的事情:预测一串单词中的下一个词。但未来必须是这些分支融合在一起。经济学的均衡视角至关重要,但适应性视角也至关重要。
No, we don't have to simulate it in that case; you can actually write equations and calculate equilibria. It's a Stackelberg game and you can actually find the equilibria. But in other cases you would simulate. The point is, a lot of my machine learning colleagues don't know much about fixed-point algorithms and finding equilibria and how they shift as you shift various parameters. That's economics stuff. Machine learning people are really good at optimization, but this is not an optimization problem. There are all these algorithms in other branches of mathematics that find Pareto frontiers and do it statistically as a function of size of various markets and populations. And it's kind of amazing in this era that the two have almost never met. The economists never had a lot of data to inform their design of their market, so they just wrote down a bunch of equations and made rational assumptions and then found equilibria mathematically. And the machine learning people never thought about the equilibria; they just had a lot of data and used it to do the obvious thing: predict the next word in a string of words. But the future has to be that those branches come together. The economics equilibria perspective is critical, but the adaptive perspective is also critical.
数据和您所说的那种知识有什么区别?
What is the difference between data and the kind of knowledge that you're talking about?
社会知识非常短暂,非常当下。我走在哥本哈根的街道上,那里有各种各样的小市场,有什么商品、什么价格、我可能喜欢什么。这一切都非常短暂。这是思考这一切的更好方式。你不能仅仅收集足够的数据就知道那个走在街上的人会来买这个产品。我们所有的决定和选择,即使是在接下来的 10 秒内会发生什么,都无法用足够的数据来覆盖。所以你必须更谦虚一点:我有很多无知,但这并不意味着我不能建立一个安全的系统,比如一个市场,人们可以进来,不被欺骗,获得价值,并且它可以随着时间演变。它可以以各种方式转变,我不需要成为顶上的神,设计人类价值函数,让它特别符合人类真正想要的东西。不,它必须是一个允许自下而上的偏好以人类当下想要的方式表达的系统。系统尊重这些东西,并希望更多地了解它们,在当下使用它们,也许保留一些。也许一切都是短暂的,然后消失。人们对数据似乎有一种巨大的天真:即使你有艾字节的数据,你也会错过所有细节,而这些细节可能正是某类决策中最重要的东西。即使是基于海量数据的 AlphaFold,在某些查询上也表现不佳。他们会修补那些问题,它会做得更好,但人们提出的新问题总是处于知识的前沿。人们没有考虑到这一点;他们想,‘我只需要取代老师,因为老师不是在知识前沿工作,而是在已知的东西里工作。’这没问题,你可以辅助老师,但好老师也知道如何迁移到知识前沿。
Social knowledge is very ephemeral and very in the moment. I walk down the streets of Copenhagen and there are all kinds of little markets out there, what's available and at what price, what I might like. It's all super ephemeral. That's a better way to think about all this. You can't just gather enough data to know that that person walking down the street is going to come buy this product. All of our decisions and choices cannot be covered by enough data to know everything about what's going to happen even in the next 10 seconds. So you have to be a little more humble: I have a lot of ignorance, but that doesn't mean I can't build a safe system like a market where people can come in, not get cheated, get value, and it can evolve over time. It can shift in ways, and I don't have to be the god figure at the top designing the human value function to make it respond particularly well to what humans really want. No, it has to be a system that permits bottom-up preferences to be expressed in the way the human wants in the moment. The system respects those things and wants to learn more about them, use them in the moment, and maybe keep some of it. Maybe it's all ephemeral and goes away. There seems to be a huge naivety about data: even if you have exabytes of data, you're going to miss all the details that are probably the main thing that matters for a particular class of decisions. Even AlphaFold, based on huge amounts of data, doesn't do well on certain queries. They'll patch those and it will do better, but the new questions people ask will always be on the edge of knowledge. People aren't thinking about that; they think, 'I just need to replace the teacher because a teacher is working not on the edge of knowledge, but back in the stuff that's already known.' That's fine, you can aid teachers, but good teachers also know how to migrate to the edge of knowledge.
当我们进行抽象和理想化时,总是会有一点损失。我对这个观察很着迷:市场在资本主义之前就存在了。这是一种自下而上的东西。
When we do abstraction and idealization, it's always a little bit lossy. I'm fascinated by this observation that markets were around before capitalism. It's this bottom-up thing.
这不是资本主义。那是让市场运作的一种方法论,但不是唯一的方法。
This is not capitalism. That's one methodology for making markets work, but it's not the only one.
没错。所以这是一种自然现象,是我们可能称之为市场的东西的涌现,它是建设性的、发散的、多样化的。然后你说,在某个地方,理论与实践可以结合:我们可以创建抽象,做一些建模,并尝试以不太有损的方式去做。
Exactly. So it's a natural phenomenon, the emergence of something we might call markets, and it's constructive and divergent and diverse. And then you're saying somewhere the rubber can meet the road: we can create abstractions, do some modeling, and try to do it in such a way that it's not too lossy.
绝对是的。人类文化创造抽象;个体人类也创造对他们有用的抽象。当这些抽象足够有用时,它们可以被传播并提升到文化中,这上下流动一直在发生。确实,系统或许可以帮助这一点。我不会完全信任系统来承担这个负担,但它可能是有帮助的。所以,确实,不仅仅是单个认知实体创造抽象,我们不应该仅仅将其具体化。
Absolutely. Human culture creates abstractions; individual humans create abstractions that work for them. When those abstractions are useful enough, they can be communicated and promoted into the culture, and that flows up and down all the time. Indeed, that's something that systems could perhaps help with. I'm not going to just trust systems to take on that burden, but it could be helpful. So indeed, it's not just the individual cognitive entity that creates abstractions, and we should just reify that.
是的,文化确实会产生抽象。你可以研究它的微观经济,也可以不研究。你只需说这些抽象是文化的一部分,它们很有用,所以留存了下来。未来,我们不会只保留旧的抽象,而是构建允许新抽象涌现的系统。这不是由神祇想好一切然后塞进去。硅谷说我们有海量数据,可以自上而下搞定一切,但他们忘了数据首先是自下而上产生的,是情境化的,由人提供的。他们也忘了这一切必须持续在微观层面进行,而这超出了他们的感知能力,我们也不希望他们过多介入我们的生活。搜索引擎之后,我觉得那是一项很棒的技术,提供了访问权限,但后来很多东西都变得非常侵入。他们要给你戴眼镜,在你家里装摄像头,了解你生活的所有细节,然后让你的生活变得更好。这个等式对我来说不成立。
Yes, it comes, but cultures create abstractions. You can study the microeconomy of that or not. You can just say those abstractions are part of the culture and they're useful. They've stayed around because they're useful. Going forward, it's not that we're going to just keep the old ones, but we're going to build systems that allow new ones to emerge. That is not the god figuring it all out and putting them in there. Silicon Valley says we've got so much data that we can do it all top-down, but they forget that first, the data came bottom-up. The data was contextual and supplied by people. They also forget that it all has to continue at a micro level that is beyond their ability to sense, and we don't want them so much in our lives. After the search engine, which I thought was fantastic, allowed access, but then a lot of it was very prime. They're going to put glasses on you, cameras in your home, know all the details of your life, and make your life better somehow. That equation did not calculate for me.
是的。所以这种观点认为文化是天空中的抽象硬盘,但文化适应性很强。我们可以删除策略。知识衰减很快,组织维护知识。真正好的知识会留存很久。在底层,我们在创造新知识。那么整个生态系统如何运作?我们如何指定好的东西留存下来,并找到新的知识?
Yeah. So this idea that culture is the abstraction hard drive in the sky, but culture is very adaptable. So we can delete strategy. Knowledge decays quickly, and organizations maintain knowledge. Really good bits of knowledge stay around for a long time. Down at the bottom, we're creating new bits of knowledge. So how does that whole ecosystem work? How do we designate good things that stick around and find new bits?
我们不会。你我不会,但那是优秀知识分子做的事。经济学、行为组织学——组织如何有效涌现?这非常有趣。不是全部,但已知很多,有些是数学的,有些是最佳实践。这些才是讨论 AI 生态系统的方式,而不仅仅是神经科学、神经元隐喻和物理隐喻,这些也是我背景的一部分,但当这些事物落地时,感觉非常欠缺。行为组织学——人们如何组织起来,不仅促进公司良好收入,也促进民主。有人讨论所有这些,但我认为他们在硅谷没什么存在感。也许这也有好处——硅谷让他们走,他们会烧很多钱,引起头痛,但也会创造搜索引擎之类的东西。我认为很多公司专注于创造价值。亚马逊和 Meta 不同。亚马逊有商业模式:把包裹送到门口,背后创造技术支持,大多有益。你需要担心劳动力市场,但这些都是值得担心的好事。仅仅创造预测的计算产物,戴上眼镜就活在他们的世界里——那不是商业模式。那是科幻梦想,可能对人类无益。
We don't. You and me don't, but that's something that good intellectuals do. Economics, behavioral organization—how do organizations effectively emerge? That's really interesting. It's not everything, but there's a lot known there, some mathematical, some best practices. Those are the ways AI ecosystems should be talked about, not just in terms of neuroscience, neuron metaphors, and physics metaphors, which were part of my heritage too, but felt so lacking when we see these things hitting the road. Behavioral organization—how people are organized into things that promote not only good revenue for companies but also democracy. There are people that talk about all these things, and I just don't think they have much presence in Silicon Valley. Maybe that's for good—Silicon Valley lets them go, they'll burn a lot of money, cause headaches, but also create things like search engines. I think a lot of companies are focused on creating value. Amazon is different from Meta. Amazon has a business model: bring packages to doors, and behind that, they create technology to support it, mostly doing good. You have to worry about labor markets, but those are good things to worry about. Just creating computational artifacts that make predictions, and if you wear goggles, you live in their world—that's not a business model. It's a science fiction dream that may not be helpful for humanity.
那么,我们能探讨一下吗?因为你提到了 Spotify 的例子。我们之前讨论过这个三层结构,现在我们有了这种奇怪的激励结构,Spotify 实际上有动力用 AI 生成歌曲,对吧?
Well, can we explore that because you gave the example of Spotify. We were talking about this three-layer thing before, and now we've got this weird incentive structure where Spotify is actually incentivized to generate songs with AI, right?
是的,他们有。我有一个项目——我是 United Masters 的科学顾问,这是一个替代方案,让音乐人保留自己的作品。United Masters 将他们与品牌和其他机会连接起来,使他们更像真正的艺术家,而不仅仅是歌曲被播放赚点小钱。Spotify 接近垄断,但没有付费的动力——存在垄断定价。人们希望市场能解决这个问题,足够多的年轻艺术家会说‘我在这里被坑了’,然后另一个服务会出现。但我们处于一个某些服务很快成为垄断的时代。我把这留给我的经济学家朋友去思考,回顾历史例子,考虑是否需要监管或其他做市机制。我不反对 Spotify,但它应该是一个更奖励艺术家的生态系统的一部分。现在,艺术家收入很少,我不相信价格是在竞争机制下设定的。有一个更广泛的宏观经济视角,看这些系统在做什么以及它们在社会中的角色。对于搜索引擎,我们很多人困惑它们如何赚钱。整个广告业务是个惊喜,至少对我来说,它变得如此巨大。根本原因是人们期望免费的东西。谷歌无法收费。我认为他们在 YouTube 上犯了个错误。当他们收购 YouTube 时,YouTube 激励创作者创作人们会看的内容。那时,一个有社会责任的谷歌会说:‘我们在这里创造了一个市场,一个生产者-消费者关系。我们必须让这个市场更有效,这样当有人观看时,他们与创作者有经济联系,激励流动,使创作者有动力创作更多。’
Yeah, they are. I have a project—I'm a scientific adviser to something called United Masters, which is an alternative that lets musicians keep their work. United Masters connects them to brands and other opportunities so they are more like real artists, not just getting their song streamed for a little money. Spotify is close to a monopoly, but it's not incentivized to pay—there's monopoly pricing. One would hope the market will fix that, that enough young artists will say 'I'm getting screwed here' and another service will emerge. But we are in an era where some services become monopolies pretty quick. I leave that to my economist friends to think through, look at historical examples, and consider if regulation is needed or other market-making mechanisms. I'm not against Spotify, but it should be part of an ecosystem that rewards the artist more. Right now, an artist gets paid very little, and I don't believe prices are set under competitive mechanisms. There's a broader macroeconomic view of what these systems are doing and their role in society. With the search engine, many of us were puzzled about how they would make money. The whole advertising thing was a surprise, at least to me, that it became so huge. The underlying thing is that people expect things for free. Google couldn't make payments. I think they made a mistake with YouTube. When they acquired YouTube, YouTube incentivizes creators to create things people will watch. At that point, a socially responsible Google would have said, 'We've created a market here, a producer-consumer relationship. We've got to make that market more valid, so when someone watches, they have an economic connection to the creator, and incentives flow so the creator is incentivized to make more.'
相反,一切都通过谷歌,然后谷歌在旁边放广告为自己赚大钱,只有一点点回馈的动机。在我看来,这是一个巨大的错误,然后 Facebook 让情况更糟。所以你与像杰弗里·辛顿和伯克利的斯图尔特·罗素这样的人发生过冲突,他们描绘了一幅画面:这项技术是递归自我改进的,它是本质性的,不是一种文化技术,而是一个独立的事物。这初看有点科幻。
Instead it was all going through Google and then Google was putting advertisers next to make a ton of money for themselves and then there's a modest incentive to give back a little bit of money. That to me was a huge mistake and then Facebook made it even worse. So you've butted up against folks like Jeffrey Hinton and Stuart Russell at Berkeley and these guys are painting a picture that this technology is recursively self-improving, that it is essential, that it's not a cultural technology, it's a thing in itself, and this seems a little bit science fiction on the first read.
非常科幻。
Very science fiction.
你怎么看?
What do you think?
我认为这是科幻,科幻对社会很重要,但以这种被推广的水平和那些声音,它真的在伤害 25 岁和 20 岁的年轻人。你知道,这些年轻人数量庞大,他们对科技感到兴奋,想建造能帮助家人和国家的产品。实际上,更多是帮助家人而非国家,老实说。他们看到了真正的机会,但却被领导者告知:我们玩过了,我们开发了一堆算法,我们只是对纯粹理解智能感兴趣。尽管他们并不理解智能,他们构建了梯度下降算法。现在你们不能做这个,因为它很危险,它很可能会毁灭人类,或者超级智能很快到来,所以没什么可做的了。这发生在你们有生之年。这太令人沮丧了,太令人沮丧了。我认为这是最困扰我的。第二部分困扰我的是,那里完全没有经济思考,零。这实际上是认知科学或神经科学的心态:我们弄清了大脑如何工作,它是梯度下降加上大量分布式神经元,这些语言模型工作得这么好,证明我们搞清楚了,否则不会这么好。我认为这很可疑。大脑远不止如此。你问神经科学家这和大脑有没有关系,他们基本会说没有。这是一个不错的比喻,一个卡通。梯度下降在大规模下有效吗?是的,远超我们想象。但它是否暴露了弱点?是的。能在某些领域修复吗?是的。你构建某些垂直领域,它们会做好事,会让数学家更快,但不会让他们失业等等。它对社会影响很大。我更担心劳动和资本的关系,而不是它决定接管。所以其余部分对我来说更接地气:下一代年轻人如何利用技术并与之合作。我不认为那样的声音真正帮助那一代人去感知他们应该做什么以及为什么。超级智能对灭绝,这是你的两个选项。该死的,那不是仅有的两个选项。有大量非常积极的事情可以在人类尺度上完成。希望足够多的年轻人能支持这一点。但他们没有足够的榜样:那些通过制造东西赚钱的人——山姆让生活更好了吗?不清楚。我认为在之前几代人中,有更多像这样的人:他们在制造疫苗之类的东西。哦,我想成为那样。而现在,不太好。
So I think it's science fiction and I think science fiction is important for society, but it's also at the level it's being promoted and those kind of voices, it's really hurting 25 and 20 year olds. You know, these young folks of whom there are huge numbers are excited about technology and they want to build things that help their family and help their country. Actually, more of their family than their country, honestly. And they see real opportunities in doing that and they're kind of being told by the leaders, well, we had our fun. We developed a bunch of algorithms. We did it and we were just interested in the pure, you know, understand intelligence. Even though they didn't understand intelligence, they built gradient descent algorithms. And now you guys, you can't do this because it's dangerous. It's going to wipe out humanity with a high probability or super intelligence will arrive soon, so there's nothing left to do. That's in your lifetime. That is so demoralizing. So demoralizing. And that thing I think that bothers me the most. I mean, the second part that bothers me is there's no economic thinking going on there. It's zero. It's really about cognitive science mentality or neuroscience. We figured out how the brain works. It's gradient descent with a lot of distributed neurons and the fact that these LMs are working so well shows that we figured it out. It wouldn't work so well otherwise. Well, I think that's dubious. We don't—the brain is way beyond. I mean you ask a neuroscientist if this has anything to do with the brain, basically they'll say no. It's a nice metaphor. It's a cartoon. Does gradient descent work at massive scale? Yeah, more than we would have ever imagined. But is it showing its weaknesses? Yeah. Can it be fixed in certain areas? Yeah. You know, you build certain verticals, they'll do good things and it'll make mathematicians go faster, but it won't put them out of business and so on. It's having a big effect on society. I worry more about labor and capital relationships than I worry about it deciding to take over. So the rest of that to me is more on the ground, sort of how does a young the next generation take technology and work with it, and I don't think that voices like that are actually helping that generation to actually perceive what they should work on and why. Super intelligence versus extinction. Those are your two options. And damn it, those aren't the only two options. There's a huge number of very positive things that can be done at human scale. And let's hope that enough of the young mentalities kind of get behind that. But they don't have enough examples of people out there who made money by making—did Sam make life better? You know, not clear. And I think in previous generations there was a little bit more, you know, here's people that are out there making things that—vaccines or whatever. Oh, I want to be like that. And right now, not so good.
我不知道能否让你当一分钟心理学家,试着理解为什么这些——我不知道是否是寻求目的,但你是否也注意到,有些人认为会有乌托邦式的未来,当你和他们交谈时,他们有非常相似的 DNA。他们也认为它是递归自我改进的,会变成超级智能等等。如果非要你说为什么——我能理解这些东西很聪明,但他们为什么相信呢?
I don't know if I could get you to be a psychologist for a minute and try and understand why these—I don't know whether it's the search for purpose, but have you noticed as well that some folks think that there's going to be a utopian future and when you speak with them, they have quite a similar DNA. So they also think that it's recursively self-improving, it's going to be super intelligence and so on. If I was to press you to say why—I can understand these things are so clever, but why do they believe it?
我的意思是,它们在某种程度上以可识别的方式很聪明。它们不是——它们是把所有人类的聪明才智用新方式包装起来。而且,我认为它总会有点缺失重点,因为它不是当下的、不是短暂的东西。但这并不意味着它不能更聪明,我觉得这没问题。对我来说,目标从来不是构建一个超级智能并让它发号施令。从来不是。我有点震惊,有些人似乎认为那一直是目标。对我来说从来不是。相反,我一开始就说过,人类是美妙的。你知道,我讨厌机器人取代我们,我不认为那会发生。人性中有太多美好,人类能产出的东西惊人地美丽、有创意和鼓舞人心。我们需要支持这一切。但问题是,另一面是我们远非完美。人们确实伤害了很多人,而且他们被赋予更多力量去做更多伤害。我们非常狭隘。我们也不理解——人们经常因为不理解对方的动机而伤害别人。他们误解了。有多少战争是因为有人不理解对方的意图,然后说,让我们先发制人轰炸他们。这经常发生。这就是人类的行为和思考方式。缺失的是对不确定性和信息信号的 appreciation,最终博弈论出现帮助人们思考,但仍然非常粗糙。看看我们的政治体系,一个年迈的江湖骗子领导一个国家——这是我们优化的人类最高决策系统。人类还有很大的改进空间。民主必须是道路。但现在的民主是几个老人坐在各个首都的圆形大厅里,大多不知道自己在说什么。我们在许多领域有非常破碎的人类系统。有一些好的——我认为大学相当好,很多公司相当好,很多各种技能的人类协会相当好,但我们有太多破碎的。所以对我来说,这就是 AI 的意义所在。
I mean they're clever in a recognizable way at some level. They're not—they're taking all this human cleverness and packaging it in a new way. And again, I think it'll kind of always be missing a little bit of the point because it's not in the moment, it's not the ephemeral stuff. But that doesn't mean it can't be even more clever, and I kind of think that that's okay. I think that for me the goal here is not to build a super intelligence and have it dictate or tell or anything like that. It never was. And I'm kind of shocked that some people seem to think that was always the goal. It to me it just never was. Rather, again, I think I said this in the very beginning. Humans are wonderful. You know, I'd hate to have robots taking over from us and I don't think that's going to happen. There's just too much good about human nature and about what humans are able to produce that are shockingly beautiful and creative and inspiring. We need to support all that. Now, the issue though is that the flip side is that we are far from perfect. People really hurt a lot of people and they are being empowered to do yet more of it. And we are very narrow-minded. We also don't understand—people hurt other people often because they don't understand their motivations. They got a misunderstanding. How many wars are created because someone didn't understand the intentions of the other side and they said, let's just proactively bomb them. That's all the time. That's how humans act and think. What's missing there is an appreciation of uncertainty and information signaling, and eventually game theory arose to help people think it through a little bit, but it's still extremely rough. And if you look at our political system, an aged charlatan leading a country—this is our optimized human system for making decisions at the highest kind. There's so much room for improvement of the human being. And democracy's got to be the way. But democracies right now are a few aged people sitting in various rotundas in various capitals, not knowing what they're talking about mostly. We have a very broken human system in many domains. We have a few that are—I think the universities are pretty good and I think a lot of companies are pretty good and a lot of human associations of various skills are pretty good, but we have so many broken ones. And so to me that's what AI is about.
AI 是关于帮助人类处理那些太难的事情,并辅助信息流动,这样人类就能真正做出他们大多数人都想做的正确决定,而不是因为了解不足而被迫做出他们害怕的错误决定。如果你从这个层面思考,机会太多了。这就是我眼中的 AI。AI 不是用计算机取代人类,也不是递归自我改进。那感觉像是一种隐喻。我们确实使用递归算法和改进算法,但我不认为它会像病毒一样失控。我们将与这些系统合作,并希望专注于纠正进化未能为人类妥善处理的一些事情,尤其是在 70 亿人口的规模下。进化可能没有为此做好准备。专注于这一点才是 AI 的意义所在。因此,我对此持积极态度,看好 AI。我对当前两种对话感到震惊:一边是那些拥有大量资金、只为建造而建造的人,另一边是反智地声称 AI 可怕、会毁灭人类的人。这就是现在公众眼中的对话,我觉得这非常有害。让我困扰的是,那些多年来从事 AI 研究的人认为我们已经走到了尽头,认为梯度下降就像大脑,你可以把多个大脑融合在一起,它就会做出无法计算的事情。那是科幻小说。无论它是否真实,都不值得思考。路径是什么?我们如何让年轻人参与做积极的事情?要讨论什么机制、教育、目标?思想领袖们并没有谈论这些语言。这在人类历史上是不寻常的,思想领袖们正朝着这两个方向分道扬镳。
AI is about helping with things that were too hard for humans and aiding the information flow so humans could actually make the good decisions that most of them really wanted to make, instead of making the bad decisions they were afraid they had to make because they didn't know enough. There's so much opportunity if you think about it at that level. That's what AI is about to me. AI is not about replacing humans with computers or recursive self-improvement. That feels like a metaphor. We work with recursive algorithms and improving algorithms, but I don't see that getting out of control like a virus. We're going to work with these systems and hopefully focus on getting right some of the things that evolution didn't quite get right for humans, especially at a scale of 7 billion. Evolution perhaps didn't prepare for that. Focusing on that is what AI can be about. So I'm positive and bullish about AI in that sense. I'm appalled by the dialogue between the people who have all the money and want to build for the sake of building, and the people who are anti-intellectual and say it's terrible and will destroy humanity. That's the dialogue in the public eye right now, and I find it so harmful. It bothers me that people who worked on it for all these years think we've reached the end, that gradient descent is like the brain, and you can take multiple brains and fuse them together and it will do uncalculable things. That's science fiction. Whether it's true or not, it's not worth thinking about. What's the path? How do we engage younger people to do positive things? What mechanisms, education, goals? The thought leaders are not talking that language. It's unusual in human history that thought leaders are heading off in these two directions.
这也很复杂,因为任何自主软件在没有直接人类监督的情况下运行都存在非常现实的安全风险。
It's also complex because there are very real security and safety risks of having any autonomous software doing things without direct human supervision.
嗯,是也不是。想想飞机。这是个经典例子。我小时候飞机失事很多,但现在大规模事故很少了。这是因为自动驾驶仪。现在飞机大多由自动驾驶仪飞行,人类在需要时介入。这种自动化与人类的结合实际上是最有效的方式。它正在改进。人类并非生来就能在空中驾驶这么大的东西,所以你可以在这方面提升人类能力。将两者结合,就能做出对大家都有益的事情。
Well, yes and no. Think about airplanes. That's the classic example. There used to be a lot of plane crashes when I was a kid, but now very few at massive scale. It's because of autopilots. Mostly planes are flown by autopilots, and the human can come in as needed. That blend of automation with human is actually the most effective way. It's improving. Humans didn't evolve to fly a big thing in the air, so you can improve upon human ability there. Put the two together, you can do something helpful for everybody.
我想在这种情况下,这是一个相当明确的问题。我们想从 A 到 B,参数是给定的。
I suppose in that case it's quite a well-specified problem. We want to go from A to B and here are the parameters.
是的。但你有空中多架飞机、云层、变化的天气模式、某个做了蠢事的人。这更容易,因为空中空间很大。在 3D 中,空间比 2D 大得多。但在 2D 中,你有所有这些汽车飞来飞去,每个国家每年有数万人死亡。这在某种程度上是一团糟,尽管它对很多人来说非常重要和有效。一个具有大量自主性和一些人类参与的混合系统才是正确的方向。但你必须从系统层面思考。仅仅把超级智能放在汽车方向盘后面是一种愚蠢的技术思考方式。
Yes. But you have multiple planes in the air, clouds, changing weather patterns, someone who did something stupid. It's easier because up in the air there's a lot of room. In 3D, there's more room than in 2D. But in 2D, you have all these cars flying around, tens of thousands of people dying each year in each country. It's a mess at some level, even though it's very important and effective for many of us. A hybrid system with a lot of autonomy and some human involvement is the way to go. But you have to think at the system level. Just putting a superintelligence behind the wheel of a car is a dumb way to think about technology.
还有希望吗?什么能让这些人改变想法?
Is there any hope? What would make these folks update?
我认为 Ilya 和其他人做了一些伟大的事情。他们构建的系统不仅被我们所有人使用,而且正在改变我们的思维方式。我是一个建造者,不是大师或思想家。也许我更擅长建造。有了现在可用的资源——不仅仅是金钱,还有整个互联网和前人做的一切——有一件事让我非常困扰,不是 Elon Musk 或 Sam Altman,而是他们只是进来从人们付出的所有努力中攫取最精华的部分。很多人感到恼火,因为事情朝着这个方向发展,而没有理解这些人当初建造这些东西的初衷。他们心中有其他目标。所以,是的,这些人中有一些建造者,但也有一些非常爱出风头的建造者。有一个我们称之为硅谷的系统,在那里,你的语言越离谱、越天马行空、越涉及物理、生物、神经科学,你就越像大师。人们喜欢那种姿态。它创造了大量财富。他们可能不在乎财富,但这让他们更加突出,这样他们就可以再开一家公司尝试其他疯狂的事情。如果失败了,那说明你有一个好主意。我不想身处那个世界。我正试图成为一名历史学家。回顾历史,有过这类事情的闪光点,但这种脱离现实的程度在人类历史上是不寻常的。这种‘我疯狂的科幻小说般的 25 岁梦想就是我余生要追求的,无论遇到什么困难’——然后某个时刻他们转变了,因为他们意识到自己并没有一个伟大的目标,花了很多钱,得到了一个不知道该怎么办的东西,还为此担忧。对我来说,这是某种不成熟的表现。
I think Ilya and others have done some great things. They built systems that all of us are not only using, but that are changing our thinking. I'm a builder, not a guru or thinker. Maybe I'm better as a builder. With the resources available now—not just money, but the whole internet and everything previous generations did—one thing that bothers me a lot about these people, not Elon Musk or Sam Altman, is that they're just coming in and taking the cream off the top from all this effort that people put in. Many of these people are annoyed that this is the direction it's taken, without appreciation of why these people were building these things. They had other goals in mind. So yes, these are some builders, but there are some very pressy builders. There's this system we call Silicon Valley where the more outrageous, far-flung, physics-biology-neuroscience-inflected your language is, the more you sound like a guru. People enjoy that posture. It creates a great amount of money. They don't care about wealth perhaps, but it allows them to be more prominent, so they can have another company that tries some other crazy thing. If it doesn't work, that's a sign of a great idea. I wouldn't want to be in that world. I'm trying to become a bit of a historian. Looking back at history, there were glimmers of these kinds of things, but this level of detachment from reality is unusual for human history. This level of 'my crazy science fiction 25-year-old dreams are all I'm going to pursue for the rest of my life, come hell or high water'—and then at some point they flip because they realize they didn't have a great goal in mind, and they've spent a lot of money and got this thing they don't know what to do with and are worried about it. That to me is a sign of a certain level of immaturity.
坦白说,回到统计契约理论,它涉及信息不对称和激励建模。很多听众都听说过博弈论。有什么区别?
Frankly, circling back to this statistical contract theory, which is about information asymmetry and modeling incentives. A lot of folks in the audience would have heard of game theory. What's the difference?
博弈论是一门数学学科,始于 20 世纪 20 年代的冯·诺依曼,有很多分支。它是一种数学思维方式。我喜欢把它比作 F=ma。它做出预测。如果我写下一个博弈,我可以计算纳什均衡或相关均衡,并预测会发生什么。在物理学中,F=ma 预测抛物线轨迹。在博弈论中,我们检查均衡是否刻画了系统和人的行为。有时是,有时不是。还有很多其他均衡:斯塔克尔伯格均衡、序贯均衡等,以及各种社会福利和遗憾构造。这是一个巨大的领域,最终会像物理学一样大,涉及战略互动。物理学的逆问题是工程学:我想建一座桥,所以反转 F=ma,从目标到设计。大多数工程领域都是逆问题。博弈论的逆问题是机制设计。机制设计说:我想要某个结果——公平、市场——我设计什么博弈来实现它?我是设计者。契约理论是机制设计的一部分,处理信息不对称。拍卖理论是另一部分,机制揭示竞拍者的价值。所以博弈论是一个丰富、有 100 年历史的学科,持续发展并提供算法思想。我职业生涯中主要是统计学家,关注不确定性,但当我涉及均衡和博弈时,博弈论是思考的一部分。
Game theory is a mathematical discipline. It started with von Neumann in the 1920s and has many branches. It's a mathematical way of thinking. One way I like to think about it is like F=ma. It makes predictions. If I write down a game, I can calculate Nash equilibria or correlated equilibria and predict what will happen. In physics, F=ma predicts a parabolic curve. In game theory, we check if equilibria characterize how systems and people behave. Sometimes yes, sometimes no. There are many other equilibria: Stackelberg, sequential, etc., and various social welfare and regret constructs. It's a huge field, eventually as big as physics, about strategic interactions. The inverse question in physics is engineering: I want to build a bridge, so I invert F=ma to go from goal to design. Most engineering fields are inverse problems. The inverse of game theory is mechanism design. Mechanism design says: I want a certain outcome—fairness, a market—what game do I design to realize that? I'm the designer. Contract theory is a part of mechanism design, dealing with asymmetric information. Auction theory is another part, where a mechanism reveals bidders' values. So game theory is a rich, 100-year-old discipline that continues to evolve and supply algorithmic ideas. I've been mostly a statistician, worried about uncertainty, but when I go to equilibria and games, game theory is part of the thinking.
你说过我们需要考虑激励、集体和不确定性量化。机器学习中有一个很棒的领域叫共形预测,由我在伦敦大学学院的教授弗拉基米尔·瓦普尼克发明。我们学习了转导置信机和奇异度度量,比如 SVM 中到超平面的距离,来计算类似 p 值的东西,并得到置信区域。
You've said we need to think about incentives, collectives, and uncertainty quantification. There's a wonderful field in machine learning called conformal prediction, invented by my professor at University College London, Vladimir Vapnik. We learned about the transductive confidence machine and measures of strangeness, like distance from a hyperplane in SVM, to calculate something like a p-value and have a confidence region.
是 E 值。实际上,我不想深入讨论 E 值的技术细节,但弗拉基米尔很了不起。我认为他是一位统计学家,也有博弈论背景,属于菲尔·道威德和大卫·布莱克威尔那一派,他们从统计学溢出做其他事情。经典地,p 值是一次性量:你有一个模型,它给出概率分布,一个结果出现,看起来不太可能,所以模型一定是错的。这就是 p 值,一个尾部概率。问题是如果你重复这样做,并沿途看最小的 p 值,那就是 p 值操纵,给出错误答案。E 值不同。它是一个非负随机变量或非负上鞅的期望。你观察证据累积,并确保每一步期望≤1。然后你有多重证据收集:如果它总是指数级≤1,它就保持在 1 以下;如果非负,它会衰减。在原假设下,这个随机过程衰减。你可以随时观察,反复观察,并保持控制。
An E-value. Actually, I don't want to get into technical talk about E-values, but Vladimir is fantastic. I think of him as a statistician with a game theory background, in the school of Phil Dawid and David Blackwell, who spilled out of statistics to do other things. Classically, p-values are a one-shot quantity: you have a model, it gives a probability distribution, an outcome arrives that looks improbable, so the model must be wrong. That's the p-value, a tail probability. The problem is if you do that repeatedly and look at the smallest p-value along the way, that's p-hacking, giving wrong answers. E-values are different. It's an expectation of a non-negative random variable or a non-negative supermartingale. You watch evidence accrue and ensure the expectation is ≤1 at each step. Then you have multiplicative evidence gathering: if it's always exponential ≤1, it stays below one; if non-negative, it decays. Under the null hypothesis, this stochastic process decays. You can look at it at any time, repeatedly, and have control.
有一种叫维尔不等式的东西,弗拉基米尔和其他人利用了它,说可以在整个路径上控制它。所以现在我们可以用一种新的方式做统计,叫做‘随时推理’。我们可以偷看、可以改变、可以收集新数据。我们可以非常自由地这样做。弗拉基米尔是这方面的领导者之一。而 e 值就是那些在特定时间由可选停止定理停止的鞅之一。你可以随时停止它。这开启了很多联系。事实上,我们的统计契约理论。什么是契约?记得它就像服务和价格。服务就像证据收集。价格也是一个随机变量。结果发现,在契约领域,激励相容性成立当且仅当在统计领域有 e 值。所以博弈论概率和激励理论之间有一个紧密的联系。所以对我来说,不确定性量化很少只是‘这是一个误差条’。那是经典统计学。它更多是关于上下文。上下文可能是一个契约或其他证据收集机制。这种思维方式让你对更广泛的证据收集类别开放。
There's something called Ville's inequality that Vladimir and others have exploited that says that can be controlled over the entire path of this thing. So now we can do statistics in a new way. It's called anytime inference. We can peek, we can change, we can gather new data. We can do this in a very liberating way. Vladimir is one of the leaders of that. And e-values are one of those martingales stopped at a particular time by the optional stopping theorem. You can stop it whenever you want. So that has opened up a lot of connections. In fact, our statistical contract theory. What is a contract? Remember it was like services and prices. Well, the services are like evidence gathering. And the price also is a random variable. And it turns out that we can have incentive compatibility in contract land if and only if e-value in statistics land. So there's a nice tight connection between game-theoretic probability and the theory of incentives. So to me, uncertainty quantification is rarely just 'here's an error bar.' That's kind of classical statistics. It's more about the context. The context might be a contract or some other evidence-gathering mechanism. This way of thinking opens you up to a broader class of evidence gathering.
非常酷。我应该说在你的论文中,你有这个三角形的图,我们现在把它放在屏幕上。但你是在说,有经济学、计算机科学和统计学。
Very cool. And I should say in your paper you have this figure of a triangle which we'll put on the screen now. But you're kind of saying there's economics, computer science, and statistics.
嗯,这些都是思维风格。我甚至用这些学科来称呼它们。几十年前,Jeannette Wing 有一篇论文谈到计算思维。它说计算机科学发展出了这些思维风格,比单纯的计算机更抽象。就像模块化、抽象、API 等等。为什么我们不教所有科学和学科的人做计算思维呢?我认为这完全正确,很棒。但很多算法并非来自那些计算机科学原理。它们来自对推理不确定性的思考:我如何收集数据来预测尚未存在的事物?以及思考激励:我如何确保激励到位?我把这两种思维称为:推理思维(不仅仅是统计学,很多领域都有推理)和经济思维(不仅仅是经济学,各种社会科学家、法律学者等)。当你把这三者放在一起,你就得到了一个很好的平台,用于培训下一代和解决我们一直在讨论的那类问题。仅仅一个领域——计算逻辑和优化——给了我们语言模型,很好,但没有给出语言模型周围的任何上下文。激励给了我们一直在谈论的所有东西。而统计学对我来说至关重要。它思考我必须犯什么样的错误,如何确保数据被控制以避免错误。我们把三者放在一起。它们也带来了合作伙伴:经济学家与行为心理学家交谈,计算机科学家与物理学家交谈,统计学家与法律人士交谈。有整个子社区聚集在一起。所以对我来说,如果你戴上这个三角形并围绕它思考,它开始成为一种思考学术的新方式。这是这个时代的博雅教育。这是核心。我的人文学科同事可能不同意——核心仍然是人文——但我只是认为它没有触及这个时代的核心智力问题,这些问题关乎数据和计算等等。但我想把要素放在适当的位置,以便以对社会负责的方式思考这些事情。
Well, these are thinking styles. I even call them by those disciplines. There was a paper by Jeannette Wing a couple of decades ago talking about computational thinking. It says computer science has developed these thinking styles that are more abstract than just computers. It's like modularity, abstractions, APIs, and all that. And why don't we teach everybody in all the sciences and disciplines to do computational thinking? I think that's totally right, that's great. But lots of algorithms don't come about from those computer science principles. They come about from thinking about inferential uncertainty: how do I gather data to make predictions about things that don't yet exist? And think about incentives: how do I make sure incentives are in place? I called those two kinds of thinking: inferential thinking (not just statistics, a lot of fields have inference) and economic thinking (not just economics, social scientists of all kinds, legal scholars, etc.). When you put those three together, you get a pretty good platform for training the next generation and for problem solving of the kinds we've been talking about. Just one of the fields—computational logic and optimization—gives us LMs, fine, but doesn't give us any context around the LM. The incentives give you the whole things we've been talking about. And statistics to me is critical. It thinks about what kind of errors I have to make, how to make sure the data is controlled so I don't make errors. We put the three together. They also bring partners: economists talk to behavioral psychologists, computer scientists talk to physicists, statisticians talk to legal people. There are whole subcommunities that come together. So to me, if you put on this triangle and think around it, it starts to become a new way to think about academia. This is the liberal arts of the era. This is the core. My colleagues in the humanities might disagree—the core is still the humanities—but I just don't think it's touching the core intellectual issues of the era, which are about data and compute and all. But I want to put the ingredients in place so that those things are thought about in a socially responsible way.
但你能把这个讲得生动一些吗?你曾著名地谈到一个语言模型,我问它对自己的答案有多自信,它往往非常模态化。所以它要么是 1 要么是 0 或者不是。有什么区别?为什么语言模型真的不知道自己的置信度?
But could you bring this to life? So you famously spoke about a language model and I'm going to ask it how confident it is about the answer, and it tends to be quite modal. So it'll either be like one zero or not. What's the difference? Why does the language model not really have any idea about its confidence?
你应该问语言模型的构建者,因为他们所做的只是预测下一个词,其中没有任何关于不确定性量化的思考。你可以嫁接一些想法,但它们很可疑——他们经常加入可疑的先验。所以你去找统计学家,人们已经这样做了:他们说,‘好吧,我可以把它当作一个黑箱,并在它周围加上共形预测。’这是一个很好的方法。不需要很多假设。是的,没错。但它做了一个可交换性假设:数据,如果你打乱它,它是一样的。虽然我认为这一切都非常关键和重要,但我倾向于更多地思考更广泛的背景。所以我在你提到的那篇文章中举了一个例子,关于一只鸭子去湖边,一只统计学家鸭子。它计算了过去一年中,湖的那一边的谷物通常是这一边的两倍——2:1 的比例。第二天鸭子需要决定去湖的哪一边。贝叶斯鸭子拥有这些概率,会做最大期望值,并以概率 1 去左边。但真正的鸭子不那样做。它们可能以 2/3 的概率去那边,1/3 去另一边。它们在对冲。但这不仅仅是对冲;对冲只是偶尔去另一边。它们实际上得到了正确的比例。解释是,你没有考虑这个不确定性的背景。不只是你,个体鸭子。可能你是在一个有很多鸭子的世界中进化而来的。如果所有鸭子都去同一边,你就错过了资源。那么有没有一种算法允许许多鸭子合作?如果它们都有同样的不确定性,那么它们可以以 2/3 的概率采样去这边,1/3 去那边。这实际上是一个更大系统的纳什均衡。
You should ask the language model builders, because all they're doing is predicting the next word, and there's not any thinking about uncertainty quantification in doing that. You can graft in ideas, but they're dubious—they're often putting in dubious priors. So you go to the statistician, and people have done that: they said, 'Okay, I can just treat it as a black box and put conformal prediction around it.' It's a nice method. Doesn't require a lot of assumptions. Yes, that's true. But it makes an exchangeability assumption: the data, if you scramble it, it's the same. While I think all that's really crucial and important, I tend to think more about the broader context. So I gave an example in that article you mentioned of a duck who goes to a lake, a statistician duck. It's calculated that over the last year there tends to be twice as much grain on that side of the lake than on this side—a 2:1 ratio. The next day the duck needs to decide which side of the lake to go to. The Bayesian duck, who has those probabilities, would then do maximal expected value and go to the left side with probability one. But actual ducks don't do that. They go probably 2/3 to that side and 1/3 to the other side. They're hedging. But it's not just hedging; hedging would just occasionally go to the other side. They're actually getting the right ratio. The explanation is that you weren't thinking about the context of this uncertainty. It's not just you, the individual duck. Probably you evolved in a world where there are many ducks. If all the ducks went to the same side, you've missed out on a resource. So is there an algorithm that allows many ducks to cooperate? If they all have that same uncertainty, then they can sample with probability 2/3 to this side versus 1/3. And that's actually a Nash equilibrium of the bigger system.
所以,思考不确定性的正确方式是,在总体的背景下,我应该如何运用我的不确定性?另一种不确定性是经济层面的,我称之为信息不对称。你知道,有些事情我不知道,而你有我不知道的专业知识,但我们要合作,我可能会给你一份合同,一个选项菜单。但即使我和你互动一段时间,我可能仍然不知道。有些事情你知道,但不会透露给我,也许你会有所保留,甚至撒点小谎。所以我不知道那是什么。那不仅仅是抽样问题,那是另一种不确定性。最后,还有一种我称之为天意的不确定性。那更像是数据库类型的不确定性。如果我想做一次医疗手术,你是医生,你查看了像我这样的病人的数据,这里显示如果这样操作生存概率是多少,那样操作又是多少。我看了之后觉得很好,但你又告诉我那些数据是十年前收集的。我就会说,好吧,我的置信区间应该扩大。经典统计学可以讨论这个。事实上,我更像一个贝叶斯主义者来思考这个问题,但经典统计学并不这样做。它只是把数据当作数据。而数据应该在一个更大的系统中流动,它应该始终带有关于数据年龄的元数据标签,并且这些信息应该被定量地纳入不确定性量化中。我们现在根本没有这样做。所以,可怜的 LLM,它们基本上没有做到以上任何一点,如果它们要开始像人类一样行事,就必须在这些方向上有所突破。我们人类很擅长处理这些,带有一点天意。哦,这是旧数据,我会打折扣。我们得到一点上下文。哦,这里有一个社会环境。我应该做同样的事情,应该随机化。哦,还有一些抽样不确定性,等等。我们几乎无缝地把所有这些整合在一起。然后我们在一个社会环境中这样做,如果我不知道如何从这里到城市的另一边,我会问一个看起来像丹麦人的人。我知道如何收集更多数据等等。所以可怜的 LLM 没有以上任何能力。那么,当你问它有多确定时,它应该说什么?据我所知,它所做的只是,过去有人在互联网上问一个人,你刚刚写下的那个方程有多确定,然后有人说,哦,我非常确定,因为这样或那样的原因。我认为它只是模仿了那些断言,但那不是不确定性下的推理。
So the right way to think about uncertainty there is that in the context of the population, how should I use my uncertainty? Another kind of uncertainty is the economic side, which I've alluded to as information asymmetry. You know, things I don't know and you have expertise I don't know about, but we're going to work together and I'll maybe give you a contract, a menu of options. But even if I interact with you for a while, I still might not know. There are things you're going to know that you're not going to give away to me, and maybe you'll hedge, you know, you'll lie a little bit. So I don't know about that. That's not just sampling. That's a different kind of uncertainty. Okay. And then finally, there's what I like to call providence. That's more like a database kind of uncertainty. If I want to do a medical operation and you're a doctor, you look at the data for people like me, here's the probability of survival if you do the operation this way versus that way. I look at that and say great, but now you tell me that data was gathered 10 years ago. And I'm going to say, okay, my confidence interval should go up. Well, classical statistics could talk about that. In fact, I'd be more of a Bayesian to think about that, but it doesn't. It just sort of treats the data as the data. And it should be in a bigger system where data is flowing around. It should always be tagged with metadata about how old it is, and that should be quantitatively brought into the uncertainty quantification. We're not doing anything like that right now. And so the poor LLMs, which are basically doing none of the above, have to strike out a little bit in all these directions if they're going to start to do what humans do. We are pretty good at getting these with a little bit of providence. Oh, it's old data, I discount that. We get a little bit of context. Oh, there's a social environment here. I should just do the same thing, should randomize. Oh, there's some sampling uncertainty, and so on. We put all that together almost seamlessly. And then we do this in a social context where if I don't know how to get from here to the other side of town, I will ask someone who looks Danish. I know something about how to gather more data and so on. So the poor LLM has none of the above. And so what should it say when you ask how sure are you? All it's doing, to the best of my knowledge, is that in the past someone asked a human on the internet how sure are you of that equation you just wrote down, and someone said, oh, I'm very sure because of this or that. And I think it just mimics those kinds of assertions, but that's not reasoning under uncertainty.
如果我们确实有了认知量化,那主要会带来什么提升?是不是关于我知道自己不知道某些事情,所以我会倾向于在那个领域进行更多的认知探索?
And if we did have epistemic quantification, what would be the main uplift from that? Is it about knowing that I don't know something, so I'm going to lean in and try to do more epistemic foraging in that area?
嗯,再次,我认为我们现在进入了统计学的领域。统计学家关心的是岛上存在哪些物种。我采样足够了吗,知道没有新物种了吗?这些都是统计学的经典领域。针对那个子群体的最优实验设计,我没有足够的数据。我在推理的背景下做出糟糕的推断和数据收集,在做出断言并重复这样做的背景下。这就是统计学长期以来关注的。所以我认为应该肯定他们处理一种主动形式的不确定性减少。但对我来说,真正的大规模不确定性减少来自于更广泛的组成部分,比如市场。如果我想开一家这样的披萨店,我需要西红柿,如果我每天都必须自己去寻找西红柿,那么那天晚上我是否能吃到披萨就很不确定。但因为存在一个市场,别人已经去寻找了,每天都有稳定的西红柿供应。我可以假设这一点成立来开我的店。我找到西红柿的不确定性降低了。因此,我可以在那之上建立并做其他事情。市场减轻了不确定性,这不是因为有人设计了最优实验设计或运行了多臂老虎机,不是直接的,而是因为市场尝试了各种东西。人们有激励去探索和利用。
Well, again, I think we're now in statistics land. The statisticians are all about what species are present on the island. Have I sampled enough to know that there's not a new species? These are classical areas of statistics. Optimal experiment design for that subpopulation, I don't have enough data. I'm making a bad inference and data collecting in the context of inference, in the context of making assertions and doing that repeatedly. That's what statistics has long focused on. So I think give them credit for handling a kind of active form of uncertainty reduction. But again, for me, real uncertainty reduction in the large comes about from much broader sets of components, like a market. If I want to have a restaurant like this for pizza and I need tomatoes, and if I had to forage for tomatoes every day, it would be pretty uncertain whether I would have pizza that evening. But because there exists a market where someone else did the foraging, there's a stable amount of tomatoes every day. I can build my restaurant assuming that's true. My uncertainty for finding tomatoes went down. Therefore, I can build on top of that and do other things. Markets mitigate uncertainty, and they don't do it because someone designed an optimal experiment design or ran some multi-armed bandit, not directly, but because the market tried various things out there. There are incentives for people to explore and exploit.
乔丹教授,非常荣幸能邀请您上节目。非常感谢。
Professor Jordan, it's been an honor having you on the show. Thank you so much.
好的。这是我的荣幸。我很享受和你聊天。
All right. It's been my pleasure. I've enjoyed talking to you.