Yejin Choi on Large Language Models: Progress, Limitations, and the Quest for Intelligence
打开互动全文版(中英对照 + 朗读 + 问答)→顶尖 NLP 研究者 Yejin Choi 探讨 GPT-4 和 ChatGPT 等大型语言模型的能力与局限,以及它们如何重塑我们对智能的理解。
Yejin Choi, a leading NLP researcher, discusses the capabilities and limitations of large language models like GPT-4 and ChatGPT, and how they reshape our understanding of intelligence.
我们今天的嘉宾是 Yejin Choi。Yejin 在康奈尔大学获得博士学位,目前是华盛顿大学的教授,也是艾伦人工智能研究所的高级研究主任。她是自然语言处理乃至人工智能领域世界顶尖的研究者之一。她的研究获得了许多奖项,包括 NeurIPS、ACL、ICML、AAAI、ICCV——几乎涵盖了所有顶级 AI 研究会议。她获奖无数,被《纽约时报》和《纽约客》报道,最近还在 TED 主舞台演讲,并获得了天才奖,正式名称为麦克阿瑟奖。Yejin,非常高兴你能来。欢迎来到我们的节目。
Our guest today is Yejin Choi. Yejin received her PhD from Cornell and is currently a professor at the University of Washington, as well as a senior research director at the Allen Institute for Artificial Intelligence. She is one of the world's leading researchers in natural language processing and more generally artificial intelligence. Her research has won many awards, including at NeurIPS, ACL, ICML, NeurIPS, AAAI, ICCV—that's pretty much all the leading AI research conferences. She's won awards, she's been featured in The New York Times and The New Yorker, she spoke at the main TED stage just recently, and she won the genius award, officially known as the MacArthur Fellowship. Yejin, so great to have you here with us. Welcome to the show.
嗯,我很兴奋能来这里。谢谢你的邀请。
Yeah, I'm excited to be here. Thank you for the invitation.
非常高兴你能来。但在开始对话之前,我想感谢我们的播客赞助商:Index Ventures 和 Weights & Biases。Index Ventures 是一家风险投资公司,投资于各个阶段的杰出创业者,从种子轮到 IPO,在旧金山、纽约和伦敦设有办事处。该公司支持多个垂直领域的创始人,包括 AI、SaaS、金融科技、安全、游戏和消费领域。就我个人而言,Index 是 Covariant 的投资方,我极力推荐他们。Weights & Biases 是一个 MLOps 平台,通过实验跟踪、模型数据集版本管理和模型管理,帮助你更快地训练更好的模型。OpenAI、Nvidia 以及几乎所有发布大型模型的实验室都在使用它。事实上,我在伯克利的所有学生和 Covariant 的同事可能都是 Weights & Biases 的重度用户。Yejin,欢迎。让我们直接进入正题。像 GPT-4 和 ChatGPT 这样的大型语言模型已经席卷了世界。似乎自从几个月前 ChatGPT 发布以来,全世界都注意到 AI 正变得极其强大。作为世界领先的自然语言处理和 AI 研究者,你正处于这一切的前沿。事实上,尽管取得了巨大的进步,你也一直在指出仍然存在的局限性。你能告诉我们这些模型到底是怎么回事,以及有哪些局限性的例子吗?
So glad to have you here. But before diving into our conversation, I'd like to thank our podcast sponsors: Index Ventures and Weights & Biases. Index Ventures is a venture capital firm that invests in exceptional entrepreneurs across all stages, from seed to IPO, with offices in San Francisco, New York, and London. The firm backs founders across a variety of verticals, including AI, SaaS, fintech, security, gaming, and consumer. On a personal note, Index is an investor in Covariant, and I couldn't recommend them any higher. Weights & Biases is an MLOps platform that helps you train better models faster with experiment tracking, model data set versioning, and model management. They are used by OpenAI, Nvidia, and almost every lab releasing a large model. In fact, probably all of my students at Berkeley and colleagues at Covariant are big users of Weights & Biases. Yejin, welcome. So let's dive right in. Large language models like GPT-4 and ChatGPT have taken the world by storm. It seems that since the release of ChatGPT just a few months ago, the whole world has taken notice that AI is becoming extremely capable. As a world-leading natural language processing and AI researcher, you're personally at the frontier of all of this. And in fact, while there has been inordinate progress, you've also been identifying remaining limitations. Can you tell us what exactly is going on with these models, and from there, what are examples of remaining limitations?
是的,大型语言模型被证明是一个非常非常强大的工具,可以泛化许多基于语言的推理问题、理解问题和各种生成问题,所有这些都使用同一个统一的基于 Transformer 的模型,当然,它是在大量互联网数据上训练的。事实证明,它确实可以学习到令人惊叹的语言模式和推理能力,从而产生非凡的实证结果。这是一个非常激动人心的时代。但让我兴奋的是,它也解锁了我们以前甚至无法想象的新智力问题,现在它让我们能够思考智能到底意味着什么,我们能识别出哪些局限性,甚至如何做到这一点。嗯,总之,这是一个非常激动人心的时代。
Yeah, so large language models turn out to be a really, really powerful tool to generalize a lot of language-based reasoning problems, understanding problems, and a variety of generation problems, all using the same one unified model based on Transformers, of course, trained on really a lot of internet data. And it turns out it can really learn amazing patterns of language as well as reasoning, such that it can enable just phenomenal empirical results. Super exciting time. But what's exciting to me is also that it unlocks new intellectual problems that we couldn't even think about before, but now it sort of enables us to think about what intelligence even means, what sort of limitations can we identify, and even how to do that. And yeah, just overall really exciting time.
你谈到这些模型让我们开始思考智能的真正含义,我想很快和你深入探讨这个问题。但在此之前,模型是在整个互联网——不是整个,而是很大一部分——上训练的,这到底意味着什么?在大量互联网数据上训练模型意味着什么?
You talk about these models allowing us to start thinking about what intelligence really means, and I want to dive into that with you shortly. But before that, models trained on the entire—not entire, but a large swath of the internet—what does that even mean? What does it mean to train a model on a lot of data from the internet?
是的,这个问题在某种意义上是对规模极限的探索。所以我非常欣赏 OpenAI 和其他一些人所做的是,看看如果你在一个非常大的模型上训练互联网上大量可用的数据会发生什么。事实证明,它可以捕捉到我们以前认为语言模型不可能捕捉到的模式。尽管学习目标归结为像预测下一个词这样简单的事情,但这些模型学到的表示是如此强大,以至于可以稍微调整一下,用于各种不同的下游语言问题或下游 NLP 应用。现在,大多数人心中最流行的应用可能是 ChatGPT,这个应用就是你与一个以某种方式读取了大部分互联网的神经网络模型聊天,并基于此给你回复,对吧?谁能想到在如此多的互联网数据上训练实际上会产生有意义的回复,因为很多人把垃圾信息放到互联网上,而不知何故,这个模型筛选了所有信息,内化了一些东西,并且在大多数情况下似乎能进行有意义的对话。
Yeah, so the question in some sense is the quest for the limit of the scale. So what I really appreciate about what OpenAI and some others did is to see what happens if you train a very large model on the vast amount of data available on the internet. And it turns out it can pick up on patterns that we thought would be impossible for language models to pick up on. And although the learning objective boils down to something as simple as predicting which word comes next, the representation that is learned by these models is so powerful that it can then be massaged a little bit so that it can be used for a variety of different downstream language problems or downstream NLP applications. Now the most popular application in most people's minds is probably ChatGPT, where the application is literally you chat with the model that somehow is a neural network that has read most of the internet, and based on that gives you responses, right? Who would have thought that training on so much of the internet would actually lead to meaningful responses, because a lot of people put garbage on the internet, and somehow this model sifts through all of it, internalizes some things, and has a meaningful conversation most of the time, it seems.
是的,谁能想到在如此多的互联网数据上训练实际上会产生有意义的回复,因为很多人把垃圾信息放到互联网上,而不知何故,这个模型筛选了所有信息,内化了一些东西,并且在大多数情况下似乎能进行有意义的对话。
Yeah, so who would have thought that training on so much of the internet would actually lead to meaningful responses, because a lot of people put garbage on the internet, and somehow this model sifts through all of it, internalizes some things, and has a meaningful conversation most of the time, it seems.
是的,从某种意义上说,既是也不是:如果你只进行互联网数据训练,那么模型通常不会做出 ChatGPT 能做的那些惊人事情。所以魔力来自于之后发生的事情。因此,在所谓的预训练之后,也就是在这个神经网络上训练大量原始互联网数据之后,接下来会发生的是,还有大量的额外训练,以微调的形式,基本上是在一些人类问答示范上进行监督训练。所以推测是,如果没有这样的数据,仅靠 RLHF(基于人类反馈的强化学习)不会那么成功,因为强化学习的探索可能无法发现 ChatGPT 使用的那些刻板的律师式语言,你知道,那不是互联网网络数据通常训练的语言。所以这非常有帮助,就像任何其他强化学习案例一样,当你一开始进行基于示范或行为克隆或基于示范的模仿学习,然后开始强化学习时,它可以学得更快。同样,这里也发生了这种情况,加上大量的基于人类反馈的强化学习,这样它实际上可以生成对最终用户来说更安全、更社交愉悦和得体的回复。所以这些都是预训练之后发生的事情。
Yeah, so yes and no in the following sense: if you just do the internet data training, then usually the model doesn't do the amazing things that ChatGPT can do. So the magic comes from what happens afterwards. So after this so-called pre-training, which is about training this neural network on a lot of raw internet data, what happens is that there's a bit of—well, I shouldn't say a bit of—there's a lot of extra training in the form of fine-tuning, which is basically supervised training on some human demonstrations of question answering. So speculation is that without such data, RLHF—so reinforcement learning with human feedback—alone is not going to be as successful, because the exploration of reinforcement learning might not be able to discover the stereotypical lawyer-like language that ChatGPT uses, you know, that's not the normal language that internet web data necessarily is trained on. And so it really helps a lot, like any other reinforcement learning case, when you do demonstration-based or behavioral cloning or demonstration-based imitation learning at the beginning and then start the reinforcement learning, then it can learn much faster. So likewise, there's that going on, plus a lot of this reinforcement learning with human feedback, so that it can actually generate responses that are a little bit safer to consume and a little bit socially more pleasant and proper to consume for end users. So this is all sort of what happens after pre-training.
最近,John Schulman 做了一个关于 ChatGPT 系统的演讲,他提出的框架,我很好奇你的看法,这个框架对你来说是否有道理。他的观点是:你打开整个互联网进行预训练,目的是让模型基本上知道所有存在的东西,但现在它完全不知道什么才是真正重要的。然后微调告诉它——不是任何新知识——而是从你已经知道的一切中,在对话时使用这些东西。你认同这个观点吗?你觉得有什么不同吗?
Recently, John Schulman gave a talk about the ChatGPT system, and the way he presented it, I'm curious about your take on this and if that framing makes sense to you. His take was: you turn on the entire internet for pre-training, and the purpose of that is for the model to essentially know everything that's out there, but now it has no clue what actually matters. And then the fine-tuning tells it—not any new knowledge—but just of everything that you already know, use these things when you have a conversation. Does that resonate with you? Does it seem different to you?
我认为这是一个非常好的框架,描述了通过强化学习发生的事情。另一种可能从稍微不同的角度说同样事情的方式是,预训练给了模型很多知识,但它不知道如何在对话上下文中恰当地使用这些知识。微调和 RLHF 帮助它学习交互的风格、语气和相关性。
I think it's a really good framing of what's happening through reinforcement learning. Another way to maybe say the same thing but slightly from a different angle might be that essentially the pre-training gives the model a lot of knowledge, but it doesn't know how to use it appropriately in a conversational context. The fine-tuning and RLHF help it learn the style, tone, and relevance for interaction.
从某种意义上说,学习目标从一开始就不太对,因为人类学习从来不是预测下一个词是什么。我们会抽象出重要的东西,并试图理解它。在这个过程中,我们还能忽略那些不可信的信息,比如阴谋论。例如,你读到阴谋论时不会突然相信它,而是会思考。同样,当你阅读 arXiv 上的论文时,你也不会盲目相信所有内容。所以人类有能力反思所读内容,并构建关于科学事实和世界运作方式的真实知识模型。而目前的预训练学习目标并没有真正鼓励神经网络以纠正的形式学习知识。因此,我们不得不对模型进行一些调整或校准,使其朝着正确的方向发展,而这正是强化学习目前能够处理的问题。
The learning objective was not quite right to begin with, in some sense, because human learning is really never about predicting which word comes next. We really abstract away what's important and try to make sense of it. In doing so, we are also able to ignore information that's likely not to be trusted, like conspiracy theories. For example, if you read one, you're not going to suddenly believe it; you are going to think about it. Similarly, when you read papers on arXiv, you don't necessarily trust everything you read. So humans do have the capability of reflecting on what they read and then building a true knowledge model about scientific facts as well as how the world works. Whereas currently, the pre-training learning objective is not really encouraging neural networks to learn knowledge in a corrective form. So we have to sort of massage the model a little bit or calibrate it into the right direction, and that's what reinforcement learning currently is able to handle.
这非常有趣。我从未这样想过,这似乎开辟了非常令人兴奋的研究方向。因为这是一个重大的开放性问题:如何重新设计预训练,使其追求真理、寻找正确性,决定哪些是值得保留的真实事实,哪些只是人们发布的内容?但就是这样。有些人喜欢在互联网上发布东西。这方面有什么进展吗?
That's very interesting. I haven't thought of it that way, and it seems to open up very exciting research directions. Because it seems like a big open question: how do you reformulate the pre-training such that it's truth-seeking, looking for correctness, deciding what might be good to retain as a true fact versus what is just stuff people put out? But that's all there is to it. Some people like to put out things on the internet. Any progress on that front?
我只能说进展不大,但这是我经常思考的问题。尤其是当我思考常识挑战时,它与关于世界的百科全书式知识有相似的挑战。知识通常有两种:一种是知道哪个演员哪年出生——我从来记不住,得用谷歌或必应搜索。另一种是关于世界如何运作的一般性常识。例如,录音时需要好麦克风,并保持房间安静。这就是关于录音的常识知识。即使没有明确写出来,我们也能推断出大量这样的知识。所以从某种意义上说,我喜欢这个挑战:仅基于互联网上的原始数据或 YouTube 视频,抽象出这些关于世界如何运作的未言明的常识知识,并让算法尝试提炼隐藏的隐含知识,使其更明确,以便人类检查。这就是下一个前沿。
I would just say not very much, but it's something that I think a lot about. Especially when I think about common sense challenges, it shares similar challenges compared to encyclopedic knowledge about the world. Knowledge generally has two flavors: one is just knowing which actor was born in what year—which I never remember, I have to Google or Bing search to find out. And then there's just general, generic knowledge about how the world works. For example, when you record, you need a good microphone and you want to keep the room quiet. This is our common sense knowledge about recording. Even if it's not written somewhere out loud, we figure out a vast amount of such knowledge. So in some sense, I love the challenge of being able to abstract away these unspoken common sense knowledge about how the world works based on just the raw internet data or YouTube videos that we can find on the web, and have an algorithm that can try to distill the hidden implicit knowledge and make it more explicit so that humans can inspect. That's the next frontier.
这就是下一个前沿。我非常欣赏你的工作的一点是,你在某种意义上拥有一种超能力,能够剖析当前模型还做不到的事情。你有一些漂亮的例子,比如 GPT-4,即使比之前的模型大得多,仍然会犯一些连三岁小孩都不会犯的错误。所以这是一个巨大的反差:GPT-4 能做我们 80% 的工作,却会犯三岁小孩都不会犯的错误。你脑海中有什么好的例子可以分享吗?GPT-4 尽管能力惊人,但仍然会失败的地方?
That's the next frontier. One thing I really love about your work is that you have a superpower in some sense of dissecting what these current models cannot do yet. You have these beautiful examples where GPT-4, the latest model, even much larger than previous ones, still makes mistakes that maybe even a three-year-old wouldn't make. So it's a big contrast: GPT-4 can do 80% of our work, yet it makes mistakes that three-year-olds wouldn't make. Are there any examples that are top of your mind that you can share where GPT-4 still fails despite all its amazing capabilities?
我最近在 TED 演讲中用的一个有趣例子是:假设你要在太阳下晒五件衣服,它们完全干透需要五个小时。那么五件衣服在太阳下完全干透需要五个小时。如果我想同时晒更多衣服,比如 30 件,需要多长时间?GPT-4 回答 30 小时。实际上还有不少例子。我还有很多例子,但很多不适合 TED 演讲那种需要简单例子的风格。不过我要指出,找这些例子有点随机。过去我用 GPT-3 时,对它会失败的地方有很好的直觉。现在用 GPT-4,我其实不知道;它为什么失败或为什么成功非常随机。所以连我也要花不少时间才能找到那些失败案例。
One of the funny examples I used in my recent TED talk was: suppose you're going to dry five clothes in the sun, and it took them five hours to dry out completely. So five clothes took five hours to dry out in the sun completely. How long would it take if I want to dry more clothes together, like 30 clothes? GPT-4 said 30 hours. There are quite a few actually. I had a lot more examples, but many of them didn't fit into the TED talk style presentation where examples should be simple. But I should note that defining these examples is a bit random. In the past, when I worked with GPT-3, I had a pretty good intuition where it would fail. These days with GPT-4, I actually don't know; it's very random why it fails or why it doesn't. So even I have to spend a good amount of time to find those failure cases.
我非常感谢你花时间,因为我认为做研究、理解下一步该做什么的最好方法,就是真正理解哪些问题还没有解决。现在很多人可能会争辩说:当然,它不知道 30 件衣服也只需要五个小时,因为你可以把它们都挂在一起,不需要一件一件地晾干。很多人可能会说,这种答案只需等待 GPT-5,在更多互联网数据上训练,也许在训练过程中加入更多或大量的人类反馈,但本质上还是更大规模的相同流程。这在某种意义上就是大多数人期望的做法,并期望它变得更好。你怎么看?这能带我们走多远?有没有一个极限?
I really appreciate you spending the time, because I think the best way to do research, to understand what to work on next, is to really understand what's not solved yet. Now a lot of people might argue: sure, it doesn't know that 30 clothing items also take five hours because you can hang them all next to each other; you don't need to dry them sequentially. A lot of people might say that kind of answer just wait for GPT-5, train on yet more internet data, maybe a bit more or a lot more human feedback in the training process, but essentially the same procedure at larger scale. That is in some sense what most people expect to be done and then expect it to be even better. What do you think about that? How far is that going to get us? Is there a limit to how far that can get us?
谁知道呢?我是说,老实说,谁知道呢?但我猜测我们不知道常识的深度。它可能非常非常深。问题是,人类不需要额外训练就能回答这些让 GPT-4 头疼的问题。难道这不会激发你的好奇心吗?也许除了在训练汤里加入更多 RLHF 和更多数据之外,还有另一种智能的解决方案。我个人认为,你用这些失败例子做越多的 RLHF,性能就会越好。你将能够覆盖互联网用户发现的越来越多的错误案例角落。但问题是:你怎么确定你覆盖了所有?我们永远无法知道。事实上,我推测总会有一些人们尚未发现的 GPT-4 失败的其他角落。但即使这能解决问题,我们一开始就喜欢这个解决方案吗?因为 RLHF 数据不是开放的,目前只有一家公司拥有它。这意味着任何其他开发预训练神经语言模型的人都必须以这种蛮力方式重建整个系统来覆盖所有角落案例。我们喜欢这个解决方案吗?我个人相信,每当有一个解决方案时,就会有多个解决方案——因为很可能存在另一个可能更好的解决方案,而且第一个解决方案很可能不是最优的。
Who knows? I mean, honestly, who knows? But I'm speculating that we don't know the depth of common sense. It might be really, really deep. And the thing is, humans do not need additional training to answer any of these three questions that give a hard time to GPT-4. So doesn't this really trigger your curiosity that there might be an alternative solution to intelligence other than putting more RLHF and putting more data into the training soup? I personally believe that the more you do RLHF with these failure examples, the better performance you will get. You will be able to cover more and more corners of these error cases that internet people find out. But the question is: how do you know for sure that you covered everything? We will never know. In fact, there will always be, I speculate, some other corners that people haven't found out where GPT-4 fails. But even if that does solve it, do we like the solution in the first place? Because this RLHF data is not open, so only one company currently owns it. That means any other people developing pre-trained neural language models have to reconstruct the whole thing to cover all the corner cases in such a brute-force way. Do we like that solution? One thing I personally believe is that whenever there's a solution, there are solutions—plural—because it's likely that there's another solution that's potentially even better, and it's likely that the first solution is not the optimal one.
我喜欢这个观点。我们确实看到过……
I like that. And we've definitely seen...
这在研究历史上经常上演:一旦某个团队发表论文取得了某项成果,即使细节不全,也常有其他团队跟进、超越。研究就是这样进步的。但现在有点一家独大的味道,因为正如你所说,并非所有数据都对其他人开放,成本太高。而且当前打造最强大 AI 所用的算力如此庞大,只有极少数机构拥有。所以也许你希望存在一种需要更少算力的解决方案?你怎么看?
That play out historically in research: once one group puts out a paper achieving something, even if not all the details are there, often another group matches it or outperforms it. It's amazing how research progress gets made. But a bit of a king here right now because, as you said, not all the data is available to others, so that's expensive. And the amount of compute used for the current approach to the most capable AI is so large that very few organizations have that compute. So maybe your hope is that there's a solution that requires less compute. What do you think?
是的,我非常欣赏 OpenAI 的一点是,他们证明了类似的事情是可能的。这非常令人兴奋,因为它让我更加确信存在一种使用更少算力、更小模型的更好解决方案。肯定有,我们只需要找到它。但重要的是,人们应该尝试不同的想法,而不是做同样的事情——一味地扩大规模。我有点担心,现在很多公司都在做同样的事情。与此同时,学术界的教授或学生们在这种大模型面前感到有些绝望,觉得这是一场规模游戏,因为他们认为自己没有资源去产生影响,只能再写一篇提示工程论文,这真的很可悲。
Yeah, so one thing I really appreciate that OpenAI did was demonstrating that something like that is possible at all. This is super exciting because that really increased my belief that there's a better solution using less compute, smaller models. There's gotta be one; we just need to find it. But it's really important that people actually try different ideas, as opposed to trying to do the same thing—scaling things up. And I worry a little bit that a lot of companies now try to do the same thing. Meanwhile, professors in academia or students in academia feel a little bit hopeless in this large model, like a scale game, because they feel like they don't have the resources to be able to make an impact other than writing yet another prompt engineering paper, which is really sad.
确实,学术界有很多人在思考如何与当前的工业界竞争,尤其是在自然语言处理领域,因为这个领域受影响最大。通过训练大模型,能力提升如此显著。那么你的看法是什么?我是说,你会去工业界训练更大的模型吗?你认为学术界有没有更令人兴奋的事情可做?
Definitely, there are a lot of people wondering in academia how to compete with current industry efforts, especially in the natural language processing space, because that's the space that's been most affected. Capabilities have shot up so dramatically by training large models. So what is your take? I mean, are you going to go to industry to train larger models? Do you think there are things to do in academia that are even more exciting?
我不知道。我总想尝试一些不同的东西。可能会失败,但我宁愿尝试,也不愿随大流。我确实有一些经常使用的方法,在追求能够真正提升小模型性能的替代方案方面相当有效。当然,这些都是学术论文,不可能通过其中一篇就立刻造出能与 ChatGPT 竞争的 GPT。但大致来说,我发现有两个方向非常富有成效。一个是推理时算法。我最初被计算机科学吸引就是因为算法,而人工智能曾经非常注重算法,但现在由于强调规模,算法几乎被抛到脑后了。但我发现,推理时算法可以真正挖掘出神经语言模型的隐藏潜力,而不必完全依赖规模。当你使用推理时算法时,确实可以大幅提升性能,有时甚至不需要对目标任务进行监督训练。我不敢说这对每个任务都成立,但我们在多种任务上都发现效果很好。好到我们有一篇论文叫《The Neurologic Asterisk》,最近在 NAACL 获得了最佳论文奖。所以这是一个富有成效的方向:推理时算法。顺便说一句,这样做的好处是,它让学生们觉得有智力上的满足感,更愿意研究算法。他们永远不需要做超参数调优,那可能很无聊。只要他们正确实现了算法,它通常就能立刻工作,而且往往适用于不同的任务设定。所以我们觉得这总体上非常令人兴奋。
I don't know. I'm always a little bit like I want to try something that is different. Probably I will fail, but I'd rather try than going with the trendy bandwagon. So I do have some recipes that I tend to resort to that have been pretty fruitful in pursuing alternatives that can actually improve the performance of smaller models. Of course, these are sort of academic papers; it's not the case that through one of these I can create GPT that can compete with ChatGPT right away. But roughly speaking, there are two things that I found super fruitful. One is inference-time algorithms. I was drawn to computer science in the first place because of algorithms, and AI used to be heavy on algorithms, but now it's just sort of gone out the window thanks to the emphasis on scale. But I found that inference-time algorithms can really bring out the hidden potentials of neural language models without only resorting to scale. When you resort to inference-time algorithms, you can certainly boost performance dramatically, sometimes even without having to do supervised training on your target task. I wouldn't say this is always true for every single task, but for a variety of tasks we found that this works well. So well that we had a paper called "The Neurologic Asterisk" that won a best paper award at NAACL recently. So that's one fruitful direction: inference-time algorithms. The nice thing about this, by the way, is that it feels intellectual and more pleasing for students to work on algorithms. They never need to do hyperparameter tuning, which is potentially boring. So long as they implement the algorithm correctly, it does tend to work right away, and it tends to work on different task formulations as well. So we found that it's generally quite exciting.
在你讲下一个之前,你提到了推理时算法。能稍微详细解释一下吗?这是什么意思?我自己并不直接从事 NLP 领域,所以实际上不知道该怎么想象。我在想束搜索或束搜索的改进,比如并行生成多个样本并选择最佳的那个。这样想对吗,还是别的什么?
Before you go to the next one, you said inference-time algorithms. Can you elaborate on that a little bit? What does that mean? I don't work in the NLP space that directly myself, so I actually don't know what to imagine. I'm thinking beam search or improvements to beam search where you generate multiple samples in parallel and select the best one. Is that the right thing to think about, or is it something else?
是的,束搜索或贪心解码曾经是许多 NLP 应用的事实标准。然后,让我稍微回溯一下,先谈谈采样算法。我们在 2019 年发表了一篇论文,题为《神经文本退化的奇特案例》,就在 GPT-2 之后。我们注意到,GPT-2——当时引起广泛关注的独角兽文本——并不是基于束搜索,而是基于采样,即所谓的 top-k 采样。这引发了我们的好奇心:为什么要采样?为什么不找 argmax?为什么不用束搜索找到更好的文本?所以我们进行了调查,发现如果你试图寻找 argmax——神经语言模型中概率最高的序列——那么你会得到退化的文本,破碎到非常可笑。每当这种文本出现时,我们都笑得停不下来,因为你简直不敢相信。起初我以为这一定是代码中的 bug,但结果发现这是神经语言模型的真实行为:给如此不自然的破碎文本分配极高的概率分数,这些文本可能从未出现在训练数据中。但学习到的语言模型确实隐藏着这种退化行为。论文解释了原因等等。这引发了许多后续研究,但我们当时提出的一个具体解决方案是 top-p 采样。基本思想是,当你观察预测下一个词的概率分布时,几乎就像在看一个原子和原子核。原子核很小,但占据了绝大部分概率质量。所以它只占很小一部分。从某种意义上说,是一个很小的有效词汇集吃掉了所有的概率质量。如果你从中采样,突然就能得到令人惊叹的文本。质量提升之大令人难以置信,可见这有多重要。这就是我开始品尝修改神经网络解码时算法的乐趣的时候,然后你会得到截然不同的性能。我有一系列这样的论文,试图稍微修改概率分布。但最近,我们研究了所谓的“神经逻辑解码”,基本思想是,如果你给我任何关于包含或排除哪些词的逻辑约束,我们可以以更原则性的方式将它们纳入解码过程。
Yeah, so beam search or greedy decoding used to be the de facto standard for a lot of NLP applications. Then, let me backpedal a little bit and first talk about sampling algorithms as well. We had this paper titled "The Curious Case of Neural Text Degeneration" in 2019, right after GPT-2. We noticed that GPT-2, the unicorn text that caused a lot of attention back then, was not based on beam search; it was based on sampling, what they called top-k sampling. That triggered our curiosity quite a bit because why do you sample? Why not find the argmax? Why not use beam search to find a better text? So we investigated that and found that if you try to look for the argmax—the best probability sequence out of your neural language model—then you get degenerate text, so broken that it's very funny. Whenever this sort of text comes out, we start laughing so much because you cannot believe it. Initially I thought this must be a bug in your code, but it turns out this is the true behavior of neural language models: assigning extremely high probability scores to such broken text that is so unnatural it was probably never in the training data. But the learned language model does have this sort of degenerate behavior hidden. The paper provides why that's the case, etc. It triggered a lot of follow-up research, but one particular solution we proposed back then was top-p sampling. Basically, the idea is that when you look at the probability landscape for predicting which word comes next, it's almost as if you're looking at an atom and the nucleus inside an atom. The nucleus is tiny but occupies the vast majority of the probability mass. So it's a tiny little proportion. In some sense, it's a small working vocabulary that's eating up all your probability mass. If you sample out of it, then you get amazing text suddenly. The quality improvement is quite mind-blowing how much this actually matters. So this is when I started tasting the joy of just modifying the decoding-time algorithm out of the neural network, and then you get vastly different performance. I had a variety of such papers in which I tried to modify the probability distribution a little bit. But more recently, we looked at what we call "neurologic decoding," where basically the idea is that if you give me any logical constraints for which word to include or exclude, we can incorporate them into the decoding process in a more principled way.
不包含以逻辑形式指定的内容,更具体地说是合取范式,但那是技术细节。然后我们可以编写一个复杂的束搜索,在生成流畅英文文本的同时尝试满足所有约束。事实证明,很多自然语言处理应用都可以做类似的事情。有时你想生成一些输出文本,并希望融入某些关键词,或者希望融入或避免某些你不想要的关键词,那么你可以指定这些。即使在机器翻译中,有时也会有语法屈折规则,你想将其作为逻辑规则融入。所以我们做了很多应用场景,甚至包括机器翻译,并且我们可以证明,在各个方面你都能立即提升性能。在某些情况下,即使在无监督的现成 GPT-2 上使用神经逻辑解码,也能比基于束搜索的有监督模型表现更好。甚至有时较小的网络可以胜过较大的网络。这是一个真正出乎意料的实证结果:在较小的无监督模型上使用神经逻辑解码,可以比在较大的有监督网络上应用束搜索表现更好,具体取决于目标任务。
Not include uh specified in logical forms more concretely conjunctive normal form, but that's a technical detail. Then we can write a sophisticated beam search that can try to satisfy all your constraints while producing fluent English text. And it turns out a lot of NLP applications can do something like this. You know, sometimes you want to generate some output text and want to incorporate some keywords, or want to incorporate or avoid some other keywords that you don't want to incorporate. Then you can specify this. And even for machine translation, sometimes there can be grammar inflection rules that you want to incorporate as logical rules. So we've done a lot of application scenarios, even including machine translation, and we could demonstrate that across the board you can improve the performance right away. And in some cases, even using neuro-logic decoding on top of unsupervised off-the-shelf GPT-2 can do better than your supervised model based on beam search. And so much so that sometimes a smaller network can do better than a larger network. This is really an unexpected empirical result: neuro-logic on top of a smaller unsupervised model can do better than beam search applied to a larger supervised network, depending on the target task.
感谢你详细阐述了推理时的算法。你说还有第二个让你非常兴奋的方向,是什么?
Thanks for elaborating on the inference-time algorithms. You said there was a second direction that you're very excited about. What's that?
数据。AI 模型的质量取决于它所喂的数据。增强数据的一种方法就是增加数量:添加更多数据,或者雇人编写演示数据,然后为 RLHF 提供更多点赞/点踩。这些都是社区依赖的更直接的方法。但我认为我们在这方面还有更多可以探索的。我们的一些研究正在调查 AI 生成的数据。通常 AI 生成的数据不是很好,但例如,对于常识知识蒸馏,甚至对于完全不同的任务如摘要,我们探索了 AI 生成的数据,这些数据经过一系列不同的 AI 处理,基本上是过度生成,然后对其他 AI 的生成进行批评,这几乎就像强化学习探索结合奖励网络过滤掉一些糟糕的探索。所以它有一个真实的解释。但无论如何,我们发现如果你真正探索好的数据,那么质量比数量更重要。如果做得正确,我们发现有多篇论文属于这种性质,但我们发现你可以提升较小模型的性能,使其匹配甚至超越 GPT-3 DaVinci(GPT-3 DaVinci 的指令版本)在某些应用场景下的表现。
Data. So AI models are as good as the data that it's fed on. And one way to enhance the data is just quantity: just add more, or hire people to write demonstration data, and then more thumbs up/thumbs down for RLHF. So these are more straightforward methods that the community relies on. But I do think that there's so much more that we can explore there as well. Some of our research is investigating AI-generated data. Usually AI-generated data is not very good, but for example, for common sense knowledge distillation or even for entirely different tasks such as summarization, we explored AI-generated data that goes through a collection of different AIs, basically over-generating and then creating criticism on other AIs' generation, so that it's almost like reinforcement learning exploration combined with a reward network filtering out some of the bad explorations. So it has a real interpretation to it. But anyhow, we found that if you actually explore really good data, then quality is more important than quantity. If you do it correctly, we found that there are multiple papers of this nature, but we found that you can enhance the performance of a smaller model to match or win over GPT-3 DaVinci, the instruct version of GPT-3 DaVinci, for some of these application scenarios.
太棒了。在我们的对话中,有一个概念反复出现,那就是常识。你一直在暗示这个概念。你说常识是什么意思?它是否仍然缺失于当前的模型中?
That's amazing. One of the things that keeps coming back in our conversation here is common sense. You keep alluding to that concept. What do you mean when you say common sense, and is it still missing from these current models?
是的,我将常识定义为大多数人都共享的关于世界如何运作的一般知识。这里的关键词是“大多数人”。所以常识不是普适知识,它涉及多样的事物,包括社会常识知识、物理常识知识等等。较大的模型确实拥有更多常识;它们拥有的量令人印象深刻。所以这非常令人兴奋,因为我们可能终于能够在这方面取得重大突破。但相当奇怪的是,它在一些简单案例上会随机失败。所以它的表现令人兴奋,同时也非常令人好奇,为什么它会在儿童都能正确回答的简单问题上失败。这是一个我思考了很久的话题,但我主要是在 2017 或 2018 年开始研究它。当时这似乎是一个相当疯狂的想法。如今我认为人们已经习惯了反复听到这个词,所以有些人也将其作为自己的主要研究方向。但当我开始时,它被认为是一个不能说的词,因为说出来会让人觉得你可能疯了,太雄心勃勃了。是这么说吗?当时被认为太雄心勃勃了。是的,可能显得太雄心勃勃了。以至于在写论文时,有时我的合著者告诉我不要用“常识”这个词,就是不要用,因为它会引发人们错误的情感反应。所以我有一段时间甚至避免说这个词。但后来我意识到一点:我开始研究它的原因是,我意识到人们认为它太雄心勃勃是基于 70 年代和 80 年代的失败。我们怎么知道他们在 70 年代和 80 年代尝试了一切呢?那时他们没有神经网络,没有众包,没有相同类型的算法或学习算法或推理算法。但除此之外,我看到的另一个问题是,70 年代和 80 年代的研究完全基于逻辑形式。我的第一反应是,我们必须摆脱所有这些。
Yeah, so I define common sense as general knowledge about how the world works that most people share. The keyword here is 'most people'. So common sense is not universal knowledge, and it's about a diverse spectrum of things, including social common sense knowledge, physical common sense knowledge, and all of the above. The larger models do have more of it; they have a really impressive amount of it. So that's really exciting because we might be able to finally give a major crack at it. What's quite curious though is that it randomly fails on some of these simpler cases. So it's both exciting for its performance, and also really curious why it fails at such easy cases that children would be able to answer correctly. It's a topic that has been in my mind for a long time, but I started working on it primarily let's say 2017 or 18. Back then it seemed relatively like a crazy idea to presume. These days I think people got used to hearing this word on and on, so some people actually work on it as a major research direction for themselves as well. But when I began, it was sort of considered as the word not to be spoken, because it sounds like one might be crazy by saying so, too ambitious. Is that the way to put it? It was considered too ambitious at the time. Yeah, maybe it came across as too ambitious. So much so that when writing a paper, sometimes my co-authors told me not to use the word 'common sense', just don't use the word, because it triggers the wrong emotional reaction from people. And so I was avoiding even saying the word for some time. But then here's one realization I had: the reason why I started working on it is because I realized that the reason why people thought it's too ambitious is based on the failures in the 70s and 80s. And how do we know that they tried everything in the 70s and 80s when they didn't have neural networks, they didn't have crowdsourcing, they didn't have the same kinds of algorithms or learning algorithms or inference algorithms developed? But above and beyond, one problem that I saw was the fact that only research in the 70s and 80s were based on logical forms. My first reaction to that was that we got to get rid of all of that.
嗯,你确实在这个方向上取得了很大进展,而且我认为现在越来越多的人对研究常识感到兴奋,将其视为在这些模型中实现真正通用常识的关键缺失部分之一,而这显然仍然普遍缺乏。但与你刚才所说的相关,你还声称语言是推理的最佳媒介。而似乎存在其他被发明出来擅长推理的替代方案,比如数学形式证明语言,或者逻辑或概率推理系统,这些是由人类设计用于推理的。然而你却说实际上我们应该直接用语言推理。为什么?
Well, you've definitely started making a lot of progress in this direction, and I think also more and more people are currently excited to work on the topic of common sense, seeing it as one of the key missing pieces to have a real general common sense in these models, which is clearly still lacking in a general way. But related to what you just said, you have also claimed that language is the best medium for reasoning. And it seems there are alternatives out there that were invented to be good for reasoning, like mathematical formal proof languages, or logical or probabilistic reasoning systems, which are designed by humans to be good for reasoning. And yet you are saying actually we should just reason directly in language. Why?
是的。我不知道为什么我对此着迷,这是一个难以坚持的立场,但这是我的推理。数学方程的问题在于它们很优美、正确,但它们的表达能力太狭窄了。与人类实际能够获取、交流并推理的信息量相比,方程和逻辑形式都存在这个问题。所以当我思考信息的范围时,我担心逻辑形式或方程是用于交流某些推理的惊人人类发明,它们也是推理的有力工具,但并不能涵盖一切。特别是如果我想在常识方面取得突破,我觉得语言可能非常重要。
Yes. I don't know why I'm attracted to this, and it's a difficult position to hold, but here's my reasoning about that. The problem with mathematical equations is that they're beautiful, they're correct, but they're too narrow in the expressive power that they have. This is a problem for both equations and logical forms, compared to the amount of information that humans can actually acquire and then communicate and then reason about, and so forth. So when I just think about the scope of information, I worry that logical forms or equations are amazing human inventions for communicating certain kinds of reasoning. So it's a powerful vehicle for reasoning as well, but it does not encompass everything. And especially if I wanted to make a dent at common sense, I felt that language might be really important.
我们想覆盖人类级别的常识,而不是猫或狗级别的常识。但如果要针对人类级别的常识,我的推测是语言很重要。话虽如此,我很好奇真正的数学家或理论家会怎么看待我的说法。我想到了至少一个人,他同意作为理论家,他无法脱离语言思考。事实上,如果你看任何理论论文,都是自然语言写的。不存在只能通过方程或逻辑形式推理和描述一切的情况。所以语言是推理中非常重要的一环,即使是数学,也得到了理论家的确认。
We want to cover human-like common sense, not like cat or dog level common sense. But if we want to target human-like common sense, then my speculation was the language is important. Having said all of this, I was kind of curious what actual mathematicians or theoreticians might think about my claim. I thought about at least one person who agreed that as a theoretician, he cannot think without language. In fact, if you look at any theory papers, it's all covered in natural language. There's no such thing where you can only reason through and describe everything through equations or logical forms. So language is a really important aspect of reasoning, even for math, confirmed by a theoretician.
嗯,我很喜欢这个观察。而且,当你读论文时,里面肯定有很多语言,至少是为了让其他人更容易理解,对吧?读一篇全是方程的论文非常困难。事实上,我不知道这是否完全可行。
Well, I love that observation. Also, when you look at the papers, there's definitely a lot of language in there, at least for other humans to be able to consume it more easily, right? It's very hard to consume a paper that's just equations. In fact, I don't know if it's all that feasible at all.
现在,这可能是一个比常识更艰巨的任务,而且你一直在思考。你目前的想法是:机器能学习道德吗?
Now here's maybe an even taller order than common sense, and you've been thinking about it. What's your current thinking: can a machine learn morality?
是的,我对此非常兴奋,原因有很多,但其中一个原因是 AI 安全。我想知道,如果不真正学习规范和价值观,AI 是否可能对人类真正安全,这包括道德规范,无论我们喜欢与否。道德规范、社会规范或文化规范之间的界限相当模糊,它们是共享边界的独特概念。但无论如何,机器能学会吗?谁知道呢?人类是否学得很好也不清楚。所以机器能比人类做得更好吗?我不知道。但让我感兴趣的是,当我比较这些挑战与常识时,它们并没有很多共同点。常识规则,比如基本的常识规则:鸟通常能飞,然后有无数例外。所以鸟在睡觉时不能飞,鸟死了不能飞,鸟是雏鸟时不能飞,鸟在笼子里不能飞。而谈到道德规范,让我着迷的是,一般来说,你我都同意我们不应该偷窃。但如果你在街上看到一个人受了重伤,血流不止,旁边正好有一家药店,你想给这个人买绷带,但队伍很长。你是排队,还是先处理情况然后稍后付款?这可能听起来像编造的例子,我们仍然不应该偷窃。但社会规范的另一个方面是,它们更像是没有广泛共识的规则。这里有一个例子:这是悉尼·莱文的研究,她研究了人们何时会打破规则,比如不插队。当你排队时,通常的礼仪是应该排到队尾,对吧?但如果是在厕所排队,而有人有医疗状况呢?也许插队是可以的。如果是在排队买药,有人头痛得厉害,或者排队买水,有人感觉很不舒服呢?可能他们插队是可以的。有很多情况下规则可以被打破,而你认为是否可以打破规则与你的价值观有很大关系。美妙之处在于人类有不同的价值观,所以我们需要纳入价值多元主义。总的来说,这看起来是一个计算问题,但非常混乱。我发现思考这个问题非常令人兴奋。
Yeah, so I'm quite excited about this for different reasons, but one reason being AI safety. I wonder whether it's even possible for AI to be truly safe to humans without really learning norms and values, and that includes moral norms as well, whether we like it or not. The boundary between moral norms, social norms, or cultural norms are rather blurry in the sense that these are distinct concepts that share boundaries. But anyway, can machine learn that? Who knows? It's not clear if humans learn that all that well either. So can machines do even better than humans? I don't know about that. But what's intriguing to me is that it doesn't share a lot of commonalities when I compare the challenges with common sense. In the sense that common sense rules, basic common sense rules like birds can generally fly, and then there's this endless list of exceptions. So birds cannot fly if they're sleeping, birds cannot fly if they're dead, birds cannot fly if they're newborn, and birds cannot fly if they're in a cage. And when it comes to moral norms, what's fascinating to me is that in general, you and I agree that we shouldn't steal. But what if you are on the street and you see somebody majorly injured and bleeding profusely, and there's a pharmacy right there, and you want to buy a bandage for this person, but the line is so long. Do you wait in the line, or do you first do something about the situation and then pay them later? Maybe this may come across as a made-up example, and we still shouldn't steal. But the other aspect of social norms is that they're more like rules that are not as widely agreed upon. Here's one rule example: this is work by Sydney Levine, and she studied when people bend rules, like not cutting lines. When you wait in line, usually the etiquette is that you should go to the back of the line, right? But what if it's a line for a bathroom and someone has a medical condition? Maybe it's okay to cut the line. What if it's a line for medication and somebody has a major headache, or a line for water and somebody's feeling really sick? Probably it's okay for them to cut the line. There are so many cases in which rules can be bent, and whether you think it's okay to bend the rule or not has a lot to do with your values. The beauty of this is that humans have different values, so somehow we need to incorporate value pluralism. All around, this comes across as a computational problem, but such a messy one. And I found that really quite exciting to think about.
这确实是一个难题,而且非常重要。我喜欢你给出的例子,因为它们触及了难点。学习基本规则,当然,语言模型应该能记住基本规则。但理解何时可以打破这些规则的细微差别,或者至少很多人可能认为可以,这很快就变得非常微妙。而且,互联网文本肯定无法自己搞定,我想。
It's a difficult problem for sure, and really important. I love the examples you're giving because I think it really gets to the hard part. Learning just the basic rules, sure, a language model should be able to memorize the basic rules. But then understanding the nuances of when it might be okay to bend those rules, or when at least many people might consider it okay, that gets very, very nuanced very quickly. And definitely internet text won't get it right on its own, I imagine.
是的,不行。
Yeah, nope.
现在每个人都在谈论更大、更大的模型。拥有更小的模型有什么好处吗?我的意思是,人类大脑有特定的大小,也许生物学限制了大脑能有多大。但你怎么看?我们人类是否拥有最佳大小,超过这个大小就没有帮助了?还是越来越大最终会达到超人类?
Now everybody talks about larger, ever larger models. Are there ever benefits to having smaller models? I mean, human brains have a certain size, and maybe biology constrained us of how large our brains can be. But what do you think? Do we as humans have the optimal size, and going larger than that is not going to help anymore? Or going larger and larger will end up superhuman?
是的,总的来说,我可能同意在其他条件相同的情况下,越大越好。但可能在其他条件不同的情况下,比如你要做出选择:要么让模型更大,要么在推理时算法上投入更多算力,或者制作更好的数据。那么可能有一个较小的模型可以做得和较大模型一样好。无论如何,我的推测是存在一个规模的“金发姑娘区”,太小不好,太大也可能不必要。有一个建议的合适数据量,能够实现惊人的智能。
Yeah, in general I probably agree that larger is better, with everything else equal. But it might be that if everything is not equal, meaning you have a choice to make: either you make the model larger, or you spend more compute for inference time algorithms, or making better data. Then it might be that there's a smaller model that can do as well as a larger model. In any case, my speculation is that there's a scale Goldilocks zone, such that too little is not good, too large may not be necessary either. There's a suggested right amount of data to be able to achieve amazing intelligence.
在我读博士期间,差不多 20 年前,上 NLP 课程时,人们会谈论句法、语法,以及 NLP 研究的一些核心内容:我设计的这个算法,可能加入了一些学习,能否把句子解析成主语和动词,什么是名词短语和动词短语?这在当时几乎被视为 NLP 努力的顶峰。而现在我几乎听不到有人提这个。为什么当时那么重要,现在却没人关心了?它会卷土重来吗,还是说那在当时就是一条岔路?
One of the things that when I was taking NLP classes in my PhD days, almost 20 years ago at this point, people would talk about syntax, grammar, and some of the core things that people were studying in NLP was: can this algorithm that I just designed, which maybe has some learning sprinkled into it, also parse a sentence into the subject and the verb, and what's a noun phrase versus a verb phrase? And that was seen almost like the pinnacle of NLP efforts at the time. And I hear simply nothing about it these days. How could it be so important back then and people don't care now? Is it going to make a comeback, or was that kind of a side track at the time?
是的,我读博士时,句法分析是 NLP 会议上最大的方向之一。所以变化确实很大,因为现在几乎没人做句法分析了。而且我从来对句法分析不太兴奋,所以我很高兴它不再是 NLP 领域的中心了。原因如下:我以前是那样想的,我要推翻我过去的想法。但我以前那样想是因为我觉得句法分析过于关注句子结构的一些基本信息,忽略了……
Yeah, when I was a PhD student, parsing was one of the largest tracks at NLP conferences. So certainly the change is really quite something, because these days hardly nobody does parsing. And I was never that excited about parsing, by the way, so I was very happy to see that it's not the center of the NLP field anymore. And here's the reason why: I used to think that way, meaning I'm going to flip what I used to think. But I used to think that way because I felt that parsing focused too much on some of the basic information of the sentence structure, ignoring...
语言学家所说的更务实的含义是什么?所以有语法、语义,还有语用学。语用学很大程度上是关于对未言明内容的推理。例如,当你读一个食谱时,食谱说‘烤半小时’,它可能没告诉你‘在哪里烤’,但你知道,当然是在烤箱里烤。你想想,还能在哪里?然后‘烤什么’?没说的是什么?嗯,很可能就是面团或者你正在准备的任何其他东西。所以你可以轻松地省略这些动作的论元。在人类实际语言中,函数的论元可以很容易地被省略,而传统的句法分析完全忽略了这一点。或者也许我应该退一步说,有些人担心过这种省略论元的问题——术语叫‘省略’——但很多人在实践中并没有真正研究它。而且无论如何,没有足够的监督数据来训练任何模型。所以我一直对长期以来这个领域如此狭隘感到不满。但现在,既然神经网络在理解语言的许多方面如此出色,我想知道现在是否真的是重新审视句法分析的时候了,但第一次真正攻克所有那些先前努力无法处理的语用理解的困难方面。因为对于这些情况,人工标注的数据极其昂贵,而且不可能收集到。但我想知道是否有可能通过算法生成这样的数据或注释,再结合人工验证。如果这是可能的——这完全是推测——如果可能,它也可能有助于让更小的模型更好地理解语言或文本。但这就像是在推测之上叠加推测。但总的来说,我认为尝试与主流不同的事情是好的。谁知道呢?
What linguist was described as more pragmatic meaning? So there's syntax, there's semantics, and then there's pragmatics. Pragmatics is a lot about reasoning about what's not spoken. For example, when you read a recipe, when the recipe says 'bake for half an hour,' it may not tell you 'bake where,' but you know, of course, bake in the oven. You know, where else? And then even 'bake what'? What is not said? Well, probably there's this dough or whatever else you're preparing in the recipe. So it's okay for you to drop these arguments of an action. Arguments of a function can be easily dropped in actual human language, which traditional parsing completely ignored. Or maybe I should step back a little bit and say that some people worried about this dropped arguments—offensive terminology for this is ellipsis—but a lot of people in practice didn't really work on it. And in any case, there was not enough supervised data to train any model. So I just was a bit unhappy about the fact that it's so narrowly scoped for a long time. But now that neural networks are so good at understanding a lot of aspects of language, I wonder if now could be really the time to revisit parsing, but for the first time really attack all these difficult aspects of pragmatic understanding that previous efforts couldn't really handle. Because for those cases, the human-annotated data was unbelievably expensive and it was impossible to gather. But I wonder whether it's possible to algorithmically induce such data or annotations mixed with human verification. And if that's possible—this is all very speculative—if that's possible, it might also help enable smaller models that can understand language or text in general better. But this is like speculation built on top of another speculation. But in general, I think it's good to try things that are different from the mainstream. So who knows?
是的,我喜欢。推测之上叠加推测,偏离主流——这是获得惊人结果的最佳方式,但我想也常常是完全没有结果的最佳方式。这样做是有风险的。
Yeah, I love it. Speculation on top of speculation, deviating from the mainstream—best way to have a surprising result, but also I guess often best way to have no results at all. It's risky to go yeah.
现在从更基础的研究转向应用领域。当你看到这些语言模型正在发生的事情,如此根本性的进展正在取得,有没有一些真正让你个人兴奋的实际应用,要么现在已经发生,要么你认为可能很快会使用语言模型实现?
Now switching gears from the more fundamental research to the application space. When you look at what's happening with these language models, such fundamental progress being made, are there some real-world applications that really excite you personally, either already happening now or maybe that you see are likely gonna happen soonish using language models?
我非常兴奋地看到更多多模态应用可能成为现实,比如帮助机器人更好地规划或导航,与从未见过的新物体更好地交互,或者通过语言界面更好地理解图像或视频。然后在文本领域,思考现在可以在这些神经语言模型之上构建哪些新应用是非常令人兴奋的,比如教育辅助工具,通过辅导服务引导学生。可汗学院目前正在做的那种事情真的非常鼓舞人心。另外,我从牛津大学的道德哲学家约翰·塔西乌拉斯那里了解到,他还专攻法学。他提到,有大量人口无法获得法律服务,因为他们没有经济能力。但即使他们使用一些公共服务,排队也可能太长,所以很多时候他们根本无法对任何事情进行合法投诉。而能够服务于这些人群似乎很棒。最近我遇到了一位艾森豪威尔奖学金获得者——在我送出的奖学金获得者中——很抱歉我记不起他的名字了,但他告诉我,在非洲,有一些人群无法真正去看医生,但可以通过电话界面获得关于他们生病孩子当前状况的基本建议,从而受益匪浅。所以我觉得,哇,可能还有很多积极的用例有待开发。这真的很令人兴奋。
I'm so excited to see more multimodal applications being potentially part of the reality, like helping robots to plan better or navigate better, interact better with new objects that it's never seen before, or understanding images or videos better using language interface. And then perhaps in the text domain, it's very exciting to think about what new applications could be now built on top of these neural language models, such as education assistant tools to guide people or students through a tutoring service. The sort of stuff that Khan Academy is currently working on is really very inspiring. And then I also learned from John Tasioulas, a moral philosopher at Oxford, and he also specialized in jurisdiction or laws. So he mentioned something that there are a lot of large populations of people who don't have access to laws or service for laws because they just don't have the financial means. But even if they were to use a bit of a public service, their line may be too long, so that a lot of the times they just have no way to complain legally about anything. And somehow being able to serve those populations seems to be amazing. And then recently I met someone who's Eisenhower Fellowship winner—among those I send our fellowship winners—I'm so sorry that I'm blanking on his name, but he told me about how in Africa there could be some populations who cannot really go see a doctor but really could benefit from a phone interface that gives them some basic advice about their current situation with their sick child, for example. So I felt that wow, there's potentially really a lot of positive use cases yet to be developed. That's really exciting.
现在如果你考虑其中一些用例,可能也需要更多研究。因为如果我想象像医疗用例这样的事情,我的意思是,目前我们还不能完全信任这些模型的回答。可能比没有回答好,但也可能不是。我的意思是,也许推迟行动比采取错误行动更好。所以似乎还有一段路要走。你不同意吗?
Now if you think about some of those use cases, it probably needs more research too. Because if I think about things like the medical use case, I mean we can't fully trust the responses of these models at this point. It might be better than no response at all, but maybe not. I mean maybe it's better to hold off on taking action than to take a misguided action. So it seems a bit of a ways to go. Would you think otherwise?
是的,我很高兴你跟进这个问题。所以存在一个重大风险,尤其是一些医生开始依赖 ChatGPT,因为它确实令人印象深刻——它在很多病例中诊断得很好。然后医生开始依赖它,突然之间这种疏忽可能导致大问题。但这不仅仅是医生依赖 GPT-3.5 或 GPT-4 层面的问题。这对公民来说也是一个实际问题,因为他们可能只是用 ChatGPT 自我诊断,然后那就是你的医生,因为那看起来足够好了。所以我认为这种部署相当复杂和微妙,取决于我们针对什么人群以及什么目的。无论如何,我们需要提高更广泛社区的人工智能素养,让他们意识到这还不能完全信任。而发现缺陷可能会变得越来越难,这意味着人们会更加依赖它。所以我认为一个关键的研究方向是系统性地发现谬误在哪里,局限在哪里。但这是一个开放的研究问题。嗯,也许几年后就不再是了,取决于你和其他人的进展。
Yeah, I'm so glad that you follow up on that. So there's a major risk, especially if some doctors start relying on ChatGPT because it does impress—it does do well on diagnosis in a lot of cases. So then doctors start relying on it, and then suddenly there's this oversight that can cause a big problem. But this is not just a problem at the level of a doctor's reliance on GPT-3.5 or GPT-4. This is actually a real problem for citizens as well, because they might just self-diagnose using ChatGPT, and then that's your doctor because that seems like good enough. So I think this deployment is quite complicated and nuanced, depending on what population we target and for what purpose. And in any case, we kind of need to increase the level of AI literacy in the broader community so that they realize that this is not something to be fully trusted yet. And finding the flaws might become harder and harder, which means people will rely on it even more and more. And so one of the critical research directions in my mind is some sort of systematic discovery of where the fallacies are, where the limitations are. But it's an open research question right now. Well, maybe not anymore a few years from now, depending on your progress and others.
我们谈了很多关于人工智能、自然语言处理研究等等。但让我们暂时回到更早的时候。你在哪里长大?小时候你在忙什么?你是怎么进入人工智能领域的?
We talked a lot about AI, NLP research and so forth. But let's go much further back for a moment. Where did you grow up? What kept you busy as a little kid? And how did you end up in AI?
是的,说来话长。我来自韩国。成长过程中,我想我很早就知道自己被 STEM 领域吸引。我喜欢玩那些可以建造、构造或拆解的东西,比如组装和拆卸东西。每当家里有电子设备送到时,我会非常高兴,看看是否需要组装,如果需要,我就去组装。所以我知道自己真的被科学和数学吸引。顺便说一句,我喜欢这些,但讨厌英语。所以我的英语真的很差。现在我不得不补上英语,以便在美国生存和工作,但那是韩国的事了。但我当时不知道计算机科学这个领域。所以当我差不多要决定申请哪所大学时,我决定也许学计算机工程。我对计算机了解不多,或者说根本不知道。
Yeah, such a long story. So I'm from South Korea. Growing up, I think I knew early on that I was drawn to STEM fields. And I loved playing with things that I can build or construct or deconstruct, like assemble things and disassemble things. And whenever some electronic devices get delivered in the house, I would be so happy to see if I can assemble it if there's assembly needed. So I knew that I was really drawn to science and math. I loved that and hated English, by the way. So my English was really bad. Now I had to catch up on this so I can get by and do my job in the US, but this is back in Korea. But I didn't know about computer science as a field. So when I was about the time to decide where to apply for college, I decided I'm gonna do maybe Computer Engineering. I never knew much about computer or didn't know.
我其实不太懂编程,但觉得它很酷,所以就申请了。读计算机工程期间,我主要设计 CPU、做电路设计和电路编程,不是做 AI 需要的那种编程。我也从没上过 AI 课。我不知道自己还能读研,因为我成长的文化环境让我觉得,去韩国那所大学读书已经学得太多了。周围人都在泼冷水,说这太过分了。他们开玩笑说,读这个专业我既不是女人也不是男人,是第三性别。所以我想,要么是我学得太多了,要么是父母负担不起,或者两者都有,我得找工作。为了找工作,我注意到有软件工程这个职业,就自己业余学编程,因为专业课程不教。后来碰巧,2000 年左右互联网泡沫,微软在美国招不到工程师,所以那一年他们破例从韩国等国家招人。但我不会英语,没法申请。我拼命学了四个月英语,才有机会面试。面试时英语很烂,但编程还行,通过了。于是我在西雅图微软开始了软件工程师的职业生涯。那时我完全不知道自己会做 AI。在韩国,AI 被认为是寒冬。我主要做高性能软件编程,那是我的专长,在微软做得很开心。但两年后我开始不安,因为我负责的代码越来越多,不确定能否做一辈子。学习曲线平了,工作越来越轻松,但我需要新的冒险。于是我决定读博。当然,周围人都劝我别去,说这主意太糟了,尤其是 AI——‘你开玩笑吧?AI 是寒冬啊!’所以我对博士没抱太大期望。我只是想要一段能探索 AI 的时间,出于某种说不清的原因,我必须这么做。我打算几年后回到软件工程师的岗位,但后来没回去。
I didn't know much about programming but it sounded cool, so I applied to it. During my computer engineering degree, I mostly designed CPUs, circuit design, and programming for circuit design, not the kind of programming you need for AI. I never took an AI class either. I didn't know I could go to grad school because I grew up in a culture where I got the impression that I studied too much by going to that particular university in Korea. I got discouragement from left and right that this was too much. They used to joke that I was neither woman nor man by doing this major, a third gender. So I assumed for one reason or another I studied too much, or my parents couldn't afford it, or both, so I had to find a job. For finding a job, I noticed there was something called software engineering, and I studied programming on my own as a side hobby because the classes didn't teach it in my major. Then it so happened that Microsoft ran out of engineers in the US because of the dot-com bubble in 2000 or slightly before. So only for one year they were going to hire from random countries outside the US, such as Korea. Except I couldn't speak English, so I couldn't really apply. For four months I studied English so hard that I might have a chance to do the interview. I did the interview with very basic English skills, but my coding was okay, so I passed the coding interviews. I started my career as a software engineer in Seattle working for Microsoft. Still, I had no clue that I would eventually do AI. If anything, back in Korea, AI was considered to be in a winter time. I was doing more high-performance software programming, and that was my specialty. It was so fun for me to do that at Microsoft. But after two years, I started becoming very restless because I started owning a lot of code, and I wasn't sure if I could do this for the rest of my life. The learning phase kind of plateaued, and I could get away working less and less, but I needed a different adventure. That's when I decided maybe I could do a PhD. Of course, people around me tried to talk me out of it. They said it was such a bad idea, especially AI: 'Are you kidding me? Did you know that AI is in a winter time?' So I didn't expect too much out of my PhD. I just wanted that period of time when I could explore AI for some mysterious reason that I couldn't explain, but I had to. Then I was going to be ready to return to the same job as a software engineer after those few years, but that didn't happen.
是啊,你没回去。是什么让你没回去呢?
Yeah, you didn't go back. What made you not go back?
哦,事情是这样的:AI 太难了。读博期间,我决定做 NLP,尤其是因为它看起来比 AI 其他领域更不发达。否则我考虑过基于逻辑的 AI。还好我没选那个。当时有人发表看法。教科书第一版出来了,从零到一,那是个激动人心的时刻。然后 Jurafsky 的第一本教材出来了,我一看,有个错别字。我想,‘嗯,我能改这个错别字。也许这个领域对我来说有点机会做点什么,而不是太成熟的领域。’但到我快毕业时,次贷危机爆发了,没有工作。我准备好了却不能马上毕业,得等着。最后毕业时,我基本上只有一个面试和一个大学的工作机会,另一个是博士后。两者之间,我想也许可以试试石溪大学的教授职位。就这样开始了。那会儿挺有意思的,NLP 还不是热门领域,工作机会很少。但我想继续做下去,真的很有趣。我就是不想回去。这对我来说太有趣了。
Oh, what happened was that AI was so hard. During my PhD, I decided to do NLP, especially because it seemed even less developed than anything else in AI. Otherwise, I was considering logic-based AI. I'm glad I didn't do that, by the way. I'm glad I didn't. People had opinions. The textbook version one came out, and from zero to one it was an exciting time. Then Jurafsky's first textbook came out, and I'm looking at it and there's a typo. I thought, 'Hmm, I can correct the typo. Maybe this is a good field for me to have some chance to do something at all, as opposed to something too developed.' But about the time I was graduating, the subprime mortgage crisis happened, so there was no job. I couldn't graduate right away when I was ready; I had to wait. Then when I was finally graduating, I had basically one interview and one offer at a university, otherwise a postdoc offer. Between the two, I figured maybe I'll try this professor job at Stony Brook. So that's how it began. Still, it was a funny time because NLP was not a hot field at all yet, so jobs were really limited back then. But I wanted to keep going; it was really fun. I just didn't want to go back. This was too fun for me.
石溪大学真是慧眼识珠,发现了你的远见和专长。如果他们当时没雇你,谁知道呢,也许这个领域就错过你了。所以很感谢他们。当然,后来的事大家都知道了。现在你引领着方向,人人都想雇你。事情变化真快。更不用说现在人人都想做 NLP 了。世界完全不同了。我们已经聊了很多,但我还想再问你一个问题,Yejin。你显然工作很努力,但你会做些什么来远离工作、真正放松呢?
Well, hats off to Stony Brook for identifying your vision and expertise. I mean, if they hadn't hired you at the time, who knows, maybe we would be missing out on you in the field. So thankful to them. Of course, the rest is history. Now you're leading the way, and everybody would want to hire you. Things can change quickly. Not to mention everybody wants to work in NLP now. It's a completely different world. Now we've already covered a lot, but I'm hoping to ask you one more question, Yejin. Obviously you work a lot, but what are some things you do to get yourself away from work to really relax?
我做什么呢?好问题。说实话,我工作得很开心。读博时我做过很多别的事,比如烤面包。我有一堆爱好。顺便说,我面包烤得不好,别让我烤。但现在,我伴侣很担心我的长期健康,所以他会把我从电脑前拉走,带我去徒步,让我离开电脑。当然,我还会带手机。但我发现我现在挺喜欢爬山这类活动,所以周末有空就会去。另外,有些半工作相关的事情也让我放松,比如读道德哲学或认知科学的书。我觉得这些特别有意思,希望能多读一些。但我也说不出什么特别炫酷的爱好来炫耀。嘿,爬山对我来说就很不错了。
What do I do? Good question. To be honest, I'm having so much fun. I used to, during PhD, I did a lot of other things like baking bread. I had a range of hobbies. By the way, my bread is no good, so don't ask me to bake anything. But these days, I guess I could say my partner is very much concerned about my long-term health, so he pulls me out of my computer and tries to put me on a hiking trail so that I'm departed from my computer. But of course I have my phone. But I discovered that I actually like these mountain activities these days, so I try to do that on weekends when I can. Otherwise, there are semi-work-related things that also feel relaxing, such as reading books on moral philosophy or cognitive science. I find those just so exciting, and I wish I could do more of it. But I don't know, I can't really say anything fancy as a hobby that I can brag about. Hey, going into the mountain sounds pretty good to me.
是啊,听起来很不错。非常感谢你抽出时间。很高兴能邀请到你。
Yeah, that sounds pretty good to me. Thanks so much for making the time. Really glad to have you on.
谢谢,非常开心。感谢邀请。
Yeah, thank you so much. This was fun. Thanks for having me.