大型语言模型的实际工作原理

How Large Language Models Actually Work

古斯塔夫·瑟德斯特伦 Gustav Söderström · Spotify 研发 · 2023-07-31 · 约 91 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Spotify联合总裁用通俗语言解释生成式AI和LLM背后的理论。

Spotify's co-president explains the theory behind generative AI and LLMs in simple terms.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 31)

全文 · Full transcript(中英对照)

引言与目的 Introduction and Purpose

Gustav

大家好,我叫 Gustav Söderström,是 Spotify 的联席总裁。应同事们的邀请,我要为你们所有人——从工程师到 Spotify 的高管——做一次关于 AI 的深度讲解,重点讲这种新型生成式 AI,并试着解释这些东西到底是怎么运作的。为什么我们会有像 ChatGPT 这样的服务,能让你创作整部小说;或者像 Stable Diffusion、Midjourney 这样的服务,能仅凭文字或白噪声就生成美丽的图像甚至音乐?

Hey everyone, my name is Gustav Söderström. I'm co-president at Spotify. I was asked by my colleagues to do a deep dive on AI for all of you, from engineers to executives of Spotify, specifically on this new type of generative AI, and try to explain how these things actually work. How is it that we have services like ChatGPT where you can create an entire novel, or services like Stable Diffusion or Midjourney that can create beautiful images, even music, out of just text or white noise?

Gustav

对我们这些科技行业的员工和高管来说,理解这一点几乎就是我们的本职工作。但我认为,即使你不在科技行业,也几乎有义务去理解当下正在发生的事情,因为这是一件大事。一百年后人们仍会谈论 2023 年,因为就在这一年,计算机开始通过图灵测试,也就是说,它们能让不知道自己在跟计算机系统还是真人对话的人误以为它们是真人。

Now, for us as employees and executives in the tech industry, it is quite literally our job to understand this. But I think that even if you're not in the tech industry, it is almost an obligation to understand what is going on right now, because this is a big thing. People will talk about 2023 a hundred years from now, because this is the year that computers started passing the Turing test, meaning that they could pass for being human to someone who doesn't know if they're speaking to a computer system or to a human.

Gustav

我认为一百年后人们仍会谈论这件事,就像我现在跟孩子们谈论近一百年前原子被分裂一样。我觉得如果我活在 20 世纪 30 年代原子分裂的时代,我会希望理解当时发生了什么。我不希望 20 年后才意识到那件事发生了,而我当时就在场却浑然不觉。所以我认为,对每个人来说,至少对已经发生的事情及其运作方式建立起一种直觉,是很重要的。

And I think people will talk about this a hundred years from now, just as I'm talking to my kids about splitting the atom almost a hundred years ago. And I feel that if I were there, you know, in the 1930s when we split the atom, I would have liked to understand what was going on. I wouldn't have liked to have realized 20 years later that it happened and I was there but I was actually unaware. So I think it's important for everyone to sort of at least get an intuition for what it is that has happened and how it works.

Gustav

所以我在一开始就大胆地向你们承诺:听完这次演讲后,你会觉得自己确实理解了正在发生的事情,即使你不太懂数学。不幸的是,我发现要真正掌握正在发生的事情相当困难,我认为部分原因在于此。萧伯纳提出了“针对外行的阴谋”这个概念。它的意思其实是,外行可以指神职人员阶层,而它真正的意思是,任何职业都倾向于为其他人进入该职业设置障碍。这可能是故意的,比如围绕该职业设立认证机构和规则体系;但也可能不那么刻意,只是围绕该职业创造了非常复杂的词汇和行话。

So my bold upfront promise to you is that after this presentation, you will feel like you do understand what is going on, even if you don't know a lot of math. Unfortunately, I've found that it's pretty hard to actually get a grip on what is going on, and I think it's partially because of this. So George Bernard Shaw created this notion of conspiracies against the laity. And what it means is really, so laity could be the priest class, and what it means really is that any profession tends to raise barriers towards other people entering that profession. This could be deliberately, by creating certification authorities and rule systems around this profession, but also less deliberately, but just creating very complicated vocabulary and lingo around this profession.

Gustav

你们都知道我的意思。比如金融或法律,你常常会觉得它们看起来非常复杂、难以理解。而当你真正理解之后,你会问自己:他们为什么不能直接这么说呢?往往词汇本身就让它显得比实际更难。其实,这通常并非有意为之。事实是,专业群体倾向于创造专业词汇,因为这样他们彼此之间在更高层次上交流更高效,但这也为其他人理解正在发生的事情设置了障碍。

You all know what I mean. You take something like finance or legal, and you often get this feeling that it seems very complicated and hard to understand. And then when you actually do understand it, you kind of ask yourself, like, why couldn't they just have said that? Often the vocabulary itself makes it seem harder than it actually is. Now, it's often not actually deliberate. It is a fact that specialized groups tend to create specialized vocabularies because it's more effective for them to talk to each other at sort of a higher level, but that also creates these barriers to understand what is going on for everyone else.

Gustav

我认为,让它看起来比实际更复杂的问题在于,人们混淆了理论与实践。虽然让这些东西真正运转起来的实践非常复杂,确实需要大量数学,但理论其实并不复杂。我认为,完全可以在不理解实际问题的情况下,对正在发生的事情建立起直觉。那么现在,让我们试着看看能否揭穿这个“阴谋”。

I think the problem that makes it seem more complicated than it is is that people confuse theory with practice. And while the practice of getting these things to really work is very complicated and actually does require a lot of math, the theory actually isn't. I think it's entirely possible to build intuitions about what is going on without having to understand the practical problems. So let's try to see if we can expose this conspiracy now.

Gustav

对于你们当中熟悉这些内容的人来说,你们会看到我在各处走了不少捷径,实际数字和百分比并不总是完全合理,但请你们包容我,因为我力求简单,只求足够真实,以建立大体正确的直觉。准备好了吗?我们开始吧。

For those of you who know this stuff well, you will see that I take a bunch of shortcuts here and there and that the actual numbers and percentages don't always make full sense, but you'll have to indulge me as I'm trying to keep it simple and just keep it true enough to create largely correct intuitions. Are you ready? Let's go.

什么是LLM? What is an LLM?

Gustav

那么什么是 LLM?LLM 代表大型语言模型,它是驱动 ChatGPT 这类服务的东西。所以在这一部分之后,希望你们能理解 ChatGPT 实际上是如何工作的。但我们必须经历一系列步骤。第一步是理解,你如何让一台本质上只理解数字的计算机真正理解语言。

So what is an LLM? Well, LLM stands for large language model, and this is the thing that powers something like ChatGPT. So after this section, hopefully you'll understand how ChatGPT actually works. But there are a bunch of steps we have to go through. The first step is to understand how you even get a computer that literally only understands numbers to actually understand language.

Gustav

其实,要大致理解发生了什么,有一个相当直观的方法。假设你作为一个人,拿起英语词典,从第一个词开始,比如“ace”,然后是“amazing”,接着是“appreciative”、“aromatic”,一直到最后。英语词典里大约有 60 万个词,你给它们编上号。第一个词“ace”是 1,第二个词“amazing”是 2,第三个词“appreciative”是 3,依此类推。

Well, there is a pretty straightforward intuition to understand roughly what's going on. So let's say that you, as a human, take the English dictionary, you just start at the first word, so 'ace', then 'amazing', then 'appreciative', 'aromatic', all the way to the last word. And there's about 600,000 words in the English dictionary, and you just give them a number. So the first word 'ace' has number one, second word 'amazing' has the number two, third word 'appreciative' has the number three, and so forth.

Gustav

那么现在,你完全可以拿一个英语句子,比如“hey how are you”,然后查每个词对应的数字。所以对计算机来说,像“hey how are you”这样的句子并不是真正的句子,它只是一串数字。比如“hey”这个词对应数字 25,“how”对应 30,“are”对应 5,“you”对应 75。

So now you can literally take a sentence in English, let's say 'hey how are you', and you can just look up every word and see what number it has. So for a computer, a sentence like 'hey how are you' is not actually a sentence, it's just a sequence of numbers. So the word 'hey' has, for example, the number 25, the word 'how' has the number 30, the word 'are' has the number 5, and the word 'you' has, for example, the number 75.

语言模型简介 Introduction to Language Models

Gustav

现在你可能会觉得这些数字很小,但这只是为了简化。所以这只是一个查找表:你有一长串 60 万个单词,每个单词对应一个唯一数字,你只需把句子翻译成数字序列。

Now you may see that these are very small numbers, but that's just to keep it simple. So it's just a lookup table: you have this long list of 600,000 words, a unique number for every word, and you just translate a sentence into a sequence of numbers.

Gustav

那我们来更详细地讲讲:这到底是怎么运作的?我们先从一个超级简单的模型开始。假设我们有一个 Excel 表格,每一行是英语词典里的一个单词,每一列也是英语词典里的一个单词。这样,对于每个单词,我们都有一个百分比,表示下一个词出现的概率。

So let's go into a bit more detail: how does this actually work? Well, let's start with this super simple model. Let's say that we have this Excel sheet where we have every word in the English dictionary as a row, and then we have every word in the English dictionary as a column. So for every word, we have a percentage for how likely every next word is.

Gustav

假设我们输入单词“are”,电脑得到数字 5。现在语言模型的任务就是判断:“are”后面最可能出现的下一个词是什么?更准确地说,根据我在互联网上见过的所有单词或所有数字序列,数字 5(代表“are”)后面最可能出现的下一个数字是什么?

So let's say we get the word 'are' — the computer gets the number 5. And now the job of the language model is to say: what is the most likely next word after 'are'? Or more correctly, what is the most likely next number after the number 5, which represents the word 'are', according to all the words or all the sequences of numbers I've seen on the internet?

Gustav

你可以想象,作为人类,你可能会给很多词分配相同的概率。可能是“you”,因为句子可能是“嘿,你好吗”;也可能是“they”,也可能是“things”,因为句子可能是“最近怎么样”;也可能是“fine”,因为句子可能是“他们很好”。因为你只有一个词作为上下文,很难猜对。你根本不可能猜对,信息不够。但你可以猜得不错。

So you could imagine that as a human, you might give an equal percentage to a lot of words. It could be 'you' because maybe the sentence is literally 'hey, how are you', but it could be 'they' — it could be 'things' because the sentence may be 'how are things'. It could be 'fine' because the sentence may be 'they are fine'. Because you only have one word to guess from, one word of context is going to be hard to guess correctly. You literally can't guess correctly; you don't have enough information. But you can make a good guess.

Gustav

所以像“you”“they”“things”“fine”这些词可能有相同的概率,但有些词正确的可能性很低,比如“animals”——这在互联网上是非常少见的句子。虽然可能发生,但确实不常见。

So maybe these things like 'you', 'they', 'things', 'fine' have the same percentage, but you're going to have words that have a very low chance of being right, like 'animals' — that's a very uncommon sentence on the internet. Most likely it happens, but it's just uncommon.

Gustav

所以再说一遍,电脑并不是在猜单词。电脑会说:根据我在互联网上见过的所有文本,或者更准确地说,根据我在互联网上见过的所有数字序列,按照这种从单词到数字的转换,在数字 5 之后我通常看到数字 75(代表“you”),或者在数字 5 之后我通常看到数字 42(代表“they”),或者在数字 5 之后我通常看到数字 97(代表“things”)。

So again, through the computer, it's not guessing words. What the computer would say is: based on all the text I've seen on the internet, or more correctly, based on all the sequences of numbers that I've seen on the internet, according to this translation from words to numbers, after the number 5 I usually see the number 75, which represents 'you', or after the number 5 I usually see the number 42, which represents 'they', or after the number 5 I've usually seen the number 97, which represents 'things'.

Gustav

所以它做的事情和你作为人类会做的事情非常相似。作为人类,你对这些百分比有直觉。而且我觉得很容易理解:如果电脑能查看维基百科和整个互联网上的所有统计数据,它也能对下一个词的概率和统计有很好的把握。

So it's doing something very similar to what you would do as a human. And as a human, you have intuitions about the percentages. And I think it's easy to understand that if a computer could just look at all the statistics over all of Wikipedia and all the internet, it could also have a pretty good idea about the percentages and the statistics about the next word.

Gustav

好了,这就是我们用一个 60 万行、60 万列的表格能做到的极限。现在我们假设多一个上下文词。比如不再是“are”,而是“how are”,电脑看到的数字是 30 和 5。这时概率就变了。

All right, so this is as far as we get with this one column of 600,000 rows, or one 600,000 rows and 600,000 columns. Now let's say that we get one more word of context. So instead of 'are', we get 'how are', which the computer again sees as the numbers 30 and 5. Well, now the percentages change.

Gustav

对你来说,当只有“are”这个词时,“you”“things”“they”“fine”的概率差不多。但当你看到“how are”两个词时,突然“you”的可能性更大——现在可能达到 50%,因为在互联网上“how are you”比“how are they”更常见。所以如果你要猜,你可能会选“how are you”。

For you as a human, when you just had the word 'are', 'you', 'things', 'they', 'fine' were equally likely. But when you get two words 'how are', all of a sudden 'you' is probably more likely — maybe it's 50% likely now, because it's more common to say 'how are you' on the internet than 'how are they', for example. So if you had to guess, you'd probably go with 'how are you'.

Gustav

而且“how are you”或“how are they”比“how are things”更常见。再比如“fine”,当只有一个词时它很常见,因为“they are fine”里会有“are fine”——但没人说“how are fine”。所以现在“fine”的概率变得非常低。

And it's even more likely to say 'how are you' or 'how are they' than 'how are things', perhaps. And if you take something like 'fine', which was pretty likely when you just had one word because 'are fine' could happen if you say 'they are fine' — no one says 'how are fine'. So now all of a sudden, the word 'fine' has a very low percentage.

Gustav

把它转换成数字,这意味着大型语言模型现在有两个数字作为依据。就像你作为人类,有两个词时能猜得更准一样,大型语言模型根据它在互联网上见过的所有数字,有两个数字时也能猜得更准。

Translating this to numbers, what it means is the large language model now has two numbers to guess from. And just like you as a human would guess much better with two words, the large language model, based on all the numbers it's seen on the internet, is going to guess much better with two numbers as well.

Gustav

你可以想象接下来会怎样:如果有三个数字或三个词呢?现在上下文是“hey how are”,或者数字 25、3、5。这时你的猜测会变得非常准。“hey how are you”现在非常可能——比如 70%。“hey how are they”——可能更低。你可能会说“how are they”,但你很少说“hey how are they”,因为他们不在你面前。所以这可能降到 5%。

You can imagine what comes next: what if you had three numbers or three words? So now the context is 'hey how are', or the numbers 25, 3, 5. Now your guess is going to start to get really good. 'Hey how are you' is now very likely — let's say it's 70%. 'Hey how are they' — maybe even less likely. You might say 'how are they', but you seldom say 'hey how are they' because they are not in front of you. So that might go down to 5%.

Gustav

而像“hey how are things”这种,之前比“how are they”可能性略低,现在会飙升,因为“hey how are things”在互联网上很常见。而“hey how are fine”和“hey how are animals”的概率则非常低。

And something like 'hey how are things', which was a bit less likely than 'how are they', now shoots up because 'hey how are things' is something you would see on the internet quite often. And then 'hey how are fine' and 'hey how are animals' have very low percentages.

Gustav

所以关键在于:作为人类,你获得的上下文越多——你拥有的词越多——你的猜测就越准,你对下一个词的把握就越大。大型语言模型也是如此:它获得的数字越多(数字代表单词),它对下一个数字的猜测就越准。

So the whole point is: the more context you get as a human — the more words you have — the better your guess, the more sure you're going to be about the next word. And it's the exact same thing for a large language model: the more numbers it gets (the numbers represent words), the better its guess is going to be on the next number.

Gustav

所以从某种意义上说,大型语言模型所做的就是:它只是猜测下一个词,或者更准确地说,根据前面的数字猜测下一个数字。所以我们称它们为“大型语言模型”其实有点用词不当。它们或许应该叫“大型数字模型”或“大型序列模型”,有时候确实有人这么叫。

So in a sense, this is all that a large language model does: it just guesses the next word, or more correctly, the next number from previous numbers. And so it's actually a bit of a misnomer that we call them 'large language models'. They should probably be called 'large number models' or 'large sequence models', which they are called sometimes.

Gustav

因为事实证明,你可以把单词转换成数字,但一切都可以是数字。你可以把像素转换成数字——实际上像素本身就是数字——所以你可以把像素或图像输入大型语言模型,根据之前的像素(之前的像素数字,即 RGB 值),它会非常擅长猜测下一个像素。

Because it turns out that you can turn words into numbers, but everything is numbers. You can turn pixels into numbers — actually pixels are numbers — so you can put pixels or an image into a large language model, and based on the previous pixels (the previous pixel numbers, the RGB values), it is going to get very good at guessing the next pixel.

Gustav

你可以输入音频样本,比如某人说话的音频——那些就是数字。如果你有这些数字,并基于这些数字进行训练,大型语言模型会非常擅长猜测音频序列中的下一个样本。

You can put audio samples, for example from someone speaking — those are just numbers. And if you have those numbers and you train on those numbers, a large language model will get very good at guessing the next sample in an audio sequence.

Gustav

所以记住:它不是语言模型,实际上只是一个数字模型。世界上的一切都可以转换成数字。所以任何序列——只要你有大量数据——这些模型就能学习这些序列的统计规律,并正确猜测下一个数字。

So remember: it's not a language model; it's actually just a number model. And everything in the world can be translated into numbers. So anything that is a sequence — if you have lots of data — these models can learn the statistics about those sequences and correctly guess the next number.

Gustav

所以从某种意义上说,你现在真正理解了大型语言模型的作用。但这里有一个你应该理解的问题:如果这么简单,为什么以前没有实现?

So in a sense, now you actually understand what a large language model does. But there is a problem here that you should understand: why hasn't this happened before if it's so simple?

Gustav

嗯,尽管理论上很简单——所谓“只是统计”,只是猜测下一个数字——但事实证明,对于很长的上下文组合、很多单词,猜测下一个数字的计算量非常大。

Well, even though it's simple in theory — it's quote unquote 'just statistics', just guessing the next number — it turns out that guessing the next number for long combinations of context, for many words, is very computationally intensive.

Gustav

所以如果我们回到你的 Excel 表格,其中每个单词都有对词典中其他每个单词出现概率的百分比——我之前说英语词典大约有 60 万个单词,但我们简化一下,只取最常用的 5 万个。

So if we go back to your Excel sheet, where for every word you have a percentage on how likely every other word in the dictionary is — now I said there were about 600,000 words in the English dictionary, but let's just simplify it and just take like the most popular 50,000.

问题的规模 The Scale of the Problem

Gustav

因为 5 万大致就是一个大型语言模型实际使用的词汇量大小。所以我们把它缩小了。现在你有 5 万行单词,以及对应每个单词的 5 万列。所以那是 5 万乘以 5 万,实际上是一个相当大的 Excel 表格。大概有 25 亿个单元格之类的,所以是个相当笨重的 Excel,但还是有点可行的。单元格很多,但可行。但这只是针对一个单词的情况。记住,当你只有一个单词可以猜测时,你的猜测会相当糟糕。

Because 50,000 is roughly the size of a vocabulary that a large language model actually uses. So now we made it smaller. You now have 50,000 rows of words and 50,000 columns for every such word. So that's 50,000 times 50,000, which is actually a pretty big Excel sheet. It's like two and a half billion cells or something, so pretty unwieldy Excel, but it's still sort of doable. It's a lot of cells, but it's doable. But this is just with one word. And remember, when you just have one word to guess from, your guess is going to be pretty bad.

Gustav

现在假设你有两个单词可以猜测。那么你将有 5 万行乘以 5 万列,然后再乘以 5 万,因为你有两个单词的组合。所以现在你的电子表格里有大约 125 万亿个单元格,这已经超出了可解决的范围。

Now let's say that you get two words to guess from. So now you're going to have 50,000 rows times 50,000 columns, and then times 50,000 again because you have combinations of two words. So now you have something like 125 trillion cells in your spreadsheet, and it's getting beyond what is solvable.

Gustav

所以事实证明,虽然我说的理论看似简单,但实际上实践非常困难。在实践中进行这些统计真的非常非常难。

So it turns out that while the theory I said is deceptively simple, the practice is actually very hard. It's just very, very hard to do these statistics in practice.

进入Transformer Enter the Transformer

Gustav

这时 Transformer 登场了。2017 年,谷歌的机器学习科学家发表了一篇论文,名为《Attention Is All You Need》。论文提出了一种特定的机器学习架构,叫做 Transformer,就是你在图片上看到的这个。别担心,你完全不需要理解这张图。如果你说“哦,我认出来了,这是 Transformer”,那它看起来会很酷。

Enter the Transformer. There's this paper that came out from machine learning scientists at Google in 2017, called 'Attention Is All You Need.' It suggested a specific machine learning architecture called the Transformer, which is what you see in this image. And don't worry, you don't have to understand the image at all. It's just going to look cool if you say, 'Oh, I recognize that it's a Transformer.'

Gustav

你真正需要理解的唯一一点是,他们找到了一种巧妙的方法来解决如何拥有大量上下文的问题。事实证明,Transformer 可以处理数千个单词。实际上,它可以处理数万个,最近甚至可以处理大约十万个单词作为上下文,仅仅是为了猜测下一个词。

The only thing you really need to understand is that they managed to find a clever way of solving this problem of how do you have a lot of context. Turns out a Transformer can handle thousands of words. In fact, it can handle tens of thousands, and recently even something like a hundred thousand words as context, just to guess the next word.

Gustav

所以想想看。十万个单词。这 literally 意味着,作为人类,你的任务是读完整本书,然后把最后一页的最后一个词藏起来,但你有整本书作为上下文来猜测缺失的词。你的猜测会非常出色,对吧?你几乎会完全猜对,因为你有如此多的上下文。即使那个缺失的词非常不寻常,比如说故事结尾是“然后她去了……”,也许正确的词是“冥王星”,但你会知道,因为你了解小说的其余部分是关于太空的,而她正在去那里的路上,对吧?所以即使是非常不寻常的猜测,你也能做出正确的判断,因为你有上下文。

So think about that. A hundred thousand words. That literally means that the job that you would have as a human is you get to read an entire book, and you just hide the last word on the last page, but you have the entire book as context to guess the missing word. And your guess would be amazing, right? You would be almost completely correct in that guess because you would have so much context. Even if that missing word was something very unlikely, let's say that the story ends and then she went to... and maybe the correct word is Pluto, but you would know that because you know that the rest of the novel was about space and she was on her way there, right? So you would be able to make even very unlikely but correct guesses because you have context.

注意力机制 Attention Mechanism

Gustav

这就是 Transformer 机器学习模型解决的问题。这个架构,以及论文之所以叫《Attention Is All You Need》的原因,是因为它解决这个问题的方式是让模型在数学上真正地“注意”,对不同单词赋予不同的权重。它不会对所有单词一视同仁,所以问题不会像你的 Excel 表格那样爆炸式增长。它可以根据猜测来关注不同的单词。

This is the problem that the Transformer machine learning model solved. So this architecture, and the reason the paper is called 'Attention Is All You Need,' is because the way it solves this is that it allows the model to literally pay attention mathematically, to put different weights on different words. It doesn't put the same weight on all the words, so the problem doesn't blow up in the same way as your Excel sheet did. It can pay attention to different words depending on the guess.

Gustav

如果你想了解更多,你可以去阅读相关资料。但你真正需要理解的唯一一点是,Transformer 模型以一种非常巧妙的方式解决了上下文问题。如果你想深入研究,别害怕。数学并不难。实际上主要是线性代数,一点点微积分,但不是难的部分。这不是量子力学。所以如果你想深入,那就去吧。但这就是大家都在谈论的东西。这就是 Transformer 模型。它只是让机器能够以互联网规模进行这些统计。

Now, if you want to know more about this, you can read about it. But the only thing you really need to understand is that the Transformer model solves the context problem in a very clever way. And if you want to go into this, don't be afraid. The math isn't that hard. It's actually mostly a bit of linear algebra, a little bit of calculus, but not the hard stuff. This is not quantum mechanics. So if you want to go there, go there. But this is what everyone is talking about. This is what the Transformer model is. It's simply allowed machines to do these statistics on internet scale.

带上下文的例子 Example with Context

Gustav

好了,我们现在已经走得很远了。现在我们有了这个非常擅长猜测缺失单词的东西。让我们以这个句子为例:“我的狗叫 Ben。他是一只大爪子的大狗。Ben 喜欢和我玩接球。”假设大型语言模型试图隐藏“fetch”这个词。对,它停在“play”这个词上,然后需要猜测下一个词。

All right, so now we've come pretty far. So now we have this thing that is very good at guessing the missing word. So let's take this sentence for example: 'My dog's name is Ben. He's a big dog with large paws. Ben likes to play fetch with me.' Let's say that the large language model tries to hide the word 'fetch.' Right, so it is on the word 'play' and it's supposed to guess the next word.

Gustav

你可以想象,在只给你“play”这个词的模型中,最好的猜测可能是“play soccer”。这是互联网上最常见的词。也许“play basketball”。“play fetch”相当罕见。但现在有了上下文,语言模型可以说:“谁在玩?是 Ben。”然后它甚至更进一步回溯,说:“我的狗叫 Ben。他是一只大狗。”所以现在它知道 Ben 是一只狗。现在突然之间,有了这个上下文,最可能的下一步猜测可能就是“fetch”,因为那是跟随这个数字序列最常见的词或数字。

Well, you can imagine in the model where you just get the word 'play,' the best guess would maybe be 'play soccer.' It's the most common word on the internet. 'Play basketball' maybe. 'Play fetch' is pretty uncommon. But now that you have context, the language model can say, 'Who is playing? Well, it's Ben.' And then it goes even further back and says, 'My dog's name is Ben. He's a big dog.' So now it knows that Ben is a dog. And now all of a sudden, with that context, the most likely next guess probably is 'fetch,' because that's the most common word or number to follow that sequence of numbers.

Gustav

所以现在我们有了这个基于注意力的 Transformer,它非常擅长猜测缺失的单词,并且可以在互联网上的整个文本语料库上训练自己。所以我们快到了。现在我们可以生成语言了。如何生成语言呢?假设我们有这个句子:“数字 30 和 5 怎么样。”

So now we have this attention-based Transformer that is very good at guessing the missing word, and it can train itself on the entire corpus of text on the internet. So we're almost there. So now we can generate language. How do you generate language? Well, let's say that we have this sentence: 'How are the numbers 30 and 5.'

自回归模型如何生成文本 How Autoregressive Models Generate Text

Gustav

现在,训练这个模型时,你做的只是根据它在互联网上看到的一切,说出最可能的下一个词是什么。那么,它很可能会说“how are”之后最可能的下一个词,或者 35 之后最可能的数字是“you”。好,那么你加上“you”这个词,然后把你刚生成的文本放回模型里,说:“现在,根据这段文本‘how are you’,最可能的下一个词是什么?”模型会说:“好,在‘how are you’之后,我认为最可能的下一个词是‘I’。”好,我们加上它,然后我们再把这句话喂回给它,问:“在‘how are you I’这句话之后,最可能的下一个词是什么?”它很可能会说最可能的词是“am”。然后我们加上它,再把它喂回去,问最可能的下一个词是什么,很可能是“fine”。这就是你构建句子的方式,你可以一直这样继续下去。所以,你使用这些语言模型时,输入一些文本、一些数字,然后让它填出最可能的下一个数字,把它再喂回去,得到最可能的下一个数字,再喂回去,如此往复。顺便说一句,这就是所谓的自回归模型。如果你听过这个词,它就是指模型根据上下文,尝试猜出最可能或其中一个最可能的下一个词。

Now all you do when this model is trained is you say what is the most likely next word according to everything you've seen on the internet. Well, it's probably going to say the most likely next word after 'how are' or the most likely next number after 35 is 'you'. Okay, so then you add the word 'you', and then you take the text you just generated and put it back into the model and say, 'Now given this text, how are you, what is the most likely next word?' And the model says, 'Okay, after how are you, I think the most likely next word is I.' Okay, so we add that, then we take the sentence and feed it back into itself again and say, 'After the sentence how are you I, what is the most likely next word?' It's probably going to say something like the most likely word is 'am'. And then we add that, and now we feed this into itself and ask what is the most likely next word. It's probably 'fine'. And this is how you build sentences, and you can just keep going forever. So what you do with these language models is you put in some text, some numbers, and you just ask it to fill out the likely next number, take that, feed it in again, have the most likely next number, take that, feed it in again, forever. This is what is called an autoregressive model, by the way. If you ever hear that word, it just takes the context and tries to guess the most likely or one of the most likely next words.

补全文本与再现网络内容 Completing Text and Reproducing Internet Content

Gustav

所以现在我们可以做这样的事。比如,你可以把莎士比亚文本的一半喂给大型语言模型,对语言模型来说,这又是一长串数字,但它在互联网上见过这些数字。互联网上有很多莎士比亚文本。所以它会说:“嘿,我认识这些数字,我知道接下来是什么。”所以如果它只是不断取最可能的下一个数字,你得到的东西要么非常接近,要么实际上就是这部莎士比亚戏剧的其余部分。所以这很有用。现在我们有了这个东西,它可以接受一段文本,并用非常可能的内容补全它,如果它存在于互联网上,往往不仅是可能,甚至是一模一样的内容,因为如果你看到很多相同的莎士比亚戏剧,那些确切的词就是最可能的下一个数字,对吧?

So now we can do something like this. You could, for example, feed in half of the Shakespeare text into the large language model, which to the language model again will be a long series of numbers, but it will have seen these numbers on the internet. There's a lot of Shakespeare text on the internet, Shakespeare or Shakespearean text. So it's going to say, 'Hey, I recognize these numbers, I know what comes next.' So if it just takes them, takes the most likely next numbers again and again, you are going to get something that is either very close or quite literally actually the rest of this Shakespeare play. So that's pretty useful. Now we have this thing that can take a piece of text and complete it with something very likely, often actually if it existed on the internet, not just very likely but even the exact same thing, because if you see a lot of the same Shakespeare play, those exact words of the Shakespeare play are the most likely next numbers, right?

创造性与温度 Creativity and Temperature

Gustav

但创造力呢?我们刚才说 ChatGPT 不只是重复你在互联网上看到的东西。我们说它能写出从未存在过的诗,能写出新的文本。这怎么做到?嗯,我们需要再深入一层来理解,但请耐心听,我们快讲完了。现在你需要理解一个叫“温度”的概念。所以我们之前说的是我们有“how are”这些词,或者对计算机来说是数字 30 和 5。我们说我们取最可能的下一个词,也就是“you”。那是在温度为 0 的情况下。假设 0 意味着你取最可能的下一个词,然后取之后最可能的词,是“I”,再之后最可能是“am”,再之后最可能是“fine”。所以你会得到这个句子,而且每次都会得到这个句子,因为这些词总是最可能的。所以有一种误解,认为大型语言模型是概率性的,意味着它们每次给出不同的答案,因为当你使用 ChatGPT 时,即使同一个问题,每次确实得到不同答案。但实际上,如果温度为 0,大型语言模型是非常确定性的。如果你取最可能的下一个词,统计当然不会变,你总是会得到完全相同的句子。在这些大型语言模型中发生的是,当你希望它们有点创造性,而不是无聊地重复你从互联网上已知的内容时,你实际上做了不同的事。你可以选择一个非常可能但不是最可能的词。所以可以这样想:与其取最可能的词,在这个例子中是“you”,而且每次都是“you”,你不如在最可能的词周围稍微随机化。所以你会选择一个可能的词,但不一定是最可能的。你仍然会选择一个非常可能的词,这很重要,因为如果你去掉“非常可能”这个条件,句子在语法上仍然会通顺,因为这个词是可能的,而且在语义上仍然有意义,因为在机器学习模型的情况下,这个词或数字是可能的。所以假设我们不取最可能的词“you”,而是取第二可能的词,比如“day”。那么你实际上创造了从未在互联网上存在过的文本。统计上是可能的,句子会通顺,因为你选择的是在“how are”之后统计上会出现的词,但它不是互联网上已有内容的复制品。从不同的定义来说,它是新的。这就是所谓的提高温度。你提高温度越多,你就在最可能的词周围随机化得越多。你可以想象,如果你保持在顶部附近,它仍然会相当相似,但会是新的,不会在互联网上存在过,除非偶然,但它不是复制品,而是新颖的,但会非常接近互联网上已有的内容。你可以想象,温度越高,随机化越多,离最可能的词越远,模型就会变得越有创造性。但如果走得太远,它就会开始选择不太可能的词。所以如果你把温度调得太高,它可能会说“how are animals”。所以毫不夸张地说,如果你把温度调得太高,模型就会从看起来非常有创造性变成看起来有点精神错乱和疯狂。我认为这对人类来说是一个非常有趣的类比。往往最富有创造力的人类处于非常有创造力和有时看起来疯狂之间的边界上。也许这是巧合,也许不是,也许他们只是比我们其他人温度更高。总有一天我们会知道。

But what about creativity? We just said that ChatGPT doesn't just repeat the things you've seen on the internet. We just said that it can write poems that never existed, that it can write new text. How does that work? Well, we need to go one level deeper to understand this, but bear with me because we're almost there. Now you need to understand something called temperature. So what we said previously was that we have the words 'how are' or to the computer the numbers 30 and 5. And we said that we took the most likely next word, which is 'you'. That's if the temperature is zero. Let's imagine that zero means you take the most likely next word, and then you take the most likely next word after that, which is 'I', most likely next word after that, which is 'am', most likely next word after that, which is 'fine'. And so you're going to get this sentence, and you're actually going to get this sentence every time because these words will always be the most likely. So there's this misunderstanding that large language models are probabilistic, meaning that they give different answers every time, because when you use ChatGPT you actually do get different answers every time even for the same query. But in reality, large language models are very much deterministic if the temperature is zero. If you take the most likely next word, of course the statistics don't change, you will always get the exact same sentence. What is happening in these large language models is that when you want them to be a little bit creative and not just boring, not just repeat what you already know from the internet, you actually do something different. What you can do is you can pick something that is very likely but not the most likely. So think of it as instead of taking the word that is the most likely, which in this case would be 'you', and it will be 'you' every time, you randomize a little bit around the most likely words. So you're going to pick one of the likely words but not necessarily the most likely. So you're still going to take a word that is very likely, and that's important because if you take away that it's very likely, the sentence is still going to make sense grammatically because the word is likely, and it's still going to make sense semantically because the word or number in the case of the machine learning model is likely. So let's say that we take not the most likely word 'you' but we take the second most likely, which is 'day'. So now you actually created text that never existed on the internet. The statistics are likely, the sentence is going to make sense because you're picking something that statistically comes after 'how are', but it's not a copy of what existed on the internet. It is by different definition something new. And this is called raising the temperature. The more you raise the temperature, the more you sort of randomize around the most likely words. And you can imagine that if you stay very close to the top, it's going to still be quite similar, it's going to be new, it will not have existed on the internet unless by chance, but it's not a copy of what existed on the internet, it will be novel, but it's going to be very close to what exists on the internet. And you can imagine that the more you raise the temperature, the more you randomize, the further you go from the most likely, the more creative the model is going to get. But if you go too far, it's going to start to pick words that are unlikely. So if you raise the temperature too much, it might just say 'how are animals'. And so quite literally, if you raise the temperature too much, the model goes from seeming very creative to starting to seem a bit unhinged and insane. And I think it's a very interesting analogy to humans here. It tends to be the fact that the most creative humans are somewhere on the borderline of very creative and sometimes they just seem crazy. And maybe that's a coincidence, maybe it's not, maybe they just have a higher temperature than the rest of us. We'll know someday.

从GPT-3到ChatGPT From GPT-3 to ChatGPT

Gustav

好了,所以现在我们有了一个大型语言模型,它不仅可以用最可能的文本补全一段文字,而且如果你把温度调高一点,它实际上可以用从未存在过的新颖内容来补全。所以现在我们很接近了。这实际上是 GPT-3,大约一年半前发布的。那有点像“打了类固醇的自动补全”。它可以接受一些文本,并非常可信地补全它,如果你提高温度,它可以用从未写过的东西来补全,对吧?但它不是 ChatGPT。对于那些能访问它的人来说,它可以说是“只是打了类固醇的自动补全”。所以一种思考方式是……

Alright, so now we have a large language model that can not only complete a piece of text with the most likely text, but if you raise the temperature a little bit, it can actually complete the text with something novel that never existed before. So now we're pretty close. This is actually GPT-3, which was about a year and a half ago that came out. That was sort of an autocomplete on steroids. It could take some text and it could complete it very believably, and if you raise the temperature, it could complete it with something that was never written before, right? But it wasn't ChatGPT. For those of you who had access to it, it was quote unquote just an autocomplete on steroids. So a way to think about...

从基础模型到聊天助手 From Base Model to Chat Assistant

Gustav

我们现在处于什么阶段,而 OpenAI 和其他公司大约一年半前又处于什么阶段?

Where are we now, and where were OpenAI and other companies about a year and a half ago?

Gustav

在 GPT-3 阶段,它属于基础模型。你可以把它想象成一个孩子或一个不太靠谱的青少年。它拥有大量关于世界的知识,能补全文本,能生成从未存在过的文本,但你无法真正引导它或格式化它。你希望拥有的是像 ChatGPT 这样的东西,它不只是补全句子,而是真正回答你的问题。你还希望引导它避开某些领域,谈论其他领域。比如,你可能希望它不回答如何制造炸弹的问题。

At the GPT-3 stage, it's sort of a base model. You can think of it as maybe a kid or an unhinged teenager. It has a lot of knowledge about the world, it can complete text, it can generate text that never existed, but you can't really steer it or format it. What you'd like to have is something like ChatGPT, where it doesn't just complete sentences but actually answers your questions. You'd also like to steer it to stay away from certain areas and talk about others. Maybe you'd like it to not answer questions about how to create a bomb, for example.

Gustav

那么我们如何达到最后那个阶段呢?

So how do we get to that last stage?

Gustav

这需要两样东西:一种叫监督微调(SFT),另一种叫基于人类反馈的强化学习(RLHF)。又是针对外行的阴谋——SFT 和 RLHF——这正是让外行远离你专业领域的方法。让我深入讲讲它实际的含义。其实并不复杂。

That requires two things: something called supervised fine-tuning (SFT) and something called reinforcement learning with human feedback (RLHF). Again, conspiracies against the laity—SFT and RLHF—that's exactly how you get people to stay away from your profession. Let me dig into what it actually means. It's not very complicated.

Gustav

我们先从监督微调开始。

Let's start with supervised fine-tuning.

Gustav

我们说,我们现在拥有的是一台能接收一串数字并猜测最佳后续数字的机器。所以它就像你的 iPhone 自动补全的加强版。现在想象你从互联网上拿一堆恰好是问答格式的文档——有一个问题,然后有人回答了。如果你让模型自动补全这些内容,它会学到一种模式:每次看到问题,那种类型的数字后面总有答案。所以如果你幸运的话,即使是基础模型,如果你向它提出一个问题,而它在互联网上见过大量问答对话,它可能真的会给你一个答案,仅仅因为这是互联网上很常见的语言模式。你可以说:“问:地球的宽度是多少?”它可能会说:“答:某个数字。”对吧?但它也可能只是用另一个问题来回答,因为这在互联网上也很常见。在它见过的所有文档中,有问-问的情况,比如没有答案的数学测试。所以它有点像拥有一些关于问答的知识,但并不是确定性的。有时你让它自动补全正确的结构,能得到你想要的结果;有时则不然。

We said that what we have now is a machine that can take a sequence of numbers and guess the best next numbers. So it's your iPhone autocomplete on steroids. Now imagine you take a bunch of documents from the internet that happen to be Q&A documents—there's a question and someone else answered it. If you train the model on autocompleting that, it will learn the pattern that every time it saw a question, those types of numbers, there was always an answer. So if you're lucky, even with the base model, if you post a question to it and it has seen a lot of question-answering dialogues on the internet, it may actually give you an answer just because it's a pretty common language pattern on the internet. You could say, "Q: What is the width of the earth?" and it could say, "A: some number," right? But it could also just answer with another question, because that's also common on the internet. In all the documents it's seen, you have question-question, for example like a math test with no answers. So it's kind of like it has some knowledge about question-answering, but it's not deterministic. Sometimes you get what you want if you ask it to autocomplete the right structure; sometimes you don't.

Gustav

那么监督微调有什么不同呢?

So what does supervised fine-tuning do differently?

Gustav

基础模型是自监督的——它确实通过在互联网上逐词(或逐 token)隐藏来训练。顺便说一句,对于懂行的人,一个 token 并不完全等于一个词,但为了本次演讲的目的,我们假设一个 token 就是一个词或一个数字。所以基础模型在互联网上的所有文本上训练,即数万亿 token 的文本。现在你做一些不同的事情:你创建一个小型监督数据集。“监督”是什么意思?意思是,与自监督不同,这是由人类监督的。你创建一批文档,这些文档正是你想要的格式——总是问题后跟答案。你创建的数量,相对于互联网来说很小,比如 10,000 个这样的文档。所以你让人类去创建 10,000 个问答文档,一问一答。然后你拿这个在互联网上训练过的基础模型,进行所谓的微调。你多训练它一点;你不是从头开始——它已经训练过了——你只是在那个总是问答的数据集上多训练一点。然后,凭直觉,语言模型保留了所有基础知识,但学会了这种行为。所以它会开始把所有内容都当作问题的答案来回答,仅仅基于最后那个额外的数据集。你几乎可以认为它有点像记住了你最后做的事情,如果你过度强调了“一切都应该是问题的答案”,它就会开始模仿那种行为。所以现在它拥有所有世界知识,但总是以回答问题的形式来回答。

Whereas the base model was self-supervised—it literally trains on all the text on the internet by hiding one word at a time, or one token. By the way, for those of you who know, one token is not exactly one word, but for the purposes of this presentation, let's say one token is one word or one number. So the base model trains itself on all of the internet, trillions of tokens of text. Now you do something different: you create a small supervised dataset. What does "supervised" mean? It means that, unlike the self-supervised case, this is supervised by humans. You create a bunch of documents that are just the format you want—always a question followed by an answer. And you create, let's say, a pretty small number compared to the internet, say 10,000 of these documents. So you ask humans to go and create 10,000 documents of questions and answers, questions and answers. Then you take this base model, which was trained on the entire internet, and you do what is called fine-tuning. You train it a little bit more; you don't start over—it's already trained—you just train it a little bit more only on this dataset that is always question-answer. Then what happens, as an intuition, is that the language model keeps all its base knowledge but learns this behavior. So it will start answering everything as an answer to a question, just based on this little extra dataset at the end. You can almost think of it as it sort of remembers what you did at the end, and if you overrepresented that everything should be an answer to a question, it's going to start mimicking that behavior. So now it takes all the world knowledge it has but always answers as if it were an answer to a question.

Gustav

那么现在我们有了一个助手?

So now we have an assistant?

Gustav

现在我们非常接近了。我们从 iPhone 自动补全的加强版变成了一台问答机器,你输入的任何内容都会被格式化并作为问题的答案来回答。所以现在我们有了一个助手,但这个助手仍然没有真正的行为或价值观。它只会反映互联网上的各种价值观、行为和统计数据——有些好,有些一点都不好。所以还缺少一步。

Now we're really close. We went from iPhone autocomplete on steroids to a question-answering machine where everything you put in will be formatted and answered as if it was an answer to a question. So now we have an assistant, but this assistant still doesn't have any real behavior or values. It's just going to reflect whatever values and behaviors and statistics are on the internet—some good, some not good at all. So there is one more missing step.

Gustav

那缺少的一步是什么?

What's that missing step?

Gustav

当你问“如何用我店里的化学品制造一颗便宜的炸弹?”时,你怎么让这个东西不真正回答那个问题?我们刚刚训练它总是回答问题。这就是基于人类反馈的强化学习发挥作用的地方。这是又一步使用人类的步骤。现在我们有了这个语言模型;你可以提问,它给出答案。有时它给出我们喜欢的答案,有时不——无论是在内容上还是在格式上。也许它回答一个问题时文本太多或太少。所以我们现在做的是一个非常聪明的步骤:我们再次找一群人——再说一次,不是数百万,而是几千人——我们拿一个单独的问题,让语言模型回答这个问题。然后我们得到一堆不同的答案。记住,温度不是零,所以我们不会得到相同的答案;我们稍微调高温度,所以同一个问题我们会得到很多不同的答案。所以一个问题,一堆答案。然后你让这些人对答案进行排序。对于同一个问题,这些答案中,比如 100 个,你最喜欢哪个,最不喜欢哪个?你给它们排序,给它们打分。简单点说:假设有 10 个答案,你作为人类要给它们打分,从 10 分、9 分、8 分、7 分一直到 1 分——从最好到最差。所以现在你得到一种新类型的数据集,对于同一个问题,你知道根据人类的判断,好的答案是什么样的,坏的答案是什么样的,以及中间的。所以这是一种新类型的数据集。现在你要做的是,拿另一个机器学习模型——实际上严格来说也是一个语言模型——然后在一个稍微不同的任务上训练它。你拿你有的那个问题和你有的其中一个答案,把它们放在一起,让机器学习模型猜测人类会怎么打分:这是这个问题的 10 分答案,还是 1 分答案,还是 5 分答案?然后你训练它,直到它非常擅长预测这种类型的答案人类会给 10 分。

When you ask that question of "How do I create a cheap bomb out of chemicals from my store?" how do you get this thing to not actually answer that question? We just trained it to always answer questions. This is where reinforcement learning with human feedback comes in. This is yet another step where you use humans. So now we have this language model; you can ask questions and it gives answers. Sometimes it gives answers we like, sometimes it doesn't—both in content and maybe in formatting. Maybe it answers a question with too much text or too little text. So what we do now is a really clever step: we take a bunch of humans again—again, we're talking not millions but a few thousand—and we take a single question and ask the language model this question. Then we get a bunch of different answers. Remember, the temperature is not zero, so we don't get the same answer; we raise the temperature a little, so we're going to get a lot of different answers for the same question. So one question, a bunch of answers. Then you ask these humans to rank the answers. For this same question, which of these, let's say 100 answers, did you like the most and the least? And you rank them, you give them a score. Let's make it simple: let's say it's 10 answers, and you as a human are supposed to score them from 10 points, nine, eight, seven, all the way to one—best to worst. So now you get a new type of dataset where, for the same question, you know what a good answer looks like according to a human, and what a bad answer looks like according to a human, and in between. So this is a new type of dataset. Now what you do is you take another machine learning model—actually technically also a language model—and you train it on a slightly different task. You take the question you had and one of the answers you had, put them together, and ask the machine learning model to guess how a human would have scored it: was this the 10 answer to the question, or the one answer, or the five-point answer? Then you train it until it gets really good at predicting that this type of answer a human would have scored 10.

奖励模型与强化学习 Reward Model and Reinforcement Learning

Gustav

于是你构建了一个叫奖励模型的东西,这个模型擅长根据人类提供的监督数据来猜测人类会如何评价这个回答。

So you build something called a reward model, a model that is good at guessing, based on supervised data from humans, what the human would have thought about this answer.

Gustav

好,现在我们快讲完了。我们有大型语言模型,我们通过监督微调让它总是以被提问的方式作答,现在它能产生回答,但回答质量参差不齐。现在我们有了另一个模型,即奖励模型,它能查看一个回答,并说出人类会如何评价它,是好是坏。现在你只需把这两个模型接在一起,然后放手让它运行。这就是强化学习部分。现在它能再次自我提升。它拿一个问题,生成一个回答,给自己的回答打分,然后说:“这不好,我应该做得更好,我应该这样做。”它给那个回答打分,然后说:“根据人类标准,这很好,我应该多做这类。”所以现在它能针对奖励模型自我训练。这就是所谓的强化学习,而且这是一个封闭系统,不需要人类参与。所以你可以这样做数百万次、数千万次、数亿次,直到它变得非常擅长回答,不仅符合人类想要的格式,而且在风格上也符合,甚至可以说,如果你选择你想要的价值观,它也能符合。

Okay, so now we're almost there. We have the large language models, we supervised fine-tune it to always answer as if it was asked the question, and now it produces answers, but it produces answers of varying quality. Now we have this other model, which is a reward model, that can look at an answer and say what a human would have thought about it, if it was good or bad. Now you just take these two models, you hook them together, and you just let it go. This is the reinforcement learning part. Now it can surprise itself again. It takes a question, generates an answer, scores its own answer, and says, "That was bad, I should do better, I should do this." It scores that answer and says, "That was good according to human, I should do more of this." So now it can train itself against the reward model. This is what it's called the reinforcement learning, and this is a closed system where you don't need humans. So you can do this millions, tens of millions, hundreds of millions of times until it gets really good at answering, not just in the format that a human wants, but also in the style and literally if you choose the values that you want.

Gustav

所以我觉得这很有趣,因为人们会问:“这个模型在想什么?它的价值观是什么?”而事实是,在基础层,模型的价值观只是整个互联网的平均值,即互联网认为的好与坏。但强化学习这一步实际上注入了某种行为,而这实际上是由一小群人决定的。因此,这个模型的行为在很大程度上取决于这一环节。

So I think this is interesting because people ask like, "What does this model think? What are the values?" And the truth is in the base layer, the values of the model are just an average of the entire internet. It is what the internet thinks good and bad. But the reinforcement learning step is actually what inserts a certain behavior, and that is actually a pretty small group of humans. So that is where a lot of the responsibility lies for how this model behaves.

为何令人惊讶? Why the Surprise?

Gustav

好了,这就是 ChatGPT。现在你明白它是如何工作的了。你从历史一路走到了 2023 年,你明白了这些东西是如何通过图灵测试的。这并不复杂,对吧?至少你可以想象,如果你有无限的时间和无限大的 Excel 表格,你作为人类会如何解决这些问题。那么现在有一个问题:为什么每个人都如此惊讶,为什么可以说没有人预见到它的到来?有些人声称他们预见到了,但大多数人没有。即使是机器学习科学家总体上对这一切发生得如此之快也感到非常惊讶。机器学习模型本身并不新鲜。2017 年的 Transformer 架构显然是一项创新,但语言模型已经存在,语言建模也已经存在很长时间了。那么,是什么让包括专家在内的人们感到惊讶呢?

All right, that is ChatGPT. Now you understand how it works. You've gone from back in history all the way to 2023. You understand how these things are passing the Turing test. It wasn't that complicated, was it? At least you can imagine how you as a human would solve these things if you had infinite time and an infinitely big Excel sheet. So now one question is: why is everyone so surprised, and why did, quote unquote, no one see it coming? Some people claim they did, but most didn't. Even machine learning scientists are very surprised in general about how quickly this happened. And the machine learning models themselves aren't that new. The Transformer architecture in 2017 was clearly an innovation, but language models have been around, and modeling language has been around for a long time. So what was it that surprised people here, including the experts?

Gustav

嗯,是规模和速度。预测下一个词的概念并不是最近才发明的;人们尝试这样做已经很长时间了。当 Transformer 架构出现时,大规模实现变得更容易。但不太明显的是,仅仅把这种简单的事情做得更多,就会开始产生全新的行为,即所谓的涌现行为。所以你可以看到这些大型语言模型在某些事情上并不擅长,比如数学,然后在不改变架构的情况下,仅仅通过扩大规模,它突然就开始擅长这些事情了。这令人惊讶,因为你会得到这些涌现能力,而许多人认为这些能力需要某种新的数学或架构创新。所以仅仅是规模本身就能提高性能,这一点非常令人惊讶。

Well, it was scale and speed. The notion of guessing the next word was not something that was recently invented; people have been trying to do this for a long time. And when the Transformer architecture came along, it became easier to do it at scale. But what wasn't obvious was that just doing this simple thing much more would start giving completely new behaviors, what is called emergent behaviors. So you could see these large language models not being very good at certain things like math, for example, and then without changing the architecture, just by scaling it up, all of a sudden it started getting good at things. And that was surprising, that you would sort of get these emerging capabilities that many people thought would require some new mathematical or architectural innovation. So just that scale alone improved performance was very surprising.

Gustav

另一件我认为让大多数机器学习从业者感到惊讶的事情是这种创造力。温度的概念对机器学习从业者来说并不新鲜,它已经存在很久了。但同样,概念是存在的,但即使是机器学习从业者也惊讶于它实际上能扩展到看起来非常像人类创造力的东西。我们是否就是这样运作的,这一点非常非常未知,但结果看起来确实像我们做的。所以要么我们获得了我们所拥有的那种创造力,要么我们设法模拟了我们所拥有的那种创造力,而这让人们感到惊讶。

The other thing that I think is surprising to most machine learning people is this creativity thing. Now the temperature notion was not surprising to people in machine learning; it's been around forever. But again, the concept was there, but even machine learning people were surprised that it actually scales to something that looks very much like human creativity. It is very, very much unknown if this is what we do, but certainly the result looks like what we do. So either we got the type of creativity that we have, or we managed to simulate the type of creativity that we have, and that surprised people.

Gustav

最后,缺失的是引导它的能力。监督微调赋予它某种行为,使其成为助手,然后强化学习使其能够实际使用。你知道,GPT-3 作为自动补全工具确实很酷,但它有点失控。监督微调使界面变得可用,因为它变成了一个问答机器,但仍然有点失控。然后基于人类反馈的强化学习使其在现实中变得实用。所以这些是关键。我认为这里真正有趣的不是我们惊讶于我们最终设法破解了智能,以及它有多复杂。让人们惊讶的实际上是相反的:我们有点破解了智能,而且它如此简单,几乎是挑衅性的简单。

And lastly, the thing that was missing was this ability to steer it. The supervised fine-tuning to give it a certain behavior, to be an assistant, and then the reinforcement learning to be able to use it practically. You know, GPT-3 was really cool as an autocomplete, but it was a bit unhinged. And supervised fine-tuning made the interface workable because it was a Q&A machine, but it was still unhinged. Then reinforcement learning with human feedback made it practically useful in reality. So those were the unlocks. And I think what is really interesting here isn't that we're surprised that we finally sort of managed to crack intelligence and how complicated it was. What surprises people is actually the opposite: that we kind of cracked intelligence and it was so simple. It's almost provocatively simple.

统计鹦鹉与自我反思 Statistical Parrots and Self-Reflection

Gustav

所以大约一年前,在 ChatGPT 或 GPT-3 出现的时候,这些模型经常被称为“只是统计鹦鹉”,这是一个贬义词。实际上,在 GPT-2 时期,人们说:“这些只是统计鹦鹉”,意思是它们只是像我说的一样,鹦鹉学舌地重复互联网上的统计数据。比如,“这不是真正的智能。它并不令人印象深刻,实际上可能有用,但不算惊艳。”但随着 GPT-3、GPT-3.5、ChatGPT、GPT-4 的出现,“这些东西难道不只是统计鹦鹉吗?”这个问题变成了“天哪,如果我们只是统计鹦鹉呢?”所以这真的让我们照见了自己,让很多人开始思考自己是什么,我认为这非常令人兴奋。

So about a year ago, around ChatGPT or GPT-3, these models were often called, quote unquote, "just statistical parrots," and that was meant as a derogatory term. Actually, with GPT-2, they were saying like, "These are just statistical parrots," meaning that they just parrot back, as I said, statistics from the internet. Like, "This isn't real intelligence. It's not very impressive, actually. Maybe useful, but not impressive." But then as GPT-3 came along, GPT-3.5, ChatGPT, GPT-4, this question of "Aren't these things just statistical parrots?" turned to "Oh crap, what if we are just statistical parrots?" And so it really puts a mirror to ourselves, and it gets a lot of people to start thinking about what they are, which I think is very exciting.

回顾与向量简介 Recap and Introduction to Vectors

Gustav

好了,我们完成了第一部分。我们实际上了解了什么是大型语言模型以及 ChatGPT 是如何工作的。现在我认为你对实际发生的事情有了和大多数人一样好的直觉。所以现在希望你对类似 GPT 的东西如何工作、为什么它这样工作、它如何理解问题和答案,以及为什么当你问它如何制造炸弹时,它实际上不会告诉你,而是说“作为大型语言模型,我不会回答这个问题”有了至少一个直觉。这就是基于人类反馈的强化学习部分。

All right, so we did the first part. We actually went through what a large language model is and how ChatGPT works. And now I think you have as good an intuition as most people about what it is that actually happened. So now hopefully you have at least an intuition for how something like GPT works, why it works the way it does, how it can understand questions and answers, and why when you ask it how to create a bomb, it doesn't actually tell you, and rather it says, "As a large language model, I'm not going to answer that question." This is the reinforcement learning with human feedback part.

Gustav

好了,所以你明白了语言在某种程度上只是统计,因为语言可以表示为数字,也许对我们来说甚至也是如此,而数字只是统计。但我想教你另一件我觉得很酷的事情,同样被搞得过于复杂了。它实际上并不难理解,但一旦你理解了,你会有点震惊。你可能听说过一些叫做向量、向量空间、嵌入、编码或分布式表示的东西。所有这些花哨的词汇,但你不一定进一步理解它们是什么。我将向你解释它们是什么。我将再次向你展示这是一个非常直接的概念,但是……

All right, so you understand that language is sort of just statistics, because language can be represented as numbers, and maybe they are even to us actually, and numbers are just statistics. But I want to teach you one more thing that I think is really cool, again made overly complicated. It's not actually that hard to understand, but once you understand it, your mind is a little bit blown. So you may have heard about something called vectors, or vector space, or embeddings, or codes, or distributed representations. All of these fancy words without necessarily further understanding what they are. I'm going to explain to you what they are. I'm going again to show you that it's a very straightforward concept, but...

词向量与语义维度 Word Vectors and Semantic Dimensions

Gustav

依然非常非常酷。所以我们说过,语言可以被表示为数字,你可以简单地给字典里的每个词一个编号。但事实证明,与其只给一个词一个数字(这当然可行),你还可以做得更好一点。所以你可以取一个词,用三个或四个数字来表示它,而不是仅仅说“国王”这个词是数字 54。我这么说是什么意思呢?让我给你演示一下。

Still very very cool. So we said that language can be represented as numbers and you can simply give every word in the dictionary its own number. But it turns out that instead of giving a word just one number, which you can do, you can also do a little bit better than that. So you can take a word and you can represent it, instead of just saying the word 'king' is the number 54, you can say I'm going to use like three numbers or four numbers to represent the word 'king'. What do I mean with that? Let me show you.

Gustav

让我们取一个非常简单的世界。假设我们生活在一个只有三个维度的宇宙里。只有三个维度。事物要么是“王权性”,要么是“男性气质”,要么是“女性气质”。这是一个非常简化的宇宙,一切事物都只有三个维度。那么在这个宇宙里,你可以取一个词,比如“国王”,然后你可以说,与其说“国王”是数字 29,不如说这个词里含有多少这些维度。所以你可以说“国王”这个词几乎有 100% 的王权性,也就是 0.99。在统计学里没有什么是百分之百的,所以 1 表示 100%,0 表示 0%。所以 0.99 意味着几乎 100% 的王权性,对吧?因为国王几乎总是王室的。但“国王”这个词在男性气质上也非常高,几乎接近 100%。我们不知道,但历史上我们认为大多数国王到目前为止都是男性。然后还有女性气质,这个维度很低,可能不是零——没有什么是绝对的零——但在统计上平均较低,所以是 0.05。所以现在,对于“国王”这个词,我们不再用一个单一的数字,而是有三个数字,对应着这个词在某些维度上含有多少某种属性。

Let's take a very simple world. Let's say we live in a universe that only has three dimensions in it. Only three dimensions. Things are either royalty, they are masculinity, or they are femininity. It's a very simplified universe. It only has three dimensions to everything. So now in this universe, you can take a word, for example like 'king', and you can say, instead of just saying 'king' is the number 29, you can say how much of these dimensions are in the word 'king'. So you can say that the word 'king' has almost 100% royalty, so 0.99. Nothing in statistics is a hundred percent, so 1 means 100%, 0 means zero percent. So 0.99 means almost 100% royalty, right? Because the king is almost always a royal. But the word 'king' is also very high on masculinity, it's almost 100%. We don't know, most kings so far have been men, historically we think. And then you have femininity, which is low, probably not zero, nothing is ever zero, but low statistically on average, so 0.05. So now instead of a single number for the word 'king', you have three numbers which correspond to some dimensions of how much of something that word is.

Gustav

现在让我们看另一个词,比如“女王”。所以“女王”也大约有 100% 的王权性,对吧?女王几乎总是王室的。它在男性气质上非常低,所以假设是 0.05,几乎为零,而在女性气质上非常高,比如 0.98。让我们再看一个词,“女人”。女人可能是王室成员,但从统计上看,在整个群体中这个比例相当低,几乎为零。男性气质非常低,女性气质非常高。让我们看一个像“公主”这样的词。“公主”很有趣,因为它几乎总是王室的,王权性非常高。它通常不具有男性气质,而女性气质非常高。好了,所以现在我们有这四个词,它们不是用单一数字描述的,而是用代表它们在某个维度上含有多少属性的数字来描述的。在这个简单的宇宙里,只有三个维度:王权性、男性气质和女性气质。

Now let's take another word, let's take 'queen'. So 'queen' is also about 100% royalty, right? A queen is almost always royal. It is very low in masculinity, so let's say 0.05, almost zero, and very high on femininity, 0.98 for example. Let's take another word, 'woman'. So a woman could be royalty, but statistically across the population pretty low, it's almost zero percent. Very low masculinity and very high on femininity. Let's take a word like 'princess'. So 'princess' is interesting because it is almost always royalty, it's very high on royalty. It is usually not masculine and very high on femininity. All right, so now we have these four words described not as single numbers but as numbers that represent how much they are of some dimension. And in this simple universe, only three dimensions: royalty, masculine, and feminine.

Gustav

所以现在这些词不是用单一数字描述的,而是用多个数字描述的,这些数字实际上代表了这些词在某个维度上含有多少某种属性。在这个简单的宇宙里,我们只有这三个维度:某物有多少王权性、男性气质和女性气质。但你可以很容易想象,就像我们之前做的那样,不用这三个维度,而是直接把整个英语词典作为维度。所以也许你可以有一个“年龄”维度,你可以说“国王”这个词,如果 1 是 100% 的老年,0 是青年,那么国王可能在年龄上是 0.7,即 70%,他是年老的。“女王”可能是 0.6,平均年轻一点。“女人”字面上是 0.5,即 50%,正好在中间。而“公主”通常年轻,所以可能是 0.1。你可以就这样遍历整个英语词典,作为人类,你可以凭直觉尝试给英语词典中每个词在多大程度上含有其他每个词的属性打一个百分比。这说得通吗?

So now we have these words not described as a single number but as several numbers, and these numbers actually represent how much of something, of some dimension, are in these words. In this simple universe, we just have these three dimensions: how much something is royalty, masculine, and femininity. But you could easily imagine, as we did before, that instead of these three dimensions, you literally take the entire English dictionary as dimensions. So maybe you could have a dimension that is age, and you could say the word 'king', if 1 is 100% old age and 0 is young age, a king is maybe 0.7, 70% in terms of age, he's old. 'Queen' maybe 0.6, a bit younger on average. 'Woman' literally 0.5, 50%, right in between. Whereas a princess would be usually young, so maybe 0.1. And you could just go down the English dictionary, and as a human you can intuitively try to put a percentage on how much of every word in the English dictionary is in every other word of the English dictionary. Does that make sense?

Gustav

所以你取“国王”这个词,在最坏的情况下,你要拿这 60 万个词,然后试着给“国王”在多大程度上是王权性、男性气质、女性气质、年龄、汽车等属性打一个百分比,这些属性会是 0%。所以实际上大多数属性会是 0%。但你可以想象一个非常长的向量:我们有一个百分比,表示每个词在多大程度上含有其他每个词的属性。所以那会是一个有 60 万个数字的向量。而实际上,你并不是这么做的。这些模型通常有大约 1000 个维度,它们会挑选最有用的维度。我稍后会再讲它是如何挑选的。但如果你在想“60 万个维度似乎很笨重”,知道这一点是有好处的,确实如此。你会有大约 1000 个维度来描述每个词以及这些词中含有多少某种属性。但为了从实践回到原理,让我们再次回到我们那个只有三个维度的简单宇宙。

So you take the word 'king', and you take, in the worst case, these 600,000 words, and you try to put a percentage on how much is 'king' royalty, masculinity, femininity, age, car, things that would be zero percent. So most of these would actually be zero percent. But you can imagine a very long vector: we have a percentage for how much of every word is in every other word. So it would be a vector that is 600,000 numbers. And in reality, that's not how you do it. These models usually have about 1,000 dimensions, and they sort of pick the most useful dimensions. And I'll come back to how it picks them later. But it could be good to know, if you're wondering like 'seems unwieldy with 600,000', that's true. You would have about a thousand dimensions that describe every word and how much of something is in those words. But for purposes of principles from practice again, let's go back to our simple universe where there are only three dimensions.

Gustav

所以现在我们有这些词,它们被描述为含有这三个维度的多少。现在我们可以做一些非常酷的事情:我们实际上可以用语言做数学运算,因为它们被表示为数字。让我给你演示。现在我们有这个三维的宇宙。我们取一个词,取那个词的向量。我们取“国王”,它几乎有 100% 的王权性,几乎 100% 的男性气质,以及几乎 0% 的女性气质。然后我们直接减去一个“男人”。所以我们取“国王”并减去“男人”。让我们在这里做一下数学。王权性发生了什么?我们有“国王”的 0.99 王权性减去“男人”的 0.01 王权性,这意味着 0.98 的王权性。我们有“国王”中 0.99 的男性气质减去“男人”中 0.99 的男性气质,所以现在我们有 0% 的男性气质。我们有 0.05 的女性气质减去 0.05 的女性气质,所以我们有 0% 的女性气质。所以我们取了一个“国王”,从“国王”中减去了“男人”,我们得到了一个新的词向量。你认为这是什么词?什么是 100% 王权性但没有性别?那就是王权性,纯粹的王权性。好的,所以我们得到了一个新词。

So now we have these words that are described in how much they have of these three dimensions. Now we can do something really cool: we can actually do mathematics with language because they're represented as numbers. Let me show you. Now we have our universe with the three dimensions. We take a word, we take the vector for that word. We take 'king', which was almost 100% royalty, almost 100% masculinity, and almost zero percent femininity. And then we simply literally subtract a man. So we take a king and we subtract a man. Let's do the math here. What happens to the royalty? So we had 0.99 royalty on the king minus 0.01 royalty on the man, that means 0.98 royalty. We had 0.99 masculinity in the king minus 0.99 masculinity in the man, so now we have zero percent masculinity. And we had 0.05 femininity minus 0.05 femininity, so we have zero percent femininity. So we took a king, we subtracted the man from the king, and we got a new word vector. What word do you think this is? What is it that is 100% royalty but it's genderless? It is royalty, pure royalty. Okay, so now we got a new word.

Gustav

如果我们取纯粹的王权性,然后加上一个“女人”,会发生什么?让我们来做数学。我们有 0.98 的王权性加上“女人”中另外 0.02 的王权性,这意味着我们真的达到了 100% 的王权性。我们有 0% 的男性气质加上 0.01 的男性气质,这几乎就是 0% 的男性气质。我们有 0% 的女性气质加上 0.999 的女性气质,所以几乎 100% 的女性气质。所以现在我们有了这个新的词向量,它几乎有 100% 的王权性和几乎 100% 的女性气质。那是什么?是“女王”。所以现在突然间,你可以取一个“国王”,减去“男人”,加上“女人”,然后你就得到了“女王”。所以现在你确实是在用词做数学运算。这就是为什么向量如此有趣和有用,因为它们编码了它们含有多少某种属性。所以这可能非常有用。我将向你展示如何有用。但首先,一个问题可能是:在实践中你会怎么做?在理论上,同样,作为人类,你可以坐在那里猜测这些百分比。实际上,我们刚刚就是这么做的。所以直觉上,如果你能做到,计算机可能也能做到。但你会如何用统计方法来做呢?

What happens if we take pure royalty and we add a woman? So let's do the math. We have 0.98 royalty plus another 0.02 royalty in the woman, that means we literally get to 100% royalty. We had zero percent masculinity plus 0.01 masculinity, it's almost zero percent masculinity. We had zero percent femininity plus 0.999 femininity, so almost 100% femininity. So now we have this new word vector which is almost 100% royalty and almost 100% femininity. What is that? A queen. So now all of a sudden you can take a king, you can subtract the man, you can add a woman, and then you have a queen. So now you're quite literally doing math with words. And this is why vectors are so interesting and useful, because they encode how much they are of something. So this can be really useful. And I'm going to show you how. But first, a question might be: how would you go about doing this in practice? In theory, again, you as a human could just sit and guess at these percentages. Actually, we just did. So intuitively, if you can do it, the computer could probably do this. But how would you do it statistically?

Gustav

嗯,再说一次,有互联网。所以这里有一种方法。假设你取整个互联网,或者也许维基百科,对于你想学习的任何词,在这个例子中,焦点词是“学习”,你只需要说其他词在句子中离它有多近。所以例如,在这个句子中:“一种用于学习高质量分布式向量的高效方法”,你可以看到……

Well, again, there is the internet. So here's one way of doing it. Let's say that you take the entire internet, or maybe Wikipedia, and for any word you want to learn, in this case for example the focus word is 'learning', you just say how close the other words are to it in a sentence. So for example, in this sentence: 'An efficient method for learning high quality distributed vector', you can see...

词向量及其学习 Word Vectors and Their Learning

Gustav

你解释了像“for”和“and”这样的词靠近“learning”,而“inefficient”和“distributed”则很远。这是怎么做到的?

So you've explained that words like 'for' and 'and' are close to 'learning', but 'inefficient' and 'distributed' are far away. How does that work?

Gustav

对,所以“for”和“and”正好在“learning”周围,因此计算机给它们高分,因为它们在文本中确实靠近“learning”。而“inefficient”和“distributed”则更远,所以得分较低。如果你采用这个简单方法——遍历互联网上的所有文档,对每个词看哪些词在它附近出现——那些词可能相关,所以得到高百分比。在所有文档中远离该词的词则得到低百分比。所以,就像其他方法一样,你有一种可扩展的统计方法来学习这些统计量和向量。现在你不仅直观理解了词向量是什么,还知道了如何根据词在句子中与其他词的接近程度,自动学习互联网上每个词的向量。

Right, so the word 'for' and 'and' are right around the word 'learning', so the computer will give them a high score because they are literally close to the word 'learning' in the text. Whereas the words 'inefficient' and 'distributed' are further away, so they get a lower score. If you take this simple method—go through all the documents on the internet and for each word, see which words appear very close to it—those are probably related, so they get a high percentage. Words that are far away in all documents get a low percentage. So, just like with the other methods, you have a scalable statistical method to learn these statistics and these vectors. Now you have an intuition for not just what a word vector is, but also how you could automatically learn word vectors for every word on the internet, based on how close it is to other words in sentences.

向量空间与相似性 Vector Space and Similarity

Gustav

那么这些向量有什么用?它们如何帮助我们理解语言?

So what's the use of these vectors? How do they help us understand language?

Gustav

事实证明,如果你遍历比如整个维基百科并这样做,你会发现,在一个只有“王权”、“男性气质”和“女性气质”三个维度的简单三维世界里,“King”的向量在“王权”上几乎为 1,在“男性气质”上几乎为 1,在“女性气质”上几乎为 0。你可以把它想象成在这个空间中指向某个方向的向量,或者处于某个位置。相似的词——共享许多维度——会指向同一方向或彼此靠近。例如,“king”和“queen”很接近,因为它们在“王权”上都高,并且共享其他维度。如果你能直观理解词在向量空间中可能接近或远离,那么句子(词的组合)也可以通过将词向量相加来表示。因此你可以衡量句子之间的接近程度。例如,“Lion is the king of the jungle”会接近“Tiger hunts in the forest”,因为狮子和老虎是相似的动物,在“动物”维度上都高,而“丛林”和“森林”也相似。另一方面,“Everybody loves New York”则会远离。所以你有办法将语言转化为多个数字,从而看出哪些词或句子相似或接近。在现实中,你会有 60 万个维度(词典中每个词一个),但概念与三维相同。

Well, it turns out that if you go through, for example, all of Wikipedia and do this, you'll find that in a simple three-dimensional world with dimensions like royalty, masculinity, and femininity, the vector for 'King' is almost one on royalty, almost one on masculinity, and almost zero on femininity. You can think of it as a vector pointing in a certain direction in this space, or as being in a certain place. Words that are similar—sharing many dimensions—will point in the same direction or be close to each other. For instance, 'king' and 'queen' are close because both are high on royalty and share other dimensions. If you can intuit that words can be close or far in this vector space, then it's intuitive that sentences, which are combinations of words, can be represented by summing the vectors of their words. So you can measure how close sentences are to each other. For example, 'Lion is the king of the jungle' would be close to 'Tiger hunts in the forest' because lion and tiger are similar animals, both high on the animal dimension, and jungle and forest are similar. On the other hand, 'Everybody loves New York' would be far away. So you have a way to turn language into multiple numbers, allowing you to see which words or sentences are similar or close. In reality, you'd have 600,000 dimensions (one for each word in the dictionary), but the concept is the same as three dimensions.

歌曲与推荐系统示例 Example with Songs and Recommendation Systems

Gustav

你能举个具体例子让这更清楚吗?

Can you give a concrete example to make this clearer?

Gustav

当然。我们取另一个只有三个维度的世界:摇滚、古典和电子舞曲(EDM)。现在我们拿歌曲:披头士的《Here Comes the Sun》、贝多芬的《致爱丽丝》、Avicii 的《Levels》和皇后乐队的《波西米亚狂想曲》。我们为每个维度分配百分比。《Here Comes the Sun》几乎是纯摇滚,所以可能 98% 摇滚、0.02% 古典、0.01% EDM。《致爱丽丝》几乎没有摇滚(0.01%),大量古典(99%),几乎无 EDM(0.05%)。《Levels》摇滚少(0.02%),古典少(0.01%),但 EDM 多(99.9%)。《波西米亚狂想曲》很独特:它有很多摇滚(99%)也有很多古典(99%),但 EDM 不多(0.05%)。现在,如果你想象像这样给 Spotify 上所有歌曲打分,你就能看出哪些歌曲在向量空间中接近。听《致爱丽丝》的用户可能对《波西米亚狂想曲》感兴趣,因为它们在古典维度上都高,但不会对《Levels》感兴趣,因为它没有任何共同维度。这实际上就是 Spotify 等推荐系统的工作原理。

Sure. Let's take another world with three dimensions: how much something is rock, classical, and EDM. Now we take songs: 'Here Comes the Sun' by The Beatles, 'Für Elise' by Beethoven, 'Levels' by Avicii, and 'Bohemian Rhapsody' by Queen. We assign percentages for each dimension. 'Here Comes the Sun' is almost pure rock, so maybe 98% rock, 0.02% classical, and 0.01% EDM. 'Für Elise' has almost no rock (0.01%), a lot of classical (99%), and almost no EDM (0.05%). 'Levels' has little rock (0.02%), little classical (0.01%), but a lot of EDM (99.9%). 'Bohemian Rhapsody' is unique: it has a lot of rock (99%) and also a lot of classical (99%), but not much EDM (0.05%). Now, if you imagine scoring all songs on Spotify like this, you can see which songs are close in vector space. A user who listens to 'Für Elise' might be interested in 'Bohemian Rhapsody' because they both score high on classical, but not in 'Levels' because it shares no dimensions. This is actually how recommendation systems like Spotify work.

向量表示与推荐 Vector Representations and Recommendations

Gustav

互联网上的所有文档,你会如何对歌曲做类似的处理?

All the documents on the internet, how would you go about doing this for songs?

Gustav

你几乎希望有大量文档,其中歌曲彼此接近。播放列表就是这样的。Spotify 有几十亿个这样的播放列表。如果你把播放列表看作一个句子,你可以直接取中间的歌曲,看它和列表中其他歌曲有多接近。或者你把所有播放列表看作一个大文档,如果歌曲在同一个播放列表里,它们很可能彼此接近。所以同一播放列表中的歌曲会相互获得高分。这样你就能看到如何构建一个向量,了解 Spotify 上每首歌在多大程度上与其他歌曲相关,以百分比表示。现在你有了 Spotify 上每首歌的向量表示。因为它是向量,存在于这个多维空间中,你就可以做推荐了。

You almost wish that there would be a lot of documents where songs are close to each other. That's what a playlist is. So Spotify has a few billion of these playlists. And if you think of this playlist as a sentence, you can literally take the song in the middle and say how close is this song to all the other songs in this playlist. Or if you think about all the playlists as one big document, if they're in the same playlist, they are probably close to each other. So songs that are in the same playlist would score high relative to each other. So you can see how you could build a vector where you kind of understand how much of every song on Spotify is in every other song on Spotify as a percentage. And now you have a vector representation of every song on Spotify. And because it's a vector and it lives in this multi-dimensional world, you can do recommendations.

Gustav

所以 Spotify 上的品味档案,如果简化一点,其实就是你听过的所有歌曲以及这些维度的总和。你会得到一个分数,表示你听的所有歌曲中古典音乐占多少。你把这些加起来,除以歌曲数量,就得到该用户的古典乐分数。对摇滚、电子舞曲、爵士乐也做同样处理。现在你就有了该用户的品味档案。你就能理解该用户在向量空间中的位置。你可以说这两个用户向量几乎相同,他们彼此接近,音乐品味相同。

So a taste profile on Spotify is actually, if you simplify a little bit, just all the songs that you listen to and all those dimensions added together. So you get a score for how much classical is there in all the songs you listen to. You sum that up, you divide by the number of songs, and then you have a classical score for that user. You do the same for rock, the same for EDM, the same for jazz. And now you have a taste profile for that user. And now you understand where that user is in the vector space. And you can say that these two users, they have almost the same vectors, they're close to each other, they have the same music taste.

Gustav

所以现在你不仅理解了什么是向量、向量空间嵌入、编码、分布式表示——这些都是一回事,就是我刚才展示的,只是用不同名字让它看起来更难——而且你还理解了它为什么有用。你实际上还顺带理解了推荐系统的工作原理。

So now not only do you understand what vectors are, and vector space embeddings, codes, distributed representations—it's all the same thing, it's all what I just showed, just different names to make it seem harder than it is—but you also understand why it's useful. And you actually happen to accidentally understand how a recommendation system works.

神经网络与图像生成 Neural Networks and Image Generation

Gustav

好的,现在你 hopefully 理解了什么是大型语言模型、GPT 如何工作、什么是词向量以及它为什么有用。但我也承诺解释如何从文本或噪声生成图像,甚至从噪声生成音乐。为了做到这一点,我们需要更深入一步。我将以简化的方式解释神经网络到底是什么。再说一次,实践中非常复杂,理论上并不那么复杂。

All right, so now you hopefully understand what a large language model is, how GPT works, what a word vector is, and why it's useful. But I also promised to explain to you how it is that you can make images from text or images from noise, and even music from noise. So in order to do that, we need to go one step deeper. I'm going to explain in a simplified way what a neural network actually is. And again, in practice very complicated, in theory not that complicated.

Gustav

神经网络大致基于我们大脑中的生物神经元,卡通画里看起来像这样。黄色部分上有这些小分支,叫做树突。这些是输入。假设它们接收来自视网膜的电信号。光线进入眼睛,变成电信号,进入这些树突。这些信号会合并。假设这个特定神经元在寻找你眼前的垂直线或水平线。当它接收到这些电信号模式时,这些树突会得到一个值。它会达到某个阈值,说:嘿,我觉得我看到了一条垂直线。然后它会沿着轴突向右发送一个脉冲,传到下一层神经元。神经元做的就是这些:接收几个输入信号,合并它们,如果达到某个阈值,它就会说:嘿,我看到了什么,然后发送信号或脉冲。

So the neural network is loosely based on the biological neuron that we have in our brain, and it looks something like this in a cartoon. So you have these little arms on the yellow part called dendrites. Those are the inputs. Let's say that they get electrical signals from your retina. So light hits your eyes, they're electrical signals, they go into these dendrites in the yellow part. These things combine. So let's say that this particular neuron is looking for vertical lines or horizontal lines in front of your eyes, right? And so maybe this particular neuron, when it gets a pattern of these electrical signals, these dendrites, they get a value. It's going to hit some threshold that says, hey, I think I'm seeing a vertical line here. And then it's going to send a spike along this axon to the right that goes to the next layer of neurons. This is all that a neuron does: it takes a few input signals, it combines them, and if it hits a certain threshold value, it's going to say, hey, I'm seeing something here, and it's going to send a signal or a spike.

Gustav

计算机科学家做的是对生物神经元进行非常简化、理想化的数学版本,称为人工神经元。这里的箭头相当于树突。你有 A1、A2、A3。这些可能来自眼睛的电信号。在计算机世界里,它们可能是相机的像素值。它们进来后,会乘以这些叫做 W1、W2、W3 的权重。我稍后会讲。然后,你取电信号,乘以权重。如果这个细胞体,中间那个圆圈,达到某个阈值——比如细胞说:我觉得我看到了一条垂直线——它就会向右发送一个脉冲或信号,叫做 Z。所以这是生物神经元所做事情的非常简化的版本。

So what computer scientists did was they did a very simplified, idealized mathematical version of the biological neuron, called the artificial neuron. So the arrows pointing in here are the equivalent of the dendrites. So you have A1, A2, A3. These would be the electrical signals from the eyes. In a computer world, they would be the pixel values from a camera. They come in, they get multiplied by these things called W1, W2, W3, which are weights. I'll talk about that later. And then, as you take the electrical signal, you multiply them by the weights. If this cell body, the circle in the middle, hits a certain threshold—for example, the cell says, I think I see a vertical line—it is going to send a spike or a signal to the right called Z. So it's a very simplified version of what the biological neuron does.

猫分类器示例 Cat Classifier Example

Gustav

好的,如果你没有完全理解,别担心,你会在实践中看到。现在假设我们有一张猫的图片。为什么不呢?互联网上到处都是猫。假设你拿相机或摄像机对准这只猫。计算机看到的是像素值。记住,计算机里一切都是数字。像素值只是数字。也许我们简化一下,假设是灰度图,0 代表黑色,255 代表白色,中间是灰色阴影。所以只是数字。现在你放一堆这样的人工细胞,每个输入得到一个像素值。你可以想象,也许最上面的那个神经元,如果它看到像这样的对角线,它就会触发,说:嘿,我看到一条对角线。然后它下面的神经元,也许它在寻找另一个方向的对角线,只有看到那个它才会触发并向网络其余部分发送值。

Okay, if you didn't fully get that, don't worry, you'll see it in practice. Now let's say that we have a picture of a cat. Why not? The internet is full of cats. Now let's say that you take a camera and you take a photo or a video camera and you point it towards this cat. What the computer is going to see are pixel values. Remember, the computer, everything is numbers. The pixel values are just numbers. Maybe we simplify it and say it's a grayscale, so number zero is black and the number 255 is white, and in between there's shades of gray. So just numbers. So now you put up a bunch of these artificial cells, and each of these inputs gets a pixel value. And now you can imagine that maybe the top neuron there is going to spike if it sees maybe a diagonal line like this, right? And it says, hey, I see a diagonal line. And then the neuron below it, maybe it's looking for a diagonal line in the other direction like this, and it's only going to spike and send a value to the rest of the network if it sees that.

Gustav

现在你添加第二层网络。这就是为什么它们被称为深度神经网络。你添加越来越多的层。现在第二层神经元可以从第一层获得脉冲,并说:第一层看到了这样的对角线和那样的对角线。第二层神经元只有在第一层同时看到这两者时才会触发。所以也许这实际上是猫耳朵尖的形状,也许这条线是胡须的形状。你再加一层,它合并所有这些信号。在这一层的最后,可能非常深,最后一个神经元会说:我看到了所有这些层的输入,或者实际上是一只猫。现在你就有了所谓的猫分类器。

And now you add a second layer of network. This is why they're called deep neural networks. You add more and more layers. And so now the second layer of neurons, it can get a spike from the first layer and says, the first layer saw a diagonal line like this and a diagonal line like this. And the second layer neuron is going to spike only when it sees both of those in the first layer. So maybe this is actually the shape of the tip of a cat's ear, and maybe this line is the shape of a whisker. And so you go one more layer, and it combines all of these signals. And at the very end of this layer, which can be very deep, the last neuron would actually say, I see all the inputs from all of these layers, or what is actually a cat. And now you have what is called the cat classifier.

Gustav

那么这是如何工作的呢?如果你看这个,你几乎可以直觉地认为,如果你有所有这些权重,W1、W2、W3,如果你有无限的时间,你可以想象坐在那里调整所有这些数字,使这些神经元恰好达到阈值,并针对猫的形状从许多不同方向触发,对吧?所以有些神经元在寻找耳朵尖,有些在寻找眼睛、鼻子等等。你可以想象,如果你有无限的时间,你可以调整网络中的所有参数,使它只在有猫的时候一路触发到最后,而在有汽车、飞机或其他东西时不触发。所以,原则上并不难,实践中相当复杂。你会如何调整这些参数?因为这样的网络中可能有数十亿,现在甚至接近一万亿个参数。但我们还在……

So how does this work? Well, if you look at this, you can almost intuit that if you have all of these weights, the W1s and W2s and W3s, if you had infinite time, you could imagine sitting and tweaking all of those numbers so that these neurons happen to hit that threshold and spike exactly for the shape of a cat from many, many different directions, right? So some of these neurons are looking for tips of ears, some are looking for eyes and noses and so forth. You could imagine that if you had infinite time, you could tweak all of these parameters in the network so that it only spikes all the way back to the end when there is a cat, but not when there is a car or an airplane or anything else. So again, in principle not that hard, in practice pretty complicated. How would you tweak these parameters? Because there can literally be many billions, almost up towards a trillion now, of parameters in a network like this. But we're on the...

反向传播与神经网络训练 Backpropagation and Neural Network Training

Gustav

所以在理论层面,事情很简单。这里有一点你应该理解。

So at the theory level, things are simple. There is one thing you should understand here.

Gustav

科学家很久以前就提出了一个叫反向传播的东西。它的意思是,他们找到了一种方法,让模型自己学会所有正确的参数。你可以这样想:你有很多猫的图片和不是猫的图片。与其让人类坐下来自己调整这些参数,不如把所有这些参数初始化为随机数,完全是随机的。所以你想,第一次给这个网络看一张猫的图片时,它肯定会完全搞错,因为按定义是随机的,对吧?但这也意味着,碰巧有时候你给它看猫,它真的会猜是猫。然后你通过反向传播这个机制说,嘿,等一下,你碰巧猜对了,保留这些值,实际上把所有 W 值往那个方向移一点,因为你对了。当你猜错的时候,你就反过来,往另一个方向移。然后你对着成千上万、几十万、上百万张图片反复这么做,你问这是猫吗?它说是,你就说好网络,再移一点,把这些值都往那个方向再移一点。然后你给它看一架飞机,它说是猫,你就说坏网络,把所有值往另一个方向移一点。你不断强化它。如果你这样做几百万次,最终所有这些参数,因为你强化了好的行为,会通过多层结构找到图像中代表猫而不是狗或其他东西的精确形状组合。这就是神经网络的工作原理。再说一次,理论上很简单,尽管在实践中很难,花了很长时间才实现。

Scientists actually a long time ago came up with something called back propagation. What that means is they found a way for the model to teach itself all the right parameters. The way to think about this is you have a lot of pictures of cats and things that are not cats. Instead of a human sitting and tweaking these parameters themselves, you initialize all these parameters as random numbers. It's completely random. So if you think about it, the first time you show this network a cat image, it's going to get it completely wrong, by definition random, right? But it also means that by random chance, sometimes when you show the cat, it will actually guess that it was a cat. Then what you do is through something called back propagation, you say, hey wait a minute, you happened to guess correctly. Keep those values. In fact, move all the W values a little bit in that direction because you were right. And when you guess it's wrong, you do the opposite: you move them in the other direction. Then you do this for literally tens of thousands, hundreds of thousands, millions of images, where you say, is this a cat? And it says yes. Then you say, good network, move a little bit, move all these values a little bit more in that direction. Then you show it maybe an airplane, and it says it's a cat. Then you say, bad network, move all the values a little bit in the other direction. You just keep reinforcing it. If you do this millions of times, eventually all of these parameters, because you reinforce the good behavior, are going to end up finding exactly the combination of shapes in this image through multiple layers that represent the cat but not a dog and not anything else. So this is what a neural network does. Again, it's quite simple in theory, even though it was hard and took a long time to do in practice.

智能即压缩 Intelligence as Compression

Gustav

所以理解这一点很重要,因为当你理解了这种取数字、相乘、看是否超过阈值、再取那些数字相乘的概念,你就能理解这些图像生成网络等是如何工作的了。还有一个我觉得很有趣的概念:智能就是压缩。这被当作事实陈述,但并非已证实的事实,而是很多人持有的一种理论,即思考智能的一种方式是压缩。我这么说是什么意思?

So this is important to understand, because when you understand this notion of taking numbers, multiplying them, seeing if they go over a threshold, then taking those numbers, multiplying them, now you can understand how these image generation networks and so forth actually work. Here's another concept that I think is very interesting to understand: intelligence is compression. Now this is stated as a fact; it's not a proven fact, but it is a theory that a lot of people have, that one way to think about intelligence is as compression. What do I mean with that?

Gustav

回到直觉。如果你和一个非常懂某件事的人交谈,他们通常很擅长解释,对吧?而如果你和一个不太懂的人交谈,他们很难解释清楚。所以懂的人能用简单的方式解释。这通常意味着,如果他们能简单解释,就比不能简单解释的人理解得更好。从这里你已经能看到压缩的影子了,对吧?可能那个人在这个问题上花了大量时间,学会了把所有细节压缩成真正重要的东西,并深入理解。然后突然他们就能解释了。实际上互联网上有个叫 Hutter 奖的东西,那是一个竞赛,要求你把整个维基百科尽可能压缩,包括它自己的解压器,实际上,同时不丢失维基百科中的信息。因为想法是,为了有效压缩维基百科,压缩它的系统必须对世界有很多理解。要压缩你必须聪明;你必须很好地理解世界的维度才能压缩世界。就像我说的,这是直觉:能简单解释事物的人通常理解得更多。所以压缩它的系统必须发展出理解。把智能看作压缩信息能力的一个副作用、一个必要的恶。在这个 Hutter 奖中,你每压缩维基百科一个百分点,实际上就能赚钱,而且是无损压缩,不丢失任何信息。但总的来说,概念是如果你能压缩某物并保留大部分价值,你可能就真正理解了它,因为你用更少的信息表示了同样的东西。而用更少的信息表示某物,需要压缩它的系统有理解力。

Well, back to intuition. If you speak to someone who knows something very well, they're usually very good at explaining it, right? Whereas if you speak to someone who doesn't know something very well, it's hard for them to explain it. So the person that knows something well can explain something in a simple way. That usually means that if they can explain it in a simple way, they understand it better than if they can't. Already there you can see that there's something around compression, right? Probably that person spent a lot of time on this problem and they learned to take all the details and compress it into what it actually means and understand it deeply. Then all of a sudden they're able to explain it. So there is actually even this thing on the internet called the Hutter Prize, which is a competition where you're supposed to take all of Wikipedia and try to compress it as much as possible, including its own extractor, actually, without losing the information in Wikipedia. Because the idea is that in order to compress Wikipedia effectively, the system that compresses it is going to have to understand a lot about the world. You have to be smart in order to compress; you have to understand the dimensions of the world really well to be able to compress the world. And like I said, this is intuitive: humans that can explain things simply usually understand more. So the system that compresses it had to develop understanding. To think of intelligence as a side effect, a necessary evil of being able to compress information. So in this Hutter Prize, you can actually make money for every percentage that you can compress Wikipedia in that case losslessly, without losing any information. But in general, the concept is if you can compress something and retain most of the value, you probably understood it really well, because you could represent the same thing with less information. And representing something with less information kind of requires understanding on the part of the system that compresses it.

实际示例:自编码器 Practical Example: Autoencoder

Gustav

好,让我们更实际一点。记住,对神经网络来说,一切都是数字。语言是数字,像素是数字,样本是数字,DNA 序列是数字。任何东西都是数字。如果我们拿这个看起来有点奇怪的神经网络,假设我们拿一个句子,比如“一只猫从窗户跳出去”,这会被表示为一个数字序列,所以是 1、2、3、4、5、6 六个数字。然后我们让这个神经网络接收这六个数字,通过神经元相乘、相加,用更少的数字表示,然后再相乘、相加,在中间它只用三个数字来表示这六个数字。我知道这看起来奇怪,但等等。所以我们强迫网络用六个数字表示完整句子,却只用三个数字表示。现在你让网络做相反的事:从这三个数字开始,再次相乘,试图把它变回 5、6 或 7 个数字,在这个例子中是 1、2、3、4、5 五个数字。所以我们所做的就是告诉这个网络:这里有一个数字序列,对我们来说意思是“一只猫从窗户跳出去”。你必须把这六个数字压缩成三个数字,然后不添加任何新信息,仅从这三个数字扩展回同一个句子,或者尽可能接近。所以整个训练任务就是:取一个句子,压缩它,然后尝试重建同一个句子。在完美世界里,输出会是“一只猫从窗户跳出去”,和输入完全一样。所以你反复训练这个网络,尝试压缩这些数字并重建相同的数字。网络不可能完美重建相同的数字,因为当你从六个数字变成三个数字时,你会丢失信息。所以按定义你丢失了信息,从六个到三个至少丢失了那三个数字。所以它会尽力而为。这意味着它必须选择中间这三个数字,回到维度。它必须选择最能描述世界的三个维度、三个数字。我这么说是什么意思?意思是,如果你看“一只猫从窗户跳出去”这样的东西,也许中间这三个数字代表的东西,不完全是猫,那太详细了,但也许是宠物,还有……

All right, let's get a little bit more practical there. So remember, to a neural network, everything is just numbers. Language is numbers, pixels are numbers, samples are numbers, DNA sequences are numbers. Anything is a number. So if we take this neural network that looks a bit funky, let's say that we take a sentence like "a cat jumping out of a window", which then again will be represented as a sequence of numbers, so one, two, three, four, five, six numbers. Then what we do is this neural network just takes those six numbers and through these neurons it multiplies and adds and has it represented by fewer numbers, then multiplies and adds again, and in the middle it only gets three numbers to represent those six numbers. I know it seems weird, but hang on. So we force the network to take six numbers that represent the full sentence and represent it with only three numbers. Now you ask the network to do the opposite: from these three numbers again it multiplies and tries to turn it back into five, six, or seven numbers, in this case one, two, three, four, five numbers. So all we did was we told this network that here's a sequence of numbers that to us means "a cat jumping out of the window". You have to compress those six numbers to three numbers and then expand without any new information, just from those three numbers, back into the same sentence or as close as possible that you can get. So the entire training task here is to take a sentence, compress it, and try to recreate the same sentence. So in a perfect world, the output would be "a cat jumping out of a window", the exact same as the input. So you train this network again and again to try to compress these numbers and recreate the same numbers. And it's not going to be possible for the network to recreate the same numbers perfectly, because when you go from six to three numbers, you will lose information. So by definition you lost information, you lost at least those three numbers from six to three. So it's going to do the best it can. And that means that it's going to have to pick these three numbers in the middle, back to dimensions. It's going to have to pick the three dimensions, the three numbers that sort of best describes the world. What do I mean with that? Well, it means that if you look at something like "a cat jumping out of a window", maybe these three numbers in the middle, they represent something like not quite a cat, that's too detailed, but maybe a pet and...

嵌入与压缩 Embedding and Compression

Gustav

也许第二个数字代表某种进出某物的状态,第三个数字代表某个实体或房屋。所以输出结果在概念上会与你输入的句子相似,但又不完全相同,因为信息丢失了。比如你输入“一只猫从窗户跳出去”,输出的可能是“一只宠物离开房子”或“一只狗离开房子”,对吧?因为系统必须做出选择,它必须丢失信息,必须抽象化,必须挑选最重要的维度来尽可能做好,它必须压缩。所以这意味着,中间这个嵌入编码,也就是向量,如果你在大量句子上正确训练,它就会找到最能代表所训练世界的数字,在这里就是互联网上的文本,对吧?对这个网络来说,这是一个文本世界。这叫做嵌入。你输入一个句子,把它从六个数字压缩成三个,如果训练正确,网络会选择正确的维度,这些维度能提供关于世界的最多信息,从而最好地完成训练任务。所以,这看起来是个很没用的任务,尤其是对语言来说。为什么你想得到同一个句子的不同版本呢?但也许现在有个更容易理解的例子。这对理解扩散模型很重要,所以我们才讲这些。

Maybe the second number represents something like going in and out of something, and the third number represents something like an entity or a house. So what you will get on the output is something that is similar conceptually to the sentence you put in, but not quite the same, because information was lost. So if you put in a cat jumping out of a window, maybe what you get out is a pet leaving the house or a dog leaving the house, right? Because the system had to pick, it had to lose information, it had to abstract it, it had to pick the most important dimensions to do as well as it could, it had to compress. So what this means is that this thing in the middle, the embedding code, again the vector, hopefully if you do this right over a lot of sentences, is going to find the numbers that are the best representation of the world that is being trained on, which in this case would be the text of the internet, right? It's a textual world for this network. This is called embedding. So you take the sentence and you embed it from six numbers into three, and the network, if you're training correctly, is going to choose the right dimensions that give the most information about the world that completes the training task the best. So again, it seems like a pretty useless task, especially for language. Why would you want to get a different version of the same sentence out? But maybe an example that would be easier to understand right now. This is important to understand for the diffusion models, that's why we're going through it.

Gustav

如果你转而考虑图像或视频,你可以想象左边这里有一张图像或一段视频,占用大量空间。对于图像,你知道可以压缩它,即使丢失很多信息,它仍然足够好,对吧?所有图像格式都是这么做的。它们压缩图像,这样你就能用比原始图像少得多的数据在互联网上传输。所以创建这种压缩的一种方法是,你拿一张图像,比如左边的一只猫,你强迫网络接收所有这些数字,所有像素数字,在中间用更少的数字来表示它们,然后尽力重建同一只猫的图像。如果你做得好,它就能很好地重建出几乎正确的图像。如果你教会网络尽可能做好,它就会找到中间需要保留的最重要维度,以便人类认为这仍然是同一只猫的图像。所以网络会自己学会压缩到最重要的维度。比如,人类往往更关注图像中的低频信息,而不太关注高频信息,所以它可能会学会保留一些低频信息,而不是高频信息,等等。但细节并不重要。整个概念是,它接收一堆数字,在这种情况下是图像像素数字,你强迫网络尝试用更少的数字选择最佳表示,同时尽可能保留原始图像中的信息,然后重建它。如果你这样做,你就有了一个压缩算法,你可以拿一张图像,但不用在互联网上发送图像,只需嵌入它,发送中间这段代码,接收方拥有网络的另一端,然后解码回完整图像。这样你就节省了大量带宽。这叫做自编码器,实际上互联网上一些压缩算法就是这么做的。所以现在你不仅理解了什么是向量,什么是词向量,还理解了如何实际创建这个向量,以及什么是嵌入——一个通过训练接收文本、图像或其他内容,自动为你创建这个向量的网络。而且重要的是,它会自动选择最佳维度。你不是用世界的六十万个维度,而是强迫它选择少得多的维度。通过强迫它选择更少,你迫使网络真正变得智能,至少根据某些智能的定义,在它如何选择这些维度方面。

If you instead think about images or maybe video, you can imagine on the left side here that you have an image or a video that takes a lot of space. And when it comes to images, you know that you can compress it and actually lose a lot of information and it's still good enough, right? This is what all image formats do. They compress the image so that you can send it over the internet using much less data than the actual original image. So one way to create that compression would be that you take an image, for example of a cat on the left, you force the network to take all those numbers, all those pixel numbers, and represent them there in the middle with much fewer numbers, and then just try to recreate the same image of the same cat as well as it can. And if you do this right, it's going to be very good at recreating almost the right image. And if you teach the network to do it as well as possible, it's going to find the dimensions that are most important to keep there in the middle in order for humans to think that this is still the same image of a cat. So the network is going to learn itself to compress to the most important dimensions. For example, humans tend to care a lot about low-frequency things in images and not as much about high frequencies, so it will probably learn to keep some of the low frequencies and not the high frequencies, etc. But it's not important with the details. The whole concept is it takes a bunch of numbers, in that case the image pixel numbers, you force the network to try to pick the best representation with much fewer numbers that still keeps as much as possible about the information in that original image, and then recreate it. So if you do that, now you actually have a compression algorithm where you can take an image, but instead of sending an image over the internet, you just embed it and you just send this code in the middle, and the receiver has the other side of this network and then decodes it back into the full image. So now you saved a lot of bandwidth. This is called an autoencoder, and it is actually how some compression algorithms are done on the internet. So now you not only understand what a vector is, what a word vector is, but you also understand how you would actually create this vector and what an embedding is, which is a network that through training takes a piece of text, an image, or something, and automatically creates this vector for you. And it automatically, importantly, chooses the best dimensions. Instead of having six hundred thousand dimensions of the world, you force it to choose much fewer. And in forcing it to choose fewer, you force the network to actually become intelligent, at least according to some definitions of intelligent, in how it picks those dimensions.

扩散模型 Diffusion Models

Gustav

好了,我们快讲完了。那么图像、音乐生成、视频等等呢?既然你已经理解了所有这些,要理解像 Stable Diffusion 或 Midjourney 这样的东西是如何工作的,你只需要理解最后一点,那就是扩散模型的概念。扩散模型,我认为在概念上也非常直观。还记得这个神经网络吗?假设你拿一张图像,左边这张叫 t0,你做的就是加一点噪声。从左边第一张图到第二张图,你在这张图上加了一点噪声,然后训练一个神经网络,让它简单地找到并去除这些噪声。如果你看这些图像,第一张和第二张之间的差异非常小。我想你凭直觉就能明白,只要花时间,你就能去除这些噪声。那神经网络为什么不能呢?这看起来并不难,因为你只是加了一点点。所以这是第一步。你可以想象有一个单独的神经网络,它只是把第二张图上的噪声去除,回到第一张图。现在你拿第二张图,它有一点噪声,你再加一点噪声,然后训练一个网络,让它只去除这额外的噪声,也就是从第三张图回到第二张图,而不是一路回到第一张图。它只是一步一步来。每一步都只是去除一点噪声。然后你拿第三张图,再加一点噪声,训练一个网络去除那部分噪声,如此继续。当我说“去除”时,你训练网络识别出这是添加的噪声,然后从图像中减去它,以重建原始图像。希望你能直观地理解这一点。所以你有一个网络,可以在任何阶段接收图像并去除那额外的少量噪声,因为仔细想想,每一步都是一个看似简单的小任务。但到最后,你添加了太多噪声,图像中只剩下纯噪声,不再是图像了。所以在最右边,任务实际上是从完全噪声中去除最后添加的噪声,变成倒数第二张几乎完全噪声的图像。

All right, we're almost there. So what about images, music generation, video, etc.? There's only one last thing you need to understand now that you understand all of this, in order to understand how something like Stable Diffusion works or Midjourney or something like this, and that is this notion of diffusion models. The diffusion models, it's actually also something that I think conceptually is very intuitive. So remember this neural network, right? Let's say that you take an image, the one to the left here called t0, and what you do is you just add a little bit of noise. So you go from the image on the left to the second image on the left. You add a little bit of noise on top of this image, and now you train a neural network to simply try to find that noise and remove it again. And if you look at those images, the difference between the first and the second image is very small. It is intuitive, I think to you, that you could remove that noise if you just had time. And so why couldn't a neural network? It doesn't seem that hard because you just add a little bit. So that's the first step. So think of it almost as you have a separate neural network that just removes a little bit of noise from the second back to the first image. Now you take the second image with a little bit of noise and you add a little bit more noise, and now you train a network to just remove that additional noise, like the image that you added in the third, the noise that you added to the third image back to the second, not all the way back to the first. It's just one step at a time. So every step here is just removing a little bit of noise. And now you take that third image, you add a little bit more noise, you train a network to remove just that noise, and you keep going. You take that image, you add more noise, you train the network to remove that additional noise, etc. And when I say remove, you train a network to identify this was the noise added and simply deduct it from the picture to recreate the original image. So hopefully it's intuitive that you could do this. So you have like a network that can take an image at any stage and remove that additional little noise, because if you think about it, at every stage it was a deceptively simple little task. But at the end, you've added so much noise that there is pure noise in the image, there is no image anymore. So to the right, the task is actually to remove the last piece of added noise from complete noise to almost complete noise in the second-to-last image.

Gustav

好了,这看起来又是个愚蠢的任务。为什么你要拿一张完好的图像,慢慢破坏它,然后训练网络一次去除一点噪声呢?嗯,真正酷的是,现在你有了这个……

All right, so again this seems like a stupid task. Why would you take a perfectly good image and destroy it slowly and train a network to remove a little bit of noise at a time? Well, the really cool thing is now you have this...

扩散模型如何从噪声生成人脸 How diffusion models generate faces from noise

Gustav

你能解释一下扩散模型是如何从纯噪声生成图像的吗?

Could you explain how diffusion models can generate images from pure noise?

Gustav

当然。一个从好图像开始的网络学会去除一点点噪声,使其稍微变差,然后再移动一点点。但如果你把这个网络倒过来运行——不是用最左边的那个从好图像中去除少量噪声的神经网络,而是从最右边的那个从纯噪声图像中去除少量噪声的神经网络开始——你可以从随机噪声开始,然后反向运行网络。本质上会发生的是:你训练了这个网络,让它拼命寻找图像中像人脸一样的噪声。所以它拿到这张完全噪声的图像,即使那里什么都没有,它也会说:‘嘿,我受过训练,要在这里找到像人脸一样的噪声。’或者简单说:‘我受过训练,要在这里找到人脸的粗略轮廓。我觉得我看到了一些东西。’然后它会去除一点噪声,这实际上让它看起来更像一张脸。

Sure. A network that starts with a good image learns to remove a little bit of noise, making it slightly worse, and then it moves a little bit again. But if you take this network and run it backwards—instead of taking the neural network that is furthest to the left, which removes a little bit of noise from a good image, you start with the one to the right that removes a little bit of noise from a pure noise image—what you can do is start with just random noise and then run the network backwards. What will happen is essentially this: you've trained the network to desperately look for face-like noise in an image. So it takes this complete noise image, and even though there's nothing there, it's going to say, 'Hey, I've been trained to find face-like noise in here,' or to simplify, 'I've been trained to find the rough outlines of a face in here. I think I see something there.' And it's going to remove a bit of noise that actually makes it look a little bit more like a face.

Gustav

然后在下一阶段,下一个网络——即使是同一个网络,但你可以把它看作独立的——拿到那张图像说:‘我也受过训练,去除噪声以找到这里像人脸一样的噪声并去除它。我觉得我在里面看到了脸的轮廓。’尽管在第一阶段那里什么都没有,只是噪声,但在第二阶段,那里实际上有一点脸了,因为网络本身去除了那些让它看起来更像脸的像素。

Then in the next stage, the next network—even though it's the same network, but think of it as separate—takes that image and says, 'I've also been trained to remove noise to find face-like noise here and remove it. I think I see the outlines of a face in there.' Even though in the first stage there was nothing there, just noise, in the second stage there is actually a little bit of a face there, because the network itself removed exactly the pixels that made it look more like a face.

Gustav

下一步说:‘嘿,等等,我看到了脸的轮廓。我要去除这个噪声,这会让它看起来更像一张脸。’第三张图像接手说:‘哦,我清楚地看到了脸的轮廓。我确切知道该去除什么样的噪声,才能让它看起来更像一张脸。’然后你继续下去,到最后,一直走到最右边,你就会从纯噪声中创造出一张脸——也就是说,一张实际上从未存在过的脸。

The next step says, 'Hey wait a minute, I see the outlines of a face. I'm going to remove this noise, which is going to make it look even more like a face.' The third image takes over and says, 'Oh, I clearly see the outlines of a face there. I know exactly what kind of noise I should remove to make this look even more like a face.' And you just keep going, and at the end, all the way to the right, you will have created a face out of pure noise—meaning a face that actually never existed.

Gustav

需要明确的是,这不是对已存在面孔的复制。你训练这个扩散模型在数百万张不同的面孔上去除噪声,所以它没有学会创造特定的面孔;它学会了创造一般的面孔,并在第一张随机噪声图像中寻找一般的类人脸噪声。这就是你如何得到像‘这个人不存在’这样的网站——每次你访问它,它都会生成非常逼真的人脸,而这些人在现实中从未存在过。

To be clear, it's not a copy of a face that existed. You train this diffusion model to remove noise across millions of different faces, so it didn't learn to create a specific face; it learned to create general faces and is looking for general face-like noise in this first random noise image. This is how you get to something like this site called 'This Person Does Not Exist'—every time you go there, it literally generates very believable faces of people that never existed.

Gustav

对于那些对此了解更多的人来说,这个特定的网站实际上并不使用扩散模型;它使用了一种叫做 GAN 的东西,即生成对抗网络,它出现得更早。但原理是一样的。扩散模型在某种程度上已经取代了 GAN。

Now, for those of you who know more about this, this particular site actually doesn't use a diffusion model; it uses something called a GAN, which came before—the generative adversarial network. But the idea is the same. Diffusion models have sort of taken over from GANs.

文本条件与嵌入 Text conditioning and embeddings

Gustav

所以现在我们知道了如何从纯白噪声中生成至少是人脸的东西。但你答应过要解释的不只是如何创造一样东西,而是如何做到像 Stable Diffusion 或 Midjourney 那样的事情。在这些服务上,你能做的不仅仅是得到同一事物的许多版本;你可以输入文字,要求它生成,例如,一张宇航员在月球上骑马的图片。那是怎么做到的?这种文本条件控制是如何工作的?

So now we know how to generate at least faces out of pure white noise. But you promised to explain not just how to create one thing, but how to do something like Stable Diffusion or Midjourney does. On these services, you can do more than just get many versions of the same thing; you can put in text and ask it to generate, for example, a picture of an astronaut riding a horse on the moon. How does that work? How does this text conditioning work?

Gustav

嗯,它是一个扩散模型,所以它做了我刚才展示的事情,但它还做了更多。这就是为什么你刚刚学习了向量。让我们回到‘智能即压缩’这个概念,以及你实际上如何对文本进行条件控制。我给你看过这个我们曾说看起来相当无用的网络:当你拿一个句子,你压缩那个句子中的数字——希望以某种方式捕捉该句子重要维度的更少数字——然后你尝试再次将其扩展为同一个句子。我们曾说它相当无用,但现在它将被证明相当有用。

Well, it is a diffusion model, so it does what I just showed you, but it does something more. This is why you just learned about vectors. Let's go back to this thing of intelligence is compression and how you actually condition on text. I showed you this network that we said looked pretty useless: when you take a sentence, you compress the numbers in that sentence—the fewer numbers that hopefully somehow capture the important dimensions of that sentence—and then you try to expand it to the same sentence again. We said it was pretty useless, but now it's going to turn out to be pretty useful.

Gustav

一旦你训练了这个网络,你可以做的是切掉右边部分,只保留这部分。所以现在你有一个机器,你可以给它一个英文句子,并要求它嵌入——将其压缩成尽可能好地代表该句子内容的数字。

What you can do once you train this network is you cut off the right part and you just keep this part. So now you have a machine that you can give a sentence in English and ask it to embed it—to compress it to these numbers that represent what is in that sentence as well as possible.

Gustav

现在让我们想象你是一个服务,比如社交网络或搜索引擎,拥有大量图像及其标题的示例。例如,可能是一张猫盯着你的图片,标题是‘一只猫盯着我’。你可以做的是:你可以拿那个标题‘一只猫盯着我’,然后拿这个编码器。你可以拿那个句子——它同样只是一串数字,比如一、二、三、四、五个数字——然后你可以要求这个网络将其压缩成这三个数字,希望它们能捕捉到‘一只猫盯着我’的含义。

Now let's imagine that you're a service, say a social network or search engine, that has a lot of examples of images and captions to those images. For example, maybe a picture of a cat staring at you and the caption 'a cat staring at me.' What you can do is the following: you can take that caption, 'a cat staring at me,' and you can take this encoder. You can take that sentence, which again is just a sequence of numbers—so one, two, three, four, five numbers—and you can ask this network to compress it into these three numbers that hopefully capture what it means to be a cat staring at me.

Gustav

现在我们拿扩散模型。我们拿属于这个标题的图像,像之前一样,我们添加一点噪声,然后更多噪声,然后更多噪声,直到它变成完全噪声。然后我们在中间放入这个神经网络,它将尝试去除噪声。这正是我们对人脸所做的,对吧?所以如果我们只是这样做,我们将构建一个总是找到盯着你的猫的扩散模型,这不是我们想要的。我们想要的是可引导的东西。

Now we take the diffusion model. We take the image that belongs to this caption, as before we add a bit of noise, then a bit more noise, then a bit more noise until it's complete noise. And we put in this neural network in between that is going to try to remove the noise. This is exactly what we did with the faces, right? So if we just did this, we're going to build a diffusion model that always finds cats staring at you, which is not what we wanted. We wanted something that is steerable.

Gustav

所以我们要再做一步。我们要拿我们有的另一个网络,它接受句子‘一只猫盯着我’,将其从这五个数字嵌入到这三个数字——代码,粉色的东西。当这个扩散模型试图在第二张和第一张图片之间去除噪声时,我们将给它这三个数字作为线索。记住,我们给它一张猫盯着你的图片,同时给它代表‘一只猫盯着你’的三个数字。

So we're going to do one more step. We're going to take this other network we had that takes the sentence 'a cat staring at me,' embeds it from these five numbers into these three numbers—the code, the pink thing. And as this diffusion model is trying to remove the noise between the second and the first picture, we're going to give it these three numbers as a clue. Remember, we're giving it the picture of a cat staring at you and we're giving it the three numbers that represent a cat staring at you.

Gustav

你现在可以把扩散模型看作有一个关于它在寻找什么类型噪声的线索。它不仅仅是在寻找一种类型的噪声;它会说:‘嘿,这三个粉色的数字,我以前见过。这意味着这里可能有一只猫,或者这里可能有类猫的噪声。’然后你这样做,在下一步你给它同样的线索——你仍然在寻找类猫的噪声。下一步你给它同样的线索。

You can think of the diffusion model now as having a clue about what kind of noise it's looking for. It's not just looking for one type of noise; it's going to say, 'Hey, these three pink numbers, I've seen them before. It means there's probably a cat in here, or there's probably cat-like noise here.' And then you do that, and at the next step you give it the same clue—you're still looking for cat-like noise. The next step you give it the same clue.

Gustav

记住,对神经网络来说,一切都是数字。像素是数字,但这些句子也是数字。它并不真正理解一个是图像,另一个是文本。它只是说:‘我见过这些数字。当我看到这些数字,上面还有这三个数字时,里面总是有类猫的数字,或者类猫噪声的数字。所以让我试着看看我能否在这里找到一只盯着我的猫。’

Remember, to a neural network everything is numbers. The pixels are numbers, but these sentences are also numbers. It doesn't really understand that one is images and the other is text. It's just saying, 'I've seen these numbers. When I saw these numbers with these three numbers on top, there was always cat-like numbers in there, or cat-noise-like numbers in there. So let me try to see if I can find a cat in here staring at me.'

Gustav

所以你这样做,例如,对于这张‘一只猫盯着你’的图片,然后你有其他类似的图片,但带有不同的标题,等等。

So you do this, for example, for this image 'a cat staring at you,' and then you have other similar images but with different captions, and so on.

扩散模型与文本条件 Diffusion models and text conditioning

Gustav

不同的描述文字,比如不是一只猫盯着你,而是另一张图是猫跳出窗户。现在顶部的这个编码器会把那句话嵌入,那句话有六个词,变成三个数字。这三个数字可能很相似——因为里面都有猫,对吧?但也有跳跃的概念,而没有凝视的概念,有窗户的概念等等。所以这个向量在维度上会相似,但略有不同。然后你对那张图做扩散。同样,网络在尝试去除噪声时,有这三个数字作为线索,这样它就能理解它在寻找什么样的噪声。

Different captions, so instead of a cat staring at you, maybe another image is a cat jumping out of a window. And now this encoder at the top is going to embed that sentence, which is one, two, three, four, five, six, seven, six words, six numbers, into three numbers. These three numbers would probably be similar—there's cats in there, right? But there's also the concept of jumping, and there's not the concept of staring, there's a concept of window, and so forth. So this vector will be similar in the dimensions but a little bit different. And now you do the diffusion on that image. So again, the network has this clue, these three numbers, as it's trying to remove the noise, so that it can understand what kind of noise it's looking for.

Gustav

所以记住,在第一个过程中,我们总是去除同一种噪声,类似人脸的噪声。但现在我们做的是扩散模型,它会得到线索,知道它在寻找什么样的噪声。所以我们这里非常专注于猫,但这可以是任何东西。可以是一张飞机的图片,文字写着“一架飞机在飞行”,然后它会学习飞机类噪声是什么样的,或者说去除飞机类噪声是什么样的。

So remember, in the first process we always removed the same kind of noise, face-like noise. But now we're doing the diffusion model that gets a clue for what kind of noise it's looking for. And so we're very focused on cats here, but this could be anything. It could be a picture of an airplane and the text saying 'an airplane flying', and then it's going to kind of learn what airplane-like noise looks like, or the removal of airplane-like noise.

Gustav

而你能做的是,实际上可以拿一首歌这样的东西,对吧?歌曲是音频波,但事实证明,你可以拿一段音频,把它转换成所谓的声谱图。声谱图是歌曲的可视化表示。它大致说明在某个时间点,歌曲中每个频率有多少,而幅度并不重要。你只要想象你可以拿一段音频转换成声谱图,就能把音频表示为图像。

And what you can do is you can actually take something like a song, right? A song is an audio wave, but it turns out you can take a piece of audio and you can transform it into what is called a spectrogram. So a spectrogram is the visual representation of a song. It kind of says how much of every frequency is in a song at a certain time, and the magnitude of that doesn't really matter. If you just imagine that you can take a piece of audio and transform it into a spectrogram, you can represent audio as an image.

Gustav

所以现在你能做的是,你可以拿一首歌,比如披头士的《Ob-La-Di, Ob-La-Da》。你可以拿那首歌的声谱图,然后拿文字描述,字面上就是“披头士的”或“一首披头士的歌”。你把它嵌入,所以现在代码会以某种方式表示歌曲的概念、披头士的概念,以及其他一些东西。我们不完全知道,它会找到这个句子的最佳表示。所以现在你把那句话——“披头士的《Ob-La-Di, Ob-La-Da》”——作为线索给这个扩散网络,而它正在尝试找到这个声谱图。所以你要教它,首先找到声谱图,但也要教它,如果你给它看许多不同音乐类型的声谱图,它就能学会特定类型的声谱图。

So now what you can do is you can take a song, for example 'Ob-La-Di, Ob-La-Da' by The Beatles. You can take the spectrogram of that song, and you can take the description in text, literally 'by The Beatles' or maybe 'a song by The Beatles'. You embed that, so now the code is going to somehow represent the concept of a song, the concept of The Beatles, and some other things. So we don't fully know, it's going to find the best representation of this sentence. And so now you're giving that sentence, 'the song 'Ob-La-Di, Ob-La-Da' by The Beatles', as the clue to this diffusion network as it is trying to find this spectrogram. And so you're going to teach it to, first of all, find spectrograms, but you're also going to teach it, if you show it many spectrograms with many different types of music, a certain type of spectrogram.

Gustav

那么现在会发生什么,一旦你在数百万种不同类型的图像上训练这个网络,如我所说,猫、狗、飞机、音乐图像,随便什么,你现在可以把这个网络翻转过来,从纯白噪声开始。那里没有任何结构,图片里字面上什么都没有。现在你拿一个从未存在的句子,比如说“一首披头士风格的艾维奇歌曲”。

So what happens now is, once you train this network on millions of different types of images, as I said, cats, dogs, airplanes, images of music, whatever you want, you can now take this network, you can turn it around, and you can start with just pure white noise. So there is no structure in there, there's literally nothing in this picture. And now you take a sentence that never existed, let's say 'an Avicii song in the style of The Beatles'.

Gustav

而且当你训练了这个网络,这个编码过程希望能捕捉到世界的维度。所以我们会知道艾维奇大致是什么、意味着什么,我们也大致知道歌曲是什么,它应该生成声谱图而不是猫的图片,以及披头士代表什么、那些声谱图长什么样。所以它会把这句话嵌入到这个代码中,在我们简化的例子里只是三个数字。实际上不是三个数字,而是更多数字,但为了简单起见。现在我们拿这个纯白噪声,把这个代码作为线索给扩散模型,告诉它它在寻找什么样的噪声。

And as you've trained this network, this encoding process hopefully will have captured the dimensions of the world. So we'll know whatever sort of what Avicii is and what it means, and we kind of know what a song is, that it should do a spectrogram now and not a picture of a cat, and sort of what The Beatles represents and what those kinds of spectrograms look like. And so it's going to embed this into this code that, in our simplified example, is just three numbers. It's not three numbers in reality, it's more numbers, but for purposes of simplicity. And now we take this pure white noise, we give this code as a clue to the diffuser model for what kind of noise it's looking for.

Gustav

而这里有趣的是,现在我们告诉它去寻找一种实际上从未存在过的噪声——“一首披头士风格的艾维奇歌曲”那种噪声。它会竭尽全力在这纯白噪声中寻找那种结构。所以它会去除噪声,让它看起来更像一首歌的声谱图,或一首披头士风格的歌。在下一步,我们给它同样的线索,说,再努力一点,真正去寻找披头士风格的艾维奇歌曲的结构,再去除一些噪声。我们一遍又一遍地做。

And the interesting thing again here is, now we're telling it to look for kind of noise that actually never existed before. 'An Avicii song in the style of The Beatles' kind of noise. And it's going to try its darndest to try to find that kind of structure in this pure white noise. So it's going to remove noise that makes it look a little bit more like a spectrogram of a song, or a song in the style of The Beatles. And at the next step, we give it the same clue and say, try harder, try to really look for the structure of an Avicii song in the style of The Beatles, remove some more noise. And we do it again, and we do it again.

Gustav

如果你感兴趣,实际上这些扩散模型大约有 50 个这样的步骤。在最后,第 50 步,你会得到一个从未存在的歌曲的声谱图,希望是一首披头士风格的艾维奇歌曲。当你把它从声谱图转换回纯音频,现在你终于到了。希望你能评判,告诉我你的想法。

And if you're interested, in reality these diffusion models have about 50 of these steps. And at the end of the line, at the 50th step, you're going to get a spectrogram of a song that never existed, that hopefully is an Avicii song in the style of The Beatles. When you transform it back from a spectrogram into pure audio, so now you're finally there. Hopefully you'll be the judge, tell me what you think.

Gustav

你对如何仅从文本甚至白噪声中创造出新小说、诗歌、图像甚至音乐有了一些直觉。希望你觉得我们成功揭穿了这个阴谋论一点。非常感谢你的关注。

You have some intuitions about how it is actually possible to create new novels, poems, images, even music out of just text or even white noise. And hopefully you feel like we managed to debunk this conspiracy a little bit. Thank you very much for paying attention.

互动版:逐字朗读 + 针对本期提问 →