From Trading to AGI: Mark Chen on Research Taste and Replication
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 首席研究官 Mark Chen 分享他从高频交易到 AI 研究的历程,强调复现对培养研究品味的重要性,以及 AlphaGo“第 37 手”带来的启发。
OpenAI's Chief Research Officer Mark Chen shares his journey from high-frequency trading to AI research, the importance of replication for developing research taste, and the inspiration from AlphaGo's 'move 37'.
嘿,各位,欢迎来到 Latent Space 烹饪系列节目,我们邀请创始人和研究人员,让他们自由发挥。今天,我们有一位非常特别的嘉宾,OpenAI 的首席研究官 Mark Chen。欢迎你,Mark。
Hey guys, welcome to the Latent Space cooking series where we invite founders and researchers and just let them cook. Today, we have a very special guest, the chief research officer of OpenAI, Mark Chen. Welcome, Mark.
谢谢你邀请我,Alan。
Thanks for inviting me, Alan.
嗯,谢谢你来。首先,这一切的灵感来源于一个故事:马克·扎克伯格会做汤来吸引研究人员,作为回应,你也给研究人员带了汤。这是真的吗?发生过吗?有效吗?
Yeah, thank you for coming. I mean, to begin, this all started from the inspiration after hearing the story that Mark Zuckerberg would make soup to try to poke researchers and in response, you brought soup to researchers. Is this true? Did this happen? Did it work?
哦,你知道,这绝对是个真实的故事。我确实给我们的研究人员带过汤。我觉得那件事后来平息了一点。我想我们最终占了上风,但没错,在 AI 参与的疯狂中,这仍然是个非常有趣的故事。
Oh, you know, it's absolutely a true story. And I have brought soup to our own researchers. I think that matters calmed down a little bit. I think we came out on top, but yeah, still a very funny story in the craziness of how AI is involved.
你多久做一次饭?你熟悉烹饪吗?
How often do you cook? Is it something you're familiar with?
嗯,你知道,我确实喜欢烹饪,但没有太多时间经常做。所以,我通常每晚都有工作晚餐,也许在 AGI 之后,这会成为我的爱好。我一直开玩笑说,等一切结束后,我要开个面摊。
Well, you know, I do enjoy cooking, but I don't have the luxury of doing that so often. So, I usually have a work dinner every night of the week and, you know, maybe post AGI, this is going to be my hobby. I've always joked that I'm going to start a noodle stand once it's all over.
是啊,是啊,你知道,AGI 之后,希望那还在,但很好。我想看看我们面前的东西,你大概知道我们在做什么吗?
Yeah, yeah, you know, post AGI, hopefully that'll still be there, but great. And I guess looking at what we have in front of us, do you have an idea generally of what we're probably making?
呃,韩式豆腐汤,也许?
Uh, Korean tofu soup, maybe?
是啊,是啊,差不多就是这个。所以,我们受你给研究人员带汤的故事启发,正在做韩式豆腐炖菜,还有虾,我来煮。你准备好了吗?
Yeah, yeah, that's generally what it is. So, we inspired off of the story of you bringing soup to researchers. So, we're making a tofu Korean stew and we have prawns that I'll be cooking. Are you ready to go?
好,开始吧。开始吧。
Yeah, let's do it. Let's do it.
太好了。好的,我们首先应该把蔬菜分开,然后切一下。基本上,我们要做的就是切掉带泥土的部分,然后分开。你可以做这个。同时,我想我可以多问问你的背景。你以前是个交易员。甚至 Sam,我想去年四月也发推说,如果你是高频交易员,应该考虑加入 OpenAI,因为要构建 AGI。那么,你认为交易员和研究员之间有关系吗?还是说这只是一个技术性强、竞争激烈的领域,很多优秀员工可能来自这里?
Great. Okay, so the first thing we should probably do is we'll separate the veggies and then we can cut them. And basically, what we want to do is just cut the dirty part off with the dirt and then, yeah, separate that across. So, you can do that. And while that's going, I guess I could ask more about your background. So, in a previous life, you were once a trader. And even Sam, I think, last year in April also tweeted about how if you're a high frequency trader, you should consider joining OpenAI because, you know, build AGI. So, do you think there's a relation between being a trader and being a researcher, or do you think it's just like a very technical and competitive area where a lot of great employees can come from?
我认为最重要的是,很多研究员一开始并没有接受过机器学习或 AI 研究的正规训练。我们非常相信培养人来做这件事。真正的难点在于创造性解决问题和跳出框框思考的能力。并不一定需要博士学位,尽管那确实带来了宝贵的技能。特别是交易,我不认为它有多特别。我有点把它看作是,我们有伟大的数学家加入,有伟大的物理学家加入,但交易是那种很难破解的东西。你无法欺骗现实世界,对吧?这是一个很难优化的指标。而且它有很多特点,比如这是一个非常注重细节的领域。这是一种残酷的硬优化,榨取系统的每一分价值。其中一些技能是可以迁移的。
I think really the most important thing is there are a lot of researchers who just started out without a formal training in machine learning or AI research. We've very much believed in training people up to do this. I think the real hard thing is the ability to creatively solve problems and think outside of the box. It's not so much that you have to do a PhD, even though that does bring a valuable skill set. With trading in particular, I mean, I don't know that it's that special of a profession. Like, I kind of think of it as, you know, we've had great mathematicians join, we've had great physicists join, but trading is something where it's very unhackable. You can't kind of cheat the real world, right? It's a hard metric to optimize. And there's also a lot of characteristics like it's a field where attention to detail really matters. And it's kind of the brutal hard optimization and squeezing out the juice of a system. And some of those skills transfer over.
明白了。嗯,我想对于那些想进入研究领域但没有博士学位的人,你认为他们可以学习哪些主要特质或东西来培养研究品味?因为我觉得这是进入这个领域的主要部分,可能对他们来说很陌生。
Got you. Yeah, and I guess for people who want to get into research, who let's say don't have a PhD, what do you think are the main attributes or things that they can learn to develop research taste? Because I guess that's the main part of getting into this field that may be very foreign to them.
是的,我觉得这有点被高估了。这是你需要培养的东西,但我发现最好的机制就是复现。所以,我认为你应该找那些你真正敬佩的论文,然后尝试完全复现它。我记得很多应用都让我印象深刻。比如,2018 年有 ResNet、Pixel CNN,我通过尝试精确复现训练曲线,达到论文中暗示的训练损失或困惑度,学到了很多。你会看到很多人们不常谈论的技巧,但一旦你深入几层,就能学到这些技巧。而且,真正让我进入这个领域的第一件事是 AlphaGo 对阵李世石。我认为那是很多人的转折点。它鼓舞人心。我真正追求的第一个大项目是:我能让 DQN 工作吗?
Yeah, I mean, I think it's a little bit overrated. It is something you have to develop, but the best mechanism I've found for developing that is really just replication. So, I think you should take papers that you really look up to and just try to fully replicate it. I think a lot of applications stood out in my mind. You know, back in 2018, there was ResNet, there were Pixel CNNs, and I learned so much just trying to replicate the training curves exactly, get to the exact amount of training loss or perplexity that the papers hinted towards. You just see a lot of techniques that people don't really talk about, but once you dive in a couple of layers deeper, you learn those techniques. And yeah, I think really the first thing, too, that got me into the field was when AlphaGo played Lee Sedol. And I think that was a turning point for so many people. And it was inspirational. And the first big project that I really went after was: can I get a DQN working?
是啊,没错。我记得是第 37 手之类的,看那场比赛发生、看到所有进展,以及我们今天达到的水平,尤其是研究方面,真是太疯狂了。
Yeah, that's true. I think it was move 37 or something like one of the games that was pretty insane watching it happen and seeing all that development and seeing also where we have gotten to today, especially with the research.
我的意思是,现在几乎每个领域都能看到“第 37 手”,这难道不疯狂吗?数学里有,计算机科学和编程里也有。我觉得甚至,很多人今年年初醒来就发现,天哪,智能体在我的职业中发挥作用了。他们基本上意识到这些模型可以为他们做长期有意义的任务。
I mean, isn't it crazy that you're seeing move 37s in almost every field now? It's like there's move 37s in math, there's in computer science and coding. I think even yeah, just it feels like a lot of people woke up at the start of this year and were like, man, agents are working in my profession. And they're essentially realizing that these models can just do long horizon meaningful work for them.
是啊,没错。看到这些确实令人印象深刻,我自己也在工作中使用。但好的,接下来我们可以做切洋葱,很简单。我们只需要把它切丁。你认为有没有 RL 更难突破的工作?比如,编程可能更容易,因为很多上下文可以通过代码库或你正在做的工作获得,但假设你尝试做初级顾问的工作,上下文有点分散,可能更难。你如何看待这些不同的场景?有没有办法评估哪种方法合适?
Yeah, no, that's true. It is very impressive to see, like I'm even just using it in my own work. But okay, the next thing we could do is just as simple as cutting the onion. So, what we have to do here is just like dicing it. Do you think there's jobs that RL maybe will have a much harder time to kind of break into? So, for example, coding maybe easier since a lot of context is accessible either through the code bases or even the work you're trying to do, but let's say if you're trying to do the job that a junior consultant may do, where all the context is a little scattered, maybe it more difficult. How do you view through like those different scenarios? Is there a way that you kind of assess what can be the right approach?
是的,我认为 RL 传统上在那些更主观而非客观的领域面临阻力。所以,举个例子,创意写作就是如此。
Yeah, I mean, I think RL's traditionally had headwinds when it's come to fields that are more subjective than objective. So, if you kind of think of one example of this is creative writing.
比如,你拿两篇创意写作,两位专家可能会有截然不同的看法。
Where, you know, you could take two pieces of creative writing and two experts can have wildly different opinions.
对。
Yeah.
所以,正是这些难以评分的领域,强化学习最难以直接应用。我知道很多人正在开发在这些场景中应用强化学习的技术,但就目前而言,只有在那些有铁一般事实的领域,比如数学和计算机科学,你可以证明对错,才能真正看到它蓬勃发展。
So, it's these fields where things are hard to grade. Um where you know, RL has the least amount of ability to kind of go and um and directly apply there. I know a lot of people are developing techniques to apply RL in these um these settings, but um for now, it's just where there's cold hard truth, things like math and computer science, where you can prove it correctly or wrong. Um that's where you kind of see it really taking off.
对。这其实引出了一个关于评估这些领域的想法。随着模型变得越来越强大,甚至达到饱和,比如解决国际数学奥林匹克问题。
Yeah. No, that actually brings up a thought on in terms of evaluating those fields. So, um you know, as models get much much more powerful and even saturate, for example, solving like the IMO questions.
对,对。
Yeah, yeah.
你怎么看待评估超人类智能?当它变得如此擅长,甚至能完成只有顶尖 0.01%的人类才能做到的事情时,我们如何突破智能的前沿?
Um how do you view evaluating like superhuman intelligence? Like it gets to a point where it's so good at things that even the top what, 0.01% of humans can do, that like, you know, how can we push past that frontier of intelligence?
这确实有点疯狂,我觉得很大程度上这归结于与现实世界的交互。当我们思考如何超越编程竞赛这类事情时,我认为最初的方向是转向真实世界的研究。我们看到模型在发现新定理和推动硬科学前沿方面变得非常擅长。但即使到今天,这也不再令人惊讶。我们几乎已经习以为常,认为这些模型能解决非常困难的问题,做出贡献,甚至在不同领域之间建立新颖且有洞察力的联系。所以,我认为编码协作是一个很好的测试领域,检验我们的模型能否在高上下文和真实世界的长期任务中学习。
No, it's kind of it's kind of crazy, and I feel like um a lot of it centers in on in on kind of interfacing with the real world. And um when when we've thought about how to evolve past things like programming contests in the past, um I think a lot of the initial direction we took was you should move it to real-world research, right? And we've seen that the models uh they've gotten a lot better at uh just kind of discovering novel theorems and uh pushing the frontiers of of hard sciences. But even today, right? That's no longer a surprise. I think like you we almost take it for granted now that these these models can solve very very difficult problems. They can make contributions and even kind of draw relationships between um fields that, you know, that are that are novel and insightful. So, I think you know, we we think of coding co-working as really a domain for that that tests whether our models can learn in high-context settings and in real-world long-horizon settings.
明白了。好的,有道理。既然你切完了所有蔬菜,我们可以进行下一步,就是炒菜。我们可以用电感应炉,之前见过,功率很大。我来打开它。我们就炒一下。
Got you. Okay. Yeah, that makes sense. And since you're done with all the vegetables, we can now do the next step, which is sauteing it. So, yeah, we can use the impulse stove, which we've seen before and very powerful. Let me just turn it on. Um and yeah, so we'll just saute it
好。
Great.
加点油。把锅放在前面的灶眼上,然后
with some oil. So, yeah, put the pan in the front burner and then
这炉子真酷。
These are cool stoves.
对,你可以倒点油。然后,对,一大勺就好。完美。然后我们可以打开炉子。按一下,然后转动旋钮。好,完美。等它加热的时候,我们可以等一下,然后加入蔬菜。不过,我想多聊聊研究方面的观点。有没有一些普遍接受的观点是你不同意的?比如预训练已死,或者语言模型永远无法实现 AGI。我觉得有很多说法很模糊,显然还没有被证实。从你在 OpenAI 领导研究的角度来看呢?
Yeah, you can use oil to pour some in. And then, you can also Yeah, just a good dollop. Perfect. And yeah, and then we can turn on the stove. So, just press it. And then, yes, spin the knob. Great. Perfect. And then while that heats up, we can just wait and then add the vegetables. But yeah, I guess more so on views for research, are there I guess, you know, commonly accepted ideas that are, you know, you disagree with, whether it be like pre-training is dead or language models will never get us to AGI. I think there's a lot of takes out there that are very ambiguous and obviously haven't proven out yet. And I guess from your perspective as like the research like leading things in OpenAI.
嗯。
Yeah.
比如这些?
Like things like those?
我坚信指数增长和缩放定律。所以,任何看衰的观点我都强烈反对。关于预训练已死,有趣的是,这种说法大概只是最近一两年才开始广泛传播。在 LLM 发展史上,人们多次说过类似的话。总有一些瓶颈,人们说无法突破,但我们总能找到某种技术,无论是更好的工程还是新的研究洞察,帮助我们突破边界。所以,我认为这只是同样的过程:更精细的研究工程、数据工程和规模扩张。它总能解锁下一步的扩展能力。它已经持续了将近十个数量级,没有理由不继续下去。
I mean, I I firmly believe in exponent being on the exponential and in scaling laws. So, I think any of these bear takes, I fairly strongly disagree with. Um you know, when when it comes to pre-training is dead, I I mean, I think the the funny thing is this narrative only started spreading more widely after, let's say, the last one or two years or so. Right? In many times uh in the history of uh developing LLMs, people have been saying this, right? And you know, um there there've always been some some bottlenecks that people will Well, you can't scale past this because of this bottleneck. Um and we've always found some kind of technique, whether it be better engineering or some new research insight that helps you break past the boundary. And so, I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling. And it always unlocks that next ability to scale further. So, I I mean, it's held for, you know, almost 10 orders of magnitude, but there's no reason it should not keep keep holding.
对,很有道理。那么,在帮助你突破规模的研究赌注中,有没有早期大家都说行不通的具体想法?
Yeah, that's a very fair point. I guess on research bets that have helped you scale beyond, were there specific ideas that you can even remember in the early days that everyone was was saying that this is not going to work?
嗯,我认为推理就是最大的例子之一。我们向世界推出的第一个突破是 O1,但起步并不容易,因为当时的世界是预训练加后训练的范式,那看起来非常有前景。
Well, yeah, I mean, I think of reasoning as one of the biggest examples of this. Yeah, and um you know, the the first breakthrough that we launched to the world here was O1, but it wasn't easy to get that off the ground cuz one the world we were back living in back then, it was one where pre-training plus post-training, right? That felt like such a promising paradigm.
对。
Yeah.
所以,即使在 OpenAI 这样的公司,人们自然会问:既然有能用的机器,为什么还要做别的?从根本上说,这要归功于 Yakub、Ilya 等许多对这个领域有坚定信念和远见的人,我们才开始认真推动这件事。即便如此,也花了很多引导才让整个公司支持这个基本赌注。
Um and so, even at a company like OpenAI, you would have people ask naturally, why do something when you have a machine that works? And fundamentally, you know, it's to the credit of you know, Yakub, Ilya, many of the people who really had conviction and vision in the space um that we started pushing on this in earnest. And even then, it took a lot of steering to get the whole company behind this as a as a fundamental bet.
明白了。
Got you.
对。
Yeah.
你如何培养激励研究人员的能力?因为我认为这是很重要的一部分:下很多赌注,有些不会成功,但依然建立团队的信任,让他们知道最终有些会带来幂律效应。
And how do you kind of develop that ability to motivate researchers? Cuz I assume that's a big part of you know, taking a lot of bets and some won't pan out, but still building the trust in the team to know that eventually some of these will actually have no power law effects.
OpenAI 很酷的一点是,研究感觉像是一个精英体制。研究经理通常是过去做过最好研究的人。所以很多引导可以是自上而下的。如果你的经理说“我坚信这是前进的道路”,人们通常会认真考虑。毕竟,你长期尊重这个人的研究品味和执行能力,现在他对这个想法很兴奋。这肯定会被考虑。所以,有很好的自上而下的引导。同时,OpenAI 另一个很酷的地方是自下而上的元素。我们喜欢被说服自己错了。有人可以拿出确凿的证据。很多这样的事情最终成为我们研究路线图的核心部分。有些东西没人刻意引导,但某个基层研究员有强烈的信念。看到这些也让人非常高兴。
You know, what's what's really cool about OpenAI is um research it feels like a meritocracy. So, um often times the research managers are the people who um do the actual have done the best research in the past. And so I think a lot of steering can come top-down, right? Like if your manager says, "Hey, you know, I'm like really convinced this is the path forward." Um, generally people will take that into heavy consideration, right? It's like, you know, this person who you've respected for their research taste and execution for so long is like now very excited by this idea. Um, it's it's definitely something that uh that um yeah, you you you people take into account. So, I think there there's good top-down steering. At the same time, you know, I think one really cool thing about OpenAI is um there bottom-up elements. Like we like to be convinced that um uh you know, that we're wrong, right? And and someone can just come with cold, hard evidence. And many things like that have turned into core parts of our research roadmap. Just things that no one was really kind of trying to steer, but some researcher on the ground had a heavy conviction in. Um, and and that's also a really good big delight to see.
对,绝对。
Yeah, no, absolutely.
我最近听到一个采访,你说内部研究路线图基本没变,尽管我们看到模型开发和其他公司都在进步。你们多久评估或重新评估一次?是主动行动吗?我猜不会因为其他模型出来就做很多被动决策。但你是怎么思考这个过程的,尤其是周围一切都在不断变好的情况下?
I heard in a recent interview that your internal research roadmap hasn't really changed even through all we've seen with model development and other companies. How often do you assess or reassess it, and do you act proactively? I assume it's not a lot of reactive decision-making as other models come out. But how do you think through that process, especially as everything around you continues to get better?
高层研究路线图应该是稳定的。人们需要一些基础,一条通往我们正在构建的东西的路径。我很高兴我们坚持了一段时间。但实现细节会随时间变化。顺序、相对资源分配和具体步骤都很重要。我们有一些时间点会迫使重新考虑,比如在做算力分配时。那是质疑我们是否把算力用在了最高优先级的赌注上的时候。
The high-level research roadmap should be stable. People need something to ground in, a path to what we're building. I've been happy that we stayed the course for a while. But implementation details can change over time. The sequencing, relative resourcing, and exact steps on the ground matter. We have points in time that force us to reconsider, like when we do compute allocation. It's a time to question if we're putting compute to use on the highest priority bets.
你能更清楚地解释一下高层和实现细节的区别吗?高层是像 AGI 这样的总体目标,还是更具体一些?
Could you clarify more what you mean by higher level versus implementation detail? Is high level as general as AGI, our North Star, or more granular?
在最高层面,我们有一个专注于预训练的团队,给模型世界知识。我们关注强化学习,教模型如何推理和串联洞察。然后是对齐和后训练。我们关注每个领域主线的 Scaling,以及能解锁不同或更激进 Scaling 特性的新赌注。
At the very highest level, we have an org focusing on pre-training, giving models world knowledge. We focus on RL, teaching models how to reason and chain insights. Then alignment and post-training. We look at scaling the mainline in each domain and new bets that unlock different or more aggressive scaling properties.
我听说每一到两个月你们会审查大约 300 个研究项目。鉴于很多有才华的研究人员会提出想法,你们如何优化决策,决定哪些要加倍投入?
I heard that every one to two months you go through about 300 research projects. How do you hone decision making on what to double down on, given many talented researchers provide ideas?
本着专注的精神,我们在 OpenAI 真正聚焦赌注,并做更直接的算力分配。我不喜欢微观管理经理;我授权他们,但把大量算力给大赌注。也给他们灵活的计算池用于他们相信的事情。我们把每个团队的三到五个赌注与主要研究路线图挂钩,然后让经理和团队负责人去执行。
In the spirit of focus, we're really focusing our bets at OpenAI and doing more directive compute allocation. I don't like micromanaging managers; I empower them but give big swaths of compute to big bets. Also give them flexible pools for things they believe in. We tie a small number of bets, say three to five from each org, to the main research roadmap, then let managers and org leads take it from there.
对于有潜力的研究人员,在面试中有什么特定迹象可以识别潜力?还是只看他们之前的研究?
For rising researchers, in an interview setting, are there specific tells to identify potential? Or is it just looking at previous research?
在有人加入 OpenAI 之前,这是个难题。最好的研究经理通过与许多研究人员合作培养直觉:他们说的话、提出的想法,是否切中要害或与你的想法一致。这是一种直觉检验。但一开始很难判断。六到十二个月内,谁有最强的发展轨迹就很清楚了。看过很多人经历研究发展,你会培养出直觉。不是每个研究人员都一样;有不同类型的影响力:有人快速实现清晰的想法,有人提出疯狂但又不那么疯狂的想法,以不同方式说服你。
It's a hard problem before someone comes to OpenAI. Best research managers develop intuition from working with many researchers: the things they say, ideas they bring up, whether they hit the same mark or match your own thinking. It's a gut check. But it's hard to tell out of the gate. Within six months to a year, it's clear who has the strongest trajectory. Having seen many people go through research development, you develop intuition. Not every researcher is the same; there are different types of impact: those who implement clear ideas quickly, and those who come up with crazy but not too crazy ideas that convince you in a different way.
你觉得顶尖工程师和顶尖研究人员有相似之处吗?顶尖工程师能把一个想法的苗头带到生产。在研究中,是从提出想法到交付,还是只关注研究而不考虑最终设计?
Would you say there are similarities between top engineers and top researchers? Top engineers take an iota of an idea to production. In research, is it coming up with the idea all the way to delivery, or focusing solely on research without considering end design?
在研究中,前进的道路往往不清晰。研究人员的区别在于他们多常被指向正确方向,以及有好的研究品味。在工程中,模式有效;原则可能相似。对于研究,是拥有好品味、说服别人你的工作有前景,并将其整合到核心研究路线图中的能力。
In research, the path forward is often unclear. What differentiates researchers is how often they're pointed in the right direction and having good research taste. In engineering, patterns work; principles can be similar. For research, it's the ability to have good taste, convince others your work is promising, and integrate it into the core research roadmap.
一个非常有趣的方面是评估。更具体地说,有没有出现过这样的情况:通过直觉检查感觉模型很好,但在实际基准测试中表现很差?或者你认为两者高度相关,如果 Sweet Bench Pro 分数高,那么对编码任务的直觉检查也会很高?
One aspect that seems very interesting is evals. More specifically, have there been instances where through vibe checks it seemed really good, but on the actual benchmark it performed very poorly? Or do you think it's heavily correlated that if your Sweet Bench Pro is a high number, then your vibe check on coding tasks is also really high?
不,不。我认为存在这种现象。内部我们称之为“bench maxing”。你可能会过度拟合某些分布,这并不能反映你的泛化能力。简单的方法就是拿一个基准测试,找到非常相似的实例,然后过度训练。除此之外,另一个可怕的事情是经典黄金标准基准测试的数量很少。
No, no. I think there is this phenomenon. Internally, we call it 'bench maxing'. You can overfit onto certain distributions, and it won't reflect how well you generalize. Easy ways to do this are taking a benchmark and finding very similar types of instances and overtrain on them. Beyond that, the other scary thing is the number of canonical gold standard benchmarks is low.
是的。
Yeah.
我们确实处于评估危机中。所有我们熟知的优秀评估,比如 SAT,都已经饱和了。我们需要找到新的好方法来对模型进行基准测试。像 Codex 这样的工具的一个好处是它们能够快速迭代评估。我们可以让一个人快速构建一个高质量的评估。另一个有趣的事情是,部署模型可以让你看到人们在使用它们时的评估情况。在数学、编码和软件领域,你可以感受到它们在哪里失败,以及它们能处理的任务范围。
We really are in an evals crisis. All the great evals we know, like the SAT, are fully saturated. We need to find good new ways to benchmark models. One great thing about tools like Codex is they enabled fast iteration of evals. We can have one person quickly put together a high-quality eval. Another interesting thing is deploying models lets you see them eval as people use them. In math, coding, and software, you get a sense of where they fall over and the task horizon they can handle.
是的,这很有帮助。你如何平衡在基准测试中表现良好但不进行 bench maxing?你想保持诚实,不欺骗系统,但如果你的分数低于竞争对手,消费者可能会认为模型不好。你如何平衡这种矛盾?
Yeah, that's helpful. How do you balance doing well on benchmarks but not benchmark maxing? You want to be honest and not cheat the system, but if you have lower scores than a competitor, consumers might think the model isn't good. How do you balance that dichotomy?
你必须在具有代表性的评估组合上操作,并始终投资于创建新的评估。有一种理念是,一旦评估发布到世界上,它就不再是一个好的评估。此外,与外部组织合作创建评估也有帮助。在许多困难的数学和科学评估中,我们与外部组织合作制定了黄金标准。有一个有趣的理念是将创建评估的团队与优化模型本身的团队分开。这样你就不会让他们产生共谋激励。评估团队试图构建对模型来说困难的评估,创造一种对抗性过程,这样你就不会欺骗自己。两个团队之间的激励以正确的方式对齐。
You have to operate over representative mixtures of evals and always invest in creating new evals. There's a philosophy that once an eval is out in the world, it's already not a good eval. Also, partnering with external organizations to create evals helps. In many hard math and science evals, we've partnered with external organizations to craft gold standards. There's an interesting philosophy of separating the teams that create evals from the teams that optimize the models themselves. That way you don't co-incentivize them. The evals team tries to build evals that are hard for the models, creating an adversarial process where you don't cheat yourself. The incentives are aligned in the right way between the two teams.
是的。你是否也参与构思过程或决定与第三方合作开发哪些评估?
Yeah. Do you also contribute and help in the ideation process or deciding what evals to work on with a third party?
是的。Yacoub 和我做的很多工作都涉及引导评估的方向。我们会注意到某些差距或我们想要的能力。前沿的每一项能力都是一个评估——你需要某种评估来衡量你是否准确地引出了你想要的东西。这需要很多引导,让每个人在评估上达成一致也是一项艰巨的工作。
Yeah. A lot of the work Yacoub and I do involves steering the direction that evals go. We notice certain gaps or capabilities we want. Every capability on the frontier is an eval—you need some eval that measures if you elicited exactly what you want. It takes a lot of steering and getting everyone on the same page with evals is tough work.
你在之前的采访中说 Yacoub 是个很有趣的人。你有没有一些没分享过的和他一起工作的趣事?你还说你们很合拍,所以研究讨论很高效。在有趣的反面,有没有什么……
You said in a previous interview that Yacoub is a very funny guy. Do you have any fun stories you haven't shared about working with him? You also said you align very well, so discussions on research are efficient. On the opposite side of being funny, are there things that you...
是的。你问到一个有趣的故事。他昨天给我讲了个笑话,我觉得很好笑。在很多方面我们共同管理研究工作。显然有个研究员来找他说:“感觉我现在有了一支非常愚蠢的 IOI 形式主义大军。”Yacoub 回答说:“这感觉就像我已经身处其中的情况。”他就是非常尖刻又幽默。
Yeah. You asked about a funny story. He told me this joke yesterday, which I thought was very funny. In many ways we jointly manage the research efforts. Apparently some researcher came up to him and said, 'It feels like I now just have an army of really dumb IOI formalists.' And Yacoub was like, 'That feels like already the situation I'm in.' He's just brutally sarcastic and funny.
是的,太好了。在工作场所拥有幽默感很好,尤其是在推动前沿重要工作时。这让我想到一个奇怪的情况:模型在 IMO 或 IOI 上表现很好,但在人类容易做的平凡任务上却挣扎。你怎么看?
Yeah, that's great. It's great to have humor in the workplace, especially when pushing the frontier on important work. That brings to mind a weird scenario: models can perform very well on the IMO or IOI, but struggle with mundane tasks a human can easily do. How do you feel about that?
是的。最终,对模型来说直观的东西往往对人类来说并不直观。有很多关于“锯齿状前沿”的类比,模型在某些事情上天生擅长,这基于它们看到的数据或我们更容易教给它们的东西。
Yeah. Ultimately, what's intuitive for the models is often not intuitive for humans. There's a lot made of the 'jagged frontier' analogy where there are some things the models are inherently good at, based on the data they see or what we can teach them more easily.
嗯,我其实觉得,很多问题归根结底还是上下文,对吧?模型没有那么多上下文可以利用……视觉当然是人类生物学上更自然的能力。所以,是的,我认为模型在某些能力上比人类强,反之亦然,存在一种锯齿状的能力分布。但我也认为,上下文——也就是能够从单个任务中学习经验并应用到未来任务——这种能力是现在很多 AI 研究者正在努力的方向。不过,这对人类来说确实很自然。
Um I actually think, you know, a lot of it boils down to also just context, right? The models don't have a lot of context that it can... Vision, of course, is something that's more naturally biologically wired for humans. And so, yeah, I think there are just certain kind of jagged capabilities that models are better at than humans and vice versa. But I also think, you know, context, just being able to take a single task, learn lessons from it, and apply them to future tasks, that capability is something that, you know, a lot of people are in AI working towards right now. But it's, yeah, very natural for humans.
是的。关于上下文这一点,很多人会举一个很简单的例子,就是增加上下文窗口来提供更多示例,让模型表现更好。但你认为,实际上要实现这一点是否更复杂?因为即使有大的上下文窗口和很多上下文,也可能出现冗余,甚至像人们说的上下文腐烂。那么,你如何处理这个过程呢?
Yeah. And on the context point, a very low-hanging fruit example that many people will say is just to increase the context window to provide more examples so the model can perform. But do you think I assume there's more complexity on how to actually enable, since even with a large context window and a lot of context, there could be bloat or even just a lot of like context rot, as people have said. So, how do you go through that process of...
是的,我的意思是,我认为解决超长程学习有一种经典方法,就是简单地增加上下文窗口,对吧?这说得通。但实现长上下文和很好地实现长上下文是有区别的,就像你说的。有很多类似大海捞针的评估来衡量这一点。但除此之外,我认为还有很多工程和研究上的捷径可以走。比如,现在很多编码产品都有压缩功能,对吧?你可以压缩见解或工作状态之类的东西,这就能绕过很多用原生长上下文必须构建的极其困难和昂贵的原语。
Yes, so I mean, I think there's kind of the canonical way you would solve for very long horizon learning, which is, you know, you just naively increase your context window, right? And I mean, that makes sense. I think there's a difference between implementing long context and implementing long context well, like you said. And there is a lot of kind of like needle-in-a-haystack style evaluations to measure that. But I do think beyond that, there are also a lot of, in some sense, like engineering and research shortcuts that you could take. So, like, many coding products today have features like compaction, right? Where you can compress kind of either insights or working state and stuff like that, you know, it just shortcuts a lot of the brutally difficult and expensive primitives that you have to build with just native long context.
明白了。很好。
Got you. Great.
好了,现在我们要做有趣的部分了。我们把火调小一点,再加一点油到锅里。然后我们用喷枪烧一下虾,让味道更浓。
Okay, now we're going to do the fun part. So, let's lower the heat a little bit and then add a little bit more oil to the pan. And then we'll torch the shrimp the prawns to get a little more flavor in there.
好的。
Yep.
我先在我的锅里做一下,给你看看效果。
So, I'll first do it on my pan to show you what it looks like.
但是……一次学习。
But... One shot learning.
是的,确实。
Yeah. Indeed.
等等,我没倒波本酒。
Wait, I didn't pour any bourbon.
好的,等等。我们倒四分之一杯。好。然后倒进去。关火。然后喷烧。
Okay, wait. So, let's pour like a fourth. Okay. And then pour it in. Heat is off. And then torch it.
太棒了。很好。
Awesome. Great.
好的。
All right.
好的,我觉得我懂了。
Okay, I think I got this.
好的,是的。你想自己做吗?好。
Okay, yeah. So, do you want to do it your... Yeah.
是的,是的。然后……
Yeah, yeah. And...
所以,倒到四分之一杯的一半,然后你有了之后,我可以把这个给你。
So, pour it to like half of the fourth cup and then once you have that, I can give this to you.
完美。很好。所以,火关了。然后关掉之后,你可以再打开。
Perfect. Great. So, it's off. And then once that's off, you can turn it back on.
完美。好的。
Perfect. Okay.
然后现在,你想拿着这个,按下这个按钮点火吗?好。
And then now, do you want to hold on this and just press this button to fire it up. Yeah.
酷。是的,完美。很好。好的。我在烧它。火有点小,但还行。很好。然后我们可以再开火。
Cool. Yeah, perfect. Great. Okay. I'm burning it. It's a little light, but yeah. Great. And then we can turn on the heat again.
好的。
Okay.
然后我们就把酒精煮掉。
And then we'll just cook off the alcohol.
好的。很好。很好。很好。
Okay. Great. Great. Great.
很好。你感觉怎么样?
Great. How are you feeling?
你知道……
You know...
很好。很好。基本上都在煮。所以,好的。
Great. Great. Basically there cooking everything. So Okay.
太棒了。
Awesome.
是的,我想就研究思路和方向而言,你觉得还有很多唾手可得的成果,或者通过优化已实现工作的小部分就能大幅改进的想法吗?还是说现在人们必须进行很多全新的押注式研究?
Yeah, and I guess in terms of research ideas and what to work towards, do you think there's still a lot of low-hanging fruit or ideas that can still be improved a lot through just optimizing small parts of already implemented work or do you think right now there has to be a lot of research that are completely new bets that people take?
嗯,是的,这是个非常好的问题。我觉得有新赌注,但可能不多。是的,从某种意义上说,希望你觉得 AGI 即将到来,对吧?我认为每个人都看到这些模型变得越来越强大,如果你真的想象这带来的影响,我们正越来越接近一个模型能提出更多创新的世界。
Um yeah, that's a really great question. I feel like there are new bets, but probably not that many. Yeah, in some sense like hopefully you feel like, you know, AGI is coming soon, right? And I think everyone sees that these models are getting really capable, and I think if you really imagine the implications of that, we're getting closer and closer to a world where the models can come up with more of the innovation funder out.
是的,如果它们能进行自我维持的研究,这是我们研究工作设定的主要目标之一。
Yeah, if they can kind of do self-sustaining research, this is one of the big or goals that we've set for for our research work.
所以我认为,真正重要的是在那个时间点之前有没有大的赌注?
And so I think like you know, what really matters is are there big bets before that point in time?
明白了。
Got it.
而且我认为窗口期很短,但还有一些相当重要的想法我们在尝试。
And I think the window is small, but there are still like some pretty significant ideas we're trying out.
是的。我的意思是,有些研究人员说过,要达到 AGI,我们还需要两三个突破,比如持续学习或其他想法。你同意这种观点吗?还是说你觉得没那么夸张,不需要三个完全不同的范式?
Yeah. I mean, there've been some researchers who have stated that to get to AGI, we still need, let's say like two or three more breakthroughs, if you like continual learning or some other ideas. Do you follow that same view perspective, or do you think it's kind of more so like not as drastic as coming up with like three completely different paradigms?
嗯,是的,我的意思是,我不知道。我不确定我是否认同那个框架。比如持续学习是一个你必须解锁的基本原语。是的。有很多不同的技术。我不知道……是的,我认为我们正在尝试很多它的局限性。我不知道什么算突破,什么不算,但我认为显然有很多尝试的机会,而且我很确定它们会成功。
Um yeah, I mean, I don't know. I don't know if I have that same framing. Like continual learning is a basic primitive that you have to unlock. Yeah. There's so many different techniques. I don't know that yeah, I think, you know, we're trying a lot of in the limitations of it. I don't know what would consider as a breakthrough versus not, but I think there clearly many shots on goal, and I'm pretty sure they'll work.
很好。好的。
Great. Okay.
嗯。
Mhm.
很好。好的,虾基本上好了。
Great. Okay, so the shrimp is basically done.
太棒了。
Awesome.
好的。你想再点火烧一下,让颜色更深吗?
Okay. So, do you want to do the flambé thing again to get more color?
来吧。是的。
Let's do it. Yeah.
好的。我……把火开大一点,然后我再加点油。
Okay. I'll... Turn on the heat a little bit, and then let me get some more oil.
很好。
Great.
是的,因为你想得到一些深色。我觉得我的已经有了。给你。很好。然后我们希望能再来一次。好的。
Yeah, cuz you want to get like some dark color. I think mine has. And here. Great. And then we can hopefully get another shot. Okay.
完美。等一下。
Perfect. minute.
是的。你想先来吗?
Yeah. Do you want to go first?
嗯……
Um...
好的。
Okay.
同样的啤酒量。
Same same amount of beer.
同样的量。
Same amount.
我们希望能成功,因为我觉得火候不够……是的。
We could hopefully get cuz I think the heat wasn't as... Yeah.
好的。我觉得我们倒进去了。
Okay. I think we put it in.
是的。我觉得……好了。我们按下按钮。
Yeah. I think... There we go. Let's press the button.
很好。
Great.
好了。
There we go.
成了。
There it is.
口感不错。
It's good texture.
是的,确实。而且增加了很好的风味。来,我尝尝。应该不错。
Yeah. Indeed. And it's like add some good flavor. Here, let me try it. It'll be good.
好的。
Yep.
是的,看看。
Yeah, let's see.
很好。好了,看看会是什么样子。
Great. All right, let's see what it's going to look like.
哦。成了。好了。
Ooh. There it is. All right.
但我们到了最后阶段。
But we're in the final stretch.
好的。
Okay.
虾都熟了。
shrimp all cooked.
是的。
Yep.
有点火。所以,现在我们可以……是的。
Some fire. So, now we can kind of... Yeah.
稍微煮一下,然后我们应该把蔬菜加到水里。
Cook it off a little bit and then we should add our veg to the water.
我对你的多任务处理能力印象深刻。你知道,我认为这其实是我们的模型需要改进的一个方面。比如它应该能同时处理这样的任务,还能和你聊天。
I'm impressed by your multitasking abilities. You know, I think that's actually one thing we need our models to get better at. Like it should just be able to do some thread like this and also just have a conversation with you at the same time.
对对对。
Yeah. Yeah. Yeah.
不,这也让我想到。你认为图像、音频、视频甚至文本都应该统一在一个模型下,还是说它们会各自突破,比如专门的音频模型?
No, that also reminds me. Do you think images and audio and video and even text like that should all be one under one model or do you think it'll break through like specific specialized like audio model or
嗯,是的。我认为对于一个研究实验室来说,统一在一个模型下有很多好处。比如你只需要维护一个基础设施栈。我认为同时维护和扩展多个基础设施栈的成本不容低估。所以,我认为有很多好处,比如你在基础栈上做一些核心研究,然后这些成果可以迁移到你想要的任何模态或任务上。因此,我们强烈倾向于尽可能保持架构的统一。
Well, yeah. I mean, I think for a research lab, there are a lot of advantages for it to be under one. You just have to maintain one infrastructure stack for instance. I think the cost of maintaining and scaling many infrastructure stacks at once is something that you shouldn't underestimate. So, I think there are a lot of benefits to just like you know, you do some core research in your fundamental stack and that just carries over to whatever modality or whatever thing that you want. So, I think there's a strong bias for us to keep it in as few different architectures as possible.
明白了。
Got you.
嗯。
Yep.
很好。
Great.
不,这很有道理。我认为架构本身也是一个经常被忽视的因素。
No, that makes a lot of sense. I think the architecture as well is something that isn't often considered.
是的。
Yeah.
这非常重要。但我经常看到的一个词,你也提到过,就是“氛围研究员”。你知道,我们有“氛围程序员”。
It's very important. But one term that I've been seeing a lot that you've also kind of mentioned is a vibe researcher. You know, we have vibe coders obviously.
对对对对。
Yeah, yeah, yeah, yeah.
但我想问,对于“氛围研究”,你认为最终状态是什么?你认为“氛围研究员”的主要价值在于提出正确想法的研究品味,还是在于实际执行和跟进研究?
But I guess on vibe researching, what do you think is like the end state? Do you think the main value out of a vibe researcher is just the research taste of coming up with the right idea or do you think it's more so the execution of going through and following through on the actual research?
嗯,是的,我认为我们正在快速走向这样一个世界,对吧?我认为无论是 OpenAI 还是其他实验室,你开始看到很多工作变得以编排为主,对吧?比如研究人员提出想法,模型足够强大,可以自己执行实现。所以我认为,当谈到提出想法与执行的价值时,两者仍然重要,但我感觉市场正在转向:你只需要提出大量想法,然后模型可以为你执行和编排。我认为这将是未来做研究的方式。我们之前也说过,模型还没有很好的品味。这就是为什么你仍然需要研究人员来提出想法。教模型好的品味很难。我们注意到了这一点,但在加速研究方面,已经看到了切实的好处。
Um yeah, so I think we're actually moving towards this world very quickly, right? I think both that OpenAI and our other labs, you're starting to see a lot of the work become mostly orchestration-focused, right? Like the researchers coming up with the ideas and the model's great enough to do the implementation execution by itself. So I think when it comes down to like you know, is the value of coming up with ideas versus execution? Both are still important, but I just feel like there's this market shift towards just kind of being able to come up with a lot of ideas and then the model can actually do the execution and orchestration for you. So I think that's going to be the future of doing research. We also said earlier, you know, like the models don't quite have the taste yet. And that's why you still need the researchers coming up with the ideas. It's going to be hard to teach the models good taste. We noticed that, but in terms of actually accelerating the research, there's clear tangible benefits already.
是的。你认为模型在研究品味上会达到与人类相当的水平吗?
Yeah. Do you think they'll ever be parity in terms of research taste with models?
我认为会。当我们看我们的三年路线图时,最终目标是让模型能够进行端到端的研究,而其中一部分问题就是让模型拥有好的品味。你给它一个通用基准或类似的东西,它就能找到正确的解决方案。是的。
I think so. I mean, when we look at our kind of three-year roadmap, right? The end goal that we want to reach is one where the models are just doing end-to-end research and I think a part of that problem is just being able to have the model come up with good taste. You point at some generic benchmark or something and it finds the right solutions. Yeah.
是的,这很有帮助。关于 OpenAI 人类进行的研究,你们如何处理那些结果不佳的研究赌注的事后分析?我猜很多都是下大赌注,有些不会成功。
Yeah, no, that's helpful. And in terms of research done by humans at OpenAI, how do you guys go about the postmortem process of let's say a research bet that didn't turn out well? Because I assume that a lot of it is taking these vast bets and some don't turn out well.
嗯,我认为这是 OpenAI 优势的重要组成部分,因为我认为我们与其他实验室的一个区别是我们承担了很多高风险赌注。我认为这正是我们能够长期保持在最前沿的原因。但这也意味着有些赌注不会成功。一个必然结果是,当赌注失败时,你不能自欺欺人地认为它还会成功,而要果断放弃。所以,我认为你必须做出一些判断,对吧?比如回顾一下,嗯,这个想法当时很有前景,但实际上它没有我们想象的那么重要,有其他更好的方法,或者我们发现了其他东西。但我认为这些工作很多也很有成果。我们意识到,即使有时人们未能证明一项技术,他们的报告也非常重要,因为这些想法往往很自然,你可以让很多人避免经历同样的痛苦。
Well, I would say that is a big part of OpenAI's alpha because I think one thing that differentiates us from other labs is we take a lot of high-risk bets. And I think that's what's allowed us to stay at the frontier so consistently over time. But it also means that some of the bets are not going to pan out. And a hard corollary of that is when a bet doesn't pan out you have to not delude yourself into thinking that this is something that will work and kind of disconnect from it. So, I think there are certain calls that you have to make, right? Like look back and be like, well, this was a promising idea at the time, but actually it's less important than we thought, there's some other approach that works better or something we discovered. But I think much of that work is also very fruitful. So, what we realized is even sometimes when people fail at proving out a technique, their write-ups are very important because they'll often be a natural idea and you can perhaps save a lot of people from going through the same pain.
是的,这很有帮助。
Yeah. Well, that's helpful.
那么,当谈到这种对失败的积极看法时,你如何平衡这一点?比如一个研究员下了很多赌注,连续赌,但都没有成功。因为我认为在某个时候,你希望研究员最终能做出真正有益的贡献,而不是只下那些可能不会带来好结果的赌注。
So, I guess when it comes to this positive view on failure, how do you balance that with, you know, a researcher, let's say, takes a lot of bets, consecutive bets, and none of them pan out? Because I assume at a certain point you'd want a researcher to eventually have contributions that are actually beneficial compared to only taking bets that maybe pan out to being not good space.
根据经验,我确实见过一些人陷入这种情况。但我也遇到过几个案例,赌注一个接一个失败,就在你快要沮丧的时候,突然出现了一个巨大的成功。这种情况发生过足够多次,所以这实际上取决于想法本身是否合理?它们可以很雄心勃勃,但必须合理。有一种人就是会尝试很多这样的想法,这没问题,因为他们处于风险较高的前沿。
Just through experience, I've definitely seen some people fall into this. But I've also had several cases where, you know, it's just like bet after bet, it doesn't pan out, and just when you're at the brink of frustration, you have something that's like a mega hit. And this happened enough, so it really just depends on kind of are the ideas themselves sound? They could be ambitious, but they still have to be sound. And there's a certain kind of person who will just take a lot of those ideas, and it's okay because they're somewhat on the riskier frontier.
嗯。
Mhm.
但他们只需要偶尔证明其合理性就行,对吧?这有点像用交易的眼光看世界,但确实,这只是我们的期望,对吧?他们需要创造价值。
But, they only have to justify it once in a while for it to make sense, right? It may be like a very trading like kind of lens on the world, but yeah, it's just on our expectation, right? Like they need to add value.
是的,这很好。好了,我们基本完成了,现在是收尾工作,你可以尝尝你的汤,如果不够咸就加酱油,如果太咸就加点水稀释。让我们看看最终成果。味道怎么样?
Yeah, no, that's great. Okay, so we're basically assembled, now it's the finishing touches, so you can taste your soup, and then we just add soy sauce if it's not salty enough, and if it's too salty, we can add some water to lower this down. But, let's see our final creation. How is it?
还不错。
It's pretty good.
不错?
Good?
我接受。我的很好。
I'll take it. Mine is good.
嗯,我的需要加点水。你能把水递给我吗?
Yeah, mine needs a little bit of water. Could you pass me the water?
好的,当然。
Yeah, absolutely.
太好了。那么,感觉怎么样?
Great. So, how was that? How did that feel?
嗯,我觉得这是学生蒸馏。你显然比我做得好。
Um I think this is a student distillation. You're clearly better than I am at this.
不不不。
No, no, no.
我觉得你做得非常好,尤其是虾和棕榈心。闻起来很棒。哇,听起来不错。我想更广泛地说,我有点好奇,有没有什么研究方向或话题你觉得现在被高估或低估了?你会怎么归类?
I feel like you did a very good job, especially with even with the shrimp and palm hearts. Yeah. Smells great. Wow, okay, that sounds good. I guess just more generally, I'm kind of curious, are there any reason research or topics that you think are right now overrated and underrated? Like what would you categorize under?
嗯,我认为如果你还持有“预训练已死”的观点,那我觉得预训练绝对还没死,它被低估了。嗯,老实说,我认为产品和思考最终用途,以及如何将研究中构建的所有原语与现实世界中的智能体用例联系起来,这也被低估了。我觉得你真的不能凭空构建一切而不与实用性挂钩。
Um Well, I think if you still have a pre-training is dead view of the world, um I think I think uh pre-training is definitely uh yeah, yeah, yeah, not not dead. It's it's um it's underrated. Um Yeah, and honestly, um I think products and kind of thinking about end uses and you know how you tie all the primitives you build in research to you know real agentic use cases in the world. That's also underrated. I think um you really can't just kind of build everything in a vacuum and not connect things to utility.
是的,说得很好。
Yeah. No, that's great.
太棒了。我觉得我们可以品尝了。那么,我们要试试吗?
Awesome. I think we are ready to taste. So, do we want to give it a go?
来吧。
Let's do it.
我们可以挪一下,我觉得应该摆盘。你想把这个盘子拿过来吗?
We can move this and I think there should be some plating. Yeah, do you want to take this plate over here?
好的。
Okay.
好的。虾看起来很棒。其他东西都在这里。太好了。然后我们可以用这些锅。所以,我们有虾了。
Okay. Shrimp looks great. But everything else is here. Great. And then we could just use these pots. So, we have our shrimp.
嗯。
Yep.
你想尝尝吗?干杯。干杯。让我看看。可能有点甜。有点太甜了。因为我们找到了弥补的方法。
Do you want to try it? Cheers. Cheers. Let's see. Maybe a little sweet. That's a little too sweet. Cuz we found a way to make it up.
对我来说,我喜欢甜的东西。
For me. I like sweet things.
不错。好的,太好了。嗯,肖恩,你想来尝尝我们的松露汤吗?
It's good. Okay, great. Um Sean, do you want to come and taste our truffle soup?
顺便说一句,那太好吃了。你们闻不到,但嗯
That was so good by the way. You guys can't smell it, but um
我们就假装你是一个想升职的研究员。
We'll pretend you're a um researcher that's trying to get a promotion.
我是扎克,他是……你知道的,想让你……所以你想拿那边的勺子吗?
And I'm Zach and he's a you know trying to get you so do you want to grab the spoon over here?
汤真的会左右这里的决定。汤的质量。
Soup is really going to sway the decision here. The soup quality of soup.
是的,太好了。
Yeah. Great.
我来。
I'll do it.
好吧。这里的艺术方向是什么?
All right. What was the artistic direction here?
嗯,艺术方向。嗯,是模仿。我认为模仿中有伟大的艺术。
Um artistic direction. Well, it was mimicry. I think there's great art in mimicry.
是的,就让他们煮。
Yeah. Just letting them cook.
嗯哼。
Mhm.
好。
Good.
哇。是的。
Wow. Yeah.
浓郁。
Strong.
我的意思是,我觉得有一件事我绝对喜欢,就是咸味和香料搭配在一起,还有那种海鲜味。
I mean, I think um one thing I I definitely like like savory and spice that goes together, but then also like the sort of sea seafoodness.
嗯哼。
Mhm.
嗯
Um
是的,真的融进去了。
Yeah. Kind of really goes into it.
太好了。
Great.
好的。我的味道很突出。
Okay. Mine's mine's very dominant.
赢家还是什么?
Winner or what?
不,不,不。你应该先尝尝,然后
No, no, no. You're you're supposed to try it and then
不,你来选赢家。
No, you do pick a winner.
好的。
Okay.
这就像评估?
This is like evals?
是的。
Yes.
嗯。
Yeah.
所以,这很贵。
So, it's expensive.
外部评估。
External evals.
好的,我得说我觉得这里面水太多了。我觉得
Okay, I got to say I feel like there's too much water in this. I think like the the
我觉得也是锅的问题,因为这是个很大的锅。
I I think it's also a pot. Cuz this is a big pot.
是的。等等,好的。嗯,我得说我选这个。
Yeah. Wait, okay. Um I I would say I I have to go for this.
好的。好的。
Okay. Okay.
就像你是我们尊敬的客人,但我想保持客观。
Just like you're our respected guest, but I want to be objective.
当然。当然。是的。是的。是的。
Of course. Of course. Yes. Yes. Yes.
嗯,而且我觉得浓度和味道真的
Um and like yeah, the density I think really a flavor really
我放一半水?一半的水?
I'll do half water? Half the water?
可能吧。
Probably.
好的。听起来不错。
Okay. That sounds good.
我的意思是,我觉得这很个人化,对吧?
I mean I think it's very personal, right?
是的,我也觉得这很个人化。味道,你知道,你提到过
Yeah, I think it's also very personal. Taste, you know, you were mentioning
你经常做饭。
You do a lot of cooking.
味道。
Taste.
不。不。不。
No. No. No.
好的。我知道几个食谱。我严格照着做。如果你告诉我“哦,稍微改一下做法”,我就完全不知道了。我完全不知所措。
Okay. I know a couple recipes. Um I follow them to the T. I can't Like if you tell me, "Oh, cook something slightly different." I have no idea. I'm completely lost.
对。对。
Right. Right.
是的。忙做饭。
Yeah. There's busy cooking.
GPT 可以告诉你。
GPT can tell you.
嗯哼。是的。是的。哦,我不撒谎。我查过几次 ChatGPT。
Mhm. Yeah. Yeah. Oh, I I'm not going to lie. I kind of looked up in ChatGPT a couple of times.
所以对你来说,就像“哦,只是准备。”但
So for you, it's like, "Oh, just as prep." But
没关系。但是,是的。
No worries. But yeah.
很高兴你来做客。我觉得你一直在研究品味方面引领领域,看到你的工作很棒。所以希望这很有趣。
It was It was great having you. I feel like you're always leading the field with a lot of research taste as well, and it's great seeing your work. So hopefully It was fun.
是的,非常有趣。非常感谢。
Yeah, a lot of fun. Thank you so much.