从特种部队到 AI:Dan Biderman 的创业之路

From Special Forces to AI: Dan Biderman's Journey

丹·比德曼 Dan Biderman · Latent Space · 2026-07-13 · 约 50 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Engram 联合创始人兼 CEO Dan Biderman 一边烹饪地中海风味牛肉丸,一边分享他从以色列海军特种部队到创立 AI 研究公司的历程。

Dan Biderman, co-founder and CEO of Engram, shares his journey from Israeli naval special operations to founding an AI research company, while cooking Mediterranean beef meatballs.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 11)

全文 · Full transcript(中英对照)

开场与介绍 Opening and introduction

Host

我现在太投入烹饪了。抱歉,我有点手忙脚乱。整场节目我都假装自己是烹饪专家,但这简直太疯狂了。嘿,各位,欢迎来到 Latent Space 烹饪秀,我们邀请创始人和研究人员,让他们一展厨艺。今天,我们有一位非常特别的嘉宾,Engram 的联合创始人兼 CEO Dan Biderman。

I'm too absorbed in the cooking now. I'm sorry, I'm fried up. The entire show I pretended to be the cooking expert here. But it has to be pure pure crazy. Hey guys, welcome to the latent space cooking show where we invite founders and researchers and let them cook. Today, we have a very special guest, co-founder and CEO Dan Biderman of Engram.

Dan Biderman

很高兴来到这里。

Nice to be here.

Host

是的,感谢你的到来。我听说我们是通过 Jack 认识的,他说你厨艺很棒。再次祝贺你完成 9800 万美元的种子轮融资。你经常做饭吗?频率如何?我猜你一直在研究和构建东西。

Yes, thank you for coming. I heard we got connected through Jack who was saying that you're a very good cook. And again, congrats on the $98 million seed round. Do you cook often? How frequent is it? I assume you're constantly researching and building.

Dan Biderman

是的,自从创办公司以来,这种烹饪确实少了一些。但烹饪其他东西还是有的。烹饪是我从小就开始做的事情。我的父母做饭,我的祖父母也做饭,这对我来说是一种放松、消化和思考问题的方式。

Yeah, so I would say since I started the company, there's a little bit less cooking of this type. Cooking other things, but cooking is something I've been doing since I was a kid. My parents cook, my grandparents used to cook and it's a way for me to kind of relax and digest things and think things through.

Host

太好了。嗯,我很兴奋。今天有点不同,我们实际上要让 Dan 带领我们完成一道菜谱。你想透露一下我们在做什么吗?

Great. Yeah, well I'm very excited. Today is a little different where we're actually going to have Dan lead us through one of the recipes. Do you want to reveal to us what we're making?

Dan Biderman

是的,我想我们可以做牛肉丸配一些蔬菜,让它们浸在某种白酱里,加白葡萄酒和鸡汤。这是一种相当不错的地中海式肉丸,我最近在特拉维夫看望父母时,用家里的食材偶然发现的。所以我想和你们一起重现这道菜,看看感觉如何。

Yeah, so I was thinking we can make beef meatballs with some vegetables and have them swimming in some kind of white sauce with white wine and chicken stock. It's a pretty nice Mediterranean type meatballs that I've kind of stumbled upon recently when visiting my parents in Tel Aviv and cooking stuff with what they had at home. So I thought it would be nice to redo this with you and see how it feels.

Host

太好了。我很兴奋。你想开始吗?

Great. Well, I'm very excited. Do you want to kick us off?

Dan Biderman

我们要炒洋葱。我来切洋葱,再切点大蒜,然后做成丸子,再煎一下。

We're going to fry the onion. I'm going to chop it and going to chop some garlic and then we'll make them into balls and fry them up.

Host

煎一下?太好了。好,那我们开始吧。我们俩同时进行。不过,我想先问一下,我对你的背景很好奇。我记得你想成为教授,而且你还有以色列军队或特种部队的经历?

Fry them? Great. Okay, well let's get started. We'll both do it in parallel. But yeah, I guess to start, I'm kind of curious about your background. So you wanted to become a professor if I recall correctly and you also have experience in the Israeli military or special forces?

Dan Biderman

是的。

Yeah.

Host

这怎么让你变成了一个人工智能研究公司的创始人?

How did that turn you into becoming a founder of a research AI research company?

Dan Biderman

是的,我想事后看来很容易把点连起来,但当时并非如此。我从小并没有计划成为旧金山的一位创始人,研究人工智能这个热门话题。

Yeah, so I guess it's easy to connect the dots looking backward, but in real time it was not. I didn't grow up planning to be this founder in SF working on this hot topic in AI.

Host

嗯。

Mhm.

Dan Biderman

事情就这样发生了。我在特拉维夫长大,正如你所说,我曾是一名军官,从事特种作战,特别是海军特种作战,这非常有创业精神。基本上就是发现和识别我们可以做的疯狂事情,并向人们证明我们应该扩大资源和人员来实际执行。所以这需要大量的研究、大量的推销,然后大量的努力才能把想法变成现实。

Things came to it. So I grew up in Tel Aviv and as you said was an officer and working on special operations, specifically naval special operations, which was very entrepreneurial. It was basically finding and identifying crazy things we could do and prove to people that we should scale resources and people to actually do them. So it was a lot of researching, a lot of pitching, and then a lot of hard work to make things real.

Host

是的。

Yeah.

Dan Biderman

之后,我上了大学,在以色列学习认知神经科学,这是一个很有趣的项目,实际上很多以色列小说家、教授、科学家,甚至厨师,比如以色列厨师 Yotam Ottolenghi,都比我早很多年参加过这个项目。我去了那里,对大脑、人类行为、统计学产生了浓厚的兴趣,然后我去了纽约攻读计算神经科学博士学位。这个领域是一个固执的社区,在神经网络成为硅谷热门话题之前的几十年里,他们一直关注神经网络。我在那里与一些统计学家、物理学家和生物学家一起研究,试图理解如何利用神经网络来理解大脑和动物行为。在看到 ChatGPT 后,我决定更深入地研究大语言模型领域。于是我开始在 Mosaic 工作,在那里我研究了 LoRA,并结识了很多优秀的人。然后从那里开始,我与斯坦福大学的 Chris Ray 和 Scott Linderman 合作,他们俩都是我目前公司的联合创始人,在那里我遇到了我的联合创始人 Sabri 以及我的其他联合创始人 Jack 和 Jesse,他们在康奈尔大学和伯克利大学的实验室研究非常相似的主题。

And then after I finished this I went to university and studied cognitive neuroscience in Israel in an interesting program where actually a lot of Israeli novelists and professors and scientists and even chefs, Yotam Ottolenghi, the Israeli chef, went to this program many years before me. And so I went there, got really curious about the brain, curious about human behavior, curious about statistics and then I went to New York to do a PhD in computational neuroscience which is a community that's been one of those stubborn communities that cared about neural networks for many decades before it was a hot topic here in the valley. So I was there and studied with some statisticians and physicists and biologists trying to understand how neural networks can be used to understand the brain, to understand animal behavior. And then decided that I wanted to go deeper into the LLM world after seeing ChatGPT. So, started working at Mosaic where I worked on LoRA and met a bunch of great people. And then from there went to work with Chris Ray at Stanford and Scott Linderman who both are co-founders of my current company and where I met my co-founder Sabri and my other co-founders Jack and Jesse who work on very similar topics from their labs in Cornell and Berkeley.

Host

是的,这看起来非常注重研究,而且充满机缘巧合。稍微回顾一下你的背景,以色列特种部队有什么特别之处吗?因为我想 Wiz 的联合创始人也是前特种部队成员,对吧?

Yeah, no, it seems like a very research-focused and also serendipitous. Going back a little on your background, is there something about the Israeli special forces? Because I think it was the Wiz co-founders, right? Who were also ex-special forces.

Dan Biderman

是的,我们与 Assaf、Wiz 的 CEO Eran Rabinovitch 合作非常密切,他是我的导师。那里确实有一些有趣的事情。我想最重要的是,当你 18 岁时,你要加入情报部门或承担责任。有趣的是,你不仅仅处于学生模式,不只是埋头做考试和项目。相反,你要与成年人互动,拥有资源,为你的预算之类的事情争论。这迫使你建立一些社交技能和一点成熟度,我从中受益匪浅。你知道,世界上很多国家选择不让每个人都参军是有原因的。这需要很长时间,而且是一个一揽子交易。我在以色列军队服役的年份不同,不是最近的年份。

Yeah, we work pretty closely with Assaf, Eran Rabinovitch, the Wiz CEO, who's a mentor to me. And yeah, there's interesting things going on. There, I would say the most important thing is actually when you're 18 it's joining the intelligence or taking on responsibility. The interesting bit about it is that you're not just in a student mode, you're not just heads down doing exams and projects. Rather, you interact with grown-ups and you have resources and you argue for your budgets and things like that. And it forces you to kind of build a little bit of social skills, a little bit of maturity, which I benefited from. And you know, there's a reason many countries in the world choose not to have everyone go to the army. It's a long time and it's a package deal and I was in the Israeli army in different years, not the ones that are the recent ones.

Host

明白。

Got you.

Dan Biderman

所以,我并不是建议每个人都去参军,但我想说的是,它会迫使你思考自己个性的其他方面,而不仅仅是智力方面。而且总的来说,我认为以色列文化是一个你可以多次获得机会的地方。如果你在高中不是最优秀的,你仍然可能在军队中获得一个好职位。如果你真的很优秀,大学会为你打开更多的大门。即使你在军队中没有做过什么特别有趣的事情,上了大学后发现自己对科学和工程有深厚而良好的倾向,你也可以成为最成功的人,我们有一些以色列人就是各个学术和技术领域的领导者。

So, it's not like I recommend everyone goes to the army, but I would say that it forces you to kind of think about other aspects of your personality, not just the intellectual ones. And also generally, I would say Israel as a culture is a place where you can basically get multiple shots at goal. If you're not the best in high school, you still might have a good position in the military. And if you're really good, more doors open up for you for university. And even if you didn't do anything super interesting in the military and got to university and found out you have good and deep inclinations in science and engineering, you can be the most successful person and we have some of those Israelis who are leaders in various academic and technological places.

Host

嗯,那太好了。

Well, that's great.

Dan Biderman

所以,我想说,是的,基本上就像发牌多次一样,更强调社交方面,更强调成熟和团队合作。但同样,这也是有代价的,你开始做事情时年龄更大。所以我 24 岁才开始上大学,这在美国大学里相当于博士预科的中期。但就是这样。

So, I would say yeah, it's mostly like you can basically the cards are dealt multiple times, more emphasis on social aspects, more emphasis on being mature and being a team player. But again, it comes at a price that you get older when you get to things. So, I started college when I was 24, which is middle pre-PhD in an American university. But yeah.

Host

这都说得通。我认为它给了你很多机会,这是一种非常美好的文化。现在快进到今天,回到 Ngram。

That all makes sense. I think it's great how it gives you a lot of shots on goal and is a very beautiful culture. And now fast forwarding to today, back to Ngram.

动机与背景 Motivation and background

Host

是什么促使你关注上下文这个问题?是特定的突破吗?是开源模型的表现?还是上下文欺诈、长周期智能体的问题?或者只是你个人对这个领域充满热情?

What kind of prompted you to focus on this issue with context? Were there specific breakthroughs? Was it the performance of open-source models? Or even just the issue of context fraud, long horizon agents? Or was it just that you're very passionate about this area as a whole?

Dan Biderman

好,我马上回答。我们能先热一点油来煎……

Yeah. I'll answer that in a second. Can we heat up some oil here to fry up the...

Host

好的,我们开始吧。可以打开电磁炉。

Yes, let's start getting going. So, we can turn on the impulse stove.

Dan Biderman

太好了。我的博士研究方向是半监督学习,这个领域现在不太流行,但我认为它会再次变得至关重要。

Great. So, yeah, my PhD was focusing on a field that's not super in vogue today, but I think will become super crucial again, which is semi-supervised learning.

Host

嗯,好的。

Mhm. Okay.

Dan Biderman

你只有很少的数据,没有无限的万亿 token 数据。你只有一小部分样本,却要从中高效地外推和学习,实现更通用的能力。所以核心就是数据效率。我整个学术生涯的所有论文几乎都有同一个图:x 轴是某种成本或资源,y 轴是准确率,类似帕累托曲线。这就是我的成长背景——效率,如何用更少的资源做更多的事。

You have very little data. You don't have infinite trillion tokens of data. You have a very small set of examples. And from them you want to extrapolate and learn efficiently and achieve things more generally. So, how to be data efficient. My whole academic work was basically all of my papers had the same plot where on the x-axis I had some cost, some resource, and on the y-axis I had accuracy, a sort of Pareto curve. And that's been my upbringing, efficiency, how to do more with less.

Host

嗯。

Yeah.

Dan Biderman

后来我去了斯坦福 Chris Ray 的实验室,在那里研究智能体,提出了一些叫“minions”的想法。我们探讨如何让智能体以经济的方式与海量知识库交互,并引入了一些涉及子智能体的模式,这比它流行起来要早一些。

And I went to Chris Ray's lab at Stanford and there worked on agents, on ideas called minions, and asked the question of how can we have agents interface with very large corpora of knowledge in a way that's economical. We introduced some patterns that involve sub-agents a bit before it was popular.

Host

哦,抱歉。我们也可以把这个放这边。

Oh, sorry. We can also put this over here as well.

Dan Biderman

从效率和成本的角度出发,我们开始问自己:是否存在一种更高效的数据交互方式,而不仅仅是智能体编排?我们能否利用训练的魔力,让模型更高效、更快,并且用更少的 token 运行?其中一些想法是由 Jesse、Sabrina 和我并行开发的。共同的主题是:如果你有一个非常大的知识库,让模型提前学习它——让它自己提问、自己测验、尝试解决问题,然后用梯度下降训练它,就像预训练模型一样——那么你就可以创建非常紧凑的上下文表示,之后可以加载。我们称之为“卡带”。这些就像知识胶囊,可以加载到模型或从模型中卸载,是一种描述模型在语料库中世界的脑状态,压缩了 1000 倍。使用它们时,你可以用更少的 token,更少困惑,更准确。

And from that perspective of efficiency and of cost, we started asking ourselves: is there a more efficient way to interface with data that's not just agentic orchestration? Can we actually use the magic of training to have models that are efficient, faster, and can operate with fewer tokens? Some of these ideas were developed in parallel by some of us, Jesse, Sabrina, myself. The common theme was that if you have a very large corpus of knowledge and you allow the model time to study it in advance, to basically ask itself questions, give itself quizzes, try to solve problems, and then train it using gradient descent in the same way you would pre-train a model, then you can create very compact representations of your context that can later be loaded. We call those cartridges. These are like capsules of knowledge you can load in and out of the model, like a brain state that describes the model's world in the corpus in a way that's 1,000x more compressed. When you use them, you can basically use far fewer tokens, be less confused, and more accurate.

Host

是任务特定的卡带吗?

Are task-specific cartridges?

Dan Biderman

它们可以是语料库特定的,比如公司内部文档;也可以是任务特定的,比如你想学习某项技能。但关键是,在很多情况下,超越文本表示是有意义的。比如我们现在做饭,这里有 Sam Sifton 和 French Laundry 的烹饪书,都是很棒的书籍,我十分尊重。但如果说因为我们有书、能读,就声称我们和 Sam Sifton 或 Thomas Keller 一样精通烹饪,那是不对的。我们可以读这些书,机械地重复所有步骤。当前的 LLM 就像每次都是第一次进厨房,读着教科书做菜,测量所有东西,但它们没有厨师那种捏盐、揉面的直觉。我们通过训练和创建这些卡带所追求的,正是模型中的这种直觉——超越笔记和食谱,达到那种能让你想出下一个菜谱、做出从未有过的创新和推演的智能。

So they can be corpus-specific like documents inside of a company. They can be task-specific like you want to learn a certain skill. But the idea is that in many cases it makes sense to go beyond textual representations. So when we're cooking now and we have all these great textbooks or cookbooks here, Sam Sifton and a book from French Laundry. These are amazing books that I respect, but to say that because we have the books and we can read them, to say that we deeply know how to cook to the same level of Sam Sifton or Thomas Keller, it's incorrect. We can read those things and we can repeat all their steps in a very robotic way. Current LLMs are like coming into the kitchen first time every time, reading the textbook, cooking the dish, measuring everything, but they don't have the intuition of a chef that's pinching salt and kneading dough. So the kind of thing we're after with this training and creating those cartridges is this kind of intuition in the models that goes beyond notes and recipes to the kind of intelligence that allows you to come up with the next recipe, something that hasn't been explored before, to do the next move and the next extrapolation.

Host

嗯。我想深入探讨一下构建这种直觉。你能更具体地说明它和提取有何不同吗?比如用烹饪的例子,如果你从烹饪书中提取所有有用的笔记和章节,帮助你理解做一道菜的所有复杂性。仅仅提供这些,我想那类似于 RAG,只获取那些片段,然后在上下文中理解并输出。

Yeah. I guess double-clicking on building this intuition, could you kind of contextualize it more on how it would be different from extracting, let's say, using the cooking example, if you get all the useful notes and useful sections from a cookbook that will help you actually understand all the complexities of making a dish. Just providing it, I guess that would be similar to RAG, getting just those chunks and then understanding that, I guess, in context and then providing an output.

Dan Biderman

是的,这是一个很好的方法。我们并不认为不需要做笔记。世界上所有最伟大的厨师都有笔记、书籍和日记,记录实验中什么有效、什么无效。但他们也有大脑、双手和舌头,能提醒他们某道菜的味道如何,什么有效、什么失败,什么容易、什么困难。所以在我们的所有工作中,我们从未说过文本表示无用。我们一直在使用它们,构建那些 wiki 和知识库。同时,我们说的是,在这些之上的一层——以参数数量形式存在的直觉和学习层——才是人类厨师的完整经验。世界上最好的厨师,是所有笔记加上一个能阅读这些笔记、执行并创新、且不必反复查阅相同笔记的神经系统。有些菜他们不需要重写。当前的问题是……所以我说,我们想要两全其美。每个知识工作者,如果不能写笔记、不能记录当天的事件,就会处于劣势;但如果每晚都清空他们的大脑,他们也会处于严重劣势。所以我们想要两全其美。文本表示可以带你走很远。

Yeah. So, this is an excellent method. We do not take the bet that no notes need to be taken. All the greatest chefs in the world, they have notes and they have books and they have diaries and they document what worked and what didn't work with their experiments. But they also have brains, they also have hands and a tongue that can remind them how something tastes and what worked and what failed and what was easy and what was hard. So in all of our work, we never say that textual representations are useless. We basically use them all the time and we construct those wikis and knowledge bases. At the same time, what we say is that the layer above those, the layer of intuition and learning that is in the form of numbers of parameters, is the full experience of the human chef. The best chef in the world is all the notes combined with a nervous system that reads those notes and can implement them and innovate them and maybe don't return to the same notes over and over again. There's some dishes where they don't need to rewrite them. And the current problem is... So I was saying that we want the best of both worlds. Every knowledge worker, if they can't write notes and they cannot document the events of the day, they would be at a disadvantage, but if you wipe their brain every evening, they would also be at a severe disadvantage. So we want to have the best of both worlds. The thing is that textual representations can take you a long way.

18个月知识规模 Knowledge Scale in 18 Months

Dan Biderman

但我们在思考的是,看看现在知识被创造的速度——智能体代表知识工作者工作,生成制品、代码、文档、演示文稿。我认为人们还没有完全理解 18 个月后他们将面对的知识工作空间有多大。18 个月后,许多公司可能会拥有数万亿 token 的内部公司数据、专有数据。我说的是数万亿。这听起来夸张,但我认为如果它们真正是 AI 原生,这并非不可能。

But the thing we're thinking about is, look at the rate at which knowledge is being created now with agents working on behalf of knowledge workers, creating artifacts, code, documents, presentations. I think that people don't fully comprehend the size of the knowledge workspaces they will deal with in 18 months. In 18 months, many companies would have maybe trillions of tokens of internal company data, proprietary data. I'm talking about maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.

Host

数万亿 token 是什么概念?就像我两三年前在 Mosaic 时,我们称之为互联网规模的数据、预训练数据。所以想象一下,每家公司都有基本上达到互联网规模的数据。

And what are trillions of tokens? Like when I was at Mosaic 2-3 years ago, we called this internet-scale data, pre-trained data. So imagine every company has data that is basically internet-scale data.

Dan Biderman

是的,也许今天有些公司还没有达到这个规模。也许今天你可以用文本表示走得很远。它们也是可解释的,非常好。但我只是认为,在某个规模上,即使这些文本表示也会变得难以构建。如果你有数万亿 token,你如何精确地创建一个 wiki 或索引,并始终保持更新?你创建的知识库会有多大?用对你公司一无所知的前沿模型来处理它会有多昂贵?

Yeah, and maybe today some companies don't have it. And maybe today you can go relatively far with textual representations. They're also interpretable. They're very good. But I just think at a certain scale, even those textual representations will be hard to make. If you have trillion tokens, how do you create exactly a wiki or an index of those that you keep updated all the time? How big is this knowledge base that you create? How expensive will it be to process it with frontier models that know nothing about your company?

Host

那么,使用前沿模型会很昂贵,是因为每次都要从头开始吗?这是主要问题吗?

So, is it just that using frontier models will be expensive because you start from scratch every time? Is that the main issue?

Dan Biderman

昂贵的一个因素是,你重读更多内容,消耗更多 token。这是其一。但第二点是,对于 18 个月后在这些大型知识库中的智能体式任务,以不明确的方式向模型提问越来越多的问题,我怀疑模型的准确性会下降。上下文欺诈现象,对吧?模型必须读更多,它会变得更不准确。而且我们知道,即使在 1000 万上下文窗口规模下,这种情况也会持续。

There's the element of expensive because you reread more things, you consume more tokens. That's one. But two is, for the agentic tasks of 18 months from now inside those major repositories of knowledge, asking the models more and more things in underspecified ways, I suspect that the accuracy of the models would go down. The phenomenon of context fraud, right? The model has to read more, it will be less accurate. And we know this will remain the same even at the 10 million context window scale.

长周期智能体与上下文限制 Long Horizon Agents and Context Limitations

Host

好的,有道理。我们赶紧去洗个手,马上回来。好了,我们回来了。手都干净了。那么,下一步是什么?就炸肉丸吗?

Okay, that makes sense. Let's go wash our hands real quick and we'll be right back. Great, and we're back. Hands are all clean. So, now what's the next thing? Just fry the meatballs?

Dan Biderman

是的,我们来炸这些肉丸。

Yeah, let's fry those meatballs.

Host

好的。你想打开这个吗?按一下然后旋转就行,很直观。我们倒点油。好的,开始吧。我们可以把这些放进去,同时,我想更多谈谈关于长周期智能体的问题,它试图解决的主要问题是什么?所以,现在如果我以最有力的方式呈现反对观点,也许公司内部数据实际上并没有大到需要将其放入权重的程度。比如,仅仅使用 RAG 或特定模型,甚至更便宜的模型(因为开源模型在处理许多任务时也非常高效)有什么问题?

Okay. Do you want to turn this on? You just press it and then turn it. Very intuitive. And we'll put some oil. Okay. Let's go. We can put these in and while we do that, I guess, more on the question of long horizon agents, what's the main issue that is trying to solve? So, right now if I were to steel man the opposing view, maybe internal company data isn't actually big enough where I'd need to put it into the weights. Like, what's the issue with just having RAG or having specific models or even cheaper models since open source models are also very performant to handle a lot of tasks.

Dan Biderman

我会说,研究界有一个棘手的问题:你能想出一个例子,其中只有权重内训练有效,而上下文学习会失败吗?结果发现,很难设计出这样的例子。对于你给出的每个例子,都会有人问:“那么,如果下一个寓言有 1000 万上下文窗口呢?”在我看来,根据我的科学训练,我认为所有这些关于持续学习和记忆的问题,本质上都是长上下文问题的伪装。如果模型能看到整个公司的数据,并且原则上拥有无限的上下文窗口,那么限制是什么?限制有两个方面。一是,即使在非常小的规模上,我们也知道,你给模型提供的上下文越多,它就越困惑。这被称为上下文腐烂。所以你可以向模型输入一定数量的 token 而不出错,但这并不意味着模型能够以整体方式推理它们。这是其一。

I would say a thorny question in the research community is: can you come up with an example where only in-weights training would work where in-context learning will fail? And it turns out it's very hard to devise such examples. For every example you give, someone can ask, 'Well, what happens if the next fable has a 10 million context window?' The way I see it, in my scientific upbringing, I see all of these questions of continual learning and memory as questions of long context in disguise. If the models could see a whole company's data and in principle would have this infinite context window, what then is the limitation? The limitation is twofold. One is that we know even at very small scales that the more context you feed to the model, the more confused it gets. It's called context rot. So you can feed in a certain number of tokens into the model and not get an error, but that doesn't mean the model can reason in a holistic way about them. That's one thing.

Host

那么压缩的问题是什么?

And what's the problem with compaction?

Dan Biderman

压缩技术每天都在进步。对于不了解的听众,压缩是指模型自己管理上下文,驱逐某些 token 并保留其他 token。压缩是有效的。但同样,当你进入更长的周期时,压缩本质上是有损的。你丢弃一些,保留另一些。我认为这是一个正确的方法。但它也非常确定:要么进,要么出。在当前版本的压缩中,当会话进行得很深时,你可能会感到困惑,也可能会遗忘。所以,我们认为压缩将是故事的一部分。我们认为故事的另一部分是某种神经记忆痕迹,它也是有损的。它驱逐一些,保留一些,但不是在文本表示中,而是在权重表示中。

Compaction is improving by the day. For those in the audience who don't know what it is, models actually managing their own context, evicting certain tokens and keeping others. Compaction works. That too, when you go into a longer horizon, compaction by definition is lossy. You discard some and keep other things. I think it's a correct way to go. But it's also very deterministic. Either you're in or you're out. In the current versions of compaction, when very deep into the session, you can get confused and you can get forgetful. So we think compaction will be part of the story. We think another part of the story is some sort of neural memory trace, which too is a lossy thing. It evicts some and keeps some, but not in the text representation, in the weights representation.

持续学习的重要性 Why continual learning matters

Dan Biderman

但我想说,我们试图解决的主要问题,或者说需要持续学习的主要原因,是 token 效率和成本。这在我们去年底创办公司时还不是问题,但现在变得更加紧迫。第二个问题是,如果你能用更少的资源做同样的事情,当扩展到非常大的资源时,突然就能承担以前不可能完成的任务。

But I would say the main thing we're trying to solve, or the main thing for which continual learning is needed, is token efficiency and cost, which is a major issue that wasn't actually an issue when we started the company late last year and became more urgent now. And problem number two is if you can do the same thing with fewer resources, when you scale up to very large resources, suddenly you can take on tasks that were previously not possible.

Host

嗯。

Yeah.

Dan Biderman

更长的视野,更强的适应性,我们还没完全做到。我们现在专注于第一个部分:让这些模型用更少的 token 在大型上下文中推理,并且减少混淆。但我们认为最终,解决科学、工程、国防等领域非常困难的任务,将涉及在长周期任务中进行某种形式的基于梯度的更新。有人称之为测试时计算或测试时训练,只是同一事物的不同名称。我认为我们从预训练中已经有一个存在性证明:你可以非常高效地将大量信息压缩到极少的数字中。我们喜欢举的例子是,如果你拿一个 Llama 70B 模型,加载一篇来自维基百科的文章,只有几十 KB,让模型阅读它,模型在阅读这几十 KB 时的脑状态在 GPU 的 HBM 上大约是 80 GB。这太疯狂了。而整个模型的参数在 B16 精度下大约是 140 GB。所以那大约 100 GB 的权重(有些失真)代表了整个互联网。而这篇关于 Taylor Swift 的文章在 GPU 上消耗的内存数量级相同。所以这非常低效。这是一个系统问题。这是 KV 缓存这个怪物,世界上最聪明的人正试图从芯片层面以及软件和内核层面解决它。但从系统角度来看,我们正在做的事情的另一种看法是,如果更技术地说,我们不是在模型反复读取语料库时做那些预填充,而是某种程度上摧毁预填充。我们在其他时间扩展训练算力,这样我们可以将内容加载到模型中,然后立即开始解码或少量预填充。这与数据中心的建设趋势相辅相成,即分离预填充和解码,并在不同的专用卡上执行。这也是我们最初对此感兴趣的部分原因。

Way more long horizon, way more adaptive, and we're not quite there yet. We're focusing right now on the first component, which is getting these models to reason on large context with fewer tokens and do it in a way that's less confused. But we think that eventually, part of the solution for very hard tasks in science, engineering, defense, and all that stuff will involve some form of gradient-based updates during these long horizon tasks. Some people call this test-time compute or test-time training; these are just different names for the same thing. I think we have a proof of existence from pre-training that you can pack a lot of information in very few numbers very efficiently. So the examples we like to give is that if you take a Llama 70B model and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this, the brain state of the model when reading this few tens of kilobytes is like 80 gigabytes on the HBM of the GPU. An insane amount. And the entire set of parameters of this model would be like 140 or so gigabytes at B16. So those hundred or so gigabytes with some distortion represent the entire internet. And this one article about Taylor Swift is like the same order of magnitude memory consumption on the GPU. So it's highly memory inefficient. That's a systems problem. That's the KV cache monstrosity that the smartest people in the world are trying to solve from the chip side and from the software and kernel side. But yeah, so I would say another way to look at what we're doing from that angle, from the systems angle, is basically if we can get a bit more technical, instead of doing those prefills where the model is reading and reading a corpus, we are kind of destroying prefill. We're scaling training compute in some other times, so we can load the thing into the model and can just immediately start decoding or prefill a little bit. And this kind of goes hand in hand with trends in how data centers are built out, disaggregating prefill and decode, and doing this on different specialized cards. And this is part of our initial interest in this thing.

Host

好的,看来我们这里基本准备好了。我觉得是时候喝点白葡萄酒了。嗯,也许我们该这么做。是的。那么,当您谈到产品时,有什么具体的例子吗?

Okay, so it seems like we're mostly ready in here. I think it's a good time for our white wine on both of them. Yeah, maybe we want to do that. Yeah. And do you have any tangible examples on this when you talk about your product?

Dan Biderman

所以,如果你想想,比如一家企业,或者投资银行、律师事务所。他们可能有很多客户事务,很多……

So, if you think about, for example, an enterprise firm, or investment banking, a law firm, or investment banking. They can have many client matters, many...

Host

就像大量的知识工作。

Like a lot of knowledge work.

Dan Biderman

很多。是的,很多客户事务。他们有客户,客户做融资、并购之类的事情,还有贷款、交易。

A lot. Yeah, so many client matters. They have clients, clients do financing, mergers and acquisitions, and things like this, and take loans, and do deals.

Host

要戴上这个吗?还是先不?

Want to put this on or not yet?

Dan Biderman

不,先不。我们想让它……

No, not yet. We want to let it...

Host

让它稍微降一点。

Let it reduce a little.

Dan Biderman

好的。

Okay.

Host

嗯。

Yeah.

Dan Biderman

所以,例如,这就是我们与 Harvey 合作处理的事情,非常大的文件系统。智能体可能会遇到很多查询,要么是人类问它们的,要么是智能体必须解决的,这些是那种不易通过 RAG 搜索的隐性难题。比如,如果你想问“我们今年还有哪些并购交易没完成?”要真正解决这个问题,你必须逐个客户事务去查。

And so, for example, these are the kind of things we work with Harvey, very large file systems. And there are many queries that agents might run into, either humans ask them or the agents have to solve them, which are these kinds of ambient hard questions that are not easily searchable with RAG. For example, if you want to ask like 'Which M&A deals haven't we completed this year?' So to actually solve this problem, you have to go client matter by client matter.

Host

嗯。

Yeah.

Dan Biderman

阅读所有文件。你无法在任何一个地方读到“未完成”。你必须抓住要点。

Read all the files. You can't read in any place that it was not completed. You have to take the gist.

Host

你必须理解……

You have to understand...

Dan Biderman

事情还没结束,还没完成。现在你可以用前沿模型和压缩来解决这些任务。但当你让它们这样做时,它们会消耗数千美元,去回答我们认为每个员工都能轻松回答的简单问题。所以这只是一个例子,但这类整体性问题,你读遍所有东西才能抓住要点,但无法在一处找到答案。这种整体大于部分之和的查询。这就是训练的魔力所在,也是 Ilya 等人通过预训练向我们展示的魔力,对吧?你在这些权重中学习了整个互联网,然后模型突然就能推理、泛化、内插和外推到新事物,这就是我们想要的知识。我们训练整个互联网,而不是仅仅对整个互联网做 RAG 或把它放进目录然后实时读取,这并非巧合,因为我们确实认为从大量知识中学习能以某种方式在……中创建这些关联。

Thing hasn't been closed. Thing hasn't been completed. And now you can solve these tasks with frontier models and compaction. And when you ask them to do so, they will consume thousands of dollars for queries that we think are harmless that every employee in the company would be able to answer. So this is just one example, but these kinds of holistic things where you can get the gist if you read everything, but you can't find the thing in one thing. The kind of queries where the whole is greater than the sum of its parts. So that's where this kind of magic of training comes in, and that's the magic that Ilya and others have shown us with pre-training, right? You learn the entire web in these sets of weights and suddenly the model can infer things, can generalize, can interpolate and extrapolate to new things, and that's the kind of knowledge we want to do. And it's not a coincidence that we're training on the entire web and we're not just doing RAG over the entire web or putting it in the catalog and reading it in time, because we do think that learning from a lot of knowledge somehow creates these associations in the...

Host

明白了。

Got you.

Dan Biderman

所以,一部分是当下的具体问题。另一部分则是赌注:18 个月后,数据的规模将需要我们从预训练中已知的方法。

So, some of it is concrete problems of now. Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.

Host

明白了。那么,你们是否在针对特定公司做参数高效的微调,比如基于语料库的 LoRA?

Got you. And so, are you doing parameter-efficient fine-tuning for specific companies like LoRA based off of the corpus?

Dan Biderman

所以我们的长期目标是……

So our ambition in the long term...

Host

还有类似的东西。

And also stuff like that.

Dan Biderman

是的。所以我们的长期目标是,每个人都拥有一个模型,或模型的一部分,或一组权重,代表他们的知识和专长,从他们身上学习;他们与模型相处的时间越长,模型对他们就越好;他们给模型的数据越多,模型对他们就越好。

Yeah. So our ambition in the long term is that every person has a model, or a part of the model, or a set of weights that represents their knowledge, their expertise, learns from them, that the more time they spend with the model, the better it gets for them, the more data they give the model, the better the model is for them.

Host

嗯。

Mhm.

Dan Biderman

是的。

Yes.

Host

嗯。

Mhm.

Dan Biderman

他们控制那些权重。那是他们的。

They control those sets of weights. It's theirs.

Host

还有权利。

And the rights.

Dan Biderman

而且他们有动力去……现在让它沸腾吧。

And they're incentivized to... Let's let it boil over now.

Host

嗯。

Yeah.

Dan Biderman

所以,他们有动力去……我们可以把香料放进去。

So, they're incentivized to... We can put the spices in.

Dan Biderman

所以,他们有动力去让它变得更好,就像你养电子宠物一样。你越用心,它就越开心。这就是我们想用模型创造的东西。所以这是终极的持续学习。我们认为在极端规模上,在单个人类规模上……

So, they're incentivized to make it better in the same way that you would work with a Tamagotchi. The more you nurture it, the happier it is. And that's the kind of thing we would like to create with the model. So that's the ultimate continual learning. And we think on the extreme scale, on the single human scale...

Host

嗯。

Mhm.

基础设施与个性化 Infrastructure and Personalization

Dan Biderman

事实证明,这不仅仅是一个研究问题,还是一个重大的基础设施问题。从长远来看,我确实认为这些东西最终会在人们的设备上运行。

Turns out this is not just a research problem, it's a major infrastructure problem. And in the long, long term, I do think these things will actually run on people's devices.

Host

嗯。

Yeah.

Dan Biderman

我们现在看到,个人电脑的新硬件已经很快接近能够对接近万亿参数的模型进行推理的能力。

And we're seeing right now the new hardware on personal computers is already soon approaching the ability to run inference on close to trillion parameter models.

Host

嗯。

Mhm.

Dan Biderman

这对那种个性化来说会非常有趣。但在短期内,我们确实认为那种包含大量专业知识、非常密集的大型知识库可以在企业中找到。

Which will be very interesting for that kind of personalization. But in the shorter term, we do think that the kind of large corpora of knowledge, very dense with a lot of expertise, can be found in the enterprise.

Host

嗯。

Mhm.

Dan Biderman

那是人们花费大部分时间的地方,也是人工智能实际应用最多的地方。

That's where people spend most of their time. That's where AI is actually being used the most.

Host

对,就是这样。

Yeah, and then that's it.

Dan Biderman

所以我们瞄准那里,那是我们的赌注。在那里,我们的赌注是像 LoRA、cartridges、memory layers 等参数高效微调方法——我们作为研究团队贡献了这些——实际上可以表示知识,并且可以与其他方法结合,比如上下文管理、可追溯的方法,让人们能够使用、理解和审计。

So, we go there and that's our bet. And there, our bet is that parameter efficient fine-tuning methods like LoRA, like cartridges, like memory layers, different things we contributed to as a research team, can actually represent knowledge and can be combined with other methods that are like context management, traceable methods that people can use and understand and audit.

Host

嗯。

Yeah.

Dan Biderman

所以,这是一个组合。

So, it's a combination.

Host

明白了。

Got you.

Dan Biderman

所以,这就像厨师有菜谱、有日记,还有从每次烹饪中学习的神经系统。

So, again, it's like the chef that has the cookbook and it has the recipes, has the diary, it also has the nervous system that learns from every session.

Host

明白了。

Got you.

Dan Biderman

所以,这里正在沸腾,一切都很高效。

So, it's boiling here. It's all really efficient here.

Host

嗯,非常强大的炉子。

Yeah. Very powerful stove.

Dan Biderman

是的,非常好。

Yeah, man. Very good.

Host

然后中火、小火。

And then medium, low.

Dan Biderman

然后这个,也混合一下。对,然后你放一些……

Then one of this one. Also, let's mix it up. Yeah, so then you put some...

Host

嗯,我放点盐。我们要做黄米饭,感谢 Alan。

Yeah, I'll put some salt in. We're going to have yellow rice. Courtesy of Alan.

Dan Biderman

是的。

Yes.

Host

太好了,我们快好了。让它们再煮一会儿。

Great. Well, we're nearly there. Yeah. We'll let these guys cook for a little bit.

内部知识与外部知识 Internal vs External Knowledge

Host

针对你们正在攻克的难题,你们如何确定哪些应该放在权重内部,哪些仍然应该用 RAG 处理,以及如何编排这些?

With what you guys are attacking, how do you determine what should live inside of the weights, what should still be handled with RAG, where things should be orchestrated?

Dan Biderman

是的。所以,我认为这对我们和其他所有人来说都是一个重大的开放性问题。这不仅是人工智能和初创公司中的开放性问题,也是人类记忆研究从一开始就存在的开放性问题。比如,什么样的知识应该内化,什么样的知识应该外化。作为一个人,记住你看到的一切有意义吗?拥有这种能力的人往往并不享受它。

Yeah. So, this is a major open question, I would say, for us and for everyone else. And it's not just an open question in AI and in a startup, it's an open question in the study of human memory from its inception. It's like, what kind of knowledge should be internalized and what kind of knowledge should be externalized. Does it make sense for you to remember everything you've seen as a person? People who have that often are not enjoying that capability.

Host

嗯。

Yeah.

Dan Biderman

而且它可能非常分散注意力,有时甚至很可怕。所以一定程度的遗忘是健康的。有些事情你想按原样记住,有些事情你想记在文本里。所以我认为这是一个开放的研究问题。解决方法是训练模型自己管理。这是我们活跃的研究领域。让模型知道,没有任何明确的监督信号,就能判断哪些东西可以从大脑中提取,哪些东西最好记在笔记里。这与数据中某些部分的显著性、它们重复的频率以及从大脑中知道这个事实后能做什么有关。你可以想象,记住今晚酒店的房间号不如记住伴侣的电话号码或你的地址重要,对吧?

And it can be very distracting. It can be at times very scary. So a certain amount of forgetting is healthy. Certain things you want to remember in the way they are, and certain things you want to put in text. So I would say it's an open research problem. The way to work on it is to train models to manage it themselves. And that's an active area for us. Have the model know, without any explicit supervision signal, to determine this kind of stuff I can pull from my brain, and that kind of stuff I rather keep in notes. Relate to the saliency of certain parts in the data, how often the frequency at which they repeat, and the affordances of what you can do when you know this fact from your brain. You can imagine that remembering the room number in a hotel for tonight is less important than remembering your partner's phone number or remembering your address, right?

Host

嗯,所以那是更高信号的数据。

Yeah. So that's higher signal data.

Dan Biderman

更高信号的数据。但问题是,如果你开始手动启发式地决定哪些放进去、哪些不放,那就成了打地鼠游戏。

Higher signal data. And now the thing is if you start manually heuristically saying this is in, this is out, then it becomes a whack-a-mole.

Host

好的。

Okay.

Dan Biderman

每个人和每个企业都有不同的数据,你很难轻松地挑选哪些放进去、哪些不放。所以圣杯是让模型自己学习。让它带着一个笔记本操作,可以记笔记。让它带着一个大脑——一个关联性的参数高效的东西——可以从中读取。让它决定什么时候使用哪个。并通过无约束的训练来实现这一点。我们正在努力,还需要更多突破,但我认为这就是梦想。模型学会什么该放进去、什么该拿出来,而我们则退到一边。

Every person and every enterprise has different data, and you can't really easily pick and choose what goes in and what goes out. So the holy grail is to have the model learn for itself. Have it operate with a notebook where it can take notes. Have it operate with a brain, associative parameter efficient thing that it can read from. And have it decide when to go to each. And do this with training in an unconstrained way. And we're working on it, and more breakthroughs are needed, but I think that is the dream. That the model learns what comes in and what comes out, and we get out of the way.

Host

明白了。所以那是你希望实现的最终目标,让它完全自主。没有人类介入告诉它该取什么。

Gotcha. Okay. And so that's the end goal you want to achieve where it's completely autonomous. So there's no human in the loop to tell it what to fetch.

Dan Biderman

人类介入可以是用户。如果用户选择说保留这个或排除这个,我们希望模型都听从用户。

The human in the loop can be the user. And if the user chooses to say keep this in or keep this out, we would like them all to listen to the user.

Host

嗯。

Yeah.

Dan Biderman

但我们不想依赖用户,对吧?所以用户反馈是我们学习的东西,我们可以隐式地学习,但我们不希望有人监督每一步,因为那不是人们喜欢使用语言模型的方式。但我要说,我们正在构建的模型,不同于其他模型——你点赞或点踩,基本上是在帮助提供商在下一个版本中给你更可用的东西。

But we would not like to depend on the user, right? So, user feedback is something we can learn from, and we can learn from implicitly, but we don't want a person supervising every step because that's not the way people enjoy using language models. But I would say the kind of models that we're building, unlike other models where you do thumbs up and thumbs down, and you're basically helping the provider maybe in the next version give you something that's more workable.

Host

嗯。

Yeah.

Dan Biderman

在这里,如果你点赞或点踩,或者说了什么,你知道有人会根据你的话扩展算力,有人会去练习以更好地满足你的要求。这就是我们想要达到的目标——与用户建立信任,让他们感到被倾听,并且他们在打造一个更好的模型,不是对每个人都更好,而是对他们更好。

Here, if you give a thumbs up or thumbs down, or you say something, you know that someone's going to scale compute on what you said, and someone's going to go and practice to get better at what you said. And this is kind of the thing we want to get to, like building trust with the user that they're listened to, and they're making a model that's better, and it's not generally better for everyone. It's better for them.

Host

明白了。所以好处是它是一个更紧密的循环,相比之下,像通用模型提供商那样,界面上有点赞点踩,你不知道那是否真的会贡献到反馈中。

Got it. So, the benefit is it being a tighter loop compared to the like if say the general model provider, the UI having a thumbs up, thumbs down, you don't know if that's actually going to contribute to the feedback.

Dan Biderman

是的,循环更紧密,我们使用不同的机制。如果我们允许自己使用训练机制,我们可以把它敲进去,没有不确定性。另一件重要的事情是,用户告诉你的并不都是绝对真理,对吧?不是所有人,包括我自己,都是爱因斯坦,我们可能会对模型说一些我们认为正确而模型错误的话。而且模型会越来越好,它们会知道越来越多我们不知道的事情。所以模型在某种程度上必须学习和理解,并辨别哪些反馈有价值,哪些应该忽略。

Yeah, there's tighter loop, and we use different machinery. If we allow ourselves to use the machinery of training, we can hammer that in. There's no uncertainty about it. And another important thing to say is not everything that a user tells you is ground truth, right? Not all of us, including myself, are Einsteins, and we can say things to the model where we think we're right and the model is wrong. And increasingly, the models will get better, and increasingly, they'll know more things than we do. So, the model in some way has to learn and understand, and kind of discern which feedback is valuable and which feedback should be ignored.

Host

嗯。

Mhm.

Dan Biderman

我认为圣杯是退到一边,如果你定义了正确的训练目标,就让模型自己去学习。

I think the holy grail is to get out of the way and have the model learn it if you define the right objectives for training.

Host

明白了。好的,所以不要妨碍模型。

Got you. Okay, so get out of the model's way.

自主知识工作的挑战 Challenges in autonomous knowledge work

Host

有没有哪些时刻真正向你展示了还需要什么,或者还有什么可能?因为要达到模型能自主处理这一切的最终状态,似乎非常雄心勃勃。而且我认为,多智能体设置,甚至现在,仅仅约束解码仍然存在问题。所以我想,当涉及到知识工作和特定企业时,风险似乎高得多。而且你不能犯太多错误。那么,有没有一些时刻,或者甚至是你现在看到取得巨大进展的研究课题,能够让你实现这一目标?

Have there been any moments that have really kind of shown you what is still needed, or what's still possible? Because it seems very ambitious to be in an end state where a model autonomously can handle this all. And I think like multi-agent setups and even right now, too, just constrained decoding still has problems. And so I guess when you get to knowledge work and to specific enterprises, it seems a lot higher stakes. And you can't make as many mistakes. And so have there been moments or even research topics that you're kind of seeing great progress in right now that can get you there?

Dan Biderman

是的,我想说我们会在未来几周和几个月分享更多结果,我们目前对很多内容保密。但我想说,主题是,你正在看到的,而且我们不是唯一看到这一点的,但我们正在密切关注的是关于 token 效率的行为,即模型能够去到需要去的地方并更快地解决问题的能力。基本上,这是一种智能观,变得更聪明意味着用更少的精力解决越来越难的问题。这就是我们追求的目标。其中很多以效率和速度的形式体现出来,但我们很快就会分享更多。

Yeah, I would say like we will share more of our results in the coming weeks and months and we're kind of keeping a lot of it, yeah. But I would say that the theme is, and the kind of thing you're seeing and we're not the only one seeing it, but I would say we're looking very closely into it, is behaviors around token efficiency, the ability of the models to go where they need to go and solve things faster. And basically it's this kind of view of intelligence where getting smarter means exerting less energy to solve increasingly harder problems. And these are the kind of things we're going for. And a lot of it comes into life in the form of efficiency and speed, but we will share more of that soon.

Host

好的。我想在 token 效率这一点上,你是否也重点考虑模型路由?因为我假设可能有一些更便宜的模型可以以更低的成本完成同样的工作,但有些任务可能用一个更便宜的开源模型需要价值 100 美元的 token,而一个更智能的模型可能一次就能快速完成,成本更低。

Okay. I guess on the token efficiency point, do you also heavily consider like model routing? Because I assume there may be some cheaper models that may do the same job for much less, but there may be some tasks that it may take a much cheaper open-source model like $100 worth of tokens when a much smarter model may be able to do it in one shot very quickly for much cheaper.

Dan Biderman

是的,我想说路由是一个有趣的事情。很多企业和计算机科学家都在研究它,这是有原因的,因为我们有模型对很多事情来说都过于强大了。是的,你不需要 Fable 来告诉你煮米饭要放多少盐。同时,当你试图解决超出你能力范围的问题时,你需要知道何时去求助 Fable。所以我认为路由是一个很好的方向。路由也不容易,作为一个研究过它的人。

Yeah, I would say that routing is an interesting thing. And there's a reason so many enterprises and computer scientists are looking into it, because we have models that are overkill for many things. Yeah, you don't need Fable to tell you how much salt to put in rice. At the same time, you do need to know when to go to Fable when you are trying to crack something that's above your pay grade. So I think routing is a great direction. Routing too, it's not that easy, as someone who's worked on it.

Host

或者,以你看到路由难点的经验来看,主要挑战是什么?

Or what are the main challenges having the experience of seeing the difficulty of routing like?

Dan Biderman

我认为路由肯定会成为解决方案的一部分,而且我认为很多人,不仅仅是我,都说解决方案是多模态的。不是 N-gram 接管,只有一个模型,你教它东西,然后你就可以关闭星门。那不是我们的方法。解决方案将涉及某种形式的路由。在路由的那部分,我们的想法是,就像你最好的朋友或你的员工,你不能解雇一个一直盯着你肩膀看的人,他能告诉你看哪里、关注什么,甚至能以更有针对性的方式、带着更多上下文去问 Fable,然后 Fable 可以花几天时间工作,为你解决一个非常高风险的任务。所以我认为路由很棒。它就像任何其他研究领域一样。这是一个未解决的问题。已经取得了一些进展,但还需要更多工作才能真正以合适的成本在合适的时间将任务路由到合适的模型。而且模型在不断更新。新版本不断推出。所以是的。

I think routing will be part of the solution there for sure, and I think many people, not just myself, say the solution is multimodal. It's not N-gram taking over, there's one model and you teach it things and you can close the stargate. That's not our approach here. The solution will involve some form of routing. And in that part of routing, our thing would be like, you know, your best friend or your employee you cannot fire someone who's been looking over your shoulder the whole time and can tell you where to look, where to focus, and can even go out and ask Fable things in a more targeted way with more context that Fable can then go and work on it for days and solve a very high stakes tasks for you. So I think routing is great. It's just like any other research area. It's an unsolved thing. There's been progress on it but a lot more work is required to actually route things to the right model at the right time at the right cost. And models are continually updated. Versions are coming out. So yeah.

Host

是的。好的,这说得通。似乎还有很多悬而未决的问题,甚至你正在做的一些工作也是保密的,这很合理,而且已经发布得更多了。我想现在转向团队。

Yeah. Okay, that makes sense. It seems like there's still a lot of open questions and even some of the work you're doing is confidential which makes sense and has already been released more. I guess now shifting towards the team.

Host

是的,和这样一群混合背景的人一起工作感觉如何?比如有些是你学习时的教授,有些是你遇到的、已经离开或完成了博士学业的人,比如 Jack。是的,拥有这样一个多元化的群体。

Yeah, how is it like working with such a mixed group of folks like some who had professors while you were doing studying and some that you also met and have left their PhDs or finished their PhDs as well like Jack. And yeah, having such a diverse group.

Dan Biderman

所以,我想说,我们团队有很多优势。我不确定多样性是其中之一。所以,如果你在看这个,我们有很多研究员。我们可以稍微多样化一点。我想说,对于我们要解决的这一细分领域,即记忆和持续学习,将知识塞进权重,我认为我们的团队是这方面最专业的团队。而且我们的团队很有趣。你知道 Jack。Jack 是个有趣的人。Jack 教会了我们很多关于如何清晰思考、如何清晰地向内部和外部世界表达自己的东西。Jessie 也是,方法非常互补。她思考了很多关于人类和 AI 如何在机器和人类的闭环中协作的问题。Sabreena 和我更偏向系统方面。Scott 和我研究统计学。所以,都是研究员,在这方面并不多样,但正如你所说,我们的倾向是多样的。我们中有些人更偏数学,有些人更偏系统,有些人更像是 AI 领导者。所以,这很有趣。我们尝试的方法是,让那些经验丰富的博士与那些加入我们公司的、崭露头角的、厉害的新人配对。例如,来自斯坦福的 Shiraz 和来自伯克利的 Dhruv。他们有研究背景,写过论文。但他们带着很大的势头进入这个领域。我们喜欢把他们与在这个领域待了几年、有一些直觉、能提醒他们避免钻牛角尖的人配对。我认为这种不同专业水平、不同思维新鲜度的强大组合,让我们的地方变得有趣。而且我想说,从 25 年秋冬的第一天起,从文化上讲,那段时间总是前沿实验室招聘和 AGI 焦虑的疯狂时期,我想说我们所有人都以非常清醒的方式进入了创业世界,知道这里不仅仅是一个研究俱乐部。这里不仅仅是一个期刊俱乐部。缺少的是产品,产品需要分发。你必须通过销售人们喜欢的东西来赢得参与的权利。所以我们也在为此非常努力地工作。

So, I would say, there are many strengths to our group. I'm not sure diversity is one of them. So, if you're watching this, we have a lot of researchers. We could diversify a little bit. I would say, for this niche that we're trying to solve, which is memory and continual learning, taking knowledge, shoving it into weights, I think our team is the most specialized team in that kind of thing. And our team is fun. You know Jack. Jack is a fun guy. And Jack is teaching us a lot of things about how to think clearly, communicate ourselves clearly internally and to the external world. Jessie in the same way, very complimentary kind of approaches. She thought a lot about how humans and AI work together in this closed loop of a machine and human. Sabreena and I were kind of bit more on the systems side. And Scott and I worked on statistics. So, it's all researchers and not diverse in this kind of way, but it is, as you said, diverse in our inclinations. Some of us are more mathy, some of us are more systemsy, some of us are more like AI leaders. And so, it's been interesting. The way we try to do this is to take those PhDs with a lot of experience and pair them up with like those kind of up-and-coming, rising, cracked types that have joined our company. For example, Shiraz or Dhruv from Stanford and Berkeley respectively. They have research backgrounds and have written papers. But they're kind of entering this field with a lot of momentum. And we like to pair them up with someone who's been around the field for a few years, has some intuitions, can warn them from the rabbit holes. And this I think powerful combination of different levels of expertise, different levels of freshness of thought is making our place an interesting one. And I would say culturally all of us from day one in the fall winter of '25, which is always a crazy time in terms of frontier lab recruiting and AGI anxiety, I would say all of us entered into the startup world in a very sober way knowing that it's not just a research club in here. It's not just a journal club in here. The thing that's missing is products and products need to be distributed. You have to earn the right to play but by selling things that people love. So we're working very hard on that as well.

联合创始人趣味问答 Fun questions about co-founders

Host

有道理。我想问个更有趣的问题。在创始人中,你觉得谁对食物最有品味?我猜你们有时会点外卖或者出去吃。有没有谁特别突出?

That makes sense. I guess more of a fun question. Out of the founders, who do you think has the best taste in food? I assume you guys sometimes order food or even go out. Are there some that stand out to you?

Dan Biderman

嗯。我会说我的联合创始人 Sabrina,他一直……他热爱美食。他在公司里很有名,是个“海陆大餐”爱好者。

Yeah. I would say my co-founder Sabrina, he's consistently... he loves good food. He's very well known in the company as someone who's the surf and turf guy.

Host

明白了。

Got you.

Dan Biderman

他午餐和晚餐都会吃这个。和他一起用 DoorDash 点餐对我很有启发。我觉得 Jack 和 Jesse 品味也不错。

He will have it for lunch and dinner. Being on DoorDash with him has been an inspiration for me. I would say Jack and Jesse have good taste as well.

Host

品味也不错。

Good taste as well.

Dan Biderman

是啊。和他们在一起很开心。我们现在在一个像公寓的小办公室里,所以经常一起吃午餐和晚餐。有点像一家人,你知道吗?在家里你会见到父母和兄弟姐妹。有时候会觉得太多,但你永远不会忘记那段时光。

Yeah. It's fun to be with them. We're now in a small office like an apartment, so we have all the lunches and often dinners together. It's a bit like a family, you know? At home you see your parents and siblings. Sometimes it's too much, but you're never going to forget that period.

Host

听起来确实很有趣。你给他们做过饭吗?

No, it does sound very fun. Have you cooked for them yet?

Dan Biderman

我做过吗?可能不够多。也许我应该做。也许这个周末我会做点。

Have I? Maybe not enough. Maybe I should. Maybe this weekend I'll cook some.

Host

好的。

Okay.

Dan Biderman

他们在做别的事。他们在研究上“烹饪”,不过是的。

They're cooking other things. They're cooking stuff on the science, but yeah.

Host

没错。他们在研究上“烹饪”。

That's true. They're cooking on the research.

Dan Biderman

是的,在研究上。嗯。

Yeah, on the research. Yeah.

Host

太好了。好了,米饭应该好了。

That's great. Okay, the rice should be done.

Dan Biderman

应该好了。

Should be done.

Host

我们要尝尝吗?看看。第一口。

Do we want to give it a try? Let's see. First bite.

Dan Biderman

不错。挺好的。

It's good. It's nice.

Host

哦,是的。确实,底部更熟一些。你也可以让它再焖一会儿。

Oh, yeah. It is, yeah. The bottom's definitely more cooked. You can also probably let that sit.

Dan Biderman

这些家伙……哦,可能好了。我们可以稍微打开一点。

These guys are... Ooh, probably ready. We can perhaps open it up a little bit.

Host

嗯。

Yeah.

Dan Biderman

让它收收汁。我们可以让那些家伙再收几分钟汁。然后也许我们现在就可以用手撕开那些东西。

Let it concentrate a bit. We can let those guys concentrate for a couple minutes here. And then maybe we can cut those things with our hands right now.

行动号召与招聘 Call to actions and hiring

Host

你们有什么号召行动吗?你们在找特定的人吗,招聘哪类研究员?

Do you have any call to actions? Are you guys looking for specific folks, hiring for types of researchers?

Dan Biderman

嗯,我想说我们正在构建的是那些持续学习的系统。显然,研究方面有很多开放问题:如何在不破坏模型的情况下做到这一点,如何以经济高效的方式做到,从哪些数据学习,等等。但这也是一个极其雄心勃勃的基础设施问题。如果你真的相信未来会有数万亿 token 的数据,甚至个人数据。如果你真的相信我们可以达到为每个人和团队提供参数高效适配器的水平,那么你突然会想到涉及数百万个不同端点的部署,这些端点存储在不同地方,需要高效地从磁盘读取到 HBM,然后在推理时使用,还要交换和更新。如果事情顺利,这个东西将拥有巨大的算力足迹,并带来许多关于系统以及以新方式平衡 AI 工作负载的新问题。所以,我认为那些能享受其中并帮助我们很多的人是那种 LM 性能工程师、研究工程师,他们知道如何让东西“跑得快”。我们有一些这样的人。我们有 Cade Daniel,他曾是 Databricks 的推理负责人之一,也是 vLLM 的核心贡献者之一。我们都有点系统倾向,但我们认为基础设施工程师,那些知道如何工作和构建大型 API 和数据库的人,会在解决其他地方找不到的问题上度过非常愉快的时光。总的来说,我们始终对聪明、有创造力、思维开阔并致力于解决有趣问题的人持开放态度。

Yeah, I would say what we're trying to build, which is those systems that continually learn. Obviously, there are many open problems on the research side. How you do this without destroying the model, how you do it in a cost-efficient way. What data do you learn from? And stuff like that. But it's also an extremely ambitious infrastructure problem. If you truly believe in the possibility that there's going to be trillions of tokens of data, or even personal data. And if you truly believe that we can get to the level where we have those kinds of parameter-efficient adapters for every person and team, you suddenly think about deployments that involve millions of different endpoints stored in different places that need to be efficiently read from disk to HBM. And then use that inference time. And swapped and updated. It's going to be... If things work out for us, this thing will have a massive compute footprint and many new questions on systems and balancing of AI workloads in new ways. So, the kind of people that I think could enjoy them and help us a lot are those one LM kind of performance engineers, research engineers who know how to make things go burr. We have some of them. And we have Cade Daniel who was one of the inference leads at Databricks and one of the core contributors of vLLM. And we are all kind of systems inclined, but we think infrastructure engineers, people who know how to work and build those large APIs and databases, I think would have very fun time working on questions that they can't find in other places. And generally, we are always open to smart and creative people who think out of the box and are committed to working on interesting problems.

Host

太好了。这似乎是一个非常创造性的组合,既有研究人员,也有像你说的关心基础设施的人。

That's great. It seems like a very creative mix of researchers and people who also care about the infrastructure as you said.

Dan Biderman

是的。我还想说点别的……我刚刚在说话的时候就在想,这种训练和持续学习的很多用例涉及……我们再稍微加一点。

Yeah. And I also wanted to say something... I was just thinking about it while I was talking before that a lot of the use cases for this kind of training and continual learning involve... Let's give it a teeny bit more.

Host

那里少一点液体。

A bit less fluids there.

Dan Biderman

好的。好吧,我们再来一次。所以,对我来说,原则是任何形式的效率和智能都不能真正脱钩。有时人们认为,如果你构建更高效的东西能省钱,那么你就不在高端类别,你在做更便宜的产品。而这一点和智能,这完全是错误的,对吧?所以,你能用更少做更多,长期就能解决更宏大的任务。因此,我认为当前 AI Scaling 的范式是用更多做更多,它把我们带到了很远的地方,并且将继续是构建智能的有价值方式。但我认为下一个范式,许多实验室的领导者都看到了,它涉及用更少做更多的元素,以处理更长周期的任务和更困难的问题。所以,我认为超越企业、超越效率,那是我希望去的地方。如果我们解决了现在面临的挑战。

Okay. All right, let's do it again. So, the point for me as the principle is any kind of efficiency and intelligence cannot really be decoupled. Sometimes people think if you're building something that's more efficient that can save you dollars, therefore you're not in the premium category, you're making the cheaper product. And that and intelligence, this is just purely wrong, right? So, the more you can do with less, the more ambitious tasks you can solve longer term. So, I think that the current paradigm of scaling with AI has been doing more with more and it took us extremely far and it will keep being a valuable way to build intelligence. But I think the next paradigm many leaders of the labs are seeing it involves certain element of doing more with less to take on longer horizon tasks and harder problems generally. So, I think going beyond enterprises and going beyond efficiency, that's where I hope to go. If we solve the challenges that we're facing right now.

Host

有道理。我认为两者兼顾非常重要。尤其是像你说的,你想做更宏大的任务。

That makes sense. I think thinking of both is just very important. Especially as you want to, like you said, do more ambitious tasks.

Dan Biderman

是的。

Yeah.

Host

好的。我们要把菠菜拌进去吗?还是让它……

Okay. Should we mix in the spinach? Or just let it...

Dan Biderman

嗯,你可以试着拌一下。试试看。它会蒸软一点。

Yeah, you can try mix it up. Let's try it. It will steam and soften a little bit.

Host

我可以切这个柠檬。嗯。然后直接把柠檬汁挤进去?

I could cut this lemon. Yeah. And then just squeeze the lemon in?

Dan Biderman

嗯。

Yeah.

Host

这是酱汁,对吧?

It's the sauce, right?

Dan Biderman

嗯。

Yeah.

Host

好的,挤这个。这个少一点,但很棒。人们在哪里可以找到你?

Okay, go on squeeze this one. This one has a little less, but amazing. And where can people find you?

Dan Biderman

人们在哪里可以找到我们?

Where can people find us?

Host

嗯。好了,各位。

Yeah. All right, guys.

Dan Biderman

他们可以在 Ingram.com 找到我们。他们可以在 NSF 找到我们。他们可以给我写信,Dan@Ingram.com,讨论事情。嗯,我想说,随着我们扩大公司涉及工程、产品和业务的不同部分,有很多事情要谈,超越了前沿 AI 类型的研究,我们有很多东西要向聪明人学习。

They can find us at Ingram.com. They can find us NSF. They can write to me Dan at Ingram.com to talk about things. And yeah, I would say as we're scaling up different parts of the company that involve engineering, that involve product and business, there's a lot of things to talk about beyond the frontier AI type research and there's a lot for us to learn from smart people.

Host

太好了。太棒了。我很兴奋。我们要尝尝吗?

Great. Awesome. I'm excited. Should we try it?

Dan Biderman

嗯。

Yeah.

一起烹饪 Cooking together

Host

我向 Neo lab 的每位负责人发出挑战,请你们来我在 Noe Valley 的家,和我一起做饭。我相信我们一定能从彼此身上学到很多。

I challenge every leader of the Neo lab to come to my house in Noe Valley and cook things with me. And I'm sure we can learn a lot from each other.

Dan Biderman

好的,干杯。味道不错吧?

All right, cheers. Good?

Host

嗯。你觉得呢?做得不错,对吧?

Mhm. What do you think? It turned out well, no?

Dan Biderman

是的,非常好。哇,真的很好吃。

Yeah, very good. Wow, that's very good.

Host

嗯,我觉得有点波斯风味。配上黄米饭,太棒了。

Um, maybe a little bit Persian, I would say. With yellow rice, so great.

Dan Biderman

我非常喜欢。都很棒。Sean 呢?米饭太棒了。这次的米饭非常非常好吃。其实,我打算录完之后把整份都吃完。

I'm a big fan. All good. Sean? Rice is amazing. Rice is very, very amazing on this one. Actually, I'm just going to finish the whole thing after this.

Host

是啊。

Yeah.

Dan Biderman

嗯。肉丸是不是也稍微软了点?

Mhm. Are the meatballs a bit soft as well?

Host

是的。

Yeah.

Dan Biderman

嗯。

Mhm.

Host

好极了。你们知道吗,整期节目里我都假装自己是烹饪专家。

Sweet. So you guys know, in the entire show I pretended to be the cooking expert here.

Dan Biderman

好吧。

Okay.

Host

这简直太疯狂了。开个玩笑。

It has to be pure crazy. Just kidding.

Dan Biderman

这大概能打八分,九分吧。

This is like an eight out of ten, nine out of ten.

Host

嗯。

Mhm.

Dan Biderman

很好。不过,嗯,我觉得差不多就是这样了。做得挺不错的。我是说,感觉怎么样?有趣吗?你喜欢吗?

Great. But yeah, no, I think that's basically it. It turned out pretty well. I mean, how was it? Was it fun? Did you enjoy it?

Host

这是我做过的最有趣的播客。让人有家的感觉。

The funnest podcast I ever had. Makes you feel at home.

Dan Biderman

确实。

That's true.

Host

聊起天来也更轻松。

Easier to talk about things.

Dan Biderman

再次感谢你们的到来,希望你们玩得开心。

Thank you again for coming, and hopefully it was a fun time.

Host

非常有趣。

It was super fun.

Dan Biderman

我的权重已经更新了。

My weights have been updated.

Host

那就好。谢谢各位。

That's good. Thanks, guys.

互动版:逐字朗读 + 针对本期提问 →