Omnigents: The Meta Harness for AI Agents
打开互动全文版(中英对照 + 朗读 + 问答)→Matei Zaharia 讨论了将编码智能体和自定义智能体融合到名为 Omnigents 的统一平台中,强调了可移植性、安全性和协作的需求。
Matei Zaharia discusses the convergence of coding agents and custom agents into a unified platform called Omnigents, highlighting the need for portability, security, and collaboration.
在进入今天的内容之前,我有个小消息要告诉听众们。谢谢你们。如果不是你们选择点进来收听我们的内容,我们就不可能为大家带来你们如此渴望的 AI 工程科学和娱乐内容。几乎每天都有赞助商找上门,但幸运的是,有足够多的你们订阅了我们,让我们能够无广告地持续运营下去,我们想保持这样。但我只想请大家帮一个忙。你们能做的最有力且完全免费的事情,就是点击那个订阅按钮。这是我唯一会请求你们的事,它对我以及每周努力为大家带来《In Space》的团队来说意义重大。如果你们订阅了,我保证我们会永不停止地让这个节目变得更好。现在让我们开始吧。来自 Databricks 的 Matei Zaharia,欢迎来到《In Space》。
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering science and entertainment contents that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you and it means absolutely everything to me and my team that works so hard to bring the In Space to you each and every week. If you do it, I promise you we'll never stop working to make this show even better. Now let's get into it. Matei Zaharia from Databricks. Welcome to In Space.
谢谢邀请。
Thanks for having us.
非常感谢。
Yeah, thanks so much.
感谢你抽出时间。你们的 Databricks 数据 AI 峰会正在进行。你刚才告诉我,你们举办的第一届峰会只有 50 人参加。
Thanks for taking time out. You have your Databricks Data AI Summit going on. You were just telling me how the first summit that you guys ran was just 50 people.
嗯。是的,我想那是在伯克利的一个小型聚会。我们组织了一些教程,教大家使用 Spark。
Mhm. Yeah, it was a little meet-up at Berkeley, I think. We put together these tutorials and just teach people Spark.
是啊。现在显然,我想 headline 数字是全球 10 万人,3 万人到场。这真是个疯狂的社区。我刚看了主题演讲。Ali 简直……你当时就知道 Ali 会成为一个这么出色的 CEO 吗?他是个很棒的演讲者。你怎么看?
Yeah. You know, obviously now it's like I think the headline number is like 100,000 people around the world, 30,000 in person. It's a crazy community. Well, I mean I just saw the keynote. Ali is just... Did you know that back when, was it obvious that Ali would be such a great CEO? Like he's a great presenter. What do you think?
我的意思是,在我们这群创始人中,很明显他会是最擅长这个的。结果也确实很棒。他在经营公司时涉猎了众多领域。他会深入研究,变得和专家一样博学。即使招不到那个人,他也会学习足够多的财务、销售等知识,然后从那里出发。
I mean, I think among our group of founders, it was clear that he'd be the best at this. And yeah, it turned out great. And he's ramped up on so many topics running a company. He would just go in and study and become as knowledgeable as the experts. Even if you can't hire the person, you know, learn enough about finance and sales and whatever it was. And then go from there.
他显然智商和情商都很高,但今天的 Ali 和 10 年前的 Ali 很不一样。我认为他付出了很多努力才达到这个水平。
I mean, he's obviously very high IQ and very high EQ, but it wasn't like Ali today is quite different from Ali from like 10 years ago. I think he put in a lot of work to get to this point.
是的。对我来说,他最吸引人的地方是他很幽默。而且拿数据、严肃话题和安全之类的东西开玩笑很难。
Yeah. I mean, to me the most appealing thing about him is that he's funny. And it's hard to make jokes about data, about serious topics and security and what have you.
哦,是的。确实如此。
Oh, yeah. That's for sure.
所以,你们发布了一大堆东西。我简单提一下,因为我们不会覆盖所有内容。Omnigents,你的宝贝。El Tap,你的宝贝。你们的 Dream Engine。我们还会谈到 Genie、Customer League。你们收购了 Panther、Open Sharing,还有 Unity AI Gateway。其中很多我认为是 Databricks 会做的事情,是路线图的一部分。你们这个领域的每个人都在做类似的事情。但我认为你们两位可能正在引领这个领域中最独特、最具差异化的两个项目。也许我们先从 Omnigents 开始,然后再深入。
So, you guys launched a whole bunch of things. I'll just briefly name-check the stuff because we're not going to cover everything. Omnigents, your baby. El Tap, your baby. Your Dream Engine. We're also going to cover Genie, cover Customer League. You acquired Panther, Open Sharing, and there's Unity AI Gateway. A lot of these I think are things that you would expect a Databricks to do. It's like part of the roadmap. Everyone in your category has similar things. But I think probably the two of you are leading the two most unique and differentiated initiatives in the landscape. Maybe we'll start with Omnigents and then we'll go into it.
我确实认为很多人都在探索这种元框架的概念。是什么让你想到这个的?
I do think that a lot of people are exploring this sort of meta harness concept. What led you to it?
是的。实际上有几条线汇聚在一起,我认为这是需要新东西的好迹象。一方面,内部有各种编码智能体。我们有一个非常棒的开发基础设施团队。他们构建了一个叫 Isaac 的东西,基本上是 Claude Code 和 Codex 的包装器,让你可以在网页沙箱中或开发机器、笔记本电脑上使用它们。然后他们还在那里添加各种功能,我们看到更高级的工程师正在用大量智能体构建自己的工作流,甚至在上面构建自己的 UI 等。另一方面,我们自己在构建智能体,我们在研究团队(我共同领导)发布了一个叫 Genie 的数据科学智能体。我们还为各种事情构建了很多内部智能体,然后还有所有客户的智能体,它们都遇到了这样的问题:'哦,我需要每隔几个月切换模型和框架。' 此外,如果你不能与他人共享会话、有历史记录、有搜索以及所有这些协作层,智能体就完全没用。我从这两个背景思考了一下,起初人们觉得奇怪,为什么你要把编码智能体和自定义智能体放在同一个东西里,但我说这基本上是同样的问题,你只是想构建能让你交付智能体的东西,如果你关心安全的话可以控制它,并使其可移植。然后我们做了一些实验原型。我们说,是的,实际上我们可以让它工作,然后我们就真正构建了它。
Yeah. There were actually a couple of converging lines, which I think is a good sign that you need something new. So, on the one hand, there's all the coding agents internally. We have a really great dev infra team. They built something called Isaac that's basically a wrapper on Claude Code and Codex and lets you use them either on the web in sandboxes or just on your dev machine or on your laptop or whatever. And then they were adding all kinds of stuff there and we saw all the more advanced engineers were building their own workflows with tons of agents and they were building their own UIs and stuff on top of even on top of that. And then the other one was that us building agents, we shipped this data science agent called Genie on the research team which I co-lead basically. We also build a lot of internal ones for various things and then we have all the customer ones and all of them were running into this thing of like, 'Oh, I need to switch model and harness and so on every few months.' Plus the agent is completely useless if you can't share sessions with someone and have history and have search and all this layer on top of it for collaboration. I thought a bit about it from both contexts and at first people thought it was weird and like why are you doing coding agents and custom agents in the same thing but I said it's basically the same problems and you just want to build the stuff that lets you deliver the agent, maybe control it if you care about security, and make it portable across things. And then we prototyped some things as experiments. We said, yeah, actually we can make it work and then we built it for real.
我在想,这种架构是否与你过去的职业生涯有相似之处。你知道,我总觉得很多东西实际上都可以追溯到操作系统。很多操作系统又追溯到数据库,或者反过来。
I'm wondering if this kind of, let's call it architecture, maps to anything in your careers in the past. You know, I always think about how a lot of things actually just tie back to operating systems. A lot of operating systems tie back to databases or the other way around.
所以我认为它很大程度上与网络协议有关,比如互联网协议。我们还做了数据共享方面的工作,可能大多数观众不知道,除非他们……
So the thing I do think it ties a lot to like network protocols, you know, internet protocol. We also did stuff with data sharing which is probably most viewers won't know unless they...
是的,开放协议是共享的术语。Open sharing。
Yeah, open protocol is the term for sharing. Open sharing.
Open sharing。是的,比如你有一家公司,维护着某种表,比如沃尔玛之类的。他们有库存和每家店的销售数据。然后你还有供应商,他们希望在你需要的时候正好生产更多东西并发货。所以他们希望实时访问你的表。
Open sharing. Yeah, so it's like you have a company, you maintain some kind of table like, let's say Walmart or something. They have the inventory and what's been sold in each store. And then you also have suppliers and they would love to produce more things and ship them exactly the moment you need them. So they would love real-time access to your table.
所以,与其发邮件、传 Excel 表格或打电话,为什么不能实时共享那个表格的视图呢?然后他们查询,与自己的数据关联,再决定发送什么。这就像今天你可能会问:既然我们能这么快地“氛围编程”任何东西,为什么还需要设计协议、API 或软件?为什么不能按需“氛围编程”呢?但实际上,对于这种多方以不同速度构建、仍需要上层协调的互操作性场景,你确实需要设计和构建它。这让我想起智能体之间、用户与智能体以及工具之间的对话。
So instead of sending emails around or Excel sheets or phone calls, why can't you share a view of that table in real-time with them? Then they query, join it with their data, and decide what to send. It's one of these things where you might ask today, since we can vibe code anything so fast, why do we even need to design protocols or APIs or software? Why can't you just vibe code things on demand? But actually for this type of interoperability where multiple parties moving at different speeds are building stuff and you still want some layer on top to coordinate, you do want to design it and build it. So it reminds me of agents talking to each other and users talking to agents and tools.
我们还有其他评论或不同观点吗?
Do we know of any other comments or alternative viewpoints?
顺便说一句,我们确实就这个问题辩论过。我们谈了很多关于 Matter 的好处。我记得在决定做这件事的时候,我对 Matei 说:“嘿,碰巧有那么一周,我从醒来到睡觉都在不停地编码。我盯着我的云会话、Codex 会话,特别烦人的一件事是必须一直开着笔记本电脑。”
I think by the way, we had a debate on exactly this. We said the benefits with Matter a lot. And I think around the time we decided to do this thing, I was telling Matei, "Hey, it just happened to be there's a particular week that I was coding non-stop from the moment I woke up to the moment I went to bed. I was looking at my cloud sessions, my Codex sessions, and one of the things particularly annoying was having to keep my laptop open."
嗯。
Mhm.
我当时正开车去看医生,我记得,因为我想确保整个事情继续运行。顺便说一句,听到你这么说真让人欣慰,因为我在想我是不是个小丑,在做这种事……是啊,说实话,我一边开车一边把笔记本电脑连到手机上。把它放在旁边,每次遇到红灯,我就看看笔记本上发生了什么。我觉得这太荒谬了。
I was actually driving to a doctor's appointment and I remember because I want to make sure the whole thing continues working. By the way, it's so comforting to hear you say that because I'm like I don't know if I'm a clown and I'm doing this or like... Yeah, yeah, like honestly, I was driving and I was tethering my laptop to my phone. Keeping it on the side, whenever I hit a red light, I started looking at what's going on on my laptop. And I just felt that was ridiculous.
是啊。
Yeah.
感觉就像回到了编程的黑暗时代。
It felt like we went back to the dark ages of programming.
我是说,你从这个编码时代获得的生产力是惊人的,但是……你听说过云吗?
I mean, the productivity you gain from all this coding age is amazing, but... Have you heard of cloud?
这让我抓狂。
It was crazy to me.
我们当时在做的是沙盒,还是在那之前?
Was the thing we were working on the sandboxes or was this before that?
是沙盒。
It was a sandbox.
好的。所以你当时……
Okay. So you were...
所以我从一个非常不同的角度切入。我想:“嘿,我们要有不会关闭的云沙盒。你可以很快得到一个,但不仅用于运行智能体会话,也用于开发。”所以那周我亲自在构建它,过程中遇到了所有这些问题。然后我写了一份文档,列出了我对实际环境应该做什么的愿望清单。我觉得他最后几乎实现了每一条。
And so, I was approaching from a very different angle. I wanted to, "Hey, we're going to have cloud sandboxes that actually don't shut down. You can get one very quickly, but not just for running agentic sessions. It's actually also for running development." So, I was actually personally building that that week and through building that, I ran into all these issues. And then, I wrote a document for my case: "Here's my wish list of what the actual environment should do." And I think he actually ended up implementing almost every single one of them.
是啊,我记得 Reynold 说过,因为我的第一个原型只有和智能体聊天,他说:“我必须能打开一个 shell,就像我自己的 shell,列出文件、跟踪它们之类的。”所以我当时……
Yeah, I remember Reynold saying, because my first prototype of this had just chats with your agent and he said, "I have to be able to open a shell, like my own shell, and list files and tail them and stuff." So, I was...
这是 SSH 到大型机吗?
Is this an SSH into a mainframe?
是啊,实际上它就有这个功能。
Yeah, actually it has that.
跟踪我的日志。
Tailing my logs.
对,对。
Yeah. Yeah.
另外,我觉得我还问过一件事……我仍然用 Cursor 的唯一目的就是渲染 Markdown 文件。
And also, another thing I think I asked was... I still use Cursor for the sole purpose of rendering markdown files.
嗯哼。是的。
Uh-huh. Yes.
“给我一种查看 Markdown 文件并正确渲染的方法。我不再需要单独的工具了。”是啊,我觉得你也把这个功能加进去了。
"Give me a way to see my markdown files and render them properly. I don't need a separate tool anymore." Yeah. I think you also built that in.
我们做到了。是啊,我们有很多工程师在搭建自己的“氛围编程”环境。但他们都说了另一件事:“嘿,我为自己构建了很棒的东西,但团队里其他人都没法用,因为我没有服务器来协作。”这就是为什么我们尝试设置 Omnigen,这样你就可以有一个服务器,并在里面设置安全机制。所以,你可以用 Google 登录之类的,然后真正安全地共享东西。这就是为什么我们看到很多其他智能体遇到问题,比如人们以为他们原型了一个很棒的智能体,但因为安全团队的原因,它不被允许连接一些非常重要的数据。所以,是的。
We did that. Yeah. Yeah, we had a lot of engineers building their own vibe coding setup. But then, the other thing they all said is, "Hey, I built something that's amazing for me, but no one else on the team can use that because I don't have a server to collaborate." And this is why we tried to set up Omnigen so you can have a server and have the security set up in there. So, you know, log in with Google or whatever and actually securely share stuff. And that's why we've seen a lot of other agents hit things like people think they prototyped an awesome agent, but it's not allowed to connect to some really important data or whatever because of the security team. So, yeah.
是啊,现在……对于在 YouTube 上观看的朋友们,我们将展示一张结构图,然后简单聊聊架构。我想让大家理解,因为当我们谈论软件时,它可能非常抽象。而这就是我们实际在说的东西。你基本上在开源中构建了整个平台,有一个运行器组件和服务器组件,带有你设计出的统一 API。还有其他元素,显然你可以插入所有这些持久化层和计算层。这是一个完整的云。这就是 Agent Cloud。
Yeah, at this point... So, for those watching along on YouTube, we're going to bring up an image of the structure here and we can talk through a little bit of the architecture. I think I just want to have people understand because when we're talking about software, it can be very abstract. And here's actually what we're talking about. You've worked out in open source this entire platform basically and there's a runner component and server component with a sort of uniform API that you've figured out. Any other sort of element and obviously you can plug in all these persistence layers and compute layers. This is a whole cloud. This is Agent Cloud.
是啊,它有这些组件可以配合使用。很多操作发生在你部署智能体的机器上。所以,无论你在上面有什么,都可以运行。但没错,它是你想要托管协作智能体和拥有那个服务器的最小化东西。我们开源它的原因之一是,任何构建智能体的人都能得到一个可以开始使用并定制的应用,我们在 Databricks 也看到了这一点。比如有人做了一个不错的智能体应用,然后其他团队会问:“哦,我能直接用你的来做我的智能体吗?”
Yeah, it's got these components to work with it. A lot of the action happens on the machine where you deploy your agent to. So, whatever you've got on there, you can run. But yeah, it's sort of the minimal thing you want to have hosted like collaborative agents and to have that server. And one of the reasons we open sourced it is anyone building agents, this gives them an app they can start with and customize, which we were seeing in Databricks too. Like someone would make a nice agent app and then other teams would ask, "Oh, can I just use yours for my agent?"
是啊,我想我们有五六个不同的智能体框架,由各个团队构建。它们做的事情都差不多。
Yeah, I think we had like five or six different Agentic frameworks built by every different team. They all do more or less the same thing.
是啊,基本上人们想要拿一个在 Forkit 里能用的东西,你不如直接开源一个。是啊,这也是另一个问题,对 Databricks 来说很有趣:你选择开源什么,选择专有什么?我是说,这要追溯到 Spark,对吧?
Yeah, you need to basically people want to take something that works in Forkit and you might as well have something open source. Yeah, which also was another question that is interesting for a Databricks like what do you choose to open source, what do you choose to make it proprietary? I mean this goes back to Spark, right?
对。
Yeah.
所以,开源某样东西的一个原因是,如果你认为它是一个能产生网络效应的层,它会从许多人的协作中受益。例如,对于 Spark,我不知道你是否了解,当 Spark 出现时,我们也非常注重让你在上面拥有库。以前有不同的分布式计算引擎用于机器学习和图计算。我们说它们都应该成为你可以组合的库,并且我们也让添加数据源连接器变得非常容易。然后我们受益,因为我们没有时间为一千种不同的数据库和文件格式编写连接器。但我们可以直接使用人们制作的连接器,当然他们也从加入这个生态中受益。
So, I mean one of the reasons to open source something is if you think it's a layer that will actually have some network effect. It'll benefit from many people collaborating on it. So, for example, with Spark, I don't know if you know when Spark came out, we also focused a lot on letting you have libraries on top. So, like they used to be different distributed computing engines for machine learning and graph computation. We said they should all be libraries that you can compose, and we made it super easy to add connectors to data sources, too. And then we benefit because we don't have the time to write connectors to a thousand different databases and file formats. But we can just use the ones people make, and of course they benefit from joining this thing.
所以,这就像其中之一。另一种思考方式是,我可以想象,我们的东西不是开源的。我们有某种智能体托管服务,但它不是开源的,然后还有一个开源的。长期来看,哪个会赢?在这里,因为人们编写集成有好处,所以会是那个。然后还有其他事情,你根本无法以开源形式交付,而是公司做的事情。比如,你如何确保你的流式作业或基于湖的数据库不会在晚上丢失所有数据?这需要一个运营团队来维护。没办法,它必须是一项服务。所以,作为一家公司,我们希望确保我们非常擅长那些基础设施服务,然后在你构建的上层方面尽可能开放。
So, that's like one of these. Another way to think about it is I can imagine, you know, our thing wasn't open. We had some kind of agent hosting thing, but it's not open, and then there's an open one. Which one's going to win in the long run? So, here, because there is this benefit from people writing integrations, it'll be that. And then there are other things that you just can't even deliver as open source that are things the company does. Like, for example, how do you make sure your streaming jobs or your lake-based database doesn't lose all your data at night? Well, that requires an operational team that's going to sit there. There's no way, it has to be a service. So, we want to make sure as a company we're really good at those infra services, and then we're as open as we can in terms of what you build on top.
从好处来说,我认为我们已经看到了拉取请求和生态系统集成,尽管它只是在周六发布的。
I mean, speaking from a benefits, I think we're already seeing pull requests and ecosystem integration, even though it was only released on Saturday.
是的,周六。对,所以有人……
Yeah, Saturday. Yeah, so someone...
让我们看看发生了什么。是的。
Let's see what's going on. Yeah.
是的,你可以看看合并行。我今天早上问了一个传奇人物关于……
Yeah, you can look at the merge lines. I actually asked some legend this morning about the...
已经有 400 次合并了?
400 merges already?
是的,我认为相当多……我猜大约一半不是来自我的团队。但例如,有人添加了在 Kubernetes 上运行的支持。人们添加了许多云沙箱。所以,这可以启动一个云沙箱并在其中运行你的智能体,这对于共享也很好,因为它不像在你的笔记本电脑上,有人在那里运行恶意代码。嗯,所以,是的,许多初创公司已经加入了这些,我们预计会看到更多。我们还有更多的智能体框架,Cursor、CLI 和 anti-gravity 也是。
Yeah, I think quite... I would guess around half are not from my team. But for example, someone added support for running it on Kubernetes. People added many cloud sandboxes. So, this can launch a cloud sandbox and run your agent in there, which is great for sharing, too, because it's not like on your laptop and someone's running sky code on there. Um, so, yeah, many startups have put those in and we expect to see more of them. We also have more agent harnesses already, Cursor, CLI, and anti-gravity, also.
是的。
Yeah.
这一切都很美好。你知道,我觉得上次发生这种情况时,是现代数据栈的兴起。我不知道它是否那么有用。我其实对你的事后分析有点好奇。我想大多数人会同意它终于死了。但也许这催生了一个新的现代 AI 栈,做着同样的事情。
That's all beautiful. I, you know, I feel like the last time this happened, there was the rise of the modern data stack. I don't know if it was that useful. I'm actually kind of curious in your postmortem. I think most people will agree that it is finally dead. But maybe this arises to a new modern AI stack that does the same thing.
我不知道。
I don't know.
我的意思是,我认为现代数据栈是一个相当有用的东西,可能直到今天也是如此。我想也许对于那些不了解历史的听众来说。我认为现代数据栈实际上分解为:你需要一个层来摄取数据,你需要一个层来转换数据,然后所有这些都运行,然后你需要一个层来可视化数据。所有这些都运行在某种数据仓库上,或者后来,就像我们做数据仓库一样,还有湖仓一体。我认为这些概念都非常强大且非常有用。它某种程度上启用了许多工作负载。人们最终遇到的是一个统一和整合的问题。就是,“嘿,你真的需要把所有这些切成不同的部分,并与这么多不同的供应商和平台合作,才能完成一个非常简单的可视化吗?”对吧?所以,我认为随着时间的推移,每个人都开始意识到客户在推动我们。我们开始意识到这一点,所以我们开始构建越来越多的功能并尝试整合。最终,现在客户不必担心需要连接五个不同的系统才能生成一个图表。但我认为诚实地讲,类似的事情可能正在发生,你需要连接多少个不同的框架才能制作一个非常简单的智能体。
I mean, I think the modern data stack was a pretty useful thing, probably even up until this day. I think what maybe for the audience who don't actually understand the history. I think the modern data stack is effectively decomposed into: you need a layer to ingest the data in, you need a layer to transform your data, then all of this is run, and then you need a layer to maybe visualize your data. And all of this runs on some sort of data warehouse or later on, as we're doing data warehouse, also lakehouse. I think those concepts are all very powerful and very useful. It's sort of enabled a lot of workloads. What people eventually run into is kind of a question of unification and consolidation. It's, "Hey, do you really need to chop all of this into different pieces and work with so many different vendors and platforms in order to get a very simple visualization done?" Right? So, I think over time everybody started realizing that customers are pushing us. We start to realize that, so we start building more and more capabilities and trying to consolidate. And at the end of the day now, customers don't have to worry about having to hook up five different systems in order to produce a chart. But I think honestly something like this is probably happening in how many different frameworks do you want to hook up together in order to produce a very simple agent.
明确一下,我认为核心是这个在所有框架之上的通用 API。所以这个 API 基本上是这样的:你有一个智能体会话,你可以发送一条消息或一个文件。基本上,这就是你可以发送的内容,然后你会得到这些流,因为它正在流式传输文本或进行工具调用。或者你可以发送的另一件事是告诉它取消或轮次。这就是 API。现在,我们做的是我们可以让你在 Claude Code、终端中运行、Codex、Phi、OpenAI SDK 等所有这些东西之上获得这个。我们将它们全部映射到同一个接口。所以,如果你自己构建智能体编排器,你就需要自己维护这个。然后每当 Claude 更改其 API 时,你都必须调整你的东西,否则它会丢失一些消息。所以这就是值得维护的东西。在此之上,我们构建了一些应用。我认为我们构建了一个很酷的 UI 之类的东西,但那是……我们还构建了安全和控制部分,我对此感到兴奋。但它是那个通用接口。所以我们不试图成为一个栈。事实上,你可以在这个服务器之上插入你自己的 UI。这是我们非常关心的一个用例,因为我们想在自己的产品中使用它。
Just to be clear, I would say the core of this is this common API on top of all the harnesses. So the API is basically like: you've got an agent session and you can send in a message or a file. Basically, that's what you can send in and then you get out these streams as it's streaming text or as it's doing tool calls. Or the other thing you can send in is you can tell it to cancel or turn. So that's the API. Now, the thing we did is we could get you that on top of Claude Code, running in a terminal, Codex, Phi, OpenAI SDK, all that stuff. We map them all to that same interface. So that is something that you'd have to maintain yourself if you built your own agent orchestrator. And then whenever Claude changes its API, you got to tweak your thing or it's going to lose some messages. So that's the thing that's valuable to maintain. Then on top of that, we built a few apps. I think we built a pretty cool UI and stuff, but that's and we built the security and control piece which I'm excited about. But it's that common interface. So we don't it doesn't try to be a stack. And in fact, you could plug in your own UI on top of this server. That's one of the use cases we care a lot about because we want to use this in our own products.
是的,它应该无处不在。
Yeah, it should be everywhere.
是的。
Yeah.
我认为其中一件对我来说非常有趣的事情是,首先,我会努力做所有事情,而不称它为现代 AI 栈,因为我认为我们有一个名字。但就像是的,所以第一个告诉我关于计算沙箱的人是来自 Neon 的 Nikita。因为很多人认为 Neon 是,嗯,它是无服务器 Postgres,具有计算和存储分离以及即时分支等功能。但实际上,每个数据库公司也是一家计算公司。
I think one of those things that is really interesting to me is well, first of all, I'll endeavor to do everything and not call it the modern AI stack because I think we have a name. But like yes, so one of the first people that told me about compute sandboxing was Nikita from Neon. Because a lot of people think about Neon as, well, it's serverless Postgres with the separation of compute and storage and instant branching and all those things. But actually, every database company is also a compute company.
是的。
Yeah.
所以他实际上向我展示了他的整个沙箱解决方案。我不认为他发布过它。
And so he was actually showing to me his whole sandboxing solution. I don't think he ever launched it.
所以我们的沙箱解决方案,我们能够如此快速构建它的原因是因为我们意识到,如果你只是采用实际的湖仓一体架构并从中移除数据库,顺便说一句,这来自你。现在有一些差异。例如,在支持这种特定工作负载的系统中,拥有本地持久性很重要。因为你希望你的状态持久化。你的库,你不必每次都安装你的库。对吧?而 Neon 架构,由于存储与计算分离,你不需要持久的本地磁盘。所以有一些差异。但最终,是的,它是……
So our sandbox solution, the reason we could have built it so quickly was because we realized if you just take the actual lake base architecture and remove the database from it, by the way it was coming from you. Now there are some differences. For example, in the ones that support this particular workload, it is important to have local persistence. Because you want your state to persist. Your libraries you don't have to install your library every time. Right? Whereas the Neon architecture, because of the separation of storage from compute, you don't need persistent local disk. So there's some differences. But at the end of the day, yeah, it's...
是的,所以这就是当你运行一个编码沙箱时。就像如果我使用那个,我们在 Databricks 内部有开发基础设施,有几十 GB 的数据,仅仅是为了我构建的所有源代码和工件之类的东西,我希望下次能恢复它们。所以,是的。
Yeah, so this is when you run a coding sandbox. Like if I use that we had the dev infra internally at Databricks, there's many, many tens of gigabytes of data just for all the source code and artifacts and stuff that I built and I want that to come back next time. So but yeah.
在节目开始前,我们讨论了一些可能在采用方面令人惊讶的统计数据。可能是内部的,也可能是外部的,无论想到什么。只是为了让人们印象深刻,这个规模正在发生。
Before the show, we were talking about some statistics that might be surprising at the adoption. It could be internal, could be external, whatever comes to mind. Just to impress people the scale this is happening.
在分析方面,我们每天大概在三个云上启动 5000 到 6000 万个虚拟机。所以我们是最大的算力编排器之一,至少在 CPU 算力方面是这样。
So we on the analytics side, I think we launch maybe 50 or 60 million virtual machines a day across our three clouds. So we're one of the biggest compute orchestrators out there. That's for sure for CPU compute.
嗯。
Yeah.
所有这些流程处理的数据量,我觉得是 EB 级。我开玩笑说,取决于你在哪个时区,通常在你吃早餐之前,Databricks 当天就已经处理了 EB 级的数据。而在 Neon 上,它现在每天启动大约 1300 万个数据库,这非常有趣。
And all of those processes, I think exabytes of data. I joked about depending on which time zone you are, typically before you have breakfast, Databricks would have processed exabytes of data already on that day. And on Neon it's actually pretty interesting—it's launching I think 13 million databases a day now.
嗯,对我来说那是个大……
Yeah, that to me was like a big...
你是什么意思?
What do you mean?
而且其中很多要归功于智能体和分支实验。因为我们让启动数据库变得非常容易和快速,这多亏了 Nikita 的团队。这正在改变人们使用数据库的方式。
And then a lot of those were thanks to agents and branching experimentation. Because we made it so easy and so quickly—thanks a lot to Nikita's team—to launch databases. It's changing the way people use databases.
嗯。好的,我们一会儿再深入聊数据库,但我想先确认一下 Omnigen 相关的内容。你提到你对安全和控制方面感到兴奋。很多公司现在都在解决这个问题,还有成本方面。你发现了什么?
Yeah. Okay, we're going to go into more database talk in a bit, but I want to make sure we close up anything on Omnigen. You mentioned you're excited about the security and control side. A lot of companies are figuring that out right now, as well as the spend side. What have you found there?
是的,我花了很多时间与内部用户——开发者、安全团队、经理——以及很多客户交流。有几件事。首先,很明显的一点是,在安全方面,可用性和安全性之间存在张力。现在人们使用很多编码智能体的方式非常基础,比如你可以告诉它哪些工具模式允许或不允许。就像“是”或“否”。但这会让你陷入非常棘手的境地。举个例子,我的智能体应该能读取一些机密文档吗?或者它应该能安装来自 NPM 的新包吗?这些包可能被入侵。是或否。也许我想允许它。我的智能体应该能发布内容到公司网站吗?如果我在更新网站代码,那么是的。但它应该能同时做这两件事吗?这样它就可以获取机密文档,被提示注入并泄露它?可能不行。所以我们决定需要的是有状态的,或者我们称之为上下文策略,即跟踪会话的状态。这不是“它是否被允许推送到营销网站?”而是“嘿,如果它做了有风险的事情,比如安装了一个一天前发布的 NPM 包,或者读取了一千份机密文档,那么就不允许。否则,也许可以。”这就是一个移动权衡的例子。所以通过拥有更强大的引擎,它既更安全又更有用。这需要跟踪会话。另一个有趣的点是,它执行的是非常底层的事件,你希望有一些库在上面解析它们。例如,我们在内部有一个 Google Drive 上的 MCP 服务器,它有 60 个 API 调用。我怎么知道哪些会与互联网上的内容共享文档,哪些不会?这很烦人。所以我们在 Omnigen 中设计了策略层,使其成为函数,并且你可以拥有库。有人可以创建将底层事件映射到高层事件的东西,然后你针对产生的高层事件编写策略。这与 Panther 有关。Panther 会对此有所帮助。Panther 在事件处理方面有类似的想法,并且它基于 Python,而不是奇怪的自定义语言。这更像是实时的。
Yeah, so I spent quite a bit of time talking to internal users—developers, security team, managers—and also lots of customers. And there are a few things. First of all, one thing that immediately became obvious is for security, there's this tension between usability and security. The way people do a lot of coding agents today have very basic things like you can tell me which tool patterns are allowed or disallowed. It's like yes or no. But that puts you in a very tough spot. Just as an example, should my agent be able to read some confidential documents? Or should it be able to install new packages from NPM, which might be compromised? Yes or no. Maybe I want to allow it. Should my agent be able to publish stuff to the company website? Well, if I'm updating the code on the website, yes. But should it be able to do both? So it can grab a confidential document and be prompt injected and leak it? Probably not. So the thing we decided we need is stateful or what we call contextual policies, where you keep track of the state of that session. It's not like 'is it allowed to push to the marketing site or not?' But like 'hey, if it did a risky thing, like it installed a one-day-old package from NPM, or it read a thousand confidential docs, then no. Then don't do it. Otherwise, maybe it's okay.' That's one example of moving that trade-off. So it's both more secure and more useful by having a more powerful engine essentially. This requires tracking sessions. The other piece that was interesting there is there are these very low-level events it's doing and you want some libraries on top that parse them. For example, we have an MCP server on Google Drive internally. It's got 60 API calls. How do I know which of those will share a document with stuff on the internet and which ones won't? It's annoying. So we designed in Omnigen the policy layer so that it's functions and you can have libraries. Someone can make something that maps the low-level events to high-level ones and then you write a policy about the high-level things that came out. And that was related to Panther. Panther will help with that. Panther is kind of a similar idea on the event processing side and it's Python-based versus a weird custom language. This is sort of more as in real time.
这些事情正在发生。嗯。
Those things are happening. Yeah.
所以是的,但这些是酷的地方。我认为上下文或有状态的部分,以及它可以作为库的方式——这也是我们将其开源的原因之一,因为其他人会编写库,我们和我们的客户可以使用它们。最后一点,因为它是有状态的,我们跟踪的状态之一就是你在该会话中花了多少钱。所以我遇到过——我让一个智能体调试某个东西,它花了 500 美元,因为它决定读取大量日志文件并消耗大量 token。但我可以直接说,‘好的,启动一个子智能体来做这个,限制花费 5 美元。如果需要更多,就问我是否允许。’因为我们在该会话中进行了计数,它会弹出告诉我,‘好的,你花了 5 美元。你想继续吗?’
So yeah, but these are the cool things. I think the contextual or stateful part and then the way it can be libraries—and that was another reason to make it open source, because others will write libraries and we and our customers can use them. And the final thing, because it's stateful, one of the states we track is how much you spent in that session. So I can have—I've had like I asked an agent to debug something and it spent $500 because it decided to read a lot of log files and burn a lot of tokens. But I can literally say, 'Okay, launch a sub-agent to do this and cap it to spending $5. Ask me for permission if it needs more.' And because we're counting that within that session, it'll pop up and tell me, 'Okay, you spent $5. Do you want to go on?'
所以这里不仅仅是上下文。Matei 在过去五年里,大部分时间都在 Databricks 架构 Unity Catalog,这是数据治理层。
So I'm more than context here. Matei spent the last five years, a lot of his time was architecting Unity Catalog at Databricks, which is the governance layer for data.
没错。嗯。
That's right. Yeah.
这有点像把那一层的专业知识与这里所有的 AI 治理结合起来。
And it's sort of combining expertise at that layer together with all the AI governance here.
是的。嗯,但我也花了很多时间被编码智能体惹恼,还踩过雷。而且作为 CTO,我不想登上头条新闻,比如‘我安装了一个奇怪的 NPM 包,泄露了所有代码。’所以我特别多疑,但我也没什么时间。所以我不想坐在那里批准,比如‘你想运行一个 20 行的 bash 脚本吗?是还是否?’这就是为什么我花了很多时间思考,如何让它尽可能安全又不烦人。
Yes. Yeah, but I also spent a lot of time being annoyed by coding agents and getting bombs. And also as the CTO, I don't want to end up on the front page as 'I installed some weird NPM package and leaked all the code.' So I'm especially paranoid, but also I have very little time. So I don't want to sit there approving like, 'Do you want to run a 20-line bash script? Yes or no?' So that's why I spend a lot of time figuring out like, how can I make it as safe as possible and not annoying.
嗯。安全——我们称之为安全性——是比 token 上限或 token 预算更大的担忧吗?哪个更像……
Yeah. Is safety and let's call it security a bigger concern than token maxing or token budgets? Which one is like...
哦,是的,两者都存在。我的意思是,我不知道。我想这取决于你是什么类型的公司。所以我认为有些公司预算有限,他们真的很在意这个。或者,我的意思是,你可能是 Uber,但仍然会担心,你知道的。
Oh, yeah, they're both there. I mean, I don't know. I guess it depends on the type of company you are. So I think some companies like the budget is limited and they really care about that. Or I mean, you can be Uber and still be concerned, you know.
嗯,哦,是的,完全同意。嗯,嗯,嗯。如果你们……嗯,对我们来说,作为云提供商,安全绝对是至关重要的。这是最重要的事情。而 token 上限,我们目前还不太担心,但我见过——例如,我和一些咨询公司聊过。他们有大约 10 万名员工,都在为客户编码。如果每个人每月多花 1000 美元,那就不妙了。你知道,我们只有几千名工程师。
Yeah, oh yeah, totally. Yeah, yeah, yeah. If you have... Yeah, for us security is absolutely critical as a cloud provider. It's the most important thing. And token maxing, we're not so worried about it yet, but I've seen—for example, I talked to some consulting companies. They have like 100,000 employees who are all coding for customers. If those each spend like an extra $1,000 a month, that's not fun. You know, we have like only a few thousand engineers.
Databricks 的政策是什么?是无限的吗?
What's the policy in Databricks? Is it just unlimited or...
无限,但我们确实使用自己的产品来分析跟踪记录等,并且我们有一个团队负责优化和查看是否有人在做奇怪的事情。实际上,我们仅通过分析当前的跟踪记录就获得了一些非常酷的见解,比如哪些模型更擅长 Rust 而不是 TypeScript 之类的。所以是的,至少在我们的代码库中是这样。
Unlimited, but we do use our own product to analyze the traces and stuff, and we have a team that's looking to optimize and to see if anyone's doing something weird. And we actually had some really cool insights just from analyzing current traces, like which models are better at say Rust versus TypeScript or whatever. So yeah, at least in our code base.
嗯,太棒了。
Yeah, amazing.
显然,我得问一下关于 token maxing 的问题。我觉得这很关键,但安全和控制更重要,要找到一个合理的层次,既能有一些自主性,又不能太多。
Obviously, I have to ask the token maxing question. Obviously, I think it's a key thing, but security and control above that, and figuring out a sane layer where you can have some autonomy, but not too much.
是的,我们希望让它变得非常简单。作为工程师,你应该能设置这个东西。所以在 Omnigen 中,你可以让你的智能体为自己设置一个策略来做这件事。
Yeah, and we want to make it super easy. As an engineer, you should set the thing. So, in Omnigen, you can ask your agent to set up a policy on yourself to do this.
如果有什么我应该展示的,我没在 GitHub 上看到,但你知道,这是……
If there's anything I should be showing, I don't see it on the GitHub, but you know, this is...
它已经放在文档里了。你可以稍后查看。如果想看,就去文档里找“上下文策略”部分。
It's put in the docs there. So, you can look at it later. Just look in the docs on contextual policies if you want to see.
我只是想给人们展示一下策略。
I just like to show people policies.
对,如果你想深入了解,这就是你要看的地方,对吧?
Yeah, if you want to follow up on this, this is exactly where to look, right?
嗯。
Yeah.
是的。这些策略的由来是这样的:我写了一篇文档,里面列了 10 个想法,就像你在做的东西。那是我根据人们的需求列出的愿望清单,然后我跟团队说:“嘿,你们能在发布时至少实现其中五个吗?”结果他们把所有十个都做出来了。
Yeah. The story of these is like I just wrote a doc with like 10 ideas for things before as you were working on them. That was like my wish list of things people asked, and I told the team, 'Hey, can you do like at least five of these for the launch?' And then they just got back with all of them.
哦,哇。
Oh, wow.
所以,你可以想出更多策略,但其中一些只是作为示例。实际上,你可以拦截智能体发出的任何事件,然后你可以选择阻止、强制它询问用户,或者允许,并且你可以更新状态来跟踪信息。
So, you can come up with more, but some of them are just meant to be examples. Really, you can intercept any event the agent is making, and you can then either block or force it to ask the user, or allow, and you can update state to keep track of stuff.
是的,因为最终,我把你视为一个系统设计师,你让人们可以接入,对吧?这就是你做事的一贯方式。
Yeah, because ultimately, I think of you as a systems designer, you let people plug in, right? That's the whole modus operandi of what you do.
是的。我们也很注重可组合性。比如,别人能否编写一个供其他人使用的库,而这正是我们想要实现的。
Yeah. We care a lot about composability too. Like, can someone else write a library that others use, which this is meant to enable.
这里还有一种“开箱即用”的理念,可能和你做 Spark 的方式非常相似,就是你可以直接开始使用。
There's also a batteries included philosophy here, probably very similar to how you did Spark, which is you could just start using.
是的,没错。它必须在某些方面开箱即用,然后你可以在上面构建自己的东西。我们不想……但在 Spark 中,如果你只是想读取一个表或做聚合操作,它应该开箱即用就很棒。
Yeah, that's right. It has to be good out of the box at certain things, and then you can build your own things on top. We don't want to... but in Spark, if you just want to read a table or do an aggregation, it should be awesome out of the box.
人们如果想了解 Omnigen,应该看你的主题演讲,浏览 GitHub 和文档。如果他们想贡献或在这个生态系统上构建,你会指出哪些地方是参与进来最有杠杆效应的?
People want to catch up on Omnigen, they should watch your keynote, go through the GitHub and the docs. If they wanted to contribute or build on this ecosystem, where would you call out as the most high-leverage places to get involved?
是的,加入 Discord 和 GitHub。我们的团队在那里监控。人们要求的一些东西我们自己已经构建了。有些我们正在与他们合作构建。还要告诉我们你想如何使用它,因为我认为特别是对于开发者,每个人都希望它能按自己的方式工作。一个真正好的开发者工具必须听取所有方式的反馈,并找出抽象层以及如何让人们自定义。所以我们很乐意听到如果你觉得“嘿,我不希望它这样工作”,就告诉我们。我们真的只是想建立那个跨智能体的兼容层,然后让你在上面做事情。
Yeah, get involved in the Discord and GitHub. Our team is there monitoring. Some of the things people ask for we just built ourselves. Some of them, we're collaborating with them to build. Also tell us how you would like to use it, because I think especially for developers, everyone wants it to work their own way. A really good developer tool has to hear the feedback on all the ways and figure out the abstractions and how to let people customize. So we love to hear if you think, 'Hey, I don't want it to work this way,' tell us. We really just want to get that compatibility layer across agents and then let you do stuff on top.
从创业的角度来看,我是一个创始人,看到了一个机会,想和你聊聊。你对初创公司有什么期望,你希望有人正在做什么?
Is there any, in terms of the startup side, I'm a founder, I see an opportunity, I want to get in front of you. What's your request for a startup that you wish someone was working on?
哦,关于初创公司。
Oh, for a startup.
是的。比如,你有自己的初创公司,做得很好,但如果你不在做自己的初创公司,有什么明显应该做的事情……
Yeah. Like, you have your own startup, it's doing well, but if you weren't working on your own startup, what is obvious that you should...
你显然也建议了很多初创公司。
You advise many startups too, obviously.
我的意思是,作为一个拥有很多工程师的公司,任何能帮助我理解人们如何使用编码智能体以及花费,还有质量,或者像“你应该写这个技能”或“你应该写这个东西”或“你的智能体在涉及这个服务的任务上表现很差”或“去花时间”,那都会很好。
I mean, I do think just as a company with a lot of engineers, anything that helps me make sense of how people are using coding agents and spend, but also quality, or like 'you should write this skill' or 'you should write this thing' or 'your agents are really horrible at tasks involving this service' or 'go spend time.' That would be nice.
是的,我发现最接近的是 Get AI 这个团队。
Yeah, the closest I found is this team Get AI.
嗯。哦,酷,是的。
Mhm. Oh cool, yeah.
他们一开始说“我们只做代码和人类归属”,但基本上是在那之上构建分析层。我确实认为有很多像 Artificial Analysis 这样的公司,显然做得很好。所以会有人做。我认为这首先是咨询师的领域,但随后人们会真正构建软件,比如编码智能体的管理平面。
They started with 'we will just do code and human attribution' but they're basically building the analytics layer on top of that. I do think there are a bunch of like Artificial Analysis is obviously doing super well with their stuff. So there will be people. I think this is like the domain of consultants first, but then people will actually build software that is like the management plane for coding agents.
是的,我认为那里会有很多见解。你在其他领域也有这种情况。
Yeah, I think there'll be a lot of insights there. You have it in other areas.
好的,那么另一件大事是你的梦想引擎。如果你想讲讲 OLTP 的故事。
Okay, well, and then the other big thing is your dream engine. If you want to tell the story of OLTP.
我们的背景是……我会让人们听我们和 Ankur Goyal 的那期节目,我们讨论了 SingleStore、HTAP 以及所有那段历史。
And our background was... I'm going to make people listen to our Ankur Goyal episode where we talked about SingleStore, HTAP, and all that history.
是的,是的。OLTP 的想法其实很简单。人们听过 Ankur 关于 HTAP 的演讲,它实际上就是数据库的世界。抱歉,这里可能需要注入很多背景。数据库的世界……
Yeah, yeah. The OLTP idea is actually pretty simple. So, people have heard of Ankur's talk about HTAP, it's effectively the world of databases. Sorry, there's like maybe a lot of context that needs to be injected here. The world of databases...
为了成为那个我强迫人们学习数据库的播客,各位。你不能只用 markdown 文件来 vibe coding。
To be the database podcast that I'm forcing people to like learn your databases, guys. You cannot vibe code with just markdown files.
这是系统技术中最基础的东西之一。但是,数据库的世界实际上大致分为两半。一边是我们所说的 OLTP 数据库,它们是事务性的,想想你的 Postgres、MySQL、Oracle 数据库。另一边是我们所说的分析型数据库,有时可能被称为 OLAP。区别在于,在 OLTP 上,你通常可能对某个事件运行一些事务,查找特定的一行,更新那一行,对吧?这是一种非常面向行的数据结构。而在分析型数据库中,你试图对数据进行推理,试图计算:“嘿,我每家店的收入是多少?我的网站每天表现如何?”然后你最终可能想在上面运行机器学习来预测:“嘿,我未来的销售情况会怎样?”它们的架构非常不同,每个人都从 OLTP 数据库开始,因为每个应用,当它变得足够严肃时,需要的不仅仅是 markdown 文件,你需要一个数据库。你不想丢失数据。你想要一些事务一致性。但是,一旦你想对数据进行推理,如果你只有 100 行数据,在 Postgres 或 MySQL 上运行可能没问题。但是,一旦你有更多数据并想运行更复杂的分析,这些分析可能会压垮你的 Postgres 数据库。
It's one of the most important fundamentals of systems technologies out there. But, the world of databases is effectively split into roughly two halves. There's what we call OLTP databases, which are transactional, and think of your Postgres, your MySQL, your Oracle databases. And the other side is what we call analytics, and sometimes might refer to the term OLAP. The difference is on OLTP, you typically have maybe run some transactions on an event that looks up one specific row, we update that row, right? It's a very row-oriented data structure. And on analytics, you're trying to reason on the data, you're trying to compute, 'Hey, what's my revenue per store? How's my website doing every day?' And then you eventually want to probably end up running machine learning on it to predict, 'Hey, how will my sales be going in the future?' They are so very different architecture, and everybody starts with OLTP databases because every app, when it becomes serious enough, needs more than markdown files, you need to have a database. You don't want to lose your data. You want to have some transactional consistency. But, once you want to reason on the data, if you only have like 100 rows, it's probably okay to run it on your Postgres or your MySQL database. But, once you have more data and want to run more complicated analysis, the very analysis might crush your Postgres database.
所以你开始把数据取出来……复制到分析系统中。
So you start doing getting data out of the... Replicate them into the analytic systems.
是的,对一些人来说 Elasticsearch 是个大……嗯,有些数据确实会进入 Elasticsearch 做日志分析。我们很多客户显然会用 Databricks 来运行更复杂的任务。还有一个术语叫 CDC。
Yeah, which is for people Elasticsearch is like a big... Yeah, so some of them actually get into Elasticsearch for log analysis. A lot of our customers obviously get into Databricks to run more sophisticated things. And there's this term called CDC.
CDC,变更数据捕获。
CDC, change data capture.
它的作用是读取数据库的 bin log。如果你不懂 bin log 是什么,没关系。它就是一个数据的小增量,然后基于这个增量在分析端重建数据库的状态。但 CDC 非常痛苦。它基本上是行业标准,每个人都在用。但结果就是……我觉得很多数据工程师会因为管道问题凌晨 3 点被叫醒。
And what it does is it reads the bin log of the database. And if you don't understand what bin log is, fine. But it's a little delta of the data and then reconstructs based on the delta the state of the database on the analytic side. But CDC is a very painful thing. It's basically standard in the industry. Everybody uses it. But it ends up being... I think many data engineers end up being woken up at 3:00 a.m. because of some pipeline thing.
你知道,我的解释是,每个人都……你知道,靠做 CDC 就能成为一家 50 亿美元的公司。
You know, my explanation is like everybody is like... you know, became a $5 billion company just doing CDC.
是的,没错。CDC 是一个非常……它是现代社会最无聊但也最基础的操作之一。但它太脆弱了,我们开玩笑说它应该叫“持续数据损坏”,因为你可能在 OLTP 数据库上修改了 schema,然后 CDC 管道无法处理 schema 变更,接着一切就都乱了。
Yeah, exactly. CDC is a very... it's one of the most boring but one of the most fundamental operations powering modern society. But it's so brittle that we joke it should be called continuous data corruption because you might change your schema on your OLTP database and then the CDC pipeline fails to handle the schema change. And then everything goes out.
我的意思是,你可以用各种技巧,比如加一些版本控制之类的。但……
I mean, there's all sorts of tricks that you can do like you add in some versioning or whatever. But...
是的,但总的来说非常复杂。比如我在主题演讲中问观众,如果喜欢自己的 CDC 管道就举手。大概只有两个人举手。
Yeah, but it's very in general very complicated. Like I think at my keynote I asked the audience to put up their hand if they love their CDC pipeline. Only like maybe two people put it up.
所以如果 SingleStore……大概十年前,我认为行业有过这个想法:“嘿,如果我建一个能同时处理两种工作负载的数据库会怎样?”
So if SingleStore... like about maybe a decade ago I think the industry had this idea, "Hey, what if I built a single database that can handle both workloads?"
顺便说一句,每个数据库从业者都一直梦想着这个。
Which, by the way, every database person ever has always dreamed about this.
是的,是的,这是数据库工程的圣杯。为什么不建一个能同时做这两件事的系统呢?但结果往往是很多妥协。
Yes, yes, this is the holy grail of database engineering. Why not build a single system that can do both of this? But it ends up just being a lot of compromises.
首先,我认为第一个问题是,嘿,每个……他们说 Postgres 有庞大的生态系统,对吧?你想使用为 Postgres 构建的工具。而 Spark 也有庞大的生态系统,有很多你想用的库。如果你创建一个新东西,你没有生态系统。你往往会创建一个更小的专有 API,结果两边都缺。而且在性能上也很难做到与任一边相当。所以最终两边都做不好。
One, I think one of the first issues is that hey, each... they say Postgres has a massive ecosystem. Right? You want to be using the tools built for Postgres. And Spark, for example, had a massive ecosystem. There's a lot of libraries you want to use. If you were to create a new thing, you don't have an ecosystem. You tend to create a new smaller proprietary API and you're lacking both. And it's also very difficult to make it performance-wise to be comparable on either side. So it ends up actually sucking at both.
而我们 L-TAP 的整个想法,显然是对 HTAP 这个术语的戏仿,我们认为这是 HTAP 的正确做法。HTAP 想为两者构建一个统一的引擎。我们认为通过统一存储就能获得 99% 所需的功能。只要有一个统一的存储层,如果你的 Postgres 数据库以列式格式写入数据,那么所有分析任务就可以直接读取这些数据,没有任何延迟。对吧?中间没有管道,所以所有数据都能立即用于推理分析。
And our whole idea of L-TAP is kind of obviously a word play on the term HTAP, is that we think this is HTAP done right. HTAP wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage. And just have a single storage layer. And once you have the single storage layer, if your Postgres databases are writing data in a column-oriented format, everything analytics can just go read that data directly without any delay. Right? There's no pipeline in between, so all the data would immediately be available for reasoning analytics.
我想我之前跟一些客户说过,嘿,我们讨论过这对智能体非常有用。实际上我自己一开始也不太相信,尽管我们写了那个定位。但昨晚我和一个澳大利亚客户吃饭,他们告诉我,哦嘿,我们的一大问题是,我们有来自服务的所有日志,看到 SLA 下降想调查,但那些智能体根本无法理解实际数据库中发生了什么。我们看到的只是数据库和服务的产品遥测数据。如果你能了解,比如谁在下订单,发生了什么,他们具体在做什么,就能让那些智能体强大 10 倍。所以现在我完全相信我们自己的理念了。
I think I was telling some customers earlier, hey, when we talked about this is going to be super useful for agents. I actually at first didn't really believe in it myself, even though we wrote that positioning. But then last night I was having dinner with an Australian customer and they actually told me, oh hey, one of the big issues we have is we have all these logs from our services and we see SLA dips and want to investigate. But then there's no way for those agents to even understand what's going on in the actual databases themselves. All we see is just like product telemetry of the database and the services. You would actually make those agents 10 times more powerful if you understand, for example, who's actually placing those orders. What is happening? What exactly are they doing? So now I'm actually sold on our own message.
是的。
Yeah.
我认为这基本上能让你获得 HTAP 圣杯的几乎所有好处,那就是让数据立即可用于推理分析。
I think it's really kind of... it gets you basically almost all of the benefits of the HTAP holy grail, which is hey, make the data available immediately for reasoning analytics.
是的,我认为你知道,人类通常很聪明,希望有能力访问和查询任何东西。即使在工作时,他们也需要历史记录和上下文,而上下文从哪里来?那就是分析工作负载。
Yeah, I think you know in the way that humans are generally intelligent and want to have the ability and access to query anything. Even while they do the work, they also need history and they need context and like where else do they get context? That's an analytical workload.
正是。
Exactly.
是的,我记得我们的数据库出问题时,工程师说,我不能在上面跑一个大查询来看发生了什么,因为那会让数据库崩溃,造成更严重的伤害。而 L-TAP 就能解决这类问题,因为你可以启动一个完全独立的机器集群来做分析,不会让仍在提供服务的数据库过载。
Yeah, and I remember when we had incidents with our databases and the engineer said, well, I can't just run a giant query on it to see what's going on because that's going to bring down the database and hurt it even more. Like that's the kind of stuff that this gets rid of because you spin up a whole separate fleet of machines that's doing the analytics. You're not overloading the main database that's still trying to serve stuff.
所以,这已经是一个梦想很久了。为了实现今天的目标,需要做些什么?你知道,是的。我觉得你已经宣布过好几次变体了,但都没有 L-TAP 这么清晰。
So, this has been a dream for a while. What had to get done in order to get to today? Like, you know, yeah. I feel like you have announced variants of this several times. But it wasn't as clear as L-TAP.
是的。
Yeah.
我觉得 L-TAP 就像是……好了,我们搞定了,伙计们。
I think L-TAP is like a... like, okay, we've got it guys.
我和 Meta 的一个人聊天,他问我,嘿,有什么陷阱?为什么现在能实现了?我认为现实是,我们花了大量时间在 lake base 架构上。我的意思是,显然很多来自 Neon 团队,也就是存储与计算分离。结果发现只差一小步。从那个架构到 L-TAP 的想法,就是我们在 Neon 架构和 lake base 架构中,以行式格式将数据写入开放数据湖。但那里我们用的是 Postgres 页面。实际上,Ali 和我花了很多时间争论,嘿,我们能不能改成以列式格式写入?我们正在争论,然后有一天,我们一个非常聪明的工程师过来说,嘿,我刚做了原型,它成功了。
I was talking to somebody at Meta and then he was asking me, hey, what's the catch? Why is it possible now? And I think the reality is we took a lot of time to actually work on the lake base architecture. I mean, obviously a lot of it came from the Neon team, which is hey, separation of storage from compute. And it turned out it was just a tiny little step away. Going from that to this L-TAP idea, which is hey, we just in the Neon architecture and in the lake base architecture, we're writing data in row-oriented format to the open data lake. But in there, we're writing in Postgres pages. Actually, Ali and I were spending a lot of time debating, hey, can we actually just change that to write in column-oriented format? And we're just debating and then one day, one of our engineers, who's actually super smart, came in and said, hey, I just prototyped it, it works.
等等,原型什么?
Wait, prototype what?
原型不是以行式格式将数据存储在数据湖中。
Prototype instead of storing the data in the data lake in the row-oriented format.
比如 Postgres 页面,以 Parquet 格式写入。
Like Postgres pages, write them in Parquet.
是的。他观察到,嘿,我们的存储集群有很多空闲的 CPU。我们可以用这些 CPU 来做从行式到列式的转码,行式适合 OLTP,但列式适合分析。所以我们在那时做转码。事实上,一旦你转码了数据,数据压缩得更好。所以对于那些写入 S3 或其他数据湖(如对象存储)的服务,你可以写得更快,因为现在数据更小了。
Yeah. And he just made the observation that hey, our storage fleet has a lot of extra idle CPUs. And we could use those CPUs to do the transcoding from row to column, where row is good for OLTP, but column's good for analytics. So let's do the transcoding at that time. And as a matter of fact, once you transcode the data, the data compresses better. So from those services writing to, for example, S3 or other data lake like object stores, you can actually write them faster because now they are smaller.
所以没有额外开销,性能也没有妥协。
So there's no overhead, it's no compromise in performance.
对,反正我们集群里有多余的 CPU。
Yeah, but because we had extra CPUs anyway in the fleet anyway.
所以争论结束了。这是科技界的经典场景:大量争论,然后有人直接动手做了原型,结果成功了。
So the debate ended. I mean, it's one of the classics of tech: a lot of debate, but then somebody actually went ahead and just tried to prototype it and it worked.
但像这样对公司战略意义重大的事,我以为会有个启动会、设计文档之类的。完全没有。我们开了很多会争论,从基本原理上争论是否可行,然后有人就直接动手做了。
But something this strategic and important to the company, I expect there to be a kickoff thing, like a design doc. Nothing like that. He just... We were debating in many meetings, and then we were just debating whether it's possible or not from first principles, and then somebody just did it.
对,如果你营造出让大家能这么做的环境,那就太好了。Omni John 也有点类似。我觉得如果我只是写个文档说‘我们可以一起做这个’,大家就会想‘那这个呢?那个呢?’但如果你实际去试,就有帮助。然后如果有真实用户去测试,他们使劲用还能正常工作,或者像这个例子,你有实际工作负载,知道负载长什么样,就可以直接测试同样的模式。
Yeah, I mean, if you set yourself up so people do that, that would be great. And that happened a bit with Omni John too. I think if I just had a doc and like 'we can make these together', everyone would think 'oh, what about this? What about this?' But then if you try it out, it helps. And then if you have real users and they bash it and it's still working, or in this case, if you have the workload, you know what the workload looks like, you can just test the same pattern.
技术本身很酷,但最重要的是创新文化。你不需要征求我的许可,不需要走整套正式流程,直接去做就行。
Tech aside, which is very cool, this is like the most important thing: the culture of innovation. And you don't have to ask my permission, you don't have to do a whole formal process, just do it.
嗯,尤其是现在,我觉得有了 AI,做这种事其实更容易了。
Well, especially these days, I think with AI, it's actually easier to do that.
我觉得你很真实。我见过很多大公司的高管,规模大了事情就会变慢。你肯定也感觉到了,但你们 somehow 有一群核心人员是例外。
I think you are very real. I mean, I've met a lot of C-suite of large companies, and I think that at scale, things slow down. I'm sure you felt it already, but somehow you have this core of people that are exempt.
怎么做到的?
How?
我觉得我们招聘和共事的人非常优秀,这是很重要的一点,同时也要给他们授权,而且我们自己也花大量时间深入一线,这也很重要。
I think we hire and work with really good people, and that's a very important part of it, and empowering them, but also spending a lot of time maybe us in the trenches matters a lot too.
对,首先,人们会逐渐适应在大公司工作,这有帮助。我们要确保他们知道可以尝试新事物、解决争论,并且有很多先例可以参考,或者直接发布 beta 版。另外,作为一家公司,尽管规模很大,我们并没有发布很多产品。我们尽量保持产品线非常连贯。这其实是公司的核心理念:与其像亚马逊那样搞 20 个服务才能搭起分析和机器学习栈,不如只有一个服务,所有组件共用同一套 API、同一套语义、同一份数据。这需要统一。然后我们每次只增加一个东西,比如我们增加了存储层 Delta Lake——以前我们不做存储——然后增加了 SQL,再增加了机器学习平台。所以不要做太多,但要把每件事做好,这也有助于保持可控。
Yeah, I think first, people can kind of adapt to being in a larger company, so that helps. And we want to make sure they know that they can try stuff and settle debates, and have a lot of examples of how it was done before, or launch a thing in beta or whatever. And then the other thing I do think as a company, despite the size, we don't launch that many products. We try to keep it pretty coherent. That was actually the whole sort of theory of the company: instead of having like 20 Amazon services you need to set up an analytics and machine learning stack, you just have one, and it's the same API, the same semantics across all of them, the same copy of the data. So that requires unification. And then we basically added one more thing at a time, like we added storage with Delta Lake — we didn't used to do any storage — then we added SQL, we added machine learning platform stuff. So yeah, don't do too many, but do those things well, and that also helps keep it manageable.
对。
Yeah.
我们大力倡导的另一件事是:不要试图一开始就面面俱到,而是想办法增量地、快速地做。比如我们的很多产品,都是在几周内构建出来的,然后我们会问——通常我问开发团队的第一个问题是:目标客户是谁?你和谁合作?你和他们直呼其名吗?你跟他们发短信吗?我认为这种紧密的反馈循环非常重要。
The other thing we kind of encourage a lot is instead of building to boil the ocean for everything, let's figure out how to do it incrementally, how to do it very quickly. Like many of our products, they're built in the span of weeks, and then we go to... Hey, usually my first question to whoever team is building is: who's the target customer? Who are you working with? Are you on a first name basis with them? Are you texting with them? I think having that very tight loop.
你能再举一个类似这样启动的例子吗?我想给你一些背景。
Can you bring up another launch that comes to mind with this kind of thing? I just want to give you some background on that way.
那个……
The...
客户是谁?
Who's the customer?
对,其实这更偏向内部项目,因为我们自己开发人员要用。基本上整个 AI 团队都拿到了访问权限并在使用,我们从一开始就确保它能在我们巨大的单体仓库上正常工作。我们给了他们一些基础设施,大量的 token 容量。所以所有开发者都在用。我们还有其他例子。我不确定……这是个公开的故事,但……
Yeah, I'm even was more of an internal thing actually, because we would use that for our developer. Yeah, basically the whole AI team got access to it and was using it, and we made sure it works from the beginning with our internal code base, which is a monorepo that's like enormous. We gave them some infrastructure. We gave them lots of token capacity. So it's all the developers. Yeah. We had others. I don't... This is a public story, but...
我本来想问市场开放共享之类的,但我不太记得哪些是公开提到过的。
I was going to ask marketplace open sharing all of them had... I just don't remember exactly which ones publicly referenced it.
对,还有其他例子。公司非常早期的时候,我们做了 Delta Lake,也就是事务性存储层。当时我们最大的客户说:‘我需要一个云上的东西,因为如果我们的网络其他部分被攻破,这个东西必须独立出来存储和查询事件。’然后他们跟我们沟通。他说:‘这是每秒的事件率,这是我想要的时效性。你们能做到吗?’那个负载比我们当时任何工作负载都大得多。我们的工程师 Michael Armbrust 专门为此工作,让它跑起来。一旦为这个客户搞定了,对其他所有客户也就都适用了。那是公司早期,大概成立四年左右。
Yeah, they had other... Very, very early in the company there was like Delta Lake, which is the transactional storage layer we did. We had our largest customer at the time said: 'Okay, I need something in the cloud, because if the rest of our network is compromised, this thing needs to be separate to store and query the events.' And then talked to us. He said: 'Okay, this is the rate of events per second. This is the freshness I want. Can you do it?' So that was way larger than any workload we had, and we had our engineer working on that, my Michael Armbrust, and he worked just to make this work. And once it worked for them, it worked for everyone else. Yeah, this was early in the company, probably like four years in or so.
2018 年?
2018?
对,17、18 年。
Yeah, 17, 18.
对,比如 clean room,本质上是一种在不共享底层数据的情况下共享数据的方式,只允许特定操作。这些最初其实只为两个客户做的。我觉得行业里有一种看法:‘如果你过度适配一两个客户,对你很不利。’但我认为过度适配的坏处远小于好处。如果你野心太大想一口吃成胖子,那才是更大的问题。
Yeah, clean room, which is basically how you share data in a way without sharing underlying data but you allow specific operations. Those were done effectively initially just for two customers. I think the industry has a sense of: 'Hey, maybe if you overfit to one or two customers, it's going to be really bad for you.' But I think the downside of overfitting is much smaller than the upside itself. And if you sort of try to be too ambitious and boil the ocean, it's a much bigger problem.
对。
Yeah.
因为你可能最终一个客户都没有。
Because you might end up actually having no customer.
对,那更可能是结果。然后你可以从那里转向。我确实认为存在不好的客户,有时候你应该放弃他们。
Yeah, that's the more likely outcome. Then you can sort of pivot from there. I do think there is such a thing as a bad customer that sometimes you should fire.
确实存在。有时候如果你……嗯,我觉得我们可能遇到的一个挑战,也是很多新一代 AI 公司可能遇到的,就是科技公司与非科技公司或传统企业非常非常不同。如果你只针对科技公司优化一切,那么将产品扩展到科技公司之外可能会非常困难。
They could exist. Sometimes if you drive... Well, one of the challenges I think we probably see, and maybe many AI newer generation companies are seeing, is that tech companies are very, very different from non-tech companies or traditional enterprises. And if you optimize everything just for tech companies, you might have very challenges scaling them outside of tech companies.
好,你经常想到的前三大差异是什么?
Okay, what are like the top three differences that you always think about?
治理是很大的一个。
Governance is a big one.
对,我觉得很大的差异包括安全、数据隐私、治理等等。所以通常如果你在构建某种 B2B 或开发者工具,你最大的市场是企业,但企业非常不同。一家已经存在了比如 30 年、有某种 IT 系统的公司,他们有大量遗留系统,或者处于受监管的行业。
I think yeah, a big one is like security, data privacy, governance, all that stuff. So usually if you're building some kind of B2B or developer tool, your biggest market is going to be enterprises, but it's just very different. A company that's existed for like, you know, it's had some form of IT for like 30 years. They have so many legacy systems or they operate in a regulated space.
相比之下,初创公司甚至较新的科技公司,一切都是新的、近乎完美的。所以确实不同,如果你从未与企业合作过或身处其中,你就不会了解这一点。
Whereas a startup or even a more recent tech company, everything is new and sort of pristine. So yeah, it's just different, and if you've never worked with enterprises or been in one, you just won't know about it.
而且采购流程可能大不相同,实际上涉及的利益相关者要多得多。
And the procurement process is probably quite different. There's actually far more stakeholders.
这是其一。另一个有趣的点是,我认为一些科技公司的人会说:“哦,我可以自己建。”然后你就……
That's one. Another piece that's interesting is I think some tech companies, people will say, "Oh, I can build that myself." So then you go...
我觉得没人会这么说 Databricks。
I don't think people say that about Databricks.
不,他们会的。他们会的。他们会的。
Yeah, they do. They do. They do. They do. They do.
是的,我的意思是,这取决于团队等因素。但另一方面,许多企业会说:“实际上,我永远不想涉足构建这种东西的业务。比如,我不想因为某个怪胎书呆子搞不定流式管道而宕机。”这就是我的意思。
Yeah, I mean, yeah, and it depends on the teams and things. But on the other hand, many enterprises say, "Actually, I never want to be in the business of building that. Like, I don't want my... whatever, I'm a retailer or something, I never want to be down because some weird nerd couldn't get streaming pipelines working." That's what I'm...
是的,说实话,这让他们成为很好的客户,对吧?
Yeah, this makes them great customers, to be honest, right?
但你必须明白,没有在那里工作过是很难理解的。你可能体会不到。
But you have to understand that it's hard without having worked there and stuff. You may not appreciate.
听着,我认为他们都很好。别误会。他们有不同的挑战。但很多科技公司肯定更倾向于自己动手。
Look, I think they're all great. Don't get me wrong. They have different challenges. But many of the tech companies, for sure, there's a lot more DIY.
另一方面,有些人是其领域的专家。比如他们在造飞机、设计药物等等。他们只想要一座通往知识的桥梁,不想学习数据库之类的东西,不管我们觉得它多酷,甚至普通软件工程师可能觉得读一点也挺有趣。他们就是不想知道。他们只会说:“我有一个巨大的矩阵,里面是临床数据。我怎么对它进行聚类之类的?”所以是的。
On the flip side, you have people who are very much experts in their domain. Like they're building airplanes, designing medicines, whatever. And they just want a bridge to the knowledge where they don't want to learn databases or whatever, as cool as we think it is, even as interesting as the average software engineer might think it is to read a little bit. They just never want to know. They just say, "I have a giant matrix with my clinical data. How do I cluster it or whatever?" So yeah.
是的,没错。
Yeah, that's true.
好的,那么我实际上想展开讲讲这个“梦想引擎”愿景。这一切将走向何方?
Okay, so then I wanted to actually build out the sort of dream engine vision. Where does this all lead?
我们大概几年前意识到的一件事是,实际上每一个数据库引擎,尤其是分析型数据库引擎,都差不多有十年历史了。几乎所有有一定影响力的系统都大约有十年历史。它们最初都针对非常具体的狭窄用例。随着时间的推移,它们越来越成功,野心也越来越大。然后它们试图支持越来越多的用例。但支持这些用例的最快方式往往是在最初的设计上打补丁,而这些设计并非为这些用例而生。然后你勉强能支持它们。不知不觉中,经过十年的有机演化,它变成了一个巨大的……这包括 Databricks。我认为很少有公司或系统有勇气说:“让我们从头开始。让我们回到绘图板,基于我们今天所知道的一切——经过十年的工作流和可能数十亿美元的收入——重新设计。让我们尝试从头重写,并确保它能工作,能支持所有这些用例。”所以我们开始做了。但这是一个非常雄心勃勃的项目。顺便说一句,你可以在维基百科上搜索“第二系统综合征”。
So one of the things we realized maybe a couple of years back is that actually every single database engine out there, especially on the analytic side, are kind of a decade old. Pretty much everything that has reasonable traction is about a decade old. And they all started targeting some very specific narrow use cases. And then over time, they've become more and more successful. They've grown in their ambition. And then they try to support more and more use cases. But the fastest way to support those use cases tends to be hacks around the initial design that were not for those use cases. And then you can kind of support them more or less okay. And before you know it, after 10 years of organic evolution that way, it becomes a gigantic pile of... and that includes Databricks. And very few companies or systems, I think, have the guts to say, "Let's go start from scratch. Let's go back to the drawing board and design knowing everything we know today after a decade of workflows and probably billions in revenue. Let's attempt to rewrite it from scratch and actually make sure it will work and can support all these use cases." So we started doing that. But it's a very ambitious project. By the way, you can search on Wikipedia this thing called "second system syndrome."
是的,我知道。或者叫“第二系统效应”。
Yeah, I know that. Or second system effect.
每个开发者都必须知道第二系统……
Every developer must know what a second system...
基本上就是你构建第一个东西,效果很好,第二个注定失败,因为太雄心勃勃。然后你问别人……你知道,你以为自己无所不知,然后你会说:“这次我要设计一个完美的系统。”
It's basically you build your first thing and it works out great, and the second one's bound to fail because too ambitious. And then you ask someone... you know, you think you know everything and then you're like, "I'm going to design the perfect system this time."
是的。结果发现它并不完美,然后开始失败,你太雄心勃勃,永远无法发布,然后你就完蛋了。实际启动这个项目的工程团队非常出色。我认为我们雇佣了地球上一些最好的数据库工程师到 Databricks,他们非常出色。谢天谢地,这不是他们的第二个系统。他们中的许多人过去已经构建过不止两个系统。
Yeah. And it turned out it's not perfect, and then it starts failing, and you're too ambitious, never launch, and you get killed. The engineering team that actually started this, they were brilliant. I think we hired some of the best database engineers on the planet into Databricks, and they were brilliant. Thank god it's not their second system. Many of them have built more than two in the past.
不错。
Nice.
但他们仍然担心这一点。“嘿,从头构建一个数据库引擎,我认为传统观点是需要大约 5 年才能成熟。这将是一个非常长期的项目。它可能会失败。”我想其中一位工程师开玩笑说:“嘿,也许我们干脆叫它‘Stream Engine 项目’。”如果以联合创始人的名字命名,也许我们就不会被取消或扼杀。但我认为他们构建了一些非常了不起的东西。他们回到了……他们从范式的角度改变了数据库引擎的构建方式。通常当你构建一个数据库引擎时,你会阅读大量学术论文,尝试理解最新的算法和数据结构,然后把它们组合起来,看看是否有效。这也有很高的失败风险,因为任何在纸上看起来很好的东西可能在 70% 的工作负载下表现良好,但在另外 30% 下却适得其反。他们实际上构建了一个更像“数据库工厂”的东西。所以他们花了更多时间构建这个工厂,而这个工厂利用了我们过去十年的追踪数据。我认为他们在追踪表中计数了千万亿个数据点。
But they were still worried about this. "Hey, building a database engine from scratch, I think the conventional wisdom is going to take like 5 years to mature. This will be a very long-term project. It could fail." I think one of the engineers kind of joking and said, "Hey, maybe we just call it Project Stream Engine." If we name after co-founder, maybe we don't get cancelled or killed. But I think they built something pretty remarkable. They went back to... they changed the way the database engines were built from a paradigm point of view. Usually when you build a database engine, you read a lot of academic papers, you try to understand the latest algorithms and data structures, and you put them together and see if they work or not. And there's a high risk of failure there also, because whatever looks really good on paper might work out in 70% of the workloads, but then backfires on the other 30%. They actually built more of a factory for building the databases. So they spent more time building this factory, and the factory takes the decade of traces we have. I think they count as a quadrillion data points in the trace table.
你们没有丢弃任何数据?还是你们采样了?
You don't drop anything? Or you see sample?
我们当然会采样,但数据量仍然巨大。他们利用这些数据构建了一个模型。一个机器学习模型,不是大语言模型。这个模型可以非常快速地告诉我们,任何算法和任何实现对于任何特定类型的查询会有什么样的性能,而且保真度非常高。基于此,他们可以选择最有可能帮助处理不同类型工作负载的算法和数据结构。
We for sure sample, but there's like massive amount of things. And they use that to build a model. Like a machine learning model. Not an LLM, a machine learning model. It can very quickly tell us how any algorithm and any implementation will perform for any specific type of queries with very high fidelity. And based on that, they can pick the most likely algorithm and data structure that will actually help with the different kinds of workloads.
嗯。
Mhm.
无论是在运行时还是在实现时。
Both at runtime as well as at implementation time.
嗯。
Mhm.
因为有无穷无尽的……
Because there's like unlimited number of...
我的意思是,听起来你们想路由到不同的数据结构。
I mean it sounds like you want to route to different data structures.
是的,我的意思是,如果你想一想,一个单一的数据库有很多东西是集成实现的,但你需要确保它们彼此都能很好地协同工作,而且对于任何给定的操作,可能有不止一种实现。所以我们实际上让它……现实是,一个在极低延迟下表现极好的算法,可能并不适合扫描 PB 级数据。大多数情况下,吞吐量和延迟之间存在权衡。
Yeah, I mean if you think about it, a single database has many things implemented together, but you want to make sure they all work well with each other, and then for any given operation there might be more than one implementation. So we make it actually... reality is, an algorithm that works super well for very low latency might not work very well for scanning through petabytes of data. Most often there's a trade-off between throughput and latency.
关键维度有哪些?比如规模、吞吐量、延迟,还有什么别的吗?
What are the key dimensions like scale, throughput, latency, what what what anything else?
还有数据的分布。数据的稀疏程度,这非常重要。你多久会碰到相同的数据?有多少不同的值等等?这些都很重要。比如不同值的数量基本上会影响聚合的内存消耗。你的哈希表,到某个点就会有一个哈希表。
And the distribution of data. How sparse the data is. That matters a lot. How frequently do you hit the same data? How many distinct values and stuff like that? Those things matter a lot. Like number of distinct value basically impacts the memory consumption of your aggregation. Your hash like at some point there's a hash table.
我打算在我的文章里把这些都列出来,因为我真的很想要一个分类法。对我来说分类法非常有用,因为它涵盖了所有应该考虑的东西。
That somebody I'm going to in my write-up I'm going to try to list all these out because I really want a taxonomy. To me taxonomies are so helpful because it covers everything they should think about.
我觉得如果你真的想列出来,可能有一百万个不同的特征。
I think if you actually try to list it out probably like a million different features.
我总是想要那种,比如,给我 12 个。
I always want like okay give me like 12.
呃,你知道,就像大概 40 年前 Oracle 的一篇论文提出了分布式系统的八大谬误。那种东西非常有用。
Uh you know like a someone did like I think a Oracle paper in like 40 years ago did like the This is the eight fallacies of distributed systems. That kind of thing is super useful.
对,就是这样。就像,好的,把这八个想清楚。
Yeah, that's it. It's like, okay, think through these eight.
但让我给你一个非常奇怪的例子,它实际上对性能有深远的影响,比如你的字符串是 ASCII 还是包含 Unicode?我应该怎么编码?我的意思是字符串是最复杂的数据类型。
But let me give you a very weird example, but it actually has a profound implication on performance, which is like it's your string is ASCII or does it have Unicode in it? How should I encode this? I mean strings are the most complex data types.
所以,比如如果字符串非常密集,你实际上可以把每个字符串转换成一个……想象一下我要做聚合,不用哈希表,你可以用一个数组,因为如果你的字符串足够密集,只有 256 个选项,你就不需要哈希表,直接做数组查找就行。
So the that like for example if strings are super dense, you could actually convert every string into a like imagine I have to do a aggregation instead of having a hash table, you could actually have an array because if your string is dense enough, if you only have 256 options, you don't need a hash table. You can just do array lookup.
对,比如国家代码之类的。
Yeah. Like a country code or something.
所以实际上那个模型里可能有数百万个特征。但利用这些,他们基本上可以优先考虑那些在实践中真正有影响的算法。其中很多都非常反直觉。那些你以为效果很好的东西,实际上在实践中并不那么好。但更重要的是,在运行时,你可以调度正确的算法和结构。
So it's actually like probably millions of uh features in that model. But using that, they can one basically prioritize the different algorithms that might actually impact in practice. And many of them are very counterintuitive. It isn't actually things that you think it might work super well actually don't work that well in practice. But also more importantly at runtime, you can dispatch the right algorithm and structure.
我在听这个愿景。我觉得 Databricks 在渐进式演进方面做得非常好。你们有没有在某个点必须硬切换到新系统?
I'm listening to the dream. I feel like Databricks is doing a really good job with the incremental evolution. Do you have to hard cut to a new system at any point or like
我们设计成可以渐进式推进。所以首先我们发布一个新的端点。但这涉及到更广阔的领域。我们想做的是一方面,通过设计,这个新引擎应该能够完成我们以前能做的所有事情,并且做得更好。对吧?特别是“更好”这部分指的是非常低延迟的工作负载,可以在几十毫秒内完成。但我们希望以增量能力逐步推出,这样就不需要花 5 年时间才能看到曙光。
We designed it in a way that it can be incremental. So first we're releasing a new endpoint. But this goes to the broader ocean versus What we wanted to do is one to the by design, this new engine should be able to do everything we're able to do before and better. Right? It's been particular the better part refers to very low latency low latency workloads that can finish in tens of milliseconds. But we want to roll it out incrementally with incremental capabilities so it doesn't take like 5 years to actually see the light at the end of the tunnel.
我觉得这是一项艰巨的任务。我不知道还能怎么说。我对任何新型工作负载和新数据库都非常感兴趣。显然,我想我已经表明了自己有点数据库极客。事务型数据库,抱歉,会计数据库,比如 TigerBeetle,你见过吗?
I think that's a heroic task. I don't know what other way to say it. I am really interested in any sort of new workload and new databases. I mean, obviously, I think if I've maybe established that I'm a little bit of a database nerd. The transactional databases, sorry, the accounting databases, like the TigerBeetles. I don't know if you've seen those.
他们是做什么的?
What do they do?
复式记账会计数据库。就是专门用来建模金融账户和信用系统的。
Dual entry accounting database. Like it's just meant to really model like financial accounts and credit systems and
这是一个非常特定的……
It's like a very specific
非常高吞吐量,对。不,所以当你谈到每个人如何从一个东西开始,然后扩展,再添加其他东西时,正是如此。我最近采访了 TurboFifo 的 Simon,也是一样。还有 Chroma。2023 年的所有向量数据库公司突然都变成了通用的 blob 存储。
very high throughput, yeah. Yeah. No, it's so when you were talking about how everyone like starts with a thing and then they scale up and then they tack on other things. It's exactly that. And then I used recently interviewed Simon from TurboFifo, same thing. And Chroma as well. Like they all the vector database companies of 2023 all are suddenly now just we're just generally general storage of blob storage.
尤其是数据库从来就不应该是一个独立的类别。
Especially a database should have never been a separate category.
我觉得这曾经是一个激进的观点。现在却成了常规智慧。什么应该是一个独立的类别?如果一切都变成了 ELT,那……
I think that used to be a hot take. Now it's like the conventional wisdom nowadays. What should be a separate category? You know, if everything becomes ELT, like what's
我认为 ELT 的论点是,我们不是在查询层合并数据库,我们只是在合并存储层。对,我认为这是非常重要的一点。而且我们实际上认为将查询层合并成一个像 HTAP 风格的数据库是没有意义的。另外,很多人觉得如果只有一种查询语言就好了,不用操心 Postgres SQL 和 Spark SQL,为什么不能只有一种?但我不认为这对智能体是个问题。智能体在 Postgres SQL 或 Spark SQL 上都非常流利,永远不会混淆。只要数据在那里并且可访问,智能体就能做得很好。这可能是……
I think the thesis of ELT is we're not collapsing the databases at the actual query layer. We're just collapsing the storage layer. Yeah. And that's a I think a very important part. And we actually don't think it makes sense to collapse the query layer into a single like HTAP style database. And part of it By the way, the other thing I think a lot of people had is hey, it would be nice if there was only one query language I have to worry about instead of worrying about Postgres SQL and maybe Spark SQL, why not just one? But I don't think that's an issue for agents. Agents are very eloquent in Postgres SQL or Spark SQL. It's never going to get confused. As long as the data is there and it's accessible, agents will do fine. That that might have been
对,而且……
Yeah, and the
5 年前对人类来说可能是个问题。
5 years ago might have been a problem for humans.
这个问题随着时间的推移也可能出现,但这引出了如何渐进式做事的问题,对吧?比如我们意识到现在不需要它。我们不需要解决那个问题就能从当前的 Delta 应用中获得很多价值。
That could arise over time also, but it should and this is leads to how to do things incrementally, right? Like you we realize you don't need it right now. We don't need to solve that problem to have a lot of value from from the current Delta app.
好的,我要用一些更劲爆的话题来结束这期播客。每个人都接受了存储和计算分离的理念,并试图构建云。我从 Snowflake 那里也听到过同样的说辞。你们是如何在他们失败的地方成功的?
Yeah, okay. I'm going to end the pod with a little bit of more of sort of spicier things. Everyone has like had the receive within a separation of storage and compute and try to build you know the the clouds. I had the same pitches from Snowflake. How have you succeeded where they failed?
这问题有点尖锐。
Rough.
好吧,我尊重他们是竞争对手。客观上说,你们已经超越了它们。
Well, I mean respect that they are a competitor. Objectively, you have outpaced them.
从你的角度来看,核心洞察是什么?你们只是走了不同的方向?
What is the core insight from your point of view that you guys just went different directions?
可能最大的根本区别是。两家公司差不多同时起步,都转向了云端,都专注于存储与计算分离的架构。但最大的区别之一是开放性。Databricks 从来没有专有格式,对吧?我们从开放生态系统开始,从 Parquet 开始,然后演进到 Delta 和 Iceberg 等等。这是一个大问题。我认为这非常重要。另一个是 AI。我的意思是,在 2022 年 10 月 ChatGPT 出现之前,我们一直把 Databricks 定位为机器学习加数据。很多平台都是围绕机器学习用例构建的。显然 AI 有点不同。Matei 在这方面花的时间比我多得多。但整个平台从来没有让我们觉得我们只是一个数据基础设施平台。
Probably the biggest fundamental difference. Both companies started around the same time. Both went to the cloud. Both focused on storage from compute architecture. But the biggest difference one is open. Like Databricks had never had a proprietary format, right? We started with the open sort of ecosystem. Started with Parquet and then evolved into Delta and Iceberg and all that. It's like one big thing. I think that matters a lot. The other one is AI. I mean before 2022 October 2022 when ChatGPT came out, we had always pitched Databricks as a machine learning plus data. And a lot of the platform were built with machine learning use cases in mind. And obviously AI is a little bit different. And Matei is like spent far more time there than I do. But the whole platform was we never felt hey we're just a data infrastructure platform.
就像 Databricks 独有的那样,对。
Like Databricks only yeah.
我认为他们最初的想法是:‘好吧,我们只管理最有价值的数据,并让它变得非常快。为此,我们会有自己的存储,通过引擎优化。然后我们只针对经理和财务人员查看的那一小部分数据,让它的服务速度极快。’那是另一个领域。而我们则从批量处理和数据摄入开始。你有 JSON 日志文件,不管什么,我们做那种大规模处理,因为 Spark 就是干这个的——大规模批处理。然后我们把数据保存在开放格式中。可能慢一些,但数据已经在那里,下游可以消费。结果发现,从那个擅长规模化、摄入和低成本的批处理系统出发,更容易创建出兼具速度和易用性、面向商业用户的小数据版本。
I think they started with a mindset like, 'Okay, we'll manage the most valuable data and try to make it really fast. For that, we'll have our own storage, optimized with the engine. Then we'll target the small amount of data that managers and finance people look at and make that super fast to serve.' It was a different space. Whereas we started with bulk processing and ingest. You have JSON log files, whatever, we do that large-scale stuff because that's what Spark was for—large-scale batch processing. Then we keep the data in an open format. It might be slower, but it's already out there, you can consume it downstream. It turned out that it's easier to go from that batch system, which is great at scale, ingest, and low cost, and create versions with the speed and features of the easy-to-use smaller data for business users.
然后进行优化。
And then optimize.
是的,从开放和大规模开始。从某种意义上说,我们处于他们的上游。有一段时间,我们互相列为合作伙伴,因为你可以用 Databricks 做数据摄入和计算,然后用 Snowflake 提供表服务,实现可视化和速度。那很棒。然后我们都意识到客户在问:‘为什么我需要另一个东西?为什么我不能直接查询你的表?’我们说:‘不,我们在这方面很糟糕。请用我们的合作伙伴做 SQL 仓库。’然后他们意识到:‘等等,这么多计算正在向上游转移到这个东西里。’
Yeah, start open and start large. In some sense, we started upstream of them. There was a time when we both listed each other as partners because you could use Databricks for ingest and compute, then serve the tables out of Snowflake for visualization and speed. That was great. Then we both realized customers were telling us, 'Why do I need this other thing? Why can't I just query your tables?' And we said, 'No, we're horrible at that. Please use our partner for the SQL warehouse stuff.' Then they realized, 'Wait, so much of the compute is moving upstream into this other thing.'
你们不得不进入彼此的领域。
You have to go into each other's territory.
但我认为我们确实是从更大的范围和开放开始的。这很重要。对于那种幽灵企业来说,如果你的公司已经存在了 30 年,你经历过被 Oracle 锁定和各种疯狂的事情。如果你是 CTO,正在为未来搭建架构,你会想选择一个开放的基础。理想情况下,你只想用一种方式管理公司数据,而不是七种不同的系统。
But I think we did start with the bigger scope and with the open thing. That's important. As a kind of ghost enterprises, if your company has existed for 30 years, you've experienced being locked into Oracle and all kinds of crazy things. If you're the CTO setting up the architecture for the future, you want to pick a foundation that's open. You only want one way to manage data in your company ideally, not seven different systems.
但我认为数据格式已经赢了。现在每个企业都想把数据放在开放数据格式中。但当时这很有争议。五六年前,Snowflake 的一位联合创始人写了一篇博客叫《明智地选择开放》,基本上是反对的。
But I think the data format has won. Now every enterprise wants to put data in open data format. But it was very controversial back then. Five, six years ago, one of the Snowflake co-founders wrote a blog called 'Choosing Open Wisely', which basically argued against it.
是的,是的。
Yeah, yeah.
我想他们可能已经删掉了。你现在得找存档。
I think they might have taken it down. You have to find the archive now.
哦,现在它永远不会消失了。不,它还在。我喜欢只有你们才有的视角,因为你们经营公司。谢谢你容忍这个。这是一个不可思议的视角。
Oh, it's never going away now. No, it's still there. I love the perspective that only you guys will have because you run the company. Thank you for indulging this. It's an incredible perspective.
也许最后一个问题。在你说话的时候,我觉得我必须给 Ali 很多赞誉。他是一位了不起的 CEO。他是智商、情商、技术痴迷、执行力、商业头脑的完美结合。而且他也是创始人,这让他更容易动员和执行。我想就是这样。那么,你有 Ali,他们不喜欢吗?
Maybe one last one. As you were talking, I think I have to give Ali a lot of credit. He's an incredible CEO. He's the perfect combination of IQ, EQ, technology obsession, execution, business acumen. And he's also a founder, which makes it a lot easier for him to mobilize and execute. I think that's it. So, did you have Ali and they don't like okay?
嗯,还有很多其他事情,但我认为 Ali 发挥了相当大的作用。
Well, there's a whole lot of other things, but I think Ali played a pretty big role.
我以为会有一些他贡献的技术选择。
I thought there was going to be some technical choice that he contributed to.
他推动了很多这些事情。在岔路口,他推动了一个方向,然后证明那是正确的方向。
He pushed for a lot of these. There were forks in the road where he pushed for one way and then it became clear that was the right way.
需要写一整本书关于你们八个人如何合作。我想已经有一些人物特写。第二个问题,又不是一个清晰的问题。Mosaic。我们社区很多人对 Databricks 的模型故事很好奇。当你们收购 Mosaic 时,事情是:‘好吧,我们可以做微调。我们要做内部模型。’他们有 Mosaic 模型。但看起来你们没有这样做。你们似乎更倾向于 Alt App 和 Harness 之类的东西。这背后的故事是什么?
There's a whole book that needs to be written about how the eight of you worked together. I think there have been profiles. Second one, not a clear questioning again. Mosaic. A lot of people in our community are curious about the model story of Databricks. When you bought Mosaic, the thing was, 'Okay, we can do fine-tuning. We're going to do in-house model.' They had the Mosaic models. And it seems like you're not doing that. You're going towards more of the Alt App and the Harness stuff. What's the story there?
当 Mosaic 开始时,它以早期发布开源大语言模型而闻名。实际上在那之前,他们在做其他事情——优化训练系统。他们有世界上最快的图像模型训练栈。然后他们决定做大语言模型,这很聪明。他们在 ChatGPT 之前就进入了。所以他们有一些最早的开源大语言模型。
When Mosaic started, it was well known for releasing open source LLMs early on. Actually before that, they were doing other things—optimizing training systems. They had the fastest image model training stack in the world. Then they decided to do LLMs, which was smart. They moved into it before ChatGPT. So they had some of the first open source LLMs.
我们为 MPT-7B 采访了 John Franco 和 Abby。
We interviewed John Franco and Abby for MPT-7B.
是的,没错。所以我们决定,尽管我们发布了一个开源模型 DBRX,并且规模超过了 Llama 3,但我们决定专注于下一步:如何让一个非常聪明的模型变得有用。对我们来说,这关乎自动化,让它非常擅长查询数据。这就是我们称之为 Genie 的第一方智能体。它就像一个虚拟数据科学家。想象一下,有人已经对你公司的一切了如指掌,知道所有的机器学习库、数据库、网络上的东西,你可以问他们问题。那是我们想先做的事情。所以这意味着不要那么专注于训练某个前沿模型,而是构建一个使用外部模型或微调定制组件的系统。不过我们仍然在做相当多的模型训练,一直在采购大量 GPU。有几个地方我们在做这件事。
Yeah, exactly. So we decided, even though we launched an open source model DBRX and went up to above the Llama 3 scale, we decided to focus on the next step: how to make a very smart model useful. For us, it was about automating how to make it very good at querying data. That's the first-party agent we call Genie. It's like a virtual data scientist. Imagine someone who already knows all the stuff in your company inside out, knows all the machine learning libraries, data libraries, web stuff, and you can ask them questions. That's what we wanted to do first. So that meant not focusing as much on training some frontier model, but building a system using external models or fine-tuned customized components. We're still doing quite a bit of model training though, procuring lots of GPUs all the time. There are a few places where we're doing that.
一是,有很多高容量用例,如果你有一个专用模型,它比任何通用模型都要好得多。一个很好的例子是理解文档,比如 PDF、Word 文档,解析它们。如果你试过,会很沮丧,因为你把它发给 Claude 之类的,它几乎能搞定但会出错,而且超级贵——你把一张图片塞进去就消耗了大量 token。所以我们的团队构建了一个文档视觉模型,它接收一页文档,返回一个包含所有组件的漂亮 JSON。它非常有竞争力,可能比那些前沿模型便宜 100 倍,而且效果更好。这实际上是由一位来自 DeepMind 的研究员完成的,他是 Adapt 的联合创始人,很早就在做 LLM Scaling,但专注于这个方向。
One is, there are many high volume use cases where if you have a specialized model, it's just so much better than any of the general models you get. A nice example is understanding documents like PDFs, Word documents, parsing them. If you've ever tried that, it's frustrating because you send it to Claude or whatever, it almost gets it but gets some things wrong, and it's super expensive—you just burned a huge amount of tokens plopping an image in there. So our team built a document vision model that takes a page and gives you back a nice JSON with all the components. It's very competitive, probably 100x cheaper than those frontier models and still better. That's actually done by one of the researchers who came from DeepMind, was a co-founder of Adapt, very early LLM scaling person, but focused on this.
还有 Anthropic、Commission 也是,是的。
And Anthropic, Commission also, yeah.
还有 UC Berkeley,实际上,我那里的一名研究生写了一篇关于 advisor models 的论文。我想是在那些东西出来之前。我是说,我确信其他人同时也有这个想法,但这非常有帮助。所以是的,我们今天在主题演讲中展示了一些东西……
And UC Berkeley, actually, one of my grad students there wrote a paper called advisor models. I think before those came out. I mean, I'm sure others had the idea at the same time, but that's something that helps a ton. So yeah, we actually showed some stuff just today at the keynote on...
是 Parf 吗?哦,你认识 Parf?
Is it Parf? Oh, you know Parf?
Parf,是的,是的,Parf 是……
Parf, yeah, yeah, Parf is...
他在我的活动上演讲。他在 Continual Learning Bench 演讲。
Speaking at my thing. He's speaking at Continual Learning Bench.
是的,是的,对。我是他在 Adapt 的顾问之一,是的。
Yes, yes, yeah. I'm one of his advisors at Adapt, yeah.
我采访过他弟弟 Chai,因为他也在 Adapt。
Interviewed his brother, Chai, cuz he's also at Adapt.
是的,是的,是的。
Yeah, yeah, yeah.
那家人非常聪明。
That family's very smart.
是的,是的,是的。他们很棒,是的。所以我们正在做这些,随着我们在第一方智能体中获得经验,我们也在与客户一起做。我的感觉是,定制模型实际上会随着时间的推移变得越来越容易。这就是我们发现的,因为基础模型更聪明,所以它们已经在强化学习中生成更好的轨迹,而强化学习就是从自己过去的轨迹中学习。而且合成数据生成现在好得多、容易得多。我们有仅使用开源模型的流水线。比如同一个模型生成训练环境并自我训练,在某个任务上击败 Opus 和 GPT 5.5 等。所以我确实认为它会加速。训练算法的易用性只会随着时间的推移而提高。问题在于它何时进入主流。而不是像我们做的专用文档解析那样,需要一个硬核的 LLM 研究员,什么时候才能变得足够简单,任何人都可以塞进一些东西并描述一个任务。
Yeah, yeah, yeah. They're awesome, yeah. So yeah, we're doing some of that and as we get experience with these in the first-party agents, we're also doing them with customers. So my feeling is customizing models is actually going to get way easier over time. That's what we're finding because the base models are smarter, so they generate better traces in RL already, and then RL is about learning from your own past traces. And synthetic data generation is way better, way easier now. We have pipelines just using open-source models. Like the same model generates training environments and trains itself and beats Opus and GPT 5.5 and stuff at a task. So I do think it's going to pick up. The ease of training the algorithms is only going to go up over time. There's a question of when it crosses into mainstream. Instead of just the specialized document parsing thing we did, where you need a hardcore LLM researcher, when does it get easy enough that anyone can plop in some stuff and describe a task.
是的。
Yeah.
嗯,你知道什么让它变得容易吗?接口。还有统一的 API。因为显然如果它不可互操作,你就无法切换。
Well, you know what makes it easy? Interfaces. And unified APIs. Because obviously if it's not interoperable, then you cannot switch.
这就是我们在 Omnigent 和可组合智能体上看到的。比如你可以有带专用模型的子智能体,然后你可以训练整个系统。我认为那也会很有帮助。是的。
That's what we're seeing with Omnigent and the composable agents. Like you can have sub-agents with specialized models, and then you can train the whole thing. I think that'll help a lot, too. Yeah.
最后我想说的一点,实际上,我在安排这个顺序,所以我还挺自豪的。Satya 在谈论这个。几周前我在 Microsoft Build 采访了他。然后他写了那篇文章,我相信你看过,关于整个移动构建前沿生态系统。我和他交谈时,他听起来更像一个 Databricks 的 CEO。
The last thing I was going to leave, actually, I'm sequencing this, so I'm actually kind of proud of myself. Satya is talking about this. I interviewed him at Microsoft Build a couple weeks ago. And then he wrote this essay, which I'm sure you've seen, which is the whole mobile building frontier ecosystem. He sounded, when I was talking to him, more like a Databricks CEO.
嗯哼。
Uh-huh.
有没有一种理论,比如 token 作为知识产权,构建上下文,你知道,他基本上说除了数据之外一切都是新石油,或者上下文是新石油。你们以前听过类似的说法。
Is there a theory of, I guess tokens as IP, building up the context, you know, he basically said everything but data is the new oil or context is the new oil. Some version of that that you guys have heard before.
是的,我同意。我认为你拥有的数据,随着你围绕它获得更好的技术,你可以在你的领域内用它做更多事情。这甚至不仅仅是 AI。即使当人们开始实时收集数据时,我记得所有电力公司都安装了智能电表之类的东西,所有汽车制造商都开始安装传感器和摄像头。任何技术都会让数据更有价值,并给你带来一些优势。任何帮助你用它做事情并做出决策的东西。AI 也是如此。你拥有所有这些只是闲置的数据。现在你可以让一个智能体自动告诉你。例如,不是我因为客户投诉才发现产品中的某个功能坏了,而是智能体告诉我,‘我注意到没有人再上传文件了,因为他们遇到了错误之类的。’正如你在 Raiden 中看到的,作为一家数据库公司,因为我们拥有所有查询的历史记录、所有表布局以及它们如何工作,我们可以非常快速地构建一个新的引擎,它很好,我们相信它会很好。所以我认为这是对的。问题在于它具体会如何落地,但我确实认为 Sajid 谈到的定制模型会随着时间的推移变得越来越容易。
Yeah, I agree. I think that the data you have, as you get better technology around that, you can just do more in your domain with it. It's not even just about AI. Even when people started collecting stuff in real time, like I remember all the power companies put smart meters and stuff, and all the car manufacturers started putting sensors and cameras. Any technology makes data more valuable and can give you some advantage. Anything that helps you do something with it and make some decisions. And AI is the same way. You had all this stuff that's just sitting there. Now you can have an agent automatically tell you. For example, instead of I discovered a feature in my product is broken because a customer complained, the agent tells me, 'I noticed no one is uploading files anymore because they got errors or whatever.' And as you saw with Raiden, as a database company, because we have all the history of the queries and all the table layouts and how they work, we can build a new engine very quickly that is good and we're confident it's going to be good. So I think this is right. The question is exactly how it will land, but I do think custom model customization which Sajid talked about is going to get easier over time.
是的。
Yeah.
顺便说一句,这就是我提起模型问题的原因,因为他们有他们的 MEI 之类的东西,而你们没有。那是我心里想问的。
Which is why, by the way, I brought up the model thing because they have their MEI things and you guys don't. That was the mental question.
是的,我们确实有,我们正在做基于强化学习的微调即服务,与许多客户合作。我们基本上没有,你知道,我们有预览客户,我们有一个通用的东西叫 AI runtime,它就像我们按需提供 GPU 集群,里面有一个软件栈,让训练变得容易。所以我们并没有……但这已经存在一段时间了。我们已经有 GPU 计算一段时间了,很多 Mosaic 栈都用于帮助扩展它。
Yeah, we do have, we're doing RL fine-tuning as a service with a bunch of customers. We don't have basically, you know, we have preview customers and we have a general something called AI runtime that's like we got GPU clusters on demand with a software stack in there that makes it easy to do training. So we didn't sort of but that's existed for a while. We've had GPU compute for a while and that's where a lot of the Mosaic stack went to help scale that.
是的。
Yeah.
但是是的,我们发现这些合作,有些……有两种类型的客户。有些只想要 GPU 和库来输入输出数据并进行监控。这就是 AI runtime 的作用。然后还有一些人说,‘嘿,你能和我一起工作,构建评估、构建合成数据吗?’
But yeah, we found that the engagements, some of the... There's two types of customers. Some who just want GPUs and libraries to get data in and out and monitors. So that's what AI runtime is. And then there's some that say, 'Hey, can you actually work with me, build evals, build synthetic data,'
是的。更像是前部署解决方案架构师。
Yeah. The more forward deployed solutions architects.
这就是我们在做的。随着更多事情从定制过渡到非定制。但这就是今天的状况。
That's what we're doing. And as more things will transition from being custom to not. But that's sort of how it is today.
回到最初的问题,我认为我们的一个论点是,一旦你把数据放到合适的位置,AI 模型就会变得相当不错。通用智能体已经相当——我记得你提到过 AGI 已经到来。它们有相当不错的推理能力。实际上,我认为许多传统软件将被这种新范式重写,也就是把数据准备好,然后在上面加个智能体,魔法就会发生。
Going back to the original question, I think one of the theses we have is actually that once you get the data in the right place, the AI models are becoming pretty good. The generic agents are fairly—I mean, I think I heard you talking about AGI is already here. They have pretty good reasoning capabilities. Actually, I think many of the traditional software will be sort of rewritten with this new paradigm, which is just get the data to be there, and then they slap some agent on top. Magic will come out.
是的。
Yeah.
但没有正确的数据,你做不到这一点。这实际上是我们进入安全和客户数据平台领域的方法。
But without the right data, you can't really do that. And it's actually our approach going to security and our approach to going to the customer data platform space.
是的。
Yeah.
我们在 Data and AI Summit 上发布了两个产品,一个针对安全团队,另一个针对营销团队。这些领域已经有很多现有技术。我们的方法就是:“嘿,一旦你把数据导入,有了智能体在上面,一切都会容易得多。”
We launched two products at Data and AI Summit. One targeting sort of security teams and the other one targeting marketing teams. And those all have a lot of existing technologies out there. And our approach is just, "Hey, once you get the data in, everything is a lot easier with agents on top."
是的,是的。
Yeah. Yeah.
你们是太棒的嘉宾了。我真的很喜欢这次讨论,喜欢既能深入技术层面,又能探讨文化和战略。希望这不是我们最后一次聊天。祝贺你们迄今为止取得的成功。
Well, and you guys have been fantastic guests. I just love this discussion. I just love the ability to dive in on the tech side, but also culture and strategy. I hope this isn't the last time we chat. I mean, congrats on all the success so far.
谢谢,也祝贺你的成功。
Thank you. Congrats on your success also.
嗯。
Yeah.
是的。David 实际上在支持我的活动,我运营一个会议。我参加 Data AI Summit 很久了,我记得在 2022 年,它大概是 90% 数据和 10% AI。我当时就想,“好吧,我们需要一个 90% 都是 AI 的社区活动。”不是所有人都这样。
Yeah. I mean, David's actually supporting my event, which is a—so I run a conference. And I've been an attendee of Data AI Summit for a long time. And I noticed that it was like, this is back in 2022, it was like 90% data and then 10% AI. And I was just like, "Well, okay, we need a community thing that is like just 90% AI." Which not everybody is.
是的,这样很好。
Yeah, yeah, that works.
所以 Databricks 会参加这个会议。看到你们打造出除三大云之外最有趣的云,真是令人惊叹。你们发展得太快了。最深刻的一次——我不是 VC,但在电视上扮演过。比如 Ben Horowitz 跟你们讨论公司方向时,他说“不到 1000 亿不要卖”,大概是这个意思,对吧?
So yeah, Databricks will be at the conference. And it's just amazing to see you guys build out the most interesting cloud that I have ever seen outside of the big three. It's amazing how far you've grown. One of the most insightful—I'm not a VC but I play one on TV. Like Ben Horowitz when he was talking to you guys advising you on just like where is this company going? He was like, "Don't sell until 100 billion" or some version of that story, right?
他说公司应该值一万亿美元,你 100 亿就卖是低估了。
It was like the company should be worth a trillion dollars. You're underselling it for 10 billion.
他可不是对每个人都这样说的。
And like he doesn't do that for everyone, you know.
出于某种原因,我觉得他看到了愿景,也看到了你们无限的跑道。
For some reason, I think he saw the vision but also the infinite runway that you have.
我们很幸运有 Ben,他是大力支持者。
We're lucky to have Ben. Yeah, he's a big supporter.
太棒了。好的,非常感谢。
Yeah, amazing. Okay, well, thank you so much.
好的,非常感谢,Sweaks。
All right, thank you so much, Sweaks.