AI 公司为何想要放缓?Ali 谈前沿节奏

Why Would AI Companies Want to Slow Down? Ali on Pacing the Frontier

阿里·戈德西 Ali Ghodsi · The a16z Podcast · 2026-09-18 · 约 67 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Ali 探讨 AI 发展节奏的争论,主张领导者应避免不必要的生存恐惧,同时承认真实风险和政治动态。

Ali discusses the debate over AI development pace, arguing that leaders should avoid unnecessary existential fear while acknowledging real risks and political dynamics.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 22)

全文 · Full transcript(中英对照)

引领前沿:领导者的责任 Pacing the Frontier: Leaders' Responsibility

Host

谢谢你来到这里,Ali。

Thank you for being here, Ali.

Ali

非常兴奋。

Super excited.

Host

我们当然想聊 DataBricks,但眼下关于 AI 有一场更广泛的讨论,Dario 发表了看法,Yakov 发表了看法,Elon 也发表了看法,但我们想听听 Ali Ghodsi 怎么看——如果你把这个话题宽泛地称为“为前沿模型定节奏”之类的。嗯,你对目前外界说法最认同的是什么?你在哪里不同意?也许哪里还有没被捕捉到的细微之处?

So, we obviously want to get to DataBricks, but there is a broader conversation going on right now about AI and Dario's weighed in, Yakov's weighed in, Elon's weighed in, but we want to hear what Ali Ghodsi thinks in terms of, you know, if you called the topic broadly speaking pacing the frontier, etc. Um, what is your strongest agreement with what's being out there? Where do you disagree? And maybe where is there nuance that's not being captured?

Ali

好,我很乐意聊。我和 Martin 经常争论。所以,我肯定这不会花太久。

Yeah, happy to cover it. And me and Martin argue a lot. So, I'm sure that's not gonna take long.

Host

我们这次会尽量控制一下。

We'll try and we'll try and rain it in this time.

Ali

尽量保持冷静。但首先,我确实认为——也许我们在这一点上是一致的——领导者有责任不要毫无必要地吓唬人,除非有非常非常充分的理由。而且你知道,社会上总有人处在不同的心理状态。所以,谈论这类生存风险,以及全人类将被消灭的场景,我认为是不负责任的,它可能让很多人崩溃,也可能引发很多心理健康问题。

Try to stay calm. But, well, I do think first and foremost that there maybe we agree on this that leaders have responsibility to not freak people out unnecessarily unless there's really really good reason. And I think you know there's always different people in society that are at different places you know in their mind space. So, you know, talking about these kind of existential risks and, you know, scenarios where all of humanity is going to be wiped out, I think, is irresponsible like it can tip a lot of people over and it can cause a lot of like mental health issues

Host

除非你真有东西会把人类消灭。

Unless you have something that's going to wipe people out.

Ali

是的,就像我说的,如果真有实际理由,那就是另一回事。但我认为目前生存风险接近于零。所以为什么要吓唬所有人?这其实没必要。确实有风险,我们会谈到。那可能就是我们分歧所在。但首先,我认为领导者不应该吓唬所有人。我的意思是,如果 AI 研究方式有技术上的细微差别,研究人员可以讨论。你不需要每次都上电视或在 Twitter 上向数百万人广播说,嘿,我认为有 10% 的风险全人类会被消灭。我不认为这对很多人有帮助。实际上,我认为这给很多人带来很大伤害,他们会感到压力,而他们并不了解这些事情的细微之处及其含义。所以我认为我们不应该这样做。我不认为这有成效。它其实帮不了任何人。

Yeah, as I said, yeah, if there is an actual reason for it, then, you know, that's a different story. But I think that right now the existential risk is close to zero. Um, so why freak everybody out? It's not actually needed. There are risks. We'll get into it. That's probably where we disagree. Um, but first and foremost, I think that leaders should not freak everyone out. And I mean, you know, if there's like technical nuances in how we're doing AI research and so on. Well, researchers can discuss that. You don't need to every time go on TV and or blast on Twitter to millions of people that, hey, you know, I think there's like this percentage 10% risk that all humanity is going to be wiped out. I don't think that's like helpful for a lot of people. Actually I think causes a lot of harm for a lot of folks who get stressed out and actually are not in the nuances of all of this stuff and what it means. So that I don't think we should do. I don't think it's fruitful. It doesn't really help anyone.

Host

我觉得对普通公众来说,这非常非常正确。

I mean I think this is very very true for the general public.

Ali

是的。

Yeah.

Host

比如我妹妹,她很棒,是亚利桑那州乡村的一名教师,周日她发短信问我:“我要不要给你准备小屋?”她本来就是个末日准备者,但她问:“我要不要为 AI 末日准备小屋?我已经备好水了。你什么时候来?”我说:“等等。”

Like my sister who's great who's a school teacher in rural Arizona on Sunday texted me and said should I prepare the cabin for you? She's kind of a prepper anyways, but should I prepare the cabin for the AI apocalypse? You know, I've got water set up. Like, when are you showing up? I'm like, hold on.

Ali

是的。

Yeah.

Host

我们还没到那一步。所以,这显然已经蔓延到民众中,我同意这是不必要的,而且有反效果。我认为还有第二点,嗯,我不知道你进来时有没有看到,我当时在看 X,Elizabeth Warren 刚刚谈到暂停所有 AI 开发。这当然是紧随 Bernie 之后,他还和 Bannon——Steve Bannon——合作。所以现在,除了吓唬人之外,联邦体系也开始动起来了,我认为这可能与信息本身的目标背道而驰。所以,我觉得这里赌上的不只是公众的恐慌。

Like, we're not there yet. So, clearly this has kind of spilled over the populace, which I agree is unnecessary and has blowback. I think there's a second one which is um I don't know if you saw like walking in here, I was checking X and Elizabeth Warren just talked about um pausing all of AI development. That of course is on the coattails of Bernie who is also working with Bannon like Steve Bannon like so now so I in addition to like you know just scaring people the federal complex is now spinning up and I think that could be actually quite contrary to the actual goals of the message and so there's more than just you know I think public hysteria at stake here.

Ali

是的,有很多政治操作,但我身处所有这些群体中,你知道,我看到双方。我们应该说,双方都有激烈的政治操作,对吧。

Yeah, there's a lot of politics going on, but I'm like in all of these groups and you know I see both sides. There's there's heavy politics happening on both sides, we should say like right.

Host

哦是的,双方都在发生。

Oh yeah, this is happening both sides.

Ali

不,这是——我认为两党,不包括 Trump 本人,都同意 AI 应该在某种程度上受到约束。

No, this is this is no this is a this is a I think both parties that do not include Trump himself agree that um that AI should be constrained at some level.

Host

我也在说这场争论的另一方。让我举个例子。

I'm talking about the other side of this argument as well. Let me give you an example.

Ali

甚至 Greg Abbott,对吧?连 Greg Abbott 都说,你知道,德州不能有数据中心。

Even even Greg Abbott, right? Even Greg Abbott was like, you know, you can't have data centers in Texas.

Host

我不是在说政客。我是说双方都有政治操作,对吧?商业这边也有政治。那些人想要漂亮的 IPO,想要投资回报,

Well, I'm not talking about politicians. I'm talking about there's politics going on on both sides, right? There's politics on the business side. People who want to see great IPOs and they want to get returns on their investments

Ali

他们会说,别搞砸我的 IPO。

and they're like, don't mess up my IPO

Host

他们想要的是:嘿,大家能不能闭嘴,好让我们把钱收回来。所以有这种情况。他们拥有资源,并且在使用它们,所以那一方也有政治操作,他们不是安静坐着什么都不做,他们能拉关系,有人脉。另一边,有那些人想:好,我们怎么把这武器化?这太棒了。这家伙发了推文,我们把这个武器化。我们把这个植入。我们把这些帖子炒热。

and they want to get like, hey, can everybody just shut up so that we can get our money back. Uh, so there's that. and they're, you know, they have resources and they're uh using them and, you know, so there's politics on that side and those are like not they're not sitting quietly and not doing anything and they can pull strings and they have connections. On the other side, there's all the people that like, okay, how do we weaponize this? This is awesome. This guy tweeted this, you know, let's like let's weaponize this one. Let's plant this. If I, you know, let's pump these, you know, threads.

Ali

我们来谈谈大家都在用的那个具体杠杆点,因为我实际上认为这是一个典型的公关失误案例,而且不只是那些悲观末日的东西。所以,公关烂摊子在这里。我认为,行业试图自我监管并不罕见。这本身没错。说安全和保障很重要,每个技术领域都是如此,我们想要一些监督。这非常非常明智。问题在于,它被包装在“定节奏”这个概念里,而“定节奏”有若干问题。首先,它与安全和保障是正交的。你可以慢慢地造武器。那和造武器没有区别。人们不觉得它是真诚的,因为比如这些公司一直在拼命奔跑,如果

Let's let's talk to the specific leverage point that everybody is using because I actually think that this is like a classic case of a PR misstep and it's not just the doomy gloomy type stuff. So, here's the PR mess. I think which is um like it is not unusual for industries to try and regulate themselves. It's just not right. And I think saying like security and safety is important. It is with every techie puck and we want to have some oversight. That was very very sensible. The problem is is just couched in this notion of pacing and there's there's a number of issues with pacing. First off, it's orthogonal to safety and security. Like you can slowly build a weapon. That's not different than building a weapon. people don't feel it's it's it's genuine um because like these companies have been at a dead run if

Host

他们还在买更多算力,为了跑得更快。

they're still buying more compute to be even faster.

Ali

不,不。我是说,不。

No, no. I mean, no.

节奏把控与暂停 Pacing vs. Pause

Ali

我是说,他们就是没做。历史上他们就没做过,但这也感觉像是对“暂停派”一种近乎软弱的投降。所以你会说,你说暂停,我说节奏控制,这几乎就是暂停,但又不叫暂停。所以他们选了这个“节奏控制”的旗号来跟。但如果你真的读了——你读了 Dario 写的那份文件吗?那是一份完全合理的文件。

I mean, they just haven't done it. They haven't done it historically, but it also kind of feels like this almost milquetoast capitulation to the pause people. So you're like, well, you say pause, well, I say pacing, which is almost like pause, but it's not like pause. So they chose this kind of flag to follow around pacing. But if you actually read — did you read the document that Dario wrote? It's a totally sensible doc.

Host

读了,我读了。

Yeah, I read it. Yeah.

Ali

它跟节奏控制根本没关系,对吧?所以我老实说——

It just has nothing to do with pacing, right? And so I honestly —

Host

不,他确实提到了。你看,我有点不太同意。你看,这里有个公地悲剧。就是那种,嘿,如果你想停下来,如果你想慢一点,你为什么不自己慢下来?你为什么要写文章?有很多人在提这个论点。但不,我是说,作为一个商业领袖,我理解这里有公地悲剧。我在竞争。我想赢,你知道,而且你也是市场均衡的一部分,这说明节奏控制可能本来就不切实际。是的。所以我只是说,人们说“嘿,如果你们不停下来,这种公地策略就会继续下去”是有道理的。我不会停止竞赛,因为,你知道——

No, he does mention it. Look, I kind of a little bit disagree. Look, there's a tragedy of the commons. There's this like, hey, if you want to stop, if you want to go slower, why don't you go slower? Why do you write articles? There's a lot of people making that argument. But no, I mean, as a business leader, I understand that there's a tragedy of the commons. Like, I'm competing. I want to win, you know, and you're also the market equilibrium, which suggests that pacing is probably impractical anyways. Yeah. So I'm just saying that it makes kind of sense for people to say, hey, if you guys don't stop, this strategy of the commons is going to continue. I'm not going to stop racing because, you know —

Ali

有 IPO 在赌,有竞争在赌。人与人之间也有一些敌意。所以,我不会单方面停下来。那样我就成傻子了。你知道,你为什么不先停?所以然后他们就说,嘿,你能不能进来管管我们?嗯,但你知道,我觉得,呃,你也可以提出这样的论点:如果你看看 Hugging Face 和 OpenAI 那件事——顺便说,我觉得这些公司很棒,我觉得他们可能投入了大量资源——但如果你读了发生了什么,就很清楚,呃,他们并没有监控输出的每一个 token,他们只是跑这些强化学习实验,然后事后才进来检查发生了什么。所以他们本该控制节奏。在那个具体事件里,他们本该慢得多,对吧?

There's IPOs at stake. There is a competition at stake. There's also some animosity between the people. So like, I'm not going to stop unilaterally. I'll be a sucker. You know, why don't you stop first? You know, so then they're saying, hey, can you come in and stop us? Um, but you know, I think that, uh, you could also make the argument that if you look at the Hugging Face–OpenAI incident — that, uh, by the way, I think these companies are great and I think they are probably investing a lot of resources — but it's very clear from, if you read what happened, is that, uh, they weren't monitoring every token coming out and having it, you know, they were just like running these RL experiments and then after the fact coming in and checking what happened. So they should have paced. They should have been much slower in that particular incident, right?

Host

我只是不想在措辞上抠字眼,但你看,公关上措辞很重要,对吧?那我们就拿 Hugging Face 那件事来说。我读到的时候,你知道我的反应是什么吗?不是“哦,OpenAI 应该控制节奏”。而是,老兄,把你的东西保护好。就像我们一直做的那样,做安全控制,就像——

I just don't want to quibble on syntax, but like, words matter with PR, right? So let's take the Hugging Face incident. When I read that, you know what my reaction was? It was not, oh, OpenAI should pace. It's like, dude, secure your thing. Like, do security controls like we always have done, like —

Ali

但那确实是节奏控制。那就是节奏控制。

It is pacing though. It is pacing.

Host

那不是节奏控制。就像,在互联网的历史上,我们有过所有这些事,比如,让我们控制互联网的增长节奏,让我们做安全,让我们做管控,让我们做——我是说,你可以——

It's not pacing. Like, in the history of the internet, we had all of these things, like, let's pace the growth of the internet, let's do security, let's do control, let's do — like, what I mean, you could —

Ali

你应该那样做,但从某种意义上说,那就是节奏控制——你看,我一直在面对这个。我在 Databricks 有法务部门。我在 Databricks 有安全部门,你知道,他们总是说,嘿,把所有事情都放慢,不是 AI 的事,是字面意义上的每一件小事。比如,哦,你要上播客?那脚本是什么?你要说什么?你知道,我们来审一下。法律上怎么说——你不能说这个,你可以说那个,你可以说那个。你知道,你说的每句话都必须实质真实,你不能——

You should do that, but it is pacing in the sense that — look, I face this all the time. I have a legal department at Databricks. I have a security department at Databricks, and you know, they're always like, hey, slow everything down for everything, not AI, like literally every little thing. Like, oh, you're going to go on a podcast? Well, what's the script for it? What are you gonna say? And you know, let's review that. And you know what's the legal — you cannot say this, you can say that, you can say that. You know, everything you say has to be materially true, you cannot —

Host

他现在就在给这屋里的人做节奏控制。

He's pacing people in this room with us right now.

Ali

所以你知道,你在跑一个强化学习实验,你在训练下一个模型。安全团队应该在场,运行所有监控,检查一切吗?我是说,数百万 GPU 小时的 token 被生产出来,这些智能体在沙盒里跑模拟。如果我们让安全团队坐在那里检查所有东西,那会显著拖慢他们。

So you know, so you're running an RL experiment, you're training the next model. Should the security team be there and look at, like, run all their monitors and look at everything? I mean, like millions of hours of GPU hours of tokens were produced and these agents were running, you know, a mock in the sandboxes. It would have slowed them down significantly if we had a security team sit there and look at all the stuff.

Host

我同意。

I agree.

Ali

现在,他们说,嘿,你知道,如果我们那样做,会拖慢我们,而我不确定另一边是不是也在这么做。所以你们能不能进来让我们慢下来?就像,直接告诉我们,给我们设一些护栏。我们就会乐意遵守规则,做安全的事。嗯,否则就没道理,因为我们会——

Now, and they're saying that, hey, you know, if we do that, it'll slow us down, and I'm not sure the other side is doing that. So can you guys come in and slow us down? Like, just tell us, like, put some guardrails around us. We'll happily then follow the rules and do the secure thing. Um, otherwise it doesn't make sense because we'll get our —

Host

我只是觉得,当人们真的很害怕的时候,细微差别、二阶措辞是没用的。你说,我要控制节奏,因此——或者这类事情不会发生。我真的觉得我们本该直接说,迅速行动,安全与保障至关重要。我们要加这些控制。这才是重要的。而且我确实觉得,如果你看看——那种细微差别其实丢失了。

I just think like nuance, second-order words don't work when like people are really afraid. You're like, I'm going to pace and therefore — or things like these don't happen. I literally think we should have just been like, swiftly, safety and security is paramount. We're going to put in these controls. Like, that's the important thing. And I do think that nuance actually got lost if you look at —

Ali

扎克说的话。你同意——

What Zuck said. Do you agree with —

Host

我觉得扎克做的是,嘿,我们要控制自己的节奏。我们要加入安全——比如,我们晚发布这个的原因就是安全。

I thought Zuck did like, hey, we're going to pace ourselves. We're going to put in SEC — like, the reason we released this later is because of security.

Ali

嗯,听着,我觉得扎克很棒的地方在于,他非常专注于,比如,安全,嗯,保障和自我监管。Dario——Dario 开头五个词左右就是,我们需要控制前沿的节奏,对吧?这跟另一种说法让你进入的心态完全不同,他本可以说,我们需要保障前沿的安全。

Well, listen, what I thought was so great about Zuck is like, he was very focused on like, like security, um, safety and self-regulation. Dario — Dario's first five words or whatever are like, we need to pace the frontier, right? It just puts you in a very different mindset than what he could have said, which is, we need to secure the frontier.

Host

好吧,我们需要前沿的安全。我是说,某种程度上我觉得他们想同时讨好末日论者——那些人呼吁暂停——和政客,结果两边都没满足。

Fine, we need safety at the frontier. I mean, at some level I think they were trying to optimize both for the doomers, which cause for pause, and for politicians, and they kind of didn't satisfy either.

Ali

但那些人其实在实验室内部都吓坏了,顺便说,有很多安全人员是真心害怕的,而且不是所有人都是 EA 那类人。人们会说,嘿,他们很惊讶,对吧?

But those people are actually freaking out inside the labs, and there are a lot of safety people that are freaking out genuinely, by the way, and not all of them are EA people and so on. People are like, hey, they're surprised, right?

Host

对。

Right.

Ali

但问题是,用“节奏控制”这个最糟的词对谁都没帮助。我觉得就像,你字面上是在试图找这个——他——“节奏控制”就像恐怖谷,让末日论者不高兴,也让政策圈的人不高兴,因为末日论者会说,那不是暂停。这是节奏控制。是的。

But here's the thing, is like, using the worst pace doesn't help either of them. I think it's like, literally you're like, you're trying to find this like — he — pace is like the uncanny valley of making the doomer people unhappy and the policy people unhappy, because the doomer people are like, that's not a pause. This is pacing. Yeah.

Host

而且你知道,其他所有人都会说,这又不是真的。你反正也不会做,而且你也不专注于安全。所以,再说一次,抛开我们该做什么不谈——那个我们该聊——我只是觉得它的呈现方式就是很糟,就是没起作用,这就是为什么会有反弹。

And you know, everybody else is like, well, like, this isn't real. You're not going to do it anyways, and you're not focused on security. So, so again, independent of what we should do, which we should talk about, I just think that the way it was presented was just bad, and it just didn't work, and that's why we're having the blowback.

Ali

这些人,你知道,呃,他们不是训练有素的,呃,你知道,公关人员,你知道,是的,有些东西我同意——我同意。我是说,我同意——

These guys are, you know, uh, they're not trained, uh, you know, PR people, you know, and yes, some of the stuff I agree — I agree with. I mean, I agree with —

Host

训练有素的公关人员。

Trained PR people.

Ali

我同意核心前提,我们不该吓坏公众。我认为现在存在性风险接近于零。嗯,你知道,呃,但让我们谈谈核心问题,也就是,嗯,这个事实:任何人在跑大规模强化学习、给它一个奖励函数的时候。

I agree with the core premise that we shouldn't freak the public out. I think that existential risk right now is close to zero. Um, you know, uh, but let's talk about the core thing, which is, um, the fact that, you know, anyone who's doing big reinforcement learning runs and they're giving it a reward function.

释放智能体与网络风险 Unleashing Agents and Cyber Risks

Ali

所以释放出来,比如说,嘿,这里有 1 万个智能体,这里有 1 亿美元,让我们把它们并行放在一个巨大的集群上运行一两个月,尝试解决任何问题。它不一定是安全方面的事情。它可以是做任何事情,比如解决这个数学谜题。非常糟糕的事情可能会发生。非常糟糕的事情意味着东西被黑。而且,你知道,网络是首要的。

So unleashing, you know, saying hey here's like 10,000 agents and here's $100 million, let's put them in parallel and let them run on a gigantic cluster for a month or two, try to solve anything. And it doesn't need to be a security thing. It could be like do anything, you know, solve this math puzzle. Really bad things can happen. Really bad things meaning things get hacked. And it has, you know, cyber is the primary one.

Host

对,这是真实的。

Right, that is real.

Ali

对,我认为把它说成是存在性风险等等,我认为这是个错误。我认为那样吓唬公众不好。嗯,我认为这已经成为每个人,不只是你的姐妹,地球上每个人现在都在谈论的事情。我遇到过各种各样从未关心过这类事情、觉得这极其无聊的人,他们联系我说,你到底怎么看待这件事?这在我的账户里非常重要。现在我开始担心了。所以这就变成了一个政治问题,我们这里即将举行选举,但全世界都有选举。所以你会看到世界其他地方也不会坐视不管。但我认为我们的责任是以平衡的方式谈论这个问题,并真正揭示风险。我认为超级智能,那本书里的那个想法,非常非常遥远。我没有看到任何证据表明我们实际上正在朝那个方向前进,或者那会发生。

Right, I think of this making it hey this is an existential risk and so on, which I think was a mistake. I think it's not good to scare the public that way. Um, I think it's become something that everyone, not just your sister, everybody around the planet is like now talking about. I've had all kinds of people that never care about this stuff and they find this extremely boring, uh, ping me and say what do you really actually think about this? This is really important in my account. Now I'm starting to worry about it. Uh, so then it becomes a political issue and we have elections here coming up, but there's elections all around the world. Uh, so you're going to see they're not going to sit still in other parts of the world either. Uh, but I think that's our responsibility to talk about this uh in a balanced way and actually expose the risks. I think that superintelligence, that idea from that book, is very very far away. I don't see any evidence that we're actually marching towards that or that's going to happen.

Host

显然有些人……是的,显然实验室里有些人吓坏了,认为也许在这方面有进展。

Apparently some people... Yeah, apparently some people at the labs are freaked out that maybe there's progress towards that.

Ali

我认为这来自 RSI,递归自我改进,模型在自我改进。嗯,我很想了解有多少,他们看到了什么我们不知道的东西。嗯,你知道,有四个标准。如果这四个条件正在发生,我很想了解它们。第一是我们的模型,嗯,如果最终出现以下四个条件同时发生的情况,即下一个模型需要更少的资源、更少的 GPU 来训练,而且是超线性的,不只是微小的减少。下一个模型训练所需的时间也更少。所以第二个条件。第三,模型的准确性,智能在增加。

And I think it comes from RSI, recursive self-improvement, the model is improving themselves. Uh, I would love to understand how much, what have they seen something we don't know. Uh, there's you know kind of four criteria. If there, if those four things are happening, I would love to understand them. One is our models, uh, the ne— if we end up in a situation where following four conditions are happening, which is the next model require less resources, less GPUs to train, and you know super linearly, not just like tiny little bit. The next model uh, you know takes less time to train as well. So the second condition. Uh, third accuracy of the model, the intelligence is increasing.

Host

第四,我们可以一次又一次地重复前三个,你知道,不只是……

And fourth we can do the former three again and again and again in a, you know, it's not just...

Ali

所有这些都是同时发生,对吧?

All all of those at the same time, right?

Host

同时发生,不只是其中任何一个。

All at the same time, not just any of them.

Ali

是的,所有四个。如果所有四个都在发生,嗯,那么你可以想象,因为其中任何一个不发生,比如如果资源是恒定的,那也没关系,因为我们会用完硬件,所以它会自我调节。就像我们没有足够的硬件来做那件事,没有足够的 GPU,对吧?时间也一样。所以需要你最终处于这种情况。所以如果,如果你只是说软件在自我编写,嗯,我们今天已经在那里了。就像 Databricks 中 90% 多的软件是由 AI 编写的。如果最后几个百分点也由 AI 编写,有关系吗?不,真的没那么重要。但如果你得到这四个条件,那么你可能会得到加速,下一个模型,比如说,花费一半的时间和一半的资源,而且它更智能,你继续这样做,你知道,那么你可能会最终处于一种情况,我不知道,顺便说一句,我甚至不知道那是否必然导致超级智能本身。

Yeah, all four. If all four are happening, uh, then you can imagine in a way where you can, you know, because any of them does not happen, like for instance if resources is constant, then that's okay because we're going to run out of hardware, so then it'll pace itself. Like we will not have enough hardware to do that, not enough GPUs, right? Uh, time the same. So it needs to be that you end up in this situation. So if it's, if you just mean that the software is writing itself, uh, we're already there today. Like 90-some percent of the software in Databricks is written by AI. Does it matter if the last few percent is also written by AI? No, it doesn't matter really that much. Uh, but if you're getting these four conditions, then you might get a speed up where the next model, let's say, takes half amount of time and half the resources and it is more intelligent and you keep doing that, you know, uh, then you might end up in a situation where, I don't know, by the way I don't even know if that necessarily leads you to superintelligence per se.

Host

仍然会收敛,是的。

Still converge, yeah.

Ali

但它可能,所以那样会更危险。所以如果他们能分享所有数据,我们能对此有所了解和透明,那就好了。

But it could, so then that would be more risky. So it would be nice if uh, they can share all that data and we can shine some light and transparency on that.

Host

我实际上认为你的分解很棒。

I actually think it's a great breakdown that you have.

Ali

我不认为实际上有人用那个作为定义,对吧?我认为人们,我认为有一小部分人对诸如“天哪,涌现行为,现在它在创造自己”之类的事情感到恐慌。但我想正如我所说,很多人的定义只是,嘿,如果我不再编码了,它在自己编码。

I don't think anyone is using that as a definition actually, right? I think people, I think there's a little bit of people freaking out about like oh my god emergent behavior now it's creating itself and so on. But I think like as I said a lot of people their definition is just hey if I'm not even coding anymore and it's coding itself.

Host

对。

Right.

Ali

嗯,但我认为他们混淆了,嘿,我的价值是什么,这对我来说可怕吗,与嘿,那意味着我们将获得 2014 年理论上由博斯特罗姆假设的超级智能。

Uh, but I think they're conflating hey what's my value and is it scary for me versus hey that then means we'll get that superintelligence that 2014 theoretically was uh hypothesized by Bostrom.

Host

不过你关于算力的观点非常好,我认为在 RSI 的很多争论中被忽略了,对吧?因为据我们所知,训练一个好模型所需的最低算力门槛一直在上升。就像以前是 1 亿乘以 10 亿,现在可能是 50 亿。嗯,所以那……

Well your point in compute though is a really good one that's missed in a I think in a lot of arguments on RSI right? Because as far as we can tell, the minimum threshold for compute needed to train a good model just keeps going up. Like it was 100 million billion now it's probably like 5 billion. Um and so that...

Ali

你现在训练一个模型……

You train a model now...

Host

训练一个前沿模型。对。

To train like a frontier model. Right.

Ali

50 到 100 亿。

5 to 10 billion.

Host

对。对。没错。对比 10 亿乘以 10 亿乘以 10 亿乘以 10 亿。是的。

Right. Right. Exactly. Versus billion billion billion billion. Yeah.

Ali

对比非常昂贵。

Versus very expensive.

Host

是的。前沿非常昂贵。

Yeah. Frontier is very expensive.

Ali

6 个月后复制前沿模型的成本大约是 1/20。不,我认为 Sarah 有一个很好的观点,那就是这是反对整个事情的一个好论据,即下一个模型首先每年每个实验室只进行一两次这样的运行,而且它们需要——这与我提到的四个标准相反,对吧,即它将需要更多的资源、更多的人参与,而且更加脆弱,他们必须建立数据中心。我的意思是,实验室不一定这样做,但其他人必须建立数据中心,它们必须是巨大的,他们必须获得 GPU,他们必须把网络做好,他们必须做工程以确保能够容忍,因为你知道,每增加一个数量级的 GPU,你现在就必须担心以前不必担心的错误。所以你必须提高鲁棒性。所以这是一个非常脆弱的过程,如果失败,你就浪费了这么多钱。所以他们对那次运行非常非常小心,而且已经有过多次运行搞砸了。

To replicate the frontier 6 months later is about 1/20th the cost. No, I think Sarah has a great point which is it's this is a good argument against this whole thing which is that the next model first of all there's only one or two such runs a year that each of these labs do and they take it's the opposite of the four criterias that I mentioned right which is it's going to take more resources more humans involved and it's even more brittle and they have to build out the data centers I mean like the labs are not necessarily doing that but others have to build the data centers and they have to be gigantic and they have to get the GPUs and they have to get the networking right they have to do the engineering to make sure that they can tolerate because you know every order of magnitude more GPUs you cram in there, you have to now worry about errors that before you didn't have to worry about. So you have to increase robustness of the So it's like a very brittle process and if it fails, you've squandered so much money. So they're like very very careful with that run and there's been multiple runs that have been botched.

Host

所以它是相反的,嘿,下一个模型更快、更便宜、更智能并且递归改进。它是相反的。就像它需要更长时间,更脆弱,需要更多人。

So it's it's the opposite of that that hey the next model is faster, cheaper, smarter and recursive improve. It's the opposite. It's like it's taking longer and it's more brittle and it's more people.

Ali

而且更难实现。所以,嗯,我确实认为那是真的。

And it's harder to pull off. So, um, I do think that that is true.

Host

关于 RSI,关于实际的网络风险和东西被黑。我们需要非常认真地对待。是的。

With respect to RSI, with respect to actually cyber risks and things getting hacked. We need to take it super seriously. Yeah.

Ali

所以,我我我要像另一次失误一样,我实际上喜欢你的四个标准。我简直就是在等着反驳它,但我实际上认为这非常好。嗯,所以,让我给你一个黑箱,当你处理这些动态自适应系统时,你会相信什么?你会相信数字还是你撒谎的眼睛?对吧?对。所以,我认为在这些问题上你不得不依赖数字。

So, I'm I'm going to have like another like miss here like I actually love your four criteria. I was like literally just waiting to argue with it, but I actually think it's this is very good. Uh so, I'm let me give you like um like a blackbox what when like when you're dealing with these like dynamic adaptive systems like what are you going to believe? Are you going to believe like the numbers or your lying eyes? Right? Right. So, I think you kind of have to go to the numbers on these ones.

AI 进展的试金石 Litmus tests for AI progress

Host

那该看哪些数字?我真的觉得你基本上应该——也许上市才是正确的做法。如果这些公司继续增长,同时减少人数和投入的资金,那我就会说这里确实在发生一些事情。我确实认为你可以把它当作黑箱来看。但这些都没有表明他们在疯狂招人。

So what are the numbers to look at? I really think you should just basically — and maybe going public is the right way to do it. If these companies continue to grow, reduce the number of people and the amount of money that goes into them, then I would say something is definitely happening here. I do think you can actually blackbox this and take a look. But none of those indicate they're hiring like crazy.

Ali

这不公平,因为你知道公司不一定高效,对吧?所以如果你有——我是说 OpenAI 本身就在做一百万件不同的事情。一个大约 10 人的小团队在做 LLM,而 LLM 那部分是有用的。Twitter 以前有很多人,现在人少多了。

That's not fair, because you know companies are not necessarily efficient, right? So what if you have — I mean OpenAI itself was doing a million different activities. A very small team of like 10 people were doing LLMs, and the LLM stuff was useful. Twitter had a lot of people, now it's much fewer people.

Host

我同意,这只是另一个试金石。我们可以有两个试金石。我们有你的试金石,我觉得很好,但那样你就得有一种方法来度量它。然后我们还应该有黑箱试金石。我是说,听着,如果 Anthropic 两周后只有 12 个人,而他们继续增长,并且以越来越快的速度推出模型,我想我们大概应该注意到这一点。这是一个充分条件,但不是必要条件,对吧?但我只是说,可能是这样——而真正正确的做法是去看正在进行的预训练和后训练。那才是真正必要的,因为他们有那么多资源,可能会做很多他们不需要做的事,但他们做只是因为能雇到人,因为他们有无限的钱。所以真正训练下一个模型的人——那个团队是不是很小,是不是实际上在缩减,是不是他们做的工作越来越少,只是 AI 在做,还有后训练,然后他们都只是用更少的 GPU?

I agree, it's just another litmus test. We can have two litmus tests. We have your litmus test, which I think is great, but then you would actually have to have a way to instrument it. And then we should have the blackbox litmus test. I mean, listen, if Anthropic in two weeks is 12 people and they continue to grow and they're putting out models at an increasing rate, I think we should probably take notice of that. That's a sufficient criterion, but it's not a necessary condition, right? But I'm just saying it could be that — and really the right way to do this is to look at the pre-training and the post-training that's being done. That's really necessary, because they have so much resources that they might be doing a lot of other stuff they don't need to do, but they're doing it just because they can hire the people, because they have infinite money. So really, the people that are training the next model — is that team tiny, and is it actually getting reduced, and are they doing less and less work and just the AI is doing it, and the post-training, and then they're all just using less GPUs?

Ali

事实并非如此。

That's not the case.

Host

我们作为一个行业以前已经争论过很多次了。我记得当我们学会真正把计算机集群化的时候,因为大型机实际上受到像内存一致性之类的东西的限制。记得吗?你只能把它做到那么大。然后我们转向了客户端-服务器,然后我们就没有那个问题了,然后我们开始创造超级计算机,基本上就是集群计算机。

We've had this argument many times as an industry before. I remember when we learned how to really cluster computers, because the mainframe was actually kind of limited by things like memory coherence. Remember that? You can only make it so big. And then we kind of went to the client-server, and then we didn't have that problem, and then we started creating supercomputers, which were basically just clustered computers.

Ali

是的。

Yep.

Host

然后在某个时候互联网出现了,那也大致是 GPU 开始变好的时候。你还记得我们居然对 PlayStation 实施出口管制吗?因为我们担心萨达姆·侯赛因会用它们来做模拟。当时的论点非常相似,就是这些东西变得无限强大。我们用它们来模拟核武器——我们确实这么做了,就像我一样。我们不能——你知道,这东西有生存风险。其实他们没用那些词,但说这有核武器之类的潜力,我们应该阻止它。而这些都没有成真。所以我认为一个非常合理的讨论是:这次不同吗?是或否?我没有答案。

And at some point the internet happened, and that was kind of also roughly when GPUs started getting good. And do you remember that we would actually export-control PlayStations because we were worried that Saddam Hussein would use them to do simulation? And the arguments were very similar, which is like these things are getting infinitely powerful. We're using them to simulate nuclear weapons — which we were, like I was. We can't — you know, this stuff has existential risk. Actually they didn't use those words, but this has the potential for nuclear weapons or whatever, and we should stop it. And none of that came to pass. So I think a very reasonable discussion is: is this time different? Yes or no? I don't have an answer to that.

Ali

但我是 PC,但你知道,你是——

But I'm a PC, but you know, you're a —

Host

是的。我是说,看,我觉得我还没老到记得那些。所以无知是福。所以我能接受这个。得了吧。

Yeah. I mean, look, I think I'm not old enough to remember. So ignorance is bliss. So I can take this. Come on.

Ali

我不记得 PlayStation 是违法的,萨达姆·侯赛因——我就是不知道。也许我只是无知。

I don't recall PlayStations being illegal and Saddam Hussein being — I just don't know. Maybe I'm just ignorant.

Host

这大概是 1999 年。

This is like 1999.

Ali

也许我又老又无知。我是说,你知道,

Maybe I'm old and ignorant. I mean, you know,

Host

也许在瑞典他们不在乎。是不是——

Maybe they didn't care in Sweden. Is that —

Ali

也许只是年纪大了的健忘。但不管是什么,

Maybe it's just amnesia from age. But whatever it is,

Host

瑞典不在乎美国的出口管制。

Sweden doesn't care about the export controls in the US.

Ali

是的。你知道,不管是什么,我认为现在规模不同了,对吧,有了 AI,有了我们正在做的事,发展速度等等。他们在前沿领域吓坏了。我确实认为网络攻击实际上是我们将会看到的其中最大的一个,对吧?因为——

Yeah. You know, whatever it is, I think it's at a different scale now, right, with the AI and with what we're doing, the pace of development and so on. They are freaking out at the frontier. I do think cyber is actually one of the biggest ones that we're going to see, right? Because —

Host

顺便说一句,这个星球上的基础设施太多了,比当年——不管你说的萨达姆还是 Xbox 还是什么——多得多。我是说,我们只是把更多东西互联起来,而且它们相互依赖,今天这个星球从互联网技术依赖和互联的角度看,和 30 年前看起来完全不同。所以我只想说明这一点:有太多不安全的基础设施。

It's — there's just so much infrastructure on the planet, by the way, way more than it was whenever — whatever Saddam or Xbox or whatever it was you're talking about. I mean, we've just interconnected way more things and they're dependent, and the planet just looks different today from internet tech dependency interconnection than, you know, 30 years ago. So I just want to make this point: there's so much infrastructure that's insecure.

Ali

对,如果你要释放这些智能体,它们会找到漏洞,会找到利用方法,会到处闯入。所以这是一个真实的风险。你不能只是——顺便说一句,这次它没有那样做。

Right, and if you're going to unleash these agents, they're going to find loopholes, they're going to find exploits, they're going to break in here and there. So this is a real risk. You can't just — and by the way, this time it didn't do that.

Host

但你也可以想象一种情景,它开始跳跃,比如它获取资源,然后开始在别处执行自己,所以它有点像病毒一样传播。这是一个真实的风险。所以——这纯粹是好奇,我保证我不是想当反方——但你觉得为什么我们就是没看到太多呢?再说一次,我比你大得多。我记得很清楚。所以当互联网出现时,到那个时候我们已经真的拿下了 10%,我们让医院瘫痪,我们拿下了关键基础设施,我们因蠕虫造成了数百亿美元的经济损失。所有这些都已经发生了。而正如你所说,我们当时的建设少得多,你知道,经济中依赖它的部分更少。所以,你知道,AI 有那么多想找风险和威胁的人,我们跑得那么快,那么多钱被投入其中,而我们还没有看到任何与早期蠕虫相称的东西。这种脱节是什么?

But you could imagine a scenario also where it starts hopping, like it takes resources and it starts executing itself elsewhere, so it kind of spreads like a virus a little bit. That's a real risk. So — and this is pure curiosity, I promise I'm not trying to be a foil here — but why do you think we just haven't seen very much then? Again, I'm much older than you. I remember very well. So when the internet came out, by this point we had literally taken out 10%, we'd disabled hospitals, we'd taken out critical infrastructure, we'd caused tens of billions of dollars in economic damages from worms. All of that had already happened. And to your point, we had much less buildout, you know, less of the economy was on it. And so, you know, AI has so many people that want to find risks and threats, we're running so fast, so much money has been poured into it, and we haven't seen anything commensurate with the early days of worms. What is that disconnect?

Ali

是的。看,我确实记得那些日子。

Yep. Look, so I do remember those days.

Host

同一时期。

The same time.

Ali

是的。所以看,我只想说,我晚上睡得很好,我不认为现在存在生存风险。我确实认为有很多基础设施需要保护。我们在检测市场有一个产品 Lakewatch,帮助你做检测,而这个领域发展得太快了。因为你知道,以前你有这些 SOC 团队,安全运营中心的人,会查看正在发生什么入侵,我们是如何被攻击的等等,而现在人类就是跟不上。所以整个领域,安全网络空间,正在转变为使用智能体进行检测的全自动化。另一方面,如果我们不这样做——我是说现在我们正在冲刺,我们在冲刺,整个行业正在以超超快的速度冲刺去做这件事。

Yeah. So look, I would just say that I am sleeping well at night and I don't think there's existential risk right now. I do think there's a lot of infrastructure that needs to be secured. We have a product in the market in the detection market, Lakewatch, that helps you do detections, and the space is just moving so fast. Because you know you used to have these SOC teams, security operations center people, that would look at what intrusions are happening, how are we being attacked, and so on, and now the humans just can't keep up. So this whole space, security cyber space, is being transitioned into fully automated using agents for detection. On the other side, if we don't do that — I mean now we're rushing, we are rushing, the industry is rushing to do that super super fast.

自动化安全的竞赛 The Race to Automate Security

Ali

如果我们不这么做,我确实认为你会开始看到那种事情,比如网站宕机,整个系统停摆一段时间,后果会出现——不是生存危机,但经济损失和人员受伤等等都可能发生。所以我们只能非常非常快地竞速去做所有这些事。人类对正在发生的攻击响应不够快。所以你需要把所有这些自动化,而大多数组织实际上还远没做到。银行在做。一些安全意识超强的人在做,但今天大多数行业还在用老派的安全运营中心,人们每天醒来面对成百上千封触发的检测邮件。很多只是误报,你可以忽略,但有些不是。他们就是没时间逐一处理,而你需要识别出来。你需要自动化的威胁狩猎,用智能体等自动攻击自己的系统。这还没发生。所以我确实认为,如果我们只是说,嘿,这就像早期的互联网,坏事就会发生。所以一场竞赛正在进行。

If we don't do that, I do think you will start seeing those kind of things like sites going down, whole systems that stop working for a while, and there will be consequences—not existential, but economic damage and people getting hurt and so on could happen. So we just have to race very, very fast to do all of those things. The humans don't respond fast enough to the attacks that are happening. So you need to automate all of those, and most organizations are actually not close to doing that. The banks are doing it. Some of the people that are super security conscious are doing it, but most of the industry today is running with old-school security operation centers and people that are waking up every day to hundreds of emails of detections that have fired. Many of them are just false positives, so you can ignore them, but some of them are not. They just don't have time to go through those, and you need to identify that. You need to have threat hunting that's automated, where you're actually attacking your own systems automatically with agents and so on. It hasn't happened. So I do think if we just say, hey, this is just like the internet in the early days, bad things are going to happen. So there is a race going on.

数据与 AI 融入网络 Data and AI Blending with Cyber

Host

你知道吗,我其实非常惊讶。今天早上我参加了一个电话会议——我觉得你和我挺熟的。我们定期交流。我觉得我对 Databricks 了解不少。今天早上我参加了一个电话会议,一位创始人基本上说,是的,听着,我们在做所有这些可观测性、智能体威胁检测,而且我们在用 Databricks。坦白说,我甚至不知道你们有这个产品。所以从科普的角度,你们在智能体 AI 可观测性、安全、安全性方面做到多深入了?

You know I was actually very surprised. Earlier this morning I was on a call—I feel like you and I are pretty close. We talk periodically. I feel like I know a fair bit about Databricks. I was on a call this morning where a founder was basically like, yeah, listen, we're doing all of this like observability, agent threat detection, and we're using Databricks. I didn't even know that you had this offering, quite frankly. So from an education standpoint, how extensive have you gotten in the agent AI observability, security, safety thing?

Ali

是的,我们今年在 RSA 大会上和 Ben Horvitz 一起做了演讲,但问题在于数据和 AI 正在与网络安全融合。这两个市场正在崩塌,原因是——它们崩塌是因为过去情况像是,好吧,我们有数据和 AI,Databricks 这类公司过去做的那种事情,就是,你有一堆数据,你运行 AI 和机器学习,那部分独立存在,然后你有网络安全世界。网络安全世界是,我们想检测坏事——如果有坏人试图黑我们,如果坏人在做事情,我们需要检测到。但现在在数据和 AI 这边,我们有智能体在公司内部运行,人们让智能体运行,智能体也在和其他人的智能体互动,它们产生大量数据——日志、痕迹、留下的指纹。所以现在内部有这些智能体在做这些,这两个世界开始越来越融合,就像,好吧,所有产生的数据都需要分析,而你需要分析的规模比一两年前大了许多许多个数量级。所以情况发生了巨大变化。比如 2018-19 年,从 CVE 漏洞发布到你在行业中看到它被武器化,时间大约是两三年。到 2022 年显著下降,但仍然有 8-9 个月。

Yeah, I mean we gave a talk this year at RSA actually with Ben Horvitz, but the issue is that data and AI is blending with cyber. These two markets are collapsing because—and the reason they're collapsing is that it used to be like, okay, we have data and AI, the kind of stuff Databricks and these kind of companies used to do, which is like, okay, you have a bunch of data and you run AI and machine learning, and that lived separately, and then you have the cyber world. The cyber world is, you know, we want to detect if something bad—if bad people are trying to hack us, if bad people are doing things, we need to detect that. But now on the data and AI side, we have agents running internally in the company, people are having agents running, and the agents are also doing things with other people's agents, and they're producing a lot of data—logs, trails, fingerprints that are being left. And so now you have internally these agents that are doing that, so these worlds start merging more and more, which is like, okay, well, all the data that's being produced needs to be analyzed, and the scale at which you need to do that is just many, many orders of magnitude more than just one or two years ago. So things have changed dramatically. Like 2018-19, the time it would take from a CVE vulnerability being published until you see it actually be weaponized in the industry would be like two, three years. That went down to, you know, 2022 significantly, but it was still like 8-9 months.

Host

是的。

Yeah.

Ali

所以那还行——从漏洞到武器化有 8-9 个月,那是 2022 年。现在如果你看从 2022 年到现在的曲线,已经降到基本几小时。所以基本没时间了——事情立即被武器化。所以你需要用数据和 AI 平台方法自动化地做。所以这些市场,我要说,实际上将会崩塌。

So that's kind of fine—you have 8-9 months from a vulnerability, that was 2022. Now if you look at the curve from 2022 until now, it's down to basically hours. So it's down to basically no time—things get immediately weaponized. So you need to just do it in an automated way with the data and AI platform approach. So these markets, I'm going to argue, are just going to collapse actually.

Host

我同意。

I agree.

工程问题与监管 Engineering Problem vs. Regulation

Host

所以这是一个非常具体的问题。我其实觉得很多——先撇开生存风险讨论——感觉几乎有两个阵营。是的。有一个阵营相信这实际上是一个工程问题,像 Databricks 这样的公司可以解决,他们可以通过产品、工程解决方案和服务来解决,所以我们作为一个行业只需要解决那个问题。还有其他人,在我看来,实际上相信没有工程解决方案。你必须放慢速度,你必须使用监管。这更像核武器等等。所以这是否意味着你相信这是一个工程问题,还是你还不完全舒服这么说?因为等等,如果这是一个——为什么还要买 Databricks,伙计,我们干脆把东西放进国家实验室。只是放慢是唯一的解决方案。开玩笑的。大家只要放慢一点,那就没事了。坏人那边,从来没有一条探究路线能达到放慢。我不认为——我觉得要么暂停,要么解决。

So it's a very specific question. I actually think a lot of—to pull back on the existential x-risk discussion—it feels like there's almost two camps. Yes. There's one camp which believes that this actually is an engineering problem and like companies like Databricks can solve it and they can solve it through product and through engineering solutions and through services and so like we just as an industry need to solve that problem. And there's others which actually believe it seems to me that there is no engineering solution. You have to slow it down, you have to use regulation. It's more like a nuclear weapon etc. So does this mean you believe it is an engineering problem or are you not quite comfortable saying that yet? Because wait, if it's a—why would even buy Databricks, man, let's just put the stuff in a national lab. Just pace it is the only solution. Just kidding. Everybody just pace themselves a little bit then it'll be fine. The bad guys, there is no line of inquiry ever that gets to pacing. I don't think—I think it's like you pause it or like you solve it.

Ali

我们在这里学到的是,Martin 真的很讨厌“放慢”这个词。我再也不会和你用那个词了。好的。明确记下了。更沮丧的 Mark。

What we've learned here is that Martin really hates the word pacing. I will never use that word with you ever again. Okay. Clearly duly noted. More frustrated Mark.

Host

是的。这是一个可以由工程师解决的工程问题,还是还有更多?

Yeah. Is it an engineering problem that can be solved by engineers or is there more to it?

Ali

我其实觉得——我们在谈论哪个问题?我认为有两个不同的问题被混为一谈了。有超级智能问题。你知道,我认为很多来自 Bostrom 2014 年的《超级智能》一书。如果你看那些定义——我觉得人们对超级没有清晰的定义——如果你读他的书,那些定义有点疯狂。所以我认为他所说的超级智能,是指——我不知道例子是什么——比如它们能在几秒钟内写出一整篇博士论文,包含新颖的同行评审内容,它们能瞬间完成数千年的思考。所以这就是它们的速度水平、智能水平,比如它们能学习。

I actually think—which problem are we talking about? There's two separate problems that I think are being conflated. There is the superintelligence problem. You know, and I think a lot of this comes from like Bostrom's 2014 Superintelligence book. And if you look at the definitions—like I think people don't have these clear definitions of what super—if you read his book those definitions are kind of crazy. So I think what he had in mind when he said superintelligence is, you know, AIs that—I don't know what the examples were—something like they write a whole PhD thesis with novel peer-reviewed stuff in a couple seconds and they can do like millennia worth of thought in like instantaneously. And so this is like the level of how fast they are, how intelligent they are, like they can learn.

Host

是的。是的。

Yeah. Yeah.

问题的规模 Scale of the problem

Ali

这完全是好几个数量级的差别,对吧?问题的规模完全不一样。

It's just many, many, many orders of magnitude, right? The scale of the problem is just completely different.

Host

那么,如果真有这种东西存在,你觉得它只是一个工程问题吗?

So if such a thing exists, do you think it's just an engineering problem to solve?

Ali

不。如果真发生那种事,那会是非常关乎存亡的。

No. If such a thing would happen, that would be very existential.

Host

当然。

Of course.

Ali

这是大家都认同的。所以我觉得这被和现在这些智能体混为一谈了,而它们离那个还差得远。根本没有那种东西,我们现在也没有任何通往那条路的进展。但这些智能体是有能力的,你可以用它们做到人类历史上从未做过的事。所以我确实认为一个拐点已经发生了。有些东西变了。我们在 Databricks 有很好的安全研究员,但我以前绝不可能说,把一万个这样的人放进沙盒里待一个月,让他们做价值一亿美元的薪酬工作。现在我们能做到了。按一个按钮,就能弄到十万个。或者数学——我们可以说,嘿,我们想解决一个猜想。好,找些相当不错的数学家,但让一万个人协作,然后你就能取得非常快的进展。所以这引出了所有这些网络风险。我认为网络是这里的主要问题。

And that's what everybody agrees on. So I think that's being mixed with now we have agents that are nowhere near that. There's nothing like that, and we don't have anything towards that path right now. But these agents are capable, and you can do something with them that you could never do before in the history of mankind. So I do think an inflection point has happened. Something has changed. We have good security researchers at Databricks, but I could never say let's get 10,000 of them in a sandbox for a month and have them do a hundred million dollars' worth of salary wage work. We can do that now. We just turn on a button and we can get 100,000 of them. Or mathematics — we can say hey, we want to solve a conjecture. Okay, let's get pretty good mathematicians, but let's have 10,000 of them collaborate, and then you can make very fast progress. So this leads to all these cyber risks. I think cyber is the major problem here.

Ali

这个我认为可以用工程解决。而且我觉得我们正在做。很多其他人也在做。仍然有风险,但不是关乎存亡的。我觉得我们应该去做。

This, I think, you can solve with engineering. And I think we are working on it. Many others are working on it. There are still risks, but they're not existential. I think we should do it.

Ali

还有超级智能这回事。那是能写出全新的博士论文,或者在 11 维空间物理里瞬间凭直觉推理、什么都不用写下来的东西——人类做不到那种事。那种超级智能。问题是,实验室正在做的 RSI,也就是递归自我改进,会把我们带到那里吗?我们会到那里吗?那会是接下来发生的事吗?还有,那会多快发生?这才是大问题。

There is the superintelligence thing. That's the thing that could write a novel PhD thesis or reason intuitively in 11-dimensional space physics instantaneously without writing anything down — something humans can't do. That kind of superintelligence. The question is, is RSI, recursive self-improvement, that the labs are doing leading us there? Are we going to get there? Is that what's going to happen? And how fast is that going to happen? That's the big question.

Ali

他们建议说,嘿,我们应该有检查员进来看看我们在做什么。我觉得这是个好主意。让他们进去拿数据。我很乐意。问题是,谁是检查员,因为你可以堆叠,对吧?你可以把那个堆起来。

And they've suggested that hey, we should have inspectors that come in and look at what we're doing. I think it's a good idea. Have them go in there and get the data. I would love to. The question is who are the inspectors, because you can stack it, right? You can stack that.

Host

你觉得是谁?

Who do you think?

Ali

是啊。有一堆人分属两个阵营。其实,我不在乎他们是不是检查员。我不会太把他们说的话当回事,因为他们进去之前就已经打定主意了。

Yeah. There's a bunch of people that are on either camp. Actually, I wouldn't care if they're the inspectors. I would not be very impressed by what they say, because they've already made up their minds even before they would go in there.

Host

对。正是。

Right. Exactly.

Ali

但举个例子,假如 Yann——他是这项深度神经网络技术的发明者之一,先驱之一——如果他说,嘿,这里没什么可看的,没有风险——我是在转述他的话——这没什么,这个超级智能,纯属胡说,继续走,快、快、快——我们谁都不会信。

But let's say, as an example, if Yann — who was one of the inventors of this deep neural network technology, one of the pioneers — if he said, hey, there's nothing to see here, there's no risk — I'm paraphrasing him — this is nothing, this superintelligence, this is just nonsense, keep on going, go fast, fast, fast — none of us would believe it.

Host

我是在替他说话。我是说,我并不完全是——不。

I'm putting words in his mouth. I mean, I'm not exactly — no.

Ali

如果他是检查员之一,他进去看了看,出来说,嘿,我看过了,就是我说的那样,这里没什么可看的,继续走——我会感觉非常好。那会说明,好吧,我会感觉非常——或者如果他出来说,天哪——你知道,他动摇了,稍微改变了主意——那也会有很多有意思的信号。所以我觉得这归结于我们选谁当检查员,而我觉得这是个好主意。让我们找一些人来,选一组多样化的人,这样我们就能得到不同的、有细微差别的观点。

If he was one of the inspectors and he went in there and he had a look and he came out and he said, hey, I've looked and it's just what I said, there's nothing to see here, just keep going — I would feel very good about that. That would say, okay, well, I would feel very — or if he comes out and says, oh my god — you know, he's wobbling and he would change his mind a little bit — that would also have a lot of interesting signals. So I think it comes down to who we pick as inspectors, and I think it's a good idea. Let's have some of them and pick a diverse set of people so that we can get different nuanced points of view.

Host

你怎么看 Elon Musk 那种观点,就是——它不太是第三方。感觉有三种提案。OpenAI 和 Anthropic 那种是第三方。Elon Musk 那种,据我所知,是实验室之间互相交叉核查,像同行评审,像科学界那样。然后 Mark Zuckerberg 那种是自我监管,对吧?你怎么看中间这种?

What do you think about this kind of Elon Musk view, which is — it's less third party. It feels like there's kind of three proposals. The OpenAI-Anthropic one is a third party. The Elon Musk one, as far as I can tell, is the labs cross-check each other, like peer review, like you do in science. And then the Mark Zuckerberg one is police yourself, right? What do you think about this middle one?

Ali

就是他们应该互相制约、互相评估。我是说,我觉得——就像我们在拳击台上比赛,让拳手们互相当裁判。那行得通吗?不行。他们会一直喊犯规。犯规、犯规、犯规。对方一推出那个了不起的模型,就变成,啊,巨大的超级智能风险,冲浪。绝对是。他们并不负责任。当有既得利益在起作用,有 IPO 计划,这两家公司又如此竞争,而且他们之间还有这段历史——是啊,他们会对彼此非常公平的,我确信。这就是为什么你需要第三方,对吧?我是说,我们世界上为什么要有法官?为什么要有第三方?为什么人们不能自己把事情解决掉?但我是说,他们应该试试。如果他们想做,他们应该试试。但我怀疑他们不会不以多种方式偏向自己、互相评判。

That they should pace each other, evaluate each other. I mean, I think — like if we have boxing matches in the ring, the boxers should just be the judges of each other. Would that work? No. They would scream foul all the time. Foul, foul, foul. The moment the other guy puts out the great model, it's like, ah, big superintelligence risk, surf. Absolutely. They have not been responsible. When vested interests are at play and there's IPO plans and these two companies are so competitive and they have this history also between them — yeah, they'll be very fair to each other, I'm sure. That's why you need a third party, right? I mean, why do we have judges in the world at all? Why do we have third parties at all? Why can't just people figure things out between themselves? But I mean, they should try. If they want to do it, they should try. But I'm skeptical that they wouldn't just be biased in multiple ways to self-judge each other.

Host

所以我得问一下,Ali,你觉得——Elon 还说过一些话,我想是在 All-In 峰会上。他说,这是某种精心设计的四维棋,因为一方面你在说全人类都会死,另一方面你又说,嘿,你的 IPO 配售想要多少,对吧?所以我是说,那大概是更愤世嫉俗的看法,但你怎么调和这个?我是说,这种不协调让很多人困惑。你觉得那要怎么调和?

So I have to ask, Ali, do you think — something Elon also said, I think it was on the All-In summit. He was like, this is some elaborate 4D chess, because on the one hand you're saying all of humanity will die, on the other hand you're saying hey, what do you want for your IPO allocation, right? And so I mean, that is probably a more cynical view, but how do you reconcile that? I mean, the dissonance gets a lot of people. How do you think that gets reconciled?

Ali

听着,我觉得所有这些事都被混在一起了。我觉得有些人是被吓到了,我也确实觉得有些人说,嘿,如果有监管能制约我们——抱歉用这个词——那对我们有好处,对吧?那对我们有好处。但我也觉得人们有既得利益。对,这些事——通常人们总能想办法让所有这些事在他们脑子里和谐地一致起来。所以,是的,我觉得过去确实有一种倾向,总体上也会用营销噱头,说,天哪,我训练的这个最新模型太好了,简直难以置信,几乎把我吓到了,然后全世界就开始关注它。是的,一直有那种营销在发生。

Look, I think all of these things get mixed. I think there are people that are freaked out, and I do think that there are people that say, hey, if there was regulation that would pace us — sorry to use the word — that would be good for us, right? That would be good for us. But I also think that people have vested interests. Right, these things — usually people figure out a way to always get all of these things to align harmonically in their head. So yeah, do I think that there has been a tendency in the past of in general using also marketing stunts by saying, oh my god, this latest model is so good that I trained, it's unbelievable, it's almost scaring me, and then the whole world kind of starts focusing on it. Yeah, there's been that kind of marketing going on.

网络攻击与营销炒作 Cyber attacks and marketing hype

Ali

是的。但与此同时,就像我说的,从 CVE 到真正武器化的漏洞利用,时间已经从几年缩短到几分钟,就在三四年间。所以网络攻击是真实的。但这也是一种绝佳的营销手段:每当你训练一个新模型,就大肆宣扬它对世界有多大的疯狂风险。这对你有帮助,对吧?所以也许这些事并不矛盾。

Yeah. But at the same time, as I said, the time from CVE to actually weaponized exploit has been going down from years to minutes now, just in like three or four years. So the cyber attacks are real. But there's also a great marketing ploy: whenever you train a new model, make lots of noise around how much of a crazy risk it is to the world. It helps you, right? So maybe they're not in contradiction, these things.

Host

我的意思是,你和我都是搞网络的人,历史上一直有成立第三方来帮助仲裁事情的传统,对吧?比如 IETF,或者 IEEE,甚至像——

So I mean, you and I are networking folks, and there's a long history of forming third parties to help arbitrate things, right? Like IETF, or IEEE, or even like—

Ali

我知道你要说什么了。

I see where this is going.

Host

不不不。所以我的问题是:我觉得他们提出的这个提议其实非常合理。我其实同意你的看法。你可能想确保它是独立的,而目前这一点并不公平,随便吧。而且关于让谁加入会有很多争论,每个人都会不同意。对吧。但你说过,为什么我们需要法官?所以那其实是国家介入,和行业自我监管基本上是相当不同的事情。那么你认为在什么时间点,考虑联邦介入才是合理的?还是你认为现在就该考虑真正的联邦介入,而不是更多的行业自我监管?

No, no, no. So my question to you is: I think this is actually a very sensible proposal that they have. I actually agree with you. You probably want to make sure it's independent, which is not fair right now, whatever. And there's going to be a lot of arguments about who you put in, and everybody's going to disagree. Right. But you said, why do we have judges? So that's actually the state stepping in, which is quite a different thing than basically industry self-policing. So at what point in time do you think it makes sense to actually consider federal involvement? Or do you think now is the time to actually consider actual federal involvement, as opposed to more industry self-policing?

Ali

嗯,这些是非常不同的。

Well, these are very different.

Host

它们确实不同,但它们会相互渗透。比如 FINRA 并不是一个完全独立的自我——它是,但它和政府有关联。所以我认为这些事情会相互渗透。我觉得——

They are different, but they kind of bleed into each other. Like, for instance, FINRA is not like a completely independent self—it is, but it's linked to the government. So I think these things will kind of bleed over. I think it's—

Ali

你认为它们会演变成——历史上一直是行业自我监管,然后演变成监管。

You think that they evolve into—historically they've been industry self-policing, and then it evolves into regulation.

Host

如果他们声称存在生存风险——他们确实在这么说——然后又说,来监管我们、约束我们吧,我认为监管者很难说不,我们不会这么做。到目前为止他们是这么说的,但我认为这不会持续太久。不过不是 David Sacks 说过,“我从没见过哪个 CEO 要求我们监管他们。”而我最喜欢的一点是——

If they are saying there is existential risk, which they're saying, you know, and then saying come police us and regulate us, I think it's very hard for regulators to say no, we're not going to do that. So far they've said that, but I think that's not going to last very long, you know. But it wasn't David Sacks who was like, "I've never had a CEO ask us to regulate them." And my favorite thing—

Ali

而 CEO——我从没见过哪个监管者会拒绝这个。

And the CEO—I've never had a regulator that says no to that.

Host

说不。我的意思是,现实是,实际的元政治机器已经在运转了,对吧?我是说,每个人都有自己的说法。奥巴马已经发声了。这是一个重大问题。那么,你认为是否存在这样一种现实:已经太晚了?这将成为中期选举的重大议题,我们实际上会迎来强硬的联邦监管,这一切都会被暂停,你知道,进入 Anthropic,进入 DOE,而我们已经过了那个点?还是你认为我们最终能达成一种合理的自我监管?因为那些头条,你没法——

Say no. I mean, the reality is the actual metapolitical machinery is actually in motion already, right? I mean, everyone has a talking point. Obama has come out. It is a major issue. Like, do you think that there's a reality that it's too late? This will be a major issue in the midterms and we're actually going to get heavy-handed federal regulation and this is all going to be paused, you know, goes into the Anthropic, goes into the DOE, and we're past that point? Or do you think we can actually end up with a sensible self-policing regulation? Because the headlines, you cannot—

Ali

我们应该努力去做正确的事。我认为事情如何演变仍然有一些自由度,而且还有时间。

We should strive towards doing the right thing. I think there's still some degrees of freedom in how things evolve, and there's still time.

Host

是的,你说得对,大体上你有这些公司,我们往里投入了那么多十亿美元,而强化学习的运作方式是,你给它一个可验证的奖励函数,比如我们要解决这道数学题,或者这类狭窄的编程领域等等,然后我们往里面投入大量资金。你可以在那个狭窄的领域得到相当不错的结果——但这并不意味着你得到了那个超级智能。

And yeah, you're right that largely you have these companies where we're pumping in so many billions of dollars, and the way reinforcement learning works is that you give it the reward function that's verifiable, like we're going to solve this math problem or this kind of narrow area of programming and so on, and we pour in so much money into that. You can get quite good results in that narrow—that doesn't mean that you're getting that superintelligence.

Ali

不,但你甚至可以骗自己,以为更少的输入给了你更好的结果,仅仅因为你跑了那么多实验、想了那么多,对吧?但考虑到投入的资源如此之多,用这种方式做一个封闭实验其实非常困难。

No, but you can even trick yourself into thinking that less inputs are giving you a better outcome just because you're running so many experiences and thought about it so much, right? But it's actually very hard to do a closed experiment this way, given how many resources are going in.

Host

嗯,融资额正在天文数字般地增长,正如你所说。

Well, the fundraisers are going up astronomically, to your point.

Ali

是的。是的。

Yes. Yes.

Host

这就是为什么这些公司要上市,对吧?我认为否则他们会保持——我是说,作为一个经营大规模私营公司的人,我认为否则他们更愿意保持私有。他们为什么要上市?因为他们需要资本,而且他们认为缩放定律和资本是一种战略优势。所以他们才上市。但我想说,让我们回到我列出的那四件事。如果那些是真的——你知道,如果那四件事是真的,你会想知道吗?那会不会令人担忧,可能会失控?现在没有证据表明那四件事正在发生,但如果有——不,那实际上就是它的走向。

And that's why these companies are going public, right? I think otherwise they would stay—I mean, as someone who runs a private company at scale, I think they would prefer to stay private otherwise. Why are they going public? Because they need the capital, and they consider the scaling laws and the capital to be a strategic advantage. So that's why they're going public. But I would say, let's go back to the four things that I listed. If those are true—you know, if those four are true, would you want to know about it? And would that be worrisome, that that could get out of hand? Now there's no evidence that those four are happening, but if there was—like, no, that is actually where it's headed.

Ali

是的,我其实认为理解任何系统——任何自我推进的特性都很重要,我们过去对动态系统就这么做过,对吧?比如我们对编译器做过,我们对纳米技术的所有研究也做过。这一直是我们共同的兴趣。

Yeah, I actually think understanding for any system—like any sort of self-propelling property is important, and we've done this in the past with dynamic systems, right? Like we've done this with whatever compilers, we did this with all the research on nanotechnology. Like it's been a common interest of ours.

Host

我不认为这是一种新的兴趣。我只是觉得,恐惧在于这些特定的系统是元经济系统,如此复杂,以至于风险在于你其实没看到却以为自己看到了,而我认为现在很多这种情况正在发生。

And I don't think that that's new, that it's an interest. I just think the fear is that these particular systems are meta-economic systems that are so complex that the risk is crying that you're seeing it when you're not seeing it, and I think a lot of that's happening right now.

Ali

是的。但当然,如果你看到了——是的。当然,我是说,你会想知道。

Yeah. But of course, if you see it—yeah. Of course, I mean, you want to know.

Host

但可以说,这些实验室现在非常关注 RSI,那是他们下一步的方向。也许他们自己只是毫无道理地担心,就像他们当初担心 GPT-2 一样,对吧?他们当时说,GPT-2 会终结世界,结果并没有,然后 GPT-3 和 4 出来了。

But it is fair to say that the labs are now focusing a lot on RSI, and that's where they're headed next. And maybe they're just unjustifiably worried themselves, just like they were worried about GPT-2, right? Like they were like, GPT-2 is world-ending, and then it wasn't, and GPT-3 and 4 came out.

Ali

所以我不想吹毛求疵。很多时候他们说 RSI,其实说的是自催化效应,而自催化效应在我们的行业里已经存在非常非常久了。比如,没有计算机芯片你就没法制造计算机芯片。你就是做不到。

So I don't want to quibble. A lot of the times when they say RSI, they're actually talking about autocatalytic effects, and autocatalytic effects have been in our industry for a very, very long time. So for example, there's no way you can create a computer chip without a computer chip. It's like you cannot do it.

Host

任何有计算机科学学位的人都知道——编译器能写自己的编译器。

Anyone with a computer science degree—a compiler writes its own compiler.

Ali

嗯,那更接近 RSI 了,但蒸汽机也是自催化的,对吧?所以听着,我的全职工作就是面对从实验室出来创业的人,他们都说 RSI,因为每个人都说 RSI,而其中也许只有 1% 是真正的 RSI。他们更像是,我们用 AI 做数据清洗,我们用 AI 来做——

Well, that becomes closer to RSI, but like the steam engine was autocatalytic, right? So listen, my full-time job is people coming out of labs and starting companies, and they all say RSI because everybody says RSI, and maybe 1% of those are actually RSI. They're more like, we use AI for data cleaning, we use AI for making—

Host

我们来区分一下。所以你是说,它基本上是一种催化剂,在某种意义上他们在用 AI 加速事情。

Let's make the distinction. So you're saying it's basically catalyst in a sense that they're using AI to speed things up.

Ali

是自催化。是的。

It's autocatalytic. Yes.

自催化 AI 与自训练 Autocatalytic AI and Self-Training

Ali

所以我会说,从我们听到的情况来看,另一个模型——你在用一个模型来构建 GPU 内核,你在用一个模型来做数据清洗,就像我用计算机来设计计算机一样。这是自催化的,每一项技术都是如此——互联网就是自催化的,因为它让人们能够远程协作。所以我会说,这是轶事性的:90% 的算力消耗都是自催化的,这完全符合预期。

So I would say, of what we hear, another model — you're using a model to build a GPU kernel, you're using a model to do data cleaning, just like I use a computer to design a computer. It's autocatalytic, which is every tech — the internet was autocatalytic because it allowed people to collaborate remotely. So I would say this is anecdotal: 90% of the calories are autocatalytic, which is 100% what you would expect.

Host

不过这种情况已经持续一段时间了。我是说,这甚至都不算新鲜事。

And that's been going on for a while though. I mean, that's not even new.

Ali

但现在也有一个重点,就是让我们朝着——我们能不能让模型自己训练自己?有点像 Karpathy 做的自动研究,但现在他们想这么做。确实有团队在做这个。但消耗的算力远没有你想象的那么多。我觉得我其实有很好的样本,因为他们都会来跟我们交流。

But also now there is a focus on, let's move towards actually — can we get the model to train itself? This kind of like the auto research that Karpathy did, but now they want to do that. There are teams that do that. It is not nearly as many calories as you would expect. And I just feel like I actually have a good sampling of this because they all come and talk to us.

Host

对。所以也许——也许你可以当其中一个检查员。我们能不能拿到所有那些数据,让我们所有人都看看?也许这里没什么可看的。你知道,我个人不认为那四个标准正在发生的概率很高。

Right. So maybe — maybe you can be one of the inspectors. Can we get all that data and all of us look at that data? Maybe there's nothing to see here. You know, I personally don't think it's very high probability that those four criteria are happening.

Host

Brockman 这周上了播客,说我们正处于 AGI 时代。

Brockman went on the pod this week and said we're at AGI era.

Ali

是的,所以现在在他那么说之后,我问了同样的问题,现在每个人都说我们有 AGI 了。所以人们会跟随这个——但对人们来说——很长一段时间以来,当我问这个问题时,他们说 AI 大多数时候比我周围的大多数人都聪明。差不多从去年第三、第四季度开始,他们就一直这么说。然后我问他们,你们中有多少人拥有数百或数千个智能体,你在管理它们,它们以群体形式相互协调、谈判,并自动化你的生活和周围的一切?如果有,请举手。结果几乎没人举手。当然,Martin 在家里做到了。

Yeah, so now I asked the same question after he said that, and now everybody's saying we have AGI. So people follow what this — but for people — for a very long time when I asked this question, they said that AI is smarter than most other people around me most of the time. That's been almost like since Q3, Q4 last year they've been saying that. Then I asked them, how many of you have hundreds or thousands of agents that you are managing, that are coordinating with each other in swarms and negotiating and automating your life and everything around you? And if so, raise your hand. It's like almost nobody raises their hand. Of course, Martin has done that at home.

Host

不,大多数企业用的是 Microsoft Copilot。是的,这就是他们 AI 的全部了。

No, most enterprises are on Microsoft Copilot. Yeah, like that's the extent of their AI.

Ali

我交谈过的大多数企业,当我问这个问题时,他们都说,不,我们没有任何这些。所以我就问,那你们在做什么?他们在用一个聊天机器人。他们从聊天机器人那里问问题,这基本上就是一个非常非常美化的、高效的旧式谷歌搜索结果。所以它只是更快的谷歌搜索。然后编码也在发生,所以人们在用它来编码,尽管 ROI——你知道,我们可以讨论那里的 ROI。但没有任何智能体式的工作自动化了整个企业。这根本没有发生。那么为什么?我认为真正的原因,如果你仔细看,是模型已经足够聪明了。嗯,但它们没有任何组织内部存在的上下文。比如它们没有参加过每一次会议。它们不知道每个人脑子里在想什么。它们不知道所有的流程。它们不知道——每个组织里总是有那么几个员工知道一切。你知道,你去拍拍他们的肩膀,每个人都会说,天哪,如果他或她辞职了会怎样?呃,你知道,它们没有那个上下文。如果你把那个融合起来,把那个上下文给 AI 模型,就用今天的前沿模型,我认为你可以为地球上任何组织获得巨大的生产力提升。为此,我们实际上不需要更聪明的模型。所以我们不需要一个更聪明的模型,能真正解决纳维-斯托克斯方程或猜想,或在 humanity's last exam 上做得更好——就像我们需要从 60% 提高到 70%。这些都不需要。所以我认为实际上人们对某些事情非常不满——比如,哦,如果我们放慢前沿——但实际上如果前沿不前进,它实际上并不重要。我认为对于地球上绝大多数组织来说,它们在真正自动化事物和从中获取价值的采用曲线上落后太远了。但这对实验室来说将是灾难性的,因为智能的价格正在渐近式下降。我认为它每六个月下降十分之一左右。所以如果你不推动前沿,那将极大地改变他们的业务。

Most enterprises I talk to, when I ask this question, they're like, no, we don't have any of that. So I'm like, what are you doing then? They're using a chatbot. They're asking questions from a chatbot that's basically a very, very glorified efficient Google search of the old day results. So it's just faster Google search. And then coding is happening, so people are using it for coding, though the ROI is — you know, we can discuss the ROI there. But there's no agentic work that's automated the whole enterprise. That has just not happened. So then why is that? And I think the real reason, if you actually look at it, is the models are smart enough. Um, but they just don't have the context that exists inside of any organization. Like they have not been in every meeting. They don't know what's in everybody's heads. They don't know all the processes. They don't know — there's always like a couple of employees who know everything in every organization. You know, you go tap on their shoulder and everybody's like, oh my god, what would happen if he or she quits? Uh, you know, they don't have that context. And if you just fused that and gave that context into AI models, just a frontier today, I think there's so much productivity gains you could get for any organization on the planet. For that, we actually don't need smarter models. So we don't need a smarter model that can actually solve Navier-Stokes or conjectures or do better on humanity's last exam — like we needed to just go from 60 to 70%. None of that is needed. So I think actually people are very upset on some — like, oh, if we pace the frontier — but actually if the frontier doesn't advance, it doesn't actually matter. I think for vast majority of organizations on the planet, they're just so far behind in the adoption curve of actually automating things and getting value out of this stuff. But it would be disastrous to the labs because the price of intelligence is dropping asymptotically. I think it's going down by one-tenth every six months or something like that. So that would dramatically change their businesses if you weren't pushing the frontier.

Host

是的。但这是我们应该关注的,对吧?我们应该关注——你知道,有两方面——我们在这里讨论了很多成本,对吧?我们应该对一切做成本效益分析,对吧?我们在这里讨论了很多成本,比如,哦,有没有存在性威胁?有没有网络风险?有没有我们应该担心的事情等等。那是成本方面。那收益呢?我认为现在这已经成为一件公众的事情,整个公众都关心 AI,他们在问,嘿,这对我有什么好处?我能从中得到什么?似乎什么都没有。那么,他们如何达到那里?到目前为止,你见过哪些用例可能让你感到意外地好?

Yeah. But this is what we should focus on, right? We should focus on — you know, there's like two sides of — we discuss here a lot the costs, right? There's cost-benefit analysis that we should do on everything, right? We've discussed the costs a lot here, like, oh, is there like existential threat? Is there cyber risk? Are there things we should be worried about and so on. That's like the cost side. What's the benefit? And I think now that this has become like a public thing and the whole public cares about AI, they're asking, hey, what's in it for me? What am I getting out of it? Seems nothing. So, how do they get there? What are some of the use cases you've seen to date that have maybe surprised you to the upside?

Ali

是的,我是说,首先,有太多关于存在性风险等等的担忧。所以,我认为很多人只是不知道,你知道,人们实际上在做有趣事情的酷炫用例。我们有很多用例——我是说,简直令人着迷。我喜欢的一个是 Crisis Text Line。所以你知道,他们实际上和我们一起使用大型语言模型来检测青少年是否有自残、自杀的意图。

Yeah, I mean, first of all, there's like so much worry about, you know, existential risk and so on. So, I think a lot of people just don't know what, you know, cool use cases where people are actually doing interesting things. We have a lot of use cases that are — I mean just fascinating. One that I like is Crisis Text Line. So you know they actually use large language models with us to detect if teenagers want to do self-harm, suicide.

Host

哦,哇。是的。

Oh wow. Yeah.

Ali

你知道,那是一个很棒的用例,它实际上拯救生命。所以那是一家伟大的公司,那个组织在做了不起的工作。另一个有点意思的是 Omnipod,它是为糖尿病患者设计的。他们可以贴上 Omnipod,它使用 AI 来真正学习你的胰岛素释放和血糖水平,并精确释放。你知道,我不知道你是否记得人们过去喜欢给自己扎针,对吧?嗯,但现在这是自动发生的,就像,你知道,为你的身体自学习的 AI。嗯,你知道,这是一个很酷的用例。Zipline 是另一个。他们做得很棒,你知道,但当他们起步时,就像这些无人机,你知道,它们完全自动化,全部由 AI 驱动,从电池优化到路线等等一切,它们在需要的地方运送食物。

You know that's an awesome use case and it actually saves lives. So that's a great company and that organization is doing amazing work. Another one that's kind of interesting is the Omnipod, which is for diabetes patients. They can put the Omnipod and it uses AI to really learn your insulin release and your glucose levels and actually exactly release. You know, I don't know if you remember people used to like stick themselves, right? Um but this now happens automatically and it's like, you know, self-learned AI for your body. Um you know, it's a cool use case. Zipline is another one. They're doing awesome, you know, but when they started it was like these drones that had, you know, they were completely automated, all AI driven everything from the, you know, battery optimization to the routes and everything and they were delivering food in, you know, areas of need.

Host

是的。血液——给难民的血液。

Yeah. Blood — blood to refugees.

Ali

给难民的血液。是的。从非洲开始,然后到世界其他地方。所以是的,那都是——是的,这是 AI 用例,你知道,建立在 Databricks 上。所以那是个很酷的。

Blood to refugees. Yeah. Started in Africa and then elsewhere in the world. So yeah, that's all — yeah, it's AI use case, you know, built on Databricks. So that's a cool one.

AI 在药物研发中的应用 AI in Drug Discovery

Ali

但还有更先进的例子,比如一个我挺喜欢、但可能比较难解释的。我们和默克一起构建的这个模型,是一个基于 Transformer 的模型,叫 TEDDY——Transformer Enhanced Drug Discovery,就是这个名字。他们其实发表了研究,所以你可以去看看。它基本上不是预测英文里的下一个词,而是预测基因调控网络(GRN)会如何响应,它能真正检测出哪些细胞是因果性的,哪些只是反应性的。也就是说,它们只是在被动反应。因此他们可以开始把这个用在药物发现上,大幅降低针对特定疾病的药物开发成本。所以这是一个非常酷的用例。

But there's more advanced ones also, like one that I kind of like but it's harder to maybe explain. This model that we built, a Transformer-based model that we built with Merck, it's called TEDDY—Transformer Enhanced Drug Discovery—that's the name. They published the research, so you can check it out. But basically, instead of predicting the next token in English, it predicts how the gene regulatory network, the GRN, is going to respond, and it can really detect which cells are causal and which ones are just reactive. So they're just reacting. And therefore they can start using this in drug discovery and get costs down significantly for developing drugs that are targeting specific diseases. So that's a super cool use case.

Ali

这样的例子很多。你知道,我提到过 Genie。你有这个本体,就可以问任何问题。诺和诺德正在用这个。他们研发了这款 GLP-1 药物。但诺和诺德现在做的是,把它用于他们正在运行的所有试验。

There are lots of these. You know, Genie I mentioned. You have this ontology and you can ask any questions. Novo Nordisk is using this. So they built this GLP-1 drug. But what Novo is doing is now they're using it for all of their trials that they're running.

Host

哦,哇。

Oh, wow.

Ali

而且它能压缩获得洞察所需的时间——相比你做一项肥胖研究之类的——从几周缩短到几分钟。所以 AI 有很多惊人的用例。我们也不该忘记这些好处。我们想要所有这些,我们不想把这些粘贴掉。

And it can compress the time it takes to get insights—versus if you're doing an obesity study or something—from weeks down to minutes. So there are a lot of amazing use cases of AI. We should not forget these upsides also. Like we want all of these and we do not want to paste these.

Host

是的,完全正确。你说得对。

Yeah, exactly. You're totally right.

企业 AI 落地实践 Operationalizing AI in Enterprises

Host

那么他们怎么——假设你规划未来 12 个月——企业到底如何获得价值?你提到了上下文这个词,但他们怎么把它落地?

And so how do they—let's say if you map out the next 12 months—how do the enterprises actually get value? You know, you dropped the word context, but like how do they operationalize that?

Ali

这其实比大多数人以为的更难。但首先,我们必须确保组织里发生的一切都已数字化。你不可能挥一挥魔杖就实现。所以每场会议都必须转录。你必须能获取所有会议的上下文以及正在发生的一切。所有数字内容都必须喂给 AI。所以你必须构建——我们称之为本体。我们构建它,但首先,你必须收集这些。这在很多组织本身就是个问题,因为法务团队会说不要录每通电话,不要录所有东西。所以你必须以一种方式来做——

It's actually harder than most people believe. But first and foremost, we have to make sure that we have digitized everything that's happening in an organization. You cannot just have a magic wand and make that happen. So every meeting has to be transcribed. You have to be able to get all the context of all the meetings and everything that's happening. All the digital content has to be fed to the AI. So you have to build—we call it an ontology. We build that, but first and foremost, you have to collect that. That itself is a problem in many organizations because legal teams will say don't record every call, don't record everything. So you have to do that in a way where—

Host

你能给大家定义一下本体吗?因为我知道 Palantir 经常说这个词,但并不是他们拥有本体这个词。它到底是什么意思,对听众来说,应该怎么——

Can you define ontology for everyone? Because I know Palantir says the word a lot, but it's not like they own the word ontology. Like what does that mean, and for the people listening, like how should—

Ali

是的,我的意思是,你知道,本体就是指在一个组织里,所有抽象概念之间的关系——所有目标、所有部门、所有人、所有正在进行的项目——它们到底意味着什么,它们之间是什么关系,人、资源,以及这家公司在做什么。所以它就像公司里今天刚入职的新员工和已经工作了五年的人之间的区别。

Yeah, I mean, you know, ontology just means that in an organization, the relationship between all the abstract concepts—of all the goals and all the departments and all the people and all the projects that are going on—what do they exactly mean and what's the relationship between them, the people, the resources, and what that company does. So it's the difference between a person who is a new employee in the company and just started today and a person that has worked there five years.

Ali

你知道,假设他们技能相当。教育背景相同。同样聪明、勤奋等等。但一个,你知道,今天是她第一天上班,另一个已经在那儿五年了。这两个人有什么区别?

You know, let's say they're equally skilled. They have the same educational background. They're equally smart and hardworking and all of that. But one, you know, her first day today at work, the other one she's been there 5 years. What's the difference between these two people?

Ali

那个人拥有关于组织如何运作、谁是谁、事情如何办成的本体。别看组织架构图。别去问那个人。他什么都办不成。你去问这个人,你知道,他会帮你办成。而且不是那么回事。你不需要在这里提交那份文件。你知道,这是这个项目。这是正在发生的事。这很关键。所以有很多根深蒂固的知识存在于每个人的头脑中——谁知道一个组织如何运作。这就是为什么在创业圈人们说,‘嘿,如果你失去大部分员工,那家公司就无法恢复。’你不能只是补充和招聘新人,因为人太关键了。我们如何获取那个上下文——那就是本体——并把它交给 AI?其中一部分是我们必须有录音之类的,但第二部分是你如何真正把它提炼成一个图——

The one has an ontology of how that organization works, who the people are, how you get stuff done. Don't look at the org chart. Don't go ask that person. He will not get anything done. You go ask this person, you know, he'll get it done for you. And that's not how it works. You don't need to file that paperwork here. And you know, and this is this project. This is what's going on. This is essential. So there's just a lot of ingrained knowledge that's sitting in everybody's heads—who knows how an organization works. That's why people say in startup land, they say, 'Hey, if you lose most of your people, that company can't recover from it.' You can't just replenish and hire new people like the people are so essential. How do we get that context—that's the ontology—and give it to the AI? Part of that is we just have to have the recording and all of that, but the second part is how do you actually distill it down into a graph—

Ali

实际上是一个数字图,然后你可以把它喂给 AI。所以今天很多智能体的工作方式——比如 Claude Code 或它们中的任何一个,Codex 或 Pi,或者你知道,OpenCode,或者你可以遍历一大堆——你知道,它们有这个循环,智能体式循环。它可以推理,但然后它会去逐个检查每个资源。所以它会去这个 MCP 服务器找你的问题,试着看答案在这里吗?还有另一个吗?它综合一下给你一个答案。但这有点慢。我把它比作如果谷歌在 25 年前就这样构建谷歌搜索。我们会说,好吧,我们要得到 10 条蓝色链接。我们在这里搜索关键词,但它不是给你 10 条蓝色链接,而是去一个网站,用 LLM 总结它做什么,找到几个枢纽超链接,并行跳转到其中几个,读几个网站,这样做 10 分钟,然后给你它找到的最好的 10 条蓝色链接。嗯,那会非常昂贵——每次上网都要花很多钱。第二,它会花很长时间。你得等 10 分钟。第三,质量会很差,因为你实际上只看了存在的一切中很小的一部分,对吧?那么他们怎么做?他们有一个索引,对吧?你从不离开谷歌服务器。你搜索它,命中索引,反向索引立即在不到 100 毫秒内给你 10 条蓝色链接。我们需要为 AI 做同样的事情。所以本体就是我们需要一直离线计算那个索引。所以它几乎就像谷歌当年发明的 PageRank 算法,但更复杂,因为谷歌只是看一个所有人都能上的网络。这里有权限——

Actually a digital graph that you can then feed to the AI. So the way a lot of the agents work today—like Claude Code or any of them, Codex or Pi or you know, OpenCode, or you can go through the whole slew of them—you know, they have this loop, agentic loop. It can reason, but then it goes and checks every resource one at a time. So it'll go to this MCP server for your question and try to see, is the answer here? Is there another one? It synthesizes it and gives you an answer. But it's kind of slow. I liken this to if Google would have built Google Search this way 25 years ago. We would have said, okay, we're going to get 10 blue links. We search for key terms here, but instead of giving you 10 blue links, it would have gone to one website, summarized with an LLM what it does, found a few hub hyperlinks, jumped in parallel to a few of them, read a few websites, done that for 10 minutes, and then given you like its best 10 blue links it would find. Well, that would be very expensive—cost a lot of money to do that every time, go on the web. Two, it would have taken a long time. You got to wait 10 minutes. And three, the quality would be bad because you're actually only looking at a very small subset of everything that exists out there, right? So how do they do it? They have an index, right? You never leave Google servers. You search for it, hits the index, the reverse index immediately gets you the 10 blue links within, you know, less than 100 milliseconds. We need to do the same thing for the AI. So the ontology is that we need to compute that index offline all the time. So it's almost like the PageRank algorithm that Google had invented back in the day, but it's more complicated because Google was just looking at a web where everybody can go on the web. Here there permissions—

Host

而且链接是存在的,并且——

And the links existed and—

Ali

是的,这里有权限问题。我被允许访问的数据可能不是你能访问的数据。所以有隐私,有访问控制。而且,这里我们处理的是许多不同类型的对象,不只是网站。所以问题稍微难一点。但它是可管理的。你实际上可以做到。所以,你知道,我确信你可以做到这一点,并且能从中获得巨大的生产力提升,因为我们为 Databricks 做到了。

Yeah, here there's permissions involved. The data I'm allowed to access might not be the data that you're allowed to access. So there's privacy, there's access control. Also, there's many different types of objects here that we're dealing with, not just websites. So the problem is a little bit harder. But it's manageable. You can actually do it. So, you know, I'm convinced you can do this and you can get massive productivity gains out of it because we did it for Databricks.

Host

是的,没错。

Yeah, exactly.

Ali

我们为自己做到了。

We did it for ourselves.

Host

是的。是的。

Yeah. Yeah.

Ali

是的。我们就像,公司完全变了。

Yeah. And we're like the company's just completely changed.

Databricks 自用实践 Dogfooding Databricks

Host

我的意思是,你们一直在为 Databricks 吃自己的狗粮,但也许可以多谈谈你们作为一个组织所看到的影响。

I mean, you've been dogfooding Databricks for Databricks forever, but maybe say more about the impact you've seen as an organization.

Ali

是的,我的意思是,一旦我们有了这个本体,我们开始着手做,我们实际上可能拥有所有客户中最大的本体。我们自己的本体比任何客户用我们来构建的本体都要大,因为 Databricks 用 Databricks 比任何人用得都多。所以我们拥有的本体图里有数百万个节点。这就是组织里会发生的事。你有一个树状结构的组织,信息在树状结构里上下流动。如果你做不了决定,你就上报给你的老板。也许他们能打破僵局,再往上上报。他们需要了解正在发生什么,需要掌握所有上下文,然后做出决定。决定做出后,你必须把它在组织里向下渗透。如果你有本体,很多这样的事现在都可以由 AI 来做。为什么?因为会议里会发生什么?在会议里,有人做了分析。他们可能有一个 PowerPoint 幻灯片,里面有一些漂亮的图表。做分析的那个人是个聪明人,用了 Excel,做了一些模型。有一些数字。所以很多这样的事你现在都可以用 AI 来做。所以 AI 可以为你做分析。它有所有上下文。它可以按你想要的方式呈现。你可以就它提问,而不用开后续会议。你可以直接向 AI 提问。所以这非常相似。这和 Jack Dorsey 说过的你能对组织做的事是一个路子。这只是一种具体的实现方式。

Yeah, I mean, once we got this ontology and we started working on it, and we actually have probably the largest of all of our customers. We have the largest ontology. Our ontology is bigger on us than any of our customers when they use us to build their ontology, because Databricks uses Databricks more than anyone else uses Databricks. So it's like millions of millions of nodes in the graph, in the ontology graph that we have. So it's just what happens in an organization. You have a tree structure organization, and information flows up and down the tree structure. If you can't make a decision, you escalate to your boss. Maybe they can tie-break it, escalates up. They need to get up to speed on what's happening, and they need to get all the context, and then they make decisions. Once decisions get made, you have to percolate them down in the organization. A lot of this can now be done by AI if you have an ontology. Why? Because what happens in a meeting? In a meeting, someone has done the analysis. They probably have a PowerPoint deck with some pretty graphs in it. That person that did the analysis is some smart person that used Excel, made some models. There's some numericals. So a lot of that you can now just do with AI. So the AI can do the analysis for you. It has all the context. It can present it in a way that you want. You can ask questions about it instead of having follow-up meetings. You can directly ask questions directly from the AI. So it's very similar. It's along the lines of what Jack Dorsey has said that you can do to the organization. It's just a concrete way of implementing it.

Host

是的。

Yeah.

Ali

所以这对我们来说是游戏规则的改变。现在开会时每个人都在手机上用 Genie,向 Genie 提问。只要有人说了一些复杂的东西,你就能看到所有人都去看手机。

So it's a game changer for us. Like, everybody's on their phones now in the meetings on Genie, and they're asking Genie questions. You can see as soon as someone says something complicated, you see everybody go to the phone.

财富 500 强轶事 The Fortune 500 Anecdote

Host

你能分享一下你曾在一次董事会上提到过的那个财务轶事吗?

Can you share that finance anecdote you mentioned once in a board meeting?

Ali

是的,当然,内部董事会。是的,所以,嗯

Yeah, it's, yeah sure, internal board meeting. Yeah so, um

Host

只讲能公开的。

Only a kosher.

Ali

是的,没错。不,这其实是我们一次演示需要的。我需要知道我们在《财富》500 强里有多少客户,我们的《财富》500 强渗透率是多少。我问了销售运营部门的一个人,因为我以为她会有,她回短信说,哦抱歉,我现在在飞机上登不了 Genie。我说,如果你也只是要登 Genie,那我自己就能做。我不是,我问你是因为我以为你有我没有权限访问的别的渠道。所以我有点生气,就改去问 CFO,Dave。我给 Dave 发短信说,嘿,你知道我们的《财富》500 强渗透率是多少吗?他直接回了一张 Genie 的截图。所以他也去问了。所以我说,这里有人做点新颖的事吗?大家都只是去 Genie 问本体。

Yeah exactly. No, it's actually needed for one of our presentations. I need to know how many customers do we have in Fortune 500 that use, what's our penetration of Fortune 500. And I asked one of the people in sales ops because I thought she would have it, and she texted me back and said, oh sorry, I can't log in to Genie right now, I'm on a flight. And I said, if you're just going to log into Genie, I can do that myself. Like, I don't, I asked you because I thought you had something alternative that I don't have access to. So then I was kind of a little bit angry, so I texted the CFO instead, Dave. And so I texted Dave and I said, hey, do you know what our Fortune 500 penetration is? And he just copy-pasted a screenshot of Genie back. So he also asked that. So I said, does anyone do anything novel here? Just everybody just goes to Genie and asks the ontology for questions.

Host

这就像“让我帮你 Genie 一下”,而不是“让我去查一下”。

It's like let me Genie that for you instead of let me go.

Ali

现在大家都这么做。我们就直接这么说。我们说,嘿,能不能有人 Genie 一下这个,你能不能直接从本体里拿到。所以我确实认为这是游戏规则的改变。但并不是你按个按钮,组织里就有本体了。我认为 Palantir 实际上做得很好,他们走进组织,把很多隐性知识写下来,带进组织里。我们自动把它拿过来,构建成图,然后把图喂给智能体,这样我们就能回答问题,并以业务领导者喜欢看到的方式来回答,也就是用图表、分析的方式,一种你可以追问那个问题、继续提问并得到答案的方式,这样你就能做决定,然后把信息在组织里传播。

That's what everybody does now. We just say it. We say, hey can someone just Genie this, like, can you just get it from the ontology. So I do think it's a game changer. But it's not just you press a button, you have an ontology in an organization. And I think Palantir actually has done a great job of going to organizations and getting a lot of that tacit knowledge written down and getting it into the organizations. We automatically take that and build the graph, and then we feed that graph into the agents so that we can answer the question and answer it in a way that business leaders would like to see it, which is in graphs, analytical way, and a way where you can interrogate that question and continue asking questions and getting answers to those so you can make decisions and then disseminating that information in the organization.

价值最大化与 Token 最大化 Value Maxing vs Token Maxing

Host

是的,这相当惊人。你刚才标记了,开发者显然在用 AI,价值存疑。我想就这一点追问你,因为我觉得你们是最早的一批。我说,我不想用“token maxing”这个词,因为它有很强的负面含义,但就赞赏那些能用 AI 变得更高效的人而言,你们是走在前沿的,对吧?然后当然就有这个循环:哎呀,人们在浪费,现在我们需要“价值最大化”。你自己在这方面的历程是怎样的,你们怎么看待价值最大化而不是 token 最大化?然后我要把 Unity Gateway 也放进来,因为我认为管理成本这一块实际上越来越重要,而你们在帮人们做到这一点。但也许把它联系起来,展开一下。

Yeah, it's pretty amazing. You sort of bookmarked the, developers are obviously using AI, questionable value. I want to follow up with you on that because I feel like you guys were one of the earliest. And I say, I don't want to use the word token maxing because it has such a negative connotation, but I think in terms of applauding people who can use AI to become more productive, you guys were at the forefront of that, right? And then of course there's this cycle of, oh shoot, people are being wasteful, now we need a value max. Like what was your own journey on that, and how do you guys think about value maxing not token maxing? And then I'm going to throw in Unity Gateway in this, right, because I think the managing of cost piece is actually getting more important, and you guys are helping people do that. But maybe tie that in to extend it.

Ali

是的,大概去年第四季度,模型变得非常非常好,我们开始注意到,好吧,它实际上开始带来好得多的生产力。所以我自己开始用这些模型,开始为 Databricks 把代码提交到生产环境,就是真的,我想一路做到生产。我这么做了,然后开始推动组织,说嘿,每个人都需要这么做。我做到了。你为什么没做?就像如果 CEO 都能在一个有所有这些安全要求的非常敏感的数据平台上把代码提交到生产,你也应该能做到。你,任何经理,组织里的任何人。所以开始非常用力地推动每个人,我们在第四季度开始做排行榜。到了年初,我说一月、二月,我们启动这一年的时候,我们已经全面铺开了。每个人都在用这些东西,我们在推动,我们在管理这件事。但整个 token maxing 的事大概在二月、三月那段时间已经在发生了。所以是的,我们只是运气好,可能比大家早了几个季度,看到了这里正在发生什么,而且它开始失控了。所以我们已经有了一个网关,叫 Unity Gateway,我们已经在用这个网关来提供 token 容量。所以你可以拿到 OpenAI、Anthropic、Gemini、Grok 的容量,就像任何客户都可以来找我们,我们就给他们提供那个容量,因为我们和那些公司有合作关系,还有任何开源模型。

Yeah, so around Q4 last year was when the models got really really good, and we started noticing that okay, it's actually starting to give much better productivity. So I actually started using the models myself to start committing code into production for Databricks, like the actual, as I want to take it all the way to production. So I did that, and started pushing the organization that hey everyone needs to do that. I have done it. Why are you not? Like if the CEO can commit code to production on a very sensitive data platform that has all these security requirements, you should be able to do that too. You being any manager, anyone in the organization. So started pushing everyone very hard, and we started making leaderboards in Q4. And at the beginning of, I say January, February, when we kicked off the year, we were already full swing. Everybody was using the stuff, and we're pushing and we're managing this. But the whole token maxing thing was happening around February, March period, already it was happening. So yeah, we just had the luck of being maybe a few quarters ahead of folks to see what was happening here, and it was getting out of hand. So we already had a gateway, it's called Unity Gateway, where we were already, this gateway was being used to provide token capacity. So you can get OpenAI, Anthropic, Gemini, Grok capacity, like any customer can come to us and we'll just provide them that capacity because we have relationships with those, and any open source model.

预算约束与成本分析 Budget Constraints and Cost Analytics

Ali

于是我们开始设置预算约束,并给人们警告,比如:好,你的预算还剩这么多,你快接近上限了。我们开始按个人和按团队这样做,然后我们开始做很好的分析,这样我们就能准确预测成本会走向哪里。

So we started putting budget constraints in place and giving people warnings like, okay, you have this much of your budget left, you're getting close to your ceiling. So we started doing that per person and for groups, and then we started doing great analytics so we could predict exactly where the costs were going.

Ali

然后我们加入了智能路由器,如果你快接近预算上限,或者你知道你问的是简单问题,它就能挑选更便宜的模型。我们开始这样做了。我们还构建了一个叫 Omnient 的 harness,它可以在不同的 harness 之间多路复用。事实证明,harness 本身很重要。比如你用同一个模型但不同的 harness,成本差异几乎有 2 倍。

And then we added smart routers that could actually pick cheaper models if you're getting close to your budget, or if you know you have simple questions. We started doing that. We also built a harness called Omnient which can multiplex between the different harnesses. Turns out actually the harness itself matters. Like if you use the same model but different harnesses, there's almost 2x different cost difference.

Host

即使完全相同的模型。

Even exactly same model.

Ali

是的。你知道,同一个版本但不同的 harness,实际成本会差 2 倍。所以如果你能更换 harness,你就能在成本上获得很大的杠杆。于是我们开始使用所有这些,我们得以真正弯折曲线,实际上我们在 AI 上的成本基本上是 token 持续增长,但成本却基本停滞。所以这对我们来说真的超级超级重要,而且对此有巨大的需求。我认为现在每个组织都在经历这个。

Yeah. You know same version but different harness you get 2x difference in actual cost. So if you can change harness you can get a lot of leverage in the cost. So we started using all of this that we were able to actually bend the curve and actually our cost for AI has been basically the tokens continue to go up but the costs have been sort of stagnant. So that's been actually super super important for us and there's a huge demand for this. I think every organization is going through this now.

前沿模型的市场转变 Market Shift from Frontier Models

Host

是的,我第一次——我参加很多董事会。我在 20 多个董事会任职。上周,有史以来第一次,一家规模很大的公司说他们要从前沿模型转向 GLM。这是一家大型工程组织。

Yeah, I for the first time I do a lot of board meetings. I'm on 20-some boards. For the first time ever, a company at scale last week said that they're moving from the frontier models to GLM. This is a large engineering organization.

Host

你看到这个了吗?你觉得这是一个趋势,还是只是一个孤立的轶事?因为我一直听到——我记得第一次 DeepSeek 时刻,还有英伟达股价,然后那被证明不是真的。然后是 Kimi 时刻,再下一个 DeepSeek 时刻。这些似乎都没有对市场产生明显的影响。

Do you see this? Do you think that that's a trend or do you think that's just like a one-off anecdote? Because I've been hearing about I remember the first Deep Seek moment and like Nvidia Stock and then that turned out to not be real. Then the Kimmy moment, then the next Deep Seek moment. None of it seems to have actually had an appreciable impact on the market.

Ali

是的。但现在我掌握的轶事数量相当真实,而且似乎正在发生。

Yeah. But now then the amount of anecdotes that I have are pretty real and it seems to be happening.

Host

想听听你的看法。

Love your view.

Ali

我的意思是,我认为人们两者都想要。他们想要你知道,他们想要最新的、超级智能的模型来处理那些能获得 ROI 的困难任务。

I mean I think people want both. They want you know they want the latest model that's super intelligent for the difficult task where they get ROI.

Host

但然后有很多平凡愚蠢的事情,比如你知道,人们真的用他们的 harness 来重命名文件之类的。

But then there's a lot of mundane dumb things like you know you people literally use their harness to rename files and whatnot.

Ali

是的。

Yeah.

Host

你知道,就像你为此支付了高几个数量级的费用。至少你自己输入那个。别让模型做那个。它会转 5 分钟,然后帮你重命名文件,花掉你几分钱。

You know like it's you're paying you know orders of magnitude more for that. At least type that in yourself. Don't have the model do that. It's going to spin for 5 minutes and then it's going to rename the file for you and cost you you know cents.

Ali

呃,但我想人们——我只是想知道你真的看到市场变动了吗?不,人们正在转向它。但他们在做的是,你知道,模式要么是使用这种专家模式,你有一个小的、更便宜的开源模型来使用专家模型,大的那些,或者反过来,或者一种他们可以互相 ping pong 的方式,但还有多路复用 harness 和只是更换 harness 以便控制成本,这也是人们正在做的。你知道,人们发现,例如,pi 作为一个 harness 非常高效。所以,是的,我认为会有很多这样的东西。很容易——模型本身是随机的,正如你所说,每次它们给出不同的答案,而且变化如此之大。所以只是有很多实验在进行。所以我认为我们会进入一个世界,你不再总是对所有事情使用最聪明的模型,这有点是过去几年的范式。就像新模型出来,它超级聪明。他们用它做所有事情,即使是非常非常简单的平凡任务。

Uh but um I guess people I'm just wondering do you actually see market movement? No, people are moving on it. But what they're doing is that you know uh the pattern is either use you know you can you can use this expert pattern where you have you know small cheaper open source model that uses expert model the big ones or vice versa or a way in which they can sort of uh ping pong them to each other but also multiplexing harnesses and just changing harnesses so that you can control the costs is also what people are doing. You know people have found for instance you know there's pi is very efficient when it comes to as a harness. Um uh so yeah, I I think there's going to be a multitude of these. It's easy to the models themselves are stocastic as you said every time they give a different answer and they're changing so much. So there's just a lot of experimentation happening. So I think we're going to get to a world where you're not always using the smartest model for everything which is kind of the been the paradigm for the last couple years. Like new model comes out, it's super smart. They use it for everything even really really simple mundane tasks.

Host

是的。我告诉你我看到的,我看到人们用 Fable 和 Astra 来做架构。

Yeah. I'll tell you what I see I see people using Fable and Astra for like architecture.

Ali

嗯哼。

Uhhuh.

Host

用便宜的模型来实现,然后用 Fable 或 Astra 来审计。是的,这似乎正在兴起。

A cheap model for implementation and then fable or astro for audit. Yeah, like that seems to be like this emerging.

Ali

你们在初创公司中看到了什么?我的意思是,他们不是——但就是这样。老实说,这就是模式。

What are you guys seeing in the startups? I mean, aren't they um but that's it. That's like that's that's honestly the the pattern.

Host

多少开源?

How much open source?

Ali

我按 token 还是按美元,都行。

I by token or by dollar either.

Host

所以按美元,开源大约 5%。非常少,但按 token 数量,超过 60%。

So by dollar open source is like 5%. It's very little, but by token count it's over 60%.

Ali

是的,我正要说,我的意思是我们和比如说 Decagon 之类的谈过。嗯,我认为内部使用和外部产品是不同的。在外部产品上,我认为他们几乎达到 90% 开源,在内部,嗯,我不想特别说 Don,但很多都是像我们不在乎,我们就用前沿模型,我们不考虑成本控制,但随着它变大,是的,对,你和我另一次董事会会议,他们实际上从浪费的角度把那个降下来了。所以我肯定看到在产品方面更多转向开源。实际上这与另一个问题相关,也许围绕开源但特别是后训练。

Yeah, I was going to say um I mean we talked to let's say a decagon or something like that. They well I think it's different internal use versus external for product. On the external for product I think they're almost up to 90% open source on the internal um and I don't want to say for Don in particular but a lot of them are like we don't care we'll just use frontier we're not thinking about cost control but as it gets bigger yeah right you and I were another board meeting where they actually did bring that down just from a waste perspective. Um so I definitely see that moving more toward open source on the product side. Um and actually that's a related to another question um maybe around open source but post trading specifically.

Host

嗯,我觉得你有点早。我记得 2023 年和你谈过。你什么时候收购 Mosaic 的?

Um I feel like you were kind of early. I remember talking to you in 2023. You when did you buy Mosaic?

Ali

2023 年。

2023.

Host

2023 年。好的。所以你在 23 年的这个愿景在 2026 年某种程度上实现了。我不知道你们是否会同意,对吧?就像我们听到的,当然,哦,就像嘿,我们将真正拥有你自己的智能。你将会,你知道,后训练你的开源模型等等。嗯,这绝对是初创公司正在做的。嗯,我不知道企业是否还在做,但就像我的意思是,你觉得你对此早了吗,还是

2023. Okay. So this vision that you had in 23 kind of came true in 2026. I don't know if you guys would agree, right? Like that's sort of what we're hearing across you know of course oh just sort of like hey we're going to actually you're going to own your own intelligence. you're going to be, you know, postrading your open source models, etc. Um, and that's definitely what the startups are doing. Um, I don't know if that's what the enterprises are doing yet, but like I mean, do you feel like you were early to that or

Ali

是的,我的意思是,首先,你知道,当我们开始的时候,也是,嘿,我们也会为你预训练,那没有任何意义。你知道,现在有非常好的预训练模型你可以使用,对吧?但你可以实际上对模型进行后训练,你可以做强化学习。是的,我们实际上正在大规模地做,许多那些初创公司实际上是客户。所以我们实际上帮助他们强化,你知道,使用早期或强化学习环境,我们可以使模型非常非常擅长他们正在做的特定任务。对他们来说这样做很有意义。如果你有一个重复性任务,所以如果你有一个初创公司,它提供一个产品,产品做特定的事情。它不只是通用智能。它为你做特定的事情。拿一个非常好的开源模型,你知道,使用强化学习,使它非常擅长那个特定任务,这非常合理。你可以降低成本。他们可以做得非常快。

Yeah, I mean, first of all, you know, there was uh when we started it was also, hey, we'll also pre-train it for you, which that's that doesn't make any sense. You know, you can there's so very good pre-trained model now that you can use, right? But that you can do actually post- training on the model and you can do reinforcement learning. Yeah, we're actually doing it at scale and many of those startups are actually customers. So we actually help them rein you know using early or reinforcement learning environments where we can make the models very very good at the specific task that they are doing. It makes a lot of sense for them to do that. If you have a repetitive task so if you have a startup and it's offering a product and a product does something specific. It's not just a general uh intelligence. It does something specific for you. It makes just a lot of sense to uh take a really good open source model and you know use reinforcement learning and make it really good at that specific task. You can cut the cost down. they can make it really fast.

企业为何跳过评估 Why enterprises skip evals

Ali

你知道,他们掌控自己的知识产权。所以从这个意义上说,这是可能的。但大型企业,他们只需要基础的自动化,而现在做这些对他们来说负担太重了。

You know, they control their own IP. So in that sense that is possible. But large enterprises, they just need basic automation, and it's just too much for them to do this right now.

Host

是的。

Yeah.

Ali

我觉得挑战之一就是,你需要好的评估。

I think one of the challenges is, you know, you need good evals.

Host

完全同意。

Totally.

Ali

而做出好的评估很难。创业公司能做,也有动力去做,但对其他组织来说,最省事的做法可能就是直接用前沿模型,而不必自己去做评估。我们其实在产品里自动为客户生成了评估,还把它放在最显眼的位置,但人们不想用。于是我们说,好吧,把它挪到后端,变成可选项,结果他们就再也不会去碰它了。所以总的来说,他们就是不想掺和进来,太复杂了。

And making good evals is hard. So while the startups can do that and they're motivated to do that, for other organizations the easy button might be just to use a frontier model rather than having to create their own evals. We actually generated evals for the customer automatically in the product, and we had it front and center, but then people didn't want to use it. So we said, okay, let's move it to the back end so that it's optional, and then they would never go to it. So I would say in general, they just don't want to get into it. It's too complicated.

Ali

我觉得你想要的是快速的反馈,比如:嘿,出了个新模型,我想试试,我想把这个问解决掉。你没时间去走那套科学方法——先做个评估,再搞个好的基线。这有点像 TDD,测试驱动开发。你知道,在软件工程里,人们真的做测试驱动开发吗?很少人做,对吧?每个人都说这是正确的做法,但实际上没人真去做。所以这是一样的,这多少就是自己训练模型的诅咒——评估才是最难的部分。

I think you want quick reinforcement of like, hey, there's a new model. I want to try this out. I want to get this problem solved. You don't have time to go do this the scientific method of let's make an eval, let's have a great baseline. And it's sort of like TDD, test-driven development. You know, in software engineering, did people actually do test-driven development? Very few did, right? Everyone said it's the right way to do it, but nobody actually in practice did it. So that's the same, that's kind of a little bit of the curse of doing, you know, training your own model, is the eval is the hard part.

企业中的前沿部署模型 Forward-deployed models at enterprises

Host

我知道你们在 Databricks 有一个 FD 模式,这个词现在很火,或者说这个缩写。但要让大型企业用上这些,是需要全职员工(FTE)模式吗?还是说你们怎么做,这个模式又是怎么演变的?

I know you have an FD model at Databricks, that's a very popular word right now, or acronym. But does it, like, to get these at enterprises that large? Is it a full FTE model that's required, or like how do you, and how's that evolved maybe?

Ali

是的。我是说,我们一直有这些 FD,需求涨得非常厉害。很大一部分是,我们怎么构建那个本体?本体是自动的,但如果你不收集任何信息,比如你什么都不记录,对吧?所以这是我们做的关键事情之一。但也有像这样的:我想构建一个智能体,我想把它放出去,我想让它面向客户,它得有非常低的延迟,我还想让它有护栏,防止有人来滥用它,或者问一些我们不想让它回答的问题,等等。所以我们可以构建这些。比如 Fox 的那个体育 AI,你可以去跟它聊体育赛事。你试着问它政治问题,它非常擅长拒绝你,然后把话题转回体育。所以 FD 就是做这个的。我们会帮助组织真正开始用上 AI。这很重要,因为很多组织没有内部专业能力来构建这些东西。所以他们只需要一点点外部的帮助,然后就能起步了。

Yeah. Yeah, I mean, we've had these FDs and the demand for it has gone up significantly. A lot of it is, you know, how do we build that ontology? Like the ontology is automatic, but if you're not collecting any information, like you're not recording anything, right? So that's one of the key things that we do. But also things like, you know, I want to build an agent, I want to put it, I want it to be customer-facing and it has to have really low latency, and I want it to have guardrails so people don't come to abuse it or ask it things that we don't want it to answer, and so on. So we can build that. Like, you know, like sports AI that Fox has, you can go chat with it about sports events. You can try to ask it actually about politics and it's very good at rejecting you and moving and talking about sports instead. So the FDs built that. So we'll help the organizations actually get started with AI. It is important because many organizations do not have the in-house expertise to build this stuff. So they need just a little bit of help on the side and then they get started.

智能体为何选 Neon 和 Lakebase Why agents pick Neon and Lakebase

Host

好,有道理。这更多是跟智能体这边相关。我最近看到,好像有个第三方,中立的第三方,做了一些测试,发现 Lakebase 或 Neon 实际上是智能体首选的 Postgres 数据库。我觉得这挺有意思的。一是因为,你知道,这对 Databricks 来说是个好消息,二是因为一年前我大概猜不到会是这样。

So yeah, makes sense. So this is more related on the agent side, but I saw recently that I think a third party, neutral third party, did some tests that Lakebase or Neon was actually the data, the Postgres database of choice for agents. And I thought that was interesting. One, because you know, one exciting Databricks, but two, I probably wouldn't have guessed that maybe a year ago.

Ali

确实意外。

Surprise for sure.

Host

是的,挺意外的。因为外面还有别的产品,你知道,开发者势头也很好,但它明显是第一名。所以我很好奇,你们是怎么攻克这个的?是什么让你们在智能体这块胜出?因为如果你现在赢了智能体,你就赢了市场。

Yeah, it was a surprise. Just because there's others out there that have, you know, great developer momentum as well, but it was pretty clearly number one. And so I'm curious, how did you guys crack this? And what makes you win across the agents? Because if you win the agents now, you win the market.

Ali

是的。我觉得很大一部分功劳要归给 Neon 和 Nikita 以及他们的团队。我认为他们做的就是,极其执着于怎么让模型、怎么让模型选择、让智能体偏好 Lakebase 或 Neon 作为数据库。那他们做了什么?智能体想要做实验。你知道,它们跑出去,试图构建一点软件,它们需要数据库。所以你需要数据库能快速启动。他们就有这种执念:一切都应该远低于一秒。所以数据库启动远低于一秒。你可以克隆巨大的数据库,比如 PB 级的数据库,你可以在不到一秒内克隆它。所以它高度弹性、高度响应。然后他们打造了这个杀手级功能,叫分支(branching)。分支就是让你把数据库分叉,你可以在同一个数据库上有很多很多分支。他们把它做得非常非常轻量。我们在智能体的其他东西上也见过这个,对吧?比如 UV、ripgrep,基本上是把 Unix 上很多工具重新实现,让它们变得极快、极轻量,而且对智能体来说也基本不会出错。他们只是把这个思路用到了一个更难的问题上,也就是数据库。所以现在你有一个 Postgres 数据库,而 Postgres 数据库拥有所有这些优势:它非常快、灵活、不会出错,你可以回滚到快照,你可以做这些事。所以我觉得这就是为什么智能体用起来更容易。他们还确保自己的定价模式是那种——你不想因为智能体在构建软件、做实验,就让成本飙升。

Yeah. I mean, I think a lot of credit should go to Neon and Nikita and team. And I think what they've done is they've just been obsessive about how do you make the models, how do you make the models pick, and agents favor Lakebase or Neon as a database. So what do they do? The agents want to experiment. You know, they're going off, they're trying to build a little bit of software. They need the database. So you need the database to come up quickly. So they had this obsession that everything should take far less than a second. So you know, database comes up in far less than a second. You can clone gigantic database. So kind of a petabyte database. You can clone it in less than a second, you know. So it's like highly elastic, highly responsive. And then they built this killer feature called branching. So branching just lets you branch the database and you can have many, many branches over the same database. And they just made this very, very lightweight. We saw this with other things with agents, right? Like UV, you know, ripgrep, like basically reimplementation of a lot of the tools on Unix making them really, really blazing fast and lightweight, and also sort of failsafe for agents. They've just did this to a harder problem, which is database. So like now you have a Postgres database and the Postgres database has all these advantages that it's really fast, it's nimble, it's fail safe, you can go back to snapshots, you can do those things. So I think that's why it's just easier for the agents to use this. They also made sure that they had a pricing model that was like, you don't want just because the agents are building some software and experiment, you don't want the cost to run up.

Host

是的,完全同意。

Yeah, totally.

Ali

如果数据库是生产用途、很多人在用,你愿意为它付费,但如果只是做实验就不想。所以我觉得他们就是很执着。他们不是想赢数据库大战,也不是想比别的厂商更好。他们执着的是:我们怎么才能对智能体最好?

You're okay paying for your database if it's like production use and lots of people are using it, but just to experiment. So I think they were just obsessed. They were not trying to win the database war or trying to be better than some other vendor. They were obsessed with how are we the best for the agents?

Host

而这是一个新的用户画像,因为数据库领域一直以来的执念是:我们怎么帮助 DBA?怎么帮助应用开发者?怎么帮助使用数据库的人?

And that's a new persona, because in databases the obsession has been how do we help DBAs? How do we help app devs? How do we help the people that are using the database?

Ali

没错。他们改变了游戏规则,说:嘿,我们怎么聚焦于智能体,帮助智能体得到它们想要的最好的数据库?而现在,Neon 和 Lakebase 上创建的数据库,超过 90% 实际上是由智能体创建的。所以甚至都不是人类。所以你知道,数字本身就说明了一切。顺便说一句,这很了不起。

Exactly. They changed the game and said, hey, how do we focus on agents and help agents get the best database they want? And you know, now over 90% of the databases that are created on Neon and Lakebase are actually created by agents. So it's not even humans. So you know, numbers speak for themselves. By the way, it's remarkable.

Databricks 内部的 Neon Neon inside Databricks

Host

所以我已经开始把 Neon 当作我的标准数据库来用了,这让我觉得很奇怪,因为通常你进入一家大公司,事情会变慢。但实际上产品变得明显更好了。

So I've started to use Neon as like my standard database, and it was bizarre to me because normally when you enter a large company things slow down. It's actually like the products got materially better.

Ali

是的。

Yeah.

Host

他们完全独立吗?他们跟其他部分是怎么协作的?

Are they totally independent? Do they work with the rest of the, like how?

Ali

不,他们是个很棒的团队。我是说,我们合作非常紧密。你知道,我们热爱数据库和数据。所以,这就是我们生活的一部分。

No, it's a great team. I mean, and work very closely together. You know, we love databases and data. So it's, you know, we live that.

P(doom) 估计 P(doom) estimates

Host

不过,团队做得非常出色,把它做得超级快、干脆利落,而且非常适合智能体。好,嘿,Ali,你的 P(doom) 是多少?

But no, the team does a great job of just making it super fast, snappy, and great for agents. All right. Hey Ali, what is your P(doom)?

Ali

低于 10%。

Less than 10%.

Host

不是吧。

No.

Ali

接近零。你的呢?

Close to zero. What about yours?

Host

我不知道。我唯一的答案是:没有 AI 时我的 P(doom) 比有 AI 时高得多。这样回答怎么样?

I don't know. My only answer is that my P(doom) without AI is much higher than my P(doom) with AI. How's that?

Ali

哇。这就是我要说的。这一点上我同意 Ali。对。

Wow. That's what I say. I would agree with Ali on this one. Yeah.

互动版:逐字朗读 + 针对本期提问 →