Mistral AI Founder on Scaling, DeepMind Lessons, and the 7B Model
打开互动全文版(中英对照 + 朗读 + 问答)→Arthur Mensch 讨论 AI 扩展的障碍、DeepMind 的经验教训以及 7B 模型成功的原因。
Arthur Mensch discusses barriers to scaling AI, lessons from DeepMind, and why the 7B model was a hit.
准备好了吗,Arthur?我太兴奋了。JC 很久以前就介绍我们认识了。我认识你有一阵子了,一直想促成这次对话,所以非常感谢你今天来参加。
Ready to go Arthur? I am so excited for this. JC introduced us quite a long time ago now. I've known you for a while, I've been wanting to make this happen for a while, so thank you so much for joining me today.
谢谢你邀请我。这是我的荣幸。
Thank you for having me. It's a pleasure.
这是我的荣幸,朋友。但我想从这个问题开始:你的父母或老师会如何描述小时候的 Arthur?我总是对最优秀创始人的性格特质感到好奇。他们会如何描述 10 岁的 Arthur?
The pleasure is mine, my friend. But I want to start: what would your parents or teachers have described the young Arthur? I'm just always intrigued by the characteristics and traits of the best founders. How would they have described a 10-year-old Arthur?
我想我有点好奇,也有点固执,而且对兄弟们不太好。但后来有所改善。我也是他们中的老大。我不知道,你应该问他们,但我想他们应该有美好的回忆,希望如此。
I guess I was a bit curious and a bit stubborn, I should say, and not very nice to my brothers, I think. But that improved over time. I was the eldest of them also. I don't know, you should ask them, but I think they have good memories, hopefully.
你知道吗?可惜你妈妈不在我们的参考名单里,所以我们错过了。但 JC 提供了一些很好的评论。所以我也想从你的成长经历开始:你第一次接触 AI 是什么时候?你是个在法国长大的孩子,最初是如何接触到 AI 和机器学习的?那个最初的激情点是什么?
Do you know what? Sadly your mother wasn't in our reference list, so we missed that one out. But you know, JC provided some great commentary. So I do want to start also, you know, growing up, what was your first exposure to AI? You're a kid in France, how did you first get exposed to AI and machine learning, and what was that first passion point?
那是 Andrew Ng 让直升机倒飞,翻转过来。这是一个不容易解决的控制问题,我不确定它是否真的与 AI 相关。我想他说的是他用神经网络来控制这一切。但这是我记得的最近的——嗯,第一次有人向我展示当时可以用机器学习做什么的记忆。那大概是 2013 年。
That was Andrew Ng flying a helicopter backwards, a flipped around. It's a control problem which is not easy to solve, and I'm not sure if it was really AI related. I think he was saying that he was using a neural network to control all of this. But that's the latest I remember of—well, the first memory of me being shown what you could do with machine learning at the time. And so that was in 2013, I think.
最近,你在 DeepMind 待了两年半、三年。我想问,你从那段经历中最大的收获是什么?它们如何影响了你对构建 Mistral 的看法?
Most recently, you spent yeah, two and a half years, three years at DeepMind. Can I ask what are the biggest takeaways for you from that experience and how did they impact how you think about building Mistral?
一个五人团队比一个五十人团队更快,除非你把五十人团队组织成十个充分解耦的五人团队。我认为这是我在 DeepMind 通过艰难方式学到的一个发现,也是我们以略微不同的方式组织科学团队创建公司的原因。这也是为什么我们知道有机会用更小的团队做有趣的事情。
A team of five is faster than a team of 50, except if you organize the team of 50 to be 10 teams of five that are sufficiently uncoupled. And that's I think one finding that I learned the hard way at DeepMind, and the reason why we created the company in a slightly different way in terms of organization of the science team. That's also the reason why we knew we had a chance to do interesting things with a smaller team.
我能问一下吗,充分解耦——难道不会损失效率吗?或者这些孤岛之间没有泄漏,实际上这种孤岛反而造成了低效?
Can I just ask, sufficiently uncoupled—do you not lose efficiency, or is there not a leakage between those silos, and it creates actually inefficiency by having such silos?
你必须共享一些东西,所以你们共享基础设施、共享代码库、共享发现。但你知道,当你涉及——我们在做通用模型,而通用模型需要向不同方向演进,所以你需要让它们说不同的语言,让它们能够编码、做数学、推理。你需要为它们添加多模态能力。所有这些事情都是松散耦合的。如果你使用相同的框架进行优化、数据、训练,那是有用的,但你不希望你的团队整天开会协调。而且这实际上很难搞清楚。我认为到目前为止我们管理得相对不错,尽管团队只有 25 人,所以实际上并不太有挑战性。它会变得越来越有挑战性。但没错,这就是我在 DeepMind 记得的东西。一开始效果很好。Gemini 有点太慢了,我认为他们后来恢复得足够好。但没错,我们优化了团队,使其尽可能快,尽可能快地交付。
You have to share some things, so you share the infrastructure, you share the code base, you share findings. But you know, when you involve—we're doing general purpose models, and general purpose models you need to evolve them in different directions, so you need to make them speak different languages, you need to make them be able to code, be able to do mathematics, be able to reason. You need to add multimodality to them. All of these things are loosely coupled. It's useful if you use the same framework for optimization, for data, for training, but you don't want to have your team spend their entire day in meetings for coordination. And it's actually pretty hard to figure out. I think so far we've managed to scale it relatively well, although the team is only 25 people, so that's actually not super challenging. It will become more and more of a challenge. But yeah, that's what I remember from DeepMind. It worked very well at the beginning. Gemini was a bit too slow, and I think they recovered sufficiently well since. But yeah, we have optimized the team to be as fast as possible and to ship as fast as possible.
离开 DeepMind 创办 Mistral 是一个容易的决定吗?你知道,你在 DeepMind,世界上最好的 AI 机构之一,周围有一些不可思议的人才。这是一个容易的决定吗?带我去那个你决定离开去创办或联合创办 Mistral 的时刻。
Was it an easy decision to leave to start Mistral? You know, you're at DeepMind, one of the best institutions in the world for AI, with some incredible talent around you. Was it an easy decision, and just take me to that moment when you decided to leave to found or co-found Mistral.
所以这不是一个从零到一的决定,不是一个二元决定。你开始想,我有 10% 的倾向离开,然后这个比例增长,在某个时刻你跨过门槛,你说好吧,既然我已经足够决定了,我不能再多待几天,否则对我的同事不公平。这就是你开始的方式。你说没有回头路。
So it's not a zero-to-one decision, it's not a binary decision. You start to think like I'm 10% leaning on leaving, and then it grows, and at some point you cross the threshold and you say okay, well I guess now that I'm sufficiently decided, there's no way in which I stay more than a few days, otherwise I wouldn't be fair with my colleagues. And so that's how you get started. You say there's no turning back.
对你来说,那个点是什么?
And what was that point for you?
我的那个点大概是去年三月——哦不,更晚,更晚,对,大概是去年三月底,就在周五之前。我是说我在周五决定离开,周一就辞职了。
My point for me was probably around March—oh no, more, more yeah, around end of March last year, which preceded the Friday. I mean I decided to leave on Friday and I resigned on Monday.
我喜欢这样。所以你确实——如果你决定辞职就不能再待下去了。否则不太公平。
I love that. So you did—you can't stay if you've decided to resign. It's otherwise not very fair.
不,我完全同意你。
No, I totally agree with you.
现在我想按时间顺序来谈。我和你的许多顾问、投资者甚至用户都聊过,我想从第一个模型 Mistral 7B 开始,它在一段时间前发布,是最受欢迎的模型之一。你认为它为什么这么受欢迎?你认为你做对了什么,你从中学到了什么?
Now I do want to run this with some chronology. I spoke to so many of your advisors, investors, users even, and I want to start with actually the first model, Mistral 7B, being one of the most popular released a while ago now. Why do you think it was so popular? What do you think you did so right, and what did you learn from that?
我认为它有两个目的。第一个是展示在压缩模型方面有很大的空间,所以从科学角度来看,这是一个好的发现,对社区也是一个好的学习。它也填补了模型效率与性能权衡空间中的空白,那里确实缺少一些东西。7B 的大小允许你在 MacBook 或智能手机上高效运行模型,我们让它足够智能,以至于仍然有用。所以之前已经有 7B 模型,但它们不够好,无法做有趣的应用。通过瞄准这个特定空间,我们立即与开发者对话,因为开发者喜欢在游戏 GPU 或 MacBook 上运行的普通开发者。所以它引起了很多好奇和采用,因为它是性能与效率空间中的一个缺失点。
I think it served two purposes. So the first was to show that there was a lot of slack in compressing models, and so from a scientific perspective it was a good finding and a good learning for the community. It also filled the gap in the efficiency-to-performance trade-off space of models, where there was definitely something missing. 7B is the size that allows to run efficiently a model on your MacBook or on your smartphone, and we made it sufficiently smart so that it was still useful. So there were already 7B models before, but they weren't good enough to do interesting applications. And so by targeting this specific space, we talked to the developers immediately, because developers like the casual developers running on a gaming GPU or on their MacBook. So it created a lot of curiosity and adoption because it was a missing spot in the performance-to-efficiency space.
我完全理解你。当你回顾其中的教训以及它如何影响未来的发布时,有什么特别突出的吗?
I totally get you. When you look at lessons from that and how it impacts future releases, any that really stand out for you?
我想它教会我们,人们对效率而非规模扩张有浓厚兴趣,所以这就是为什么我们继续瞄准非常高效的模型,比如 Mixtral 8x7B 和最近的 Mixtral 8x22B,确保在一定的成本和大小下,我们达到市场顶尖性能。所以这一直是我们瞄准效率的主要动力,同时也在向越来越大的模型扩展。
I guess it taught us that there was a lot of interest for efficiency rather than scaling, and so that's why we continued targeting very efficient models with the Mixtral 8x7B and more recently Mixtral 8x22B, ensuring that for a certain cost and for a certain size we were reaching the top performance of the market. And so that has been our major motivation to target efficiency, while simultaneously scaling to larger and larger models.
我在节目前和 Sarah Guo 聊过,她说核心问题,我认为是你……
I spoke to Sarah Guo before the show and she said the core question that I think is you...
你现在觉得有足够的现金吗?我想初创公司总是在融资。今天 Mistral 最大的障碍是什么?
Do you feel like you have enough cash now? I guess the startup is always fundraising. What are the biggest barriers to Mistral today?
我们仍然受算力瓶颈制约,但那是因为我们没有太多算力。我们有 1500 块 H100,是竞争对手的百分之几。
We are still bottlenecked by compute for sure, but that's because we don't have much of it. We have 1.5k H100s, which is a few percent of competitors.
你没有更快地扩展算力,这是一个错误吗?
Was it a mistake for you to not scale that quicker?
你实际上无法更快地扩展那么多。你不能在种子轮就筹集 20 亿美元。我的意思是,至少在 2023 年你不能。
You can't really scale that much quicker. You can't raise like $2 billion on the seed round. I mean, at least you couldn't in 2023.
直说吧,在关注效率和效率前沿的情况下,规模重要吗?
Bluntly, with the focus on efficiency and the efficiency frontier, does scale matter?
嗯,规模很重要,因为如果你投入更多的训练算力,你可以让模型更压缩。所以确实需要一些算力来压缩模型。但规模不是唯一的配方,不是唯一的成分。你需要规模,但也需要合适的数据,否则你会遇到数据质量限制。你需要合适的训练技术。你需要弄清楚一些事情。我的意思是,人们称之为算力倍增器。如何在不增加算力成本的情况下获得效率提升,因为算力很贵?所以我们在 Mistral 做的事情之一就是尝试收获这些算力倍增器。
Well, scale matters in the sense that if you spend more training compute, you can make the models more compressed. So you do need to have some compute to compress models. But no, scale isn't the only recipe, the only ingredient to the recipe. You need to scale, but you also need to have proper data, otherwise you reach some data quality limit. You need to have proper techniques for training. You need to figure out a few things. I mean, people call it compute multipliers, I guess. How do you actually make some efficiency gains that are not costing you compute, because compute is expensive? And so one of the things that we do at Mistral is to try and harvest those compute multipliers.
我能问一下吗,在不增加算力成本的情况下,我们还能挤出多少效率?我完全外行。我们还有很多可以榨取的吗,还是已经在边际改进了?
Can I ask, in that chasm of efficiency gains without costing more compute, is there much more efficiency we can eke out? I'm totally naive here. Is there a lot for us to eke out, or are we working at marginal improvements already?
我认为这是一个开放性问题。我相信有。我相信我们可以制造出在特定规模下更好的模型。但这和另一个问题一样开放:你能在相同数据上通过扩大模型和延长训练来找到更好的模型吗?这些都是需要发现的东西。此外,在过程中,你可以尝试预测最终能达到的性能,但你需要实际尝试。所以这确实是一个研究领域。你需要做研究,需要尝试。这就是我们一直在做的。
I think it's an open question. I believe there is. I believe we can make models that are much better for a certain size. But it's as open a question as: can you find a much better model on the same kind of data by making it bigger and training it for longer? It's things you need to discover. Also, on the way, you can try and predict the kind of performance you will achieve at the end of the day, but you need to try it out. So it's really much a research field. You need to do the research and you need to try things. And that's what we've been doing.
我问过 Sam Altman 这个问题:模型格局的最终状态是什么?大多数人会说它会商品化,实际上会有 12 个玩家,然后是一场逐底竞争。在你看来,模型的最终状态是什么?你怎么看待商品化问题?
I asked Sam Altman this question: what is the end state for the model landscape? Most people say it'll become commoditized and actually there will be 12 players and it'll be a race to the bottom. What is the end state for models in your mind, and how do you think about the commoditization question?
我认为最终状态是拥有一个更成熟的开发者平台,具有更多功能,允许定制,允许制造服务于特定目的的低延迟模型,允许评估它们并随时间改进。所以模型只是很小的一部分——我的意思是,它是核心部分,但仍然是应用的一小部分。当你部署一个面向用户的应用时,你希望确保它工作正常,确保延迟随时间降低,质量随时间提高。所以我认为最终状态是:模型实际上将成为任何 AI 应用开发者的起点。它们需要被工具包围,基本上需要一个生命周期管理平台。这是我们开始构建的东西。通用模型有点无差异化,但你需要为你的应用创造的差异化来自你投入的数据、收集的用户反馈以及你用来确定应用应该做什么的智能。这完全没有商品化。没有配方能让你从通用模型变成在特定任务上超级好且优于其他所有模型的模型。我认为这是拼图中缺失的一块,这也是我们在产品方面投入力量的一个方面。
I think the end state is to have a more developed developer platform with more features that allows customization, allows making low-latency models that serve a certain purpose, allows evaluating them and improving them over time. So the model is only a tiny part—I mean, it's a central part, but it remains a tiny part of an application. And what you want to do over time, when you deploy an application that you expose to users, is to ensure that it works, ensure that its latency reduces over time, ensure that its quality increases over time. So I think that's the end state: models are effectively going to be a starting point for any AI application developer. They need to be surrounded by tools, by a lifecycle management platform basically. And that's one thing that we started to build. General-purpose models are a bit undifferentiated, but the differentiation that you need to create for your application comes from the data you put into it, the user feedback that you gather, and the intelligence that you have to figure out what the application should be doing. And that is not commoditized at all. There's no recipe that allows going from a general-purpose model to a model that is super good and better than all others at your specific task. And I think this is a missing piece in the puzzle, and that's one of the aspects where we're putting our strength on the product side.
Sam 和 Brad 前几天说,模型实际上还不太好,质量需要大幅提升。当前模型质量的最大限制或瓶颈是什么?需要改变什么才能改进?
Sam and Brad said the other day that models just aren't actually that good yet, and they need to improve a lot in quality. What are the largest constraints or bottlenecks on model quality today, and what needs to change for them to improve?
我认为数据质量是一个限制。如何确保模型利用整个世界的知识,并确保模型沿着一条路径学习越来越复杂的东西?这是一个非常重要的部分,我认为它被忽视了。显然还有算力,但考虑到我们手头的数据量,算力不再是瓶颈。对于文本到文本模型,瓶颈更多的是数据。所以问题是:如何精炼数据,如何向模型本身提供非常高质量的数据以随时间改进?我认为在这种情况下,与提升模型性能相关的一个瓶颈是如何评估这些性能。你需要有非常好的评估,针对非常具体的主题。例如,你希望模型擅长在医院帮助诊断,但用法语。很多时候,你与你拥有的数据相比有点超出领域。这就是你应该识别差距并尝试填补的地方。所以推动模型能力也变成了一个关于映射它们失败之处并找出改进方法的问题。例如,它们在数学上失败。如何改进它们的数学思维?如何改进它们证明定理的方式?这个问题的答案与如何改进法语医疗诊断的答案非常不同。
I think the data quality is a constraint. How do you ensure that the model leverages the entire world knowledge and ensures that the model follows a certain path toward learning more and more complex things? That's a very important part, and I think it has been a neglected part. There's obviously compute, but given the amount of data we have at hand, compute is no longer the bottleneck. The bottleneck is more the data at that point, if you look at text-to-text models. So the question is: how do you refine the data and how do you feed very high-quality data to the model itself in order to improve it over time? And I think in that setting, one bottleneck associated with bringing better model performance is the question of how do you evaluate these performances. You need to have very good evaluations that target very specific topics. For example, you want the model to be good at helping diagnosis in hospitals, but in French. And often times, you're a bit out of domain compared to the data you have. That's where you should identify a gap and try to fill it out. So pushing the model capabilities becomes also a question of mapping where they're failing and figuring out ways of improving it. For instance, they're failing at mathematics. How do you improve their mathematical thinking? How do you improve the way they demonstrate theorems? And the answer to this is very different from the way you answer the question of how to improve a medical diagnosis in French, for instance.
我们会看到能够回答大量非常复杂问题的大规模通用模型,还是你认为我们会看到更多垂直特定、更小、更具体的模型,这些模型更加垂直对齐?
Will we see large-scale generalized models that are able to answer huge sets of very complex problems, or do you think we'll see much more vertically specific, smaller, more specific models that are much more vertically aligned?
我们相信,实际上这些垂直模型不会直接存在;它们将由应用制造者构建。因为制造一个在特定任务上非常出色的低延迟模型的唯一方法是去掉通用方面。通用模型有点臃肿;你可以思考一切。但如果你希望你的模型深入思考一个特定主题,以便在你的 AI 应用中调用它,同时保持低延迟的良好用户体验,你需要专业化。
We believe that, and actually these vertical models are not going to be out there; they are going to be built by the application makers. Because the only way you can make a low-latency model that is super good at a specific task is to get rid of the general-purpose aspect. A general-purpose model is a bit bloated; you can think about everything. But if you want your model to think thoroughly about a specific topic so that you can call it in your AI application while maintaining a good user experience with low latency, you need specialization.
那么你在那个世界中扮演什么角色?抱歉问这么直白的问题,但如果具体的模型创建发生在应用层,那种价值也发生在那里,你在其中扮演什么角色?
What role do you play in that world then? I'm sorry for asking blunt questions, but if it's actually in the application layer where that specific model creation happens, where that kind of value occurs, where do you play in that?
制造一个专门的模型是非常困难的工作。所以它实际上与你创建预训练模型的方式紧密相关。因此,提供工具,让开发者能够以万无一失的方式创建定制模型,这些模型在任务上表现非常好,但不需要难以找到的专家知识,这绝对是我们坚持的事情。
It's a very hard job to make a specialized model. So it's actually very tied to the way you create a pre-trained model. And so bringing the tools that allow to do it in a foolproof way—allowing developers to create customized models that are performing very well at their task but that doesn't require expert knowledge, which is hard to find—is definitely something where we're insisting.
我今天是一名投资者,我很高兴你刚刚说价值会在应用层积累,因为我看着并担心坦白说一切都会被我们提到的一些玩家碾压。你怎么回答这个问题?
I'm an investor today, and I'm pleased that you just said that there will be value accrued at the application layer, because I look and I worry that bluntly everything is going to get steamrolled by some of the players that we mentioned. How do you answer that question?
你真心建议我该怎么做?有两个相反的方向。第一个是模型越来越好,所以只要你有数据并且对用例有很好的理解,创建垂直化应用会越来越容易,前提是你有合适的工具。这让我觉得应用层会变得越来越薄。但另一方面,模型也在变得更便宜,因为我们能够压缩它们,并在效率上做了很多改进。所以实际上,加上模型层的竞争压力,每单位智能的价格肯定会下降。所以有这两个方面:能力提升和价格压缩。一方面应用层变薄,另一方面模型部分变大。对我们来说,我们采取的方法是假设模型部分仍然足够大,我们需要在其上构建这个平台,因为这样我们才能实现所有对你来说有趣的垂直应用。
How would you advise me genuinely? There's two opposing directions. The first is that the models are getting better and better, so creating a verticalized application, as long as you have the data and a good understanding of the use case, is going to be easier and easier if you have access to the tools that facilitate it. That would make me think the application layer is going to grow thinner and thinner. But then there's also the fact that the models are getting cheaper because we manage to compress them and make a lot of improvements on their efficiency. So effectively, this plus the competitive pressure on the model layer means the price per intelligence unit is definitely going to reduce. So there are these two aspects: growing ability and compressed price. On one side, the application layer grows thin; on the other, the model part grows in. For us, the approach we are taking is to assume that the model part is still going to be big enough and that we need to build this platform on top of it, because that's where we will enable all the vertical applications that will be interesting for you.
你如何看待这种定位和品牌?因为还有其他玩家更直接地说,‘嘿,我们要主宰很多不同的垂直领域,有点让人害怕。’你怎么看?
How do you think about that positioning and brand? Because there are other players who are much more direct in saying, 'Hey, we're going to dominate a lot of different verticals and kind of be afraid.' How do you think about that?
我们不是一家垂直化公司。我们创立 Mistral 是为了给开发者带来价值,给开发者带来自由。当我们开始时,基本上只有一个 API,很快有两个,生成式 AI 领域开始看起来会围绕少数几个玩家集中化。我们采取了这种平台方法,我们制作的模型和技术,允许开发者拥有它、修改它,所以给开发者和 AI 应用制造者带来自由,我认为这是尽可能广泛地分发生成式 AI 的最佳方式,这也是我们公司的目标:让前沿 AI 进入每个人的头脑,这是我们创立的原因。我认为我们在这方面做得很好。我们还有很多事情要做,但开源部分,我相信,是社区的一个良好推动力,让人们意识到他们可以通过修改模型本身来构建非常有趣的技术,而不是依赖少数几个提供商的 API。
We are not a verticalized company. We started Mistral to bring value to developers and to bring freedom to developers. When we started, there was basically one API out there, soon two, and the field of generative AI was starting to look like it would be very centralized around a couple of players. We took this platform approach where the model and technology we are making, we allow developers to own it, modify it, and so bringing freedom to developers and AI application makers is, I think, the best way to distribute generative AI as widely as possible, which is our objective as a company: making a quitus, bringing frontier AI into everyone's head, is the reason why we started. I think we did a good job at it. We still have a lot of things ahead, but this open source part was, I believe, a good enabler for the community and made people realize they could build very interesting technology by modifying the models themselves instead of depending on the APIs of a couple of providers.
AI 开发者关心什么?每个人都在推特上说,‘哦,你看到 X 这周的表现比 Y 上周的表现好吗?’他们关心什么:效率、规模、成本?是什么驱动了使用和决策?
What do AI developers care about? Everyone kind of gets on Twitter and goes, 'Oh, did you see X's performance this week is better than Y's performance last week?' What do they care about: efficiency, scale, cost? What drives that usage and decision making?
他们当然关心成本。他们关心定制化,能够修改模型。在这方面,我认为我们只是触及了可能性的表面。微调作为首选解决方案,可能有点过于底层,不是我们应该做的。他们关心能够部署在任何地方。所以他们在特定的空间、特定的云上操作。他们可能是在本地操作,可能有边缘设备要部署,他们希望把技术放在那里。所以他们也关心可移植性,这反过来提供了数据控制。当你将 LLM 和 AI 连接到知识库或与特定业务相关的任何东西时,它们变得非常有用。在这方面,它成为你应用中非常敏感的部分,因为它能看到一切,你所有的数据。所以企业,例如,确实关心确保他们拥有的专有数据在完全安全的环境中被访问。这就是为什么我们将平台部署在 Azure 和 AWS 上,提供他们所需的安全层。
They care about cost for sure. They care about customization, being able to modify the models. On that aspect, I think we are only scratching the surface of what can be done. The fine-tuning aspect that has been the go-to solution is probably a little too low-level from what we should be doing. They care about being able to deploy anywhere. So they operate in a certain space, in a certain cloud. They might be operating on-prem, they might have some edge devices to deploy to, and they want to be able to put their technology there. So they also care about portability, which in turn offers data control. LLMs and AI become very useful when you connect them to knowledge bases or anything related to a certain business. In that respect, it becomes a very sensitive part of your application because it sees everything, all the data you have. So enterprises, for instance, do care about ensuring that the proprietary data they have is accessed in something that they can completely secure. That's the reason why we deployed our platform on Azure and AWS, for instance, bringing the security layer that they need.
我们马上要谈到企业。我能问一下:品牌在这个领域重要吗?当我们考虑建立品牌时,无论是开发者采用品牌还是企业品牌,品牌是这个领域采用率的重要决定因素吗?
We're going to get to enterprise. Can I just ask: does brand matter in this segment? When we think about building brand, both in terms of developer adoption brand and corporate brand, is brand a large determinant of adoption in this segment?
品牌似乎至关重要。这是我们一路走来学到的。人们使用某些模型是因为它们以优秀著称。你不可能评估所有东西。所以有某种形式的社区支持非常重要。我们采用的开源分布式模型方法促成了我认为至少是一个知名品牌。我们相信这绝对会很重要。品牌之所以重要,是因为信任在这个领域很重要,而开源在提供可信品牌方面带来了信任。
Brand seems to be critical. This is something we have learned along the way. People use certain models because they are known to be good. You can't afford to evaluate everything out there. So having some form of community backing is super important. The approach we took with open-source distributed models has contributed to what I think has become at least a known brand. We believe it's definitely going to be important. Brand is important because trust is important in that domain, and open source brings trust in terms of providing a trusted brand.
你提到了‘开源’这个词。我接下来会谈到这个。我只想提一下,你也提到了成本。一个难题:在基于 LLM 的产品中,如何、何时以及谁会实现边际收入超过边际成本?
You mentioned the word 'open source' there. I'm going to get to that. I do just want to touch on that you mentioned cost also. Hard question: how, when, and who will make marginal revenue that exceeds marginal cost in LLM-based products?
你应该告诉我。你是投资者,对吧?所以我想你有自己的模型。这意味着我什么都不知道。好吧,我可以告诉你目前谁利润率最高,但可能会随时间变化。目前谁利润率最高?英伟达。在那一点上,云提供商基本是成本价。LLM 提供商不是成本价,希望如此,但利润率已知低于典型的软件利润率。AI 应用制造商,其中一些最常用的,似乎利润率不错。我认为这将是一个相当动态的领域。正如我所说,模型的能力使得制作应用的成本越来越低。我不认为边际成本和该技术最重要部分(即基础层)的利润率会变为零,因为否则肯定会出现公平性问题。
You should be telling me. You're the investor, right? So I guess you have your own model. That means I know nothing. Okay, I can tell you who is doing the most margin at the moment, but it's probably going to evolve over time. Who is doing the most margin at the moment? Nvidia. At that point, the cloud providers are pretty much at cost. LLM providers are not at cost, hopefully, but the margins are known to be lower than typical software margins. AI application makers, some of them, the ones that are most used, seem to be doing a pretty good margin. I think it's going to be quite a moving space. As I've said, the capacity of models makes the cost of making an application lower and lower. I don't think there's any way in which the marginal cost and the margin of the most important part of that technology, which is really the foundational layer, become zero, because otherwise there's definitely going to be a fairness problem.
你说的公平性问题是什么意思?跟我谈谈。
What do you mean by the fairness problem? Talk to me about that.
通常,价值倾向于集中在最困难的部分和最具防御性的部分。一段时间以来,它一直集中在基础模型上。我认为它显然会随着时间演变,没有模型不会消失或演变。那将仍然是大部分创新发生的地方,也是大部分(至少是重要部分)累积价值所在的地方。价值将在那里累积。
Usually, value tends to accrue where most of the difficult part is and most of the defensibility is. It has been on foundational models for a while. I think it's obviously evolving with time, and there's no model that isn't disappearing or evolving with time. That will remain the part where most of the innovation will be made and where most of the, well, at least the significant part of the accrued value will be. The value will accrue there.
如今创建一家基础模型公司实际上有多大障碍?
Is there actually much of a barrier to creating a foundational model company today?
我知道这在很多方面是一个非常宽泛且愚蠢的问题,但现在有这么多不同的参与者,每天都有新的出现。门槛是不是每天都在降低?
I know that's a really broad, stupid question in many respects, but you have so many different players now and new ones popping up every day. Is the barrier just reducing day by day?
我不这么认为。要在那个领域保持相关性是一个非常困难的话题。你需要在成本效率、性能和 BYO 方面占据主导地位。目前只有少数公司处于有利位置。你可以尝试做一些事情,但如果它不相关,如果它被另一个模型或技术严格主导,那么你就有问题了。有几个很难面对的障碍:你需要筹集足够的资金,拥有足够的算力来保持相关性,拥有知道如何训练模型的人——这仍然是一种稀缺资源——然后你需要一个好的品牌,因为竞争非常激烈。这不是凭空而来的。所以我认为市场上仍然有很多防御性,尽管有很多噪音,但这是不同的。
I don't think it is. To be relevant in that space is a very hard topic. You need to be dominating on cost efficiency, performance, and the BYO front. There are only a few companies that are currently well positioned. You can try to do something, but if it's not relevant, if it's strictly dominated by another model or technology, then you have a problem. There are a few barriers that are pretty hard to face: you need to raise sufficient capital, have enough compute to be relevant, have people who know how to train models—which is still a scarce resource—and then you need a good brand because it's highly competitive. This doesn't come out of thin air. So I think there's still a lot of defensibility in the market, although there is a lot of noise, which is different.
你认为算力成本下降的速度有多快?因为如果你看看这些事情——你提到了算力成本、人才获取和品牌——如果我们大幅降低算力成本,就像许多人认为的那样很快就能实现,那么你就能获得人才和品牌。这两者更容易实现。
How quickly does the cost of compute go down, do you think? Because if you look at those things—you said cost of compute, access to talent, and brand—if we drastically bring down the cost of compute, like many think we will very quickly, you've got access to talent and brand. The two of those are more doable.
算力成本仅基于硬件成本就会随着时间的推移而降低。如果你按照英伟达的路线图,对于相同的浮点运算量,每两年大约降低 30%。另一件增加的事情是算法的效率。如果你看看三年前我们训练模型的方式和今天训练模型的方式,我认为我们可能已经取得了大约 100 倍的算法改进。所以这可能是过去几年实际取得大部分收益的地方。显然,算力成本确实在降低,但它降低的速度并没有护城河快。所以我们更看好效率,我认为在这方面还有很多改进空间。
The cost of compute reduces over time just based on hardware cost. It reduces around 30% every two years if you follow Nvidia's roadmap for the same amount of flops. The other thing that increases is the efficiency of algorithms. If you look at the way we train models from three years ago and the way we train models today, I think we have probably made around 100x algorithmic improvement. So that's probably where most of the gains were actually made in the last years. Obviously, the cost of compute does reduce, but it doesn't reduce faster than the moat. So our bet is more on efficiency, where I think there's a lot of improvement that can still be made.
鉴于英伟达在那里的突出地位,正如你提到的,英伟达是收益的来源,这直白地说是不是最重要的事情之一?不仅仅是与核心提供商——比如 AWS、英伟达或这些参与者之一——的关系质量,这难道不是当今成功的核心决定因素吗?
Given Nvidia's prominence there, and Nvidia being the one where the gains are, as you mentioned, is bluntly one of the single most important things? Not simply the quality of your relationship with the core provider—being AWS, or being Nvidia, or being one of these players—is that not the core determinant of success today?
我想这是一个重要的方面,是的。AI 层对云提供商和英伟达存在战略依赖。竞争也在加剧,但这确实很重要。当你开发软件时,了解硬件提供商也很有用,因为他们可以帮助你针对硬件进行优化。当你向企业销售开发者平台时,通过他们通常的提供商(恰好是云提供商)来推广该平台也很有用。所以那里肯定有一些重要的合作要做。
I guess it's an important aspect, yes. There is a strategic dependency from the AI layer on the cloud providers and on Nvidia. The competition is heating up as well, but it's effectively important. It's effectively useful when you develop software to also know the hardware provider because they can help you optimize for the hardware. It's useful when you're selling your developer platform to enterprises to bring that platform through their usual provider, which happens to be a cloud provider. So there's definitely some important collaboration to be made there.
当像亚马逊向 Anthropic 投资 20 亿美元之类的事情发生时,这难道不是一种交易,即 Anthropic 随后在亚马逊上花费 18 亿美元并返还给他们?你明白我的意思吗?这是不是有点用词不当?看起来像是资金循环。
When like Amazon invests $2 billion in Anthropic or whatever it was, is that not just like a trade where Anthropic then spends $1.8 billion on Amazon and returns it back to them? Do you see what I mean? Is it not a bit of a misnomer? It looks like you're round-tripping.
我不特别了解那笔交易,但从双方的角度来看,这是有道理的。
I don't know about that deal particularly, but it makes sense from both perspectives.
我能问一下,开源 LLM 的无限可用性如何影响上述问题的答案——边际成本和边际收入?它改变了很多吗?
Can I ask how does the unlimited availability of open source LLMs impact the answer to the above—being marginal cost and marginal revenue? Does it change much?
它将价值稍微提升到模型本身之上。它将价值转移到平台和定制部分,这确实是我们所期望的,并且它加速了这一过程。
It moves the value a little higher than the model itself. It moves the value to the platform and customization part, which is really something that we're expecting, and it accelerates that process.
你一开始完全开源,非常向社区开放。现在你有小模型开放,然后大模型封闭。我说得对吗?
You started off completely open source, very much open to the community. Now you have small models open and then larger ones closed. Am I right?
我们现在也有大模型是开放的,因为我们发布了——我的意思是,取决于大小阈值——但 A* 22B 实际上按任何标准都相对较大。
We also have large models that are open now, because we released—I mean, depends on the threshold for small and large—but the A* 22B is actually relatively large by any standard.
那么关闭一些模型的决定背后是什么?仅仅是一个需要赚钱的商业案例吗?
What was behind the decision then to close some models? Is it just a business case where you need to make money?
机会主义地,有一个机会可以利用该资产作为我们销售的东西来发展业务。我们仍然在商业模型的基础上发展业务,所以这也是巩固与云提供商战略关系的好方法。这种情况将继续下去。我们仍然打算成为开源部分的领导者,并拥有一些可以许可的独特资产,以及一些开发者可以使用的独特平台。
Opportunistically, there was an opportunity to grow the business using that asset as something that we were selling. It's still the case that we're growing our business on top of commercial models in particular, so also a good way of cementing some strategic relationships with cloud providers. And it's going to continue to be the case. We still intend to be a leader in the open source part and to have some unique assets that we can license and to have some unique platform that developers can use.
当你突然有一些封闭的模型并开始建立企业团队时,你觉得这有多难?作为创始人,你如何看待研究团队和销售团队之间的平衡,以及确保两种文化融合?
How do you think that it's hard when you suddenly have some closed and you start building an enterprise team? For you as a founder, how do you think about that balance between a research team and a sales team and making sure that the two cultures come together?
我认为重要的一点是建立同理心。确保科学团队也理解用户面临的问题。这有助于改进科学,因为归根结底,我们制造的通用技术只有在识别用例时才是通用的。所以确保科学团队对产品和业务团队有相对直接的接触,实际上很重要,可以让他们了解模型在哪里失败以及如何显著改进。另一方面,市场推广团队必须理解这是一个非常技术性的销售动作,因为你销售的不是产品,而是为产品提供动力的东西。所以你需要告诉客户这些东西应该如何被用来实际创造对业务有价值的东西。这只能通过市场推广团队的强力赋能来实现。所以这是一个挑战。他们不在同一个规模上运作:科学团队的周期是几个月,市场推广团队周期更短,速度更快。但我认为到目前为止,我们已经成功招募到一些有技术兴趣的市场推广人员和有商业兴趣的技术人员。我认为这就是确保最终不会出现孤岛的方法。
I think some important thing is to create empathy. Ensure that the science team also understands the problems that the users are facing. It improves the science because at the end of the day, the general-purpose technology we're making is only general purpose if you identify the use cases. So ensuring that the science team has some relatively direct exposure to the product and to the business team is actually important to make them understand where the model is failing and how it could be improved significantly. On the other side, the go-to-market team has to understand it's a very technical sales motion because you're not selling the product, but you're selling something that is going to power the product. So you need to tell the customer how these things should be used to actually make something that brings value to the business. That only goes through strong enablement of the go-to-market team. So it's a challenge. They don't operate on the same scale: the science team has cycles of several months, the go-to-market team goes faster with shorter cycles. But I think so far we've managed to recruit go-to-market people that have some technical interest and technical people that have some business interest. I think that's how you ensure that you don't have silos at the end of the day.
当我们进入企业领域时,我担心的一个问题是,品牌在企业方面非常重要,而且他们已经与微软有现有协议。我担心实际上产品或模型质量不如分销重要。微软只是将现有客户与新产品的附加。你如何看待这个核心挑战?我担心错了吗?
One of my worries with this space as we move into enterprise is that brand matters so much in terms of enterprise actually, and they already have existing agreements with Microsoft. I worry that actually product or model quality doesn't matter as much as distribution. Microsoft just tacks on existing clients with new products. How do you think about that as a core challenge? And am I wrong to be worried?
你认为开源是否已为企业做好准备,还是企业已为开源做好准备?他们对此足够关心吗?
Do you think open source is ready for Enterprise, or do you think Enterprise is ready for open source? And do they care about it enough?
这取决于企业。一些早期采用者已经在生产中使用大量模型,所以它们肯定足够准备好了。现在,要让他们进入大规模生产的下一个阶段,我认为他们仍然缺乏一些围绕管理、负载均衡和定制模型的产品。你可以用 DIY 解决方案来做,但要使其健壮且可扩展并不容易。此外,要提高定制模型的质量,配方有点难设定。所以技术最先进的企业肯定准备好了,而且已经有很多使用开源模型的生产用例。但要扩大采用,肯定还需要一些工具推向市场。
It depends on the Enterprises. Some have been early adopters and are using a lot of models in production, so for sure they're ready enough. Now, to bring them to the next level of large-scale production, I think they're still lacking some product around managing, load balancing, and customizing the models. You can do it with DIY solutions, but making it robust and scalable is not easy. Also, to increase the quality of custom models, the recipes are a bit hard to set. So the most technically savvy Enterprises are definitely ready, and there are many use cases in production using open source models. But to widen adoption, there's definitely some tooling to be brought to the market.
如今每家企业都在会议室里问自己的 AI 战略是什么。你给他们什么建议,他们应该问什么问题?
Every Enterprise today is in a boardroom asking what their AI strategy is. What do you advise them, and what questions should they be asking?
不要再想着如何以 AI 为前提改变所有产品,利用非常智能的智能体的存在。假设这种存在,然后反向推导理解组织层面的后果。不要把生成式 AI 看作是提高文字处理生产力的方式,而是看作彻底改变核心业务运营方式的手段。这通常涉及大量定制模型,以创造五年后当所有人都将技术应用于核心业务时你所需要的差异化。
Stop thinking about how they're going to change all their products using AI as a premise, using the existence of very clever agents. Assume that presence and work backward to understand the consequences in terms of organization. Don't think about generative AI as a way to increase productivity in word processing, but rather as a way to completely change how you operate your core business. That usually involves taking models and customizing them heavily to create the differentiation you'll need in five years when everyone has adopted the technology in its core business.
我担心我们严重高估了近期采用率,而低估了 10 到 20 年后的采用率。你认为情况是这样吗?你是否担心许多企业(尤其是欧洲企业)的惰性?
My concern is that we drastically overestimate adoption in the near future and underestimate it in the 10-20 year future. Do you think that's the case, and do you worry about the lethargy of many enterprises, especially in Europe?
科技领域的一个普遍现象是,你总是高估速度而低估影响。今天可能也是如此,但略有不同,因为即使在欧洲,也有高管支持推动生成式 AI 解决方案。与美国市场相比有延迟,但并不显著。挑战在于这项技术可以采取多种形式,因此专注于将带有 AI 的特定产品推向市场是一个优先级挑战。一旦他们更多地尝试现成解决方案,并意识到有开发者平台可以让他们无需内部雇佣昂贵的 AI 科学家就能做到,事情就会变得更容易。我们预计未来几年这将加速。
It's a general phenomenon in tech that you always overestimate the speed but underestimate the impact. It's probably occurring today, but it's slightly different in that there's some executive support for pushing generative AI solutions even in Europe. There's a delay compared to the US market, but it's not very significant. The challenge is that this technology can take many forms, so focusing on specific things to bring to market with AI is a prioritization challenge. It will become easier once they try off-the-shelf solutions more and realize there are developer platforms that allow them to do it without hiring expensive AI scientists in-house. We expect this to accelerate in the coming years.
你认为我们仍然在实验预算的游戏中,还是正在进入核心预算?
Do you think we're still playing in the experimental budget game, or moving into core budgets?
这取决于企业。在客户支持方面,AI 的应用非常明显,正在进入核心预算。在其他许多功能以及电信和医疗等行业的核心应用中,仍处于实验阶段。但我认为明年会有所发展。
It depends on the Enterprise. It's moving into core budget for customer support, where the application of AI is pretty obvious. It's definitely moving into core budget there. It's still at the experimental stage in many other functions and for core applications in industries like telecom and healthcare. But I think it will evolve in the next year.
在算力和人才之外,构建企业级产品也很昂贵。来自 Lightspeed 的 Paul 提到,与 OpenAI 和 Anthropic 等竞争对手相比,你们筹集的资金少得多。在一个资本等于算力等于模型质量的世界里,Mistral 如何跟上并保持相关性?
Building out Enterprise is expensive on top of compute and talent. Paul from Lightspeed mentioned how much less capital you've raised compared to competitors like OpenAI and Anthropic. In a world where capital equals compute equals quality of model, how does Mistral keep up and stay relevant?
资本与算力相关,算力与质量相关,但并不完全依赖于此。提供同类最佳模型有很大的机会,因为它们可能足以解决某些用例。这就是我们发力的地方,同时也在规模扩张方面努力。要保持相关性,你需要让你的技术团队保持积极性,为此你需要给他们提供实验平台,让他们做出新发现并推动科学进步。这就是你需要算力的地方,此外还要随时间推移扩展模型。我们像每家公司一样在增长算力,但我们相信不需要以同样的速度增长,因为沿途出现了许多与算力无关的障碍。我们认为我们可以扩展,而且我们也相信在效率方面,我们已经处于有利地位,并且正在加强这一地位。
Capital is correlated to compute, and compute is correlated with quality, but it's not completely dependent on it. There's a strong opportunity for providing models that are the best of their class because they might be sufficient to solve certain use cases. That's where we're playing, in addition to playing on the scaling part. To stay relevant, you need to keep your technical team motivated, and to do that you need to give them the experimental bed they need to make new discoveries and progress science. That's where you need compute, in addition to growing the model over time. We are growing our compute like every company, but we are convinced we don't need to grow at the same rate because there are many barriers that are not compute-related appearing along the way. We think we can scale, and we are also convinced that on the efficiency front, we are already well positioned and strengthening that position.
Mistral 今天面临的最大障碍是什么?
What are the biggest barriers to Mistral today?
我们的算力提供商有一些延迟,这已经成为一个障碍。我们仍然受算力瓶颈制约,因为我们没有多少算力。我们有 1500 块 H100,是竞争对手容量的百分之几。这绝对是一个瓶颈,将在未来几个月显著改善。
We've had a few delays with our compute providers, which has been a barrier. We are still bottlenecked by compute, because we don't have much of it. We have 1.5k H100s, which is a few percent of our competitors' capacity. That's definitely a bottleneck that will improve significantly in the coming months.
事后看来,没有更快地扩展是个错误吗?你希望自己当时扩展得更快吗?
With hindsight, was it a mistake not to scale quicker? Would you wish you had scaled quicker?
你无法真正更快地扩展,因为你不能在种子轮融资 20 亿美元,至少在 2023 年不能。你只能以那样的速度招聘、扩展基础设施以管理更多 GPU,以及筹集资金。存在一些很难克服的加速约束,这基本上是创业的第一性原理。
You can't really scale that much quicker because you can't raise two billion on a seed round, at least you couldn't in 2023. You can only hire that fast, scale your infrastructure to manage more GPUs that fast, and raise capital that fast. There are acceleration constraints that are pretty hard to fight, and they are pretty much the first principles of starting a business.
关于 Scaling(规模扩张)的限制和资金:资金来源重要吗?是欧洲资金、沙特资金还是美国资金,这有关系吗?你觉得这重要吗?
About the scaling constraints and cash: does it matter where your cash comes from? Does it matter if you have European funded, Saudi funded, US funded? Did you think that matters?
我认为治理很重要。对于像我们这样的年轻公司来说,重要的是由创始人控制,因为有很多东西需要发明,愿景只能由他们——也就是我们——来承载。我们有非常好的治理条款,简单而清晰,使我们成为一家营利性公司,通过发展业务来推动科学前沿。我们非常看重这一点:能够控制公司,适当利用我们的投资伙伴在全球不同地区——美国、欧盟——发展,这也至关重要。在世界上对 AI 兴趣浓厚的其他地区也是如此。所以这确实重要,因为我们希望合作伙伴是支持性的、长期的,因为我们身处一个快速发展的领域,价值——我们还不确切知道价值会在哪里积累。因此,在融资时,灵活和聪明绝对是必要条件。
I guess governance matters. What is important for a young company like us is to be under the control of the founders, because there's a lot of things to be invented and the vision can only be carried by them, by us. We have very good governance terms, a very simple and clean governance that makes us a for-profit company that is growing a business to actually push the science frontier. And this is something that we're very attached to: being able to control the company, leverage our funding partners appropriately to grow in different parts of the world, in the US, in the EU, it has been critical as well. And in different parts of the world where there's a lot of interest for AI. So it does matter in the sense that we want to have partners that are supportive and long-term, because we are in a field that is fast moving, where value, we don't know yet exactly where the value will accrue. So being flexible and being smart is definitely a requirement when you raise money.
你会接受沙特或中国的资金吗?
Would you take money from Saudi or China?
好问题。这取决于条款。中国对我们来说有点难,甚至在中国运营都很困难。我的意思是,我们不在中国运营,因为除非你是一家非常非常大的公司,否则你无法同时在美国和中国运营。所以你需要做出一些选择。
Good question. It depends on the terms. China is a bit hard for us; it's even hard to operate in China. I mean, we don't operate in China because you can't really operate in the US and China without being a very, very large corporation. So you need to make some choices.
你认为欧洲在 AI 方面有多大机会?我知道这听起来有点宿命论和失败主义,你可能会说‘Harry,闭嘴’,但问题是:你认为欧洲在 AI 方面有多大机会,我们需要做些什么才能在欧洲建立起一个严肃的 AI 产业?
What chance do you think that Europe has in AI? I know it sounds deterministic and defeatist, and so you might be like 'Harry, shut up', but it's like: what chance do you think Europe has in AI, and what does it take for us to stand up as a serious AI industry in Europe?
我认为它的机会在于这是一场革命,正在改变我们做软件的方式。因此,像每一次革命一样,它为新的参与者打开了大量机会。没有理由不能有一个在欧洲创立并能够快速成长的参与者。这就是我们赋予自己的使命。我们有人才;资本可以跨洋流动,问题不大。我们有市场;市场肯定比美国更碎片化,生态系统——数字原生生态系统——肯定更小,但它存在并且正在增长。所以有本地业务发展的机会。在人才方面,我们可以雇佣 23-24 岁的年轻人,在四个月内让他们入职,他们的表现和硅谷的任何软件工程师一样好。所以这里的人相当有才华。因此,如果我们能留住他们,并说服他们不去美国,我们就有很多机会。
I guess the chance it has is that it's a revolution and it's changing the way we do software. And so, as every revolution, it opens a lot of opportunity for new actors. And there's no reason why there shouldn't be an actor that was created in Europe and that could grow pretty fast. And that's the mission that we gave ourselves. We have the talent; capital can cross oceans without too much problem. We have the market; the market is more fragmented than in the US for sure, the ecosystem, the digital native ecosystem is definitely smaller, but it exists and it's growing. So there's local opportunity for business development. On the talent side, we can hire 23-24 year old people that we can onboard in four months, and they operate as well as any software engineer in the Valley. So people are quite talented here. And so if we manage to keep them and to convince them not to go to the US, we have a lot of opportunities.
当我们回顾计算机、移动、云这些核心技术的转变时,运作方式通常是欧洲将控制权让给了美国,然后仅仅通过向美国公司征税来获取对我们公民的访问权。如果持失败主义态度,现在有什么不同吗?我的意思是,欧洲正在为 60 年代没有建立风险投资体系付出代价,然后晚了 40 年甚至 50 年才建立起来。顺便说一句,一个肮脏的秘密是,欧洲的风险投资生态系统是由美国资金资助的。是的,过去是,我认为现在很大程度上仍然是。有政府机构在填补空缺,但很大程度上是用不太好的参与者来填补。而欧洲最好的提供者大多由美国资金支持,背后是美国顶级机构。
When we look at computers, mobile, cloud, the kind of core technology shifts, the way that it's worked is Europe has kind of ceded control to the US and then just taxed US companies for access to our citizens. If one's being defeatist, is it different now? I mean, Europe is paying the price of not setting up a VC system in the 60s, and then setting it up like 40 years later or even 50 years later. And by the way, the dirty secret is that the VC ecosystem in Europe is US funded. Yeah, it was, I think it is still, honestly, in large part yes. There's government institutions which are backfilling it, but largely backfilling it with bad players who aren't very good. But the best providers in Europe are largely US-funded, backed by top US institutions.
好的,我认为,正如我所说,建立一个生态系统需要时间。所以你有层层叠叠的企业家和投资者。美国有 60 到 70 年的风险投资历史。我认为欧洲最多只有 20 年。这意味着,建立生态系统需要时间,这是不可压缩的时间。它还需要一些意志力。我认为现在我们正在看到这种意志力。我们看到企业家在创建公司,我们看到他们不去美国。所以我认为一切都是积极的。只是需要时间。我坚信我们会做出一些有趣的事情。
Okay, I think, as I've said, it takes time for an ecosystem to build. So you have layers of entrepreneurs and investors that stack on top of each other. And the US has 60 or 70 years of venture capital investments. I think Europe has only 20 years maximum. And that's, I mean, it takes time, it takes an incompressible time to build an ecosystem. It takes also some willpower. And I think now we're seeing that willpower. We are seeing entrepreneurs creating companies, we are seeing them not going to the US. And so I think that's everything positive. It just takes time. And I'm adamant that we'll manage to do something interesting.
在工程方面,随着公司扩张,你觉得你有足够深度的人才库可以招聘吗?
On the engineering side, do you feel like you have the depth of talent pool to hire from as you scale?
在工程和 AI 方面,我们确实有。不过我们在美国有一个团队,负责特定主题。对于高级 AI 科学家,你在硅谷比在法国更容易找到。对于初级 AI 科学家,法国、波兰、英国有大量人才。我认为这是该地区的优势之一。
On the engineering side and on the AI side, we do. We have a team in the US though, which is working on specific topics. For senior AI scientists, you find them more in the Valley than in France. For junior AI scientists, there's a wealth of talent in France, in Poland, in the UK. And that's, I think, one of the strengths of the area.
你在融资时,与欧洲投资者和美国投资者交谈有很大不同吗?
When you were raising money, was it very different speaking to European investors versus US investors?
我想在种子轮,不,没有太大不同,因为那是种子轮。对于 A 轮,那是一轮更大的融资,我们,我的意思是,欧洲基金的结构不适合做我们提出的那种交易。所以我们甚至没有太多对话,因为他们无法理解需要进行的投资规模,而我们还是一家没有收入的公司。是的,我认为欧洲缺少的是能够以高度信念进行大额押注的增长基金。这反过来应该会随着时间的推移而改善,尤其是如果我们能够利用欧洲的财富,并将其更多地引导到这些增长基金中,而不是像现在这样。
I guess in the seed round, no, it wasn't that different because it was a seed round. For the Series A, which was a bigger round, we, I mean, European funds weren't structured to do the kind of deal that we were proposing. So we didn't even have a lot of conversation because they just couldn't get their head around the investment that needed to be made, whereas we were a pre-revenue company. Yeah, I think what is lacking, and it's related to the ecosystem part, in Europe are growth funds that are able to take huge bets with lots of conviction. And that in turn should improve over time, especially if we manage to use European wealth and channel it more into those growth funds than it is today.
是的,我认为你比我更有希望。这不会发生。未来几年我们不会看到更多欧洲增长基金建立起来,肯定不是未来三到五年。
Yes, I think you have more hope, you know, I think you have more hope than me on that one. That is not going to happen. We are not going to see many more European growth funds be built in the next few years, for sure not in the next three to five.
是的,这取决于一些政治决策。我认为这取决于资本供应和对未来欧洲生态系统能够与其他大型生态系统竞争的信心。这是一个先有鸡还是先有蛋的问题。但在某个时候,你确实需要——我的意思是,如果政治愿意推动,如果几家公司证明你实际上可以在欧洲拥有快速成长的公司,那么这可以被推向正确的方向。这就是我们正在努力做的。我并不太悲观。我觉得你太悲观了。你应该来法国;你应该来法国。我想你会变得更乐观。
Yeah, it hinges on a few political decisions. And I think it hinges on supply of capital and belief in a future European ecosystem that can contend with other large ecosystems. It's a chicken and egg problem. But at some point you do need to, I mean, this could be nudged into the right direction if politics wants to do it, if a couple of companies show that you can actually have companies that grow fast in Europe. And that's what we're trying to do. I'm not too pessimistic. I find you too pessimistic. You should come to France; you should come to France. I think you would get more optimistic.
你知道吗,如果一个巴黎人告诉我我太悲观了,那我真的需要更乐观一点。我的问题是:你提到 Scaling(规模扩张)的速度是最难的事情,伙计,就是和公司以同样的速度成长。作为 CEO,随着公司如此快速的扩张,你自己成长中最难的事情是什么?
Do you know what, if a Parisian is telling me that I'm too pessimistic, then I really need to be more optimistic. My question to you is: you mentioned that the speed of scaling is the hardest thing, dude, is scaling with your company at the same speed. What was the hardest thing about yourself scaling as CEO with such speed of scaling of the company?
需要学习的东西。我的意思是,我们和 Ganot 一起在工作中学习。实际上,你会遇到组织上的挑战。你如何……
Things to learn. And I mean, we are learning on the job with Ganot. It's effectively you have organizational challenges. How do you...
你如何确保 45 个人能良好沟通?在代表公司、业务发展和交易方面,你如何管理时间?如何在竞争噪音和方向变化的不确定性中保持方向并让团队保持平静?
How do you ensure that 45 people communicate well together? How do you manage your time in terms of representation, business development, and deal making? How do you maintain direction and keep the team calm despite competitive noise and changing direction due to uncertainty?
我不认为我做得很好,但我们正在积极寻找信息来源来学习新事物。
I don't think I'm doing it properly, but we are actively trying to find sources of information to learn new things.
如果你能在成为 CEO 和创立 Mistral 的前一晚打电话给自己,你会给自己什么建议?
If you could call yourself the night before you became CEO and founded Mistral, what advice would you give yourself?
也许可以更分阶段地进行产品开发和市场推广。我们在没有任何可卖的东西时就开始了市场推广。这确实创造了品牌知名度,但先多开发产品可能会更简单。由于这是一个快速发展的领域,我们同时启动了所有事情,组织有些欠缺。现在我们正在巩固。事后看来,我可以给自己一些关于雇佣谁或何时雇佣的战术建议。但一年前的战略没有太大变化。我们意识到需要更多资本、强大的产品,并且要快速进入美国。一年前知道这些也不会有太大帮助。
Maybe stage the product development and go-to-market a bit more. We started go-to-market when we had nothing to sell. It did create brand awareness, but it might have been simpler to develop the product a bit more first. Since it's a fast-moving field, we started everything together with some lacking organization. Now we are solidifying it. In hindsight, I could give myself tactical advice on who to hire or when. But the strategy from a year ago hasn't changed much. We realized we needed more capital, a strong product, and to go to the US quickly. Knowing that a year ago wouldn't have helped much.
你现在觉得资金充足吗?
Do you feel you have enough cash now?
初创公司总是在融资。在这个领域,未来几年投资将超过收入,因为你需要扩张并保持前沿公司的相关性。收入在增长,但研究开发的速度应该快于市场推广的速度。
Startups are always fundraising. In this field, investments will exceed revenue by design for years to come because you need to scale and stay relevant as a frontier company. Revenue is ramping up, but the speed of research development should be faster than go-to-market development.
你最尊重和钦佩哪家 AI 公司?
Which AI company do you most respect and admire?
它们都取得了成果。我们最近对 Cohere 的新模型感到惊讶。OpenAI、Anthropic 和我在谷歌的朋友们也做得很好。这是一个竞争格局,我们尊重所有公司。我们都朝着相同的更高目标努力。
They all delivered. We were surprised by Cohere recently with their new models. OpenAI, Anthropic, and my friends at Google are also doing a good job. It's a competitive landscape, and we respect all of them. We all work in the same direction with higher goals.
现在创办一家 AI 公司(比如 Holistic)是不是太晚了?
Is it too late to start an AI company now, like Holistic?
我不建议进入基础模型业务。但如果说没有新竞争者出现并击败我们的机会,那就太傲慢了。
I wouldn't recommend going into the foundation model business. But I'd be arrogant to say there's no chance for a new competitor to arise and beat us.
当今世界你最担心什么?
What worries you most in the world today?
全球变暖。地球变暖并寻找解决方案的竞赛。AI 是解决方案的一部分,因为它带来更多控制和效率。但这是一场生存竞赛,所以我们应该更加警觉。
Global warming. The race of the planet heating up and finding solutions. AI is part of the solution because it brings more control and efficiency. But there's a race for survival, so we should be more aware.
过去 12 个月里,你改变最大想法的是什么?
What have you changed your mind on most in the last 12 months?
很多我从未测试过的管理前提。最大的一点:透明反馈非常有用。以几乎完全透明的方式运作帮助我们成长。
Many management premises I had never tested. The biggest one: transparent feedback is super useful. Operating in an almost fully transparent manner has helped us grow.
在扩张 Mistral 过程中,最出乎意料的挑战是什么?
What has been the most unexpectedly challenging in scaling Mistral?
我们必须管理的需求量大,超出了我们的处理能力。品牌成功——人们知道我们——有点出乎意料。我们知道会被注意到,但没想到人们会这么快开始使用我们。
The amount of demand we had to manage, which is too high for what we can handle. And the brand success—people knowing us—was a bit unexpected. We knew it would be noticed, but not that people would start using us that fast.
你做什么来让自己冷静?
What do you do to calm down?
我跑步、骑自行车。我想我的伴侣会骂我,但我试着照顾我的女儿。
I run, I cycle. I think my partner will yell at me, but I try to take care of my daughter.
你最近当了父亲。你现在知道什么你希望当初就知道?
You recently became a father. What do you know now that you wish you'd known when you first had your daughter?
我不知道照顾小孩需要这么多精力。
I had no idea that you needed so much energy to care for a small child.
你为什么认为 AI 将在未来 10 年改变世界?AI 融入一切的未来社会是什么样子?
Why do you think AI will take the world in the next 10 years? What does the future look like with AI in everything?
它正在显著改变人们的工作方式。它要求人们更有创造力,带来超越自动化的价值。这是就业市场的结构性变化,因此需要快速在培训和教育方面进行调整,以便人们了解在有 AI 的日常工作中需要做什么。
It's changing the way people work significantly. It requires being more creative and bringing more value beyond what can be automated. It's a structural change in the job market, so adaptation in training and education is needed quickly so people know what's expected in their daily jobs with AI.
你认为对工作被取代的恐惧被严重夸大了吗?
Do you think fears of job replacement are grossly overexaggerated?
我认为是的,取决于你和谁说话。工作肯定会被取代,一些被替换,一些被创造。我们正在将人类提升到更高的抽象层次。机器能够以类似人类的方式理解和回答。与我们对计算机所做的相比,这不是范式转变。但向更高抽象提升的速度可能达到了历史上前所未有的水平,这意味着社会的适应将面临挑战。
I think they are, depending on who you talk to. Jobs will be displaced for sure, some replaced, some opened up. We're moving humanity to a higher level of abstraction. Machines can understand and answer in a humanlike fashion. This is not a paradigm shift compared to what we did with computers. But the speed of elevation toward higher abstraction is probably at an unmatched rate in history, meaning society's adaptation is going to be challenging.
要更具挑战性,它需要被预见。最后一个问题:我们在 2034 年做一期节目,也就是 10 年后。如果一切顺利,Mistral 那时会在哪里?
To be more challenging, it needs to be anticipated. Final one for you: we do a show in 2034, 10 years time. If everything goes right, where's Mistral then?
Mistral 会拥有一些非常相关的模型,包括商业和开源版本,并且会有一个非常强大的开发者平台,让你能够完成创建 AI 应用所需的一切。所以那将是一个不错的成就。
Mistral has some very relevant models, commercial and open source, and it has a very strong developer platform that allows you to do everything you need to create your AI application. So that would be a good achievement.
Arthur,我非常享受这次访谈。谢谢你容忍我快速切换许多不同方向。你非常有耐心,是一位出色的嘉宾。非常感谢你,我的朋友。
Arthur, I've so enjoyed doing this. Thank you for putting up with me going in many different fast-moving directions. You've been incredibly patient and a brilliant guest. So thank you so much, my friend.
谢谢你邀请我。
Thank you for hosting me.