The Resurgence of Open Source Model Customization
打开互动全文版(中英对照 + 朗读 + 问答)→Jeffrey Morgan 探讨了由编码代理驱动的开源模型采用激增,以及微调趋势的回归。
Jeffrey Morgan discusses the surge in open source model adoption driven by coding agents and the return of fine-tuning trends.
欢迎来到 Light Con 的新一期节目。今天,我们邀请到了 Ollama 的联合创始人兼 CEO Jeffrey Morgan。Ollama 是在本地和云端运行开源 AI 模型的最简单方式,拥有 900 万开发者用户、178,000 个 GitHub Star,并且被 85% 的财富 500 强公司使用。这意味着 Jeff 对前沿 AI 技术、在基准测试中得分最高的模型,以及开发者真正下载并持续使用的模型都非常了解。Jeff,欢迎来到 LightCon。
Welcome to the new episode of Light Con. Today, we speak with Jeffrey Morgan, co-founder and CEO of Ollama, the easiest way to run open source AI models locally and in the cloud. Ollama is used by 9 million developers, has 178,000 GitHub stars, and is used by 85% of Fortune 500 companies. This means that Jeff knows a lot about cutting-edge AI technology, models that score highest on benchmarks, and models that developers actually download and continue to use. Jeff, welcome to LightCon.
谢谢邀请。
Thank you for the invitation.
当前 AI 的前沿技术是什么?你观察到了哪些趋势?
What is the current cutting-edge technology in AI, and what trends are you observing?
我认为最大的变化是向开源模型的转变,尤其是在企业环境中。这表现为美国和中国的模型混合使用,主要由编码智能体、AI 助手以及 OpenClaw 或 Hermes 等协作环境中的用例驱动。
I believe the biggest change is the transition to open source models, especially in the corporate environment. This is emerging as a mix of models developed in the U.S. and China, and is primarily driven by use cases in coding agents, AI assistants, and collaborative environments such as OpenClaw or Hermes.
由于 Jeff 亲身经历了大量 token 的流动,他对人们实际使用哪些模型以及这些变化如何发生有着非常准确的数据。你在关注哪些趋势?
Because Jeff has personally experienced the flow of numerous tokens, he has very accurate data on which models people are actually using and how those changes are occurring. What trends are you looking at?
是的,如你所知,Ollama 最初是在 MacBook、Nvidia、AMD 和 Intel 等硬件上运行开源模型的一种方式。今年早些时候,我们推出了 Ollama Cloud。目前,主要是中国模型,但世界各地的公司都在使用它们。尤其是美国和德国是使用开源模型 token 的主要国家。
Yes, as you know, Ollama started as a way to run open source models on hardware such as MacBooks, Nvidia, AMD, and Intel. And we launched Ollama Cloud earlier this year. Currently, there are mainly Chinese models, but companies around the world are utilizing them. In particular, the United States and Germany are among the major countries using open source model tokens.
公司采用开源模型的原因仅仅是为了降低成本吗?还是有其他原因?
Is the reason companies adopt open source models simply for cost reduction? Or is there another reason?
成本是开源模型能解决的最大问题。然而,每家公司都有更好地控制 AI 并将其定制以适应其业务的愿景。这正是公司的目标。成本是一个可以在短期内解决的问题,但它能让公司根据其独特的用例定制模型。
Cost is the biggest problem that open source models can solve. However, every company has a vision to better control AI and customize it to fit their business. This is exactly the goal of a company. Cost is a problem that can be solved in the short term, but it allows companies to customize models to their unique use cases.
你能告诉我是否有大公司以这种方式采用了开源模型吗?
Could you tell me if there are any large corporations that have adopted an open source model in this way?
昨天有一篇关于 AT&T 的好文章,提到 AT&T 已经将其 40% 的 token 使用量转移到了开源模型上。目前,他们主要使用美国和欧洲的模型,但据说他们也在考虑中国模型。
There was a good article about AT&T yesterday, and it said that AT&T has already switched 40% of its token usage to an open source model. Currently, they mainly use US and European models, but it is said that they are also considering Chinese models.
你们使用什么工作流程?
What workflow do you use?
我们主要使用编码智能体。从每个开发者或用户的 token 使用量爆炸式增长来看,大部分似乎源于编码智能体。在三月和四月,OpenClaw 变得流行,随后 Hermes Agent 项目取得成功,使得非开发者也能在长时间内自动化许多任务。它已被应用于金融、支持、营销和销售等多个领域。
We mainly use coding agents. Looking at the explosive growth in token usage per developer or user, it appears that most of it stems from coding agents. And in March and April, OpenClaw gained popularity, followed by the success of the Hermes Agent project, which made it possible for non-developers to automate many tasks over a long period of time. It has become available for use in various fields such as finance, support, marketing, and sales.
有一个很酷的图表显示了 OpenClaw 的指数级增长。
There was a cool graph showing the exponential growth of OpenClaw.
是的,这是今年早些时候 Ollama Cloud 上按开发者统计的 token 使用量图表。如你所知,这是显示开发者每周平均使用 token 数量的图表。最初,这是按开发者汇总的数据。
Yes, this is a graph showing token usage by developer on Ollama Cloud earlier this year. As you know, this is a graph showing the average amount of tokens developers use per week. At first, it was a figure compiled by developer.
所以,虽然这个图表可能看起来显示的是 Ollama 的整体增长率,但实际上它是每个用户的增长率。你说得对。
So, while this graph may appear to show Ollama's overall growth rate, it is actually the growth rate per user. You're right.
这是显示每个用户每周使用多少 token 的图表。有两个主要的转折点。一个是在年初,由于编码智能体而激增。例如,Kimi、GLM 模型和 Minimax 的发布。终于,能够运行编码智能体的开源模型出现了。而在四月,由于 OpenClaw 而出现了巨大增长。不仅开发者,而且世界各地的人们都可以将难题委托给开源模型并完成工作。当然,在弄清楚使用哪些工具和检索哪些数据的过程中,会消耗大量 token。得益于开源模型,上下文窗口的大小从 128,000 激增到超过一百万。所以所有这些都使得爆炸式增长成为可能。
This is a graph showing how many tokens are used per week on an individual user basis. There are two major turning points. One is that it surged at the beginning of the year thanks to coding agents. Examples include the release of the Kimi, GLM models, and Minimax. Finally, open source models capable of running coding agents have emerged. And in April, it saw tremendous growth thanks to OpenClaw. It has become possible for not only developers but also people around the world to entrust difficult problems to open source models and complete the work. Of course, a huge amount of tokens is consumed in the process of figuring out which tools to use and what data to retrieve. Thanks to the open source model, the size of the context window has surged from 128,000 to over one million. So all of these things made explosive growth possible.
在引入核心任务类型用例之前,token 数量大约是那个量的五倍,大约 1500 万。看来四月份 OpenClaw 使用量的激增就是那个规模。总体而言,可能增加了 10 到 20 倍,甚至更多。在整个 Ollama Cloud 中,自今年年初以来增长了 150 倍。这真是太惊人了。
Before the introduction of core task type use cases, the number of tokens was about five times that amount, roughly 15 million. It seems the surge in OpenClaw usage in April was on that scale. Overall, it has likely increased 10 to 20 times, or perhaps even more. Across the entire Ollama Cloud, it has increased 150-fold since the beginning of this year. It is truly amazing.
这里最有趣的一点是,对开放模型的需求激增了。对吧?直到 2024 年和 2025 年,大规模开放模型大多以定制模型的形式提供。这是一种将 Deepseek 或 Kimi 等现有模型进行微调以适应用例的方法,就像 cursor 一样。通过这种方式,我们能够大规模提供服务。然而,今年年初,立即可用的开放模型的提供开始真正出现。
The most interesting point here is that the demand for open models has surged. Right? Up until 2024 and 2025, large-scale open models were mostly provided as customized models. It was a method of taking existing models like Deepseek or Kimi and fine-tuning them to fit the use case, like a cursor. In that way, we were able to provide services on a large scale. However, the provision of immediately usable open models began to appear in earnest early this year.
自从我开始这个播客以来,这种现象似乎周期性发生。例如,在 2024 年初,人们对微调自己的定制模型非常感兴趣,但此后这种兴趣逐渐消退,一旦下一个模型发布,所有的努力似乎都会白费或被埋没。然而,现在这种趋势似乎又回来了。你是在最近距离观察整个情况的人。你认为我们目前正在经历另一个周期,还是认为这次会持续下去?
It seems like this phenomenon occurs periodically since I started this podcast. For example, in early 2024, there was a lot of interest in fine-tuning one's own custom model, but since then that interest has faded, and it seemed like all the effort would be in vain or buried once the next model was released. However, it seems that trend is returning now. You are watching this whole situation from the closest proximity. Do you think we are currently going through another cycle, or do you believe it will continue this time?
我认为这些模型的发布速度越来越快,因此要跟上并发布所有模型变得越来越困难。
I think that the release speed of these models is getting faster and faster, so it is becoming increasingly difficult to keep up with and post about all of them.
你的意思是,最新的封闭前沿模型的发布速度比以往任何时候都快?
You mean that the latest closed frontier models are being released faster than ever before?
开源领域也在以同样的方式加速。例如,DeepSeek Flash 模型仅今年夏天就已经更新了三次。过去它每六个月更新一次。
The open source sector is also getting faster in the same way. For example, the DeepSeek Flash model has already been updated three times this summer alone. It used to be updated every six months.
是的,所以差距似乎在逐渐缩小。
Yes, so it seems that the gap is gradually narrowing.
我认为这就是为什么训练定制模型变得更加困难。另一方面,工具越来越好,使得希望微调模型的团队能够跟上最新技术。
And I think that is why training customized models becomes more difficult. On the other hand, tooling is getting better and better, enabling teams looking to fine-tune models to keep up with the latest technology.
此外,我们正在进入一个阶段,AI 安全问题正成为前沿研究实验室中日益重要的问题。因此,最先进的实验室可能会放慢开发速度,以解决对齐和隔离问题。与此同时,开源模型和开放权重模型持续增长和改进。
Furthermore, we are entering a point where AI safety issues are emerging as an increasingly important issue in cutting-edge research laboratories. Therefore, state-of-the-art laboratories may slow down development speeds to solve alignment and isolation problems. Meanwhile, open source models and open weight models continue to grow and improve.
是的,我从网络安全的角度审视了 GLM-53 模型的发布和能力,确实令人印象深刻。这为安全和治理领域的新兴公司和成熟企业都提供了重要机遇。
Yes, and I looked at the announcement and capabilities of the GLM-53 model from a cybersecurity perspective, and it was really impressive. This offers significant opportunities for both startups and established companies in the security and governance sectors.
正如我提到的 AT&T 文章所示,采用开放模型的主要障碍是安全和安保问题。
As can be seen in the AT&T article I mentioned, the main obstacle to adopting the open model is security and safety issues.
嗯,然而,根据我与欧美客户交流的经验,如果安全问题能够解决,对采用中国制造的模型实验室的整体反应是积极的,这对企业来说非常鼓舞人心。
Well, however, based on my experience speaking with customers in Europe and the U.S., the overall reaction to adopting Chinese-made model laboratories is positive if safety issues can be resolved, which is very encouraging for companies.
这里有一个有趣的点:Hugging Face 实际上不得不使用开放模型来检测 Frontier 上的黑客攻击。
There is an interesting point here: Hugging Face actually had to use an open model to detect hacking on Frontier.
我们经常被问到的一个问题是:“开放模型在哪些方面比前沿模型的封闭模型更有优势,可以在哪里使用?”
One of the questions we are frequently asked is, "Where can open models be used, which are much more advantageous than Frontier's closed models?"
不见得。
No see.
一个关键用例是安全测试,即验证软件的安全性。
One of the key use cases is security testing, that is, verifying the security of software.
是的,例如,如果你要求 Claude 对产品进行渗透测试,Claude 会拒绝。
Yes, for example, if you ask Claude to perform a penetration test on the product, Claude refuses.
你说得对。一般来说,情况确实如此。
You're right. Generally speaking, that is the case.
另一方面,Hugging Face 拥有全新的安全研究模型,能够进行渗透测试。
On the other hand, Hugging Face has completely new security research models that enable penetration testing.
你说得对。如你所知,还有定制训练的模型,旨在更灵活地操作以解决这些问题。
You're right. As you know, there are also custom-trained models designed to operate more flexibly to solve these problems.
基础模型也提供安全训练,但它们往往在区分善意和恶意用例时更加谨慎。
Basic models also provide safety training, but they tend to be a little more careful in distinguishing between good-faith and malicious use cases.
Olama 的伟大之处在于,作为这些模型的核心分发渠道,模型开发者会在正式发布前首先联系我们,协调发布日程等。因此,你可以收到即将推出的模型的预览,并在新模型发布时经历巨大的增长。
The great thing about Olama is that, as it is the core distribution channel for these models, model developers contact us first before the official launch to coordinate launch schedules and such. Thanks to this, you can receive previews of upcoming models and experience tremendous growth whenever a new model is released.
你能更详细地告诉我,大规模运行这样一个系统是什么感觉吗?
Could you please tell me in more detail what it is like to operate such a system on a large scale?
是的,当然。正如我之前提到的,模型发布的速度越来越快。因此,我们为成功的初始模型发布制定了一个剧本。有很多事情需要妥善处理。首先,你需要检查模型是否受支持,以及是否有首选的推理引擎。这个过程通常需要几周时间,以确保速度、准确性和与参考的兼容性。
Yes, of course. And as I mentioned earlier, the speed of model releases is getting faster and faster. So, we developed a playbook for a successful initial model launch. There are many things that need to be done properly. First, you need to check if the model is supported and if there is a preferred inference engine. This process typically takes a few weeks to ensure speed, accuracy, and compatibility with references.
此外,找到合适的用例和工具以便开发者和用户能够利用这个模型也很重要。通常,这些模型提供新功能。今天早上,Deep Seek 宣布推出其首个多模态模型。从大规模模型 LLM(长期模型)的角度来看,虽然我们之前提供了小规模 OCR 模型,但这个多模态模型的发布实现了多种用例。
In addition, it is important to find appropriate use cases and tools so that developers and users can utilize this model. Generally, these models provide new features. This morning, Deep Seek announced the launch of its first multimodal model. From the perspective of a large-scale model LLM (Long-Term Model), while we previously provided small-scale OCR models, the release of this multimodal model has enabled a variety of use cases.
Deep Seek 最近也开发了自己的工具来支持这一功能。因此,第二步是找到合适的工具并准备有效执行这个模型。然而,每个模型都是不同的。每个模型有不同的任务、架构变化和工具调用机制。而正确实现所有这些确实非常困难。
Deep Seek has also recently developed its own tools to support this feature. Therefore, the second step is to find a suitable tool and prepare to effectively execute this model. However, every model is different. Each has different tasks, architectural changes, and tool calling mechanisms. And implementing all of this properly is really difficult.
最终,最重要的是在发布前对最终产品运行基准测试,以验证它是否按照研究团队的规格工作。因此,你在整个生态系统中必须扮演的一个非常重要的角色是建立明确的标准,以确保每个新模型和每个新工具调用的最佳使用,并保持所有模型的一致性。这是一项非常困难的任务。
Ultimately, the most important thing is to run benchmarks on the final product before launch to verify that it works as specified by the research team. Therefore, one of the very important roles you must play in the entire ecosystem is to establish clear standards to ensure optimal usage for each new model and each new tool call, and to maintain consistency across all models. This is a very difficult task.
是的,我们试图将这三者打包在一起。首先是框架。例如,无论是现成的框架、支持模型使用的 SDK,还是为这个模型设计的现有框架。幸运的是,现在有许多优秀的开源框架可用。Codex 框架也是开源的。Open Code 是 Olami 用户中最受欢迎的工具之一,也是一个出色的工具。
Yes, we try to package the three together. The first is the harness. For example, whether it is a ready-made harness, an SDK that supports the use of the model, or an existing harness designed for this model. And fortunately, there are many excellent open source harnesses available now. The Codex harness is also open source. Open Code is one of the most popular tools among Olami users and is an excellent tool.
第一步是正确构建 Open Code,但下一步是将其与模型打包以确保可用性和可靠性。在云环境中,你必须确保足够的容量,因为初始增长阶段自然是增长最显著的时期。重要的是硬件和供应商。这正是需要大量合作的地方。例如,你必须通过让推理提供商优化模型、与 Nvidia 等合作伙伴合作或利用 Apple Silicon 堆栈来确保模型运行快速。这是因为无论模型性能多么出色,如果速度慢,就无法提供良好的用户体验。你必须考虑所有这三个因素来打包,但通常,如果你能提前几周访问模型,你就很幸运了。在模型发布前的 24 小时内,大量工作密集进行。这真是一片混乱。
The first step is to properly build Open Code, but the next step is to package it with the model to ensure availability and reliability. In a cloud environment, you must secure sufficient capacity, as the initial growth phase is naturally the period when the most significant growth occurs. And the important thing is the hardware and the supplier. This is precisely where a lot of cooperation is needed. For example, you must ensure that the model runs fast by having inference providers optimize the model, collaborating with partners such as Nvidia, or utilizing the Apple Silicon stack. This is because no matter how excellent the model's performance is, it cannot provide a good user experience if it is slow. You must package it by considering all three of these factors, but generally, you are lucky if you can access the model a few weeks in advance. A lot of work is carried out intensively during the 24 hours immediately preceding the model launch. It is truly a chaotic situation.
这就像旧的操作系统。就像所有驱动程序、硬件甚至应用程序层都必须紧密集成的时代。你就像连接这一切的粘合剂。
It is just like an old operating system. Just like the days when all drivers, hardware, and even the application level had to be tightly integrated. You are like the glue that connects all of this.
尽管操作系统类比是一个常见的陈词滥调,但我认为这是一个恰当的类比,因为虽然有硬件驱动程序和推理层的提供商,但从应用程序运行时到框架的操作,一切都必须连接起来。将所有这些整合在一起是一个非常复杂的问题。所以真正做好它非常困难。然而,随着时间的推移,组件被开发出来,使这个过程能够快速执行、测试和发布。它提供了一个通用运行时,可以连接任何框架或任何模型。这正是我们为开发者扮演的角色。
Although the operating system analogy is a common cliché, I think it is an appropriate analogy because, while there are hardware drivers and providers of the inference layer, everything from the application runtime to the operation of the harness must be connected. Putting all of this together is a very complex problem. So it is really difficult to do it properly. However, over time, components are developed that enable this process to be carried out, tested, and released quickly. It provides a common runtime that can connect any harness or any model. That is exactly the role we play for developers.
你们确实是杰出的工程师,从 Apple Silicon 的早期到现在的 DGX,你们对一切都进行了微调。你们是如何建立起如此坚实的技术团队的?
You are truly outstanding engineers, and you have fine-tuned everything from the early days of Apple Silicon to the current DGX. How were you able to build such a solid technical workforce?
我们的大多数团队成员不是 AI 研究人员。他来自 VMware、Docker 和其他网络公司。因此,有一种趋势是现有的计算问题在推理领域被重新发明,我们主要关注那个领域。然而,我相信最终一切都归结为开发者体验。我的意思是,当开发者调用 API 并收到一个 token 时会发生什么,以及那个过程是什么样的。如今,在那个领域发生的事情越来越多。而要正确实现这一点,堆栈的所有层必须无缝协作。
Most of our team members are not AI researchers. He comes from VMware, Docker, and other networking companies. Therefore, there is a trend where existing computing problems are being reinvented in the inference domain, and we are mainly focusing on that area. However, I believe that ultimately, everything comes down to the developer experience. I mean what happens when a developer calls an API and receives a token, and what that process is like. These days, more and more things are happening in that area. And to properly implement this, all layers of the stack must cooperate seamlessly.
我认为开放模型面临的困难之一是,它们的发展速度不如前沿模型研究实验室发布的模型。正如 Jensen 提到的,有一个类似应用的 5 层结构,意味着从模型、基础设施、推理、芯片,甚至能源,一切都准备好了,让开发者从第一天就能使用。我们也打算在开放模型中重现这种环境。当然,虽然可能无法实现每一层,但支持编排是可能的。这关乎构建一个层级结构。是的,开放模型可以有超过 5 层的层级。例如,模型层级中有许多元素。不仅有模型权重,还有各种编排组件,每个模型在其中独特地展现优势。还有开发者 API 层。如今的 5 层结构(也称为“5 层蛋糕”)隐藏了在模型和应用层之间构建的众多机会。
I think one of the difficulties with open models is that they are not developing at the same rapid pace as those released by cutting-edge model research labs. As Jensen mentioned, there is a 5-stage structure like an app, meaning everything from models, infrastructure, inference, chips, and even energy is prepared so that developers can use it right from day one. We intend to recreate that very environment in the open model as well. Of course, while it may not be possible to implement every layer of the stack, it is possible to support orchestration. It is about building a hierarchical structure. Yes, the open model can have more than 5 levels of hierarchy. For example, there are many elements in the model hierarchy. There are not only model weights but also various orchestration components in which each model uniquely demonstrates strengths. And there is also a developer API layer. Today's 5-layer structure (also known as the '5-layer cake') hides numerous opportunities to build between the model and application layers.
这是个有趣的观点。你认为那些隐藏的层有可能分离出来,成为独立的公司或供应商吗?
That's an interesting point. Do you think there is a possibility that those hidden layers could separate and become independent companies or suppliers?
当然。Anthropic 平台团队在他们的演示中涵盖了三个重要点。第一是知识管理,也就是将公司数据和上下文连接到模型的方法。第二是调整问题。如你所知,当从 Claude 应用或 Codex 发送请求时,云中和本地的多个子智能体被执行。这里出现了调整问题。最后是执行问题。虽然经常以沙箱的概念来讨论,但这些智能体正越来越多地转向云端。这部分需要解决的计算问题非常大。如果我们看看现有的云业务和当前的云业务,得益于开源,每个问题的最佳公司已经出现。如果你想想早期的云产品,它们的形式是所有功能捆绑在一起,比如 Heroku 或 Google App Engine。然而,开发者最终更偏好的是每个领域专业化的最佳产品。一切捆绑或分离的现象就像一种海上贸易。你可以以前沿模型为例。前沿模型实验室希望完全限制在自己的围墙内,包括托管智能体、上下文和记忆层。另一方面,小型科技公司或开源开发者不想被限制在这样的模式中。所以我们将以不同的方式开发产品,看看哪种方法最终获胜会很有趣。
Of course. The Anthropic platform team covered three important points in their presentation. The first is knowledge management, that is, the method of connecting company data and context to the model. The second is the adjustment issue. As you know, when a request is sent from the Claude app or Codex, multiple sub-agents in the cloud and locally are executed. An adjustment problem occurs here. Finally, there is an execution problem. Although it is often discussed in terms of the concept of a sandbox, these agents are increasingly moving to the cloud. The computing problem that needs to be solved in this part is very large. If we look at the existing cloud business and the current cloud business, thanks to open source, the best companies for each problem have emerged. If you think of early cloud products, they were in the form where all functions were bundled together, like Heroku or Google App Engine. However, what developers eventually came to prefer were the best products specialized in each field. The phenomenon of everything being bundled or separated is like a kind of maritime trade. You can take the Frontier Model as an example. Frontier Model Lab wants to stay completely confined within its own walls, including managed agents, context, and memory layers. On the other hand, small tech companies or open source developers do not want to be confined to such a mold. So we will develop the product in a different way, and it will be interesting to see which method ultimately wins.
我认为双方结合可能会获胜。
I think the combination of both sides will probably win.
目前,许多与开源模型相关的公司正在涌现。只有几十家开放模型提供商能够提供这样的 token。新的稀缺性——也就是当前的问题——不仅仅在于 token。例如,是如何将智能体从 A 编排到 B 的问题。这引发了许多新的系统和工程问题,对个人开发者来说太难解决了。开发者不可能构建所有这些层。有趣的是,编码智能体本身正在显著进化,表面上,这是模型面临最糟糕境地的时期。过去存在这些市场进入壁垒的原因是,创建经过适当测试和维护、真正满足用户需求的软件太难了。但如果这些壁垒消失了会怎样?也许我们生活在一个有《古兰经》、Markdown 文件作为 TypeScript 运行、供应商锁定不再存在的时代。例如,你可能在某个时候使用 OpenAI 的记忆系统,事实上,即使智能体不断向自己的记忆系统发送数据,它也能无问题运行。它与 MCP(记忆控制框架)配合良好,并且没有供应商依赖,因为它经过了充分的运行时测试。我们今天知道的许多 harness 将被集成到模型中。然而,随着模型的核心循环和连接功能被集成到模型本身,像记忆这样分离出来的部分就出现了。我们通常使用的指导原则如下。关于状态保存的问题,例如存储相关问题,最终不能直接反映在模型中。模型经过训练,如你所知,训练每月进行一次,但它不会用最新数据更新。因此,数据存储是一个非常大的问题领域,甚至可能影响模型层。例如,像凭证这样的安全相关信息永远不会下放到模型层。开源模型没有封闭模型提供商默认提供的所有安全工具,但这是必须采用的一个非常重要的因素,尤其是企业。
And currently, numerous companies related to open source models are emerging. There are only a few dozen open model providers capable of offering such tokens. The new scarcity—that is, the current problem—lies in more than the tokens. For example, it is the problem of how to orchestrate agents from A to B. This gives rise to numerous new systems and engineering problems, which are too difficult for individual developers to solve. It is impossible for a developer to build all these layers. The interesting point is that coding agents themselves are evolving significantly, and superficially, this is the time when the model faced its worst situation. The reason these market entry barriers existed in the past was that it was too difficult to create software that is properly tested and maintained, and actually meets user needs. But what would happen if these barriers disappeared? Perhaps we are living in an era where there is the Quran, Markdown files run as TypeScript, and vendor lock-in is no longer present. For example, you might use OpenAI's memory system at some point, and in fact, the agent can operate without problems even if it continuously sends data to its own memory system. It works well with MCP (Memory Control Framework) and has no vendor dependencies because it has undergone sufficient runtime testing. Many of the harnesses we know today will be integrated into the model. However, as the core loops and connectivity functions of the model are integrated into the model itself, parts that separate out like memory emerge. The guidelines we generally use are as follows. Issues regarding state saving, for example, storage-related problems, cannot ultimately be directly reflected in the model. The model is trained, and as you know, training is conducted every month, but it is not updated with the latest data. Therefore, data storage is a very large problem area that can affect even the model layer. For example, security-related information such as credentials will never go down to the model layer. Open source models do not have all the security tools that closed model providers offer by default, but this is a very important factor that must be adopted, especially by enterprises.
我很好奇,在企业环境中,关于封闭和开源模型之间的平衡,你预期的最终或未来状态是什么,特别是在成本方面。
I am curious to know what final or future state you anticipate regarding the balance between closed and open source models in an enterprise environment, particularly in terms of cost.
我认为起初,每个人似乎都把全部预算投入到 Anthropic 或 OpenAI。开源目前显然在增长,但似乎对 Anthropic 的支出也在随之增加。所以看起来一切都在共同增长。这种趋势会持续下去,还是会稳定在一种状态,即一半预算投资于封闭源前沿模型,另一半投资于开源?还是另一个方面?在我们看来,公司内部大部分 token——大约 80% 到 90%——将投资于开放模型。当然,这并不意味着 80% 到 90% 的预算投资于开放模型。事实上,我认为开放模型社区真正做得好的是降低成本,让更多人能够使用。所以,你可能最终只支付开放模型成本的 10% 到 20%,但你所有的 token 都将通过开放模型使用。对吧?随着 token 以这种方式变得更加丰富,更多的用例成为可能。你不是限制团队成员的 token 访问,而是实际上授予他们更多访问权限。我认为最困难的任务最好在最先进的研究机构处理,那里聚集了最优秀的研究人员。而中间阶段的问题可以通过结合使用开放和封闭模型来解决。在我看来,大多数软件和模型可能会以开放模型的形式提供。有趣的是,随着模型变得更加强大,理想的做法是将权力委托给最小的模型来决定何时切换到开放模型,但拥有该模型的研究机构认为用户不会想这样做。
I think that at first, everyone seemed to be pouring their entire budget into Anthropic or OpenAI. Open source is clearly growing right now, but it seems that spending on Anthropics is increasing along with it. So it looks like everything is growing together. Will this trend continue, or will it stabilize into a state where half the budget is invested in the closed-source frontier model and the other half in open source? Or is it a different aspect? In our opinion, most of the tokens—about 80 to 90%—will be invested in open models within the company. Of course, this does not mean that 80 to 90 percent of the budget is invested in the open model. In fact, I think what the open model community is really doing well is lowering costs to make it accessible to more people. So, you might end up paying only 10 to 20% of the open model costs, but all your tokens will be used through the open model. Right? As tokens become more abundant in this way, many more use cases become possible. Instead of restricting team members' token access, you are actually granting them more access. I think it is best to handle the most difficult tasks in a state-of-the-art research institute where the best researchers are gathered. And the problems at that intermediate stage could be solved by using open and closed models together. In my opinion, most software and models will likely be provided in the form of open models. The interesting point is that as the model becomes more powerful, it is ideal to delegate authority to the smallest model to decide when to switch to an open model, but the research institute that owns the model thinks that users would not want to do that.
在我看来,所有研究机构在很多方面都聚焦于一个目标:如何为客户提供服务。客户将决定是否通过前沿模型在路由器上处理诸如调度或复杂编排等任务。然而,大多数细节任务可以通过开放模型来处理。我见过许多项目将这两者结合起来。无论是 Sakana AI 还是 OpenRouter,都有很多通过结合这两者取得非常好效果的案例。这与人类组织并无太大不同。例如,就像一家律师事务所拥有一个合伙人律师和几个助理律师,合伙人律师将工作委派给助理律师。在云计算中也可以观察到类似的现象。人们在使用内部开发的软件和云提供商提供的软件的混合体。例如,AWS 有自己的专有数据库 DynamoDB,但许多客户将其与 PostgreSQL 数据库一起使用。最终,客户似乎会使用这两种数据库的组合。我认为这是一个非常普遍的模式。
In my opinion, all research institutes are focusing on one goal in many respects: how to provide services to customers. The customer will decide whether to handle tasks such as scheduling or complex orchestration on the router through the frontier model. However, most detailed tasks can be handled through open models. I have seen many projects that combine these two things. Whether it is Sakana AI or OpenRouter, there are many cases where very good results have been achieved by combining these two. This is not much different from human organizations. For example, it is like a law firm having a partner attorney and several associate attorneys, where the partner attorney delegates work to the associate attorneys. A similar phenomenon could also be observed in cloud computing. A mix of in-house developed software and software provided by cloud providers is being used. For example, AWS had its own proprietary database, DynamoDB, but many customers used it alongside PostgreSQL DB. Ultimately, it appears that customers will use a combination of the two databases. I think this is a very common pattern.
Apple Silicon 之所以出色,正是因为这种设计方法。它配备了针对各种工作负载(如图像处理和音频处理)的专用加速器。这从 PC 时代就延续下来,当时存在独立声卡、显卡等。既然我们谈到了 Apple Silicon,我们是否也谈谈本地托管模型?既然你在基于云的模型和运行在笔记本电脑上的本地托管模型两方面都开展大规模业务,我很好奇你在这两种环境中看到了哪些变化,以及你认为未来会发生什么。
The reason Apple Silicon is outstanding is precisely because of this design method. It is equipped with accelerators specialized for various workloads, such as image processing and audio processing. This has continued since the PC era, during which standalone audio cards, graphics cards, and the like existed. Since we're on the topic of Apple Silicon, shall we talk about the local hosting model as well? Since you are running a large-scale business in both a cloud-based model and a locally hosted model running on a laptop, I am curious to know what changes you are seeing in these two environments and what you think will happen in the future.
我认为这是一个非常有趣的话题,因为它显示出类似于封闭和开放模型之间冲突的模式。我们同意这个观点,我们的客户也建议,对于简单任务,本地运行可以降低延迟并降低每个词元的成本。最终,你最终会先购买硬件,然后与云模型结合使用。我们多年来拥有的下一代硬件的最大优势是它处理 200 亿到 400 亿甚至 1200 亿参数模型的能力。Quinn 3.8 38B 在编码性能方面与 Opus 4.6 一样好。
I think this is a very interesting topic as it shows a pattern similar to the conflict between closed and open models. We share this opinion, and our customers have also suggested that for simple tasks, running them locally reduces latency and lowers the cost per token. Ultimately, you end up purchasing the hardware first and using it in conjunction with the cloud model. The biggest advantage of the next-generation hardware we have had for several years is how well it handles models with 20 billion to 40 billion, or even 120 billion parameters. Quinn 3.8 38B is as good as Opus 4.6 in terms of coding performance.
真的吗?
Is that right?
基准测试结果就是这样。这真是太神奇了。它不仅能运行在内存最小的 MacBook 上,还能运行在商店里内存第二小的 MacBook 上,这真的很了不起。
That is how the benchmark results turned out. It is truly amazing. It is really amazing that it can run not only on the MacBook with the least memory, but also on the MacBook with the second least memory in the store.
用户实际上是如何利用这一点的?我很好奇你使用本地模型和云托管模型的目的是什么。
How are users actually utilizing this? I am curious about the purposes for which you use the local model and the cloud hosting model.
如果我们看看使用了哪些模型,在美国和中国训练的模型在本地环境中得到了均衡使用。当然,不仅原始的 Lama 模型,像 DeepMind 的 Gemma 模型这样的优秀模型也适合本地环境。然而,从用例来看,编码智能体在大型云模型中大多最有效。这是因为编码智能体必须解决非常困难的问题,比如编写代码和执行测试。另一方面,整体问题解决过程相对简单的用例,如文档处理工作流,在本地环境中运行得非常高效。因此,在这种情况下,采用混合执行模式,相对容易或简单的任务在本地运行,路由器确定“此任务需要大型云模型”。我认为这对客户意味着,远离开源模型可以显著提高成本节约。这不仅是因为在云中运行的成本低,而且现在还可以在用于商业目的的硬件上几乎免费运行。最终,关于云编码智能体模型,我们可以看到目前主要消耗的是中国模型。另一方面,本地模型是美国、欧洲和中国模型的显著混合。
If we look at which models are used, models trained in the US and China are being used in a balanced manner in the local environment. Of course, not only the original Lama model but also excellent models like DeepMind's Gemma model are suitable for local environments. However, when looking at use cases, coding agents are mostly most effective in large-scale cloud models. This is because coding agents have to solve very difficult problems, such as writing code and performing tests. On the other hand, use cases where the overall problem-solving process is relatively simple, such as document processing workflows, run very efficiently in a local environment. Therefore, in such cases, a hybrid execution model is used where relatively easy or simple tasks are run locally, and the router determines that "a large-scale cloud model is required for this task." And I think what this means for the customer is that moving away from the open source model leads to significantly higher cost savings. This is because not only is the cost of running in the cloud low, but it can now also be run virtually for free on hardware purchased for business use. Ultimately, regarding Claude Coding agent models, we can see that Chinese models are currently being consumed primarily. On the other hand, the local model is a significant mix of American, European, and Chinese models.
是的,如果你比较这两张图,真是令人惊叹。对于本地模型,美国和中国模型几乎持平。然而,在云托管模型的情况下,美国显示出 100% 的中国模型,就好像 x 轴的颜色被改变了一样。
Yes, if you compare these two graphs, it is truly amazing. For local models, the US and Chinese models are almost tied. However, in the case of the cloud hosting model, the US is showing a 100% Chinese model, as if the color of the x-axis had been changed.
基本上,我们需要更多能够开发大规模模型的美国研究机构。这正是这张图所显示的。而且如你所知,这一趋势的第一波始于 NeMo Triton Ultra 模型的发布,我真的很期待。
Basically, we need more U.S. research institutes capable of developing large-scale models. That is exactly what this graph shows. And as you know, the first wave of this trend began with the launch of the NeMo Triton Ultra model, and I am really looking forward to it.
NVIDIA 公司真的很有趣,因为他们的战略远非开展新的软件业务或销售代币。他们似乎对发布大量开源软件和支持生态系统非常感兴趣。多亏了这一战略,我们似乎也能在硬件领域保持领先。
The company NVIDIA is really interesting because their strategy is far from starting a new software business or selling tokens. They seem to be very interested in releasing a lot of open source software and supporting the ecosystem. And thanks to this strategy, it seems we can also stay ahead in the hardware sector.
在我看来,确实如此。最终,视频领域最令人惊讶的方面是它正在围绕开源模型构建生态系统,无论是硬件还是模型。例如,我听说目前正在开发的 DGX 工作站计算机计划配备 GB300。
In my opinion, that is the case. Ultimately, the most surprising aspect of the video field is that it is building an ecosystem centered around open source models, whether for hardware or models. For example, I heard that the DGX station computer currently under development is scheduled to be equipped with the GB300.
那声音不是很大,对吧?
That sound isn't that loud, is it?
是的,非常吵。
Yes, it is very noisy.
你知道它多少钱吗?
Do you know how much it costs?
我还不太确定。看起来 Gary 可能会住在那里。对吧?我查了一下。所以,非常慢地运行前沿模型大约要花费 20 万到 30 万美元,对吧?
I'm not sure yet. It looks like Gary will probably live there. Right? I looked it up. So, running the Frontier model very slowly costs about 200,000 to 300,000 dollars, right?
是这样吗?
Is that right?
我认为它比那更有竞争力。是的。而且它可以高速运行,性能优于前沿模型,但价格范围与标准工作站计算机没有显著差异。
I think it is much more competitive than that. Yes. And it can run at high speeds with performance superior to the Frontier model, yet the price range is not significantly different from that of a standard workstation computer.
哦,天哪,那不可能是对的。
Oh my, that can't be right.
如果你看看我们的许多客户,比如银行或工业企业,他们已经在每位工程师的办公桌上配备了 NVIDIA 工作站 GPU。有些地方甚至有数万台。所以这就是为什么反应是“啊,我也想要一个。”我们长期以来一直在销售用于 CAD 工作的这类产品,这是早期的 RTX A6000 型号。这是下一代型号。
If you look at a significant number of our customers, such as banks or industrial firms, they already have NVIDIA workstation GPUs on every engineer's desk. Some places even have tens of thousands of units. So that's why the reaction is, "Ah, I want one too." We have been selling products like this for CAD work for a long time, and this is the early RTX A6000 model. This is the next-generation model.
是的,是的,是的。我想我应该给 Jensen 发封邮件。
Yes, yes, yes. I guess I should send an email to Jensen.
这是我在 GTC 上听到的。你在这里看到的是在 GTC 上发布的产品,我们是首批收到 DGX Spark 的人之一,还有 Elon Musk 和其他几个人。这款产品可以放在桌子上,提供 128GB 的集成内存,为 20B 到 120B 的模型提供动力。然而,这只是能够驱动更大模型的全新硬件产品线的开始。
This is something I heard at GTC. What you see here is the product unveiled at GTC, and we were among the first people to receive the DGX Spark, along with Elon Musk and a few others. This product can be placed on a desk and provides 128GB of integrated memory to power models from 20B to 120B. However, this is just the beginning of a completely new hardware lineup capable of driving larger models.
你说过,如果我购买多台这款产品并将它们连接起来,我也可以驱动 400B 模型,对吧?
You said that if I buy multiple units of this product and connect them, I can drive the 400B model as well, right?
是的,当然。
Yes, of course.
这些产品支持超高速网络连接,因此可以像小型数据中心机架一样堆叠在桌面上。Mac Mini 不也是这么用的吗?
These products support ultra-high-speed network connectivity, so they can be stacked on a desk like small data center racks. Isn't the Mac Mini used that way too?
是的,没错。不过,这可以称之为量产版本。苹果和英伟达在下一代硬件上取得了显著的飞跃,这真的很有趣。这些硬件是为这些模型设计的,并且构建得高效运行。所以,可以说这个平台是你现在应该购买的产品。
Yes, that is correct. However, this can be called a mass-produced version. It is really interesting that Apple and Nvidia have made a remarkable leap forward with next-generation hardware. This hardware was designed for these models and is built to operate efficiently. So, this platform can be said to be the product you should buy right now.
学生可以用 Apple Studio 来运行,但如果你只是想让它正常工作,最好使用 DGX Spark。根据我们的测试结果,两者在性能方面都非常有竞争力。好的。所以看来很多人最终是根据有什么产品可用以及他们想要什么技术栈来做决定。
You can run it with Apple Studio for students, but if you just want it to work properly, it is better to use DGX Spark. Based on our test results, both are very competitive in terms of performance. All right. So it seems that many people ultimately decide based on what product is available and what technology stack they want.
我认为它的技术栈非常成熟,因为通过苹果的 MLX 项目,它取得了显著进展,使得大语言模型不仅能在 Mac Studio 上运行,也能在更小的 Mac 上运行。当然,DGX Spark 的技术栈也非常棒。我们很高兴能与英伟达合作。我认为个人桌面将迎来一场精彩的复兴。
I believe it has a very mature technology stack, as it has made remarkable progress through Apple's MLX project to enable LLM to run not only on Mac Studio but also on smaller Macs. And of course, the DGX Spark stack is also really great. We are very pleased to partner with NVIDIA. I think a wonderful renaissance is coming to personal desktops.
关于 Ollama 的发展历程,有趣的是我们从本地开始。对编码智能体的需求显然在云端,但最终会回到本地环境。因为如果你的桌面上有像 GB300 这样的 GPU,你就能以测试执行或代码编辑器更改的速度来实现编码循环。你们可能还记得使用 GitHub Copilot 时自动补全功能在 100 毫秒内出现的那种体验。这种体验将回到桌面环境,我认为这是一个从本地开始、走向云端、回归本地,最终走向两者结合使用的旅程。
And what's interesting about Ollama's journey is that we started in the local area. The demand for coding agents is clearly on the cloud side, but eventually, it will return to the local environment. Because if you have a GPU like the GB300 on your desktop, you would be able to implement coding loops as fast as the speed of test execution or code editor changes. You all probably remember the experience of the autocomplete feature appearing in 100 milliseconds when using GitHub Copilot. That experience will return to the desktop environment, and I think this is a journey that starts locally, goes to the cloud, returns locally, and eventually moves toward using both together.
运行 Ollama Cloud 需要大量的 GPU;目前 GPU 市场的状况如何?
Running Ollama Cloud requires a massive amount of GPUs; what is the current state of the GPU market?
我认为价格波动非常迅速,供需波动性很高。所以,如你所知,从初创公司的角度来看,最终很难获得运行最新模型所需的 B200 和 B300 GPU。幸运的是,有优秀的推理提供商建立在此基础上。所以我们正在见证巨大的需求。
I believe that price fluctuations are very rapid and the volatility of supply and demand is high. So, as you know, from a startup's perspective, it is ultimately very difficult to secure the B200 and B300 GPUs needed to run the latest models. Fortunately, there are excellent reasoning providers built on this. So we are witnessing tremendous demand.
我能否获得所有必要的 GPU,还是增长会受到可用 GPU 数量的限制?目前的情况如何?
Will I be able to secure all the necessary GPUs, or will growth be limited by the quantity of GPUs available? What is the current situation?
幸运的是,我们通过与各种提供商合作共同利用 GPU 来满足需求,但这是一项需要大量精力和时间的任务。你必须仔细考虑运行哪个模型、在哪里运行、所需的速度是多少、在哪个区域运行,以及延迟会对客户产生多大影响。在这个技术栈中有许多难题需要解决。
Fortunately, we are able to meet demand by collaborating with various providers to jointly utilize GPUs, but this is a task that requires significant effort and time. You must carefully consider which model to run and where, what the required speed is, which region to run it in, and how much latency will affect the customer. There are many difficult problems to solve in this stack.
而像 Open、Router 和 Llama 这样的开放代码项目最有趣的地方在于,最终用户开发者只需注册即可立即使用,无需为 B200 或 B300 等设备讨价还价,也无需以超过 24 个月的预估成本购买 GPU。
And the most interesting thing about open code projects like Open, Router, and Llama is that end-user developers can use them immediately just by signing up, without the hassle of negotiating prices for equipment like the B200 or B300, or purchasing GPUs with an estimated cost over 24 months.
YC 下一批的招募已经开始,你梦想创业吗?请在 ycombinator.com/apply 申请。越早开始越好,仅仅写申请就能让你的想法更上一层楼。现在,回到视频,假设你是一位刚刚创立 AI 公司且融资不多的初创创始人。那么,如果你想以低成本利用尽可能多的 token,如何在有限的预算下取得最大成果?我想听听你的建议。
Recruitment for YC's next cohort has begun, so are you dreaming of starting a startup? Apply at ycombinator.com/apply. The earlier you start, the better, and simply writing the application can take your idea to the next level. Now, going back to the video, let's assume you are a startup founder who has just established an AI company and hasn't raised much funding. So, if you want to utilize as many tokens as possible at a low cost, how can you achieve maximum results with a limited budget? I would appreciate your advice.
像 DeepSeekFlash 这样的新型模型就是一个很好的例子,我认为未来会出现相当多的这类模型。这些模型是一个非常重要的指标,因为不仅每个 token 的成本非常低,而且每个任务的成本也很低。在我看来,这类模型将最先实现无限 token 的概念。你还记得 ChatGPT 吗?你可以每天无限使用,而不必担心 token 用量。我认为我们最终会回到那个时代,但这需要付出很多努力来定制训练模型架构,以处理大规模 token 使用。
New types of models like DeepSeekFlash are a good example, and I think quite a few of these models will come out in the future. These models are a really important metric because not only is the cost per token very low, but the cost per task is also low. In my opinion, these types of models will be the first to realize the concept of unlimited tokens. Do you remember Chat GPT? You could use it unlimitedly every day without having to worry about token usage. I think we will ultimately return to that era, but it will require a lot of effort to custom train model architectures to handle large-scale token usage.
今年早些时候,我们热切希望开源模型能达到人工智能的前沿水平。我们缩小了这一差距。现在,最先进的闭源模型与开源模型之间的差距不到三个月。然而,下一个要解决的挑战是极致效率。例如,发现 GPT Luna 模型对客户来说价格非常有竞争力,这是一个巨大的优势。我听到很多客户说,多亏了这一定价策略,他们才能在团队中广泛采用。我认为这种现象也会出现在开放模型中。
Earlier this year, we eagerly hoped that open source models would reach the cutting edge of artificial intelligence. And we narrowed that gap. Now, the gap between state-of-the-art closed models and open source models is less than three months. However, the next challenge to be addressed is extreme efficiency. For example, it was a huge advantage to find out that the GPT Luna model was very price-competitive for customers. I have heard many stories from customers saying that thanks to this pricing policy, they were able to adopt it widely within their teams. I think this phenomenon will also appear in open models.
这种趋势已经出现在开放模型中。我认为 Deep Seek Flash 模型尤其引领了这一趋势。
Such a trend is already appearing in open models. I think the Deep Seek Flash model, in particular, is leading this trend.
再看一下 Rama Cloud 的模型配置分析,增长率最高的领域是 Deep Seek 模型,这得益于 Deep Seek Flash 模型的推出。因此,一种新型的 flash 模型已经出现,它能够处理 80% 的工作,速度非常快,而且价格非常实惠。这种新型模型将支持各种用例。
Looking again at the model configuration analysis of Rama Cloud, the area showing the highest growth rate is by far the Deep Seek model, which is thanks to the introduction of the Deep Seek Flash model. As such, a new type of flash model has emerged that is capable of handling 80% of the work, is very fast, and is also very affordable. This new type of model will enable various use cases.
是的,这些模型将成为处理所有简单重复任务的核心模型。
Yes, these models will be the core models for handling all simple repetitive tasks.
你说得对。你将不再需要担心请求数量或 token 用量。你会更倾向于尽可能多地消费。这是因为这些模型不仅能解决最困难的问题,也能解决困难的问题。
You're right. You will no longer need to worry about the number of requests or token usage. You will have a much stronger tendency to consume as much as possible. This is because these models can solve not only the most difficult problems, but also difficult problems.
如果我们重新考虑前面提到的协调层,通过协调这些 flash 模型相互配合来执行各种任务,我们可以获得与更大模型提供的类似优秀结果。因此,使用这些廉价的模型不仅提高了可访问性,还加快了执行速度,并允许处理更大量的数据。此外,通过连接这些模型,你可以构建通过编排解决的新问题。这对新初创公司、现有的推理提供商以及解决工作流问题的大型企业来说都是一个非常有趣的领域。
If we reconsider the coordination layer mentioned earlier, by coordinating these flash models with one another to perform various tasks, we can obtain excellent results similar to those that a larger model can provide. Therefore, using these inexpensive models not only improves accessibility but also speeds up execution and allows for the processing of larger amounts of data. In addition, by connecting these models, you can build new problems that are solved through orchestration. This is a very interesting field for new startups, existing inference providers, and large enterprises solving workflow problems.
最终,连接这些模型将非常有用,而且无需担心底层成本。
Ultimately, connecting these models will be very useful, and there will be no need to worry about the underlying costs.
是的,如你所知,当我们第一次开始谈论 AGI(通用人工智能)时,甚至在这个播客上也有过某种辩论。我认为许多 AI 研究人员曾说:“一个巨大的、神一样的模型会出现并做所有事情。”然而,到目前为止似乎并非如此。
Yes, as you know, when we first started talking about Artificial General Intelligence (AGI), there was a sort of debate, even on this podcast. I think many AI researchers said, "A giant, god-like model will appear and do everything." However, it doesn't seem to have gone that way so far.
当然,如果你需要入侵 NSA,你可能需要像 Mythos 这样的人工智能,但对于大多数用例,比如编排或小规模模型,我认为任务组织实际上提高了可重复性和可靠性,并提供了各种方法,以实际可行的成本运作。
Of course, if you needed to hack the NSA, you would need artificial intelligence like Mythos, but for most use cases, such as orchestration or small-scale models, I believe that task organization actually increases repeatability and reliability and provides various methods to operate at a practically feasible cost.
那么,哪个更好:最好的模型、几个小型专用模型,还是更简单的模型?
So, which is better: the best model, several small, specialized models, or a simpler model?
到目前为止,后者似乎更占主导地位。我相信对于大多数客户用例,存在一个模型变得足够好的水平。你可以继续使用那个水平的智能。虽然模型可以变得更快、拥有更好的架构或增加新功能,但没有必要依赖最好的模型。然而,肯定会有需要最强大模型的用例。这样的模型将来会继续出现,可能性非常有趣,但有时也可能令人恐惧。
So far, the latter has appeared to be more dominant. I believe that for most customer use cases, there is a level at which the model becomes sufficiently good. And you can continue to use that level of intelligence. While models can become faster, have better architectures, or add new features, there is no need to rely on the best model. However, there will certainly be use cases where the most powerful model is needed. And such models will continue to emerge in the future, and the possibilities are very interesting, but at times they may also be frightening.
然而,我相信开放模型展现其真正价值的一般用例是,当你遇到一个可以通过开放模型解决的难题,即使它不是商业中最具挑战性的问题。一个声称拥有顶级性能的开放模型会出现吗?
However, I believe the general use case where an open model demonstrates its true value is when you reach a difficult problem that can be solved through an open model, even if it is not the most challenging problem in business. Will an open model boasting top performance appear?
我认为有可能。我们看到一些公司利用尖端技术取得了有趣的进展,比如智谱 AI 和 GLM。我们也不能忽视 Kimi 模型,它已成为 Web 开发领域的最佳模型。这在市场上引发了新浪潮,并创造了激烈的竞争格局,而不仅仅是缩小差距。
I think it is possible. We are seeing interesting developments in companies utilizing cutting-edge technology, such as Zhipu AI and GLM. We also cannot leave out the Kimi model, which has established itself as the best model in the field of web development. This caused a new wave in the market and created a fierce competitive landscape rather than simply closing the gap.
这种竞争格局似乎让事情变得更加激动人心。这可能是一个有点敏感的问题,但你对与此相关的地缘政治方面有什么看法?
This competitive landscape seems to make things even more exciting. This might be a somewhat sensitive question, but what are your thoughts on the geopolitical aspects related to this?
地缘政治问题通常源于模型的来源。然而,你与客户和用户相处的时间越多,模型在安全环境中如何运行、在哪里运行以及以何种方式运行就越重要。
Geopolitical problems often originate from the source of the model. However, the more time you spend with customers and users, the more important it becomes how, where, and in what way the model runs in a secure environment.
嗯,但我认为美国客户能够使用在美国训练的模型非常重要。
Well, but I think it is very important that US customers can use models trained in the US.
我们咨询两类主要客户。第一类客户不太关心模型的来源,只对模型在哪里执行感兴趣。然而,对于每一个这样的客户,都有另一个客户说模型的来源非常重要。这是因为数据就是数据。这不仅仅是安全问题;模型如何沟通也很重要。我们都经历过与模型沟通的过程,模型可能会经历像机器人一样说话,然后变得更友好的阶段。这方面也很重要。然而,我相信最重要的是确保一个能够从头到尾理解数据源的模型。NeMo 模型在这方面非常出色。你可以直接检查和分析构成这个模型的元素。
We consult with two main types of customers. The first category consists of customers who do not care much about the source of the model, but are only interested in where the model was executed. However, for each such customer, there is a customer who says that the source of the model is very important. This is because data is data. It is not just a security issue; how the model communicates is also important. We all go through the process of communicating with models, and models may go through stages where they speak like robots and then speak more friendly. This aspect is also important. However, I believe the most important thing is to secure a model that can understand the data source from beginning to end. The NeMo model is excellent in this regard. You can directly examine and analyze the elements that made this model.
这是因为人们正在积极利用开源模型执行关键任务。例如,你可能看到过一篇关于芬兰发电厂的在线帖子,该电厂使用 Llama 模型构建了一个分析系统,用于检测电压波动并维持电力供应。正是因为这个原因,模型的来源非常重要。我们如何确保对于如此重要的任务,即使一个中国模型托管在美国,也没有可能引发问题的隐藏陷阱?就像谜团中的候选人一样。
This is because people are actively utilizing open source models for mission-critical tasks. For example, you may have seen an online post about a power plant in Finland that built an analysis system using the Llama model to detect voltage fluctuations and maintain power supply. It is precisely for this reason that the source of the model is very important. How can we be sure that for such an important task, even if a Chinese model is hosted in the U.S., there are no hidden pitfalls that could potentially cause problems? Just like the Candidate in the mystery.
实际上有没有涉及候选人推理的案例?我现在想不起来。如果有,我肯定会听说。然而,媒体文章中有一些部分没有很好地覆盖。我的意思是,大公司(包括财富 500 强公司)的 IT 和安全团队到底有多强大。这个问题已经很熟悉了,因为开源软件,尤其是常见应用,有成千上万的依赖项。这不是一个新问题。即使是一个依赖项也可能导致整个应用出现严重的安全问题。供应链污染攻击非常严重。这个问题已经存在了几十年。从这个意义上说,它并不新鲜。只是有点更不透明,因为很难深入研究模型。
Were there actually any cases involving a candidate in a man's deduction? I can't think of anything right now. If it had been there, I would have definitely heard it. However, there are parts that are not well covered in media articles. I mean exactly how strong the IT and security teams are at major companies, including Fortune 500 firms. This problem is already familiar because open source software, especially common applications, have thousands of dependencies. It's not a new problem. Even a single dependency can cause serious security issues for the entire application. Supply chain contamination attacks are really serious. This is a problem that has existed for decades. In that sense, it is not new. It is just a bit more opaque because it is difficult to delve into the model.
它是确定性的。
It is deterministic.
它绝对是确定性的。据说,如果模型通过安全检查得到适当验证,大多数问题都可以解决,至少从我们客户那里听到的是这样。
It is definitely deterministic. It is said that if the model is properly verified through safety inspections, most issues can be resolved, at least from what we hear from customers.
你能谈谈 Ollama 的创立吗?你在 Docker 生态系统中成长,许多人希望拥有像 Ollama 那样似乎永远持续的强大品牌竞争力。
Could you tell me about the beginning of Ollama? You have grown within the Docker ecosystem, and many people hope to have strong brand competitiveness that seems like it will last forever, like Ollama.
不,那只是……一个非常特殊的情况。感觉就像我骑在一个巨大的油井上。
No, it was just ... a really special situation. It felt just like I was riding on top of a giant oil well.
你能为从事石油勘探的人谈谈这个吗?你在 2021 年和 Jared 一起工作过吗?
Could you tell me a little about that for the people engaged in oil exploration? Did you work with Jared in 2021?
是的,我和我的联合创始人之前在 Docker 开发了 Docker Desktop。所以我很清楚什么是出色的开发者体验。然而,在创立 Ollama 后的最初几年里,我专注于根据我为开发者设计出色体验的经验,弄清楚要解决什么问题。
Yes, my co-founder and I previously developed Docker Desktop at Docker. So I knew very well what a great developer experience was. However, during the first few years after founding Ollama, I focused on figuring out what problems to solve, based on the experience I had gained designing great experiences for developers.
你带着一个完全不同的想法申请了 YC,对吗?
You applied to YC with a completely different idea, right?
是的,没错。
Yes, that's right.
你还记得 2021 年冬天申请 YC 时的口号是什么吗?
Do you remember what the slogan was when you applied to YC in the winter of 2021?
我认为当时还没有明确决定。我想我们决定重新专注于为容器和 Kubernetes 创造出色的桌面体验。我记得我在申请中写的内容。大概是 KiteMatic 或 Docker Desktop for Kubernetes。是的,它实际上和 Docker Desktop 一样。我认为它是一个优秀的 Kubernetes 工具。
I don't think it was clearly decided. I think we decided to refocus on creating a great desktop experience for containers and Kubernetes. I remember what I wrote on the application. It was something like KiteMatic or Docker Desktop for Kubernetes. Yes, it was practically the same as Docker Desktop. I think it is an excellent Kubernetes tool.
这正是我作为第二次创业者面临的困难之一。Michael 和我认为我们在很多方面试图过度设计这个想法。然后,在接下来的两三年里,例如,以 Ollama 为例,他们在 2021 年参加了 YC,直到 2023 年 7 月获得 A 轮投资后才推出。那是在完成 YC 之后。这个过程是一个寻找能够满足开发者的客户问题的旅程。
That was exactly one of the difficulties I faced as a second founder. Michael and I think we tried to over-design the idea in many ways. And then, over the next two to three years, for example, in the case of Ollama, they participated in YC in 2021 and only launched in July 2023 after securing Series A investment. It was after finishing YC. The process was a journey to find customer problems that could satisfy developers.
我认为实际上很幸运,我尝试了各种想法,并在 2023 年改变了方向。那一年,Llama 发布,开创了称为开源模型的新趋势。
I think it was actually fortunate that I tried various ideas and changed direction by 2023. That year, Llama was released and created a new trend called the open source model.
这就是为什么名字叫 Ollama。
That's why the name is Ollama.
不一定。
That isn't necessarily the case.
啊,我明白了。
Ah, I see.
一般来说,根据我们的经验,无论你是否想到 subreddit 'Local Llama',Ollama 实际上意味着开放模型。选择这个名字时,它并不是从现有模型取来的。它是动物名和 LLM 的组合。
Generally, in our experience, whether you think of the subreddit 'Local Llama' or not, Ollama actually means an open model. When choosing the name, it wasn't taken from an existing model. It's a combination of an animal name and LLM.
是的,我认为创造这样一个角色很重要。那么,“什么是一个好的角色名字?”开放模型有点吓人,你知道。我们需要一个像吉祥物一样的东西。
Yes, I think it was important to create such a character. So, 'What would be a good character name?' Open models are a bit scary, you know. We need something like a mascot.
差不多一个。
About one.
是的。
Yes.
GitHub 还是不听我的建议,没有在跟羊驼相关的活动上带一只羊驼来。
GitHub still isn't listening to my advice to bring a llama to the llama-related event.
我们吸引了 Benchmark 等优秀投资者的投资,但在 2023 年,我们现在所处的未来甚至还没有开始。
We attracted investment from great investors like Benchmark, but in 2023, the future we are in now hadn't even begun.
我很好奇当时的融资过程和愿景。
I am curious about the investment attraction process and vision at that time.
是的,我们在 2022 年与 Benchmark 合作。那是 Dolly Park 时代,也是在 ChatGPT 之前。在融资路演时,我们强调了我们在改善开发者体验和解决安全问题方面的努力。
Yes, we partnered with Benchmark in 2022. It was the Dolly Park era, and before ChatGPT. When pitching to attract investment, we emphasized our efforts to improve the developer experience and solve security issues.
Benchmark 的合伙人 Peter 是我们之前在开发 Docker 时就认识的人,因为他是 Docker 的 A 轮投资者。因此,与我们存在理由相关的人力资源非常重要。
Peter, a partner at Benchmark, is someone we have known since we were developing Docker previously, because he was a Series A investor in Docker. Therefore, human resources that were with our reason for existence were very important.
解决 Kubernetes 中的 SSO 问题——也就是实际解决问题——并不是我们真正的热情所在。我相信我们非常幸运能遇到一位理解我们价值观和目标的合伙人。因为当时,他们更看重我们的愿景,而不是我们试图解决的问题。
Solving the SSO problem in Kubernetes—that is, actual problem-solving—was not our true passion. I believe we were truly fortunate to meet a partner who understands our values and goals. Because at the time, they valued our vision more than the problem we were trying to solve.
啊,既然你是 A 轮投资者,那是在转型之前。对吗?
Ah, since you were a Series A investor, that was before the pivot. Is that right?
我不知道。我在看 Cloud Token 模型系列的图表,从这张图来看,Olama 的故事似乎始于 2026 年 2 月,并呈爆炸式增长。但有趣的是,它实际上始于 2021 年。
I didn't know. I was looking at the graph by Cloud Token model family, and looking at this graph, it seemed like the Olama story started in February 2026 and grew explosively. But it's really interesting that it actually started in 2021.
在长期没有取得实质性成果的项目上工作后,突然一切爆炸式增长,那是什么感觉?这对联合创始人、团队和员工的心理有什么影响?那段经历是怎样的?
What was it like when everything suddenly grew explosively after working on a project that hadn't produced proper results for a long time? How did it affect the psychology of the co-founder, the team, and the employees? What was the experience like?
说实话,我有点害怕。有几个原因,但首先,你知道有时候即使你不停地打电话试图为开发者或客户解决问题,事情还是不顺利吗?在那种时刻,其他问题比‘我们是否出现在媒体上’或‘项目是否成功’这样的想法更重要。‘我们真的在解决某人的问题吗?’我产生了疑问。
Honestly, I was a little scared. There are a few reasons, but first of all, you know how there are times when things don't go well even if you keep talking on the phone trying to solve a problem for a developer or a customer? In such moments, other issues feel more important than thoughts like 'whether we appear in the media' or 'whether the project succeeds.' Are we really solving someone's problem? I had a question.
就像在荒野中迷失方向,我们的指南针是什么?客户是很好的指南针,但当指南针看不见时就更可怕了。因为很多时候,人们知道自己想解决的问题,但不知道具体是什么问题。
As if lost in the wilderness, what is our compass? Customers serve as an excellent compass, but it is scarier when the compass is invisible. This is because there are many cases where people know the problem they want to solve, but do not know specifically what that problem is.
我和联合创始人创办这家公司,是因为我们过去有经营一家被 Docker 收购的公司的经验。那是创始团队。Arnar Star 说:“我想通过解决开发者真正挣扎的部分来提供最佳体验。”然而,花两年时间试图找到那个问题真的很可怕。因为我有 10 多名团队成员,所以更难。
My co-founder and I started this company because we had experience running a company in the past that was acquired by Docker. It was the founding team. Arnar Star said, "I want to provide the best experience by solving the parts that developers really struggle with." However, the two years spent trying to find that problem was really scary. It was even harder because I was with more than 10 team members.
我非常感谢那些在我尝试各种想法时一直在我身边的团队成员。
I am truly grateful to my team members who stayed by my side while I was trying out various ideas.
如你所知,一个鲜为人知的事实是,我们已经将重点从 Kubernetes 安全转向桌面开发者安全。这是我们之前没有提到的重要转折点。我们基于这种形式开发了一个模型。
As you know, a little-known fact is that we have shifted our focus from Kubernetes security to desktop developer security. This is an important turning point that we have not mentioned before. And we developed a model based on that form.
当模型发布时,我们说:“这是一次飞跃,但之所以可能,是因为我们知道开发者面临什么问题,以及他们想要什么样的感觉。”然而,通过 LLM,我终于清楚地意识到了那个问题。
When the model was launched, we said, "It was a leap, but it was possible because we knew what problems developers were facing and what kind of feeling they wanted to have." However, through the LLM, I finally clearly realized that problem.
运行 Llama 模型真的很难。我想:“啊,这就是问题所在。”当我真正让它运行起来的那一刻,真是太神奇了。就像从 0 到 1 的转变。
Running the Llama model was really difficult. I thought, "Ah, this is a problem." And the moment I actually made it work was truly amazing. It was just like a transition from 0 to 1.
Olama 的开发过程中有几个转折点,但最重要的一个一定是转向开发托管 LLM。我很想听听那个转折点背后的故事。它是怎么开始的?是你们头脑风暴时突然冒出来的吗?当你们想到那个想法时,公司里的每个人都立刻认为这是正确的方向吗?还是有过重大争论,直到实际实施后才变得清晰?
There were several turning points in the development process of Olama, but the most important one must have been the move toward developing a hosted LLM. I am curious to hear the story behind that turning point. How did it start? Did this just come out while you were brainstorming ideas? When you came up with that idea, did everyone in the company immediately think it was the right direction? Or was there a major debate, and did it only become clear after it was actually put into practice?
我们聚在一个房间里交谈。当时,团队分散在多伦多和帕洛阿尔托,但现在主要在帕洛阿尔托。那时,我们决定放弃一切,从头开始。我在想,如果我现在加入 YC,我会怎么做。
We gathered in one room and talked. At that time, the team was split between Toronto and Palo Alto, but now it is mainly in Palo Alto. At that time, we decided to give up everything and start over from scratch. I'm wondering what I would do if I joined YC right now.
在与 LLM 用户交谈时,我们发现了两个主要问题。我尝试使用开源 LLM,也尝试解决 LLM 本身的问题。
While talking with LLM users, we discovered two major problems. I tried using open source LLMs, and I also tried solving the LLM itself.
第一个问题是,我们能否通过构建一个可以访问和托管任何模型的网关来无缝使用它。当时,我把它看作是 LLM 的一种细分。我认为这样想是个好方法。现在这已经演变成了巨大的路由器想法,但这只是开始阶段。我认为这是一个巨大的机会。
The first problem was whether we could use it seamlessly by building a gateway that could access and host any model. At the time, I thought of it as a kind of segment for LLM. I think thinking that way is a good approach. And that has now evolved into the huge router idea, but it is just the beginning stage. I think it's a tremendous opportunity.
另一个问题是,我们的团队主要由 XV 开发者组成,而 X Docker 开发者说:“我们知道如何运行系统”,并试图帮助使用开源模型构建系统。而我们是系统开发者。
Another problem was that our team consisted mostly of XV developers, and the X Docker developers said, "We know how to run the system," and tried to help build the system using an open source model. And we were system developers.
是的,我是一名系统开发者。所以,我们花时间在内部认真反思团队。由于安全团队与开发者工具团队是完全不同的销售团队,我认为我们早就应该这样做了。
Yes, I was a system developer. So, we took the time to seriously reflect on the team internally. Since the security team is a completely different sales team from the developer tools team, I think we should have done that a long time ago.
在这样做的过程中,我们得出了“让我们试一试”的结论。我们决定尝试在两周内发布 Ollama 的第一个版本。就在两周快结束的时候,Llama 2 发布了。所以,我们决定:“好吧,我们发布吧。”我们是行动导向的。
While doing so, we reached the conclusion that "let's give it a try." We decided to try releasing the first version of Ollama within two weeks. And just as the two weeks were almost over, Llama 2 was released. So, we decided, "Okay, let's launch it." We were action-oriented.
回想起来,一切都在两周内发生。从构思想法到发布,甚至获得了比之前产品多得多的用户。在那之前,两年里,我只过度思考客户和产品,但实际什么也没做出来。
Looking back, everything happened in two weeks. From conceiving the idea to launching it, and even securing far more users than the previous product. Before that, for two years, I only thought excessively about customers and products, but I couldn't actually produce anything.
我第一次在 Reddit 上听说 Ola Ma。你不知道那是你的公司,对吧?
I first heard about Ola Ma on Reddit. You didn't know it was your company, did you?
我只是在看标志,但我对运行本地模型律师事务所感兴趣。我想那是一个本地 LLM 的 subreddit,那里有很多帖子称赞 Ola Ma。“哦,是 YC 公司吗?”我这么想。后来我发现那是一家不同的公司。它没有在我们的内部系统中注册。
I was just looking at the logo, but I was interested in running a local model law firm. I think it was a local LLM subreddit, and there were a lot of posts there praising Ola Ma. "Oh, was it YC Company?" I thought so. I found out later that it was a different company. It was not registered in our internal system.
在和 Jared 谈话时,我说:“哦,对了,说到安全问题,我们有个叫 Ola Mara 的东西。”我记得谈过这个。
While talking to Jared, I said, "Oh, right, speaking of security issues, we have something called Ola Mara." I remember talking about it.
我和 Jared 谈话时在楼梯上遇到了他。他穿着 T 恤或某种纪念品,所以我问:“呃,大家,你们是 Olla 公司的吗?”我问。“是的,”他回答。感觉就像遇到了摇滚明星。
I was talking to Jared and ran into him on the stairs. He was wearing a T-shirt or some kind of souvenir, so I asked, "Uh, everyone, are you from the Olla company?" I asked. "Yes," and answered. It felt just like meeting a rock star.
把握时机一定很难,但我认为如果你更早做出这个大胆的尝试会更好。YC 时代是最合适的时机,但当时不可能。
It must have been really difficult to get the timing right, but I think it would have been better if you had made this bold attempt much earlier. The YC era was the most appropriate time, but it was impossible back then.
因为 Ollama 根本不存在。没错。Llama Llama 还没有发布。
Because Ollama didn't even exist. That's right. Llama Llama hadn't been released yet.
我认为它是最早迅速达到 GitHub 10 万星标的项目之一。你记得吗?
I think it was one of the first projects to reach 100,000 stars on GitHub very quickly. Do you remember?
那真的很快。
It was really fast.
我不太记得具体花了多长时间,但比 Docker 或 Kubernetes 快得多。正如你所说,一切突然开始顺利起来,我们开始快速增长,但作为创始人,我真的很困惑,因为我无法解释到底为什么会这样。我认为这是解释产品市场契合度的最佳方式。
I don't remember exactly how long it took, but it was much faster than Docker or Kubernetes. As you mentioned, everything suddenly started going well and we began growing rapidly, but as a founder, I was really bewildered because I couldn't explain exactly why that was happening. I think that is the best way to explain product-market fit.
产品市场契合度也有几个阶段。如你所知,我们今年早些时候才开始通过 Llama 的云服务产生收入,但看到人们迷上产品后,一切瞬间就变了。
There are several stages of product-market fit as well. As you know, we only started generating revenue with Llama's Cloud earlier this year, but seeing people get hooked on the product, everything changed in an instant.
如果在 YC 时代就能这样做就好了,但在很多方面是不可能的。当你想到这一切时,真的很惊人。从 Reddit 上几个对在家自己组装电脑感兴趣的极客开始,仅仅两年时间,85% 的财富 500 强公司都在使用我们的产品。就像个人电脑从家酿计算机俱乐部开始用了 10 年才普及一样,我们似乎只用了 12 到 18 个月就广泛分布了。
It would have been good if we had done it this way during the YC era, but it was impossible in many ways. When you think about all of this, it's really amazing. From just a few geeks on Reddit interested in building their own computers at home, in just two years, 85% of Fortune 500 companies are using our product. Just like how PCs became popularized in 10 years starting from the Homebrew Computer Club, it seems like it only took 12 months—18 months, to become widely distributed.
是的,这是最让我们惊讶的部分之一。开源模型的早期用户是那些把摆弄这摆弄那当爱好的人,他们想:“哇,我不知道还能这样!”我也这么认为。
Yes, that is one of the parts that surprised us the most. The early users of the open source model were people who tinkered with this and that as a hobby, thinking, "Wow, I didn't know something like this was possible!" I thought so.
然而,这种快速增长有两个原因。首先,任何人都可以免费开始并在任何地方运行它。这对财富 500 强公司的 IT 开发团队来说是一个极其有用的因素。因为无需请求许可即可使用。
However, there are two reasons for this rapid growth. First, anyone can start for free and run it anywhere. This is an incredibly useful factor for IT development teams at Fortune 500 companies. This is because there is no need to request permission to use it.
所以,对业余用户非常实用的东西也很快被公司内部的开发者采用。这就像命运一样发生了。
So, what was very useful for hobby users was quickly adopted by developers within the company as well. It happened just like fate.
例如,在数据库领域也能看到类似的情况。像 MongoDB 这样为开发者而生的数据库,已经迅速扩展到企业环境。这种转变非常容易,因为 LLM 是一种无状态方法。现在,随着我们转向云端,需要考虑的事情更多了。从客户的角度来看,有经济方面的考量,安全问题自然存在,还必须考虑其运行环境。然而,开放模型最大的优势——无论是业余爱好者还是 IT 开发者都喜欢的一点——就是你可以立即开始使用,无需任何许可。
For example, you can see similar cases in databases as well. Databases that started for developers, like MongoDB, have rapidly spread to enterprise environments. This transition was very easy because LLM is a stateless method. Now, there are many more things to consider as we move to the cloud. From the customer's perspective, there are economic aspects, security issues naturally exist, and consideration must also be given to the environment in which it will be operated. However, the biggest advantage of the open model, which everyone liked whether they were hobbyists or IT developers, is that you can get started right away. No permission was needed.
创收也是一个有趣的话题,我们聊聊这个吧?
Revenue generation is also an interesting topic, so shall we talk about it?
所以,在 2023 年,随着 Llama 2 突然流行起来,用户激增,GitHub 星标多达 10 万,很明显我们发现了了不起的东西。然而,实际使用它的是 Reddit 上的极客,完全没有产生任何收入。没有明确的方法从这些 Reddit 极客身上赚钱。实际上花了两年时间才找到商业模式,而且巧合的是,这与 Docker 当年面临的情况完全吻合。
So, in 2023, as Llama 2 suddenly gained popularity, with a surge in users and as many as 100,000 GitHub stars, it seemed clear that something great had been discovered. However, the people who actually used it were Reddit geeks, and no revenue was generated at all. There was no clear way to generate revenue from these geeks on Reddit. It actually took two years to find a business model, and coincidentally, it matched the situation Docker faced exactly.
那两年你是怎么想的?你担心吗?团队成员问:“我们该怎么处理商业模式?”你有没有问过“你们想过怎么创收吗?”这个问题?
What did you think during those two years? Were you worried? The team members asked, "What should we do with the business model?" Did you ask the question 'Have you been thinking about how to generate revenue?'
我相信开源模型总有两条创收之路,能同时对公司、开发者和客户都有利。其中之一是专注于隐私的 AI 产品,Ollama 开始以开源版本开发这个产品。然而,我们觉得总有一个时间点,Ollama 模型(例如工具调用功能)在最初发布时并不会立即被充分利用。换句话说,Ollama 仍然常常只用于最先进的模型中。
I believe there are always two ways for the open source model to generate revenue in a way that is good for the company, developers, and customers alike. One of them was an AI product focused on privacy, and Ollama started developing this product as an open-source version. However, we felt that there is always a point in time when Ollama models, for example, the tool calling feature, are not utilized immediately upon their initial release. In other words, Ollama was still often used only in state-of-the-art models.
重申一下,这可能有点想多了,我认为开放模型的产品市场契合度不如封闭模型高。
To reiterate, and this might be overthinking, I believed that open models do not have as high a product-market fit as closed models.
从哲学上讲,我们当时想抓住那个机会。我相信那个时刻已经到来,随着今年基于开放模型的编码智能体的出现。这是因为开放模型终于能够满足 AI 使用率最高领域的需求。未来有很多保护个人信息并安全使用的机会,许多财富 500 强公司已经采用了 Ollama。然而,我们问自己:“我们能为客户解决的最重要的问题是什么?”
And philosophically, we wanted to seize that opportunity at that time. I believe that time has come with the emergence of coding agents running on an open model this year. This is because open models have finally become able to meet the demand in the fields with the highest AI usage. There are many opportunities to protect personal information and use it safely in the future, and many Fortune 500 companies have already adopted Ollama. However, we ask ourselves, "What is the most important problem we can solve for our customers?"
他问过这个问题。地域方面是一个重要部分,但从来不是全部。关键是如何利用开源模型解决最困难的问题。
He asked the question. The regional aspect was an important part, but it never felt like the whole story. The key was how to utilize open source models for the most difficult problems.
所以我知道我必须在某种程度上等待,我也知道我必须等到市场成熟一些。同时,等待的危险是什么?嗯,如果你不小心,你最终会创造一种人们不考虑创收的文化。这也是我们在 Docker 时代学到很多教训的地方。创收不是优先事项。得益于团队过去的经验教训,我们对此非常清楚。
So I knew that I had to wait to some extent, and I also knew that I had to wait a little until the market matured. At the same time, what are the dangers of waiting? Well, if you aren't careful, you end up creating a culture where people don't think about generating revenue. This is also a lesson we learned a lot during the Docker days. Revenue generation is not a priority. Thanks to the lessons learned from previous experiences as a team, we were well aware of this.
然而,另一个重要因素是与客户的持续沟通。当开源项目成功时,最大的风险之一是将用户群和客户视为互联网上的一群乌合之众。这是一个非常危险的想法。
However, another important factor is continuous communication with customers. One of the biggest risks when an open source project succeeds is viewing the user base and customers as a simple crowd on the internet. This is a very dangerous idea.
你必须亲自会见客户,识别他们的需求,了解他们在做什么。我想知道他们六个月后想做什么,他们的故事是什么。我相信这正是我们在过去几年里应该做得更多的领域,我们现在正在那个领域投入大量精力。
You must meet customers in person, identify their needs, and know what they are doing. I wanted to know what they wanted to do in six months and what their story was. I believe that is exactly the area we should have done more of over the past few years, and we are now putting a lot of effort into that area.
你加入 YC 时是第二位创始人,对吧?我很好奇是什么让你决定去 YC 的。
You were the second founder when you joined YC, weren't you? I was curious about what made you decide to go to YC.
实际上,我们在这个问题上也纠结了相当长一段时间,但根本没有必要。我应该直接说:“我当然必须去 YC。”创业真的是一种非常孤独的经历,无论你的联合创始人有多出色。我的联合创始人 Michael 是我第一家公司的联合创始人,也是我在滑铁卢大学的室友。尽管如此,我还是忍不住感到孤独。尽管当时是疫情期间,但每周与 Jared 和其他创始团队成员交谈帮助我减轻了孤独感。我认为这是非常重要的一部分。
Actually, we also agonized over this issue for quite a while, but there was no need to. I should have just said, "Of course I have to go to YC." Starting a company is a really lonely experience, no matter how great your co-founder is. My co-founder Michael was a co-founder of my first company and was also my roommate at the University of Waterloo. Still, I couldn't help but feel lonely. Even though it was the pandemic, just talking with Jared and the other members of the founding group every week helped me feel less lonely. And I think that is a really important part.
当然,当我们最终搬到这里,疫情结束后,人脉网络真的很棒。能够遇到基于开源模型或各种 AI 技术创业的人真是太好了。
And of course, when we finally moved here and COVID ended, the network was really amazing. It was really great to be able to meet people starting businesses based on open source models or various AI technologies.
因为我在疫情前在滑铁卢大学就认识许多参加 YC 的创始人,所以我多少预料到会发生类似的事情。他们说 YC 的核心是彼此见面。那确实是非常重要的部分,既然我们知道这一点,就没有必要犹豫。此外,重要的是你可以避免许多不需要重复的错误。我喜欢 YC 社区的原因是创始人之间会透明地分享这些方面。
Since I knew many of the founders who participated in YC at the University of Waterloo before COVID, I somewhat expected that something like that would happen. They said that the core of YC is meeting each other. That was a really important part, and since we knew that, there was no need to hesitate. Also, it is important that you can avoid many mistakes that do not need to be repeated. The reason I like the YC community is that the founders transparently share those aspects with each other.
我仍然与 Docker 的创始人保持联系,他投资了我们的公司。在之前几代公司所经历的困难中,我们可以讨论那些不一定需要重蹈覆辙的,或者那些行之有效且可以应用于未来的。是的,仅仅避免重蹈覆辙就能带来巨大帮助,因为有时这样的错误可能会毁掉公司。
And I still keep in touch with the founder of Docker, who invested in our company. Among the difficulties experienced by previous generations of companies, we can talk about those that do not necessarily need to be repeated, or those that were effective and can be applied to the future. Yes, simply avoiding repeating those mistakes can be a great help, because sometimes such mistakes could have ruined the company.
是的,确实如此。
Yes, that is correct.
有一句名言说,我们团队中有相当多成员来自成功的公司。Docker 目前取得了巨大的成功。我们的一些团队成员从 VMware 早期就一直在那里。当然,一切事物都必然有起有落。我认为,让拥有不同经验的人聚集在一起,分享真正有效的方法,这一点非常重要。基于这些经验和专业知识积累的诀窍非常有帮助。
There is a famous saying that a significant number of our team members come from successful companies. Docker is currently achieving tremendous success. Some of our team members have been with VMware since its early days. Of course, everything is bound to have its ups and downs. I think it is really important for people with diverse experiences to gather in one place and share what actually worked. The know-how accumulated based on such experience and expertise is of great help.
啊,我正好想到这一点。
Ah, I just had that thought.
我经常和公司里 18、19 岁的年轻人混在一起,有时他们会问:“我该怎么办?”他们这样问我。所以我一直在想:如果你没有在大量观众面前展示过出色技能,或者没有在一个产品市场契合度明确的团队中工作过,那你绝对应该尝试一下。无论是一个月还是三个月,你都能在那三个月里学到很多东西,因为你会知道什么是好的。如果你没有那种经验,当然也不是不可能。即使是 YC 出来的人最终也能做到,但那要困难得多。从一开始就看到它实际运作,并想着“这就是 bug 数据库的样子”、“发布就是这样做的”、“所需的质量标准是这个水平”、“我们彼此要求的标准是这个水平”,与亲身经历这个过程之间有着巨大的差异。从一个拥有大量此类经验的联合创始团队开始,会有巨大帮助。尤其是在通过 AI 获得巨大力量的情况下,这些经验在形成坚定价值观方面发挥着重要作用。
I hang out a lot with 18 or 19-year-olds at my company, and sometimes they ask, "What should I do?" They asked me that. So I've been thinking: if you haven't had the experience of showcasing great skills to a large audience or working in a team with a clear product-market fit, you should definitely give it a try. Whether it's one month or three months, you will be able to learn a lot during those three months, because you come to know what is good. If you don't have that experience, of course it's not impossible. Even those from YC eventually manage to do it, but it is much more difficult. There is a huge difference between seeing it actually working from the start and thinking, "This is what a bug database looks like," "This is how a release is done," "The required quality standards are this level," and "The standards we demand of each other are this level," and actually experiencing that process firsthand. Starting with a co-founding team that has a lot of that kind of experience is a huge help. Especially in a situation where immense power has been gained through AI, such experiences play a significant role in forming firm values.
我可以在某些领域提供帮助,但我不会强迫你承担所有责任。想想软件是如何运作的。例如,如果你想想 llama,两年前创建的 llama 软件可能仍在野外游荡。如果有人迷上了你的软件并持续使用两年,它还能正常工作吗?
I can help in some areas, but I am not forcing you to take on all the responsibility. Think about how software works. For example, if you think about llamas, llama software created two years ago is likely still roaming the wild. If someone gets hooked on your software and continues to use it for two years, will it still work properly?
嗯,我希望它能更新到最新软件或作为云服务提供。然而,我认为那种经验已经深深印在我身上。我们的团队有很多来自 VMware 或 Nicira 的资深工程师,所以我们有丰富的经验。但与此同时,我认为从上一代 DevOps 和基础设施中学到的许多经验在 AI 世界中已不再适用。
Well, I hope it gets updated to the latest software or provided as a cloud service. However, I think that experience is ingrained in me. Our team has many veteran engineers from VMware or Nicira, so we have plenty of that experience. However, at the same time, I think there are many lessons learned from previous generations of DevOps and infrastructure that are no longer valid in the AI world.
啊,我明白了。你能告诉我哪些点不再适用吗?
Ah, I see. Could you tell me which points are no longer valid?
我经常引用的一个例子是“平台即服务(PaaS)”概念存在的时代。
An example I often cite is the era when the concept of 'Platform as a Service (PaaS)' existed.
是的。
Yes.
2010 年代的 Heroku 等公司就是典型例子。你知道,Docker 最初也是作为 PaaS 起步的。从初创公司的角度来看,有一种看法认为,作为其他系统之上的一个层会让自己变得脆弱,但在 AI 世界中,情况完全不是这样。事实上,向上攀登技术栈有时可能更好。这是因为我们更接近客户。基础设施领域也是如此。就像制作 llama 时必须锻炼各种肌肉一样。另一点是,这种 LLM(学习领导力模型)绝不是完美的。在系统世界中,我们希望一切都按设计精确运行。我们还必须经过测试和验证。然而,根据定义,LLM 并非如此。这不是 bug,而是一个特性。
Companies like Heroku in the 2010s are prime examples. Docker started as a PaaS, too, you know. From a startup's perspective, there is a perception that acting as a layer on top of other systems makes one vulnerable, but in the world of AI, that is not the case at all. In fact, climbing the stack can sometimes be better. It is because we become closer to the customer. The same applies to the infrastructure world. It is just like having to train various muscles while making a llama. Another point is that this LLM (Learning Leadership Model) is by no means perfect. In the system world, we hope that everything works exactly as designed. We also have to go through testing and verification. However, by definition, LLM is not like that. That is not a bug, but a feature.
是的。
Yes.
我认为实际上有点犹豫不决反而更好。在团队建设方面也是如此,得益于 AI,现在许多问题不再需要像 10 年前那样庞大的劳动力。想想客户支持管道是什么样的,提供云服务的过程是什么样的。AI 时代确实是一个完全不同的世界。
I think it is actually better to be a little indecisive. In terms of team building as well, thanks to AI, there are now many problems that do not require a large workforce, unlike 10 years ago. Think about what the customer support pipeline looks like and what the process of providing cloud services looks like. The AI era is truly a completely different world.
当并非所有工程师都确切知道代码如何运作时,如何构建服务?
How can a service be built when not all engineers know exactly how the code works?
这正是目前的情况。所以,我们的团队在从 2000 年代的基础设施 1.0 到 2010 年代的云,再到现在的 AI 领域的过程中,学到了新的教训。许多现有的规则也被打破了。事实上,这相当于为生态系统的双方提供了巨大的服务。从最终用户的角度来看,你可以体验到真正干净的服务,token 生成干净,API 易于理解且合乎逻辑。然而,反过来,如果没有像 Ollama 这样的支持层,你可能会遇到诸如“啊,默认的推理提供商在我输入这样的特定参数时给出了奇怪的错误”或“它期望 JSON 格式,但甚至没有文档记录”之类的问题。这种情况让人感觉快要疯了。智能体们尽力以某种方式解决它,但他们别无选择,只能几个小时地撞墙,直到问题解决。与此同时,你最终会想:“这真是一次糟糕的体验。”
That is exactly the situation right now. So, our team is learning new lessons as we have progressed from Infrastructure 1.0 in the 2000s to the cloud in the 2010s, and now to the field of AI. Many of the existing rules are broken as well. In fact, it amounts to providing tremendous service to both sides of the ecosystem. From the end user's perspective, you can experience a truly clean service where tokens are generated cleanly and the API is easy to understand and logical. However, conversely, if there is no support layer like Ollama, you may encounter problems such as, "Ah, the default inference provider gives a strange error when I input a specific parameter like this," or "It expects a JSON format, but it isn't even documented." This is a situation that feels like it's driving me crazy. The agents try their best to solve it somehow, but they have no choice but to bang their heads against the wall for hours until it is resolved. In the meantime, you end up thinking, 'This is a truly terrible experience.'
所以,我想推理提供商正在帮助修复这些根本性的 bug。
So, I suppose inference providers are helping to fix these fundamental bugs.
是的,这是我们工作的一部分。而且我认为在开放模型环境中最大的机会之一是策展。将碎片化的模型、推理技术、云服务和工具整合在一起,使它们正常工作,是一项非常重要的任务。这是因为,正如你提到的,最终开发者想要构建软件。他们想要构建一些东西。他们想要构建下一个公司,他们想要构建下一个应用。我认为这正是我们发挥作用的地方,但许多优秀的服务也扮演着重要角色。例如,Open Router 就是一个很好的例子,因为它提供各种模型。开发者不需要注册超过 100 个提供商;他们只需要使用一个。他们可以在一个地方付款。Open Code 也是如此;你可以用一个工具集成所有模型。这为想要测试新技术的开发者提供了非常强大的体验。
Yes, that is part of what we do. And I believe one of the biggest opportunities in an open model environment is curation. Bringing together fragmented models, inference technologies, cloud services, and harnesses to make them work properly is a very important task. This is because, as you mentioned, end developers want to build software. They want to build something. They want to build the next company, they want to build the next application. I think that is exactly where we play a role, but numerous excellent services also play an important part. For example, Open Router is a good example in that it provides various models. Developers don't need to sign up for over 100 providers; they just need to use one. They can pay in one place. The same applies to Open Code; you can integrate all models with a single harness. This provides a very powerful experience for developers who want to test out new technologies.
审查以下模型,看看它们是否更适合你的用例。换句话说,通过策展整合各种模型和提供商,并实际提供有效的解决方案,往往存在困难。
Review the following models and see if they are a better fit for your use case. In other words, there are often difficulties in integrating various models and providers through curation and actually delivering effective solutions.
感谢你的参与。时间到了。谢谢邀请。
Thank you for being with us. Time is up. Thank you for the invitation.