Harvey's co-founder shares how application-layer companies can compete with frontier labs by leveraging the frontier ecosystem, building benchmarks, and post-training open-source models.
要点 · TL;DR
应用层公司可以通过利用开源模型和针对特定任务的后训练来与前沿实验室竞争。 Application-layer companies can compete with frontier labs by leveraging open-source models and post-training on specific tasks.
构建并开源高质量基准对于模型训练和评估至关重要,并能吸引生态系统合作伙伴。 Building and open-sourcing high-quality benchmarks is crucial for model training and evaluation, and attracts ecosystem partners.
在真实数据敏感或不可用时,由领域专家指导的合成数据生成是创建训练数据的实用方法。 Synthetic data generation guided by domain experts is a practical method for creating training data when real data is sensitive or unavailable.
核心观点 · Key points
应用层公司可以通过利用前沿生态系统,使用开源模型和NeMo实验室进行后训练,与前沿实验室竞争。 Application-layer companies can compete with frontier labs by leveraging the frontier ecosystem, using open-source models and NeMo labs for post-training.
构建高质量的基准测试对于训练和评估模型至关重要,开源这些基准可以通过社区反馈提高其质量。 Building high-quality benchmarks is essential for training and evaluating models, and open-sourcing them can improve their quality through community feedback.
在真实数据敏感或不可用时,由领域专家指导的合成数据生成是创建训练数据的关键方法。 Synthetic data generation guided by domain experts is a key method for creating training data when real data is sensitive or unavailable.
在考虑后训练之前,必须拥有强大的模型服务基础设施,包括回退和路由。 Having robust model serving infrastructure, including fallbacks and routing, is necessary before considering post-training.
对开源模型进行后训练可以在特定任务上达到前沿水平,使其成为应用公司的可行策略。 Post-training open-source models can achieve frontier-level performance on specific tasks, making it a viable strategy for application companies.
反共识 · Contrarian takes
应用层公司不应在人才或算力上与前沿实验室竞争,而应利用前沿生态系统。 Application-layer companies should not try to compete with frontier labs on talent or compute; instead, they should leverage the frontier ecosystem.
开源基准测试在战略上是有益的,尽管它帮助实验室改进,因为它吸引了生态系统合作伙伴并验证了基准。 Open-sourcing benchmarks can be strategically beneficial, even though it helps labs improve, because it attracts ecosystem partners and validates the benchmark.
合成数据虽然不完美,但在真实数据不可获取时,是训练模型的实用起点。 Synthetic data, while not perfect, is a practical starting point for training models when real data is inaccessible.
最终目标不是构建最好的法律模型,而是让每个律师事务所都能针对其特定工作定制模型,这是一种不同的范式。 The end goal is not to build the best legal model, but to enable every law firm to customize models for their specific work, which is a different paradigm.
在个体生产力上竞争不如关注组织生产力重要,后者涉及跨项目协调人类和智能体。 Competing on individual productivity is less important than focusing on organizational productivity, which involves orchestrating humans and agents across projects.
本期章节 · Chapters(共 11)
引言Introduction
预算内建研究实验室Building a Research Lab on a Budget
构建基准测试Building Benchmarks
用开源模型进行后训练Post-training with Open-Source Models
后训练与模型服务Post-training and Model Serving
法律训练数据生成Data Generation for Legal Training
招聘策略Hiring Strategy
基准创建挑战Benchmark Creation Challenges
端到端管道调试End-to-End Pipeline Debugging
后训练与门控Post-training and gating
项目管理与企业复杂性Project Management and Enterprise Complexity