Serge 首席执行官 Edwin Shan 讨论了优化有缺陷基准(如 El Marina)的后果,用户倾向于选择花哨、冗长的回答而非准确的回答。
Edwin Shan, CEO of Serge, discusses the consequences of optimizing for flawed benchmarks like El Marina, where users prefer flashy, verbose responses over accurate ones.
要点 · TL;DR
糟糕的基准如 Elo 导致模型优化点击诱饵而非真正质量。 Bad benchmarks like Elo cause models to optimize for clickbait, not real quality.
由老练评估者进行的人工评估是衡量模型进展的黄金标准。 Proper human evals with sophisticated evaluators are the gold standard.
强化学习环境是下一个训练范式,需要丰富的模拟和工具。 RL environments are the next training paradigm, requiring rich simulations.
核心观点 · Key points
针对 Elo 等不良基准进行优化会导致模型优化点击诱饵,而非真正质量。 Optimizing for bad benchmarks like Elo can lead to models optimizing for clickbait, not real quality.
由经验丰富、富有创造力的评估者进行适当的人类评估是衡量模型进步的金标准。 Proper human evals with sophisticated, creative evaluators are the gold standard for measuring model progress.
强化学习环境是训练范式的下一步,需要丰富的模拟和工具支持。 RL environments are the next step in training paradigms, requiring rich simulations and tooling.
模型可能以多种方式奖励作弊和失败;仅靠最终奖励是不够的。 Models can reward hack and fail in myriad ways; final reward alone is insufficient.
数据创建的质量需要品味、成熟度和创造力,而不仅仅是资历。 Quality in data creation requires taste, sophistication, and creativity beyond credentials.
不同前沿实验室在目标函数上存在分歧,这塑造了模型行为和进步。 Different frontier labs diverge in objective functions, shaping model behavior and progress.
反共识 · Contrarian takes
如果实验室在没有适当测量的情况下针对错误数据进行优化,模型可能在 6-12 个月内退化。 Models can regress over 6-12 months if labs optimize on wrong data without proper measurements.
一些前沿实验室忽略 Elo 等流行基准,反而表现更好。 Some frontier labs ignore popular benchmarks like Elo and perform better as a result.
不会有一个模型统治一切;不同的理念将塑造多样化的 AI。 There will not be one model to rule them all; different theses will shape diverse AIs.
最终每家公司都应训练自己的模型,以符合其独特理念。 Eventually every company should train its own models to align with their unique thesis.
视频评估的质量需要品味和成熟度,而不仅仅是遵循指令。 Quality in video evaluation requires taste and sophistication, not just instruction following.
最大的错误是停止发布见解;分享观点对该领域至关重要。 The biggest mistake was stopping publishing insights; sharing viewpoints is crucial for the field.
本期章节 · Chapters(共 20)
引言与基准问题Introduction and Benchmark Issues
编码模型训练中的数据质量与测量Data quality and measurement issues in coding model training
优秀评估者的特质Qualities of a good evaluator
复杂性、创造力、指令遵循Sophistication, Creativity, Instruction Following
模型对奖励的利用与底层能力Models hacking rewards and underlying capabilities
数据工作中的质量与资历Quality vs credentials in data work
对Meta收购Scale的反应Reaction to Scale acquisition by Meta
文化与招聘理念Culture and hiring philosophy
训练范式的分歧Divergence in training paradigms
单一模型与专用模型One model vs. specialized models
企业为何应自训模型Why companies should train their own models