OpenAI 首席研究官 Mark Chen 分享他从高频交易到 AI 研究的历程,强调复现对培养研究品味的重要性,以及 AlphaGo“第 37 手”带来的启发。
OpenAI's Chief Research Officer Mark Chen shares his journey from high-frequency trading to AI research, the importance of replication for developing research taste, and the inspiration from AlphaGo's 'move 37'.
要点 · TL;DR
规模定律仍然有效,预训练并未消亡。 Scaling laws still hold; pre-training is not dead.
强化学习在数学和编程等客观领域表现出色,而非主观领域。 RL excels in objective domains like math and coding, not subjective ones.
研究品味通过复现关键论文来培养。 Research taste is developed through replication of key papers.
核心观点 · Key points
缩放定律仍然成立;预训练并未消亡。 Scaling laws still hold; pre-training is not dead.
强化学习在数学和编程等客观领域表现出色,而非主观领域。 RL excels in objective domains like math and coding, not subjective ones.
研究品味通过复现关键论文来培养。 Research taste is developed through replication of key papers.
AGI 正在逼近;模型很快将能进行端到端研究。 AGI is approaching; models will soon do end-to-end research.
评估面临危机;经典基准已饱和。 Evals are in crisis; canonical benchmarks are saturated.
跨模态的统一架构降低了基础设施成本。 Unified architectures across modalities reduce infrastructure cost.
反共识 · Contrarian takes
预训练并未消亡;尽管屡次被预言,规模扩张仍在继续。 Pre-training is not dead; scaling continues despite repeated predictions.
强化学习在创意写作等主观领域因评分困难而举步维艰。 RL struggles with subjective fields like creative writing due to grading difficulty.
经常失败的高风险押注是 OpenAI 保持前沿成功的关键。 High-risk bets that often fail are key to OpenAI's frontier success.
模型缺乏品味;仍需要人类研究人员来产生想法。 Models lack taste; human researchers still needed for idea generation.
仅靠长上下文窗口不够;还需要压缩技术。 Long context windows alone are insufficient; compaction techniques are needed.
评估团队应与模型构建团队分离,以避免作弊。 Eval teams should be separate from model-building teams to avoid cheating.
本期章节 · Chapters(共 17)
开场与汤的故事Introduction and Soup Story
从交易员到研究员From Trader to Researcher
第 37 手与强化学习挑战Move 37s and RL Challenges
难评分领域的强化学习 vs 数学/科学RL in hard-to-grade fields vs. math/science
研究路线图的稳定性Stability of Research Roadmap
决策与专注Decision Making and Focus
识别新兴研究者Identifying Rising Researchers
顶尖工程师与研究员的相似之处Similarities Between Top Engineers and Researchers
评估危机与刷榜Evals crisis and benchmark maxing
Yacoub 的趣事Funny stories about Yacoub
背景与能力Context and Capabilities
烹饪插曲Cooking Interlude
研究方向与 AGIResearch Directions and AGI
烹饪终章Cooking Finale
多任务与模型能力Multitasking and Model Capabilities
被高估与低估的研究课题Overrated and Underrated Research Topics