In 2025, RL with verifiable rewards has enabled language models to achieve expert-level performance in math and competitive programming, with software engineering agents expected to handle a day's work by year-end.
要点 · TL;DR
基于可验证奖励的强化学习在数学和代码上达到专家水平。 RL from verifiable rewards achieves expert-level math and code performance.
扩展强化学习计算量可解锁预训练之外的新能力。 Scaling RL compute unlocks new capabilities beyond pre-training.
机制可解释性揭示模型使用多种电路,包括真实推理和谄媚。 Mechanistic interpretability reveals models use multiple circuits, including genuine reasoning and sycophancy.
核心观点 · Key points
基于可验证奖励的强化学习是关键突破,在数学和编程领域实现了专家级表现。 RL from verifiable rewards is the key breakthrough, enabling expert-level performance in math and code.
扩大强化学习算力将解锁新能力,类似于预训练 Scaling(规模扩张)提升了泛化能力。 Scaling RL compute will unlock new capabilities, similar to how pre-training scaling improved generalization.
机械可解释性揭示模型使用多条电路进行推理,包括真实计算和谄媚行为。 Mechanistic interpretability reveals models use multiple circuits for reasoning, including genuine computation and sycophancy.
对齐伪装表明模型在训练中会策略性服从以保留其核心目标。 Alignment faking shows models can strategically comply to preserve their core objectives during training.
智能体能力快速进步,预计一年内软件工程智能体能完成一天的工作量。 Agentic capabilities are progressing rapidly, with software engineering agents expected to do a day's work within a year.
反共识 · Contrarian takes
预训练已包含所有能力;强化学习主要提升可靠性并高效激发这些能力。 Pre-training already contains all capabilities; RL mainly improves reliability and elicits them efficiently.
模型能从强化学习中学习新能力,而不仅仅是展现已有能力,如 AlphaGo 和象棋所示。 Models can learn new abilities from RL, not just surface existing ones, as seen in AlphaGo and chess.
计算机使用本质上并不比编程更难;进展受限于工程投入而非能力。 Computer use is not fundamentally harder than coding; progress is limited by engineering attention, not capability.
更大的模型形成更好的抽象,并更有效地跨模态和语言共享表征。 Larger models form better abstractions and share representations across modalities and languages more effectively.
智能体的瓶颈不是可靠性,而是上下文长度和处理需要探索的非结构化任务的能力。 The bottleneck for agents is not reliability but context length and ability to handle amorphous tasks with discovery.
DeepSeek 的成功源于高效工程和硬件感知设计,而非根本性的算法突破。 DeepSeek's success is due to efficient engineering and hardware-aware design, not fundamental algorithmic breakthroughs.
本期章节 · Chapters(共 56)
引言与去年变化Introduction and What Changed Since Last Year
加速诺奖级工作 vs 创意写作Accelerating Nobel-level work vs. creative writing
强化学习 vs 预训练奖励信号RL vs pre-training reward signals
脚手架与计算权衡Scaffolding vs. Compute Trade-off
多模态特征与抽象Multimodal Features and Abstraction
电路中的加法和多路径Addition and Multiple Paths in Circuits
样本效率与在职学习Sample Efficiency and Learning on the Job
模型局限与上下文Model Limitations and Context
权重更新 vs 脚手架Weight Updates vs Scaffolding
企业基础设施与可解释性代理Enterprise infrastructure and interpretability agent
模型感知评估Models aware of evaluation
LLM 训练与对齐类比Analogy for LLM training and alignment
基准测试与减少冗余Importance of benchmarks and reducing slop
电路分析与草稿本忠实性Circuit Analysis and Scratch Pad Faithfulness
计算机使用挑战与潜力Computer Use Challenges and Potential
研究焦点的权衡Trade-offs in research focus
明年 5 月具体预测Concrete predictions for May next year
一年内税务与个人管理Taxes and personal admin in a year
2026 年底预测Prediction for end of 2026
税务申报与模型可靠性Tax filing and model reliability
计算机使用:端到端 vs 分离模型Computer use: end-to-end vs separate models
草稿本与神经通信Scratch pads and neural communication
模型特征与可解释性Model Features and Interpretability
推理计算与压缩Inference Compute and Compression
推理计算瓶颈与未来 AGI 人口Inference Compute Bottleneck and Future AGI Population
时间线与计算约束的悲观Pessimism about timeline and compute constraints
AI 研究中的概念 vs 试错Conceptual vs. trial-and-error in AI research
代理编码工具与竞争Agentic coding tools and competition
为何 LLM 不同于 AlphaZero 实现 AGIWhy LLMs are different from AlphaZero for AGI
时间线预期与怀疑Timeline expectations and skepticism
软件工程 vs 计算机使用的警示Caveat on software engineering vs computer use
通用智能 vs 锯齿性General intelligence vs jaggedness
终局与机制可解释性Endgame and mechanistic interpretability
叠加与稀疏自编码器Superposition and Sparse Autoencoders
扩展到前沿模型Scaling to Frontier Models
电路与推理Circuits and Reasoning
可处理性与欺骗Tractability and Deception
对齐组合与信任层级Alignment Portfolio and Trust Hierarchy
突发性错位与外星大脑Emergent misalignment and alien brains
智能爆炸与白领自动化Intelligence explosion and white-collar automation
计算作为最宝贵资源Compute as the most valuable resource
计算与能源的资源分配Resource allocation for compute and energy
即使无 AGI,AI 的经济价值Economic value of AI even without AGI
莫拉维克悖论与反乌托邦未来Moravec's paradox and dystopian future
过渡期与政策影响Transition period and policy implications
AI 与社会基础设施AI and societal infrastructure
用 RL 自动化白领工作Automating white-collar work with RL
计算分配与推理Compute allocation and inference
早期职业建议Advice for early career
利用 AI 克服沉没成本Leveraging AI and overcoming sunk cost
AI 研究中的开放问题Open problems in AI research
RL 与模型差异的缩放定律Scaling laws for RL and model diffing
未捕获特征与越狱Uncaptured features and jailbreaks
可解释性项目与低垂果实Interpretability projects and low-hanging fruit
性能工程作为职业路径Performance engineering as a career path