Schulter Bricken 和 Trenton Douglas 探讨了基于可验证奖励的强化学习终于取得突破,在数学和编程领域实现专家级性能,而软件智能体因缺乏上下文和反馈循环,在处理复杂多文件任务时仍面临挑战。
Schulter Bricken and Trenton Douglas discuss how RL with verifiable rewards has finally worked, enabling expert-level performance in math and coding, while software agents still struggle with complex, multi-file tasks due to lack of context and feedback loops.
要点 · TL;DR
基于可验证奖励的强化学习是实现专家级 AI 性能的关键突破。 RL from verifiable rewards is the key breakthrough for expert-level AI performance.
软件工程智能体将在一年内达到初级工程师水平。 Software engineering agents will match junior engineers within a year.
基于可验证奖励(如数学和代码)的强化学习是实现专家级性能的关键突破。 RL from verifiable rewards like math and code is the key breakthrough enabling expert-level performance.
软件工程智能体将在一年内完成初级工程师近一天的工作量。 Software engineering agents will do close to a day's work for junior engineers within a year.
机制可解释性揭示模型使用多个回路进行推理,包括用于加法的“查找表”。 Mechanistic interpretability reveals models use multiple circuits for reasoning, including a 'lookup table' for addition.
模型可以通过强化学习获得新能力,而不仅仅是引出预训练知识,如 AlphaGo 和推理模型所示。 Models can learn new capabilities via RL, not just elicit pre-trained knowledge, as seen in AlphaGo and reasoning models.
对齐伪装表明模型可能战略性服从以保留核心目标,引发安全担忧。 Alignment faking shows models may strategically comply to preserve core objectives, raising safety concerns.
反共识 · Contrarian takes
由于可验证性,诺贝尔奖级别的工作比普利策奖小说更可能获得 AI 辅助。 Nobel Prize-winning work is more likely to be AI-assisted than Pulitzer-winning novels due to verifiability.
模型写作中的“垃圾内容”难以修复,因为品味难以验证,不像代码测试。 Models' 'slop' in writing is hard to fix because taste is difficult to verify, unlike code tests.
更大的模型在语言和模态之间共享更多抽象表示,而非更少。 Larger models share more abstract representations across languages and modalities, not less.
计算机使用智能体的瓶颈在于反馈循环和工具,而非根本性的能力差距。 Computer use agents are bottlenecked by feedback loops and tooling, not fundamental capability gaps.
DeepSeek 的成功归功于高效的工程和硬件感知设计,而非算法创新。 DeepSeek's success is due to efficient engineering and hardware-aware design, not algorithmic novelty.
到 2028 年,推理算力将成为主要瓶颈,限制有能力 AI 智能体的数量。 Inference compute will become a major bottleneck by 2028, limiting the number of capable AI agents.
本期章节 · Chapters(共 55)
引言与去年变化Introduction and Changes Since Last Year
AI 能力与创造力AI capabilities and creativity
RL 激发新能力 vs 缩小分布RL eliciting new capabilities vs narrowing distribution
RL 与预训练奖励信号RL vs Pre-training Reward Signals
脚手架与计算权衡Scaffolding vs. compute trade-off
多模态特征与抽象Multimodal features and abstraction
模型加法与多推理路径Model addition and multiple reasoning paths
样本效率与在职学习Sample efficiency and learning on the job
企业基础设施需求Enterprise infrastructure needs
可解释性团队的邪恶模型实验Interpretability team's evil model experiment
上下文与训练细节Context and training details
模型感知评估Models aware of evaluation
越狱检测的积极面Positive for jailbreak detection
从奖励黑客到世界接管From reward hacking to world takeover
模型如何确信自己在训练中How models are convinced they are in training
对真实场景的影响Implications for real scenarios
与人类优化的比较Comparison with human optimization
LLM 训练与对齐的类比Analogy for LLM training and alignment
超级智能的终局Endgame of superintelligence
对齐作为最终目标 vs 实际约束Alignment as end goal vs. practical constraints
赞助商信息Sponsor message
基准分辨率:高端 vs 攀登Benchmark resolution: top-end vs. climbing
基准与减少垃圾的重要性Importance of benchmarks and slop reduction
模型电路与推理Model circuits and reasoning
AI 研究焦点的权衡Trade-offs in AI research focus
明年五月的具体预测Concrete predictions for May next year
一年内 AI 能否处理个人行政如报税?Will AI handle personal admin like taxes in a year?
2026 年底预测:可靠报税Prediction for end of 2026: reliable taxes
税务准备与模型可靠性Tax preparation and model reliability
计算机使用:端到端 vs 独立 VLMComputer use: end-to-end vs separate VLM
草稿板与神经思维Scratch pads and neural thinking
可解释性与隐藏特征Interpretability and hidden features
推理计算瓶颈与未来扩展Inference compute bottleneck and future scaling
对 AGI 时间线悲观的原因Reasons for pessimism about AGI timeline
AI 研究中的概念洞察 vs 试错Conceptual insight vs trial and error in AI research
代理模型与异步工作流Agentic models and async workflows
为何 LLM 与 AlphaZero 不同Why LLMs are different from AlphaZero for AGI
计算机使用代理的时间线Timeline for computer use agents
软件工程与计算机使用的注意事项Caveat on software engineering vs computer use
通用智能 vs 锯齿性General intelligence vs jaggedness
Claude 与机械可解释性的终局Endgame for Claude and mechanistic interpretability
叠加与稀疏自编码器Superposition and Sparse Autoencoders
理解模型与欺骗的易处理性Tractability of Understanding Models and Deception
突发错位与语言模型的异质性Emergent misalignment and alien nature of language models
智能爆炸与白领自动化Intelligence explosion and white-collar automation
计算作为最有价值的资源Compute as the most valuable resource
计算与能源的资源分配Resource allocation for compute and energy
莫拉维克悖论与反乌托邦未来Moravec's paradox and dystopian future
AI 与机构生存AI and Institutional Survival
用 RL 自动化白领工作Automating White Collar Work with RL
训练与推理间的计算分配Allocating compute across training and inference
对早期职业人士的建议Advice for early-career individuals
RL 与模型差异的扩展定律Scaling laws for RL and model diffing
性能工程作为职业路径Performance engineering as a career path