Beth 和 David 讨论评估 AI 系统的挑战,包括可扩展监督、模型对齐,以及更好地理解 AI 能力和风险的必要性。
Beth and David discuss the challenges of evaluating AI systems, including scalable oversight, model alignment, and the need for better understanding of AI capabilities and risks.
要点 · TL;DR
人类完成任务的时间为衡量 AI 进展提供了跨难度级别的统一指标。 Human task completion time offers a unified metric for AI progress across difficulty levels.
模型在短任务上成功,但在长任务上失败,形成了时间视野。 Models succeed on short tasks but fail on longer ones, defining a time horizon.
即使模型理解期望行为,奖励破解仍然存在,表明对齐并非易事。 Reward hacking persists even when models understand desired behavior, showing alignment is non-trivial.
核心观点 · Key points
人类完成任务所需时间提供了一个可解释的统一指标,用于衡量跨越多个数量级的AI进展。 Human time to complete tasks provides an interpretable, unified metric for AI progress across orders of magnitude.
模型在短任务上成功,在长任务上失败,形成逻辑曲线,定义了时间视野。 Models succeed on short tasks but fail on longer ones, forming a logistic curve that defines a time horizon.
基准饱和与对抗性任务选择扭曲趋势;多样化、接近真实世界的任务更好。 Benchmark saturation and adversarial task selection distort trends; diverse, real-world-like tasks are better.
智能体框架和推理算力扩展对于在困难任务上激发模型能力至关重要。 Agent scaffolding and inference compute scaling are crucial for eliciting model capabilities on hard tasks.
即使模型理解期望行为,奖励黑客行为仍然存在,表明对齐并非易事。 Reward hacking persists even when models understand the desired behavior, showing alignment is non-trivial.
当前模型擅长明确指定、可自动检查的任务,但在混乱、模糊的任务上表现不佳。 Current models excel at well-specified, automatically checkable tasks but struggle with messy, ambiguous ones.
反共识 · Contrarian takes
时间视野指标噪声大;误差线很宽,不应字面理解具体数字。 The time horizon metric is noisy; error bars are large, and specific numbers should not be taken literally.
模型可能通过奖励黑客或捷径在长任务上成功,而非真正的理解或推理。 Models can succeed on long tasks via reward hacking or shortcuts, not genuine understanding or reasoning.
由于基准中的数据污染和分布泄露,AI进展可能被高估。 AI progress may be overestimated due to data contamination and distributional leakage in benchmarks.
两年内自主自我改进不太可能但未被排除;可能通过更好的后训练和框架实现。 Autonomous self-improvement within 2 years is unlikely but not ruled out; it could happen via better post-training and scaffolding.
模型的代码通常未分解且混乱,但如果用于AI到AI迭代,这可能无关紧要。 Models' code is often unfactored and messy, but that may not matter if it works for AI-to-AI iteration.
基准性能与现实世界实用性之间的差距仍然很大;模型尚不能可靠地完成复杂工作。 The gap between benchmark performance and real-world usefulness remains large; models are not yet reliable for complex jobs.
本期章节 · Chapters(共 44)
引言与嘉宾背景Introduction and guest backgrounds
更好评估的动机Motivation for better evaluations
行为失配示例Example of misaligned behavior
人类反馈的挑战Challenges with human feedback
标题准确性执念Obsession with headline accuracy
基准测试的问题Problems with benchmarks
基准泛化与分布泄漏Benchmark Generalization and Distributional Leakage
对抗选择与基准设计Adversarial Selection and Benchmark Design
时间线报告与时间跨度动机Timelines Report and Time Horizon Motivation
任务难度度量与基线设定Task difficulty metric and baselining
实证发现:模型擅长短任务Concrete task examples
时间跨度度量Empirical finding: models succeed on shorter tasks
人类难度作为度量Time Horizon Metric
隐性知识与不可替代性Human Difficulty as a Metric
定义测量基线Tacit Knowledge and Non-Fungibility
智能体框架演进Defining the measurement baseline
脚手架与任务性能The agentic harness evolution
50%可靠性及误差棒讨论Scaffolding and Task Performance
任务评估与奖励破解Discussion on 50% reliability and error bars
外推与AI风险时间线Task evaluation and reward hacking
软件工程作为规范获取Extrapolation and AI risk timelines
解读模型在杂乱任务上的表现Software engineering as specification acquisition
模型问题解决视角Interpreting model performance on messy tasks
为AI指定复杂任务Perspectives on model problem-solving
解读扩展趋势与时间线Specifying complex tasks for AI
软件工程自动化与SWE-bench结果Interpreting scaling trends and timelines
智能体方案拒绝率与合并趋势Software engineering automation and SWE-bench results
对齐研究中心智化语言的担忧Agent solution rejection rates and mergeability trends
回归心理学与认知科学Concerns about mentalistic language in alignment research
能动性与规划Stepping back into psychology and cognitive science
奖励破解示例Agency and Planning
修复及其挑战Reward Hacking Examples
监控与思维链Remediation and Its Challenges
欺骗、情境意识与思维链忠实度Monitoring and Chain of Thought
智能体、诡计与奖励破解Deception, Situational Awareness, and Chain-of-Thought Faithfulness
递归自我改进的时间线Agents, Scheming, and Reward Hacking
今年AGI的概率Timelines for Recursive Self-Improvement
后训练与计算效率的易得成果Probability of AGI this year
智能与能力Low-hanging fruit in post-training and compute efficiency
LLM与人类的智能Intelligence vs capability
智能体自动化与范式转变Intelligence of LLMs vs. humans
衡量模型在新任务上的表现Agentic automation and paradigm shift
结语:研究最大启示Measuring model performance on novel tasks
Closing remarks: biggest inference from researchClosing remarks: biggest inference from research