两位 AI 研究人员讨论长上下文窗口如何让模型在上下文中学习语言的能力超越人类,使其在信息整合方面达到超人类水平。
Two AI researchers discuss how long context windows enable models to learn languages in-context better than humans, making them superhuman in information integration.
要点 · TL;DR
长上下文窗口通过解决接入问题大幅提升 AI 智能。 Long context windows dramatically boost AI intelligence by solving the onboarding problem.
AI 智能体受限于可靠性,而非长程任务表现。 AI agents are held back by reliability, not long-horizon task performance.
核心观点 · Key points
长上下文窗口被低估了;它们解决了上手问题,并显著提升了模型智能。 Long context windows are underhyped; they solve the onboarding problem and boost model intelligence dramatically.
上下文学习类似于梯度下降;注意力机制可被视为对上下文数据执行梯度下降。 In-context learning resembles gradient descent; attention can be viewed as performing gradient descent on in-context data.
AI 智能体未能起飞主要归因于可靠性问题,而非长周期任务表现。 AI agents haven't taken off mainly due to reliability issues, not long-horizon task performance.
大多数智能是模式匹配;推理源于记忆中的层级关联。 Most intelligence is pattern matching; reasoning emerges from hierarchical associations in memory.
AI 研究进展受限于算力和品味,而不仅仅是工程努力。 AI research progress is bottlenecked by compute and taste, not just engineering effort.
反共识 · Contrarian takes
在典型上下文长度下,二次注意力成本通常被密集 Transformer 中的 MLP 成本主导。 Clauderatic attention cost is often dominated by MLP cost in dense transformers for typical context lengths.
对于互联网数据的复杂性,模型严重欠参数化,导致叠加现象。 Models are dramatically underparameterized for the complexity of internet data, leading to superposition.
思维链推理可能不可信;模型可能产生误导性的推理过程。 Chain-of-thought reasoning may not be trustworthy; models can produce misleading rationales.
语言进化得易于儿童学习,使其成为 LLM 的高效表示形式。 Language evolved to be learnable by children, making it an efficient representation for LLMs.
随着长上下文和自适应计算实现动态专业化,微调可能变得过时。 Fine-tuning may become obsolete as long context and adaptive compute allow dynamic specialization.
本期章节 · Chapters(共 58)
引言与背景Introduction and Context
可靠性是智能体的关键Reliability as key for agents
进步梯度与样本效率Progress Gradient and Sample Efficiency
规模扩展:模型大小 vs 每次调用计算量Scaling Up: Model Size vs. Compute per Call
存储原始信息 vs 推理Storing Raw Information vs. Reasoning
残差流与大脑类比Residual Stream and Brain Analogy
注意力与小脑回路类比Attention and Cerebellar Circuit Analogy
推理 vs 模式匹配Reasoning vs Pattern Matching
记忆与想象的联系Memory and Imagination Link
福尔摩斯与演绎推理Sherlock Holmes and Deductive Reasoning
LLM 中的联想与推理Associations and reasoning in LLMs
智能爆炸机制Intelligence explosion mechanism
日常研究流程Daily research workflow
研究优先级与工程技能Research prioritization and engineering skill
进化优化与类脑方案Evolutionary optimization and brain-like solutions
计算与品味成为瓶颈Compute and taste as bottlenecks
计算分配:研究 vs 扩展Compute allocation and research vs scaling
AI 加速 AI 研究AI speeding up AI research
合成数据与推理轨迹Synthetic data and reasoning traces
机器学习进步的进化视角Evolutionary perspective on ML progress
智能爆炸的窄窗口框架Narrow window framing of intelligence explosion
计算扩展与大脑比较Compute scaling and brain comparison
过参数化 vs 欠参数化与蒸馏Overparameterization vs underparameterization and distillation
蒸馏 vs 从头训练Distillation vs. Training from Scratch
思维链作为自适应计算Chain of Thought as Adaptive Compute
思维链中的隐写与可解释性Steganography and Interpretability in Chain of Thought
思维链可解释性Chain of Thought Interpretability
模型专业化与微调的未来Future of Model Specialization and Fine-Tuning
语言进化与 LLM 成功Language Evolution and LLM Success
多模态学习与正迁移Multimodal Learning and Positive Transfer
代码训练与推理改进Code Training and Reasoning Improvement
模型推理的证据Evidence of reasoning in models
职业反思与快速进步Career reflections and rapid progress
能动性与主动性Agency and Taking Initiative
招聘故事与背景Hiring Story and Background
背景与通往谷歌之路Background and Path to Google
指导与系统算法理解Mentorship and Systems-Algorithms Understanding
与谢尔盖·布林共事与主动性Working with Sergey Brin and Being Agentic
谷歌内部人脉的影响Impact of personal connections at Google
招聘能动性与世界级工作Hiring for agency and world-class work
职业道德与高杠杆工作Work Ethic and High-Leverage Work
大脑组织与特征空间Brain Organization and Feature Space
定义特征与特征分裂Defining features and feature splitting
推理回路与模型可解释性Reasoning circuits and model interpretability
识别恶意特征与跨模型普遍性Identifying malicious features and cross-model universality
特征普遍性与对齐乐观Feature Universality and Optimism about Alignment
特征分裂与可扩展性Feature Splitting and Scalability
心智理论与欺骗特征Theory of Mind and Deception Features
视觉 Transformer 中的专业化与分支专业化Specialization in Vision Transformers and Branch Specialization
绑定操作与注意力Binding operations and attention
超级智能与 GPT-7 部署Superintelligence and GPT-7 deployment
深层中的抽象特征Abstract features in deep layers
角色锁定与人类心理学Persona lock-in and human psychology
AI 部署中的开放性与反馈Openness and feedback in AI deployment
谷歌的公交因素与领域动态Bus factor at Google and field dynamics
关注谁与内部叙事Who to pay attention to and internal narratives