Lex Fridman 与 Sebastian Raschka 和 Nathan Lambert 探讨 AI 最新突破,涵盖 DeepSeek 时刻、开源模型以及中美实验室的竞争格局。
Lex Fridman discusses the latest AI breakthroughs with Sebastian Raschka and Nathan Lambert, covering the DeepSeek moment, open-weight models, and the competitive landscape between US and Chinese labs.
要点 · TL;DR
缩放定律仍然有效,但重点从预训练转向后训练和推理缩放。 Scaling laws persist but shift from pre-training to post-training and inference scaling.
RLVR 是后训练中实现推理和工具使用的关键突破。 RLVR is the key post-training breakthrough for reasoning and tool use.
中国开源模型因宽松许可而流行,但美国模型在质量上仍领先。 Chinese open-weight models gain traction via permissive licenses, but US models lead in quality.
核心观点 · Key points
缩放定律在预训练中仍然成立,但低垂果实已被摘取;当前进步主要来自后训练和推理缩放。 Scaling laws still hold for pre-training, but low-hanging fruit is gone; gains now come from post-training and inference scaling.
基于可验证奖励的强化学习(RLVR)是后训练中最大的突破,实现了推理和工具使用。 RLVR (reinforcement learning with verifiable rewards) is the biggest post-training breakthrough, enabling reasoning and tool use.
DeepSeek 等中国开放权重模型因宽松许可和强劲性能而受欢迎,但美国模型在质量上仍领先。 Chinese open-weight models like DeepSeek are popular due to permissive licenses and strong performance, but US models still lead in quality.
Transformer 架构仍占主导;创新在于注意力机制调整(MLA、GQA)和混合专家模型。 Transformer architecture remains dominant; innovations are in attention tweaks (MLA, GQA) and mixture of experts.
数据质量和筛选对预训练至关重要;合成数据和精心混合可提高效率。 Data quality and curation are critical for pre-training; synthetic data and careful mixing improve efficiency.
反共识 · Contrarian takes
预训练缩放并未消亡,只是目前性价比低于推理缩放和强化学习。 Pre-training scaling is not dead; it's just less cost-effective than inference scaling and RL for now.
中国公司发布开放权重模型是为了扩大影响力,而非单纯商业目的;这一趋势可能持续数年。 Chinese companies release open-weight models to gain influence, not just for business; they may continue for years.
Qwen 模型上的 RLVR 提升可能部分源于数据污染,而非纯粹的能力解锁。 RLVR gains on Qwen models may be partly due to data contamination, not pure skill unlocking.
LLM 生成的摘要常丢失语气和核心洞见;人类撰写的内容保留独特价值。 LLM summaries often lose voice and core insights; human-written content retains unique value.
凡事依赖 LLM 可能阻碍学习;刻意挣扎和离线学习对成为专家仍然必不可少。 Using LLMs for everything may hinder learning; deliberate struggle and offline study are still essential for expertise.
本期章节 · Chapters(共 101)
引言与嘉宾Introduction and Guests
DeepSeek 时刻与国际竞争DeepSeek Moment and International Competition
中国开源模型的可持续性Sustainability of Open-Weight Models from China
中国开源模型与市场激励Chinese open-weight models and market incentives
DeepSeek 的定位与竞争DeepSeek's position and competition
中国企业的不同激励Different incentives among Chinese companies
用户习惯与多订阅User habits and multiple subscriptions
2025-2026 赢家预测Predictions for 2025 and 2026 winners
模型使用偏好Model usage preferences
GPT-5.2 长上下文与中国模型GPT-5.2 long context and Chinese models
用 LLM 编程:工具与经验Programming with LLMs: tools and experiences
书籍与从零学习Books and learning from scratch
用 LLM 的阅读习惯Reading habits with LLMs
开源 LLM 格局Open LLM models landscape
工具使用及其意义Tool Use and Its Significance
开源模型爆发原因Reasons for Open Model Explosion
开源模型的有趣想法Interesting Ideas from Open Models
KV 缓存与注意力机制调整KV Cache and Attention Mechanism Tweaks
Transformer 架构概览Transformer Architecture Overview
从 GPT-2 到如今的演进Evolution from GPT-2 to Today
后训练与系统级改进Post-training focus and system-level improvements
预训练、后训练与推理的缩放定律Scaling laws across pre-training, post-training, and inference
预训练缩放与成本Pre-training scaling and cost
后训练与发布周期Post-training and release cycles
看好所有缩放方向Bullish on all scaling directions
预训练、中训练与后训练定义Pre-training, mid-training, and post-training definitions
预训练的合成数据与数据提取Synthetic data and data extraction for pre-training
数据质量与算力Data Quality vs. Compute
提升数据质量Improving Data Quality
意外的高质量数据源Unexpected High-Quality Data Sources
训练数据的保密与许可Secrecy and Licensing of Training Data
数据护城河与领域缩放Data as Moat and Domain-Specific Scaling
LLM 摘要中的声音与洞察Voice and Insight in LLM Summaries
为不同用户设计 AIDesigning AI for diverse users
科技巨头与 AI 声誉Big Tech and AI reputation
在 AI 中寻找自主性Finding agency in AI
自主性与忽视 AIAgency vs ignoring AI
开发者对 AI 的享受度调查Survey on developer enjoyment with AI
平衡享受与 AI 使用Balancing enjoyment and AI use
AI 作为结对程序员AI as a pair programmer
延迟满足与 AIDelayed gratification and AI
AI 学习的黄金区间Goldilocks zone of learning with AI
后训练与 RLVRPost-training and RLVR
推理缩放与思维链Inference Scaling and Chain-of-Thought
后训练配方与 RLVRPost-Training Recipe and RLVR
中训练与推理轨迹Mid-training and reasoning traces
RLHF 作为收尾RLHF as finishing touch
RLVR 的算力需求Compute requirements for RLVR
RLVR 与 RLHF 缩放RLVR vs RLHF Scaling
教育与学习推荐Education and Learning Recommendations
从 Transformers 库学习Learning from Transformers Library
学习路径与研究机会Learning Path and Research Opportunities
后训练主题概览Overview of Post-Training Topics
培养研究品味Developing research taste
教育的短暂数字窗口The brief digital window in education
角色训练研究的算力需求Compute requirements for character training research
权衡:新颖想法与职业路径Trade-offs: novel ideas vs. practical career paths
职业建议:学术界与工业实验室Career advice: academia vs. industry labs
学术界与工业界Academia vs. Industry
硅谷泡沫与炒作Silicon Valley bubble and hype
文生图缩放与替代架构Text-to-image scaling and alternative architectures
文本扩散模型与自回归模型Text Diffusion Models vs. Auto-Regressive Models
工具使用的未来Future of Tool Use
工具使用中的开源与闭源模型Open vs. Closed Models in Tool Use
持续学习的定义与重要性Open vs Closed Models and Tool Use
持续学习与上下文学习Continual Learning Definition and Importance
记忆机制Continual Learning vs In-Context Learning
长上下文的创新Memory Mechanisms
世界模型与 LLMInnovations in Long Context
机器人技术与缩放World Models and LLMs
机器人挑战与安全Robotics and Scaling
AGI 与 ASI 时间线Robotics challenges and safety
AI 与软件开发未来Timelines to AGI and ASI
用 LLM 的软件工程AI and Software Development Future
经济影响与工具使用Software engineering with LLMs
需要新想法Economic Impact and Tool Use
能力平台期与经济影响Need for New Ideas
LLM 用于学习与个性化建议Plateauing Capabilities vs. Economic Impact
AI 初创公司的整合LLMs for learning vs. personalized advice
IPO 与 AI 公司未来Consolidation in AI startups
API 市场竞争IPOs and future of AI companies
Meta 路径与 Llama 未来API Market Competition
Llama 演进与内部问题Meta's Path and Llama's Future
扎克伯格角色与开源辩论Llama's Evolution and Internal Issues
社区反弹及其影响Mark Zuckerberg's Role and Open Source Debate
美国填补 Llama 空白Community Backlash and Its Impact
Adam 项目:起源与使命US Filling the Gap Left by Llama
政府与行业支持The Adam Project: Origin and Mission
AI 行动计划与开源Government and Industry Support
教育与人才培养The AI Action Plan and Open Source
开源模型与全球访问Education and Talent Development
集中化与开放性Open Source Models and Global Access
英伟达主导与硬件趋势Centralization vs. Openness
英伟达创新与黄仁勋角色Nvidia's Dominance and Hardware Trends
网络与算力缩放Nvidia's innovation and Jensen's role
神经网络作为突破Networking and Compute Scaling
未来世界:机器人与设备Neural Networks as Breakthrough
失业的社会影响Future World: Robots and Devices
对未来的希望Social impact of job loss
后奇点场景中的人与 AIHope for the future
Human vs AI in a post-singularity scenarioHuman vs AI in a post-singularity scenario