Skill Issue:代码智能体与自动化研究
Skill issue: code agents and autoresearch
安德烈·卡帕西 Andrej Karpathy · No Priors 播客 · 2026-03-20 · 约 67 分钟 · 原视频 ↗
打开互动全文版(中英对照 + 朗读 + 问答)→
本期速览 · Overview
代码智能体还差在哪,以及通往自主研究那条「绕圈」的路。
Where coding agents fall short, and the loopy road to autonomous research.
要点 · TL;DR
- 嘉宾自去年 12 月起未写一行代码,全权委托给智能体。
Speaker hasn't typed code since Dec, delegates all to agents. - 瓶颈从打字速度转向指挥智能体的技能。
Bottleneck shifts from typing speed to skill in directing agents. - 自动研究移除人类环节,能发现人类遗漏的改进。
Auto-research removes humans to find improvements they miss.
核心观点 · Key points
- 自去年 12 月以来,我没写过一行代码;所有工作都委托给智能体。
Since December, I haven't typed a line of code; I delegate everything to agents. - 瓶颈不再是打字速度,而是你指挥智能体的技能。
The bottleneck is no longer typing speed but your skill in directing agents. - 自动研究证明,将人类移出循环可以发现人类忽略的改进。
Auto research shows that removing humans from the loop can find improvements humans miss. - 大语言模型智能参差不齐:在可验证任务上出色,但还在讲多年前的烂笑话。
LLMs have jagged intelligence: brilliant at verifiable tasks, but still tell the same bad joke from years ago. - 行业应重构为智能体优先的工具,用 API 取代应用,智能体作为粘合剂。
The industry should refactor for agent-first tools, where APIs replace apps and agents are the glue. - 开源模型落后前沿几个月,但这有利于去中心化。
Open-source models lag behind frontier by months, but that's healthy for decentralization.
反共识 · Contrarian takes
- 自去年 12 月以来我没写过一行代码,全交给智能体了。
I haven't typed a line of code since December; I only delegate to agents. - 大语言模型就像天才博士生和十岁小孩的结合体。
LLMs are like a brilliant PhD student combined with a 10-year-old. - 模型并未跨领域泛化智能;讲笑话的质量毫无提升。
Models are not generalizing intelligence across domains; joke quality hasn't improved. - 前沿实验室正在积极自动化淘汰自己的研究人员。
The frontier labs are actively automating away their own researchers. - 我们应该为智能体写 Markdown,而不是为人类写 HTML 文档。
We should write markdown for agents, not HTML docs for humans. - 互联网上不可信的智能体集群可能超越前沿实验室。
An untrusted swarm of internet agents could outcompete frontier labs.
本期章节 · Chapters(共 18)
- 引言与 AI 精神病 Introduction and AI psychosis
- 新工作流与技能问题 The new workflow and skill issue
- 令牌吞吐量与技能问题 Token throughput and skill issue
- Claude 的创新与个性 Claude's innovations and personality
- 家养小精灵 Claude Dobby the house elf Claude
- 自然语言家居自动化 Home automation with natural language
- 通过自主代理增加杠杆 Increasing leverage through autonomous agents
- LLM 精神病与自动研究 LLM Psychosis and Auto-Research
- 锯齿状与模型局限 Jaggedness and Model Limitations
- 不可信工人池与验证 Untrusted Pool of Workers and Verification
- 计算力作为新货币 Compute as the New Currency
- 就业数据分析与 AI 影响 Jobs Data Analysis and AI Impact
- 数字与物理世界加速 Digital vs Physical World Acceleration
- 前沿实验室内外的影响与对齐 Impact and alignment in frontier labs vs. outside
- 开源接近前沿与可持续性 Open source proximity to frontier and sustainability
- 开源与前沿模型 Open-source vs frontier models
- 数字与物理世界机遇 Digital vs Physical World Opportunities
- MicroGPT 与 LLM 简化 MicroGPT and boiling down LLMs
阅读全文双语转录 →