斯坦福大学教授 Yajin Choy 探讨了强化学习、微调等 AI 方法归根结底都依赖于数据,并指出当前方法可能无法解决如治愈癌症等未知问题。
Stanford professor Yajin Choy discusses how AI methods like reinforcement learning and fine-tuning all boil down to data, and why current approaches may not lead to solving unknown problems like curing cancer.
要点 · TL;DR
数据是根本瓶颈,所有 AI 方法最终都归结为数据。 Data is the fundamental bottleneck; all AI methods reduce to data.
当前 AI 只是内插人类知识,而非超越。 Current AI interpolates human knowledge, not transcends it.
溯因推理是超越归纳和演绎的科学发现关键。 Abductive reasoning is key to scientific discovery beyond induction and deduction.
核心观点 · Key points
数据是根本瓶颈;所有方法最终都归结为数据。 Data is the fundamental bottleneck; all methods reduce to data.
当前 AI 方法只是对人类知识进行插值,而非超越。 Current AI methods interpolate human knowledge, not transcend it.
强化学习需要精心设计才能超越监督微调。 Reinforcement learning requires careful engineering to outperform supervised fine-tuning.
开放研究对 AI 民主化和造福全人类至关重要。 Open research is crucial for democratizing AI and benefiting all humanity.
溯因推理是超越归纳和演绎的科学发现关键。 Abductive reasoning is key to scientific discovery beyond induction and deduction.
跨机构合作是抗衡封闭前沿实验室的必要条件。 Collaboration across institutions is needed to compete with closed frontier labs.
反共识 · Contrarian takes
RLVR 可能降低 pass@k 性能,即使最佳样本提升。 RLVR can decrease pass@k performance even as best-of-n improves.
随机或错误奖励仍可提升某些模型的 RL 性能。 Random or incorrect rewards can still improve RL performance on some models.
中国公司现在主导开源贡献,吸引人才回流。 Chinese companies now lead open-source contributions, attracting talent back to China.
人类知识并非全部真理;AI 需要探索未知知识。 Human knowledge is not the universe of truth; AI needs to explore unknown knowledge.
证伪比提出假设对科学进步更关键。 Falsification is more critical than hypothesis generation for scientific progress.
如果我们不行动,AI 可能最终服务 AI 或人类服务 AI。 AI could end up serving AI or humans serving AI if we don't act.
本期章节 · Chapters(共 10)
引言与研究焦点Introduction and Research Focus
主题演讲:RL 与数据挑战Keynote Themes: Challenges in RL and Data