A discussion on defining intelligence for AGI, using the ability to write a thought-provoking novel as a benchmark, and evaluating current measures like perplexity and benchmarks.
要点 · TL;DR
写小说比基准测试更能衡量 AGI。 Writing a novel is a better AGI test than benchmarks.
推理计算是扩展模型智能的关键。 Inference compute is key to scaling model intelligence.
学术界需要更多 GPU 才能与产业竞争。 Academia needs more GPUs to compete with industry.
核心观点 · Key points
智能很难定义;写一本发人深省的小说是 AGI 的一个好代理指标。 Intelligence is hard to define; writing a thought-provoking novel is a good proxy for AGI.
推理算力对模型能力越来越重要且被低估。 Inference compute is increasingly important and underestimated for model capability.
学术界需要更多算力才能竞争;大学应大力投资 GPU。 Academia needs more compute to compete; universities should invest heavily in GPUs.
评估是学术界无需大量算力就能产生高影响的研究领域。 Evals are a viable high-impact research area for academia without massive compute.
AI 很快将端到端撰写和评审论文,提升同行评审质量。 AI will soon write and review papers end-to-end, improving peer review quality.
反共识 · Contrarian takes
扩展推理算力可以从同一基础模型获得任意更智能的模型。 Scaling inference compute can yield arbitrarily more intelligent models from the same base.
通过强化学习完全去除人类先验理论上可行但不实用。 Removing human priors entirely via RL is theoretically possible but not practical.
当前用单一数字评估模型已过时;推理算力应放在 x 轴上。 Current model evaluation with a single number is outdated; inference compute should be on the x-axis.
在多样化数据上预训练使模型能发现新算法如归并排序。 Pre-training on diverse data enables models to discover novel algorithms like merge sort.
开源模型很有价值;闭源模型留下的空白应被填补。 Open source models are valuable; the gap left by closed models should be filled.
本期章节 · Chapters(共 12)
定义智能Defining Intelligence
智能的定义与衡量Defining and Measuring Intelligence
伊利亚主义 vs 侏儒主义与推理计算Ilyaism vs. Gnomeism and Inference Compute
搜索中的启发式与人类先验Heuristics and Human Priors in Search
LLM 的生物合理性Biological Plausibility of LLMs
AI 中的简单算法与先验Simple Algorithms and Priors in AI
关于 AI 与推理计算的反主流观点Contrarian View on AI and Inference Compute
学术界与工业界的算力差距Academia vs Industry Compute Gap
学术界与工业界对基准测试的影响Impact of Academia vs Industry on Benchmarks
开源模型与 DeepSeekOpen Source Models and DeepSeek
个人成长与 Libratus 竞赛Personal Growth and the Libratus Competition