2024 年图灵奖得主 Richard Sutton 和 Andrew Barto 探讨他们在强化学习方面的开创性工作,即智能源于试错和奖励的理念。
2024 Turing Award winners Richard Sutton and Andrew Barto discuss their foundational work in reinforcement learning, the idea that intelligence emerges from trial, error, and reward.
要点 · TL;DR
强化学习通过试错学习,而非来自教师。 Reinforcement learning learns from trial and error, not a teacher.
时序差分学习利用预测随时间的变化作为误差信号。 Temporal difference learning uses prediction changes over time as error signals.
年轻研究者应追随热情而非潮流,以做出有影响力的工作。 Young researchers should follow passion, not fashion, for impactful work.
核心观点 · Key points
强化学习是从试错中学习,而非从教师那里学习。 Reinforcement learning is learning from trial and error, not from a teacher.
时序差分学习利用预测随时间的变化作为误差信号。 Temporal difference learning uses changes in predictions over time as error signals.
理解心智是关键目标;AI 有助于重建其基本功能。 Understanding the mind is a key goal; AI helps recreate its basic functions.
年轻研究者应追随热情,而非潮流。 Young researchers should follow their passion, not fashion.
当前 AI 系统如大语言模型在部署后不从经验中学习。 Current AI systems like LLMs do not learn from experience after deployment.
反共识 · Contrarian takes
强化学习尽管显而易见,却在 AI 领域被忽视了几十年。 Reinforcement learning was neglected in AI for decades despite being obvious.
神经元可能是从后果中学习的目标导向智能体,而非仅遵循赫布规则。 Neurons may be goal-directed agents that learn from consequences, not just Hebbian.
误差校正(监督学习)常被误认为强化学习。 Error correction (supervised learning) is often confused with reinforcement learning.
AI 中的潮流,如神经网络,已多次兴衰。 Fashion in AI, like neural networks, has risen and fallen multiple times.
心理学中的认知革命扼杀了早期的强化学习努力。 The cognitive revolution in psychology extinguished early reinforcement learning efforts.
基于人类反馈的强化学习并非真正的从世界经验中学习。 Reinforcement learning from human feedback is not true learning from world experience.
本期章节 · Chapters(共 10)
开场Introduction
早期生涯与重新发现强化学习Early career and rediscovering reinforcement learning
跨学科背景与神经科学类比Interdisciplinary Background and Neuroscience Parallels
时序差分学习详解Explanation of Temporal Difference Learning
衡量学习效果Measuring Learning
对工程实践与社会效益的担忧Concerns about engineering practices and societal benefit
给年轻研究者的建议:避开潮流,追随热情Advice for young researchers: avoid fashion, follow passion
强化学习的历史与潮流周期History of reinforcement learning and fashion cycles
关于追随热情与影响力的最后思考Final thoughts on following passion and impact
最终问题:计算机科学之外的 Andy 与 RichFinal question: Who are Andy and Rich outside computer science?