Jerry 和 Rohan 讨论如何从经验中学习不仅限于强化学习,以及当前 Transformer 架构为何可能成为 AI 进展的瓶颈。
Jerry and Rohan discuss how learning from experience extends beyond reinforcement learning, and why the current transformer architecture may be a bottleneck for AI progress.
要点 · TL;DR
Transformer 在持续学习上受限,AGI 需要新架构。 Transformers are limited in continual learning; new architecture needed for AGI.
架构是关键瓶颈,而非数据和算力扩展。 Architecture is the key bottleneck, not scaling data and compute.
强化学习不是从经验中学习的唯一方式。 Reinforcement learning is not the only way to learn from experience.
核心观点 · Key points
Transformer 在持续学习和测试时适应方面存在局限,需要新的架构。 Transformers are limited in continual learning and test-time adaptation, requiring a new architecture.
架构是更智能 AI 的关键瓶颈,而不仅仅是扩展数据和算力。 Architecture is the key bottleneck for smarter AI, not just scaling data and compute.
强化学习并非从经验中学习的唯一方式;更好的算法将会出现。 Reinforcement learning is not the only way to learn from experience; better algorithms will emerge.
结合预训练和强化学习的端到端优化对于真实世界性能是必要的。 End-to-end optimization combining pre-training and RL is necessary for real-world performance.
大型实验室因市场竞争而避免探索替代架构,为初创公司留下空间。 Big labs avoid exploring alternative architectures due to market competition, leaving room for startups.
反共识 · Contrarian takes
Transformer 并非最终状态;它们无法实现 AGI 所需的持续学习。 Transformers are not the end state; they cannot achieve continual learning needed for AGI.
强化学习效率低下,并非从经验中学习的终极形式。 Reinforcement learning is inefficient and not the ultimate form of learning from experience.
当前 AI 进展的主要瓶颈是架构,而非缩放定律。 Architecture, not scaling laws, is the primary bottleneck for AI progress today.
大多数架构研究的规模太小;需要大量算力才能看到潜力。 Most architectural research is done at too small a scale; large compute is needed to see potential.
像 Shampoo 这样的优化算法被低估了,与架构结合可带来巨大收益。 Optimization algorithms like Shampoo are underappreciated and can yield large gains combined with architecture.
AGI 需要模型无需人类循环即可自我改进,而当前方法无法做到。 AGI requires models that self-improve without human loops, which current approaches cannot do.