一位 Transformer 共同发明者讨论离开过度饱和领域,探索连续思维机器等新架构,以及 AI 研究中研究自由的重要性。
A Transformer co-inventor discusses leaving the oversaturated space to explore new architectures like the Continuous Thought Machine, and the importance of research freedom in AI.
要点 · TL;DR
连续思维机实现自适应计算和内部推理,在新任务上超越 Transformer。 Continuous Thought Machines enable adaptive computation and internal reasoning, surpassing Transformers on novel tasks.
Transformer 的主导地位造成局部最优,抑制了替代架构的探索。 Transformer dominance creates a local minimum, stifling exploration of alternative architectures.
研究自由和自下而上的探索是突破性 AI 创新的关键。 Research freedom and bottom-up exploration are key to breakthrough AI innovations.
核心观点 · Key points
当前 AI 研究陷入以 Transformer 为主导的局部最优,限制了对新架构的探索。 Current AI research is stuck in a local minimum dominated by Transformers, limiting exploration of novel architectures.
自适应计算和内部顺序推理对于更类人的智能至关重要。 Adaptive computation and internal sequential reasoning are crucial for more human-like intelligence.
连续思维机器(CTM)无需显式惩罚即可自然展现自适应计算和良好校准的不确定性。 The Continuous Thought Machine (CTM) naturally exhibits adaptive computation and well-calibrated uncertainty without explicit penalties.
研究自由对突破性创新至关重要;商业压力常常扼杀创造力。 Research freedom is essential for breakthrough innovations; commercial pressures often stifle creativity.
当前大语言模型缺乏真正的理解和推理能力,常在变体数独等新任务上失败。 Current LLMs lack true understanding and reasoning, often failing on novel tasks like variant Sudoku puzzles.
反共识 · Contrarian takes
Transformer 过于成功,形成了‘技术捕获’,阻碍了对替代架构的探索。 Transformers are too successful, creating a 'technology capture' that discourages exploring alternative architectures.
缩放定律和蛮力可以掩盖根本性的架构缺陷,例如对结构的糟糕表示。 Scaling laws and brute force can mask fundamental architectural flaws, such as poor representation of structure.
生物启发,如神经元同步,可以带来更高效和可解释的 AI 系统。 Biological inspiration, like neuron synchronization, can lead to more efficient and interpretable AI systems.
最好的 AI 研究来自自下而上的探索,而非自上而下的宏伟计划或商业目标。 The best AI research emerges from bottom-up exploration, not top-down grand plans or commercial objectives.
当前的强化学习方法在需要罕见、新颖推理飞跃的任务上失败,表明需要新方法。 Current RL methods fail on tasks requiring rare, novel reasoning leaps, suggesting a need for new approaches.
本期章节 · Chapters(共 23)
Transformer 过饱和与探索需求Transformer oversaturation and need for exploration
扩展进化搜索与创立 SakanaScaling evolutionary search and founding Sakana
研究中的自由哲学Philosophy of freedom in research
LLM 的受众捕获Audience capture by LLMs
超越 Transformer 的困难Difficulty of moving beyond Transformers
扩展与表征Scaling and Representation
AI 研究与代理的未来Future of AI Research and Agency
连续思维机器简介Introduction to Continuous Thought Machines
连续思维机器简介Introduction to Continuous Thought Machines
序列性质与规划Sequential Nature and Planning
固定步数与上下文窗口Fixed Number of Steps and Context Window
自举机制Self-bootstrapping mechanism
自适应计算与步数敏感性Adaptive computation and sensitivity to steps
神经元级模型与同步Neuron-level models and synchronization
扩展与子采样同步Scaling and subsampling synchronization
指数衰减与多时间尺度Exponential decay and diverse time scales
连续思维机器 vs Transformer 推理Continuous Thought Machine vs Transformers for Reasoning
自适应计算时间与涌现行为Adaptive Computation Time and Emergent Behavior
时间约束下的跳跃行为Leapfrogging behavior under time constraints
共享内存与并行代理扩展Scaling out with shared memory and parallel agents
数独基准:变体谜题推理Sudoku bench: a reasoning benchmark with variant puzzles