AI2 发布 Almo 3 系列,包含完全开放的模型、数据和配方,包括推理模型,美国开源 AI 与中国 AI 巨头竞争加剧。
AI2 releases Almo 3 family with fully open models, data, and recipes, including reasoning models, as US open-source efforts compete with Chinese AI powerhouses.
要点 · TL;DR
AI2 完全开放的 Almo 3 以完全透明挑战闭源 AI。 AI2's fully open Almo 3 challenges closed-source AI with full transparency.
中国开源模型填补了 Meta 转向 Llama 后美国的开源真空。 Chinese open models fill the US open-source vacuum after Meta's Llama shift.
后训练是一门艺术;SFT、DPO 和 RL 工具对未来模型至关重要。 Post-training is an art; SFT, DPO, and RL tooling are critical for future models.
核心观点 · Key points
预训练提供昂贵的初始化;强化学习摘取低垂果实,但两者缺一不可。 Pre-training provides a costly initialization; RL captures low-hanging fruit, but both are essential.
开源 AI 需要完全透明:数据、中间检查点和配方,而不仅仅是权重。 Open-source AI needs full transparency: data, intermediate checkpoints, and recipes, not just weights.
中国开源模型如 Qwen 和 DeepSeek 因战略开放和市场动态而占据主导地位。 Chinese open models like Qwen and DeepSeek dominate due to strategic openness and market dynamics.
后训练是一门艺术:监督微调和直接偏好优化带来巨大收益,而强化学习工具对未来的模型至关重要。 Post-training is an art: SFT and DPO yield large gains, while RL tooling is critical for future models.
AGI 将通过混乱的共同进化到来,而非突然的奇点;进步是平稳但受限的。 AGI will arrive via messy co-evolution, not a sudden singularity; progress is smooth but constrained.
反共识 · Contrarian takes
Meta 的 Llama 转向后,美国开源真空由中国实验室填补,而非美国实验室。 The US open-source vacuum after Meta's Llama shift is filled by Chinese labs, not American ones.
Qwen 模型中的预训练数据污染可能夸大强化学习研究结果,质疑可重复性。 Pre-training data contamination in Qwen models may inflate RL research results, questioning reproducibility.
直接偏好优化的收益来自选中和拒绝示例之间的对比,而非绝对答案质量。 DPO gains are from contrast between chosen and rejected examples, not absolute answer quality.
长上下文扩展更多关乎架构而非数据;糟糕的设置无法通过好数据修复。 Long-context extension is more about architecture than data; bad setup cannot be fixed by good data.
基于可验证奖励的强化学习比基于人类反馈的强化学习更难,因为存在内核不匹配和数值不稳定等系统问题。 RLVR is harder than RLHF due to systems issues like kernel mismatches and numerical instability.
本期章节 · Chapters(共 24)
引言与模型发布Introduction and Model Release
数据与预训练细节Data and Pre-training Details
长上下文数据与开源Data for long context and open source
AI 开源:风格与承诺Open source in AI: flavors and commitment
2025 回顾:开源 AI 大事记2025 recap: key events in open source AI
中国开源模型生态Chinese open-source model ecosystem
中国开源策略与美国回应China's Open Source Strategy and US Response
OLMo 起源与 AI2 开源文化Origins of OLMo and AI2's open-source culture
AI2 项目概览AI2 Projects Overview
预训练与后训练概述Pre-training vs Post-training Overview
预训练方法与数据整理Pre-training methodology and data curation
中训练与尾部修补Mid-training and tail patching
长上下文:重要性与技术挑战Long context: importance and technical challenges
长上下文对模型性能的影响Impact of long context on model performance