Nathan 讨论了 AI 突破如何可能创造出比苹果和微软大十倍的公司,并强调了 AI 开发中开放性的重要性,以避免企业垄断并确保广泛理解。
Nathan discusses how AI breakthroughs could create companies ten times larger than Apple and Microsoft, and emphasizes the importance of openness in AI development to avoid corporate capture and ensure broad understanding.
要点 · TL;DR
RLHF 仍是微调主流,但安全对齐易被破坏。 RLHF remains dominant for fine-tuning, but safety alignment can be easily undone.
AI 开发的开放性对防止企业垄断和确保广泛可及性至关重要。 Openness in AI development is vital to prevent corporate control and ensure broad access.
GPT-4 生成的合成数据可有效用于 RLHF 的偏好数据。 Synthetic data from GPT-4 effectively generates preference data for RLHF.
核心观点 · Key points
基于人类反馈的强化学习(RLHF)是最流行的微调技术,并因行业投资而可能持续存在。 RLHF is the most popular fine-tuning technique and will likely persist due to industry investment.
AI 开发的开放性对于防止企业垄断和确保广泛理解至关重要。 Openness in AI development is crucial to prevent corporate capture and ensure broad understanding.
来自 GPT-4 等模型的合成数据对偏好数据非常有效,如 UltraFeedback 所示。 Synthetic data from models like GPT-4 can be highly effective for preference data, as seen with UltraFeedback.
AI 系统的安全性是系统级属性,而不仅仅是模型产物;需要多层防护。 Safety in AI systems is a system-level property, not just a model artifact; it requires multiple layers.
音频界面将降低 AI 使用门槛,尤其对教育和非英语使用者。 Audio interfaces will lower barriers to AI use, especially for education and non-English speakers.
反共识 · Contrarian takes
微调可以用很少的算力撤销安全对齐,但考虑到调整幅度小,这并不令人惊讶。 Fine-tuning can undo safety alignment with minimal compute, but this is not surprising given the small nudge.
基于人类反馈的强化学习(RLHF)中标注者的分歧是信号而非错误,反映了多方面的偏好。 Disagreement among human labelers in RLHF is a signal, not a bug, reflecting multifaceted preferences.
由于安全和分布偏移,家用类人机器人还很遥远;远程操作可能是实用的过渡方案。 Humanoid robots in homes are far off due to safety and distribution shift; teleoperation may be a practical interim.
AI 研究往往更像炼金术而非科学,直觉和规模取代了严格的假设检验。 AI research often resembles alchemy more than science, with intuition and scale replacing rigorous hypothesis testing.
大模型中的涌现能力可能部分来自基准测试噪声底限的测量假象。 Emergent capabilities in large models may partly be measurement artifacts from benchmark noise floors.
本期章节 · Chapters(共 25)
0. 开放性的介绍与动机Introduction and Motivation for Openness
1. RLHF 与 OpenAI 模型规范RLHF and OpenAI Model Spec
2. Zephyr 与 DDPOZephyr and DDPO
3. Zephyr 与 DPO 突破Zephyr and DPO breakthrough
4. Tulu 2 与指令微调Tulu 2 and instruction tuning
5. AI2 作为混合工作场所AI2 as a hybrid workplace
6. 广告插播Ad break
7. 加州大学伯克利分校背景Background at UC Berkeley
8. 机器人学与 LLM 交叉Robotics and LLMs intersection
9. 可访问性与教育Accessibility and Education
10. 偏好学习的合成数据Synthetic Data for Preference Learning
11. RLHF 中人类偏好的主观性Subjectivity in Human Preferences for RLHF
12. RLHF 的历史基础:亚里士多德与 VNM 效用定理Historical Foundations of RLHF: Aristotle and VNM Utility Theorem
13. RLHF 的脆弱性:微调后安全性不稳健Fragility of RLHF: Safety Not Robust to Fine-Tuning