Karina Nguyen, AI researcher at OpenAI, discusses how teams build products, the skills needed as AI advances, and why she moved from engineering to research.
要点 · TL;DR
模型训练是艺术与科学的平衡,合成数据为后训练提供无限任务。 Model training balances art and science, with synthetic data enabling infinite post-training tasks.
评估是瓶颈,模型饱和基准的速度快于新基准的创建。 Evaluations are the bottleneck as models saturate benchmarks faster than new ones are created.
随着 AI 自动化硬技能,创造力和管理等软技能变得更有价值。 Soft skills like creativity and management become more valuable as AI automates hard skills.
核心观点 · Key points
模型训练更像艺术而非科学,需要在有用性和无害性之间谨慎平衡。 Model training is more art than science, requiring careful balance between helpfulness and harmlessness.
合成数据为后训练生成无限任务,克服了预训练中的数据墙。 Synthetic data enables infinite task generation for post-training, overcoming the data wall in pre-training.
评估是瓶颈;模型饱和基准的速度快于新基准的创建。 Evaluations are the bottleneck; models saturate benchmarks faster than new ones are created.
随着 AI 自动化硬技能,创造力、倾听和管理等软技能变得更加宝贵。 Soft skills like creativity, listening, and management become more valuable as AI automates hard skills.
产品开发从规格转向用 AI 原型设计,通过评估定义正确行为。 Product development shifts from specs to prototyping with AI, using evals to define correct behavior.
反共识 · Contrarian takes
模型并未耗尽数据;通过强化学习的后训练扩展提供了无限任务。 Models are not running out of data; post-training scaling via RL offers infinite tasks.
教模型自我认知(如没有身体)可能与函数调用产生混淆。 Teaching a model self-knowledge (e.g., no physical body) can cause confusion with function calls.
AI 在创造力和美学上挣扎,因为高质量示例稀缺。 AI struggles with creativity and aesthetics because high-quality examples are scarce.
策略是 AI 可以通过综合数据和生成计划而擅长的领域。 Strategy is something AI can excel at by synthesizing data and generating plans.
Anthropic 注重工艺和优先级;OpenAI 更具创新性和冒险精神。 Anthropic focuses on craft and prioritization; OpenAI is more innovative and risk-taking.
基于像素(视觉感知)的智能体比基于语言的任务难得多。 Agents operating on pixels (visual perception) are much harder than language-based tasks.
本期章节 · Chapters(共 27)
引言与嘉宾背景Introduction and Guest Background
关于模型创建的误解Misunderstandings about model creation
数据墙与合成数据Data wall and synthetic data
合成数据及其作用Synthetic data and its role
画布行为的合成数据Synthetic Data for Canvas Behaviors
确定性与人工评估Deterministic and Human Evaluations
产品开发转变与评估Product Development Shift and Evals
提示作为原型设计Prompting as Prototyping
Anthropic 构建的功能Features Built at Anthropic
画布与任务如何诞生How Canvas and Tasks Emerged
设计工具规格与原型Designing Tool Specs and Prototyping
项目人员配置与时间线Project Staffing and Timeline
合成数据与规模化Synthetic Data and Scaling
研究者与模型设计师的角色Role of Researchers vs Model Designers
让模型更智能Making Models Smarter
世界与工作的未来变化Future changes in the world and work
AI 时代的软技能与硬技能Soft skills vs hard skills in AI era
AI 用于战略与自我提升AI for strategy and self-improvement
未来技能:软技能与管理Skills for the future: soft skills and management
Anthropic 与 OpenAI 的差异Differences between Anthropic and OpenAI
早期 Anthropic 与产品形态反思Reflections on early Anthropic days and product form factors
语音克隆与内容转化Voice Clones and Content Transformation
聊天作为灵活界面Chat as a Flexible Interface
Operator:使用计算机的 AI 代理Operator: AI Agent Using a Computer