Anthropic 可解释性研究员 · Interpretability Researcher, Anthropic
Anthropic 机制可解释性研究员,研究叠加、字典学习与电路追踪,试图搞清大模型内部到底发生了什么。
Mechanistic interpretability researcher at Anthropic, working on superposition, dictionary learning and circuit tracing to understand what actually happens inside large models.
在 AI Podcast 查看 TA 的全部内容 →