Anthropic 联合创始人 · 可解释性先驱 · Co-founder, Anthropic; interpretability pioneer
Anthropic 联合创始人,领导其可解释性研究。开创了神经网络的机制可解释性——电路、特征,以及「逆向工程」模型内部到底在计算什么。
A co-founder of Anthropic who leads its interpretability research. Pioneered the mechanistic interpretability of neural networks — circuits, features and the effort to reverse-engineer what models actually compute.
在 AI Podcast 查看 TA 的全部内容 →