AI in Medicine: Magic, Judgment, and the Future
打开互动全文版(中英对照 + 朗读 + 问答)→临床医生兼 AI 研究员 Jonathan Chen 博士探讨 AI 如何支持医学中的人类判断力,与魔术进行类比,并分享真实临床实践中的见解。
Dr. Jonathan Chen, a clinician and AI researcher, discusses how AI can support human judgment in medicine, drawing parallels with magic and sharing insights from real clinical practice.
欢迎来到斯坦福医学院的《未来医学内观》。乔纳森,欢迎来到未来医学。你刚刚在我们的 grand rounds 上面对满座听众演讲,不仅展示了 AI 的前沿理念和你自己的大量工作,还表演了魔术。那么,我们就从那里开始吧。医学还是魔术,哪个先来?你是怎么走到今天的?作为斯坦福的教员,你是如何同时实践这两者的?
Welcome to Stanford Department of Medicine's Inside Look at the Future of Medicine. Well, Jonathan, welcome to the future of medicine. You just came from our grand rounds being a speaker to a packed house presenting not just cutting-edge ideas in AI and a lot of your own work, but magic. So, let us start there. Medicine or magic, which came first? How did you get there? How did you arrive at today as a faculty member here at Stanford who practices both?
呃,谁会疯狂到去尝试所有这些?嗯,我不知道。我 12 岁时玩过一点魔术,就像很多 12 岁的男孩一样。几个月后我就不玩了,因为没人可表演。我姐姐可不会看第二遍,对吧?所以我就放弃了。我学了三个蒙特牌戏法、一些魔术牌,一些基本的东西,但可能没什么能让你印象深刻的。都是些小孩子的东西。
Uh, who would be crazy enough to try all of the above? Um, I don't know. I played a little bit of magic when I was a 12-year-old boy, as many 12-year-old boys do. And then after a few months, I stopped doing it because I had nobody to perform for. My sister wasn't going to watch it more than one time, right? So, I just gave it up. I learned a three-card monte, some trick decks, a few basic things, but probably nothing you'd be impressed by. It's kind of kid stuff.
然后我成为医生,更多是因为我父母希望我这样。这里也有点故事。跟我们说说吧。我上大学出奇地早,13 岁就开始了。我不知道你怎么样。我 13 岁时,根本不知道长大后想做什么。
And then I became a doctor somewhat more because my parents wanted me. There's a little bit of story of that as well. Well, tell us. And I started college unusually young. I was 13 years old when I started college. And I don't know about you. When I was 13, I did not know what I wanted to be when I grew up.
是啊。如果你小时候问我,我会说我想当个单口喜剧演员,因为我喜欢逗人笑,那感觉很好。我从没说过想当医生或科学家之类的。
Yeah. If you had asked me when I was a kid, I would have said I wanted to be a stand-up comedian because I like to make people laugh. And that felt good. I had never said I wanted to be a doctor or be a scientist or something like that.
你 13 岁就上大学了?
Did you end up at college at 13?
洛杉矶的加州州立大学洛杉矶分校有一个项目。我想它最初是心理学系的一个实验。他们研究初中里早熟的孩子。你想参加这个项目吗?然后就这么延续下去了。我参加了一个测试,他们说:“你考得还行。要不要来夏季学期?”现在有一个完整的项目。如果你表现不错,不仅在学业上,还有情感上的准备,你就可以通过那个项目直接全日制上大学。
So, there's a program in Los Angeles at Cal State LA. It actually started, I think, out of their psychology department like an experiment, basically. They studied precocious kids in junior high. And do you want to basically be in this program? And it just perpetuated. I took a test. They said, "You did okay. Why don't you come in for summer quarter?" And now there's a whole program. If you do okay, not just academically, but kind of emotional readiness, you could just start college full-time through that program.
好吧,但这是在那些 13 岁就被考虑上大学的人当中,你考得还行。
Okay, but this is you did okay among the people who are being considered for college age 13.
对,对。是的。所以这个尺度有点不同。你明白了。并不是说我有什么极端的,但一路上就是这种动态。不过,那也非常重要和强大,因为并不是只有我一个人在巨大的大学校园里。基本上是我和大约 30 个其他孩子。所以我有一群同伴可以一起玩。几年后,我转学到了 UCLA。那真的很艰难。很艰难,因为现在我是一个 15 岁的孩子,独自住在巨大的大学校园的宿舍里。那非常具有挑战性。它迫使我飞快地成长。哦,光学习已经不够了。我可以学习并通过所有考试,哦,那已经不够了。你必须很快地弄明白生活。
Right. Right. Yes. Yes. So, the scale is a little different. You got it. Not that I have anything extreme, but that kind of dynamic along the way. But, that was also very important and powerful because it's not like it was just me amongst a giant college campus. It was me and like 30-ish other kids essentially. So, I had a cohort to hang out with. A couple years after that, I transferred to UCLA. That really was tough. That was tough because now I'm a 15-year-old kid living on a gigantic college campus by myself living in the dorms. That was very challenging. It kind of forced me to grow up real fast. Oh, it's not enough just to study anymore. I can study and ace all the exams and oh, that's just not enough anymore. You have to figure out life real quick.
是啊,对于一个 15 岁的孩子来说,那是个挑战。那么,然后你毕业了。
Yeah. Yeah, for a 15-year-old, that's a challenge. So, then you graduated.
嗯,我从 UCLA 毕业,参加了所有医学预科俱乐部。我考了 MCAT。但我那时年纪大一点了。我离开那里时已经 19 岁了。是啊。但后来在最后一刻,我听到所有这些说法:“嗯,你要当医生了。你必须对治愈有热情。这必须是你的生命使命。”而我想,我才 18、19 岁。我什么感觉都没有,好吗?我基本上还是个孩子。我还没准备好把一生奉献给这个。我需要去体验一些别的东西。而我真正的书呆子热情,我就是个书呆子,对吧?我就是喜欢学习。我喜欢学东西。我是个电脑书呆子,对吧?我喜欢编程。我就是很擅长。所以,我实际上做了几年软件工程师,做软件开发,和生物学或医学完全无关。那时正是互联网泡沫和破裂的时期。
Um, I graduated from UCLA and I was in all the pre-med clubs. I took the MCAT. But I was a little older than that. I was a ripe 19 by the time I was out of there. Yeah. But then kind of at the last minute, I hear all this stuff like, "Well, you're going to be a doctor. You got to have like a passion for healing. It's got to be like your life's mission." And I'm like, I'm like 18, 19 years old. Like I feel nothing, okay? I'm a kid, basically. I'm not ready to commit my life to this. I need to go have some other experiences. And my real underlying nerd passion, I'm just a nerd, right? I just like to study. I like to learn things. I'm a computer nerd, right? I like to program. I'm just very good at it. So, I actually worked as a software engineer, software development, nothing to do with biology or medicine for a couple years. The dot-com bubble and burst back in that era.
你那时编程吗?你用什么电脑,用什么语言?
Were you programming back then? What kind of a computer were you using and what language were you using?
哦,天哪。我的意思是,我在课堂上学习 C++,但当时业界大部分是 Java 编程。然后互联网成了个东西,做三层应用,用 HTML、JavaScript、Java。现在看起来过时了。事情就是这样演变的。
Oh, gosh. I mean, I was learning C++ in class, but most of industry at the time was Java programming at the time. And then the internet, which has become a thing, making three-tiered applications, HTML, JavaScript, Java at the time. Now it would seem antiquated. It's just kind of things have evolved.
对。但那让我对行业有了很多看法。就像,“哦,哇。突然,当我 19 岁就能赚六位数时,我妈妈就不那么在意医学院了。”
Right. But that gave me a lot of perspective about industry. It's like, "Oh, wow. Suddenly, when I can make six figures as a 19-year-old, my mom didn't care that much about medical school."
但干了一年后,公司在互联网泡沫破裂期间裁掉了三分之一的员工,对吧?包括我自己。我又找了另一份工作。还行。但这让我有了更多看法,比如,“嗯,哦,这是一份好工作,我也很擅长,但我能预见到我不确定自己能否找到一份满意、能让我快乐的长远职业。我还会再做这个十年。”
But then after a year of that and then the company laid off a third of the company during the dot-com bubble burst, right? Including myself. And I got another job. It was fine. But it gave me more perspective like, "Well, oh, this is a good job and I'm good at it, but I can foresee I'm not sure I'm going to find a rewarding long-term career I'm going to be happy with. I'm still doing this for 10 years."
你出身技术、学术还是医学家庭?
A technical or academic or medical family?
真不是。所以我才有点像是妈妈希望成为家里第一个医生的那种情况。我爸爸是电气工程师,有计算机科学硕士学位。所以我的很多书呆子气就是从那里来的。
Really not. That's why I'm sort of the mom's aspiration to be the first doctor in the family kind of situation. My dad was an electrical engineer with a computer science master's. So, that's where I get a lot of that kind of nerdiness from.
是啊。但这种组合非常奇怪。然后我的老朋友说:“你为什么不重新考虑医学院?”我说:“什么?”哦,我已经把那些都抛在脑后了。我不会回去的。我还得回去求以前的教授写推荐信。哦,不,算了吧。是我当时的女朋友。
Yeah. But this combination was very weird. Then my old friends said, "Why don't you consider medical school again?" I'm like, "What?" Oh, I left that all behind. I'm not going to go back. I have to go back and go beg my old professors for letter of recommendation after I left. Oh, no, forget it. It was my girlfriend at the time.
她说:“不,乔纳森,你绝对应该申请医学院。你应该这么做,因为人生中你会后悔没尝试过的事,而不是尝试过的事。”有意思。我一听到这个,
She said, "No, Jonathan, you should totally apply for med school. You should do it because in life you will regret what you didn't try, not what you did try." Interesting. Once I heard that,
给你施加负罪感。
Putting the guilt trip on you.
就像在伤口上撒盐。哦天哪。好吧,现在我必须去了。但我说我只申请 MD-PhD 项目,而且 PhD 必须是计算机科学。这听起来确实很疯狂,但重点就在这里。如果我要做,就必须疯狂到与众不同。否则,我还不如继续做我的工作。如果那样做,我本可以过得更轻松、更舒适。对吧。所以,MD-PhD 加计算机科学。是的。有不少地方提供这个,但也不是很多。
like so I just twisting the knife. Oh my gosh. Well, now I have to. But I said I'm going to apply for MD PhD programs only and the PhD has to be computer science. And I think that does sound crazy, but that's the point. If I'm going to do it, it has to be crazy enough to be something different. Otherwise, I might as well just stay in my job. I could have lived a much easier, more comfortable life if I had done that. Right. Right. So, MD PhD with computer science. Yes. A fair number of places do that, but not that many.
不多。不多。我要感谢加州大学尔湾分校。他们对计算机科学学院非常开明。而且他们确实还有另一个 MSTP(医学科学家培训项目)学生,比我早几年,也在做同样的组合。而大多数地方虽然会讨论,但都用奇怪的眼神看着我,好像说:“什么?是啊。你不应该学生物化学、免疫学、神经科学吗?为什么读计算机科学的 PhD?这不合适。”所以我真的很感激他们觉得这个组合有意义。尽管当时我甚至不确定它意味着什么。我只是觉得它会在 10 年、20 年后有意义。我还不知道具体是什么。
Not many. Not many. I give credit to University of California, Irvine. They were very open-minded to that school of computer science. And they actually literally had another MSTP student, a medical scientist training program student, a couple years ahead of me who was doing that same combo. Whereas most places would talk about it, but most places they looked at me strange like, "What? Yeah. Shouldn't you be getting a like biochemistry, immunology, neuroscience? Why would you get a PhD in computer science? That doesn't fit." And so, I really appreciated that they thought that combination meant something. Even though at the time I wasn't even sure what it meant. I just I think it'll mean something 10, 20 years from now. I don't even know what yet.
那是在 90 年代吗?
And this was in the '90s?
大概是 2000 年。就是 2001、2002 年。是的。所以互联网已经出现了。就像你说的,你在公司里做互联网编程,Java、JavaScript 之类的。但我们当时还在成熟过程中,我觉得世界也在逐渐理解互联网是什么。之前刚刚发生了互联网泡沫破裂。
It was about 2000. It's just 2001, 2002. Yeah. So, the internet had happened. Like you said, you were programming in the company internet stuff, Java, JavaScript, whatever. But we were really just maturing and our sense I guess the world was sort of understanding what the internet was. We hadn't had the crash had sort of happened just before
是的。刚刚开始。好的。那么,然后你就去医学院了。
Yes. We were starting to happen. Yes. Yeah. Okay. So, and then you were heading to medical school.
所以,我最后读了联合学位项目,我觉得那会很有吸引力。我不太清楚如何把各部分结合起来,但我拿到了计算机科学的 PhD,我喜欢研究有趣的极客问题。
So, then I ended up doing the joint degree program and I thought that'd be compelling. I didn't exactly know how to put the pieces together, but I got the PhD in computer science that I like to work on interesting nerd problems.
就像这里的项目一样?所以你基本上先完成 PhD,然后再做临床。有点像
Is it like the program here? So, you finish the PhD pretty much first and then do clinical. You sort of
是的。中间是交错的。两年临床前课程,PhD,然后是实习轮转。
Yes. It's staggered in the middle. Two years pre-clinical, PhD, and then clerkships after that.
你在 PhD 期间做了什么样的计算机科学工作?
And what kind of computer science work were you doing during your PhD?
现在来看,那实际上可以称为基于规则的专家系统、化学的机器学习应用。我当时做了很多有机化学的工作。那是我的主题领域。我从行业经验中学到的一课是,你确实需要强大的技术经验和技能,否则一切都只是空谈,对吧?但是,当你把它与对某个领域的深刻知识结合起来时,那就非常强大且独特了。没错。现在跳到后面,医学才是我真正关心的主题领域。当时我说:“嗯,化学。我觉得它有趣又好玩。”简而言之,我制造了能帮你做有机化学作业的 AI 系统。
Nowadays, it would actually be called like rules-based expert systems, machine learning applications for chemistry. I was doing actually a lot of organic chemistry. That was my subject domain. And a lesson I did take from my industry experience is like you do need strong technical experience and skills, otherwise it's all just talk, right? But, when you do that combine it with the deep knowledge of a subject domain, now that's really powerful and unique. Right. And now skipping ahead like medicine is really my subject domain that I care a lot about. At the time, I said, "Well, chemistry. I think that's fun and interesting." And in so many words, I made AI systems that could do your organic chemistry homework for you.
讽刺的是,如果你能帮你做作业,比如帮助药物开发,那也意味着它可以教你如何做作业。所以,虽然不是我预期的,但我 PhD 的主要产出之一是一个教育系统,被全球学生用了十多年来学习有机化学。有意思。我女儿现在正在为此受苦。
The irony is like but if you can do your homework for you and help with like pharmaceutical developments, for example, that means it could also teach you how to do your homework. And so, not what I expected, but one of the main outputs of my PhD was a education system that was used for over a decade by students around the world to learn organic chemistry. Interesting. And my daughter is currently suffering through it now.
每个人都想要这方面的帮助。是的。我们现在有其他工具在帮忙。你说得对。我们肯定会谈到那个。
Everybody would want help with that. Yes. We have some other tools that are now helping us. You got it. And we'll certainly get to that.
然后你做了临床工作,成为了一名医生。这让很多人感到惊讶,包括我自己。我其实没打算做住院医。我从来没想过要当执业医生。我当时想:“好吧,我只是获取知识,但我会毕业。也许我会做咨询。我会回到行业,很明显我会以某种形式回到行业。只是哪个行业、什么形式的问题。”然后我做了第三年的实习轮转,内科轮转。我惊讶地发现自己非常喜欢。我没想到会喜欢。天哪,你要在底层,病人总是向你抱怨。谁想处理这个?但到了那里,我觉得:“这其实很酷。你在团队中工作,一起解决有意义的问题。你做的事情有明确的影响。人们听你的话,你必须做对。这是一种非常有趣的应用专业知识,必须在那里体现出来。抽象理论不重要,你不能泛泛而谈。你要么开药,要么不开。你如何做出那种艰难的决定?我觉得这非常吸引人。所以我决定做住院医,就这样我第一次来到了斯坦福。第三次是幸运的。本科没成,医学院 PhD 成了。但最终,内科住院医。非常感激能匹配到斯坦福。从那以后我就一直在这里。
So, then you did your clinical work, became a doctor. Which was surprising to many, including myself. I actually did not intend to do residency. I was never intending to be a practicing doctor. I was like, "Well, I just get the knowledge, but I'll graduate. Maybe I'll do consulting. I'm going to go back It's so obvious I'm going to go back to industry in some form or another. It's kind of which industry that's going to be in what format." And then I did my clerkships like third year, internal medicine rotation. And it surprised me how much I liked it. I did not expect to like it. Oh, man. You're going to be at the bottom of the totem pole, have sick people complain to you all the time. Oh, who wants to deal with this? And then I got there. I was like, "This is actually really cool. You work in a team. You know you're working on good problems together. It's something that matters. And what you do has very clear impact. People are hanging on your words. You got to do it right. And it's actually a very interesting kind of applied expertise that has to manifest there. It doesn't matter what the abstract says. You can't talk in generalities. Either you prescribe the medicine or you don't. How do you make that kind of a tough decision? I thought it was very compelling. And so, then I did decide to do residency and that's how I first arrived at Stanford. Third time was a charm. Undergrad didn't work out. Med school PhD worked out. But, finally internal medicine residency. Very grateful to have a match at Stanford. I've been here since.
是的,我们很高兴你加入医学系,成为我们新成立的“计算医学部”的一员。你现在担任多个角色,但我想最好先从一些在全球媒体上报道过的、非常有影响力的工作开始。你一直在思考现代 AI 架构在医学中的应用。特别是,可能最引人注目的头条是,当你比较人类医生加 AI 工具与单独 AI 工具的表现时,AI 工具似乎胜过了医生。也就是说,医生拉低了 AI 的表现。请谈谈那项工作、一些注意事项以及它对我们意味着什么。
Yeah, we've been very happy to have you here in the Department of Medicine, part of our newly named Division of Computational Medicine. You have a number of different roles now, but I think it would be great to start first of all with just some of the really impactful work that's been covered in the press around the world. That you've been doing thinking about the application of modern-day architectures of AI to medicine. And in particular, maybe the thing that hit the biggest headline was this idea that when you compared human doctors plus AI tools to the performance of the AI tool alone, it appeared the AI tool outperformed the doctor. So, the doctor was pulling down the performance of the AI, but talk about that work and some of the caveats and what it means for us.
当然,那绝对不是我们预期的。你知道,当我第一次看到 GPT-4 的预览版,事情正在爆发时,我想:“天哪。”我真的觉得我需要扔掉一半的研究计划。这个东西超越了我认为我能做到的很多事情,而且进展比我预期的快得多。然后我转向了实证研究,但我甚至懒得去检查多项选择题,因为我知道一百个人都能做。那太容易了,而且也不是我们在意的,对吧?作为医生,你立刻知道那并不重要。
Sure, that was definitely not what we expected. You know, when I first saw a preview version of GPT-4 as things were blowing up, I was like, "Holy smokes." I literally felt like I need to throw away half of my research program. This thing has leapfrogged so many things that I thought I was capable and it just has moved much faster than I expected. And pivot there, but then empirically I didn't even bother checking multiple-choice questions because I know a hundred people can do that. It's so easy to do and also that's not what we would care about, right? As a doctor, you instantly know that's not what really matters.
不是说小故事就是全部,但它能把问题展开到另一个深度。所以,我们来看一个复杂的病例推理,真正的专家共识评分非常难做,但我们的教育工作者掌握了解锁它的关键。结果并不是我们预期的。我们原本的假设是:“哦,你可以在 UpToDate 上查资料,或者用这个看起来很聪明的 AI 工具,我打赌医生会表现得更好。我们可以证明这种组合非常棒。那该多酷啊?”但我们发现并非如此。我们发现它们并没有带来多大改变。这对我们来说非常意外,真的让我们开始问不同的问题。技术很神奇,但显然它不是答案。我的意思是,这一点其实一直都没变。
Not that the vignettes are the whole thing that matters, but that would unpack it to another level of depth. So, let's see a complex case reasoning really expert consensus grading very hard to do, but that's our educators with key to unlocking that. And it is not the result we expected. We thought our hypothesis was "Oh, you could look up stuff in UpToDate or you could use this AI tool that really seems really smart and I bet the doctors are going to be even better. We can show this combination is so great. How cool would that be?" And that's not what we found. We found they didn't make that much difference. Very surprising to us and really made us ask different questions. Oh, technology is amazing, but clearly it is not the answer. I mean, that actually has been true all along.
这些医生当时已经对工具有一定了解了吗?
And these were doctors who were already somewhat familiar with the tool?
所以,这是一个关键点。我们做那项研究的时候,还处于非常早期的阶段。那是两年半以前,三分之一的医生这辈子从未碰过 ChatGPT 或聊天机器人,另外三分之一可能只用过一两次。很明显,很多人不知道它是什么,也不知道怎么用,或者即使知道,他们也不信任它,对吧?我同意。我不知道你是否记得第一次与 AI 聊天机器人对话的感觉。那感觉很怪异。你不知道自己能问什么,它能说什么。如果你从未有过那种体验,那是一种非常奇怪的感觉。很难知道如何有效使用它。
So, this is a key. At the time we did that study, it was pretty early days. It was 2 and 1/2 years ago and so a third of the doctors had never touched ChatGPT or chatbot before in their life and the other third maybe used it once or twice. So, it was clear a lot of them they did not know what it was and they did not know how to use it or if they did, they did not trust it, right? And I agree. I don't know if you remember the first time you had a the chat with the AI chatbot. It felt weird. And you didn't know what can I ask? What what could it say? It was just a very bizarre feeling if you never had that experience. It's hard to know how to use it effectively.
哦,我的意思是,几十年来计算机科学中一直有图灵测试这个概念,对吧?然后突然之间我们几乎跳过了它。它几乎在一夜之间变得无关紧要。但对于医学来说,我认为它显然是主要应用之一。当然,这正在影响我们的整个世界。
Oh, I mean, for decades in computer science this idea of the of the Turing test right? Was uh and suddenly we were leapfrogged past it almost. It almost became an irrelevance overnight. Uh but for medicine where I think it's been clearly one of the major applications. I mean, of course, this is affecting our entire world.
是的。但我认为对于互联网来说,健康并非立即成为最有趣和最重要的应用之一。我认为对于我们现在看到的语言模型,它确实走到了前沿,而且情况已经改变。我们再也看不到没有咨询过语言模型的患者了。两年后,如果世界上还有医生不知道如何使用语言模型,那也显得不太可信了。嗯,但你认为具体是什么原因呢?是使用系统时的天真,还是你真的认为人类拖累了模型推理的某些方面?
Yes. But I think with the internet it wasn't immediately obvious that health was one of the most interesting and important applications. I think with the language models that we're seeing now, it really is to the forefront and certainly it's been changed. We don't see patients anymore who haven't already consulted a language model. The idea 2 years later that there would be a doctor who doesn't anymore in the world how to use the language model is also we would be one that isn't that credible anymore. Um but what exactly do you think it was that that naivety about just how to use the the system or or do you genuinely think the human was dragging down some element of the way the reasoning within the model worked?
哦,天哪,我在我们的一项后续研究中谈到了一点,那项研究卡在了预印本同行评审中,因为同行评审过程跟不上技术发展。我们几个月前就完成了研究。嗯,存在锚定和谄媚现象,对吧?如果你说:“嘿,AI,我有个病人呼吸急促。你觉得是充血性心力衰竭吗?”它可能会说:“当然,好主意,医生。充血性心力衰竭听起来不错。”而它可能并没有给出本可以给出的客观评估。事实上,我们多次通过实证证明了这一点。为什么它会这样?因为并非所有听众都一定具备这些语言模型架构的技术背景。我的意思是,为什么模型会谄媚?
Uh gosh, I I I talked a little bit in one of our follow-up studies that stuck in preprint peer review cuz the peer review process can't keep pace with technology. We we finished the studies you know, months ago. Um there is anchoring and sycophancy that happens, right? If you say, "Hey AI, I got this patient with a shortness of breath. You think it's congestive heart failure?" It'll probably say, "Sure, good idea, doctor. Congestive heart failure sounds great." When when maybe it's not giving you the objective assessment it could have. In fact, we've empirically shown that many times that um And why does it just for cuz not everybody I think who listens to this necessarily has a technical background in the way the architecture of these language models work. I mean, why are the models sycophantic?
好——好吧。在最基本的层面上,我喜欢称它们为“增强版自动补全”。它只是阅读了互联网上的大量文字,然后猜测下一个词来填充句子。但如果它只是那样做,实际上用起来并不有趣,因为它只是编造出一些看似合理的句子,但那不是我问的问题。那不是我问的问题。所以,他们做了监督微调和基于人类反馈的强化学习。这意味着什么?它会输出 10 个不同的答案。你有没有注意到,如果你问聊天机器人同一个问题,它会给出 10 个不同的答案?然后人类说:“我更喜欢这个答案。这个更好地回答了我的问题。我觉得这个更连贯。”所以,这很好。因此,它给出了更连贯的答案,但这也意味着人类在灌输他们喜欢的价值观。而人类真的很喜欢别人同意他们,喜欢同伴同意他们,不管对错。人们真的很喜欢那种感觉。是的。所以,这将它引导到了它所在空间的某个特定部分,在那里它只想取悦人。
O- okay. A couple At the very base level, I like to call them autocomplete on steroids. It's just read the internet in so many words and it just can guess the next word filling in a sentence. But if it just did that, it actually is not very interesting to work with cuz it just it just makes up sentences that sort of make sense, but that's not the question I asked. That's not the question I asked. And so, they did supervised fine-tuning reinforcement human feedback. What does that mean? It'll output 10 different answers. Have you noticed that if you ask the chatbot the same question, it gives you 10 different answers? And then humans said, "I like this answer better. This is answering my question better. This one I feel is more coherent." And so, that's good. So, it gives more coherent answers, but that also means humans are instilling their values of what they like. And humans really like it when people agree with them and when peers agree with them whether it's right or not. People really like that feeling. Yeah. So, this leads it then into a particular part of its its sort of the space it it lives in where it uh only wants to please.
是的。是的。嗯,但这到了它无法区分我需要作为答案提供的事实和它通过强化学习反馈循环被灌输的那种取悦欲望的程度。嗯,所以,我认为,这是你曾经建议过的,我们稍后会谈到你的其他教育角色,但当我们与这些模型互动时,有一个非常实用的建议:先让它运行,不要给它任何关于答案的先验概率或偏见。而且,不管怎样,很多这些概念,我喜欢说它们并不是新概念。你喜欢人类,我们的过程也会这样做。我不知道你是否遇到过急诊室收治病人,他们告诉你:“嘿,这个病人是收缩性心力衰竭或加重。”也许是这样。80% 到 90% 是对的,但有时等等,什么?不,这个病人是肺炎。他们完全锚定在了错误的东西上。不过,我认为有时这取决于信息来源,我们可能不想透露我们的怀疑程度。取决于另一端是谁,有些信息来源你知道,它们的先验概率有时你从稍微不同的起点开始。
Yes. Yes. Uh but that goes as far as it doesn't separate like the facts that I need to provide as the answer from its its sort of sense of wanting to please that is imbued in it by this reinforcement learning feedback cycle. Uh so, I think I mean, this is something that you have advised and again, we'll come to your other education role in in a little bit of time, but wh- as we're interacting with these models, so there's a very practical piece of of advice is to let it operate kind of first uh without giving it some pre- pre-probability of an idea, a bias that you have about what the answer might be. And and for what it's worth, a lot of these concepts I like to say it's it's not really a new concept. You like humans and our process do that, too. I don't know if you get an admission from the emergency room, they tell you, "Hey, this patient has systolic heart failure or exacerbation." Like it maybe. And 80 to 90% of that is correct, but sometimes wait, what? No, this patient has pneumonia. They they totally have anchored on the wrong thing. Although, I think sometimes it depends on the source and we we might might want not want to reveal the level of skepticism we have. Depending on the depending on who's on the other end of the Some sources you know, their their pretest probability Sometimes you start with a slightly different
而且是以合理的方式。事实上,AI 的另一个危险是自动化偏见。我的意思是,它经常是正确的,确实如此。你最终会一直浪费时间反复检查。哦,但当事情真正重要时,它也会让你陷入糟糕的境地。我发现,在诊所里使用的数字抄写员,它当然非常清晰、写得很好,标点符号很棒,大写字母位置正确,逗号也很好,到处都是破折号。你知道,它说得如此之好,以至于我认为我们的大脑习惯于阅读经过人类检查的、写得非常好的文本,所以我们可能在潜意识层面将其关联起来。
And in a reasonable way. In fact, the other danger with AI is this automation bias. It's like I mean, it's right so often which it is. You you eventually waste your time double-checking all the time. Oh, but that gets to you in a bad situation, too when when it starts to really matter. I f- I find this with the uh uh digital scribe that we have available in the clinic, it is of course so articulate and so well written and the punctuation is great and the capital letters in the right place, the commas are good, em dashes everywhere. You know, it it it speaks so well that I think our brains we're used to reading text that is very well written when it has been checked humans and and so, we sort of associate probably at a subconscious level
是的。我们现在必须重新审视这一点。完全同意,我认为在我之前的演讲中,我更强调了像虚构这样的问题。照片级逼真的图像、非常雄辩的散文,这些都不再是可靠的真理指标了。
Yes. And we have to revisit that now. Absolutely absolutely agree and I think that's in my prior talks I emphasized more where it's like confabulations. It's photo-realistic images, very eloquent prose, those are no longer reliable indicators of truth anymore.
你的判断力比以往任何时候都更重要,更不用说那些偶然的错误了。还有一些恶意行为者,他们直接试图用看起来很可信的东西来欺骗你。我们正在进入一个相当可怕的世界,我其实不知道我们目前是否完全知道如何应对。说到这个,你还展示了一些非常有趣的数据,很多模型在基准测试中接受测试,这些测试有对错之分,是多项选择题,不太贴近现实世界。但我认为你很好地指出,在现实世界中,我们当然没有多项选择题,也没有包装整齐的问题。但在医学中,我们有一个原则:首先不伤害。所以你一直在探索这一点。
Your judgment matters even more than it did before. And let alone accidental errors. There are bad actors out there who are straight-up trying to scam you with things that look very credible. And it's actually a kind of scary world we're emerging into that I actually don't know that we fully know how to cope with it at this point. I mean, talking of that, you presented also some really interesting data around, you know, many of these models have been tested on benchmarks where there's a kind of right answer and a wrong answer. It's multiple choice. It's not very real world. But I think you make the point very well that in the real world, we certainly don't have multiple-choice questions. We don't have neatly packaged questions. But we do have a rubric within medicine that we should first do no harm. And so, you've been exploring this a little bit.
是的,这里有一个非常棒的团队,你知道,经典情况是团队做了工作,而我的名字也出现在上面。这一切始于我们实际上在尝试做我们的电子或控制台系统。比如,让我们原型化一个提供该控制台的 AI。那会是什么样子?但嘿,伙计们,在我们将其发布到现实世界之前,我们能确保它至少是安全的吗?至少是安全的。让我们找一些代表性案例,确保这个东西不会给出实际上会伤害人的答案。更不用说它在那时是否有用了。这不是一个容易回答的问题。在这项刚刚作为预印本发布的研究中,准确性和伤害不是一回事。不是。我敢打赌你有很多非常聪明的学生、住院医师、实习生,他们非常聪明,知道一切,然后有人说,‘哇哇哇,那很危险。别那么做。’这不是因为他们不聪明。而是像,‘哦,我不应该那样做吗?那很重要吗?’他们没有任何依据来做出这种区分。所以,这需要一种非常不同的、经过深思熟虑的评估,结果显示它与准确性本身并不相关。
Yeah, there's a great fantastic team here who, you know, it's always this classic that the team did the work and somehow my name's on there, too. This started, we were actually trying to do our electronic or console system. Like, let's prototype an AI providing that console. What would that be like? But hey guys, you know, before we unleash that in the real world, can we just make sure it's at least safe? At least safe. Let's get some representative cases and make sure that the thing doesn't give answers that would actually harm somebody. Let alone is it even useful at that point. And that's not a trivial question to answer. And in this study, which is just emerging as a preprint right now, accuracy and harm are not the same thing. It's not. I bet you have a lot of really bright students, residents, trainees and they're super smart. They know everything and somebody like, 'Whoa whoa whoa, that's dangerous. Don't do that.' And it's not because they're not smart. It's like, 'Oh, was I not supposed to do that? Was that important?' Like they have no basis to make that distinction. So, that required a very different deliberate type of evaluation that showed that it was not correlated with accuracy alone.
是的。我经常思考我们现在在医学中应用 AI 所走的路与自动驾驶汽车世界之间的相似之处。有很多不同,但也有非常多好的类比和相似之处。我想我在你展示“不伤害”模型时就想到了这一点。但作为最低标准,我们必须从那里开始。事实上,我们可能应该将准确性置于“不伤害”之下。我认为这实际上有点像临床试验的进程。在检查疗效之前,我们能确保不会意外地害死人吗?一旦你做到了,现在让我们看看真正的疗效,而我们在某种程度上是跳跃前进的。让我们感到惊讶的是,在我们刻意测试之前并不知道,比如,‘哦,对于这些棘手的问题,虽然不是巨大的伤害,但也不是微不足道的。’10% 到 20% 的情况下,即使是前沿模型有时也会犯,而真正有趣的是,伤害往往在于它们没有做某事。没有推荐实际上需要处理的重要事项。
Yeah. I often think about the parallels between the world we're now walking down with AI in medicine and the autonomous vehicle world. Now, there's many differences, but there's also very many good analogies and similarities. And I think I was thinking of that as you were presenting the 'do no harm' model. But that as a minimum bar, we have to start with that. And in fact, we probably should subordinate accuracy to do no harm. I think that actually is kind of like a clinical trial progression. Before you check the efficacy, can we just make sure we're not accidentally killing people? Once you've done that, now let's see the real efficacy and we're kind of jumping ahead to some degree. And what's kind of surprising us, but we didn't know until we deliberately tested like, 'Oh, for these kind of tough questions, it's not a huge harm, but it's not a trivial amount either.' 10 to 20% of times even frontier models will sometimes do or, and what's really interesting, often the harm was that they didn't do something. Didn't recommend something that was actually important to address.
是的,我认为我们经常忽略,因为我们太关注于“作为之罪”。确实,我们应该关注那些,但我们往往不擅长考虑“不作为之罪”。这对于计算分析数据科学工作来说通常非常困难,对吧?我们检查阳性预测值、精确度。比如,如果我找到了,我知道我找到了。但是,你漏掉的所有疾病、所有诊断呢?它们从未被记录。关键是没有人把它们写下来,因为他们从未发现它们。我认为这是一个很好的提醒教训。再次强调,这不是一个 AI 的教训。这就像一个人生教训,对吧?这是一个临床实践教训:什么都不做并不意味着没有伤害。当你本应看到并采取行动时,你实际上可能造成很大的伤害。
Yeah, I think that we often miss because we're so interested in the sort of sins of commission. And it is true that we should pay attention to those, but we often are not very good at accounting for the sins of omission. And a very tough one for computational analytic data science work in general, right? We're checking positive predictive value precision. Like, well, if I found it, I know I found it. But, what about all the diseases, all the diagnoses you missed? They were never charted. The whole point nobody ever wrote them down because they never found them. And I think that's a great reminder lesson. Again, this is not an AI lesson. This is like a life lesson, right? It's a clinical practice lesson that doing nothing doesn't mean doing no harm. You can actually do a lot of harm when you should have done something when you saw it.
另一个值得花几分钟探讨的领域是 RAG,即检索增强生成。你在演讲中简要提到了它。但是,我们正在从一个世界转变,在那个世界里,这些模型本质上是下一个词预测器,它们建立在海量数据和词汇之上,并通过强化学习来取悦你。我们谈到了所有这些。这有一定的可靠性。但是,这种模式也与大量捏造有关,例如为医学问题捏造参考文献。然而,随着我们前进,有很多很好的例子,包括你提到的一家名为 Open Evidence 的公司。现在,进入一个更可能的世界,前沿模型在工作时也会这样做。它们更可能去查找来源,然后利用它们的语言操作技能直接从该来源带回数据。所以,请谈谈这种演变,总体上我们都在经历,但特别针对医学 AI。
Another area would be great to explore for just a few minutes is what falls under the term RAG for retrieval augmented generation. You did talk about it briefly in your talk. But, we are moving from a world where we have these models essentially as next token predictors where, you know, they're built on this huge massive of data, massive words, and are trying to please you because of the reinforcement learning. All of that part that we talked about. And there's a certain reliability with that. But, it's also that mode has also been associated with a lot of the fabrication, fabricating references for example for medical questions. But, as we've moved and there's great examples including a company you mentioned called Open Evidence. Now, to a world where it's much more likely and the frontier models do this too at work. They're more likely to go to a source and then use their skills in manipulating language to bring back the data directly from that source. So, talk a little bit about that evolution that we're sort of all in generally but specifically for medical AI.
这是一个伟大的迁移。我过去常常谈论,哦天哪,虚构、幻觉。这东西就像得了科萨科夫综合征。它只是在编造词语,尽管毫无意义。它并不是想对你撒谎。它不知道什么是谎言。它只是在过程中填补空白。我决定今天不讨论这个,因为它不是一个已解决的问题。
It's a great migration. I used to talk about, oh gosh, confabulate, hallucinate. This thing just has a ready Korsakoff syndrome. It's just making up words even though it doesn't make any sense. It's not trying to lie to you. It doesn't know what a lie is. It's just filling in the blanks as it goes. I decided not to cover that today because it's not a solved problem.
是的。
Yeah.
但是,这已经不那么成问题了,因为大多数模型都内置了专门的 RAG,即检索增强生成。就像,你问一个事实性或证据性的问题,我会去查找那个证据,并为你阅读文档。我认为这是一种更加扎实的方法。你仍然不能完全依赖它。但好处是,你不必仅仅猜测它在说什么。就像,哦,让我自己去仔细检查那篇文章,因为它很重要。而且我认为,讽刺的是,作为临床医生,我们处于一个很好的位置来处理这些模糊性,对吧?有时是对的,有时是错的,有时有来源,有时没有。对于重要的事情。但是,就像在现实生活中,你的顾问在病房里总是给你很多建议,不是吗?大多数时候他们是对的,但有时,等等,这关系到我们是否进行手术。所以,这个值得仔细核对。至于抗生素用 5 天还是 7 天,你知道吗,可能差不多。我不会太担心。所以,你知道什么时候重要,然后你知道在重要的时候如何深入挖掘。
But, it's much less of an issue because most of the models have deliberate RAG, retrieval augmented generation, baked in. Where it's like, well, you're asking it a factual or evidence-specific question. I will go look for that evidence and kind of read a document for you. And I think that is a very much more grounded approach. You still can't rely on it perfectly. But, what is also nice is you don't just have to guess what it's saying. It's like, oh, let me go double-check that article myself because it matters. And I think as ironically, as clinicians, we're in a very good position to address a lot of these vagaries, right? Sometimes it's right, sometimes it's wrong, sometimes it has a source, sometimes it doesn't. For things that matter. But, just like in real life, your consultant gives you a lot of advice all the time on the wards, don't they? And most of the time they are right, but sometimes they wait a minute, it really matters if we go to surgery or not. So, this one is worth double-checking. Whether we have the antibiotics for 5 or 7 days, you know what, it's probably about the same. I'm not going to worry too much about it. So, you know when it matters and then you know how to dig in deep when it does.
是的,我认为这是非常重要的一点,因为你提到链接也直接在那里。
Yeah, and I think that is a really important part of this because you said that the link is also right there.
是的。所以你可以直接从模型生成的摘要(无论你对其可信度如何)跳转到原始来源。没错。然后如果你要做医疗决策——我认为目前我们主要就是这样使用前沿模型的。而且正如你指出的,我们的住院医生和许多同事倾向于使用 Open Evidence,这是一个专门基于医疗数据构建、面向医生使用的语言模型。不过,我认为大多数患者和世界上其他人当然还是倾向于使用前沿模型。
Yeah. Yeah. So, you can go directly from a summary from a model that you may or may not believe to some extent directly to the source. Yes. And then if you're going to make a medical decision and I think that's how we mostly use the frontier models at this point. And I think as you pointed out, our residents and many of our colleagues tend towards using Open Evidence, which is a specific language model built really from medical data with doctors' use in mind. Although, I think most of our patients and most of the rest of the world tend to use, of course, the frontier models.
嗯,我认为 Open Evidence 需要 NPI 号码才能使用。他们试图限制访问。这个领域还有很多其他参与者。
Well, I think Open Evidence requires an NPI number to get into it. They're trying to restrict it. And there are many other players in that space as well.
限制在有提供者号码的医生范围内。是的,考虑到这一点,我想部分原因是出于责任考虑,因为否则很明显所有这些模型基本上都在提供医疗建议,对吧?它们都有免责声明:不用于医疗建议。你不能将其用于任何医疗目的。但得了吧,这显然就是人们在做的事情。这就造成了一种非常奇怪的紧张关系:到底谁对做出的决定负责?
Restrict it to doctors with a provider number. Yeah, so with that in mind, I imagine some of that's for liability purposes, because otherwise it's very obviously all of these models are basically giving medical advice, right? They have this disclaimer: not for medical advice. You cannot use it for any medical purpose. But give me a break. That's clearly what people are doing. And it creates that very odd tension of who's really responsible for the decisions being made.
显然,在斯坦福,我们非常迅速且乐于接受 AI 的潜在益处,包括在我们的医疗系统中。你提到工程团队上周确实和我一起在医院查房。
We've been obviously very quick and ready to embrace the potential benefits of AI here at Stanford, including in our healthcare system. And you mentioned the engineering team who actually visited with me on rounds in the hospital last week.
是的,没错。没错。
Yeah, that's right. That's right.
非常有趣。但这个团队开发了一个工具,允许语言模型在封闭且安全的条件下访问单个患者的记录,然后利用我们一直在讨论的这些工具——真正的优势:整合数据、获取原始数据、带回答案。但在单个患者的背景下做到这一点意义重大。我想也许一些收听播客的人不一定意识到,尤其是我们的住院医生花了多少时间在电子健康记录的不同屏幕之间点击。所以也许谈谈这个——这看起来像一场革命。我的意思是,这看起来像一场革命。先谈谈电子健康记录层层叠加造成的问题,以及这可能代表的潜在解决方案。
Really fun. But the team for a tool that allows a language model under closed and safe conditions to access actually a single patient's record and then use these kind of tools we've been talking about—the real strengths: assimilating data, going to get to the raw data, bring back answers. But to do that in the context of a single individual is huge. And I think for maybe some of the people who listen to the podcast here don't necessarily realize how much time especially our residents spend clicking through different screens in our electronic health record. So maybe just talk about that—that seems like a revolution. I mean, it seems like a revolution. Just talk about the problem first of all created by the layers of electronic health records, and the potential solution that this might represent.
无论好坏,我们的实践中有太多这样的部分,对吧?你知道,Abraham Verghese 总结为“iPatient”。我们是在和患者交谈,还是主要只是在电脑前打字?电子病历,无论好坏,是医学中许多事情的核心枢纽。ChatGPT 刚出现时,我想到的第一件事就是希望它能连接到病历上。但由于各种隐私、安全和其他原因,这实际上非常困难。但经过几年,团队一直在努力整合。现在我们有了,至少是它的初步形式。非常强大,因为我不知道人们是否意识到我们之前做过研究,点击研究。比如,受训者、医生在电脑上花最多时间做什么?实际上不是写病历,也不是下医嘱。他们大部分时间是在回顾病历。他们只是查找信息,试图阅读记录、理解内容、检查结果,然后整理。最后在最后一刻,他们从不同地方总结。所以每一条信息之间,可能有五六次点击。我最近因为另一个原因在回顾一个医疗案例。他们给了我 5000 页文档。顺便说一句,其中 4900 页毫无价值。但仅仅筛选和理解这些内容就花了数小时的工作。这是我们给医学生第一年非常常见的任务:他们就说,你去仔细查看那些记录,总结一下发生了什么。
For better or worse, so much of our practice, right? You know, Abraham Verghese sums up the iPatient. Most of our—do we talk to the patient or do we mostly just type at the computer? And the electronic medical record, for better or worse, is the central hub for so much of what happens in medicine. And the first time ChatGPT emerged, the first thing I thought is I want that attached to the medical record. But for all sorts of privacy, security, and other reasons, it's actually very difficult to do that. But for a couple of years, teams have been putting it together. And now we have that, at least the first forms of it. Very powerful, because I don't know if people realize we've done the study before, the click studies. It's like, where do trainees, where do doctors spend most time on the computer? It's actually not writing notes. It's not putting orders. Most of their time is chart review. They're just looking stuff up and trying to read notes, making sense of it, checking results, and they're collating it. And then at the last minute, then they summarize it from different places. So in between each piece of information, there might be five or six clicks. I literally was reviewing a medical case recently for another reason. They sent me 5,000 pages of documents. And most 4,900 of those pages were worthless, by the way. But just sifting through that and making sense of it was hours and hours of work. And a very common job we give to medical students for their first year: they're like, why don't you scour those records and summarize what's happening?
我认为我们经常让受训者担任这个角色,因为我们变得非常擅长找到埋藏在病历各个部分的数据片段。
I think we often use our trainees in that role because we become really, really good at finding bits of data that are buried in various parts of the chart.
因为那可能是关键,而且常常就是关键。比如,等等,那场手术是什么时候?他们有没有接受化疗?或者这就是为什么你会咨询感染科关于抗生素史的原因,对吧?
Because that could be the key and often is the key. It's like, wait, when was that surgery? Did they get the chemo or what? Like this or this is why you consult infectious disease for a history of antibiotics, right?
我们需要长长的培养结果、抗生素史。哦,当你把它摆出来,现在就很有意义了。
We need the long, long form cultures, antibiotics. Oh, it's when you laid it out, it makes so much sense now.
我们拿它开玩笑,但这实际上是一件严肃的事情。
We joke about it, but it is actually a serious thing.
这是一项真正的技能,一项非常有价值的技能。
It's a real skill, a real valuable skill.
疾病记录,因为你知道那个团队会非常详细地回顾既往病史。现在还不完美,还没有完全准备好投入实际使用,但这有点像第一个原型。你可以看到它的发展方向,而且它会达到那个目标。它非常强大且令人兴奋。我希望看到越来越多,而且我认为我们会看到。AI 可以完成很多这样的工作,然后我们可以回到:好的,谢谢你的总结和核心事实。所以现在我可以真正思考如何实际处理患者的病例了。
Diseases note because you know that team will have gone through the past medical history in real detail. And now not perfect, not exactly ready for prime time, but it's kind of like the first prototype. You can see where it's going and it's going to get there. It's very powerful and exciting where it's happening. I would love to see more and more and I think we will. AI could just do a lot of that and we can get back to: okay, thank you for the summary and the core facts. So now I can actually think about how to actually approach the patient's case now.
是的。我有——我的意思是,有几个例子——它叫 Chat EHR,是斯坦福的版本。我们的住院医生叫它 Chatter,这名字不错。
Yeah. I had—I mean, there's a few examples of—I mean, it's called Chat EHR. It's the Stanford version of this. Our residents call it Chatter, which is pretty good.
嗯,但我遇到过很多情况,它真的救了我。我最近在一个移植会议上介绍几个患者,不知怎么地需要总结病例,但只有大约 10 分钟来收集信息。不可能。是的。如果你想想患者的那种复杂性。通常如你所说,你会写感染科记录,或者至少找到 H&P,即病史和体格检查。
Um, but I've had a number of situations where it really has kind of saved me. I was recently presenting a couple of patients at a transplant meeting and one way or another ended up needing to summarize the case but having only like 10 minutes to try to collect the information. Impossible. Yeah. If you think about that sort of complexity of a patient. And normally as you said, you would do the infectious diseases note or at least you find the H&Ps, which is the history and physical.
机会更大。
Better chance.
当患者入院时,你希望有人已经总结了整个既往病史,或者至少从某处复制粘贴了。
When the patient comes into the hospital, you hope that somebody has summarized the whole past medical history or at least cut and pasted it from somewhere.
是的。在这个案例中,找不到。
Yes. In this case, it was not to be found.
哎呀。那个患者来自另一家机构。没有人记录完整的病史。但我们知道它在电子病历里。现在到处都是 Epic。100 份记录,对吧?
Uh-oh. That patient came from another facility. Nobody had taken the entire medical history. But we knew it was in the electronic record. Now it was the Epic everywhere. 100 notes, right?
正是。而我只有 5 分钟。
Exactly. And I had 5 minutes.
所以,现在这还不是即时的。它确实需要几分钟来思考和收集;你给的上下文越长,耗时越长。但它救了我一命,因为它能给出一个漂亮的患者时间线总结。这类事情它做得非常好。我甚至想进一步扩展。它节省了你的时间,这很好。但它还有潜力——而且我——我妻子是病理学家。她正在那里做一个原型。
So, and now this is not an instant thing yet. It does take a few minutes to think and gather; the longer context you give it, the longer it takes. But it saved my bacon because to just say a beautiful timeline summary of the patient. That sort of thing it does very well. And I would even extend that further. It's saved you time, which is good. But also it has the potential—and I've—my wife is a pathologist. She's doing a prototype there.
就像,哦,它抓住了人类没发现的东西。比如,这个病人有过血液恶性肿瘤吗?我不知道。仔细查看的人说,我没看到。然后说,不,等等,9 个月前有一条记录,确实有。这完全改变了现在的解读。所以,我认为这是一种非常强大的能力,我们甚至还没有触及到它能在实践中带来益处的皮毛。
It's like, oh, it caught things that the humans didn't. Like, has this patient ever have a hematologic malignancy? I don't know. The person who scoured there was like, I don't see it. It's like, no, wait, there was a note from 9 months ago and it did. That totally changes the interpretation now. So, I think it's a very powerful capabilities that we haven't even barely scratched the surface of what a change will make in beneficial ways for practice.
是的,我认为很大程度上是因为深度学习时代先来了,你知道,我们习惯了谷歌照片告诉我们,这是你的狗,这是你的货车,那是卡车,那是飞盘。我们有点习惯了这样的想法:影像可能会被 AI 的魔力所覆盖。但我认为现在我们看到,实际上所有的数据都是如此。
Yeah, and I think that so much I think because the deep learning era for AI came first where you know, we were used to Google Photos showing us yeah, this is your dog, this is your van, that's a truck, that's a Frisbee. We were sort of used to the idea that imaging could be something that would come under the spell of AI. But, I think now we're seeing that really all of the data.
绝对。我的意思是,这就是为什么当我看到 GPT-3 和 3.5 出现时,我心想,天哪。这实际上我认为会改变世界。而很多其他东西——说实话,我引用最多的论文,无论好坏,是说机器学习在医学领域被过度炒作。那是一篇两页的观点文章,是我被引用最多的东西。但在这里,我觉得,好吧,也有炒作,人们会有点过度认为 AI 是魔法。它就是一个营销流行语。但这里确实有颠覆性的能力。我认为互联网是最接近的类比。这并不意味着每个人都会失业,但它会影响几乎所有事情,基本上任何与读写有关的事情,而这几乎就是一切。
Absolutely. I mean, that's why when I saw the emerging GPT-3 and 3.5, I was like, holy smokes. This actually I think is going to change the world. Whereas a lot of stuff I literally my most cited paper for better or worse is that machine learning is way overhyped in medicine. This is a two-page perspective. It's the most cited thing I have. But, here I'm like, okay, there's hype too and people will think AI is magic a little bit too much. It's a marketing buzzword is what it is. But, there are very disruptive capabilities here. I think the internet is the closest analogy. It doesn't mean everyone's going to lose their job, but it affects everything basically anything does that has to do with reading and writing which is like everything basically.
是的。这是你的愿景吗?比如说,这个问题我经常被问到。我相信你也被问到同样多。就像,这些模型现在是最差的。对吧。对吧。带我们回顾一下,一切始于——当然 2017 年是 Transformer 论文,但真正开始是 GPT-3 到 3.5。那是在 2021、2022 年。大概三年前。所以现在问你五年后的事感觉有点荒谬,但你知道,也许你的论文那时已经发表了。我希望如此。至少能收到第一轮审稿意见。是的,但你认为未来几年会怎样?
Yeah. And is that your vision for let's say, this is something I get asked a lot. I'm sure you get asked it just as much. Like, this is the worst these models are ever going to be. Right. Right. Take us and all started I mean, of course 2017 was the Transformer paper, but really it started with GPT-3 two two and a half to three. And for that we're talking about 2021, 2022. It's like three years ago. So it feels ridiculous for me to ask you about five years in the future, but but you know, maybe your paper will be published by then. I hope so. At least get the first review back. Yeah, but what do you think just for the next few years?
哦天哪,这很难,对吧?因为在这个时代,五年后看看我们五年前的样子。我们甚至不会谈论这个。我们还在讨论风险预测模型,这完全没问题,但我一直知道它们的范围有限。五年内,这是一个简单的预测,对吧?我很确定大多数医生会使用环境记录员,或者至少那会是一种非常常见的常规技术。它甚至不会再显得新奇。现在几乎已经不算新奇了。其他对未来的预测,我最近刚填了一个宾果卡。我很惊讶还没有针对大型科技公司因 AI 伤害人(比如给出糟糕的医疗建议或有害内容)的集体诉讼。对于创意领域,有关于文本使用、电影导演的诉讼,但你知道,好莱坞和出版界有很多。是的,是的。有很多集体诉讼败诉,但还没有针对伤害的。对于医疗伤害,我真的很惊讶还没有发生。但越来越多地,我认为它会嵌入我们的系统,变成那种你直接忘记的东西。它甚至不再显得新奇。它会像——如果你的孩子还能知道一点区别,对吧?就像你和我一样,在互联网时代之前就存在,对吧?但你的孩子能理解没有互联网的世界吗?不,他们不能。也许他们勉强能理解没有智能手机的世界,但会发展到它看起来只是例行公事。每当你读写任何东西时——也就是每当你与计算机交互时——AI 就会在你肩膀上方提供指导。它真的会抽象化我们与所有事物交互的本质。
Oh gosh, that is tough, right? Cuz five years in this kind of epoch look at where we were five years ago. We wouldn't even be talking about this. We'd still be talking risk prediction models, which is perfectly fine, but I always knew they were limited in their scope here. Within five years, this is an easy prediction, right? I'm pretty sure most doctors will be using ambient scribes or at least that'll be a very common routine technology. It won't even seem novel anymore. It borderline isn't already. Other predictions for the future, I just filled out a bingo card for this recently. I'm surprised there's not a class action lawsuit against Big Tech over AI harming people because they got medical bad medical advice or harmful things. There are for the creative the use of text use of movie directors, but you know, there's a lot within Hollywood and then publishing. Yes, yes. There's a lot there are class action losses, but not yet for harm. For medical harm, I'm really kind of surprised it hasn't happened yet. But more and more I think it's going to be embedded in our systems and it's going to be the kind of thing it's just like you just forget. It doesn't even seem novel anymore. It's going to be Well, if your kids will still know a little bit of a difference, right? Like you were there before the internet age as I was, right? But can your kids comprehend a world without the internet? No, they cannot. Maybe they could borderline comprehend the world without a smartphone, but it'll get to the point where it just seems routine. It's in every anytime you read or write anything, which is like anytime you have to interact with computer. AI is going to be right over your shoulder giving directions. And it'll really abstract the nature of the way we interact with all things.
是的。我认为它也会变得更安全,我的意思是,它必须安全,我认为它应该安全,也会安全,但有一个过程。是的。所以你最近有了一个新职位:AI 教育主任。我们显然是一个教育机构。我们有医学院。我们有医学生。我们有医师助理学生。我们有很多护士和其他受训者。我们有住院医师项目研究员。所以我们正在为未来培养整整一代医生,同时也提供继续医学教育。是的,是的,是的。是的。但这不是一个小工作,但我认为我们很有信心。我们很自豪,从我们科室的角度来看,你是做这件事的人。你究竟如何规划围绕 AI 的课程?
Yeah. And I think it'll get safer as well and we'll I mean, it has to and I think it should and it will, but there's a process to get there. Yeah. So you recently had a new position director of AI education. We're obviously an educational facility. We have school of medicine. We have medical students. We have PA students. We have a lot of nurses and other trainees. We have residency program fellows. So we're training a whole generation of doctors for the future as well as providing continuing medical education for Yeah, yeah, yeah. Yeah. But that's not a small job, but I think we have a lot of confidence. We're very proud that you know, from our department perspective that you're the one doing this. How do you even go about planning a curriculum around AI?
哦天哪,这是高级副院长 Rina Thomas 找上我的。我听说 Rina 说,你应该找这个,这是……然后 Thomas 博士说,Jonathan,你不是来给建议的。你是来申请这份工作的。哦,好吧。我们不是要你推荐别人。这很棒。我认为这实际上是很多主题的很好融合,对我来说做这件事并不疯狂,因为如果你看我的历史,实际上很合适。我并不是刻意要做这个。你的有机化学教育,我有点忍不住。我刚工作时有一个幸运饼干,上面说你可以在教育方面做得很好。我说不,我努力不想成为教育者,但你是教授,所以按定义你就是教育者。然后我意识到,每当我跟年轻人说话时,我总会开始教点什么。所以我也知道不只是我。我把你诊断为教授。基本上表型就在那里,我显然找到了一群人。有很多临床信息学研究员,以前和现在的。所以现在我有副主任 Don Yao、Shivan Vadak、现任研究员 Aidan Sadeghipour 和其他许多人。我不能全点名,因为太多了。整个团队聚集起来,他们都对教育充满热情。我可以给你大纲,但实际上是做课程设计、规划、回顾文献、看有什么证据、建立框架,而不是重新创造每个资源,因为我们不需要。我们可以重复利用很多。他们在这方面做得很好。其中一个最好的,我想是 Mitra Halacani。她是一个一年级学生。她在今年六月我们第一次 AI 医学教育研讨会上做了最好的演讲之一,因为她的视角——她只是一个学生。她不是来做研究、科学或当教授的。她说,“嘿,这是我发现的三个工具,没人告诉过我。”
Oh gosh, it's chance to senior associate dean Rina Thomas tapped me for this one. I hear Rina, you should look for this this is and Dr. Thomas like Jonathan, you're not here for advice. You're here to apply for this job. Oh, okay. We're not looking for you to recommend someone. And it's great. I think this is actually a very nice coalescence of so many themes and it actually isn't crazy for me to do is if you look at my history actually really kind of fits. I wasn't trying to do this. You're organic chemistry education I kind of couldn't help it. I had a fortune cookie when I first in my job like you could do well in education. Like no, I'm trying so hard not to be an educator, but But you're here for and you're a professor, so you are by definition an educator. And then I realized I kind of can't help whenever I talk to a young person, I start teaching something. So I also knew it's not just myself. I diagnose you as a professor. Basically the phenotype is there and I clearly found a crew. There's many clinical informatics fellows, former and current. So now I have associate directors Don Yao, Shivan Vadak, a current fellow Aidan Sadeghipour and multiple others. I can't name them all because there's just so many. There's a whole crew that got assembled who had education as their passion and I can give you the outline, but really doing curricular design, outlining, reviewing the literature, what the evidence is there, having a framework, not recreating every resource cuz we don't have to. We can reuse a lot. They've done a great job with that. One of the best ones actually I will call Mitra I think Halacani. She was a first year student. She gave one of the best things at our first ever AI in medical education symposium this past June cuz her perspective she's just she's a student. She's a student. She's not here to do research or science or be a professor. She's like, "Hey, these are the three tools I figured out and nobody told me."
我自己试过这些工具,它们帮我做闪卡、练习面试。我觉得这很有说服力。她展示了‘我是学生,这是我以前无法做到的学习方式’。所以你们有框架、有团队,这很好。那么你最担心哪些原则或问题,又希望确保通过课程传递给医学生和受训者?
I just tried them myself and they helped me make flashcards, helped me do practicing interviewing. And I thought that was very compelling. She just showed this is I'm a student. This is how I helped myself learn in ways that weren't possible before. So you have a framework, you have a team, which is great. And what are the sort of principles or the things you worry most about and want to make sure that you're passing on through the curriculum to medical students and trainees?
天哪,确实有一些真正的困境,我没有简单的答案,我们整个学校的课程组也没有达成共识。比如,学生应该被允许在作业中使用 AI 吗?斯坦福大学的默认答案是假设不允许,除非老师明确允许。我反对这一点:这行不通。这不能成为默认政策,因为它无法执行,而且把学生置于一种非常尴尬的过度诱惑情境中:荣誉准则说我要遵守规则,但我知道所有同学都在用。那我该怎么办?所以我们转向了:作业只是练习。我们甚至不打算批改,因为 AI 太容易完成了。关键应该是闭卷考试。但确实存在一些紧张关系,这有点棘手。我在想该在节目中说到什么程度,但我认为这些都是我们应该面对的真实问题。
Oh gosh, there are some real dilemmas that I don't have easy answers for and we have not come to consensus amongst the broader curriculum group at the whole school level. For example, should students be allowed to use AI on their homework? The default answer at Stanford University is assume it is disallowed unless your instructor says so. And I push back: that's hopeless. That cannot be the default policy because it's unenforceable and it puts students in a very awkward undue temptation scenario where the honor code says I have to follow the rules, but I know all my classmates are doing it. So what am I supposed to do? So we've shifted towards: homework is just practice. We're not even going to bother to grade it because it's so easy for AI to do it. It's got to be the closed book exams. But there are some real tensions and this gets a little dicey. I'm thinking about how much to say on air here, but I think these are real issues we should confront.
我们一些医学生的医学推理作业做得非常好,表现很棒。但到了闭卷考试时,他们却考得很差,比往年差得多。发生了什么?很明显。他们用 AI 做作业,所以根本没费心去真正学习。
Some of our medical student classes on the homework for medical reasoning, they're killing it. They're doing great. But when it came to the closed book exams, they did not do very well. Much worse than in prior years. What happened? It's obvious. They used AI to do their homework and so they never bothered to actually learn.
天哪,不。AI 可以是你遇到过的最好的老师,但如果你滥用它,你就错过了重点。作业的目的不是完成作业,而是让你挣扎,这种挣扎才是让你学习并内化的关键。整个教育体系,不仅仅是医学,整个教育,一夜之间被颠覆了,我们还在努力适应,但人类机构无法跟上技术发展的速度。我认为其中一些我们很挣扎。我们经常用计算器的例子。我们不强迫孩子做长除法,但也许我们应该;他们仍然应该学会,这样到时候他们才能拥有计算器的能力。但我认为有一系列基础医学技能,如果你一开始就使用这些模型——它们的能力远超任何计算器——这个比喻就不成立了。如果你从一开始就没学会这些技能,那么你真的会失去一些东西,最终我们到达一个我们不希望看到的地方。
Oh gosh, guys, no. AI could be the best teacher you have ever had, but if you misuse it, you miss the point. The point of homework isn't to do the homework. It's to make you struggle and that struggle is actually what should make you learn and make this innate. And the whole education system beyond medicine, just education in general, really got turned upside down overnight that we're still trying to adjust to it, but a human institution cannot move at the pace that technology does. I think some of it we struggle with. I think we often use the calculator example. We don't force kids to do long division, but maybe we should; they should still learn that so they can, when the time comes, have the ability of the calculator. But I think there are a set of fundamental medical skills that if you start off with using these models at the level they're at, which is beyond what any calculator could do, the metaphor breaks down. And if you never learn them to begin with, then you really do lose something and we end up in a different place than we would want to be.
我认为这是真的,绝对是真的。我喜欢这样看:因为我们不得不制定学生使用政策,以及我们要让受训者——不仅是医学生,还有住院医师、专科培训医师,甚至继续医学教育——怎么做?即使是主治医生,我们也在实践。我们看了其他学校的政策,我喜欢这个原则:一旦你证明了自己能独立完成,就可以使用这些工具。这是个好原则,只是很难执行。有争议:你会允许实习医生或三年级医学生用环境抄写员来写病历吗?问题是,如果他们从未自己写过——AI 总是替他们做——他们怎么知道如何写计划?我认为这很紧张,也确实有问题。我认为最终我们希望每个人都处于核心。AI 记住一切?谁在乎?为什么要那样做?我们要有判断力,我们要做患者互动,但如果你从未学过知识,你就无法对事物做出判断。
I think that is true. I think it's absolutely true. I like looking at because we had to define our student use policy and what we're going to have our trainees — not just medical students, but residents, fellows, and also continuing medical education, right? Even as attendings, we're all practicing as well. We looked at other schools' policies and I like the principle: you can use these tools once you've demonstrated you could have done it on your own. That's a great principle. It's just hard to enforce. There's debate. Would you allow an intern, an MS3 to use ambient scribe to write their note for them? It's like, well, how will they ever know how to write a plan if they've never — AI always does it for them? I think that's quite a tension. And it does have an issue. I think ultimately we want everybody to be at the heart of it. What the AI remember everything? Who cares? Why do you want to do that? We'll have judgment. We'll do the patient interaction, but you can't have judgment about something if you've never learned the knowledge.
我还担心,因为模型显然是在人类创造的文本上训练的,这些文本来自科学实验等,我们希望它们呈现世界的真相。如果我们所有生成的文本最终都变成 AI 生成的,那就成了一个循环。是的,这就是 AI 垃圾的衔尾蛇,蛇咬自己的尾巴,而且显然已经发生了。我认为很多构建前沿模型的技术公司会说,现在的问题不是获取更多数据,而是获取更高质量的文本,这就是为什么许多报纸在起诉它们。就像‘是的,我们知道你在用我们的东西,我们知道你喜欢,因为我们生产高质量写作,对吧?’我怀疑会发生的是,文本以及我们说话和写作的方式,甚至在没有意识到的情况下,会变得越来越同质化,对吧?你现在就能检测出来,对吧?当有人给你写一封 AI 生成的邮件时,你多少能看出来。现实是,有多少次你没注意到,但有时你注意到了。
I also worry because obviously the models are trained on text that has been created by humans and that came from science experiments, for example, that we hope present truth about the world. If all of our generated text ends up becoming AI generated, there's a circle. Yes, this is the AI slop ouroboros, the snake eating its own tail, and it's clearly already happening. I think a lot of the tech companies building the frontier models, you know, they would say it's not about getting more data at this point. It's about trying to get higher quality text in there, which is why many newspapers are suing them. It's like, 'Yeah, we know you're using our stuff and we know you like it because we produce high quality writing, right?' I suspect what will happen is the text and the way we probably talk and write, without even realizing it, will become more and more homogenized, right? You can detect it now, right? When somebody writes you an AI generated email, you can kind of tell. The reality is how many times did you not notice, but sometimes you notice.
嗯,这周我发现了一个维基百科页面,专门列举语言模型的标志,有些像你说的很明显。比如‘delve’、‘critical’、‘deep’这些词,还有破折号。破折号臭名昭著。我以前喜欢破折号。现在我主动删掉所有看到的破折号。我过去停止使用它们。我是我认识的唯一喜欢破折号的人。这是 AI 还是 Ashley 医生?我不知道。我再也分不清了。
Well, there's a whole Wikipedia page I discovered this week that is dedicated to tells from language models, and some of them, like you say, are obvious. There's the word 'delve' and 'critical', the word 'deep', and the em dash. The em dash is notorious. I used to love em dashes. I actively delete anyone I see. I used to stop using them. I was the only one I knew that liked em dashes. Is this AI or is this Dr. Ashley? I don't know. I can't say anymore.
这让我个人感到难过,但这是一个挑战。作为教授或主治医生,我们在很多地方需要生成文本。有时学生帮我们,但如果这些都是捷径,那么我们最终会进入一个一切都由生成的世界。确实如此,我认为这肯定会同质化我们的阅读和写作风格,因为太多内容都是生成的。但我不确定这是否是一个好的类比。也许这让我们抽象到更高的层次,从而更接近实际的思考和意义。这不是完美的类比,但想想编程语言:这就是计算机科学家喜欢它的原因,对吧?谁想在计算机里写字节码?那太疯狂了。哦,我们有汇编语言,看起来有点像单词。哦,那很麻烦。让我们用 Pascal 或 Fortran,看起来更像在写句子。让我们用 C++,用 Python。现在让我们直接用英语写,让 AI 为我生成底层代码。这是一个很好的抽象,但我认为当涉及到正常人类互动时,它以一种我们以前从未需要处理的方式侵入了这个领域。
It's so — I feel personally sad about that, but it's a challenge. There are so many places where we, as professors or attendings, asked to generate text. Sometimes our students help us, but if these are shortcuts, then we end up in a world where everything is generated. It is and I think it will certainly homogenize our style of reading and writing because so much will be there. But I don't know if it's a good analogy. Maybe that abstracts us to the level where we can get more to the actual thinking and meaning in a way. This isn't the perfect analogy, but thinking about programming languages: this is why computer scientists love this, right? Oh, who wants to write bytecode in a computer? That's crazy. Oh, we have assembly language. So it looks sort of like words. Oh, that's a pain. Let's get like a Pascal or Fortran. So it looks more like you're writing sort of sentences. Let's get C++. Let's get Python. Now it's let's just write in English and have AI generate the lower level code for me. That's a great abstraction, but I think when it comes to it, it's treading on the domain of normal human interactions in a way that I don't think we've had to deal with before.
是的,我认为这确实有点不同。
Yeah, I think that really is a bit different.
你的演讲以我最喜欢的一句名言开场:任何足够先进的技术都与魔法无异。所以,让我们回到魔法。因为你是我们认识的人中唯一一个在讲座、演讲和表演中既包含昨天刚出的最新数据——关于一个非常热门的话题,也就是 AI——又穿插了魔术表演的人。你让我们大开眼界。你是怎么开始玩魔术的?你 12 岁就开始玩魔术了吗?
You began your talk with one of my favorite quotes, which is any sufficiently advanced technology is indistinguishable from magic. So let's get back to the magic. Because you are the only person that I think any of us know who provides lectures, talks, shows that include both really up-to-date data from yesterday around a very topical subject, which is AI, but also intersperses that quite literally with magic tricks. And you entertained us greatly. How did you get into magic? You were into magic when you were 12?
有一点吧。小时候其实没怎么玩。大概五六年前又重新捡起来了。
A little bit. I didn't really do that as a kid. I picked it up again maybe 5, 6 years ago a little bit.
这是疫情期间的事吗?
Was this a pandemic thing?
算是疫情期间的事吧。当时我家老大大概八九岁。我们去看了一场街头魔术表演,有人从盒子里变出一只兔子,他的脸一下子亮了起来,尖叫起来。不知道你有没有见过孩子真正体验到那种惊奇的样子。那真的很神奇。那个魔术本身并不神奇,神奇的是看到别人有那种体验。我当时想,‘哇,我孩子喜欢魔术。我应该给他表演一个魔术,就像我小时候那样。’我给他买了一套基础的儿童魔术道具。我在给他和我同事的小儿子表演一个魔术。正表演着,我同事看了一眼说,‘Jonathan 的手法还得练练。’
It was kind of a pandemic thing. My oldest child at the time, he was probably 8 or 9. We went to a street magic show and somebody pulled a rabbit out of a box and his face just lit up and he squealed. I don't know if you've seen a child really experience that wonder. It's a really magical thing. That magic trick is not magical. It's seeing somebody have that experience. I thought, 'Oh, wow, my kid enjoys magic. I should show him a magic trick just like I had when I was a kid.' I bought him a little basic kid set. I was showing him and my colleague's young son a trick. As I was doing it, my colleague looks over and says, 'Jonathan should work on his sleight of hand.'
什么?你刚才说什么?你竟敢这么说?我表演魔术还收到反馈了?我不过是想逗你孩子开心,你还挑刺?好吧,那我就去学点手法魔术,再多下点功夫。这成了跟学生互动的好玩方式,也让我看起来像个书呆子教授,但又平易近人。后来在新冠疫情期间,这事就一发不可收拾了。大家都关在家里,什么也做不了。有人学烤酸面包,我花了整整 3 个月学一个非常高阶的魔方魔术,今天在台上只展示了一小部分。然后就失控了。我参加了好几场比赛,还赢了奖,甚至靠表演魔术接过付费演出。
What? What did you just say to me? How dare you? I'm getting feedback on my magic? I'm just trying to entertain your child and you're giving me this flak? So I will go learn some sleight of hand magic and do a little bit more with it. It was a fun way to interact with students and make me look like a nerdy professional, but approachable too. And it kind of spiraled out of control during the COVID pandemic. We're all locked indoors, we can't do anything. Some people learned how to bake sourdough bread. I literally spent 3 months learning how to do a very advanced Rubik's Cube magic trick and I showed only a very little bit of that on stage today. And it's spiraled out of control. I've been in and won multiple competitions. I've been paid gigs just to perform magic.
你在拉斯维加斯表演过,我记得没错吧?
You performed in Vegas, I do believe.
哦,没错。是的,我确实在拉斯维加斯的主舞台上表演过。那是最近在斯坦福的一个会议上,一个新颖的关联。不是我刻意追求的,但它挖掘出了一些童年的志向和内心的小孩,而且我觉得这也有助于我更好地做演讲。玩魔术能让你学到很多同理心,因为关键不在于我在做什么,而在于我必须回答:‘你在我做这个的时候在想什么?’因为我不想让你想错方向。我想让你想的是这个,对吧?所以你必须真正理解另一个人在说什么、在想什么。‘如果我在这里挥一下手,就会让你做某件事。’引导叙事、设定预期,我觉得这是一个非常强大的组合。几年前,大学医疗合作伙伴邀请我做一场关于 AI 和医学的主题演讲。本来是那样,但他们那周的主题是‘医学的魔法’。我想,‘其实我也会变魔术。能不能作为附加节目?’这要怪我妻子。其实我要感谢我妻子,因为她说:‘为什么不呢?为什么要把它们分开?你可以把它们结合起来。如果你在演讲中间穿插一些魔术呢?那太疯狂了。谁会那么做?太荒谬了。他们会把我笑下台的。’
Oh, that's right. Yes, I have performed on the main stage in Vegas before. That was a novelty at a conference recently at Stanford, that connection. Not something I was trying to do, but it's unearthing some childhood aspirations and inner child, and also it helps me better with presentations, I think. You learn a lot of empathy when you do magic because it's all about who cares what I'm doing. I have to answer, 'What are you thinking while I do this?' because I don't want you thinking the wrong thing. I want you to think of this thing, right? So you have to really understand what another person is saying and thinking. 'If I brush my hand here, it's going to make you do something.' And directing this narrative, setting up expectations, I found that a very powerful combination. A few years ago, the university medical partners invited me for a keynote on AI and medicine. That was supposed to be, but they also had the theme about magic of medicine for that week. I thought, 'You know, I actually can perform some magic too. Can I do that as a bonus?' And I blame it on my wife. Actually, I'm thanking my wife because she said, 'Hey, why not? Why is that separate? You could do them together. What if you had some magic in the middle of your talk? That's crazy. Who would do that? That's ridiculous. They're going to laugh me off the stage.'
但如果有一个主题上的联系——我觉得确实有——那会非常有趣,也非常有说服力。尤其是在生成式 AI 兴起的当下,你分不清什么是真的了。这就是图灵测试。那是真人还是聊天机器人?不,那张图片看起来太真实了。那段我的视频全是假的,全是 AI 生成的。所以我再次用魔术作为那个算法。哇,看起来是真的吧?但你仍然需要判断力来分辨真假。这是一个非常强大的组合。Jonathan,我们非常高兴你能加入我们系。我很高兴来到这里。为你以学术方式所做的工作感到自豪,真正带领我们安全地走向未来,当然还有教育和下一代。很高兴来到这里,不管怎样,我认为作为医师科学家,这是我使命的一部分。我很高兴能在斯坦福医学院接触到我们所有的人,也尽可能接触到更广的人群。感谢你参与我们关于医学未来的讨论。谢谢你,Ashley 博士。
But if there's a thematic theme and I think there is, that could be very fun and could be very compelling. Especially in that prior area with generative AI emerging, like you can't tell what's real anymore. It's a Turing test. Is that a real human? Is that just a chatbot? No, that image looks so real. That video of me, that was all fake. That was all AI generated. And so I'm using magic again as that algorithm. Boy, does that look real? But you still have to have that judgment to tell the difference. It's a really powerful combination. Jonathan, we're so happy to have you in our department. I'm very happy to be here. Proud of the work that you've done in an academic way and really bringing us into the safely into this future and of course in education and in the next generation. So glad to be here and I for better or worse consider part of my mission as well as being a physician scientist. I'm so glad to be in Stanford Medicine reaching all of our people, but also reaching beyond to everybody else that we can. Thanks for joining us on the future of medicine. Thank you, Dr. Ashley.