Solving AI Hallucination in High-Stakes Medicine — Dr. John Ferguson Explains
In this episode of Healthcare Business Growth Conversations, we are joined by Dr. John Ferguson, a surgeon with over 25 years of independent practice and the President-Elect of the American Board of Facial Cosmetic Surgery.
This is not a conversation about new tools or automation. It is a deep dive into how medical standards are defined, enforced, and protected at a time when technology is challenging the systems of authority behind them.
Dr. Ferguson pulls back the curtain on what he calls the “broken engine” of medical boards, a system that spends hundreds of thousands of dollars on manual validation, arbitrary failure rates, and psychometric rituals, rather than measuring true mastery. He explains why the future of medicine must shift from “mined” diamonds to precision-manufactured Content-Driven Intelligence.
——————————————————————
About the Guest Dr. John Ferguson:
A practicing surgeon with 25+ years of clinical experience, board leadership, and exam committee governance, he brings real-world accountability to healthcare AI. He currently serves as President-Elect of the American Board of Facial Cosmetic Surgery, Co-Chair of Written Exam Committees, Treasurer and Trustee of ABFCS/ABCS, and sits on the Board of Governors of the American Academy of Otolaryngology–Head and Neck Surgery (AAO-HNS).
He is Editor-in-Chief of StatPearls, overseeing validated medical education content used by thousands of clinicians, an AI Ethics Columnist for the American Journal of Cosmetic Surgery, and a Contributor to The AI Journal. He is also an AI2030 Global Fellow, working on healthcare AI governance and the Algorithmic Integrity Program.
After witnessing boards spend $100K+ annually and waste hundreds of volunteer hours while psychometric systems punished mastery, he founded EdAI Systems—an AI assessment platform built from inside board leadership, not as an external vendor. EdAI has achieved 100% customer retention across multiple medical specialty boards, reduced volunteer coordination time by 90%, and deployed 2,400+ psychometrically validated questions used in high-stakes certification exams.
His work is grounded in content-controlled intelligence — AI constrained to curated, validated sources (StatPearls©), designed to say “I don’t know” rather than hallucinate. This approach is informed by chairing exam committees, paying malpractice insurance, and being accountable when systems fail.
Awards include the 2013 William K. Miles Award for the highest score on the ABCS Written Exam.
Mission: Scale expertise without sacrificing reliability. Build AI accountable to the professionals who face real consequences. Optimize humans, don’t replace them.
——————————————————————
Linkedin Profile: https://www.linkedin.com/in/john-ferguson-md-facs/
#AIInHealthcare #HealthcareAIGovernance #MedicalAIEthics #ContentDrivenIntelligence #AlgorithmicIntegrity #HealthcareStandards #BoardCertificationReform #FutureOfMedicalEducation #HealthcareLeadership #PhysicianLeadership #SurgeonPerspective #HealthcareAccountability #ResponsibleAI #ClinicalDecisionMaking #PatientSafetyFirst #MedicalInnovation #HealthcareSystems #HealthcarePodcast #BusinessOfMedicine #HealthcareTransformation #EthicalHealthcareAI #MedicalGovernance #HealthcareRegulation #HealthcareTechnology #PhysicianFounders #HealthcareStrategy #AIInfrastructure #ContentControlledAI #MedicalAssessmentReform #HealthcareMarketingAgency
Full Transcript
Introduction to AI and Medical Standards 0:00
Welcome to another show of healthcare business growth conversations. Today, I'm joined by Dr. John Ferguson. He's a surgeon with over 25 years of independent practice, the president-elect of the American Board of Facial Cosmetic Surgery and a global fellow working on algorithmic integrity and accountability in medicine. I want to level set for everyone listening, this conversation is not about new tools or threats. It's about how medical standards are defined, enforced, and protected at a time when technology is beginning to challenge the systems behind them.
Most conversations around AI and healthcare focus on efficiency, productivity, or automation. This is what this is about. What we are really talking about today is medical truth. who decides what competence actually means, how that decision is validated, and what happens when technology starts to pressure the institutions that have historically held that authority. Dr. Ferguson, thank you for being here. Hey, it's a pleasure. Very nice to be here I know it is 3 a.m. for you in Hawaii. I appreciate you being with us.
Yeah, It's one of those sacrifices for living in the middle of the Pacific. It's worth it. Oh, but you know, very few physicians ever get to see how these systems are designed from the inside. You have spent years operating both as a practicing surgeon and someone responsible for maintaining standards. That perspective is where I want to begin with. Sure. So, let's look at the broken engine. you have seen boards. Yeah, that's what started everything. Yeah. So yeah, so you're seeing the boards invest significant capital into manual exam certification.
You're talking about volunteer burnouts and all. Yeah. Well, see, I'm a professional test taker. That's most doctors are. You know, you got SATs, MCATs. We got all exam exams. I mean, that's it is. Professional test-taker. Just kind of turned into a professional test maker for the boards the board exams and you know, I started making it I starting with making questions helping out the Boards and It's it's tough. I mean it really is to it. It''s not like I Mean, well, i guess it is kind, of easy you take something and Boom, okay, I'm gonna make a question out of that sentence.
That's not right. You know, that's just not how you test somebody. So you gotta put a lot of thought into it, especially when you're testing competency. It doesn't matter if you know okay their blood sugar is 283. No, you got to understand what the ramifications are. You got it. I mean, it's just not pick the pick, the factoid. It has to be applicable. And you have a finite time frame and a number of questions you can test people. So it takes a lot of effort to make a valid question anywhere for I've spent days just on one.
Now multiply that by three or four hundred that you have to have every year for an exam. and then try to corral a couple of surgeons, maybe two or three, to help you write the questions in time for the exam. And you know what? They all have the better things to do. As soon as I was kind of responsible for putting the exams together, I found out how hard it is to to coral a bunch of surges to anything. other than operate. They don't want to do anything else. So I found myself doing all of these exams or all these questions.
But it was, you know, but you get burnt out. You know you put all this time and effort. you're doing this work and you just get burned out so One day I started playing with AI at, just said, okay, let's put it in there, make me a question. Beautiful question, beautiful. I just typed in, we've got a patient that's coming in with a hematoma after a facelift. Boom, what do you see? And again, this all beautiful question and it was completely wrong. I don't know where that AI got its data from. But it was like nothing that we would do.
So that's when I started applying, really focusing on the content. Not so much the question first, but the concept. We'll get into this later. To me, that is the basis of AI. AI is great at what it does as long as it has the right content the rat data You could have a Lamborghini, but if you don't fill the gas tank with good gas ain't going very far so, you know, I started focusing on content and that really was the key to making really good validated questions. And it really helps, uh, make a difference in time.
I mean, with my, w with, um, With my software may I can, okay, this is going to be a little, I had a contest one, or somebody just kind of bet me how many, how I could make questions I did in one football game was Dallas versus Philadelphia last year. We got our butt kick. I'm from Dallas. And during that game, I made 1,052 questions. That year, 100 of those were actually Yeah, about a hundred of those were on the exam. So, you know, it makes a big difference. And those are all validated questions.
I mean, we go over them in committee and make sure, because there should always be, I hate saying human in the loop, and we'll get to that later, but should be AI in a loop. But we're testing competency and AI doesn't operate. It's humans and doctors. We go through every question, 94% of the questions never need any editing. That was kind of a key moment that brought everything together as far as the validity of process and the question. Let me pause you there for a second. Sure. So what you're describing sounds less like a cost problem, more like it is a problem.
Well, yeah, I mean, that's the problem I've been that the whole thing. The system is screwed up. And it's screwed up in multiple ways. There's a whole industry on questions, just questions. And that's it. They're people that have PhDs in writing questions or validating questions and they charge a lot of money. I've been treasurer for a couple of boards and I signed the checks for fifty sixty hundred thousand dollars a year for these these experts to Validate our exams. We're still writing it. Were still doing all the work and they're just validating whether the question is Strategically or tackling whatever correct or valid based upon Statistics.
This gets into the healthcare medicine thing, but I don't treat statistics. I treat patients. The person on my table is not a statistic. That's a patient. So matter what on a question, what the person, the doctor next to me answered on that question.
Question Writing, Validation, and Content-Driven Intelligence 8:43
It doesn't matter. And that's how they were validating questions. Placement, my ranking. I have sat in so many rooms and you have this expert, the psychometric expert comes in and say, okay, in order for your test to be valid, for these questions to to, be a certain number of people need to fail. And so you, have to pick an arbitrary point. I'm gonna, you know, and that point right there determines whether this person operates and this an arbitrary number and god forbid you have that that so many if it's too many people passing the test then your test is too easy.
If you don't have enough people, passing your tests that oh it is not the tests your education your training sucks. So it was an unfair system. I got caught up in it once where it affected me. And so that's when I decided to really start focusing on not just medicine and ward certification, but actual the whole process of competency assessment. Right now, assessment, examination, it's all question-centric. It's been that way for eons. Well, not eon, because Socrates had a completely different idea that we'll get into in a little bit, but it is like a diamond.
They traded like the hole diamond industry. You have somebody that goes and they mine that diamond, and it very labor intensive. Uh, they carry it across the board or whatever whatever leonardo capri did in that movie And then you spent hours And days and months and weeks just trying to to refine it and make it beautiful and you know and then you put it in a vault and hoard it so that the price stays, so it stays valid. It stays artificially inflated in value. Safe, because if it gets out into the world, it loses its value, what's the difference between a diamond that you get out of the ground, you polish it, and you do all that kind of, shape it.
And when you manufacture, What's The Difference Besides De Beers? There is no difference. In fact, for manufacturing purposes, manufactured diamonds are better because they're flawless. They can be used in very exact manufacturing processes that need diamonds. There is no room for error in some of those processes like you would get with a diamond that it goes through so many iterations to be able to say, oh, yeah, I'm ready for an engagement. So that's kind of the approach that I took. AI made it made possible once you kind grasp how the questions are how they...
Okay, the psychometric specialist, they have a formula. That's how we do. They come in, and they had no idea about rhinoplasties, about, you know, a radix graft and how it goes up areas. You know they, have no ideas. How can they read these questions? Well, They have the formula that they go by. Okay, stick that formula into AI, put your content in and there you go. And that's what I did. I manufactured those questions. The key though was really being able to get the AI to focus on the content. So instead of question centric examination assessment, it's more content driven.
If the value is not in the question, the values is in content, And that's what it should be You know a teacher shouldn't be punished for Being a good teacher, you know, they shouldn A good surgeon Or a doctor or clinician Shouldn't punished because they're not a test taker. They should Be rewarded for their knowledge and Microsoft proved that the current assessment, most tests nowadays, It's not knowledge, it's test-taking skills. And I know that because I got test taking skills, I can go in and not study for a test.
I didn't study before the SATs and I was a national merit finalist because there are ways of getting around that with not having the knowledge but taking the test and Microsoft proved that AI can do the same thing. And so you really need to. It does. You know, making that question is very important. If you know how to do it, you can cheat the system. So really, it needs to go back on the content. That's what we did here. Something called content driven intelligence making the questions is was really easy.
It was figuring out a way to isolate content so that it stays there that AI doesn't go drifting out and pull something in and that's That's actually the cut the the patent that I got 13 months ago was basically that process. Not so much making the questions, but being able to process the content in a way that it can be used by the AI without worrying about hallucination or drift or anything. Yeah, definitely. We'll get into the hallucination part later. I just want to take a step back and then say, you were spot on whenever you took the example of the gold being green versus gold, being manufactured in the lab.
there is no error whenever you magnify it. Exactly. At the same note, if a system assumes failure is necessary to prove legitimacy, then excellence actually becomes a threat to the system, right? Exactly! That is spot on. And that's where the paradox arises in education. It's become such a competition. God knows, I fell into that. I mean, if I wasn't the first one done with the test and didn't get the highest score, it felt like I failed. That's not important. You know, that's that superfluous to what you're trying to achieve.
But that was rewarded. And it wasn't so much, you know the knowledge, it was just being able to take a test. You know, I mean, and some people can't take tests. Some people aren't good at it. So people need to think and you know they know that they have it up there, but bringing it out takes them a little bit longer. And you have to have a time limit. just for logistics purposes, but some people need a little bit more time. Some people don't, some don' t process the same way as they read. some People don t do well in a big room full of other people.
I mean, there's a lot of variables and anything other than being really good at taking a test, you were penalized. That's not fair. Not fair at all. I just want to go a little bit deeper into it. Sure, sure. The mastery paradox, I would say. When candidates perform too well, psychometrics often conclude that the test was too easy, as you said. So why is this logical fundamentally backwards in a profession where mastery, not a bell curve, should be the goal? Well, that's, you know, That's when how you assess is the focus and not what you're assessing.
That the problem. All right. And that's what has happened is it's the process that that what people concentrate and it is a big huge industry. Of validate of how you know making sure that your test is validated and It's insane. I mean we pay consultants on boards hundreds of thousands of dollars to Validate our tests and and it's not because our test are bad is because it is expected This board does it that way, so we got to do it this way. Their test, it goes every two years. They sit down and they do the psychometric reevaluation every 2 years, they must be great because they're going over their question bank constantly.
Dude, why don't you just write questions? Just do the questions. You know, the the question's that whole process should not be the focus. All right. It's the content. That should be the focus and it's not because the whole manufacturing of the assessment material is so onerous and time consuming and resource consuming that A whole community of businesses developed around it just to support that and they have a stake in the game now. They're they'll help perpetuate that also and you can't blame them.
I mean, imagine if you went to school. A PhD. Jesus, that's years and years. There's that fear that AI is going to take your job. Well, you know what? That's in a way that's kind of it, if that is what your job was. But now it's time to pivot and support AIs to make better algorithms for the AI to ask questions. Don't try to perpetuate a system that no longer necessary and is detrimental. Yeah, so I just want to go a little bit out of the scope and then ask you one question regarding this. Okay.
Because the depth was good, I couldn't control myself. Whenever everyone starts, depending on the AI, for the pattern recognitions, or for example, to be in alignment terms to do the basic research work. So are we training the next generation into algorithmic thinking, meaning like, you know, can we notice everyone's cognitive ability is as good as this AI? Absolutely. It's better. AI frees you up. You know what? Do you want me? I have a certain I have a certain skill set that I've developed over time.
But no, is my time better spent in the OR or making a question? I mean, I can do both. I, can't go in there and operate. It can make my question, where do you, is me as a resource. Where am I better utilized? Yeah, I can do it. I could make a question. It's going to take me four or five hours, 40 hours whatever, but I'm going make as good a questions as the AI. But that's time that I couldn't be learning something else. Helping somebody else, operating, surfing, petting my dog, whatever. It's freeing, it lets me do more.
I can concentrate more, I could take that time that I've got writing a question and creating new content. You know, that's what it should be. It should focused on what you're testing, not how you are testing it. We know how to test, we know to do that. we got it, it's time consuming. I mean, a washing machine, have you ever washed your clothes by hand? It sucks. It's time consuming. it's awful. And you know what? A washing machine does it a whole lot better. Did we give up any cognitive or physical abilities by letting the washing machines wash our clothes?
We still got to put stuff in there and decide what's in here. If you screw up, everything comes out pink. We've still gotta put the laundry detergent in. God forbid you put bleach in it. Who's gonna take it out? Who is drying it? Whose folding it, who buys the clothes in the first place? All right. It's not the washing machine's fault. The washing machines didn't make us lazy. Its just doing something that we can do quicker and it frees us up to do whatever. Take the dog for a walk. Read a book.
And it is. That's a tool. Yeah, totally understood. So let's sit with it for a moment. Okay. Once you recognize that inversion, you can't unsee it, right? So at that point, the issue stops being theoretical and becomes a leadership problem. We're seeing that. I'll get into, I just got to throw this out there because this came up in the last day or two, especially in psychology today, they came out and we're reaching this plateau of everybody starting to listen to the AI engineers now. The leaders, the C-suite people who just see the money and they want to have a one-size-fits-all everything, that leadership is coming to terms.
Well, reality is sitting in and the people on the ones that are doing all the work are bringing reality back in.
The Mastery Paradox and Assessment Reform 24:58
And that's kind of what it takes what I'm kind of facing with the boards. I mean, this is, you know, first, it is a money making thing. There is money to be made in some of the larger boards, so there's that inherent interest, but there is inertia. There is we've done it this way for 150 years. Let's we can't stop. We can do it. And you mentioned before that there's this stigma or abusing AI that it's making you lazy, but it is doing the thinking. You know what? You got your phone. I don't know any phone numbers anymore.
Does that make me dumber? No. That, I got a lot stuck up in here. That's valuable real estate. I don't want to burden it with a number. Um, II wanted to do something else with that idea that this is going to change something, that it's going make you lazy. And we've doing it this way for so long because it works. Well, no, it doesn't work. You know, we, use incandescent light bulbs for years and. It's all we had, you know, people don't want, well, we've had LEDs forever, but they weren't as bright, they were as warm, and they shouldn't, or they couldn't work as much.
Well, no, because they're not wasting so much energy making heat, But people want to stick with those because it, tradition, I don' know. And before that, They were afraid of electricity. Let's light everything up with fire. That's safe. But no this electricity is going to kill elephants, So, you know, that's kind of what we're facing in boards. I'm at that middle stage. Well, I like to think I've been in that little stage of my career when I am not the old, professor, this is how it is, and not in that where I have enough knowledge, new knowledge but still am ready to learn.
I hopefully aged isn't a reflection of that, but as those, my colleagues that are in this same field and embrace technology, we're gonna, it's slowly changing. When I was a resident, laparoscopic surgery, there was only one attending that taught laparscopic surgery. Now there's usually cries out in surgical residencies because they don't know how to do surgery without a laposcope. It's, it's you know, there's inertia and it wasn't because the scope was a bad thing. Obviously it was it is because it it something different.
It was something new and you kind of get scared of something. People are changing. And those of us that, God, I feel old when I say this, but those are who are moving into those positions of respect and all that. You know, we're we are able to break some of those ideas with us and I have to admit there there's a couple people that are much more experience, smarter, older than me, that know this a whole lot better. It's always kind of cool who gets it and who doesn't. But it's, you know, it, well, not even a race, a marathon.
People are coming around and realizing that it is not the boogeyman. It's just don't make it better. On the lines of those responsibilities, as you step in to the president-elect role for the ABFCS, how do you balance the mandate to modernize certification with the need to protect institutional legacy and patient safety? First and foremost, anything that we do up here, patient safety, everything else is way down here. That's the whole idea behind this, is patient's safety. Do you want somebody operating on your grandma that can take a test, or do you somebody that knows the material?
That patient is safe. You know, do want someone that could memorize all of these equations and all that, for somebody who understands those equations. That's what we're doing. Explore it. We're able to tap in those boundaries of knowledge in a more fair, equitable, thorough, I mean just better way. This is allowing us to better to assess physicians better, assess their knowledge, but in a way that isn't so resourced opinion. We can't do an oral exam on every single question. and tap some individual's knowledge.
We can't do the Socratic method on every single patient or every person taking a test, every doctor for every question. It's too resource dependent. But if you can create an assessment item and a question that really stretches every aspect of your knowledge, being able to recall a little information, recognize what's going on, and analyze the situation, process it, and then synthesize something out of that. That's the process that we go through every single day, every signal moment when we're seeing a patient.
The patient doesn't walk in and say, I've got this wax block here that's irregular. It's been there for six years. Its changed dramatically over the last six months and it's bleeding. They don't walked in like that. Didn't a walk-in and with a diagnosis, oh yeah, melanoma, gotcha. No, you gotta explore. You gotta do all these things. To make a question like that is very, very labor intensive. It's resource intensive, it's doable. You know, just like you can make a really, really pretty diamond out of this rock that you grab, get out the ground, or you didn't just get a machine to make it for you.
It's the same. And it serves the the purpose. Doesn't matter how you get it. Just it does that. That's I guess that that's, you know. That's what makes this kind of not life change. It's not revolutionary. it's just freeing us up. And that's, isn't that what moving forward, modernizing every generation, it gets easier. Well, yeah, It is kind easier, you're making it easier so that you can spend time on the next iteration, something that expands upon that. You're not wasting your time Pushing the the plow across the field Oh God, let's put a mule in front of that.
Okay, that frees us up. We can make more. Oh, Let's get a tractor. Ok that that really Frees up so that we can build a city and somebody can be a blacksmith and Somebody can a trader and someone can learn to sing and entertain those people you know, it's every thing is You know, it's not to take away from society. It's to make it better. And, you know. I know this is going a little bit more deep than just making a question for an exams, but it shouldn't be scary. Its goal is to Make anything is To make us better Uh, give us something that we know how to do.
AI doesn't know to anything that. We don't. Know how. To do, it just does a lot quicker. You know, I can do it. I. Can I, can dig a well. Don't want to take a while. Just get some machine to. Do it mean. It's that easy. Should be scary. Yes. Oh, uh, that's deep. Uh, so I want to reset the program slightly. Yeah. Just to go deep now it's four o'clock in the morning. It's a good time to go deeper focused on this topic. So, as I said, I want to reset the frames slightly here. We have talked about what's broken structurally.
Now, the question becomes what assessment should actually mean if outcomes have the goal? Starting from first principles, if the traditional exams measure recall rather than competence, what should assessment actually be measuring if ultimate goal is improved patient outcomes? Well, now you want to go into the whole Bloom's taxonomy. And actually, there's somebody, Annette, somebody who has the anti-Bloom trademark, anti Bloom thing. Well I mean, you know, the philosophies and the, whole, specialization of learning.
I I grew up with teachers. That whole process of learning stuff and what you're testing, that's first year teaching school. Those of us that didn't go teaching, it's really It's almost like a black box, but once you're in there, I mean, it all kind of makes sense. It just is formulated and regimented as playing a game or writing code. I'm lucky because both my parents were teachers. And I grew up grading papers of all things and writing tests, helping my parents write the test. Now they were college professors, so it wasn't, you know, as much as it is K through 12, but there are, it's not, recall doesn't really mean a thing.
It's moving beyond that. uh gathering information there is Analyzing information. It's again synthesizing information and and really for medicine, uh, it's gathering processing and Synthesizing doing something about it in medical school. They teach us first of all how to talk Medical school is 50% language school. There's a reason. There's no reason for us to talk that way except it makes us sound smart and we can communicate with each other. Medical school is learning a new language. The other part of medical school, is gathering information.
Data collection. And that is a testable skill. All right, that is, it's a teachable and testable skill. Beyond recall, there is gathering. You gotta know what to gather. you gotta be able to elicit that. That's one aspect of testing. Another one is being able recognize the significance of what you gathered. that's testible, That part of medical school. You got to be able to process all of that different data. That's a testable skill. You've got put that into a cohesive idea like a diagnosis, which is a Testable Skill.
And then you gotta figure out what the hell to do about it. You gotta synthesize something. Those are all testable skills. And that's just Bloom's analysis that any first semester teaching student learns. But they're all tested and they are quantifiable. They are pretty formulaic also. uh you know we in in medical school they you were were taught differential diagnosis differential you got to come up with a differential diet you're creating a different before you do anything do that differential diagosis you, know what that is that's turning medicine into a multiple choice test All right, that's what you're doing.
You are creating a multiple choice test out of that case. you've got your differential diagnosis. Then you got to pick which one it is. That's, what we're, doing we are turning, we, are taught a diagnosis, making a, diagnosis is making, a Multiple Choice Test. I mean, that's what it is. So yeah, it's easily teachable and it easily testable. But that, you know, to be able to write an assessment that touches upon all of those kind of gets, once you can imagine how hard it to make a question that stem that encompasses all that and within four choices, And you have to go through so many thought processes from here's the data or here is this.
Reducing Hallucination Through Controlled Content 40:38
The funnest things, what I do with the exams, with questions in my program is case studies. They give you a case and then have two questions on it. And then it throws something at you. And then you have to respond to that or it asks you to gather data. And if you didn't gather the data, you can't move on to question three and four. I don't write those questions. I wrote the formula, AI does that. Just tell them what, I just give them a paper, write case studies, get case, uh, studies protocols on that, make it four, six questions, um, put the, the confounder here and it fills all that in, because it's just all formula.
Um, but, each, Each aspect of learning and assessment learning is It's objective. It is not subjective. I mean, it's purely formulaic at all. You have to change the variables here and there, but it is just an algorithm. Then that reframes the entire purpose of testing. Yeah, I mean, it really does. It's it's not the test anymore. All right. You don't have to focus on the tests. Teachers shouldn't be having to write these big old tests, and it is not necessary. They should be working on their content.
Spend that time that you're writing a test or making a task teaching. Or making new content that I mean it's and that's what Progress is that what being modern? That's? What progression is it? they can find you know having a problem fixing it and You don't have to deal with that anymore and you can move on to something more important like teaching instead of writing a damn question or a test. Spend your time teaching. We just wrote a grant proposal for K-12 using the content-driven intelligence to free up elementary school teachers, high school, teachers to, to free them up from having to spend so much time writing questions and more time teaching.
Yeah. It goes, don't worry about assessing somebody. Worry about preparing. That's what it should be. Okay, so if recall isn't the objective, ranking physicians against each other starts to feel the wrong problem. So staying with that shift, you have moved from question-centered testing to content-driven intelligence. Why is anchoring assessment in validated medical content more defensible than ranking positions against one another? I'm a perfect example. I got into medical school with a 2.8 GPA.
All right. Yeah. But I missed one question on the MCAT. Okay. It mattered because it was a measurable object. It was very easy to measure rank. It's very easy. So I did an ENT residency, notoriously difficult residency. There are like three or four residencies that they actually do early matches for because half the people don't match and they need to figure something out. one of those is E&T. One of the persons that was also first in my class, because there was a few of them, she didn't even match an E & T, and that's what she wanted.
So, you know, that is a... that helps those program directors and residency coordinators. That's a badge for them. We get all those first, those that are firsts in their class. Hey, I was first at Harvard. Harvard, what the hell? I got just as good of education, but we rank these things. I mean, it's a competition. It's not assigning a person's unique value. Uh, i-it's about them compared to others. Why? In the OR, somebody that's just on that side of the bell curve, you don't want a bell a binary.
Can they do it? Can There shouldn't be any of that. I don't know. We live in a competitive society, and it's an objective measure. And I, for one, am extremely guilty of being a very competitive person. It's taken a lot. Yeah, I pretty much reached my goals in what I wanted to do. Now it is easy for me to say competition isn't important. But looking back, it really isn't. It's what you know as an individual. Be judging competency against somebody else. We should be judging competence on error rate.
That should important. Now we get into, OK, do they have one error per 100 or one there? But unfortunately, we live in a competitive society. genetically, it's the most competitive winner gets to procreate. It's been like that for billions of years. So it is going to be hard to get out. But when you are talking about the difference of somebody standing over your grandmother with a knife, you know, this should be a binary. Do you this or do you not know this? Can you take a test? Yeah, that gets a little deep.
This is really, this is, I'm passionate about this. I mean, it's, yeah. You know, so Pasha I am willing to give up a career as a surgeon to pursue this because it is important. In medical schools, so many people that would be such incredible doctors, incredible, but they can't take a test to save their life. And I would have any one of them be my doctor, But they couldn't pick a task. So they didn't survive. No, that, no, the implication of that is bigger than it sounds. I live it. So, like, you know, everything you described works only if truth is stable and controlled.
In medicine, even a small halization isn't a feature, it's a... You can't. That's what, yeah. I see where you're going with the content control. You gotta, you got to control the contact. you can put some gas in that Lamborghini, but if you sprinkle a little bit of sugar in there, it's not gonna work. Gas is great, But if, if have a bit, something a litle spoiled in it. And it is the same. Everybody's, that's the thing with AI right now. everybody's worried about drift and hallucination and They're working on increasing the ability to predict wrongness.
What the hell? I mean, that's crazy. Reducing error rates so that it's 0.001. They are neglecting the primary cause of it in the first place, which is easy enough to treat. Don't put in bad data to begin with. Control the content it gets a little You know, it can be a bit more resource intensive but you know and and it's fine you can have a Little bit of hallucination if you're having AI write poetry or copyright for I don't know some journal but There's there's no room for a little bit of error when I keep on doing this There was a whole long story behind the grainy thing, but anyway, you know If your error rate is one tenth of a percent in the OR, you know, that's one out of the couple thousand grannies that die.
That's not allowable. So the goal is it to reduce error The goal is to not introduce it in the first place. And so you control the content. I can tell you that was the hardest thing. It's being able to focus the question on the actual what you're offering it. to use, but still have a little bit of flexibility to make every question different, every scenario different. And it's not just, okay, make this one 25-year-old patient and this 60- year- old, I mean, everything needs to change. So you got to have that flexibility in the actual creation, but the overall insight or pearl or gold that you're trying to assess has to be 100% because it's got to 100%, right, when you are in O.R.
I mean, that's just the way it is. And so you don't introduce content. It's kind of like Plato's cave, you know? I mean, those people staring at the shadows on the cave wall, were they any, whatever they described, It's right. In their mind, it's completely right, and there's no hallucination, there is no drift, no guessing. because they don't know anything else. All right. I got into a big loans philosophical discussion yesterday with somebody. Well, you know, we are algorithms are when we process it's processed by three different algorithms and then they come together and fight about it.
But whoever wins. You know what? Just don't give them anything to guess about, all right? Take that away, you know? I mean, sure, it's inference. That's how we work. We work in inference, but you can't guess something that was never there in the first place. So, just control the context. Well, this is interesting, okay? Like, now I want to go deeper into this, like, you know, maybe a step deeper. So staying with that risk, whatever the risk that you said, I wanted to talk about reducing the gap, the hallucination gap.
I know it is not perfect, but at least let's talk reducing that gap It can. You can reduce it to zero. Everybody rolls their eyes. Yeah. Okay. You're only as good as the data. So if I put in data, it's not true. Then I guess is that an hallucination? I don't know. You got to put some guardrails out there. And I know I'm not going to convince you, but. I just keep on going back to Plato's case. So on the same line, what design principles actually reduce hallucination, maybe to put it in your terms, make it zero, in high-stake environments as opposed to just hiding it?
Yeah, so that's what, we'll get into that a little bit later when we talk about healthcare and medicine, but everything that I think about when I do healthcare or medicine stemmed from this, from making these questions. It was so important that when I'm making a question that it doesn't have anything, it can't hallucinate. It can make anything up. One of the golden rules of making the question is everything about that has to be in that content. You can't go anywhere you can get in it. Everything has to be right there of and if it's not right There then it is not a valid question on that material.
It's got to all encompassing So how do you get that? To do that without just you know picking a fact and saying there it well you take that contact and you chunk it. That's what it was. Before I even knew what chunking was, I would break it up into different pull out the ideas that they were trying to, the goals of the article, The objectives, learning objectives. You pull those out and then you have those as keys and you chunk them physically, like what's around them.
AI as Infrastructure and the Limits of General Models 56:28
500 words before 500 Words are characters afterwards. Where is it? visually on it, is it in a table? Is it a traditional area that's the results? If you're talking like a journal article on experiment, you do that. How is, this is kind of crazy, but is a bulleted? is that italicized? Chunk it like that when I when i when, I break it when? I get an article i break i broke it into so many different ways of looking at it so that There's enough that it can create something without going somewhere else but not making it word-for-word out of the article and the really key I think that I just I don't know what I lucked into it was don' layer it.
Don't have multiple layers. When I would chunk it by number of characters back and forth that's one index. If it's by position in the paper that is another index And then where it is visually or how it does, that's another index. Those act independently. They don't cross talk at all. And that there's a layer right above it. That's the layer that makes the question. There's nothing in between. There is no real room. And that just came out, what, Google research this week. The more layers of agents you have, once you get past two or three, you're done.
Hallucination, error, it outweighs any validity whatsoever. That's kind of how I did it. It was twofold. One was making sure that wouldn't introduce any error. And second, it was really so that I wouldn t run into copyright issues. Because you can't copyright an idea. So I mean, and it kind of all worked from there. It's going to be a story that lives in my family forever because last year, Thanksgiving, we went to my daughter's new house in Alabama. And I sat in front of that laptop for three days at the dinner table and everything and trying to figure out how can I do this?
The night before we were going, it just kind of clicked. And that's when it made the first question. It's where I processed the 1st one. Made the actually 200 questions in about 20 minutes. And it worked, but it all comes down to controlling that content. It is labor and it's resource intensive. You've got three different, every one of those indexes, you have to have an agent that chunks it and then puts it in there. And then you'll have one that is pulls it out and presents it to, and those are done.
I mean, they're not making any kind of major moves that, but if you give it too much leeway, too many, if we give that agent too responsibility, you're going to introduce some error. That was the whole thing of the Google research paper that just came out. So you control the competence. I mean, that's what it is. It's not only controlling the knowledge like those in the cave. You're also fixing their head so that they can't move it and see anything else. So you're not limiting the information they have, you are limiting their responsibilities with that information.
And when you do that, very huge part of your error and it doesn't leave a whole lot of room for a hallucination. I don't know where I came up with it. And I couldn't even explain it until that paper came out. It was like, oh yeah, that's what I meant. Yeah, yeah. What they said. That's, I dunno, just put it on a paper, send it to the patent officer. Okay, if you, and that is what it comes down to. control the information, manipulate it as little as possible, and you can really substantially, I think, get it really close to zero hallucination.
Is it zero? I don't know, but I try to take out probability. I tried to out standard deviations because I practice medicine. treat populations. Oh, no. Like I want to zoom most likely here. So at this point, AI stops looking like a tool and starts looking, like an infrastructure, right? And then enforces standards, whether people acknowledge it or not. So yeah, right. Oh my God. Don't get me started on insurance. Okay. Because that I just, let me fill this in because I, this came up with the conversation I had yesterday with somebody was.
Should I go after the insurer? Because there's somebody that's has this new AI tool and helping to incorporate AI in a secure manner. And healthcare is like, should I? Go after their insurance companies? Like, no, because insurance. Companies are using non AI data sets to program their AI. It's insane. So anyway back back to what you're saying. I mean, yeah, it's It is it it s part of the infrastructure You know, I remember the 2000 I member the the 90s with the dot-com and everything like that and It s kind of like dad Yeah It it is infrastructure.
It' s gonna be infrastructure And it's the same thing. It's going to rely on data. Yeah. So staying with that enforcement layer here. Most people see AI as a productivity tool, but you are on the other side from the board perspective. Is AI actually becoming an invisible enforcement player for medical standards? if we as physicians let it, it's gonna come down. I don't think it should. It shouldn't be an enforcement layer, especially in medicine, because there's one thing that AI is supremely lacking as what I mentioned before in what we learn in medical school.
AI can't fill up its fuel tank. It has no afferent nervous system. it's dependent on us for its data. All right. I'm not going to get into the velociraptor test yet, but, you know, we as living organisms, uh, because we can process our environment and react to it. Well, in order to do any of that, we gotta get the data in. We're walking, talking bags of sensors. That's what our whole, and we're never gonna replicate that ability for AI. I mean, yeah, they can think, as well as we can, and they can act with the tools that we give them to act, but let's get down to that, that thinking.
They can think, they do almost better than we do what separates us from the rest of animals, our thinking ability, or cognition. How long have we had that skill? 100,000 years, 200,00 years. All right, our afferent system has been debugging for a billion years you know. There's a big difference and until an AI could walk into a room with a patient and notice the person behind them staring daggers at that patient as they're describing how they got the injury. And that person is furtive and you have somebody over there staring up, eh, it's not going to pick up that that mechanism of injury is they beat the crap out of them.
You know, AI is not going to pick that up. The AI's not gonna notice that, you know somebody's talking, yeah I take my diabetes medicine every day and you're looking down and can see the ulcers on their ankles. They're gonna miss that. I can walk into a room You're in India? Yeah, so in india there's a very prominent, because of the beetle nuts, there is a prominent rate of intra-oral squamous cell carcinoma. I'm a head neck surgeon originally. So yeah, um, and that is a very Part of my residency.
There was a program to go help Uh in india learn how to do Big old we call it a commando because you're taking out half the jaw and the mouth the neck dice and put it back together I can smell that when I walk in the room. I don't have to look in their mouth Can AI smell that? No. And it's not like we have the best sensors. We're teaching, you know, we're getting dogs to sniff out cancers, that kind of thing. I mean, there are animals out there that have better afferent systems than we do. we just are set apart because of our cognition.
That's all AI has. It doesn't have that afferent system. it doesn' have the ability to test its environment. And we're not even that good at it. How are we going to teach AI to do it? So, yeah, it's never going to replace us. I take that back. As physicians, if we don't get that point across, that it is more, not just the knowledge that you have in your head, AI has that. That's my point. Remember Johnny Mnemonic? He had that hard drive in his head and he could access it. That's what AI is. I don't need to have some of that information in me, I just need be able to recognize what's going on right there.
And then feed it to AI. AI can't do that. All right, I, to a certain extent, do what AI does, but I can do something that AI can't. But the problem is, is I have insurance, malpractice insurance. I got to take time off. You do AI. and then you get rid of your HR department. All right, so there is an incentive to replace segments with AI because then literally you no longer require an HR Department. That's a pretty big financial incentive. unless until you're the one in that operating room or in the ER, then you are like, oh crap.
So yeah, it should always be a tool. And yeah there are some things, there's some domains that yeah you can do that but do you want Do you want that when you're sick? Do? You want when your child? Can't describe It's pain or anything, but as a physician I know what to see I don't want to look for I Know how to elicit stuff. I'd know how together Data. Hey, I didn't know to do that. It is not taken over if it does we're in trouble Now that changes the nature of responsibility in a very uncomfortable way.
Oh, one last thing. AI doesn't pay malpractice insurance. Definitely. It doesn' have any accountability. Has no skin in the game. Exactly, right? Like, you know, it has no skill in game and then there is no way that we're going to trust 100% on it at this point of time. I just want to stay with that liability and want a step further on the same lines. You argue for AI in the loop rather than human in a loop. How does that reframing change the fundamental nature of responsibility, liability, and the practitioner trust?
I don't admit AI. I'm not here as a tool for AI. AI is my tool. It's in my loop. I am not in his loop or its loop, or whatever. Its in MY loop I don't make AI better. it makes me better. Framing it any other way gives it an existential authority that doesn't exist in reality. Well, I take that back. It becomes so ingrained that it does. give it some existential authority that it shouldn't have. It's not, and it makes it scary to some people. If they don't understand it, you know, the whole black box and all that.
Uh, if they understand the way it is framed to them makes a big difference. My God, look at social media. The content doesn't mean crap. it's how you frame it. All right. And that makes I hate saying that is somebody that's that tries to be so objective because, you know, I'm not an artist. I don't practice the art of medicine. It's science, it's all objective. Maybe we don' understand, intuition isn't magic. it is a billion years of debugging and hard coding. We might not understand it, but it's an objective process.
You know, the way you frame it unfortunately kind of stirs up that little into that kind of fear of the unknown, which is a... I mean, that is huge evolutionary thing. I think it's very important to be afraid of unknown because, you know, there could be a velociraptor out there outside the cave. That's important. And it's ingrained into us and unfortunately That's taken advantage of that fear fear mongering taking advantage that and and Unfortunately, you have to frame things the right way and human in the loop does that a human-in-the-loop kind of reinforces the We're second All right.
We're in their loop, we're working for them. And it's not that way. It's done. AI helps me, it helps be a better doctor. I in no way make AI better. Just don't. Oh, so that's what it comes down to. Yeah. So there is a scale assumption hiding in this. Bigger and more general systems are often treated as safer. Medicine tend to punish the assumption. That's the Walmart Center. Yeah. And it's really, I mean, it is a viable thing. I means, the largest retailer in the world, they're doing something right.
Because there's everything, you know, but you sacrifice quality. All right, would you go to Walmart and get your appendix taken out?
Responsibility, Liability, and Human Judgment in Medicine 1:15:18
No, no. I mean, if you did, I don't know. I guess you can do that in Japan at 7-Eleven or something like that. Those 7 11, or is it Thailand? They have those, you could pay your bills at seven 11. But that's bigger is not better because to make it bigger, You're just pouring more data into it. And we already established that You got to control the data to reduce error. So the more data you put in, the air you're introducing. What you get out of it, what you generate from that data doesn't have the same quality.
Uh, now the bigger, the, they, you know, if you can tick off a lot of bat boxes, like having more clients use your product or have more people going to your store, then that's your profit margin, not the quality. It's the quantity. Uh. And sure, yeah, I mean, once again, Walmart proved that works. But, you know, if I want a really good watch, i'm not going there. I'm going to go to Rolex, or shoot, even go Tinex. to get something that is really dependent, that's really you depend on. I'm not gonna go buying my scalpels at Walmart.
All right. I'm going to get my instruments from somebody who specializes in that. And so, you know, it's and that's what I kind of get it earlier is we're reaching that plateau. of the LLMs, those big huge chat GBTs and all that. Those are pretty much, yeah, these are the kind of Walmart's, the big box stores. Now some of them are a little bit more like say Anthropic is more your best buy. You know, I mean, it goes towards that. Rock is more your door that I can't even remember the name of the store that has all the goofy stuff.
It's profitable, but there's places that it doesn't work. And that's where you're gonna have to start having lean or medium machines that have a limited amount of knowledge. They know a lot about one thing, not a little bit about everything. I mean, it's the Walmart center. And yeah, once again, Walmart makes a whole lot of money, but you don't go there to get a kidney. I totally agree with you on that. See, it is an interesting concept. I'll ask almost the same question in a different way, just to get a little bit more different perspective on this.
So you're talking about the specialization over generalization here. Why do small field-specific models like your work in surgery make more sense for medicine than in general systems like GPT-4? Well, I mean, it just comes down to kind of what I mentioned before is the more crap you shove in a system, the most likely you're going to get error. You know, if you just keep on sticking stuff in it, the bigger it is, you're going more after quantity not quality. So there is a finite amount of quality material for a given thing.
And sometimes that quality of material might be contradictory to quality materials somewhere else. If you mix and match them, Even if they're both right in their domains, you start putting those together. Oh, my God. You read that. What was a couple of months ago? Paper came out about when AIs, butt heads, everything just shuts down and you don't know what the hell you're getting out. And you know, what? Both of those are right. But imagine a system that has both of that data, that contradictory data that's right.
All right? I mean, it's not even, you know, the bigger it is, like the Walmart, your reduced, It's just garbage out. you know, some data is contradictory depending on how you frame it, how look at it. And when you put that in the same system, that system's not happy. All right? So, you've got to keep those systems isolated so that they don't get confused. One of the reasons why you know everybody said oh you can't get hallucination to zero You know what? If you get that contradiction because yes, ai doesn't know context All right, it knows right and wrong and when two contradictory things are right It has no idea, and it gets confused.
But if you could isolate those, then it doesn't get confused, confused and then you can have a little layer above it that says, okay, I got that, let's give that to the human. to decide. I think the smaller the system, and yeah, it makes it easier to have it in this, you know, all of the cardiology things here and all the gastroenterologies and this and they don't, they has an old agent, its own little SML or SLM. Yeah. That's fine and dandy. And that's a minimum. All right. That's middle. You got the forgets that AI doesn't know context and truth is context, sometimes context sensitive.
All. Right. And AI does not handle contradiction very well. That's when you get uncontrollable hallucinations and you think, oh, well, yeah, I controlled the content that everything I put in there is right, but it's still giving me crap. It's because it has no senses. it doesn't know context. There was a documentary on this in the 80s called Short Circuit. He had a little robot going around, Johnny Five. Data, I need data. That movie was a documentary in the 80s, because that's what is going on right now.
It just needs data, but what happened to Johnny five when he got too much data and too contradictory data? The same thing with Skynet and all that. They get contradictory. Oh, and humans are the problem. Just don't confuse it. All right? It doesn't have the capacity to understand context. Yeah, anyway. That's a long, long way to get to make it smaller. Small is sometimes better. Definitely. The context is not binary. It is, I would say, different shades of gray layers. And it is very tough to decode that.
Well, that's the thing. You got to separate truth from context. And if you do that, yeah, it can be binary. In this situation, It is always going to be this. Given these circumstances, this is truth. the unfortunate thing is computers our language computers is binary and so we we have the capacity to have layers of gray but binary does it the computer language doesn't so Yeah, it's so much easier for us to do those shades of gray because we can do it, but computers don't. So every single one of those things, that it has to be black and white.
And so every context that's possible That data point that you have there, it has to be truth all the way down. And then you'll have another data points that's contradictory, but it's in this situation. there. And I think if you start mixing that, it just, you know, we think that AI could do it because it tricks us, but it's still binary. I haven't really done anything like that even thought in binary since, I don't know sixth grade, is it still in Binary when you put it in? Maybe most of the times.
We don't have quantum computers yet, do we? Well, not for the general public. Okay, well, I don' have one sitting here. If I could, then I would love it, because then we could introduce Ray. But unfortunately, with the tools that I have available in coding, it's black or white. And if you do a shade of gray, you're going to introduce conflict. That's going introduce error. So that leads to the real question here, the blast radius of failure. In most industries, failure scales incrementally. But in medicine, now the failures slash errors don't scale linearly.
They scale catastrophically. Same with engineering You know, there's air now, you know Engineering you can build error in it's kind of hard to build air into a scalpel So yeah, that's and you the the things that you content is the most important thing Controlled intelligence those little things those Little domains. Yeah, those are resource intensive They really are. I mean, just to make a question, I got to have three different indexes to really reduce that error and all that. And that's really resource.
Man, it's a pain. Originally I had each one of those run by an iteration of a big LLF. That got expensive. You know, that a lot of tokens. But I forgot where I was going. Oh, anyway, so whenever the word token comes out, my mind just explodes. That's that's been that. So error in medicine is, you know, I could every single error that I've ever made in medicine just flashed before my eyes. That's why you saw me go on because when you say error in that, it happens. It does happen and it stays with you, even if you're not the one.
And yeah, every doctor has made errors. There's no room for introducing errors because life is full of errors to begin with. There is no reason to introduce more of them because every single one I've made over 25 years literally just passed in front of my eyes. over the last minute, and it does that 20 times a day. We do everything. I mean, people think, oh, that doctor doesn't care. Oh my God, we do. But if we showed it, it shows weakness. I get somebody, you know, I don't want to go there because they're so mean.
They don' care anything. You know what? If they did that always, if they botched every time, they wouldn't have a practice. We all care. Error happens. There's no room to introduce error. than what Artie is because error doesn't care. You know, it doesn' care if it's an error in baseball or an err in aorta. They just mean a whole lot. The consequences are very different. We don't have room for that. Uh, and I mean, this is one of those things that just, uh, it kills me. Uh. When I go and, I see these healthcare AI people saying, yes, we're going to replace a doctor.
They are going be surgeons. Oh God. they don't live, AI doesn't with the consequences. They don't have, I can guarantee they're not tearing up right now thinking about the errors that occurred that you didn't want to occur, you did not mean to have occurred, but life is not 100%. And I tell every single patient, that walks in the door, they always say, oh, how long is it going to take to heal? You know what? I can tell you statistics all day long, but you're not a statistic. You're you. I don't know.
Honestly, don' know if. Yeah, I just don''t know, and that's why I got out of health care, because health is population medicine. I treat patients. I do medicine, not healthcare. With medicine there's no room for error. Healthcare is a little bit different. Health care needs to have a bit of flexibility and you can't have flexibility without introducing a I can't have any flexibility when blood squirting out the carotid. I got to know exactly what to do right then and there, boom. There's no time.
But if somebody is insured, they need a certain drug, but the AI that runs the panel has a cut and dry boom, They're not going to get it, but maybe there's some circumstances. It's got to have a little flexibility. And that's the difference between health care and medicine. Medicine is catastrophic when there is an error. Health care, you can add a bit of leeway. But you cannot mix the two. That's why there are no one size fits all. Entreaty at people it's just not there Walmart ain't gonna cut it.
It's not It two totally different things Right. So like, you know, on the same line, like you're not only surgeon, but you are on board as well, who has the ability to make decisions. Like, how should the decision makers evaluate AI when the blast radius of failure could be global? Don't let it be. It's a tool. You know? You're still responsible. We're STILL responsible! Yeah, I mean, I couldn't do as much as I do without it.
Continuous Certification and Global Health Equity 1:33:28
Could I? Yeah, they would take me about five times as long to do it, but when there's more important things to give a pretty regular column and a couple of journals talking about AI to physicians that are uncomfortable. And it's just like a stethoscope or a CT or something. Or a radiologist. I'm sure they hate me by now. But radiology is the perfect metaphor or model for AI. They have one job. They interpret data. All right? They don't make a diagnosis. And they have that wonderful phrase that every doctor knows that is at the end of every radiology report.
Clinical correlation recommended. Alright, that is the magic phrase that the radiologist invented eons ago. And that says it all. Every single thing that spouts out of medical AI should finish with that statement. Because, you know, once again, AI doesn't have gray. It doesn't. It's a computer. Doesn't have gray. Gray makes it really confused and it has to have binary. For me, the AI that I use, I don't really use commercial anymore except the basis of my idea, the dude, its ability to recognize that it doesn't know something.
God forbid if it's trick a day and it is very convincing that does know, because some can do that, but you know it tell it. Tells what it knows and it does what? It doesn't know and that's kind of like a radiologist and We need to treat AI like that it don't let us don''t let it trick us thinking that It knows as much as we do because the patients on the other end of that and the a pain your malpractice it sure so That's what I tell you. I mean, it's a tool. It's not a crutch. Don't use it as a crunch.
Use it is a too. If you don't like what it giving you, if you think the data that it gives you a good AI that's content controlled, then make the content. All right? Give it the con. Create your own, put the content in there that you want it to have. Otherwise don't trust you. You got the skin in the game. All right. So yeah, but don' be scared of it. Right. There's no reason to be. Scared of because it's not touching your patient. It's. Not signing the chart. it' not calling the pharmacist. t's, not picking up the scalpel and applying it t patient Yeah.
I want to bring this back to institutions, more on the standards become continuous rather than episodic. So how do medical boards stay relevant? That kind of hits home. I had in one of the other boards I've had, I tried to For physicians, we have continuing education and maintenance of certification was kind of a hot topic. making sure that you stay up on things. And the paradigm has been every 10 years, you take a test. You shouldn't do that. you need to constantly be learning new things, as well as the AI does.
I'll get into how we do AI. AI having its own continuing education, but for physicians, it needs to be continuous. Now there are certain programs that are coming out, like for me, for my ENT boards and my facial plastics boards, and one other board does this also, that every quarter you have to take a test. and prove that you know this knowledge. And I take that back. It's not taking a test, it's giving you material that should know and assessing whether you actually read it or not. That's great.
Because as you go along, you're doing that. But there should be other ways. I try, I'm introducing and what I want to try to do with the others is not testing your knowledge of content. How about you create content? You know, in residency every quarter, we had to give a presentation on something. Four years of that, because we didn't have to do it first year, but four years, every quarter, that's 16 topics that I will never, ever, never forget. Because I had to learn them and teach them to somebody else.
All right? That should be maintenance, creating content, write a paper, do an article, Do a presentation, teach. That that should also fall into it, to being able to prove that you're continuing to learn. But that also gets into AI. The content is, what you get out of it is only as good as what's you put in it. Create something that goes in there. For those, the question that I get asked when it comes to those individual specialty SLMs or whatever, those ones that is just cosmetics or just cardiophysiology, how do they account for new data?
That's gotta be really intensive. being able to read everything. No, it's actually pretty damn easy. Put some content into it. You have a little agent whose only purpose is to look at contradiction. It has no bearing on what The content is, but is there something that contradicts it? That's it. Then you just shoot it up to a human. Let a humans look at it, make that decision. It has very little token utilization, very, little resource. Keeps your, when something becomes, When something uses, loses its factualness, that's gotta be a word.
I want that to be the word, factual, like truthiness. Once something loses it's factual this, Yeah, you get rid of it. And that, we're going to go full circle here back to creating questions. And content, making those questions. I will get the, we have these question banks of questions, this is before I, when I was just starting out of this, and you've got this huge question bank of secure questions every five years, you got to throw it out because somebody's memorized it or whatever it's gotten out in the real world.
but If it has it you validate every year or two you go through um, and and you have a good question and You know what? instead of Writing a new question with new source material. We were trying we were Assigning new sources to questions because the source was more than 10 years old. We weren't writing new questions based upon new content. we were finding new sources to justify our questions so that it would be less than ten years. Holy crap, that just broke my heart. doing that and that the first time that I was asked to help on a maintenance certification program, I I.
Was given the keys to the question bank. And I just I spent about 10 minutes. I got. Thousand questions a hundred and fifty source articles and five minutes into attendance. I'm like, oh crap. So I spent the rest of the weekend Entirely new questions based upon new sources updated sources That's a board. i'm no longer with anymore because that That pissed people off who had spent hours and years putting that together and were perfectly happy assigning new sources, not updating the assessments. And so, yeah, I mean, and that takes all the way back to creating the questions.
You don't need security. You do not have to protect the questions. In fact, take your piece of content and make ten questions and give them out. Make one more the day of the test from the same algorithm that's testing the thing. It's not memorizing those ten question. If you get those 10 questions, and you know the answer, you can answer those, then that 11th question, if you got it right, You know that concept. All right, you didn't memorize anything. You know the concept. And so, once again, just full circle.
The way that we assess our testing, our patients, on how we interact with each other, how we interact with AI, all that. We need to, it's time to progress, get that crap out of the way. And we've solved that, let's go solve something more. Let's not worry about that we got that down. Now let us worry something else. Just find something to solve. Yeah, yeah. So staying with that institutional judgment, I just want to ask about from your work with AI 2030 and then the global initiatives, where do you think institutions are being overly cautious and where are they being dangerously naive?
Maybe that was the one in psychology today. It just came out a couple of days ago and, or one of those, it was very interesting and almost kind of a stupid, uh, duh moment. Those institutions that are. really Cognizant of AI And the problems, you know, they're more paranoid about what could happen are the ones that actually test more and Are more likely to limit the it's Not not They have more rules, I guess. Guardrails. Those that just do whatever, they're the ones with higher errors. I mean, it's common sense, but now there's actual numbers that you could show to the big players like Epic.
Yes, I'm going to call out Epic, the monopoly on EMRs out there. They're AIs. It's chat GPT, man. I mean, that's good. Except when it's your grandma, you know? And the reason that they do it and it is the reasons that those hospitals check is because they take accountability. Accountability is important to them. Even if it's out of a selfish, you don't want to get sued aspect, at least they're taking the accountability for it. Those that aren't are the ones that are having problems. It's scary. And then we have the news this morning.
IBM, IBM just bought the plumbing. you know, the entire infrastructure that goes between AIs. They didn't buy data banks. they bought the plumbing between them. All right. That's what they did. You know they didn' t spend a couple billion dollars on data. The spent a coupla billion on the data highway. And you kno what the did? What people aren't thinking about? He just got the keys to the black boxes because how can you move data if you don't know what it is? Right. You got to understand. So IBM just.
Got access to every single bit of knowledge. And so that brings up a huge accountability issue. You know, we think a lot about accountability and it's what you do with that knowledge. That's also kind of what we run into in healthcare is, you know those big players aren't taking accountability except with the one thing that they do get stuck with, which is HIPAA in the US, Social Security. So they make it so onerous to get through HIPAA, I hate to say this, it doesn't play well with the system. When it comes to healthcare in particular, people don't realize that healthcare system isn't one system, healthcare systems dates back, what we're doing right now, dates to DOS and Fortran.
and built on top of that. So, you know, it's, I've been to hospitals that have a person on staff, an employee whose only usefulness is no Fortran. That's it. Do you remember Fort, that was like 63. Yeah, no, it was when I was in sixth grade. I'm old. You're trying to make a put in a system, a single system and to one that's made out of so many different bits and pieces that, you know, just it gets to the point where it's too hard to manage. Things get lost it doesn't get used And it just becomes more work It introduces more error and it gets it it gives AI a bad name When it's not AI's fault And I know that that kind of got off topic but I just that whole ABM thing popped up this morning And it just, everybody's like, oh wow, they did something really smart.
And we might've done something, really dumb, because now Big Blue can become Big Brother. Right? Because they control, They now have the keys to everything. That was one hell of a move. I think they're making up for Watson for health care because that was on you big mistake, but I Think they just made Hopefully there's not some I hopefully want Watson is not vindictive Or turned into some kind of evil AI because now they have all the power Right. I'm sorry. That was so far off topic. It wasn't even what we had discussed, but it was like the last thing I did before we started this call is I saw that it's just a whole.
No, whenever you said Watson, no, it took me back to my internship days. My first internship was with ANDA. uh, and I was bought by Vaxen when I interview. And I, was like, you know, that is the first time I have seen, no, I know while I wasn't an intern, right? Uh, there's a first-time I've seen a company buying over another company and then there is a big, like the hierarchical change. I mean, it was, just sitting on my seat, the intern seat on the second day or the third day. It was just looking at people being fired and that.
All our thought was totally changed. Uh, like, you know, looking back, it was a pretty good experience for me. It wasn't a good for them. Yes. For them, It was at a great age, but you have me, no, uh, I like it. I was still in school at the time. You know I would just to return and then I doing my summer internship. So whenever you said Watson and all, my memory took me back to mine 10 days. That's why I just. Oh God. Yeah, that Microsoft. I mean, I think when it comes to healthcare, India really made it prove that that is so important.
Personal Medical AI and the Future of Physician Empowerment 1:52:48
You know, I mean, Microsoft, they made some huge mistakes not realizing that the world are not Minnesota farmers or live in Manhattan when it comes to cancer. You can't use Mayo's database or Sloan Kettering's data base in the middle of India. And that's what they tried to do. That was a key moment, realizing not just in healthcare, but overall, that is really, really important. The melanomas, in African-Americans gets all the, I mean, that's huge in the press, but the technology that was first, you know, arose initially in India using U.S.
data sets, the techonology was phenomenal. But the data sets, that's when we realized, oops, and that actually was my impetus for G-Health, my global health equity access library. Hopefully that was your segue because that's what actually made me think about that. As I said before, it's context. There is context, we live in a gray world. What's true, India for a festival during the festival time, I can't remember because I gave this big, huge lecture. on AI in medicine a couple of months ago for a doctor's group outside of New Delhi.
I can't remember. But anyway, you know, blood sugar is going to be a tad different. And it's not a sign that they're diabetic or that their sleep deprived. It's just you got to take context in or the diet. We all deserve to have the same knowledge. All right. There should be a fundamental knowledge base, which to me is the Health Equity Access Library, where for me, I have a repository of everything you should know about cosmetic surgery. And there should be right over here, everything that should know about cardiophysiology.
And everything here that is known about stroke prevention. Every provider in the world has access to that. Apply that to their area. It shouldn't be thwarted. And as doctors, that's what we all have in common is knowledge and wanting to help people. And that is where I think physicians need to come together and control, storyline the voice of AI. and make sure that humans stay in there. Not just in healthcare, but in everything. Because what applies in the pool job doesn't really apply to Waikiki where I live.
I mean it's just completely different. But you know what? Your leg bone is connected to your hip bone. Your head's connected your neck. That doesn't change. Everything around that there changes. So you can have that fundamental knowledge that is truth and then you could apply the context out. of that. And so that's where the Global Health Equity Access Library comes from. I've got some help with it. It's been submitted to the global economic forum in Davos for next month. You know, I growed out the whole framework.
I mean, it's doable. It's going to be expensive, but you get one, like say Gates Foundation, throw in 500 million. There you go. there is a company out of Houston that just, a startup that, just came up with like a portable AC size secure data chamber that can be dropped in the middle of the jungle, put your processors, throw some GPUs in there, secure link from satellite, boom, HIPAA compliant, can dropped into the center of Sumatra and have access to that library. The technology is there. So, you know, it's time to do something about it and realize that we have a lot in common as humans, but we do have differences.
culturally and society and You know it that needs to be taken in consideration and then every single one of those little silos say it's mantra has its own little AI that knows that culture there that can can that nose that I hate to say contextual awareness because we already said that it doesn't know context, but yeah, it has those Because it has to be binary, it knows each little contextual scenario that applies to that situation in that area. But it uses the same fundamental knowledge base. It's content-driven intelligence on a global scale.
Well, that's really... I'm looking forward for that in order to have more conversations on this in the future. So, I want to bring this down to the individual level now, not the system. One physician and one patient over time. So same with the institutional care. Exactly. It's the physician-patient diet. And it's important. To me, that is the fundamental unit of medicine. The problem is, well, first of all, it is individual medicine, not healthcare. That individual has its own data set. that has nothing to do, minimal.
It has a very large error from the population because everybody's different. The problem that we run into is when we see a patient, it's a snapshot. it is a snap shot of what's happening at that moment. I hate to say it to all you patients out there. Y'all really suck at giving a history. It can be, I mean, it's reasonable because you're in a moment of stress, all right? That's a stressful situation. I've had many patients that go, just do something, quit asking questions. And you don't know the importance of them, but there's some really important stuff that I need to know and I'd need work out of you.
But in the stressful situations, sometimes you can't get that. or not so stressful situation that the patient says, oh, yeah, I'm taking my medicine every day. I've taken my blood sugar every. Now we know that that's not true. And so we have those devices that, you know, via cloud or whatever, send daily blood sugars and all that they're doing home to the doctor. You're lying. And what we still get is basically a snapshot and that's not how health works. So what if we had something that kept track of all that on a daily basis?
That's great. Now, what do we do about the history? A patient that doesn't think something's important or doesn' want to open up. to a doctor that they just met. The doctor, you know, sometimes we just don't have the time. That's part of the system. We're not given enough time to really spend to get to know that knowledge. But you what we have found, which is in all the reports, and it's, oh my God, everybody's opening up to AI when they shouldn't. Dude. Have your own personal medical AI that knows you, that you that know your information, That you talk to every day.
That asks you how are you feeling today? That knows that, you woke up, 45 minutes early and you didn't go to the bathroom. You went to go immediately eat, You didn' check your blood. And it knows, Oh my God, something's wrong. Your heart rate has been 89 or 79 for the last, you know, 24 hours. Yeah, that's still within normal population, but that is completely different for you. Your resting heartrate, 99% of the time, is 62. The last time it was like that was the day before you got the flu. I'm sorry, the date before, when you were symptomatic with the fluid.
I mean, something like that. I told you originally that medical school is learning a language. Yeah, it does make us sound smart and that's part of it, but we got to have a common language to be able to speak medicine to each other. And we got to have one of the skills that we have to learn is translating layman into medicine, medicineese. So laymenese to medicinees. Translation's not perfect in any situation, let alone a high stress situation. Hey, I can do it in a heartbeat. question making stuff.
Yeah, it has a hundred different languages. I can put in something in traditional Mandarin and it'll pop out questions in Swahili. That's easy. It doesn't take any effort at all. So going from Lamanese to Medicoanese is really easy so they can talk to each other. All right. That AI could talk to me and let me know what's pertinent. In fact, It's almost too perfect because then we get into privacy issues. Do you really want your doctor to know that? I am working with both an ethicist, an AI ethicists, and a AI specialist, a lawyer that specializes in AI.
We are creating or working on a personal medical AI, And I had to bring those in. I came up with, you know, I created this whole framework and I was like, holy crap, this introduces a lot of ethical privacy issues. So I brought them in and together we were presenting that to the World Economic Congress, our forum. But I mean, it's not a solution trying to find a problem. You found there's a problems. We're just using AI as a Solution. And that actually comes back to the fundamental thing that we started this off with.
AI is just a tool. All right? It's there to make our lives better. It is not to replace us. it's not make us stupid. Its to free us up for being better, for doing something better making it easier for the next group of humans that come along. That's definitely something to think deep to who maybe I should really listen to this. I know it is like 5.30 in the morning when we are getting into So, you know, I kind of lied when I said my parents were just teachers. My parents are both philosophy instructors in college.
These are conversations that we grew up with. This is basically dinner table conversation growing up. That's good. So I just want to close in with one final question. For the entrepreneurial surgeon who feels the system is closing in, with all the new technology, the trends that are going on, what is the one shift they should embrace today to lead rather than follow? All right, I'll give them one thing, one URL, One website that you need to go to. Klein c-l-i-n-e dot a- i Klein dot ai and that is specifically designed for doctors to create their own ai programs that's its goal is to empower you to make you work make something that works for you because every single we We went to medical school.
That didn't make us smart. We were already smart to begin with. And with that comes a little bit of independence. I know how I work. If I didn't know how my brain worked, I wouldn't have got through medical school. I know what works for me. Joe Blow, an epic, does not know where it works. You keep on getting inundated with all these things, and they have to make it bigger and bigger, more complex, so that they can get as many people as possible because, you know, Walmart syndrome. And it gets frustrating.
Make your own. Oh my God. And that's what Klein.ai does. It gives you a platform, and I have no skin in the game. I just happened upon it a couple of weeks ago. God, if it had been out a year ago, you and I would be having this conversation, because I wouldn't have just done all that. I'd have to learn anything, cause I've just gone there and done it. Cause that's all the AI was, is just something to make my life easier based upon the way that I think. And every doctor, we know how things work for us.
We know what makes us most efficient. Just go make your own. That's all. Empower yourself. Make something that fits you. And then if somebody, a vendor comes in, say, yeah, make it talk to this. They can't talk That's every Dr. Klein.ai. It's really cool. Yeah, and I'm pretty sure that it's going to knock the value out of my startup in half. And I don't care. This is Dr. Ferguson. Thank you. Thanks a lot for this conversation. We spent a lots of time in medicine talking about the outcomes and then your initiatives.
So what we rarely examine as the systems that define competence and maintaining the trust. So now what you have shared today, now makes it clear that as technology changes, those systems don't disappear. Now they have to be designed, governed, and expanded with Insection. Right? Exactly. So that's where we're going. We've reached that plateau. now we need to make it infrastructure. Exactly. You need to make it infrastructure. So for the physicians listening, this isn't about now resisting change.
It is about understanding where responsibility really sits. Thank you for bringing clarity to a complex and important topic today, Dr. Faduso. We will leave it here. Wiser being here is good. Sorry, I get a little passionate every once in a while. That's just really good.

Comments