24 July 2025 · 19 min

Governing GenAI in Healthcare: Regulating LLMs in Clinical Settings

GenAI is reshaping medical workflows—but our regulatory tools aren't ready.

In this episode, we explore:

We unpack frameworks from recent white papers and discuss what compliance will look like in the real world.

Your company here. This podcast is looking for its first sponsors: reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsorship options and rates →

Transcript

Automated transcript of the audio; it may contain errors.

Host 1: The world of healthcare is really being transformed, um, almost at breakneck speed by artificial intelligence. We're talking specifically generative AI, large language models. We're already seeing things like, you know, automated clinical workflows, even personalized diagnostics. It honestly feels like we're getting a glimpse into the future. But here's where it gets, uh, really tricky. While this tech is just rocketing forward, the rules, the regulations that are meant to keep us safe and make sure things actually work, they're struggling to keep up.

Host 2: Yeah, what's truly fascinating here is that we're kind of standing at this unique point in time. The, uh, the massive opportunities these GenAI and LLM tools offer in healthcare, they're also throwing up some significant challenges for existing regulatory frameworks. It's a real dance, you know, between innovation and oversight.

Host 1: Exactly. So, today our deep dive is all about navigating this incredibly complex world of AI governance in health and medicine. We're going to zero in on the critical gaps in how we currently regulate AI, look at some, well, some really innovative solutions being proposed, and try to understand why global collaboration is just absolutely essential for the safe and equitable advancement of these powerful tools. Our mission today is basically to give you a clear roadmap through this rapidly evolving landscape. We're drawing insights from leading research, including a recent paper focusing on AI governance in health and medicine. Let's jump in.

Host 2: Okay, let's try to unpack this. On one hand, you've got GenAI and LLMs offering these incredible possibilities, revolutionizing clinical documentation, helping with diagnostics, but the research we looked at makes it pretty clear these tools just don't fit neatly into the existing medical device rules like the total product life cycle or TPLC. So, what exactly is this TPLC approach, and why is it sort of clashing with GenAI and LLMs? Right, so the TPLC approach is, uh, essentially how regulators evaluate and monitor a medical device. It covers everything from the earliest development stages right through market approval, and then well into its performance out in the real world, post-market. It's designed to allow for quick improvements, but with safeguards built in. Think of it as the foundation for how agencies like the US FDA or the UK's MHRA oversee AI and machine learning medical devices.

Host 1: Mm.

Host 2: But LLMs, they're a completely different animal. Their outputs aren't deterministic, meaning, you know, they can give different answers even to the same prompt. Plus, they have these incredibly broad functions, and they integrate into healthcare systems in really complex ways. All this fundamentally challenges that established, more traditional method.

Host 1: That makes sense.

Host 2: It's almost like trying to regulate, I don't know, a constantly shape-shifting cloud with rules that were designed for a solid physical object like a car.

Host 1: That's a great analogy for the conflict here. So, what are the specific things about LLMs that make them such a, well, a poor fit? The sources pointed to a few big ones, starting with the data itself. These things are trained on absolutely vast, incredibly diverse datasets, often just scraped directly from the internet, which makes it almost impossible to thoroughly check the training data for quality, or even where it came from.

Host 2: Precisely. And, beyond just the sheer volume and the obvious privacy concerns, there's this critical issue of data provenance. It's really about knowing where the data originated and if it's actually reliable. There was a big audit, looked at over 1,800 text datasets on Hugging Face, which is, you know, a super popular platform where developers share AI models, and it found really alarming rates of license omission, like 70% of datasets were missing license info, and actual errors in 50% of the licenses that were listed.

Host 1: Wow, that's high.

Host 2: Yeah. And this isn't just some academic footnote. It means an AI tool you might potentially use for your health...

Host 1: Mhm.

Host 2: ...could be built on questionable, maybe even biased data, with no clear way to trace it back. And then when you add open-source models where third parties can modify them, it just further obscures the data authenticity. It gets incredibly hard to track the origins of the information the model learned from.

Host 1: Okay, so the data is one major headache, and then there's the output itself. You mentioned non-deterministic, like in medicine, if a lab test gives you a blood pressure reading, you expect it to be the same reading if you run it again right away. But with these LLMs, it sounds like asking the same medical question might get you a slightly different answer each time, or maybe even a hallucination, you know, something completely made up. The research points out that we don't even have a good handle on how often these AI-generated false or misleading bits of information actually happen. That unpredictability must be a nightmare for regulators trying to ensure safety.

Host 2: Oh, it is, absolutely. And this brings up a really fundamental question: How do we even define these general-purpose LLMs and all their possible applications as medical devices in the first place?

Host 1: Mm.

Host 2: Whether existing rules apply like the EU's Medical Device Regulation, the MDR, or the US FDA's rules for software as a medical device, SaMD, it all hinges on the AI model's intended use. But when an LLM can do so many things—transcribe, summarize, and maybe suggest medical information—where exactly do you draw that line?

Host 1: That's a really fascinating gray area. The paper gave a great example, these LLM tools that can listen to a patient-physician interaction, transcribe it, summarize it into clinical notes, but they might also interpret clinical data and suggest medical information within that note. How do different regions classify that? It sounds like it could fall into multiple buckets.

Host 2: Yeah, for that specific clinical documentation scribe example, the US FDA generally classifies it as not a medical device. They have these administrative support exclusions. But, um, the UK, they consider it probable that it is a medical device.

Host 1: Mm.

Host 2: And the EU says it's possible, but it depends on the specific medical purpose and the risk involved.

Host 1: So, completely different takes.

Host 2: Right, this really obvious lack of clear international agreement on what counts as a medical device for LLMs, and how to classify the risk, it creates huge ambiguity. It's challenging for developers trying to innovate, and it's challenging for regulators trying to keep everyone safe. It's a very fragmented landscape right now.

Host 1: Okay, so given all that ambiguity, how do regulators even start to evaluate the performance of these LLMs? The research we looked at highlighted that, unlike traditional diagnostic tests where you have clear, quantifiable metrics for accuracy, LLM evaluation often ends up relying on more subjective qualitative assessments. This seems like a massive hurdle if you're trying to guarantee patient safety.

Host 2: It is indeed. And while LLMs can show really impressive performance on certain isolated tasks like, surprisingly, scoring really well on medical licensing exams, which is impressive in itself, studies like one by Hager and colleagues reveal significant limitations when you put them into real-world clinical decision-making scenarios. They did a comprehensive evaluation of leading open-access LLMs and found that all the models performed worse than actual human clinicians in making diagnoses, interpreting lab results, and sticking to established treatment guidelines. Wow. That's a really powerful aha moment right there. It basically means a high score on a standardized test doesn't automatically translate to safe, effective clinical use. And connecting this to the bigger picture, the source also really emphasized the critical need for robust evaluation of LLM bias. What kinds of biases are we actually talking about here, and why is it so crucial for patient safety? Bias evaluation is absolutely essential if we want to avoid these powerful tools accidentally making existing health inequities worse. The research highlights concerns like, um, inaccuracy across different identities like race or gender, or generating stereotypical language, or even potentially withholding opportunities or resources disproportionately from certain patient groups based on the AI's output.

Host 1: That's serious.

Host 2: Very serious. To prevent this, we need approaches like adversarial testing. That's basically where you actively try to provoke the model, find its weak spots, find potentially harmful or biased output before it gets widely used in clinics. You have to actively look for the problems.

Host 1: Okay, so we've talked quite a bit about the challenges before these tools get approved, but what happens after they're out in the world? Monitoring them, enforcing regulations, that sounds incredibly complex, too. The source mentioned issues like tracking off-label use—using, you know, using the AI for things it wasn't approved for—and the real difficulties in relying on people voluntarily reporting problems or adverse events.

Host 2: It's kind of similar to how we track side effects of new drugs after they're approved, what we call post-licensure pharmacovigilance. We have extensive systems for that. But with LLMs, it's arguably even more complex just because they're so pervasive and used in so many different ways. Imagine, uh, an AI scribe making a mistake, a hallucination, or maybe redacting important information inappropriately, and that error gets carried forward into patient records, creating downstream problems with no clear origin.

Host 1: Right, a hidden error.

Host 2: Exactly. Yeah. And then you add in these accelerated market approval processes like the FDA's 510(k) pathway, that's based on showing substantial equivalence to a device that's already on the market. This can lead to something called predicate creep.

Host 1: Predicate creep, that's a really striking term. It sounds like a slow, subtle drifting away from what was originally approved as safe. Can you maybe give us a quick example of how that works?

Host 2: Yeah, sure. Imagine, uh, regulating a car. The first model gets approved, then the next year's model has a tiny, unreviewed change, maybe slightly different brake pads. It gets cleared because it's substantially equivalent to the first one. Then the next year, maybe a new engine component, also cleared based on previous version. Then tweaked software. Each step is seen as only slightly different from the one before it, but over several years, you end up with a fundamentally different car, right? And the accumulated changes were never rigorously reevaluated together for safety. That's predicate creep.

Host 1: Okay, I see.

Host 2: And with LLMs, which can change and update even faster, that potential for drift without thorough re-review is a real concern, especially when it impacts health decisions.

Host 1: And now with open-source LLMs getting really good, almost catching up to the big proprietary ones, and their costs dropping like a rock, access barriers are just plummeting. What does that mean for enforcement when potentially anyone can deploy these powerful tools? It sounds a bit like the Wild West.

Host 2: It really does feel that way sometimes. It means you could see widespread adoption of LLMs for medical and even non-medical purposes, often without any formal approval or solid evidence behind them. This makes enforcement incredibly challenging for regulators. We're already seeing developers release unapproved, public-facing LLM applications that pretty clearly meet the criteria for being a medical device. This rapid, kind of uncontrolled deployment makes traditional regulatory oversight really, really difficult to apply effectively.

Host 1: And beyond just the regulatory framework, the source also dives into these broader ethical considerations that maybe go beyond current oversight, things like, you know, the low trust because of hallucinations we talked about, the real potential for making health inequities worse through embedded bias, and just fundamental patient privacy concerns. What's the broader impact here, especially thinking about the relationship between patients and clinicians?

Host 2: Well, the impact on things like patient autonomy, their dignity, and just the fundamental trust in that patient-clinician relationship, that's a major concern. Conversational AI can be really powerful for patient education, for instance, but it needs incredibly careful fine-tuning and continuous monitoring after launch to reduce the risks of cognitive biases, things like confirmation bias where maybe a patient or doctor only seeks out AI information that confirms what they already believe, or automation complacency, where clinicians might just over-rely on the AI without critical thinking.

Host 1: Right, just trusting the machine too much.

Host 2: Exactly. And interpreting ethical principles here is also highly subjective. You absolutely need a multidisciplinary approach. You need bioethicists, regulators, the actual users, patients and clinicians, and the manufacturers, all working together to navigate these really complex trade-offs.

Host 1: So, okay, after laying out all these quite significant challenges, what are the solutions? This is where it gets really interesting. The paper talks about regulatory science and some innovative approaches. What does that actually look like in practice for LLMs?

Host 2: Right. Regulatory science is basically about developing new tools, new standards, new approaches to assess product safety and performance more effectively, especially for novel technologies like AI. For LLMs, adaptive regulatory approaches seem key. This might mean being, uh, perhaps less restrictive initially, but having clear mechanisms to reinstate restrictions quickly if harm starts to emerge. It's a bit like the accelerated approval pathways we have for some new drugs. There's also this concept of predetermined change control plans. This would allow developers to make certain predefined modifications to their AI model without needing a whole brand new approval process each time, acknowledging how quickly these models evolve.

Host 1: That sounds practical. And what about regulatory sandboxes? That sounds like a really fascinating idea.

Host 2: Yeah, regulatory sandboxes are essentially controlled environments. They're outcomes-oriented tools where new services or digital health tools like LLM applications can be tested with real users, but under specific constraints, and often with fewer immediate regulatory hurdles. We're seeing examples like the UK's MHRA AI Airlock or Singapore's IMDA Privacy-Enhancing Technology sandbox. These are incredibly valuable because they let developers work closely with regulators right from the early stages. They can identify and mitigate risks together through a kind of trial-and-error evidence-based process.

Host 1: So, learning by doing, but in a safe space.

Host 2: Exactly. You could even imagine developing global regulatory sandboxes to study things like international interoperability and figure out how different countries' rules might work together. It's a space for controlled, collaborative innovation.

Host 1: That seems incredibly proactive. What other kinds of shifts might be needed? The source mentioned a potential convergence of regulatory frameworks, maybe moving towards a software as a medical service, or SaMS, model.

Host 2: Yes. That's a really interesting point. As LLM-powered AI agents get more sophisticated, especially those with reasoning capabilities, they start acting less like simple assistive tools and more like, well, service agents with a degree of autonomy, almost like a skilled human worker performing a service. This suggests that regulation might need to shift its focus. Instead of just regulating the tangible software product, it might need to focus more on the ongoing service delivery. For instance, in Singapore, they're actually bringing health services regulation under the same authority, the Health Sciences Authority, that previously only regulated health products. This leads to greater convergence, recognizing that the AI service is what needs oversight, not just the initial code.

Host 1: That makes a lot of sense, especially given how pervasive AI is becoming. And speaking of pervasive, the source also raised the critical importance of looking at the entire AI supply chain. That recent massive global IT outage caused by a single software update from a cybersecurity company, CrowdStrike, that really hammered home how vulnerable systems can be.

Host 2: Oh, absolutely. That was a stark reminder. AI is getting embedded everywhere, in operating systems, in other software tools, even within hardware components used in healthcare. So, ensuring continuity of care means we have to scrutinize potential vulnerabilities across that entire healthcare delivery process. From how data is collected and used to train models, to how the AI is deployed, updated, and monitored, we really need global collaboration here to get better visibility into all the components in the supply chain, develop frameworks for testing these critical pipelines, share best practices for business continuity planning, and, importantly, build local AI development capabilities in different regions to reduce over-reliance on just a few large commercial tech providers.

Host 1: Building resilience.

Host 2: Exactly. And this also means aligning government policies and health system investments to promote diversity in the technology we use.

Host 1: So, bringing this all together, what does this mean for a global call to action? The International Medical Device Regulators Forum, or IMDRF, was mentioned as an example of an existing global effort.

Host 2: Indeed. Organizations like the IMDRF and other global regulatory research groups are absolutely pivotal. They foster that crucial international collaboration needed for harmonization, getting countries onto the same page with regulations. A really pressing issue they're tackling now is the lack of robust standardization for these LLM-based tools. We need concepts like model cards, which are basically like nutrition labels for AI models, transparently reporting their performance characteristics, limitations, and biases, and also data cards to clearly describe the datasets used for training, their diversity, their provenance.

Host 1: So transparency tools.

Host 2: Precisely. Various checklists and assurance labs are popping up around the world, but these international groups need to help align and consolidate these efforts. We need a consistent, effective approach globally, not just a patchwork quilt.

Host 1: And finally, connecting this back to the bigger picture, how do all these efforts, the regulations, the sandboxes, the standardization, how do they ultimately help advance the collective goal of health equity?

Host 2: That's maybe the most important question. LLMs have this profound potential to either bridge existing health gaps or, unfortunately, make them much worse. The research clearly shows that biases can creep in and accumulate throughout the whole process: data collection, processing, modeling, especially racial, gender, geographic, and linguistic biases that are often already present in the training datasets scraped from the internet. There's been testing using hypothetical clinical scenarios, and it's revealed clear gender and racial biases in LLM responses. These might not show up easily in controlled clinical trials, but they become really critical public health concerns when these tools are deployed in the real world.

Host 1: So the potential for harm is real.

Host 2: It is. The ultimate goal has to be ensuring intentional inclusion. We need perspectives from low- and middle-income countries, LMICs, actively included in these global discussions. It's about democratizing resources, extending support, and making sure that digital and AI interventions are accessible, affordable, safe, and effective for everyone. For example, some studies in LMICs have actually shown benefits like LLM assistants helping people manage diabetes, but those same studies also highlighted the critical need to address challenges like poor data quality and integrating these tools with the existing, often resource-limited healthcare infrastructure.

Host 1: Wow. This deep dive has really shown us just how complex, how dynamic this whole world of AI governance in healthcare is right now. We've covered a lot, from that fundamental mismatch between traditional product regulation and these new AI tools, to the need for innovative adaptive frameworks, diving into those crucial ethical questions, and seeing why securing the entire AI supply chain is so vital.

Host 2: Yeah, the future of AI in healthcare really isn't just about the cool technological advancements. It's profoundly about how we, collectively, regulate it, how we work together globally to ensure safety, efficacy, and equity for diverse populations everywhere.

Host 1: So what does this all mean for you listening in? While LLMs are undeniably continuing to transform healthcare at an amazing speed, the true measure of our success, well, it won't just be in their incredible innovation. It'll be in our collective ability to ensure they are developed and deployed ethically, safely, and equitably for everyone everywhere. As you think about this rapidly evolving landscape, maybe consider: What's one new question this deep dive has sparked for you about how AI might reshape your own healthcare experiences in the future?