16 May 2025 · 13 min
Can We Trust AI in Healthcare? Unpacking the National AI Code of Conduct
In this episode, we delve into the National Academy of Medicine's draft AI Code of Conduct, exploring its implications for healthcare. We discuss the proposed principles and commitments designed to ensure the ethical, safe, and effective integration of AI in health and biomedical sciences. Join us as we unpack the framework aiming to guide stakeholders toward responsible AI adoption in healthcare settings.
Transcript
Automated transcript of the audio; it may contain errors.
Host 1: Welcome to the deep dive. Today we're tackling artificial intelligence. It's really changing the game in health, isn't it?
Host 2: Absolutely, healthcare, biomedical science, [laughter] touching touching everything.
Host 1: We've pulled together some material looking at AI's potential, yeah, but also the, um, the considerations around it.
Host 2: And the goal for you listening in is to get a solid handle on this. We want to cut through the noise, give you the key things to think about with AI and health without, you know, getting bogged down in jargon.
Host 1: Exactly. I mean think about it, it feels like just yesterday, 2022, when things like ChatGPT really burst onto the scene. Right? And the speed since then, oh. Wow, it's kind of breathtaking.
Host 2: It really has, which brings huge possibilities, but also, uh, some pretty big questions.
Host 1: Which is why we're zeroing in on this recent commentary from the National Academy of Medicine. They're proposing an initial framework, like a sort of AI code of conduct for health.
Host 2: Yeah, and that's what we'll unpack in this deep dive. We'll look at the, um, the upside, how AI could really improve things.
Host 1: But also the pitfalls.
Host 2: But also the pitfalls, the challenges. We want to understand the path they're suggesting for using AI responsibly in these really critical areas.
Host 1: Okay, let's get into it and, you know, to set the scale, Stephen Hawking once said AI could be the biggest event in the history of our civilization.
Host 2: That's a powerful statement.
Host 1: Huge, right.
Host 2: Yeah, but when you apply it to health, something so fundamental, it starts to feel very real. People have talked for years about this idea of a learning health system.
Host 1: Right, I remember that term. The Institute of Medicine defined it back in, what, 2007?
Host 2: Exactly, the idea being a system that's constantly learning, feeding new knowledge right back into practice to improve care continuously.
Host 1: And we've seen bits and pieces working towards that, haven't we? Like, expert systems, clinical decision support tools.
Host 2: For sure, machine learning has been chipping away at it. Those were important steps, definitely incremental progress.
Host 1: But modern AI feels different.
Host 2: That's what the NAM highlighted in a 2022 publication. They called it Artificial Intelligence in Healthcare: The Hope, the Hype, the Promise, the Peril. Catchy title.
Host 1: Uh, yeah.
Host 2: But it made the point that today's AI isn't just about small improvements, it's potentially transformative. It could fundamentally change how we approach health.
Host 1: But the title also hints at the peril. That same report raised flags, right, fairness, safety.
Host 2: Privacy, yeah. Absolutely critical concerns. And one really core issue is this idea of, um, data-driven objectivity.
Host 1: Ah, the idea that if it comes from data, it must be unbiased.
Host 2: Exactly, but it's not that simple. The data AI learns from, humans create it. Humans collect it. Humans decide what's important.
Host 1: So, it reflects our biases.
Host 2: It can, yeah. Think about historical medical data. Maybe certain groups weren't included as much in trials. If you train an AI on that data, it might not work as well for those underrepresented groups. It could actually make health disparities worse. There was some important work by Obermeyer and others back in 2019 showing exactly this.
Host 1: Gotcha, so the AI is only as good or as fair as the data it learns from.
Host 2: Precisely, it's reflecting the world it sees in the data, warts and all.
Host 1: What other risks did that 2022 NAM report mention?
Host 2: Well, things like misdiagnosis, that's a big one, or, uh, recommending unnecessary tests or treatments leading to inefficient resource use.
Host 1: Mhm.
Host 2: Privacy breaches, obviously, with sensitive health data. And then there's the impact on the workforce.
Host 1: You mean becoming too reliant on AI?
Host 2: Yeah, potentially de-skilling clinicians or maybe they become less vigilant because they assume the AI is always right. Matheny and colleagues talked about this.
Host 1: And these AI models, they don't just stay the same, do they? They keep learning.
Host 2: That's another layer of complexity. They adapt based on new data, which is good, they can improve. But it can also make it harder to understand why they're making certain recommendations. Their internal logic can shift, become more of a, you know, a black box.
Host 1: Right. Okay. So, things are moving fast. There are clear benefits, but also real risks. And the tech itself is evolving. Makes the case for guard rails pretty strong.
Host 2: Definitely. And as Hutson noted, the arrival of things like ChatGPT really put an exclamation point on the speed of change since 2022. Everyone in health, doctors, hospitals, patients, developers needs to learn and adapt like now.
Host 1: And we're not starting from scratch when thinking about rules, are we? Healthcare already has quality principles.
Host 2: Exactly. This push for AI guidelines connects directly to the core ideas of that learning health system we mentioned, principles laid out in IOM reports going back to 2000, 2001.
Host 1: Like making care safe, effective, patient-centered.
Host 2: Timely, efficient, equitable.
Host 1: Yeah.
Host 2: Yes, and over time, things like transparency, accountability, and security got added, all super relevant for AI.
Host 1: So, we have a foundation, and the NAM commentary acknowledges other AI guidelines already exist.
Host 2: Oh, yeah. Lots of groups have put out frameworks: government bodies like the White House, big tech companies like Google AI, academic outfits like MITRE,
Host 1: Even earlier efforts.
Host 2: Right, like the Asilomar AI Principles from 2017. They were thinking about broader AI safety and ethics even then.
Host 1: Okay, so if these guides exist, why the need for this new NAM framework? What's missing?
Host 2: Well, a couple of things. One big issue is, um, alignment. The existing frameworks don't all point in the same direction, or they use different language for similar ideas. That makes it hard for, say, a hospital system to adopt a single, coherent approach.
Host 1: Lack of convergence.
Host 2: Exactly. And some are pretty high-level, maybe too abstract, difficult to translate into day-to-day practice.
Host 1: Like be transparent, okay, but how, specifically with this complex algorithm?
Host 2: You got it, and that ties into another challenge that the commentary flags: explainability, especially with these newer, more complex models like LLMs.
Host 1: The black box problem again.
Host 2: Yeah, it's hard to demand transparency when even the creators might not fully grasp every step of the model's reasoning. It's not like asking a human doctor for their rationale.
Host 1: Okay, so then NAM is trying to bridge these gaps. Their proposed AI code of conduct, what's the core idea?
Host 2: The main goal is to build trust, trustworthy AI in health. They want to do that by offering a harmonized set of principles and some pretty straightforward rules.
Host 1: Two main parts you said.
Host 2: Right. First, the code principles. These are the bedrock values, directly build on those learning health system principles: think engaged, safe, effective, equitable, efficient, accessible, transparent, accountable, secure, and adaptive.
Host 1: Okay, the core values, and the second part?
Host 2: The code commitments. These are meant to be the simple, actionable rules for applying those principles in the real world. Interestingly, they draw on ideas from complex adaptive systems theory, or CAS.
Host 1: CAS theory, can you break that down simply?
Host 2: Sure, CAS looks at systems with lots of interacting parts like healthcare. A key idea is that you don't need super complicated rules for every single interaction. Sometimes, a few simple, widely understood rules guiding individual actions can lead the whole system toward a desirable state.
Host 1: Ah, so the commitments are like those simple rules.
Host 2: That's the idea. Make them clear, make them broadly acceptable, so everyone involved: developers, doctors, patients, policy makers has a shared reference point for making responsible choices as things happen.
Host 1: Equip people with awareness and guidance in the moment.
Host 2: Exactly, and the hope is this accelerates progress towards AI in health that ticks all the right boxes: safe, effective, ethical, equitable, reliable, responsible, all that good stuff.
Host 1: Sounds like they did their homework reviewing what's already out there.
Host 2: They did. They looked at a big 2022 literature review by Ciala and Wang that identified key traits of responsible AI: human-centeredness, fairness, transparency, and so on.
Host 1: And then compared their framework to others.
Host 2: Yep, 56 different documents: scientific papers on AI principles from 2018 to 2023, guidance from medical societies for doctors using AI, and US federal government policies up to mid 2023.
Host 1: A broad sweep. What jumped out from that comparison? What was common and what was maybe overlooked?
Host 2: Fairness and transparency came up a lot, pretty consistent themes.
Host 1: Okay.
Host 2: But other important aspects like inclusivity, sustainability, really focusing on human needs, were less common in the documents they reviewed. And, crucially, they felt that things like accountability, data protection, ongoing assessment, and safety needed more emphasis than they found in many existing frameworks, which is why their principles tie back so strongly to the learning health system concepts.
Host 1: Did they look beyond the US?
Host 2: Yes, they brought in international guidance, too: WHO, UN, EU, OECD. Generally, those aligned well with the proposed code principles.
Host 1: Any differences internationally?
Host 2: One interesting point was the stronger emphasis in international documents on environmental protection or efficiency. That theme was pretty much missing from the US focused literature they looked at.
Host 1: Interesting. So, this review helped pinpoint areas needing more attention. What were the main gaps they identified?
Host 2: Three big ones stood out. First, inclusive collaboration,
Host 1: Mhm.
Host 2: meaning getting everyone involved across the whole life cycle of an AI tool: developers, clinicians, patients, administrators, community groups, diverse voices from different sectors, backgrounds, roles.
Host 1: Why is that so crucial?
Host 2: Well, think about it: to make sure you're solving the right problems, using appropriate data, avoiding bias, integrating the tool smoothly, training users properly, monitoring it effectively, and being clear about who's responsible for what. Without that broad input, you risk developers building things based on their own assumptions, which might not work for everyone, or could even be discriminatory.
Host 1: Makes sense. What was the second gap?
Host 2: Ongoing safety assessment. We touched on this. AI development, especially the advanced stuff, is moving way faster than our usual safety checks and regulations.
Host 1: Right, it's not like approving a standard medical device.
Host 2: Exactly. These adaptive AIs can change over time. You get model drift. What works safely on day one might not work the same way 6 months later, or in a different hospital setting. We need ways to monitor that continuously.
Host 1: And there are bigger picture safety concerns, too.
Host 2: Yes, potential societal impacts: Could AI worsen health disparities? Could it lead to monopolies in health tech? Could it impact healthcare jobs and worker power? There are efforts like the White House executive order in 2023, but the commentary really pushes for building a deep safety culture, not just ticking regulatory boxes.
Host 1: Okay, and a third gap, the one highlighted more internationally.
Host 2: Right, efficiency or environmental protection. It takes a ton of energy to train and run these big AI models, massive data centers.
Host 1: The carbon footprint of AI.
Host 2: Exactly. And resource used too, minerals for the hardware, water for cooling the data centers, it's significant.
Host 1: So, an environmental cost to this efficiency promise.
Host 2: Potentially, yeah. And the commentary points out this just wasn't really discussed much in the documents they reviewed, especially the US ones. They argue it needs to be part of the conversation about responsible AI.
Host 1: So, considering these gaps, they develop the draft principles and commitments. Can you quickly recap those?
Host 2: Sure. The 10 code principles are the values drawn from the learning health system: engaged, safe, effective, equitable, efficient, accessible, transparent, accountable, secure, and adaptive.
Host 1: The what.
Host 2: Right, and the six proposed code commitments are the how, the actionable rules: focus on human health connection, benefits distribute fairly, involvement, partner with people, workforce well-being, support clinicians, monitoring, share performance info openly, and innovation, keep learning, improving.
Host 1: All aimed at building trust.
Host 2: That's the goal: foster trust in AI within the health sector.
Host 1: It feels like a solid starting point. What happens next with this NAM framework?
Host 2: Well, they clear this is a draft. They're actively asking for feedback right now.
Host 1: From anyone?
Host 2: From anyone interested, yeah. They'll convene working groups to flesh things out, test the framework using real-world case studies involving patients, hospitals, developers, and they'll consult with regulatory bodies and professional groups.
Host 1: And the end goal?
Host 2: To release a final AI code of conduct framework along with recommendations on how to actually implement it, monitor it, and keep refining it over time.
Host 1: And listeners can weigh in.
Host 2: Absolutely. The commentary has a link for submitting comments on the draft. We'll definitely put that in our show notes so you can find it easily.
Host 1: This really does feel like a critical juncture, doesn't it? We have this powerful tool
Host 2: with huge potential,
Host 1: and a real need to steer it carefully.
Host 2: It requires collective action, you know, intentional choices about how we want AI to shape the future of health.
Host 1: Well, it seems this proposed code of conduct from the NAM is a really important piece of that puzzle. It's trying to lay out a path that balances innovation with responsibility,
Host 2: aiming to advance health and well-being for everyone, hopefully.
Host 1: So, as we wrap up, maybe the thought to leave with you, our listener, is this: What's your role, or the role of people you know, whether as a patient, a clinician, a developer, a citizen, in making sure these kinds of principles and commitments actually stick? How do we embed them in our everyday interactions with AI in health?
Host 2: That's the question, isn't it? How does it translate from a framework on paper to reality on the ground? Big implications there.