1 August 2025 · 22 min
AI in the ER: Can and should AI Save Lives Under Pressure?
Emergency rooms run on speed, pressure, and life-or-death decisions. Can artificial intelligence really help?
In this episode, we explore how AI is reshaping emergency medicine—enhancing diagnosis, predicting patient outcomes, and streamlining critical decision-making in real time. Based on a cutting-edge report, we break down the Map–Measure–Manage framework that defines how AI tools can support clinicians at the bedside.
You’ll learn:
How AI is already being used to read scans and triage patients
Where predictive algorithms are improving outcomes—and where they still fall short
What stands in the way: data silos, regulation, and medicolegal risk
Why AI won’t replace emergency physicians—but might become their sharpest tool
This is essential listening for clinicians, technologists, and anyone tracking how AI intersects with real-world patient care.
Transcript
Automated transcript of the audio; it may contain errors.
Host 1: Imagine walking into an emergency room, maybe 10, 15 years from now. What's different? How are they figuring out what's wrong when, you know, every single second counts? And maybe the big question is, are doctors still the ones calling all the shots, or something else?
Host 2: Yeah.
Host 1: Well, something else helping out. Today, we're really going deep on this powerful force. It looks set to completely reshape healthcare, and especially the fast-paced world of emergency medicine. And that force is, of course, artificial intelligence. Our conversation today is sparked by a really fascinating article. It's called The AI Future of Emergency Medicine, written by Dr. Robert J. Petrella. And our goal here is simple: to give you a shortcut basically to understanding the huge changes AI is likely to bring, the uh the hurdles it's facing, and what this all really means for you and your healthcare down the line. Dr. Petrella lays things out across three stages, which kind of overlap, but are distinct: map, measure, and manage. We'll dig into those. But first, you know, we throw around terms like AI all the time, they can get a bit fuzzy. So just Just really quick: AI, artificial intelligence, that's the big umbrella, right? Computers doing smart, human-like things. Now, you've got machine learning, ML. That's where the systems actually learn from data. They tune themselves, create their own rules almost, instead of just following pre-programmed ones. And deep learning, DL, that's a type of ML, uses these complex, layered networks. You hear about neural networks, that's often DL. And finally, the ones everyone's talking about now, large language models, LLMs. Think GPT, Bard, that sort of thing. They're a kind of deep learning, huge models that chew on language prompts and generate responses. Okay, so with that foundation, let's jump into Dr. Petrella's first stage, mapping. This is all about finding the problems and figuring out, okay, how could AI potentially solve this? And the article points out medicine's actually been doing this for a while already. You know, for years, EDs have had these uh more traditional rule-based systems.
Host 2: Oh yeah, clinical decision support. Things like checking for drug interactions or allergies, maybe helping with a first pass on an ECG reading, even voice recognition for doctors dictating notes. That's been around.
Host 1: Right, but now we're seeing things get much more sophisticated, aren't we?
Host 2: Absolutely.
Host 1: Yeah.
Host 2: What's really exciting now is how much further we're pushing with the machine learning-based AI. These algorithms are already stepping up in decision support. A big one is radiology, interpreting scans. Super critical and maybe a radiologist isn't right there, you know, middle of the night.
Host 1: Okay.
Host 2: We're seeing actual products out there like Aidoc that helps spot brain bleeds fast, or GE Healthcare's Critical Care Suite and Nanox's Health PNX. They're designed to find a collapsed lung, a pneumothorax, on chest X-rays. And this isn't just lab stuff. For certain tasks, the accuracy of these AI systems is getting really close to human radiologists. Some hospitals are even using AI to flag the abnormal scans, push them to the top of the pile.
Host 1: Wow. That sounds, well, potentially game-changing for actual diagnosis and maybe predicting how a patient might do. What are some of the really promising ways it's impacting care right now?
Host 2: Well, the article mentions quite a few interesting areas. For instance, there are AI tools already FDA-approved for acute ischemic stroke, and they seem to be genuinely cutting down the time it takes to get to treatment, like mechanical thrombectomy.
Host 1: Which is crucial for stroke, yeah.
Host 2: Absolutely critical. Saves brain tissue, improves outcomes. We're also seeing really cool work with ML algorithms identifying patients at high risk for nasty stuff like, uh, patients with febrile neutropenia who might get drug-resistant infections, or predicting who's likely to develop sepsis. You might have heard of tools like Sepsis Watch or the Epic Sepsis Prediction Model. These are trying to catch it early.
Host 1: Sepsis is a huge killer, so anything there is big.
Host 2: Huge. And there's a lot of research going into using AI for ED triage, you know, sorting patients when they arrive. At least one algorithm is apparently in use, though we're still waiting on the big validation studies for that one. And it's not just diagnosis. People are building AI models to suggest specific treatments too, like recommending the best antibiotic for a particular case.
Host 1: Okay, that sounds incredibly powerful, but how are these AI models actually helping? Are they just saying diagnosis X? Or is there more to it? And this gets us to something really key you mentioned, model interpretability. Can you sort of unpack that for us? Why is it so important to know how the AI reached its decision, especially in the ED?
Host 2: Yeah, that's a That's really a fundamental question here. The The article points out two main ways AI can help clinicians. One is like a black box, you put data in, you get an answer out, but you don't really know why. The other way is AI helping doctors actually create better clinical decision rules, ones they can understand. So, model interpretability is just that. Can the model explain itself? Can it show the connections it found in the data in a way a human gets? Some models are naturally easier to understand, like simple decision trees, but then you have these incredibly powerful deep learning models, including the LLMs, and they often act like those black boxes. It's a bit like, you know, Daniel Kahneman's System 1 and System 2 thinking?
Host 1: Thinking, Fast and Slow.
Host 2: Yeah, exactly. These black boxes are kind of like System 1, fast, intuitive, pattern matching, but often hard to explain the how. And why this matters so much to you as a patient and definitely to doctors is that we generally want to understand the reasoning behind big medical decisions, especially when the stakes are really high, right?
Host 1: Like life or death decisions.
Host 2: Precisely. Or deciding on a treatment with serious side effects. The article uses the HEART score for chest pain as an analogy. We know higher age increases risk, so the score goes up. That makes sense. But what if some opaque AI model said, 'Actually, being older lowers the risk here,' and gave no reason? We'd be completely stuck, unable to verify its logic.
Host 1: You wouldn't trust it.
Host 2: You wouldn't. There's a study mentioned about predicting low oxygen in kids. It was accurate, but it completely failed to tell the doctors which kids were most at risk or why. That's a huge problem. When models are juggling hundreds, maybe thousands of data points, it's just a daunting analytical challenge, as the article puts it.
Host 1: Okay, I definitely see the problem with the black box. Now, what about these newer large language models like GPT-4? There's been a lot of buzz, maybe hope, that they could sort of self-explain by just talking through their reasoning. Is that panning out? Are they solving the interpretability issue?
Host 2: Well, people are certainly looking into it. It's a really active research area, but honestly the results so far are pretty mixed. Yeah, LLMs can generate text that looks like an explanation, but often those explanations are, frankly, implausible or inconsistent. They don't quite hold up. Plus, they have this known issue of hallucination, just making stuff up,
Host 1: Mhm.
Host 2: or giving advice that's just plain wrong, medically speaking.
Host 1: Right, I've heard about hallucinations.
Host 2: Yeah.
Host 1: Not great in a medical context.
Host 2: Definitely not. There was a Stanford study mentioned that found the advice from GPT-3.5 and GPT-4 didn't really line up well with what their own expert human informatics teams recommended. So the takeaway is LLM use in emergency medicine is really, really early days. The article says only nine published papers as of mid-2023. That number's going to explode, no doubt, but we're just starting. But you know, besides making decisions, there's maybe a third way AI could help here: just summarizing existing medical knowledge relevant to a case, maybe pulling up differential diagnoses or treatment options for the doctor to consider. That seems less fraught with the black box issue.
Host 1: Okay, so we've explored the potential, the mapping stage, but like you said, just because something could work doesn't mean it does work in the chaos of a real emergency department, which brings us squarely to the measurement stage, really proving these tools work for actual patients.
Host 2: Yeah, and this is where we hit what the article calls the AI chasm. It's this gap, sometimes a huge gap, between how well a model does in testing using data it's familiar with and how it performs out in the wild.
Host 1: Okay, the AI chasm.
Host 2: A classic example given is the Epic Sepsis Prediction Model. Lots of hospitals use it, but a big study in 2021 found it really didn't perform as well in practice as people had hoped. And the numbers really highlight this chasm. In 2021, researchers found over 19,000 AI studies related to clinical care. 19,000.
Host 1: Dang.
Host 2: But guess how many were rigorous randomized controlled trials? 41. And only 51 studies actually involved AI making decisions in real clinical settings. 51 out of 19,000.
Host 1: That's startlingly low.
Host 2: It tells you we are just scratching the surface of this measurement stage. We're really just beginning to properly test these things. Now, it's not all bad news. The article mentions during COVID-19 machine learning helped quickly pinpoint an existing drug, baricitinib, as a potential treatment, and that was later validated in trials. So it can happen. We just need way, way more of that rigorous real-world validation.
Host 1: And getting that validation seems incredibly hard, especially because of issues with the data itself, right? What are the main roadblocks there?
Host 2: Yeah. Oh, there are quite a few. First, just data fragmentation. Your medical info might be spread across different hospitals, clinics, labs. It's not all in one neat place.
Host 1: Right.
Host 2: Then there's data locality. An AI trained on patients in Boston might not work as well in, say, rural Montana. Different populations, different common conditions, maybe even different ways doctors practice. Data representation is another one. We just have less data available for certain groups, kids, for example, so AI might not be as well trained for them.
Host 1: That could lead to gaps or biases.
Host 2: Exactly. And speaking of bias, these models often learn from data created by humans, like radiology reports. If there were errors or ambiguities in those original reports, the AI learns those, too. It inherits our mistakes, basically.
Host 1: That makes sense. And you mentioned bias itself being a major headache.
Host 2: A huge one. Think about genetic databases. They heavily overrepresent people of European descent. So AI trained on that might not work as well for people with different genetic backgrounds. The article had this really stark example from COVID-19 prediction models. One model learned that patients lying down were sicker.
Host 1: Lying down?
Host 2: Yeah. Why? Because it was trained partly on ICU data where patients are usually lying down. Now, that correlation is technically true in that dataset, but it's completely useless, even misleading, for diagnosing someone walking into the ED.
Host 1: Wow, that's subtle, but really problematic.
Host 2: Totally. And another subtle thing is sometimes the most important clue is just the doctor's gut feeling, their general impression that, you know, the patient just looks terrible.
Host 1: Right, that experience factor.
Host 2: Yeah, but that stuff rarely gets systematically recorded in the electronic medical record. So how do you train an AI on it? You could try video recording patients, maybe, but then you run headfirst into massive privacy issues.
Host 1: And that brings us right to those broader privacy concerns. It's a minefield: whose data gets used, should people have to opt in, but if they do, does that skew the data? Do you pay people for data? What are the ethics there?
Host 2: Ah, these are enormous questions, and honestly society hasn't figured them out yet. There are large data networks like PCORnet and ACT Network in the US that pool data from millions of patients. They help tackle the fragmentation and locality problems, but uh they have different rules about privacy and security. The US has generally been more hands-off about how companies use data compared to Europe, for example. The big data breaches and now AI are forcing a rethink. We have HIPAA, of course, but only a handful of states have specific laws governing automated decision-making using personal data.
Host 1: So the rules are still being written.
Host 2: Very much so. But, you know, the potential benefits of using large datasets for AI research are so huge, that might be the thing that ultimately pushes us to find solutions: better ways to anonymize data, maybe LLMs helping handle different data formats more easily. Necessity might drive invention here.
Host 1: Okay, so data, privacy, huge hurdles. But even if we figure those out, there's the whole regulatory piece. How do agencies like the FDA even approve something that can learn and change over time? That seems like another major challenge.
Host 2: It absolutely is. And the FDA has been dealing with medical software for decades, since like 1995, so they're not totally new to this. But the sheer pace and complexity of AI now, regulators are struggling to keep up. They actually tried this thing called a precertification program for digital health apps, trying to approve developers rather than every single app update, but it just didn't work. Innovation was moving too fast.
Host 1: Interesting.
Host 2: Think about those symptom checker apps you can get on your phone. They could be risky if they give bad advice, right? Initially, regulators were pretty cautious. A study in 2022 found they were generally okay at ruling things out, but not great at ruling things in, kind of like a layperson's accuracy overall. And it's not just the FDA. Other agencies like the FTC and NIST are trying to set some ground rules for commercial AI things like needing explainability. But they're still working out the specifics for reliability, accuracy, safety. NIST has this AI Risk Management Framework, but it's not mandatory. It's, well, you said it, it's kind of the Wild West still.
Host 1: Wild West indeed. Okay, so we've mapped the potential, we've looked at the huge challenges in measuring if it works and regulating it. This all leads, hopefully, to the management stage. This is where AI tools actually get woven into the daily fabric of the ED, maybe even guiding the whole process. What could that future realistically look like?
Host 2: It's tough to say exactly, but a big theme is medicine likely becoming more proactive, more preventative, detecting things before they become emergencies. But, let's be real, accidents happen, sudden illnesses happen, we'll always need emergency services. I picture a future with more embodied AI, not just software, but AI built into robots, using computer vision, understanding speech, really interacting with the physical environment.
Host 1: Like robots helping at the bedside?
Host 2: Maybe. Performing diagnostics, assisting with procedures. It sounds sci-fi, but elements are already being developed.
Host 1: That's a compelling vision.
Host 2: Yeah.
Host 1: But getting there, huge obstacles. You touched on legal issues earlier. That medico-legal question, who gets sued if an AI makes a mistake and hurts someone? The developer? The doctor who used it? The hospital?
Host 2: That is probably the biggest unanswered question right now, and it's holding things back. There's almost no case law to guide us. Traditionally, there's this idea of the learned intermediary, basically. If the drug or device maker warns the doctor about risks, the maker might be off the hook, and the doctor is responsible for using it correctly. But courts haven't really wanted to apply standard product liability to software. Is software a product? Is it a service? Is it speech? And the article makes a great point: current laws might actually discourage doctors from using AI if it suggests something different from the standard guidelines, cuz that feels riskier legally.
Host 1: Which kind of defeats the purpose of personalized AI medicine.
Host 2: Exactly. So now there's a push, maybe software developers should have more liability if their AI is giving complex medical advice. It's a massive debate with huge implications for whether doctors and hospitals feel safe adopting these tools. Without clear rules of the road, everyone's hesitant.
Host 1: Yeah, absolutely. And besides the legal fog, what about just the practical, nuts-and-bolts challenges of using AI in the ED? The time pressure must be immense.
Host 2: Oh, incredible pressure. You need answers now. The article uses the example of tension pneumothorax, a collapsed lung pressing on the heart. You need sensors, analysis, diagnosis, probably within a minute or two. There's no time for the AI to churn slowly. And ED clinicians have to trust that AI recommendation instantly. It's not like an oncology meeting where a tumor board can discuss a complex case for an hour. In the ED, you often have to act immediately based on the information you have. That requires immense trust in the AI's speed and accuracy.
Host 1: And even if an AI is accurate when it's first deployed, does it stay accurate? You mentioned model drift.
Host 2: Right. AI aging, it's a real thing. The world changes: patient demographics shift, diseases evolve, think about new virus strains or even just how medicine is practiced changes. The example given was pre-COVID models for predicting hospital admissions or sepsis. When the pandemic hit, suddenly they started throwing up way more false alarms because the patient mix and symptoms were totally different.
Host 1: So the models got out of date.
Host 2: Exactly. They need updating, retraining. But a lot of these AI tools are proprietary black boxes sold by companies. The hospital might not have much control over updates. This is where open-source AI development could be really helpful, letting hospitals maybe tweak and retrain models more often, using their own local, current data.
Host 1: This keeps circling back to that black box problem, doesn't it?
Host 2: Uh-huh.
Host 1: And maybe a deeper question about AI's actual reasoning skills. You mentioned common sense earlier, like knowing someone with two broken arms can't use crutches, that seems basic to us, but hard for AI.
Host 2: It really is. Deep learning is fantastic at that System 1 pattern recognition, boom, recognize that face, spot that anomaly on the scan. But System 2, thinking slow, logical, analytical reasoning, the kind you need for scientific discovery or complex planning, that's still a major challenge for AI. Now, modern LLMs like GPT-4 show some reasoning ability. I mean, scoring in the 88th percentile on the LSAT is nothing to sneeze at.
Host 1: No kidding.
Host 2: But it's still brittle, especially with common-sense stuff. It seems like they're mimicking reasoning based on statistical patterns in the vast amounts of text they were trained on, rather than truly understanding logical rules. If they get better at genuine reasoning, maybe they could explain themselves better, too. But this leads to this almost philosophical question, right? There's this mathematical idea, the Universal Approximation Theorem, that basically says a complex enough neural network can, in theory, learn to predict almost anything accurately if you give it enough data. So, if an AI could tell you with incredible accuracy the absolute best course of action, does it really matter if we don't fully understand how it figured that out?
Host 1: That's a tough one. What do you think?
Host 2: Well, right now, today, patients and doctors are generally not okay with a machine making critical decisions without a human involved, and without some understanding of the why. Interpretability feels essential, especially for high-stakes choices. But you could imagine a future, maybe decades from now, where we've just gotten used to AI recommendations being incredibly reliable, even if we can't fully follow the logic. Maybe trust could build over time based purely on outcomes. It feels weird now, but maybe. Think about that trauma patient example again: AI says 78% survival with a chest tube, 33% without, but maybe the doctor's gut feeling based on experience was the opposite. Without a clear, verifiable explanation from the AI, that doctor is in a really difficult spot. Trust requires more than just a number. It requires some shared understanding, and that's the hurdle.
Host 1: It really is. So, wrapping your head around all this, who or what is ultimately going to shape this future? Is it doctors, tech companies, patients? The article mentions the American Medical Association prefers 'augmented intelligence,' not 'artificial,' to stress the human remains central.
Host 2: Yeah, they emphasize AI assisting humans, not replacing them. But there's likely going to be a tradeoff, right? Efficiency versus human control. If AI genuinely gets better and faster at certain tasks, it's natural that human involvement in those specific tasks might decrease. It's like the self-driving car analogy: if they become statistically much safer, maybe eventually human driving gets restricted in some areas.
Host 1: Okay, so what could that mean for jobs? Are there estimates?
Host 2: There are some pretty eye-opening ones. Goldman Sachs estimated something like 28% of tasks in healthcare are potentially exposed to AI automation. McKinsey got more specific, suggesting about a third of tasks done by doctors and nurses could potentially be automated. They figure this could free up maybe 12% of a physician's time, maybe 8% of a nurse practitioner's time by 2030. That's a significant amount of time potentially shifted.
Host 1: Wow. So what's left for the clinician? Will they still be the final decision-maker?
Host 2: I think so, yes. Especially for applying judgment to things outside standard guidelines, you know, incorporating the patient's specific feelings, their values, their preferences, dealing with brand new drugs or treatments AI hasn't seen much data on, or handling situations where a patient refuses a recommended procedure. That uniquely human element of care, understanding the person, not just the disease pattern. And there's this other intriguing, maybe slightly unnerving idea that the article raises. If AI gets really good at predicting what works even without explaining why, could our fundamental understanding of medicine actually fall behind our ability to treat things effectively?
Host 1: Huh. Prediction outpacing insight.
Host 2: Exactly. We already use some drugs, lithium, acetaminophen, even some anti-seizure meds like levetiracetam, where we don't fully understand their precise mechanisms of action, but we know they work. The article suggests maybe future breakthroughs will come from AI spotting patterns we just don't grasp initially. We might find solutions first and figure out the why later. It's a different way of doing science, potentially.
Host 1: Okay, so to kind of pull this all together from our deep dive today, we've seen the incredible potential AI has in emergency medicine, helping diagnose, treat, manage flow. But we've also seen the huge AI chasm, the gap in real-world proof. We've wrestled with the black box problem and why understanding the why, interpretability, feels so crucial right now. We've also noted that LLMs, while exciting, are still very early stage and prone to errors in this context. And then there are these massive, tangled challenges around liability, regulation, data privacy, and bias. It seems clear there's an unavoidable tension, this tradeoff between relying on AI for efficiency and keeping the human clinician firmly in the loop.
Host 2: Definitely. So maybe the thought to leave you with as you mull this over is this: if AI reaches a point where it consistently provides better predictions, better treatment recommendations, leading to better outcomes, but it can't fully explain its reasoning in a way we grasp, will we as a society eventually make that trade? Will we choose optimized results over our deep-seated need for human understanding in medicine? And what does that choice, if we make it, imply about how we value knowledge versus outcomes in this coming digital age? It's something we'll all likely have to confront.