5 October 2026 · 18 min

FDA Clears Bayesian's AI Sepsis Tool: Analyzing the Data Behind the 510(k)

The FDA has granted 510(k) clearance to the Bayesian Health Sepsis Flagging Device, a continuous AI monitor that uses 115 electronic health record inputs to predict sepsis risk. Maya and Sam break down the 10-page clearance document, exploring the device's clinical validation study, its 79.4% positive percent agreement, and what an 11.7% positive predictive value means for alert fatigue.

Key points

Source: Bayesian Health Sepsis Flagging Device - U.S. Food and Drug Administration, 2026

This spot is available. Reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsor the show

This episode is an AI-generated conversation summarising a public document; the hosts' voices are synthetic. It is for information only and is not medical advice. Always refer to the original source.

Your company here. This podcast is looking for its first sponsors: reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsorship options and rates →

Transcript

Maya: Imagine a hospital AI tool securing FDA clearance even though it missed its own primary effectiveness target by a fraction of a percent. Well, the FDA just cleared the Bayesian Health Sepsis Flagging Device, an AI model that looks at 115 patient parameters to predict sepsis, despite it hitting 79.4% instead of its 80% target for positive agreement.

Sam: That is fascinating, and it brings up so many questions about how the FDA actually weighs clinical evidence for artificial intelligence. We are looking at the official 510(k) Premarket Notification summary and clearance letter for this device, published by the U.S. Food and Drug Administration on April 30, 2026.

Maya: Before we dive in, a quick reminder that our voices are AI-generated. We are bringing you a summary of a publicly available regulatory document. This one is a big deal for clinicians, health system IT leaders, and anyone building predictive algorithms.

Sam: Sepsis prediction has always been the holy grail, and frankly, the graveyard, for a lot of clinical decision support tools. So what exactly is this document, Maya, and why does it matter right now?

Maya: This document is the official summary of safety and effectiveness from the FDA. It outlines exactly what the Bayesian Health Sepsis Flagging Device is, how it was tested, and the data that led the FDA to say it is substantially equivalent to legally marketed devices. It is classified as a Class II device under the regulation for software devices to aid in the prediction or diagnosis of sepsis.

Sam: Got it. And let us be specific about what this tool actually does. The document describes it as a continuously monitoring artificial intelligence and machine learning-based Software as a Medical Device. It is designed to be used by healthcare providers to aid in the early detection and risk prediction of sepsis developing within 24 hours.

Maya: Exactly. And the continuous aspect is key here. It does not just run once when a patient is admitted. It sits in the background, integrated into the electronic health record, or EHR. Within an hour of new data becoming available in the EHR, the device updates its prediction.

Sam: What exactly is it looking at? You mentioned 115 parameters earlier. That sounds like a massive amount of data to process continuously for every single patient.

Maya: It is. The core inputs include patient history, comorbidities, the chief complaint documented at presentation to the emergency department, laboratory measurements, vital signs, procedures, medications, and consult orders. It takes all of that and, if the risk is high enough, it outputs a single flag displayed in the EHR that reads Sepsis Risk High.

Sam: Just Sepsis Risk High. Okay, so it does not give a percentage score or a probability scale to the end user? It just gives that binary flag?

Maya: According to the output description in the document, yes, the output is the Sepsis Risk High flag displayed in the EHR, along with a summary of clinical factors potentially contributing to the patient being flagged. But the FDA is very clear on its intended use. It explicitly states that this flag should not be used as the sole basis to determine the presence of sepsis or the risk of developing sepsis within 24 hours.

Sam: Right, it is adjunctive information. It is meant to be used alongside clinical assessments and other laboratory data. So, who is this actually used on? Is it running on every single person who walks through the hospital doors?

Maya: Not everyone. The intended population is adult patients, meaning those 18 years old and older, upon emergency department presentation or hospital admission. It runs throughout the duration of the patient's stay in acute care settings. But the exclusions are just as important as the inclusions.

Sam: Let us hear the exclusions. Where did they not validate this?

Maya: The document clearly states the device has not been validated in pediatric patients, labor or maternity care unit patients, patients who have had surgery in the past 90 minutes, or those admitted for trauma who have not been stabilized.

Sam: That 90 minutes post-surgery window is really interesting. I imagine the body's inflammatory response to a surgical incision looks an awful lot like the systemic inflammatory response of early sepsis. The algorithm would probably throw a lot of false positives if it were running on people rolling straight out of the operating room.

Maya: That is exactly the kind of clinical nuance these algorithms have to navigate. It is also why the FDA requires a predicate device comparison. To get a 510(k) clearance, you have to prove your device is substantially equivalent to something already on the market. In this case, Bayesian Health compared their tool to the Sepsis ImmunoScore made by Prenosis Inc.

Sam: How do the two devices compare? Are they basically the same algorithm?

Maya: They share the same intended use, which is assisting healthcare providers in sepsis risk assessment through AI and machine learning analysis of EHR data. But there are significant technological differences. The Sepsis ImmunoScore uses up to 22 predetermined inputs, whereas Bayesian uses 115. Furthermore, the predicate device requires a blood culture to be ordered as part of the evaluation for sepsis before it runs.

Sam: Wait, really? So the older device only flags patients after a doctor has already suspected an infection enough to order a blood culture? That seems like it misses the whole point of an early warning system.

Maya: That is a major distinction. The Bayesian device is independent of a blood culture order. It applies to all adult hospitalized patients in the included settings, regardless of whether sepsis is suspected. It is casting a much wider net. The FDA noted this difference, stating that while the Bayesian device is a continuous monitor and does not rely on a blood culture order, these modifications do not change its intended use or raise different questions regarding safety and effectiveness.

Sam: Casting a wider net means you are going to catch more fish, but you are also going to pull up a lot of seaweed. Which brings us to the performance data. How did they actually prove that this thing works safely in the real world?

Maya: They ran a multi-site clinical validation study. And the numbers here are robust. They evaluated 7,732 hospital encounters from 7,298 unique patients. This took place across 3 clinical sites, which encompassed 4 distinct hospitals. They made sure to include academic, urban, and suburban hospitals so the population would be representative of practices across the United States.

Sam: Okay, 7,732 encounters. How do you actually prove that the algorithm got it right? Sepsis is notoriously difficult to diagnose retrospectively. Two different doctors looking at the same chart might disagree on whether a patient actually had sepsis, or when exactly it started.

Maya: You hit on the exact reason why validating these tools is a nightmare. To solve this, the study used a multi-tiered adjudication process. First, they used preliminary computer-assisted adjudication. This step looked for cases that were clearly negative, meaning they did not show any signs of infection and sepsis-related organ dysfunction based on criteria identified by infectious disease physicians.

Sam: So the computer filtered out the obvious non-sepsis cases first. What happened to the cases that were ambiguous?

Maya: Any encounters that were not screened as clearly negative were considered possibly septic. Those cases went to human physician adjudicators. The physicians reviewed the charts to determine the final status and the exact onset times. Crucially, these physician adjudicators were blinded to the outcome of the preliminary computer-assisted adjudication.

Sam: But what if the computer was wrong in the first step? What if it labeled a subtle sepsis case as clearly negative?

Maya: They accounted for that. They ran a verification bias analysis. A portion of the encounters that the computer labeled as clearly non-septic were given to the physician adjudicators anyway to verify that the computer was not missing things. It adds a layer of statistical rigor.

Sam: That is thorough. So, they have their gold standard truth established. How did the Bayesian algorithm perform against that gold standard?

Maya: This is where the numbers get really interesting. They measured performance at two levels: the encounter level and the flag level. Let us start with the encounter level. A true positive encounter was defined as an encounter that received any device flag occurring within 24 hours of the adjudicated sepsis onset.

Sam: And the target metrics? What did they have to hit for the FDA to be happy?

Maya: The pre-specified acceptance criteria for the study were that the lower bounds of the 95% confidence interval had to exceed 80% for both Positive Percent Agreement, or PPA, and Negative Percent Agreement, or NPA.

Sam: Did they hit the 80%?

Maya: For Negative Percent Agreement, yes, comfortably. They achieved an encounter-level NPA of 89.5%. Out of 7,499 non-septic encounters, the device correctly stayed quiet for 6,715 of them.

Sam: So it correctly identified a vast majority of the non-septic encounters. But what about the Positive Percent Agreement? Did it find the actual sepsis cases?

Maya: It found 185 out of the 233 septic encounters. That results in an encounter-level PPA of 79.4%. The 95% confidence interval was 74.2 to 84.6. So, the lower bound was 74.2, and the point estimate was 79.4%. Both are below the 80% target.

Sam: Wait, the document specifically says the acceptance criteria required the lower bounds of the confidence interval to exceed 80%. If the lower bound was 74.2, they missed the endpoint entirely. How does a device fail its own pre-specified acceptance criteria and still get FDA clearance?

Maya: The document addresses this directly. It states: Although the device did not meet the pre-specified acceptance criteria for encounter-level PPA, the totality of device performance, including flag-level performance, was used to establish substantial equivalence to the predicate.

Sam: Totality of evidence. That is a phrase that gives regulatory folks a lot of flexibility. It basically means, yes, we missed the rigid threshold, but when you look at the whole picture, the device is still safe and effective. So what was the flag-level performance that saved the application?

Maya: Because this is a continuous monitor, a single patient encounter might generate multiple flags. So they looked at every single prediction independently. Throughout the study, the device generated 1,895 flags. Out of those, 221 flags were generated during the clinically relevant timeframe, which is within the 24 hours surrounding the adjudicated sepsis onset.

Sam: Okay, so 221 true positive flags out of a total of 1,895 flags fired. That brings us to Positive Predictive Value, or PPV. What was the PPV?

Maya: The flag-level PPV was 11.7%. And it is crucial to note that this PPV is tied to the observed sepsis prevalence in the study, which was approximately 3%.

Sam: With an 11.7% Positive Predictive Value, you have to wonder about the remaining false positive flags. If nurses and doctors are bombarded by alerts that are mostly false positives, human nature dictates that they will eventually start ignoring the alerts. How does the FDA reconcile an 11.7% PPV with clinical utility?

Maya: The document explains that the low PPV is a function of the prevalence and the intended use. Because this tool acts as a dragnet for allcomers to the Emergency Department where the prevalence of sepsis is low, around 3%, the PPV naturally drops. They contrast this with the predicate device, which had a higher observed sepsis prevalence because it was only used on patients already suspected of having sepsis.

Sam: That makes mathematical sense. If you only test people who look sick enough to need a blood culture, your prevalence goes up, and your PPV goes up. By screening everybody, Bayesian lowers its PPV, but arguably provides more value by catching people the doctors haven't noticed yet.

Maya: Precisely. The document concludes that considering this difference in patient population, the flag-level performance of the Bayesian device was found to be substantially equivalent to the predicate. The FDA is acknowledging that an 11.7% PPV is acceptable for a broad, adjunctive screening tool, as long as it is not the sole basis for clinical decisions.

Sam: It is a trade-off. You accept a higher false positive rate to ensure you do not miss the subtle, early cases of sepsis. But getting a model like this into the EHR is not just about clinical accuracy. There is a whole technical and safety framework they had to satisfy. The document mentions the Software Level of Concern is Enhanced.

Maya: Yes, it meets the Enhanced level of concern according to the June 2023 FDA guidance document. The predicate device was classified as Moderate under an older 2005 guidance. Because of this Enhanced status, Bayesian had to submit extensive documentation on software verification and validation.

Sam: What kind of testing are we talking about behind the scenes?

Maya: They performed unit, integration, and system-level tests in compliance with IEC 62304. They did non-clinical performance testing of the algorithm, which involved simulating the impact of measurement variability by perturbing the inputs. They also evaluated the impact of observed and simulated input feature missingness.

Sam: Simulated missingness. That is huge. In the real world, EHR data is messy. A patient might be missing a lab result, or a vital sign hasn't been entered yet. The algorithm has to know what to do when some of those 115 parameters are just blank.

Maya: Exactly. They tested different rates of input feature missingness and feature imputation to see how it affected the model scores and overall performance. They also had to pass human factors testing to assess the comprehension and usability of the interface, ensuring it is safe and effective for the intended users. And, unsurprisingly, they had to provide cybersecurity risk analysis and testing.

Sam: Given all the recent ransomware attacks on health systems, a cloud-based web application seamlessly integrated into the EHR is a massive target. It is good to see cybersecurity explicitly called out in the clearance summary.

Maya: The clearance letter itself, signed by the Division of Renal, Gastrointestinal, Obesity, and Transplant Devices, notes that Bayesian must comply with the Quality Management System Regulation, including design controls and corrective actions. But perhaps the most forward-looking part of this document is the post-market requirement.

Sam: Right, because AI models can drift. Medical practices change, coding standards change, and suddenly a model's performance might degrade over its market lifetime. What exactly is the FDA making them do?

Maya: Bayesian developed a post-market performance management plan to monitor the device over its market lifetime. The plan outlines procedures for data collection, analysis methods, and monitoring both encounter-level and flag-level performance. If performance degradation is identified, the plan outlines the process for investigating the sources of the changes between the validation data and the real-world environment.

Sam: That is the holy grail of AI regulation. Not just clearing it once, but making sure the manufacturer is legally obligated to watch it degrade and fix it. They also have to communicate the device's performance to users if things go wrong.

Maya: Yes, they have to assess the impact on safety and effectiveness and perform corrective actions if required. This is the new reality for software as a medical device. You do not just ship it and forget it. You are on the hook for its performance indefinitely.

Sam: Let us zoom out for a second. What does this clearance mean for the boardroom and the ward? For a health system IT leader looking to buy an AI tool, this sets a very specific baseline. If you are buying a sepsis tool, you should be asking for its encounter-level PPA, its flag-level PPV, and its post-market monitoring plan.

Maya: And for the clinician on the ward, it means you might see a new alert in your EHR. But the FDA explicitly states this flag should not be used as the sole basis to determine the presence of sepsis. It is an aid. If the flag goes off, you still have to evaluate the patient, look at the labs, and make a clinical judgment.

Sam: It is also a clear signal to developers. The FDA is willing to look at the totality of evidence. You can miss a pre-specified statistical target like that 80% PPA threshold and still get cleared if your overall clinical utility is sound, particularly if you are screening a broad population with a low prevalence of disease.

Maya: That is a great memorable takeaway. Regulatory clearance is not always a pass-fail math test. Context, prevalence, and clinical workflow matter just as much as the raw confidence interval.

Sam: To recap, the FDA has granted 510(k) clearance to the Bayesian Health Sepsis Flagging Device. It uses 115 parameters to continuously monitor adults for sepsis risk, achieving an encounter-level PPA of 79.4% and an NPA of 89.5%. Because it screens broadly, it operates with a positive predictive value of 11.7%, generating flags that clinicians must weigh alongside standard care.

Maya: A quick reminder that the source document is linked in the show notes, and that this is not medical advice. We will see you next time.