9 October 2026 · 14 min
Company Spotlight: DeepHealth Saige-Dx - Using Prior Exams to Drop False Positives
DeepHealth’s Saige-Dx recently received FDA 510(k) clearance for analyzing digital breast tomosynthesis (DBT) mammograms. We unpack the public summary to see how this AI tool uses prior screening exams to reduce false positive alerts, and break down the exact testing demographics behind the clearance.
Key points
- Saige-Dx is an AI software cleared as a concurrent reading aid to identify soft tissue lesions and calcifications in DBT mammograms.
- This new clearance expands upon the predicate device by adding compatibility with Siemens systems and the ability to process prior exams.
- Processing prior exams was validated to reduce false positive marks per image by dropping false positive bounding boxes when a matching prior finding is identified.
- Performance was evaluated in two retrospective, standalone studies involving 2,425 cases and 684 cases, proving non-inferiority in AUC.
- The FDA clearance strictly labels Saige-Dx as an adjunct tool, warning that decisions should not be made solely based on its analysis.
Source: FDA 510(k) summary K253825: Saige-Dx (DeepHealth) - U.S. Food and Drug Administration, 2026
This episode is an AI-generated conversation summarising a public document; the hosts' voices are synthetic. It is for information only and is not medical advice. Always refer to the original source.
Transcript
Maya: The FDA just cleared an AI tool that looks at a patient's past mammograms to actively drop false positive alerts on their current scan. Welcome to AI in Medicine - Smart Summaries. Our voices are AI-generated, and today we are summarizing a public document published by the U.S. Food and Drug Administration: the 510(k) summary for DeepHealth's Saige-Dx.
Sam: This is a Company Spotlight episode. Just to be absolutely clear upfront, this is based strictly on the FDA's public 510(k) summary document, not on any marketing or promotional material from the company itself. We are looking only at what the regulators cleared.
Maya: Right. We are going to cover exactly what Saige-Dx does, how it works, the real numbers from their retrospective testing datasets of 2,425 cases and 684 cases, and what it actually means when a device is cleared based on substantial equivalence.
Sam: Let us start with that concept of substantial equivalence, because it sets the stage for everything else. The summary states this is a 510(k) premarket notification. That means DeepHealth had to prove their device is substantially equivalent to a legally marketed predicate device.
Maya: Exactly. The predicate device they compared it to is actually an earlier version of their own software. Specifically, it is their own Saige-Dx software, which was cleared under a previous submission, number K251873. The summary notes that this predicate has not been subject to a design-related recall, and that no reference devices were used in this submission.
Sam: So they are building on their own foundation. What exactly is the software intended to do, and who is the intended patient population?
Maya: The indications for use describe it as a software device that analyzes digital breast tomosynthesis, or DBT, mammograms. It is designed to identify the presence or absence of soft tissue lesions and calcifications that may be indicative of cancer. The intended patient population is women from a screening population undergoing screening mammography.
Sam: And how does the radiologist interact with it?
Maya: It is intended to be used as a "concurrent reading aid" for interpreting physicians on screening mammograms with compatible DBT hardware. So the AI runs its analysis, and the interpreting physician sees those outputs while they are conducting their own review of the scan.
Sam: The summary goes into detail about the inputs and outputs. It takes a set of x-ray mammogram DICOM files from a single DBT study. But it is not just looking at the 3D images, is it?
Maya: No, it analyzes both the DBT image stacks and the accompanying 2D images. Those 2D images can include full field digital mammography and, or, synthetic images. It processes all of that and then outputs bounding boxes circumscribing any detected findings.
Sam: And it assigns a score to those findings. They call it a Finding Suspicion Level, which indicates the degree of suspicion that the finding is malignant.
Maya: That is right. And it does not stop there. It uses the results of that finding-level analysis to generate a Case Suspicion Level, which indicates the degree of suspicion for malignancy across the entire case.
Sam: How does all that data actually get back into the doctor's viewing workstation? Does it require a whole new software interface?
Maya: The summary says it is designed to fit in parallel to the standard-of-care workflow. Saige-Dx encapsulates the results into a DICOM Structured Report object. That contains the markings that can be overlaid on the original mammogram images using a standard viewing workstation. It also creates a DICOM Secondary Capture object that contains a summary report of the Saige-Dx results.
Sam: Okay, so if the predicate device already did a lot of this, what are the differences between this new subject device and the predicate? Why did they submit a new 510(k)?
Maya: There are two major additions here. First, the subject device now supports an additional manufacturer, which is Siemens, to expand system compatibility. Second, and arguably more interesting for workflow, it includes the capability to process prior exams.
Sam: Let us talk about that capability to process prior exams. AI alert fatigue is a massive issue in radiology. If an AI flags every benign cyst that has been stable for five years, it just creates noise.
Maya: Exactly. The summary specifically states that processing prior exams was validated as a way to "reduce the number of false positive marks per image". It achieves this by dropping false positive bounding boxes, and by highlighting bounding boxes when a matching prior finding is identified.
Sam: That is fascinating. The AI looks back, recognizes the finding was there on the prior scan, and actively drops the false positive bounding box on the current scan.
Maya: Yes, and they note that this feature also expands the supported image view types. DeepHealth had to prove to the FDA that these differences do not alter the safety or effectiveness of the device for its intended use.
Sam: So how did they prove it? The document details two retrospective, blinded, pivotal standalone performance studies. Let us unpack the first one, which assessed the Siemens-acquired DBT screening mammograms.
Maya: This first study evaluated a dataset of 2,425 DBT screening mammograms. Their primary endpoint was to demonstrate non-inferior performance of the dataset with versus without the Siemens data. The performance was calculated in terms of area under the curve, or AUC.
Sam: They provide a very thorough breakdown of those 2,425 cases in Table 1. And since AI bias is a huge topic for health system leaders right now, we should look at the demographics of that testing data.
Maya: Agreed. In the Siemens pivotal standalone study, the mean patient age was 59.3, with a standard deviation of 10.9. The ages ranged from a minimum of 35.2 to a maximum of 94.5.
Sam: And for patient race in that dataset of 2,425 cases, they report 1,024 were White, 751 were Black or African American, and 116 were Asian. There were also 8 American Indian or Alaska Native, 11 Multiple, 4 Native Hawaiian or Other Pacific Islander, 44 Other, and 467 Unknown.
Maya: For ethnicity, 1,748 were Not Hispanic or Latino, 152 were Hispanic or Latino, and 525 were unknown. They also broke it down by breast density, which is critical for mammography. There were 1,090 cases with density B, 1,009 with density C, 172 with density D, and 154 with density A.
Sam: Even though it was the Siemens pivotal study, the dataset still heavily featured other manufacturers to prove adding Siemens didn't break anything. The manufacturer breakdown was 1,402 Hologic cases, 514 GE cases, and 509 Siemens cases.
Maya: And we can see the complexity of the exams in the modality breakdown. Out of 2,425 cases, 1,514 were a combination of DBT plus 2D synthetic plus full field digital mammography. Another 493 were DBT plus 2D synthetic, and 418 were DBT plus full field digital mammography.
Sam: The conclusion for that first study was that it met the pre-specified performance criteria, supporting safety and effectiveness for use on DBT exams acquired from Siemens systems. What about the second study?
Maya: The second study was the Priors Pivotal Standalone Study. This dataset consisted of 684 patients for whom both a current and a prior DBT screening mammogram were available. The primary endpoint here was to demonstrate non-inferiority in performance, calculated in terms of AUC, sensitivity, and specificity, with versus without priors.
Sam: Table 1 also breaks down those 684 cases. Interestingly, there were 0 Siemens cases in this priors dataset. The manufacturer split was 380 Hologic and 304 GE.
Maya: That is right. The summary notes earlier that the prior exams capability was validated on Hologic and GE systems. For demographics in this priors study of 684 cases, 515 were White, 74 were Black or African American, and 44 were Asian. The mean age was slightly higher at 62.2, with a range from 39.0 to 90.0.
Sam: The breast density for the priors dataset skewed toward denser tissue compared to the first study. There were 399 cases with density C, 169 with density D, 94 with density B, and only 22 with density A.
Maya: Both studies met their pre-specified criteria. The document also mentions that subgroup analyses were performed as secondary assessments. These demonstrated similar standalone performance trends across breast densities, ages, race and ethnicities, manufacturers, lesion types, and sizes.
Sam: Whenever we talk about retrospective AI testing, the biggest question is how they defined what was actually cancer. How did they establish the ground truth for these 2,425 cases and 684 cases?
Maya: The summary states that ground truth for each case was established to determine the reference standard cancer cases. All truthing radiologists were experienced breast imagers, and they were completely blinded to the AI outputs.
Sam: What exactly were those truthers looking at?
Maya: They confirmed the cancer status. And for the cancer cases, they identified and localized all malignant lesions. They documented the pathology classification, the lesion type, and the location based on imaging, radiology reports, pathology reports, and other relevant clinical documentation.
Sam: That is a rigorous truthing process. But we have to point out what this document does not show. Both of these were standalone, retrospective performance studies.
Maya: Yes, that is a critical limitation for any hospital buying this. Standalone means the software's performance was evaluated against the ground truth, not by measuring how human radiologists actually performed while using the software in a live clinical setting.
Sam: The summary claims the software can "help improve reader performance, while also reducing reading time". But the pivotal studies listed here evaluated AUC, sensitivity, and specificity in a standalone environment. The document does not provide data from a prospective clinical trial proving the exact amount of reading time saved in real-world use.
Maya: Which leads directly into the warnings and precautions stated in the document. It explicitly states that Saige-Dx is an adjunct tool and is not intended to replace a physician's own review of a mammogram. It warns that decisions should not be made solely based on analysis by Saige-Dx.
Sam: Let us talk about the clearance letter itself from Dr. Yanna Kang at the FDA. Saige-Dx is classified as a Class II device, with the product code QDQ. It is a software-only device.
Maya: The letter lists the design and development standards they followed. This includes ISO 14971 from 2019 for the application of risk management to medical devices, and IEC 62304 from 2015 for medical device software life cycles.
Sam: They also conducted verification testing, which included software unit testing, software integration testing, system testing, and regression testing. This testing confirmed the software has no unintentional differences from the predicate device.
Maya: The FDA letter also reminds the company about the Quality Management System Regulation. It points out that regardless of whether a change requires premarket review, device manufacturers must review and approve changes to device design and production under specific ISO 13485 clauses.
Sam: Those clauses cover design controls, nonconforming product, corrective action, and preventative action. The letter also emphasizes that clearance does not mean the FDA has determined the device complies with other federal statutes, like medical device reporting for adverse events.
Maya: What do you see as the biggest takeaway here for health-system leaders and AI developers?
Sam: For me, the most memorable takeaway is the explicit validation of AI dropping false positives by evaluating prior exams. Moving from a single-snapshot analysis to longitudinal analysis is a huge step for making AI practical on the ward. Radiologists do not just need to find more lesions; they need less noise from stable, benign findings.
Maya: I agree. Processing prior exams to drop those false positive bounding boxes directly attacks the problem of alert fatigue. And by expanding compatibility to Siemens hardware, they are making the tool viable for larger, mixed-fleet hospital systems.
Sam: But as always, buyers need to remember that this clearance is based on retrospective, standalone data proving substantial equivalence. You still have to monitor how it impacts your own clinical workflow and diagnostic accuracy in real life.
Maya: Let us quickly recap. DeepHealth received 510(k) clearance for Saige-Dx, an AI software that acts as a concurrent reading aid for DBT mammograms. The new clearance validates processing prior exams to reduce false positives, and expands compatibility to Siemens systems, backed by two retrospective studies of 2,425 and 684 cases.
Sam: If you want to look at the exact demographics and performance study details, we have linked the full FDA 510(k) summary document in the show notes.
Maya: And a final reminder: this podcast is for informational purposes only. It is not medical advice, and it is not an endorsement of any company or product. Thanks for listening.