8 October 2026 · 23 min

Can Your Smartwatch Spot Sleep Apnea? Unpacking AI and Photoplethysmography

Obstructive sleep apnea affects an enormous global population, but traditional sleep studies are expensive and inaccessible. Maya and Sam unpack a new meta-analysis from the Journal of Medical Internet Research evaluating whether AI-powered photoplethysmography—the optical tech in many wearables—can accurately diagnose the condition. They explore the real-world accuracy of these models, the difference between deep learning and traditional machine learning, and what this means for the future of digital health screening.

Key points

Source: Accuracy of Artificial Intelligence in Diagnosing Obstructive Sleep Apnea Using Photoplethysmography: Systematic Review and Meta-Analysis - Journal of medical Internet research, 2026 (CC BY)

This spot is available. Reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsor the show

This episode is an AI-generated conversation summarising a public document; the hosts' voices are synthetic. It is for information only and is not medical advice. Always refer to the original source.

Your company here. This podcast is looking for its first sponsors: reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsorship options and rates →

Transcript

Sam: What if your smartwatch could accurately screen you for a widespread sleep disorder? Before we dive in today, a quick reminder that our voices are AI-generated, and this is a summary of a public document.

Maya: That is exactly the question at the heart of the research we are unpacking today. The document is a systematic review and meta-analysis evaluating the accuracy of artificial intelligence in diagnosing obstructive sleep apnea using photoplethysmography. It was published in the Journal of Medical Internet Research.

Sam: For anyone building digital health tools, designing wearables, or running a sleep clinic, this is a highly relevant read. We are talking about whether AI algorithms analyzing simple optical sensors can replace, or at least supplement, complex sleep lab tests.

Maya: To understand why this matters, we have to look at the sheer scale of the problem. Obstructive sleep apnea, or OSA, is described in the text as a debilitating sleep-breathing disorder. It is defined by recurring episodes of partial or complete upper airway collapse during sleep.

Sam: And the clinical criteria for that collapse? The document states it results in reduced or absent airflow lasting for at least 10 seconds.

Maya: Yes, and that reduction in airflow is associated with either cortical arousal or a decline in blood oxygen saturation. It is a massive global health issue, highlighting the urgent need for cost-efficient diagnostic and treatment approaches.

Sam: Diagnosing that massive population is currently a major bottleneck. The document notes that overnight polysomnography, or PSG, is touted as the gold standard diagnostic tool. But there are major hurdles to getting that test.

Maya: Exactly. While PSG provides reliable diagnostic accuracy, the authors point out it is constrained by high operating costs, limited accessibility in outpatient settings, patient inconvenience, and significant logistical and labor-intensive demands. It also requires manual intervention and expert oversight.

Sam: So the tech industry and the medical field have been searching for automated alternatives. That is where AI and photoplethysmography come in. Photoplethysmography is a mouthful, so we will call it PPG moving forward.

Maya: The document describes PPG as a noninvasive, cost-effective, and portable optical technique used to provide real-time measurements of blood oxygen saturation in the microvascular tissue bed.

Sam: Basically, it is the technology behind the sensors you see on the back of a smartwatch or a standard clinical pulse oximeter. And preexisting studies have differed in opinion on how accurate AI-based diagnosis of OSA using this PPG data actually is.

Maya: Which is exactly why these researchers decided to conduct a comprehensive synthesis of the evidence. They wanted to evaluate the diagnostic accuracy of these AI models using a Bayesian bivariate random-effects meta-analysis.

Sam: Let's break down how they found their data. They didn't just look at a couple of papers. They searched a wide range of databases from inception all the way to November 3, 2025.

Maya: They conducted a very comprehensive search. They looked through PubMed, Embase, Scopus, Web of Science, and IEEE Xplore. They started with 12,579 records in total.

Sam: And how did they narrow that down? Because 12,579 is a massive number of initial records. They had to have some strict inclusion and exclusion criteria.

Maya: They did. After removing duplicates, they had 8566 nonduplicated records screened at the title and abstract stage. Subsequently, 42 studies were screened for full-text eligibility. They excluded case reports, case series, reviews, meta-analyses, letters, and conference abstracts.

Sam: Did they exclude certain types of patients as well?

Maya: Yes, they excluded pediatric studies and animal studies. They also excluded foreign language studies, studies with incomplete data, and studies focusing on individual apneic events without patient-level classification. Ultimately, they included 13 studies.

Sam: Those 13 studies comprised a total of 9983 participants. Let's talk about the demographics of those participants. Who is actually represented in this data?

Maya: The mean age of the patients in these studies ranged from 42.0 to 69.4 years. As for gender breakdown, the percentage of male participants ranged from 46.4% to 74.0%. And importantly, all studies used data from patients recruited from outpatient sleep clinics.

Sam: And geographically? Where are these studies coming from? Because that always impacts how well the AI models will generalize to the broader global population.

Maya: There was some geographical spread, but it was limited. There were 5 studies from China, 2 from the United States, and 1 each from the United Kingdom, the Netherlands, Israel, Taiwan, South Korea, and Thailand.

Sam: So there is definitely an underrepresentation of certain geographic regions, which the authors later highlight as a limitation. Let's talk about how these studies diagnosed the disease to establish a baseline. What was the reference standard?

Maya: All studies used overnight polysomnography, or PSG, as the reference standard for diagnosing OSA, except for 1 study that used home sleep apnea testing, or HSAT. And the primary criterion for classification was the apnea-hypopnea index, known as the AHI.

Sam: Before we get to the topline numbers, how did they evaluate the risk of bias across these 13 studies? We need to know if the data going into the meta-analysis is actually solid.

Maya: They used the QUADAS-2 tool to evaluate the risk of bias and the applicability of the diagnostic accuracy studies. That tool assesses 4 domains: patient selection, index test, reference standard, and flow and timing.

Sam: And the verdict from that QUADAS-2 assessment?

Maya: Out of the 13 studies, the risk of bias was listed as low for 7 studies and unclear for 6 studies. They also used the GRADE framework to assess the overall quality of pooled evidence. Under that framework, the overall evidence quality was deemed moderate.

Sam: Alright, moderate evidence quality, 13 studies, 9983 participants. Let's hit the topline findings. How accurate is AI trained on PPG at diagnosing sleep apnea?

Maya: The meta-analysis found that AI models trained on PPG achieved a pooled sensitivity of 79.6% and a pooled specificity of 76.5%, compared to conventional diagnosis.

Sam: A sensitivity of 79.6% and specificity of 76.5%. That sounds like reasonable accuracy for a potential screening tool, but the authors mention the credible intervals around those numbers are quite wide, right?

Maya: Exactly. For sensitivity, the 95% credible interval was 55.5% to 93.8%. For specificity, the 95% credible interval was 48.2% to 94.0%. That wide range suggests substantial between-study variability.

Sam: That is a massive spread. If the specificity can swing anywhere from 48.2% to 94.0%, the practical experience in a clinic or for a consumer could be wildly different depending on the specific model or dataset. What is driving that variability?

Maya: A major factor they explored through meta-regression analysis was the severity of the sleep apnea itself. They stratified the data by categorical AHI thresholds. An AHI of 5 or more events per hour represents mild OSA. 15 or more represents moderate, and 30 or more represents severe disease.

Sam: And this is where the findings get really interesting, and frankly, a bit counterintuitive. How did the AI's performance change as the disease got more severe?

Maya: The overall specificity actually increased with greater categorical AHI severity cutoffs. For an AHI of 5 or more, the pooled specificity was 63.6%. For an AHI of 15 or more, it jumped to 81.8%. And for an AHI of 30 or more, it reached 85.1%.

Sam: So the system gets much better at correctly identifying the negative cases—the healthy people—when the threshold for having the disease is set to severe. But what happens to the sensitivity? Does it also get better at finding the sick patients?

Maya: No, it does the opposite. The overall sensitivity decreased with greater categorical AHI severity cutoffs. For an AHI of 5 or more, sensitivity was 87.2%. But for 15 or more, it dropped to 79.7%. And for 30 or more, it fell further to 76.7%.

Sam: Wait, really? The AI produces a higher rate of false negatives in more severe cases? It is actually worse at detecting severe sleep apnea than mild sleep apnea? That feels entirely backward.

Maya: It does seem backward, but the authors offer a fascinating physiological explanation. Severe OSA is inherently heterogeneous. It comes with variable phenotypes, variable event morphology, and substantial interindividual variability in desaturation dynamics.

Sam: Meaning that once the disease is severe, patients' bodies react in highly unpredictable ways?

Maya: Exactly. They also point to variability in autonomic response, vascular tone, and comorbid cardiovascular conditions. All of this complex physiological noise in severe cases may influence PPG detection of OSA, leading to that drop in sensitivity to 76.7%.

Sam: But on the flip side, why does the specificity go up to 85.1% for severe cases? Why is it easier to rule out the healthy people when comparing them against severe patients?

Maya: Because with increasing severity of OSA, the oxygen desaturations become more frequent, deeper, and longer lasting during apneic episodes. These alterations in the microvascular tissue bed are more pronounced and easily detected by PPG technology. This improves the between-group separability, making it easier for the AI to tell who definitely does not have the severe markers.

Sam: So the AI might be better at distinguishing individuals with and without OSA than it is at accurately classifying the severity of the OSA itself once someone has it. That is a crucial takeaway for developers trying to build diagnostic algorithms.

Maya: Precisely. The divergent trends in specificity and sensitivity across severity strata reflect the complex dynamics of harnessing AI to diagnose OSA using PPG. They note that further refinement is needed, particularly to improve the ability to identify mild cases accurately.

Sam: I want to pivot to the types of AI being used. The term 'AI' is thrown around a lot, but this meta-analysis actually breaks down the performance by model type. They looked at 7 studies using deep learning and 6 studies using traditional machine learning.

Maya: Yes, they performed a meta-regression of the AI model type. And they clearly distinguished deep learning models, such as convolutional neural networks, from traditional machine learning approaches.

Sam: What exactly do they classify as traditional machine learning in this context?

Maya: They list decision trees, hyperdimensional models, linear regression, logistic regression, local binary patterns, support vector machines, and combinations of these methods.

Sam: Quick note before we carry on. This spot is open for a sponsor. If your company builds or sells AI for healthcare and wants to reach the clinicians, health-system leaders and industry teams who listen to this show, the link to our sponsorship page is in the show notes.

Maya: And now, back to the document.

Sam: And the performance difference between those traditional methods and the deep learning models was stark, at least on one metric. The deep learning models achieved a higher specificity of 82.9%.

Maya: Right, compared to the traditional machine learning models, which only achieved a pooled specificity of 63.6%. That finding was supported by a positive coefficient of 1.02. Deep learning relies on multilayered neural networks to capture more complex patterns within the data.

Sam: So deep learning is significantly better at ruling out false positives. Did it also beat traditional machine learning on sensitivity?

Maya: Interestingly, no. There was no clear difference in sensitivity between the model types. The credible interval for that comparison ranged from -0.56 to 0.52, meaning it included the null value. But the higher specificity alone highlights the potential utility of deep learning models in OSA diagnosis.

Sam: The authors explain that deep learning significantly improves the detection of subtle physiological changes by enabling the automatic extraction of intricate, high-dimensional features directly from the PPG recordings, right? It does not depend on manually engineered features like traditional machine learning might.

Maya: Exactly. This ability to autonomously identify relevant patterns in large datasets allows deep learning models to capture complex relationships that may not be immediately apparent through conventional methods.

Sam: Let's talk about the hardware itself. Does it matter if the patient is wearing a sleek consumer smartwatch versus a medical-grade pulse oximeter clamped to their finger in a clinic?

Maya: They specifically looked at the type of sleep device, differentiating wearable smart devices from clinical oximeters. And the result was that the type of device used was not clearly associated with sensitivity or specificity. All corresponding 95% credible intervals included the null value.

Sam: That is massive for the consumer tech industry. If wearable smart devices are not clearly performing worse than clinical oximeters in these AI models, the door for democratized screening at home is wide open. But how the models were validated does seem to matter.

Maya: Validation strategy definitely played a role. They found 4 studies used a random split, 3 adopted cross-validation, 5 conducted external validation, and 1 used cross-validation and external validation for separate data cohorts.

Sam: And how did that impact the numbers? Because we know external validation is generally the gold standard for proving an AI model actually works in the real world.

Maya: Studies that used external validation were associated with better sensitivity of 87.2% compared to studies using cross-validation, which had a sensitivity of 80.1%. However, for specificity, cross-validation was actually associated with higher estimates compared to external validation.

Sam: It's a bit of a mixed bag, and the authors note that the limited number of studies using each validation method might preclude definitive conclusions. But they firmly state that external test sets are preferred for evaluating AI algorithms, as they better assess performance when applied to data from different sources.

Maya: That is a critical point for any company bringing a diagnostic AI to market. You must test it on external data. Relying purely on internal test sets can introduce significant blind spots.

Sam: I want to dive into the clinical application section, specifically how they used Bayes theorem. They mapped out what these accuracy numbers actually mean for a patient depending on where they live and the prevalence of sleep apnea in their population.

Maya: This is a very practical part of the document. They calculated posttest probabilities based on different pretest probabilities, or baseline population prevalences. They looked at regions with an approximate prevalence of 15%, 30%, and 60%.

Sam: Let's walk through those. They note that a 15% baseline prevalence is seen in places like Iceland, Indonesia, and the United Arab Emirates. If a patient from one of those regions gets a positive result from this AI on a PPG sensor, what does that actually mean?

Maya: In that 15% prevalence scenario, a positive diagnostic test result would yield an average posttest probability of 37.4%. So if the AI flags them, there is roughly a 37.4% chance they actually have OSA, though the credible interval stretches from 15.9% to 73.4%.

Sam: So in a low-prevalence area, a positive test is a helpful signal, but it is far from a definitive diagnosis. What about regions with a 30% baseline prevalence, like the United States, the United Kingdom, and Samoa?

Maya: With a population prevalence of 30%, a positive test result pushes the average posttest probability up to 59.2%, with a credible interval of 31.5% to 87.0%.

Sam: And for regions with a massive 60% baseline prevalence? The document lists Singapore, Switzerland, and France in this category.

Maya: In those regions, a positive test would yield a posttest probability of 83.6%. And conversely, we have to look at the negative predictive value. What happens if the AI says you do not have sleep apnea?

Sam: Right, how confident can a clinician be in a negative result?

Maya: If you get a negative test result, your average posttest probability of having OSA drops to 4.5% in the low prevalence regions, 10.3% in the medium prevalence regions, and 28.6% in the high prevalence regions.

Sam: So if you are in a high-prevalence area like Singapore, and the smartwatch says you are fine, there is still a 28.6% chance you actually have sleep apnea. That really underscores why this is currently viewed as a screening tool, not a final diagnostic replacement.

Maya: Absolutely. The authors specifically state that in instances where a patient’s pretest probability exceeds the regional prevalence—like if they are actively showing symptoms or clinical signs—the posttest probability is likely to be even higher. Clinical context is still king.

Sam: Before we look at the limitations, how does this meta-analysis compare to prior work in the field? Have other researchers looked at wearables and AI for sleep apnea?

Maya: Yes, they reference a few key studies. Osa-Sanchez et al. did a prior descriptive systematic review and found strong promise for wearables, but they didn't perform a pooled meta-analysis. More recently, Abd-Alrazaq et al. conducted a systematic review and meta-analysis of 38 studies evaluating wearable AI.

Sam: And what did that group find across those 38 studies?

Maya: They found that AI-powered wearable devices achieved a mean accuracy of 86.9%, with a sensitivity of 93.8% and a lower specificity of 75.2% in detecting OSA.

Sam: So this current study provides a more focused evaluation, specifically zooming in on photoplethysmography and adding rigorous Bayesian hierarchical frameworks and subgroup analyses.

Maya: Correct. The authors believe this is the most comprehensive pooled analysis to date on the diagnostic accuracy of AI algorithms using strictly PPG for OSA diagnosis.

Sam: Let's dig into the limitations. No study is perfect, and we have already mentioned the wide credible intervals and the underrepresentation of certain geographic groups. What else do the authors warn about?

Maya: A major limitation is the relatively small number of included studies—just 13—which typically feature small sample sizes. They warn that outliers from these small sample sizes could disproportionately affect the pooled results.

Sam: They also mentioned a potential bias because these studies heavily relied on convenience samples from sleep clinics.

Maya: Exactly. If you are only testing the AI on patients who are already sitting in a sleep clinic waiting room, your data might not perfectly apply to the general population walking around with a smartwatch. It limits the applicability of the existing studies.

Sam: What about patient weight? Body Mass Index is inextricably linked to sleep apnea. Did the models factor that in?

Maya: They noted there was insufficient evidence to examine how BMI may influence the diagnostic accuracy outcome. High BMI is a recognized risk factor for OSA and is associated with greater disease severity. A thorough understanding of the relationship between BMI and PPG-based AI models is still missing.

Sam: That feels like a massive open question for developers. If your AI performs differently depending on a patient's BMI, clinicians need to know that before deploying it. What about publication bias? Did they find evidence that negative studies were just not getting published?

Maya: They tested for that rigorously. While visual inspection of the funnel plots suggested possible asymmetry, the Deeks test did not suggest publication bias.

Sam: They also ran sensitivity analyses simulating unpublished studies, right?

Maya: Yes, they evaluated distinct mechanisms of publication bias with varying probabilities of unpublished studies up to 60%. And the summary receiver operating characteristic curve remained nearly unchanged. They concluded that even if a majority of studies had remained unpublished, the overall conclusions would not have been affected.

Sam: That adds a lot of confidence to the findings. So, where does the field go from here? The document explicitly mentions a device called WatchPAT by ZOLL Itamar as an established home sleep apnea testing modality.

Maya: Yes, WatchPAT is established for patients with a high pretest probability. The authors suggest that given the encouraging performance of these PPG-based AI models, they warrant future direct comparison with established platforms like WatchPAT to see if they offer comparable or incremental diagnostic value.

Sam: They also lay out a roadmap for future research. What should hardware engineers and data scientists be focusing on?

Maya: Future research needs to carefully evaluate the trade-offs of incorporating additional signal channels into AI-driven PPG systems. They need to balance diagnostic accuracy against signal complexity, assess both hardware and operational costs, and promote user compliance and comfort.

Sam: And above all, they stress the need for interpretability and generalizability of the AI models. They strongly advocate for incorporating prospective, real-world data to enhance the credibility of AI-assisted OSA diagnosis.

Maya: Ultimately, the conclusion is that AI models trained on PPG data demonstrate reasonable diagnostic accuracy. Considering the immense logistical and economic challenges of current sleep testing, these AI-driven models could serve as a vital, cost-effective screening tool, pending that further validation.

Sam: It is a massive step toward democratizing sleep health, but developers must respect the clinical complexities, particularly around disease severity and validation testing. That is a wrap for today's deep dive. You can find a link to the full meta-analysis in our show notes.

Maya: And as always, a reminder that we are analyzing research for educational purposes. This podcast is not medical advice. If you have concerns about sleep apnea, please consult a healthcare professional. Thanks for listening.