7 October 2026 · 18 min

Does AI in Eldercare Actually Save Money? A Hard Look at the Health Economics

Health systems are eager to adopt AI for aging populations, but does the technology actually save money? We unpack a scoping review of 40 economic evaluations to reveal why most AI tools are deemed cost-saving on paper, but often lack the real-world validation to prove it.

Key points

Source: Assessing the Value for Money of AI-Assisted Technologies for Older Adults: Scoping Review of Economic Evaluations - Journal of medical Internet research, 2026 (CC BY)

This spot is available. Reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsor the show

This episode is an AI-generated conversation summarising a public document; the hosts' voices are synthetic. It is for information only and is not medical advice. Always refer to the original source.

Your company here. This podcast is looking for its first sponsors: reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsorship options and rates →

Transcript

Sam: Out of 40 economic evaluations of AI in older adult care, 21 say the technology is outright cost-saving, and another 17 say it is cost-effective. That is an overwhelming majority telling health systems this technology is absolutely worth the money.

Maya: It is a massive endorsement on paper, which is exactly why we need to dig into how those numbers were calculated. We are looking at a scoping review assessing the value for money of AI-assisted technologies for older adults, published in the Journal of medical Internet research.

Sam: And just a quick note before we dive in: our voices are AI-generated, and this episode is a summary of that public document.

Maya: So, let's set the stage. Population aging is accelerating worldwide. Health care systems are facing a growing demand for chronic disease management, multimorbidity care, and long-term support. Naturally, everyone is looking at digital health interventions, specifically AI, to operationalize and support integrated care models.

Sam: Right, because translating frameworks for older populations into routine practice demands scalable, resource-efficient tools. We need solutions that can reduce provider workload while enabling proactive care. But how do we know if these AI tools actually make financial sense?

Maya: That is where Health Technology Assessment, or HTA, comes in. HTA provides a structured approach to evaluate not just clinical effectiveness, but the economic, organizational, and equity implications. This scoping review aimed to map the existing literature on the economic evaluations of AI technologies specifically for older adults.

Sam: And they found 40 studies published between 2018 and 2026 that met their inclusion criteria. Where were these studies actually taking place?

Maya: They were conducted across 16 contexts, showing a broad geographical spread, though the United States was the most frequently studied context. Most of the studies looked at community or general-population samples, while 14 focused on older adults with specific conditions.

Sam: Okay, so before we get into the actual dollars and cents, how good were these 40 studies? Are we talking about rigorous health economics or back-of-the-napkin math?

Maya: The authors used a tool called the Criteria for Health Economic Quality Evaluation, or CHEQUE, to appraise the studies. The CHEQUE tool has 24 items assessing methodological rigor and 24 items assessing reporting quality. And the scores were generally high. The mean methodological quality score was 85.7/100, and the reporting quality score was 85.4/100.

Sam: A mean score of 85.7/100. That sounds like an A grade. So we can trust the findings?

Maya: Not quite. Here is the massive catch. Despite those high scores, model validation was not reported in 21 of the 40 included studies. Neither internal nor external validation.

Sam: Wait, really? More than half of these economic evaluations didn't report whether their models were actually validated?

Maya: Because the CHEQUE checklist covers many aspects, like clearly describing the intervention setting, the modeling approach, cost outcomes, and data sources. Most studies did a great job adhering to core principles of economic evaluation reporting. But without model validation, confidence in these findings is constrained. It is difficult to assess whether model predictions accurately represent real-world clinical and economic outcomes.

Sam: So if they aren't validating the models, estimates of costs and health outcomes might be subject to considerable uncertainty. That feels like a huge red flag for a payer trying to make a reimbursement decision.

Maya: It is a major limitation. And it gets more complicated when you look at exactly what these AI tools are supposed to be doing. The review mapped all these studies against the World Health Organization’s Integrated Care for Older People framework. It is called the ICOPE pathway.

Sam: I'm assuming ICOPE breaks down the patient journey into different stages?

Maya: Exactly. It is a 4-step pathway. Step 1 is basic screening and assessment. Step 2 is in-depth assessment. Step 3 is personalized care planning. And Step 4 is implementation or monitoring.

Sam: So where are all these AI companies placing their bets? Where is the evidence concentrated?

Maya: It is heavily skewed toward the beginning of the pathway. Most evaluations examined AI for screening and diagnosis, specifically 32/40 studies. So that covers Step 1, which had 24 studies, and Step 2, which had 8 studies. AI here primarily supports diagnostic judgment and early risk identification by analyzing medical images or other data.

Sam: Analyzing images. So we are talking about radiology and eye scans?

Maya: Yes, cancer was the most frequently studied disease area with 17 studies. We also saw ophthalmologic conditions with 6 studies, cardiovascular conditions with 6, and musculoskeletal conditions with 6.

Sam: Okay, so if 32 out of 40 are just diagnosing and screening, what about the actual care? Managing the disease, keeping people out of the hospital?

Maya: Evidence for the later stages of the pathway remains critically lacking. Only 3 evaluations examined Step 3, personalized care planning, where AI supports data-driven risk stratification and decision support. And only 6 studies addressed Step 4, implementation and long-term monitoring.

Sam: That is fascinating. The ICOPE framework covers the whole trajectory of older adult care, but the AI economics literature is almost entirely focused on just finding the disease, not managing it. Why is there such a massive imbalance?

Maya: It comes down to methodology. For technologies used in Step 1 basic and Step 2 in-depth assessment, 96.9% relied on model-based methods. Things like Markov models, decision trees, and discrete-event simulations. Cost-utility analysis was the predominant approach for these stages, used in 26 studies.

Sam: Let me make sure I understand. A Markov model basically predicts what state a patient will be in over time, right? So you use a model because the benefit of a screening tool is early detection, which means you have to extrapolate the long-term health and economic outcomes far beyond whatever observed data you have.

Maya: Precisely. But in contrast, for interventions in Step 3 personalized care planning and Step 4 implementation and long-term monitoring, trial-based evaluations were more common. These technologies are embedded within ongoing care processes. They generate observable short- to medium-term outcomes, like changes in disease control or health care usage, that can be directly measured.

Sam: And it's a lot harder, and more expensive, to run a real trial measuring emergency hospitalizations over two years than it is to build a decision tree model on a computer. That explains the gap.

Maya: It does. And this reliance on modeling directly impacts how they calculate the actual cost of the AI itself. For the assessment-focused technologies, cost estimation methods varied considerably. Only 18.8% of assessment-focused studies used trial-based estimates derived from observed resource use.

Sam: So wait, if they didn't use observed resource use, where did they get the cost of the AI?

Maya: AI-related costs were frequently obtained from assumptions, manufacturer quotes, or expert opinion. Some studies approximated the costs using the price of comparable conventional procedures, or by averaging prices of existing AI tools on the market.

Sam: Manufacturer quotes and assumptions? That is wild. 'Hey expert, what do you think this algorithm should cost to run?' What kind of numbers were they actually plugging in for these costs?

Maya: The reported per-patient or per-procedure AI costs for machine learning-based systems typically ranged from US $0 to $20. Though a few technologies exceeded US $100.

Sam: Hold on. US $0? How does an AI cost zero dollars to implement and run in a hospital?

Maya: Three studies reported AI-related costs of US $0. In one study, AI-related costs were just excluded in the economic analysis. The other two assumed zero cost because the platform was open source, or the platform cost was considered negligible and therefore omitted.

Sam: If you exclude the cost of the technology, of course it's going to look cost-effective! What about the studies that looked at the later stages of care? Were their cost estimates any better?

Maya: Yes, AI cost estimation in care planning and implementation studies was generally more empirically grounded. Four evaluations used trial data, allowing costs to be obtained directly from trials or manufacturer data, typically ranging from US $0 to US $24.6 per patient. These studies also provided clearer descriptions of the cost structure.

Sam: Quick note before we carry on. This spot is open for a sponsor. If your company builds or sells AI for healthcare and wants to reach the clinicians, health-system leaders and industry teams who listen to this show, the link to our sponsorship page is in the show notes.

Maya: And now, back to the document.

Sam: What do you mean by cost structure? What else is there besides buying the software?

Maya: They distinguished between implementation costs, like algorithm development, model training, and validation, and operational costs. Operational costs include maintenance, software updates, technical support, and administrative oversight. Plus hardware, licensing, and personnel time.

Sam: That sounds way more realistic than just assuming an open-source model costs zero dollars to maintain in a clinical setting. Okay, so let's get back to that massive headline stat. 21 studies said it was cost-saving, 17 said it was cost-effective. What exactly is the difference in this context?

Maya: An intervention was classified as cost-saving when it was less costly and more effective than the comparator. It is dominant. And 52.5 % of the included studies fell into this category. They improved health outcomes and reduced total costs.

Sam: And cost-effective?

Maya: Interventions were classified as cost-effective when they were not cost-saving, but they met the study-specific willingness-to-pay threshold. Meaning, it might cost more, but the health gains justify the extra expense based on whatever threshold the health system uses. Even AI interventions with relatively high cost, above US $100, yielded cost-effective or even cost-saving conclusions.

Sam: So if almost everything was a winner, what about the losers? The paper mentioned two evaluations reported their AI interventions as not cost-effective. What went wrong with those two?

Maya: The underlying drivers differed substantially between the two. The earliest included evaluation in this review was published in 2018. Mervin et al evaluated the cost-effectiveness of PARO, which is a socially assistive robotic seal used for dementia care. They compared it to a plush toy without AI.

Sam: A robot seal versus a regular plush toy. I can guess how the economics played out.

Maya: Right. The intervention was associated with relatively high incremental costs and only marginal health gains, resulting in an unfavorable cost-effectiveness result.

Sam: Makes sense. What was the second one?

Maya: The second was from Lin et al, evaluating an AI-based screening for diabetic retinopathy compared with manual grading. The lack of cost-effectiveness here appeared to be driven by contextual factors, particularly low labor costs. The low labor costs reduced the relative advantage of AI-based screening in replacing human labor.

Sam: Now that is a crucial point for health system leaders. If your local human labor is inexpensive, the AI won't save you money. The economic advantage of AI diminishes when there is limited potential for cost substitution. So context is absolutely everything.

Maya: Exactly. The review synthesized the key factors that drive value for money. For screening and diagnostics, cost-effectiveness is primarily driven by algorithm performance, meaning sensitivity and specificity, AI-related costs, and the underlying disease epidemiology. Higher sensitivity generates quality-adjusted life year gains by avoiding downstream complications.

Sam: But that only works if the disease has a high prevalence or high mortality, right? If you deploy an expensive AI screening tool in a low-prevalence setting, you aren't going to get those massive downstream savings.

Maya: That is absolutely correct. And for AI technologies supporting individualized care strategies and disease management, the drivers shift. Cost-effectiveness there is shaped primarily by patient characteristics, adoption efficiency, local facility settings, and long-term disease trajectories.

Sam: Adoption efficiency is a great term. What exactly does the review mean by that?

Maya: It captures the extent to which AI tools are integrated into routine workflows without generating excessive transaction costs. Key parameters include clinician uptake, training requirements, interoperability with electronic health records, alert fatigue, and algorithm update cycles.

Sam: Alert fatigue! Yes. If the AI flags everything and the clinician has to spend twice as long reviewing the alerts, your adoption efficiency plummets, and your operational costs skyrocket. Did these economic models account for that kind of human-AI interaction?

Maya: Rarely. The review specifically calls out that dynamic features of AI systems were rarely considered or modeled. Things like learning curves, human-AI interaction, and performance drift over time. Performance drift occurs when AI models are applied in new populations or settings, potentially reducing effectiveness relative to initial validation studies.

Sam: Which brings us back to the fact that 21 of the 40 studies didn't report model validation. If you don't validate, you don't know if you have performance drift. You're just projecting perfect performance into the future.

Maya: Right. And there is another massive variable that these models often ignore: equity. The review found that equity or distributional considerations of AI were rarely taken into account.

Sam: I would imagine AI could either be a great equalizer or a massive divider. What did the few studies that mentioned it say?

Maya: A few studies mentioned that AI technologies can improve health equity by increasing access to advanced imaging tests in resource-constrained areas. But one study explicitly raised concerns about implementing AI in rural settings, arguing that disparities between and within countries will probably widen.

Sam: Because if the intervention requires a smartphone, high-speed internet, and a certain level of digital literacy, older adults with lower income or cognitive decline might be systematically less able to benefit. You might improve average outcomes, but you widen the inequality gap.

Maya: Exactly. Individual-level characteristics, including education, resistance to technology use, and concerns about data privacy, mediate the usability of these tools. If a model assumes 100% patient engagement, its cost-effectiveness numbers are going to be wildly optimistic.

Sam: Let's talk about how the researchers defined older adults. Did they just use 65 and older?

Maya: Actually, the review defined older adults as individuals aged 50 years and older to capture early functional decline and multimorbidity risks. But they did run a sensitivity analysis restricted to studies involving populations aged 65 years and older.

Sam: And did the results hold up when they looked only at the 65 and older group?

Maya: They did. Among the 14 included studies in that analysis, 85.7% were reported as cost-saving or cost-effective. The base-case conclusions remained robust, showing that AI performance, disease epidemiology, and implementation costs were still the most influential determinants.

Sam: So the fundamental drivers of value don't change, whether you draw the line at 50 or 65. I want to circle back to the perspectives these evaluations took. You mentioned societal perspective earlier. Who exactly were these models built for?

Maya: For the basic and in-depth assessment interventions, 13 studies adopted a societal perspective, 10 used a health care payer perspective, 6 used a national health system perspective, and 3 used a health sector perspective.

Sam: A societal perspective must include things outside the hospital walls, right?

Maya: Yes, the societal perspective has the broadest scope. It encompasses direct medical costs, like hospitalizations and medications, plus nonmedical costs like transportation and informal care. Crucially, it includes indirect costs, such as lost productivity and caregiver time.

Sam: That makes sense, especially in elder care where informal caregiver burden is huge. But if a hospital buyer is looking at a model built on a societal perspective, they need to realize those savings are accruing to society, not necessarily to the hospital's bottom line.

Maya: That is exactly why methodological inconsistency is a recurring criticism in existing reviews. There is actually a specially designed reporting standard called CHEERS-AI for health economic evaluations of AI technologies, but its uptake remains limited. Without a unified way to report workflow integration costs, payers struggle to compare findings across studies.

Sam: Alright, so if I'm a health-system leader or someone building medtech, what is my ultimate takeaway from this 40-study review?

Maya: For companies building AI: You must move beyond modeled assumptions and expert opinion. Prioritize the collection of real-world implementation and maintenance cost data. For health-system leaders: AI value in older care is highly contextual. An algorithm that is cost-saving in one model might not be in your hospital if your labor costs, workflow integration, or patient disease burden look different.

Sam: It is not a magic bullet; it is a labor substitute that needs real-world validation. To read the full scoping review and dig into the methodology yourself, you can find the link in our show notes. And as always, a reminder that this podcast is for informational purposes only and is not medical advice.