6 October 2026 · 20 min
Inside the MHRA AI Airlock: Testing the Boundaries of Medical Software
We unpack the MHRA's AI Airlock Sandbox Phase 2 Programme Report to explore how the UK regulator is pressure-testing rules for adaptive algorithms. From generative AI drifting out of scope to algorithms outperforming gold-standard pathologists, learn what happens when modern tech meets traditional regulatory frameworks.
Key points
- Generative AI can drift beyond its intended purpose over time; testing showed out-of-scope performance in approximately 39% of real-world notes without active guardrails.
- Real-world precision for AI systems can be extremely high (e.g., 0.989), which ironically reduces the effectiveness of human oversight as users become complacent.
- When AI diagnostics detect features beyond human capability, such as finding 32% more mitoses than pathologists, regulators face challenges defining the gold standard for performance evaluation.
- Pre-market testing often fails to replicate real-world conditions, making rigorous post-market surveillance essential for lifecycle management.
- Combining multiple low-risk features, like longitudinal memory and personalized biomarker tracking, can inadvertently shift a wellness app into a regulated medical device.
- Predetermined Change Control Plans offer a pathway to safely manage iterative AI updates, provided they are structurally linked to continuous real-world performance monitoring.
Source: AI Airlock Sandbox Phase 2 Programme Report - Medicines and Healthcare products Regulatory Agency, 2026
This spot is available. Reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsor the show
This episode is an AI-generated conversation summarising a public document; the hosts' voices are synthetic. It is for information only and is not medical advice. Always refer to the original source.
Transcript
Sam: Welcome to AI in Medicine Smart Summaries. Just a quick note up front: our voices are AI-generated, and this episode is a summary of a publicly available document. And today, we are looking at a wild finding about how medical software behaves in the wild. When you let a large language model run in a real clinical setting without active guardrails, it can start producing outputs outside its intended design in approximately 39% of clinical notes.
Maya: That is exactly why the Medicines and Healthcare products Regulatory Agency, or MHRA, published the AI Airlock Sandbox Phase 2 Programme Report. It was published in June 2026 and updated in July 2026. This report tackles the fundamental problem with regulating artificial intelligence in healthcare. Existing frameworks were designed for traditional, static medical technologies. They were not built to answer the complex regulatory questions raised by adaptive systems.
Sam: I am so excited for this one. Developers of complex, adaptive AI tools rarely get to see exactly how regulators evaluate algorithms that do not fit traditional frameworks. This report is basically a backstage pass to the future of medical device regulation in the United Kingdom. And it fits into a much larger national strategy. The UK Government has been pushing its AI Opportunities Action Plan, alongside the 10 Year Health Plan called Fit for the Future.
Maya: Yes, that 10 Year Health Plan outlines 3 key strategic shifts for modernizing the NHS. They want to move healthcare from hospital to community, from analogue to digital, and from sickness to prevention. The AI Airlock is the mechanism to ensure the regulatory system can safely support those shifts. The Phase 2 program ran from April 2025 to March 2026, building on the pilot program from April 2024 to March 2025.
Sam: The scale of interest is telling. Supported by the Department of Health and Social Care and the Regulatory Innovation Office, Phase 2 received 51 applications from industry. What is fascinating is who is building this technology. A massive 76% of applications came from micro and small enterprises. Only 8% were from large enterprises, and another 8% from universities. Government entities, namely the NHS, made up the rest.
Maya: That makes perfect sense. Small startups are usually the ones pushing the technological envelope, but they are also the ones who struggle most with regulatory uncertainty. The report also notes that 68% of the products were in the operational prototype or pre-clinical stage, while 14% were already post-market, meaning they were recently launched or within the first 3 years of release.
Sam: The developers actually got to indicate what specific regulatory challenges they were facing. Out of 22 identified challenges, 12% struggled with clinical or performance evaluation, 13% with human-AI interactions, 8% with post-market challenges, and 7% with the scope of intended purpose. So the MHRA structured a portfolio approach for Phase 2.
Maya: Exactly. They selected a cohort of 7 candidates to tackle 3 main regulatory challenges. They also used a tiered approach with different testing environments. They had a simulation environment for focused roundtable workshops, a virtual or research environment for controlled data testing, and a real-world environment for testing in intended purpose settings without impacting actual healthcare outcomes.
Sam: Let us start with the first of the 3 challenge areas: the scope of intended purpose and validation. This is basically asking how you define and enforce what an AI system is supposed to do, especially when its functionality can evolve over time.
Maya: Right. The intended purpose is the foundation of device regulation. It determines your risk classification, your evidence requirements, and your post-market obligations. But generative AI creates a huge problem here. Traditional medical devices are deterministic. A scalpel does not spontaneously decide to act like a stethoscope. But a large language model can shift its behavior based purely on how a user interacts with it.
Sam: And that brings us back to that wild statistic from the hook. They ran a case study with an ambient voice technology product called TORTUS. They wanted to see what happens when the tool shifts from just transcribing speech to actually influencing clinical decisions.
Maya: The TORTUS team tested 2 safety mechanisms in both a controlled virtual environment and in a real-world shadow run. In the real world, without active guardrails, they detected evidence of possible out-of-scope outputs in approximately 39% of clinical notes. When they activated the guardrails, that rate fell to 20%.
Sam: Wait, really? Even with guardrails active, 20% of the notes still had out-of-scope outputs? That feels like a massive headache for a manufacturer trying to prove their device stays within its regulatory boundary.
Maya: It is a huge challenge. It proves that model behavior alone can generate functionality outside the stated purpose. The report states clearly that disclaimers function as passive controls, but they do not prevent a generative system from producing out-of-scope content. You have to design active functional constraints into the software.
Sam: They also found a fascinating mismatch between virtual testing and real-world deployment. They tested a safety feature called the Context Guardrail, which was designed to prevent adversarial prompt injection. In virtual testing, it worked great and reduced scope drift. But in the real world, it had no significant effect.
Maya: Because real doctors in real clinics are not trying to hack the ambient scribe with adversarial prompts. They are just trying to get through their patient list and document the encounter. The sandbox committee pointed out that manufacturers often assume synthetic or adversarial testing is enough, but they risk optimizing for edge cases that rarely happen, while missing the messy variability of routine clinical use.
Sam: Which leads directly to the problem of human oversight. If the AI is mostly doing a great job, humans get lazy. The report explicitly calls human oversight a lifecycle variable with regulatory significance. We cannot just assume a human is checking the work perfectly forever.
Maya: Exactly. TORTUS achieved a real-world precision score of 0.989. When a tool is that consistently accurate, users become familiar with it and may apply less scrutiny to the outputs. This natural shift in behavior can reduce the effectiveness of human-in-the-loop arrangements over time. A safety protocol might look highly effective during the first week of deployment, but lose its value as users blindly trust the system.
Sam: So how do regulators handle that? Do you penalize an AI for being too good and making humans complacent?
Maya: You do not penalize it, but you have to monitor user behavior as part of your post-market surveillance. You might track review time or edit depth as indirect signals of engagement quality. And you also have to benchmark the AI against actual human performance in the same domain. Some experts thought an automated scribe should have fewer than 1 error in 100.
Sam: But clinicians pushed back on that, right? They pointed out that their own transcription accuracy cannot be guaranteed to exceed 99%. So holding the AI to a standard of absolute perfection when the human baseline is flawed doesn't make practical sense. It depends on the severity of the error, not just the frequency.
Maya: Precisely. A combined framework evaluating error frequency alongside severity is much better than a single precision metric. Now, another case study in this scope challenge area was a product called Nu and Aegis by Numan. This is an AI conversational and monitoring system for weight loss management. They tested this in a simulation workshop to explore borderline qualification.
Sam: Right, the boundary between a wellness app and a medical device. This is a nightmare for digital health developers. If you give generic advice about diet, you are a wellness app. But what happens if you add memory to the chatbot and it starts remembering your previous health markers?
Maya: That was the key finding. Features that are low risk individually can completely change the product when combined. When Numan combined clinician-reviewed biomarker data, longitudinal memory, and condition-specific personalization within a clinical pathway, the workshop participants felt the product crossed the line from wellness support to a regulated medical device.
Sam: Because semantic memory allows the AI to accumulate a longitudinal record of symptoms and biomarker trends. It shifts from just generic educational personalization to actual medical monitoring. And the crazy part is, the manufacturer might not even realize they crossed that line just by adding a seemingly basic memory feature to their code.
Maya: Which is why the MHRA recommends periodic reviews that consider feature interactions, not just individual additions. Manufacturers need to look at the whole product, the actual use, and the functional design, not just the claims on the label. The report notes that stating a user appears to have iron deficiency anaemia is obviously diagnostic. But simply noting a low haemoglobin value might also dictate clinical behavior depending on the context.
Sam: So context is everything. The wording alone cannot save you from medical device classification if the tool acts like a diagnostic aid. That takes us to the second of the 3 challenge areas, which focuses on performance evaluation for AI-powered in vitro diagnostic devices. Or IVDs. This is primarily looking at AI systems that analyze digital pathology slides.
Maya: Yes, and the regulatory tension here is fascinating. Traditional IVD regulation assumes that once a test is deployed, its performance remains stable. But AI IVDs are highly sensitive to pre-analytical variability. Differences in staining intensity, scanner type, image resolution, or even the pre-processing pipeline can completely alter what the machine learning model detects.
Sam: They had a case study here with Panakeia, evaluating a tool called PANProfiler Colorectal. One of the big questions they tackled was what metrics actually matter. We always hear about sensitivity and specificity, but the report points out that traditional diagnostic accuracy metrics rely on having a perfect gold standard comparator.
Maya: And in digital pathology, the gold standard is often just human agreement, which is imperfect. So the sandbox explored whether agreement metrics might be more appropriate than strict diagnostic accuracy. They also introduced a concept called test replacement rate. This is the proportion of cases where the AI gives a definitive result and could theoretically replace the need for additional laboratory testing.
Sam: Quick note before we carry on. This spot is open for a sponsor. If your company builds or sells AI for healthcare and wants to reach the clinicians, health-system leaders and industry teams who listen to this show, the link to our sponsorship page is in the show notes.
Maya: And now, back to the document.
Sam: That is a huge metric for hospital operations and budget planning. But the most mind-bending case study in this area was a simulation with a product called Octopath. They brought up a scenario where their AI detected 32% more mitoses than human pathologists.
Maya: Yes, and they used molecular staining to prove that a high proportion of those extra detections aligned with biological markers of mitotic activity. The model demonstrated approximately 91% sensitivity, but the fact that it saw things humans simply missed creates a major regulatory paradox.
Sam: Right, because if the AI consistently detects beyond the range of human review, you cannot use human review as the gold standard to validate the AI anymore. How do you regulate a device that breaks your baseline metric?
Maya: The report suggests that the appropriate comparator and the metrics themselves might need to be reconsidered in these advanced cases. Octopath also highlighted how real-world population shifts cause notable week-to-week variation in performance. Statistical thresholds alone might not capture clinically relevant stability, so regulators have to rethink what stable performance actually means.
Sam: They also discussed something called a pan-cancer or modular approach to validation with Octopath. Since their underlying AI model demonstrated strong and broadly consistent diagnostic performance across breast cancer, neuroendocrine tumours, and melanoma, they asked if validation could be extended to new indications without starting a full revalidation process from scratch.
Maya: Which sounds incredibly efficient, but the MHRA noted that the specific requirements to do that successfully require a lot more exploration. You cannot just assume a model that works on breast tissue will safely generalize to other tissue without a structured evidence framework. And this ties into the need for robust, independent datasets.
Sam: Right, the PANProfiler case study proposed a 7-domain framework for a national reference dataset. It covers intended-use alignment, ground truth derivation, real-world laboratory variability, borderline cases, demographic representativeness, statistical evaluation design, and governance. You basically need a pristine, nationally governed dataset to test these tools fairly.
Maya: Exactly. And they also emphasized that explainability for AI IVDs has to be tailored. Regulators need global explainability to understand how the model was developed and its limitations. But clinicians need local explainability to understand why a particular result was produced in a specific case. Both are required for safe clinical use and scientific validity.
Sam: That brings us to the third regulatory challenge area: Predetermined Change Control Plans, or PCCPs, and post-market surveillance, or PMS. This seems to be the holy grail for software developers who want to update their algorithms without getting bogged down in red tape.
Maya: It really is. A PCCP is essentially an agreement with the regulator up front. You define the specific scope of modifications you anticipate making to your AI, the evidence you will gather, and the thresholds for acceptable performance. If your changes stay within that plan, you do not need to go through a full regulatory reassessment every single time you update a prompt or tweak a workflow.
Sam: But the report points out a major hurdle. At the time of this Phase 2 testing, there were no formal PCCP provisions in UK regulations. During reporting, they did publish a draft statutory instrument for the 2026 regulations, but it only applies to products with Approved Body oversight, not Class I devices.
Maya: Which leaves a lot of developers in a grey area. But the sandbox still tested how these plans should work. The NHS ran a case study on a tool called Safe Summarisation for patient discharge summaries. They showed that a PCCP is useless unless it is structurally linked to continuous post-market surveillance.
Sam: Right, the real-world evidence from your post-market surveillance has to be the trigger that either allows a PCCP change to proceed, pauses it, or blocks it entirely. You cannot just schedule updates on a calendar. They have to be driven by actual performance data monitoring things like hallucinations rates, groundedness, and completeness.
Maya: And Safe Summarisation found that assessing whether a change is significant cannot be determined just by looking at the technical modification. Changing a foundation model to one within the same family might be a minor PCCP event, but if the reasoning architecture differs, it is a major event. Significance is highly contextual and depends on the clinical role.
Sam: Another case study in this area was DeepX AI, which does skin cancer lesion triage. They highlighted a really scary blind spot. Overall acceptable performance can mask material performance gaps in specific patient subgroups, like underrepresented skin tones.
Maya: This is a classic algorithmic bias problem. DeepX AI explored whether post-market surveillance could identify these gaps, and whether a PCCP could allow the manufacturer to rapidly remediate the model to improve equity, without having to pause deployment entirely or go through a massive regulatory resubmission.
Sam: But that raises tough questions. If UK demographic data is scarce, can you use synthetic data or international data to validate the fix? And how do you prove that improving performance for one subgroup does not accidentally degrade it for another? Those are the real-world tensions developers face.
Maya: Those are exactly the open questions the sandbox surfaced. We also saw a case study from Eye2Gene focusing on rare diseases. In rare inherited retinal diseases, patient populations are tiny, and outcomes take a long time to track.
Sam: Which makes setting statistical thresholds nearly impossible. The Airlock consistently found that statistically significant change and clinically important change are not the same thing. If your clinical ground truth is uncertain, defining reliable post-market surveillance triggers is incredibly difficult.
Maya: That is a brilliant summary. Performance expectations have to be grounded in clinical relevance, not just statistical math. Now, to see how well all of this sandbox testing actually worked, the MHRA brought in an independent research agency called Woodnewton to evaluate Phase 2.
Sam: And Woodnewton used a mixed-method approach, right? They pulled in ethnographic observations, qualitative interviews, and quantitative surveys to assess the program's effectiveness.
Maya: Yes. The feedback was really positive overall, especially regarding the trusted space for open regulatory dialogue. But stakeholders did have some demands. Feedback clustered around 5 themes for improvement. They want better timings and process efficiency, less administrative burden, clearer terminology, better deployment of expertise, and most importantly, tangible outputs.
Sam: The timeline piece is interesting. The Phase 2 engagement period was only 6 months long. Airlock technical experts highlighted that testing timeframes between 6-12 months would be much more appropriate to gather meaningful data.
Maya: Exactly. But the biggest takeaway from the evaluation was that innovators do not just want to learn in a sandbox; they want the rules to actually be updated. The long-term credibility of the AI Airlock depends on translating these insights into systemic regulatory guidance and policy change.
Sam: Which leads directly to the recommendations outlined in the report. The MHRA needs to update software qualification and intended purpose guidance, specifically addressing adaptive systems. They need to develop guidance on performance metrics for AI-powered IVDs with practical worked examples. And they need to provide clarity on how to implement Predetermined Change Control Plans.
Maya: There was also a strong recommendation to consider ecosystem-wide approaches to ensure responsible deployment for products that do not even qualify as medical devices. Even if a product sits outside the MHRA's formal regulatory remit, buyers and deployers need transparency and accountability mechanisms to ensure it is used safely in healthcare settings.
Sam: So where does this all go next? The pilot program ran from April 2024 to March 2025. Phase 2 ran from April 2025 to March 2026. What is the roadmap moving forward?
Maya: The Department of Health and Social Care awarded the AI Airlock funding to develop the program further over the next 3 years. Phase 3 will focus on 3 key pillars: Enhancement, Efficiency, and Effectiveness. They want to experiment with more agile testing cycles and prioritize translating the sandbox evidence into permanent regulatory change.
Sam: It is such a necessary initiative. AI moves so much faster than regulation, and sitting in a room alongside the developers to figure out how to measure a system that keeps changing is the only way we keep patients safe without crushing innovation. We cannot regulate 2030 technology with rules designed decades ago.
Maya: Exactly. For me, the most profound insight is that a product's intended purpose is a continuous lifecycle obligation, not just a label you slap on at market entry. Developers must design active constraints, and regulators must demand rigorous real-world monitoring.
Sam: A crisp recap. As a reminder, the source is linked in the show notes, and this is not medical advice.