8 October 2026 · 18 min
Beyond Algorithms: How AI Agents Could Reshape Multimodal Cancer Diagnosis
This episode explores a pivotal paper on the shift from narrow AI tools to feedback-driven AI agents in oncology. We discuss how these systems orchestrate multimodal data, the current evidence from real-world and simulated studies, and what hospital leaders need to know about implementation risks and governance.
Key points
- AI agents differ from foundation models by using a feedback-driven loop to maintain state, select tools, and revise plans under constraints.
- Current evidence heavily supports component-level infrastructure, such as data layers and diagnostic modules, but prospective clinical validation of full agents is still lacking.
- Real-world data integration remains complex; one study successfully linked data for over 170,000 patients across 11 cancer types, yet local mapping and governance are always required.
- Clinical translation demands strict operational resilience, including explicit latency budgets, bounded retries, and clinician authority over final decisions.
- The most viable near-term application is transparent and traceable clinical decision support, particularly for multidisciplinary team case preparation, rather than autonomous diagnosis.
Source: AI Agents for Multimodal Oncology Diagnosis: Toward Transparent and Traceable Clinical Decision Support - JMIR cancer, 2026 (CC BY)
This spot is available. Reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsor the show
This episode is an AI-generated conversation summarising a public document; the hosts' voices are synthetic. It is for information only and is not medical advice. Always refer to the original source.
Transcript
Maya: In a recent simulation of realistic oncology cases, an AI system managed to use the correct clinical tools 87.5% of the time, reached accurate clinical conclusions 91% of the time, and accurately cited medical guidelines 75.5% of the time.
Sam: That sounds like a massive leap forward. Today we are digging into a paper titled AI Agents for Multimodal Oncology Diagnosis: Toward Transparent and Traceable Clinical Decision Support, which was published in JMIR cancer.
Maya: And before we get started, we need to let you know that our voices are AI-generated, and this episode is a summary of a publicly available document.
Sam: So let us start with the big picture. Why focus specifically on oncology? What makes cancer diagnosis the testing ground for this kind of advanced AI?
Maya: The sheer scale and complexity of the disease. The document points out that cancer remains a leading cause of morbidity and mortality worldwide.
Sam: And we know that treatment delays across several cancer indications are directly associated with higher mortality.
Maya: Exactly. Timely diagnosis is critical. But modern oncology is intrinsically multimodal. To make a diagnosis, clinicians have to combine imaging, pathology, laboratory results, molecular assays, and longitudinal electronic records. And they have to account for incredible heterogeneity, genomic, histopathological, microenvironmental, and temporal.
Sam: Right, all those different pieces of data usually live in completely separate information systems across different medical specialties. The fragmentation creates a massive cognitive burden for the doctors trying to piece it all together.
Maya: Which is exactly why there is such a push for integrated workflows. Now, task-specific AI has actually done pretty well lately. We have algorithms for bounded tasks like lesion detection, pathology classification, molecular modeling, and risk prediction.
Sam: But the paper argues those are too narrow, right? A model that just detects a lung nodule on a scan is completely disconnected from the iterative reasoning a doctor uses to actually diagnose cancer.
Maya: Exactly. Most deployed models address a single bounded input-output task rather than the full diagnostic process. This paper is arguing for a shift toward AI agents as auditable orchestration layers for multimodal oncology diagnostics.
Sam: Okay, let us define our terms here because the tech industry uses the word agent for everything right now. How is this paper specifically defining an AI agent, and how is it different from just a large language model or a multimodal foundation model?
Maya: The authors are very deliberate about this. They define an AI agent as a feedback-driven system that maintains task state, selects among governed tools, observes results, and revises its plan under explicit safety constraints.
Sam: So a multimodal foundation model might be able to process an image and text at the same time, but that does not automatically make it an agent.
Maya: Correct. A multimodal foundation model just supplies reusable representations or generation across data types. The paper also separates agents from retrieval-augmented generation, or RAG, which just adds external evidence to generation. And it separates agents from fixed workflow orchestration.
Sam: Fixed workflow orchestration is just linking predetermined modules together in a set sequence, right? The paper uses a mathematical notation for a conventional fixed pipeline, meaning the sequence of functions does not change at runtime.
Maya: Spot on. An AI agent, instead, maintains a state and selects an action from an allowed tool set for a clinical goal. It observes each tool's result and updates its state before choosing the next action. The practical distinction lies in conditional planning, tool selection, observation, and revision within a governed loop.
Sam: So if it pulls a pathology report and realizes a key genomic test is missing, it can actually change its sequence, request the missing evidence, or escalate to a clinician according to the evolving state.
Maya: Exactly. The authors operationalize this as a plan-act-observe-update-check loop. The check phase evaluates if the evidence is sufficient, if it is consistent across modalities, and if further action remains within the authorized scope. If there is a gap that cannot be supplemented, it triggers clinician escalation.
Sam: I like that fail-safe. But before we get too deep into the theory, I want to talk about how the authors evaluate the current state of the technology. They use a specific evidence hierarchy throughout the paper, do they not?
Maya: They do, and this is crucial for anyone buying or building these systems. They applied 3 evidence labels. The first is agent-level evidence, which evaluates a complete loop that plans, uses tools, and produces an integrated output.
Sam: That was the simulation you mentioned at the very beginning, right? The external evaluation by Ferber et al.
Maya: Yes. That study looked at 20 realistic multimodal oncology cases. They reported 87.5% correct tool use, 91% correct clinical conclusions, and 75.5% accurate guideline citation. But the authors caution that this is agent-level evidence from simulated cases, a useful benchmark but not prospective clinical validation.
Sam: Got it. So what are the other evidence labels?
Maya: The second is component or infrastructure evidence, which evaluates a model, retrieval method, data pipeline, or governance mechanism that an agent could invoke. The third is prospective propositions, which are plausible clinical workflows that still require system-level and prospective validation.
Sam: That hierarchy makes sense. It prevents hospital leaders from seeing strong performance on one narrow component and mistakenly thinking they are buying a clinically validated autonomous agent.
Maya: Exactly. Now, the authors organize the field around 4 connected stages. The first stage is multimodal data collection, the governed assembly of patient-specific evidence.
Sam: Which sounds easy until you remember that this data is scattered across hospital information systems, laboratory information systems, PACS for imaging, electronic medical records, pathology archives, and genomic platforms.
Maya: Right, and all those systems differ in format, ontology, temporal granularity, and access control. To show it is possible, the paper cites a single-center oncology data supply chain that linked clinical, genomic, and imaging data for more than 170,000 patients across 11 cancer types.
Sam: More than 170,000 patients. That is a massive data engineering effort. Did they extract a lot of data points per patient?
Maya: They extracted more than 800 features per case. They used extract-transform-load, or ETL, natural language processing, and quality control procedures. But again, the authors classify this as infrastructure evidence. It supports the feasibility of a callable data layer, not the validation of an agent.
Sam: So an agent would sit on top of that infrastructure. It would use secure interfaces to identify and retrieve data for a specific diagnostic question. I imagine standards like Health Level Seven Fast Healthcare Interoperability Resources, or HL7 FHIR, and DICOM for imaging help reduce site-specific translation.
Maya: They do, but the paper stresses that local mapping and governance remain necessary. And every input should carry its source, acquisition and update times, missingness status, and a freshness threshold. Contradictory or stale records should trigger clarification rather than silent fusion.
Sam: Silent fusion sounds incredibly dangerous in oncology. You do not want the system quietly blending an outdated biopsy result with a brand new scan. So once the agent has collected the data, what is the next stage?
Maya: Multimodal data preprocessing. This converts heterogeneous inputs into representations downstream tools can use. Imaging varies by scanner. Digital pathology depends on tissue preparation, staining, and slide digitization. Clinical text often contains duplication and institution-specific shorthand.
Sam: And in a fixed workflow, you just run the same preprocessing sequence every time. But an agentic workflow could select modality-specific tools based on the file type, quality, and clinical purpose.
Maya: Yes, it might invoke pathology stain normalization, or use pathology foundation models that provide reusable component representations. It might use attribution tools like gradient-weighted class activation mapping to support region-level inspection, though the paper notes an attribution map is not itself a clinical explanation.
Sam: What happens when the preprocessing fails? Like if an optical character recognition tool simply cannot read a scanned fax from an outside clinic?
Maya: Quick note before we carry on. This spot is open for a sponsor. If your company builds or sells AI for healthcare and wants to reach the clinicians, health-system leaders and industry teams who listen to this show, the link to our sponsorship page is in the show notes.
Sam: And now, back to the document.
Maya: This is a great point. The system should not fabricate a normalized record. It needs a bounded policy. It should retry, switch to a validated alternate parser, return a partial result with an explicit missingness flag, or request human review. This provenance-aware preprocessing creates the conditions for fusion without concealing upstream uncertainty.
Sam: Which brings us to the third stage: multimodal fusion and representation learning. Fusion is not just throwing imaging, pathology, and genomics into a single folder, is it?
Maya: Not at all. Imaging describes macroscopic phenotype, pathology reveals tissue architecture, molecular assays characterize biological drivers, and longitudinal records provide clinical context. Multimodal models can combine these views, but many assume all modalities are harmonized and simultaneously available.
Sam: And in the real world, you might have the pathology report on Tuesday, but the genomic testing is pending until next week.
Maya: Exactly. An agent could select a fusion strategy based on what is actually available, mark molecular evidence as unavailable, and revise the synthesis later. To show what is possible at the component level, the paper discusses a study by Omidi that integrated molecular layers.
Sam: What kind of data were they fusing in that one?
Maya: They combined RNA sequencing and microRNA sequencing profiles. The analysis included 78 adrenocortical carcinoma samples from The Cancer Genome Atlas adrenocortical carcinoma, and 250 normal adrenal samples from the Genotype-Tissue Expression project.
Sam: And did they use deep learning to integrate that?
Maya: Actually, random forest models performed best for predicting microRNA-messenger RNA associations. The predicted interactions were cross-referenced with external databases like TargetScan and miRTarBase. But the authors warn that batch effects, cohort imbalance, cross-cohort heterogeneity, and the absence of independent validation prevent this from being seen as validation of a multimodal agent.
Sam: Right, it is just an in silico systems biology module. Component-level evidence. Are there any studies looking at multimodal fusion on a larger patient population?
Maya: Yes, they cite an explainable multimodal real-world oncology AI model that evaluated 15,726 patients across 38 solid cancer types. That one actually had external validation in 3288 patients with lung cancer. It estimated patient-level marker contributions and prognostic interactions.
Sam: That is a much stronger validation base. But again, the paper notes this supports component-level representation and explanation tools, not autonomous clinical orchestration. Fusion outputs must pass an evidence-adequacy check before moving to the final stage.
Maya: Which brings us to stage four: diagnostic decision support. The authors are extremely clear here. The purpose is not autonomous diagnosis or treatment selection.
Sam: If it is not making the diagnosis autonomously, what exactly is the output of this final stage?
Maya: The goal is a structured, reviewable diagnostic summary that states the supported findings, unresolved conflicts, missing evidence, uncertainty, and provenance in a form clinicians can audit.
Sam: Okay, that sounds highly practical for a clinician. What kind of evidence do we have for decision support modules right now?
Maya: The paper highlights a recent single-center feasibility study that evaluated Gemini 2.5 Pro and ChatGPT 4o on 80 Turkish-language prostate-specific membrane antigen positron emission tomography-computed tomography reports, or PSMA PET-CT for short.
Sam: Just 80 reports? That is a pretty exploratory sample size. What were they asking the models to do with them?
Maya: The prompts incorporated the American Joint Committee on Cancer staging, the Chemohormonal Therapy Versus Androgen Ablation Randomized Trial for Extensive Disease criteria, and few-shot examples. The models assigned T, N, and M categories and disease volume classes.
Sam: And how did they perform?
Maya: The overall task accuracies were 93.8% for Gemini 2.5 Pro and 91.3% for ChatGPT 4o. But errors clustered around equivocal findings and different interpretive thresholds, including overstaging of ambiguous lesions. The authors conclude it supports an LLM-based report-to-staging component that requires expert oversight.
Sam: So it definitely does not validate a complete AI agent. And we cannot ignore the hallucination risk when using language models for decision support.
Maya: Absolutely. The authors cite a multimodel simulation of adversarial hallucination. They found that fabricated details could induce hallucinated clinical elaboration despite simple mitigation attempts. That is why agent evaluation has to include contradiction handling, safe refusal, and workflow burden, rather than just testing how fluently the model speaks.
Sam: They also mentioned multimodal oncology chatbot benchmarking on 79 image-containing clinical cases. That showed you can measure image-text oncology reasoning, but a chatbot benchmark does not establish workflow orchestration or safety.
Maya: Right. So in practice, how would this actually be used in a hospital today? The paper proposes using an agent to prepare a multidisciplinary team, or MDT, case.
Sam: A tumor board. That makes perfect sense. Those meetings require synthesizing massive amounts of unstructured data from different specialties.
Maya: Exactly. An agent could link a suspicious lung lesion to pathology, molecular markers, smoking history, and prior imaging. It could rank competing diagnoses, identify a missing confirmatory test, and provide an evidence-linked summary to the MDT. But the clinician retains authority over the final judgment.
Sam: Which leads us to the challenges and future directions. If a health system wants to implement this, or a software company is building it, what are the hard operational requirements? The paper mentions operational resilience and graceful degradation.
Maya: Yes, real-time MDT use requires an explicit latency budget for retrieval, preprocessing, inference, and rendering. If the system hangs, time-outs should activate circuit breakers rather than indefinite retries.
Sam: That is vital for workflow feasibility. A doctor cannot sit there waiting for several minutes for an API to respond.
Maya: Exactly. Institutions might need to precompute stable features or choose smaller local models when the expected benefit does not justify the delay. And they have to manage the knowledge life cycle. Oncology guidelines and drug knowledge change constantly.
Sam: So you cannot just train the model once and let it run. How do they suggest keeping the agent synchronized with current guidelines?
Maya: Retrieval stores should preserve source dates, jurisdiction, guideline version, and effective period. Updates require staged ingestion, regression tests against reference cases, canary deployment, and rollback criteria. A response should identify the version used and warn when local policy conflicts with an external guideline.
Sam: That reduces, but probably cannot entirely eliminate, legacy hallucinations where the model remembers an old standard of care. This all points to precision oncology governance.
Maya: Precisely. Trustworthy AI principles require fairness, traceability, usability, robustness, and explainability. But for precision oncology, you have to add access control, deidentification, audit logging, incident response, continuous monitoring, model change control, and explicit accountability across users, institutions, and manufacturers.
Sam: Those controls are prerequisites for a clinical service, not evidence that the service improves outcomes. Prospective comparison with current clinical practice is required before broader implementation claims are justified.
Maya: Which brings us to the memorable takeaway. AI agents should be understood as auditable, feedback-driven orchestration systems. The defensible near-term goal is a resilient human-agent workflow that exposes provenance, uncertainty, failures, guideline versions, and escalation decisions. The opportunity is transparent and traceable clinical decision support rather than autonomous cancer diagnosis.
Sam: That is a brilliant wrap-up. If you want to dig deeper into the studies and the proposed agent architecture, you can find a link to the source document in the show notes.
Maya: And please remember, this podcast is for informational purposes only and is not medical advice. Thanks for listening to Smart Summaries.