Medical AI has a proof problem
Kayla Secrest was a newly fledged doctor beginning a residency programme at a Michigan hospital when she had her first frustrating encounter with AI.
Managers had implemented an algorithm designed to scan patients’ records every 15 minutes and send an alert if red flags for sepsis were detected. It was the kind of innovation that the AI industry likes to highlight: taking a routine but vitally important task, and automating it.
But hopes it would lighten her load were soon dashed. “Over the course of a few months, I realised, ‘Gosh, this alert is coming up for every single one of my patients,’ and I would start to glaze over,” she recalls, likening it to the boy in the fable who cried “wolf”.
Eventually, she and her colleagues lost faith in the tool and no longer rushed to patients’ bedsides when the notifications popped up. Only later did she discover it had been rolled out to hundreds of hospitals without extensive testing in the real-life hurly-burly of a ward or emergency department.
Huge claims have been made for AI’s impact on healthcare, and many hard-pressed doctors and hospitals hope it can offset rising demand for their services.
But the experience of Secrest and others suggests a gap between the work of AI developers and medical professionals: there is often little independent oversight of what happens once an AI tool passes from the artificial environment of a lab into the unpredictable world of frontline healthcare.
Experts believe that, in some cases, this absence of scrutiny poses an active risk to patients. It can also make it harder to prove that specific tools improve treatment or justify the investment that they represent.
Jess Morley, an associate research scientist at the Yale Digital Ethics Centre and formerly AI subject lead for the UK’s Department of Health and Social Care, says tech companies tend to focus on “statistical validation” — looking at the accuracy, rather than the clinical efficacy, of a tool.
“But what we really care about in healthcare is that second bit: we want to know that it actually makes an impact on patient outcomes,” she says. “And there is really shockingly poor evidence that AI actually makes an impact on patient outcomes.”
Advocates for AI argue that it is being adopted at a faster rate in healthcare than other sectors, and that it is improving care partly by lightening the administrative load on doctors, such as by automating note-taking. “It’s already been very apparent for a decade or more now that we’ve lost our clinicians to typing,” said Kimberly Powell, head of healthcare at AI chip designer Nvidia, in June.
OpenAI said AdventHealth, a US hospital network, reported an 80 per cent reduction in time spent on certain administrative tasks using ChatGPT for Healthcare, while an independent review of its work with Penda Health in Kenya showed a 16 per cent reduction in diagnostic errors among clinicians using its AI assistant.
Other potential applications include monitoring patients using wearable devices, interrogating large datasets to predict the likelihood of future disease and speeding up the drug development process.
In some areas, such as diagnostics and imaging, evidence of AI’s impact is already irrefutable. Eric Topol, who runs the Scripps Research Translational Institute, cites a 2024 study showing that colonoscopies carried out by gastroenterologists with the assistance of AI detected substantially more polyps than those conducted without the technology.
But rather than adopt tools with a proven record, Topol says healthcare managers have been captivated by “the wow factor” of generative AI. That advance, heralded by the launch of ChatGPT in 2022, enabled the technology to summarise voluminous medical literature to help a doctor make a decision, for example.
Gaps in regulation and a dearth of performance data make it hard to determine whether the rewards of generative AI outweigh the risks. The imperative now should be to “establish that benefit-to-harm ratio . . . in real-world medicine”, he suggests.
Topol sees little appetite among leading AI groups to measure performance against data collected after a model was developed, to check that results can be replicated in active use. “The tech titans are good at making models and testing and validating them [in the lab],” he says. “But by and large . . . they haven’t really demonstrated their enthusiasm or financial backing of the kind of [real-world] trials we’d like to see.”
The transparency deficit
Andrew Wong, a clinician and researcher, investigated the performance of the flawed sepsis tool while at the University of Michigan and, along with colleagues, worked with the developer to improve its performance.
He says companies often make highly specific claims about how AI tools improve workflow and generate financial savings, yet are under no obligation to release information that would make it easier to verify such supposed benefits.
This could include “how the model was trained, what kind of dataset was used, what methodology did they use to develop it, and where did they validate the model to come up with the reported performance statistics”.
Without such transparency, says Wong — now director of clinical AI research at the University of Utah — the work of testing the tools, whose effectiveness can depend on factors such as patient demographics and hospital working practices, will fall to individual institutions. Not all have the resources to do it.
Excessive hype about the transformational impact of AI tools can start surprisingly early in the development process, according to Constanza Andaur Navarro, assistant professor in the Department of Data Science and Biostatistics at Utrecht’s University Medical Centre. She and colleagues analysed papers describing new AI models and found a high propensity for what she terms “spin”. Of the 21 abstracts that recommended a particular tool should be used in daily practice, 20 “lacked any external validation of the developed models”, the research found.
Approaches to the regulation of AI in healthcare differ around the world. But according to Stephen Gilbert, professor of medical device regulatory science at the University of Dresden, in one respect they are similar: surveillance once a tool has entered the market tends to consist of “sanitised reports written by the manufacturer, which have no raw data”, meaning it cannot easily be interrogated by third parties.
A report published this year by Stanford University found the Food and Drug Administration authorised 258 AI medical devices in 2025, most of which did not require new clinical trials. Only 2.4 per cent of the approved devices were supported by the kind of trial data considered mandatory before a drug can reach patients.
Some believe these studies underline the importance of rethinking long-established approaches to ensuring the safety of medical products for an era of generative AI.
Randomised controlled trials are designed to evaluate a molecule or a device such as a blood pressure cuff that will not change once in clinical use. But the algorithms that power large language models are changing all the time, says Morley at Yale University.
Some regulators are taking pains not to block the new technology even as they look at higher guardrails.
“We need to get these AI tools into the hands of clinicians more quickly and more widely than we have done up to now, because the potential to improve healthcare . . . is really, really large,” says Lawrence Tallon, chief executive of the UK’s Medicines and Healthcare products Regulatory Agency.
The MRHA’s National Commission into the Regulation of AI in Healthcare proposed measures this month including conditional authorisation for new tools until they have proved they are safe and effective, and continuous monitoring of those already in use.
The real-world monitoring the commission has recommended would help to build a robust body of evidence to show AI is making a clinical difference on multiple fronts, Tallon argues, although its conclusions have not yet been formally accepted.
He also highlights recent impressive results of a cancer vaccine developed by Merck and Moderna with help from AI as proof of the technology’s potential “at some of those more advanced ends, [to] really radically improve outcomes for patients with cancer [and] rare diseases”.
Similar thinking is evident in the US. Gilbert, the regulatory expert at the University of Dresden, has been collaborating with the FDA on a pilot programme looking at how AI models in areas such as cardiovascular disease, diabetes and mental ill health — chronic conditions that will eventually afflict almost every American — may be able to win regulatory approval on the basis of evidence generated in real-world use.
But some clinicians warn that regulation is not yet keeping pace with the growth in AI capability, a rate that is without precedent in the fields of medicine and healthcare.
John Paul Jeans, an NHS consultant anaesthetist and entrepreneur, cites evidence from Model Evaluation and Threat Research, a non-profit research organisation that tests advanced AI systems and which measured how long a software task an AI system could complete unaided at a 50 per cent success rate. It has found that since 2023, task length has been doubling roughly every four months.
“The test for any regulation being written now is whether it still delivers for the public after three doublings of AI capability,” he says. “If it’s written for the tool sitting in front of us, it could well be out of date before it’s published.”
Another concern is whether hospitals and other facilities have the resources and expertise to evaluate AI tools. Paige Nong, at the University of Minnesota School of Public Health, led research in 2024 that found only a little over half of hospitals surveyed had evaluated tools for potential bias before using them for resource allocation or patient care.
“A small rural hospital that’s already relatively understaffed may just not have the capacity to conduct that kind of evaluation,” she says, fuelling concerns that the onward march of AI risks compounding health inequalities.
Ziad Obermeyer, a physician and researcher at the University of California, Berkeley, came across this problem after studying a “risk stratification” AI tool being used for resource allocation at his own hospital. He discovered it treated past expenditure as a proxy for future demand, with the result that people who needed healthcare but in the past had failed to get it — such as ethnic minorities and non-English speakers — continued to be overlooked.
Overall, hundreds of millions of patients are having decisions about the sums spent on their care influenced by such tools, he says.
“The scale at which these things are already being adopted is just enormous, and I don’t think many people know that.”
Elusive data
There may be technical hurdles to monitoring performance successfully. Amy Abernethy, a former principal deputy commissioner of the FDA, says gathering more evidence from clinical settings will require infrastructure of a kind that has traditionally proved difficult to assemble.
She says the irony is that “we need algorithms and AI . . . to actually help us clean up the data so that we can evaluate algorithms, as well as other healthcare delivery interventions, with a much higher level of granularity”.
Legislation to ensure records can readily be shared across healthcare authorities, and to limit their legal liability for handling this data, would also be needed, she suggests.
Patient data tends to be siloed within hospitals and carefully guarded to protect privacy, Obermeyer argues. “Regulators have a hard time saying to the companies, ‘Well, you claim this algorithm is great, go show me how it does on a dataset that your algorithm has never seen’ . . . because those data are very hard to find,” he adds.
He co-founded a company, Dandelion Health, in an attempt to circumvent this problem. It has worked with hospitals across the US to build a “very diverse and large” dataset that AI developers can use to test their algorithms.
Other researchers are experimenting within academic settings to refine more effective ways of generating data that can demonstrate an AI tool’s impact on patients’ health.
Sammy Chouffani El Fassi, based at Duke University, says the need for developers to back up their assertions about the performance of healthcare tools is increasingly recognised. He is working with researchers across industry and academia to develop applications that will help surgeons make better decisions.
“We’re not trying to find out ‘does this tool work?’” El Fassi says. “We’re asking the question, how will this tool impact the patient’s outcomes?”
He adds: “We’re trying to get out of the theoretical, and into real clinical impact.”