Skip to main content

The FDA now authorizes a new AI-powered medical device roughly every 33 hours. But most of those devices reach patients without disclosing who they were actually tested on, or how well they perform across different groups of people.

Artificial intelligence in healthcare has moved from a research curiosity to a regulated product category faster than almost any other medical technology in recent memory. Between 1995 and 2014, the US Food and Drug Administration authorized fewer than two AI-enabled medical devices a year on average. Between 2023 and 2025, that figure reached 264 a year, a roughly 146-fold increase in authorization rate over three decades. Put differently: a technology that regulators once reviewed a handful of times a year is now being cleared for clinical use at a rate of one new device every day and a half.

 

This growth is not evenly spread. Radiology alone accounts for somewhere between three-quarters and four-fifths of all authorized AI/ML devices, largely because medical imaging offers exactly the kind of large, structured, well-labelled dataset that machine learning systems are good at learning from. Cardiology and neurology follow a distant second and third. The vast majority of these tools are decision-support systems: software that flags an image, a scan or a signal for a clinician’s attention, rather than fully autonomous diagnostic systems making a call on their own.

Fast regulatory approval is not the same as thorough validation

The speed of this growth is, on its own, not necessarily a problem. Software-based medical devices can iterate and improve faster than a physical device ever could, and a lighter-touch regulatory pathway is part of what has allowed genuinely useful diagnostic tools to reach clinics quickly. The trouble is what that speed has come at the cost of. A detailed analysis of the devices authorized by the FDA in 2024 found that while the regulatory paperwork itself was almost universally in order, 97.5% of devices cited an identifiable predicate device, the underlying evidence supporting how well these tools actually work, and for whom, was disclosed far less consistently.

 

Only 29.2% of devices authorized in 2024 reported both sensitivity and specificity, the two basic statistics that describe how often a diagnostic tool correctly identifies a condition and how often it correctly rules one out. Fewer still, 15.5%, disclosed any demographic data about the population the device was actually tested on. Just 16.7% used a Predetermined Change Control Plan, the FDA’s newer mechanism for pre-specifying how a device is allowed to change after deployment, which matters enormously for tools that continue learning from new data once they are in clinical use. Barely half included a basic cybersecurity statement.

This is the practical shape of the “algorithmic bias” problem that gets discussed a great deal in the abstract and rather less in the specifics. An AI diagnostic tool trained and validated predominantly on one demographic group, without that fact being disclosed, can perform noticeably worse for patients outside that group: a skin cancer detection algorithm trained mostly on lighter skin tones, a pulse oximetry algorithm calibrated on a narrow age range, a cardiac imaging tool validated in a population that does not reflect the ethnic or socioeconomic diversity of the patients who will eventually use it. None of these are hypothetical failure modes. They are documented ones. A widely cited 2019 study published in Science found that a commercial algorithm used across US hospitals to identify patients needing extra care systematically underestimated the needs of Black patients relative to white patients with the same level of underlying illness, because the algorithm had been trained to predict healthcare costs rather than health need, and unequal historical access to care meant Black patients had generated lower costs for a given level of illness. The bias was not the result of anyone deliberately excluding a group from training data. It emerged from a reasonable-looking modelling choice interacting with an unequal healthcare system, which is precisely the kind of failure that demographic performance reporting is designed to catch before deployment, not after.

What the 2024 FDA authorization data shows is that the infrastructure to systematically catch problems like this before a device reaches patients is, for the large majority of tools, simply not yet in place.

How this compares outside the United States

The FDA is the most closely tracked regulator in this space, partly because its authorization data is unusually accessible, but the underlying dynamic is not a uniquely American one. The European Medicines Agency and the UK’s Medicines and Healthcare products Regulatory Agency (MHRA) are both still in the relatively early stages of building comparable frameworks for AI-enabled devices, and neither publishes authorization data in a form that allows the kind of year-by-year tracking possible with the FDA list. That is not necessarily because European or UK oversight is any laxer in practice. It partly reflects a genuinely difficult regulatory problem: AI/ML systems that continue learning from new data after deployment do not fit neatly into a device-approval model built around a single, fixed product being cleared once and sold unchanged for years. Regulators everywhere are, to varying degrees, adapting a framework designed for stable physical devices to a category of product that is explicitly designed to keep changing.

For the NHS specifically, this puts a considerable amount of weight on procurement and local evaluation processes to do the validation work that a device’s regulatory clearance alone does not guarantee. An AI tool authorized for the US market, tested predominantly on a US population, is not automatically equally reliable when deployed in an NHS trust serving a different demographic mix, and the gap in the underlying evidence is exactly what the FDA’s own 2024 data shows is still, for most devices, undisclosed even in the market where it was authorized.

Why this matters for health economics, not only for clinicians

There is a tendency to treat AI validation as a purely technical or regulatory question, something for biomedical engineers and FDA reviewers to sort out. But the economics of this gap are significant, and they run in both directions.

On one side, AI tools that work as advertised offer a genuinely attractive economic case: faster, cheaper, more consistent interpretation of scans and test results, at a moment when radiology and other diagnostic specialties face severe workforce shortages in most health systems, the NHS included. A validated AI triage tool that catches urgent findings faster, or that lets one radiologist safely cover a larger caseload, has real value in a system straining under demand.

On the other side, a device that performs unevenly across patient groups, and whose performance gaps are not disclosed because they were never systematically measured, creates a different kind of cost: missed or delayed diagnoses concentrated in whichever groups the device was not adequately tested on, which in practice tends to track the same lines of ethnicity, age, sex and socioeconomic status that already drive disparities in health outcomes elsewhere. An AI tool deployed at scale without adequate demographic validation does not remove inequality from a health system, it automates and scales whatever inequality already exists in the data the tool was trained on. That is a substantially harder problem to detect and unwind after the fact than it would have been to require and check for before authorization.

There is also a slower-moving financial risk worth naming: liability. As AI tools become embedded in more clinical pathways, the question of who is accountable when an undisclosed performance gap contributes to a missed diagnosis, the manufacturer, the health system, the clinician who relied on the tool, remains largely untested in most jurisdictions. Weak evidence standards at the point of authorization do not just create a clinical risk. They create a deferred legal and financial one, for exactly the institutions least able to absorb it.

The gap regulators are trying to close

To be fair to the FDA, the direction of policy travel here is toward tighter requirements, not looser ones. Predetermined Change Control Plans, cybersecurity statements and demographic reporting expectations have all been introduced relatively recently, and their low uptake in the 2024 cohort partly reflects the fact that many of these devices were already in the regulatory pipeline before the newer expectations were finalised. It is a reasonable bet that the 2026 and 2027 authorization cohorts will look meaningfully different on these measures than 2024 does.

But the core tension is unlikely to disappear on its own. Authorization volume is growing at a rate, 264 devices a year and rising, that will keep outpacing the depth of scrutiny any regulator can realistically apply to each one, particularly for the roughly 95% of devices cleared through the faster 510(k) pathway rather than the more rigorous De Novo or full premarket approval routes. The practical question for health systems adopting these tools, the NHS very much included as it expands AI-assisted diagnostics across imaging and beyond, is not whether to use AI, but whether procurement and deployment decisions are asking for the demographic and performance evidence that authorization alone does not guarantee. Regulatory clearance is a floor, not a validation that a tool works well for the specific population a given hospital or clinic actually serves.

Sources:

  • “Three Decades of FDA Authorizations of AI/ML-Enabled Medical Devices” (2026)
  • “Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration in 2024: Regulatory Characteristics, Predicate Lineage, and Transparency Reporting,” PMC (2025)
  • FDA AI/ML-Enabled Medical Devices list.

 

Joan Madia