Blog

Resources

AI
Machine Learning
FDA
Regulatory
Medical device
SaMD
Europe

AI in Mammography Screening: What a New Meta-Analysis of FDA-Cleared and CE-Marked Systems Shows

What 42 studies say about FDA-cleared and CE-marked AI mammography systems

August 11, 2026
cosm logo
Cosm
AI mammography screening meta-analysis
Share this

On August 5, 2026, the European Journal of Radiology published a systematic review and meta-analysis led by University of Sydney researchers that addresses a gap recent reviews of AI in breast screening had left open: it pools the evidence specifically for commercially available, regulator-authorized AI systems in mammography screening. The authors identified 12 FDA-cleared and 15 CE-marked AI products for breast cancer detection, then synthesized 42 studies published between 2019 and 2024 (the search covered January 2018 to July 2024) across digital mammography (DM) and digital breast tomosynthesis (DBT). All 42 studies evaluate systems from just eight vendors, so only a subset of the identified products has published performance evidence at all. Below is what the evidence shows, and what it means if you are developing or deploying one of these systems.

What the review covered

The included systems are the ones regulatory and clinical teams know well: Transpara (ScreenPoint Medical, evaluated in 21 of the 42 studies), Lunit INSIGHT MMG, Mia (Kheiron), ProFound AI (iCAD), Vara, Mammoscreen, Saige-Dx, and cmTriage. Studies were grouped by the AI's role in the screening workflow: standalone reader, reader aid, triage, independent reader within double reading, and additional reviewer of non-recalled cases. That workflow framing, rather than the raw accuracy numbers, is where the practical value of this paper sits.

The headline numbers

In digital mammography, the pooled AUC of standalone AI was 0.89 (95% CI: 0.85 to 0.92), based on 12 studies and 976,256 examinations, while the pooled sensitivity of 76.3% and specificity of 89.6% came from 16 studies and 1,608,644 examinations. That is broadly comparable to radiologists: sensitivity comparable to double reading and higher than single reading, with specificity broadly similar to double reading but lower than single reading. In DBT, pooled standalone AI AUC was 0.90 versus 0.85 for radiologists, although the DBT evidence base is much thinner.

The number that deserves equal attention is the heterogeneity: I-squared exceeded 97% for the principal pooled digital mammography estimates (AUC, sensitivity, and specificity), with wide prediction intervals. Exploratory analyses suggested vendor differences explain more of the variability than screening setting or cancer definition, but substantial residual heterogeneity remained. In plain terms: performance is context-dependent, and a pooled AUC does not tell a screening program, or a regulator, how a specific system will perform on a specific population at a specific threshold.

The autonomy gradient: where the evidence is strong and where it is not

The review's most useful contribution is a GRADE-style certainty assessment across workflow roles. The pattern is consistent: certainty of evidence declines as the AI's role becomes more autonomous.

Triage: the strongest case

AI-based triage, allocating exams to single or double reading based on AI risk scores, has the strongest operational evidence, including the Swedish MASAI randomized controlled trial, which reduced radiologist workload by 44.3% with a numerically higher cancer detection rate (6.1 versus 5.1 per 1,000; p = 0.052). Updated MASAI results cited in the paper reported a 28% increase in cancer detection alongside a 44.2% workload reduction. A Danish implementation cut workload 33.5% with a higher detection rate. Retrospective simulations consistently show 50 to 70% workload reductions with non-inferior sensitivity.

Independent reader: promising, prospectively tested

The prospective ScreenTrustCAD trial in Sweden (55,581 women) showed AI plus one radiologist was non-inferior, in fact slightly better, on cancer detection compared with standard double reading, while cutting initial reading workload by about 50%. In DBT, a retrospective analysis found that substituting the second reader with AI detected approximately 95% of the cancers found by human double reading at half the reading workload.

Reader aid and additional reviewer: real but mixed

AI decision support improved radiologist AUC and sensitivity consistently in laboratory reader studies, but real-world effects varied with workflow: specificity and PPV gains in single-reading settings, detection and PPV gains with modest recall increases in double-reading settings. A prospective Hungarian study of AI as an extra reviewer of non-recalled cases added 0.7 to 1.6 cancers per 1,000 with minimal additional recalls; of the additional cancers detected, 83.3% were invasive and 47% measured 1 cm or less.

Full replacement: not supported

No prospective study has evaluated AI as a full replacement for radiologists in routine practice. The authors grade certainty for fully autonomous reading as very low, and state plainly that current evidence does not support replacing radiologists with standalone AI.

A governance framework built for regulators and buyers

The paper closes with a seven-stage implementation and governance framework: define intended use and governance, technical verification and local validation, workflow and threshold selection, workforce preparation, silent-mode piloting, controlled rollout, and continuous monitoring, with two explicit go/no-go decision gates and a revalidation loop triggered by performance drift, software updates, equipment changes, or population change.

Two details of the framework deserve particular attention. The entire pipeline sits under six standing governance structures that span the AI lifecycle: clinical ownership, regulatory compliance, privacy and cybersecurity, equity and transparency, incident reporting, and vendor accountability. And the final stage requires sites to predefine pause, rollback, and withdrawal criteria: the conditions under which the AI comes out of the screening programme are written down before it goes in, alongside version control and an audit trail for every software update.

The regulatory framing matters. The authors note that approval is a step toward implementation, not a guarantee of consistent long-term performance, because systems are validated under specific conditions and vary across populations, equipment, protocols, and software versions. They point to FDA's Predetermined Change Control Plan (PCCP) guidance, finalized in December 2024, under which modifications specified in an FDA-authorized PCCP can be implemented without a new marketing submission for each change, and to the EU AI Act, which extends oversight of high-risk AI beyond initial CE marking. One timing note the paper does not dwell on: the AI Act's high-risk obligations for AI embedded in regulated products such as medical devices are not yet applicable, with the compliance date expected to move to August 2, 2028 under the provisionally agreed Digital Omnibus (see our breakdown of the new timeline). The authors' recommendation to sites: monitor cancer detection rate, recall, interval cancers, PPV, arbitration rates, and workload impact continuously.

What this means for developers

The takeaways below are our reading of the paper's implications for AI device companies, not findings of the meta-analysis itself.

Build your evidence around a workflow role, not just standalone accuracy. The review is explicit that standalone performance and workflow performance are not interchangeable. Strong standalone AUC does not translate automatically into screening program benefit, which depends on workflow design, radiologist-AI interaction, and implementation strategy. If you are selling a triage configuration, generate triage evidence.

Expect local validation as a condition of adoption. With I-squared above 97% on the principal pooled estimates, buyers have every reason not to take them at face value. The framework in this paper, silent-mode piloting and predefined go/no-go criteria, is likely to become the template for procurement. Developers who arrive with a local validation protocol, reference datasets, and threshold calibration support will move faster than those who arrive with a marketing deck.

Treat threshold selection as a regulated design decision. Most included studies used vendor-default thresholds; others used site-specific calibration. Recall, workload, and arbitration burden all move with the operating point. Documenting how thresholds should be selected, constrained, and monitored, within regulatory limits, is part of the product.

Plan the PCCP and post-market surveillance together. The paper's monitoring loop (drift, software updates, revalidation) is exactly the territory a well-constructed PCCP and post-market surveillance plan should own. Version churn is real: several included systems were evaluated across multiple versions, and results are version-specific.

Expect to be governed as a vendor. Vendor accountability is one of the framework's six standing governance structures, and deploying sites are expected to keep version-controlled audit trails and predefined withdrawal criteria. Contracts, support models, and update processes should anticipate customers who can trace every outcome to a software version and reserve the right to pause or remove the system against criteria set before go-live.

Do not assume DM evidence carries to DBT. The authors support implementation decisions more strongly in DM than DBT. If tomosynthesis is in your intended use, the evidence gap is an opportunity: prospective DBT workflow data is scarce and valuable.

Caveats

The evidence base has real limitations, which the authors document carefully: 39 of 42 studies were retrospective, 95% carried high or unclear risk of bias in flow and timing, follow-up periods were often too short to catch missed cancers, and applicability concerns affected 74% of studies because laboratory workflows deviate from clinical practice. Geographic diversity is limited, with most studies from Europe. None of this undermines the direction of the findings, but it explains the cautious certainty grades.

The bigger picture

This meta-analysis lands in a broader shift: regulators, purchasers, and professional societies are converging on lifecycle oversight of AI devices, where initial clearance or CE marking is the entry point and continuous local evidence is the operating requirement. It pairs well with the validation methods work we covered in Validating AI Segmentation Without a Gold Standard and the lifecycle expectations in IMDRF's Draft N93 Technical Framework.

You can download it from our resources library or access it at the publisher via doi.org/10.1016/j.ejrad.2026.113136 (it's under a CC BY 4.0 license)

How Cosm Can Help

Cosm helps AI/ML medical device companies design evidence strategies that survive contact with buyers and regulators: FDA regulatory strategy and submissions, PCCP construction, clinical validation planning, post-market surveillance design, and quality systems built for iterative software. If you are bringing an imaging AI product to market, or defending one that is already there, reach out at info@cosmhq.com or visit cosmhq.com.

Disclaimer - https://www.cosmhq.com/disclaimer