The 12% was non-inferior, and P was 0.41
MASAI is the best-evidenced AI deployment in medicine and the headline everyone quoted describes a result the trial did not claim.
TL;DR. MASAI is a randomised, controlled, non-inferiority, single-blinded, population-based screening-accuracy trial of 105,915 Swedish women, run at Lund University with three protocol-defined analyses published in Lancet Oncology, Lancet Digital Health and The Lancet. It is the strongest evidence for any AI deployment this corpus has examined. It found a 44.3% reduction in screen-reading workload, a 29% increase in cancer detection without more false positives, and sensitivity of 80.5% against 73.8% at identical specificity. And the headline figure everybody quoted needs reading carefully. Interval cancers were 1.55 per 1,000 against 1.76, a 12% lower rate, 82 cases against 93. Proportion ratio 0.88. P = 0.41. The trial was designed to show non-inferiority and it did. It did not demonstrate that AI reduces interval cancers, and the difference reported as a benefit is not statistically distinguishable from chance.
---
Status: established, and the primary literature is open about exactly this. The Lancet paper reports the proportion ratio and P value in its abstract, describes itself as a non-inferiority trial in its title, and the authors' own framing is careful. The inflation happens downstream, including in a vendor press release. This article does not dispute the trial. It disputes how one of its numbers travelled.
---
What the trial is
MASAI randomised 105,915 women one-to-one: 53,043 to AI-supported screening and 52,872 to standard double reading without AI.
The AI performed two functions. It triaged whether a scan required single or double reading by radiologists. And it acted as detection support, highlighting suspicious findings.
Three protocol-defined analyses have been published.
*The first, in Lancet Oncology in 2023, assessed clinical safety in the first 80,033 participants and concluded AI-supported screening was safe because the cancer detection rate did not decline despite a 44.3% reduction in screen-reading workload.*
*The second, in Lancet Digital Health, reported a 29% increase in cancer detection without an increase in false positives.*
*The third, in The Lancet in January 2026*, is the interval cancer analysis, published after all participants completed two-year follow-up in December 2025.
Three pre-specified analyses, a registered protocol, single blinding, population-based recruitment and a named funder. By the standard of everything else in this corpus, this is what good looks like.
The number, and what it means
Interval cancers are breast cancers diagnosed between screening rounds, or within two years of the last scheduled screening, that were not detected at screening. They carry higher breast-cancer-specific mortality than screen-detected cancers, which is why they are the endpoint that matters rather than detection counts.
The result: 1.55 per 1,000 participants in the AI arm against 1.76 per 1,000 in the control arm. 82 interval cancers against 93.
A 12% lower rate. Proportion ratio 0.88. P = 0.41.
That P value is the article.
A P of 0.41 means the observed difference is well within what chance would produce if the two arms were identical. The trial did not demonstrate that AI-supported screening produces fewer interval cancers.
What it demonstrated is what it was designed to demonstrate: non-inferiority. The AI arm was not worse, and the trial's own title says "non-inferiority" and its abstract reports the P value.
Which means the honest summary is a good result stated precisely: AI-supported screening achieved a 44.3% workload reduction and higher sensitivity without a detectable increase in the cancers that screening misses.
That is genuinely valuable and it is not what "12% fewer interval cancers" conveys.
How the number travelled
A vendor press release described the finding as a "non-inferior 12% reduction in the rate of interval cancers with 27% fewer aggressive cancers."
Note that the word "non-inferior" is present. The release is technically accurate and the phrase "non-inferior 12% reduction" is close to self-contradictory in ordinary reading: a reduction that is non-inferior is a reduction that was not shown to be a reduction.
A university release said "fewer missed cancer cases" and gave the counts, 82 against 93.
Secondary coverage described AI reads leading to "fewer interval breast cancer diagnoses than with standard double reads."
And one summary stated the trial "demonstrated that AI-supported mammography reduces interval breast cancers by 12%."
Each step is slightly stronger than the last and none is a fabrication. The counts are right, the percentage is right, and the qualifier attenuates until it is gone.
This is selective transmission with a statistical qualifier rather than a scope condition, and it is the cleanest instance the corpus has recorded, because the discarded material is a single number printed in the abstract.
What the trial did establish, stated precisely
Worth separating carefully, because the criticism above should not obscure a strong result.
A 44.3% reduction in screen-reading workload, with cancer detection not declining. In a system with a radiologist shortage, halving the reading burden is the finding with the clearest operational value.
A 29% increase in cancer detection without an increase in false positives, which is a genuine sensitivity gain rather than a threshold shift.
Sensitivity of 80.5% against 73.8% at identical specificity of 98.5%, consistent across age and breast density subgroups. The subgroup consistency matters, because sensitivity gains that concentrate in easy cases are the usual failure mode.
And non-inferiority on the endpoint that carries mortality risk, which is what makes deployment defensible rather than merely attractive.
None of that requires the interval cancer difference to be real. The case for AI-supported screening in MASAI rests on workload and sensitivity, both measured with precision, and on safety demonstrated at the endpoint that would have stopped it.
What the trial did not measure
A clinical summary from the Association of Breast Surgery names three gaps before widespread adoption in a programme such as the NHS Breast Screening Programme.
Cost-effectiveness. Not measured.
Overdiagnosis risk. Not measured, and this is the substantive one. A 29% increase in cancer detection is a benefit only to the extent the additional cancers would have progressed. Screening programmes detect indolent disease that would never have caused symptoms, and treating it causes harm. A detection increase is ambiguous on its own, which is why interval cancers and mortality are the endpoints screening trials are judged on.
And long-term outcomes including mortality reduction. Not measured, and not measurable at two years.
That last gap is the structural one. Screening exists to reduce deaths, not to find cancers, and the relationship between the two is exactly what a screening trial has to establish. MASAI measured detection, workload and interval cancers at two years. The mortality question requires a decade.
Which is not a criticism of the trial. It is what the trial can support, and it is less than the coverage implies.
Against the corpus's own baseline
The medicine article found that of 1,524 FDA-cleared AI medical devices, 1.6% cite clinical trial data.
MASAI is what the other 1.6% looks like at its best: three pre-specified analyses, a hundred thousand participants, population-based recruitment, an endpoint chosen because it carries mortality risk, and results published in journals that would have published a null.
And the contrast produces an uncomfortable observation. The device with the strongest evidence in the field still had its most-quoted finding overstated in transmission, by a chain that included its vendor, a university communications office and mainstream coverage.
Which suggests the evidence problem and the communication problem are separate. Territory 10 found the evidence base commissioned and concluded that better evidence would help. Here the evidence is excellent and the public claim is still stronger than the finding.
Better trials fix what is known. They do not fix what is repeated, and this corpus has now seen both failures in the same subject.
The attenuation, step by step
Tracing the exact wording at each hop makes the mechanism visible in a way describing it does not.
| Source | Wording | Qualifier present |
|---|---|---|
| The Lancet paper | Non-inferiority trial, proportion ratio 0.88, P = 0.41 | Full |
| Vendor release | "non-inferior 12% reduction in the rate of interval cancers" | Word retained, sense lost |
| University release | "fewer missed cancer cases", 82 compared to 93 | Counts correct, statistics absent |
| Trade coverage | "AI reads led to fewer interval breast cancer diagnoses" | None |
| Summary article | "demonstrated that AI-supported mammography reduces interval breast cancers by 12%" | None, and "demonstrated" added |
Five rows, and no row contains a false statement about the counts.
Row two is the interesting one. "Non-inferior 12% reduction" retains the technical word and destroys its function, because a reader parses it as a reduction that is also good rather than as a difference that was not established. The qualifier survives as vocabulary and dies as meaning.
Row five adds a verb the trial did not support. "Demonstrated" is a claim about evidential status, and the trial demonstrated non-inferiority.
What this shows is that the failure is not one bad actor. It is a gradient, and each participant made a small, defensible compression of the previous step. The vendor kept the word. The university kept the counts. Coverage kept the direction. The summary kept the conclusion.
And the corrective is correspondingly small. Carrying "P = 0.41" or "not statistically significant" costs four words at row two, and every subsequent row inherits it.
What this territory is testing
Territory 11 opened by asking whether better evidence produces better answers, and two articles in, the result is more specific than expected.
The scribe trials showed that good evidence bounds a claim without collapsing a range, because the effect genuinely varied by setting. Better measurement moved the question rather than closing it.
MASAI shows something different. The evidence is not merely good; it is close to the ceiling of what is achievable in this domain. And the public claim is still stronger than the finding, by a mechanism that operates entirely downstream of the research.
Which separates two problems the corpus has been treating as one.
The evidence problem is that most AI claims rest on commissioned work, self-report or nothing. It is real, it is the subject of several territories, and better trials solve it.
The transmission problem is that a finding loses its conditions as it travels. It is equally real, it operates after publication, and better trials do not touch it. MASAI is the proof: three Lancet-family papers, a registered protocol, careful author framing, and a summary in circulation stating the trial demonstrated something it explicitly did not test.
The corpus has spent five territories arguing for better measurement. This one suggests measurement is necessary and addresses only half the failure, and the other half is a reading problem with no institutional owner at all.
Three things this establishes
A P value is a load-bearing qualifier and it does not travel. 0.41 appears in the abstract of a paper in The Lancet. It appears in almost no description of the result, and without it the sentence changes from "not worse" to "better."
Non-inferiority and superiority are different claims and read identically in a headline. A trial designed to show something is not worse, which succeeds, produces a descriptive difference in the favoured direction roughly half the time by chance. Reporting that difference as a finding converts a safety result into an efficacy claim.
And the strongest evidence in the field did not prevent it. Three Lancet-family publications, a registered protocol and careful author framing were not sufficient, because the failure happened after publication and none of those mechanisms operates there.
What it does not establish
That MASAI is weak. It is the best-evidenced AI deployment this corpus has examined and the workload and sensitivity findings are precise and important.
That AI-supported screening should not be adopted. Non-inferiority on interval cancers with a 44.3% workload reduction is a strong operational case, and the trial supports it.
That the interval cancer difference is absent. A P of 0.41 means not demonstrated, which is different from disproved. The point estimate favours AI and the trial could not resolve it.
And nothing about mortality. No trial of AI screening has run long enough, and the corpus should not be read as claiming otherwise in either direction.
What is unresolved
Whether the interval cancer benefit is real. It would require either a larger trial or longer follow-up, and the point estimate is encouraging.
What the overdiagnosis cost is. A 29% detection increase carries an unquantified overdiagnosis burden, and no analysis in MASAI addresses it.
Whether mortality moves. The endpoint screening exists for, measurable in a decade.
And whether the workload reduction survives deployment. 44.3% in a trial with trained readers and protocol discipline is not automatically 44.3% in routine practice, which is the external validation question every previous territory has raised.
Why non-inferiority trials are structurally prone to this
The problem is not specific to MASAI and it is worth generalising, because non-inferiority is the default design for AI in medicine.
A superiority trial asks whether the new thing is better and reports a result that is either significant or not. The finding and the framing align: a null result reads as a null result.
A non-inferiority trial asks whether the new thing is acceptably close to the old one, and succeeding means demonstrating an absence. That is a harder thing to write a headline about, and it produces a specific temptation.
Because the point estimate almost always favours one arm. With two arms and a real difference of zero, the observed difference lands in the new treatment's favour roughly half the time by chance alone. In those cases a trial that succeeded at non-inferiority has a descriptive number pointing the right way, sitting unused, in a paper whose actual conclusion is "not worse."
And it gets used. Not dishonestly, and not usually by the authors, but by every downstream party that needs a sentence with a direction in it.
Which means the failure mode is predictable in advance. A non-inferiority trial with a favourable point estimate will be reported as a superiority finding, and the frequency of that is roughly the frequency with which chance favours the intervention, which is half.
Why this matters for AI in medicine specifically. Non-inferiority is the natural design when the claim is efficiency rather than efficacy: AI that reads faster, at equal accuracy, is the pitch for most clinical deployments. Workload reduction is the benefit and safety is the requirement, which is exactly a non-inferiority structure.
So the corpus should expect this pattern to recur across ambient documentation, triage, imaging and decision support, and the tell is always the same: a trial whose title contains "non-inferiority" and coverage that reports a percentage improvement.
The four-word correction is available at every hop. "Not statistically significant", or the P value itself, costs nothing and every subsequent summary inherits it.
The counter-argument
Criticising a press release is a small target. The trial is excellent, the authors were careful, the qualifier is in the abstract, and holding a research programme responsible for how a vendor's communications team phrased a summary is not obviously fair. This article spends considerable length on a distortion that a reader of the paper would never encounter.
A P value is not the only evidence. The point estimate favoured AI, the direction was consistent with the sensitivity finding, and the 27% reduction in non-luminal A cancers suggests a mechanism. Treating P = 0.41 as though it establishes no effect is exactly the misuse of significance testing that statisticians have objected to for decades, and this article edges toward it.
Non-inferiority was the right design. The clinically important question was whether halving the reading workload was safe, not whether AI outperformed radiologists. Judging a trial for not demonstrating superiority it was not designed to test inverts the criticism.
And the overdiagnosis objection applies to screening generally. Mammography programmes have carried an unquantified overdiagnosis burden for forty years, and requiring an AI trial to resolve what the underlying programme never resolved sets a standard the comparator does not meet.
The short version
MASAI randomised 105,915 Swedish women, 53,043 to AI-supported screening and 52,872 to standard double reading, with three protocol-defined analyses in Lancet Oncology, Lancet Digital Health and The Lancet. It is the strongest evidence for any AI deployment this corpus has examined.
It found a 44.3% reduction in screen-reading workload with no decline in cancer detection, a 29% increase in detection without more false positives, and sensitivity of 80.5% against 73.8% at identical 98.5% specificity, consistent across age and density subgroups.
And the interval cancer result was 1.55 per 1,000 against 1.76, or 82 cases against 93, a 12% lower rate, proportion ratio 0.88, P = 0.41.
The trial was designed to show non-inferiority and it did. A P of 0.41 means the difference is well within chance. It did not demonstrate that AI reduces interval cancers, and the qualifier attenuated through a vendor release, a university release and secondary coverage until one summary stated the trial "demonstrated that AI-supported mammography reduces interval breast cancers by 12%."
Each step was slightly stronger and none was a fabrication.
What the trial did not measure is also worth stating: cost-effectiveness, overdiagnosis, and mortality. A 29% detection increase is a benefit only to the extent the extra cancers would have progressed, and screening exists to reduce deaths rather than to find cancers.
Which produces the uncomfortable observation. The corpus found 1.6% of 1,524 cleared AI medical devices cite trial data. MASAI is the other 1.6% at its best, and its most-quoted finding was still overstated in transmission. Better trials fix what is known and not what is repeated.
Common questions
What is the MASAI trial? A randomised, controlled, non-inferiority, single-blinded, population-based screening-accuracy trial of 105,915 Swedish women, run from Lund University, with 53,043 randomised to AI-supported mammography screening and 52,872 to standard double reading without AI. The AI both triaged whether a scan needed single or double reading and acted as detection support by highlighting suspicious findings. Three protocol-defined analyses have been published, in Lancet Oncology, Lancet Digital Health and The Lancet.
What did it find? A 44.3% reduction in screen-reading workload with no decline in cancer detection, a 29% increase in cancer detection without an increase in false positives, and sensitivity of 80.5% against 73.8% at identical specificity of 98.5%, consistent across age and breast density subgroups. On interval cancers, the rate was 1.55 per 1,000 against 1.76, or 82 cases against 93, with a proportion ratio of 0.88 and a P value of 0.41.
Why does the P value matter? Because it changes what the interval cancer finding means. A P of 0.41 indicates the observed difference is well within what chance would produce if the two arms were identical, so the trial did not demonstrate that AI-supported screening reduces interval cancers. What it demonstrated is non-inferiority, which is what its title says it was designed to test: the AI arm was not worse. That is a strong safety result and it is a different claim from a 12% reduction.
Was anyone misrepresenting the trial? Not fabricating, and the qualifier attenuated at each step. A vendor press release described a "non-inferior 12% reduction", which is technically accurate and close to self-contradictory in ordinary reading. A university release said "fewer missed cancer cases" with the counts. Secondary coverage said AI led to fewer interval diagnoses. One summary stated the trial demonstrated a 12% reduction. The counts are correct throughout and the statistical qualifier disappears.
So is AI-supported screening a good idea? The trial supports it, on grounds that do not depend on the disputed number. A 44.3% workload reduction in systems facing radiologist shortages, a genuine sensitivity gain consistent across subgroups, and non-inferiority on the endpoint carrying mortality risk together make an operational case. The case rests on workload and sensitivity, both measured precisely, and on safety demonstrated where failure would have mattered.
What did the trial not measure? Cost-effectiveness, overdiagnosis risk, and long-term outcomes including mortality. The overdiagnosis gap is substantive: a 29% increase in cancer detection is a benefit only to the extent those additional cancers would have progressed, and screening programmes detect indolent disease whose treatment causes harm. Mortality is the endpoint screening exists for and requires roughly a decade, which no AI screening trial has yet run.
How does this compare with the rest of medical AI? It is the exception. Of 1,524 FDA-cleared AI medical devices, 1.6% cite clinical trial data. MASAI represents that minority at its best: pre-specified analyses, a hundred thousand participants, population-based recruitment, an endpoint chosen for its mortality relevance, and publication in journals that would have printed a null result.
What is the strongest objection to this article? That treating P = 0.41 as establishing no effect is itself a misuse of significance testing, which statisticians have objected to for decades. The point estimate favoured AI, the direction is consistent with the sensitivity finding, and a reported 27% reduction in non-luminal A cancers suggests a mechanism. A second objection is that non-inferiority was the correct design, since the clinically important question was whether halving reading workload was safe, and criticising a trial for not demonstrating superiority it was not built to test inverts the argument.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study Gommers, Hernström, Josefsson et al., The Lancet 2026;407(10527):505-514 The third protocol-defined analysis: 1.55 against 1.76 interval cancers per 1,000, proportion ratio 0.88, P = 0.41, and the non-inferiority design stated in the title.
- AI and Breast Cancer Screening at a Crossroads: Insights from the MASAI Trial Radiological Society of North America, 2026 The arm sizes of 53,043 and 52,872, the non-inferiority framing, and the assessment of what the trial established for screening workflow.
- AI-supported mammography screening results in fewer aggressive and advanced breast cancers Lund University via EurekAlert, January 2026 The earlier analyses: 44.3% screen-reading workload reduction in Lancet Oncology 2023, and the 29% increase in cancer detection without more false positives in Lancet Digital Health.
- The Lancet publishes final results from the first randomized controlled trial in Breast AI ScreenPoint Medical press release, January 2026 The vendor wording analysed here, including the phrase non-inferior 12% reduction, the 27% figure for non-luminal A cancers, and the sensitivity comparison of 80.5% against 73.8% at 98.5% specificity.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- International Energy Agency (2025), Energy and AI — a worked example in the report that undercuts the framing its projections are used for. :: https://www.iea.org/reports/energy-and-ai Selective Transmission
- Epoch AI (2025), LLM inference prices have fallen rapidly but unequally across tasks — a hundredfold range and a contamination caveat published alongside the rate that circulated. :: https://epoch.ai/data-insights/llm-inference-price-trends Selective Transmission
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — the canonical case, and the source of the 7% figure. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 External Validation
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why a developer-selected evaluation cannot license a general claim. :: https://arxiv.org/abs/2111.15366 External Validation
Related articles
- Adding the doctor to the model changed nothingA randomised trial found physicians did better with an LLM than with conventional resources. Its second comparison, reported in the same abstract, found the model alone did just as well as the model plus the physician.
- Zero-click is 60%, or 22.4%, from one providerTerritory 9 opens on what cheap generation does to information. Every method agrees the traffic is falling. None agrees on how far, and the headline metric differs threefold within one dataset.
- 99.98% is a tracking rate, and the field is animationNVIDIA's motion controller is a real advance that routes around the data problem Territory 7 identified. The number attached to it measures something narrow, and it gets weaker the further it travels from the field it was measured in.
- The EU delayed the part with no standardsThree AI Act obligations take effect today and the widely reported headline says the opposite. What was deferred, what was not, and why the split falls exactly where it does.