SDAV Insight

Healthcare AI has passed its hardest test.

MASAI followed 105,915 Swedish women for two years to measure what AI-supported mammography missed, rather than what it found. Its results are unusually precise about what was established and what was not — a distinction worth keeping when the lesson is applied outside healthcare.

Die Insights erscheinen auf Englisch.

On 29 January 2026, The Lancet published the final results of the MASAI trial, the first randomised controlled trial of artificial intelligence inside a national breast cancer screening programme. Between 12 April 2021 and 7 December 2022, 105,934 women attending four screening sites in south-west Sweden — Malmö, Lund, Landskrona and Trelleborg — were randomly allocated to AI-supported screening or to standard double reading by two radiologists; 19 were excluded, leaving 105,915 in the analysis. The trial is registered as NCT04838756 and complete.

Most AI results describe what a system found. This one describes what it missed, two years later.

The trial measured the failure that screening programmes actually fear

The primary outcome of MASAI was the interval cancer rate: breast cancers diagnosed between two scheduled screening rounds, or within two years of the last screen, that were not picked up at screening. It is the endpoint that cannot be improved by finding more of anything: recalling more women detects more cancers, but only reading better leaves fewer cancers to surface in the interval.

The design was a non-inferiority trial with a 20% margin: the researchers pre-committed to what would count as “not meaningfully worse” before the data existed, then waited two years for the answer; the registry records the study as completed on 12 August 2025. That sequence — a fixed question, a fixed bar, an independent outcome and a wait — is what much commercial AI evidence lacks.

The system was placed in the reading queue, not in the diagnosis

The system was Transpara version 1.7.0, of ScreenPoint Medical, Nijmegen. It produced an examination-level malignancy risk score on a ten-point scale: examinations scoring 1 to 9 went to a single radiologist, only those scoring 10 to two. For examinations scoring 8 to 10, the radiologist also saw computer-aided detection marks highlighting suspicious regions.

That is the entire intervention. No image was reported by a machine, and no woman was recalled or cleared without a radiologist. Only the routing of work changed.

The effect on capacity was easy to audit: 61,248 screen readings in the AI-supported group against 109,692 in the control group, a 44.2% reduction in screen-reading workload, consistent with the 44.3% in the 2023 interim safety analysis in The Lancet Oncology.

The published numbers separate what was proved from what was observed

Interval cancers occurred at 1.55 per 1,000 participants in the AI-supported group (95% CI 1.23–1.92, 82 cases) and 1.76 in the control group (1.42–2.15, 93 cases). The proportion ratio was 0.88 (95% CI 0.65–1.18; p = 0.41). Read strictly, that establishes non-inferiority; it does not establish superiority, and the trial does not claim it. The differences in the character of those interval cancers — fewer invasive (75 against 89), fewer T2 or larger (38 against 48), fewer non-luminal A, the more aggressive subtypes (43 against 59) — are reported descriptively.

Two results in the same analysis reached statistical significance. Sensitivity was 80.5% (76.4–84.2) in the AI-supported group against 73.8% (68.9–78.3) in the control group, p = 0.031, an effect the authors report as consistent across age and breast density, and present for invasive cancer but not for cancer in situ. Specificity was 98.5% in both groups, p = 0.88 — the system found more without becoming less discriminating.

The 2025 secondary analysis in The Lancet Digital Health fills in the operational picture from the same cohort. Cancer detection was 6.4 per 1,000 screened against 5.0, a ratio of 1.29 (1.09–1.51, p = 0.0021). The false-positive rate was unchanged, at a ratio of 1.01 (0.91–1.11, p = 0.92), and the positive predictive value of recall improved, at a ratio of 1.19 (1.04–1.37, p = 0.012). More detection, the same false alarms, fewer than half the readings.

Certified AI is narrow, and Europe has already scheduled its own clock

The regulatory record makes the same point. The US Congressional Research Service reported on 10 June 2026 that the Food and Drug Administration has authorised roughly 1,450 AI-enabled devices, most of them in radiology, cardiology and neurology and most through the 510(k) pathway, and that it does not appear to have authorised any generative-AI-enabled device.

The AI that has cleared medicine’s evidence bar is not the AI that dominates general discussion. It is narrow, task-specific, version-numbered software with a defined intended use — which is why it can be tested.

In the European Union, AI that is a safety component of — or is itself — a product covered by the harmonisation legislation in Annex I of the AI Act is high-risk; that annex includes Regulation (EU) 2017/745 on medical devices. After the political agreement on the AI Omnibus, the European Commission gives the timetable as 2 August 2028 for high-risk systems embedded in those products and 2 December 2027 for the Annex III use cases. A Swiss supplier selling into the EU can read those as procurement dates rather than compliance dates: conformity questions arrive in tenders well before a deadline applies.

Swiss screening carries the same bottleneck, and has not yet reached this question

Switzerland runs organised breast screening through cantonal and regional programmes, listed canton by canton by Swiss Cancer Screening, which records for the remaining cantons that no organised programme exists. Their design is the MASAI control arm. The national quality standards, built on the European guidelines, require every image to be read twice, independently, by two qualified radiologists, and set minimum annual reading volumes: 2,000 readings for at least one of the two, 1,000 for the second, 3,000 as the desirable standard. Cantonal programmes go further: in Thurgau, each radiologist assesses at least 3,000 mammograms a year.

The volume is growing. Swiss Cancer Screening reports that more than 12 million mammography images passed through the programmes’ shared platform in 2025, around 9% more than in 2024, and coverage is widening: the Aargau foundation confirmed at the end of 2025 a start in 2026, with discussions continuing in Lucerne. The financing frame moved at the same time: the TARDOC outpatient tariff entered into force on 1 January 2026, and the Federal Council was expected to adopt a National Cancer Plan 2026–2032 in the summer of 2026. The association’s 2025 report covers the platform, the tariff negotiations and the new programmes; it does not discuss artificial intelligence.

The United Kingdom answered with another trial rather than deployment. On 4 February 2025 the Department of Health and Social Care announced EDITH — nearly 700,000 women across 30 sites, £11 million from the National Institute for Health and Care Research — to test whether AI can safely replace one of the two readers each screening currently needs.

The transferable part is the method, not the medicine

Very few companies read mammograms. Almost all have a queue where scarce, certified or expensive judgement is spent on volume that does not need it — credit files, engineering approvals, claims, quotations, quality inspections, contract review. MASAI is a template for testing that honestly, at company scale.

  • Locate the triage point, not the task. The gain came from deciding which examinations needed two experts, not from automating the expert. Ask which items in your queue genuinely need senior judgement, and whether anything sorts them today.
  • Name the failure you fear before you start. Interval cancer is the missed case that surfaces later. Its business equivalents — the defect that reaches the customer, the underpriced quotation, the contract clause nobody flagged — are usually measurable, and usually unmeasured.
  • Set a non-inferiority bar in advance, with a margin and a period. Deciding beforehand what counts as “no worse” prevents a pilot from being judged on its best anecdote.
  • Keep the decision with the person and the record with the system. Every recall in the trial was a radiologist’s; version numbers, thresholds and dates were logged, which is what made the result defensible.

Three markers belong in the corporate calendar: the EDITH results, which test the same mechanism at another scale; the Swiss National Cancer Plan 2026–2032, which will show whether requirements written around double reading are revisited; and 2 August 2028, when the European obligations for AI inside regulated products take effect.

MASAI does not show that AI outperforms radiologists. It shows something narrower and more useful: a well-scoped system, tested against the outcome that matters and kept underneath human judgement, can hold quality constant while releasing a large share of expert capacity. That is a management result before it is a medical one, available to any company prepared to define its own interval cancer.

Newsletter

Bleiben Sie über unsere Aktivitäten informiert.

Veranstaltungen, Publikationen und Neuigkeiten aus dem Verband – einige Male im Jahr, in der Sprache Ihrer Wahl.