Clinicians override 49% to 96% of alerts
Five years after the sepsis model was externally validated and found wanting, the sector-level evidence on clinical prediction alerts is process markers and no high-quality mortality signal.
TL;DR. Alert override rates of 49% to 96% are documented across clinical decision support studies. A systematic review screening 3,393 articles and extracting from 44 on predictive models actually implemented in practice found sepsis alerts in emergency departments with sensitivities from 10% to 100%, specificities 78% to 99% and positive predictive values from 5.8% to 54%. It found some evidence for improved process markers such as time to antibiotics, improved length of stay in two studies, and one low-quality study showing improved mortality. No high-quality study showed a mortality difference. And the sepsis model's own record has moved. A 2021 external validation reported AUC 0.63, sensitivity 33% and PPV 12%. A 2024 validation across two county emergency departments reported sensitivity 41.5% and PPV 31.4%, with its authors concluding that a random alert achieves similar specificity.
---
Status: strong systematic review evidence, and the central finding is a null. Sources are peer-reviewed: a systematic review of implemented predictive models, external validations of the Epic Sepsis Model in JAMA Internal Medicine and JAMIA Open, and a JAMA Network Open study of alert volume. One figure, on the share of models lacking external validation, is secondary and labelled where used.
---
The number that governs everything else
Override rates between 49% and 96% are documented across clinical decision support systems.
That range is the subject. A model with perfect sensitivity and specificity, whose alerts are overridden nine times in ten, delivers whatever value survives being ignored.
And overriding is rational. A system generating alerts a clinician has learned are usually wrong is one where dismissal is the correct response, and the learning happened through experience rather than through carelessness.
Which relocates the problem. The sepsis article established that a specific model performed poorly on the population that mattered. The sector-level finding is that model quality is necessary and a long way from sufficient, because the delivery mechanism has a failure mode that operates regardless of the model's accuracy.
Alert fatigue is desensitisation to repeated notification. It is well documented, it predates AI by decades in drug interaction warnings, and a predictive model is a new source of alerts entering a channel that was already saturated.
What the implementation literature actually found
A systematic review screened 3,393 articles and extracted data from 44 describing predictive models integrated into electronic health records and implemented in clinical practice.
The most common domains were thrombotic disorders and anticoagulation at 25%, and sepsis at 16%, with the majority conducted in inpatient academic settings.
For sepsis alerts in emergency departments it reports sensitivities from 10% to 100%, specificities from 78% to 99%, and positive predictive values from 5.8% to 54%. Negative predictive value was consistently high at 99% to 100%.
A range from 10% to 100% sensitivity is not a performance estimate. It is a statement that the category contains systems doing entirely different things, and any figure quoted from within it describes one implementation.
On outcomes, the review found some evidence for improved process-of-care markers including time to antibiotics. Length of stay improved in two studies. One low-quality study showed improved mortality.
And no high-quality study showed a difference in mortality.
The named implementation challenges were alert fatigue, lack of training, and increased work burden on the care team. Three organisational factors, which is the finding Territory 10 reached across nine subjects arriving in a clinical setting.
The sepsis model's record, updated
Worth tracing, because the corpus covered the original finding and the record has moved in both directions.
A large external validation published in 2021 reported an AUC of 0.63, with sensitivity of 33% and positive predictive value of 12%, roughly seven false alarms for each true positive.
A 2024 retrospective external validation across two county emergency departments in Houston, covering all adult patients through 2023, reported sensitivity of 41.5%, PPV of 31.4% and NPV of 97.7%.
PPV rising from 12% to 31.4% is a real improvement, and it may reflect the setting, the threshold, the population or the version rather than the model getting better.
The authors' conclusion is nonetheless blunt: the alerting fails to achieve meaningful sensitivity, and a random alert achieves similar specificity and negative predictive value.
Which is the comparator point. A test with 97.7% negative predictive value sounds strong until compared against the base rate, and in a population where most patients do not have sepsis, predicting "no sepsis" for everyone achieves a similar figure.
The finding that should have changed deployment practice
*A study published in JAMA Network Open examined 24 hospitals in the early pandemic.*
Total sepsis alerts per day increased 43% in the three weeks after each hospital's first COVID-19 case, while total hospital census decreased by 35%.
Sepsis alerts rose from 9% to 21% of all alerts.
The model did not change. The population did, and nothing in the deployment caught it.
The University of Michigan paused Epic's sepsis alerts entirely in April 2020 in response to complaints about over-alerting.
That is distribution shift with a measured operational consequence, and it is the clearest case in this corpus of why post-deployment monitoring is not optional. A model validated on one patient mix, deployed unchanged, met a different mix and produced 43% more alerts to 35% fewer patients.
The review authors' own conclusion was that AI algorithms need careful monitoring after deployment, particularly during dramatic shifts in hospital resources and patient acuity.
The regulatory position, and why it matters here
The FDA article established that the agency's device framework governs post-clearance modification and that clearance itself runs largely on substantial equivalence.
Many EHR-embedded predictive models sit outside that framework entirely.
Algorithms embedded in an electronic health record do not always require FDA clearance, and are frequently positioned as clinical decision support, which is a category with a lighter regulatory position than a diagnostic device.
The consequence is that a model deployed across hundreds of hospitals may have had no external validation before deployment, which is what the 2021 study was: an independent evaluation performed after the fact by researchers who chose to run it.
One secondary source puts the share of clinical prediction models lacking external validation at 94%, which should be read as indicative rather than precise.
And the proprietary position compounds it. The original external validation was conducted without access to the model's internals, which is the aggregate evidence gap with an additional barrier: not merely that nobody owns the sum, but that the components are not inspectable.
The chain from model to outcome
Setting out the stages shows where the evidence stops, and it stops early.
| Stage | Evidence | Status |
|---|---|---|
| Model discriminates | AUC 0.63 to 0.90 depending on study | Measured, highly variable |
| Alert reaches clinician | Delivered into a saturated channel | Not the bottleneck |
| Clinician reads it | Override rates 49% to 96% | Measured, and the bottleneck |
| Action changes | Time to antibiotics improved | Some evidence |
| Patient outcome changes | No high-quality mortality difference | Null |
Five stages, and the evidence thins at each one.
The first stage is where almost all research effort goes, because discrimination is cheap to measure retrospectively and does not require deploying anything.
The third stage is where the value is lost, and it is measured by a different literature that model developers do not generally read.
And the fifth stage is where the question was.
Which explains a pattern this corpus keeps finding without naming the mechanism. Effort concentrates at the stage that is easiest to measure, not at the stage that determines the outcome, and the two are usually far apart.
The practical version for a health system is narrow. Improving a model's AUC moves stage one. Reducing the number of alerts a clinician receives moves stage three, and stage three is where 49% to 96% of the value is currently going.
Both are engineering work. Only one is exciting, and the sector-level evidence suggests the unexciting one has the larger available gain.
What the corpus has now found twice about monitoring
The pandemic alert finding is the second clean case in this corpus of a model degrading because the world moved, with a measured consequence.
Here: 43% more alerts to 35% fewer patients, from an unchanged model meeting a changed population, detected because researchers happened to look and because clinicians complained loudly enough that one university paused the system.
And in the enterprise territory, deployments that passed a pilot failed weeks later when the manual review nobody had documented as part of the system quietly stopped.
The shared structure is that neither failure was detectable from inside the system. The model reported nothing unusual. The alert rate was only anomalous relative to census, which is a comparison nothing in the deployment was computing.
Which suggests the monitoring that matters is relational rather than absolute. Not "is the model performing as validated", which it was, but "has the ratio between what it emits and what the environment contains changed."
That is a cheap thing to compute and an uncommon thing to compute. Alerts per patient-day against historical baseline is a single query, and it would have surfaced the pandemic finding in the first week rather than in a retrospective study across 24 hospitals.
Three things this establishes
A delivery channel can defeat any model. Override rates of 49% to 96% mean the marginal value of accuracy improvement is bounded by whether anyone reads the output. A better model in a saturated channel is a better model nobody sees.
Five years of sector evidence produced process markers and no mortality signal. Improved time to antibiotics, length of stay in two studies, one low-quality mortality result and no high-quality one. That is the honest state of the evidence for the most-deployed category of clinical AI.
And the population moved without anyone noticing. 43% more alerts to 35% fewer patients, from an unchanged model, is the case that establishes post-deployment monitoring as a requirement rather than a recommendation.
What it does not establish
That clinical decision support does not work. Process improvements are real, high negative predictive value has clinical uses, and the review found genuine benefits alongside the null.
That the sepsis model is unchanged. PPV rising from 12% to 31.4% between validations is substantial, whatever caused it.
That override rates are all alert fatigue. Some overrides are correct clinical judgement on a genuinely inapplicable alert, and the studies do not consistently separate them.
And nothing about any specific hospital's implementation. Performance varies by setting to a degree the 10% to 100% range makes obvious.
What is unresolved
Whether any high-quality trial shows a mortality benefit. The review's finding is a null on the outcome that matters, and the studies that would settle it have not been run.
What proportion of overrides are appropriate. Distinguishing fatigue from judgement requires per-alert adjudication that almost nothing does.
Whether monitoring is now standard. The 2020 alert-volume finding was published, widely covered, and no subsequent audit establishes whether health systems monitor alert rates against patient mix.
And what the 41.5% sensitivity means for the version deployed today. External validations lag deployments, and the most recent published figure describes a system as it stood in 2023.
What a health system could measure this quarter
Three things, all computable from data already held, none requiring a vendor or a study.
Alerts per patient-day, against a rolling baseline. This is the query that would have surfaced the pandemic finding in week one instead of in a retrospective study across 24 hospitals. It is one line of SQL and a threshold.
Override rate by alert type, by unit, by shift. The 49% to 96% range is a literature figure. Every hospital has its own number and most do not compute it, which means they cannot tell whether a given alert is a control or a formality.
And time-to-action for alerts that were acted on. The process markers the review found improved are the ones with a mechanistic link to outcome. Measuring them locally converts a literature claim into a local one, and it is the only stage of the chain where a health system can see its own effect.
None of these evaluates a model. They evaluate a deployment, which is the thing the health system controls and the vendor does not.
And the asymmetry is worth stating. A hospital cannot improve a proprietary model's discrimination. It can silence an alert type, change a threshold, route by unit, or stop entirely, and one university did exactly that in April 2020 on the basis of clinician complaints rather than a metric.
Complaints worked and took months. A rolling ratio would have worked and taken days.
The counter-argument
High negative predictive value is more useful than this article allows. A test that reliably identifies who does not have sepsis has real triage value, and dismissing 97.7% NPV by comparing it to the base rate applies a standard that would invalidate most screening tools.
Override is not necessarily failure. A clinician who dismisses an alert because they have already acted on the underlying concern has used the system correctly, and reading override rates as evidence of ineffectiveness assumes the alert was the only path to the action.
Process markers are not a consolation prize. Time to antibiotics in sepsis has a documented relationship with mortality, so improving it is a mechanistically grounded benefit, and demanding a direct mortality trial for every intervention with a validated surrogate is a standard clinical medicine does not generally apply.
And the pandemic finding is an extreme case being used as a general one. A once-in-a-century shift in patient acuity is precisely when any model would degrade, and citing it as evidence for routine monitoring requirements generalises from the least representative period available.
The short version
Override rates of 49% to 96% are documented across clinical decision support, which bounds the value of any accuracy improvement by whether the output is read.
A systematic review screening 3,393 articles and extracting from 44 implemented models found sepsis alerts in emergency departments at sensitivities of 10% to 100%, specificities 78% to 99% and PPVs of 5.8% to 54%. A range that wide is not a performance estimate.
On outcomes it found improved process markers including time to antibiotics, improved length of stay in two studies, one low-quality study showing improved mortality, and no high-quality study showing a mortality difference. Named implementation challenges were alert fatigue, lack of training and increased work burden.
The sepsis model's record has moved. A 2021 external validation gave AUC 0.63, sensitivity 33%, PPV 12%. A 2024 validation across two county emergency departments gave sensitivity 41.5%, PPV 31.4%, NPV 97.7%, with the authors concluding that a random alert achieves similar specificity and negative predictive value.
And the clearest deployment finding is about the population rather than the model. Across 24 hospitals in the early pandemic, sepsis alerts per day rose 43% while total census fell 35%, and alerts rose from 9% to 21% of all notifications. The model did not change. One university paused the alerts entirely.
Much of this sits outside the FDA framework, because EHR-embedded algorithms positioned as clinical decision support do not always require clearance, which is why the external validations that exist were performed after deployment by researchers who chose to run them.
Common questions
How often are clinical alerts overridden? Between 49% and 96% across documented studies of clinical decision support systems, a range that includes drug interaction warnings and other interruptive alerts as well as predictive models. That figure bounds the value of any model accuracy improvement, since a system whose alerts are dismissed nine times in ten delivers whatever value survives being ignored.
What does the implementation evidence show? A systematic review screening 3,393 articles and extracting from 44 describing predictive models actually implemented in clinical practice found the most common domains were thrombotic disorders and anticoagulation at 25% and sepsis at 16%. For emergency department sepsis alerts it reports sensitivities from 10% to 100%, specificities from 78% to 99% and positive predictive values from 5.8% to 54%, with negative predictive value consistently 99% to 100%.
Is there a mortality benefit? Not on the current evidence. The review found some evidence for improved process-of-care markers including time to antibiotics, improved length of stay in two studies, and one low-quality study showing improved mortality. It found no high-quality study showing a mortality difference. That is the honest state of the evidence for the most widely deployed category of clinical AI.
Has the Epic sepsis model improved? The published record has moved. A large external validation in 2021 reported an AUC of 0.63 with sensitivity of 33% and positive predictive value of 12%, roughly seven false alarms per true positive. A 2024 retrospective validation across two county emergency departments covering all adult patients through 2023 reported sensitivity of 41.5%, PPV of 31.4% and NPV of 97.7%. The PPV improvement is substantial and may reflect setting, threshold, population or version. The authors nonetheless concluded that the alerting fails to achieve meaningful sensitivity and that a random alert achieves similar specificity and negative predictive value.
What happened during the pandemic? A study of 24 hospitals published in JAMA Network Open found that total sepsis alerts per day increased 43% in the three weeks after each hospital's first COVID-19 case, while total hospital census decreased by 35%, and sepsis alerts rose from 9% to 21% of all alerts. The model did not change; the patient population did. The University of Michigan paused Epic's sepsis alerts entirely in April 2020 in response to over-alerting complaints. The researchers concluded that algorithms need careful monitoring after deployment, particularly during dramatic shifts in patient acuity.
Why do these models often lack external validation before deployment? Because many EHR-embedded predictive models sit outside the FDA device framework. Algorithms embedded in an electronic health record do not always require clearance and are frequently positioned as clinical decision support, a category with a lighter regulatory position than a diagnostic device. One secondary source puts the share of clinical prediction models lacking external validation at 94%, which should be read as indicative. The independent validations that exist were performed after deployment, by researchers who chose to run them, in one case without access to the model's internals.
Is overriding an alert a failure? Not necessarily, and the studies do not consistently separate the cases. A clinician who dismisses an alert because they have already acted on the underlying concern has used the system correctly. What the override range does establish is that the delivery channel has a failure mode operating independently of model accuracy, and that a better model entering a saturated channel is a better model nobody reads.
What is the strongest objection to this article? That process markers are not a consolation prize. Time to antibiotics in sepsis has a documented relationship with mortality, so improving it is a mechanistically grounded benefit, and requiring a direct mortality trial for every intervention with a validated surrogate is a standard clinical medicine does not generally apply. A second objection is that the pandemic alert-volume finding describes the least representative period available and is being generalised into a routine monitoring requirement.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- External validation of the Epic sepsis predictive model in 2 county emergency departments JAMIA Open, 2024 The 2023 retrospective validation across two Houston county emergency departments: sensitivity 41.5%, PPV 31.4%, NPV 97.7%, and the conclusion that the alerting fails to achieve meaningful sensitivity while a random alert achieves similar specificity and negative predictive value.
- Epic's sepsis algorithm may have caused alert fatigue with 43% alert increase during pandemic Fierce Healthcare, on the JAMA Network Open study The 24-hospital finding: sepsis alerts per day rising 43% in the three weeks after each hospital's first COVID-19 case while total census fell 35%, alerts rising from 9% to 21% of the total, and the University of Michigan pausing the alerts in April 2020.
- Alert fatigue measurement in clinical decision support: a systematic review Ray, Wilson et al., JAMIA, 2026 The systematic review of implemented predictive models: 3,393 articles screened and 44 extracted, sepsis alert sensitivities from 10% to 100% and PPVs from 5.8% to 54%, improved process markers, and no high-quality study showing a mortality difference.
- Overall performance of a drug-drug interaction clinical decision support system PMC The documented alert override range of 49% to 96%, and the observation that overriding both irrelevant and clinically relevant alerts compromises the safety effect the system was deployed for.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — a widely deployed system whose aggregate performance required an independent study nobody was obliged to run. :: https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307 Aggregate Evidence Gap
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — how a category's composition determines what its aggregate can support. :: https://arxiv.org/abs/2111.15366 Aggregate Evidence Gap
Related articles
- 72 seconds, or 30 minutes, and both are trialsTerritory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials, CONSORT-AI reporting, peer review, and effect sizes that still differ by a factor of twenty-five.
- The sepsis model caught 7% of what clinicians missedA sepsis warning system ran at hundreds of US hospitals before anyone outside the vendor validated it. The external check found the number that matters is not the one being reported.
- Adding the doctor to the model changed nothingA randomised trial found physicians did better with an LLM than with conventional resources. Its second comparison, reported in the same abstract, found the model alone did just as well as the model plus the physician.
- Phase I improved. Phase II did not.AI-designed drugs clear safety trials at well above industry rates. At the stage that tests whether a drug works, the sources contradict each other, and the more careful ones report no advantage at all.