Home/Blog/Safety & governance/The sepsis model caught 7% of what clinicians missed
The sepsis model caught 7% of what clinicians missedThe only population an early-warning system exists for.100%missed by clinicians7%also caught by the modelOF SEPSIS CASES CLINICIANS MISSEDThe only population an early-warning system exists for.
The only population an early-warning system exists for.

The sepsis model caught 7% of what clinicians missed

A sepsis warning system ran at hundreds of US hospitals before anyone outside the vendor validated it. The external check found the number that matters is not the one being reported.

TL;DR. The Epic Sepsis Model scores hospital patients every fifteen minutes and alerts clinicians when it thinks sepsis is developing. It was running at hundreds of US hospitals before anyone outside the vendor published an independent check. When researchers at Michigan Medicine finally ran one across 38,455 hospitalisations, they found an area under the curve of 0.63 against the 0.76 to 0.83 the vendor had cited. It missed two thirds of sepsis cases while alerting on 18% of all patients, roughly 109 alerts for each true case. And the number that actually matters: of the sepsis patients clinicians had not already identified, it caught 7%. That is the only population an early-warning system exists for, and it is not the number anyone was reporting.

---

Status: established. Primary source: Wong et al., External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients, JAMA Internal Medicine 2021;181(8):1065-1070, with the accompanying editorial in the same issue. Figures are from that paper. The vendor's response is included below.

---

Sepsis kills. Early recognition allows treatment that measurably reduces mortality, and recognising it early is genuinely hard, which is why an automated warning is an attractive idea.

The Epic Sepsis Model takes around 80 clinical data elements from the electronic health record, vital signs, labs, comorbidities, demographics, and produces a risk score every fifteen minutes. Above a threshold, it fires an alert.

By 2021 it was deployed at hundreds of US hospitals.

No independent validation had been published.

What the check found

Researchers at Michigan Medicine ran a retrospective cohort covering every adult admitted between 6 December 2018 and 20 October 2019: 27,697 patients across 38,455 hospitalisations, with sepsis in about 7%.

At the vendor's own recommended alerting threshold:

Area under the curve: 0.63, against the 0.76 to 0.83 the vendor had reported.

Sensitivity 33%. It missed two thirds of sepsis cases.

Alerts on 18% of all hospitalised patients, which works out at roughly 109 alerts to find one true case.

Positive predictive value 12%. Nearly nine in ten alerts were wrong.

The number that was not being reported

Every figure above is a property of the model. The one that determines whether the system is worth having is different.

An early-warning system is not there to agree with clinicians. If a doctor has already recognised sepsis, an alert saying so changes nothing. The entire value of the tool is in cases the clinical team would otherwise have missed.

On that population it identified 7%.

That is the operational figure, and it is not the one in vendor materials, not the one in procurement discussions, and not the one an AUC captures. A model can have respectable discrimination overall while contributing almost nothing on the only subgroup it was bought for, because its correct predictions concentrate on the obvious cases a nurse would flag anyway.

This is the general form of the problem, not a quirk of sepsis. Any decision-support system deployed alongside competent humans should be measured on incremental contribution, not on standalone accuracy. Almost none are.

Why nobody checked first

This is the structural finding and it is more useful than the numbers.

The model is not a regulated medical device. It is clinical decision support embedded in an electronic health record, which in the United States sat outside the device authorisation pathway. Nothing required evidence of performance before deployment.

It was proprietary. The scoring logic was not public, so a hospital could not evaluate it analytically before buying.

And it was distributed as a feature. It arrived with the record system rather than as a separate procurement decision, which meant many hospitals turned it on without the review a standalone purchase would have triggered.

Those three together produce deployment at hundreds of sites with no evidence at all, and none of them is a technical failure. This is the same shape as the finding in AI in medicine, where 1.6% of cleared devices cited trial data: the evidence requirement was absent, so the evidence was absent.

The vendor's account

Included because the record should carry it.

The vendor's position was that the model exists to catch harder-to-recognise patients rather than obvious ones, pointed to prior research showing the model could predict sepsis, and stated that customers have complete transparency into it.

The first part is the strongest version of the defence and it is also what the 7% figure directly measures. If the purpose is the harder cases, then performance on the harder cases is the test, and that is the number the external validation reports.

Following publication, the vendor overhauled the algorithm and began recommending that hospitals train it on their own patient data before clinical use. That recommendation is a substantive change, and it concedes that a model shipped pre-trained on a national population was not adequately calibrated to any particular hospital.

What later checks found

A second external validation, using 2023 data from two county emergency departments across 145,885 encounters, found sensitivity of 14.7% and positive predictive value of 7.6% at the recommended threshold.

Different population, different definition of sepsis onset, different years. The direction is the same.

Four things this establishes

Deployment scale is not evidence. Hundreds of hospitals running a system tells you about procurement, not performance. It is routinely cited as though it were validation.

Vendor-reported and independently measured performance can differ substantially, and where the model is proprietary the buyer cannot tell which they are looking at.

Alert burden is a clinical harm, not a usability complaint. At 109 alerts per true case the rational response is to ignore the alerts, which degrades response to every other alert in the system. A system with poor precision does not merely fail to help; it consumes the attention that would have caught the case unaided.

And the right denominator is the cases humans miss. Standalone accuracy measures agreement with an existing process. Incremental contribution measures whether the system is worth its cost, and it is nearly always the smaller number.

What it does not establish

That the model never helps. A sensitivity of 33% means it identified a third of sepsis cases, some earlier than clinicians would have. The question is whether that offsets the alert burden, and the study does not resolve it.

That the current version performs the same. The algorithm was overhauled after publication and the recommendation now is local training. Whether that closed the gap is a separate question requiring separate validation.

That electronic health record vendors are uniquely at fault. The absent requirement is regulatory, and it applies to a whole category of clinical decision support rather than one product.

And that harm occurred in any specific case. This is a measured performance shortfall in deployed software. No individual patient outcome is attributed to it in the record, which distinguishes it from the other entries here.

Where this sits in the record

Seven cases in, this one differs from the rest in a way worth naming.

Moffatt, Zillow, the Dutch benefits scandal and Williams all describe harm that occurred. A refund refused, a write-down taken, families destroyed, a man arrested.

This one describes harm that cannot be counted. Sepsis missed by both a clinician and an alert leaves no artefact saying an algorithm failed. The patient deteriorates, and the cause recorded is sepsis.

Which means this failure mode is invisible to every incident register in existence. The registers count events somebody noticed. A warning system that quietly does not warn produces no event at all. That is not an argument that the harm is small; it is an argument that the record is structurally incapable of seeing it.

What is unresolved

Whether the overhauled model performs better. No large independent validation of the revised version has been published.

Whether locally trained versions work. The recommendation is now local training. Most hospitals lack the staff to do it properly, and there is no public evidence on how many have.

How many similar systems are running unvalidated. Clinical decision support of this kind is widespread and the requirement to publish performance is still largely absent.

And what the alert burden actually cost. Nobody has measured the downstream effect of 109 false alerts per true case on response to other alerts in the same system.

The counter-argument

An AUC of 0.63 is not nothing. Better than chance, and in a condition this hard to recognise, a modest signal applied continuously across every patient may still find cases. Reporting the figure as a failure implies a standard that few clinical prediction tools of any kind would meet.

The comparison may be unfair. The vendor's figures came from different populations with different sepsis definitions. Sepsis has several operational definitions and the choice moves the numbers substantially. Some of the gap between 0.63 and 0.76 is definitional rather than a performance shortfall.

The 7% figure is the harshest possible framing. It conditions on cases clinicians missed, which is a small and unusual subgroup, and small subgroups produce unstable estimates. It is the right question and it is measured with less precision than the headline figures.

And the alert burden is a threshold choice, not a model property. Hospitals set their own thresholds within the recommended range. A site drowning in alerts could raise it, trading sensitivity for precision. That the default produced 109 alerts per case is a configuration failure shared between vendor and hospital.

The short version

The Epic Sepsis Model scores hospital patients every fifteen minutes from around 80 data elements and alerts on suspected sepsis. It ran at hundreds of US hospitals before any independent validation was published.

When Michigan Medicine ran one across 38,455 hospitalisations, it found an AUC of 0.63 against a vendor-reported 0.76 to 0.83, sensitivity of 33%, and alerts on 18% of all patients, roughly 109 alerts per true case.

And on the only population that matters, sepsis patients clinicians had not already identified, it caught 7%. An early-warning system exists for exactly that group. Agreement with doctors who have already diagnosed the patient is not a benefit.

Nobody checked first because nothing required it. Not a regulated device, proprietary so unevaluable from outside, and shipped as a feature of the record system rather than as a purchase that would trigger review. The evidence was absent because the requirement was.

A second validation on 145,885 encounters in 2023 found sensitivity of 14.7%. Different population, same direction.

This entry differs from the others in the record. Moffatt, Zillow, the benefits scandal and Williams all describe harm that happened. This describes harm that cannot be counted. Sepsis missed by a clinician and an alert leaves no artefact naming an algorithm. Which means every incident register in existence is structurally blind to it, because registers count events somebody noticed, and a warning system that quietly fails to warn produces no event.

Common questions

What is the Epic Sepsis Model? A proprietary prediction tool built into the Epic electronic health record. It takes around 80 clinical data elements, vital signs, laboratory results, comorbidities and demographics, and produces a sepsis risk score every fifteen minutes. Above a configurable threshold it fires an alert to clinicians. By 2021 it was running at hundreds of US hospitals.

What did the external validation find? Wong and colleagues at Michigan Medicine studied 27,697 patients across 38,455 hospitalisations between December 2018 and October 2019, with sepsis occurring in about 7%. At the vendor's recommended threshold the model had an area under the curve of 0.63, against the 0.76 to 0.83 the vendor had cited. Sensitivity was 33%, specificity 83% and positive predictive value 12%. It generated alerts on 18% of all hospitalised patients, roughly 109 alerts for each true case.

Why is the 7% figure the important one? Because an early-warning system exists to catch what clinicians miss. If a doctor has already recognised sepsis, an alert confirming it changes nothing. Of the sepsis patients the clinical team had not identified, the model caught 7%. That is the incremental contribution, and it is the number that determines whether the system is worth its cost. Standalone accuracy measures agreement with an existing process, which is a different and much easier test.

Why was it deployed without validation? Three things together. It is clinical decision support rather than a regulated medical device, so nothing required performance evidence before deployment. It is proprietary, so hospitals could not evaluate the logic analytically. And it shipped as a feature of the record system rather than as a standalone purchase, so many sites enabled it without the review a separate procurement would have triggered.

What did the vendor say? That the model is intended to identify harder-to-recognise patients rather than obvious ones, that prior research showed it could predict sepsis, and that customers have complete transparency into it. The first point is the strongest form of the defence, and it is also precisely what the 7% figure measures. After publication the vendor overhauled the algorithm and began recommending that hospitals train it on their own data, which concedes that a nationally pre-trained model was not well calibrated to individual sites.

Is alert fatigue a serious problem or a complaint about usability? Serious, and clinical. At roughly 109 alerts per true case the rational response is to stop reading them, and that degrades response to every other alert in the same system. A low-precision system does not simply fail to help. It consumes attention that would otherwise have been available to catch the case unaided, which can leave a unit worse off than with no system at all.

Has performance improved since? The algorithm was overhauled after the 2021 publication and the vendor now recommends local training before clinical use. No large independent validation of the revised version has been published. A separate external validation using 2023 data from two county emergency departments across 145,885 encounters found sensitivity of 14.7% and positive predictive value of 7.6%, on a different population with a different sepsis definition.

Why does this case not appear in incident registers? Because it produces no incident. A patient whose sepsis is missed by both a clinician and an alert deteriorates, and the cause recorded is sepsis. Nothing in the file says an algorithm failed to fire. Incident registers count events that somebody noticed and reported, so a warning system that quietly does not warn is invisible to them. That is a limitation of the registers rather than evidence that the harm is small.

Sources

Primary documents only. Where a claim rests on a single report, the entry says so.

  1. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients Wong, Otles, Donnelly, Krumm, McCullough, DeTroyer-Cooley, Pestrue, Phillips, Konye, Penoza, Ghous and Singh, JAMA Internal Medicine 2021;181(8):1065-1070 The external validation. Every figure in this article comes from it, including the 7% of clinician-missed cases, which is the number the paper reports and almost nobody quotes.
  2. The Epic Sepsis Model Falls Short: The Importance of External Validation Habib, Lin and Grant, JAMA Internal Medicine 2021;181(8):1040-1041 The accompanying editorial, which sets out why deployment at scale without independent validation is the structural failure rather than the performance figures. Named rather than linked because the editorial sits behind the journal paywall.

Learn the concepts

← All posts