Self-Report Gap
Where asking the operator and measuring the artefact give different answers, systematically in the operator's favour, because effort saved is felt and cost deferred is not.
When not to use it
- Where no effort is deferred, displaced or diffused, in which case the participant sees the whole outcome and the report is the measurement.
- To dismiss survey evidence, which frequently captures things no telemetry can, including whether the output was worth producing.
- Where the two instruments measure genuinely different constructs rather than the same one from different positions.
Reach for something else instead
- Paired instrumentation — run both on one population, which converts a contradiction into a measurement of the displaced cost.
- Ask about the deferred portion directly — survey the reviewer and the maintainer rather than only the author, which recovers displaced cost without telemetry.
- -
Read more on the blog
- Refactoring fell from 25% to 3.8%Survey evidence says AI improves code quality. Repository telemetry says the opposite. They are measuring different things, and the gap between them is where the productivity went.
- AI in software: 19% slower, and they felt 20% fasterA randomised trial put experienced developers 19% behind on their own repositories while they reported being 20% ahead. The researchers had expected a speedup and published the opposite.
- 72 seconds, or 30 minutes, and both are trialsTerritory 11 opens on the first subject this corpus has examined where the evidence is genuinely good. Registered trials, CONSORT-AI reporting, peer review, and effect sizes that still differ by a factor of twenty-five.
- Nine subjects, and the model was almost never the blockerTerritory 10 closes. Across nine enterprise deployment subjects the binding constraint was organisational in every one, the evidence was commissioned in almost all of them, and the reconciling study exists nowhere.
Further reading
- METR (2025), randomised trial of experienced developers on their own repositories — self-estimated 20% speedup against a measured 19% slowdown.
- Wong et al. (2021), External Validation of a Widely Implemented Proprietary Sepsis Prediction Model — reported performance against independently measured performance on the population that mattered.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Averaging a survey figure and a measured figure as though they bracketed a true value.
- Treating the gap as evidence that respondents are exaggerating, when partial visibility explains it without any misreporting.
- Assuming a longer survey window fixes it, when the cost may land on a different person entirely.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works