Proxy Decay
A measurement that genuinely worked, because of a correlation nobody wrote down, and stopped working when that correlation broke.
When not to use it
- Where the metric never worked, which is ordinary construct invalidity and needs a different remedy.
- Where the environment has not changed, in which case a long track record is genuine evidence rather than a warning.
- As a reason to distrust all established metrics, since most are stable and the claim requires identifying the specific correlation that broke.
Reach for something else instead
- Scheduled revalidation — compare against an independently measured outcome on a fixed cycle rather than when someone happens to notice.
- Publishing the assumed correlation — state what has to remain true for the metric to mean what it claims, so the assumption can be challenged directly.
Read more on the blog
- Slop scores premium 70% of the timeThe first rigorous measurement of machine-generated content in ad buying found it passes every quality check the industry uses, and passes them better than real inventory does.
- Collapse needs you to throw the old data awayThe Nature result is correct under its stated condition. The condition is that each generation discards its predecessor's data, and the realistic case has a proof going the other way.
- 61% flagged for non-native writers, 3% for nativeThe tool that would answer how much text is machine-written does not work, and its errors concentrate on a specific group for a reason that is structural rather than fixable by tuning.
Further reading
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — operationalisation standing in for the construct it was meant to represent.
- Liang et al. (2023), GPT detectors are biased against non-native English writers — a statistical signature that stopped separating the populations it was assumed to separate.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Citing a metric's long history of working as evidence that it still works.
- Reading internal consistency as evidence of validity.
- Treating a decayed proxy as a fraud problem, when the numbers are usually accurate and the inference from them is not.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works