Construct Validity
Whether a measurement actually captures the thing it claims to, rather than something correlated with it that is easier to count.
When not to use it
- To dismiss all measurement. The argument is that scores support narrow claims, not that they support none.
- Where the benchmark and the deployment task genuinely coincide, in which case the score is the thing you care about.
- As a substitute for measuring your own use case, which is the only evaluation that resolves the question for you.
Reach for something else instead
- Task-specific evaluation — measure on your own data and definition, which sidesteps the inference gap entirely.
- Held-out replication sets — build a fresh test by the original protocol to separate benchmark performance from task performance.
- -
Read more on the blog
- Why AI benchmarks mislead: contamination, gaming, saturationEvery model launch leads with benchmark scores, and buyers read them like thermometer readings. They are closer to opinion polls: directionally useful, methodology-dependent, and easy to game. Here is why the numbers mislead, from training-data contamination to Goodhart's law to saturation, and how to read them without being fooled.
- How do we measure AI progress? The benchmark problemEvery AI model launches with a table of benchmark scores that look like objective proof of progress. A growing body of evidence says those numbers are far less trustworthy than they appear, because benchmarks saturate, leak into training data, and reward the wrong thing. This is the evaluation problem, and it has quietly become one of the hardest parts of AI.
- 621,000 robots installed, virtually no humanoidsTerritory 7 opens on the question article 117 left unresolved. Industrial robotics is enormous and growing. The general-purpose machine that would move the boundary has almost no deployments.
- 2.6 million robotic surgeries, none of them autonomousThe most-deployed medical robot makes no decisions. And the largest meta-analysis of its outcomes was co-authored by the company that sells it.
Further reading
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why general benchmarks cannot carry general claims.
- Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference.
Primary sources, listed so you can check the claims on this page rather than take them on trust.
Where people go wrong
- Reading a benchmark score as a capability level rather than as performance on that benchmark.
- Assuming a harder benchmark has better validity; difficulty and validity are unrelated.
- Comparing scores across benchmarks that operationalise the same construct differently.
At a glance
Where this sits
A starting point. Nothing needs to come before it.
Computed from the prerequisite graph, not assigned. How this works