Home/Foundations/Construct Validity
Foundations

Construct Validity

Whether a measurement actually captures the thing it claims to, rather than something correlated with it that is easier to count.

Reviewed July 16, 2026Stable
Reading level: Curious
Pick your depth ↓

When not to use it

  • To dismiss all measurement. The argument is that scores support narrow claims, not that they support none.
  • Where the benchmark and the deployment task genuinely coincide, in which case the score is the thing you care about.
  • As a substitute for measuring your own use case, which is the only evaluation that resolves the question for you.

Reach for something else instead

  • Task-specific evaluation — measure on your own data and definition, which sidesteps the inference gap entirely.
  • Held-out replication sets — build a fresh test by the original protocol to separate benchmark performance from task performance.
  • -

Further reading

  • Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why general benchmarks cannot carry general claims.
  • Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference.

Primary sources, listed so you can check the claims on this page rather than take them on trust.

Where people go wrong

  • Reading a benchmark score as a capability level rather than as performance on that benchmark.
  • Assuming a harder benchmark has better validity; difficulty and validity are unrelated.
  • Comparing scores across benchmarks that operationalise the same construct differently.

At a glance

FieldFoundations
Originpsychometrics
Sibling conceptscontent and criterion validity
Core claima specific measurement cannot license a general claim
Compounded bycontamination, saturation
DifficultyIntermediate → Advanced
Flashcards for this concept · study, save or share them →
Question
Answer
1 / 4

Where this sits

A starting point. Nothing needs to come before it.

0Levelstarting point
0Needs firstnothing
0Opens upnothing further
1Areastays here

Computed from the prerequisite graph, not assigned. How this works