A benchmark is not a measure of capability in the abstract; it is a measurement instrument whose score supports — or fails to support — a specific interpretation for a specific use. This article applies the measurement-theory tradition of construct validity to AI evaluation, treating contemporary benchmarks as operational definitions whose claimed inferences are routinely under-evidenced. It catalogs four recurring validity failures, sketches the structure of a validity argument, and ends with the unresolved question of whether construct validity is achievable for general-capability evaluation or a category mistake when the construct is reduced to a single leaderboard score.
Coverage note: verified through May 6, 2026.
What construct validity is, and is not
In classical measurement theory, construct validity is the degree to which the evidence and theory available about a measurement support the interpretation of its scores as reflecting the construct that the measurement is claimed to assess. The locus classicus is Cronbach and Meehl's 1955 paper "Construct Validity in Psychological Tests," which framed construct validation as an iterative process of articulating a nomological network — a web of theoretical relationships that the construct is supposed to participate in — and then collecting evidence that the test's scores behave inside that network as the construct would predict (Cronbach & Meehl 1955). Samuel Messick's later unification of validity theory pushed this further: validity is not a property of a test but of the inferences and uses that scores license, with consequences of those uses included in the validity argument (Messick 1989).
Two implications are load-bearing for this article. First, validity does not attach to a benchmark; it attaches to a particular interpretation of a benchmark score for a particular decision context. The same dataset can be valid evidence for one inference (this model variant has not regressed on Python function synthesis from self-contained prompts) and invalid for another (this model is a better reasoner than its competitor). Second, reliability — score consistency under repeated administration — is necessary but not sufficient. A perfectly reproducible measurement of the wrong thing is still a measurement of the wrong thing.
The contemporary AI evaluation literature has begun to adopt this vocabulary, but unevenly. Benchmark designers often invoke "construct validity" as a synonym for "the benchmark is real and useful," which inverts Messick's point. Critics often invoke it as a synonym for "this benchmark is bad," which is closer to Messick but typically skips the validity argument's actual structure: what construct is being claimed, what operation produces the score, what inference is being made from the score, and what evidence supports or defeats that inference.
A useful taxonomy distinguishes construct validity from several adjacent notions that benchmark discourse tends to merge:
| Validity type | Asks |
|---|---|
| Construct | Do scores reflect the intended construct? |
| Criterion | Do scores predict an independent outcome that matters? |
| Convergent | Do related measures move together? |
| Discriminant | Do unrelated constructs stay separable? |
| Content | Does the item universe represent the construct's domain? |
| External / ecological | Do results generalize to the populations and contexts of intended use? |
| Reliability (not a validity type) | Are scores reproducible? |
These are not interchangeable boxes to check; they are sources of evidence in a validity argument. Criterion and convergent evidence support the inference that scores reflect the construct. Discriminant evidence rules out alternative explanations such as "scores measure verbosity" or "scores measure prompt-format competence." External evidence bounds the claim's domain of application. The argument either hangs together or it does not. A checklist that confirms each cell without integrating them produces compliance, not validity.
Why this matters for AI
Most public AI benchmarks are operational definitions of capabilities. A dataset, a prompt protocol, a sampling regime, an inference harness, a scoring rubric, and a tool policy together specify exactly what behavior counts as an instance of "reasoning," "coding ability," "factuality," "instruction following," "helpfulness," "safety," or "general intelligence." Once published, the benchmark is the operational referent for the construct in any debate that cites it. If the operation does not faithfully sample the construct, the benchmark does not measure what its name advertises. If the inference from score to capability is not warranted by an evidence base, the benchmark does not support the leaderboard claims it is used to underwrite.
This is the central observation of Raji, Denton, Bender, Hanna, and Paullada's "AI and the Everything in the Whole Wide World Benchmark," presented at the NeurIPS 2021 Datasets and Benchmarks track (Raji et al. 2021). The paper's argument is not that benchmarks are useless. It is that the social life of benchmarks routinely transports them from narrow, situated datasets to expansive claims about progress in vague, contested constructs — "intelligence," "language understanding," "reasoning" — without the validity argument that would license the move. The article calls this a problem of construct validity in the strict measurement-theory sense: the inference from a finite, particular dataset to a sweeping ability claim is unsupported, and often unsupportable, because the construct itself is underdefined and the evidence base is missing.
Bowman and Dahl's "What Will it Take to Fix Benchmarking in Natural Language Understanding?" reaches a parallel conclusion from a more operational angle (Bowman & Dahl 2021). Even on narrow NLU tasks, current practices fail to support the inferences benchmarks are used to license: annotation reliability is poor, statistical power is low, social bias is often unmeasured, adversarial collection introduces selection effects, and saturation arrives before the underlying construct is understood. Liao, Taori, Raji, and Schmidt's "Are We Learning Yet?" reviewed 107 ML survey papers and catalogued recurring internal- and external-validity failures in the way ML benchmarks instantiate the broader learning problems they purport to study (Liao et al. 2021). The cumulative empirical record across these critiques is consistent: AI benchmarks frequently outrun the validity arguments their published claims would require.
A prudent reader should not generalize this into "most benchmarks fail construct validity." That claim is too blunt; it conflates the benchmark with the inference. The defensible formulation is narrower and stronger: many benchmark-based capability claims lack adequate validity evidence for the inferences they are routinely used to support. The same dataset can be a valid regression instrument under a fixed elicitation regime and an invalid measure of broad capability. The validity question is always: valid for what inference, in what use?
The rest of this article works in that frame. It catalogs four recurring failure modes, distinguishes them carefully, and locates each within the structure of a validity argument. It then turns to the literature on benchmark contamination, where the 2025 evidence base is now strong enough to be treated as a primary failure mode rather than a hygiene footnote. It closes on the unresolved question of whether construct validity is tractable for general-capability evaluation, or a category mistake when the target construct is reduced to a single number.
Four recurring failure modes
The four failures most commonly attributed to construct validity in AI benchmarks are operationalization slippage, contamination, format-matching mode collapse, and missing real-world diversity. They are real, recurring, and worth naming. They are not, however, the same kind of failure. Treating them as a homogeneous list invites a confusion the validity-argument frame is supposed to prevent.
| Failure mode | Mechanism | Validity component most directly threatened |
|---|---|---|
| Operationalization slippage | Task proxy stops matching the claimed construct | Construct (the operation no longer indicates the construct) |
| Contamination | Exposure to items or near-variants breaks the score-to-capability inference | Score interpretation; sometimes construct, often internal validity and held-out assumption |
| Format-matching mode collapse | Models optimize benchmark form, prompt regularities, or answer schema rather than the target behavior | Discriminant (cannot rule out construct-irrelevant variance) |
| Missing real-world diversity | Item universe omits contexts, users, languages, interfaces, stakes, and adversarial pressure of intended use | External / ecological |
Each entry compresses substantial structure. The table is meant as a forcing function for the rest of this section: a benchmark critique that does not localize the alleged failure to a component of the validity argument is doing rhetoric, not measurement.
Operationalization slippage
Operationalization slippage is the most direct construct-validity failure of the four. The benchmark's items, prompt protocol, and scoring rule are jointly the operational definition of the construct. Slippage occurs when the operation drifts away from the claimed construct without anyone updating the claim. A canonical instance is the migration from multiple-choice science item accuracy to "scientific reasoning" claims. Multiple-choice success can be supported by surface heuristics: the position of the correct answer, the distribution of distractor lengths, the distinctive vocabulary that correct options use under crowdsourced authoring, or the prior probability of options under a generic language model. None of these are the construct; all of them can produce high benchmark scores; the validity argument that connects multiple-choice item performance to scientific reasoning has rarely been made in detail for any popular science-QA benchmark.
The pattern recurs whenever the operation is much narrower or much wider than the named construct:
- A code benchmark that scores function bodies passing hidden tests under self-contained problem statements is an operational definition of constrained Python synthesis from contest-style prompts. It is not, without further evidence, an operational definition of programming ability. The latter would require evidence about debugging unfamiliar codebases, integrating with real APIs, refactoring, reading test failures, and operating under ambiguous specifications.
- A safety benchmark that scores refusal rates on a static red-team set is an operational definition of refusal under that distribution of prompts. It is not, without further evidence, an operational definition of safety in deployment. The deployment distribution includes attempts at adversarial elicitation, multi-turn manipulation, jailbreaks composed at inference time, and user populations the benchmark does not sample.
- An "agentic" benchmark that scores end-to-end task completion in a sandboxed environment with curated tools is an operational definition of task completion under that sandbox's affordances. Generalizing it to autonomy requires a convergent evidence base across environments, tool sets, and failure modes.
The remedy is not to stop building narrow benchmarks. Narrow benchmarks are often the only kind it is possible to build well. The remedy is to keep the inference inside the benchmark's evidence envelope. Operationalization slippage is the validity failure that occurs when that discipline lapses.
Contamination
Contamination is the failure mode that has changed most rapidly during 2024 and 2025. It is also the one most likely to be misclassified. Contamination is sometimes presented as a single phenomenon — "the model has seen the test set" — when in fact several distinct mechanisms produce different validity threats:
- Direct exposure: benchmark items appear verbatim in the pretraining corpus.
- Near-duplicate exposure: paraphrased, translated, or otherwise transformed variants of items appear in pretraining.
- Public solution leakage: not the items themselves but their solutions, walkthroughs, or community discussions are present in the corpus.
- Benchmark-specific fine-tuning: post-training pressure (RLHF, instruction-tuning datasets, evaluation-harness exemplars) targets the benchmark or close cousins.
- Temporal leakage: benchmark items predate the cutoff and are increasingly likely to be reflected in future training distributions, even when no leak is intentional.
These mechanisms break different things. Direct exposure most cleanly attacks the held-out assumption: scores no longer reflect generalization to unseen items because the items are not unseen. Near-duplicate exposure attacks construct validity itself when the duplication pattern is correlated with the construct (for instance, if the duplicated items share a stylistic signature absent from the broader construct domain). Public solution leakage tends to inflate scores in ways that decouple them from any plausible interpretation as construct measurement. Benchmark-specific fine-tuning attacks both construct validity (the model is now trained on the operation, not the construct) and external validity (the score predicts behavior on this benchmark, not in deployment). Temporal leakage is the slow-rolling version of all of the above and is the reason benchmarks age.
The 2025 literature is now mature enough to treat as a primary evidence base. LiveBench, presented at ICLR 2025, was deliberately designed as a contamination-limited benchmark with frequently updated source items drawn from post-cutoff math, coding, and reasoning material, scored objectively to remove judge contamination (White et al. 2025). AntiLeakBench, also at ICLR 2025, takes a complementary approach, building benchmark items on real-world knowledge that postdates a model's training cutoff and offering an automated update protocol (Wu et al. 2025). Sander, Soulier, and colleagues' 2025 work analyzed how search-time contamination — exposure to evaluation items through retrieval rather than pretraining — can inflate scores in ways that conventional contamination audits miss (Sander et al. 2025). Han, Patil, and colleagues' contemporaneous work probed the reliability of contamination detection itself, finding that several popular detection methods produce inconsistent verdicts on the same model–benchmark pair (Han et al. 2025). Sun et al.'s ICML 2025 paper on contamination-resistant benchmark construction argued that many proposed mitigation strategies fail to balance fidelity to the original construct against genuine contamination resistance — a warning against treating "decontaminated benchmark" as a self-validating label (Sun et al. 2025).
The lesson is that contamination is a cross-cutting validity threat with several distinguishable mechanisms, not a single phenomenon. A benchmark report that addresses one mechanism (verbatim string matching, say) while leaving others unaddressed (near-duplicate semantic overlap, post-training selection, retrieval-time exposure) has not, by itself, established validity. Conversely, a benchmark cannot be dismissed as wholesale invalid on contamination grounds alone; the question is which inferences the evidence still supports. A leaderboard claim that "model A reasons better than model B" is hard to defend if the gap can be explained by exposure differences. A regression claim that "model A has not gotten worse on this fixed task" can survive substantial contamination, because the inference does not require generalization beyond the benchmark.
Format-matching mode collapse
The phrase "mode collapse to format-matching" is more imprecise than its frequency in benchmark critique suggests. In the strict generative-modeling sense, mode collapse refers to a sampling distribution narrowing to a small subset of the data manifold. As applied to benchmarking, the phrase is being used to describe something different: scores being driven by the benchmark's surface form — prompt template, answer schema, scoring conventions — rather than the construct the benchmark claims to measure. A more accurate label is construct-irrelevant variance from format adaptation, but the shorter phrase has stuck.
The validity component most directly threatened is discriminant. The validity argument needs to rule out the alternative explanation that performance reflects test-taking competence rather than the construct. The empirical signal that this alternative is in play is straightforward: scores should be substantially stable under construct-preserving perturbations of format. If reordering options, paraphrasing the prompt template, swapping the answer schema (single letter vs. JSON vs. natural language), changing the chain-of-thought elicitation, or moving from zero-shot to a different few-shot composition substantially changes scores or, worse, changes leaderboard ranks, the score is partly a measurement of format competence. The construct is not insulated from the operation as the validity argument requires.
The recent literature has documented this pattern repeatedly. Prompt-format sensitivity studies have shown rank reordering across models when only superficial features of the prompt change. Self-consistency, majority voting, and structured-output protocols can shift scores by margins comparable to the gaps between models on the leaderboard. In cases where the score is taken as evidence of the construct, this is a validity failure. In cases where the score is taken as evidence of behavior under a fixed prompt protocol, it is not — the inference has been narrowed to fit the operation. Format-matching, in other words, is a failure of validity argument hygiene more than a property of the benchmark.
Missing real-world diversity
The fourth common failure is the easiest to name and the easiest to flatten. "Missing real-world diversity" is often used as an undifferentiated complaint that the benchmark is unrealistic. As a validity failure, it is more usefully decomposed along the axes that matter for the construct's intended use:
- Task diversity — does the benchmark sample the range of tasks the construct is supposed to cover, or only a stylized subset?
- User population — are the inputs representative of the people who will interact with deployed systems, or are they written by a narrow set of authors with predictable conventions?
- Language — is the construct claimed to apply across languages? If so, is the operation multilingual? If not, is the claim explicitly scoped?
- Domain — does the benchmark cover the specialized vocabularies, conventions, and base rates of real domains, or only the domain-neutral subset?
- Interface — does the benchmark sample the affordances of deployed systems (multi-turn, tool use, retrieval, file I/O, browsing, code execution) or only single-turn closed-book responses?
- Stakes — does the benchmark capture the failure cost asymmetries of the deployment context?
- Adversarial pressure — does the benchmark anticipate the inputs that motivated agents will craft, or is it limited to cooperative authorship?
Naming the axes clarifies whether the failure is properly external/ecological (the benchmark cannot generalize to the deployment population) or content (the construct's domain is broader than the operation samples). The remedy is, again, scope discipline. A benchmark that samples one slice along each axis can still license a corresponding inference; what it cannot license is a claim that scales to the union of all slices. Salaudeen, Bommasani, and colleagues' 2025 "Measurement to Meaning" framework articulates this point as a claim-centered approach to AI evaluation: the validity question is downstream of the claim, and evaluation choices should be reverse-engineered from the claim the score is meant to support (Salaudeen et al. 2025). Bean, Liang, and colleagues' systematic review of 445 LLM benchmarks supplies the empirical companion: across the corpus, validity-evidence reporting is sparse, and the gap between what benchmarks claim and what they evidence is a structural feature of the field rather than a property of a few outlier projects (Bean et al. 2025).
Building a validity argument
The validity-argument tradition descends from Cronbach, Meehl, and Messick. Its modern form, distilled for AI evaluation by Liao and colleagues and by the more recent measurement-theory work of Freiesleben and Zezulka, treats a validity argument as five linked components (Freiesleben & Zezulka 2025):
- Construct definition. What is the named capability supposed to be? What does it include and exclude? A construct that cannot be specified beyond its name cannot be validated; one that can be specified can be evaluated for whether the operation samples it.
- Operation specification. What dataset, prompt protocol, sampling regime, inference harness, scoring rubric, and tool policy produce the score? Validity-relevant details often hide in the protocol, not the dataset.
- Claimed inference. What is the score being used to support? "Model A scores higher than model B on this benchmark" is one inference. "Model A is more capable" is another. "Model A is preferable for deployment in this context" is a third. Each inference makes different evidentiary demands.
- Evidence. What evidence supports the inference, and what evidence would defeat it?
- Scope. Where does the claim stop? What populations, conditions, and decisions does the validity argument cover, and which fall outside it?
Inside (4), several evidence types do specific work. Criterion evidence asks whether the benchmark score predicts an independent outcome that matters — downstream task success, user task completion, post-deployment incident rates, expert ratings of unscripted behavior. Convergent evidence asks whether independently designed measures of the same construct correlate with the benchmark in question. Discriminant evidence asks whether plausibly unrelated constructs (verbosity, instruction-format compliance, prompt familiarity, length of reasoning trace) remain separable from the benchmark score. Content evidence asks whether the items represent the construct's domain. Reliability evidence asks whether scores are stable under repeated administration.
Several supporting analyses are not validity types in their own right but are necessary inputs to a validity argument:
- Construct-preserving perturbation tests: do scores remain stable when only the format is varied?
- Temporal holdouts: do scores hold up on items drawn from after the model's training cutoff?
- Contamination audits: do exposure indicators correlate with score advantage?
- Disaggregated analysis: do scores hold uniformly across sub-populations of the item universe and the user space?
- Rank stability: do the orderings on the leaderboard survive resampling, perturbation, and protocol changes?
A common confusion in benchmark discourse is to treat each of these as a validity type to be stamped on the benchmark's documentation. They are not. They are evidence within a validity argument. Their value depends on what is being claimed. Rank stability under prompt format perturbation is excellent evidence for the inference "this score reflects more than format adaptation." It is weaker evidence for "this score reflects the named construct," because a benchmark can be stable under perturbation and still not be measuring what its title says it is. Contamination audits raise the floor of plausibility for any inference, but they are not, on their own, evidence for a positive interpretation of the score. Validity is the integrated argument; the audits and perturbations are its inputs.
This is the right place to address the question of whether validity is best presented as an audit checklist. A checklist is a useful artifact for benchmark builders and for downstream consumers, particularly when published validity arguments are sparse and a single page of structured questions does substantial work. The risk is that a checklist becomes a substitute for the integrated argument it is meant to summarize. The cleanest practical form is to treat the checklist as a summary instrument of an underlying argument that is documented elsewhere — in a benchmark paper, model card, or evaluation report — rather than as a self-contained verification.
The active critiques
Three lines of critique structure most of the current literature and are worth treating as substantive arguments rather than as decorative citations.
Raji et al. on benchmark over-reach
Raji, Denton, Bender, Hanna, and Paullada's 2021 NeurIPS Datasets and Benchmarks paper is the most-cited single piece in the construct-validity-for-AI literature. Its argument is that benchmarks acquire a status in the discourse of AI progress that vastly outruns what any single dataset can support (Raji et al. 2021). The paper traces the historical migration from narrow datasets (ImageNet for object recognition, SQuAD for extractive QA, GLUE/SuperGLUE for NLU) to capability claims (visual recognition, reading comprehension, language understanding) without a corresponding migration in evidence base. The structural failure is what the paper calls abstraction inflation: a finite, situated dataset is treated as a comprehensive operationalization of an abstraction whose actual scope vastly exceeds the dataset's coverage.
The paper's contribution to construct-validity discourse is to make the social mechanism explicit. Benchmarks are not validated into representing constructs by their authors; they are promoted to representing constructs by their citations. Once a benchmark functions as a leaderboard target, the inferences attached to its scores are determined by the discourse around it, not by the validity argument the authors did or did not make. This is why a benchmark designer's "this dataset is intended only as a regression test" disclaimer rarely survives downstream use. The construct validity question becomes a question about a community's interpretation of scores, and that interpretation is hard to bound from inside any single paper.
Bowman & Dahl on NLU benchmarking
Bowman and Dahl's "What Will it Take to Fix Benchmarking in Natural Language Understanding?" attacks the problem at the operational level and stays close to the data (Bowman & Dahl 2021). The paper argues that NLU benchmarking, even at its narrowest, fails to support the inferences it is used to license. Annotation reliability is poor on a substantial fraction of items, inducing noise in ground truth that current models can match or exceed without solving the underlying construct. Statistical power is low: leaderboard gaps that drive headlines and product claims are often not distinguishable from sampling noise under standard tests. Adversarial data collection, intended to harden benchmarks, introduces selection effects that produce items predictable from features unrelated to the construct. Saturation arrives before the construct is understood: a benchmark whose ceiling is reached without an associated theoretical account of what was learned is, by Bowman and Dahl's argument, a sign that the benchmark was measuring something narrower than it claimed.
The paper's contribution to construct validity is methodological. It demonstrates that even when a benchmark community is close to the data and competent at measurement, the validity argument can fail at the operational layer rather than at the level of construct definition. This implies that fixing construct validity is not only a matter of reining in marketing claims; it is also a matter of repairing the basic measurement hygiene of the benchmarks themselves.
Liao et al. on measurement theory for ML
Liao, Taori, Raji, and Schmidt's "Are We Learning Yet?" is the broadest of the three (Liao et al. 2021). The paper reviews 107 ML survey papers across application areas and synthesizes recurring failures along internal- and external-validity dimensions. The internal-validity failures concern whether benchmark experiments support the conclusions drawn within the experimental design: leakage, overfitting to test sets, weak baselines, p-hacking, and selection effects in reported results. The external-validity failures concern whether benchmark performance generalizes to the broader learning problems benchmarks are taken to instantiate.
The paper's contribution to construct validity is taxonomic. It pushes back against a tendency in benchmark critique to file every failure under "construct invalidity." Internal-validity failures are about the quality of the experiment; external-validity failures are about generalization across populations and contexts; construct-validity failures are about whether the operation indicates the construct. The same benchmark can fail any one, two, or all three of these. Naming the relevant failure mode supports cleaner remediation. Decontaminating a benchmark addresses one cluster of internal-validity threats; broadening the item universe addresses external validity; specifying the construct and tightening the operation addresses construct validity. The taxonomy is an argument against the single-bucket mistake.
Companion threads
Several supporting threads belong in the literature layer without occupying a section of their own. Salaudeen, Bommasani, and colleagues' "Measurement to Meaning" frames AI evaluation as claim-centered, demanding that evaluation choices be reverse-engineered from the claim a score is meant to support (Salaudeen et al. 2025). Bean, Liang, and colleagues' systematic review of 445 LLM benchmarks reports the structural sparsity of validity-evidence reporting at corpus scale (Bean et al. 2025). Freiesleben and Zezulka's 2025 piece imports the unified validity-theory tradition into AI evaluation directly, with explicit attention to the consequential dimension of validity that Messick emphasized (Freiesleben & Zezulka 2025). Xiao, Zhang, Lai, and Liao's earlier work on measurement theory for natural-language generation ties validity language to evaluation metric design, complementing the benchmark-level framing of Raji et al.
A coherent reading of these threads converges on a position narrower than "most benchmarks fail construct validity" but more substantive than "benchmarks are useful, don't worry about it." Benchmarks operate as operational definitions; the inferences attached to their scores routinely outrun the validity arguments their evidence base supports; the gap is structural, well-documented at corpus scale, and not closeable by any single intervention.
When this fails: the contested interior
The literature does not speak with one voice, and the disagreements are part of the topic's substance.
The first contested point is whether benchmark and measurement instrument are the right model for what AI evaluations actually are. The measurement-theory tradition assumes that the instrument exists in some independence from the system being measured. AI benchmarks frequently violate this. Benchmark items are released into the corpus from which future models are trained. Leaderboards are explicit selection environments: a model that performs better on a benchmark is more likely to be released, scaled, fine-tuned further, and built into products. Over time, the population of "models being measured" is conditioned on the population of benchmarks doing the measuring. This is not the unproblematic measurement situation Cronbach and Meehl had in mind. A position that takes this seriously holds that psychometric validity language imports assumptions that AI evaluation does not satisfy; a softer position holds that the assumptions are partly violated and that contamination-resistant designs (LiveBench, AntiLeakBench, post-cutoff item generation) are the field's attempt to restore them. The article takes the softer position: the assumptions are partly violated and partly recoverable, and the recovery is one of the field's active research programs.
The second contested point is whether contamination should be classified as a construct-validity failure or as a separate category. The argument for separate categorization is that contamination, in its primary mechanism, attacks the held-out assumption that lets a finite test sample stand in for the construct's broader item universe. If a model has memorized the item, the item's score no longer indicates anything about the construct; the score has become an indicator of memorization. This is closer to an internal-validity failure (the experiment does not support its own conclusions) than to a construct mismatch (the operation does not indicate the construct). The argument for classifying contamination under construct validity is that, in practice, the inferences benchmark scores are used to license are construct-level inferences ("this model reasons better"), and contamination breaks those inferences directly. Both readings are defensible. The article's working position is that contamination is a cross-cutting validity threat: its primary mechanism varies across the five mechanisms enumerated above, and the validity component it most directly attacks varies with the mechanism. The taxonomic discipline is to localize the failure case by case rather than to cement contamination into one slot.
The third contested point is the scope of "construct validity for general capability." The strong skeptical position holds that "general capability" is not a coherent latent construct in the psychometric sense: AI systems' observed performance is jointly produced by base model, prompt, decoding parameters, tool access, scaffolding, retrieval, refusal policy, and benchmark-specific training pressure, and is therefore not "in" the model the way a latent trait is in a person. On this reading, attempting to validate a general-capability benchmark is a category mistake. The strong recoverist position holds that "general capability" can be decomposed into nested subconstructs (mathematical reasoning under symbolic constraints, scientific problem solving under tool access, long-horizon task execution under realistic context, calibrated factual recall, instruction following under specified ambiguity), each of which can be validated, and that the aggregate construct then inherits the validity of its decomposition. The recoverist position has not yet been demonstrated by any current general-capability leaderboard. The skeptical position has not been demonstrated to be unconditional. The article's working position is intermediate and conditional: construct validity is tractable for bounded capability claims under specified elicitation regimes, currently underargued for broad aggregate claims, and plausibly incoherent for a single scalar "general capability" unless the construct is formally decomposed and validated across contexts. This is an open empirical question, not a settled philosophical one.
The fourth contested point is the proper relationship between benchmarks and engineering use. Many benchmarks are built and used as regression tests within a model family. For that use, construct validity in the strong sense is not the relevant question; reliability and within-family rank stability are. A benchmark that reliably detects within-family regressions on a fixed task distribution is doing useful work even if its score does not generalize to broad construct claims. A position that takes this seriously calls for a two-track account of benchmarks: one track for engineering instruments, where validity is local and the inferences are narrow, and another track for capability claims, where validity is global and the inferences are broad. The two tracks should not be merged. Many of the most contentious benchmark debates are about benchmarks that began life on the engineering track and migrated, through external citation, onto the capability track. The two-track account makes the migration visible and contestable.
A practical validation sketch
The validity-argument frame can produce a concrete protocol. The cheapest way for a benchmark consumer to test whether a benchmark supports the inference attached to it is a claim-centered mini-validation study. The protocol is not decisive for general capability — no small battery is — but it is decisive for many narrower claims and is decisive enough as falsification probe to expose overclaims that would otherwise survive scrutiny.
The sketch:
- Pin the claim. Write down the inference the score is being used to support, in operational terms. Not "model A is a better reasoner," but "model A is more likely to produce correct multi-step solutions to symbolic math problems under standard prompting."
- Build a private criterion set. Construct 100–300 items drawn from sources that postdate the model's training cutoff or that are otherwise unlikely to be in the corpus. Match the items to the claimed construct in content, but vary their surface features (prompt style, answer schema, distractor design, problem framing) along axes the construct should be insulated from.
- Add construct-preserving perturbations. For a meaningful subset of items, produce variants that preserve the construct (the answer is the same, the reasoning is the same) but change the surface form (paraphrase, reorder options, change format from MCQ to free response, alter the chain-of-thought elicitation).
- Evaluate a fixed model panel. Run six to ten models drawn from multiple model families under fixed elicitation protocols. Single-family panels do not exhibit the discriminant variance the analysis requires.
- Run five tests:
- Format stability. Do scores remain within sampling noise across construct-preserving perturbations? Do leaderboard ranks?
- Criterion correlation. Does the original benchmark predict the private criterion set's scores beyond simple proxies (answer length, instruction-format compliance)?
- Convergent comparison. Does the original benchmark correlate with independently designed measures of the same construct?
- Discriminant separation. Do unrelated constructs (verbosity, format compliance, prompt-genre familiarity) explain less variance than the named construct?
- Contamination indicators. Do exposure proxies (string overlap, semantic overlap, post-training selection signals, model-family release timing relative to benchmark publication) explain score advantage?
A benchmark whose original score predicts the private criterion better than simple proxies, remains stable under perturbation, correlates with independently designed measures of the construct, separates from unrelated constructs, and is not explained by exposure indicators is a benchmark for which the original construct claim has positive validity evidence. A benchmark that fails any of these tests has had a specific kind of validity evidence withdrawn; the inference must narrow accordingly.
The protocol has known limits. It cannot validate a broad construct that is not operationally specified. It cannot rule out failure modes that are not represented in the perturbations or in the criterion set. It cannot, on its own, distinguish "the construct is real and the operation captures it" from "the construct is artifactual and the operation captures the artifact." But it can, with modest effort, falsify a substantial fraction of currently active capability claims, and that asymmetry is the practical value of the validity-argument frame.
Is construct validity tractable for general capability?
The article's final question is whether construct validity is achievable for general-capability evaluation, or a category mistake when the target construct is reduced to a single leaderboard score.
The framing of the question matters. If "general capability" is understood as a single scalar latent trait that AI systems possess to varying degrees and that can be measured by a sufficiently well-designed public static benchmark, then current evidence does not support the goal. The reasons are cumulative. Benchmark scores depend on prompt format, decoding, tool access, scaffolding, retrieval, and post-training adaptation in ways that latent traits do not depend on item presentation in well-designed psychometric instruments. The construct itself shifts under deployment conditions: a model's "reasoning" looks different in a sandbox with a Python kernel and a retrieval tool than it does on a closed-book MCQ. The model population is conditioned on the benchmark population through release selection, which violates the independence assumption that lets a fixed instrument stand in as a measure across model families. Aggregate scores hide the dimensions along which models differ, collapsing convergent and discriminant structure into a single number whose interpretation requires evidence that the aggregation preserves. None of these problems are unsolvable in principle. None of them have been solved in current practice. A single-number general-capability leaderboard is, on present evidence, not licensed by an adequate validity argument.
If "general capability" is understood instead as a structured collection of subconstructs, each operationally specified and each validated under the standard machinery of criterion, convergent, discriminant, and external evidence, then the goal is tractable in principle and partly tractable in practice for some subconstructs. Bounded capabilities that have a natural operational definition — Python function synthesis from self-contained prompts under hidden tests; closed-book recall of post-cutoff factual questions under a fixed retrieval-free protocol; refusal under a documented adversarial-prompt distribution — can be validated, and a respectable subset of current benchmarks support inferences in this neighborhood. What remains undemonstrated is the integrative step: aggregating validated subconstructs into a coherent claim about general capability that survives the same evidentiary discipline.
The intellectually honest position is therefore neither benchmark nihilism nor leaderboard realism. It is claim discipline. State exactly what the score is being used to measure, what inference it supports, what evidence would falsify that inference, and what remains unmeasured. Resist the move from a narrow operational result to a broad capability claim unless the validity argument bridges the gap. Treat contamination, format-matching, missing diversity, and operationalization slippage as distinct threats to specific components of the validity argument, not as interchangeable instances of "the benchmark is bad." Distinguish the engineering-regression use of benchmarks, where local validity suffices, from the capability-claim use, where global validity is the relevant standard. Hold the general-capability question open. The literature does not yet support a settled answer, and the prudent position is to write the article as if it doesn't.
Companion entries
Core theory: - Measurement Theory - Construct Validity - Nomological Network - Reliability vs Validity - Internal Validity - External Validity
Practice: - AI Evaluation - AI Capability Evaluation - Benchmark Contamination - Leaderboard Overfitting - Format-Matching Mode Collapse - Operational Definition (Benchmark) - Evaluation Methodology - Engineering Regression Tests - Claim-Centered Evaluation - Mini-Validation Study
Counterarguments and open questions: - General Intelligence (AI) - Capability Decomposition - Category Mistakes in AI Evaluation - Two-Track Account of Benchmarks