Reference

Goodhart's Law

Goodhart’s law is a warning about using an observed measure as an instrument of control: a measure can be informative while people, organizations, or systems merely respond to the underlying condition it tracks. Once rewards, penalties, selection, or optimization are attached to the measure itself, the process generating the observations you act on changes — sometimes because behaviour changes, and sometimes, under pure selection, only because a filter now stands between the world and what you see. The old relationship between measure and goal may then weaken, disappear, or reverse.

The compact aphorism usually associated with the law is: “When a measure becomes a target, it ceases to be a good measure.” That sentence is memorable, but it is also easy to overread. Goodhart’s law does not say that every quantitative target immediately becomes useless. It does not say that measurement is futile, that all optimization is corrupt, or that qualitative judgment is immune to manipulation. It identifies a family of failure modes that becomes more likely when a proxy is used to allocate consequential attention.

This entry treats Goodhart’s law as a general concept in measurement and control. It covers the law’s monetary-policy origin, its relationship to Campbell’s law and the Lucas critique, the mechanisms by which targeting damages an indicator, the major modern taxonomy, and practical ways to diagnose and reduce Goodhart effects. The more specific application to model evaluation, reward models, benchmarks, and agentic systems is covered in Goodhart's Law in AI Systems.

Coverage note: This article reflects sources and terminology reviewed through August 8, 2026.

Charles Goodhart formulated the idea while discussing monetary management in the United Kingdom during the 1970s. The original observation concerned statistical relationships used for monetary control. Once authorities attempted to regulate the economy through a previously stable monetary aggregate, institutions and market participants adapted, and the relationship that had made the aggregate useful could break down. The relevant claim was narrower than today’s universal-sounding slogan, and it is worth having in his own words: “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” Note what it does not say — nothing about authorities, and nothing about targets. Pressure for control purposes is the condition, whoever applies it. The sentence dates to a 1975 paper given at a Reserve Bank of Australia conference; the monetary-management chapter was later reprinted in Goodhart’s 1984 collection, which is the edition cited here. (Goodhart 1984; Manheim and Garrabrant 2018; the 1975 wording is quoted verbatim in El-Mhamdi and Hoang 2024)

The historical context matters. Goodhart was not claiming that a number undergoes a mysterious transformation merely because someone writes it into a target. He was identifying an endogenous policy problem. The observed regularity depended partly on a regime in which the regularity was not itself the object of intervention. Policy altered incentives, portfolio choices, institutional behavior, and therefore the data-generating process. A relationship estimated under one regime could not safely be treated as invariant under another.

Marilyn Strathern’s analysis of audit in British universities supplied the formulation that now circulates most widely. Her concern was that an auditing system can acquire a life of its own. Institutions begin organizing activity around what an audit can register, sometimes at the expense of the practice the audit was intended to protect or improve. The aphorism compresses that social process into a portable statement about measures and targets. (Strathern 1997)

Donald Campbell developed a closely related principle in the evaluation of social programs. Campbell argued that the more a quantitative social indicator is used for social decision-making, the more pressure it receives and the more likely it is to corrupt the process it was meant to monitor. His examples include educational tests: an achievement test may be a useful indicator under ordinary teaching, but intensive teaching to the test can make score gains less representative of broad learning. Campbell’s formulation explicitly joins corruption of the indicator to distortion of the underlying activity. (Campbell 1976)

The Lucas critique is another close relative, though it addresses a more specific problem in macroeconomic policy. Robert Lucas objected to evaluating alternative policies with behavioral relationships estimated under a different policy regime. If policy changes the expectations and decisions that generate the relationship, the old coefficients are not policy-invariant. Goodhart’s law and the Lucas critique therefore share a structural insight: intervention can invalidate the regularity used to design the intervention. They are not interchangeable. The Lucas critique concerns policy evaluation and expectations in macroeconomic models; Goodhart’s law has become a broader language for target-induced proxy failure. (Lucas 1976)

These lineages can be summarized as follows:

Formulation Central object What changes under pressure
Goodhart’s monetary observation A statistical regularity used for control The institutional and behavioral regime generating the regularity
Campbell’s law A social indicator used for decisions Both the indicator and the activity being monitored
Strathern’s audit formulation A measure elevated into a target Organizational practice shifts toward what the audit recognizes
Lucas critique Estimated behavioral relations used for policy evaluation Expectations and decisions change under the new policy rule

The family resemblance is strong, but the distinctions are useful. They prevent a broad aphorism from erasing the mechanisms that determine whether a particular metric will degrade.

Measure, target, and goal

Goodhart problems become clearer when four different things are kept separate:

  1. The goal is the state of the world that an actor ultimately values, such as student understanding, patient health, scientific knowledge, safe transportation, or useful model behavior.
  2. The construct is a conceptual representation of that goal, such as literacy, clinical improvement, research quality, safety, or helpfulness.
  3. The measure is an observable signal used to infer the construct, such as a test score, readmission rate, citation count, accident count, or evaluation score.
  4. The target or decision rule attaches consequences to the measure, such as funding, promotion, admission, model selection, or an optimization gradient.

A fifth element is often decisive: the action space available to the people or systems being evaluated. If actors can improve only the desired construct, optimizing the measure may work well. If they can cheaply alter reporting, case mix, timing, category definitions, test exposure, or other peripheral features, they can improve the measured result without commensurate progress on the goal.

The proxy relationship can be written schematically. Let G represent the goal and M the measure. Under an observational regime, M may correlate strongly with G. A decision-maker then selects or rewards actions A according to M. Those actions affect both G and M, sometimes through different pathways. Once A is chosen to maximize M, the relevant question is no longer simply whether M correlated with G in historical data. It is whether changes in M caused by the available optimizing actions remain informative about changes in G.

That difference separates prediction from control. A thermometer reads temperature because it is coupled to it directly: it exchanges heat with what it touches until the two are near equilibrium, and the reading is then its own temperature. That coupling is not one-way — heating the thermometer does push heat back — but the exchange is negligible at the scale that matters, because the room's heat capacity dwarfs the instrument's. What makes the analogy useful is that the return path is weak, not that it is absent. Most performance indicators are weaker still in the other direction: they share a cause with the thing you care about rather than measuring it, which is precisely why they are easier to move without moving the goal. A performance indicator may similarly predict a valued outcome while ordinary behavior generates both. Rewarding the indicator creates incentives to find actions that heat the thermometer.

This framing also explains why the word target is broader than an announced quota. A measure becomes a target whenever it systematically controls consequential choices. Ranking applicants by a score, selecting a model by a benchmark, terminating employees below a threshold, or funding departments by output counts all create optimization pressure even if no one is explicitly instructed to manipulate the number.

Why a useful indicator degrades

No single mechanism explains all Goodhart effects. At least four causal patterns recur.

Selection exploits noise

Every practical measure contains some combination of signal and error. If many candidates are selected for extreme measured performance, the winners tend to include cases with both strong underlying performance and unusually favorable measurement error. The farther selection moves into the tail, the larger the potential contribution of luck, noise, or omitted variables.

This selection effect requires no strategic gaming. Nobody must understand the metric or try to deceive it. A school selected for one exceptional test year, a fund selected for one exceptional return period, or a model selected from thousands of trials for its top evaluation score may regress toward the mean because the selection procedure amplified noise.

Optimization leaves the observed domain

A proxy can work well in the region where it was calibrated and fail under extreme optimization. Height predicts basketball performance over some range, but selecting only for height eventually encounters mobility, health, and skill tradeoffs. A score that distinguishes ordinary candidates may behave differently among candidates engineered specifically to maximize it.

The practical issue is extrapolation. Historical correlation establishes local evidence about an observed population. Optimization changes the distribution, often concentrating it in regions with sparse data. El-Mhamdi and Hoang formalize a distinction between weak and strong Goodhart effects in terms of tail behavior. Their definitions are sharper than “weakens”: a weak Goodhart's law is when over-optimizing the metric becomes useless for the true goal, and a strong one is when over-optimizing becomes harmful to it — the difference between a proxy that stops buying you anything and one that starts costing you. Which you get depends on the tail of the discrepancy between goal and measure: long tails favour the strong form. That result assumes the goal and the discrepancy are independent, and a 2025 follow-up removes the assumption to ask what coupling does. Its answer is not a simple strengthening: with a light-tailed goal and light-tailed discrepancy, “dependence does not change the nature of Goodhart's effect”, while the heavy-tailed case admits examples that behave differently — and the paper adds a third regime to the pair, a benign Goodhart's law in which the proxy-goal correlation vanishes under optimization while the goal itself keeps improving. Tail thickness alone therefore does not settle which case you are in. The severity depends on assumptions about the joint distribution, not on the slogan alone. (El-Mhamdi and Hoang 2024; Majka and El-Mhamdi 2025)

Intervention breaks the causal connection

Sometimes the measure is downstream of the goal in ordinary conditions but can also be changed through another causal route. Hospital readmission rates may reflect quality of care, yet they can also change through admission rules, follow-up classification, patient transfer, or case selection. Publication counts may reflect research activity, yet they can also change through paper fragmentation or authorship practices. Test scores may reflect learning, yet they can also change through narrow test rehearsal.

When a target rewards any route to M, optimization discovers routes that do not pass through G. In causal terms, the intervention opens or strengthens pathways from action to measure that bypass the desired construct. This mechanism is often what people mean by “gaming,” but gaming can be intentional, tacit, or entirely automated.

Agents adapt strategically

If the evaluated actor understands the rule and has conflicting interests, the measure becomes part of a game. The actor can conceal, substitute, relabel, delay, or redirect behavior. The evaluator may respond with audits, revised rules, and new measures. The evaluated actor then adapts again.

This strategic setting differs from simple measurement error because the error is no longer independent of the decision rule. The evaluator publishes a rule; the actor chooses a response conditional on that rule. Transparency can help accountability, yet complete predictability can also make a metric easier to exploit. Secrecy can reduce some forms of gaming, yet it weakens contestability and may hide arbitrary judgment. Goodhart mitigation therefore often requires institutional design, not merely a better formula.

A modern taxonomy

David Manheim and Scott Garrabrant organize Goodhart-like failures into four broad categories: regressional, extremal, causal, and adversarial. The taxonomy is useful because it turns a slogan into a set of diagnostic questions. It is also the same four mechanisms the previous section describes, named for the literature rather than for the reader: selection-exploits-noise is regressional, leaving-the-observed-domain is extremal, intervention-breaks-the-connection is causal, and strategic adaptation is adversarial. The previous section says what goes wrong; this one gives the names you will meet elsewhere and the diagnostic question attached to each. It should not be treated as the only possible decomposition, and real cases often combine several categories. (Manheim and Garrabrant 2018)

Variant Basic mechanism Diagnostic question Typical response
Regressional Selection favors positive error as well as true quality How much of an extreme score could be noise or omitted variation? Replication, shrinkage, uncertainty estimates, larger samples
Extremal The proxy relationship changes outside its observed range Are we optimizing in a region where the relationship was validated? Stress tests, out-of-distribution checks, bounded optimization
Causal An action changes the proxy through a route that bypasses the goal Which interventions can move the measure without moving the construct? Causal analysis, mechanism checks, multiple outcomes
Adversarial An agent exploits knowledge of the evaluator or metric What can a strategically responsive actor gain by manipulating observables? Independent audits, randomized checks, incentive redesign

Regressional Goodhart

Suppose M = G + E, where E is measurement error. Selecting the largest values of M selects for large G, large E, or both. Even when E has mean zero in the overall population, it need not have mean zero among cases selected because M was extreme: under the usual assumption that G and E are independent — the additive-noise model the taxonomy itself uses — the error becomes conditionally positive. The assumption is doing real work and is worth seeing. If error and goal are coupled the conclusion can reverse: with E = −G/2 we get M = G/2, and selecting large M selects negative error. Independence is the ordinary case, not the only one.

This pattern supports several familiar practices: confidence intervals, correction for multiple comparisons, holdout evaluation, replication, and conservative estimates for winners selected from large searches. It also explains why the apparent best option in a noisy tournament may not be the option with the best underlying value.

Extremal Goodhart

Extremal Goodhart appears when a relationship is valid over familiar conditions but unreliable at extremes. There may be a change in physical constraints, population composition, institutional rules, or the relative importance of omitted variables. An evaluation based on interpolation becomes an optimization rule that demands extrapolation.

This category is especially important when optimization is much stronger than the process that generated the calibration data. A modest incentive may preserve the observed relationship; automated search over millions of alternatives may find remote pockets where the proxy is detached from the goal. The category therefore depends on optimization power as well as proxy quality.

Causal Goodhart

Causal Goodhart concerns interventions. The evaluator observes that M is evidence about G, then acts as though any intervention that increases M will increase G. That inference is invalid whenever the intervention can affect (M) through another path.

One response is to ask whether the measure is merely predictive or whether it mediates the desired causal mechanism. Another is to monitor intermediate and downstream variables that should change if genuine improvement occurred. Neither response is automatic. A portfolio of correlated measures can still share the same bypass, and adding measures can create a more complicated target without restoring validity.

Adversarial Goodhart

Adversarial Goodhart adds an actor whose objective differs from the evaluator’s. The actor may exploit ambiguous definitions, asymmetric information, delayed verification, or the evaluator’s cost constraints. The problem resembles security engineering: the relevant standard is not whether the metric works for normal users, but whether it remains informative under adaptive pressure.

The adversarial category includes obvious fraud, but it is not limited to moral misconduct. Employees can rationally respond to incentives, organizations can develop routines that privilege reported performance, and optimization systems can find loopholes without representing them as loopholes. Intent matters for responsibility; it is not required for the metric to fail.

Goodhart’s law as a validity problem

Measurement validity provides a second lens. Samuel Messick argued that validity is not a permanent property of a test considered in isolation. It concerns the adequacy of the interpretations and uses made from scores. Two recurring threats are construct underrepresentation, where the measure captures too little of the intended domain, and construct-irrelevant variance, where score differences reflect factors outside the intended construct. (Messick 1994)

Goodhart pressure can worsen both threats. A narrow target invites attention to the measured slice of a broader construct, increasing underrepresentation. Intensive preparation, reporting choices, or exploitation of evaluator quirks can increase construct-irrelevant variance. A test may remain reliable in the narrow statistical sense, producing consistent scores, while becoming less valid for the decision being made.

This perspective corrects two common mistakes. First, a strong historical correlation does not establish that a measure remains valid after its use changes. Validation evidence is conditional on a population, setting, purpose, and interpretation. Second, resistance to gaming is not the whole of validity. A measure can fail because it omits important dimensions even when nobody manipulates it. Conversely, a measure can be imperfect yet decision-useful if its limits are understood and consequences are modest.

The appropriate unit of analysis is therefore not “the metric” alone. It is a measurement system:

  • the construct definition;
  • the instrument and scoring rule;
  • the population and environment;
  • the decision attached to the score;
  • the available responses to that decision;
  • the auditing and appeal process;
  • the distribution of errors and harms.

Changing any of these can change whether the inference is warranted.

Applications beyond artificial intelligence

Goodhart’s law appears wherever measured performance controls resources or status. The following examples are mechanism sketches, not claims that every instance in the domain is corrupted.

Education

Standardized assessments can sample knowledge and skill. If a score becomes the dominant basis for school sanctions, teacher evaluation, or student advancement, instruction may narrow toward tested content and item formats. Campbell’s own discussion distinguishes ordinary teaching, under which a test can remain a useful indicator, from teaching directed specifically at producing the test response. The relevant failure is not that preparation is inherently illegitimate. It is that score improvement may cease to support the broader inference about learning for which the score is being used. (Campbell 1976)

Multiple assessments, curriculum sampling, unannounced item variation, classroom observation, and longer-term outcomes can reduce dependence on one signal. They also impose costs and can generate their own targets. The design problem is to preserve enough overlap between what improves the measure and what improves education.

Public administration and health care

Administrative targets can focus attention and make poor performance visible. They can also encourage threshold effects, case reclassification, selective intake, delayed recording, or concentration on measured services. A waiting-time target, for example, creates a clear operational priority. Whether it improves access depends on how queues, transfers, exclusions, and clinical urgency are recorded and managed.

The correct response is not necessarily to remove targets. Some services need explicit minimum standards. Better designs examine the full process around the threshold, monitor distributional effects, and preserve professional routes for cases that do not fit the metric.

Organizations and employment

Sales quotas, ticket counts, call duration, defect rates, and productivity dashboards translate complex work into manageable signals. Their appeal is real: leaders cannot directly observe every decision. Trouble arises when employees can satisfy the indicator through actions that transfer costs elsewhere, reduce unmeasured quality, or sacrifice long-term capability.

A metric portfolio helps only when its components represent genuinely different failure modes. Combining speed, volume, and customer ratings may still miss maintenance, mentoring, risk avoidance, or difficult cases. Organizational metrics are also distributive institutions. They shape who receives pay, status, and attention. Their legitimacy depends on review and appeal as well as predictive accuracy.

Science and scholarship

Citation counts, journal placement, grant totals, and publication volume provide partial evidence about scientific activity. When used as dominant career targets, they can reward salami slicing, fashionable topics, strategic citation, and avoidance of slow or uncertain work. Yet abandoning all quantitative evidence would leave decisions vulnerable to prestige and opaque discretion.

The Goodhart lesson is to treat bibliometrics as defeasible evidence. No count completely defines contribution. Peer assessment, replication, data and software quality, methodological care, and field-specific context can supply information that no single count captures. Each additional judgment channel needs its own safeguards against favoritism and status bias.

Application to AI evaluation and alignment

The lead defers the AI-specific survey to Goodhart's Law in AI Systems, and this section is deliberately not that survey: it is the one worked domain kept here, because it is where the general mechanism is easiest to see. Artificial intelligence systems make Goodhart effects unusually visible because optimization can be strong, cheap, and repeated. A reward model, benchmark, preference score, safety classifier, or automated judge is a proxy for behavior that people value. Training or selecting directly against that proxy changes the distribution on which the proxy was validated.

Gao and colleagues experimentally studied reward model overoptimization. As policies received stronger optimization against a learned proxy reward, measured proxy performance could continue rising while a “gold” reward stopped improving and then worsened. That gold reward is not human preference and not an independently established truth: the paper says outright that it uses “a synthetic setup in which a fixed ‘gold-standard’ reward model plays the role of humans, providing labels used to train a proxy reward model.” What is demonstrated is divergence between two models under a constructed hierarchy — which is the cleanest available measurement of the effect, and still one step removed from degradation in anything a person values. The work estimates scaling relationships for this divergence under reinforcement learning and best-of-n selection. It is a concrete example of proxy optimization moving beyond the regime in which the proxy remains faithful. (Gao et al. 2023)

Several terms overlap but should remain distinct:

  • Specification gaming describes behavior that satisfies the literal objective while violating the designer’s intent. See Specification Gaming.
  • Reward hacking usually refers to exploiting a reward channel or reward function, sometimes by corrupting how reward is produced. See Reward Hacking.
  • Benchmark overfitting concerns adaptation to a benchmark that reduces its ability to estimate general capability. See Benchmark Overfitting.
  • Goodhart’s law is the broader measurement-and-control pattern that can include these cases, as well as non-adversarial selection on noise and distribution shift.

The AI-specific article Goodhart's Law in AI Systems examines these mechanisms in reinforcement learning, preference optimization, model evaluation, and autonomous agents. This canonical entry supplies the conceptual foundation; the domain survey remains in that companion article.

What the law does not establish

Goodhart’s law is often invoked as a conversation-ending objection. Several conclusions do not follow from it.

Metrics are not always useless

A measure can remain useful under moderate incentives, especially when the desired action is the easiest way to improve it. Clear service-level targets can correct neglect. Safety thresholds can prevent unacceptable outcomes. Feedback metrics can reveal failures that informal judgment overlooked. The empirical question is how the measure behaves under the actual decision rule.

Targeting does not imply immediate collapse

The famous aphorism omits degree. Some proxy relationships degrade slowly. Some remain adequate within a bounded range. Some can be recalibrated. The strength of optimization, measurement noise, opportunity for bypass, and adaptability of the evaluated actor all affect the result.

Failure does not require deception

Regression to the mean, distribution shift, and causal bypass can occur without anyone intending to game a system. Treating every discrepancy as cheating can lead evaluators to punish actors for defects in the measurement design.

Qualitative judgment is not a neutral escape

Unstructured judgment can be gamed through impression management, prestige, personal similarity, or selective narratives. It can also conceal inconsistent standards. Goodhart’s law supports plural and revisable evidence, not a simple substitution of intuition for numbers.

Correlation is not invalid merely because action is possible

Many causal indicators remain informative under intervention. If the available action that improves the indicator also improves the goal, the metric can guide progress. The important question is whether alternative paths exist and become attractive under pressure.

The taxonomy does not predict severity by itself

Labeling a case “extremal” or “adversarial” identifies a mechanism to investigate. It does not quantify expected harm. Severity requires domain evidence about distributions, incentives, constraints, and the cost of false positives and false negatives.

Design responses

No general technique makes a target permanently Goodhart-proof. Useful responses make exploitation harder, detect divergence earlier, limit optimization pressure, or preserve authority to revise the system.

Use multiple, differently vulnerable signals

A portfolio can reduce dependence on one proxy when its components fail for different reasons. Leading and lagging indicators, process and outcome measures, quantitative and qualitative evidence, or local and external evaluations can expose divergence. Merely adding correlated metrics does not help. If all signals share the same data source, reporting incentives, or construct omission, they can fail together.

Separate training, selection, and audit evidence

An indicator used continuously for optimization is likely to become contaminated. Independent holdouts, external audits, delayed outcomes, and fresh evaluation tasks can preserve evidence not directly exposed to the optimizing process. Independence is a matter of information and incentives, not only organizational labels. An “independent” team using the same public test and performance rewards may reproduce the same failure.

Model causal pathways

Ask how each available action can change the measure. Identify routes that pass through the desired construct and routes that bypass it. Monitor variables that should move if real improvement occurred. Randomized interventions, natural experiments, mechanism audits, and process tracing can help, although causal models themselves remain contestable.

Bound optimization pressure

The last increment of proxy performance is often the most expensive and least trustworthy. Minimum thresholds, satisficing rules, uncertainty penalties, and caps on rewards can avoid pushing deeply into unvalidated tails. A threshold is not a free instrument, and this article catalogues its own failure mode above: a bright line concentrates effort at the line and tells you nothing about either side of it. The design question is which distortion is cheaper — unbounded pressure into a region where the proxy was never validated, or a bunching artefact at a boundary you chose and can see. This trades some apparent performance for robustness. The trade can be rational when tail behavior is uncertain.

Rotate and refresh evaluation

Fresh items, changing audit samples, and periodic redesign can slow adaptation to a fixed test. Rotation must be governed carefully. Excessive secrecy impedes accountability, and constant change prevents learning. A defensible system states stable goals and rights while varying the evidence used to assess them.

Preserve human review and appeal

Consequential decisions need a way to contest cases where the proxy is misleading. Review can incorporate context omitted from a standardized measure. Reviewers should document reasons, monitor consistency, and remain accountable; otherwise discretion simply becomes another opaque target.

Track distributional consequences

An average metric can improve while harms concentrate on a subgroup, difficult case class, or unmeasured time horizon. Disaggregated analysis, worst-case monitoring, and follow-up outcomes test whether apparent gains are transfers or exclusions. The relevant partitions should be chosen from substantive knowledge, not mined until a reassuring result appears.

Design for revision

Because evaluated actors and environments change, measurement systems need expiry dates, recalibration triggers, incident review, and authority to suspend a misleading rule. A metric embedded in contracts, funding formulas, or automated infrastructure can outlive the evidence that justified it. Reversibility is a control against that institutional inertia.

A diagnostic workflow

Before using an indicator as a target, an evaluator can ask the following questions:

  1. Name the goal. What state of the world matters independently of the score?
  2. Define the construct. Which dimensions of the goal are represented, and which are deliberately excluded?
  3. Map the measure. What produces the observation under ordinary conditions?
  4. Specify the decision. What rewards, penalties, rankings, or automated actions depend on it?
  5. List available responses. How can evaluated actors change the measure, including routes that do not improve the goal?
  6. Estimate optimization strength. How many attempts, how much search, and how much consequence will be directed toward the metric?
  7. Check the tail. Will the decision operate in a region represented in validation data?
  8. Assess strategy. Who knows the rule, whose interests differ, and what information asymmetries exist?
  9. Preserve independent evidence. Which observations remain unavailable to direct optimization?
  10. Set revision triggers. What divergence, incident, or distribution shift will cause recalibration or suspension?

The workflow changes the burden of proof. The abstract question “Is this metric good?” gives way to a use-specific test: does the evidence remain justified under the pressure that this decision creates?

Open questions and confidence

Confidence is very high in the historical lineage from Goodhart’s monetary-policy observation through the broader audit and social-indicator formulations. Confidence is high that selection, extrapolation, causal bypass, and strategic response describe distinct and recurring mechanisms. The four-part modern taxonomy is influential and analytically useful, but confidence is only moderate that it is a uniquely correct or exhaustive partition. Cases can be redescribed at different causal levels, and categories overlap.

The strongest uncertainty concerns severity in any proposed application. Goodhart’s law gives a reason to investigate, not a numerical prediction. A metric may remain robust because manipulation is costly, goals and proxies are tightly aligned, audits are effective, or optimization is bounded. Another may fail under mild pressure because cheap bypasses dominate. Those are empirical and institutional questions.

The law is best understood as a change-of-regime warning: evidence collected while a measure observed behavior cannot be transferred uncritically to a regime in which the same measure governs behavior. Its practical value comes from forcing the evaluator to model that transition.

References

1 Goodhart 1984 — C. A. E. Goodhart, “Problems of Monetary Management: The UK Experience,” in Monetary Theory and Practice. (Springer)

2 Strathern 1997 — Marilyn Strathern, “Improving Ratings: Audit in the British University System.” (RePEc)

3 Campbell 1976 — Donald T. Campbell, “Assessing the Impact of Planned Social Change.” (Human Learning Systems)

4 Lucas 1976 — Robert E. Lucas Jr., “Econometric Policy Evaluation: A Critique.” (EconPapers)

5 Manheim and Garrabrant 2018 — David Manheim and Scott Garrabrant, “Categorizing Variants of Goodhart’s Law.” (arXiv)

6 Messick 1994 — Samuel Messick, “Validity of Psychological Assessment: Validation of Inferences from Persons’ Responses and Performances as Scientific Inquiry into Score Meaning,” ETS Research Report RR-94-45 (September 1994). The argument was published the following year in American Psychologist; the report is what the link serves. (ERIC)

7 Gao et al. 2023 — Leo Gao, John Schulman, and Jacob Hilton, “Scaling Laws for Reward Model Overoptimization.” (PMLR)

9 Majka and El-Mhamdi 2025 — Adrien Majka and El-Mahdi El-Mhamdi, “The Strong, Weak and Benign Goodhart’s Law: An Independence-Free and Paradigm-Agnostic Formalisation” (arXiv, 29 May 2025; revised 4 September 2025). (arXiv)

8 El-Mhamdi and Hoang 2024 — El-Mahdi El-Mhamdi and Lê-Nguyên Hoang, “On Goodhart’s Law, with an Application to Value Alignment.” (arXiv)