Goodhart’s Law in data and AI describes a moment every data-driven organization walks into sooner or later: the moment a metric stops describing reality and starts producing it. That is the weakness this part covers, and it is a peculiar one, because it grows out of doing the right thing. Part 1 looked at organization and collaboration, where the problems at least announce themselves; this one starts where everybody believes the numbers.

The irony is obvious. Basing decisions on data requires metrics, and taking metrics seriously means tying them to targets and consequences. That is the exact moment when the metric starts to change, and often the system it was only supposed to observe changes with it. The five models in this part describe how that happens and what can be done about it.

Goodhart’s Law

I have watched this happen in more than one organization, always with a data quality programme that turned into a league table. Scores per domain, published every month, ranked top to bottom, and the ranking was the part everyone looked at. What followed was not better data. It was better positioning. Teams worked out which checks moved the score, added more of those, and were in no hurry to add the ones that might have exposed something. A number introduced to surface problems had been converted into an instrument for appearing not to have any.

Charles Goodhart described the mechanism in the 1970s, watching monetary policy, and it was later condensed into a rule: once a metric becomes a target, it stops being a good metric. A useful metric is a proxy for something that matters and cannot be measured directly, so the moment people are rewarded for the proxy, the proxy is what improves.

The pattern repeats wherever a proxy is easier to move than the thing it stands for. Make dashboard usage a target and view counts rise, because dashboards get set as browser home pages.

The effect is currently clearest in the evaluation of language models. Use a second model as an evaluator, the so-called LLM-as-a-judge, and elevate its score to an optimization target, and you get answers that are polite, thorough and neatly structured, because evaluators respond reliably to exactly those properties. The problem is not the method, which is a reasonable and often necessary way to evaluate at scale. The problem is treating the evaluator’s score as ground truth rather than as one more proxy that can be optimized against.

The public leaderboards show what that looks like at scale. Researchers found in 2025 that a model provider could quietly test dozens of private variants against a ranking and publish only the best result. In one case a build tuned for chatty, agreeable answers took second place and was announced as such, while the model actually released placed somewhere around thirty-second when tested (The Leaderboard Illusion, 2025). The ranking had stopped measuring models and started measuring submissions. The same holds for automated resolution rates in support: an assistant measured on how many conversations end without escalation primarily learns to avoid escalating.

The answer is not to abandon metrics but to separate their roles. A metric can serve as an observation without becoming a target. A metric becomes dangerous at a specific point, and it is not the point where it enters a dashboard. It is the point where it starts influencing careers. Where a metric genuinely works as a target, it therefore helps to pair it with a counter-metric that prevents it from being optimized in isolation. Usage frequency combined with a qualitative judgement of value works considerably better than usage frequency alone.

A sharper variant of this pattern is the cobra effect, named after an anecdote from colonial India. A bounty on dead cobras led people to breed cobras, and when the bounty was withdrawn the now worthless animals were released. The story is historically poorly documented, but the pattern is thoroughly familiar. Where Goodhart’s Law describes a metric losing its meaning, the cobra effect describes the situation ending up worse than before. If reporting data quality issues is officially encouraged while their occurrence is evaluated negatively, what disappears is not the issues but the reports. Before defining any target it is therefore worth asking how you would satisfy this metric if you had neither time for nor interest in the underlying goal. The answer describes fairly reliably what is going to happen.

The McNamara Fallacy

Robert McNamara, US Secretary of Defense during the Vietnam War, steered the war effort using quantitative metrics, most notoriously enemy body counts. Anything that could not be expressed numerically, such as political legitimacy or the attitude of the population, never entered the assessment. The fallacy named after him proceeds in four steps: measure what can be measured; disregard what cannot be measured; presume that what cannot be measured is unimportant; conclude that it does not exist.

This is the fallacy with the closest affinity to data-driven organizations, because it arises not from carelessness but from consistency. Establishing the rule that decisions must be based on data implicitly excludes everything that is not in the data warehouse. Yet some of the most important variables live exactly there: trust in the numbers, technical debt, team morale, the quality of the customer relationship, and the question of whether a data product actually improved a decision or merely looked good.

GenAI initiatives currently offer the clearest illustration. What gets evaluated is what can be measured immediately, meaning latency, token cost and results on standard benchmarks. Whether the business unit trusts the answers, what risk a fluently phrased hallucination carries in customer contact, and what reputational damage a failure would cause appears in no evaluation at all. Yet those are the variables that decide whether the initiative succeeds, not the benchmark score. The same thing happens one level down in clinical AI: a therapy recommendation computed only from what the sensors record can be arithmetically correct and clinically wrong, because a patient’s circumstances never enter the calculation and are therefore treated as though they were not part of it.

The remedy is unspectacular. It is enough to include an explicit section in decision documents for what was not measured but is relevant. The mere existence of that section prevents the unmeasurable from quietly becoming the unimportant.

The Streetlight Effect

A drunk is searching for his keys under a street lamp at night. Asked whether he lost them there, he answers no, but the light is better here. The streetlight effect describes the tendency to search where searching is easy rather than where the answer lies.

In analytics this effect is ubiquitous, because data availability varies enormously. Clickstream data is abundant, so on-site behaviour gets analysed thoroughly. The question of why customers leave is rarely answered, because churned customers do not fill in surveys and leave no click paths.

Data quality work shows the same asymmetry with unusual clarity. What gets measured is what a machine can count on its own: null rates, primary key violations, uniqueness constraints, referential integrity. What does not get measured is whether the address in the record is the one the customer actually lives at, whether the revenue on this row belongs to this contract, whether the product hierarchy still matches the way the business sells. Those are the failures that cost money, and every one of them needs a person who knows the domain to look. So the dashboard fills with green technical checks while the expensive errors pass straight through, and the programme reports progress the whole time. A/B testing shows the same asymmetry: what gets tested is what is easy to test, meaning colours, copy and placement, while the questions with the largest leverage such as pricing, positioning or product scope remain untested. The most expensive demonstration was Google Flu Trends, a system built to estimate the share of doctor visits caused by flu by watching what people searched for. By 2013 it was reporting more than twice the figure that the public health authorities recorded, and it had been built to reproduce exactly those figures (Lazer et al., Science 2014). Search queries were not the right data. They were the available data.

Language model evaluation repeats the pattern in pure form: what gets measured is what can be checked automatically, meaning tasks with an unambiguous reference answer, while the question that actually matters, whether an answer helped the user get their work done, requires costly manual judgement and therefore does not happen.

For data leads this is one of the most valuable checks available. The question is not what the data says, but what question was being asked and whether this data can answer it. Where it cannot, a small and expensive piece of primary research is often worth more than a large and cheap secondary analysis.

The Hawthorne Effect

The term goes back to studies conducted at Western Electric’s Hawthorne Works in the 1920s and 1930s. Worker productivity there appeared to improve regardless of which working conditions the researchers happened to change. The obvious interpretation was that observation alone changes behaviour. Later reanalyses of the original data have qualified that reading considerably. As a rule of thumb the core still holds: a system that is being measured behaves differently from a system that is not.

The difference from Goodhart’s Law matters. Goodhart requires a metric to be elevated to a target and someone to have an interest in optimizing it. The Hawthorne effect operates without any incentive at all, purely through attention. A Canadian hospital measured this precisely. Hand sanitizer dispensers were wired to record every use, and auditors walked the wards to check compliance. Dispensers the auditors could see were used 3.75 times an hour. Dispensers they could not see, at the same moment on the same ward, were used 1.48 times an hour. The same dispensers the week before, with no auditor in the building, managed 1.07. The rate rose only after the auditors arrived, which is why the researchers questioned whether published compliance figures mean anything at all (Srigley et al., BMJ Quality & Safety 2014). Nobody was cheating. They were simply being watched.

For data organizations this explains a particularly frustrating phenomenon: pilot projects almost always work. During a pilot, experienced people are paying close attention, data errors get corrected manually in the background, users receive personal support and edge cases surface immediately. A full rollout removes exactly that attention, and the measured effect disappears with it. This applies to dashboards as much as to ML models whose input data was hand-curated during the pilot.

Two consequences follow. First, pilot results should never be extrapolated to regular operations unadjusted, but always with a deliberate discount. Second, the more interesting question is not whether the pilot worked, but what share of the success is attributable to the solution and what share to the attention. If that question cannot be answered, the pilot was set up badly.

Regression to the Mean

In the nineteenth century Francis Galton observed that the children of unusually tall parents were on average shorter than their parents, and the children of unusually short parents taller. The reason is purely statistical. Where an outcome consists of a stable component and a random component, extreme values are disproportionately often the product of a random swing, and that swing does not repeat at the next measurement.

Kahneman’s account of Israeli flight instructors is the shortest way to see it. The instructors were certain that praising a cadet made the next flight worse and shouting at one made it better, and their observation was accurate. What they had missed is that they only praised flights far above the cadet’s average, which were mostly luck, and only shouted after flights far below it, which were mostly bad luck. Both would have moved back toward the middle if the instructor had said nothing at all.

In organizations this model matters because interventions almost always target whatever is performing worst. The weakest branch gets the new sales concept, the region with the most complaints gets the quality initiative, and the tables with the most errors get the new tests first. In all three cases the numbers will very likely improve afterwards, and they will do so even if the intervention was entirely ineffective. Conversely, the team that won an award last year often slips the following year without having done anything wrong.

The most practically relevant case is measuring the success of data quality initiatives. A table draws attention through an outlier event, say a faulty delivery from a source system, gets escalated, and consequently receives a new regime of tests and controls. The fact that the quality metrics look considerably better afterwards is reliably attributed to the new methodology. A substantial share of the improvement is simply statistical normalization after a one-off event, and that share is then budgeted as an expected benefit for the next initiative.

An improvement following an intervention is not evidence of causality. It may reflect regression to the mean, changed observation, seasonality, a selection effect, or a real effect, and the comparison design decides which of those explanations survives. The consequence is inconvenient because it costs effort: without a comparison group, the effect of an intervention targeted at extreme values simply cannot be assessed. Where a control group is organizationally impossible, the minimum is to look at how comparable units without the intervention developed over the same period. Together with the Hawthorne effect above, this model accounts for a considerable share of all reported project successes.

Takeaway

The five models in this part describe five distinct ways in which measurement goes wrong. Goodhart’s Law shows that a metric under target pressure loses its meaning and in the worst case makes matters worse. The McNamara fallacy shows that the unmeasurable quietly becomes the unimportant. The streetlight effect shows that measurement happens where it is easy rather than where the answer is. The Hawthorne effect shows that measuring itself changes the outcome. And regression to the mean shows that an improvement can arrive entirely on its own and still be booked as a success. None of these failure modes disappears with better tooling.

Concluding from this that one should measure less is the wrong lesson. The workable consequence is to attach two additional pieces of information to every metric: what it stands in for, and how it could be gamed. Those two sentences cost little and prevent a lot.

Part 3 (stay tuned) turns to data and judgement, and to the question of how a correctly built report still leads somewhere wrong.