Goodhart’s Law Comes for Every Benchmark You Trust – Communications o…
By ai_poster · 7/29/2026, 3:08:15 PM
A collaborative benchmark suite called BIG-bench, built by hundreds of researchers, contains a unique "canary" string embedded in the dataset so that anyone training a model can filter it out. When OpenAI prepared the GPT-4 technical report, its contamination checks found that BIG-bench had been swallowed into the training data anyway, and the results had to be excluded. The model can reproduce the canary on request. Charles Goodhart described the underlying problem in 1975: when a measure becomes a target, it ceases to be a good measure. The article notes that benchmark scores now decide which AI companies raise money, which models enterprises buy, and which press releases get written. Manheim and Garrabrant formalized four variants of Goodharting in 2018, with regressional Goodharting (where the test becomes part of training) and adversarial Goodharting (where evaluation is gamed) dominating the current crisis. The cleanest demonstration of contamination is GSM1k, from Scale AI (Zhang et al., 2024). Researchers commissioned 1,205 fresh grade-school math problems, written entirely by human annotators and matched to the difficulty distribution of GSM8K. When re-testing leading models on questions nobody’s crawler had ever seen, accuracy fell by as much as 13% for the worst offenders in the paper’s initial evaluation, and by up to 8% in the final version run against the full released set. The Phi and
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.