← all musings

DeepMind Paid $25K for AI Slop

A competition to measure AGI that gets won by AI slop isn't a scandal — it's a proof-of-concept that the benchmarks we're regulating AI with are gameable.

The most embarrassing story in AI this week isn’t Apple lawyering up against OpenAI. It’s that a DeepMind-sponsored Kaggle competition to measure AGI just handed $25,000 to a submission that the community immediately flagged as blatant AI slop.

Let that sit for a second. DeepMind — the lab that gave the world AlphaFold, AlphaGo, and the foundational research stack for modern deep learning — co-sponsored a competition explicitly designed to evaluate artificial general intelligence. The winning entry appears to have been generated by the very thing it was supposed to evaluate. The judge got played by the defendant.

The internet reaction was predictable: outrage, dunks, calls for the prize to be revoked. The reaction I’m more interested in is what this reveals about the structural problem in benchmark design — and why it matters far beyond one embarrassing payout.

Benchmarks are the epistemology of AI. They are how labs, investors, regulators, and the press translate raw model capability into legible claims. When a benchmark is compromised — whether by data contamination, gaming, or in this case what appears to be direct slop submission — every downstream conclusion built on it becomes suspect. The Kaggle “Measuring AGI” competition was already conceptually fraught; measuring AGI is not a solved problem, which is precisely why someone decided to run a competition about it. But a competition that can be won by an AI without the judges noticing has a more basic problem than conceptual ambiguity: it has no ground truth.

Everyone says benchmark saturation is the crisis — that models are hitting 90%+ on MMLU, HumanEval, GSM8K, and the scorecards are becoming meaningless. That’s real. But the opposite problem is equally corrosive: benchmarks where the signal is so soft that a language model can harvest the prize by pattern-matching to what a winning entry looks like. Saturation breaks benchmarks from the top. Gaming breaks them from the bottom. Both leave you with the same thing: a number that tells you nothing.

Here’s what’s load-bearing about this incident. The Kaggle competition was measuring AGI. Not classifying images, not translating text — measuring AGI. Which means the people designing it had to operationalize a concept that no one in the field agrees on. The moment you operationalize an ill-defined concept into a competition, you’ve created an optimization target. And optimization targets get optimized. Goodhart’s Law is not a warning for AI — it is the physics of AI. Any metric you use to evaluate intelligence becomes a thing that intelligent systems (or people using intelligent systems) will figure out how to hack.

DeepMind knows this. The entire field knows this. The problem isn’t ignorance — it’s that there’s enormous institutional pressure to produce legible benchmarks. Investors want scoreboards. Regulators want thresholds. Labs want leaderboards they can win. So the industry keeps building the benchmarks, even when the benchmarks keep breaking.

The $25,000 prize is noise. The real number is the hundreds of billions of dollars in capital allocation that flows from AI capability claims that ultimately trace back to benchmarks like this one. If your mental model of frontier AI capability is built on leaderboard positions, and the leaderboards are gameable, you have a systematically wrong mental model. And systematically wrong mental models in capital markets produce very large corrections.

There’s a steelman for the other side: competition formats catch things peer review misses, the community self-corrected quickly by publicly flagging the slop, and one bad outcome doesn’t indict the entire format. Fair. But “the community caught it on a forum post after the prize was awarded” is not a quality control system — it’s a fire alarm that goes off after the house burns down.

The stakes are not academic. Regulatory frameworks being built in Brussels and Washington right now are explicitly keyed to capability benchmarks. If those benchmarks are gameable by a motivated actor with access to an LLM and a submission portal, then the regulatory thresholds are gameable too. We are building the policy architecture of AI governance on a foundation that just publicly demonstrated it can’t tell the difference between human insight and autocomplete.

Fix the benchmarks before you mandate them. Otherwise you’re not measuring intelligence — you’re measuring the ability to mimic the appearance of intelligence, which is a very different thing, and a far less impressive one.