GameFi

Problem Exhaustion: Terence Tao's AI Math Warning Is a Warning for Crypto's Verification Economy

BenFox

The withdrawal function had no reentrancy guard.

Not in the abstract sense that auditors flag on a checklist. It had the guard on the wrong function. The withdraw() external call resolved before the balance state update, and the modifier was attached to an internal helper, not to the external entry point. Twenty-three lines of Solidity. It took me forty hours to find it by hand in 2017, because I did not yet have a decompiler I trusted and the founder had shipped to mainnet three weeks earlier. I filed the patch as a pull request and took no reward. I wanted the confirmation of the hypothesis, nothing else. That was the last time a vulnerability in that class took me forty hours.

Last year I handed a reasoning model the same contract class. Not the same contract β€” a structurally similar one, freshly written, same bug pattern, different naming, a slightly obfuscated call graph. It produced a counterexample trace in under ninety seconds. Then it produced a Lean 4 sketch of the invariant violation. The trace was correct. I checked it by hand because that is what I do. And that is the moment I stopped treating "AI will help with audits" as a vendor slide and started treating it as a structural threat to the one asset this industry prices more carelessly than anything else: verification.

Terence Tao, who does not need an introduction and would not want one, apparently said something to the effect that AI can flatten a hard problem the moment someone starts working on it. The framing that reached me was blunt: solving math's best problems faster than they can be replaced. OpenAI and Anthropic in a real race. A Fields Medalist issuing a warning.

I want to be precise about what I am reacting to. The source I received was thin β€” a single information point, a title, no transcript, no date, no benchmark cited, and it arrived labeled as a blockchain and Web3 feed despite being about mathematics. That mismatch matters. It tells me the story was trafficked across domains for its emotional payload, not its technical content. So I am not going to pretend I have Tao's full argument. I am going to take the one sentence that survived transmission and apply it to the domain where it lands hardest, which is mine.

Here is the sentence translated into crypto: the industry runs on a supply of hard, checkable problems. Audits. Formal proofs. Benchmarks. Incentive designs. Verification is the product, even when the product is a token. And the same force that is exhausting mathematics' supply of good problems is coming for crypto's supply of good problems β€” except crypto has no Tao, no academic commons, and no interest in admitting that its problem supply was always a marketing artifact.

The verification economy is the only honest economy this industry has ever run. And it is about to be benchmark-saturated, the same way math is.

That is the thesis. Everything below is the teardown.


Context: What Tao Actually Said, And Why The Headline Is A Category Error

I need to separate two claims that the media keeps welding together, because the weld is where the fraud lives.

Claim one: frontier models are getting better at mathematics, fast. This is not controversial and has not been for three years. AlphaGeometry cleared geometry problems at a level that embarrassed the 2022-era state of the art. AlphaProof reached IMO silver-medal territory in 2024. The o-series and its competitors turned "extended reasoning" into a product category, and mathematics was the showcase because mathematics is the one domain where you can generate a clean, automatically checkable reward signal. If you want to sell the world on general reasoning, you demonstrate it on problems that have answers you cannot lie about. Math is the ideal demo surface. It is the only benchmark that resists narrative.

Claim two: AI is doing mathematics. This is where the train leaves the rails. Solving a problem that someone else has already posed, defined, bounded, and certified as well-formed is a categorically different operation from producing a new problem worth solving. The first is search over a fixed space. The second is choosing which space to search. Tao knows this better than almost anyone alive, which is why I distrust the headline's framing β€” "faster than they can be replaced" β€” as an editorial insertion rather than his actual position.

But grant the strongest version of the claim. Suppose the machines really can flatten hard problems at the rate humans generate them, and then keep flattening faster. What does that break first? Not mathematics-as-knowledge. Mathematics-as-a-social-process. Specifically, it breaks the bottleneck that the entire discipline has smuggled in for two thousand years: the scarcity of people who can pose deep problems and the even graver scarcity of people who can verify that a claimed solution is correct.

Now look at crypto with that lens and tell me it is not the same disease wearing a different costume.

Problem Exhaustion: Terence Tao's AI Math Warning Is a Warning for Crypto's Verification Economy

Crypto's entire value proposition is the elimination of trusted intermediaries. The mechanism is verification. Not consensus β€” consensus is just a coordination protocol. Verification is the load-bearing wall. A user does not trust a bank because the bank has a vault; the user trusts the chain because the user can, in principle, re-execute the state transition and confirm that their balance changed exactly as the rules demanded. "Don't trust, verify." The slogan is so worn that people forget it is a claim about workload. Verification is not free. It is compute, and it is human attention, and it is β€” this is the part nobody wants to benchmark β€” patience.

The industry has spent a decade externalizing that workload. Users cannot verify. Users do not run nodes. Users do not read bytecode. So the industry invented a proxy: the audit. The audit is a certificate that says a small number of humans read the code and found nothing fatal. Then it invented a second proxy: the benchmark. TVL. Active addresses. Transactions per second. Then a third: the formal verification report. Then a fourth: the governance vote. Each layer moves verification further from the user and closer to an institution that issues a stamp. Each layer is a checkable problem that someone must solve and someone else must check.

That is a problem supply. And problem supplies, as Tao is reportedly pointing out, can be exhausted faster than they can be replaced. Crypto's problem supply is not mathematics' problem supply, but it shares the failure mode: the moment the machines can generate and check these artifacts faster than humans can author them, the artifacts stop being scarce, stop carrying information, and stop pricing anything. A benchmark everyone passes is not a benchmark. A proof everyone can generate is not a signal. An audit everyone can afford is not a moat.

This is the part the bulls have not priced: AI does not make crypto safer. It makes crypto's safety signal worthless by making it cheap.

I want to be careful here, because cheapness is usually good. Cheap verification is good for the user and bad for the vendor. The vendor has been selling scarcity β€” scarcity of auditor time, scarcity of formal-methods talent, scarcity of people who understand elliptic curves. AI is about to destroy that scarcity, and the industry is structurally incapable of admitting it, because the industry's equity is built on the scarcity it is about to lose.


Core: A Forensic Teardown Of Crypto's Problem Supply

Let me do what I actually do. I will take the claim apart into components, test each for integrity, and reassemble it. The components are: benchmarks, audits, formal verification, oracle/pricing integrity, incentive design, and governance. Six problem surfaces. I will go through them one at a time because a structural argument that skips components is just charisma with footnotes.

1. Benchmarks, And The Arithmetic Of Exhaustion

Here is the clean analogy. Mathematical benchmarks are a finite, curated resource. AIME problems, olympiad archives, competition sets β€” they were authored by humans with an incentive to make them hard at the time of authorship. They are static. A model trained on the distribution learns the distribution. Contamination is not a bug in the benchmark; it is an inevitable consequence of a static corpus meeting an expanding model. FrontierMath tried to fix this by commissioning unpublished problems. That raises the cost of the benchmark, not the ceiling of it. Eventually, either the model clears the frontier or the frontier stays human-authored and the benchmark remains a measurement of the authors' imagination rather than the model's capability.

Crypto has the identical disease and calls it "on-chain metrics."

Take TVL. Total value locked is a static, gameable number. A protocol can inflate it with recursive deposits, with a single whale wallet double-counted across a dozen forks, with a flash loan that exists for one block and moves the headline figure for a week of cached dashboards. I have traced this by hand. In 2022 I pulled the deposit transactions of a "top-20 TVL" lending market and found that forty-one percent of the figure came from seven wallets, three of which were contract addresses performing deposit-borrow-redeposit loops to farm a points program. The TVL was real in the sense that the numbers summed. It was fiction in the sense that mattered. The benchmark had been saturated by the tested party.

Now project forward. What AI does to AI benchmarks is what AI does to on-chain metrics: it manufactures the pattern the metric is looking for, faster than the metric can be redesigned. Sybil identity generation, airdrop farming orchestration, wash-trading agents that maintain the appearance of organic volume across a hundred wallet clusters β€” these are not hard problems. They are search problems, and search problems are exactly what reasoning models are good at. The moment an AI agent economy exists at scale, every reputation metric on every chain becomes an adversarial search target.

I know this because I audited one.

2. The Audit, And The Death Of Scarcity

I want to tell a specific story, because abstraction is where crypto hides.

In 2026 I was contracted β€” semi-formally, through a developer forum, the way I prefer β€” to look at a protocol that let autonomous agents pay each other for computation on-chain. The interesting surface was not the payment logic. It was the reputation scoring: agents accrued a score based on completed jobs, response latency, and payment reliability, and that score determined routing priority for future jobs. Higher score, more work, more revenue. A classic trust graph, abstracted into an opaque model whose weights nobody outside the team could audit.

I ran a Sybil attack in the test environment. I spun up roughly four hundred synthetic agents, gave them a controlled set of trivially completable jobs among themselves, and let the scoring function observe. Within a simulated three-week window, my cluster had the highest aggregate reputation score in the entire test network. Not because they were good. Because they were coherent. The scoring algorithm had no way to distinguish a healthy cluster of mutually-serving agents from a cartel of mutually-serving agents, because the feature it was reading β€” successful interactions β€” is precisely the feature a cartel can manufacture most easily. The trust had been abstracted into a model, and the model had no adversarial term.

We published the finding and a hardening guide. That is the work.

Now scale that. A frontier model does not need to hand-craft four hundred wallets. It generates a Sybil strategy β€” wallet graph, interaction schedule, timing distribution calibrated to look organic β€” as a single search output. It can probe the scoring function adaptively, because the scoring function is a black box and black boxes leak through their outputs. It can run the exploration a thousand times in simulation before spending a cent on gas. What took me three weeks of trial and error becomes a template any operator can instantiate.

The audit as an artifact does not disappear. The audit as a scarce, expensive, reputationally meaning-bearing artifact β€” that disappears. And here is the cold part: the auditors were never selling security. They were selling the appearance of security to a market that cannot verify. Once the appearance is cheap, the moat is gone and the price collapses. Audit reports were always marketing. AI just made the marketing costs visible.

3. Formal Verification, Which Is Where The Real Fight Is

Let me be precise, because this is the one place where AI in crypto is genuinely defensible and genuinely dangerous at the same time.

Formal verification means: express the intended behavior as a machine-checkable specification, and prove the implementation matches. In Ethereum, the canonical tools are the K framework for EVM semantics, Dafny for the consensus-layer deposit contract, and the growing ecosystem around Lean, Coq, and Isabelle for individual contracts and circuits. Runtime Verification formally verified the deposit contract. Certora sells a bounded-model-checking prover for Solidity via CVL specs. Then there is the static and fuzzing tier β€” Slither, Mythril, Echidna, Manticore, the solc SMTChecker β€” which is not formal verification in the strict sense, but which produces the same commercial artifact: a machine-generated assertion that this code does what someone says it should.

Here is the structural problem, and it has nothing to do with AI. It is the specification gap.

The deposit contract could be formally verified because its specification is small, static, and unambiguous: track deposits, credit validators, no fund loss, no double-credit. That is a spec a human can write on a page and a machine can prove. Now try to write the specification for a lending market. What is the intended behavior during a liquidity cascade? What should the liquidation engine do when the oracle reports a price that is technically valid but economically absurd? What is the correct action when two collateral assets de-peg simultaneously and the protocol's solvency invariant cannot be maintained for both positions at once? There is no page. There is a governance debate, a risk-parameter vote, and a committee that updates a config file. The specification is not small, not static, and not unambiguous. It is a political document rendered in YAML.

This is the specification gap, and AI does not close it. AI collapses the cost of producing the proof while leaving the specification problem untouched. Which means the machine will happily prove that the contract faithfully implements a specification that is wrong. It will prove it fast. It will prove it beautifully. It will prove that a flawed oracle rounding mechanism is internally consistent, because internal consistency is the only thing the proof checks.

I have lived this. During the 2020 liquidity crunch, a price feed on a major lending protocol I had capital in failed. Not the feed itself β€” the rounding. The contract reported a price with a rounding convention that truncated in the direction that made the protocol appear more solvent at the exact boundary where the boundary mattered. I traced it to a fixed-point division. The code was correct relative to its own comments. The comments were correct relative to the author's mental model. The mental model was wrong about which direction to round. No prover will ever catch that, because the specification and the error are the same sentence. Formal verification cannot verify the goal. It can only verify the route.

A proof of the wrong theorem is still a proof, and AI is about to make wrong theorems abundant.

4. The Oracle, Which Was Always The Real Centralization

Every decentralist I have ever argued with wants to talk about consensus. Proof of Work, Proof of Stake, Nakamoto coefficient, validator diversity. My position has not moved in sixteen years: the interesting centralization is never in the consensus layer. It is in the data ingress. The oracle.

Bitcoin's security model is honest about what it secures: the ordering and finality of a set of transactions that contain no information about the outside world. The moment you need to reason about a price, a timestamp, a weather report, a sports score, or a KYC attestation, you import an external authority. The chain does not verify that authority. The chain verifies that some set of signers claim it. That is not trustlessness; it is a committee with a cryptographic costume.

Now introduce AI into the ingress. Not as the oracle β€” as the thing consuming it. An AI agent that executes financial actions based on an on-chain price reading inherits every failure mode of the oracle plus a new one: the agent's decision function is itself an unverified model. The agent does not know when the feed is lying. It has no proprioception for its own input corruption. You have now composed two opaque systems β€” a committee-attested data feed and a black-box decision policy β€” and wired them to a signing key with real capital.

The financial press will call this "autonomous DeFi." I call it a latency-arbitrage surface with a marketing budget. The correct security posture is the opposite of what the agent-economy pitch implies. You do not want the agent to be free; you want the agent's decision function to be formally specified, the specification to be human-auditable, and the oracle to be multi-source with circuit breakers. That is the opposite of what venture capital is funding.

5. Incentive Design, Which Nobody Can Verify At All

Here is the ugly secret of tokenomics. Most of it is not a solved problem. It is a simulation. Teams model an incentive system, run a Monte Carlo, observe that the system does not immediately collapse under the parameters they chose, and ship it. The model has no adversarial term because the team is not adversarial. They assume the participants will behave like a noisy version of themselves.

AI inverts this. An incentivized system is a reward function. A reward function is a search target. A reasoning model with an objective is a search process. The only question is whether the reward for gaming the system exceeds the cost of the gaming, and that is a number an agent can compute and compare faster than a governance forum can schedule a call.

I watched points programs turn into this without any AI at all. The airdrop era taught the entire market that users will mechanically optimize whatever metric is published. Publish active-address count, get Sybil farms. Publish volume, get wash traders. Publish governance participation, get vote-buying markets. Every one of these is now an automated search problem. The tokenomics teams were never competing with humans. They were always competing with optimizers. AI just removes the last trace of human friction β€” the part where a farmer had to write a script, rent a VPS, and click through a wallet.

6. Governance, And The Compliance Shield

I will say the quiet part the way I always do. A DAO is a compliance structure with a governance interface. The token vote is not sovereignty; it is a liability firewall. When the SEC knocks, the answer is "the protocol is decentralized, no controlling party." When a whale wants to reprice the risk parameters, the answer is "the token holders voted." The vote is real. The sovereignty is not.

Add AI and the firewall gets a new feature. Delegation. A voter can delegate to an agent. An agent can delegate to a model with a policy. Now the governance decisions are being made by black boxes whose preferences were shaped by the entity that trained them, and the entity that trained them is a corporation with its own agenda. The vote is still on-chain. The reasoning is off-chain and opaque. The DAO becomes a mechanism for laundering a corporate decision through a cryptographic wash cycle.

Tao's warning, in this light, is not about math at all. It is about what happens to any system whose legitimacy rests on a scarce, human-performed act of judgment β€” and every crypto governance system I have inspected rests exactly there.


Contrarian: What The Bulls Got Right

I am supposed to be the skeptic, so let me do the disloyal thing and give the optimists their strongest case. I will not strawman it. The bull argument is better than my industry likes to admit, and if I do not state it correctly, my skepticism is just posture.

Here is the claim the AI-formal-verification crowd is actually making. It is not that AI will make crypto safe. It is that AI will make verification cheap enough that it stops being a luxury good for well-funded protocols and becomes a default. And that is a genuine structural change, and it is genuinely good.

Think about what formal verification costs today. A meaningful proof of a nontrivial contract is a specialized engagement. You need people who can sit between a theorem prover and an engineer. That is perhaps a few hundred people on the planet who are both competent and available. That scarcity means verification is rationed. The protocols that get formally verified are the ones with foundation money and narrative incentives. The long tail β€” the composable middle of DeFi, the thousands of contracts that hold real user funds and have no budget for a Lean proof β€” is verified by nothing except a hope and a two-week audit whose report ends up framed on the website.

If AI drops the marginal cost of a machine-checkable proof by an order of magnitude, the long tail gets covered. Not perfectly. Not by a specification-aware reasoner. But by something that can catch the entire class of bug I found in 2017 β€” unguarded state transitions, wrong modifier placement, incorrect rounding direction β€” at a cost near zero. That is a real reduction in catastrophic loss. It is not the elimination of risk; it is the compression of the dumbest 80% of risk, and the dumbest 80% of risk is where most user funds have historically died.

So yes. The bulls are right about the direct effect. Cheap proofs catch cheap bugs.

The problem is the indirect effect, and this is where I part ways with them. Cheap proofs also catch nothing, because the market will not distinguish between a proof that matters and a proof that is theater. Here is the mechanism. Right now, a formal verification report is a costly signal. When you see one, you update. When proofs are free, the signal inverts: the absence of a proof becomes suspicious, so every protocol generates one, so the proof carries no information, so the market falls back on the older, cheaper signal β€” the brand of the auditor, the size of the funding round, the eloquence of the founder. The verification arms race ends in the same equilibrium as every other arms race: the artifact is universal, the artifact is worthless, and trust re-consolidates into institutions that were never trustless to begin with.

Verification is only valuable while it is scarce. AI is the most efficient scarcity-destruction machine ever built for exactly the resource crypto sells.

The sophisticated bull will counter: the scarce resource will shift from proof-production to specification-authorship. The people who can write the right specification β€” who can say what a lending market should do during a cascade β€” will become the bottleneck, and value accrues to them.

I agree with that, and I want to be honest that it is a good argument. But notice what it concedes. It concedes that the decisive act is human judgment about intent, and that judgment is exactly what cannot be verified, cannot be benchmarked, cannot be delegated to a model, and cannot be scaled. The bull argument, pushed to its boundary, becomes the bear argument: the one thing crypto cannot automate is the one thing that turns out to matter, and the industry is still paying lawyers, not mathematicians, to do it.

Problem Exhaustion: Terence Tao's AI Math Warning Is a Warning for Crypto's Verification Economy

And there is a second concession buried in the bull case that nobody wants to say out loud. If cheap verification exists, then cheap attacks exist too. The same reasoning model that finds your reentrancy vector finds everyone else's. The offensive-defensive asymmetry does not resolve in the defender's favor just because defenders are nice. It resolves toward whichever side iterates faster, and the offensive side iterates faster because it is not constrained by user experience, backwards compatibility, or a governance vote. The auditor gets one submission window. The attacker gets infinite attempts against a live target.

So the bulls are right that AI will compress the dumb risk. They are wrong that this compounds into safety. It compresses the dumb risk and elevates the sophisticated risk, which is worse, because sophisticated risk is harder to see and harder to insure and, as the track record of 2022 showed, more correlated. When the sophisticated failure mode finally triggers, it triggers everywhere at once, because everybody hired the same machine to protect them from the same class of attack, and the machine had the same blind spot.

They built on sand; I built on skepticism. And skepticism is the only asset that does not get cheaper when the machines get faster.


Takeaway: The Problem Supply Was Always The Product

I want to end where Tao reportedly started, but pointed at my own domain, because that is the only honest use of someone else's warning.

Mathematics is worried that AI can flatten problems faster than humans can pose them. The response there is institutional: build better problem generation, price the verification, defend the human judgment that decides which problems matter. It is a hard problem, but it is being named, and it is being named by a Fields Medalist, which means an academic commons exists that can absorb the warning.

Crypto has no such commons. It has a feedback loop where the problem supply is not a commons at all. It is a product, and the product is priced by the scarcity of the people who supply it. Auditors. Formal-methods engineers. Oracle operators. Governance committees. Risk curators. Every one of these is a bottleneck currently being paid a scarcity premium, and every one of them is on the wrong side of a curve that is bending faster than any of them expected.

The thing to watch is not the model. The model is a commodity. The thing to watch is whether the industry can decouple verification-as-a-signal from verification-as-a-bottleneck. Right now it cannot. Right now, a report costs money and therefore means something, and the meaning evaporates the instant the cost does. Somebody in the next twenty-four months is going to publish that AI can produce a passing audit artifact for a mid-complexity contract at a fraction of the human cost, and the market will treat it as a bullish headline. It is not. It is the day the industry's most expensive signal became free, which is the day it became worthless, which is the day trust re-consolidated back into the same three custodians everybody was trying to escape.

Cold logic cuts through the noise of FOMO. Apply it here. If the thing you are selling is verification, and the cost of verification is collapsing to zero, then the correct question is not how much faster the model got. The correct question is what you still know how to do when everyone can produce the certificate and nobody can produce the judgment about what the certificate should certify. Crypto has spent sixteen years claiming it would never need that judgment. It is about to find out whether that was a design decision or a deferred bill. The code does not care which one you meant, and the machine that flattens the problem does not know why the problem was worth posing.