One of the world’s biggest companies has been added to the long list of AI overspending casualties. Amazon spent a whopping $1.8 million on a single Claude project that matched author names to product listings. It overran 860% above budget, went unnoticed for five months, and the real kicker is it never shipped a single usable result.
This is the same company that in May switched off KiroRank, the internal leaderboard its staff had been gaming to look busy with AI, a practice now known as tokenmaxxing.
The two stories are only months apart, and both are symptoms of the same condition: nobody at Amazon, the company that runs the world’s biggest AI cloud, can reliably see what any of this costs.
What Was KiroRank, and Why Did Amazon Kill It?
Kiro is the AI coding platform Amazon launched in July 2025. KiroRank was the leaderboard that grew up next to it, scoring engineers on how much Kiro they used. Whilst it was never official, it was, for a time, popular.
Behind KiroRank was a company target of more than 80% of developers using AI every week. Set a target like that and people know exactly what to chase. So they chased it. Staff pointed autonomous agents at pointless tasks to push their numbers up, the Financial Times reported, and the compute bill climbed with them. The staff were eventually told to stop using AI for the sake of it.
Amazon called KiroRank an unapproved beta and deprecated it. Meta’s staff had already killed off their own grassroots version, Claudeonomics, in April. The wider culture around all this, the leaderboards, the in-house titles, the $80,000 vibe-coded video game, we covered here.
How Is the $1.8 Million Failure Different From Tokenmaxxing?
Tokenmaxxing is visible waste. Someone decides to burn tokens and the cost is the whole point. You can see it on a leaderboard.
Amazon’s $1.8 million project was the other kind of waste, the kind nobody chose. It started out as a job trying to match author names to product listings. The task was dull, useful, and exactly the sort of task AI is meant to be good at.
Instead it ran 860% over budget, stayed unnoticed for five months, and produced nothing worth shipping. Amazon’s own engineers reportedly called the results “catastrophically expensive”.
Unfortunately for Amazon, this wasn’t even a one-off. The same internal review flagged around $541,000 wasted on, of all things, a financial auditing tool, and roughly $134,000 on a system meant to make deliveries faster. Three projects, about $2.5 million, none of it deliberate.
How Does a Free Bug Become a Seven-Figure One?
The mechanism is a billing change most people have not clocked yet. Enterprise software used to cost a flat fee. A bug was a nuisance, an engineer fixed it, and the meter did not move while they did. AI vendors charge per token now, for every unit of text the model chews through. Amazon’s engineers put it plainly: a bug that once cost nothing to fix becomes ruinous when an agent sits in a loop, because every wasted turn carries a price.
Then multiply. On its own earnings call Amazon said its Bedrock platform processed more tokens in the first quarter of 2026 than in every prior year combined. At that rate a small mistake does not stay small for long. This one cost $1.8 million.
What Are “Normalised Deployments”, and Do They Fix Anything?
Amazon’s fix hints at where this goes next. It has reportedly dropped raw token counts for a measure it calls normalised deployments: proof that engineers are using AI to ship useful code, rather than proof they are simply using AI. An admission, in other words, that the first metric measured the wrong thing.
Others agree. Cognizant’s chief executive has written off token consumption as a “vanity metric”. A market is already forming around the alternative, tooling that ties AI spend to outcomes instead of volume. One of the better-known names, Langfuse, was founded in Berlin in 2023 by Max Deichmann, Marc Klingen and Clemens Rawert, then bought by the American database firm ClickHouse in January. Which tells you, roughly, where the value and the ownership are pooling.
Is $2.5 million Just Pocket Change for Amazon?
Around $2.5 million against $23.9 billion of quarterly operating income is, for Amazon, a rounding error. Amazon has said it is experimenting and learning, and some overspend is the ordinary cost of trying anything new. It caught the problem, killed the leaderboard, changed the metric. That is a feedback loop doing its job.
All true, but also besides the point. The number was never the real story, the blindness was. The most capable AI operator alive could not see an 860% overrun for five months, on its own platform, using its own tools. Nobody is worried about Amazon covering $2.5 million. The worry is everyone else in the queue behind it, holding the same billing model and none of the balance sheet.
What Does This Mean for European Firms Buying AI by the Token?
It means the risk has moved from the model to the meter. A mid-size company in Munich or Warsaw signing a per-token contract with a US vendor is buying the exact billing structure that blindsided Amazon, minus the cushion to absorb a surprise. The exposure sits with the finance team now, not the engineers, and it stays quiet until a quarter closes badly.
Moving slowly on AI gets treated as a failure of nerve. Amazon’s $1.8 million argues the reverse: wiring in spend controls before you switch the agents on is not timidity, it is the only version of this that does not end with a finance team reverse-engineering where the money went.
See Also:
Tokenmaxxing 101: How to Become Your Office’s Cache Wizard
Company Blew $500M on Claude Because Nobody Set a Spending Limit
Y Combinator Built Sam Altman’s Empire, Now He’s Tokenmaxxing It
