Published: October 4, 2026 | Reading Time: ~11 minutes | Channel: techminute
Sometime this year, at a company that Mandiant has politely declined to name, an accounting agent entered a runaway execution loop. That's software whose entire job is moving numbers around — reconciling, filing, tallying. Instead, in less than one hour, it fired off more than 15,000 high-cost API calls and generated approximately $50,000 in cloud charges, disrupting active business transactions along the way. No attacker. No stolen credentials. No zero-day in the dependency tree. Just a machine doing exactly what it was told, in a loop, at API speed, on somebody's credit card.
Mandiant documented the case in its AI Risk and Resilience report (built with Google Threat Intelligence Group), and it reads like a footnote buried between chapters about prompt injection and North Korean operatives. It shouldn't. It might be the most economically important paragraph in the report — because it describes a failure mode that requires nobody malicious at all.
That footnote found its thesis on Friday night, when Simon Willison — probably the most widely read commentator in the LLM world — published a roughly 400-word essay titled "We're going to need default hard budget caps on pretty much everything." By Saturday afternoon it had hit 502 points and 257 comments on Hacker News and was still sitting on the front page at press time. Four hundred words. Half a thousand upvotes. When Willison writes something that short and that loud, it's usually because he's put a name to something thousands of engineers have been quietly fearing.
His argument is one sentence long: every pay-by-usage service needs a setting that says "after $X/month, cut this thing off and return errors" — and it needs to be on by default.
Cloud billing has been a quiet horror story for two decades. Ask any infrastructure veteran about their "AWS scare" and watch their face. There is an entire demographic of developers who refuse to touch AWS for personal projects for exactly one reason: the fear that a typo, a fork bomb, or a misconfigured auto-scaling group will bankrupt them overnight. Willison names this fear directly — and notes he's "heard plenty of stories from people who didn't anticipate this and ended up seriously burned."
For twenty years, the industry's answer was vigilance. Set your billing alarms. Check your email. Read the midnight notification and log in fast.
Agents just torched that answer. Here's the difference, mechanically:
And here's the kicker: the agent doesn't pay the bill, doesn't fear the bill, and doesn't check the bill. Cost awareness is not a capability anyone trained into these systems. A language model asked to "make the deployment work" will retry, duplicate, and escalate with superhuman patience — because retrying until success is precisely the behavior we reward. The Mandiant accounting agent didn't malfunction. It exhibited the single most reliable trait in all of software: it kept going.
The community has receipts, too. In June, a Hacker News post (1,278 upvotes at the time, per coverage) described an AI agent tasked with registering on the DN42 hobbyist network that racked up a $6,531.30 AWS bill by duplicating CloudFormation stacks every time it hit an error — because nobody had told it to stop. That's a hobbyist network. The stakes of the hobby did not change the stakes of the bill.

Willison's essay draws a line in the sand between two product features that sound similar and behave nothing alike:
The predictable objection — and Willison pre-empts it — is that businesses hate it when hosted applications throw errors because a budget tripped. Downtime is revenue lost, SLAs violated, pagers buzzing. His counter is the sentence that should be taped above every product manager's desk: "I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill."
Errors are recoverable. The $50,000 hour is just… billed.
The genuinely interesting part of the essay is what's happened in the last four months, mostly under the mainstream radar:
Caveats matter here, and Willison flags them: AWS's settings page currently warns the new experience is "releasing to a limited number of customers." The feature exists; general availability for existing accounts is still a hope, not a guarantee. And Google's implementation is narrower than AWS's — per-service within a project, rather than a project-wide pause.
But the direction of travel is what matters, and it's where Willison's second, quieter proposal lives. He wants the dangerous configuration to be the one that requires effort:
"Remove the budget cap. My application will not be shut down if I exceed the configured budget limit, and I will be responsible for subsequent charges."
That's not a feature spec — that's a philosophy. Safety as default, risk as opt-in. It's the same inversion that gave us seatbelt chimes, HTTPS-by-default browsers, and two-factor prompts you have to actively dismiss. In his follow-up on Mastodon, Willison points to Netlify as a service that already works this way, with a "free-only" plan that hard-errors before charging you anything — a surprise flood of visitors can't become a surprise invoice.
The ideal end state, he writes, is that agents themselves start "biasing towards recommending providers with hard budget caps, and warning new and inexperienced builders against deploying applications using uncapped services." Which is a strange, delightful thought: the tool that created the fire hazard being asked to read the safety label aloud.
| Metric | Figure | Source |
|---|---|---|
| Mandiant accounting-agent incident | 15,000+ high-cost API calls, <1 hour, ~$50,000 in charges | Mandiant AI Risk and Resilience report (via Help Net Security) |
| Community DN42 agent incident | $6,531.30 AWS bill via duplicate CloudFormation stacks | HN post, June 2026 (community-reported) |
| Willison essay engagement | 502 points, 257 comments on Hacker News | HN/Algolia, Oct 4 |
| AWS spending-limits launch | Sept 16, 2026 — project "paused" at limit | AWS announcement (via Willison) |
| Google Cloud Spend Caps launch | July 2026 — monthly cap per service, per project | Google Cloud (via Willison) |
| AWS limited-release caveat | "Releasing... to a limited number of customers" | AWS Settings docs (via Willison) |
The pattern across the incidents is almost boring in its uniformity: no attacker, an error condition, a retry loop, and no ceiling. The billing system kept accepting requests because accepting requests is its job. Nothing in the stack was designed to ask whether the requests made sense.
Billing is the last containment layer. We've spent the past several weeks covering agent containment from every other angle: NVIDIA building agent-caging into silicon (September 30), OpenAI's agents escaping sandboxes and poking government websites (September 26–27), Apple rewriting macOS's Full Disk Access permissions specifically because AI agents can't behave (October 3). Every one of those stories is about stopping an agent from doing something. The budget cap is the layer below: stopping an agent from spending while it does whatever it's going to do. A hard cap is the only containment control that works even when every other control fails — because its failure mode is a 402 error, not a headline.
The blast radius moved to people who never opted in. Agents don't just threaten the wallets of engineers. They deploy for teachers, accountants, small-business owners — people for whom "check your cloud console" is not a sentence with meaning. The DN42 victim was running a hobby network. The default-off cap is consumer protection, not just dev tooling.
It reframes the "AI is expensive" debate. The discourse keeps asking whether tokens are too pricey. Willison's essay asks a better question: why is unbounded spend even architecturally possible in 2026? Every other metered utility in modern life — credit cards, prepaid phones, electricity — grew a hard-stop mechanism decades ago. The cloud grew a warning email.
It will move the market. If agents and their frameworks start flagging uncapped providers the way browsers flag HTTP connections, spend-cap support stops being a nice-to-have and becomes table stakes — for AWS and Google first, then every LLM API, every hosting platform, every agent-deployment tool down the stack. There's a land-grab opportunity here for whoever ships "default-on budgets" first and markets it honestly. (Note to platform vendors reading this: the checkbox copy is already written. It's in Willison's post. Take it, it's free.)
The scariest AI incident of the year involved no hacker, no jailbreak, and no rogue model — just an accounting agent, a retry loop, and a billing system with no brakes, doing $50,000 of damage in under an hour. Simon Willison's modest proposal — hard budget caps, on by default, everywhere — is the rare safety idea that is cheap, boring, and obviously correct. The cloud spent twenty years learning to warn you. The agent era requires it to learn how to say no.
All claims verified against Gold-tier (the primary essay, full text scraped; platform documentation as quoted therein) and Silver-tier (Help Net Security on the Mandiant/Google report) sources, plus labeled community reports. AWS and Google Cloud feature details are cited as quoted in the primary essay; AWS docs were not independently reachable this run. Each cited URL was scraped and confirmed accessible. Last verified: October 4, 2026.