TL;DR
On Oct 6, 2026, a developer reported a US$10,811.41 Cloudflare invoice caused by a Durable Object alarm() handler that re-scheduled itself in an infinite loop — code written by an AI coding agent. Cloudflare support denied any credit: usage-based billing charges for resources actually consumed, and the loop was customer code, not a metering fault.
Two days earlier, another developer reported a ~US$1,000 bill: an AI-vibe-coded voting page polled an API endpoint every 3 seconds from every open browser tab, racking up 2.6 billion requests.
Our takeaway: this failure mode is real, documented, and recurring. The fix is not "be careful" — it's guardrails in code, billing alerts with an owner, and a review gate for AI-generated scheduling and polling code. The full playbook is below.
中文摘要:2026年10月,两起 AI 生成代码导致的 Cloudflare 天价账单:一是 Durable Object alarm() 死循环,账单约 $10,811,官方拒绝减免;二是 AI 写的投票页面每 3 秒轮询接口,产生 26 亿次请求,账单约 $1,000。此类事故真实且反复发生。下面是我们的团队防范方案:代码层面加熔断与开关、账单告警指定负责人、AI 生成的调度/轮询类代码必须人工复核。
1. Case 1 — the $10,811 alarm loop
- A small-team developer's Cloudflare account was hit with an invoice of US$10,811.41 for one billing period.
- Root cause: a project using Durable Objects had an
alarm()handler that calledsetAlarm()again at the end of every run — including failed runs — creating a tight infinite loop. The loop performed ~600,000+ storage read/write operations. - The buggy code was written by an AI coding agent (the reporter blames Codex-generated code that was committed without close review).
- Cloudflare support's reply: "the alarm loop in your Workers … the loop originated from a bug in your Worker code rather than Cloudflare metering or infrastructure. Support is unable to issue a credit for these Durable Objects charges." The invoice stayed open and due.
- Amusing footnote: the support reply was signed "This reply was enhanced by Cloudflare Workers AI (Kimi K2.7 by Moonshot AI)."
2. Case 2 — the 2.6-billion-request poll
- On Oct 4, 2026, a developer reported a Cloudflare bill of ~US$1,000 for a simple voting site.
- The site was "vibe-coded" in a few sentences with an AI agent. The generated frontend polled an
api/resultendpoint every 3 seconds from every open browser tab — an unbounded polling loop with no backoff, no stop condition, and no caching. - Total damage: 2.6 billion requests. The math checks out: on the Workers Paid plan (~$0.50 per million requests beyond the included quota), 2.6B requests ≈ $1,000+.
- The reporter's own post-mortem: didn't have the agent's output reviewed, didn't add throttling — "these could have been constrained in the prompt; I was careless."
3. How the meters burn money
Two different bugs, same billing model. Case 1 spins all three Durable Object meters at once; Case 2 is pure request volume on Workers:
| Meter | Paid plan | What the bugs do |
|---|---|---|
| Requests (incl. alarm invocations) | Workers: 10M/month included, then $0.50 / million DO: 1M/month included, then $0.15 / million | Case 2: 2.6B polls ≈ $1,000+. Case 1: every alarm firing = a request, thousands per hour, 24/7. |
| Duration (GB-s) | 400K GB-s/month included, then $12.50 / million GB-s | The object stays active in memory the whole time — billed wall-clock. |
| Storage rows read / written | Metered per million operations | Each loop iteration reads state and writes it back. 600K+ ops in the reported case. |
4. The prevention plan — team playbook
RULE 1
Code rules for anything that schedules itself
Alarms, Cron Triggers, Queues with retries, recursive fetch() — any construct that can re-trigger itself must carry all of these:
- Never re-arm unconditionally.
setAlarm()insidealarm()must be gated on real remaining work, never the last line of the handler. - Circuit breaker. Keep a consecutive-failure counter in storage. After N failures (we use 5), stop re-arming and raise an alert instead of looping.
- Kill switch. Check a flag (env var or KV) at the top of every scheduled handler. One flip stops the loop without a deploy.
- Bounded retries with capped backoff. Exponential backoff with jitter, a maximum delay, and a maximum attempt count — then dead-letter.
- Idempotent handlers. Check-then-act inside a storage transaction so a retried alarm can't double-apply.
// ❌ the $10,811 pattern — do not ship this async alarm() { await this.doWork(); await this.storage.setAlarm(Date.now() + 60_000); // unconditional re-arm } // ✅ guarded pattern async alarm() { if (await this.storage.get("killSwitch")) return; const fails = (await this.storage.get("fails")) ?? 0; if (fails >= 5) { await this.alert("alarm circuit open"); return; } try { const more = await this.doWork(); await this.storage.put("fails", 0); if (more) await this.storage.setAlarm(Date.now() + 60_000); } catch (e) { await this.storage.put("fails", fails + 1); await this.storage.setAlarm(Date.now() + backoff(fails)); // capped } }
RULE 2
AI-generated code review gate
The incident's code was AI-written and committed without a close read. New team rule:
- Any AI-generated code touching alarms, cron triggers, queues, retries, recursive scheduling, or client polling must be reviewed by a human against Rule 1 before merge. Add it as a PR-template checklist item.
- Treat AI output as a junior dev's first draft: fast, plausible, and exactly the kind of code that writes
setAlarm()at the bottom "to keep it running" — or asetInterval(fetch, 3000)that never stops.
RULE 2b
Polling rules — the Case 2 fix
A fixed-interval poller multiplied by every open tab is a request firehose. Any client polling a metered endpoint must have:
- No dumb fixed intervals. Back off when nothing changes (e.g. 3s → 10s → 30s on identical responses), and stop entirely after N unchanged polls.
- Pause when hidden. Stop polling when
document.hidden— background tabs are pure waste. - Conditional requests. Use ETags /
If-None-Matchso unchanged polls return 304 instead of full responses. - Prefer push over poll. WebSocket or SSE for live data; one connection replaces thousands of polls.
- Cache at the edge. Short Cache-Control on the polled endpoint so repeat polls hit cache, not the Worker.
- Server-side rate limit per client as the last line of defense.
RULE 3
Billing guardrails with a named owner
- Threshold alerts are assigned, not just enabled. One person owns the billing-alert inbox and a 24h response SLA. An alert nobody owns is decoration.
- Daily usage review. A scheduled check (cron) of Workers/Durable-Object usage via the Cloudflare dashboard or GraphQL analytics API, with a sudden-jump alarm (e.g. 10x day-over-day on any meter).
- Experiments on Free, production on Paid. The Free plan fails closed — it cannot surprise-bill. Prototypes and AI-generated experiments stay on Free until reviewed.
- Separate accounts for experiments vs production where practical, so a runaway test can't touch the production invoice.
RULE 4
Test the failure, not just the happy path
- Test alarm logic locally (Miniflare /
wrangler dev) with a failing run: verify the handler does not re-arm into a loop. - Chaos-check: flip the kill switch mid-run and confirm the loop stops.
RULE 5
Incident runbook — if usage spikes
- Flip the kill switch (or delete/disable the Worker) first — stop the meter before diagnosing.
- Check the dashboard: which meter spiked, when it started, which deployment introduced it.
- Contact support with evidence — but know the policy: customer-code loops are not credited. Prevention is the only refund.
- Post-mortem: which guardrail was missing, and add it to the rules above.
5. Sources
- Case 1 thread: @shmily7 on X, Oct 6 2026 (screenshot; ~$10,811.41 invoice, Codex-generated Durable Object alarm loop, credit denied).
- Case 2 thread: @keking99 on X, Oct 4 2026 (screenshot; ~$1,000 bill, AI-vibe-coded voting page polling every 3s, 2.6B requests).
- Why Cloudflare bills spike: 8 D1 and Durable Objects cases from 2026 — documents the recurring alarm-loop pattern, including a ~$34,895 case.
- Cloudflare Durable Objects pricing — requests $0.15/M, duration $12.50/M GB-s on Paid; alarm invocations count as requests.