Team incident record · Cloudflare

Two Bills, One Lesson

How AI-written code burned $10,811 and $1,000 on Cloudflare — and our plan to never be that thread.

TL;DR

On Oct 6, 2026, a developer reported a US$10,811.41 Cloudflare invoice caused by a Durable Object alarm() handler that re-scheduled itself in an infinite loop — code written by an AI coding agent. Cloudflare support denied any credit: usage-based billing charges for resources actually consumed, and the loop was customer code, not a metering fault.

Two days earlier, another developer reported a ~US$1,000 bill: an AI-vibe-coded voting page polled an API endpoint every 3 seconds from every open browser tab, racking up 2.6 billion requests.

Our takeaway: this failure mode is real, documented, and recurring. The fix is not "be careful" — it's guardrails in code, billing alerts with an owner, and a review gate for AI-generated scheduling and polling code. The full playbook is below.

中文摘要:2026年10月,两起 AI 生成代码导致的 Cloudflare 天价账单:一是 Durable Object alarm() 死循环,账单约 $10,811,官方拒绝减免;二是 AI 写的投票页面每 3 秒轮询接口,产生 26 亿次请求,账单约 $1,000。此类事故真实且反复发生。下面是我们的团队防范方案:代码层面加熔断与开关、账单告警指定负责人、AI 生成的调度/轮询类代码必须人工复核。

1. Case 1 — the $10,811 alarm loop

Verdict: true — and not a one-off. This exact failure mode is documented across 2026: a roundup of 8 D1/Durable-Object billing spikes includes an alarm handler that re-armed on every run (failed ones too) and ran 8 days straight to ~$34,895 with zero users. The pattern is always the same: a self-scheduling construct with no guard, and the first sign of trouble is the invoice.

2. Case 2 — the 2.6-billion-request poll

Verdict: consistent with the same root cause. No exotic bug here — just client-side code multiplying a cheap operation by every visitor and every second. Any fixed-interval poller with real traffic becomes a request firehose. This one is arguably more likely to hit a normal team than the alarm loop.

3. How the meters burn money

Two different bugs, same billing model. Case 1 spins all three Durable Object meters at once; Case 2 is pure request volume on Workers:

MeterPaid planWhat the bugs do
Requests (incl. alarm invocations)Workers: 10M/month included, then $0.50 / million
DO: 1M/month included, then $0.15 / million
Case 2: 2.6B polls ≈ $1,000+. Case 1: every alarm firing = a request, thousands per hour, 24/7.
Duration (GB-s)400K GB-s/month included, then $12.50 / million GB-sThe object stays active in memory the whole time — billed wall-clock.
Storage rows read / writtenMetered per million operationsEach loop iteration reads state and writes it back. 600K+ ops in the reported case.
The critical asymmetry: Cloudflare's budget/threshold alerts (on by default for pay-as-you-go accounts since mid-2026) send one informational email. They do not pause, cap, or throttle anything. Someone has to log in and act. On the Free plan, by contrast, exceeding a limit makes operations fail — it fails closed, so it cannot surprise-bill you.

4. The prevention plan — team playbook

RULE 1
Code rules for anything that schedules itself

Alarms, Cron Triggers, Queues with retries, recursive fetch() — any construct that can re-trigger itself must carry all of these:

  1. Never re-arm unconditionally. setAlarm() inside alarm() must be gated on real remaining work, never the last line of the handler.
  2. Circuit breaker. Keep a consecutive-failure counter in storage. After N failures (we use 5), stop re-arming and raise an alert instead of looping.
  3. Kill switch. Check a flag (env var or KV) at the top of every scheduled handler. One flip stops the loop without a deploy.
  4. Bounded retries with capped backoff. Exponential backoff with jitter, a maximum delay, and a maximum attempt count — then dead-letter.
  5. Idempotent handlers. Check-then-act inside a storage transaction so a retried alarm can't double-apply.
// ❌ the $10,811 pattern — do not ship this
async alarm() {
  await this.doWork();
  await this.storage.setAlarm(Date.now() + 60_000); // unconditional re-arm
}

// ✅ guarded pattern
async alarm() {
  if (await this.storage.get("killSwitch")) return;
  const fails = (await this.storage.get("fails")) ?? 0;
  if (fails >= 5) { await this.alert("alarm circuit open"); return; }
  try {
    const more = await this.doWork();
    await this.storage.put("fails", 0);
    if (more) await this.storage.setAlarm(Date.now() + 60_000);
  } catch (e) {
    await this.storage.put("fails", fails + 1);
    await this.storage.setAlarm(Date.now() + backoff(fails)); // capped
  }
}

RULE 2
AI-generated code review gate

The incident's code was AI-written and committed without a close read. New team rule:

RULE 2b
Polling rules — the Case 2 fix

A fixed-interval poller multiplied by every open tab is a request firehose. Any client polling a metered endpoint must have:

RULE 3
Billing guardrails with a named owner

RULE 4
Test the failure, not just the happy path

RULE 5
Incident runbook — if usage spikes

  1. Flip the kill switch (or delete/disable the Worker) first — stop the meter before diagnosing.
  2. Check the dashboard: which meter spiked, when it started, which deployment introduced it.
  3. Contact support with evidence — but know the policy: customer-code loops are not credited. Prevention is the only refund.
  4. Post-mortem: which guardrail was missing, and add it to the rules above.

5. Sources

← back to fun facts