Skip to content
11 min read

A Budget Alert Is Not a Budget: What a Hard Spend Cap Does to a Running Agent

Cloud and model providers now ship hard spend caps. They stop the bill, but each one also decides what your agent experiences next, and most agent code was never written for that moment.

Antonio J. del Águila

Knaisoma

Picture a finance lead who approves an agent pilot with a monthly budget and a spend alert at 80 percent. The alert is a sound control for a human who reads email. It is a weak one for a loop that can issue thousands of model calls an hour, because the message arrives after the money is spent and nobody is awake to act on it. The first thing to understand about agent cost is how much of what organizations call a budget is really a notification.

That is changing. In the space of three months, Google Cloud, OpenAI and AWS each shipped enforcement, meaning a switch that stops usage when a number is reached, and Anthropic’s API documents its own monthly caps. Simon Willison argued on October 3 that hard caps should be the default, because coding agents make it trivially easy to deploy something that bills by usage. We agree with the direction and want to add the part that the announcements skip. A cap does not make an agent safe to run unattended. It converts a cost failure into a different failure, and whether that second failure is tolerable depends on decisions your team makes before the cap ever triggers.

Rate limits do not bound what an agent can spend

Teams often assume that provider rate limits already act as a brake. They bound throughput, not money, and the gap is easy to compute from published numbers. Anthropic’s rate limit page lists, for the Build tier, 5,000,000 input tokens and 1,000,000 output tokens per minute for Claude Sonnet 5.5, and its pricing page lists that model at $2 and $10 per million input and output tokens. A client that saturated both limits would spend about $10 a minute on each side, roughly $1,200 an hour, against a Build tier monthly spend cap of $1,000. On the Start tier the same arithmetic gives about $480 an hour against a $500 cap.

Read this as a ceiling, not a forecast. The limits are maxima rather than guaranteed minimums, they apply per model, and no ordinary agent sustains them. The point is that the ceiling a rate limit allows is large enough to consume a month’s budget within an hour, so throttling cannot stand in for a budget. Cached input tokens make the input side higher still, since for most Claude models cache reads do not count against the input limit. Whatever your provider, the question to ask is how long a misbehaving loop needs to reach your monthly number, and the answer is usually shorter than your on-call response time.

Three controls that are often confused

It helps to separate three mechanisms that get lumped under cost control.

An alert notifies a person and changes nothing. OpenAI’s spend limit guide is explicit that an alert sends a notification while API traffic continues. A throttle, such as a rate limit, slows consumption without regard to cost. A cap refuses further usage once a spend figure is reached, which is the only one of the three that bounds the bill. A common arrangement is to have the first, assume the second, and never decide on the third.

The caps themselves differ in ways that matter more than their headlines suggest. This table and the details after it summarize what each provider documents today.

Provider Scope At the cap
AWS Project Project paused, data kept 90 days
Google Cloud One service in one project New requests paused, compute keeps accruing
OpenAI Organization or project HTTP 429, with lag
Anthropic Tier cap, plus limits you set 429 at tier cap, 400 at yours

Beyond the table, the details are these. AWS spend limits require a paid plan, cover up to 10 projects and were in limited rollout when we checked. Google’s cap covers four services (Gemini API, Gemini Enterprise Agent Platform, Cloud Run and Cloud Run functions), lets in-flight requests finish and enforces on estimated costs. OpenAI lets organization and project hard limits apply to the same request. Anthropic’s tier caps are $500, $1,000 and $200,000 a month, with none on the Custom tier, and you can set lower limits per organization or workspace.

Two rows deserve attention. AWS describes its limit as designed for experimentation, learning and sandbox workloads, and says it can serve production only where a brief pause of resources is acceptable. That is an unusually candid statement of the tradeoff, and it should shape where you apply it. Google’s cap pauses new requests to a named service but leaves fixed resources running, so a capped agent platform can keep paying for the machines under it. Neither behavior is a flaw. Both are things a reader assuming “cap means stop” would not expect.

What the agent experiences when the cap fires

The most important detail in the table is the last column of the Anthropic row, because it is about the client, not the invoice. According to the documentation, hitting the tier cap returns HTTP 429, the same status as a rate limit, but with no retry-after header, and the page states that retrying, including the SDK’s automatic retries, fails until access resumes. A cap you set yourself returns a 400 instead. OpenAI also returns a 429, with separate error codes for organization limits, project limits and its own tier limit, and its announcement for the feature warns users to keep that status in mind when distinguishing insufficient credits from an enforced cap.

This is our inference from those documents, not a result either provider reports: an agent framework that treats every 429 as a transient condition will keep retrying a refusal that cannot clear for days. At best the loop burns time and logs. At worst it holds a customer session open, retries tool calls with side effects, or escalates to a fallback model on a different account that has no cap at all. The same status code now means “slow down” in one case and “stop until the first of next month” in another. The only safe policy is to classify by the error code in the response body, treat budget exhaustion as terminal, and make the terminal state visible.

Mid-task refusal also leaves work half-finished. An agent that has drafted three of five steps in a refund workflow, or opened a change request it can no longer update, is in a state no one designed. Caps stop money, not consequences. This is why the reliability work we described in our piece on agent success rates matters here: a budget refusal is one more way a run fails partway, and it needs the same checkpointing and idempotency as any other.

Delay and scope: the two ways a cap misses

Even a correctly handled cap can miss its target in two ways, and both are documented. The first is delay. OpenAI says plainly that enforcement is not instantaneous and recorded spend can slightly exceed the configured amount. Google enforces on gross, estimated usage precisely because billed costs can take more than 24 hours to appear, and it states that any overage is billed as normal. Treat the configured number as a target the provider will approximately honor, not a contractual ceiling. If the true limit of your tolerance is $5,000, set the cap well below it.

The second is scope. A cap attaches to a project, a workspace or a service, while the cost of an agent attaches to a workload that crosses all three. An agent that calls a model API, a search tool and a hosted function may be covered by three different caps, one cap, or none, depending on which services each provider considers eligible. Google lists four eligible services. AWS limits apply per project, and the project is a billing construct that your architecture may not map onto. A team that wants to cap one agent but runs it in a shared project will either cap everyone or cap no one. This is a design input: give each agent workload its own project or workspace and its own key before you rely on a cap to contain it.

Where each control belongs

Because every cap trades a bounded bill for an interruption, the right setting depends on how costly the interruption is. We find it useful to sort workloads into three groups and assign controls accordingly.

Workload Primary control Platform cap role
Sandboxes and personal projects Hard cap, set low, on by default The only control
Internal agents with human users Per-run and per-user budget in the app Backstop per project
Customer-facing agents Per-task budget with graceful fallback Ceiling above normal peak

Behavior at the cap differs by group. A sandbox accepts the pause and a person raises the limit by hand. An internal agent degrades to a cheaper model or asks a person. A customer-facing agent fails closed on the risky step, pages on-call and keeps a queue.

The common thread is that the platform cap is the last line, not the first. The first line is a budget the application itself tracks per run, which is the argument we made earlier in AI agent budgets need control loops. A run that knows it has used 80 percent of its token allowance can summarize and stop on its own terms, finishing the refund step before the project pauses it. A cap that fires first gives the same containment with none of that grace.

For the customer-facing group in particular, the tension is real. Turning on a hard cap there means accepting that an unexpected spike takes a service down, which for some businesses is the worse outcome. The honest alternative is not to skip the cap but to set it high enough that normal traffic never touches it, back it with application budgets that fire much sooner, and rehearse what the service does when it trips. A cap you have never triggered in a test environment is a hypothesis.

A checklist before you turn caps on

Use this list to decide whether a workload is ready for a hard cap, and to find the gaps if it is not.

  1. Estimate the fastest plausible burn rate from your rate limits and pricing, and compare the time to your monthly cap with your on-call response time.
  2. Give each agent workload its own project, workspace or key, so a cap contains one thing.
  3. List every billable service the agent touches, and check which of them each cap covers.
  4. Set caps below the real tolerance, since enforcement lags and overages are billed.
  5. Put alerts below both application budgets and platform caps, so that a person hears first.
  6. Classify errors by code, make budget exhaustion terminal, and stop retries for it.
  7. Define the half-finished state: what the agent must leave behind, and who finishes the work.
  8. Decide who may lift a cap, how quickly, and where that decision is recorded.
  9. Trigger the cap in a test project at least once, and keep the evidence.

The position we take

Hard caps are a real improvement, and defaulting to them for sandboxes and personal projects is the right call. What we would resist is the idea that turning one on completes the job. A budget alert tells a person something happened. A cap decides what happens to a running system, and that decision belongs in your design, not in a settings page. The organizations that will handle this well are the ones that treat spend as a runtime condition their agents are written to expect, with the platform cap as the safety net underneath.

If your teams are moving agents from pilot to production, spend controls are probably still a set of alerts that nobody has tested. We help organizations estimate worst-case burn rates, design per-run budgets and cap placement across providers, and rehearse the failure path before it happens in production. Talk with us about your agent cost controls.

AIGovernanceCost ManagementAgentic AI
Share:

Stay updated

Get insights on engineering transformation delivered to your inbox.

Newsletter coming soon.