Every developer who has built on a hosted model has met HTTP 429. The status code says one thing, "too many requests," and for most of the history of web APIs that is exactly what it meant: you called faster than the account allowed, so slow down. Inference providers inherited the code, and with it the assumption that a 429 is always the caller's fault and always fixed by waiting.
That assumption does not survive contact with autonomous agents. An agent is not a person clicking a button. It is a loop that reads a repository, plans, calls tools, reads the results, and calls the model again, dozens or hundreds of times an hour, sometimes around the clock. When an agent hits a 429 it needs to know something a human never asked: did I run out of what I paid for, or did the system run out of room for this one request right now? Those are different failures with different fixes, and a provider that answers both with the same three digits is leaving the agent to guess.
This article is about the engineering behind that distinction. It uses Cobble's infrastructure as the worked example, because it is the one we know, and because we have spent most of this year making our 429s mean something.
Three things that look like a rate limit
When a request is refused, one of at least three unrelated limits has been reached.
Account limits are about what the customer is entitled to. On Cobble these are a usage budget per window, measured in dollars of inference at published rates rather than in request counts, plus a cap on parallel requests and a per-window request ceiling that stops floods of near-free calls. These limits are per account across every key the account holds. They reset on a schedule the customer can see, and nothing the fleet does changes them.
Concurrency limits are about slots. A plan allows a certain number of requests to be executing at once. An agent that fans out ten tool calls on a one-slot plan will see nine of them refused, not because any budget is spent but because there is nowhere to run them yet. The fix is to wait seconds, not hours.
Capacity decisions are about the fleet. Every model is served by a pool with a finite amount of memory and a finite rate at which it can process prompt tokens and generate output. When a request arrives, the question is not whether the caller is allowed to make it but whether the pool can complete it in acceptable time. A very large uncached prompt may need more prefill work than the pool can schedule before the request's deadline, even if the pool is otherwise idle. A burst of traffic may have every replica's queue at its limit. Neither of those is the caller's fault, and neither is fixed by the caller slowing down.
Most providers collapse all three into one status code and, at best, a Retry-After header. The agent retries, the capacity problem has not changed, and the loop burns its real budget on a request that was never going to run.
What a refusal should say
The first principle we settled on is that a refusal must say who refused and what it cost.
Every 429 the Cobble relay returns carries a source, either plan or fleet; a scope, either the whole account or this one request; a budget_charged flag, which is always false on a refusal; a machine-readable code; and a retry_after_seconds the client can trust. An account refusal names the window that is spent and the exact instant it resets. A fleet refusal says why the request could not be scheduled and shows the caller's remaining budget in the same body, so the agent can see that its plan is intact without a second call. The same split is in a response header, for clients that only look there.
That sounds like a formatting detail. It is not. It is the contract that lets an agent framework make the right decision automatically: back off until the stated reset on a plan refusal, retry with jitter on a queue refusal, shorten the prompt or reuse a cached prefix on a deadline refusal, and never retry a request the fleet has said it cannot serve as shaped.
Admission control, not queues
The second principle is that we decline early rather than queue late.
A naive serving stack accepts every request and lets the engine's queue absorb the overload. Under sustained load that produces the worst possible behavior for agents: requests sit in a queue for a minute or more, then time out, and the caller learns nothing until it is too late to adapt. Meanwhile the GPUs are busy doing prefill for requests whose callers have already given up.
Cobble's fleet router makes an admission decision when the request arrives. It estimates the work the request represents, from the size of the prompt, what part of it is likely already cached, and the output the caller asked for, and compares that against the live state of the replicas that can serve the model: how deep their queues are, how much key-value cache they have free, and how fast they are currently producing tokens. If no replica can be expected to complete the request inside its service deadline, the request is refused immediately with a reason, before any GPU time is spent on it. If a replica can, the request is routed to the one most likely to already hold its prompt prefix.
Backpressure, in other words, reaches the client as information rather than as latency. An agent that gets a fleet refusal in a few milliseconds can do something useful with the time. An agent that gets a timeout after ninety seconds cannot.
Reserve, then settle
The third principle is that a refused request is free.
Because Cobble meters plans in dollars, the relay has to decide whether a request fits the caller's budget before the upstream call and know what it actually cost after. It does this by reserving an estimate at admission and settling to the real token counts when the response completes. If the fleet refuses the request, or the upstream errors, the reservation is released in full, including the request count. Only served work is charged.
This matters more for agents than for anyone else, because agents retry. A system that charged for refused requests would make every capacity hiccup a billing event and every retry loop a slow drain. A system that does not charge for them lets an agent treat a fleet refusal as what it is: a scheduling outcome, not a transaction.
Workload classification and model-aware scheduling
Not every request deserves the same treatment, and the fleet does not pretend otherwise.
Requests are classified on the way in by the shape of their work. A short chat turn with a warm cache, a long uncached prompt, a batch-style request from a pipeline, and a heartbeat from an always-on assistant all land differently on the hardware, and the router weighs them differently when deciding admission and placement. Models are classified too: a dense flagship, a sparse mixture-of-experts model, a small high-throughput model, and an OCR model each live in a pool configured for how that architecture actually behaves under load, and the scheduler knows which pool a request is bound for before it decides whether the pool can take it.
Priority is part of this, and it is worth being precise about what it means. Higher plans are routed first when the fleet is busy. That is a capacity preference, not a speed guarantee: under real saturation it means a higher plan is refused last. We say so in the documentation because an agent operator planning around "priority" deserves to know exactly what it buys.
Why this is the right problem to solve
It would be easier to treat agents as ordinary API consumers and tune the rate limiter until the complaints stop. We think that misreads where inference demand is going. The workloads that grow are the ones that run unattended: coding agents working through a backlog overnight, research harnesses processing document collections, assistants that never log off. Those workloads are relentless in a way interactive use never was, and they are also far more cooperative, if the system gives them something to cooperate with.
An agent that is told precisely why it was refused, what it was and was not charged, and exactly when to try again will behave well. An agent that gets a bare 429 will hammer, and its operator will blame the provider. The difference is engineering on the provider's side, and it is engineering most of the industry has not done because interactive users never needed it.
Cobble is built for the agents. The budgets, the admission control, the honest refusals, and the model-aware scheduling exist because persistent agents are the workload we expect to serve most of, not an edge case we tolerate. If you run one, the fleet is already speaking your language.
The Cobble Team

