Docs / Limits & billing

Limits & billing

Windows, slots, wallet, and spillover — with real numbers.

Cobble has one billing model: your plan includes a budget of inference, and your wallet covers everything past it. No hidden fair-use policy, no surprise overages. This page explains exactly how that works, with real numbers.

The plan: a usage budget per window

Each plan includes a budget of inference per 5-hour window and per week. Budgets are measured in dollars of usage at our published per-model rates, not in requests, so a small model or a cached prompt uses far less of your budget than a large, uncached one.

PlanPer 5-hour windowPer weekParallel requestsContext
Solo$1.40$141128K
Pro$3.50$353256K
Max$8.40$846256K

Every model in the catalog is included on every plan. Larger plans get more usage per dollar: Pro includes 2.5x the Solo budget for 2x the price, Max 6x for 5x.

How the windows work. Your first request starts a 5-hour window. When it ends, the whole budget comes back at once. The weekly window works the same way over 7 days and caps the sum of all 5-hour windows: the weekly budget equals ten full 5-hour windows. A request must fit in both. Unused budget does not roll over.

What a request costs. Uncached input tokens at the model's input rate, cached input tokens at a quarter of the input rate, and output tokens at the output rate, all as published on the pricing page. On the flagship at the current rates, a typical agent step with a warm cache costs well under a cent; a fully uncached 160K-token prompt costs a few cents. The budget is charged when the request completes, from the tokens actually used.

Backstops. Each plan also has a request ceiling per 5-hour window (Solo 1,500, Pro 4,000, Max 9,000), a requests-per-minute limit, and a cap on active API keys (Solo 3, Pro 5, Max 10). These exist to stop floods of near-free requests and key sharing; normal use never reaches them.

Embeddings are the one thing not measured in dollars: each plan includes a monthly embedding token allowance (Solo 10M, Pro 50M, Max 150M), with overage billed to your wallet.

Slots: how many requests can run at once

A slot is one request executing concurrently. One slot means requests queue behind each other; three mean your agent's parallel tool calls actually run in parallel. Higher plans are routed first when the fleet is busy. That is a capacity preference, not a speed guarantee: under real saturation it only means you are refused last.

When you hit a limit

Every refusal is an HTTP 429 whose body says who refused and whether anything was charged:

codesourceWhat it meansWhat to do
USAGE_BUDGET_EXHAUSTEDplanYour 5-hour or weekly budget is spent. The body names the window and its reset time.Wait for the reset, enable spillover, or top up.
REQUEST_CEILING_REACHEDplanToo many requests in this 5-hour window.Wait for the reset.
CONCURRENCY_LIMIT_EXCEEDEDplanAll your parallel slots are in use.Retry when one finishes.
FLEET_DEADLINE_EXCEEDEDfleetThe fleet could not schedule this one request inside its service deadline.Retry, shorten the prompt, or reuse a cached prefix.
FLEET_QUEUE_FULL, FLEET_MODEL_UNAVAILABLE, FLEET_BUSYfleetThe fleet declined this one request for capacity.Retry shortly.

A source: fleet refusal is about one request, not your plan: nothing is charged and your budget is untouched. The body includes your remaining budget so you can see that. Every 429 carries Retry-After.

The wallet: everything past the plan

Your wallet holds dollar credit from top-ups. Credit you pay for never expires. Packs of $100 and above include bonus credit (10% at $100, 15% at $250) that is spent first and expires 12 months after it is granted. Custom amounts from $5 to $500 earn the same bonus tiers.

The wallet pays for two things, both at the published catalog rates:

  1. Spillover: usage past your budget, if you enable it
  2. Embedding overage past your monthly allowance

Credits-only accounts (no plan) pay catalog rates for everything, with 2 parallel requests.

Spillover: what happens past your budget

You choose, in the dashboard:

  • Spillover off (default): requests past the budget return 429 USAGE_BUDGET_EXHAUSTED with a Retry-After header until the window resets. Nothing is ever billed beyond your subscription.
  • Spillover on, with a cap you set: requests past the budget keep working, billed from your wallet at catalog rates, up to your monthly cap. Your agent never hits a wall mid-session unless you decided where the wall goes.

There is no way to spend wallet money you did not top up, and no way to exceed a cap you set.

Checking your usage

The dashboard shows both budget windows, what is left in each, when each resets, your slots in use, and your wallet balance. Programmatically, every response carries:

code
X-Cobble-Window-Unit: usd
X-Cobble-Window-5h-Limit: 1.40
X-Cobble-Window-5h-Remaining: 0.92
X-Cobble-Window-5h-Reset: 2026-10-02T19:14:05.000Z
X-Cobble-Window-Weekly-Limit: 14.00
X-Cobble-Window-Weekly-Remaining: 11.30
X-Cobble-Window-Weekly-Reset: 2026-10-07T14:09:41.000Z
X-Cobble-Window-Requests-Remaining: 1412

A refusal also carries X-Cobble-Reject-Source: plan or fleet.

Plan changes

Upgrades take effect immediately; downgrades at the next billing cycle. Usage already burned in the current windows carries over to the new plan's budgets. Lapsed subscriptions do not kill your keys: they fall back to wallet billing at catalog rates, so nothing you built stops working.

Fair use, in numbers

We do not have a hidden fair-use policy. The budgets, ceilings and slot limits above are the policy, enforced mechanically. Things we do prohibit: reselling plan capacity, sharing one subscription's keys across an organization, holding more than one self-service plan per person, and using multiple accounts to stack budgets. Accounts doing so are suspended. The binding version of these rules is Section 18 of the Terms of Service and Part B of the Acceptable Use Policy.

Your data

Cobble does not store prompts or completions, for any plan or wallet request, and does not train on them. We keep request metadata only (tokens, model, cost, status, timing). The full commitment is in Section 7 of the Terms and Section 2 of the Privacy Policy.