How do usage budgets work?
Each plan includes a budget of inference per 5-hour window and per week, measured in dollars at the per-model rates below. Your first request starts the 5-hour window; when it ends, the whole budget comes back at once. The weekly window works the same way over 7 days, and a request must fit in both. Your dashboard shows what is left in each and when it resets.
What happens when I use up a budget?
Requests get a 429 that names the window and its reset time, and nothing is charged. If you opt into wallet spillover, usage past your budget is billed from your wallet at published catalog rates up to a monthly cap you set. Wallet credit you pay for never expires; bonus credit on larger packs expires after 12 months.
How is inference cost calculated?
Every request is priced from its tokens at the per-model rates below: uncached input, cached input at a quarter of the input rate, and output. Plans burn that amount from your budget; wallet usage is charged it. A request the fleet declines for capacity is never charged and never touches your budget.
Can I change plans?
Yes, from billing settings in your account. Upgrades take effect immediately and your usage so far in the current windows carries over to the new budgets. Downgrades take effect at the start of your next billing cycle. If your subscription lapses, your API keys keep working on wallet billing at catalog rates until you subscribe again.
What is the parallel request limit?
Each plan caps how many requests can run at once across all of your keys: 1 on Solo, 3 on Pro, 6 on Max. Extra requests queue until one finishes. Higher plans are also routed first when the fleet is busy, which means being refused last under load, not faster responses.
Can my team share one plan?
No. A self-service plan is for one person and their own agents and apps. Sharing keys across a team, reselling plan capacity, or opening a second account to stack budgets is prohibited and leads to suspension. Teams that need shared capacity should contact us about team or dedicated options.
What happens to my prompts?
Nothing is kept. Cobble does not store prompts or completions for any request, on any plan, and never trains on them. We keep request metadata only (model, token counts, cost, status, timing) for billing and abuse detection. All inference runs on hardware we operate in the United States.
Where are the binding terms?
Section 18 of the Terms of Service describes plans, budgets, windows, the wallet and plan changes in full; Part B of the Acceptable Use Policy covers fair use; the Privacy Policy covers data handling. The numbers on this page come from the same registry the API enforces.