Pricing

Subscription plans with 5-hour and weekly usage budgets and wallet spillover — plus per-model API rates synced live with our catalog for pay-as-you-go usage.

Flat-rate plans

One price, a generous usage budget per window, and a dashboard that shows exactly what is left. Every plan runs on recycled GPUs in regions with access to renewable energy.

Every number on this page is the binding one. The same budgets, limits and plan rules appear in Section 18 of the Terms of Service, and there is no hidden fair-use quota.

Solo

$20/mo

For one person and their agent. Every model in the catalog, a 5-hour usage budget that resets all at once, and a weekly budget that keeps a runaway loop from eating the month. Budgets are measured in dollars of inference at our published rates, so a small model or a cached prompt barely dents them.

$1.40 of usage per 5-hour window, $14 per week

Every model in the catalog (27)

1 request at a time

128K context

10M embedding tokens / mo

Up to 3 active API keys

Most popular

Pro

$40/mo

For daily agent work. Two and a half times the Solo budget, three parallel requests so tool calls do not queue behind each other, and 256K context for repo-scale prompts. The plan most people building with opencode, Hermes or Cursor end up on.

$3.50 of usage per 5-hour window, $35 per week

Every model in the catalog (27)

3 parallel requests

256K context

50M embedding tokens / mo

Up to 5 active API keys

Max

$100/mo

For agents that never log off. Six times the Solo budget, six parallel requests, routed first when the fleet is busy. Built for always-on assistants and multi-agent pipelines that run while you sleep. If you will use more than this, talk to us about dedicated capacity.

$8.40 of usage per 5-hour window, $84 per week

Every model in the catalog (27)

6 parallel requests, routed first when the fleet is busy

256K context

150M embedding tokens / mo

Up to 10 active API keys

A La Carte

No subscription — fund a wallet, pay published rates per token. Credit you pay for never expires. Bonus credit on larger packs is spent first and expires after 12 months.

Top-up presets

$10$25$50$100$250

Custom amounts $5–$500

How limits work

  • Usage budgets — Each plan includes a budget of inference per 5-hour window and per week (Solo: $1.40 / 5h, $14 / week · Pro: $3.50 / 5h, $35 / week · Max: $8.40 / 5h, $84 / week), measured in dollars at our published per-token rates: uncached input, cached input at a quarter of the input rate, and output, as listed in the rate table below. A small model or a cached prompt uses far less of your budget than a large uncached one.
  • Fixed windows — Your first request starts the 5-hour window; when it ends, the whole budget comes back at once. The weekly window works the same way over 7 days and caps the total of all 5-hour windows. A request must fit in both. Windows are fixed, not sliding: nothing refills gradually, and unused budget does not roll over.
  • Parallel requests — Each plan caps how many requests can run at once (Solo: 1 · Pro: 3 · Max: 6, routed first). Extra agent loops queue until one finishes. “Routed first” means your requests are refused last when the fleet is busy. It is a capacity preference, not a speed guarantee.
  • Backstops — A request ceiling per window, a requests-per-minute limit, and a cap on active API keys (Solo: 1,500 requests / 5h, 30/min, 3 keys · Pro: 4,000 requests / 5h, 60/min, 5 keys · Max: 9,000 requests / 5h, 120/min, 10 keys). They stop floods of near-free requests and key sharing; normal use never reaches them.
  • When you hit a limit — You get an HTTP 429 that names the window or limit, when it resets, and whether anything was charged. A request the fleet declines for capacity is never charged and never touches your budget.
  • Wallet spillover — Off by default. Turn it on and set a monthly cap, and usage past your budget is billed from your prepaid wallet at catalog rates, up to that cap. Credit you pay for never expires; bonus credit on larger packs is spent first and expires after 12 months. Unused credit is forfeited if you close your account.
  • Plan changes — Upgrades take effect immediately; downgrades at your next billing cycle. If your subscription lapses, your keys keep working on wallet billing at catalog rates. A failed renewal gets a 7-day grace period before the plan ends.
  • One person, one plan — Plans are for a single person and their own agents. No reselling plan capacity, no sharing keys across a team, no second account to stack budgets. Teams that need shared capacity should talk to us.
  • The binding versions of these rules are Section 18 of the Terms of Service, Part B of the Acceptable Use Policy, and the Limits & billing guide. Your prompts and completions are never stored; see how we handle data.

Sustainability, included

Recycled GPUs

Built from reclaimed GPUs and server hardware — extending useful life instead of manufacturing new silicon.

Renewable-ready

Deployed in regions with access to renewable energy.

No evaporative cooling

Designed to avoid evaporative water cooling.

Per-model API rates

Pay-as-you-go token, page, and endpoint pricing — synced live with our catalog.

Generative AI

ModelProviderInput / 1M tokOutput / 1M tokCached / 1M tokPages / 1KEndpoint
Qwen3.8 27BFLAGSHIP
Alibaba Cloud$0.35/1M$2.75/1M$0.0875/1M—/v1/chat/completions
DeepSeek V4 FlashFEATURED
DeepSeek$0.191/1M$0.506/1M$0.0478/1M—/v1/chat/completions
Qwen3.8 Flash Next
Alibaba Cloud$0.15/1M$0.47/1M$0.0375/1M—/v1/chat/completions
Qwen3.6 35B A3B
Alibaba Cloud$0.15/1M$1/1M$0.0375/1M—/v1/chat/completions
Qwen3.5 9B
Alibaba Cloud$0.1/1M$0.15/1M$0.025/1M—/v1/chat/completions
Gemma4 31B
Google DeepMind$0.12/1M$0.37/1M$0.03/1M—/v1/chat/completions
Gemma4 26B A4B
Google DeepMind$0.06/1M$0.33/1M$0.015/1M—/v1/chat/completions
Gemma4 12B
Google DeepMind$0.05/1M$0.15/1M$0.0125/1M—/v1/chat/completions
Gemma4 E4B
Google DeepMind$0.06/1M$0.12/1M$0.015/1M—/v1/chat/completions
Gemma4 E2B
Google DeepMind$0.05/1M$0.1/1M$0.0125/1M—/v1/chat/completions
Ornith 1.5 35B A3B
Ornith AI$0.07/1M$0.7/1M$0.0175/1M—/v1/chat/completions
Ornith 1.5 9B
Ornith AI$0.1/1M$0.15/1M$0.025/1M—/v1/chat/completions
Laguna S 2.1
Poolside$0.1/1M$0.2/1M$0.025/1M—/v1/chat/completions
MiMo V2.6 Distill 9B
Xiaomi$0.1/1M$0.15/1M$0.025/1M—/v1/chat/completions
Ling 3.0 Tiny
inclusionAI$0.02/1M$0.11/1M$0.005/1M—/v1/chat/completions
LFM2.5 8B A1B
Liquid AI————/v1/chat/completions
Granite 4.1 8B
IBM————/v1/chat/completions
Granite 4.0 H Tiny
IBM————/v1/chat/completions
Mistral Nemo 12B
Mistral AI$0.02/1M$0.03/1M$0.005/1M—/v1/chat/completions
GPT-OSS 20B
OpenAI$0.029/1M$0.14/1M$0.0073/1M—/v1/chat/completions
Muse Glimmer 30B
Unsloth————/v1/chat/completions
Hermes Compressor
Cobble Labs$0.06/1M$0.33/1M$0.015/1M—/v1/chat/completions

OCR

ModelProviderInput / 1M tokOutput / 1M tokCached / 1M tokPages / 1KEndpoint
GLM-OCR
Zhipu AI———$0.08/v1/ocr
DeepSeek OCR2
DeepSeek———$0.08/v1/ocr

Embeddings

ModelProviderInput / 1M tokOutput / 1M tokCached / 1M tokPages / 1KEndpoint
EmbeddingGemma 300MFEATURED
Google DeepMind$0.01/1M———/v1/embeddings

Wallet top-up

Fund your wallet for spillover and pay-as-you-go usage. Credit you pay for never expires. Larger packs include bonus credit, which is spent first and expires after 12 months.

FAQ

How do usage budgets work?

Each plan includes a budget of inference per 5-hour window and per week, measured in dollars at the per-model rates below. Your first request starts the 5-hour window; when it ends, the whole budget comes back at once. The weekly window works the same way over 7 days, and a request must fit in both. Your dashboard shows what is left in each and when it resets.

What happens when I use up a budget?

Requests get a 429 that names the window and its reset time, and nothing is charged. If you opt into wallet spillover, usage past your budget is billed from your wallet at published catalog rates up to a monthly cap you set. Wallet credit you pay for never expires; bonus credit on larger packs expires after 12 months.

How is inference cost calculated?

Every request is priced from its tokens at the per-model rates below: uncached input, cached input at a quarter of the input rate, and output. Plans burn that amount from your budget; wallet usage is charged it. A request the fleet declines for capacity is never charged and never touches your budget.

Can I change plans?

Yes, from billing settings in your account. Upgrades take effect immediately and your usage so far in the current windows carries over to the new budgets. Downgrades take effect at the start of your next billing cycle. If your subscription lapses, your API keys keep working on wallet billing at catalog rates until you subscribe again.

What is the parallel request limit?

Each plan caps how many requests can run at once across all of your keys: 1 on Solo, 3 on Pro, 6 on Max. Extra requests queue until one finishes. Higher plans are also routed first when the fleet is busy, which means being refused last under load, not faster responses.

Can my team share one plan?

No. A self-service plan is for one person and their own agents and apps. Sharing keys across a team, reselling plan capacity, or opening a second account to stack budgets is prohibited and leads to suspension. Teams that need shared capacity should contact us about team or dedicated options.

What happens to my prompts?

Nothing is kept. Cobble does not store prompts or completions for any request, on any plan, and never trains on them. We keep request metadata only (model, token counts, cost, status, timing) for billing and abuse detection. All inference runs on hardware we operate in the United States.

Where are the binding terms?

Section 18 of the Terms of Service describes plans, budgets, windows, the wallet and plan changes in full; Part B of the Acceptable Use Policy covers fair use; the Privacy Policy covers data handling. The numbers on this page come from the same registry the API enforces.

Questions? Documentation · Terms of Service · Acceptable Use · Privacy