Cobble hosts a curated catalog of open-source models, quantized to FP8, served on reclaimed hardware. The catalog is the single source of truth — what /v1/models returns is exactly what the API accepts.
Discover models programmatically
bash
curl https://api.cobble.network/v1/models \
-H "Authorization: Bearer $COBBLE_API_KEY"Every agent and SDK that supports OpenAI-compatible providers can read this endpoint. Model IDs are stable — we never rename a published ID.
Recommended defaults
Don't want to evaluate the whole catalog? Start here:
| You're doing | Use | Why |
|---|---|---|
| General agents / tool use | qwen/qwen3.8-27b | The flagship: strongest reasoning and tool calling in the catalog, 128K context |
| Agentic coding | deepseek/deepseek-v4-flash-0731 | MoE coding model with repo-scale reasoning, 256K context |
| Long documents, cheap loops | qwen/qwen3.8-flash-next | 200K context at a fraction of the flagship output price |
| Coding on a budget window | qwen/qwen3.6-35b-a3b | MoE efficiency — near-flagship coding at a fraction of the compute |
| Writing, chat, roleplay | google/gemma-4-26b-a4b-it | Best prose quality per dollar of budget |
| Fast and cheap everything | mistralai/mistral-nemo | Lowest cost per token in the catalog |
| Embeddings | google/embeddinggemma-300m | Covered by your plan's embedding allowance; overage from the wallet |
| Document OCR | z-ai/glm-ocr or deepseek/deepseek-ocr-2 | Billed from the wallet, per 1K pages |
Which models come with which plan
Every model in the catalog is included on every plan. Plans differ in how much usage they include per window, how many requests can run in parallel, and context length, not in which models you can call. See Limits & billing.
