Models

Catalog of supported open-weight models across families including Meta (Llama ecosystem), Alibaba Cloud Qwen, Mistral AI, DeepSeek, NVIDIA Nemotron, IBM Granite, Nomic, and Cobble-built pipelines. Pricing assumes reclaimed GPU infrastructure, efficient vLLM serving, and open-weight licensing—benchmarked against typical marketplace rates.

Labels such as Flagship, Enterprise ready, Multilingual, and Best for RAG appear as tags on each card.

Featured picks

Featured

Qwen3.8 27B

Cobble's flagship dense model: strong reasoning, code generation and tool use, with prefix caching that keeps long agent sessions cheap.

$0.35 / 1M input tokens · $2.75 / 1M output tokens · $0.0875 / 1M cached tokens

Featured

DeepSeek V4 Flash

Efficiency-optimized Mixture-of-Experts model (284B total / 13B active) built for fast inference over a 1M-token context window.

$0.191 / 1M input tokens · $0.506 / 1M output tokens · $0.0478 / 1M cached tokens

Featured

EmbeddingGemma 300M

Google's compact multilingual embedding model for semantic search and RAG. Covered by every plan's monthly embedding allowance; overage is billed from the wallet.

$0.01 / 1M input tokens

Generative AI

Qwen3.8 27B

Alibaba Cloud

FLAGSHIP

Cobble's flagship dense model: strong reasoning, code generation and tool use, with prefix caching that keeps long agent sessions cheap.

Context window128K tokens
Throughput55 tokens/sec
QuantizationFP8
Pricing$0.35 / 1M input tokens · $2.75 / 1M output tokens · $0.0875 / 1M cached tokens
Supported endpoint/v1/chat/completions
ReasoningCodingAgents

Qwen3.8 Flash Next

Alibaba Cloud

Fast, long-context member of the Qwen3.8 family for agent loops and large documents at a fraction of the flagship output price.

Context window200K tokens
QuantizationFP8
Pricing$0.15 / 1M input tokens · $0.47 / 1M output tokens · $0.0375 / 1M cached tokens
Supported endpoint/v1/chat/completions
FastLong ContextAgents

Qwen3.6 35B A3B

Alibaba Cloud

Sparse MoE architecture delivering high quality responses with excellent cost efficiency.

Context window256K tokens
Throughput85 tokens/sec
QuantizationFP8
Pricing$0.15 / 1M input tokens · $1 / 1M output tokens · $0.0375 / 1M cached tokens
Supported endpoint/v1/chat/completions
MoeFastCost Efficient

Qwen3.5 9B

Alibaba Cloud

Low-latency utility model ideal for chatbots, summarization, and lightweight automation.

Context window256K tokens
Throughput90 tokens/sec
QuantizationFP8
Pricing$0.1 / 1M input tokens · $0.15 / 1M output tokens · $0.025 / 1M cached tokens
Supported endpoint/v1/chat/completions
FastEconomyUtility

DeepSeek V4 Flash

DeepSeek

FEATURED

Efficiency-optimized Mixture-of-Experts model (284B total / 13B active) built for fast inference over a 1M-token context window.

Context window256K tokens
QuantizationFP8
Pricing$0.191 / 1M input tokens · $0.506 / 1M output tokens · $0.0478 / 1M cached tokens
Supported endpoint/v1/chat/completions
MoeCodingAgents

Gemma4 31B

Google DeepMind

Large open model with excellent instruction following, multilingual capabilities, and coding performance.

Context window256K tokens
Throughput49 tokens/sec
QuantizationFP8
Pricing$0.12 / 1M input tokens · $0.37 / 1M output tokens · $0.03 / 1M cached tokens
Supported endpoint/v1/chat/completions
MultilingualInstruction Following

Gemma4 26B A4B

Google DeepMind

Efficient sparse variant of Gemma optimized for strong quality with lower serving costs.

Context window256K tokens
Throughput90 tokens/sec
QuantizationFP8
Pricing$0.06 / 1M input tokens · $0.33 / 1M output tokens · $0.015 / 1M cached tokens
Supported endpoint/v1/chat/completions
MoeEconomyGeneral Purpose

Gemma4 12B

Google DeepMind

Unified encoder-free multimodal model handling text, image, audio, and video with strong quality at small-model cost.

Context window256K tokens
Throughput70 tokens/sec
QuantizationFP8
Pricing$0.05 / 1M input tokens · $0.15 / 1M output tokens · $0.0125 / 1M cached tokens
Supported endpoint/v1/chat/completions
MultimodalEconomyGeneral Purpose

Gemma4 E4B

Google DeepMind

Efficient 4B-class model tuned for high-volume, low-latency workloads like classification and extraction.

Context window128K tokens
Throughput130 tokens/sec
QuantizationFP8
Pricing$0.06 / 1M input tokens · $0.12 / 1M output tokens · $0.015 / 1M cached tokens
Supported endpoint/v1/chat/completions
FastEconomyUtility

Gemma4 E2B

Google DeepMind

Ultra-light 2B-class model for massive-scale pipelines, routing, and lightweight chat at minimal cost.

Context window128K tokens
Throughput160 tokens/sec
QuantizationFP8
Pricing$0.05 / 1M input tokens · $0.1 / 1M output tokens · $0.0125 / 1M cached tokens
Supported endpoint/v1/chat/completions
FastEconomyUtility

Ornith 1.5 35B A3B

Ornith AI

Sparse MoE agentic coding model (about 3B active parameters per token) from the Ornith 1.5 family, with strong tool-use and self-scaffolding behavior.

Context window256K tokens
QuantizationFP8
Pricing$0.07 / 1M input tokens · $0.7 / 1M output tokens · $0.0175 / 1M cached tokens
Supported endpoint/v1/chat/completions
CodingAgentsMoe

Ornith 1.5 9B

Ornith AI

Compact coding model from the Ornith 1.5 family that punches above its size on agentic coding tasks at a 9B price.

Context windowTBA
QuantizationFP8
Pricing$0.1 / 1M input tokens · $0.15 / 1M output tokens · $0.025 / 1M cached tokens
Supported endpoint/v1/chat/completions
CodingCompact

Laguna S 2.1

Poolside

Poolside's software-engineering model, built for repo-scale coding agents, tool calling and long diffs.

Context window256K tokens
QuantizationFP8
Pricing$0.1 / 1M input tokens · $0.2 / 1M output tokens · $0.025 / 1M cached tokens
Supported endpoint/v1/chat/completions
CodingAgents

MiMo V2.6 Distill 9B

Xiaomi

Distilled 9B reasoning model with strong math and code performance for its size.

Context window128K tokens
QuantizationFP8
Pricing$0.1 / 1M input tokens · $0.15 / 1M output tokens · $0.025 / 1M cached tokens
Supported endpoint/v1/chat/completions
ReasoningCompact

Ling 3.0 Tiny

inclusionAI

Lightweight MoE model from the Ling family for fast, inexpensive chat, extraction and classification.

Context window128K tokens
QuantizationFP8
Pricing$0.02 / 1M input tokens · $0.11 / 1M output tokens · $0.005 / 1M cached tokens
Supported endpoint/v1/chat/completions
FastCost EfficientMoe

LFM2.5 8B A1B

Liquid AI

Liquid Foundation Model with about 1B active parameters per token: very low latency for routing, classification and edge-style workloads.

Context window32K tokens
QuantizationFP8
Pricing—
Supported endpoint/v1/chat/completions
FastMoe

Granite 4.1 8B

IBM

IBM's enterprise-oriented 8B instruct model with strong instruction following and permissive Apache 2.0 licensing.

Context window128K tokens
QuantizationFP8
Pricing—
Supported endpoint/v1/chat/completions
EnterpriseGeneral Purpose

Granite 4.0 H Tiny

IBM

Hybrid Mamba/transformer tiny model from IBM Granite 4.0 for high-throughput, low-cost pipelines.

Context window128K tokens
QuantizationFP8
Pricing—
Supported endpoint/v1/chat/completions
FastCost Efficient

Mistral Nemo 12B

Mistral AI

Beloved creative workhorse with natural prose, strong multilingual range, and dependable instruction following.

Context window128K tokens
Throughput75 tokens/sec
QuantizationFP8
Pricing$0.02 / 1M input tokens · $0.03 / 1M output tokens · $0.005 / 1M cached tokens
Supported endpoint/v1/chat/completions
CreativeMultilingualEconomy

GPT-OSS 20B

OpenAI

OpenAI's open-weight MoE model (3.6B active) with strong reasoning and tool use at very low cost.

Context window128K tokens
Throughput120 tokens/sec
QuantizationFP8
Pricing$0.029 / 1M input tokens · $0.14 / 1M output tokens · $0.0073 / 1M cached tokens
Supported endpoint/v1/chat/completions
MoeReasoningOpen SourceFast

Muse Glimmer 30B

Unsloth

30B model served for creative and conversational workloads.

Context window32K tokens
QuantizationFP8
Pricing—
Supported endpoint/v1/chat/completions
Creative

Hermes Compressor

Cobble Labs

Gemma4 26B-based context compressor that summarizes long tool histories for Hermes Agent; usable from any OpenAI-compatible client.

Context window256K tokens
QuantizationFP8
Pricing$0.06 / 1M input tokens · $0.33 / 1M output tokens · $0.015 / 1M cached tokens
Supported endpoint/v1/chat/completions
AgentsUtility

OCR

GLM-OCR

Zhipu AI

General-purpose OCR model with strong support for complex layouts and multilingual documents.

Context windowUp to 200 pages per batch
QuantizationFP8
Pricing$0.08 / 1K pages
Supported endpoint/v1/ocr
DocumentsMultilingual

DeepSeek OCR2

DeepSeek

High-accuracy OCR and document understanding model optimized for tables and technical PDFs.

Context windowUp to 250 pages per batch
QuantizationFP8
Pricing$0.08 / 1K pages
Supported endpoint/v1/ocr
PdfTablesTechnical

Embeddings

EmbeddingGemma 300M

Google DeepMind

FEATURED

Google's compact multilingual embedding model for semantic search and RAG. Covered by every plan's monthly embedding allowance; overage is billed from the wallet.

Context window2K tokens
QuantizationFP8
Pricing$0.01 / 1M input tokens
Supported endpoint/v1/embeddings
MultilingualRag

OpenAI-compatible routes where noted — see Docs · Sign up