Qwen3.8 27B
Alibaba Cloud
Cobble's flagship dense model: strong reasoning, code generation and tool use, with prefix caching that keeps long agent sessions cheap.
Catalog of supported open-weight models across families including Meta (Llama ecosystem), Alibaba Cloud Qwen, Mistral AI, DeepSeek, NVIDIA Nemotron, IBM Granite, Nomic, and Cobble-built pipelines. Pricing assumes reclaimed GPU infrastructure, efficient vLLM serving, and open-weight licensing—benchmarked against typical marketplace rates.
Labels such as Flagship, Enterprise ready, Multilingual, and Best for RAG appear as tags on each card.
Featured
Cobble's flagship dense model: strong reasoning, code generation and tool use, with prefix caching that keeps long agent sessions cheap.
$0.35 / 1M input tokens · $2.75 / 1M output tokens · $0.0875 / 1M cached tokens
Featured
Efficiency-optimized Mixture-of-Experts model (284B total / 13B active) built for fast inference over a 1M-token context window.
$0.191 / 1M input tokens · $0.506 / 1M output tokens · $0.0478 / 1M cached tokens
Featured
Google's compact multilingual embedding model for semantic search and RAG. Covered by every plan's monthly embedding allowance; overage is billed from the wallet.
$0.01 / 1M input tokens
Alibaba Cloud
Cobble's flagship dense model: strong reasoning, code generation and tool use, with prefix caching that keeps long agent sessions cheap.
Alibaba Cloud
Fast, long-context member of the Qwen3.8 family for agent loops and large documents at a fraction of the flagship output price.
Alibaba Cloud
Sparse MoE architecture delivering high quality responses with excellent cost efficiency.
Alibaba Cloud
Low-latency utility model ideal for chatbots, summarization, and lightweight automation.
DeepSeek
Efficiency-optimized Mixture-of-Experts model (284B total / 13B active) built for fast inference over a 1M-token context window.
Google DeepMind
Large open model with excellent instruction following, multilingual capabilities, and coding performance.
Google DeepMind
Efficient sparse variant of Gemma optimized for strong quality with lower serving costs.
Google DeepMind
Unified encoder-free multimodal model handling text, image, audio, and video with strong quality at small-model cost.
Google DeepMind
Efficient 4B-class model tuned for high-volume, low-latency workloads like classification and extraction.
Google DeepMind
Ultra-light 2B-class model for massive-scale pipelines, routing, and lightweight chat at minimal cost.
Ornith AI
Sparse MoE agentic coding model (about 3B active parameters per token) from the Ornith 1.5 family, with strong tool-use and self-scaffolding behavior.
Ornith AI
Compact coding model from the Ornith 1.5 family that punches above its size on agentic coding tasks at a 9B price.
Poolside
Poolside's software-engineering model, built for repo-scale coding agents, tool calling and long diffs.
Xiaomi
Distilled 9B reasoning model with strong math and code performance for its size.
inclusionAI
Lightweight MoE model from the Ling family for fast, inexpensive chat, extraction and classification.
Liquid AI
Liquid Foundation Model with about 1B active parameters per token: very low latency for routing, classification and edge-style workloads.
IBM
IBM's enterprise-oriented 8B instruct model with strong instruction following and permissive Apache 2.0 licensing.
IBM
Hybrid Mamba/transformer tiny model from IBM Granite 4.0 for high-throughput, low-cost pipelines.
Mistral AI
Beloved creative workhorse with natural prose, strong multilingual range, and dependable instruction following.
OpenAI
OpenAI's open-weight MoE model (3.6B active) with strong reasoning and tool use at very low cost.
Unsloth
30B model served for creative and conversational workloads.
Cobble Labs
Gemma4 26B-based context compressor that summarizes long tool histories for Hermes Agent; usable from any OpenAI-compatible client.
Zhipu AI
General-purpose OCR model with strong support for complex layouts and multilingual documents.
DeepSeek
High-accuracy OCR and document understanding model optimized for tables and technical PDFs.
Google DeepMind
Google's compact multilingual embedding model for semantic search and RAG. Covered by every plan's monthly embedding allowance; overage is billed from the wallet.