Generative AI
Curated catalog · Quantized & benchmarked
Models Available From Cobble
Generative AI
DeepSeek V4 Flash
FEATUREDGenerative AI
Qwen3.8 Flash Next
Generative AI
Qwen3.6 35B A3B
Generative AI
Qwen3.5 9B
Generative AI
Gemma4 31B
Generative AI
Gemma4 26B A4B
Generative AI
Gemma4 12B
Generative AI
Gemma4 E4B
Generative AI
Gemma4 E2B
Generative AI
Ornith 1.5 35B A3B
Generative AI
Ornith 1.5 9B
Generative AI
Laguna S 2.1
Generative AI
MiMo V2.6 Distill 9B
Generative AI
Ling 3.0 Tiny
Generative AI
LFM2.5 8B A1B
Generative AI
Granite 4.1 8B
Generative AI
Granite 4.0 H Tiny
Generative AI
Mistral Nemo 12B
Generative AI
GPT-OSS 20B
Generative AI
Muse Glimmer 30B
Generative AI
Hermes Compressor
OCR
GLM-OCR
OCR
DeepSeek OCR2
Embeddings
EmbeddingGemma 300M
FEATUREDGenerative AI
Qwen3.8 27B
FLAGSHIPGenerative AI
DeepSeek V4 Flash
FEATUREDGenerative AI
Qwen3.8 Flash Next
Generative AI
Qwen3.6 35B A3B
Generative AI
Qwen3.5 9B
Generative AI
Gemma4 31B
Generative AI
Gemma4 26B A4B
Generative AI
Gemma4 12B
Generative AI
Gemma4 E4B
Generative AI
Gemma4 E2B
Generative AI
Ornith 1.5 35B A3B
Generative AI
Ornith 1.5 9B
Generative AI
Laguna S 2.1
Generative AI
MiMo V2.6 Distill 9B
Generative AI
Ling 3.0 Tiny
Generative AI
LFM2.5 8B A1B
Generative AI
Granite 4.1 8B
Generative AI
Granite 4.0 H Tiny
Generative AI
Mistral Nemo 12B
Generative AI
GPT-OSS 20B
Generative AI
Muse Glimmer 30B
Generative AI
Hermes Compressor
OCR
GLM-OCR
OCR
DeepSeek OCR2
Embeddings
EmbeddingGemma 300M
FEATUREDThe Cobble difference
Not your typical inference provider
Traditional AI providers burn megawatts in massive datacenters. We built something different.
vs
Traditional
Cobble
Infrastructure
Massive datacenters
Distributed edge nodes
Power Source
Grid-dependent megawatts
Renewable-ready regions
Cooling
Evaporative water cooling
No evaporative cooling
Hardware
Proprietary enterprise GPUs
Reclaimed GPUs & servers
Carbon Footprint
High manufacturing churn
No new-silicon manufacturing
Model Focus
Full precision only
Per-model quantization
Receipts, not promises
Built to do better
Every component was sourced, recycled, and repurposed.
0
Water Usage
0%
Green Energy
0%
Recycled Hardware
-0x
Carbon Footprint
Open numbers · Open weights · Open methodology
Quantized. Benchmarked. Real.
We publish what others hide. Every model is tested, quantized, and documented.
Qwen3.8 27B
FP8Context
128K tokens
Throughput
55 tokens/sec
DeepSeek V4 Flash
FP8Context
256K tokens
Throughput
See catalog
EmbeddingGemma 300M
FP8Context
2K tokens
Throughput
See catalog
Three steps to inference
How it works
01
Choose a model
Pick from our curated selection of quantized models optimized for speed and quality on recycled hardware.
02
Send a request
Use our OpenAI-compatible API. Drop-in replacement for your existing inference pipeline.
03
Get results
OpenAI-compatible responses from distributed edge nodes running vLLM.



















