All news
Cobble news · October 2026

Why Open-Weight Models Deserve First-Class Infrastructure

Open-weight models have evolved from experimental alternatives into powerful tools for software development, research, reasoning, and autonomous agents. Yet access to them often depends on infrastructure optimized for a small selection of commercially dominant systems. Here is why diverse model catalogs matter, how different architectures create different serving requirements, and how Cobble's multi-model approach lets developers choose models for the task rather than for the limits of a provider.

A glowing green mesh of connected nodes, representing a network of open-weight models

Open-weight models have evolved from experimental alternatives into powerful tools for software development, research, reasoning, and autonomous agents. The best of them now hold their own against closed systems on the work developers actually do: reading a repository, planning a multi-step change, calling tools reliably, writing structured output that a program can parse. They are released under permissive licenses, in a range of sizes, by labs on three continents.

Yet access to these models often depends on infrastructure that was designed for something else. Most inference stacks were built around a relatively small selection of commercially dominant systems: one or two dense model families, one serving engine, one set of assumptions about prompt length and batch shape. Open-weight models get bolted onto that stack afterward, and the result is familiar. A capable model shows up on a provider's list, served at a quantization nobody chose on purpose, behind a queue tuned for a different workload, with tool calling that works until it does not.

This article is about why that happens, why it matters, and what we do differently.

A catalog is a set of commitments, not a list

It is easy to publish a long model list. It is hard to serve one well, because every model on the list is a promise about behavior: that tool calls will come back in the format the client expects, that JSON mode will actually constrain the output, that a 128K prompt will not quietly truncate at 32K, that a streaming response will carry the usage block a billing system needs.

We treat the catalog as a set of those commitments. Every model Cobble serves is run through a conformance suite before it is listed and again whenever the serving layer changes. The suite exercises the things agents depend on, including tool calling, structured output, long-context behavior, streaming, and prompt caching, and it produces a per-model matrix we use internally to decide what goes on the public list and at what context length. When a model fails a check, it does not ship with an asterisk. It waits.

That discipline is what lets us say something simple to developers: if a model is in the catalog, it works the way the API says it does.

Different architectures want different infrastructure

The deeper reason open-weight models are underserved is that they are not one thing. The catalog Cobble runs today spans dense transformers in the 20 to 30 billion parameter range, sparse mixture-of-experts models that activate a few billion parameters per token out of a hundred-billion-plus total, hybrid-attention designs built for very long context, compact edge-class models in the single-digit billions, vision-language models, document OCR models, and embedding models. Each of these has a different shape at serving time.

A dense model wants fast decode and a deep prompt cache, because its cost is dominated by generating tokens. A mixture-of-experts model is cheap per token but hungry for memory, so it rewards a serving pool that is sized for its full weight footprint rather than its active parameters. Long-context models live or die on key-value cache management: whether a 100K-token prompt can be held, reused across turns, and shared between parallel requests from the same session. Small models are a throughput problem, not a latency problem, and belong on hardware that would be wasted on a flagship. Vision and OCR models have prefill patterns that look nothing like chat.

Infrastructure built for one of these shapes serves the others badly. That is not a criticism of any particular provider. It is a statement about why "we also host open models" so rarely means "we host them well."

How Cobble serves a diverse catalog

Cobble's fleet is organized around the models, not the other way round. A few of the principles behind it:

Per-model serving pools. Each model family runs in a pool sized and configured for its architecture. The flagship dense models get the pool with the deepest cache and the most decode headroom. Sparse and long-context models get memory. Small models share a high-throughput pool where many replicas can be hot at once. Nothing is forced through a one-size-fits-all deployment.

Cache-affinity routing. Agentic workloads reuse the same prompt prefix across dozens of turns. Our router keeps a session's follow-up requests on the replica that already holds its context, so warm requests skip the work they have already done. In practice a large share of input tokens on the network are served from cache, which is why cached input is priced at a fraction of the uncached rate.

Predictive admission. Rather than queueing requests until something times out, the fleet estimates whether a request can be served within its service deadline and declines early if it cannot. A declined request is never charged. Clients get a clear, machine-readable reason, and capacity is spent on requests that will complete.

Measured capacity, not assumed capacity. Every pool is load-tested to find the concurrency at which latency degrades, and that measured figure, not a datasheet number, is what the planner and the sales side work from. When we add hardware, we measure it before we sell against it.

Multiple engines where it helps. Some models are best served by one inference engine, some by another, and some by different engines for different context tiers. The relay in front of the fleet presents one OpenAI-compatible API regardless, including the usage reporting, headers, and error semantics that routers and agents expect.

Reclaimed hardware, renewable power. All of this runs on GPUs that other operators had already written off, in facilities with access to renewable energy and without evaporative cooling. Serving open-weight models well does not require the newest silicon. It requires matching the model to the right pool and keeping that pool busy.

Why diversity is worth the trouble

It would be simpler to run one flagship and call it a catalog. We do not, because the developers we serve do not have one kind of problem.

A coding agent working through a large repository wants a strong dense model with long context and a warm cache. A classification pipeline processing a million short documents wants the cheapest competent small model with the highest throughput. A research harness extracting tables from scanned PDFs wants an OCR model with a vision fallback. A chat product wants a fast mid-size model with reliable tool calling and strict JSON. The right answer differs by task, and often by step within a task.

A diverse catalog on infrastructure that takes each model seriously lets a developer pick the model for the job and route between them freely. A narrow catalog, or a wide one served carelessly, forces the choice the other way around: pick whatever the provider happens to run well, then bend the task to fit.

What this means for the ecosystem

Open-weight models are a public good. Labs release them so that people can build on them, inspect them, fine-tune them, and run them on their own terms. That promise only holds if there is infrastructure willing to serve the long tail of them properly, with the same care a closed model gets from its own vendor.

Cobble's multi-model approach reflects a belief that developers should choose models based on the task at hand rather than the limitations of a particular provider. We will keep adding models as they earn a place in the catalog, keep publishing what each one can and cannot do, and keep building the fleet around the models rather than squeezing the models into the fleet.

If you build with open-weight models and have been waiting for them to be treated as first-class, that is what we are here for.

The Cobble Team

Sustainable inference on reclaimed hardware — built for the communities that use it.