All news
Cobble news · October 2026

The Economics of Inference: Why Efficiency Matters More Than Ever

As language models become more capable and inference demand grows, the economics of serving them are becoming the whole game. This article explores the hidden costs of inference, including electricity, GPU utilization, idle capacity, memory, and the overhead of long-running agent workflows, and makes the case that the next generation of providers will compete not just on model quality or token prices, but on the efficiency of their entire operating systems.

Layered green contour lines forming rolling terrain, representing the cost curves of inference

For most of the last three years, the story of language models was a story about capability. Each release did something the last one could not, and the question worth asking about a provider was which models it had. That question is losing its edge. The best open-weight models are now good enough for most of the work developers bring them, several providers serve the same weights, and a model that was remarkable in spring is a commodity by autumn.

What has not become a commodity is the cost of running them. Inference demand is growing faster than the hardware to serve it, agents have turned occasional requests into continuous load, and the price of a token has fallen to the point where the margin on serving one is decided by details most customers never see. The economics of inference are becoming the whole game. This article is about where the money actually goes, and why we think the providers that survive will be the ones that treat efficiency as the product.

The costs nobody lists on the pricing page

A per-token price hides a stack of costs that have very little to do with tokens.

Electricity is the obvious one, and it is less obvious than it looks. A GPU serving a model draws most of its rated power whether it is producing a thousand tokens a second or waiting for the next request. The cost that matters is not watts per GPU but watts per useful token, and that number is dominated by how busy the hardware is kept.

Idle capacity is the quiet killer. Inference traffic is bursty. A fleet sized for the peak sits mostly empty, and a fleet sized for the average falls over at the peak. Every provider lives somewhere on that line, and the position they pick sets their cost per token more than any other decision. Capacity that is paid for and not used is the single largest expense in most serving operations, and it never appears on an invoice.

Memory is the constraint that decides what can run at all. A model's weights have to fit, and then the working memory for every in-flight request has to fit beside them. Long prompts are expensive not because they take long to read but because they occupy memory that could have served other requests. Sparse models that activate a fraction of their parameters per token still need the whole model resident. Memory, not compute, is usually what a pool runs out of first.

Prefill is the cost most pricing ignores. Reading a prompt into the model is a large burst of computation that produces no output tokens. An agent that re-sends a hundred-thousand-token context on every turn is asking the fleet to redo that work every time, and a provider that cannot avoid redoing it is paying for it whether the price list reflects it or not.

Long-running agents combine all of the above. They hold memory for the life of a session, they re-send context, they run at night when nothing else does and at noon when everything does, and they retry on every error. They are the most valuable workload a provider can win and the most expensive to serve badly.

Two ways to respond

Faced with those costs, the industry's default response has been to buy its way out: newer accelerators, larger clusters, more of everything. That works, for the handful of operators who can finance it, and it has a hard ceiling. Hardware spending scales linearly with capacity; it does not change the ratio of useful work to wasted work. A fleet of the newest GPUs idling at low utilization, re-reading the same prompts and timing out under bursts, is an expensive fleet that is inefficient in exactly the same ways as a cheap one.

The other response is to extract more useful computation from the infrastructure that already exists. This is the harder path, because the gains come from a hundred decisions rather than one purchase order, and it is the path Cobble was built on. The fleet runs on GPUs that other operators had already retired. That is not a constraint we work around; it is the discipline that forced us to care about every one of the costs above, because we could not outspend them.

What efficiency looks like in practice

A few of the ways the fleet turns existing hardware into more useful tokens:

Caching what has already been computed. Agentic traffic repeats itself. The same system prompt, the same repository context, the same conversation prefix arrives turn after turn. Cobble's router keeps a session's requests on the replica that already holds its context, so warm requests skip the prefill they have already paid for. A large share of the input tokens the network serves come from cache. That is why cached input is priced at a quarter of the uncached rate: the discount is not a promotion, it is the cost structure made visible, and it rewards the clients that are cheapest to serve.

Keeping the hardware busy. Idle capacity is the largest hidden cost, so the fleet is operated at high utilization and shaped so that the load has somewhere to go. Admission control refuses work the fleet cannot complete in time rather than letting it sit in a queue consuming memory, and a declined request is never charged. Work that is accepted is work that will finish.

Matching models to pools. A dense flagship, a sparse mixture-of-experts model, and a small classifier have nothing in common at serving time. Each runs in a pool configured for its architecture, so a small model's throughput is not spent on a flagship's memory profile and a flagship's cache is not fragmented by thousands of tiny requests. Right-sizing is the cheapest efficiency gain there is, and the one most often skipped.

Measuring instead of assuming. Every pool is load-tested to the point where latency degrades, and that measured figure is what capacity planning uses. Hardware is added when the measurements say it is needed, and re-measured before anyone sells against it. Datasheet capacity is a story; measured capacity is a budget.

Power that is not fought over. Efficiency per token is one half of the energy equation; where the energy comes from is the other. Part of the fleet's power already comes from on-site solar, every completion uses a fraction of the energy of a typical industry equivalent, and no evaporative cooling is used anywhere. These are not offsets bought elsewhere. They are properties of the infrastructure, reported against real load.

Pricing that follows the cost

Efficiency only compounds if the pricing model points customers in the same direction as the infrastructure. This is where most of the industry's flat-rate plans went wrong: a price that does not move with the work invites the work that costs the most. A plan measured in requests makes a hundred-thousand-token uncached prompt cost the same as a greeting, so that is what the heaviest users send.

Cobble's plans are measured in dollars of inference at published per-token rates, per window. A small model, a short prompt, or a warm cache uses little of the budget; a large uncached prompt uses more. The customer's cheapest choice and the fleet's cheapest choice are the same choice, which means every optimization an agent operator makes to stretch their budget also lowers the cost of serving them. Aligned incentives are the most durable efficiency there is, because they recruit the customer into the engineering.

The competition ahead

Model quality will keep improving and keep converging. Token prices will keep falling toward the cost of serving them. Neither is a place a provider can stand for long.

What is left is the operating system: how much useful work a fleet extracts from each watt, each gigabyte of memory, and each GPU-hour it has already paid for; how it handles the relentless, repetitive, bursty load that agents generate; and whether its pricing makes customers partners in efficiency or adversaries of it. The next generation of inference providers will be ranked on that, whether or not it ever appears on a benchmark.

We built Cobble from hardware the industry had written off, in part because it was the right thing to do with it and in part because it taught us to run an inference business the way one will have to be run. Efficiency is not a feature we added. It is the reason the fleet exists.

The Cobble Team

Sustainable inference on reclaimed hardware — built for the communities that use it.