There is a moment in the life of every infrastructure project when the question changes. For a long time the question is "does it work?" and the answer is a demo: a model loads, a prompt goes in, a response comes out, and everyone in the room is pleased. Then the first real customer arrives, and the question becomes "does it work at three in the morning, under load, when a node is overheating and someone's agent has sent four hundred requests in a minute?"
Cobble crossed that line this year. This article is an account of what the crossing looked like: what was built, what broke, what was learned, and what standard the network is now holding itself to as it opens to a wider developer ecosystem.
Where it started
The earliest configurations were exactly what you would expect from a team bringing up reclaimed laboratory hardware for the first time. Nodes were brought online one at a time. Driver versions were pinned by trial and error until a combination worked. Models were loaded by hand, benchmarked by hand, and served through a single endpoint with no routing, no queue, and no idea what would happen if two people used it at once.
That phase produced the first real lesson. The datasheet throughput of a reclaimed accelerator is a number from another era, measured on a workload that is not yours. The only figures that mattered were the ones produced by loading each pool with the models it would actually serve, increasing concurrency until latency started to climb, and recording where that happened. Every capacity number the network uses today, for planning, for sales, and for admission control, descends from that measurement rather than from a specification.
The second lesson arrived with the second serving engine. Different model architectures want different things from the hardware: a sparse mixture-of-experts model needs its full weight footprint resident and is cheap per token once it is; a dense model wants decode speed and cache; a vision model has its own memory profile. No single engine served all of them well on the hardware available. The fleet became a set of pools, each organized around an architecture, each running the engine and weight format that suits it, and the relay in front of them became responsible for sending each request to the right one.
Building the relay
The relay is the layer a developer actually talks to: an OpenAI-compatible endpoint that authenticates the key, checks limits, picks a pool, forwards the request, streams the response, and records what happened. It sounds simple. Most of the engineering in the network lives there, and most of the lessons came from it.
Refusals must say why. Early on, every failure was a generic error. A developer could not tell whether they had hit their own limit, whether the fleet was full, or whether something was broken. The fix was a taxonomy: every rejection carries a source, a machine-readable code, whether the request was charged, and when to retry. A plan-limit refusal and a fleet-capacity refusal are different events with different correct responses, and the network now says which one occurred, in the body and in a header, on every single refusal.
Fail closed, within bounds. An enforcement layer that cannot reach its own state has to decide whether to let traffic through or stop it. Letting it through indefinitely is how abuse happens; stopping it instantly is how a brief blip becomes an outage. The network now fails open for a bounded window, then refuses with an explicit code that tells the developer the platform, not their account, is the problem.
Unknown keys are refused at the edge. A key the network does not recognize is rejected before it touches a single accelerator, with a code that says exactly why. Nothing reaches the fleet that has not been authenticated and checked against a plan.
Reserve, then settle. Metering a request after it finishes means a burst of concurrent requests can all pass the limit check before any of them is charged. The network now estimates a request's cost at admission, reserves it, and settles to the actual figure when the response completes, refunding anything the fleet refused. This was the single change that made usage limits mean what they say.
What a plan actually sells
Agent workloads changed what a subscription has to mean. An autonomous agent can consume in an afternoon what a human user consumes in a month, and a plan defined by request counts alone says nothing about the share of the fleet a customer is really getting.
Cobble's plans are therefore denominated in usage value over fixed windows, a short window for burst and a longer one for sustained use, so that the plan a customer buys corresponds to a share of the fleet they can actually have. The shape of those windows was chosen by replaying real usage patterns through the proposed rules before they went live, to confirm that ordinary workloads fit comfortably and that only the heaviest would notice the ceiling.
The principle is honesty in both directions. A plan that promises more than the fleet can deliver is a lie to the customer; a plan that quietly throttles is a lie by omission. The budget windows are shown in response headers on every request, so a developer always knows where they stand.
Models, conformance, and what gets published
Loading a model is not the same as being able to sell it. Before a model appears in the public catalog, it is tested for the behaviors developers rely on: tool calling that actually produces well-formed calls, structured output that honors the schema it was given, prompt caching that returns the right prefix, streaming that terminates cleanly. Several models that served text perfectly failed one or more of these under test, and the catalog reflects what passed rather than what loaded.
The same discipline applies to the catalog as a whole. Models that cannot be served at production quality are not listed, even when they run. Capacity that exists only as a fallback is not counted as sellable capacity. A model's price and limits are published once, consistently, in every place a developer might read them.
Monitoring and operating
A production fleet is a stream of things going slightly wrong. The operations tooling the network runs today watches each node's power draw, temperature, utilization, latency, and error rate, and treats a power drop or a hot node as seriously as a spike in failures. Every request is logged with its routing decision, its timing, and its outcome, and that log is what the usage metering, the energy accounting, and the capacity planning all read from.
Over the year, this grew into an operating loop rather than a dashboard: scheduled checks that compare the fleet's state against rules, alerts that reach a person when a rule is breached, and a morning report that summarizes health, capacity, security events, and subscription load so that the day starts with the facts. The relay also gates new subscriptions against measured capacity, closing signups automatically when the fleet cannot honor more commitments, so that growth never outruns the hardware.
The standard the network is working toward
Opening to a broader developer ecosystem means holding to standards that were optional when the customer base was small. The ones the network is working toward are concrete.
- Every refusal is explained. No developer should ever receive a generic error from the network.
- Published rules. Rate limits, budget windows, pricing, retention, and the energy allocation rules behind sustainability figures are documented in one place and match what the code does.
- Measured capacity. Every pool is load-tested on the models it serves, and nothing is sold beyond what the measurement supports.
- Zero retention, verified. Request contents are not stored after a response is served, and the network's logging and vendor configuration are reviewed to confirm it.
- Honest sustainability. Energy figures are measured shares of completion power, not offsets, and the goal is to put a per-request energy figure in the dashboard.
- Security as practice. Credentials are rotated, secrets are kept out of code and history, access is reviewed, and the morning report includes a security section every day.
None of these are finished in the sense that infrastructure is ever finished. They are the bar the network measures itself against, and the bar it intends to be held to by the developers who build on it.
What the transition taught
The hardest part of going from experiment to production was not any single system. It was accepting that every convenient shortcut from the experimental phase, the generic error, the datasheet number, the untested assumption, would eventually be found by a real workload and fail in front of a customer. Each one was replaced with something that tells the truth: about capacity, about limits, about cost, and about what went wrong.
That is what production means. Not that nothing breaks, but that when it does, the developer knows exactly what happened and what to do next.
The Cobble Team

