The physical shape of AI infrastructure is being decided right now, and it is being decided in one direction. Computation is concentrating into a small number of enormous facilities, each drawing hundreds of megawatts, built and operated by a handful of companies, and sited wherever land, power, and tax treatment line up. The model you call from a laptop in one state is most likely answered from a building in another, through a company whose decisions about pricing, availability, and data handling you have no say in.
This is presented as inevitable. It is not. It is the result of choices about what inference infrastructure should look like, and those choices can be made differently.
Cobble's position is that a great deal of inference should happen closer to the people who use it, in smaller installations, operated by organizations that are part of the communities they serve. This article makes that case.
Why centralization is winning
It is worth being honest about why the hyperscale model dominates, because the arguments for it are real.
Frontier training requires tens of thousands of accelerators in one place, connected by networks that only make sense at scale. The economics of power purchasing favor very large buyers. Operational expertise is scarce and concentrates where the most machines are. And once a company has built the facility, it is natural to serve inference from the same place.
But inference is not training. An inference request needs one model resident on a handful of accelerators, a few hundred milliseconds of compute, and a path back to the user. Nothing about that requires a hundred-megawatt building. The reasons inference is served from hyperscale facilities today are convenience and incumbency, not physics. It is the one part of the AI stack that could be distributed, and it is the part that touches users directly.
What a local installation looks like
A local inference installation, in Cobble's model, is a modest facility: a room or a small building, a few racks of accelerators, a power supply measured in tens or hundreds of kilowatts rather than megawatts, and a network connection to the region it serves. It runs open-weight models through a serving stack that routes each request to the pool best suited to it, manages a queue so that overload is refused cleanly rather than allowed to degrade everything, and meters usage so that the operator can bill for it.
Critically, it can be built from hardware that already exists. The Cobble Network runs on reclaimed computing equipment, including accelerators retired from the supercomputing ecosystem around Oak Ridge, Tennessee. A facility of this size does not need the newest silicon. It needs capable hardware, run at high utilization, with software engineered around what the hardware does well. That combination has already been proven on the network, where a completion uses roughly a third of the energy of a typical industry equivalent.
It can also be powered from sources that already exist locally. On-site solar supplies a growing measured share of the fleet's completion power today. A facility sized in kilowatts rather than megawatts can be matched to a rooftop array, a small-scale hydro installation, a municipal utility's surplus, or an industrial site's underused electrical service, without needing a new substation or a decade of permitting.
Resilience
A centralized architecture has centralized failure modes. When a major cloud region has an outage, a substantial fraction of the AI applications in the world stop working at the same moment, regardless of where their users are. A fiber cut, a power event, or a configuration error in one facility becomes everyone's problem.
A network of regional installations fails differently. Each one serves its region, and when one is unavailable the others continue. The Cobble Network already works this way: requests are routed across three physically separate sites, and the relay steers around capacity that is unhealthy or overloaded at any one of them. Fleet refusals are a first-class, explicitly labeled outcome rather than a mysterious error, so that applications can respond intelligently when a site cannot take a request. Adding a fourth site, or a fortieth, extends a mechanism that is already in production rather than inventing one.
Resilience also means something simpler. Latency to a facility in your own region is lower than to one across the country, and lower latency matters more every year as applications become agentic and make dozens of calls per task rather than one.
Data sovereignty
Where inference happens determines whose laws govern the data that flows through it. An organization in one jurisdiction sending prompts to a facility in another is subject to both sets of rules, and to the data-handling practices of a provider it did not choose and cannot inspect.
Local inference changes that. A hospital system, a school district, a municipal government, a regional bank, or a law firm can route its workloads to an installation in its own jurisdiction, operated under terms it can read and audit. Cobble's own network already operates on a zero-retention basis, with request contents not stored after a response is served. Extending that model to installations that are physically local to the organizations using them makes the guarantee concrete in a way that a privacy policy from a distant provider never can.
Local economic development
The economic benefits of a hyperscale data center flow mostly to its owner. A community that hosts one receives some construction jobs, a small number of permanent positions, and a large new demand on its grid and water supply.
A local inference installation is a different kind of economic actor. It is small enough to be owned and operated by a regional company, a cooperative, a university, or a public utility. The revenue it earns from serving inference stays in the region. The skills required to run it, in power, cooling, networking, and serving infrastructure, are skills that can be taught locally and that are valuable far beyond the facility. And the hardware it reclaims is hardware that would otherwise have left the region as scrap.
This is the heart of the broader vision. The AI economy, as currently structured, treats most of the world as consumers of services delivered from a few places. A network of local installations lets communities participate as producers: operating capacity, serving their neighbors, and earning from it.
A network, not a building
Cobble's vision is a network in which installations like these are connected by shared infrastructure: a common relay that accepts requests through a standard interface and routes them to the best available capacity, common metering and billing so that an operator can be paid for what they serve, common standards for refusal, queueing, and health so that every installation behaves predictably, and common accounting for energy so that the environmental cost of each request is measured wherever it is served.
The pieces of that network already exist. The relay, the multi-site routing, the queueing, the metering, the budget enforcement, the energy measurement, and the serving stack that runs open-weight models on reclaimed hardware are all in production across the Cobble Network's three sites today.
The next installations are being planned with the communities that will own them. Cobble is in active discussions with communities in Western Kentucky and Eastern Kentucky about building municipal inference infrastructure: locally owned installations, running on the same network, serving their own regions and earning from it. The work ahead is to make what already runs on the network the foundation those communities can build on, so that the next installation does not have to reinvent any of it.
Inference should be closer to the people who use it. The technology to put it there exists. What remains is to build the network.
The Cobble Team

