Every response from an AI API comes back with a usage block: how many tokens went in, how many came out, sometimes how many were served from cache. Developers read those numbers constantly, because they are the numbers the bill is built on.
Imagine the same block carried one more field: the electricity this request consumed, in watt-hours, measured rather than guessed. Not a fleet-wide average quoted in a sustainability report once a year, but a per-request figure that showed up in the dashboard next to latency and cost, broke down by model and by application, and could be summed over a month the way spend is summed today.
That is the direction Cobble is building toward. This article is about why it is harder than it sounds, how we think about doing it honestly, and what becomes possible once the number exists.
Why the number is hard
The naive approach is to take a server's power draw, divide by the tokens it produced, and call the result energy per token. It is wrong in at least four ways, and each one has to be dealt with separately.
Idle power. An accelerator draws a substantial amount of power doing nothing. If a node sits at low utilization for an hour and then serves a burst, who pays for the idle hour? Charge it to nobody and the per-request figure flatters the fleet. Charge it all to the burst and a single busy minute looks catastrophic. The honest answer is an allocation rule, applied consistently and disclosed: idle power is a real cost of keeping capacity available, and it belongs to the requests that capacity was kept available for. The practical consequence is that the single largest lever on energy per token is utilization. A fleet run full spreads its idle cost thin. A fleet run half empty doubles the energy cost of every token it serves, without changing the hardware at all.
Shared hardware. Modern inference engines batch many requests together on the same accelerator, interleaving them at every decode step. At any instant a device is working for dozens of customers at once. Attributing its power to any one request means attributing a share, and the fair share depends on how much of the batch that request occupied and for how long. This is a solvable accounting problem, but it means per-request energy is a derived figure, built from measured device power and measured per-request compute share, not something read directly off a meter.
Input versus output. Prompt tokens and generated tokens cost very different amounts of energy. Processing a prompt is a parallel operation that uses the hardware efficiently. Generating output is sequential, one token per step, and is where most of the time and energy of a long response goes. A cached prompt prefix costs less again, because the work was done once and reused. Any energy figure that reports a single blended number per token is hiding the thing a developer could actually act on.
Hardware efficiency. Different accelerators and different pools in the fleet have different energy per unit of work, and the same model served on different hardware consumes differently. A per-request figure has to come from the device that actually served the request, which means the measurement has to follow the routing decision rather than assume an average.
How we intend to measure it
Cobble's fleet already measures power draw and temperature per node alongside utilization, latency, and error rate, and the operations tooling already treats a power anomaly with the same seriousness as an error spike. The raw signal exists. The work is in attribution.
The approach we are building is layered. At the bottom is measured device power, sampled continuously. Above that is a per-request record of which pool and which device served the request, how long it occupied the batch, and how many prompt, cached, and generated tokens it involved; the usage metering that drives billing already records most of this. The attribution layer combines the two: a request's energy is its measured share of device power over the interval it was active, plus its proportional share of the idle power of the capacity it drew on. The result is written next to the usage record and surfaces in the same places spend does.
Two principles govern it. The figure is measured, not modeled from a datasheet or an industry average, because a reclaimed fleet with deliberately high utilization is exactly the kind of fleet where averages are most misleading. And the allocation rules are published, so that a developer comparing two providers can see what the number includes rather than take it on faith. The same discipline already applies to the sustainability figures on the network, which are reported as measured shares of completion power, with the solar contribution stated as a measured percentage and the absence of evaporative cooling stated as a fact rather than an offset.
What a developer can do with it
The point of the number is not the number. It is what changes once developers can see it.
Prompt design becomes an energy decision. If a dashboard shows that a long system prompt re-sent on every call accounts for most of an application's energy, the fix is obvious and cheap: structure the prompt so the stable prefix is cached. The developer already had a cost reason to do this. The energy figure gives them a second reason and shows them the result.
Output length gets scrutiny. Generated tokens are the expensive ones, in both dollars and watt-hours. An application that asks for verbose responses it then truncates is wasting energy twice. Seeing the generation share of each request makes the waste visible.
Model choice becomes informed. A smaller model that solves a task adequately uses a fraction of the energy of a larger one. Today that trade-off is made on cost and quality. With per-model energy in the dashboard it can be made on all three.
Workloads can be scheduled. Batch jobs that are not time-sensitive can be run when the fleet has spare capacity and when the on-site solar share is highest. A developer who can see energy per request can also see when it is cheapest in the way that matters.
Over time, this is the foundation for an application to carry an energy budget the way it carries a spending budget: a target, a trend, and the information needed to meet it.
Why it belongs in the open
There is a reason this has not been standard across the industry. Per-request energy accounting is unflattering to fleets run at low utilization, to providers whose numbers rest on offsets, and to anyone whose sustainability claims would not survive a per-token audit. Transparency is easy to promise in a report and much harder to deliver in an API response.
Cobble's position is that the figure should exist, should be measured, and should be given to the developer along with the tokens and the price. It is the same philosophy that leads the network to report rejections with their reasons, to expose budget windows in response headers, and to state its sustainability numbers as measured shares. A developer should be able to see what their requests cost, in every currency that matters.
The energy cost of a token is one of those currencies. We intend to put it on the bill.
The Cobble Team

