AI Data Center Economics · Sizing and Placement
Published September 2026
~
“The facility team owns the numbers. The workloads set them.”
Ask what makes a data center good and you get two answers. The operator names utilization, cooling, and the power bill. The customer names performance, cost, and reliability. It is tempting to treat these as opposing interests — one side squeezing efficiency, the other demanding headroom. They are not opposing. They are the same pair of decisions seen from two ends: which workload runs where, and on what.
Everything that matters inside the building — the heat at a rack, the traffic on a link, the draw on a circuit, the work returned per watt — is downstream of those two answers. Which is worth stating plainly, because it is the part most optimization effort skips: heat is not a thing a data center has. It is a thing a workload produces. Same for the draw on a circuit, the load on a link, and the gap between the capacity you installed and the capacity you can sell.
That distinction decides what optimization can accomplish. Tune the room — setpoints, fan speeds, grille layout, chilled water temperature — and you are working on the effect while the cause continues producing it. The returns are real and they are bounded, because the plant can only respond to the thermal picture it is handed. Change which workload sits where, and on what silicon, and you change the picture the plant is handed in the first place. One approach adjusts the building. The other adjusts what the building is being asked to do.
The lever is never the facility. It is the workload: where it sits, and what it sits on.
A workload carries a physical signature
We treat a workload as an abstract unit of demand: so many cores, so much memory, some storage. But every workload also leaves a physical mark. It runs with a certain efficiency, and in the running it produces heat, pulls power, and moves data. That mark isn’t one number you can slide up and down — it’s several things at once:
Heat is the thermal load it adds to its rack and its cooling zone. Power is the draw it places on a shared circuit and power domain. Network is the bandwidth it consumes and the path its traffic takes. Efficiency is how much useful work it returns per watt and per dollar. Two workloads that look identical on a capacity sheet can behave nothing alike once placed. One runs cool and quiet. The other lights up a thermal hot spot and saturates a link.
Location changes the signature
The mark is not fixed, because it depends on where the workload lands — the neighbors it sits beside, the cooling it draws from, the power domain it shares, the path its traffic takes. Move the same workload a few racks over and its effect on the building changes. Change what surrounds it and its own efficiency changes with it. The axes interact, which is why you cannot place a workload by looking at one number and stapling it to an open slot.
Two decisions: right-placed and right-sized
Placement and sizing are separate decisions, and they are easy to conflate.
Right-placed is the question of position — which rack, which node, which power and cooling domain, on which path. This is where hot spots are created or avoided, where links are balanced or bottlenecked, where a circuit runs comfortably or flirts with its limit. Right-placed puts a workload in the right seat.
But a seat is not a fit. Right-sized is the other decision — which silicon the workload actually runs on, and how much of it. A training job on an accelerator whose memory bandwidth it cannot saturate, an inference service holding a GPU tier it never draws on, a rendering job on hardware that finishes in four times the necessary time: none of those are placement failures. The workload is in a perfectly reasonable seat. It is the wrong size for the seat.
Right-placed is not enough
Sizing and placement are independent, which is exactly why treating them as one thing fails. A workload can be badly sized in a well-balanced hall — nothing around it is stressed, and it still returns less work per watt and per dollar than it should, because it never fit the silicon it was given. And a workload can be correctly sized and still run badly, because a hot spot or a saturated link around it throttles the hardware it was correctly matched to.
Both show up as the same symptom. The job is slower and more expensive than it should be. But they demand opposite corrections: one workload needs to be resized where it sits, the other needs to be moved without changing its shape at all. Resize the throttled workload and you have hidden a facility problem behind a bigger bill. Move the misfit workload and you have relocated the inefficiency without touching it.
This is why the useful step is not detection but attribution — separating a workload that is on the wrong silicon from one that is on the right silicon in the wrong conditions. Detection tells you a workload is underperforming, which is the part that is already visible on any dashboard. Attribution tells you which of the two decisions to revisit, and it is the only thing standing between a measurement and a correction. Until you can tell those apart, every corrective action is a guess.
Wrong size, wrong place: both strand capacity
Installed capacity is not usable capacity, and the gap between them is set by these decisions. Rightsizing is usually pitched as a customer concern — the tenant spends less. That framing is why operators have been slow to care about it, and it is wrong. Every one of these errors takes capacity out of circulation, and all of it lands on the operator’s side of the ledger.
Over-sizing strands capacity by holding it. A workload sitting on an allocation it never draws on has taken that silicon off the market, and no one else can use it. The rack looks committed and the hall looks full while a meaningful share of the fleet does nothing. Utilization — the number the whole business is measured on — degrades without any visible cause, because a booked GPU and a working GPU are indistinguishable on a capacity report.
Misplacement strands it a second way, and this one is harder to see, because the capacity is genuinely free. It just cannot be sold. A workload that concentrates heat in a cooling zone consumes that zone’s thermal budget, so the empty slots around it cannot be filled without breaching it. A workload sitting on a circuit already near its limit leaves the remaining power on that circuit unusable. A job whose collective traffic saturates a link makes the ports behind it worth less to anyone else. The GPUs are installed, powered, unallocated — and unsellable, because the conditions they need were spent by a neighbor. Capacity reports will show them as available inventory right up until someone tries to place a workload there.
Under-sizing stresses power and thermal. A workload squeezed onto hardware it has outgrown does not draw less overall; it pins that hardware at its ceiling and holds it there. Sustained peak utilization is the least efficient regime a device operates in, and it is the regime that generates the most heat per unit of useful work. The job also takes longer, so the thermal and power burden is not just higher but extended. Undersized workloads are a durable source of exactly the hot spots that then strand the capacity around them — which is how a sizing error becomes a placement problem two racks away.
And the tenant is paying for that at the same time, in a way that has nothing to do with the operator’s power bill. An under-provisioned workload takes longer to finish, misses latency targets, and returns a worse cost per result than it would on hardware that fit — and it does all of that even in a facility whose thermal and electrical characteristics were never affected. The two injuries are independent. Fix the cooling and the customer is still waiting; comp the customer and the hardware is still running pinned at its ceiling. Under-sizing is not one problem seen from two angles. It is two problems from one cause, and the operator carries both.
The obvious objection is that an over-sized customer is still paying for what they hold. True — and it is still the worse outcome. The same workload would have run on a smaller footprint at the same cost per result, because that is how these workloads scale. The difference is what happens to the capacity it was not using: rightsized, that silicon goes back into inventory and gets sold to a workload that can actually saturate it. The operator collects on both. Left as it is, the operator collects once and holds hardware that is contributing nothing to either party.
“Rightsizing a tenant does not cost the operator revenue. It frees inventory to sell a second time.”
That logic holds wherever demand exceeds supply, which describes the market these facilities were built for. And the tenant left over-provisioned is not a stable customer anyway — a workload quietly running at a fraction of what it pays for is a repricing conversation waiting to happen, or a churn risk, depending on who notices first.
So the operator has a first-order stake in both decisions, not a courtesy interest in them. Over-sizing strands capacity by holding it. Misplacement strands the capacity next to it. Under-sizing costs power, cooling, and hardware life, creates the conditions that strand capacity in turn, and delivers a worse result to the customer while it does so. None of it is the customer’s problem alone, and none of it is the facility’s problem alone.
Why the two sides stop competing
Get both right and temperature, bandwidth, and power spread evenly instead of piling up. The facility runs closer to its true capacity, because nothing is stranded behind an artificial hot spot, a saturated link, or an allocation nobody is using. The workload runs closer to its best, because it is not fighting its neighbors for thermal or network headroom and is not fighting hardware that was never the right match.
That is the whole point. Efficiency and performance stop being a trade-off you balance and start being the same thing. The operator gets higher effective utilization, less cooling overhead, fewer capacity surprises — more useful work out of every watt and every square foot. The customer gets consistent performance, predictable cost, fewer noisy-neighbor problems. Neither result was bought at the other’s expense. Both came from putting the right workload in the right place, at the right size.
Where naive placement fails
Most placement does none of this. It reads a capacity requirement, finds an open slot the workload nominally fits, and stops. It does not weigh heat, power, or network path. It does not ask what the workload does to its neighbors or what they do to it. It scores on fit, not on effect, and it runs once against a snapshot.
“It is not placing the workload where it belongs. It is dropping it in the first slot a capacity check does not reject.”
What it is worth
Four returns, and they compound.
Lower bills. Cooling is the largest operating cost after the IT load itself, and a plant that has to chase its worst rack runs to that rack rather than to the hall’s average. Removing the hot spots created by concentrated placement, and the sustained ceiling-pinned operation created by undersizing, lets the same plant hold the same hall at less expense. The saving is recurring and it requires no new equipment.
Better revenue. Capacity recovered is inventory recovered, and the gap between installed capacity and usable capacity is consistently found to be large. Every point of it recovered is silicon that can be sold to demand that already exists, without a single additional megawatt procured.
Happier customers. A tenant that is right-sized and right-placed gets what it is paying for — predictable performance, a cost per result that reflects the work done, and none of the noisy-neighbor behavior that generates support tickets nobody can explain. Retention in this market is won on delivered outcomes, not on rate cards, and the operator who improves a customer’s results without discounting anything has bought loyalty at no margin cost.
Happier investors. These facilities are underwritten on revenue per megawatt and on how quickly deployed capital starts earning. Recovering capacity improves the numerator without touching the denominator, and it defers the next buildout — the single largest call on capital an operator faces. A facility that sells more of what it already has is a better asset than one that answers demand by breaking ground.
None of these four are traded against each other. That is the tell that the underlying decision was right.
This is what the Serra Labs Platform does. It sizes every workload against how it actually behaves, not against what it requested — identifying the configuration that matches the goal in force, whether that is saving cost, preserving performance, or improving it, weighing the expected benefits and drawbacks of every change, and validating each recommendation against the workload’s own history before anyone acts on it. Sizing is judged in the conditions the workload actually runs in: a workload straining against its surroundings is not treated as though it were running clean, so a constrained workload is not mistaken for an oversized one, and the correction points at the right problem.