ISSUE 2026-09-27SERIES BREC 26TYPE BRIEFgarnetgrid.com
The bill is a shadow of the architecture
Insights
Most cost programmes deliver a spreadsheet rather than a smaller invoice. The difference between the two is whether the change was structural or whether it quietly depended on somebody continuing to care. This is about which reductions hold, which decay, and how to tell them apart before you spend a quarter finding out.
A billing console groups spend by service, which invites you to believe the problem is settings. It usually isn't. Every line item has two factors: the price of the thing, and the amount of the thing you consume.
The price factor is instance family, region, storage class, commitment level, support tier. It is negotiable, it is quick to work on, and it is quickly exhausted. There is a floor and you will hit it.
The amount factor is how many calls a request makes, how much data crosses a zone boundary, how many copies you keep and for how long. That is architecture, and it is decided in design reviews, often years before anyone looks at the bill.
Price work lowers the intercept once. Architecture changes the slope. A programme that only does price work will show a step down in the invoice and then resume the same growth rate. That is the failure mode this piece is about; I have no figure for how common it is.
Why the savings evaporate
Four distinct mechanisms, and they need different responses.
First, one-off reductions misread as trend changes. Deleting a decade of orphaned snapshots is real money, but it happens once. It tells you nothing about next year.
Second, savings that require continuous attention. Rightsizing is accurate on the day you do it and decays from then on, because workloads change and nobody re-runs the analysis. Any reduction whose maintenance plan is "we'll keep an eye on it" is on a timer.
Third, savings that were displacements. Work moved onto a cheaper service that now needs glue code and a person who understands it. The cloud bill fell, engineering time rose, and engineering time does not appear on the invoice you are being judged on.
Fourth, savings that consumed margin for error. Cut the headroom, cut the replica, shorten the retention, and the first serious incident reverses the decision in an afternoon. What almost never happens is the savings tracker being updated to match. Months later the dashboard still claims a reduction that was rolled back.
There is also an accounting failure underneath all of this. If savings are measured against a forecast rather than against last quarter's actual spend, the claim is unfalsifiable: the bill can rise while the programme reports success. Decide which baseline you are using before you start, and prefer the one someone can check.
You cannot cut what you cannot attribute
Attribution comes first, and it is boring, and skipping it is why most of this work produces opinions instead of decisions.
Two things are needed. One is a defensible rule for splitting shared cost, because the expensive parts of an estate are usually shared: the cluster, the database, the observability pipeline, the network. Perfect allocation is not achievable and chasing it can cost more engineering than the spend being allocated. Pick a proxy you can defend, publish it, and then keep it stable, because a rule that changes every quarter destroys the trend you were trying to read.
The other is a denominator. Cost per tenant, per job, per thousand documents indexed, per active user. Without one, a rising cost graph and a growing business are indistinguishable, and you will either panic about healthy growth or miss genuine regression.
Then put the unit cost where the people who can change it already look. A finance dashboard that no engineer opens has never changed an architecture. The number belongs next to latency and error rate, on the team's own dashboard, owned by the team.
Where the money usually actually is
Estates differ, so treat this as a list of places to point an instrument rather than a set of conclusions. What these have in common is that they are consumption problems, not price problems.
Elasticity is insurance, and commitments are a bet
On-demand and serverless pricing carry a premium per unit of work. You are buying the option not to forecast. That is a good purchase for spiky, unpredictable or brand-new workloads whose shape you genuinely do not know yet, and a poor one for flat, predictable base load, where you are paying a premium for flexibility you never exercise.
Committed-use and reserved pricing invert the trade. Lower unit cost, and you absorb the forecast risk plus a subtler cost: you have partly bought the right to keep your current architecture, because re-architecting can strand the commitment. Discount shapes and terms vary by provider and change without much notice, so do not design an estate around a figure you read in a blog post last year. Get current terms for your own account and your own region.
The usual mistake is not picking the wrong option, it is picking one philosophy for the entire estate. A committed base, an on-demand shoulder for variability, and interruptible capacity for genuinely interruptible batch is a reasonable default shape. The word doing the work there is "genuinely": most batch code claims it tolerates being killed mid-run, and a fair amount of it does not.
AI workloads break the capacity-planning habit
Metered inference does not behave like compute. Cost per request is a function of prompt size, how much context you stuff in, output length, retry policy and user behaviour. Request volume alone no longer predicts spend, so the capacity-planning instinct of "traffic times unit cost" stops working. Two teams with the same request count can sit far apart because one of them pastes an entire document into every call.
That makes the levers engineering levers rather than procurement ones. Retrieve context instead of stuffing it. Cache aggressively, including at the semantic layer if your traffic repeats. Route by difficulty, so a small model handles the easy majority and escalation is the exception. Cap output length, because unbounded generation is unbounded spend. Make retries bounded and idempotent. Batch anything that is not interactive.
On owning hardware versus renting it, this is the kind of system we build, and we still cannot give you a crossover point. It depends on your duty cycle, how long you keep the hardware useful, power and space, and whether you have someone who can operate it. Rented capacity is clearly better for lumpy, low-utilisation work. Owned capacity wins somewhere above a utilisation threshold that is specific to you. Anyone who quotes you the threshold without your numbers is selling, and that includes us.
The honest case for owned hardware is usually not a unit-cost argument anyway. It is residency, egress, and a bill that does not move. The last one has a second-order effect worth naming: when there is no per-query meter, teams experiment more freely. That is sometimes the real value and it is genuinely hard to put a number on.
Make the reduction structural, or don't bother
Budget alerts are not controls. An alert tells you that you have already spent the money. Controls are the things that make the expensive path harder than the cheap one.
In practice that means quotas and service limits set deliberately rather than inherited. TTLs on ephemeral resources with an automated reaper that actually deletes. Retention and lifecycle defaults baked into the module everybody uses, so the cheap choice is the default choice and the expensive one requires an argument. A cost estimate attached to infrastructure changes at review time, while the decision is still cheap to change. And the unit-cost graph on the team dashboard, so regressions surface in days rather than at the next quarterly review.
The general rule: fix the creation path, not the resources. Anything you solve with a clean-up sprint will need another clean-up sprint, and by then the people who understood the first one will have moved on.
Do this in order. Get attribution good enough to argue with and pick one denominator, because without those you are guessing. Do the one-off clean-ups, take the money, and then stop counting them as a programme. Choose the two structural changes with the steepest slope — on the estates I have looked at those are usually data movement and observability volume, but measure yours rather than taking that ordering from me — and push them into the defaults so the next team inherits them for free. Then judge the whole effort on the invoice three months later, compared against last quarter's actual spend rather than a projection. If you cannot produce that comparison, you did a clean-up, not cost work. Both are worth doing, but only one of them holds.