ISSUE 2026-09-27SERIES BREC 42TYPE BRIEFgarnetgrid.com

The question behind the job title

Insights

Platform engineering is not a tooling choice, and it is not the operations team with new business cards. It is a decision about which parts of your path to production become a product with an owner, a contract and a support boundary, and which parts stay with the teams shipping features. This is what that looks like in practice, including the places it reliably goes wrong.

Most people asking what platform engineering means are asking something narrower: should we have a team for this, and what should it own? The title arrived after a decade in which "you build it, you run it" was taken literally, and every product team ended up maintaining its own pipeline, its own infrastructure-as-code, its own alerting and its own opinions about secrets. That is survivable while there are only a handful of teams. Past that, the cost shows up as a long tail of infrastructure that exactly one person understands and nobody maintains, and as four different answers to "how do I get a service into production".

Platform engineering is the response to that: take the parts of the path to production that every team needs, make them somebody's product, and hold that product to a contract. It is not a synonym for Kubernetes. It is not a portal. It is a decision about where a boundary sits and who answers for it when it breaks.

A platform is an interface with an owner and a promise

There is a simple test for whether you have a platform or just a collection of shared tooling. Three questions:

If any of those is no, what you have is a toolbox and a wiki page. That is not worthless, but it will not produce the effects people expect from a platform, because the thing doing the work in a platform is the promise, not the tooling. A promise has a scope, an owner, a support expectation and a change policy. Tooling has none of those and cannot acquire them by being put in a shared repository.

The failure mode I run into most often is a rename. An operations team becomes a platform team, keeps working out of a ticket queue, and the only thing that changed is the org chart. The tell is mechanical: look at how work arrives. If most of it arrives as a request for a human on your team to go and do a thing, you are running a service desk. Self-service is the mechanism that makes a platform different, not the word on the slide.

The paved road has to be optional to be any good

Every platform team eventually wants to mandate the golden path, usually for defensible reasons: consistency, auditability, one thing to patch. Mandating it early is a mistake, and not mainly for cultural reasons. It is a measurement problem. Adoption is the only honest feedback signal you have. If teams are compelled to use the path, you cannot tell whether the path is good, and you will not find out until something expensive happens on it.

Keep the escape hatch, and instrument it. The interesting data is not how many teams are on the road. It is which teams leave, and at what point they leave. A team that adopts your build and deploy path but writes its own alerting is telling you where your product stops.

The other half of this is coverage. A paved road needs to handle the boring majority of real workloads, not the demonstration case. If the path takes a stateless HTTP service from repository to production cleanly but has nothing to say about a job that needs an accelerator, a stateful store, or an unusual ingress, then the teams with the hardest problems, which is to say the teams who would benefit most, are precisely the teams outside it. Those teams then build a parallel platform, and now you maintain two.

Abstractions should hide decisions, not information

The pitch for a platform is usually that application teams no longer need to understand the layer underneath. That holds right up until something fails. Then the error surfaces in your abstraction's vocabulary while the cause lives one or two layers down, and the developer now needs to understand both the abstraction and the thing it was supposed to replace. You have added a layer of knowledge, not removed one.

The specific mechanisms are worth naming, because they recur:

The rule that follows is narrow and useful. Hide decisions, not information. It is fine, good even, for the platform to choose the base image, the ingress shape, the retry policy. It is not fine for the platform to be the only thing that can see why a deployment failed. Pass the underlying diagnostics through verbatim alongside your translation of them, and make sure the path from a failed deploy to the real error does not route through a person on your team.

The portal is the last thing you build

A developer portal is the visible artefact, so it tends to get built first. The result is a website that lists services and links to dashboards nobody maintains, which then becomes evidence that the platform exists when in fact nothing has been paved.

The order that works is unglamorous. Pave one path end to end, for one real workload, with one team that actually wants it. Then a second. Then look at what the two have in common and factor that out. Only then does a front door have anything true to show.

One specific warning about catalogues: a catalogue is only true if something enforces its registration. Ownership metadata typed in by hand decays silently, and the tell is a service whose listed owner left the organisation months ago. Derive the catalogue from the thing that actually deploys, so that an unregistered service is an impossible state rather than a tidiness problem.

Owning the hardware moves the boundary, it does not remove it

If you run on your own machines, which is the situation we work in most often, the platform question changes shape. Cloud providers hide capacity behind a pricing model, so the API always says yes and the consequence arrives on an invoice. When you own the hardware, something has to say no, and that something is the platform.

So admission control becomes a first-class part of the interface rather than an implementation detail. Who decides whether a workload fits? What happens when two tenants both want the same accelerator? What is the behaviour when the answer is no: a refusal with a reason, or a queue, or a job that starts and dies?

The failure mode here is worth stating precisely, because it is subtle and we have seen it in our own systems. A gate that reads a signal which is not the constraint will cheerfully admit work that cannot possibly run. If the real limit is memory and the gate consults utilisation or temperature, every rejection you needed never happens, and the logs will show a healthy-looking check immediately before each failure. When you build an admission gate, the first thing to verify is that the quantity it measures is the quantity that will stop the job.

The rest is ordinary but must be owned explicitly: node lifecycle, firmware, disks, and the fact that adding capacity is a purchase with a lead time rather than an API call. None of that disappears because you wrote a nice CLI in front of it.

How to tell whether it is working

The most useful test I know involves a stopwatch and no charts. Take somebody who has not done it before and have them carry a new service from empty repository to production, with an on-call rotation attached, using only the documentation. Watch them. Do not help. Where they stop is your roadmap, and it will not be where you expected.

The four delivery measures in common use, deployment frequency, lead time for a change, change failure rate and time to restore service, are worth tracking, with one caveat: use them as a before-and-after on your own teams. I would not trust a cross-industry benchmark quoted at you without knowing the population it was drawn from, and I am not going to quote one here.

Two negative signals matter more than most dashboards. First, shadow platforms: if two teams have independently built the same capability, that is information about your backlog, not about their discipline. Second, interrupt load: if the platform team's week is mostly unplanned requests, the self-service claim is not true yet, whatever the portal says.

And on the business case, which is usually made as duplicated effort removed. Count the work you add at the boundary as well. Every abstraction is a thing to version, document and support, and every escape hatch is a second path to keep working. Whether that trade is worth it at your size depends on how many teams you have and how similar their workloads are, and I do not think there is an honest general answer.

If you are deciding this now, start smaller than feels defensible. Pick one real workload, pave its path end to end with one willing team, and write down what you will not support. Keep the escape hatch open and instrument it, so adoption stays a signal rather than a policy. Build the portal when you have two paths worth listing. And be open to the conclusion that a small organisation does not need a platform team at all; it needs one paved path, written down, with a named owner. That is a smaller thing than a platform, and it is usually the thing that was actually missing.

Talk to us about this