ISSUE 2026-09-27SERIES BREC 40TYPE BRIEFgarnetgrid.com
The question is not which tool
Insights
Observability gets bought as a platform and lost as an engineering discipline. The decisions that determine whether you can debug an unfamiliar failure at three in the morning are made long before the tool is chosen: what identifier you propagate, what you pre-aggregate away, what you alert on, and where the telemetry is allowed to go. This is a practitioner's account of where that work actually breaks.
Most observability decisions get framed as procurement: which platform, which agent, which tier. That framing hides the only test that matters. Something is wrong in production right now. You have one customer's complaint, a rough time window, and a failure nobody has seen before. Can you find out what happened to that single request, across every hop it took, without shipping new code first? If the answer is no, you do not have a problem you can buy your way out of. You have an instrumentation problem, and the tool you pick will mostly determine how expensive it is to keep not solving it.
Monitoring answers only the questions you already asked
A dashboard is a frozen hypothesis. Every panel exists because somebody once predicted a failure mode and built a view for it. That is genuinely useful, and it is not sufficient, because the incidents that hurt are the ones nobody predicted. Novel failures do not show up as a red panel. They show up as a slightly wrong number on a graph built for something else.
Pre-aggregation is where the ability to ask new questions gets thrown away. A counter incremented on every error tells you the rate. It cannot tell you that all of those errors came from one tenant, on one node, on one build, using one model. That information existed at the moment the error happened and was discarded before storage. No tool recovers it afterwards.
The alternative is to emit one wide, structured event per unit of work — request, job, generation — carrying every attribute you might later want to group by: tenant, route, model, node, build id, queue, cache outcome, retry count, whether it was a cold start. A new question then becomes a new query rather than a new deploy. That has a cost, it arrives immediately, and it is the next section.
Related, and cheap: the most valuable dashboard in any estate is the one somebody threw together during an incident. Keep those. Delete the ones nobody has opened in three months, because a wall of unread panels is how a team loses the habit of looking.
Cardinality is the engineering problem
In a metrics system, every distinct combination of label values is a separate time series. Add a user id, a request id, a raw URL path or a full model prompt hash as a label and you have not added a dimension, you have multiplied your series count by the size of that set. Memory in the scraping process, index size in the store and query latency all follow. A single label added in a pull request can multiply the cost of the whole system, and nothing in code review catches it unless somebody put a check there deliberately.
Pricing shapes vary — per ingested gigabyte, per host, per custom metric, per active series, per seat, and combinations of those — and they change, so any specific figure I gave you would be wrong somewhere. What is stable is the structural trap: the quantity that drives the bill is almost always one that developers can change without ever seeing the bill.
Two specific traps are worth knowing by name. First, percentiles do not average. Taking the p99 from each of twenty instances and averaging them yields a number that is not the fleet's p99 and has no clean interpretation. You need either the raw distribution or histograms designed to be merged. Second, histogram buckets are chosen in advance. If your largest finite boundary sits below the real tail, the reported p99 will be pinned near that boundary. It will look precise to three decimal places. It is your own configuration being read back to you.
One identifier, propagated everywhere
The single highest-value piece of observability work is usually not a tool at all. It is making sure one identifier follows a unit of work through every component that touches it, and that it appears in every log line, span and event those components emit. Without a join key, three well-instrumented services give you three unrelated stories.
Propagation breaks at seams, and the seams are predictable: message queues, scheduled jobs, batch runs, webhooks arriving back from a third party, shell scripts, and any call an agent makes on behalf of a request. Auto-instrumentation will cover inbound HTTP and your database driver and produce spans named after your framework. It knows nothing about your domain, so the spans that would actually explain an incident — which tenant, which model, which queue, which retry — are the ones you have to add by hand.
Distributed tracing also collides with clocks. Spans are stamped on the machine that produced them, so a child span can appear to start before its parent and a duration can come out negative. Treat cross-host timings as approximate unless you have measured the skew.
Then sampling. Head-based sampling at a low rate keeps that fraction of the boring traffic and precisely the same fraction of the errors, so the trace you want is probably not the one you kept. Tail-based sampling keeps traces that turned out to be interesting, but it has to buffer the whole trace before deciding, which means every span for a given trace must reach the same collector instance, and spans that arrive after the decision window are simply gone. Both approaches are defensible. Neither is free, and the second one is a piece of stateful infrastructure you now operate.
Absence of signal is not health
A healthy system and a dead exporter look identical on a dashboard. The panel says no data, the threshold never trips, and the absence reads as calm. This is the most common way monitoring lies, and it lies in the direction that lets you sleep.
Every signal you rely on needs an inverse: something that fires when the heartbeat stops, not only when the value is bad. Staleness of the newest data point is a first-class alert. So is ingest lag on the telemetry store itself, because a backlog means your dashboards are describing the past while presenting as live.
Run alerting somewhere that does not share a failure domain with the thing it watches. A monitoring stack colocated with the service it monitors will fail at the same moment, for the same reason, and your first notification will be a customer. And plan for the storage layer being the component that eventually fills the disk: give it its own volume, a retention policy you have tested by actually letting it expire something, and an owner.
Alerts need an owner and an action
Alert on what a user can feel. Errors, latency, unavailability, work not getting done. Cause-based alerts — CPU high, memory above eighty per cent, disk busy — fire routinely during healthy operation, and their real effect is to train people to dismiss the channel without reading it. Once that habit forms, the genuine page gets dismissed too.
A test I would defend in any review: if you cannot write the first action in one sentence, it is not a page. It is a dashboard panel, or it is a weekly report. Aggregate before you alert, as well, because a per-instance rule during a rolling deploy will page once per instance and teach everyone that deploys are noisy.
The uncomfortable part is that alert quality decays silently. Nobody files a ticket saying this page was useless. You have to go looking, periodically, at what fired and what anyone did about it.
Telemetry is often your most sensitive data
Logs contain identifiers. Traces contain paths and parameters. For an AI system, prompts and outputs are the customer's own business content, which means the debug telemetry can be more sensitive than the production database, because it was never designed with the same care.
Shipping that to a hosted platform is a data-processing decision, not an infrastructure one. In a system whose entire premise is that data stays on hardware the customer owns, the observability stack is frequently the largest unexamined egress path in the design. It is worth auditing before anything else, because it is the one people forget they added.
Keeping it inside the boundary is the consistent choice, and it costs something honest: you are now operating a stateful store with compaction, upgrades, retention and its own outages. Separate the audit trail from debug telemetry while you are at it — they have different retention needs and different deletion obligations, and an erasure request against a year of raw logs keyed by user id is a genuinely unpleasant project.
For private AI workloads, the signals that earn their keep are narrower than general web telemetry:
If you are starting, do it in this order. Propagate one identifier through every hop, including queues and scheduled jobs. Emit one wide structured event per unit of work with the attributes you would want to group by during an incident. Put a cardinality check in code review, because that is the only place it can be caught. Add staleness alerts before you add threshold alerts. Give every page an owner and a first action. Instrument through an open protocol so the storage and query layer can change later without touching application code. And on the two questions people most want answered — how much to retain, and whether managed or self-hosted works out cheaper — the honest answer is that it depends entirely on your volumes and your obligations, and you should measure a week of your own traffic rather than trust anyone's general rule, including mine.