ISSUE 2026-09-27SERIES BREC 27TYPE BRIEFgarnetgrid.com
The question you are actually answering
Insights
Most of the CSPM decision has very little to do with which product you buy. It turns on what an API-polling inventory engine can and cannot see, why a clean result is the easiest output to manufacture, and who is going to read the queue the thing produces. This is what actually goes wrong.
The question is almost never "which cloud security posture management tool is best". It is "do we have the capacity to act on what one of these will find, and will we be able to tell the difference between a clean estate and a blind one". A posture tool is cheap to switch on and expensive to take seriously. If you buy it without answering that second half, you will end up with a dashboard that is quietly green for reasons nobody has checked. Here is what the tooling actually does, where it misleads, and what to do first.
What the tool actually is, and what follows from that
Strip away the category names and a posture tool is three things: an inventory, a rule engine, and a queue. It authenticates to your cloud accounts with a read-only role, walks the provider's control-plane APIs, builds a picture of every resource and its configuration, evaluates a catalogue of rules against that picture, and shows you what deviates. Better tools add graph analysis over the inventory. Most now bundle vulnerability data from disk snapshots. The core is still polling an API.
Nearly every limitation follows from that sentence. The tool sees what the control plane will tell it, at whatever interval it polls, and that interval is a configuration detail that differs by resource type rather than a constant. A security group opened for a debugging session at two o'clock and closed forty minutes later may never appear at all. It sees configuration, not content: it can tell you a bucket policy permits anonymous reads, it cannot tell you whether anything inside matters. And it sees the control plane, not the guest. What is baked into the container image, what your application does with the token it was handed, whether the database engine has a second authentication path configured inside itself, all of that is out of scope.
That is not a criticism. An inventory-and-rules engine over the control plane is worth owning, because the control plane is where a great many cloud incidents are actually made. It just has to be bought as what it is rather than as general cloud security.
The snapshot scanning deserves its own note, because it is where expectations drift furthest. The tool takes a point-in-time snapshot of a volume, mounts a copy, and inventories the packages it finds. It is genuinely low-friction and it works. What it tells you is that a package was present on disk at snapshot time. Whether that code is loaded in a running process, reachable from anything, or even part of the image the scheduler is running today are separate questions, and they are the ones that decide whether a given CVE matters to you. Answering them needs a runtime sensor, in practice an eBPF one, which means a deployment programme and a conversation with whoever owns the nodes about kernel-level software in production. The agentless-versus-agent argument is usually conducted as doctrine. Treat it as scoping: which of those two sentences do you need to be able to say out loud?
The first run is an inventory, not an assessment
The first scan will produce a number large enough to alarm whoever asked for the tool, and that number is mostly a function of how many resources you have. It is usually a small set of distinct defects multiplied across a large estate: one Terraform module without encryption enabled, instantiated everywhere; one account created before the logging standard existed; one team's naming convention that fails a tagging rule.
So collapse by cause before you triage by severity. Group findings by rule and by the thing that produced the resource, and you will typically find the work is a handful of pull requests rather than thousands of tickets. This reframing is the single highest-value hour of the whole exercise, and default views rarely do it for you: grouping by rule is common, grouping by the module that produced the resource is not, because a big number is a better demo.
Then be sceptical of the severity column. Severity is the vendor's opinion of a rule in the abstract. It does not know that the unencrypted volume sits on a host in a subnet with no route out, or that the low-severity finding is on the one machine holding a signing key. Reachability and blast radius reorder the list completely, and only you hold those facts. Expect the honest first pass to involve a lot of deletion: abandoned accounts, forgotten test estates, resources no team will claim.
Identity is the real surface, and the hardest to read
Public storage gets the headlines. Escalation happens in identity. The defects that convert a small foothold into an incident are nearly always permission-shaped: a role whose trust policy names a principal more broadly than the author intended, a service identity that can pass a privileged role to a compute service, a CI identity with permission to edit the pipeline that grants it, the ability to rewrite a function's code where the function holds the permission you actually wanted.
None of those is a misconfiguration in the checkbox sense. They are legitimate grants in a bad combination, which is why flat rule catalogues miss them and why every serious vendor now sells graph analysis under some variation of the phrase attack path.
The graph is the most valuable thing in these products and the most likely to be wrong in ways you cannot see from the outside. It has to model the real permission semantics of hundreds of services, and providers ship new services faster than any catalogue absorbs them. Policy conditions are where it fails most often, and it fails in both directions: a grant that looks catastrophic but is fenced by a source address, tag or organisation condition, and a grant that looks tightly scoped until you notice the condition key can be set by the caller.
The practical defence is cheap. Before you trust the graph, hand-verify two or three paths against the actual policy documents, including one the tool says is fine. A good vendor expects this and will help. If they discourage it, you have your answer.
A coverage gap looks exactly like good news
The most dangerous output of a posture tool is a clean result you have not earned, because absence of findings and absence of visibility render identically on the same screen. Three mechanisms produce that routinely.
The read role goes stale. You attached a policy at onboarding. The provider has launched services since. If the role cannot describe a resource type, those resources simply do not appear. Ask the tool what it failed to read, not only what it found. If it cannot tell you, that is itself a finding about the tool.
Accounts nobody onboarded. Coverage follows the account list you supplied. A team that opened an account on a corporate card is invisible, and your dashboard will not look any different for it. Reconcile the tool's account inventory against the billing or organisation list, which is much harder to slip past than an onboarding process.
Suppression rot. Every exception is reasonable on the day it is made. Without an expiry and a named owner, a dashboard that is green a year later is substantially a record of what people muted. Exceptions should expire by default and come back. A tool that cannot expire them will be configured into dishonesty, not by anyone acting in bad faith, but by ordinary accumulation.
Fixes belong upstream, and prevention has a political cost
A finding fixed in the console is fixed until the next apply. If the resource came from a module, the module is the defect and the console fix is a claim your pipeline will shortly contradict. There are two places worth pushing a fix: into the code that produces the resource, and into the provider's preventative layer, where service control policies, Azure Policy and organisation policy constraints refuse rather than report.
Prevention is better and its cost is understated. A deny that fires at deploy time breaks somebody's release, quickly and visibly, and the blame attaches to whoever added the rule. So sequence it rather than declaring it. Detect first, and measure how often the rule is wrong about your real estate over a few weeks. Move it into CI as a warning. Promote to a hard deny only the small set of rules that have not yet been wrong, and only where an exception can be granted in minutes rather than days.
A guardrail with no fast exception path gets routed around, and the usual route is a new account you will not learn about until the previous section happens to you.
What you are actually buying
Read the metering definition before you look at the rate. These products are priced per something: per resource, per workload, per asset, per protected seat. The definition of that unit determines your bill far more than the headline number, and ephemeral compute is where estimates come apart. Ask in writing whether a container that lives ninety seconds counts, how a stopped instance counts, and whether one host can be billed twice across two modules of the same platform. Pricing itself moves and is regional, so take your numbers from the quote, not from an article.
Every major provider ships a first-party posture and configuration service of its own. For a single-provider estate with a few accounts, that plus some custom rules covers more of the baseline than most people assume, and it is worth pricing as your control case. Third-party tools earn their keep at multi-account and multi-cloud scale, through normalisation, the identity graph, and having one queue instead of three consoles. Be wary of cross-cloud rule parity: the same rule name can mean materially different things on different providers, so ask to see the rule as implemented on your second provider rather than on the matrix in the deck.
Then be honest about the queue before funding it. A posture tool manufactures work. If no human owns each account, findings have no addressee and the product degrades into something screenshotted once a quarter. Naming an owner per account costs no licence fee and does more for your posture than the tool will.
Finally, consider the tool as an asset in its own right. You are about to grant a vendor a read role across your entire estate and ship your configuration, and possibly your disk contents, into their platform. That integration promptly becomes one of the most valuable credentials you hold, and a plausible exfiltration path if the vendor is compromised. Scope the role, log its use, alert on use from unexpected sources, and know which jurisdiction the analysis and any retained data sit in. If you cannot say where snapshots are mounted or how long they are kept, procurement is not finished.
Some organisations conclude that the analysis should run on infrastructure they control, and for a regulated or genuinely sensitive estate that is a defensible conclusion. It is not a free one. You give up the vendor's rule catalogue and the research team maintaining it, and you take on writing and updating checks yourself. That is a real loss and it should be counted rather than waved away.
Start smaller than the vendor proposes. Onboard one account that genuinely matters, read-only, and give the first run a week of attention rather than a triage sprint. Count distinct causes rather than findings. Hand-verify two or three attack paths against the policy documents, including one the tool calls safe. Ask the tool what it could not read, and reconcile its account list against billing. Name an owner per account. Then pick the two rules you would be content to have block a deployment, and put only those in CI. If that sequence bores a prospective vendor, you have learned something useful for the price of a week.