ISSUE 2026-09-27SERIES BREC 22TYPE BRIEFgarnetgrid.com

The list is not ten equal problems

Insights

Most API vulnerability lists treat ten problems as equally likely. In practice one class dominates, and it is the class automated scanning cannot see. This is a working view of what actually gets exploited, why it survives code review, and the small number of changes worth making. No figures, because the honest ones are not public.

Every list of API vulnerabilities has ten items on it, and the list is broadly right. The trouble is that it reads as ten equal problems, and they are nowhere near equal. If you are about to spend a quarter on API security, the useful question is not what the ten things are, it is which one is going to happen to you. The honest answer, taken from how these systems break rather than how they are documented, is that it is almost always authorisation. Not injection, not cryptography, not a clever protocol attack. Someone changed an identifier in a request, and the server answered.

Object-level authorisation is the whole game

The shape is always the same. A request arrives with a valid token, hits an endpoint that requires authentication, asks for object 4312, and gets it. The endpoint checked that you are someone. It never checked that 4312 is yours.

What makes this the one that gets exploited is not difficulty. It is that nothing about the attack looks like an attack. The request is well formed, the token is real, the status is 200. There is no payload to detect, no anomaly in the traffic shape, nothing for a WAF to pattern-match. You find it by reading code, or by testing with two identities. You do not find it by watching.

It survives because authorisation in most codebases lives per-handler. That makes it a completeness problem over a set that grows every sprint, and completeness problems lose. Every new endpoint is a fresh chance to forget, and the forgetting is invisible in review, because the handler looks like all the others.

The fix that holds is structural. Stop fetching an object and then deciding whether the caller may have it. Make the owner or tenant a mandatory part of the query at the data-access boundary, so an object that is not yours does not exist rather than existing and being refused. Return 404 rather than 403; a 403 confirms the object is real, which is worth something to an attacker. Random identifiers help, but only by raising the cost of enumeration. They are not a check, and they leak through shared links, logs and referrers.

The same mistake, one layer up and one layer down

One layer up is function-level. The role check exists on the endpoints somebody remembered to guard. A normal token calls the administrative route and it works, because the button was hidden in the interface and hiding the button felt like the fix. The subtler version is per-method: GET on a resource is guarded, and the PATCH added three months later is not. The other common variant is an endpoint meant only for internal service-to-service calls, exposed through the same gateway as everything else, protected by the belief that nobody knows it is there.

One layer down is fields. In one direction the handler returns the whole model and lets the client render three properties, leaving the rest in the response body for anyone with developer tools or a proxy. In the other the request body is bound straight onto the model, so an attacker adds a property they should never control and the framework writes it for them. That was the famous mass-assignment lesson more than a decade ago, and the framework behaviour behind it still ships enabled in plenty of stacks.

Both directions need an explicit shape. Declare what a response contains and construct it, rather than serialising whatever the ORM handed back. Allowlist the properties an endpoint will accept, and reject unknown ones with an error instead of dropping them, because silent dropping hides integration bugs as well as attacks.

You are serving more API than you documented

The endpoint that gets hit is rarely the one you have been protecting. It is the old version prefix, still routed two years after everyone moved on, written before the middleware you now rely on and never retrofitted. It is a staging host on the same zone, carrying a copy of production data and none of the rate limiting. It is a debug route behind a flag whose default flipped during a refactor.

The mechanism is dull and reliable. Gateway rules, WAF policies and test suites are all written against the routes people know about. Anything outside the document is outside the rules, and outside the tests too, so it never shows up as failing either.

Measure what is served, not what is specified. Enumerate routes from the proxy configuration and from the running router, enumerate hostnames from DNS, and diff both against the specification. The difference is the output worth reading: routes serving traffic nobody documented, and documented routes that no longer exist. Give deprecation a date and a 410, because a note in a changelog does not stop a request.

Webhooks, and signature checks that do not check

Inbound webhook handlers are where I see the most genuinely dangerous code, because they get written once, they work, and nobody opens them again.

One rule before any of the specifics. A valid signature proves who sent the message; it does not make the contents authoritative. Do not take amounts, prices or entitlements out of the payload and act on them. Take the identifier, read the real state back from the provider or from your own records, and act on that.

Four implementation failures recur. We have shipped the last two ourselves and found them by reading the parser, not from an alert, which is the point: nothing in the system was unhappy.

Parameters that spend your money or reach your network

Two classes of bug get filed under performance, and treated as lower priority than they deserve.

The first is unbounded work. A list endpoint whose page size the client sets, with no ceiling. A filter parameter that maps to an unindexed column. A GraphQL schema where nesting and aliasing let a single request do an enormous amount of work, which is invisible to a limiter that counts requests. Per-IP limits do little against a distributed client, and per-token limits do little when signing up is free. The version that hurts now is any endpoint that forwards to something metered. If a request can trigger model inference, an email, a message or a third-party lookup, then a cheaply authenticated path to it is a way for a stranger to spend your budget. Meter and cap per tenant, and make exceeding a cap return a clear refusal rather than a collapse.

The second is any parameter that takes a URL: webhook registration, avatar-by-link, import-from-URL, HTML to PDF, link previews. Cloud sharpens it: the instance metadata service on the link-local address is reachable from inside, and plenty of internal services require no authentication at all on the grounds that they are internal. The defence has to be mechanical. Resolve the hostname yourself, validate the resolved address against an allowlist, connect to that address, and re-validate after every redirect, or refuse redirects entirely. Validating the string before resolution is precisely what DNS rebinding defeats. Where you can, run the fetch somewhere whose network reach is already restricted, because a code-level check is one refactor away from failing quietly.

What finds these, and the test that proves nothing

A scanner will not find the authorisation bugs. It has no model of who owns what, and that model is the entire question. What finds them is a test with two tenants in it.

Parameterise it over the route table rather than writing it per endpoint: for every route that takes an object identifier, assert that tenant A gets a 404 for tenant B's object. Then make a route with no cross-tenant case a failure of the suite, so the next endpoint added fails by absence instead of passing by omission. That is the only form of this test that survives a growing codebase.

Then prove the test can fail. Delete the authorisation check on purpose, run the suite, watch it go red, put the check back. A test never observed failing is a claim, not a control, and an authorisation test that passes because its fixtures are wrong looks identical to one that passes because the code is right.

Two smaller habits pay well: read every new handler's diff specifically for the ownership check, and log authorisation decisions rather than only requests, because after an incident you need to answer who was refused and why.

Pick the two or three object types where a single cross-tenant read would end a customer relationship. Write the cross-tenant test for those today, and break it deliberately to prove it works. Then spend an afternoon enumerating what your infrastructure is actually routing, and delete or 410 whatever should not be there. Then read your webhook verifier line by line. That is a few days of work, and it covers the failure modes that actually get exploited. Everything else on any ten-item list can queue behind it. None of it requires buying anything, which is probably why it stays undone.

Talk to us about this