ISSUE 2026-09-27SERIES BREC 24TYPE BRIEFgarnetgrid.com
Start with the question you are actually answering
Insights
Build versus buy is usually argued as a cost comparison, and that is the one framing that reliably gets it wrong. The better question is which parts of your system you are prepared to be permanently unable to change, and what it costs you when the part you cannot change turns out to be the part you got wrong. Licence fees and salaries are inputs to that question, not answers to it.
Most of these decisions are made in a meeting where someone puts a subscription cost next to a number of engineer-months, and the tidier spreadsheet wins. Both sides of that spreadsheet are wrong in the same direction: they count the visible cost of the first year and none of the structural cost of the next five. The decision that matters is about control. When you buy, you are renting someone else's prioritisation of what your system should do next. When you build, you are taking on the obligation to keep a thing alive after the person who wrote it has moved on. Ask which of those two liabilities you can actually carry for the specific capability in front of you, and the answer usually stops being close.
You are renting a roadmap, not a feature set
A vendor's roadmap is the median of its customers' demands. If your requirement sits near that median, buying is extraordinarily good value: you get years of other people's edge cases already handled, for a fraction of what discovering them yourself would cost. If your requirement sits in the tail, you will wait, escalate, be told it is on the roadmap, and wait again.
The failure mode this produces is the eighty-per-cent product. The tool does most of what you need, so you buy it, and the remaining twenty per cent gets filled with an apron of glue: a nightly export, a reconciliation script, a webhook handler that patches records the tool got wrong, a spreadsheet one person maintains by hand, an automation chain nobody owns. None of that appears as a line item. It is distributed across several teams as unattributed overhead, which is exactly why it grows unchecked and why the true cost of the purchase is invisible at renewal time.
So the useful question is not "does this tool do what we need". It is "is the part we care about most this vendor's core competence, or their integration surface". Core competence gets fixed. Integration surface gets deprecated.
Buy the things that are adversarially hard; build the thing you are paid for
Some capabilities look small and are not. Authentication is easy on day one and hard on the day someone reports a session-fixation bug, or an enterprise customer requires their own identity provider, or you need user provisioning that deactivates leavers automatically. Payments look like an API call until you meet chargebacks, retries, tax, and the reconciliation work that proves the money you think you took is the money you have.
Email deliverability is the cleanest example. The code to send a message is trivial. What you are buying from a provider is sending-IP reputation, bounce and complaint classification, feedback loops with the large mailbox providers, and someone whose job it is to argue with them when your mail starts landing in junk. You cannot compile a reputation from source, and you cannot build it quickly at any price.
Then look at the other end. Build where the domain logic is the business: pricing, allocation, matching, scheduling, underwriting, whatever your margin actually lives inside. A vendor who packages that logic has packaged the industry median, which is definitionally the opposite of an advantage. Buying your differentiator is how organisations end up competing on sales effort because their product is the same product as everybody else's.
Estimate against the meter, not the price
Never build a five-year case on a current list price. Pricing changes, varies by region, and is negotiated. What is stable and worth analysing is the shape of the meter, because the shape tells you how your bill behaves when things go well.
Per-seat pricing taxes internal adoption, which pushes you towards shared logins and fewer people having access to the thing you bought to give people access. Per-request or per-token pricing taxes success, and it also taxes retries, which is how distributed systems are supposed to work. Consumption units abstracted away from anything recognisable (credits, compute units, whatever the vendor has coined) decouple your bill from any quantity your finance team can forecast, so nobody notices the trend until it is a problem. Per-workload pricing on a platform you do not control means a background reindex can cost you real money.
The test is simple: does the meter move with a quantity you control, or one your users control? A feature where the customer decides how much compute to consume, billed to you per unit, is a business model you have to design around rather than an engineering detail. If you cannot answer that question, you do not yet know whether buying is cheaper, and no spreadsheet built on today's price will tell you.
Data gravity sets your real switching cost
The cost of leaving a tool is not the cost of replacing its features. It is the cost of reproducing its accumulated state somewhere else. A tool that holds configuration is easy to leave. A tool that holds four years of history is harder. A tool that holds four years of derived history (threads, scores, embeddings, audit trails, anything computed from events it no longer stores) may be effectively impossible to leave, because the raw material to recompute it is gone.
Do this during the trial, not at renewal: request a full export and actually run it end to end. Then check whether it contains stable identifiers you can join on, whether archived and deleted records are included, whether attachments come with it or only links to them, whether the rate limit means a full export of your real volume takes days or weeks, and whether the format is documented or simply whatever the internal schema happens to be. Exports are frequently the least-exercised path in a product, and you would rather discover that while you still have the option not to sign.
If you cannot leave, you have not bought a tool. You have accepted an ongoing dependency whose price is reset annually by somebody whose interests are not yours. That can still be the right trade. It just should not be a surprise.
What AI workloads change, and where our own bias sits
Two things are genuinely different about the current generation of bought capability. The first is that a lot of it involves sending your data to a third party, which for some categories of data is a legal and commercial decision rather than an architectural one, and gets taken above the engineering team. The second is that the dependency is non-deterministic and versioned by the supplier. Model versions get retired and replaced, behaviour shifts without a single line changing on your side, and your own evaluation suite is the only instrument that will notice. Own that evaluation harness whichever way you go; it is the part that tells you the truth.
We should be straight about our position here, because we sell private systems that run on hardware the customer owns, and that makes us an interested party. The unglamorous reality of running your own models is that you inherit capacity planning, accelerator memory limits, driver and runtime upgrades, model storage, and a quality-regression process you now have to design yourself. Self-hosting tends to win at steady, high, predictable volume and lose at low, bursty volume, because you are buying capacity rather than usage. Where the crossover sits depends entirely on your utilisation curve and your hardware, and anyone quoting you a general figure for it is guessing. The decision more often turns on data sovereignty and on wanting the behaviour to stay fixed until you choose to change it, and those are reasons worth stating plainly rather than dressing up as cost savings.
A procedure that survives contact with reality
Write the requirement as invariants rather than features. "A booking must never be double-allocated" can be tested against any candidate, bought or built. "Has a scheduling module" cannot be tested against anything.
Then, in most cases, buy first and build the seam. Put the vendor behind a thin interface of your own from the first commit, so that replacing it is a contained project rather than surgery. That seam is small, bounded work — often a few days against a simple API, longer against a large one — and it buys you an option worth far more. The usual mistake is letting vendor concepts leak into your domain model: their identifiers as your primary keys, their webhook payload shape as your internal event, their status enum in your business rules. That is what converts a replaceable component into an amputation.
Pilot on real data and the awkward case, never the demo dataset. The awkward case belongs to your most important customer, and it is the reason the tool will or will not fit.
Prefer the reversible option when the two are otherwise close. If buying is reversible and building is not, that asymmetry usually outweighs the cost difference, and the same holds the other way round.
Finally, write the exit criterion down at the point of purchase, with a review date. Something checkable: a per-unit cost above an agreed share of gross margin, or more than an agreed number of workarounds in the codebase. Counting workarounds is a leading indicator you can actually measure, which is more than most of the alternatives offer. A rising number of scripts that exist only to compensate for a tool is the tool telling you it no longer fits, in the only language it has.
Split the system into the part that is your reason for existing and the part that is plumbing. Buy the plumbing from people who do it adversarially well, because reputation, edge cases and regulatory grind cannot be built quickly. Build the narrow core where your margin lives, and accept the maintenance obligation that comes with it. Put every purchase behind a seam, and run the export before you sign. Above all, refuse to settle it with a spreadsheet of licence cost against salary cost, because that spreadsheet omits both of the things that actually get you: the glue that accumulates around a near-fit product, and the second year of a system with only one author.