ISSUE 2026-09-27SERIES BREC 20TYPE BRIEFgarnetgrid.com

The question underneath the checklist

Insights

Readiness has very little to do with models. It is about whether your organisation can be read by a program, whether it can prove who was allowed to read what, and whether anyone would notice if an answer were wrong. This is the checklist worth running, including the parts of it nobody can answer for you.

Most AI readiness checklists are procurement artefacts. They ask whether you have a data strategy, an AI governance framework and executive sponsorship, and they can be passed comfortably by an organisation that would be entirely unable to run anything useful. The question they are groping towards is narrower and harder: if a capable model were dropped into your organisation tomorrow, what could it do that a new starter with a laptop and no access could not? For many enterprises the honest answer is not much, and in the cases we have seen the reason is rarely the model.

So what follows is not really about AI. It is about whether your organisation can be read by a program, whether it can prove who was allowed to read what, and whether anyone would notice if an answer were wrong. Those three things take most of the effort. Choosing a model takes an afternoon.

Can a program answer a question about your own business?

Set aside data quality as an abstraction and run a specific test. Pick a question a competent employee answers every week by opening four different systems and applying judgement. Now ask whether a program could reach those same four systems, with credentials it holds, without a human exporting anything.

The failures here are mundane, and they are the project. The system of record turns out to be a person. One of the four systems has no usable API, so the real interface is a nightly CSV that someone drops on a share. The warehouse is a day behind and the question is about today. And the join fails, because the customer identifier in the CRM is not the one in billing, and the mapping lives in a spreadsheet on a laptop.

None of this is glamorous and all of it is load-bearing. Retrieval is plumbing. If you are not willing to fund plumbing, you are not ready, and no amount of model capability substitutes for it.

Permissions are the whole problem, not a later phase

The moment you build one index over your documents, you have built a machine that can answer a question using a file the asker was never allowed to open. In our experience it is one of the most common reasons a system that demonstrated beautifully turns out to be unshippable.

There are two workable shapes. You can filter at query time using the asker's identity, which requires identity to be present in every request and access rules you can evaluate cheaply at scale. Or you can partition the index by audience, which is much simpler, duplicates content, and drifts out of date. Most organisations end up with some of both. What does not work is deciding later.

Expect to discover that your file shares carry permissions nobody has reviewed in years. Note that the system does not leak the document, which would at least be traceable. It paraphrases it, unattributably, into an answer. That is worse.

Related and frequently missed: anything the model reads is untrusted input. If retrieval can reach a shared mailbox or a ticket queue, then anyone who can send you an email can write text into your prompt. Instructions smuggled inside retrieved content are a real attack class, not a thought experiment. The mitigation is not a sterner system prompt. It is constraining what the system is permitted to do once it has read something, which leads directly to the next question.

Decide what a wrong answer costs, first

Classify every intended use by reversibility rather than by expected accuracy. Drafting a reply is reversible; sending it is not. Suggesting a price is reversible; issuing a credit note is not. Summarising a contract for a lawyer is reversible; acting on the summary is not.

Then insist on a named human owner for each irreversible action, and design the interface so that approving is a real decision rather than a formality. A queue of forty approvals a day will be rubber-stamped by Thursday. That is a design failure, not a discipline failure, and it is predictable enough that you should design against it from the start.

Measure the status quo before you begin, too. In our experience it is common to find that nobody can say how often the current manual process is wrong, how long it takes, or what it costs. Without that, every later argument about the system's accuracy is unresolvable, because it is being compared against an imagined process that never makes mistakes.

Where it runs is a constraint set, not a preference

There are three real questions here, and they shape the architecture far more than any feature comparison will.

Can the data leave? Not as a policy aspiration. "Nothing leaves" is either an architectural fact you can demonstrate, such as a machine with no outbound route at all, or it is a clause in a contract backed by an audit report. Both are legitimate positions. They are not the same claim and should not be described with the same words.

What shape is the cost? Hosted inference is metered, typically per token, per seat, per credit or per unit of compute time, so the bill tracks usage and grows with success. Hardware you own is capital, plus power, plus someone competent to keep it running, so the cost is largely fixed and the marginal query is nearly free. Which is cheaper depends entirely on your volumes, your utilisation and what your engineers' time is worth. Vendors publish very different numbers here, those numbers change, and anyone offering you a general crossover point is guessing.

What latency, and what peaks? Owning capacity means a spike becomes a queue you have to schedule. Renting it means the spike is absorbed and shows up on an invoice. Neither is free. Choose which failure you would rather explain.

Our own bias is towards systems running on hardware the customer controls, and it is worth being explicit that this is a bias with costs. You take on operations, and you give up the largest models you have not bought. For some workloads that trade is obviously right, and for others it is obviously wrong.

Evaluation you would actually bet on

A demo is a sample chosen by the person who built it. It tells you the system can succeed, which you already assumed.

What you need instead is dull: a fixed set of real cases drawn from your own work, with answers a practitioner has agreed are correct. As a rule of thumb from practice rather than a measured optimum: a few dozen cases give you signal, and a couple of hundred give you something you can defend in a room full of sceptics. It has to be graded by someone who does the job, not by the person who bought the system, because the buyer cannot tell a confident wrong answer from a right one.

Build it early, because the thing you really need it for is regression. Every change to model version, prompt, chunking, retrieval ranking or context assembly can move behaviour in ways nobody predicts, and if you use a hosted model some of those changes will be made for you. Without a fixed set you will have opinions about whether last week was better.

Using a model to grade a model is genuinely useful for scale, with the caveat that it tends to agree with whoever wrote the rubric. Keep a human-graded subset as the anchor.

Tells that a programme is real

Past the technical questions, a few organisational signals separate programmes that ship from programmes that hold meetings. There is a specific recurring task the system is meant to change, named at the level of "first-pass triage in the underwriting team" rather than "improve productivity". Someone owns it whose own job is affected by whether it works. The current process is measured well enough to compare against. Somebody has decided in advance what result would cause the project to be stopped. And the people whose work it touches have seen it, said something sceptical, and been thanked for it.

The inverse is a budget with no decision attached to it, a steering group, and a pilot with no exit criteria. Those programmes can run for a year and produce a slide deck.

Pick one recurring decision that somebody makes every week. Write down how it happens now, how long it takes and how often it goes wrong; if you cannot, that measurement is your first piece of work. Then get a program read-only access to the systems that decision touches, with permissions enforced at query time from day one rather than retrofitted. Assemble the evaluation set out of real past cases before choosing any technology. Pick the model last, and expect to change it. And if at some point the honest answer is that the plumbing would cost more than the decision is worth, stop and say so. That is a readiness exercise succeeding, not failing.

Talk to us about this