ISSUE 2026-09-27SERIES BREC 41TYPE BRIEFgarnetgrid.com

The denominator decides the answer

Insights

Comparing OpenAI, Anthropic and Mistral on their published token rates is the wrong exercise, and it usually produces the wrong answer. The number that decides your bill is the cost of getting one task done correctly, and that is governed far more by how you build the system than by which logo is on the invoice. Here is the arithmetic that matters, the failure modes that quietly multiply it, and the parts of it we cannot honestly fill in for you.

The comparison almost everyone makes is price per million tokens. It is the figure the vendors publish, so it is the figure that ends up in the spreadsheet. On its own it is close to useless, because you do not buy tokens. You buy resolved tasks: an extraction that was correct, a draft nobody had to rewrite, a support reply that closed the ticket.

The honest unit is cost per resolved task: tokens in times the input rate, plus tokens out times the output rate, divided by the proportion of attempts that actually succeed.

That divisor is what upends most comparisons. A model at twice the rate of another, which gets it right first time and produces output that survives review, is the cheaper model. A cheap model that fails one attempt in five, and hands each failure to a person, is the most expensive option on the table, because at the volumes most teams are actually running, an hour of a person's time costs more than a day of API spend — worth checking against your own invoice, since a high-volume system can invert it.. None of that appears on a rate card.

So the first thing to accept is that the three providers cannot be ranked in the abstract. They can only be ranked against your tasks, with your prompts, at your quality bar.

Why the published rates are not comparable anyway

Even if you only cared about the rate, the rates are not measured in the same unit. Each provider tokenises text with its own tokeniser, so the same prompt becomes a different number of tokens depending on who you send it to. A price per million tokens is therefore a price per million provider-specific units, and the gap moves with your content. Dense English prose behaves reasonably everywhere. Code, tabular data, heavily punctuated text and non-English languages do not behave uniformly, and that is where the discrepancy shows up.

Then there is the shape of the pricing. All three charge separately for input and output, and output is the expensive side by a wide margin. That single asymmetry means verbosity is a budget decision, not a style preference. A system prompt that says "explain your reasoning fully" is a line item.

There are also multiple rates for the same model at once: a lower rate for cached input, a lower rate for asynchronous or batched work, different treatment above certain context lengths on some providers. A comparison that uses only the headline rate is comparing the worst price each vendor offers.

Finally, the reasoning-style models bill the thinking they do before answering as output tokens, whether or not you are shown all of it. That makes per-request cost genuinely variable and hard to forecast from a prompt alone. Check the current documentation for the specific model you are using, because this is exactly the sort of detail that changes between releases, and it is the difference between a forecast and a guess.

Where the bill actually comes from

In the systems we have looked at, provider choice has rarely been the dominant cost term and prompt architecture usually is — that is our experience across a handful of builds, not a measured survey, so treat it as a hypothesis to test on your own traffic.

The largest invisible constant is usually the tool definitions. A dozen tools with verbose JSON schemas and generous descriptions are re-sent on every single call, and can easily exceed the user's actual question. Teams review their system prompts line by line and never once read their own tool schemas as a cost surface.

The second is conversation history. If you resend the full transcript each turn, a long conversation costs roughly the sum of all its prefixes, so spend grows far faster than the number of turns. Truncation and summarisation are the levers, and both trade accuracy for money, which is a decision worth making deliberately rather than by default.

The third is retrieval width. Top-k is a cost dial disguised as a quality dial. People raise k because recall improves, then never check whether answers improved. Often they got worse, because the relevant passage is now buried in the middle of a long context.

The fourth is loops. Every retry after malformed output, every failed tool call, every re-plan is billed. An agent that averages a dozen steps has a dozen chances to burn tokens on nothing. And the distribution is heavy-tailed: the average run is not what hurts you, the pathological run that circles until the context window fills is. Put a hard token budget on each run, alert on the 95th percentile rather than the mean, and make a runaway run fail loudly instead of expensively.

Caching is a prompt-architecture decision, not a billing setting

Prompt caching is the largest legitimate discount available on the interactive path, and it is the one most often left on the floor. The mechanism is roughly the same wherever it is offered: a repeated prefix can be read back at a reduced rate, there is a cost or overhead to write the cache, and the entry expires after a window. It therefore only pays when the same prefix genuinely recurs inside that window.

The consequence is structural. Stable content goes first, volatile content goes last. Put a timestamp, a user's name, a session identifier or a freshly reordered retrieval block near the top of your system prompt and every request is a cache miss that you are also paying to write. The mechanism is unforgiving: add a timestamp to the top of a system prompt for observability and every request becomes a cache miss you are also paying to write, so if that prefix is most of your tokens a one-line change can move the whole monthly bill.

Two follow-on habits are worth building. Treat any edit to a tool schema as a cache invalidation event, because everything after it in the prefix is invalidated too, which is how a harmless description tweak shipped at nine in the morning changes the day's spend. And read the usage fields the API returns instead of assuming the cache is working. Cache hit rate deserves to be on the dashboard next to latency and error rate.

Batching deserves the same treatment. If a workload does not need an answer in the next few seconds, the asynchronous path is usually cheaper. Sorting your traffic into interactive and patient is free engineering work with a direct effect on the invoice.

Mistral changes the question, because some of it you can host

The interesting difference between these three is not three rate cards. It is that part of Mistral's lineup has shipped with downloadable weights you can run on hardware you control, while other models in it are API-only. Licences differ per model and have changed over time, so read the licence for the specific model rather than trusting a general impression, including ours.

Where weights are available, the economics change category rather than degree. An API is a marginal cost: you pay per use and nothing when idle. Owned hardware is a fixed cost: you pay whether or not anything is running. So the comparison is not about rates at all, it is about utilisation. Bursty, low-volume, unpredictable workloads favour the API, and it is not close. Steady, high-volume, forecastable workloads are where owning the machine starts to win.

The honest accounting on the self-hosted side includes more than the GPU. It includes the engineer who serves, monitors and upgrades it. It includes quantisation choices that trade quality for throughput, and the evaluation work to find out how much quality you traded. It includes headroom for your peak, not your average. It includes power, cooling and the hardware's depreciation. It includes losing the free capability upgrades you get when a provider ships a better model behind the same endpoint, because from then on improvements are a project you fund.

Where is the crossover? That depends on your volumes, your duty cycle and who you employ, and there is no honest general answer. What you can do is measure a month of your own token traffic and compute both sides with your real numbers. It is dull work and it settles the argument.

We should declare the bias: we build private systems on hardware customers own, so read our enthusiasm accordingly. The strongest reasons to self-host are frequently not cost. They are data that must not leave the building, contractual or regulatory constraints, a pinned model version that behaves the same next quarter, and independence from someone else's rate limits and deprecation schedule. If none of those apply and your volumes are modest, the API is very likely the right answer and we will say so.

Switching cost, and the one asset that is portable

Prompts are not portable in practice. They accumulate small accommodations to one model's habits: how it follows formatting instructions, how reliably it emits valid structured output, how it behaves as context grows, where its refusal boundaries sit. Move providers and you rewrite scaffolding and re-tune, then discover the edge cases again.

The portable asset is the evaluation set. Tasks, inputs, and an agreed definition of a correct answer survive any migration and are what let you answer the cost question at all. If you build one thing before choosing a provider, build that. It is also the only defence against the slow drift where a system that used to work is now quietly failing a tenth of the time and nobody has a way to notice.

Multi-provider routing, sending easy work to a small model and hard work to a large one, does produce real savings. It also produces two sets of failure modes, two sets of rate limits, an abstraction layer that leaks, and a classifier of its own that can be wrong. Above a certain monthly spend that complexity pays for itself. Below it, it is a hobby with an on-call rota.

How to actually run the comparison

The measurement is not difficult, it is just rarely done, because it is less fun than reading benchmark tables.

None of this requires a platform, a consultant or a procurement cycle. It requires a couple of days and the discipline to use your own traffic rather than someone else's benchmark.

Choose the provider on quality for your specific tasks and on the constraints that are not negotiable, such as where the data may go and how stable the model version needs to be. Then engineer the cost down, because that is where the leverage genuinely is: prompt prefix order and caching, tool schemas trimmed to what is used, retrieval width set by measurement rather than instinct, history managed deliberately, and hard token budgets on every agent loop. Measure cost per resolved task on your own tasks before signing anything, and set a date to measure it again. If your volumes are steady and large, or your data cannot leave the building, price the self-hosted option properly against those same numbers rather than against a rate card. And distrust any article, including this one, that offers you a figure instead of a method.

Talk to us about this