ISSUE 2026-09-27SERIES BREC 33TYPE BRIEFgarnetgrid.com

Reconciliation is the deliverable, not the transfer

Insights

Most migration plans answer the wrong question. They answer "how do we move the data", which is mainly a matter of bandwidth, patience and a sensible chunking strategy. The question that decides whether the project lands is different: on the morning after cutover, how will you prove the new system agrees with the old one, and what will you do when it doesn't? Nearly everything worth arguing about in a migration falls out of that one question.

Copying bytes is the boring part. The hard part is that the meaning of the old data is not written down anywhere. It lives in application code, in a handful of stored procedures, and in the heads of two people, one of whom left.

So define reconciliation before you write a line of extraction code. Row counts prove very little on their own: they pass happily while every amount is out by a factor of a hundred because one system stored pence and the other assumed pounds.

What works is a small set of invariants the business already trusts and already looks at. Total open balance by account. Record counts by status by month. Sum of postings per period per ledger. The fifty largest customers by value, in the same order, in both systems. Agree the tolerances in advance, in writing, and agree who signs.

Build that reconciliation pack as code and run it against the legacy system early, before anything has moved. If a number cannot be reproduced on the old platform, you will never be able to prove it on the new one, and you will spend cutover night arguing about which system is wrong.

Profile before you commit to a date

Migration timelines are driven by the number of surprises in the data, not by its volume. Terabytes are a procurement problem. Semantics are the schedule.

The surprises are consistent enough to list, and you will recognise most of them:

Profiling for this is mechanical and cheap: null rates, distinct counts, min and max, length and pattern histograms, and referential checks across every relationship, both the ones the schema declares and the ones only the application knows about. A fortnight of that work will change the plan more than any architecture diagram.

Then give a range, and say which specific findings would move it. An estimate produced before profiling is a guess wearing a confident face.

The cutover shape is the real architecture decision

Big bang gives you one reconciliation, one rollback point and one bad night. Phased gives you a smaller blast radius per step, at the price of keeping two systems in agreement while both are taking writes. That is the hardest engineering in the whole project, and teams usually discover it late, after the phasing has been promised to the business.

Whether phased is even available to you turns on a technical question that plans tend to skip: can the source system emit a change feed you can trust? Most legacy systems cannot. The updated_at column is maintained by some code paths and not others. Deletes leave no trace. The triggers were disabled during an incident years ago and never re-enabled.

Without a reliable feed, your options narrow to snapshot-and-diff on every cycle, which costs real compute and becomes its own project on any table without a dependable key, or freezing writes for the duration.

Get the order right. Establish what the source can tell you about change, then choose the cutover style, then set the downtime window. Done the other way round you promise a window you have no mechanism to keep.

Build the pipeline for the restart at three in the morning

It will fail partway through. Design for resumption rather than for the happy path.

Chunk by stable key ranges, not by offset and limit. Offsets shift under concurrent writes, and the result is rows silently skipped or loaded twice with no error anywhere.

Make loads idempotent: merge on a natural key so the same batch applied twice converges instead of duplicating. Keep transforms deterministic, which means no current timestamps baked into output and no surrogate keys generated fresh on each run unless you persist the mapping.

Quarantine rather than abort. Rejected rows belong in a table with the reason attached, visible and countable. Silent drops are the worst failure mode available, because the gap surfaces weeks later and by then nobody can say when the rows vanished or which run lost them.

Every manual fix has to become code. The classic incident is an engineer correcting forty records by hand during a rehearsal, production then running cleanly without that fix, and nobody noticing for a month. If the fix is not in version control, it did not happen.

Rehearse at full size, and rehearse the rollback

Subsets lie. They hide the index build that takes most of a night, the query that times out only above a certain cardinality, and the volume that fills the disk at four fifths of the way through. Rehearse on a full-size copy, end to end, with the clock running.

Time the parts nobody schedules: constraint validation, index and statistics rebuilds, the reconciliation pack itself, and the human sign-off, which will not happen at 4am simply because the plan says it does.

Then rehearse the rollback at least once, for real. A rollback that has never been executed is a paragraph, not a plan. While you are at it, write down the one-way doors. Once the legacy system is read-only and the downstream integrations have been repointed, going back may mean replaying a day of transactions by hand.

Set go/no-go criteria before the day, tied to the reconciliation tolerances you already agreed. Without them, the decision gets made at three in the morning by tired people who want to go home, and tiredness is reliably optimistic.

Decide what you are not migrating

The biggest lever on scope is the history you choose to leave behind. Old rows were created under rules that no longer exist, so each additional year of history adds transformation branches and edge cases rather than mere volume.

Usually better than full-fidelity history: migrate the active window, archive the rest into a queryable read-only store, and keep the legacy database available read-only if the licensing allows it. This needs a real conversation with finance and legal about retention and legal hold, and getting that answer will take longer than implementing it, so start it in week one.

Attachments and documents are a separate migration with their own identity mapping and their own storage behaviour. They are routinely discovered late, usually after the relational plan has been signed off.

Keep the old-to-new identifier mapping permanently. Support tickets, audits and integrations that cached the old identifiers will need it for years. Treating that table as scratch is a decision you will regret quietly, some time after everyone has moved on.

The costs that do not appear in the plan

I am not going to give you figures. Cloud and platform pricing is regional, it changes, and the total depends on your volumes and on how many reruns your data quality forces. There is no honest general number here. What is stable is the shape.

Egress is metered per gigabyte on the way out of most clouds, and a migration pays that toll once per rehearsal as well as once for real. Warehouse compute is metered per credit, per unit of processing or per query scanned, so a rerun-heavy rehearsal schedule carries a bill. Dual running means two platforms, two sets of licences and two operational burdens for as long as it lasts, and phased migrations are precisely the ones that quietly extend.

The scarcest resource is rarely compute. It is the small number of people who can look at a figure and say whether it is right. They have day jobs, and their availability, not your pipeline, sets the pace of sign-off. Book them early and in writing.

Moving onto hardware you own removes the egress meter and replaces it with capacity planning you now own. That is a genuine trade, not a free win, and it should be argued on its merits.

If you take one thing from this: do not let anyone hand you a date before the data has been profiled. The sequence that works is profile first, then write the reconciliation pack and prove it runs against the legacy system, then establish whether the source can report change, then choose the cutover style and only then the downtime window. Build loads that are idempotent and restartable, rehearse twice at full size including one real rollback, and agree go/no-go criteria in writing while everyone is still well rested. If you cannot answer how you will prove the new system is right, and what you will do on the morning it isn't, you do not yet have a migration plan.

Talk to us about this