ISSUE 2026-09-27SERIES BREC 38TYPE BRIEFgarnetgrid.com

The question underneath the question

Insights

Most migration plans start with an argument about big bang versus phased. That argument is a proxy for a harder question about dual running and arbitration, and answering it properly determines almost everything else: what you can slice, what you must rehearse, and how you will know you are finished. What follows is how we think about it, and the failure modes that tend to be waiting.

The big bang versus phased argument is usually standing in for the question nobody wants to put on a slide: for how long are you willing to run two systems as one business, and when the two disagree about what a customer owes you, which one is right? That argument is usually standing in for the question nobody wants to put on a slide: for how long are you willing to run two systems as one business, and when the two disagree about what a customer owes you, which one is right? Answer that, with a name attached to the arbitration, and the sequencing largely falls out of it. Leave it unanswered and you will settle it during cutover weekend instead, at speed, with the finance director on the phone.

The old system is the specification

Whatever documentation exists describes the system somebody intended to build. The running code describes the one your business actually depends on. Between the two sit years of accretion: a discount that rounds a particular way, an order type that skips a credit check because a manager asked for it a decade ago, a field that means one thing when it is empty and something else when it contains a single space.

These are not all bugs to be tidied away. Some of them are load-bearing. A rounding quirk becomes load-bearing the moment a customer has built their own reconciliation to match it. You discover which ones matter by changing them and seeing who complains, which is a very good argument for doing that somewhere other than a live ledger.

The old system is also the performance specification, and this is the part that gets missed. Staff complete an order line in a certain number of keystrokes, with no network round trip, against a denormalised schema on one box. A replacement can be faster on every synthetic throughput measure and still be slower on the one loop a member of staff repeats all day — count how often that loop actually runs in your operation, then measure it. Measure that loop before you promise an improvement, because if it regresses, the workarounds start immediately and you will not hear about them.

Find the consumers before you touch the data model

The integration list you are handed will be incomplete. It contains the interfaces that have owners. It will not contain the spreadsheet with an ODBC connection into a production table, the small database in the warehouse office, the overnight file that lands on an SFTP server and is collected by a bank, or the report somebody rebuilt in a BI tool by querying the transaction table directly.

You find these by instrumenting, not by asking. Database audit logs, connection-level accounting, firewall and file-transfer logs, gathered over a window long enough to catch the monthly and quarterly work. The quarterly job is the classic miss: you cut over in March and break something that only runs at the end of June, by which time nobody connects the two events.

Every consumer you find becomes a decision: repoint it, replace it, or keep feeding it. Keeping it fed through a compatibility view that preserves the old column names and the old sentinel values is often the right call, and it is almost always presented as defeat. Keeping a compatibility view alive is often cheaper than putting a renegotiation with every downstream consumer on your critical path — count your own consumers and price both options before you decide.

Slice by entity, not by module

A phased migration needs a slicing key: something you can move a subset of, run in both systems, and reconcile. The tempting slice in ERP is by module. Finance first, then inventory, then sales. It is usually the wrong one, because a single business transaction crosses all three. Slice that way and you have placed a distributed transaction boundary in the middle of your ledger, and you will spend the programme learning about eventual consistency in the one place where the business expects arithmetic.

The slices that tend to work are the ones the organisation is already shaped by: a legal entity, a country, a warehouse, a brand, a distinct customer segment. Those are naturally separable because tax, reporting and operations are already separated along the same lines. You can move one and genuinely leave the rest alone.

If you cannot name a slicing key, you do not have a phased plan. You have a big bang with extra steps, and it is worth saying so early, because it changes what you must buy: far more rehearsal, far better data repair tooling, and a cutover runbook that somebody has actually executed under failure conditions.

Shadow-running is the cheapest evidence you will ever buy

Before the new system is allowed to make a decision, let it make the decision and then throw the answer away. Feed it the same inputs as production, compute the same outputs, compare them, and act on neither. Pricing, tax, availability, allocation, posting: all of it can be run in the dark.

What comes out is a diff stream, and the diff stream is your actual project plan. Ranked by frequency and by money, it tells you what to fix next without anybody guessing. It also gives you the only defensible answer to "are we ready", which is not a percentage but a shape: this many transaction types, over this many days, at this residual difference, with every remaining difference explained and accepted in writing. An unexplained difference closed by widening a tolerance is a future incident with a date on it.

Shadow-running is also the honest answer to the test data problem. You need production-shaped data to exercise the edge cases, and anonymising ERP data tends to destroy precisely the edge cases you are hunting. Running the comparison inside the controlled environment, under whatever data protection controls apply to you, is usually better than shipping a scrubbed copy out to a development estate where it is both less useful and less safe.

One discipline to agree up front: a change budget. While you dual-run, every change applied only to the legacy system widens the gap you are trying to close. A full feature freeze for several quarters is not realistic, so decide instead that changes land in the new system, or land in both with the duplication documented, and that anything legacy-only requires somebody senior to say yes.

Old data encodes decisions, not just records

Migration is presented as an extract-transform-load exercise with a mapping document. It is really an interpretation exercise, and the interpretations need owners. What you will find, more or less reliably:

Each of these needs a human decision and a written rule, and the rules are worth more than the scripts that implement them.

Keep the old identifiers. Renumbering looks like a tidy-up and is a trap: the old reference is printed on labels, quoted in customer purchase orders, pasted into emails and typed into search boxes by people who are on the phone to a customer. Carry it as a first-class, indexed, permanently supported field, and make sure searching by it works for as long as anyone remembers it exists.

There is no rollback, only fix-forward

Every plan has a rollback step. Ask what it actually does once the new system has accepted a full day of writes. Reversing a day of trading is not restoring a backup. It is reconstructing state across a ledger, a warehouse and a set of customer communications that have already been sent. Once the new system has taken a meaningful volume of writes, the practical rollback window has closed, and it closes without announcing itself — find out where yours actually closes during rehearsal rather than assuming a duration.

Be honest about that and spend the money differently. Data repair tooling you have used in rehearsal rather than written about. A documented manual procedure for continuing to trade while the system is unusable, including who is allowed to authorise it. A named person with the authority to stop. And a rehearsal of partial failure, not only of the clean run, because the clean run is not the one you are being paid to survive.

Calendars, archives, and knowing when you are done

Two calendar facts shape ERP cutovers more than any technical choice. Do not cut over mid-period if you have any say in it. And accept that the old system stays in use longer than the plan claims, because adjustments, audits and statutory reporting all reach backwards into periods you thought you had left behind.

Decommissioning therefore rarely means deleting anything. It means the old system becomes a read-only, access-controlled artefact, with documented access and a named owner, often for years. Keeping one small machine alive for that purpose is frequently cheaper and easier to defend to an auditor than extracting a decade of records into a queryable form somewhere else. If that is the right answer for you, choose it deliberately rather than sliding into it.

Write the exit criteria before you start. The interfaces repointed and the old ones switched off. The reconciliation tolerance met and signed by someone who carries the number. A full period closed end to end in the new system. The old system read-only. Without criteria in writing, migrations end when everybody is exhausted, and exhaustion is not the same thing as done.

Three things, in this order, before anyone argues about dates. Instrument the old system and find out who is actually reading from it, including the spreadsheets and the monthly jobs. Name your slicing key, and if you cannot name one, say plainly that you are planning a big bang and fund the rehearsal accordingly. Then build the comparison harness and start shadow-running, before you write a line of migration code. Everything after that is ordinary engineering against a diff list you can see. Everything before it is guesswork with a Gantt chart on top.

Talk to us about this