Maciej Nuzia · Advisory

I develop existing systems without starting from scratch

An old system may still handle the orders, documents and data a company needs while every change raises concerns about what might break elsewhere. Before I suggest rewriting it, I inspect the code, dependencies and release process. Only then can we tell what can be changed safely in the current system.

It helps if somebody who knows this system from the inside is on the call.

The rewrite

A rewrite has costs that are easy to miss at the start

The proposal sounds clean: the old system stays up, the new one grows beside it, and one day the traffic moves across. Four things go unmentioned in it.

  • for the whole length of the rewrite the company gets nothing new. The whole effort goes into reproducing what already works
  • the old system carries on living meanwhile. The law changes, a courier changes its file format, a new tax rate comes in. Every one of those has to land in two places at once, or the new system starts drifting from the old one on day one
  • the date comes from a picture of what the old system does, drawn from outside it, and the code always holds more than anybody remembers
  • there is one cutover day, and everything has to work on it at once: the data, the integrations, the habits of the people using it

The expensive part is somewhere else. The knowledge of how the business actually works is not in the documentation. It is in the code. A rule like ‘an order from this one country takes a different route’ can live in a single condition and nowhere else, with no record left of why anybody put it there. A rewrite means recovering every rule of that kind by asking around, one by one. The ones that get missed announce themselves later, through the people the gap hurts.

A rewrite is sometimes the only way out, and there is a list further down of the situations where that is so. Before the decision is made, though, it is worth counting how much is left to do without it.

Getting into the system

Where I start in somebody else’s code

The order comes from one assumption: in a system I did not write, the most expensive change is the one where nobody can tell what else it disturbed.

01

Reading the code

Where a request begins and where it ends, what happens to the data along the way, which files changed most often over the past year. The change history says as much about a system as the code does: it points at the places that had to be gone back to.

02

Running it outside production

Until I can bring the system up on my own machine, every change is a bet. Reproducing the environment can be harder than it looks: missing versions, configuration nobody wrote down, data without which nothing will start. That is why it comes first, and why I ask about it on the very first call.

03

Characterization tests

Before I fix anything, I write down in tests what the system does today. Including the behaviour that looks like a bug, because somebody outside may have got used to it and built their own work on top. A characterization test is a record of the state as it stands. It shows exactly what my fix changes.

04

A dependency map

What depends on what: library versions, nightly jobs, outside systems reaching into the same database, reports that may only come to mind once they stop arriving. The map is what shows which parts will come out cleanly and which are bound too tightly to the rest.

05

The first changes

Small and reversible, one at a time. After each one it has to be visible whether the system still does the same thing. A big change in an unfamiliar codebase has one flaw: when something breaks, you cannot tell which part of it was at fault.

Without a rewrite

What can change without rewriting the whole system

Which of these makes sense for you follows from the dependency map and from whatever costs you the most time today. None of them requires the company to stop while the work happens.

An API on top

The old system keeps working, and a second route into the data grows beside it: one route, visible, open to change. From that point on I write the next features on that new side.

Carving out one piece

One piece comes out of the system: the one that changes most often, or the one that breaks the most things around it. It gets its own data and its own release, and everything else is left where it is. Then the same with the next piece, or a stop after the first if that is enough.

Releases that can be undone

A list of things that have to line up before anything reaches production, and a way back should the change turn out to be the wrong one. As long as a release is an event, few changes get made, and every postponed one adds to the distance left to cover. That is why I take it on early.

What turns up when you look at it from the security side

Permissions, sessions, data exposed more widely than anyone intended, libraries with publicly known vulnerabilities. The review follows OWASP. It ends in a list split between what I fix straight away and what the company knowingly lives with, now at least named.

Upgrading dependencies

Library versions and the environment it all runs in. Dull and risky at the same time, so it goes in small steps, with the characterization tests as a safety net. Whatever no longer receives security patches goes first.

Every one of these stops halfway without damage: the system is left in a working state.

Straight answer

When it is better to rewrite the system

There are situations where being careful with the old system stops paying off. Better that you learn about them now than pay me to read code that is going in the bin.

  • the technology the system stands on no longer receives security patches, and raising the version means rewriting most of the code anyway
  • the system rests on something that can no longer be bought or maintained: hardware, a licence that is no longer sold, a database driver that will not run on newer operating systems
  • what the system has to do next has drifted away from what it does today. When the way the business works is itself changing, carrying the old rules across would mean carrying answers to a question nobody asks any more
  • the code cannot be brought up outside production and nobody can reproduce the environment. Step two above then fails, and there is nothing left to build the rest on
  • the system is small enough that all this caution costs more than writing it once more

A rewrite then comes with one condition that is easy to forget in the relief of having finally decided: the old system has to be written down before it is switched off. The rules inside it are the only documentation of the process the company has.

If your situation is on that list, that is the first thing I will tell you. I would rather lose this engagement now than in the middle of careful changes that end in a rewrite anyway.

Talk

Tell me what happened during the last change

Three things are enough to start: what it is written in, roughly how old it is, and what happened the last time anybody went into that code. There are open slots in the calendar, and email works just as well.