Home  /  Insights

Discover to Support: ITIL 5’s eight-stage lifecycle, and the stage it forgot

July 25, 2026 · ITIL 5

At 08:32 on Monday 28 August 2023 — a bank holiday, the busiest travel day of the British summer — a single flight plan arrived at the UK’s air traffic control system. Within twenty seconds, the primary flight plan processing system and its hot standby had both shut themselves down.

Nobody attacked anything. No hardware failed. No change had been deployed. The system did exactly what it had been designed to do, and over 700,000 passengers had their journeys disrupted.

I want to use that morning to talk about the most substantive change in ITIL Version 5, because the two are more closely connected than they look.

What actually changed in the model

ITIL 4 gave us six value chain activities: Plan, Improve, Engage, Design and Transition, Obtain/Build, and Deliver and Support. ITIL 5 replaces that — or restates it, depending on which commentator you follow, and the sources genuinely disagree — with an eight-stage product and service lifecycle:

Discover → Design → Acquire → Build → Transition → Operate → Deliver → Support

The official descriptions are straightforward. Discover is understanding what the market needs and aligning it with strategy. Design is planning the solution through prototypes and specifications. Acquire is obtaining the resources, by buying, hiring or building. Build is coding, configuring, assembling and testing. Transition is moving the product safely into production. Operate is keeping infrastructure and systems running and monitoring performance. Deliver is making the service available for consumption, managing access and requests. Support is resolving incidents and problems and restoring normal service.

Two of these separations matter a great deal, and they are the reason I think this is a real improvement rather than a redrawn diagram.

Design and Transition are now distinct stages. In ITIL 4 they were welded together into one activity, which never matched how any organisation I have worked with actually operates. Design is a set of decisions made months before anyone touches production. Transition is a specific, high-risk, time-boxed event with its own governance, its own approvals and its own failure modes. Treating them as one thing meant that transition governance was routinely inherited from design assumptions that were no longer true. Splitting them forces the question: is the thing we are about to move into production still the thing we designed?

Operate, Deliver and Support are now three things instead of one. ITIL 4’s “Deliver and Support” bundled together keeping the platform alive, making the service available to a user, and fixing it when it breaks. Those are three different accountabilities, usually held by three different teams, measured on three different numbers. Collapsing them into one activity is precisely how organisations end up with an infrastructure team whose dashboards are green while customers cannot log in.

The trainer Ben Kalland called the lifecycle model “a clear improvement.” I agree with him, with one significant reservation.

The stage that is missing

Gil Regev of itecor made the observation that I have not seen anyone answer properly: “The lifecycle view is not a complete cradle to grave sequence. It misses the crucial activity of Decommissioning.”

He is right, and it is not a pedantic point. Ask any CIO what they actually run and you will get a confident answer about the top thirty applications and a vague one about the rest. Decommissioning is where organisations discover the server nobody owns, the interface that three business processes silently depend on, the licence that has been renewing for eleven years, and the database that holds personal data nobody can lawfully justify keeping. It is also the single most reliable source of cost reduction available to most IT organisations, and the least funded.

A lifecycle that runs from Discover to Support describes how things come into the world and how they are kept alive. It says nothing about how they leave. Given that ITIL 5 has moved IT Asset Management into the product and service management group and made “planning and managing the full lifecycle of all IT assets” its stated purpose, the omission is odd.

If you adopt the eight stages, add a ninth. Call it Retire. Then go and find out how many things in your estate should already be in it.

What NATS shows about stage boundaries

Back to the bank holiday.

The system in question is FPRSA-R, which converts flight plan data from the European format into the UK National Airspace System format, processing over 800 flight plans an hour. The flight plan that arrived at 08:32 was a transatlantic route containing a rare combination of six specific attributes. Two geographically distinct waypoints, roughly 4,000 nautical miles apart, happened to share the same identifier, sitting on either side of UK airspace, with the real UK exit point absent from part of the plan.

FPRSA-R could not reconcile the route. Rather than risk passing incorrect data to controllers — and I want to be clear that this was the safe behaviour, correctly designed — it raised a critical exception and placed itself into maintenance mode.

Then the backup received the identical flight plan, applied the identical logic, raised the identical exception, and did the same thing.

Both systems were down in under twenty seconds.

NATS’ own major incident investigation put it plainly: “The set of data in the flight plan meant that there was a grouping of six distinct attributes which, taken in combination, created a unique exception that the system was unable to process.” And: “If any one of these attributes had not been present, the flight plan would have been processed normally.”

A hot standby is not redundancy

This is the lesson I would take into every design review from now on. A standby running the same code against the same input is not redundancy against a deterministic, data-triggered fault. It is a second copy of the same decision. It protects you against hardware failure, power loss and a corrupted file system. It protects you against nothing at all in the class of failure that NATS hit.

Where in the lifecycle does that get caught? Design. Not Build, not Transition, not Operate. It is an architectural assumption, made early, that never got tested because the test that would have found it — feed both nodes the same malformed input and see what happens — is a design-validation test, not a functional one.

ITIL 5 separating Design from Transition does not by itself make anyone ask that question. But it creates a stage boundary at which somebody could be made accountable for asking it. Boundaries are where governance can attach. That is the whole practical value of a lifecycle model, and it is why the number and placement of the stages is not a cosmetic matter.

Why it took nearly six hours

Restoration came at 14:27. Almost six hours after 08:32. The reasons are a catalogue of everything that lives in the Operate, Deliver and Support stages.

Manual entry into the National Airspace System achieves 40 to 60 flight plans an hour, against an automated capacity of over 800. Traffic restrictions were unavoidable from the first minutes.

NAS holds a rolling four-hour buffer of flight plan data. Because flight plans are continuously amended, that stored data degraded as the outage ran on, and once the four hours had elapsed everything had to be entered by hand. There was a clock running on the recovery that had nothing to do with the fault itself.

Level 1 engineering — reboots and standard procedures — failed. A Level 2 specialist needed 95 minutes of travel to reach the site. Restart attempts kept failing, because unidentified messages sitting in a pending queue outside the pause queue repeatedly re-triggered the same exception.

And then the finding that I think every organisation with a critical supplier should read twice: only the original manufacturer understood the interaction between the two systems well enough to diagnose it. The supplier was not contacted until 12:39 — four hours into the incident. They identified the queued-message problem and directed the reprocessing at 12:58. Restoration followed at 14:27.

Four hours before the only party who could actually solve the problem was called.

That is not an incident management failure in the narrow sense. The incident was being managed. Bridges were open, people were working, escalation was happening within the organisation. It is a supplier management failure — an Operate-stage dependency that had never been rehearsed. Somebody, somewhere, had the supplier’s emergency contact in a contract. Nobody had a trigger that said: if the fault is in the AMS-UK interface, call Frequentis in the first thirty minutes, because we cannot fix this ourselves and pretending otherwise costs an hour a try.

The scale, and what came of it

Over 700,000 passengers were affected: roughly 300,000 by cancellations, 95,000 by delays of over three hours, and 300,000 by shorter delays. Around 1,500 flights were cancelled on the day, and disruption ran for several days afterwards as aircraft and crew ended up out of position over a bank holiday weekend.

The Civil Aviation Authority commissioned an Independent Review Panel, which reported in March 2024 with 34 recommendations. Its chair, Jeff Halliwell, did not soften the conclusion: “The incident on 28 August 2023 represented a major failure on the part of the air traffic control system.”

Note what the recommendations covered — contingency planning, earlier warning to airlines, passenger support, the incentive framework applying to NATS, and an incident management forum. Almost none of it is about the software defect. The defect was a genuinely rare data combination in a system that behaved safely. What the review went after was everything around it: the assumptions, the rehearsals, the escalation paths, the arrangements with the people downstream.

Costs to airlines and airports were described in the official report only as “substantial.” Press estimates around £100 million circulated widely, but those are third-party numbers and I would not put them in a board pack without saying so.

What to do with this on Monday morning

Four things you can do without buying anything.

Find your correlated failures. Take your three most critical services. For each, ask: does the standby run the same code as the primary? Does it consume the same input from the same source? If the answer to both is yes, you have a NATS-shaped exposure. The mitigation is not necessarily a different technology stack — it may simply be a documented, rehearsed manual path with a known throughput, and an honest number for how long your buffered data stays usable.

Put a clock on your degraded mode. NATS could run manually at 5 to 7 per cent of automated capacity, and the four-hour data buffer set a hard deadline for recovery. Do you know the equivalent numbers for your operation? Not the aspiration in the continuity plan — the measured throughput of your workaround, and the point at which it stops working.

Write supplier escalation triggers, not supplier contacts. A phone number in a contract is not a plan. The plan is a rule: these symptoms mean we cannot fix it ourselves; call this supplier within N minutes; here is who is authorised to make that call at 3am on a public holiday. Then test that the number answers.

Use the new stage boundaries deliberately. If you are adopting the ITIL 5 lifecycle, do not treat Design → Transition and Operate → Deliver → Support as new vocabulary for an old diagram. Each boundary is a place to put a question that somebody must answer before work crosses it. Is the thing we are transitioning still the thing we designed? Who is accountable when the platform is healthy and the customer still cannot get service? Those are the questions the model makes room for. It will not ask them for you.

And add the ninth stage. Retire is where the money is.

A worked example: one change, eight stages

Abstractions are easy to agree with. Here is what the model looks like applied to a single, ordinary piece of work — replacing an ageing payment reconciliation service — and what question belongs at each boundary.

Discover. What is the business actually asking for? Not “replace the reconciliation engine” but “reduce the unreconciled balance at end of day and shorten the finance close.” Numbers, owners, and the cost of the current state. Boundary question: can we state the outcome in a measure somebody outside IT already reports?

Design. Architecture, controls, failure modes, data flows. This is where the NATS lesson lives — if you specify an active-standby pair, this is the only stage at which anyone will ask whether the standby shares a failure mode with the primary. Boundary question: which failures does this design not survive, and have we written them down?

Acquire. Buy, build or hire. Supplier selection, licences, skills. Boundary question: what have we just made ourselves dependent on, and what is our escalation path to them at 3am?

Build. Code, configure, assemble, test. Boundary question: does our test coverage exercise the paths that only run in failure — or only the paths that run on a good day?

Transition. The move into production. Approvals, rollback, rehearsal, cutover. Boundary question: is the thing we are transitioning still the thing we designed, and if not, who re-approved the difference? TSB’s board never asked this one.

Operate. Keeping it running. Monitoring, capacity, patching, the DR drill. Boundary question: does our recovery capability depend on anything inside the failure domain?

Deliver. Making it available and usable — access, requests, entitlements. Boundary question: is the platform being healthy the same thing as the customer being served? Usually it is not, which is precisely why ITIL 5 separated these.

Support. Incidents, problems, known errors. Boundary question: where do known errors from this service go, and who owns them?

Retire — the stage you have to add yourself. Boundary question: what does this replace, when does that thing switch off, and who is accountable for confirming it did?

Ten minutes with that list will tell you more about a programme’s real risk than a RAG status ever has. Note that six of the ten questions are about failure, and that in most governance forums none of them get asked, because the agenda is built around delivery dates rather than stage boundaries.

One further practice worth borrowing, from Meta’s account of its own six-hour global outage in October 2021: the company credited its regular “storm drill” programme — deliberately taking services, data centres or entire regions offline — with giving it the playbook and the confidence to sequence its recovery safely, rather than restoring everything at once and triggering a second cascading failure. Storm drills are an Operate-stage practice. They are also the only reliable way to discover which of your recovery assumptions are fiction.


Sources

  • ITIL.com, ITIL Foundation (Version 5): what’s new (the eight-stage lifecycle) — itil.com
  • ITSM.tools, ITIL (Version 5) vs ITIL 4: key changes (lifecycle model, Ben Kalland quote) — itsm.tools
  • Gil Regev, itecor, A new ITIL, so what? (the decommissioning omission) — itecor.com
  • NATS, Major Incident Investigation report: 28 August 2023 flight plan processing failure (September 2023)
  • Civil Aviation Authority, Independent Review of NATS’ Flight Planning System Failure on 28 August 2023, CAP2993 (March 2024)