Of the thirty-four practices in ITIL, problem management is the one I most often find written down and least often find funded.
The pattern is consistent enough to be predictable. There is a problem management procedure. There is a problem manager, usually as part of a wider role. There is a monthly meeting. The meeting reviews a list of open problem records, most of which are older than the person chairing the meeting, and closes the ones where the underlying system has since been replaced. Nobody is doing anything wrong. It simply has no consequence.
Meanwhile incident management is fully staffed, because incidents shout.
ITIL 5 defines problem management as “the practice of reducing the likelihood and impact of incidents by identifying actual and potential causes of incidents, and managing workarounds and known errors.” That last clause is the one that carries the weight, and it is the one organisations skip.
A known error is a problem that has been analysed but not resolved. You know what is wrong. You know what happens if it is triggered. You may have a workaround. And you have decided — actively or by default — to live with it for now.
Almost every serious IT failure I have studied was a known error that nobody owned.
Twelve May 2017
The WannaCry ransomware propagated using the EternalBlue exploit against SMBv1 on unpatched Windows systems.
Microsoft had released the patch, MS17-010, in March 2017. NHS Digital issued critical alerts telling NHS organisations to patch in March and again in April 2017.
On 12 May, at least 81 of the 236 NHS trusts in England were infected or had services disrupted. 603 primary care and other NHS organisations were infected, including 595 GP practices. Five accident and emergency departments diverted patients — in London, Essex, Hertfordshire, Hampshire and Cumbria. The National Audit Office estimated that over 19,000 appointments were cancelled, of which 6,912 were identified with certainty and the remainder extrapolated.
No NHS organisation paid the ransom.
The NAO’s assessment is one of the least forgiving sentences a public auditor has ever written about an IT failure: “It was a relatively unsophisticated attack and could have been prevented by the NHS following basic IT security best practice.”
And: “All organisations infected by WannaCry shared the same vulnerability and could have taken relatively simple action to protect themselves.”
Correcting a myth while we are here
The popular version of this story is that the NHS was running Windows XP. The NAO found otherwise: the dominant problem was unpatched but supported Windows, principally Windows 7. XP was a minority of the affected estate.
That distinction matters enormously for anyone trying to learn from it. “We were running an obsolete operating system” is a capital-expenditure problem — expensive, slow, sympathetic. “We were running a supported operating system and had not applied a patch published two months earlier” is an operational discipline problem. The second is much cheaper to fix and much more uncomfortable to admit.
The NAO added a further finding that should be read carefully by anyone whose defence rests on perimeter controls: “taking action to manage their firewalls facing the internet would have guarded organisations against infection.” There were two independent ways to be safe. Neither was taken.
The missing loop
Here is the finding that turns this from a security story into a problem management story.
“The Department had no formal mechanism for assessing whether local NHS organisations had complied with their advice and guidance.” — National Audit Office
Advice was issued. Twice. Correctly, in good time, by the right body. And there was no mechanism to find out whether anybody had acted on it.
That is a known error without an owner. The vulnerability was documented. The remediation was known, published and free. What was absent was the loop that closes a known error: somebody accountable for confirming that the fix had actually been applied, everywhere it needed to be applied, and for escalating where it had not.
Issuing guidance feels like discharging a responsibility. It is not. It is the first half of one.
I have seen this exact shape in commercial organisations many times, usually in the form of a security advisory circulated to a distribution list, or a vendor patch note filed against an application whose support team was reorganised eighteen months ago. The advisory goes out. A tick appears in a compliance register. Nobody asks the second question.
And then the coordination failed too
A second-order finding from the NAO, easy to overlook and worth its own paragraph: “Many local organisations could not communicate with national NHS bodies by email as they had been infected by WannaCry.”
The channel through which the national response was to be coordinated was itself a casualty of the incident. It is the same failure mode as the Rogers incident team relying on the Rogers network, and the same lesson: any recovery capability that depends on the thing that has failed is not a recovery capability.
The NAO also found that the Department had never tested a national response to a cyber attack, so roles were unclear when one happened.
The cost nobody could calculate
The Department of Health and Social Care eventually estimated the cost at £92 million — £19 million of lost output from cancelled appointments and operations, and £73 million of IT costs to recover data and restore systems.
But read the caveat the Department published alongside it: “No data was systematically collected on the costs of recovering IT.” The £92 million was assembled after the fact from assumptions — for instance, that each of eighty severely affected trusts required five days of full-time IT specialist support.
Before that estimate existed, the NAO had stated flatly that “the Department does not know how much the disruption to services cost the NHS.”
I find this the single most useful detail in the whole case, and it has nothing to do with ransomware. An organisation that cannot price its own outage cannot make a rational investment decision about preventing the next one. Every business case you have ever seen rejected for lack of quantified benefit was rejected in an organisation that did not measure the cost of failure. The two facts are the same fact.
If you take one action from this article, make it this: the next time you have a significant incident, capture the cost while the incident is running. Staff hours, contractor time, lost transactions, waived fees, overtime, the goodwill you gave away. Not perfectly. Just consistently. Within a year you will have the only number that reliably moves a budget.
What a working problem management practice looks like
Six characteristics. None of them is exotic and all of them are unglamorous.
A known error database that people actually read. Not a list of problem tickets. A register of things that are wrong, with the symptom, the affected components, the workaround, the conditions under which the workaround stops working, and the reason it has not been fixed. Written so that a service desk analyst at 2am can find it and use it.
Named ownership per known error, at a level that controls budget. A known error owned by “Infrastructure” is owned by nobody. If nobody who can fund the fix owns the record, it is a note, not a control.
A closure loop with evidence. The NHS finding in one line. When you issue guidance, patch instruction or remediation direction, define in advance how you will confirm it happened, and who chases the gap. If you cannot verify it, you have not remediated it — you have communicated about it.
Deliberate acceptance, dated and reviewed. Living with a known error is often the correct commercial decision. It becomes dangerous only when it is a default rather than a decision. Give every accepted known error an expiry date and a named acceptor. Revisit it. Southwest Airlines described its crew optimisation software as having “a functional gap that was revealed in December” — revealed, note, not created. That gap had been accepted, implicitly, for years. Nobody ever signed for it.
Proactive problem management with time protected for it. ITIL 5’s AI Capability Model calls this pattern-finding work Cognition — “identifying patterns and hidden insights for proactive problem detection.” This is one of the genuinely valuable applications of machine learning in service management, and it is far less fashionable than putting a chatbot on the portal. Trend analysis across incident records will tell you where your next major incident is coming from. It requires somebody to have the time to look.
A standing item at a governance forum that controls money. Problem management fails when its outputs are reported to a forum that can only note them. The known errors accepted this quarter, and their expiry dates, belong in front of the same committee that approves the capital plan.
Where this sits in ITIL 5
Problem management is in the Product and Service Management practice group, and in the qualification scheme it sits in the Monitor, Support and Fulfil module with service desk, incident management, service request management, and monitoring and event management. That grouping is sound: these five are the practices that decide what your users experience day to day.
Note that the ITIL 5 practice guides themselves have not been reissued yet — the thirty-four practice publications currently in circulation are still the ITIL 4 documents, with updates scheduled for the second half of 2026. So the guidance you have is the guidance you had. That is not an argument for waiting. Nothing in this article requires a new edition.
The uncomfortable summary
WannaCry cost the NHS an estimated £92 million, disrupted a third of England’s hospital trusts, and cancelled over nineteen thousand appointments.
The patch was free. It had been available for two months. Two separate national alerts had been issued. Managing the internet-facing firewalls would have worked instead.
The failure was not technical, and it was not really a security failure either. It was the absence of a loop — the thing that turns a known error from a piece of information into a managed risk.
Go and look at your own known error register. If you do not have one, that is the finding. If you do, ask three questions of the ten oldest entries: who owns this, when did we decide to live with it, and when does that decision expire?
I have never asked those questions of an organisation and had comfortable answers.
How to run a problem review that has consequences
If your problem review meeting has no consequences, no amount of process documentation will give it any. Here is the shape of one that does. It takes an hour a month and it is entirely unglamorous.
Bring three lists, not a ticket queue.
The first is repeat incidents: anything that has recurred three or more times in the quarter, ranked by total customer-minutes lost rather than by ticket count. Ticket count flatters small, noisy faults and hides the quiet ones that take out a business process for two hours.
The second is known errors approaching or past their review date. Every accepted known error should carry an expiry. Past-date entries are the agenda.
The third is changes that failed or were backed out, with a one-line cause for each. Over six months this list will show you two or three systemic causes, and they will not be the ones people assert in the room.
Give every item one of four outcomes, and no others. Fix it, with an owner and a date. Fund it, meaning it goes to whoever holds the budget with a written cost of not fixing it. Accept it, with a named acceptor and a review date. Or close it, because the underlying system is gone. “Under investigation” is not an outcome, and an item that carries it twice should be escalated on that basis alone.
Put a number on the cost of each accepted item. Not a precise number — a defensible one. Customer-minutes lost, staff hours consumed, transactions failed, goodwill paid out. The NHS could not price WannaCry after the fact because nobody had collected the data during it. Collect it during. Within a year you will have the only evidence that reliably moves a budget, and you will be able to answer the question every finance director asks: what does this actually cost us today?
Report upward on two things only. The number of accepted known errors past their review date, and the aggregate cost of the accepted list. Those two numbers, trending, tell an audit committee more about operational risk than any RAG status. They also make acceptance visible, which is the entire point — a risk somebody has consciously accepted and priced is governance; the same risk unrecorded is an accident waiting for a date.
And protect time for proactive work. If every hour of the problem management function is consumed reacting to major incidents, you have an incident management overflow team, not a problem management practice. Ring-fence a day a fortnight for trend analysis. It is the first thing sacrificed and the last thing anyone regrets protecting.
Sources
- National Audit Office, Investigation: WannaCry cyber attack and the NHS, HC 414 (October 2017)
- Department of Health and Social Care, Securing cyber resilience in health and care: October 2018 update and parliamentary written answer on estimated costs (October 2018)
- Southwest Airlines Co., statements on crew optimisation software following the December 2022 disruption
- ITSM.tools, ITIL (Version 5) management practices (problem management definition, practice grouping) — itsm.tools
- ITSM.tools, ITIL (Version 5) explained: key changes, lifecycle, AI governance (practice guide update timing) — itsm.tools
- PMG Academy, The definitive guide to ITIL Version 5 Foundation (AI Capability Model — Cognition) — pmgacademy.com