Home  /  Insights

Stop citing the CHAOS report. Here is what the evidence on project delivery actually says

January 24, 2026 · PRINCE2 Agile

You have seen the slide. Some proportion of IT projects fail, some proportion are “challenged,” a small proportion succeed, and the numbers are always startling enough to justify whatever the presenter is selling.

Almost every one of those slides traces back to the Standish Group’s CHAOS report. I would like to persuade you to stop using it, and to give you something better to use instead — because the alternative is free, published by public auditors, and considerably more useful.

What CHAOS actually measures

The criticism is not that Standish’s data collection is sloppy, though the methodology has never been fully published. The criticism is more fundamental, and it was made in a peer-reviewed venue.

J. Laurenz Eveleens and Chris Verhoef of Vrije Universiteit Amsterdam published “The Rise and Fall of the Chaos Report Figures” in IEEE Software in early 2010. Their central finding:

“The Standish definitions of successful and challenged projects are solely based on estimation accuracy of cost, time, and functionality.”

A project is “successful” if it came in close to its estimate. That is all the measure captures.

Which produces the perversity the authors demonstrate:

“Some organizations tend to overestimate while others underestimate, so their success rates are meaningless because Standish doesn’t account for these clearly present biases.”

In one case organisation they examined, the CHAOS definitions scored the organisation as “very successful” while it was overestimating “from tenfold to a hundredfold.”

Read that carefully. Pad your estimates enough and you become a world-class project organisation. Estimate accurately and deliver something valuable slightly late, and you are “challenged.”

The measure rewards estimation conservatism and is blind to delivered value. It is exactly the wrong instrument for anyone trying to decide how to run projects, and it has shaped a generation of practitioner belief about how often projects fail.

Why it survives

Because it is convenient, and because everyone else cites it.

A statistic that says most projects fail is useful to anyone selling a method, a tool, a certification or a consultancy — my own profession very much included. I have sat in rooms where the CHAOS numbers were on the first slide and the remedy was on the second, and I have been in the audience for that presentation more times than I would like.

The honest position is less comfortable: there is no reliable, independent, generalisable statistic on project success rates. There are audited datasets for particular portfolios, and there is a large amount of vendor material that should be read as marketing.

Saying so is more credible than quoting a number you cannot defend. If a client asks you what proportion of projects fail, “nobody reliably knows, and here is why the figure you have heard is unreliable” is a better answer than a percentage.

What good evidence looks like

There is genuinely useful data. It is narrower than the CHAOS claim, which is precisely what makes it trustworthy.

The UK government’s major projects portfolio

The Infrastructure and Projects Authority merged with the National Infrastructure Commission on 1 April 2025 to become the National Infrastructure and Service Transformation Authority (NISTA). Its annual report publishes delivery confidence assessments for the Government Major Projects Portfolio.

As at 31 March 2025: 213 projects, £996 billion of whole-life costs, £742 billion of monetised benefits.

Delivery confidence:

RatingProjectsShare
Green3014%
Amber13563%
Red3115%
Exempt178%

£198 billion of whole-life costs sit in red-rated projects.

For comparison, the prior year’s IPA report covered 227 projects at £834 billion as at 31 March 2024.

Now the caveat, which is as important as the data. These are delivery confidence ratings, not method attribution. Nothing in the portfolio tells you whether a project used PRINCE2, PRINCE2 Agile, Scrum, SAFe or nothing at all. Anyone who presents the 14 per cent green figure as evidence for or against a methodology is misusing it — and I include people who would use it to sell PRINCE2 training.

What it is good for is calibration. Fourteen per cent green across the largest and most heavily governed project portfolio in the UK, with the best available assurance and the most senior sponsors, tells you that delivery confidence is hard-won everywhere. It is a useful antidote to the assumption that your organisation is uniquely bad at this.

The National Audit Office on mega-projects

The NAO published Lessons learned: Governance and decision-making on mega-projects in March 2025. Three findings are worth more than any survey:

“Where the government commits to budgets and timetables before it fully understands what is necessary to deliver the project and its intended value, we have sometimes seen significant cost increases or delays.”

“Mega-projects tend to take significant time to develop and deliver, often spanning parliaments.”

“Governance structures and processes can only go so far, particularly given the nature of mega-projects.”

That third quote is the one I would put in front of anyone about to buy a governance framework — including a client of mine. The UK’s public spending watchdog, having examined this for decades, says governance structures can only go so far. That is not an argument against method. It is an argument against expecting method to substitute for capability, honest estimating and the willingness to stop.

Individual audited cases

The most useful evidence of all is a small number of thoroughly investigated cases where an auditor established what actually happened. Audit Scotland on i6. The Auditor General of Canada on Phoenix. The Royal Commission on Robodebt. The NAO on ESN and HS2. Grant Thornton on Birmingham’s Oracle implementation.

Each of those is a properly investigated single case with named findings. They do not generalise into a percentage, and they are far more actionable than one — because each names a specific failure you can go and check for in your own organisation this afternoon.

The uncomfortable position on PRINCE2 Agile itself

I should apply the same standard to the thing I teach.

There is no published, independent effectiveness data for PRINCE2 Agile. No adoption survey with a defensible methodology, no controlled comparison, no outcome study. I looked, and I did not find one.

PeopleCert’s marketing states that more than two million professionals hold PRINCE2 certifications. That is a vendor figure, cumulative rather than current, unaudited, and with no breakdown by edition or by PRINCE2 Agile specifically. It measures certificates sold. It says nothing about outcomes.

So if you are evaluating whether to adopt PRINCE2 Agile, the honest position is this: the method is a coherent, well-structured codification of practices that experienced people have found useful, and there is no evidence base demonstrating that adopting it improves delivery outcomes. Anyone who tells you otherwise should be asked for the study.

I find that a more persuasive pitch than a percentage, and I would rather a client engaged me on that basis.

Measure your own portfolio instead

The most valuable dataset available to you is your own, and almost nobody collects it. Six measures, all obtainable, none requiring a tool purchase:

Forecast accuracy at each stage boundary, not just at the end. Plot forecast completion date against actual, at every gate. The shape of that line is the single most diagnostic thing about your organisation’s estimating culture. A line that stays flat until three months before delivery and then jumps tells you nobody was willing to escalate.

Benefits realised twelve months after closure. Very few organisations look. Those that do usually find that a third of the claimed benefits were never measured and a further third were double-counted with another project.

Proportion of projects stopped. Not failed — stopped, deliberately, by a board exercising its authority. If this is zero across several years, your governance is not functioning, whatever else it is doing.

Time from a problem first being written down to it reaching a decision-making forum. In i6 the programme board’s concerns escalated while external assurance reported green. In Robodebt a legal warning sat from 2014. Measure the lag.

Change in estimate at each re-baseline, and how many re-baselines a project has had. HS2’s Department and delivery body produced materially different estimates of the same scope — the Public Accounts Committee called that “a failure of governance and oversight.” Divergence between two internal estimates is a leading indicator.

Cost of the legacy thing you are replacing, per month of delay. ESN’s Airwave bill runs at least £250 million a year. Every programme replacing something has this number, and it almost never appears in the programme’s own reporting.

Collect those for two years and you will have better evidence about your organisation than any industry survey could give you — and it will be about you, which is the only thing that can actually change a decision.

What to say instead of the slide

When someone asks for the failure statistic, try this:

“The number you have seen is from the CHAOS report, and it measures estimation accuracy rather than success — a peer-reviewed paper showed an organisation scoring as very successful while overestimating by up to a hundredfold. There isn’t a trustworthy industry figure. What we do have is the UK government portfolio, where 14 per cent of 213 major projects are rated green, and a set of properly investigated cases that tell you exactly what goes wrong. Shall I show you the cases?”

It is a less dramatic opening. It is also the difference between being believed by a finance director and being tuned out by one.

How to read a statistic somebody hands you

Since I have spent this article attacking one number, it is only fair to give you the test I apply to any of them. Five questions, in order. Most industry statistics fail at the second.

1. What exactly was measured? Not what does it claim to show — what was the actual observation? CHAOS observes the ratio of outturn to estimate. That is a completely different thing from whether a project delivered value, and the gap between the two is where the misuse lives.

2. Who was in the sample, and who chose to be? Self-selected samples are the norm in industry surveys and they are systematically biased. Organisations that respond to a project management survey are organisations with a project management function that has time to respond.

3. Who paid for it, and what do they sell? This is not a knock-out blow — vendor research is sometimes good — but it should raise the standard of evidence you require. Apply it to me as well: I sell training and consulting, and you should read anything I write about the value of training and consulting with that in mind.

4. Is the methodology published in enough detail to be replicated? If not, the number is an assertion. Eveleens and Verhoef could critique CHAOS only by reconstructing the definitions from what Standish had published, which is itself part of the story.

5. Would the conclusion change if the definition changed slightly? If moving from “delivered within 10 per cent of estimate” to “delivered within 20 per cent” swings the headline from crisis to competence, the headline was never about reality.

A statistic that survives all five is worth quoting. In my experience roughly one in ten does.

The measure I would actually put on a board pack

If a board asked me for a single number to track project health across a portfolio, I would not give them a success rate. I would give them forecast stability: for each project, the number of months by which the forecast completion date has moved in the last twelve months, plotted against time.

It is easy to calculate from data you already hold. It is very difficult to game, because moving a forecast is a visible act. It reveals estimating culture rather than estimating accuracy — a portfolio where forecasts sit still for eighteen months and then jump six months in a single quarter is a portfolio where nobody feels able to escalate early.

And unlike any industry statistic, it is about you, which is the only kind of evidence that can change a decision.


Sources

  • J. Laurenz Eveleens and Chris Verhoef, The Rise and Fall of the Chaos Report Figures, IEEE Software, January/February 2010 — cs.vu.nl
  • NISTA, Annual Report 2024-25 (GMPP data at 31 March 2025; IPA/NIC merger) — gov.uk
  • Infrastructure and Projects Authority, Annual Report 2023-24 — gov.uk
  • National Audit Office, Lessons learned: Governance and decision-making on mega-projects, 14 March 2025 — nao.org.uk
  • Audit Scotland, i6 (March 2017); Office of the Auditor General of Canada, Phoenix Pay System (Spring 2018); Royal Commission into the Robodebt Scheme (July 2023); NAO reports on ESN and HS2; Grant Thornton on Birmingham City Council
  • PeopleCert, PRINCE2 7 Foundation certification page (“2M+” certification claim) — peoplecert.org