Home  /  Insights

When the outage is national news: major incident management at Rogers and Optus

September 19, 2026 · ITIL 5

There is a moment in every serious incident that separates organisations that recover quickly from organisations that make the evening news. It is not the moment of failure. It is the moment somebody tries to convene the people who can fix it, and discovers they cannot.

At Rogers Communications in Canada, incident staff tried to coordinate over Rogers’ own mobile and internet services. Those services were the outage. The team had to obtain SIM cards from competitors.

Sixteen months later at Optus in Australia, the company communicated with ten million customers primarily through social media — for four hours, while those customers had no internet.

Both are documented in official investigations. Both are worth studying properly, because they are not stories about exotic technology. They are stories about assumptions.

Rogers, 8 July 2022: twenty-six hours

Rogers was executing a seven-phase upgrade of its IP core network. The overall multi-phase programme had been risk-assessed as High. Phase 6, on 8 July, was scored Low — which under Rogers’ own change process meant it bypassed mandatory lab testing and higher-level approval.

Phase 6 removed an Access Control List policy filter from the configuration of the distribution routers. The intent was to allow direct distribution of wireless DNS addresses of the cloud infrastructure into OSPF.

With the filter in place, roughly 10,000 routes were being managed. Without it, a single distribution router released over 900,000 routes into the core. The core routers’ CPU and memory were exhausted. In the words of the independent assessment commissioned by the Canadian regulator: “The core network routers crashed within minutes from the time the policy filter was removed.”

Filter removed at 04:43 EDT. Core gateways failing at 04:45. All Rogers services ceased.

More than 12 million customers lost wireless and wireline service, including the Fido and Chatr brands, plus wholesale and corporate customers. Interac e-Transfer and electronic payment services were disrupted nationally — a country’s retail payments, degraded by a routing change. 9-1-1 emergency calling stopped working. The precise emergency call statistics are redacted in the public version of the report.

Service was substantially restored by 07:00 on 9 July. Approximately twenty-six hours.

Why it lasted so long

Three findings from the independent assessment, and each one is a design decision rather than a mistake.

There was no overload protection on the core routers. No configured maximum for acceptable routing data. The assessment named this as the single most consequential preventable gap: “The July 2022 outage exposed the absence of overload protection on the core network routers.” A limit would not have prevented the configuration error. It would have prevented the error from destroying the core.

The management network depended on the failed IP core. Remote engineers could not reach the management plane of the equipment they needed to fix. Staff had to be physically dispatched to sites. There was also no initial access to error logs from the failed routers, which is why root cause was not identified for roughly fourteen hours.

And the incident team could not communicate. “Rogers staff relied on the company’s own mobile and Internet services for connectivity to communicate among themselves.”

Read those three together. Each is individually defensible. Together they describe an organisation whose ability to respond was entirely contained within the thing that had failed.

Notice also what the assessment did not find. It concluded the outage was not the result of a design flaw in the core network architecture — the converged wireless and wireline design was “typical of what would be expected of such a Tier 1 service provider.” The problem was not that the architecture was unusual. It was that the recovery capability had inherited every one of the architecture’s dependencies.

The communications timeline

This is the part I would put in front of a crisis committee.

  • 04:43 — filter removed
  • 04:45 — services begin failing
  • 08:39 — 9-1-1 providers notified. Three hours and fifty-six minutes in.
  • 08:54 — first public notification, via Twitter. Four hours and eleven minutes in.

Four hours before customers were told anything, on a network carrying emergency calls.

Rogers’ CEO told the House of Commons industry committee: “We failed to deliver on our promise to be Canada’s most reliable network.” The company gave all affected customers five days of service credit, committed $10 billion over three years to network reliability and a minimum of $250 million specifically to physically separate the wireless and internet networks. Canada’s major carriers signed a memorandum of understanding on emergency roaming and mutual assistance within two months.

Optus, 8 November 2023: ten million customers

During a planned software upgrade at a Singtel internet exchange node in North America, Optus’s network received what it called an “overload of IP routing information” through an alternate peering router while the primary was under maintenance.

Approximately ninety Cisco provider-edge routers hit their BGP prefix-limit thresholds and, on exceeding them, disconnected themselves from the network.

Here is the detail that makes this case so instructive. Optus told the Australian Senate inquiry: “These self-protection limits are default settings provided by the relevant global equipment vendor (Cisco).”

Vendor defaults. Not a value someone had chosen for this network, at this scale, in this topology. And because the protective mechanism isolated the routers rather than degrading gracefully, recovery required resetting routing connectivity and, for many elements, physical reboots, followed by carefully staged reintroduction of traffic to avoid a signalling surge.

Rogers had no overload protection at all. Optus had overload protection set to somebody else’s number. Both ended up with a national outage. The lesson is not “configure limits” — it is that a protective control you have not deliberately tuned and tested is not a control, it is a default.

Scale and consequence

The outage began at 04:05 AEDT. Around 88 per cent of services were restored by early afternoon; full restoration ran into 9 November. Roughly nine hours for most, up to twelve for some.

More than 10 million customers and 400,000 businesses. Hospitals. Banks. EFTPOS terminals. Metro Trains Melbourne cancelled hundreds of trains in the 05:00 to 06:00 window. Wholesale customers of Optus — Dodo, amaysim, Aussie Broadband, Moose Mobile — went down with it.

On emergency calling there are two official figures and they measure different things, so use both carefully. The Australian Communications and Media Authority found that Optus failed to provide access to the emergency call service to 2,145 people, and failed to conduct 369 required welfare checks. The Senate committee referred to approximately 2,700 Triple Zero calls that did not reach emergency operators. Optus’s initial public figure was much lower and was subsequently revised upwards; do not cite it.

ACMA’s conclusion was blunt: “Our findings indicate that Optus failed in the management of its network in a number of areas and that the outage should have been preventable.” Its chair, Nerida O’Loughlin, framed the stakes: “Triple Zero availability is the most fundamental service telcos must provide to the public.”

The regulator imposed a A$12 million penalty, paid in November 2024. Customers received extra data — 200GB for post-paid, unlimited weekend data for pre-paid. There was no cash compensation scheme.

The CEO, Kelly Bayer Rosmarin, resigned on 20 November 2023, twelve days after the outage.

“Manifestly inadequate”

The Senate Environment and Communications References Committee reserved its sharpest language for the communications. Optus’s public communications were “manifestly inadequate” — it relied on social media between 06:30 and 10:40 while the affected customers, by definition, had no internet connection.

The Bean Review commissioned by the Australian government made eighteen recommendations, all accepted: mandatory carrier communication protocols during outages, a comprehensive Triple Zero testing regime across networks and devices, a Triple Zero Custodian with end-to-end oversight, and a review of the underlying legislation.

And then the uncomfortable postscript. Optus suffered a further Triple Zero outage in September 2025, and ACMA commenced fresh Federal Court proceedings over it in 2026. Remediation programmes are announced in the weeks after an incident, when attention is total. They are delivered over years, when it is not.

What ITIL 5 has to say, and what it does not

Incident management is defined as “the practice of minimizing the negative impact of incidents by restoring normal service operation as quickly as possible.” In ITIL 5 it sits in the Product and Service Management group, and in the certification scheme it is bundled into the Monitor, Support and Fulfil module alongside service desk, problem management, service request management, and monitoring and event management.

The framework will tell you to have a major incident procedure, to define roles, to communicate with stakeholders. All correct, all necessary, and none of it would have changed either outcome, because neither organisation lacked a procedure.

What both lacked was a rehearsed answer to a narrower question: what do we do when the failure has taken away the means of responding to it?

That question does not belong to incident management alone. It sits across incident management, service continuity management, supplier management and the Operate stage of the new lifecycle. Which is precisely why it falls between them.

Nine things to check before your next major incident

1. Out-of-band communications, tested. A bridge, a messaging channel and a contact list that do not traverse your own infrastructure — or your own product, if you are a service provider. Test it by disabling the primary, not by reading the plan.

2. Out-of-band management access. Rogers could not reach the management plane of the failed routers. If your management network rides on your production network, you have a recovery capability that fails at exactly the moment you need it.

3. Logs somewhere other than the failed thing. Fourteen hours to root cause, because the error logs were on the equipment that was down. Ship them off-box, always.

4. Deliberate overload protection, with your numbers. Every protective threshold in your estate should be a value somebody chose, documented, for this topology at this scale. Audit for vendor defaults. Optus is the entire argument.

5. Risk-scoring that cannot downgrade a critical change. A Low-risk score on Phase 6 removed lab testing and senior approval from a change to a national core network. Whatever your scoring algorithm, there should be a floor below which changes to critical infrastructure cannot fall, regardless of what the calculation says.

6. A customer communication path that works when your service does not. Status page on separate infrastructure and a separate domain. SMS if the outage is data-only. Traditional media if it is total. Optus’s failure here was not slowness — it was choosing a channel its customers could not reach.

7. A notification clock, and an owner. Rogers took three hours fifty-six minutes to notify 9-1-1 providers and four hours eleven minutes to notify the public. Set explicit targets — thirty minutes for regulators and emergency services, sixty for customers — and name the person who owns the clock. For Indian regulated entities this is not optional: CERT-In’s directions of April 2022 require specified cyber incidents to be reported within six hours of noticing them, and RBI’s Master Direction on IT Governance sets incident reporting obligations for regulated financial entities.

8. Emergency and life-safety paths tested independently. Both regulators focused hardest on emergency calling. Whatever the equivalent is in your sector — payments, dispensing, dispatch, alarms — test it as its own service, not as a by-product of the platform being up.

9. A remediation programme with a delivery mechanism, not just an announcement. Rogers’ core separation was still described by the regulator as “a work in progress” two years later. Optus had a second emergency-calling outage. The commitment is the easy part.

The common factor

Neither of these organisations was incompetent. Both are large, technically sophisticated, heavily regulated national carriers with mature processes and expensive tooling.

What they shared was a set of assumptions that had never been tested against their own failure: that the network would be available to fix the network, that the phones would work to coordinate the phones, that customers could be reached through a channel that required the service that was down.

Assumptions like those are invisible in a normal audit, because on any ordinary day they are true.

The only reliable way to find them is to ask, of every recovery capability you have: what does this depend on, and is that thing inside the failure?

Ask it about your incident bridge tomorrow.

A newer failure mode: congestive collapse during recovery

One more case, because it illustrates something the two telecom outages do not, and it is the failure mode most incident plans underestimate.

On 20 October 2025, an AWS outage began in the us-east-1 region. The mechanism was a latent race condition: DynamoDB’s DNS management uses a planner that generates plans and multiple redundant enactors that apply them. An unusually delayed enactor applied an older plan to the regional endpoint at the same moment another enactor’s cleanup process deleted that same plan as stale. The result was an empty DNS record for the DynamoDB regional endpoint — every IP address removed. In AWS’s words: “This situation ultimately required manual operator intervention to correct,” because the automation could not repair itself once the state it needed was gone.

Then came the part worth studying. EC2’s droplet workflow manager maintains leases on physical hosts and depends on DynamoDB for state checks. With DynamoDB unreachable, leases timed out across the fleet. When DynamoDB recovered, over 2.25 million droplets simultaneously attempted to re-establish leases. AWS’s summary is unusually candid: the manager “had entered a state of congestive collapse and was unable to make forward progress in recovering droplet leases.” Selective host restarts were required.

A third layer compounded it. Delayed network state propagation caused load balancer health checks to fail on newly launched instances “even though the underlying NLB node and backend targets were healthy.” Alternating pass and fail results drove automatic availability-zone failover, which removed healthy capacity and made things worse.

DynamoDB itself was recovered in two hours fifty-two minutes. Cascading impacts continued for roughly fifteen hours, across more than twenty-two AWS services — including the identity and token services, which meant the blast radius reached workloads that did not use DynamoDB at all.

Three transferable lessons.

Recovery is a load, and you have not capacity-planned it. Everything reconnecting at once is a bigger event than the failure. Ask, for your own critical systems: if every client reconnected simultaneously, would the system come up? Meta sequenced its 2021 restoration for precisely this reason.

Health checks can amplify. A check that fails during a recovery, and drives automated capacity removal, turns your resilience mechanism into an accelerant. AWS’s remediation included a velocity control on capacity removal triggered by health-check failures. That is a good idea for anybody.

Shared dependencies define your real blast radius. Authentication was the multiplier here. In your estate the equivalent is probably identity, DNS, certificate services or a message bus. Map what everything depends on, then assume it is unavailable and see what your incident plan still permits you to do.


Sources

  • Canadian Radio-television and Telecommunications Commission / Xona Partners, Independent assessment of the Rogers Communications network outage of 8 July 2022 (executive summary published 2024)
  • Tony Staffieri, testimony to the House of Commons Standing Committee on Industry and Technology (INDU), 25 July 2022
  • Australian Communications and Media Authority, investigation findings and penalty in relation to the Optus outage of 8 November 2023
  • Senate Environment and Communications References Committee, Optus network outage (report, September 2024)
  • Department of Infrastructure, Transport, Regional Development, Communications and the Arts, Review into the Optus outage of 8 November 2023 (the Bean Review), April 2024
  • Optus submission to the Senate inquiry (vendor default prefix limits)
  • ITSM.tools, ITIL (Version 5) management practices (incident management definition and grouping) — itsm.tools
  • CERT-In, Directions under section 70B(6) of the Information Technology Act, 2000, 28 April 2022
  • Reserve Bank of India, Master Direction on Information Technology Governance, Risk, Controls and Assurance Practices, RBI/2023-24/107, 7 November 2023