In most organisations, business continuity is a document. It is written by someone competent, approved by a committee, filed, and produced during audits. Its recovery time objectives were agreed in a workshop where nobody wanted to be the person arguing for a slower number. It has never been tested in a way that could fail.
In Indian banking, that arrangement stopped being viable on 1 April 2024, and the evidence that it had stopped being viable was already three years old by then.
What the regulator actually did
On 2 December 2020, the Reserve Bank of India issued an order to HDFC Bank — India’s largest private sector bank. As disclosed by the bank to the stock exchanges, the RBI directed it to stop:
“(i) all launches of the digital business generating activities planned under its programme – Digital 2.0 (to be launched) and other proposed business generating IT applications and (ii) sourcing of new credit card customers.”
The RBI also directed the bank’s board to examine the lapses and fix accountability.
The stated basis, again from the bank’s own filing, was “certain incidents of outages in the internet banking/ mobile banking/ payment utilities of the bank over the past two years, including the recent outages in the bank’s internet banking and payment system on November 21, 2020, due to a power failure in the primary data centre.”
Read the precipitating cause once more. A power failure in a primary data centre. Not a cyberattack. Not a novel exploit. A power failure — the single most anticipated event in the entire discipline of continuity planning — that the disaster recovery arrangements did not absorb.
The restriction on sourcing new credit card customers was relaxed on 17 August 2021, roughly eight and a half months later. The restriction on the Digital 2.0 programme was lifted in March 2022, about fifteen months after the order. In February 2021 the RBI had ordered a special third-party audit of the bank’s IT infrastructure.
Fifteen months during which the largest private bank in the country could not launch the digital products it had already built, because its data centre lost power and the failover did not hold.
Then it happened again, to somebody else
On 24 April 2024, using its powers under Section 35A of the Banking Regulation Act, the RBI directed Kotak Mahindra Bank to cease and desist, with immediate effect, from onboarding new customers through its online and mobile banking channels and from issuing fresh credit cards. Existing customers were to be serviced without disruption.
The RBI’s stated reasons are, to my mind, the single most important passage in Indian IT governance enforcement to date, because they read like a clause-by-clause audit finding:
“serious deficiencies and non-compliances were observed in the areas of IT inventory management, patch and change management, user access management, vendor risk management, data security and data leak prevention strategy, business continuity and disaster recovery rigour and drill, etc.”
Every one of those is an ITIL practice. IT inventory management. Change enablement. Access management within information security. Supplier management. Service continuity management — named explicitly, and named with the word drill.
The RBI went further. It said that across IT Examinations in 2022 and 2023 the bank had been found “deficient in its IT Risk and Information Security Governance”; that “for two consecutive years, the bank was assessed to be significantly non-compliant with the Corrective Action Plans issued by RBI”; and — the sentence I would put in front of any board considering a growth plan — that the bank was “materially deficient in building necessary operational resilience on account of its failure to build IT systems and controls commensurate with its growth.”
The restriction was lifted on 12 February 2025, about nine and a half months later, after corrective measures and an external audit commissioned with the RBI’s approval.
The enforcement instrument is growth, not money
This is the pattern that Indian executives need to understand, and it differs sharply from the Western model.
In the United Kingdom, the FCA and PRA fined TSB £48.65 million for its 2018 migration failure. In Australia, ACMA penalised Optus A$12 million over emergency call failures. In the United States, the DOT fined Southwest Airlines $140 million.
The RBI does not generally do this. It freezes the customer acquisition channel that the failing IT supports.
| Entity | Date | Trigger | Restriction | Lifted |
|---|---|---|---|---|
| HDFC Bank | 2 Dec 2020 | Repeated outages; data centre power failure 21 Nov 2020 | Digital 2.0 launches; new credit card sourcing | Aug 2021 (cards); Mar 2022 (Digital 2.0) |
| Paytm Payments Bank | 11 Mar 2022 | Supervisory concerns; system audit ordered | New customer onboarding | Superseded by the Jan 2024 order |
| Bank of Baroda | 10 Oct 2023 | Onboarding controls in the ‘bob World’ app | New onboarding to that app | — |
| Kotak Mahindra Bank | 24 Apr 2024 | IT Examinations 2022 and 2023; non-compliance with corrective action plans; outage 15 Apr 2024 | Online and mobile onboarding; new credit cards | 12 Feb 2025 |
A fine is a cost. It is absorbed, disclosed and moved past. Growth denial is different: it lands directly on the strategy, it is visible to every competitor, and it lasts for as long as the regulator says it lasts.
The line for your board is short. IT service quality is now a direct constraint on the ability to grow the book.
The Master Direction, and the clause with teeth
The RBI’s Master Direction on Information Technology Governance, Risk, Controls and Assurance Practices was issued on 7 November 2023 and took effect on 1 April 2024. It applies to banking companies, small finance banks, payments banks, NBFCs in the top, upper and middle layers, credit information companies, and the all-India financial institutions.
Chapter V deals with business continuity and disaster recovery. Paragraph 28(b) sets the objective: capabilities “shall be designed to effectively support its resilience objectives and enable it to rapidly recover and securely resume its critical operations (including security controls) post cyber-attacks/ other incidents.”
Then paragraph 29 does something most continuity regimes never do. It specifies the test.
“Periodicity of DR drills for critical information systems shall be at least on a half-yearly basis.”
“The DR testing shall involve switching over to the DR / alternate site and thus using it as the primary site for sufficiently long period where usual business operations of at least a full working day (including Beginning of Day to End of Day operations) are covered.”
Read that twice. Not a failover test. Not a tabletop. Run the business on the DR site for a full working day, from beginning of day to end of day operations.
On recovery objectives, regulated entities “should prioritise achieving minimal RTO (as approved by the RE’s ITSC) and a near zero RPO for critical information systems,” with a documented reconciliation methodology wherever RPO is non-zero. Paragraph 29 also requires configuration consistency between the primary and DR sites, and joint resilience testing with vendors and interconnected systems.
That last requirement is the one most organisations will fail first. Your DR site may come up perfectly and still be unable to transact, because a payment switch, a KYC provider or a credit bureau has never tested against it.
I would suggest that paragraph 29 is, in substance, the regulator’s response to a power failure in a primary data centre in November 2020. If you want to know what a regulator does with a case study, this is what it does.
Where ITIL 5 fits, and where it does not
ITIL 5 defines service continuity management as “the practice of ensuring that service availability and performance are maintained at a sufficient level in case of a disaster.” It sits in the Product and Service Management practice group.
The framework gives you the structure — the practice, its relationships with availability management, information security management and supplier management, and its place in the new lifecycle’s Operate stage. It does not give you a drill periodicity, a duration, or an obligation. The RBI does.
That is the correct division of labour, and it is worth being clear about it when you buy training. ITIL will teach your people the discipline. It will not make you compliant with anything. If you need to demonstrate a management system, ISO 22301 for business continuity and ISO/IEC 27001 for information security are the certifiable instruments; the Master Direction is the obligation; ITIL is the operating vocabulary that makes all three coherent.
Concentration risk, which nobody’s DR plan covers
One more case, because it exposes the limit of everything above.
In February 2024, ransomware operators obtained compromised credentials and used them against a Citrix remote access portal at Change Healthcare that did not have multi-factor authentication enabled. UnitedHealth’s corporate policy required MFA on externally facing systems; the recently acquired subsidiary had not been brought into line two years after acquisition. Initial access on 12 February, lateral movement and exfiltration over nine days, ransomware deployed on 21 February.
Change Healthcare processes roughly a third of US medical claims. Its outage stopped claims processing, eligibility verification, prior authorisation and pharmacy dispensing nationwide. Small and rural providers ran out of operating cash within weeks. 190 million individuals were ultimately notified — approximately a third of the US population, and the largest healthcare breach ever reported.
UnitedHealth reported a $3,090 million total 2024 impact: $2,223 million of direct response costs and $867 million of business disruption. The CEO confirmed under oath that a $22 million ransom was paid. Senator Ron Wyden’s summary at the Senate Finance Committee was: “This hack could have been stopped with cybersecurity 101.”
Two lessons that no internal DR plan addresses.
Acquisition integration is a continuity risk. A subsidiary running outside your control standards is a hole in your resilience that your own drills will never find, because your drills test your estate.
And your continuity plan probably assumes a supplier that has no alternative. There was no viable alternative clearinghouse capacity in the United States. If your recovery plan depends on switching to another provider, name them, and check that they could actually take your volume — along with everyone else who has written the same sentence.
What to do in the next two quarters
1. Run one real drill. Switch to the alternate site and run a full working day, beginning of day to end of day, on live business operations. If your organisation is not regulated by the RBI, do it anyway — it is simply the only test that tells you the truth. Expect it to fail the first time. That is the point.
2. Test with your counterparties. Payment switches, bureaus, KYC providers, settlement systems, key SaaS vendors. Joint resilience testing is explicit in paragraph 29 and it is where most failover plans quietly break.
3. Check configuration drift between primary and DR. Every organisation’s DR environment has been diverging from production since the day it was built. Reconcile it and put the delta on a dashboard.
4. Reconcile your RPO honestly. If your RPO is not zero, you must know exactly which transactions are lost in a failover and how you reconstruct them. If nobody can describe that process on a whiteboard, your RPO is aspirational.
5. Include acquisitions in scope with a deadline. Every acquired entity should have a dated commitment for adoption of your control standards, reported to the audit committee until closed. Change Healthcare is what happens when that date is indefinite.
6. Put the drill result in front of the board. Not the plan — the result. Under the Master Direction, the Audit Committee of the Board is responsible for oversight of information systems audit, and critical IT and information security issues must be reviewed by it. Directors who see only a green continuity status and never a failed drill are not being given the information they need. As an independent director I would rather see one honest failed test than four years of unblemished summaries.
The last word
The UK Treasury Committee, having written to nine major banks and building societies, published in March 2025 that they had collectively suffered at least 803 hours of unplanned outage — over thirty-three days — across at least 158 separate incidents in two years and two months. The recurring causes were third-party supplier failure and disruption arising from system changes.
Outages are not rare events at the edge of the risk register. They are a continuous background rate that most institutions have never measured and most boards have never seen.
The RBI decided that in India this would have consequences attached to it, and chose the consequence carefully: not money, but growth. That is a regulator that understands exactly what it is dealing with.
If your continuity plan has never failed a test, it has never been tested.
What a director should ask
Since much of this lands in board papers, here is what I would ask as an independent director on the board of a regulated entity — and what a satisfactory answer looks like.
“When did we last switch to the DR site and run a full working day on it, and what broke?” A satisfactory answer names a date within the last six months and lists failures. An answer with no failures means the test was not real.
“Which counterparties have we tested failover with?” Payment switches, bureaus, settlement, key SaaS. If the answer is “none,” the plan has never been tested end to end regardless of how the internal drill went.
“What is our RPO, and if it is not zero, who has seen the reconciliation procedure?” The RBI expects near-zero RPO for critical systems and a documented reconciliation methodology where it is not. A number in a policy is not a methodology.
“How many outages did we have last year, for how long in total, and what did they cost us?” Most boards have never been given this as a single figure. The UK Treasury Committee had to write to nine banks to obtain it, and the answer — over 803 hours across 158 incidents — surprised the institutions themselves.
“Which of our acquisitions are not yet on our control standards, and by what date will they be?” With a named owner per entity, reported until closed.
“Has any corrective action plan from a supervisory examination been carried forward more than once?” The RBI’s order against Kotak Mahindra Bank turned on exactly this: two consecutive years of assessed non-compliance with corrective action plans. A repeated carry-forward is not an operational detail. It is the leading indicator of regulatory action, and it belongs on the audit committee agenda as such.
None of these questions require technical knowledge to ask. They require only the willingness to keep asking until the answer stops being a status colour.
Sources
- Reserve Bank of India, Master Direction on Information Technology Governance, Risk, Controls and Assurance Practices, RBI/2023-24/107, DoS.CO.CSITEG/SEC.7/31.01.015/2023-24, 7 November 2023 (effective 1 April 2024) — paragraphs 28, 29 and 30
- HDFC Bank Limited, disclosure to the stock exchanges under Regulation 30, SEBI (LODR) Regulations, 3 December 2020; and subsequent disclosures of 17 August 2021 and 12 March 2022
- Reserve Bank of India, press release on directions to Kotak Mahindra Bank under Section 35A, Banking Regulation Act 1949, 24 April 2024; and the lifting of restrictions, 12 February 2025
- Reserve Bank of India, directions to Bank of Baroda, 10 October 2023, and to Paytm Payments Bank, 11 March 2022 and 31 January 2024
- UnitedHealth Group Incorporated, FY2024 results (cyberattack impact); testimony of Andrew Witty to the US Senate Committee on Finance and the House Committee on Energy and Commerce, May 2024
- House of Commons Treasury Committee, disclosure on IT failures at major banks and building societies, 6 March 2025
- ITSM.tools, ITIL (Version 5) management practices (service continuity management definition) — itsm.tools