Nothing fails alone: How to challenge your standard business continuity assumptions

Risks in volatile markets do not arrive in sequence, and they do not stay in their categories. Continuity planning built for single points of failure is only insuring the smallest loss on the table.
Mutisunge Zulu, CRO at Zanaco Plc, explores the consequential nature of business disruptions and how to shift from static continuity planning to continuous resilience management.
- Continuity plans must become capabilities, not documents. Traditional annual plans are no longer enough in a world of constant, interconnected disruptions. Resilience needs to be embedded into everyday operations and decision-making.
- Risks don’t stay in one category. Operational, cyber, legal, liquidity, and reputational risks can quickly evolve into one another. Organisations need to understand these cascading effects, not just manage individual risks in isolation.
- Detection speed is as important as recovery speed. The real clock starts when an issue occurs, not when it is formally declared. Early detection can significantly reduce the impact of a disruption.
- Stress test for multiple failures at once. Modern disruptions are rarely isolated. Institutions should test compound scenarios (e.g., power outages, connectivity issues, fuel shortages, and third-party failures happening simultaneously) rather than single-event incidents.
- Customer impact should drive resilience priorities. Start with the services customers rely on most, such as payments, card transactions, cash access, and onboarding, then work backwards to strengthen the people, systems, vendors, and infrastructure that support them.
- Resilience is a business investment, not a compliance exercise. Spending on recovery capabilities, cybersecurity, and operational resilience protects future value and can prevent far greater costs such as regulatory penalties, customer loss, and reputational damage.
Continuity plans: Why do they need an update?
Somewhere this quarter, a bank will discover that its continuity plan is a document rather than a capability. The discovery will not come from a regulator's thematic review. It will come at two o'clock on a Tuesday, when the grid fails, the generator starts, the fuel contract proves indexed to a shortage nobody modelled, and the core platform runs while the branches do not.
Chaos and complexity are the new business as usual.
The shocks increasingly begin in no jurisdiction of the institution's own. For example:
- A rocket fired in a Middle Eastern crisis takes a data centre offline, and with it, systems serving customers on another continent who cannot see the conflict.
- Rockets launched in Russia's war strike not only Ukrainian targets but fuel tanks and pump prices in emerging markets that import every litre they burn.
- Tariff decisions in Washington or Brussels reshape fiscal posture in Lusaka and Accra within a quarter.
- Climate change has stopped being a disclosure exercise and now shows in balance sheets and in the light switches of hydro-dependent economies.
- Technology never respected borders: a ransomware crew thousands of miles away reaches a bank's customers without crossing one. Geography has stopped being a control.
The institution absorbs it locally: the fuel levy, generator diesel, hardware still in transit, correspondent pricing that widens whenever the world feels less certain. Chaos and complexity are the new business as usual.
Most continuity plans were built for a different world. They assume a single point of failure, a finite duration and an orderly recovery. They are drafted, approved, filed and exercised annually against a scenario chosen for its tractability. That is theatre: continuity management performed on a calendar rather than embedded in business as usual. A plan that lives outside the operating rhythm will not be there when that rhythm breaks.
Reading risk in orders: Examples in how to prioritise
The first correction is analytical: boards must read risk in orders.
- A transformer failure is a first-order operational event.
- The second order is a branch queue, a surge onto digital channels the capacity model did not anticipate, authorisation latency at merchant terminals.
- The third order is customer confidence: deposit movement, a trending hashtag, a call from the supervisor.
The first order is an engineering problem; the third is a franchise problem. Institutions that plan only to the first are insuring the smallest of the three losses.
The industry has spent a generation ranking risks. It must now understand how they metamorphose into one another. Litigation is a legal exposure until judgment, at which point it becomes market risk and reprices equity and bonds within a session. A cyber intrusion is operational until payments stop settling, at which point it is a liquidity event - and a bank that is solvent but cannot move money is illiquid in the only sense a depositor cares about. Vendor risk is fraud risk in waiting: the control environment an institution actually operates is the weakest in its outsourcing chain. Each discharges into reputation, the terminal account into which every other risk is posted. And because risks transform faster than they escalate, detection is the new recovery time objective.
The clock starts not when an incident is declared but when the institution notices, and every hour before that compounds.
Add a failure mode with no precedent in the literature. Everything is AI-amplified, and agentic systems do not advise but act - opening tickets, moving money, answering customers, deciding in the gaps where a human once sat. An agent that goes rogue does not produce one bad output; it produces a sequence at machine speed, discovered only after it completes. The legal position is less ambiguous than boards assume: what a chatbot tells a customer is what the institution has told the customer. Innovation binds the board. Most are still casual about that, and their agents are already in production.
Hence continuous scenario analysis and stress testing are needed rather than the annual variety. The question is not whether the institution is capitalised against a matrix of independent exposures, but against the correlated, compounding version history keeps delivering. Capital calibrated on an assumption of independence is capital measured. So is a liquidity buffer sized for a market disruption that does not begin with a systems outage.
The board's unglamorous questions
This is where fiduciary duty becomes concrete. Directors are not accountable for the engineering of the recovery site, but for having asked whether it works. The diagnostic questions are unglamorous.
- When did this board last sit through a live simulation rather than a paper walk through?
- Is the chief risk officer's voice amplified - is risk a standing item or a standing lens?
- Does anyone present know what the institution's agents may do without a human in the loop?
Too often the conversation is suppressed at executive level. Sometimes the cause is a chief executive formed in earnings, whose instinct treats resilience spend as margin leakage. Just as often, it sits with the risk function: a chief risk officer who has not translated resilience out of control language and into business language, and has never won front-line conviction. The first failure is a governance problem. The second belongs to the risk profession, and is the more fixable.
Start continuity planning where the customer is
Prioritisation begins where the customer is. Operational resilience is the heartbeat of any service offering - unnoticed while it holds, the only thing that matters the moment it stops. The benchmark sits outside banking: a search engine is simply always there, zero disruption, no explanation required. Consumers have imported that expectation into financial services. Identify the services whose failure the customer feels first - payments, card authorisation, cash access, on-boarding - then work backwards through the chain that delivers them: people, processes, applications data, interfaces, connectivity, premises, third parties and governance. Impact tolerances mean little until that array exists.
The institutions that survive the next cycle will not be those with the most elegant documentation.
Then the hardest test: simultaneity. Standard exercises fail one thing at a time. Infrastructure-constrained markets fail several at once - grid, fuel, connectivity and a washed-out access road in one week. Test the compound scenario, without notice, and the substitutes as rigorously as the primaries. Does the generator have a fuel contract that holds when fuel is rationed rather than merely expensive? Does the alternate site depend on the same fibre ring? A backup that shares a dependency with the thing it is backing up is not a backup. It is a duplicate.
The arithmetic of prevention
All of this costs money, and the industry should stop being defensive about it. The cost is real but not an accounting cost; it is an investment in the preservation of future value. A recovery site, cyber cover, machine-learning fraud detection and a serious compliance engine carry a visible, budgetable, depreciable price. Regulatory penalties, remediation orders, deposit flight and the withdrawal of a licence do not. Institutions already hold capital against credit losses they hope never to incur. Resilience spend is the same trade, in operating expenditure rather than regulatory capital.
The institutions that survive the next cycle will not be those with the most elegant documentation. They will be those for which continuity was a habit - tested when nothing was wrong, funded when nothing was burning, understood by people who would never call themselves risk professionals. Resilience architecture is the wiring; comprehending how risks metamorphose is the current that runs through it. Without the second, the first is only cable.

