This site is part of the Informa Connect Division of Informa PLC

This site is operated by a business or businesses owned by Informa PLC and all copyright resides with them. Informa PLC's registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. Number 3099067.

Risk Management
search
CRO

The test you are meant to fail: A practical study of third-party risk management

Posted by on 04 September 2026
Share this article

Continuity exercises in most institutions are built to be passed. Where the grid, the fuel supply and the fibre fail in the same week, an exercise nobody fails is an exercise nobody ran.

Mutisunge Zulu, CRO at Zanaco Plc, shares practical ways to create holistic resilience tests that bring true value to the business:

  • Test for multiple failures, not single incidents. Real disruptions are interconnected. Power, fuel, network, supply chain, and staffing issues can fail simultaneously, making single-failure tests unrealistic.
  • Focus on shared dependencies. Redundancy only works if backup systems don't rely on the same underlying infrastructure, suppliers, networks, or people. Hidden common points of failure are often the biggest risk.
  • Validate third-party resilience, don't just assume it. Reviewing a supplier's continuity plan is not enough. Organisations should actively test critical vendors, escalation processes, and failover capabilities.
  • Run realistic, unannounced stress tests. Planned exercises often measure the plan rather than the organisation's true readiness. Testing should occur under challenging conditions, at inconvenient times, and with minimal notice.
  • Measure business impact, not just recovery times. Key metrics should include detection time, decision-making speed, customer service disruption, and transaction value affected. These metrics better demonstrate resilience outcomes and investment needs.
  • A good test should find weaknesses. If an exercise uncovers no issues, it was probably not challenging enough. The goal of testing is to identify and fix vulnerabilities before a real disruption does.

What unimpactful tests look like

Ask a bank how it tests for infrastructure failure and you will usually be shown a calendar. A date, a nominated system, a recovery site on standby, a report afterwards confirming the recovery time objective was met. Everybody passed. The exercise was designed so that everybody would. In markets where infrastructure is not a given, that is the one design flaw that matters.

Take the dependency map for a critical service and interrogate it not for what supports each component, but for what supports more than one.

A conventional test fails one thing at a time. Events do not observe the same discipline. The currency shortage that delays fuel imports is the same shortage that leaves imported transformer parts and network hardware sitting in a bonded warehouse. The rains that flood an access road also take down the tower serving the branch and keep half the staff at home. The grid, the generator and the refurbishment programme have a common parent, and single-failure testing is built on the assumption that they do not.

Reliability engineers call this common-cause failure, and their central finding transfers to banking without modification: redundancy buys only what the redundant paths do not share. Institutions test components. Events attack contexts.

Correlation lives in the supply chain

Testing for simultaneity therefore begins on paper, before anything is unplugged. Take the dependency map for a critical service and interrogate it not for what supports each component, but for what supports more than one. The candidates are always the same, and they are rarely in the risk register: the substation, the fibre ring, the last-mile provider, the cloud availability zone, the single fuel supplier on a single contract, the systems integrator who built both environments, the access road, the two or three individuals who actually know how the settlement interface behaves at month-end.

Third-party resilience is assumed far more often than it is evidenced.

The findings are uncomfortable and cheap to obtain. The primary and alternate sites draw from the same substation. The recovery site's connectivity terminates in the same exchange. Two independent payment channels traverse one aggregator two hops down. The generators at both locations are fuelled by one supplier whose own storage is a single tank. None of that requires an exercise to discover. It requires somebody to read the map with the specific intention of finding shared roots, which is a different exercise from confirming that each service has a documented alternate.

The discipline has to extend past the perimeter, where most programmes stop. A supplier's continuity plan is a document the bank has read, not a capability the bank has observed. Contracts should carry a right to test, and the right should be exercised: a joint failover with the card processor, an unannounced call to the fuel supplier's emergency line, a verified attempt to reach the cloud provider's escalation path at three in the morning rather than a certificate confirming that one exists. Third-party resilience is assumed far more often than it is evidenced.

Design backwards: Identify causes and consequences

The second correction is to reverse the direction of the test. The conventional question is what happens if the grid fails. The useful question is which combination of failures would breach the institution's impact tolerance for payments, and how far away from that combination the institution is on an ordinary Tuesday. The first question samples the space of failures. The second searches it. Only the second tells the board what it is actually exposed to, because it starts from the outcome the institution has already declared intolerable and works backwards to the conditions that produce it.

Compound scenarios should then be constructed around shared drivers rather than assembled at random. A test that fails the grid and simultaneously fails an unrelated core system is arbitrary, and the exercise team knows it. A test that fails the grid, rations fuel and delays a hardware replacement, all as consequences of the same foreign exchange position, is a scenario the treasurer can recognise and the board cannot dismiss. Election periods, flood seasons, settlement cut-offs and public holidays are cheap sources of realistic correlation. So is the maintenance window, which is the hour when the institution's own defenses are deliberately lowered.

Test unannounced, and at the worst hour

Notice destroys the test. An announced exercise measures the plan; an unannounced one measures the institution. Timing is part of the design and should be adversarial rather than convenient: month-end, payroll day, the settlement cut-off, the Friday before a long weekend, two o'clock in the morning. An institution that will only test at ten on a Saturday has learned what it can do at ten on a Saturday.

Degraded-mode operation is a skill, and skills that are documented but never practised are not held by the institution.

Then fail the substitute rather than the primary. A generator that runs for fifteen minutes every Monday under no load is not a tested generator; it is a reassured one. Run the branch on it for eight hours at full load, with the tank at the level it would actually hold on the day, and with the fuel supplier told nothing in advance. Where it is safe to do so, pull the plug rather than simulate the plug being pulled. Simulation tests the plan. Interruption tests the dependency that nobody thought to document.

Equipment is the easier half. The harder half is decision rights under degradation. Who declares an incident, and who can authorise unbudgeted spend at two in the morning without a committee? What happens when the mobile network is itself the failure and the crisis team's call tree runs on it? What happens when the incident commander is unreachable? Can a branch transact offline against agreed limits, and has anyone currently employed ever done it? Degraded-mode operation is a skill, and skills that are documented but never practised are not held by the institution. They are held by whoever remembers.

Measure what the board should fund

Measurement is where most exercises quietly fail. Recording that the recovery time objective was met tells the board almost nothing, because the objective was set against a scenario the institution chose.

The numbers worth carrying upstairs are:

  • how long until anyone noticed
  • how long from detection to decision
  • how many customer-minutes of service were lost
  • what value of transactions went unserved

Those figures translate resilience into the language of the income statement, which is the only language in which it competes successfully for funding.

The last metric is the most revealing, and almost nobody reports it: the finding rate. An exercise that produces no findings was designed to produce none. Boards should ask how many issues the last test surfaced, how many remain open a year later, and how many of those were known before the test and simply confirmed by it. A rising finding rate on a maturing programme is a sign of an institution testing harder, not one deteriorating, and directors who cannot tell the difference will reward the wrong behaviour. Better still, they should attend. A director who has watched a crisis room work at two in the morning asks a different question of next year's budget than one who has read the summary.

None of this is free, and the honest framing is that it is not meant to be. The cost of a serious testing programme is the cost of discovering weaknesses on a schedule of the institution's own choosing rather than the market's. The alternative is discovery by event, which arrives with a supervisory letter attached and no opportunity to reschedule. An institution that has never failed an exercise has not been tested. It has been rehearsed.


Read Mutisunge Zulu's take on business continuity planning.


Tackle the defining challenges facing today’s risk leaders with 1300+ of them by your side at RiskMinds.


Share this article

Sign up for Risk Management email updates

keyboard_arrow_down