When the pandemic reached Switzerland in the spring of 2020, few organisations had a pandemic plan. Plans existed where regulators or group policies demanded them: at banks and insurers, say, or at operators of critical infrastructure that had extended their BCM binders after the swine flu of 2009. These plans were not fig-leaf documents, and yet almost all of them described the same wrong situation: an influenza wave, meaning a few weeks of high staff absence and then a return to normal operations. The pandemic that actually arrived did not follow the scenario. Most of the workforce was healthy but not allowed into the building. The backup site, BCM’s classic answer to the loss of a building, had no answer to this: the problem was not the site, it was the assumption that work means gathering people in a building. And that assumption applied at the backup site just as much. On top of that, an end date from which normal operations could have resumed could not even be estimated for months.

What carried organisations through those weeks, the ones with a plan and the ones without, was something no binder contained: the ability to rebuild themselves quickly. That rebuilding looked entirely different from sector to sector. Service companies switched to working from home within days. Manufacturers, for whom that was not an option, separated shifts and developed protection schemes just to be allowed to keep working. Restaurants improvised takeaway and delivery, and where even that was not enough because the business model itself had been shut down by decree, adaptive capacity still decided what state a business would be in when the forced pause ended. Honesty requires one addition here: in Switzerland, the state absorbed a large share of the damage, fast, with a lot of money and with astonishingly little bureaucracy, from expanded short-time work compensation to Covid loans issued on a half-page form. This diminishes the achievement of the businesses less than it seems, because that too was stretching, this time by an administration extending its own procedures far beyond their normal mode.

Business continuity management and resilience engineering increasingly appear in strategy papers and tender documents as interchangeable terms. Yet they are not. Business continuity answers the question: how fast are we back, if what we could imagine occurs? Resilience engineering deals with the other question: what do we do when something occurs that we could not imagine? One discipline plans the return. The other builds the capacity to stretch.

What Business Continuity Can Do

First the positive part, because this confusion is no reason to beat up on BCM. Business continuity management is a mature craft with a standard of its own, ISO 22301. A business impact analysis clarifies which processes may stand still for how long before the damage becomes intolerable. From it follow recovery targets: the RTO sets the time within which a process must run again, the RPO how much data may be lost at most. Add scenarios, plans and exercises.

The real value of this work is that decisions are made before the crisis instead of during it. Which process comes back first and who decides that is nothing anyone wants to think through for the first time at three in the morning. An organisation that stumbles into an outage without this groundwork loses hours on questions a calm meeting could have settled. The pandemic plan of 2019 was no mistake in this sense either. For an influenza wave it would have served.

The problem only begins when this work has to stand in as evidence of something it is not.

Four Meanings of One Word

David Woods, one of the founders of resilience engineering, sorted out in a short 2015 paper what people mean when they say “resilience”. He found four different concepts under the same word. The first is rebound: a system springs back to its initial state after a disturbance. The second is robustness: a system withstands defined loads without losing its function. The third he calls graceful extensibility, the ability of a system to stretch when a surprise pushes against the boundaries of what was foreseen. The fourth is sustained adaptability, the ability to preserve that adaptive capacity across time and changing conditions.

With this sorting, the confusion becomes precisely nameable. Business continuity institutionalises the first two meanings: it defines which loads the system must withstand (robustness) and plans the way back to the known state (rebound). Resilience engineering means the last two. It asks what a system does when the disruption leaves the scenario, and what an organisation needs so that this ability to stretch does not vanish with the next reorganisation. Whoever says “resilient” and shows a BCM certificate is not merely being imprecise. They are answering a different question than the one asked.

The pandemic was the teaching case for the difference precisely because it was not an event in the sense of the plans. A data-centre fire has a time, an extent and an end; afterwards you restore. The pandemic had none of these. It shifted the boundary conditions over months and kept shifting them, and there was no known state to return to; part of the shift, such as working from home, simply stayed. The rebound logic reached into a void because no point existed to spring back to. What carried organisations was the third meaning on Woods’ list: the ability to stretch. Woods also has a name for its opposite: brittleness, the way systems that cannot yield at the edge of their design envelope break there instead.

What Resilience Engineering Builds Instead

If resilience does not live in the binder, where does it live? Erik Hollnagel has described the answer as four potentials an organisation can develop. The potential to respond, including to what no plan exists for. The potential to monitor, meaning to see weak signals before they become situations; the organisations that checked their supply chains and their working-from-home capacity in February 2020 were doing exactly that. The potential to learn, and to learn from what went well too: whoever worked out after the first wave why the improvised switch had worked could build on it in the second. And the potential to anticipate, which does not mean prediction, it means the readiness to expect surprise.

None of these potentials can be filed as a document. They live in reserves that a tight cost calculation would flag as inefficiency, in people who know their processes well enough to be able to deviate from them, and in leadership that does not treat improvisation as loss of control. That makes them hard to exhibit. A certificate can be framed; adaptive capacity only shows when it is needed.

A certificate can be framed. Adaptive capacity only shows when it is needed.

Where the Confusion Does Damage

All of this would be harmless if it were only about word choice. It is about more, at three points.

The first is evidence. Towards the board, the regulator or customers, the BCM certificate is submitted as proof of “resilience”, and the recipient usually understands more by that than rebound and robustness. Even the standard invites the conclusion: ISO 22301 carries “Security and resilience” in its title. The fallacy has the same structure as one I wrote about here last week (“What a Good Audit Can Do”): a passed audit demonstrates that an organisation can be audited, not that it is safe. A BCM certificate demonstrates the plannability of the planned cases, not adaptability in the unplanned ones.

The second point is the budget. Where resilience counts as done because the BCM is in place, adaptive capacity no longer even competes for funds. Reserves and people who master more than one role appear in no binder as an obligation and therefore lose against every efficiency round. The irony: the more tightly an organisation optimises for its defined scenarios, the more brittle it becomes at their edges. You can buy robustness and, with the same calculation, abolish the ability to stretch.

The third point is the exercises. A BCM exercise that succeeds every year according to script tests the script, and that too has its value. But the result must not be read as a statement about resistance to surprise. The difference shows in a single question: when did an exercise last take the organisation to a place where the plan did not know what to do next? Where those places are systematically avoided, the organisation is practising how to pass the exercise.

Both, But as What They Are

It takes both, kept apart. Business continuity is the hygiene layer, analogous to compliance in its relation to actual safety: necessary, valuable, and systematically not in charge of surprise. Resilience engineering is not a better version of it, it is work on the other question. An organisation that keeps the two apart can maintain its BCM binder and know at the same time that it says nothing about its ability to stretch.

The pandemic plan of 2019 belonged in the binder, and the binder has its place. But what carried organisations through the spring of 2020 was not in the binder. It lived in the people, in the reserves and in the permission to deviate from the plan. Whoever wants to give an account of their organisation’s resilience today should therefore not start by showing what they can restore. They should show where they last stretched.

Sources

  • David D. Woods – Four Concepts for Resilience and the Implications for the Future of Resilience Engineering, Reliability Engineering & System Safety 141, 2015 (main source)
  • Erik Hollnagel – Safety-II in Practice: Developing the Resilience Potentials, Routledge 2018
  • ISO 22301:2019 – Security and resilience – Business continuity management systems – Requirements
Discussion

Thoughts on this? Take it to LinkedIn.

No comments here – but discussion is welcome, where the readership is anyway.


Join the discussion on LinkedIn