Operational resilience

The tolerance clock for operational resilience

When a system fails, which important business services go outside impact tolerance, and at what hour?

The tolerance clock answers that question for each service, with the working shown. You see which services go outside tolerance, when, and why, before an incident shows you.

Scenario testing

Often the outage time in a scenario test is someone's estimate.

That estimate has to allow for the dependencies between systems, knock-on effects two or more steps away, and the backlog of work that grows while a service cannot run. Those are hard to hold in your head across a whole firm.

How it works

The clock calculates the outage from the dependencies.

You choose one element to fail, and you watch the disruption reach each service, hour by hour. For each important business service you see:

  • the hour the disruption arrives, and the route it took
  • how long the service is down, split into waiting for what it relies on, its own recovery, and clearing the backlog
  • that time against its impact tolerance
  • whether fixing the cause sooner keeps it within tolerance, or whether it stays outside even when the cause is fixed at once

Who it is for

The clock is for operational resilience leads who choose and justify scenario tests, for COOs, CROs and boards who approve the self-assessment, and for consultants who advise them.

The timings in the demonstration are estimates, graded against public outage reports where evidence exists.

To see it, ask us for a guest login.

Choose the Operational Resilience dataset, a simplified UK retail bank with 10 important business services. The guided tour shows the clock step by step.

Ask for a guest login