QgenticQgentic
Qgentic / guides / impact tolerances
Guide · Operational resilience (UK)

Impact tolerances that hold.

The two decisions that make or break an operational resilience framework are which services you call important, and how long you say you could tolerate losing each one. Get these right and the rest follows; get them wrong and everything downstream is built on sand.

FCA / PRA · Impact tolerance · ~6 min read

In short

An important business service is one a firm provides to an external end user where disruption could cause intolerable harm to consumers or pose a risk to market integrity. An impact tolerance is the maximum disruption to that service the firm could tolerate, expressed as a limit you can actually measure — usually a duration. Both start from harm: the question is who is hurt and how quickly. Setting a tolerance is what turns a judgement into something testable, because once it is set, whether an incident breached it is arithmetic on timestamps.

Start from the harm

The instinct is to list your critical applications and work outwards. Resist it. An important business service is defined by the harm its failure causes to customers or the market. A service is important if losing it could cause intolerable harm: customers locked out of their money, claims that can't be paid, trades that can't settle. Begin with the customer outcome and trace back to the systems.

The two failure modes are equal and opposite. Name your services too narrowly and genuine risk sits outside the framework entirely. Name them too broadly — every internal process an "important business service" — and you dilute the exercise until nothing is really prioritised. Aim for the handful of services whose failure a regulator, a journalist, or a customer would immediately recognise as serious.

Setting an impact tolerance you can defend

An impact tolerance is the maximum tolerable level of disruption to an important business service — most often a duration. The discipline is to set it from the point of view of harm, assuming disruption happens, then ask honestly whether you could stay inside it under a severe but plausible scenario.

  • Anchor it to external harm. "Four hours" should mean four hours is the point beyond which customer harm becomes unacceptable — not the point at which your recovery runbook happens to finish.
  • Make it testable. A tolerance you can't measure an incident against is decoration. A time, a transaction threshold, or a customer-impact count can all be checked after the fact.
  • Don't set it to what you can already do. The tolerance is the outer limit of acceptable harm; your current recovery capability is a separate question. If they're the same number, you've probably set the tolerance to flatter the capability.

Where frameworks quietly fail

The most common weakness isn't a missing service or a wrong number — it's an unmapped dependency. If an important business service leans on a third party you never traced, your impact tolerance is a guess. This is where UK operational resilience meets ICT third-party risk: the same providers that appear in a DORA register are the ones your tolerances silently depend on.

The determination is arithmetic

Once tolerances are set, judging whether an incident breached one should not be a matter of opinion. The disruption ran for a measurable time; it either exceeded the affected service's tolerance or it didn't. Assessing that from the incident timestamps — rather than a hand-set flag — is what makes the resilience record consistent and audit-ready.

How Qgentic keeps it honest

Qgentic OpRes takes your important business services, their impact tolerances, and your logged incidents, and computes the breach determination from detection-to-recovery timestamps every cycle — flagging any incident that ran past the tolerance of a service it touched. The numbers stay defensible because they are derived. See it run end to end in the live demo.

Qgentic OpRes holds your tolerances against real incident timestamps, every cycle.

See the OpRes platform Try the live demo