Why there is no single number
Downtime cost is the money an organisation loses, or spends, because a system it depends on is unavailable. It depends on what stops, for how long, when, and whether the work comes back later. An hour without a payments system on a busy afternoon is a different event from an hour without a reporting dashboard at night.
So any industry figure says more about the organisations that answered a survey than about your system. Published figures are useful as a check on the scale of the problem, as long as you know what they measured.
What one published survey measured
Uptime Institute, which researches data centre resilience, publishes an annual outage analysis. In its Annual Outage Analysis 2026, announced on 13 May 2026, it reported that 57% of respondents to its 2025 annual survey said their most recent major outage cost more than $100,000, and that for the second year running one in five reported costs above $1 million.
Read it carefully before you use it:
- What it measures. The self-reported cost of each respondent’s most recent major outage, counted against thresholds such as $100,000 and $1 million. It is not a cost per hour or per minute, and not an average.
- Who answered. Uptime describes its annual survey as covering the practices and experiences of data centre owners and operators. The outages are failures of data centres and IT services, not a field team losing mobile coverage.
- What else it found. The same release says outages linked to fibre and connectivity issues are rising and are more likely to cause extended disruptions, and that third-party IT and data centre providers, including cloud, telecommunications and colocation companies, have accounted for about two-thirds of the publicly reported outages Uptime has tracked over nine years.
That last point matters to anyone building on cloud services: your service can go down without anything in your own stack failing.
Estimate your own cost
NIST’s contingency planning guide, SP 800-34, gives a method: the business impact analysis. It identifies the processes a system supports, the impact of a disruption on each, and how much downtime each can tolerate. It suggests expressing impacts in units that mean something to the organisation, for example a “Costs” category measured in staffing, overtime or fee-related costs.
For each process the system supports, list what an outage would cost:
| Component | What to count | Where to find it |
|---|---|---|
| Lost transactions | Sales, bookings or orders that do not happen, minus those that come back later | Transaction logs by hour and day |
| Idle or diverted staff | People who cannot work, or who work on paper and re-enter it later | Rosters and pay rates |
| Recovery work | Reconciliation, data re-entry, overtime, support calls, incident response | Records of past incidents |
| Penalties and commitments | Service credits, contractual penalties, regulatory consequences | Contracts and service level agreements |
| Harder to price | Reputation, lost customers, safety | Management judgement, stated as a range |
Then work through a plausible outage, not an average one. NIST notes that the longer a disruption continues, the more costly it can become, so estimate a few durations rather than one.
Set tolerances, then balance them against recovery cost
SP 800-34 turns the impact analysis into three measures:
- Maximum tolerable downtime (MTD): the total time the system owner will accept a process being disrupted.
- Recovery time objective (RTO): how long a system resource can stay unavailable before the impact on the process becomes unacceptable. It normally has to be shorter than the MTD.
- Recovery point objective (RPO): how far back the data can be restored to, which sets how much data loss the process can tolerate.
The guide then balances the two sides. Disruption costs rise with time, while recovery options that bring a system back faster cost more to build. The point where the curves meet is different for every organisation, so it has to be worked out from your own numbers.
Where working offline changes the sum
A network outage is not always the same as a system outage. If devices can keep working while the network is down and sync afterwards, as apps designed to work without internet do, some lost transactions become delayed ones, staff keep working, and recovery work becomes reconciliation of records that changed apart, which needs sync conflict rules.
That moves cost rather than removing it. You pay for building and testing offline behaviour and for any services involved, so count both sides. Check what each option meters: Offline Protocol, for example, never meters local peer-to-peer traffic on any plan, and keeps your backend as the system of record that accepts records once it is reachable. The cost of adding offline messaging is the figure to set against the outage cost you estimated above.