Connectivity resilience

What is a single point of failure?

A single point of failure is any one component whose failure stops the whole system, because nothing else can take over its job. Some are easy to see, like a single router or internet link. Others are hidden, like two "redundant" cables in the same trench, a configuration file every server loads, or a status page hosted on the service it reports on.

Learning objectives

After reading this article you will be able to:

  • Explain why one component with no backup can stop a whole system
  • Identify hidden single points of failure such as shared paths, files and recovery tools
  • Describe how redundancy, smaller blast radius and static stability contain single points of failure

The idea

A system has a single point of failure when there is one component that everything depends on and no second component ready to do its job. When it fails, the whole system fails with it, no matter how healthy everything else is.

The term is usually applied to hardware: one router, one power supply, one internet connection into a building. But the same idea applies to anything that sits on the critical path, including software, data, a provider, or a tool that operators need during an incident.

Where they hide

NIST’s contingency planning guide (SP 800-34) asks planners to identify single points of failure that affect critical systems, and the examples it gives show how easy they are to miss.

  • Shared physical paths. If redundant links are used, NIST says they should be physically separate and not follow the same path; otherwise a single incident, such as a cable cut, could disrupt both.
  • Shared facilities. Where two network service providers are used for near-continuous connectivity, the guide says they should not share common facilities at any point, including building entries and demarcation points.
  • Shared software and data. On November 18, 2025, a feature file that Cloudflare’s traffic-routing software loads on every machine doubled in size after a database permissions change. The software’s size limit was below the new size, and because the file had been propagated to all the machines in the network, the failure was network-wide. One file had become a single point of failure for the whole fleet.
  • Shared dependencies. In Meta’s October 4, 2021 outage, its DNS servers were built to withdraw their route advertisements when they could not reach Meta’s data centres. When the backbone went down, the DNS servers removed themselves from the internet even though they were still running.
  • The tools you need to recover. During the Amazon S3 disruption in us-east-1 on February 28, 2017, AWS could not update individual service statuses on its Service Health Dashboard, because the dashboard’s administration console depended on S3. AWS afterwards changed that console to run across multiple regions.

The pattern is the same each time: two things that look independent share something underneath.

How to find them

Start from a workflow that matters, not from a diagram of servers. Walk through every step and ask what it needs to succeed: a network path, a DNS name, a credential service, a database, a configuration push, a person with access. Then ask, for each one, what happens if it is gone for an hour.

A few questions catch the hidden cases:

  • Do the “two” paths share a trench, a building entry, a provider or a power feed?
  • Does every server load the same file, flag or rule set, pushed at the same time?
  • Does the failover mechanism, the monitoring, or the status page depend on the thing that failed?
  • Does the app need a backend round trip for an action that only involves people in the same room?

How to remove or contain them

Add redundancy that does not share fate. Duplicate the component and make sure the copies do not depend on the same thing underneath. NIST describes high-availability designs that use duplicate hardware and failover software to eliminate any single point of failure, and notes they are expensive, so it suggests reserving them for systems that cannot tolerate downtime. Switching to the spare is failover; using several paths at once is multipath networking.

Limit the blast radius. In its report on the 2017 S3 disruption, AWS describes breaking services into small partitions it calls cells, to reduce blast radius and improve recovery. The opposite shows up in the Cloudflare report: one file, pushed to every machine, caused failures across the whole network.

Design for an impaired dependency. AWS uses the term static stability for a system that keeps doing what it was already doing when a dependency is impaired, even if it stops receiving updates. For an app, the equivalent is keeping the data and logic a user needs on the device, so a backend outage pauses sync instead of stopping work. Local-first vs cloud-first compares the two approaches.

Remove the dependency where it is not needed. Some workflows only involve nearby people: a handover between two shifts, a checklist shared by a team on site. If the app routes every one of those actions through a server, the server and the network path to it become single points of failure for something that never needed them. A direct device-to-device path, as in a peer-to-peer network, takes them off the critical path for that workflow.

Frequently asked questions

Is redundancy enough to remove a single point of failure?

Only if the copies do not share a weakness. NIST's guide warns that two redundant links following the same path can both be cut by one incident, and that two network providers should not share facilities, including building entries.

Is the cloud a single point of failure for a mobile app?

For any feature that needs a round trip to the backend, the backend and the network path to it are on the critical path. Features that only need nearby devices to agree can be designed so they keep working when the backend is unreachable.

Build it with Offline Protocol

This page sets out which parts of a workflow the Offline Protocol SDK handles between devices and which stay with your backend, including finishing an authorised local handoff before the backend is reachable and submitting it later.

Read What the SDK handles