Connectivity resilience

What is failover?

Failover is switching over to a redundant or standby system when the active one fails, usually automatically. It only works if three things hold at the moment of failure: the standby can carry the load, something detects the failure reliably, and the switch itself does not depend on the component that just failed.

Learning objectives

After reading this article you will be able to:

  • Describe the detect, switch and recover steps of a failover
  • Compare cold, warm, hot and mirrored standby sites on readiness and cost
  • Explain why failover fails, including switches that depend on the failure itself

What failover means

NIST defines failover as the capability to switch over automatically, typically without human intervention or warning, to a redundant or standby system when the previously active system fails or terminates abnormally. The same idea applies at every layer: a second server, a second data centre, a second internet link, or a phone moving from Wi-Fi to mobile data.

Failover is the mechanism that makes redundancy useful. A spare that nothing switches to does not remove a single point of failure.

The three parts of a failover

  1. Detect. Something has to notice that the active component is unhealthy. Amazon Route 53, for example, offers health checks that can monitor a resource such as a web server, the status of other health checks, or a CloudWatch alarm, and it can use them for DNS failover. Detection is a trade-off: react too slowly and users see errors; react too quickly and a brief blip triggers an unnecessary switch.
  2. Switch. Traffic or work moves to the standby. That might mean DNS answers pointing at a different address, a cluster electing a new leader, or a client opening its next connection on another network.
  3. Recover. The standby takes the load, and later the original is repaired. Moving back, called failback, is a separate change.

How ready is the standby?

NIST’s contingency planning guide describes alternate sites by how ready they are to take over:

StandbyWhat NIST describesReadiness and cost
Cold siteSpace and infrastructure (electric power, telecommunications connections, environmental controls) to support recoveryLeast expensive to maintain, but may need substantial time to acquire and install equipment
Warm sitePartially equipped, with some or all of the hardware, software, telecommunications and powerIn the middle of the spectrum
Hot siteSized and configured with the necessary hardware, supporting infrastructure and support personnelEquipment and staff already in place
Mirrored siteFully redundant, with automated real-time mirroring, identical to the primary in all technical respectsThe most expensive choice in NIST’s examples, but it ensures virtually 100 percent availability

The guide adds that a fixed alternate site should be in an area unlikely to be affected by the same hazard as the primary site, which is the same shared-fate problem that creates single points of failure.

Which one is right depends on how long the work can be down. NIST frames this with two targets: the maximum tolerable downtime (MTD), the total outage time the owner is willing to accept for a business process, and the recovery time objective (RTO), the longest a system resource can stay unavailable before it affects other resources, processes and the MTD.

Failover on a phone

Phones fail over between networks too. On Android, the system keeps a default network for each app, and it can change at any time; Google’s documentation gives the example of the device coming into range of a known, unmetered Wi-Fi network. When a new network becomes the default, new connections use it, and remaining connections on the previous default network are later forcefully terminated. An app that needs to react registers a default network callback and is told when the network is lost or replaced.

In other words, the operating system switches the path, but the app owns the consequences: requests in flight on the old network can fail, and the app has to retry them safely. Protocols that keep a connection alive across that switch are covered in What is multipath networking?.

Why failover fails

The switch depends on the failure. During AWS’s December 7, 2021 event in us-east-1, network congestion stopped the Service Health Dashboard tooling from failing over to its standby region as designed. The tooling meant to switch regions was itself held up by the failure it needed to escape.

The standby needs something that is broken. AWS’s static stability article argues for designs that keep working without having to make changes during an impairment. Its example is overprovisioning capacity across Availability Zones, so the service keeps operating without launching any new instances even if one zone is impaired.

It was never exercised. A failover path that only runs during real incidents has not really been tested. NIST recommends that, for high-impact systems, a full-scale functional exercise include a system failover to the alternate location.

Both sides think they are primary. If the old primary is only cut off rather than stopped, it may keep accepting work while the standby does the same. This is what a network partition does to any system with a primary: each side can keep working without knowing about the other.

Failover between paths, not only servers

The same idea works for network paths. A device with several ways to reach a peer, such as Wi-Fi, mobile data and a direct radio link, can switch to another path when one degrades. A self-healing network does this continuously by routing around failed links. The test is the same as for servers: turn the primary path off and confirm the remaining path really carries the work.

Frequently asked questions

What is the difference between failover and failback?

Failover moves work to the standby when the primary fails. Failback moves it back once the primary is healthy again. Failback is a change of its own and can fail too, so it is worth doing deliberately rather than the moment the primary looks healthy.

Is failover the same as load balancing?

No. Load balancing spreads work across several active components all the time. Failover keeps a standby ready and moves work to it when the active one fails. The two can be combined, so losing one balanced component shifts its share to the others.

Build it with Offline Protocol

The transport and routing docs describe how the Offline Protocol SDK picks among configured paths, where DORS scores eligible paths and applies switching controls, and how to test that a remaining path really carries the workflow.

Read Transport and routing