What failover means
NIST defines failover as the capability to switch over automatically, typically without human intervention or warning, to a redundant or standby system when the previously active system fails or terminates abnormally. The same idea applies at every layer: a second server, a second data centre, a second internet link, or a phone moving from Wi-Fi to mobile data.
Failover is the mechanism that makes redundancy useful. A spare that nothing switches to does not remove a single point of failure.
The three parts of a failover
- Detect. Something has to notice that the active component is unhealthy. Amazon Route 53, for example, offers health checks that can monitor a resource such as a web server, the status of other health checks, or a CloudWatch alarm, and it can use them for DNS failover. Detection is a trade-off: react too slowly and users see errors; react too quickly and a brief blip triggers an unnecessary switch.
- Switch. Traffic or work moves to the standby. That might mean DNS answers pointing at a different address, a cluster electing a new leader, or a client opening its next connection on another network.
- Recover. The standby takes the load, and later the original is repaired. Moving back, called failback, is a separate change.
How ready is the standby?
NIST’s contingency planning guide describes alternate sites by how ready they are to take over:
| Standby | What NIST describes | Readiness and cost |
|---|---|---|
| Cold site | Space and infrastructure (electric power, telecommunications connections, environmental controls) to support recovery | Least expensive to maintain, but may need substantial time to acquire and install equipment |
| Warm site | Partially equipped, with some or all of the hardware, software, telecommunications and power | In the middle of the spectrum |
| Hot site | Sized and configured with the necessary hardware, supporting infrastructure and support personnel | Equipment and staff already in place |
| Mirrored site | Fully redundant, with automated real-time mirroring, identical to the primary in all technical respects | The most expensive choice in NIST’s examples, but it ensures virtually 100 percent availability |
The guide adds that a fixed alternate site should be in an area unlikely to be affected by the same hazard as the primary site, which is the same shared-fate problem that creates single points of failure.
Which one is right depends on how long the work can be down. NIST frames this with two targets: the maximum tolerable downtime (MTD), the total outage time the owner is willing to accept for a business process, and the recovery time objective (RTO), the longest a system resource can stay unavailable before it affects other resources, processes and the MTD.
Failover on a phone
Phones fail over between networks too. On Android, the system keeps a default network for each app, and it can change at any time; Google’s documentation gives the example of the device coming into range of a known, unmetered Wi-Fi network. When a new network becomes the default, new connections use it, and remaining connections on the previous default network are later forcefully terminated. An app that needs to react registers a default network callback and is told when the network is lost or replaced.
In other words, the operating system switches the path, but the app owns the consequences: requests in flight on the old network can fail, and the app has to retry them safely. Protocols that keep a connection alive across that switch are covered in What is multipath networking?.
Why failover fails
The switch depends on the failure. During AWS’s December 7, 2021 event in us-east-1, network congestion stopped the Service Health Dashboard tooling from failing over to its standby region as designed. The tooling meant to switch regions was itself held up by the failure it needed to escape.
The standby needs something that is broken. AWS’s static stability article argues for designs that keep working without having to make changes during an impairment. Its example is overprovisioning capacity across Availability Zones, so the service keeps operating without launching any new instances even if one zone is impaired.
It was never exercised. A failover path that only runs during real incidents has not really been tested. NIST recommends that, for high-impact systems, a full-scale functional exercise include a system failover to the alternate location.
Both sides think they are primary. If the old primary is only cut off rather than stopped, it may keep accepting work while the standby does the same. This is what a network partition does to any system with a primary: each side can keep working without knowing about the other.
Failover between paths, not only servers
The same idea works for network paths. A device with several ways to reach a peer, such as Wi-Fi, mobile data and a direct radio link, can switch to another path when one degrades. A self-healing network does this continuously by routing around failed links. The test is the same as for servers: turn the primary path off and confirm the remaining path really carries the work.