Connectivity resilience

Why do network outages happen?

Network outages happen when a component that traffic depends on stops working, and the large ones spread because many services share that component. Published post-incident reports from carriers and cloud providers trace them to a short list of triggers, such as a faulty configuration change, a latent software defect, overload made worse by retries, a failed dependency such as DNS, and physical damage such as a cut cable, fire or water.

Learning objectives

After reading this article you will be able to:

  • Identify the common triggers of network outages in carrier and cloud reports
  • Explain how shared components, hidden dependencies and feedback loops turn small faults into outages
  • Describe when a US wireless outage must be reported to the FCC

What triggers an outage

The incidents below are each described in the operator’s own report, or in a regulator’s report on it. Their triggers fall into a few groups.

  • A change goes wrong. On February 22, 2024 at 2:42 AM, an AT&T Mobility employee placed a misconfigured network element into the production network during a routine night maintenance window. Three minutes later a nationwide outage began. The FCC’s report found the configuration had not been peer reviewed and was not adequately tested. On October 4, 2021, a command issued during routine maintenance at Meta, meant to assess the availability of backbone capacity, took down all the connections in its backbone network; a bug in the audit tool that should have blocked the command let it through.
  • A latent defect fires. On October 19 and 20, 2025, a latent race condition in Amazon DynamoDB’s DNS management system left an incorrect, empty DNS record for the service’s regional endpoint in us-east-1, and the automation failed to repair it. On November 18, 2025, a change to database permissions at Cloudflare caused a “feature file” used by its Bot Management system to double in size. The software that reads that file on every machine had a size limit below the new size, so it failed.
  • Load feeds on itself. On December 7, 2021, an automated scaling activity in AWS’s us-east-1 Region triggered a surge of connections that overwhelmed the devices linking AWS’s internal network to its main network. The resulting delays caused more connection attempts and retries, and a latent issue stopped clients from backing off as designed.
  • Something physical breaks. NIST’s contingency planning guide lists threats to network cabling such as cable cuts, electromagnetic and radio frequency interference, fire and water damage.

Why a small fault becomes a large outage

The trigger is often small. What makes it an outage is how far it reaches and how long recovery takes.

Shared components. Cloudflare’s bad file was propagated to all the machines in its network, so every one of them hit the same limit. A single misconfigured element at AT&T sent traffic that pushed the network into a protective mode which disconnected all devices. When one thing is used everywhere, a fault in it is felt everywhere, which is why a single point of failure matters so much.

Dependencies that hide. Meta’s DNS servers were designed to withdraw their BGP route advertisements if they could not reach Meta’s data centres, as a health safeguard. When the backbone went down they all did so, and became unreachable even though they were still running. DNS was not the trigger, but it meant the rest of the internet could not even find Meta’s servers.

Feedback loops. Google’s SRE book defines a cascading failure as one that grows over time through positive feedback: one replica fails from overload, its load moves to the others, and they fail in turn. The AWS 2021 event shows the network version: delays caused retries, and retries caused more delay. This is the network congestion problem at the scale of a data centre, and the reason clients are expected to use exponential backoff.

Recovery is its own load. AT&T rolled back its change in close to two hours, but full restoration took at least 12 hours because its device registration systems were overwhelmed by re-registration requests. Cloudflare reported core traffic largely flowing by 14:30 UTC, then spent hours handling the load of traffic rushing back, with all systems normal at 17:06 UTC. The AWS 2021 congestion also cut off real-time monitoring data, so operators had to diagnose the problem from logs.

How outages get reported

In the United States, the FCC collects outage reports from communications providers through its Network Outage Reporting System (NORS). As the FCC’s AT&T report summarises the rules, a wireless outage is reportable when, among other criteria, it lasts at least 30 minutes and potentially affects at least 900,000 user minutes of telephony or a 911 special facility. Providers must notify the FCC within 120 minutes of discovering it, file an initial report within 72 hours and a final report within 30 days.

Cloud and network providers publish their own post-incident summaries, like the AWS, Cloudflare and Meta reports cited here. They are useful because they include timelines and contributing causes, not only the trigger.

What this means for an app

An app sits on top of all of this. It reaches its backend through a carrier or Wi-Fi network, a DNS name, a cloud region and often a content delivery network, and any of them can fail for one of the reasons above. A few habits follow from the reports:

  • Treat every network call as something that can fail or hang, and retry with backoff and limits, not in a tight loop.
  • Keep work the user has done on the device until a backend confirms it, so an outage delays the work instead of losing it. How do apps work without internet? covers the patterns.
  • Decide which features need the cloud and which do not. Two devices in the same room do not need a cloud region to exchange data if the app can use a direct link between them.
  • Test the failure cases on purpose, including a backend that is down, slow, or returns errors. How do you test an app for offline behaviour? has a checklist.

Frequently asked questions

Are big outages usually cyber attacks?

None of the reports cited here were. Cloudflare wrote that its November 18, 2025 outage was not caused by a cyber attack or malicious activity, and AT&T's initial statement, quoted in the FCC report, attributed the February 2024 outage to an incorrect process used while expanding its network.

How long does an outage last?

It depends on the cause and on how long recovery takes. Cloudflare's November 18, 2025 failures began at 11:20 UTC and all its systems were normal at 17:06 UTC; the FCC found AT&T's February 22, 2024 outage lasted at least twelve hours.

Build it with Offline Protocol

The production page lists the failure boundaries to test before rollout, including internet loss while peers keep communicating, backend outage and recovery, and process restart with pending records.

Read Deploy to production