What breaks
A cloud-backed app makes a request for nearly everything: signing in, loading a screen, saving a change, sending a message, fetching configuration. When the service behind those requests fails, the phone’s network is still fine, so the app sees errors and timeouts rather than an offline state. Anything that cannot complete without that round trip stops.
The service that fails is not always the app’s own backend. Apps sit on layers they do not run: DNS, a database service, an API gateway, a content delivery network, an identity provider. When one of those layers fails, every app built on it fails at the same moment, regardless of how healthy its own code is. Each such layer is a single point of failure the app inherits without having chosen it.
Three recent outages, from the operators’ own reports
| Provider | Date | Trigger | Visible effect |
|---|---|---|---|
| AWS, us-east-1 | 19 to 20 October 2025 | A latent race condition in DynamoDB’s automated DNS management left an empty DNS record for the regional endpoint | Systems could not resolve or connect to DynamoDB; knock-on failures in EC2 launches, Network Load Balancer, Lambda and console sign-in |
| Google Cloud | 12 June 2025 | An automated quota policy update with unintended blank fields, replicated globally within seconds, hit a code path without error handling | Service Control binaries crash-looped and external API requests across Google Cloud and Google Workspace returned 503 errors |
| Cloudflare | 18 November 2025 | A database permissions change made a Bot Management feature file double in size, past a limit in the proxy software | Users saw Cloudflare error pages for customers’ sites; Turnstile failed to load, so dashboard users without an active session could not log in |
The timelines show how long the tail can be. AWS reports that the event started at 11:48 PM PDT on 19 October and that DynamoDB’s DNS information was restored by 2:25 AM PDT on 20 October, but the event did not end until 2:20 PM PDT that afternoon, because EC2’s internal systems and the load balancer health checks took hours more to recover. Google reports an incident start of 10:49 and an end of 13:49 US Pacific time, with us-central1 the last region to recover. Cloudflare reports that failures began at 11:20 UTC, core traffic was largely flowing by 14:30, and all systems were normal by 17:06.
What the reports have in common
- One shared dependency, spread fast. A DNS record, a quota policy and a configuration file are small things, but each sat underneath other services. Internal AWS services relied on DynamoDB, Service Control performs policy checks for Google’s external APIs, and Cloudflare’s traffic-routing software on every machine reads the feature file. The Google policy data was replicated globally within seconds, and the Cloudflare file was propagated to all the machines in its network.
- Recovery is slower than the fix. Once the cause was removed, backlogs and restarts became their own problem. Google reports a herd effect in us-central1 as Service Control tasks restarted and overloaded the database they depend on, because the service lacked randomized exponential backoff. Cloudflare reports a backlog of login attempts overwhelming its dashboard after the fix.
- Visibility can fail too. Google’s first incident report was posted about an hour after the crashes began because its Cloud Service Health infrastructure was itself affected, and some customers’ monitoring running on Google Cloud was failing as well.
What an app can do
An app cannot stop a provider’s outage, but it can decide how much of itself depends on one.
Keep a local source of truth. Android’s offline-first guidance describes an app that can perform all, or a critical subset, of its core functionality without internet access, and says the local data source should be the canonical source of truth that the rest of the app reads from. That keeps screens populated when the network side fails. The design is covered in how to design an offline-first architecture.
Queue writes instead of losing them. A change the user makes during the outage can be stored on the device with a stable identifier and sent when the service returns, as described in what happens to a write made offline.
Retry without making it worse. Use exponential backoff with jitter and a cap, and do not retry errors that will not succeed on a second attempt. Google’s report lists auditing its systems for randomized exponential backoff among its follow-up actions.
Degrade rather than stop. Show cached data with its age, disable only the features that need the live service, and tell the user what is unavailable. Graceful degradation covers the patterns.
Do not route local work through a distant region. If two devices in the same room need to exchange a record, a round trip through a cloud region makes that exchange depend on the region. A direct path between the devices does not. Offline Protocol is built on this split: your backend stays the system of record, and the SDK adds device-to-device paths so nearby devices can keep exchanging and recording work locally while the backend is unreachable, as its core concepts describe.