Connectivity resilience

What happens to apps when the cloud goes down?

When a cloud service an app depends on fails, every feature that needs a round trip to it stops, which can include sign-in, loading data, saving changes and sending messages. How much of the app keeps working depends on what it can do with data and connections already on the device. Published post-incident reports show failures that started in one shared dependency, such as a DNS record or a configuration file, and spread to everything built on it.

Learning objectives

After reading this article you will be able to:

  • Explain why a cloud provider failure stops every app feature that needs a round trip
  • Describe what the AWS, Google Cloud and Cloudflare reports show about shared dependencies
  • List ways an app can keep working while a cloud service is down

What breaks

A cloud-backed app makes a request for nearly everything: signing in, loading a screen, saving a change, sending a message, fetching configuration. When the service behind those requests fails, the phone’s network is still fine, so the app sees errors and timeouts rather than an offline state. Anything that cannot complete without that round trip stops.

The service that fails is not always the app’s own backend. Apps sit on layers they do not run: DNS, a database service, an API gateway, a content delivery network, an identity provider. When one of those layers fails, every app built on it fails at the same moment, regardless of how healthy its own code is. Each such layer is a single point of failure the app inherits without having chosen it.

Three recent outages, from the operators’ own reports

ProviderDateTriggerVisible effect
AWS, us-east-119 to 20 October 2025A latent race condition in DynamoDB’s automated DNS management left an empty DNS record for the regional endpointSystems could not resolve or connect to DynamoDB; knock-on failures in EC2 launches, Network Load Balancer, Lambda and console sign-in
Google Cloud12 June 2025An automated quota policy update with unintended blank fields, replicated globally within seconds, hit a code path without error handlingService Control binaries crash-looped and external API requests across Google Cloud and Google Workspace returned 503 errors
Cloudflare18 November 2025A database permissions change made a Bot Management feature file double in size, past a limit in the proxy softwareUsers saw Cloudflare error pages for customers’ sites; Turnstile failed to load, so dashboard users without an active session could not log in

The timelines show how long the tail can be. AWS reports that the event started at 11:48 PM PDT on 19 October and that DynamoDB’s DNS information was restored by 2:25 AM PDT on 20 October, but the event did not end until 2:20 PM PDT that afternoon, because EC2’s internal systems and the load balancer health checks took hours more to recover. Google reports an incident start of 10:49 and an end of 13:49 US Pacific time, with us-central1 the last region to recover. Cloudflare reports that failures began at 11:20 UTC, core traffic was largely flowing by 14:30, and all systems were normal by 17:06.

What the reports have in common

  • One shared dependency, spread fast. A DNS record, a quota policy and a configuration file are small things, but each sat underneath other services. Internal AWS services relied on DynamoDB, Service Control performs policy checks for Google’s external APIs, and Cloudflare’s traffic-routing software on every machine reads the feature file. The Google policy data was replicated globally within seconds, and the Cloudflare file was propagated to all the machines in its network.
  • Recovery is slower than the fix. Once the cause was removed, backlogs and restarts became their own problem. Google reports a herd effect in us-central1 as Service Control tasks restarted and overloaded the database they depend on, because the service lacked randomized exponential backoff. Cloudflare reports a backlog of login attempts overwhelming its dashboard after the fix.
  • Visibility can fail too. Google’s first incident report was posted about an hour after the crashes began because its Cloud Service Health infrastructure was itself affected, and some customers’ monitoring running on Google Cloud was failing as well.

What an app can do

An app cannot stop a provider’s outage, but it can decide how much of itself depends on one.

Keep a local source of truth. Android’s offline-first guidance describes an app that can perform all, or a critical subset, of its core functionality without internet access, and says the local data source should be the canonical source of truth that the rest of the app reads from. That keeps screens populated when the network side fails. The design is covered in how to design an offline-first architecture.

Queue writes instead of losing them. A change the user makes during the outage can be stored on the device with a stable identifier and sent when the service returns, as described in what happens to a write made offline.

Retry without making it worse. Use exponential backoff with jitter and a cap, and do not retry errors that will not succeed on a second attempt. Google’s report lists auditing its systems for randomized exponential backoff among its follow-up actions.

Degrade rather than stop. Show cached data with its age, disable only the features that need the live service, and tell the user what is unavailable. Graceful degradation covers the patterns.

Do not route local work through a distant region. If two devices in the same room need to exchange a record, a round trip through a cloud region makes that exchange depend on the region. A direct path between the devices does not. Offline Protocol is built on this split: your backend stays the system of record, and the SDK adds device-to-device paths so nearby devices can keep exchanging and recording work locally while the backend is unreachable, as its core concepts describe.

Frequently asked questions

Does running in more than one region prevent this?

It helps for failures confined to one region. In the October 2025 AWS event, customers with DynamoDB global tables could still use their replicas in other regions, with replication lag to and from the affected one. It does not help when the failing piece is global, as in the Google Cloud and Cloudflare reports.

Can an app tell the cloud is down rather than the phone being offline?

Not reliably from one failed request. The device may have a working network while the service returns errors or timeouts. Check the platform's connectivity signal, and treat repeated server errors as the service being unavailable rather than as the user being offline.

Sources

Build it with Offline Protocol

The core concepts page explains how a local path between devices can remain available when internet access is lost, and why protocol delivery, application acceptance and backend acceptance are tracked as separate outcomes.

Read Core concepts