Connectivity resilience

What is graceful degradation?

Graceful degradation is designing a system so that when a dependency fails or load exceeds capacity, it keeps doing its most important work at reduced quality, such as stale data, fewer results or missing extras, instead of failing completely. It turns hard dependencies, without which nothing works, into soft ones the system can do without for a while.

Learning objectives

After reading this article you will be able to:

  • Distinguish hard dependencies from soft ones when designing for graceful degradation
  • Compare load shedding with graceful degradation as responses to overload
  • Describe how a mobile app degrades gracefully when the network disappears

What it means

AWS’s Well-Architected reliability guidance puts it directly: application components should continue to perform their core function even if their dependencies become unavailable. They might serve slightly stale data, alternate data, or even no data, but the central job still gets done. The same guidance asks teams to treat a component’s failure modes as normal operation and to design workflows so a dependency failure ends in a predictable, recoverable state rather than a complete outage.

The core idea is to separate hard dependencies from soft ones. A hard dependency is one the feature cannot work without. A soft dependency improves the result but can be missing for a while. Graceful degradation is the work of moving as many dependencies as possible from the first group to the second, and deciding in advance what the user sees when each one fails. It complements failover, which moves work to a replacement rather than doing less with what is left.

That decision is partly a business one. AWS notes that when not every requirement can be met, a team has to choose which matters more: for payment processing it might be consistency, for a real-time application it might be availability.

Degrading under overload

Google’s Site Reliability Engineering book treats degradation as a response to too much load. Its chapter on handling overload describes serving degraded responses, ones that are less accurate or contain less data than normal but are cheaper to compute. Its examples are searching only a small percentage of the candidate set instead of the whole corpus, and relying on a local copy of results that may not be fully up to date instead of the canonical storage.

The chapter on cascading failures separates two techniques:

  • Load shedding drops some traffic as a server approaches overload, for example by returning HTTP 503 when too many requests are already in flight, so the server keeps doing as much useful work as it can.
  • Graceful degradation goes further by reducing the work each request needs, such as searching an in-memory subset of data or using a cheaper ranking algorithm.

The book also describes tagging each request with a criticality, from CRITICAL_PLUS down to SHEDDABLE, so the least important traffic is the first to be refused when capacity runs short.

Degrading when a dependency fails

The AWS guidance lists concrete cases:

  • A landing page that combines recommendations, top products and order status keeps showing the other sections when one upstream system fails, instead of an error page.
  • A batch writer keeps processing when one operation fails, and either reports which items failed or puts failed requests into a dead-letter queue for later retries.
  • A caller stops hammering an overloaded downstream service with a circuit breaker, letting only occasional calls through to test recovery.
  • A component falls back to cached values or sensible defaults when its parameter store is unavailable.
  • A database primary that is down still allows reads from replicas, and writes can be buffered in a queue.

It also names anti-patterns: serving no data when only one of several dependencies is down, and emptying local state after a failed refresh.

Degrading on the device

For a mobile app, one dependency that can disappear at any moment is the network. Android’s offline-first guidance makes the local data source the canonical source of truth that the rest of the app reads from, so screens still show data when the network side fails, and it describes retrying network reads with exponential backoff. An app degrades gracefully when it shows cached data with an honest age, queues the user’s changes, switches off only the features that need a live connection, and shows its sync status so people know what has and has not been sent.

Keeping the degraded path working

A degraded mode that never runs tends to break. The SRE book warns that the code path you never use is the code path that often does not work, and suggests regularly running a small subset of servers near overload to exercise it, along with alerts when too many servers enter degraded mode. AWS adds that failure pathways must be tested and should be significantly simpler than the primary pathway.

Retries need the same care. The SRE book recommends limiting retries per request, considering a server-wide retry budget, and avoiding retries at several layers at once, which multiply. AWS recommends exponential backoff with jitter and a maximum number of retries, and warns against retrying operations that are not idempotent. Without those limits, a system trying to recover can keep itself overloaded, the same feedback loop that drives network congestion towards collapse.

Frequently asked questions

Is graceful degradation the same as failover?

No. Failover moves work to a replacement, such as a standby server or another region, and aims to keep full service. Graceful degradation keeps running on what is left and accepts a reduced service. Systems use both.

Does graceful degradation apply to the web front end too?

Yes. A page that still renders its main content when a recommendations widget or analytics script fails to load is degrading gracefully. The AWS guidance gives the example of an online shop that keeps showing the rest of its landing page when one upstream system fails.

Sources

Build it with Offline Protocol

The page on what the SDK handles sets out which parts of a workflow the SDK covers and which stay with your application and backend, including finishing an authorized handoff locally and submitting it once the backend is reachable.

Read What the SDK handles