What it means
AWS’s Well-Architected reliability guidance puts it directly: application components should continue to perform their core function even if their dependencies become unavailable. They might serve slightly stale data, alternate data, or even no data, but the central job still gets done. The same guidance asks teams to treat a component’s failure modes as normal operation and to design workflows so a dependency failure ends in a predictable, recoverable state rather than a complete outage.
The core idea is to separate hard dependencies from soft ones. A hard dependency is one the feature cannot work without. A soft dependency improves the result but can be missing for a while. Graceful degradation is the work of moving as many dependencies as possible from the first group to the second, and deciding in advance what the user sees when each one fails. It complements failover, which moves work to a replacement rather than doing less with what is left.
That decision is partly a business one. AWS notes that when not every requirement can be met, a team has to choose which matters more: for payment processing it might be consistency, for a real-time application it might be availability.
Degrading under overload
Google’s Site Reliability Engineering book treats degradation as a response to too much load. Its chapter on handling overload describes serving degraded responses, ones that are less accurate or contain less data than normal but are cheaper to compute. Its examples are searching only a small percentage of the candidate set instead of the whole corpus, and relying on a local copy of results that may not be fully up to date instead of the canonical storage.
The chapter on cascading failures separates two techniques:
- Load shedding drops some traffic as a server approaches overload, for example by returning HTTP 503 when too many requests are already in flight, so the server keeps doing as much useful work as it can.
- Graceful degradation goes further by reducing the work each request needs, such as searching an in-memory subset of data or using a cheaper ranking algorithm.
The book also describes tagging each request with a criticality, from CRITICAL_PLUS down to SHEDDABLE, so the least important traffic is the first to be refused when capacity runs short.
Degrading when a dependency fails
The AWS guidance lists concrete cases:
- A landing page that combines recommendations, top products and order status keeps showing the other sections when one upstream system fails, instead of an error page.
- A batch writer keeps processing when one operation fails, and either reports which items failed or puts failed requests into a dead-letter queue for later retries.
- A caller stops hammering an overloaded downstream service with a circuit breaker, letting only occasional calls through to test recovery.
- A component falls back to cached values or sensible defaults when its parameter store is unavailable.
- A database primary that is down still allows reads from replicas, and writes can be buffered in a queue.
It also names anti-patterns: serving no data when only one of several dependencies is down, and emptying local state after a failed refresh.
Degrading on the device
For a mobile app, one dependency that can disappear at any moment is the network. Android’s offline-first guidance makes the local data source the canonical source of truth that the rest of the app reads from, so screens still show data when the network side fails, and it describes retrying network reads with exponential backoff. An app degrades gracefully when it shows cached data with an honest age, queues the user’s changes, switches off only the features that need a live connection, and shows its sync status so people know what has and has not been sent.
Keeping the degraded path working
A degraded mode that never runs tends to break. The SRE book warns that the code path you never use is the code path that often does not work, and suggests regularly running a small subset of servers near overload to exercise it, along with alerts when too many servers enter degraded mode. AWS adds that failure pathways must be tested and should be significantly simpler than the primary pathway.
Retries need the same care. The SRE book recommends limiting retries per request, considering a server-wide retry budget, and avoiding retries at several layers at once, which multiply. AWS recommends exponential backoff with jitter and a maximum number of retries, and warns against retrying operations that are not idempotent. Without those limits, a system trying to recover can keep itself overloaded, the same feedback loop that drives network congestion towards collapse.