Multi-zone is not enough: finding shared fate in resilient architecture

Multi-zone is not enough: finding shared fate in resilient architecture

Replicas across zones solve only part of the problem. Find the control plane, identity layer, and broad changes that can fail several zones together.

Spreading workloads across several availability zones is a strong starting point for resilience, but it does not automatically create independence. A system can still fail together when replicas rely on one DNS layer, control plane, configuration store, identity provider, or change that spreads too widely.

Draw failure boundaries, not only server boundaries

Alongside a network diagram, create a dependency map covering five layers: data, control, identity, connectivity, and change. For each component, record where it resides, what it needs to function, who can change it, and what happens when it slows down or becomes unavailable.

Three conditions that are often missed

  1. Correct quorum: the remaining nodes must still be able to make decisions after the loss of any one zone.
  2. Available capacity: the surviving zones need tested headroom instead of waiting for scale-out while the control plane is stressed.
  3. Intentional locality: where possible, consecutive calls should favor local paths so one zonal failure does not compound across a transaction chain.

Two forms of shared fate to hunt

The first is a shared control dependency: replicas remain alive but cannot be discovered, routed, or authorized. The second is a wide-blast-radius change: a syntactically valid policy, pipeline, or automation is applied incorrectly across every zone. Both can be invisible on diagrams that show only compute and databases.

Rehearse rather than assume

Choose one zone, one control component, and one broadly scoped configuration for separate exercises. Verify whether the application serves in a degraded mode, whether dashboards identify the correct cause, and whether operators can stop a change. Record dependencies that cannot be isolated and turn them into an architecture backlog.

Resilience is not the number of zones printed on a diagram. It is the result of understood dependencies, tested limits, and a recovery path that actually works.

Published ; updated

Related pages