Skip to content
← cd ../blog

Your cluster says 3/3. It will still go dark with one zone.

Sai Pisey4 min readKubernetes · Scheduling · Reliability

In short

  • A deployment can report 3/3 ready and still stop serving when one availability zone fails: topology spread constraints apply at scheduling time and are never re-checked.
  • kubectl survive-zone reads real pod placement and Service dependencies, removes one zone at a time, and reports which workloads are lost, degraded or impaired.
  • It only prints a fix after proving it with the upstream kube-scheduler's own Filter plugins, and its verdicts are checked nightly against a real control plane in 200 randomised scenarios.

Most teams believe they are multi-AZ because their manifests say so. The manifests describe intent. The cluster holds the placement, and the two drift apart without anyone making a mistake.

Here is a three-zone cluster where every deployment reports healthy:

terminal
kubectl get deployments
NAME            READY   UP-TO-DATE   AVAILABLE
checkout-api    3/3     3            3
session-store   1/1     1            1
web             3/3     3            3

And here is what happens to it if us-east-1a goes away:

terminal
kubectl survive-zone
Losing us-east-1a  ->  2 lost, 0 degraded, 1 impaired by a dependency
  WORKLOAD       PLACEMENT                               VERDICT
  checkout-api   us-east-1a:3                            LOST all 3 available replicas are in us-east-1a
  session-store  us-east-1a:1                            LOST all 1 available replicas are in us-east-1a
  web            us-east-1a:1 us-east-1b:1 us-east-1c:1  IMPAIRED depends on session-store

us-east-1aLOST

checkoutcheckoutcheckoutsessionweb

us-east-1b

web

us-east-1c

web

workload

kubectl says

lose 1a

checkout-api3/3LOSTsession-store1/1LOSTweb3/3IMPAIRED
Every deployment reports ready, yet five of the seven pods sit in one zone. Losing us-east-1a takes two workloads down outright and leaves the third without its dependency.

checkout-api was scheduled while us-east-1a was the only zone. The scheduler put all three replicas there, correctly, and nothing moved them when two more zones were added. Topology constraints bind at scheduling time and never again; the Kubernetes documentation says so directly: there is no guarantee they stay satisfied once pods move.

web is the more interesting row. It is spread one replica per zone, which is exactly what a linter would want to see, and it still stops serving, because the only session-store pod is in the zone that died. Zone survivability is a property of the dependency graph, not of each workload.

webus-east-1a
webus-east-1b
webus-east-1c
session-storeus-east-1a · only replica
web2 of 3 replicas aliveIMPAIRED
A spread linter passes web: one replica per zone. The survivability check follows its Service to session-store, whose only pod is in us-east-1a, so two healthy web replicas still cannot serve.

What it does

kubectl survive-zone reads real pod placement rather than manifests, removes one failure domain at a time, and reports what stops serving. It follows Services to their EndpointSlices to build that dependency graph, and it separately finds PodDisruptionBudgets that will hang a node drain.

It is read-only. It never evicts, cordons or deletes anything, and it never reads ConfigMap or Secret contents.

Advice is only printed once it is proven

Every survivability fix can make pods unschedulable. Tightening ScheduleAnyway to DoNotSchedule, or preferred anti-affinity to required, can leave replicas Pending, which is the outage the tool was run to prevent. So kubectl survive-zone fix does not print a patch until two things hold on the cluster's actual nodes:

terminal
FIX  Add an enforced topology spread constraint on topology.kubernetes.io/zone
     ├─ schedulable?  yes — us-east-1a:1 us-east-1b:1 us-east-1c:1
     └─ survives us-east-1a?  yes — no domain still reports this workload lost

ALT  Add a PodDisruptionBudget with maxUnavailable: 1
     ├─ schedulable?  yes — us-east-1a:3
     └─ survives us-east-1a?  NO — still lost losing us-east-1a
FIXEnforced topology spread on zone1 · schedulable?yes · 1a:1 1b:1 1c:12 · survives 1a?yes · nothing lostresultprinted as the patch
ALTPodDisruptionBudget, maxUnavailable 11 · schedulable?yes · 1a:32 · survives 1a?NO · still lostresultshown, marked as not a fix
Proof 1 runs the upstream kube-scheduler's Filter plugins on the cluster's real nodes. Proof 2 re-runs the zone analysis on where the replicas would land. Only a candidate that passes both is printed as the fix.

The first proof runs the upstream kube-scheduler's own Filter plugins, embedded in the binary. The second re-runs the survivability analysis on where the replicas would actually land. The PDB option is shown and marked as not solving the problem, because a budget is never consulted when a zone's nodes simply vanish: PodDisruptionBudgets only govern voluntary disruptions.

The simulation only credits Filter, never Score. Filter is what the scheduler must honour; Score is what it may prefer. A workload that ends up spread only because scoring happened to favour an empty zone has no guarantee, and crediting it would be the same false confidence the tool exists to remove.

A drain that hangs forever, with a budget that looks fine

A PodDisruptionBudget can be satisfiable on paper and still hang a drain forever. Give a workload an enforced spread constraint and a budget of one disruption at a time, then drain a zone. The first eviction succeeds. Its replacement cannot schedule anywhere, because every surviving zone would violate the skew and the draining zone is cordoned, so availability never recovers and the budget refuses every later eviction.

  1. step 1Drain us-east-1aNodes cordoned. Budget allows one disruption.
  2. step 2First eviction succeedsBudget now spent: one replica is down.
  3. step 3Replacement stays PendingEvery surviving zone would break maxSkew 1, and 1a is cordoned.
  4. step 4Every later eviction refusedAvailability never recovers, so the budget never frees up.

steps 3 → 4 loop; the drain never finishes

Budget arithmetic says the drain is safe. It is not: step 3 can never complete, so step 4 repeats forever. Seeing this needs scheduling feasibility, which is why the tool simulates the replacement pod before anyone drains.

Budget arithmetic cannot see this. It needs scheduling feasibility, so the tool simulates the replacement against the real scheduler and reports the deadlock before anyone starts the drain.

Survivability decays

The failure that matters happens with nothing changed in git. A node drains at night, pods reschedule, and a workload that was spread is now in one zone. A command someone has to remember to run will not catch that, so the tool also runs as a Prometheus exporter:

terminal
survive_workload_survives{domain="us-east-1a",workload="default/web"}  1  ->  0
survive_workloads_lost{domain="us-east-1a"}                            2  ->  3
survive_scrape_success                                                 1

cordon 2 zones · no deploy

survive_workload_survives{workload="web"}10
survive_workloads_lost{domain="us-east-1a"}23
survive_scrape_success 1 throughoutalert firing
Two zones cordoned, no deploy, no change in git. web stops surviving the loss of us-east-1a and the lost count rises. The shipped alert fires on exactly this edge, and only while survive_scrape_success stays 1.

That change came from cordoning two zones, with no deploy. The shipped alert fires on exactly this: a workload survived losing a zone an hour ago, does not now, and the exporter itself is healthy. That last condition matters, because a failed analysis and a healthy cluster would otherwise both read as zero workloads lost.

How the verdicts are checked

Unit tests encode what I believe Kubernetes does. So the verdicts are checked against a real control plane instead. A differential harness runs an unmodified kube-apiserver, scheduler and controller-manager under KWOK, predicts each outcome, then actually cordons and drains a zone through the eviction API and compares. It runs 200 randomised scenarios every night.

That harness found problems no unit test did. It caught a refused eviction being read as evidence of survival, and a drain that was only throttled being treated as permanently blocked. It also found the deadlock above, which is why that detector exists.

Something the demo found

While recording, I added an environment variable to web with kubectl set env. That triggered a rolling update, and afterwards web sat at us-east-1a:1 us-east-1b:2, with nothing in us-east-1c, despite an enforced maxSkew: 1.

During a rollout the old pods still count toward the spread, so the new ones land against a skew that is about to disappear. Once the old pods terminate, nothing rebalances. The fix is matchLabelKeys: [pod-template-hash] on the constraint, which most spread constraints do not set. An ordinary deploy quietly violated an enforced constraint, and nothing alerted.

before

1a
1b
1c
1 · 1 · 1

during the rollout

1a
1b
1c
old pods (dashed) still count

after

1a
1b
1c
1 · 2 · 0 · skew 2
New pods are placed while the old ones still count toward the spread, so they land against a skew that is about to vanish. Once the old pods terminate nothing rebalances. Adding matchLabelKeys: [pod-template-hash] makes each revision count only its own pods.

Limits

  • It models involuntary domain loss, where a PodDisruptionBudget offers no protection. A graceful drain is reported separately.
  • Dependency edges come only from EndpointSlices and literal environment values that name a Service. ConfigMap and Secret contents are never read, so dependencies configured there are invisible.
  • ExternalName and selectorless Services are shown as dependencies but never reported as impairing, because their placement cannot be proven from the cluster.
  • Remediation rungs are independent. A single replica in a single zone needs more replicas and an enforced spread together, and each rung alone is correctly reported as partial.
  • The scheduler-backed checks target Kubernetes 1.35. On another minor they are disabled with a named warning and the survivability analysis still runs.

Try it

It installs through krew and only reads from the cluster:

terminal
kubectl krew install survive-zone
kubectl survive-zone

Source, the exporter and the alert rules are at github.com/SaiPisey2/kubectl-survive.

Ready to Build Something Resilient?

Whether it's hardening your platform, untangling a messy CI/CD setup, or just talking infrastructure over a call, let's connect. I'm always up for a good systems conversation.

Book a call

$ ./contact.sh

Get in Touch

I only use these details to reply to you. Ask and I'll delete them.