Your cluster says 3/3. It will still go dark with one zone.
Sai Pisey4 min readKubernetes · Scheduling · Reliability
In short
- A deployment can report 3/3 ready and still stop serving when one availability zone fails: topology spread constraints apply at scheduling time and are never re-checked.
- kubectl survive-zone reads real pod placement and Service dependencies, removes one zone at a time, and reports which workloads are lost, degraded or impaired.
- It only prints a fix after proving it with the upstream kube-scheduler's own Filter plugins, and its verdicts are checked nightly against a real control plane in 200 randomised scenarios.
Most teams believe they are multi-AZ because their manifests say so. The manifests describe intent. The cluster holds the placement, and the two drift apart without anyone making a mistake.
Here is a three-zone cluster where every deployment reports healthy:
kubectl get deployments
NAME READY UP-TO-DATE AVAILABLE
checkout-api 3/3 3 3
session-store 1/1 1 1
web 3/3 3 3And here is what happens to it if us-east-1a goes away:
kubectl survive-zone
Losing us-east-1a -> 2 lost, 0 degraded, 1 impaired by a dependency
WORKLOAD PLACEMENT VERDICT
checkout-api us-east-1a:3 LOST all 3 available replicas are in us-east-1a
session-store us-east-1a:1 LOST all 1 available replicas are in us-east-1a
web us-east-1a:1 us-east-1b:1 us-east-1c:1 IMPAIRED depends on session-storeus-east-1aLOST
us-east-1b
us-east-1c
workload
kubectl says
lose 1a
checkout-api3/3LOSTsession-store1/1LOSTweb3/3IMPAIREDcheckout-api was scheduled while us-east-1a was the only zone. The scheduler put all three replicas there, correctly, and nothing moved them when two more zones were added. Topology constraints bind at scheduling time and never again; the Kubernetes documentation says so directly: there is no guarantee they stay satisfied once pods move.
web is the more interesting row. It is spread one replica per zone, which is exactly what a linter would want to see, and it still stops serving, because the only session-store pod is in the zone that died. Zone survivability is a property of the dependency graph, not of each workload.
What it does
kubectl survive-zone reads real pod placement rather than manifests, removes one failure domain at a time, and reports what stops serving. It follows Services to their EndpointSlices to build that dependency graph, and it separately finds PodDisruptionBudgets that will hang a node drain.
It is read-only. It never evicts, cordons or deletes anything, and it never reads ConfigMap or Secret contents.
Advice is only printed once it is proven
Every survivability fix can make pods unschedulable. Tightening ScheduleAnyway to DoNotSchedule, or preferred anti-affinity to required, can leave replicas Pending, which is the outage the tool was run to prevent. So kubectl survive-zone fix does not print a patch until two things hold on the cluster's actual nodes:
FIX Add an enforced topology spread constraint on topology.kubernetes.io/zone
├─ schedulable? yes — us-east-1a:1 us-east-1b:1 us-east-1c:1
└─ survives us-east-1a? yes — no domain still reports this workload lost
ALT Add a PodDisruptionBudget with maxUnavailable: 1
├─ schedulable? yes — us-east-1a:3
└─ survives us-east-1a? NO — still lost losing us-east-1acandidate
1 · schedulable?
2 · survives 1a?
result
The first proof runs the upstream kube-scheduler's own Filter plugins, embedded in the binary. The second re-runs the survivability analysis on where the replicas would actually land. The PDB option is shown and marked as not solving the problem, because a budget is never consulted when a zone's nodes simply vanish: PodDisruptionBudgets only govern voluntary disruptions.
The simulation only credits Filter, never Score. Filter is what the scheduler must honour; Score is what it may prefer. A workload that ends up spread only because scoring happened to favour an empty zone has no guarantee, and crediting it would be the same false confidence the tool exists to remove.
A drain that hangs forever, with a budget that looks fine
A PodDisruptionBudget can be satisfiable on paper and still hang a drain forever. Give a workload an enforced spread constraint and a budget of one disruption at a time, then drain a zone. The first eviction succeeds. Its replacement cannot schedule anywhere, because every surviving zone would violate the skew and the draining zone is cordoned, so availability never recovers and the budget refuses every later eviction.
- step 1Drain us-east-1aNodes cordoned. Budget allows one disruption.
- step 2First eviction succeedsBudget now spent: one replica is down.
- step 3Replacement stays PendingEvery surviving zone would break maxSkew 1, and 1a is cordoned.
- step 4Every later eviction refusedAvailability never recovers, so the budget never frees up.
steps 3 → 4 loop; the drain never finishes
Budget arithmetic cannot see this. It needs scheduling feasibility, so the tool simulates the replacement against the real scheduler and reports the deadlock before anyone starts the drain.
Survivability decays
The failure that matters happens with nothing changed in git. A node drains at night, pods reschedule, and a workload that was spread is now in one zone. A command someone has to remember to run will not catch that, so the tool also runs as a Prometheus exporter:
survive_workload_survives{domain="us-east-1a",workload="default/web"} 1 -> 0
survive_workloads_lost{domain="us-east-1a"} 2 -> 3
survive_scrape_success 1cordon 2 zones · no deploy
That change came from cordoning two zones, with no deploy. The shipped alert fires on exactly this: a workload survived losing a zone an hour ago, does not now, and the exporter itself is healthy. That last condition matters, because a failed analysis and a healthy cluster would otherwise both read as zero workloads lost.
How the verdicts are checked
Unit tests encode what I believe Kubernetes does. So the verdicts are checked against a real control plane instead. A differential harness runs an unmodified kube-apiserver, scheduler and controller-manager under KWOK, predicts each outcome, then actually cordons and drains a zone through the eviction API and compares. It runs 200 randomised scenarios every night.
That harness found problems no unit test did. It caught a refused eviction being read as evidence of survival, and a drain that was only throttled being treated as permanently blocked. It also found the deadlock above, which is why that detector exists.
Something the demo found
While recording, I added an environment variable to web with kubectl set env. That triggered a rolling update, and afterwards web sat at us-east-1a:1 us-east-1b:2, with nothing in us-east-1c, despite an enforced maxSkew: 1.
During a rollout the old pods still count toward the spread, so the new ones land against a skew that is about to disappear. Once the old pods terminate, nothing rebalances. The fix is matchLabelKeys: [pod-template-hash] on the constraint, which most spread constraints do not set. An ordinary deploy quietly violated an enforced constraint, and nothing alerted.
before
during the rollout
after
matchLabelKeys: [pod-template-hash] makes each revision count only its own pods.Limits
- It models involuntary domain loss, where a PodDisruptionBudget offers no protection. A graceful drain is reported separately.
- Dependency edges come only from EndpointSlices and literal environment values that name a Service. ConfigMap and Secret contents are never read, so dependencies configured there are invisible.
- ExternalName and selectorless Services are shown as dependencies but never reported as impairing, because their placement cannot be proven from the cluster.
- Remediation rungs are independent. A single replica in a single zone needs more replicas and an enforced spread together, and each rung alone is correctly reported as partial.
- The scheduler-backed checks target Kubernetes 1.35. On another minor they are disabled with a named warning and the survivability analysis still runs.
Try it
It installs through krew and only reads from the cluster:
kubectl krew install survive-zone
kubectl survive-zoneSource, the exporter and the alert rules are at github.com/SaiPisey2/kubectl-survive.