advancedProduction: HA, DR, Namespace Strategy & Cluster Hardening
An AZ outage takes out 1 of 3 control plane nodes -- what happens?
With 3 control plane nodes and 1 lost, etcd still has a quorum of 2. The cluster continues operating normally -- scheduling, failure recovery, and kubectl commands all work. The remaining 2 control plane nodes handle all API requests. Worker nodes in the affected AZ are lost (nodes become NotReady), their Pods are evicted and rescheduled to healthy nodes in surviving AZs.
This is a Pro chapter
Sign in, then upgrade to Pro or Power to unlock this and the full DevOps Mastery library.
An AZ outage takes out 1 of 3 control plane nodes -- what happens?