Kubernetes

Kubernetes - Hard

Page 1 of 4

Kubernetes Architecture - Deep Dive

Admission control is the step between "the API server authenticated and authorized this request" and "this object is actually written to etcd." Every write passes through a chain of : mutating ones can modify the object (inject a sidecar, set a default), validating ones can only accept or reject it, and mutating controllers always run before validating ones (so a mutation can't sneak past validation that would have rejected the original). Modern clusters extend this via admission webhooks - your own HTTP service the API server calls out to mid-request - which is how tools like Istio (sidecar injection) and OPA Gatekeeper (policy enforcement) hook into object creation without patching Kubernetes itself.

API aggregation lets additional API groups be served by separate services behind the same API server endpoint, extending the API surface (metrics APIs, custom resources served by their own aggregated API server) without every extension living inside the core API server binary.

Controllers and reconciliation, precisely: a controller runs a loop - watch for changes (via the watch mechanism, a long-lived connection that streams object changes instead of polling), compare desired vs observed state, take action, repeat. Informers are the client-side machinery (used by every built-in controller and most custom ones) that maintain a local cache of watched objects and dispatch change events to handler functions, so a controller doesn't re-list the entire object set on every reconciliation pass - crucial for controllers watching large numbers of objects without hammering the API server.

kubectl get validatingwebhookconfigurations
kubectl get mutatingwebhookconfigurations
kubectl api-resources                          # every resource type the API server currently serves, aggregated APIs included

etcd

is the cluster's single source of truth - every object's current state, full stop. It's a distributed key-value store using the Raft consensus algorithm: one node is elected leader, all writes go through the leader and are only committed once a quorum (a majority of nodes) acknowledges them. This is why etcd clusters are sized in odd numbers (3, 5) - a 3-node cluster tolerates 1 node failure and still has quorum (2 of 3); a 5-node cluster tolerates 2.

Losing quorum is the single most catastrophic failure mode in a Kubernetes cluster - without a quorum, etcd can't accept writes at all, which means the API server can't create, update, or delete anything, even though already-running Pods keep running untouched (the kubelet doesn't need etcd to keep containers alive, only the control plane does).

Backup and restore: etcd supports point-in-time snapshots (etcdctl snapshot save), which is the actual disaster-recovery mechanism for a cluster - restoring a snapshot rebuilds etcd's entire state as of that snapshot, meaning any objects created/changed after it are simply gone. Compaction and defragmentation are separate, regular maintenance: etcd keeps historical revisions of every key (compaction reclaims the space old revisions use), and defragmentation reclaims disk space at the storage-engine level after compaction - skipping both on a long-lived cluster is a real, common cause of etcd eventually running out of disk and refusing writes.

etcdctl snapshot save backup.db
etcdctl snapshot restore backup.db
etcdctl endpoint status --cluster -w table       # leader, DB size, raft term across all members
etcdctl endpoint health --cluster

Advanced Scheduling

The scheduler's decision for each Pod happens in two phases: filtering (which nodes are even possible - enough resources, satisfies nodeSelector/affinity/taints) and scoring (among the possible nodes, which is best - spreading, resource balance, affinity preferences). Understanding this split explains a lot of scheduling behavior that otherwise looks arbitrary: a Pod landing on a "less full" node isn't random, it's the scoring phase's resource-balance priority doing its job.

Preemption: when a higher-PriorityClass Pod can't be scheduled anywhere because the cluster is full, the scheduler can evict lower-priority Pods to make room, provided their priority is genuinely lower. This is why setting realistic PriorityClasses matters in a resource-constrained cluster - without them, every Pod has equal claim and nothing yields room for anything else.

, precisely: rather than one anti-affinity rule per pair of Pods, you declare a topologyKey (e.g. topology.kubernetes.io/zone) and a maxSkew (how uneven the distribution is allowed to get), and the scheduler actively works to keep replicas balanced across that topology domain - the standard mechanism for genuine multi-AZ resilience, since anti-affinity alone doesn't reason about how evenly spread something is, only whether two specific Pods can coexist.

topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: DoNotSchedule
    labelSelector:
      matchLabels:
        app: web

    Welcome to OpsQuiz!

    Real scenario-based DevOps questions, hands-on practice, and clear explanations for every answer.