Kubernetes

Kubernetes - Advanced

Page 1 of 4

Control Plane Internals

Every write to the cluster follows the same path through the API server: authentication (who are you - client cert, bearer token, OIDC) → authorization (RBAC decides if that identity can perform this verb on this resource) → admission control (mutating webhooks, then validating webhooks/policies) → the object is written to etcd → the change is published to every watch connection subscribed to that resource type. Every controller in the cluster, built-in or custom, is ultimately just a client sitting on one of those watch connections, reacting to changes.

Custom Resource Definitions (CRDs) let you extend the API server with entirely new object kinds that behave exactly like built-in ones - your own kind: Backup or kind: DatabaseCluster, complete with kubectl get, RBAC, and watch support, with zero changes to Kubernetes core. A CRD alone just stores structured data, though - it's inert until a custom controller actually watches it and reconciles real infrastructure to match (see Operators, below).

Conversion webhooks solve a real versioning problem: once a CRD has multiple API versions (v1alpha1, v1beta1, v1), something has to translate objects stored in one version into another when they're read at a different version - a conversion webhook is exactly that translation layer, letting a CRD evolve its schema over time without breaking clients still using an older version.


etcd - Expert

The properties that make etcd trustworthy as a source of truth all trace back to : a write is only acknowledged once it's replicated to and acknowledged by a quorum of members, which is what prevents a split-brain scenario (two partitions of the cluster both believing they're authoritative) - a minority partition simply can't commit writes at all, by construction, not by convention.

Leader election happens automatically on startup or leader failure - and every leader election briefly pauses writes cluster-wide while a new leader is established, which is part of why an etcd cluster experiencing frequent leader elections (often from network flakiness or disk I/O latency spikes) manifests as intermittent, hard-to-pin-down API server slowness rather than an obvious hard failure.

Disaster recovery in practice: a snapshot restore rebuilds etcd's state as of the snapshot moment - anything written after it is gone, which means backup frequency directly bounds your recovery point objective. Restoring is also a full-cluster operation (every etcd member needs to be restored consistently), not something done piecemeal to a single struggling member.

etcdctl endpoint status --cluster -w table    # confirms current leader + raft term across members, first thing to check for election churn
etcdctl alarm list                             # NOSPACE alarms specifically block ALL writes cluster-wide until cleared

Kubernetes Networking - Expert

eBPF-based dataplanes (Cilium being the primary example) replace both traditional CNI routing and kube-proxy's iptables/IPVS rules with programs running directly in the kernel, attached to network hooks - this eliminates the iptables rule-chain traversal that becomes a genuine performance bottleneck at high Service/endpoint counts, and enables richer, identity-aware policy enforcement (allow/deny based on workload identity, not just IP) that pure IP-based NetworkPolicy can't express.

Conntrack (connection tracking) is the Linux kernel subsystem that makes NAT-based Service routing (iptables/IPVS mode) actually work - it remembers which real backend a given connection was mapped to, so return traffic gets un-NAT'd correctly. Conntrack table exhaustion under very high connection churn is a real, specific failure mode: new connections start silently failing once the table fills, and it manifests as intermittent connection resets that look nothing like an obvious "table full" error on the surface.

MTU mismatches are a subtle, classic overlay-networking failure: if the CNI's overlay encapsulation (e.g. VXLAN) isn't accounted for in the effective MTU, larger packets get fragmented or dropped - showing up as connections that work for small payloads and mysteriously stall or hang for larger ones, which is a very different debugging path than "the network is down."


Cluster Lifecycle

Cluster bootstrapping (via kubeadm or a managed provider's equivalent) initializes the control plane, sets up the cluster CA and initial certificates, and generates the tokens worker nodes use to join. Version skew policy is a real, hard constraint on upgrades: the API server can be at most one minor version ahead of kubelets, and control plane components generally must be upgraded in a specific order (API server first, then controller-manager/scheduler, then kubelets) - skipping this order is a common cause of a cluster ending up in a partially-upgraded, inconsistent state.

API deprecations: Kubernetes formally deprecates and eventually removes old API versions on a defined timeline (a deprecated API is typically still served for at least one year / several minor releases before removal) - the real operational risk isn't the deprecation itself, it's discovering at upgrade time that some Helm chart or manifest still targets an API version that's been fully removed, which is why kubectl convert/deprecation scanning before a major upgrade is standard practice, not optional caution.

Cluster migration/replacement (standing up a new cluster and moving workloads over, vs upgrading in place) is a legitimate strategy specifically to avoid the accumulated risk of many sequential in-place upgrades - at the cost of needing a real traffic cutover plan and, if state is involved, a data migration plan too.


    Welcome to OpsQuiz!

    Real scenario-based DevOps questions, hands-on practice, and clear explanations for every answer.