Kubernetes

Kubernetes - Medium

Page 1 of 4

Advanced Pod Configuration

Resource requests and limits are how you tell the scheduler and the kubelet what a container actually needs, and what it's allowed to consume. A request is what the scheduler uses to decide placement - it will only put a Pod on a node that has at least that much CPU/memory free. A limit is a hard ceiling enforced at runtime: exceed the memory limit and the container is OOMKilled; exceed the CPU limit and it's throttled, not killed.

resources:
  requests:
    cpu: "250m"        # 250 millicores = a quarter of one core
    memory: "256Mi"
  limits:
    cpu: "500m"
    memory: "512Mi"

QoS classes fall out of how you set these, automatically - you don't set a QoS class directly:

  • Guaranteed - every container's request equals its limit, for both CPU and memory. Last to be evicted under node pressure.
  • Burstable - at least one request is set but doesn't equal its limit. Evicted before Guaranteed Pods, after BestEffort.
  • BestEffort - no requests or limits set at all. First to be evicted the moment a node is under memory pressure.

Pod lifecycle hooks (postStart, preStop) run custom logic at container start/stop - a preStop hook is the standard way to drain in-flight requests gracefully before a container actually receives its termination signal, which matters a lot during rolling updates and scale-downs.


Probes & Application Health

Kubernetes can't know if your app is actually healthy just because the process is running - probes are how you tell it.

  • - "is this container still working, or should it be restarted?" A failing liveness probe causes the kubelet to kill and restart the container. Get this wrong (e.g. probing a dependency like the database) and a transient database blip restarts your app for no reason - a liveness probe should only check the process's own health, never an external dependency.
  • - "is this container ready to receive traffic right now?" A failing readiness probe doesn't restart anything - it just pulls the Pod out of the Service's Endpoints until it passes again. This is the correct place to check dependencies (can I reach the database?) - a Pod correctly marked not-ready during startup or a downstream blip simply stops receiving traffic instead of being killed.
  • - exists for slow-starting containers: while it's running, the liveness probe is disabled, so a legitimately slow boot (a JVM app doing 40 seconds of warmup) doesn't get liveness-killed before it ever gets a chance to become healthy.
livenessProbe:
  httpGet:
    path: /healthz
    port: 8080
  initialDelaySeconds: 10
  periodSeconds: 10
readinessProbe:
  httpGet:
    path: /ready
    port: 8080
  periodSeconds: 5

Probe types: HTTP (expects a 2xx/3xx), TCP (just checks the port accepts a connection), exec (runs a command inside the container, success = exit code 0). The single most common probe mistake: using the same endpoint/logic for liveness and readiness - they answer different questions, and conflating them turns a "temporarily can't serve traffic" situation into an unnecessary restart storm.


Scheduling

Beyond nodeSelector's exact-match, Kubernetes has richer ways to express placement rules.

is nodeSelector with more expressive matching (In, NotIn, Exists) and two enforcement levels: requiredDuringSchedulingIgnoredDuringExecution (a hard rule - won't schedule without a match) and preferredDuringSchedulingIgnoredDuringExecution (a soft preference - schedules anyway if nothing matches, just tries to honor it).

Pod affinity/anti-affinity place Pods relative to other Pods, not just nodes - "put this Pod on a node that's already running a Pod with app: cache" (affinity - useful for co-locating a cache-hungry app near its cache), or "never put two Pods with app: web on the same node" (anti-affinity - the standard way to make sure a Deployment's replicas actually spread across failure domains instead of all landing on one node).

Taints and tolerations work in the opposite direction from affinity: a on a node repels Pods by default (kubectl taint nodes node-1 dedicated=gpu:NoSchedule), and only a Pod with a matching in its spec is allowed to schedule there anyway. This is how you get genuinely dedicated nodes (GPU nodes, high-memory nodes) that ordinary workloads simply can't land on by accident - affinity alone only expresses a preference/requirement from the Pod's side, it can't stop unrelated Pods from being scheduled there too.

tolerations:
  - key: "dedicated"
    operator: "Equal"
    value: "gpu"
    effect: "NoSchedule"

generalize anti-affinity into "spread these Pods evenly across zones/nodes" without needing one anti-affinity rule per pair - the right tool once you're reasoning about zone-level failure domains, not just per-node ones.

kubectl describe node node-1 | grep Taints
kubectl get pods -o wide                    # see actual placement across nodes

    Welcome to OpsQuiz!

    Real scenario-based DevOps questions, hands-on practice, and clear explanations for every answer.