Page 1 of 4
Resource requests and limits are how you tell the scheduler and the kubelet what a container actually needs, and what it's allowed to consume. A request is what the scheduler uses to decide placement - it will only put a Pod on a node that has at least that much CPU/memory free. A limit is a hard ceiling enforced at runtime: exceed the memory limit and the container is OOMKilled; exceed the CPU limit and it's throttled, not killed.
resources:
requests:
cpu: "250m" # 250 millicores = a quarter of one core
memory: "256Mi"
limits:
cpu: "500m"
memory: "512Mi"
QoS classes fall out of how you set these, automatically - you don't set a QoS class directly:
Pod lifecycle hooks (postStart, preStop) run custom logic at container start/stop - a preStop hook is the standard way to drain in-flight requests gracefully before a container actually receives its termination signal, which matters a lot during rolling updates and scale-downs.
Kubernetes can't know if your app is actually healthy just because the process is running - probes are how you tell it.
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 10
readinessProbe:
httpGet:
path: /ready
port: 8080
periodSeconds: 5
Probe types: HTTP (expects a 2xx/3xx), TCP (just checks the port accepts a connection), exec (runs a command inside the container, success = exit code 0). The single most common probe mistake: using the same endpoint/logic for liveness and readiness - they answer different questions, and conflating them turns a "temporarily can't serve traffic" situation into an unnecessary restart storm.
Beyond nodeSelector's exact-match, Kubernetes has richer ways to express placement rules.
Node affinity is nodeSelector with more expressive matching (In, NotIn, Exists) and two enforcement levels: requiredDuringSchedulingIgnoredDuringExecution (a hard rule - won't schedule without a match) and preferredDuringSchedulingIgnoredDuringExecution (a soft preference - schedules anyway if nothing matches, just tries to honor it).
Pod affinity/anti-affinity place Pods relative to other Pods, not just nodes - "put this Pod on a node that's already running a Pod with app: cache" (affinity - useful for co-locating a cache-hungry app near its cache), or "never put two Pods with app: web on the same node" (anti-affinity - the standard way to make sure a Deployment's replicas actually spread across failure domains instead of all landing on one node).
Taints and tolerations work in the opposite direction from affinity: a taint on a node repels Pods by default (kubectl taint nodes node-1 dedicated=gpu:NoSchedule), and only a Pod with a matching toleration in its spec is allowed to schedule there anyway. This is how you get genuinely dedicated nodes (GPU nodes, high-memory nodes) that ordinary workloads simply can't land on by accident - affinity alone only expresses a preference/requirement from the Pod's side, it can't stop unrelated Pods from being scheduled there too.
tolerations:
- key: "dedicated"
operator: "Equal"
value: "gpu"
effect: "NoSchedule"
Topology spread constraints generalize anti-affinity into "spread these Pods evenly across zones/nodes" without needing one anti-affinity rule per pair - the right tool once you're reasoning about zone-level failure domains, not just per-node ones.
kubectl describe node node-1 | grep Taints
kubectl get pods -o wide # see actual placement across nodes
Real scenario-based DevOps questions, hands-on practice, and clear explanations for every answer.