Scenario rounds hand you a broken cluster in two sentences and watch how you think. A Java pod keeps getting OOMKilled, every deploy throws a burst of 502s, a drain hangs for half an hour, a namespace won't finish deleting. There is rarely one right answer. The interviewer wants the order you check things in, what you would look at first, and what would change your mind. This page is for anyone facing that round, from a first platform job to a senior hire. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying your first three checks before the fix.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Read the event: kubectl describe pod says whether it's not found, unauthorized or a timeout.
Not found: compare the image path and tag with what's really in the registry.
Unauthorized: a docker-registry Secret in the pod's own namespace, listed in imagePullSecrets or on the ServiceAccount.
Timeout: the nodes can't reach the registry; test from a node.
“Old pods keep running because they pulled their image before the move and never need to pull again, so the problem is only the pull. First I'd run kubectl describe on a stuck pod and read the events at the bottom, because the error splits three ways. If it says not found or manifest unknown, it's the image name or tag, so I'd compare it with what's actually in the registry. If it says unauthorized or access denied, the pod has no working credentials. I'd check there's a docker-registry Secret in the same namespace as the pod, and that it's listed under imagePullSecrets or attached to the pod's ServiceAccount. A Secret in another namespace doesn't count. If it's a timeout, it's the network, so I'd test from a node whether it can reach the registry at all. If only some nodes fail, that points at those nodes' network or setup.”
Deleting and recreating pods over and over without reading the event that names the cause.
What happened: the node's disk crossed an eviction threshold; the kubelet removed unused images, then evicted pods.
Find the filler: container logs, emptyDir volumes, writes into the container layer, or piles of images.
Clean up: evicted pods stay behind as Failed objects until deleted.
Prevent: ephemeral-storage requests and limits, emptyDir sizeLimit, log rotation and a node disk alert.
“DiskPressure means the node's disk crossed the kubelet's eviction threshold. The kubelet first tries to free space by removing unused images, and if that isn't enough it starts evicting pods, starting with the ones using more than they asked for. Their controllers recreate them on other nodes, so the service may have coped, but the node is still sick. I'd run kubectl describe node to see the condition and when it started, then look on the node itself for what's big: container logs, emptyDir volumes, an app writing files inside its own container, or old images. Often it's one pod writing debug logs or temp files. Evicted pods stay behind as Failed objects, so I'd clean those up too. To stop it coming back, I'd give pods ephemeral-storage requests and limits, set a sizeLimit on emptyDir, make sure logs rotate, and alert on node disk well before the threshold.”
Deleting the evicted pods and moving on without finding what filled the disk.
Ask the namespace: its status conditions list the resources and finalizers still remaining.
Usual causes: an object whose controller was uninstalled keeps its finalizer, or an aggregated API is down.
Fix the cause: bring the controller back, or remove that one object's finalizer once you know what it guarded.
Last resort: forcing the namespace itself can leave disks and load balancers behind.
“A namespace deletes everything inside it first, and stays Terminating until that's done. I'd start with kubectl get namespace with -o yaml and read the status conditions, because they usually list exactly which kinds of resources are left and which finalizers they're waiting on. Two causes come up again and again. One is a custom resource from an operator that was uninstalled before the namespace was deleted, so its finalizer is waiting for a controller that no longer exists. The other is an aggregated API, like a metrics service, that's unavailable, so the namespace controller can't even list everything. I'd fix the cause: reinstall the controller long enough to clean up, repair the broken API service, or remove the finalizer on that one object after checking what it was meant to clean up. Forcing the namespace through can leave cloud disks or load balancers running with nothing tracking them.”
Removing the namespace's finalizer straight away with no idea what it leaves behind.
Events first: kubectl describe pvc usually names the reason.
StorageClass: the class the claim names exists, or there's a default when it names none.
Binding mode: WaitForFirstConsumer keeps a claim Pending until a pod using it is scheduled.
Provisioner: the CSI driver for that class is installed, healthy and supports the size and access mode.
“I'd start with kubectl describe on the claim, because the events normally say what's wrong. First I check the StorageClass. If the claim names a class that doesn't exist, or names none and the cluster has no default, nothing will ever create a disk. Next I'd look at the class's binding mode. With WaitForFirstConsumer the claim stays Pending on purpose until a pod that uses it is scheduled, so if the pod is Pending for another reason, like not enough CPU, the claim waits too and the real problem is the pod. Then the provisioner: is the CSI driver for that class installed, are its pods running, and do their logs show errors like a quota or an access mode the backend can't do, such as ReadWriteMany on a block disk. If there's no provisioner at all, someone has to create a matching volume by hand.”
Assuming a Pending claim is always broken, without checking whether it is waiting for its pod by design.
Why: nobody can confirm the old pod stopped, and a StatefulSet never runs two pods with one identity.
Confirm dead: make sure the machine is truly off, from the cloud console or the hardware team.
Release it: delete the node object, force delete the pod, or use the out-of-service taint so the volume detaches.
Placement: a zonal disk means the new pod must land in the same zone.
“This is the StatefulSet being careful. The node stopped reporting, so its pod shows Terminating or Unknown, but nobody can prove the old container stopped. A StatefulSet allows at most one pod per identity, so it won't start db-0 somewhere else while the old db-0 might still be writing to the disk. First I'd confirm the machine is really off, from the cloud console or whoever runs the hardware, because if it's only cut off from the network, a second copy could corrupt data. Once I'm sure, I'd delete the node object or force delete the pod. Newer versions also have an out-of-service taint for exactly this case. Then the volume can detach and the pod comes up elsewhere. If the disk is zonal, it must land in the same zone, so I'd check there's room there. A Multi-Attach error means the volume hasn't detached yet.”
Force deleting the pod straight away without confirming the old node is really down.
Restore: put the fix into the repo and ship it through the pipeline, or roll back as a stopgap.
Why: the pipeline applies what's in Git, so a live edit lasts only until the next deploy.
Prevent: an emergency path that still ends in Git, and less write access in production.
“First the bug. I'd find exactly what was changed last night, from the person who did it or by comparing the old ReplicaSet's spec with the new one, put that change into the repo and push it through the pipeline. If that's slow and the edit was to the pod template, I'd use kubectl rollout undo to get back to last night's version as a stopgap, then follow with the proper commit. Then the why, which is simple: the pipeline applies whatever is in Git, so anything edited by hand lasts until the next deploy. Nobody did anything wrong at midnight, the process just had no emergency path. So I'd propose one: urgent live fixes are fine, but the same person opens a pull request straight after. Longer term I'd take production write access away from most people and add a tool that flags drift between Git and the cluster.”
Blaming the person who made the hotfix, or re-applying the edit by hand again and leaving the repo wrong.
Look: helm history shows the newest revision stuck in a pending state.
Check the cluster: see what the half-finished upgrade actually changed.
Recover: roll back to the last deployed revision, then run the pipeline again.
Prevent: never let CI kill Helm mid-run; give it a timeout and automatic rollback on failure.
“Helm keeps the state of each release in the cluster, by default as Secrets, one per revision. The cancelled job left the newest revision marked pending-upgrade, and Helm refuses to start another operation on top of it. I'd run helm history on the release to confirm that, and note the last revision marked deployed. Then I'd check what's actually running, because a half-finished upgrade may have changed some resources and not others. The clean fix is helm rollback to that last deployed revision. That puts the resources back and records a new, finished revision. Then I'd rerun the normal pipeline. I'd avoid deleting Helm's release Secrets by hand unless rollback fails, and even then I'd back them up first. To stop it recurring, I'd give the CI step a longer timeout than Helm's own, and turn on the option that rolls a failed upgrade back automatically.”
Uninstalling and reinstalling the release in production just to clear the error.
Simple way: a second Deployment with the new image that carries the label the Service selects.
The split: traffic follows pod count, so one canary pod beside nineteen stable ones is about one in twenty.
Finer control: an ingress controller that supports weighted canary routing.
Decide: compare errors and latency by version, then promote or delete the canary.
“The plainest way uses only a Service and two Deployments. The Service selects on a label like app: checkout. The stable Deployment has that label and nineteen replicas, and a canary Deployment with the new image has the same label and one replica. Since the Service spreads connections across all ready pods, roughly one in twenty requests hit the canary. It's rough, because it splits by pod and by connection, not exactly, and I'd have to scale the stable side to hold the ratio. If our ingress controller supports weighted canary routing, I'd use that instead, because it splits by request and doesn't tie the ratio to replica counts. Either way I'd add a version label to metrics so I can compare errors and latency by version. Before starting I'd check both versions can share the database schema and sessions, since users will bounce between them.”
Running the canary with no way to see its metrics separately from the stable version.
Restore: re-run the last good pipeline or apply the manifests from Git.
Check scope: a deleted Deployment takes its ReplicaSets and pods; Services, ConfigMaps and claims usually survive.
Tell people: one clear message in the incident channel while you work.
Prevent: separate kubeconfigs, the context in the shell prompt, and no delete rights in production for daily work.
“First I'd get the app back, and the fastest safe way is to redeploy from source: re-run the last good pipeline, or apply the manifests straight from Git. While that runs I'd check what else went. Deleting a Deployment removes its ReplicaSets and pods, but Services, ConfigMaps, Secrets and volume claims normally stay, so data should be safe. I'd post one short message in the incident channel so people know it's being handled and don't pile in. Once it's back, I'd write it up without blame. Anyone can type into the wrong terminal window. The fixes are in the setup: a separate kubeconfig file for production that you have to pick on purpose, the current context shown in every shell prompt, and day-to-day production access that doesn't include delete.”
Hunting for who to blame while the app is still down, or rebuilding the Deployment by hand from memory.
Assume leaked: it's in history, every clone and maybe CI logs; deleting the line doesn't un-leak it.
Rotate first: a new password, rolled out to the app, then the old one disabled.
Clean up: the chart reads an existing Secret; rewriting history is a secondary step.
Prevent: a secret manager synced into the cluster, plus secret scanning in the pipeline.
“I'd treat it as leaked, because it has been for three months. Anyone with repo access, every clone and possibly the CI logs have it, so deleting the line fixes nothing on its own. The first job is rotation. I'd tell the database owner and the security team, create a new password, put it into a Kubernetes Secret that's managed outside Git, roll the app so it picks it up, check it's healthy, then disable the old password. I'd also look through the database logs for logins that don't look like ours. Then I'd change the chart so it reads the password from an existing Secret instead of from values. Rewriting Git history is a nice extra, but it can't take back copies people already have. To stop it recurring, I'd add secret scanning to the pipeline and move secrets into a proper secret manager synced into the cluster.”
Deleting the line from the file and considering the problem solved.
Why: the scheduler prefers spreading, but only as a preference; short on room, pods can land together.
Fix: topology spread constraints on hostname, and on zone if the cluster spans zones.
Backstop: a PodDisruptionBudget so a drain can't take them all at once.
“Three replicas only help if they're in different places, and the scheduler's default spreading is a scoring preference, not a rule. If the other nodes were short on room when the pods were created, all three can end up together. I'd check placement now with kubectl get pods -o wide. Then I'd add a topology spread constraint on the hostname label with a max skew of one, so replicas stay even across nodes, plus one on zone if the cluster spans zones. I'd choose between a hard rule and a soft one. Hard means a pod stays Pending rather than doubling up, which is safer but needs spare capacity. Existing pods don't move by themselves, so I'd restart the rollout to place them again. I'd also add a PodDisruptionBudget, so a node drain can't evict every replica at once.”
# inside the Deployment's pod template
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: checkout
Saying three replicas are enough for availability without checking where they actually run.
Why: a pod succeeds only when all its containers have exited; the sidecar never does.
Native fix: newer versions let a sidecar be an init container with restartPolicy Always, stopped after the main one ends.
Older clusters: the main container tells the sidecar to quit through a shared file or an endpoint.
Safety net: activeDeadlineSeconds so a stuck Job can't run forever.
“A Job counts a pod as done only when the pod succeeds, and a pod only succeeds when all of its containers have exited. The log shipper is built to run forever, so the pod just sits there after the real work is finished. On a newer cluster the clean fix is a native sidecar: I'd declare the shipper as an init container with restartPolicy set to Always. It starts before the main container, keeps running beside it, and Kubernetes stops it once the main container is done. On an older cluster I'd make the main container tell the sidecar to stop, for example by writing a file to a shared emptyDir that the sidecar watches, or calling a quit endpoint if it has one. Either way I'd set activeDeadlineSeconds on the Job, so if something gets stuck again it fails loudly instead of running for days.”
spec:
activeDeadlineSeconds: 3600
template:
spec:
restartPolicy: Never
initContainers:
- name: log-shipper
image: registry.example.com/log-shipper:2.1
restartPolicy: Always # native sidecar
containers:
- name: report
image: registry.example.com/report:1.4
Deleting the finished pods by hand every night instead of fixing why the Job can't complete.
Confirm: last state OOMKilled with exit code 137 in kubectl describe pod, not a node eviction.
Why: the limit counts everything: heap, metaspace, thread stacks, direct buffers, code cache, GC structures.
Measure: native memory tracking in a test pod, and the container's memory graph over time.
Fix: size the heap as a share of the limit with real headroom, or raise the limit; hunt a native leak if it keeps climbing.
“First I'd confirm it's really the container limit. kubectl describe pod should show the last state as OOMKilled with exit code 137, which means the kernel killed it for going over its limit, not the kubelet evicting it for node pressure. Then, the heap setting is only part of the story. The JVM also uses metaspace for classes, a stack for every thread, direct buffers for I/O, the code cache and memory for the garbage collector itself. A service with hundreds of threads and heavy networking can use a lot outside the heap. I'd turn on native memory tracking in a test pod to see where it goes. Usually the fix is to size the heap as a share of the limit, leaving real headroom, or to raise the limit. If memory climbs steadily for hours instead, I'd suspect a native leak and dig into that before just giving it more room.”
Raising the heap size, which makes the container go over its limit even sooner.
How limits work: a CPU quota per short period; a busy multi-threaded app can use it up early and then wait.
Check: the throttled periods counters from cAdvisor, or cpu.stat in the container's cgroup.
Fix: raise or remove the CPU limit while keeping the request, and match thread counts to the CPU the container really gets.
“A CPU limit isn't enforced as an average. The kernel gives the container a quota for each short period, a tenth of a second by default. If the app runs many threads, they can burn the whole quota early in a period and then sit frozen until the next one starts. The graph averages over a minute and looks calm, but individual requests hit those pauses, so the slow end gets worse. To check, I'd look at the throttled periods counter that cAdvisor exposes, or read cpu.stat inside the container, and compare throttled periods with total periods. If a large share is throttled, that's the cause. The fixes are raising the limit, or removing the CPU limit and keeping a sensible request, which many teams do for latency-sensitive services. I'd also check the runtime isn't sizing its thread pools from the node's cores instead of what the container gets.”
Reading the average CPU graph, seeing headroom, and ruling throttling out.
Did pods arrive: new replicas may be Pending with no room, waiting on the node autoscaler.
Are they serving: slow start or a strict readiness probe keeps them out of the Service.
Real bottleneck: a shared database, connection pool or downstream API that more pods make worse.
Right signal: CPU burned on retries can drive scaling that only adds load.
“More replicas only help if they're running and the pods are the bottleneck. First I'd check the new replicas actually came up. kubectl get pods might show several Pending because the nodes are full and the node autoscaler is slow or at its own cap. Then I'd check the running ones are ready and in the Service's endpoints. If all that's fine, the bottleneck is probably behind the pods. The usual one is the database: every new pod opens its own connection pool, so scaling out can exhaust connections and make things slower, not faster. I'd look at database load and connection counts, and the latency of any downstream APIs. I'd also ask whether CPU is the right signal. If pods burn CPU retrying against a struggling database, the autoscaler keeps adding pods that only add more load.”
Raising the maximum replica count and hoping, without finding where the time goes.
Why: scheduling adds up requests; if requests sit far above real use, nodes are full on paper.
Check: compare Allocated resources in kubectl describe node with kubectl top.
Right-size: set requests from observed usage plus headroom, biggest gaps first.
Keep it right: vertical autoscaler recommendations and a regular review.
“The scheduler never looks at how busy a node really is. It adds up the CPU requests of the pods already there, and if the new pod's request doesn't fit in what's left, it stays Pending. So if teams asked for two cores and use a fifth of one, nodes look full while sitting idle, and the node autoscaler keeps buying more. I'd confirm it by comparing the allocated section in kubectl describe node with real usage from kubectl top or our monitoring. Then I'd rank services by the gap between request and real use and start with the biggest. For each one I'd set the request near its busy-hour usage plus some headroom, roll it out, and watch latency and throttling. A vertical pod autoscaler in recommendation mode helps keep the numbers honest over time. I'd share the before and after with the teams so they see why it matters.”
Adding more nodes, or cutting every request by the same amount, without looking at real usage per service.
Clue: restarts across every pod at once point at the liveness probe.
Cause: the liveness endpoint checked the database, so the kubelet killed healthy processes.
Why longer: all pods restart together, warm up slowly and hit restart back-off.
Fix: liveness checks only the process; the app rides out dependency blips, and readiness is used with care.
“Restarts across every pod at the same moment make me look at the liveness probe first. I'd bet the health endpoint it calls also checks the database. When the database failed over, every pod's liveness check failed, so the kubelet killed perfectly healthy processes all at once. Then they all restarted together, spent time warming up, and some hit the restart back-off, so thirty seconds of database trouble became minutes of outage. I'd confirm it from the pod events and the probe endpoint's code. The fix is that liveness answers one question: is this process stuck? It shouldn't touch the database at all. For dependencies, the app should retry and return clear errors while the database is away. I'd be careful moving the check into readiness too, because if every pod goes unready at once the Service has no endpoints, which is also an outage.”
Raising the probe timeout and leaving the database check inside the liveness endpoint.
New pods: confirm readiness really means ready, and the rollout never drops below capacity.
Old pods: SIGTERM and removal from endpoints happen in parallel, so traffic still arrives after shutdown starts.
Fix: a short preStop sleep, SIGTERM handling that drains in-flight requests, and a grace period longer than both.
“Readiness probes protect the new pods, so I'd suspect the old ones. When Kubernetes terminates a pod, it sends SIGTERM to the container and removes the pod from the Service's endpoints at the same time, not one after the other. kube-proxy and the ingress controller take a moment to notice, so for a second or two they still send requests to a pod that has started shutting down. If the app exits the moment it gets SIGTERM, those requests fail as 502s. The fix has three parts. A preStop hook that just sleeps a few seconds, so the pod keeps serving while everyone stops routing to it. The app handles SIGTERM by finishing in-flight requests before it exits. And a termination grace period long enough to cover both. I'd still check the new pods too, because a readiness probe that passes before the app is warm gives the same symptom.”
# pod template: keep serving while routing updates, then drain
spec:
terminationGracePeriodSeconds: 45
containers:
- name: api
image: registry.example.com/api:3.2
lifecycle:
preStop:
exec:
command: ["sleep", "10"]
Blaming the readiness probe on the new pods and never looking at how old pods shut down.
The clue: a fixed five seconds is the resolver's default timeout, so a DNS query was lost.
Extra lookups: ndots 5 and the search list turn one outside name into several queries.
Lost packets: parallel A and AAAA queries over UDP can collide in connection tracking and one is dropped.
Fix: trailing-dot names or a lower ndots, a node-local DNS cache, resolver options that avoid the race.
“Exactly five seconds is a big clue, because that's the default time a Linux resolver waits before retrying a DNS query. So I'd suspect a lost DNS packet, not the API. I'd look at the pod's resolv.conf. Pods usually get ndots set to five and a list of cluster search domains, so a name like api.example.com, with only two dots, is first tried with each cluster suffix added. That's several extra queries per call, and each is a chance to lose one. On top of that, the resolver often sends the A and AAAA queries at the same moment from one socket, and on some kernels those collide in connection tracking and one gets dropped. To fix it I'd use a trailing dot on outside names or lower ndots in the pod's dnsConfig, and run a node-local DNS cache. A small loop of timed lookups inside a pod would prove it before and after.”
Blaming the outside API or adding retries in the app without noticing the fixed five-second pattern.
Why: denying all egress also blocks port 53 to CoreDNS in kube-system, so no name resolves.
Confirm: lookups time out rather than fail fast, and the policy has no egress rule for DNS.
Fix: allow UDP and TCP 53 to the DNS pods only, then add each real dependency one by one.
“Even a service in the same namespace is found by name, and that name is answered by CoreDNS, which runs in kube-system. A default-deny policy that covers egress blocks that traffic like anything else, so every lookup times out. I'd confirm it from a pod: nslookup of a service name hangs instead of failing quickly, and reading the policy shows egress is denied with nothing allowed for DNS. The fix isn't to drop the policy. I'd add a small egress rule allowing UDP and TCP port 53 to the pods labelled as kube-dns in the kube-system namespace, and nothing else. Then I'd work through each app's real dependencies, like its database and the other services it calls, and allow those one at a time. Next time I'd roll default-deny out to one namespace in a test environment first, with the DNS rule already in place.”
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-dns
spec:
podSelector: {}
policyTypes: ["Egress"]
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
podSelector:
matchLabels:
k8s-app: kube-dns
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
Deleting the policy to make the errors go away.
Detection: the node stops renewing its heartbeat, is marked NotReady or Unknown, and its pods are marked not ready.
The wait: pods tolerate the not-ready and unreachable taints for five minutes by default, then are evicted.
Trade-off: a shorter tolerationSeconds replaces pods faster but also reacts to brief network blips.
Better answer: enough replicas spread across nodes that five minutes costs capacity, not availability.
“That's the default behaviour, not a fault. When a node stops renewing its heartbeat, the control plane marks it NotReady or Unknown after a short grace period and marks its pods not ready, which is why traffic left quickly. Then it taints the node, and every pod by default tolerates those taints for five minutes. Only after that are the pods evicted and their Deployments create replacements elsewhere. The wait is on purpose: a node that drops off the network for a minute shouldn't cause every pod on it to be rebuilt. We can shorten tolerationSeconds for the not-ready and unreachable taints on services that really care, but I'd be careful, because a flaky network then causes churn. I'd rather make sure each service has enough replicas spread across nodes, so losing one node's pods for five minutes costs some capacity but never availability.”
Saying the cluster is misconfigured without knowing the default five-minute tolerations exist.
Read the message: drain says the eviction would violate the pod's disruption budget.
Check the budget: kubectl get pdb shows allowed disruptions at zero, often one replica with minAvailable 1.
Fix properly: scale up first, wait for a healthy spare, then let the drain finish.
Avoid: force deleting the pod, which skips the protection the budget exists for.
“Drain doesn't just delete pods, it asks the eviction API, and that honours PodDisruptionBudgets. If a budget says no disruption is allowed right now, drain keeps retrying, and the message usually says so. I'd run kubectl get pdb in that namespace and look at allowed disruptions. The classic case is a Deployment with one replica and a budget saying minAvailable one. That pod can never be evicted, so drain waits forever. Another cause is a budget whose other pods are already unhealthy, so there's no slack to give. The right fix is to scale the Deployment to two, wait for the new pod to be ready on another node, let the drain finish, then scale back if needed. I'd tell the owning team, because a one-replica service with a strict budget is a design problem. I wouldn't force delete the pod, since that skips the very protection they asked for.”
Force deleting the pod to get past the drain without telling anyone.
Show the impact: node graphs and the incidents their jobs caused, without blame.
Defaults: a LimitRange gives every container sensible requests even when none are written.
Fair share: a ResourceQuota per namespace, and a lower priority class or separate nodes for batch work.
Make it easy: templates or chart defaults so setting requests takes no effort.
“I'd start with a conversation, not a policy. I'd show them the node graphs from when their jobs ran and the incidents other teams had, and explain that pods with no requests are scheduled as if they need nothing, so they pile onto nodes and then fight everyone for CPU and memory. Then I'd make the right thing effortless. A LimitRange in their namespace gives every container default requests and limits even if nobody writes them, so they barely have to change anything. A ResourceQuota caps the namespace so one team can't take the whole cluster. For the batch jobs I'd suggest a lower priority class or a separate node pool, so they use spare capacity without pushing out services. I'd agree the default numbers with them and review after a couple of weeks of real usage, so the limits feel fair rather than imposed.”
Silently applying a strict quota that breaks their next deploy without talking to them.
Removed APIs: scan manifests, charts and live objects for API versions the target drops.
Order: one minor version at a time, control plane first, then nodes; kubelets never newer than the API server.
Rehearse: a staging cluster first, and a backup of etcd or the cluster before production.
Nodes and add-ons: drain in batches respecting budgets, and confirm add-ons support the new version.
“The biggest risk isn't the upgrade button, it's old manifests. Releases can remove API versions that were deprecated earlier, and a manifest using one simply stops applying. So first I'd read the release notes for both target versions and scan our Git manifests, Helm charts and the live cluster for removed API versions, then fix those while still on the old version, since the newer API versions usually exist there already. For the upgrade itself, I'd go one minor version at a time, control plane first, then nodes, because kubelets can lag behind the API server but must never be newer. I'd rehearse on a staging cluster built the same way, and back up etcd or the cluster before touching production. I'd drain nodes in small batches, respecting disruption budgets, and check workloads after each batch. I'd also confirm add-ons like the ingress controller and network plugin support the new version.”
Jumping straight to the newest version in production without checking for removed APIs.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.