Troubleshooting • Scaling • Networking • Storage • Upgrades • Helm • Security • 2026

Scenario-Based Kubernetes Interview Questions

Scenario rounds hand you a broken cluster in two sentences and watch how you think. A Java pod keeps getting OOMKilled, every deploy throws a burst of 502s, a drain hangs for half an hour, a namespace won't finish deleting. There is rarely one right answer. The interviewer wants the order you check things in, what you would look at first, and what would change your mind. This page is for anyone facing that round, from a first platform job to a senior hire. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying your first three checks before the fix.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Troubleshooting 3 questions

Easy Technical round Fresher, Mid-level Practice question

1. You moved your images to a private registry, and now new pods sit in ImagePullBackOff while old ones keep running. What do you check?

What the interviewer is really testing:
Whether you read the pull error before guessing, and know how a pod gets registry credentials.
Answer frame:

Read the event: kubectl describe pod says whether it's not found, unauthorized or a timeout.

Not found: compare the image path and tag with what's really in the registry.

Unauthorized: a docker-registry Secret in the pod's own namespace, listed in imagePullSecrets or on the ServiceAccount.

Timeout: the nodes can't reach the registry; test from a node.

Sample spoken answer:

“Old pods keep running because they pulled their image before the move and never need to pull again, so the problem is only the pull. First I'd run kubectl describe on a stuck pod and read the events at the bottom, because the error splits three ways. If it says not found or manifest unknown, it's the image name or tag, so I'd compare it with what's actually in the registry. If it says unauthorized or access denied, the pod has no working credentials. I'd check there's a docker-registry Secret in the same namespace as the pod, and that it's listed under imagePullSecrets or attached to the pod's ServiceAccount. A Secret in another namespace doesn't count. If it's a timeout, it's the network, so I'd test from a node whether it can reach the registry at all. If only some nodes fail, that points at those nodes' network or setup.”

Red flag to avoid:

Deleting and recreating pods over and over without reading the event that names the cause.

They may ask next:
  • How would you give every pod in a namespace pull access without editing each Deployment?
  • What does imagePullPolicy Always change about this failure?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

2. Dozens of pods on one node suddenly show Evicted, and the node reports DiskPressure. What happened, and how do you stop it happening again?

What the interviewer is really testing:
Whether you know the kubelet evicts pods to protect a node's disk, and can find what actually filled it.
Answer frame:

What happened: the node's disk crossed an eviction threshold; the kubelet removed unused images, then evicted pods.

Find the filler: container logs, emptyDir volumes, writes into the container layer, or piles of images.

Clean up: evicted pods stay behind as Failed objects until deleted.

Prevent: ephemeral-storage requests and limits, emptyDir sizeLimit, log rotation and a node disk alert.

Sample spoken answer:

“DiskPressure means the node's disk crossed the kubelet's eviction threshold. The kubelet first tries to free space by removing unused images, and if that isn't enough it starts evicting pods, starting with the ones using more than they asked for. Their controllers recreate them on other nodes, so the service may have coped, but the node is still sick. I'd run kubectl describe node to see the condition and when it started, then look on the node itself for what's big: container logs, emptyDir volumes, an app writing files inside its own container, or old images. Often it's one pod writing debug logs or temp files. Evicted pods stay behind as Failed objects, so I'd clean those up too. To stop it coming back, I'd give pods ephemeral-storage requests and limits, set a sizeLimit on emptyDir, make sure logs rotate, and alert on node disk well before the threshold.”

Red flag to avoid:

Deleting the evicted pods and moving on without finding what filled the disk.

They may ask next:
  • How would you find which pod wrote the most to disk on that node?
  • Why is writing into the container's own filesystem a bad habit here?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

3. You deleted an old namespace yesterday and it's still stuck in Terminating. How do you find what's holding it and clear it safely?

What the interviewer is really testing:
Whether you know namespaces wait on their contents and finalizers, and avoid the shortcut that leaves cloud resources orphaned.
Answer frame:

Ask the namespace: its status conditions list the resources and finalizers still remaining.

Usual causes: an object whose controller was uninstalled keeps its finalizer, or an aggregated API is down.

Fix the cause: bring the controller back, or remove that one object's finalizer once you know what it guarded.

Last resort: forcing the namespace itself can leave disks and load balancers behind.

Sample spoken answer:

“A namespace deletes everything inside it first, and stays Terminating until that's done. I'd start with kubectl get namespace with -o yaml and read the status conditions, because they usually list exactly which kinds of resources are left and which finalizers they're waiting on. Two causes come up again and again. One is a custom resource from an operator that was uninstalled before the namespace was deleted, so its finalizer is waiting for a controller that no longer exists. The other is an aggregated API, like a metrics service, that's unavailable, so the namespace controller can't even list everything. I'd fix the cause: reinstall the controller long enough to clean up, repair the broken API service, or remove the finalizer on that one object after checking what it was meant to clean up. Forcing the namespace through can leave cloud disks or load balancers running with nothing tracking them.”

Red flag to avoid:

Removing the namespace's finalizer straight away with no idea what it leaves behind.

They may ask next:
  • What is a finalizer, and why does Kubernetes use them?
  • How would you stop this happening the next time you uninstall an operator?
Say it in 60 seconds

Storage & State 2 questions

Easy Technical round Fresher, Mid-level Practice question

4. A new PersistentVolumeClaim has been Pending for twenty minutes and the pod that needs it won't start. How do you work out why?

What the interviewer is really testing:
Whether you know how dynamic provisioning works and can tell a real failure from a claim that is simply waiting for its pod.
Answer frame:

Events first: kubectl describe pvc usually names the reason.

StorageClass: the class the claim names exists, or there's a default when it names none.

Binding mode: WaitForFirstConsumer keeps a claim Pending until a pod using it is scheduled.

Provisioner: the CSI driver for that class is installed, healthy and supports the size and access mode.

Sample spoken answer:

“I'd start with kubectl describe on the claim, because the events normally say what's wrong. First I check the StorageClass. If the claim names a class that doesn't exist, or names none and the cluster has no default, nothing will ever create a disk. Next I'd look at the class's binding mode. With WaitForFirstConsumer the claim stays Pending on purpose until a pod that uses it is scheduled, so if the pod is Pending for another reason, like not enough CPU, the claim waits too and the real problem is the pod. Then the provisioner: is the CSI driver for that class installed, are its pods running, and do their logs show errors like a quota or an access mode the backend can't do, such as ReadWriteMany on a block disk. If there's no provisioner at all, someone has to create a matching volume by hand.”

Red flag to avoid:

Assuming a Pending claim is always broken, without checking whether it is waiting for its pod by design.

They may ask next:
  • Why might a zonal disk make the pod unschedulable after the claim binds?
  • What happens to the disk when you delete the claim, and how do you control that?
Say it in 60 seconds
Hard Technical round Senior Practice question

5. A node running your StatefulSet's database pod lost power. Twenty minutes later the pod still hasn't come back on another node. Why, and what do you do?

What the interviewer is really testing:
Whether you know a StatefulSet refuses to risk two copies of one pod, and how to recover without two writers on one disk.
Answer frame:

Why: nobody can confirm the old pod stopped, and a StatefulSet never runs two pods with one identity.

Confirm dead: make sure the machine is truly off, from the cloud console or the hardware team.

Release it: delete the node object, force delete the pod, or use the out-of-service taint so the volume detaches.

Placement: a zonal disk means the new pod must land in the same zone.

Sample spoken answer:

“This is the StatefulSet being careful. The node stopped reporting, so its pod shows Terminating or Unknown, but nobody can prove the old container stopped. A StatefulSet allows at most one pod per identity, so it won't start db-0 somewhere else while the old db-0 might still be writing to the disk. First I'd confirm the machine is really off, from the cloud console or whoever runs the hardware, because if it's only cut off from the network, a second copy could corrupt data. Once I'm sure, I'd delete the node object or force delete the pod. Newer versions also have an out-of-service taint for exactly this case. Then the volume can detach and the pod comes up elsewhere. If the disk is zonal, it must land in the same zone, so I'd check there's room there. A Multi-Attach error means the volume hasn't detached yet.”

Red flag to avoid:

Force deleting the pod straight away without confirming the old node is really down.

They may ask next:
  • Why is force deleting the pod dangerous if the node is only cut off from the network?
  • How would you design the database setup so losing one node doesn't need a person at all?
Say it in 60 seconds

Delivery 3 questions

Easy Situational round Fresher, Mid-level Practice question

6. Last night someone fixed production with kubectl edit. This morning the normal pipeline deployed and the fix vanished, so the bug is back. What do you do?

What the interviewer is really testing:
Whether you understand that the pipeline's files are the source of truth, and can fix both the outage and the habit without blame.
Answer frame:

Restore: put the fix into the repo and ship it through the pipeline, or roll back as a stopgap.

Why: the pipeline applies what's in Git, so a live edit lasts only until the next deploy.

Prevent: an emergency path that still ends in Git, and less write access in production.

Sample spoken answer:

“First the bug. I'd find exactly what was changed last night, from the person who did it or by comparing the old ReplicaSet's spec with the new one, put that change into the repo and push it through the pipeline. If that's slow and the edit was to the pod template, I'd use kubectl rollout undo to get back to last night's version as a stopgap, then follow with the proper commit. Then the why, which is simple: the pipeline applies whatever is in Git, so anything edited by hand lasts until the next deploy. Nobody did anything wrong at midnight, the process just had no emergency path. So I'd propose one: urgent live fixes are fine, but the same person opens a pull request straight after. Longer term I'd take production write access away from most people and add a tool that flags drift between Git and the cluster.”

Red flag to avoid:

Blaming the person who made the hotfix, or re-applying the edit by hand again and leaving the repo wrong.

They may ask next:
  • How does a GitOps controller change what happens to a manual edit?
  • Who should still have write access to production, and how would they get it in an emergency?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

7. A CI job was cancelled halfway through a helm upgrade. Now every new deploy fails saying another operation is in progress. How do you get deploys working again?

What the interviewer is really testing:
Whether you understand Helm keeps release state in the cluster and can recover a stuck release cleanly.
Answer frame:

Look: helm history shows the newest revision stuck in a pending state.

Check the cluster: see what the half-finished upgrade actually changed.

Recover: roll back to the last deployed revision, then run the pipeline again.

Prevent: never let CI kill Helm mid-run; give it a timeout and automatic rollback on failure.

Sample spoken answer:

“Helm keeps the state of each release in the cluster, by default as Secrets, one per revision. The cancelled job left the newest revision marked pending-upgrade, and Helm refuses to start another operation on top of it. I'd run helm history on the release to confirm that, and note the last revision marked deployed. Then I'd check what's actually running, because a half-finished upgrade may have changed some resources and not others. The clean fix is helm rollback to that last deployed revision. That puts the resources back and records a new, finished revision. Then I'd rerun the normal pipeline. I'd avoid deleting Helm's release Secrets by hand unless rollback fails, and even then I'd back them up first. To stop it recurring, I'd give the CI step a longer timeout than Helm's own, and turn on the option that rolls a failed upgrade back automatically.”

Red flag to avoid:

Uninstalling and reinstalling the release in production just to clear the error.

They may ask next:
  • Where does Helm store release information, and why does that matter for access control?
  • What happens to a release when an upgrade with automatic rollback times out?
Say it in 60 seconds
Medium Case round Mid-level, Senior Practice question

8. You're asked to send about one in twenty requests to a new version for a day before a full rollout. There's no service mesh. How would you do it?

What the interviewer is really testing:
Whether you can build a simple canary with plain Kubernetes pieces and know the limits of splitting by pod count.
Answer frame:

Simple way: a second Deployment with the new image that carries the label the Service selects.

The split: traffic follows pod count, so one canary pod beside nineteen stable ones is about one in twenty.

Finer control: an ingress controller that supports weighted canary routing.

Decide: compare errors and latency by version, then promote or delete the canary.

Sample spoken answer:

“The plainest way uses only a Service and two Deployments. The Service selects on a label like app: checkout. The stable Deployment has that label and nineteen replicas, and a canary Deployment with the new image has the same label and one replica. Since the Service spreads connections across all ready pods, roughly one in twenty requests hit the canary. It's rough, because it splits by pod and by connection, not exactly, and I'd have to scale the stable side to hold the ratio. If our ingress controller supports weighted canary routing, I'd use that instead, because it splits by request and doesn't tie the ratio to replica counts. Either way I'd add a version label to metrics so I can compare errors and latency by version. Before starting I'd check both versions can share the database schema and sessions, since users will bounce between them.”

Red flag to avoid:

Running the canary with no way to see its metrics separately from the stable version.

They may ask next:
  • Why can keep-alive connections make the split far from one in twenty?
  • What would make you stop the canary early?
Say it in 60 seconds

Security & Access 2 questions

Easy Situational round Fresher, Mid-level, Senior Practice question

9. A teammate meant to clean up dev but ran kubectl delete deployment against the production cluster, and the app is down. What do you do now, and afterwards?

What the interviewer is really testing:
Whether you restore service calmly from source, know what a delete takes with it, and fix the setup rather than the person.
Answer frame:

Restore: re-run the last good pipeline or apply the manifests from Git.

Check scope: a deleted Deployment takes its ReplicaSets and pods; Services, ConfigMaps and claims usually survive.

Tell people: one clear message in the incident channel while you work.

Prevent: separate kubeconfigs, the context in the shell prompt, and no delete rights in production for daily work.

Sample spoken answer:

“First I'd get the app back, and the fastest safe way is to redeploy from source: re-run the last good pipeline, or apply the manifests straight from Git. While that runs I'd check what else went. Deleting a Deployment removes its ReplicaSets and pods, but Services, ConfigMaps, Secrets and volume claims normally stay, so data should be safe. I'd post one short message in the incident channel so people know it's being handled and don't pile in. Once it's back, I'd write it up without blame. Anyone can type into the wrong terminal window. The fixes are in the setup: a separate kubeconfig file for production that you have to pick on purpose, the current context shown in every shell prompt, and day-to-day production access that doesn't include delete.”

Red flag to avoid:

Hunting for who to blame while the app is still down, or rebuilding the Deployment by hand from memory.

They may ask next:
  • What would you do differently if the teammate had deleted the whole namespace?
  • How would you make a destructive command in production need a second step?
Say it in 60 seconds
Medium Situational round Fresher, Mid-level, Senior Practice question

10. You notice a production database password in plain text in a Helm values file that's been in the Git repo for three months. What do you do?

What the interviewer is really testing:
Whether you treat the secret as already leaked and rotate it first, then fix the process so it can't happen again.
Answer frame:

Assume leaked: it's in history, every clone and maybe CI logs; deleting the line doesn't un-leak it.

Rotate first: a new password, rolled out to the app, then the old one disabled.

Clean up: the chart reads an existing Secret; rewriting history is a secondary step.

Prevent: a secret manager synced into the cluster, plus secret scanning in the pipeline.

Sample spoken answer:

“I'd treat it as leaked, because it has been for three months. Anyone with repo access, every clone and possibly the CI logs have it, so deleting the line fixes nothing on its own. The first job is rotation. I'd tell the database owner and the security team, create a new password, put it into a Kubernetes Secret that's managed outside Git, roll the app so it picks it up, check it's healthy, then disable the old password. I'd also look through the database logs for logins that don't look like ours. Then I'd change the chart so it reads the password from an existing Secret instead of from values. Rewriting Git history is a nice extra, but it can't take back copies people already have. To stop it recurring, I'd add secret scanning to the pipeline and move secrets into a proper secret manager synced into the cluster.”

Red flag to avoid:

Deleting the line from the file and considering the problem solved.

They may ask next:
  • How would you rotate the password without any downtime?
  • How can you keep secrets out of Git and still deploy everything from Git?
Say it in 60 seconds

Workloads 2 questions

Easy Technical round Fresher, Mid-level Practice question

11. Your service runs three replicas, yet it went fully down when a single node crashed, because all three pods were on that node. How do you stop that happening again?

What the interviewer is really testing:
Whether you know replica count alone doesn't give high availability, and how to ask the scheduler to spread pods.
Answer frame:

Why: the scheduler prefers spreading, but only as a preference; short on room, pods can land together.

Fix: topology spread constraints on hostname, and on zone if the cluster spans zones.

Backstop: a PodDisruptionBudget so a drain can't take them all at once.

Sample spoken answer:

“Three replicas only help if they're in different places, and the scheduler's default spreading is a scoring preference, not a rule. If the other nodes were short on room when the pods were created, all three can end up together. I'd check placement now with kubectl get pods -o wide. Then I'd add a topology spread constraint on the hostname label with a max skew of one, so replicas stay even across nodes, plus one on zone if the cluster spans zones. I'd choose between a hard rule and a soft one. Hard means a pod stays Pending rather than doubling up, which is safer but needs spare capacity. Existing pods don't move by themselves, so I'd restart the rollout to place them again. I'd also add a PodDisruptionBudget, so a node drain can't evict every replica at once.”

Code:
# inside the Deployment's pod template
spec:
  topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: kubernetes.io/hostname
    whenUnsatisfiable: DoNotSchedule
    labelSelector:
      matchLabels:
        app: checkout
Red flag to avoid:

Saying three replicas are enough for availability without checking where they actually run.

They may ask next:
  • How is pod anti-affinity different from a topology spread constraint here?
  • Why might a hard spread rule leave a pod Pending during a zone outage?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

12. You added a log-shipping sidecar to a batch Job. The main container finishes its work, but the Job never completes and the pod runs forever. Why, and how do you fix it?

What the interviewer is really testing:
Whether you know a Job's pod completes only when every container exits, and the ways to stop a helper container.
Answer frame:

Why: a pod succeeds only when all its containers have exited; the sidecar never does.

Native fix: newer versions let a sidecar be an init container with restartPolicy Always, stopped after the main one ends.

Older clusters: the main container tells the sidecar to quit through a shared file or an endpoint.

Safety net: activeDeadlineSeconds so a stuck Job can't run forever.

Sample spoken answer:

“A Job counts a pod as done only when the pod succeeds, and a pod only succeeds when all of its containers have exited. The log shipper is built to run forever, so the pod just sits there after the real work is finished. On a newer cluster the clean fix is a native sidecar: I'd declare the shipper as an init container with restartPolicy set to Always. It starts before the main container, keeps running beside it, and Kubernetes stops it once the main container is done. On an older cluster I'd make the main container tell the sidecar to stop, for example by writing a file to a shared emptyDir that the sidecar watches, or calling a quit endpoint if it has one. Either way I'd set activeDeadlineSeconds on the Job, so if something gets stuck again it fails loudly instead of running for days.”

Code:
spec:
  activeDeadlineSeconds: 3600
  template:
    spec:
      restartPolicy: Never
      initContainers:
      - name: log-shipper
        image: registry.example.com/log-shipper:2.1
        restartPolicy: Always   # native sidecar
      containers:
      - name: report
        image: registry.example.com/report:1.4
Red flag to avoid:

Deleting the finished pods by hand every night instead of fixing why the Job can't complete.

They may ask next:
  • In what order do native sidecars start and stop compared with the main container?
  • Would you ship logs from a sidecar here at all, or use a node-level agent instead?
Say it in 60 seconds

Health & Scaling 5 questions

Hard Technical round Mid-level, Senior Practice question

13. A Java service keeps restarting with OOMKilled, but its max heap is set well below the container's memory limit. How can that be, and what do you do?

What the interviewer is really testing:
Whether you know a container limit covers the whole process, not just the heap, and fix it with measurement instead of guesswork.
Answer frame:

Confirm: last state OOMKilled with exit code 137 in kubectl describe pod, not a node eviction.

Why: the limit counts everything: heap, metaspace, thread stacks, direct buffers, code cache, GC structures.

Measure: native memory tracking in a test pod, and the container's memory graph over time.

Fix: size the heap as a share of the limit with real headroom, or raise the limit; hunt a native leak if it keeps climbing.

Sample spoken answer:

“First I'd confirm it's really the container limit. kubectl describe pod should show the last state as OOMKilled with exit code 137, which means the kernel killed it for going over its limit, not the kubelet evicting it for node pressure. Then, the heap setting is only part of the story. The JVM also uses metaspace for classes, a stack for every thread, direct buffers for I/O, the code cache and memory for the garbage collector itself. A service with hundreds of threads and heavy networking can use a lot outside the heap. I'd turn on native memory tracking in a test pod to see where it goes. Usually the fix is to size the heap as a share of the limit, leaving real headroom, or to raise the limit. If memory climbs steadily for hours instead, I'd suspect a native leak and dig into that before just giving it more room.”

Red flag to avoid:

Raising the heap size, which makes the container go over its limit even sooner.

They may ask next:
  • What does MaxRAMPercentage do, and what value would you start with?
  • How would you tell a slow native leak apart from a limit that's simply too tight?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

14. Your API's slowest requests have got much worse, yet its CPU graph shows it well under its limit. Someone mentions throttling. How do you check, and what would you change?

What the interviewer is really testing:
Whether you understand that CPU limits are enforced in short time slices, and why an average usage graph hides the problem.
Answer frame:

How limits work: a CPU quota per short period; a busy multi-threaded app can use it up early and then wait.

Check: the throttled periods counters from cAdvisor, or cpu.stat in the container's cgroup.

Fix: raise or remove the CPU limit while keeping the request, and match thread counts to the CPU the container really gets.

Sample spoken answer:

“A CPU limit isn't enforced as an average. The kernel gives the container a quota for each short period, a tenth of a second by default. If the app runs many threads, they can burn the whole quota early in a period and then sit frozen until the next one starts. The graph averages over a minute and looks calm, but individual requests hit those pauses, so the slow end gets worse. To check, I'd look at the throttled periods counter that cAdvisor exposes, or read cpu.stat inside the container, and compare throttled periods with total periods. If a large share is throttled, that's the cause. The fixes are raising the limit, or removing the CPU limit and keeping a sensible request, which many teams do for latency-sensitive services. I'd also check the runtime isn't sizing its thread pools from the node's cores instead of what the container gets.”

Red flag to avoid:

Reading the average CPU graph, seeing headroom, and ruling throttling out.

They may ask next:
  • What's the risk of running without CPU limits in a shared cluster?
  • How do requests still protect a pod if you remove its limit?
Say it in 60 seconds
Medium Case round Mid-level, Senior Practice question

15. During a traffic spike the autoscaler took your API to its maximum replicas, but response times didn't improve at all. How do you work out why?

What the interviewer is really testing:
Whether you check that scaling actually produced capacity, and look for the real bottleneck instead of adding pods.
Answer frame:

Did pods arrive: new replicas may be Pending with no room, waiting on the node autoscaler.

Are they serving: slow start or a strict readiness probe keeps them out of the Service.

Real bottleneck: a shared database, connection pool or downstream API that more pods make worse.

Right signal: CPU burned on retries can drive scaling that only adds load.

Sample spoken answer:

“More replicas only help if they're running and the pods are the bottleneck. First I'd check the new replicas actually came up. kubectl get pods might show several Pending because the nodes are full and the node autoscaler is slow or at its own cap. Then I'd check the running ones are ready and in the Service's endpoints. If all that's fine, the bottleneck is probably behind the pods. The usual one is the database: every new pod opens its own connection pool, so scaling out can exhaust connections and make things slower, not faster. I'd look at database load and connection counts, and the latency of any downstream APIs. I'd also ask whether CPU is the right signal. If pods burn CPU retrying against a struggling database, the autoscaler keeps adding pods that only add more load.”

Red flag to avoid:

Raising the maximum replica count and hoping, without finding where the time goes.

They may ask next:
  • How would you protect the database from a sudden jump in pod count?
  • When would you scale on a custom metric instead of CPU?
Say it in 60 seconds
Medium Case round Mid-level, Senior Practice question

16. Your nodes show low real CPU use, yet new pods stay Pending with insufficient CPU and the cloud bill keeps climbing. What's going on, and how do you fix it?

What the interviewer is really testing:
Whether you know the scheduler places pods by requests, not real usage, and can right-size requests with data.
Answer frame:

Why: scheduling adds up requests; if requests sit far above real use, nodes are full on paper.

Check: compare Allocated resources in kubectl describe node with kubectl top.

Right-size: set requests from observed usage plus headroom, biggest gaps first.

Keep it right: vertical autoscaler recommendations and a regular review.

Sample spoken answer:

“The scheduler never looks at how busy a node really is. It adds up the CPU requests of the pods already there, and if the new pod's request doesn't fit in what's left, it stays Pending. So if teams asked for two cores and use a fifth of one, nodes look full while sitting idle, and the node autoscaler keeps buying more. I'd confirm it by comparing the allocated section in kubectl describe node with real usage from kubectl top or our monitoring. Then I'd rank services by the gap between request and real use and start with the biggest. For each one I'd set the request near its busy-hour usage plus some headroom, roll it out, and watch latency and throttling. A vertical pod autoscaler in recommendation mode helps keep the numbers honest over time. I'd share the before and after with the teams so they see why it matters.”

Red flag to avoid:

Adding more nodes, or cutting every request by the same amount, without looking at real usage per service.

They may ask next:
  • What's the risk of setting requests too low?
  • How would you choose the headroom for a very spiky service?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

17. The database had a thirty-second failover, and your API with ten replicas was down for several minutes afterwards while the pods kept restarting. What went wrong?

What the interviewer is really testing:
Whether you spot a liveness probe that checks a dependency, and know why that turns a short blip into a long outage.
Answer frame:

Clue: restarts across every pod at once point at the liveness probe.

Cause: the liveness endpoint checked the database, so the kubelet killed healthy processes.

Why longer: all pods restart together, warm up slowly and hit restart back-off.

Fix: liveness checks only the process; the app rides out dependency blips, and readiness is used with care.

Sample spoken answer:

“Restarts across every pod at the same moment make me look at the liveness probe first. I'd bet the health endpoint it calls also checks the database. When the database failed over, every pod's liveness check failed, so the kubelet killed perfectly healthy processes all at once. Then they all restarted together, spent time warming up, and some hit the restart back-off, so thirty seconds of database trouble became minutes of outage. I'd confirm it from the pod events and the probe endpoint's code. The fix is that liveness answers one question: is this process stuck? It shouldn't touch the database at all. For dependencies, the app should retry and return clear errors while the database is away. I'd be careful moving the check into readiness too, because if every pod goes unready at once the Service has no endpoints, which is also an outage.”

Red flag to avoid:

Raising the probe timeout and leaving the database check inside the liveness endpoint.

They may ask next:
  • What should a good liveness endpoint actually check?
  • When is it right for readiness to depend on something outside the pod?
Say it in 60 seconds

Networking 3 questions

Hard Technical round Mid-level, Senior Practice question

18. Every deploy causes a short burst of 502 errors, even though you use rolling updates and readiness probes. What's going on, and how do you fix it?

What the interviewer is really testing:
Whether you know pod shutdown and endpoint removal happen at the same time, and how to make old pods leave gracefully.
Answer frame:

New pods: confirm readiness really means ready, and the rollout never drops below capacity.

Old pods: SIGTERM and removal from endpoints happen in parallel, so traffic still arrives after shutdown starts.

Fix: a short preStop sleep, SIGTERM handling that drains in-flight requests, and a grace period longer than both.

Sample spoken answer:

“Readiness probes protect the new pods, so I'd suspect the old ones. When Kubernetes terminates a pod, it sends SIGTERM to the container and removes the pod from the Service's endpoints at the same time, not one after the other. kube-proxy and the ingress controller take a moment to notice, so for a second or two they still send requests to a pod that has started shutting down. If the app exits the moment it gets SIGTERM, those requests fail as 502s. The fix has three parts. A preStop hook that just sleeps a few seconds, so the pod keeps serving while everyone stops routing to it. The app handles SIGTERM by finishing in-flight requests before it exits. And a termination grace period long enough to cover both. I'd still check the new pods too, because a readiness probe that passes before the app is warm gives the same symptom.”

Code:
# pod template: keep serving while routing updates, then drain
spec:
  terminationGracePeriodSeconds: 45
  containers:
  - name: api
    image: registry.example.com/api:3.2
    lifecycle:
      preStop:
        exec:
          command: ["sleep", "10"]
Red flag to avoid:

Blaming the readiness probe on the new pods and never looking at how old pods shut down.

They may ask next:
  • What would you do if the image has no shell or sleep binary?
  • How do long-lived connections like WebSockets change this?
Say it in 60 seconds
Hard Technical round Senior Practice question

19. Calls from your pods to an outside API are sometimes exactly five seconds slower than normal, while the API itself looks fine. Where would you look?

What the interviewer is really testing:
Whether you recognise a DNS timeout pattern and know how pod DNS settings multiply lookups.
Answer frame:

The clue: a fixed five seconds is the resolver's default timeout, so a DNS query was lost.

Extra lookups: ndots 5 and the search list turn one outside name into several queries.

Lost packets: parallel A and AAAA queries over UDP can collide in connection tracking and one is dropped.

Fix: trailing-dot names or a lower ndots, a node-local DNS cache, resolver options that avoid the race.

Sample spoken answer:

“Exactly five seconds is a big clue, because that's the default time a Linux resolver waits before retrying a DNS query. So I'd suspect a lost DNS packet, not the API. I'd look at the pod's resolv.conf. Pods usually get ndots set to five and a list of cluster search domains, so a name like api.example.com, with only two dots, is first tried with each cluster suffix added. That's several extra queries per call, and each is a chance to lose one. On top of that, the resolver often sends the A and AAAA queries at the same moment from one socket, and on some kernels those collide in connection tracking and one gets dropped. To fix it I'd use a trailing dot on outside names or lower ndots in the pod's dnsConfig, and run a node-local DNS cache. A small loop of timed lookups inside a pod would prove it before and after.”

Red flag to avoid:

Blaming the outside API or adding retries in the app without noticing the fixed five-second pattern.

They may ask next:
  • What would you see in CoreDNS metrics if it were simply overloaded instead?
  • What's the trade-off in lowering ndots for every pod in the cluster?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

20. Security applied a default-deny NetworkPolicy to a namespace. Now every app in it fails with name resolution errors, even for services in the same namespace. What happened?

What the interviewer is really testing:
Whether you know DNS is ordinary egress traffic to CoreDNS in another namespace, and write the narrow rule that allows it.
Answer frame:

Why: denying all egress also blocks port 53 to CoreDNS in kube-system, so no name resolves.

Confirm: lookups time out rather than fail fast, and the policy has no egress rule for DNS.

Fix: allow UDP and TCP 53 to the DNS pods only, then add each real dependency one by one.

Sample spoken answer:

“Even a service in the same namespace is found by name, and that name is answered by CoreDNS, which runs in kube-system. A default-deny policy that covers egress blocks that traffic like anything else, so every lookup times out. I'd confirm it from a pod: nslookup of a service name hangs instead of failing quickly, and reading the policy shows egress is denied with nothing allowed for DNS. The fix isn't to drop the policy. I'd add a small egress rule allowing UDP and TCP port 53 to the pods labelled as kube-dns in the kube-system namespace, and nothing else. Then I'd work through each app's real dependencies, like its database and the other services it calls, and allow those one at a time. Next time I'd roll default-deny out to one namespace in a test environment first, with the DNS rule already in place.”

Code:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-dns
spec:
  podSelector: {}
  policyTypes: ["Egress"]
  egress:
  - to:
    - namespaceSelector:
        matchLabels:
          kubernetes.io/metadata.name: kube-system
      podSelector:
        matchLabels:
          k8s-app: kube-dns
    ports:
    - protocol: UDP
      port: 53
    - protocol: TCP
      port: 53
Red flag to avoid:

Deleting the policy to make the errors go away.

They may ask next:
  • Why allow TCP on port 53 as well as UDP?
  • How would you find every dependency an app needs before locking its egress down?
Say it in 60 seconds

Cluster Operations 4 questions

Medium Technical round Mid-level, Senior Practice question

21. A worker node went offline. Traffic moved off it quickly, but it took about five minutes before its pods were recreated elsewhere. The team asks why so slow. What do you tell them?

What the interviewer is really testing:
Whether you know the default tolerations that delay rescheduling, and the trade-off in shortening them.
Answer frame:

Detection: the node stops renewing its heartbeat, is marked NotReady or Unknown, and its pods are marked not ready.

The wait: pods tolerate the not-ready and unreachable taints for five minutes by default, then are evicted.

Trade-off: a shorter tolerationSeconds replaces pods faster but also reacts to brief network blips.

Better answer: enough replicas spread across nodes that five minutes costs capacity, not availability.

Sample spoken answer:

“That's the default behaviour, not a fault. When a node stops renewing its heartbeat, the control plane marks it NotReady or Unknown after a short grace period and marks its pods not ready, which is why traffic left quickly. Then it taints the node, and every pod by default tolerates those taints for five minutes. Only after that are the pods evicted and their Deployments create replacements elsewhere. The wait is on purpose: a node that drops off the network for a minute shouldn't cause every pod on it to be rebuilt. We can shorten tolerationSeconds for the not-ready and unreachable taints on services that really care, but I'd be careful, because a flaky network then causes churn. I'd rather make sure each service has enough replicas spread across nodes, so losing one node's pods for five minutes costs some capacity but never availability.”

Red flag to avoid:

Saying the cluster is misconfigured without knowing the default five-minute tolerations exist.

They may ask next:
  • How would the answer change for a StatefulSet pod on that node?
  • Where would you set tolerationSeconds, and for which kinds of services?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

22. During a node upgrade, kubectl drain has been stuck for half an hour, retrying the same pod over and over. What's blocking it, and what do you do?

What the interviewer is really testing:
Whether you know drain uses the eviction API that honours PodDisruptionBudgets, and fix the conflict rather than force past it.
Answer frame:

Read the message: drain says the eviction would violate the pod's disruption budget.

Check the budget: kubectl get pdb shows allowed disruptions at zero, often one replica with minAvailable 1.

Fix properly: scale up first, wait for a healthy spare, then let the drain finish.

Avoid: force deleting the pod, which skips the protection the budget exists for.

Sample spoken answer:

“Drain doesn't just delete pods, it asks the eviction API, and that honours PodDisruptionBudgets. If a budget says no disruption is allowed right now, drain keeps retrying, and the message usually says so. I'd run kubectl get pdb in that namespace and look at allowed disruptions. The classic case is a Deployment with one replica and a budget saying minAvailable one. That pod can never be evicted, so drain waits forever. Another cause is a budget whose other pods are already unhealthy, so there's no slack to give. The right fix is to scale the Deployment to two, wait for the new pod to be ready on another node, let the drain finish, then scale back if needed. I'd tell the owning team, because a one-replica service with a strict budget is a design problem. I wouldn't force delete the pod, since that skips the very protection they asked for.”

Red flag to avoid:

Force deleting the pod to get past the drain without telling anyone.

They may ask next:
  • Which other kinds of pods make drain refuse to start, and which flags handle them?
  • How would you write a budget for a three-replica service that still allows upgrades?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

23. One team deploys pods with no resource requests to a shared cluster, and their batch jobs keep starving other teams' services. They say requests slow them down. How do you handle it?

What the interviewer is really testing:
Whether you can put fair guardrails in place and win the team over instead of just blocking them.
Answer frame:

Show the impact: node graphs and the incidents their jobs caused, without blame.

Defaults: a LimitRange gives every container sensible requests even when none are written.

Fair share: a ResourceQuota per namespace, and a lower priority class or separate nodes for batch work.

Make it easy: templates or chart defaults so setting requests takes no effort.

Sample spoken answer:

“I'd start with a conversation, not a policy. I'd show them the node graphs from when their jobs ran and the incidents other teams had, and explain that pods with no requests are scheduled as if they need nothing, so they pile onto nodes and then fight everyone for CPU and memory. Then I'd make the right thing effortless. A LimitRange in their namespace gives every container default requests and limits even if nobody writes them, so they barely have to change anything. A ResourceQuota caps the namespace so one team can't take the whole cluster. For the batch jobs I'd suggest a lower priority class or a separate node pool, so they use spare capacity without pushing out services. I'd agree the default numbers with them and review after a couple of weeks of real usage, so the limits feel fair rather than imposed.”

Red flag to avoid:

Silently applying a strict quota that breaks their next deploy without talking to them.

They may ask next:
  • What quality of service class does a pod with no requests and no limits get, and why does it matter?
  • How would you pick the default request values for their namespace?
Say it in 60 seconds
Medium Case round Senior Practice question

24. Your cluster is two minor versions behind and must be upgraded this quarter. Nobody on the team has done it before, and some manifests are years old. How do you plan it?

What the interviewer is really testing:
Whether you plan upgrades around removed APIs and version order, with a rehearsed path and a way back.
Answer frame:

Removed APIs: scan manifests, charts and live objects for API versions the target drops.

Order: one minor version at a time, control plane first, then nodes; kubelets never newer than the API server.

Rehearse: a staging cluster first, and a backup of etcd or the cluster before production.

Nodes and add-ons: drain in batches respecting budgets, and confirm add-ons support the new version.

Sample spoken answer:

“The biggest risk isn't the upgrade button, it's old manifests. Releases can remove API versions that were deprecated earlier, and a manifest using one simply stops applying. So first I'd read the release notes for both target versions and scan our Git manifests, Helm charts and the live cluster for removed API versions, then fix those while still on the old version, since the newer API versions usually exist there already. For the upgrade itself, I'd go one minor version at a time, control plane first, then nodes, because kubelets can lag behind the API server but must never be newer. I'd rehearse on a staging cluster built the same way, and back up etcd or the cluster before touching production. I'd drain nodes in small batches, respecting disruption budgets, and check workloads after each batch. I'd also confirm add-ons like the ingress controller and network plugin support the new version.”

Red flag to avoid:

Jumping straight to the newest version in production without checking for removed APIs.

They may ask next:
  • How would you find which clients are still calling a deprecated API?
  • What's your plan if the control plane upgrade goes wrong halfway?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card