Scenario rounds hand you a broken build or a misbehaving container and watch how you think. A script that is clearly there says 'no such file', a firewall that looks closed isn't, a port is already taken, a Postgres upgrade won't start. The interviewer wants the order you check things in, what you would look at first, and what would change your mind. This page is for anyone facing that round, from a first DevOps job to a senior platform role. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying your first two checks before the fix.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Why: each CI job starts on a clean runner, so there is no local layer cache to reuse.
Check first: confirm the Dockerfile copies the lockfile and installs before copying the code, or even a warm cache won't help.
Fix: export the cache to the registry or the CI cache store and import it on the next build.
Extra: cache mounts for the package manager help on builders that persist between jobs.
“My laptop keeps its layer cache between builds, but a CI runner is usually brand new every job, so it has nothing to reuse and runs every step from scratch. First I'd check the Dockerfile order, because if the code is copied before the dependency install, even a warm cache gets thrown away on every commit. If the order is fine, I'd give CI a cache it can carry between jobs. With BuildKit I can push the cache to the registry with cache-to and pull it back with cache-from, using mode max so the intermediate stages are kept too. Most CI systems also have their own cache store that BuildKit can use. I'd compare build times before and after on a commit that changes only code. If it's still slow, I'd look at whether the lockfile changes on every commit.”
docker buildx build \
--cache-from type=registry,ref=registry.example.com/api:buildcache \
--cache-to type=registry,ref=registry.example.com/api:buildcache,mode=max \
-t registry.example.com/api:1.8.0 --push .
Blaming the CI provider's machines without noticing that a fresh runner simply has no build cache.
Context: check which folder is passed as the build context; paths are relative to it, not to the Dockerfile.
Ignore file: check whether .dockerignore excludes the file or its folder.
Outside the context: a path like ../shared can never be copied.
Case: a name that differs only in capital letters works on a Mac and fails on Linux.
“COPY can only see files inside the build context, so I'd start there. I'd check the last argument of the build command. If CI runs the build with the docker folder as the context but the file lives at the repo root, it's simply not in what Docker was sent. Next I'd open .dockerignore, because a broad pattern like a wildcard that excludes everything, or one that ignores a whole config folder, silently removes the file. Then I'd check whether the path climbs out with dot-dot, which COPY never allows. The sneaky one is letter case: my Mac's file system usually ignores case, so Config.json and config.json look the same locally, but the Linux CI runner treats them as different files. If it only fails in CI, that's where I'd look first.”
Moving the Dockerfile around at random without understanding that COPY paths are relative to the build context.
Shebang: check the first line; if it asks for bash and the image only has sh, the interpreter is what's missing.
Line endings: Windows line endings make the kernel look for an interpreter named sh plus a hidden carriage return.
Confirm: print the first line in the image through od -c and look for a stray carriage return.
Prevent: force LF endings for shell scripts in .gitattributes.
“When the script is clearly there, the missing file is almost always the interpreter named on its first line. I'd run the image with a shell and look. If the first line says bin bash and the base is Alpine, there's no bash, so exec fails with exactly this message. The other classic cause is Windows line endings. If someone edited the script on Windows or git converted it, the first line ends with a carriage return, and the kernel goes looking for a program called sh with that invisible character stuck on the end. Printing the first line through od -c shows a backslash r just before the newline. I'd fix it by converting the file to LF, and add a gitattributes rule so shell scripts always check out with LF, which stops it coming back on the next Windows laptop.”
docker run --rm --entrypoint sh myapp:dev -c 'head -1 /entrypoint.sh | od -c'
# ... s h \r \n <- carriage return: Windows line endings
# .gitattributes
*.sh text eol=lf
Rebuilding again and again, or adding more COPY lines, without checking the script's first line.
Cause: Alpine uses musl libc, and many packages ship prebuilt wheels only for glibc systems.
Effect: pip falls back to building from source, which needs compilers and is slow, and native code can behave differently.
Fix: move to a slim Debian-based image, which is small enough and uses the normal wheels.
Measure: compare final size and build time, not just the base image size.
“Alpine is small because it uses musl instead of glibc. For Python that matters, because many packages with native code publish ready-made wheels built against glibc. If there's no musl wheel for a package, pip downloads the source and compiles it, so I suddenly need compilers in the image and the build crawls. The runtime crash fits too: native libraries built or tuned for glibc can behave differently on musl, and some only get tested on glibc. I'd check the pip log to see which packages built from source. Then I'd usually switch to the slim Python image. It's a bit bigger than Alpine but installs the normal wheels in seconds, and with a multi-stage build the final image often ends up about the same size. I'd judge the switch by final size and build time together.”
Insisting Alpine is always the best choice because its base image is the smallest.
Base image: a tag like python:3.12-slim can point to a new image today; compare the digests both builds pulled.
Unpinned installs: apt, pip or npm without a lockfile can pull newer versions.
Downloads: anything fetched from a URL during the build can change too.
Fix: pin the base by digest with a bot to bump it, use lockfiles, and keep the build log's resolved versions.
“If the repo didn't change, something outside it did, so I'd list the inputs that come from the network. First the base image: a tag isn't fixed, and it may have moved to a new patch release or a newer OS overnight. I'd compare the digest yesterday's build resolved with today's, from the build logs or by inspecting both images. Next, unpinned package installs. An apt-get install without versions, or pip and npm without a lockfile, can quietly pull a newer library. I'd diff the installed package lists from the two images. Then anything downloaded by URL. Once I find the change, I'd pin it. The base goes by digest with a tool that opens a pull request to bump it, so updates still arrive but they're tested and reviewed instead of landing by surprise.”
Saying the build must be flaky and re-running it until it passes.
Measure: docker system df shows images, containers, volumes and build cache.
Logs: check the container log files; the default json-file driver doesn't rotate them.
Safe cleanup: stopped containers, dangling images and old build cache first.
Careful: never prune volumes blindly; set log rotation so it doesn't come back.
“First I'd get the space back without breaking anything, then stop it happening again. I'd run docker system df to see the split between images, containers, volumes and build cache. On servers the surprise is often logs. The default json-file driver keeps every line forever unless you set rotation, so I'd check the log files under the containers folder and find the chatty one. Stopped containers, dangling images and old build cache are usually safe to prune. I'd not prune volumes on a whim, because that's where databases keep their data. For the lasting fix I'd set max-size and max-file for logs in the daemon config, knowing it applies only to containers created after the change, so running ones need recreating. I'd also add a disk alert that fires well before the disk is full.”
{
"log-driver": "json-file",
"log-opts": { "max-size": "50m", "max-file": "3" }
}
Running a full system prune with volumes included on a production server to free space quickly.
Read the output: docker inspect shows the last check results, with exit codes and output.
Missing tool: a check that calls curl fails if a slim or distroless image has no curl.
Timing: a start period or timeout that's too short for a slow start fails the first checks.
Target: a wrong port or path inside the container.
“The container's health state records the last few checks with their output, so I'd read that before guessing. docker inspect on the container shows the Health section. Very often the output says curl isn't found. The check was copied from an example, but our slim or distroless image doesn't include curl, so the check fails every time even though the app is fine. Other things I'd look for: a timeout shorter than the endpoint needs, no start period for an app that takes a while to warm up, or the check hitting a different port than the app listens on inside the container. The fix depends on the cause. For the missing tool, I'd rather add a tiny health command built into the app than install curl just for the check, especially on a distroless image with no shell.”
Removing the health check so the container stops being replaced.
How limits work: a CPU limit is a quota per short period, often 100 milliseconds.
Throttling: many threads can burn the whole quota early in a period, then wait for the rest of it.
Check: the cgroup's cpu.stat shows how often and how long the container was throttled.
Fix: match thread pools to the limit, raise the limit, or drop the hard limit and rely on CPU shares.
“Averages hide this. A CPU limit isn't a smooth speed cap. The kernel gives the container a quota of CPU time per short period, usually 100 milliseconds. If the service has eight busy threads, a limit of one CPU is used up in about twelve milliseconds, and then every thread waits for the rest of the period. Requests caught in that wait get slow, even though the average over a minute looks low. Some runtimes also size their thread pools from the host's core count, not the limit, which makes it worse. I'd confirm by reading cpu.stat in the container's cgroup: rising throttled counts and time line up with the spikes. Then I'd size the thread pools to the limit, raise the limit, or for latency-sensitive services remove the hard limit and use CPU shares so they can burst when the host has room.”
Reading the low average CPU and concluding CPU can't be the problem.
Replicas: with three copies of the app, the job runs three times a night.
Process care: two processes need a supervisor; a cron crash goes unnoticed and signals get messy.
Cron quirks: jobs get a minimal environment and log to places no one reads.
Instead: run the same image with a different command, started by a scheduler or a cluster's scheduled job.
“I like that the job reuses our image, but I'd not put cron inside the web container. First, scaling: the moment we run three copies of the app, the cleanup runs three times a night, which might be harmless or might double-delete things. Second, the container now has two processes, so something has to supervise them. If cron dies, the web server keeps the container looking healthy and no one notices. Third, cron itself is awkward in containers. Jobs start with a very small environment, so they may not see our config variables, and their output often goes nowhere. Instead I'd run the same image with a different command, like a cleanup script, as its own container started on a schedule. That could be a scheduled job in our cluster, or a timer on the host. It then logs and fails visibly.”
Accepting cron inside the web container without thinking about what happens when the app scales out.
Numbers not names: the kernel checks the numeric user and group IDs; the names inside and outside don't matter.
Check: compare id inside the container with ls -ln on the host folder.
Fix: run the container with a matching ID, change the folder's owner, or use a named volume.
Other cause: on SELinux hosts the mount may need a relabel option.
“A bind mount shares the real host folder, and permissions are checked by numeric user ID, not by name. So if my app runs as user 1000 inside the container but the host folder belongs to user 1001, it's a stranger to that folder. A common twist is that the host path didn't exist, so Docker created it, and it's owned by root. I'd compare the output of id inside the container with ls -ln on the host folder. In development I'd usually run the container with my own user and group IDs so files come out owned by me. On a server I'd change the folder's owner to the container's user, or switch to a named volume, which Docker creates for me. If the IDs match and it still fails on a Fedora or RHEL host, I'd suspect SELinux and try the z label.”
Fixing it with chmod 777 on the folder or by running the app as root.
Cause: the mount covers the image's /app, including the node_modules installed during the build.
Check: your host folder has no node_modules, or one built for a different OS.
Fix: add a separate volume for /app/node_modules so the mount doesn't hide it.
Watch out: that volume keeps old packages until you renew it after a dependency change.
“The image did install node_modules into /app during the build. But when I mount my project folder over /app, the mount hides everything the image had there, and my laptop's folder either has no node_modules or has ones built for macOS, which breaks native modules on Linux. So the app starts and can't find its packages. The usual fix in Compose is to add a second volume just for /app/node_modules. That path then isn't covered by my source mount, and it gets filled from the image. The catch is that this volume sticks around, so after I add a package and rebuild, the container can still see the old modules. I'd rebuild with the option that renews anonymous volumes, or remove that volume. Another option is installing packages outside /app entirely.”
Deciding the Dockerfile's install step is broken and adding a second npm install at container start.
Cause: the official image runs init scripts and sets the password only on an empty data directory.
Check: the data sits in a volume that survived the restart, so setup was skipped.
Local fix: remove the volume to start fresh, knowing that deletes the data.
Shared fix: change the password with SQL and apply schema changes as migrations.
“The Postgres image does its setup once. When the container starts, it checks whether the data folder is empty. Only then does it create the database, set the password from the environment and run the scripts in the init folder. My data lives in a volume that survived the restart, so the image saw an existing database and skipped all of that. The password variable is ignored after the first run, which is why the new one fails. On my laptop with throwaway data, I'd bring the stack down with its volumes and start again. On a shared environment I'd never do that. I'd change the password with an ALTER USER statement and run schema changes through a proper migration tool, so they apply to databases that already exist.”
Deleting the volume on a shared or production environment to make the init scripts run again.
Restore: put the old image tag back; the volume is untouched, so the old version starts again.
Why: a new major version can't read the old data format directly.
Plan: back up, then dump and restore or use pg_upgrade, tested on a copy first.
Downtime: pick the method by data size and how much downtime is allowed.
“First, get the app back. The new container refused to start, so it shouldn't have touched the data. I'd put the old image tag back and start it, then check the app works. Then I'd explain the cause: a new major Postgres version can't open the old version's data files, so swapping the image is never enough. For the real upgrade I'd take a backup, then pick a method. For a small database, dumping from the old version and restoring into a new one on a fresh volume is simple and safe. For a large one, pg_upgrade is much faster but needs both versions' binaries. If downtime has to be tiny, logical replication to the new version and a quick switch works. Whatever I choose, I'd rehearse it on a copy of the volume and keep the old volume until the new one is proven.”
Deleting the volume so the new version can start clean, or letting the app keep writing during a half-done upgrade.
Cause: the app listens on 127.0.0.1, which inside the container means only the container itself.
Check: ss -ltn or netstat inside the container shows the address it bound to.
Fix: make the app listen on 0.0.0.0; many dev servers default to localhost.
Rule out: a wrong container port in the -p mapping.
“The fact that it works from inside the container tells me the app runs fine, so it's about where it listens. Each container has its own network stack, and localhost inside it means only that container. When I publish a port, Docker forwards traffic to the container's own network interface, not to its loopback. If the app bound to 127.0.0.1, that forwarded traffic finds nothing listening and the connection gets reset. I'd confirm with ss -ltn inside the container: if it shows 127.0.0.1:8080 instead of 0.0.0.0:8080, that's it. The fix is to make the app listen on 0.0.0.0, usually a host flag or setting. Flask's dev server and several front-end dev servers default to localhost, so this bites a lot of people. I'd also double-check the mapping points at the port the app really uses.”
Blaming the firewall or reinstalling Docker without checking which address the app listens on.
Cause: Docker adds its own iptables rules that forward published ports before the host firewall's input rules see them.
Right now: stop publishing the port, or bind it to 127.0.0.1 only; check Redis for a password and unknown keys.
Better: let other containers reach it over a private Docker network with no published port.
Policy: put any extra filtering in the DOCKER-USER chain and scan from outside after every change.
“Docker writes its own iptables rules. When a port is published, traffic to it is redirected and forwarded to the container, so it never passes through the input rules that tools like ufw manage. The firewall looks closed but the port is open. I'd treat it as an incident. First I'd stop the exposure by removing the port mapping or binding it to 127.0.0.1 only, then check Redis for a password, strange keys or changed config, since an open Redis gets found fast. The better design is not to publish it at all. The app reaches Redis by service name over a private Docker network. If I really need host-level filtering for published ports, it goes in the DOCKER-USER chain, which Docker leaves alone. Then I'd scan from outside to prove it's closed, and add that scan to our checks.”
Trusting the host firewall's status output and assuming a published port is blocked because the firewall says so.
Cause: Docker's default bridge uses 172.17.0.0/16, so the host now sends that range to docker0.
Check: ip route shows the 172.17 route pointing at docker0; ip route get 172.17.5.10 confirms it.
Fix: set bip and default-address-pools in the daemon config to a range the company doesn't use.
Apply: restart the daemon and recreate networks; Compose networks can clash too.
“The timing is the clue. Docker's default bridge network takes 172.17.0.0/16, and the service lives in that same range. Once Docker starts, the host has a route saying everything in 172.17 lives on the docker0 bridge, so packets for the internal service go into Docker instead of out to the real network. I'd confirm with ip route, and ip route get on the service's address would show docker0. The fix is to give Docker ranges nobody on the company network uses. In the daemon config I'd set bip for the default bridge and default-address-pools for the networks Compose and users create, then restart Docker and recreate the networks. I'd check with the network team which ranges are free, and put the same config on every Docker host so it doesn't happen elsewhere.”
Blaming the internal service or its firewall without looking at the host's routing table.
Containers: docker ps with a publish filter finds a container using that host port.
Processes: ss -ltnp or lsof shows a Postgres installed straight on the machine.
Old projects: a copy of the project in another folder runs under a different Compose project name.
Fix: stop what you don't need, or map a different host port like 5433:5432.
“Something already listens on host port 5432, so I'd find out what. First docker ps, filtered by the published port, to see if a container holds it. Often it's a leftover from another project, or the same project cloned into a second folder. Compose names projects after the folder by default, so the two copies don't see each other's containers. If no container shows up, I'd run ss -ltnp or lsof on the port, because a Postgres installed straight on my machine is a common culprit. Then I decide. If it's something I don't need, I stop it. If I need both, I change our mapping to 5433:5432, since only the host side has to be unique, and point my local tools at 5433. Inside the Compose network the app still uses 5432.”
Killing whatever process holds the port without checking what it is or who uses it.
Cause: restart restarts the same container with the config it was created with.
Fix: docker compose up -d notices the change and recreates the container.
Check: docker compose config shows the resolved value; docker exec env shows what the container has.
Build time: if the value is baked in with ARG or ENV in the Dockerfile, you need a rebuild.
“A container's environment is fixed when the container is created. docker compose restart just stops and starts that same container, so it keeps the old values. I'd run docker compose up -d instead. Compose compares the current config with what the container was created from, sees the change and recreates it. To be sure, I'd check two places. docker compose config shows what Compose resolves from the .env file, which catches typos or a shell variable overriding the file. Then docker exec with env shows what the running container actually has. If both look right and the app still shows the old value, the value may be baked into the image through ARG or ENV in the Dockerfile, or the app caches its config, and that needs a rebuild or an app restart of its own.”
Rebuilding the image over and over without checking what the running container's environment actually holds.
Revoke first: assume the token is compromised; rotate it before anything else.
Investigate: check the token's usage logs for access you don't recognise.
Remove: delete the public tags, knowing copies may already exist.
Fix the build: use BuildKit secret mounts so the token never lands in a layer, and add secret scanning.
“The token has been public for a month, so I'd assume someone has it. Step one is to revoke it and issue a new one, before any cleanup, because deleting the image doesn't un-leak anything. Anyone who pulled it can read build args from the image history. Next I'd check the package registry's logs for that token to see if anyone else used it, and tell the security team. Then I'd remove the image tags from the public registry, and make the repo private if it was never meant to be public. Finally I'd fix the build. With BuildKit I'd pass the token as a secret mount on the one RUN step that needs it. It's available during that step and never written into a layer or the history. I'd also add secret scanning on images in CI.”
# Dockerfile
# RUN --mount=type=secret,id=npmrc,target=/root/.npmrc npm ci
docker build --secret id=npmrc,src="$HOME/.npmrc" -t api:dev .
Deleting the image quietly and moving on without rotating the token.
Risk: whoever can talk to the socket can start a privileged container that mounts the host's disk.
Read-only myth: a read-only mount of a socket still lets you send any API request.
Options: a filtering socket proxy that allows only the read calls the agent needs, or a different data source.
If kept: pin and trust the agent image, and write the accepted risk down.
“I'd agree with the review. The Docker socket is the full Docker API. Anything that can send requests to it can start a new container with privileged mode and the host's root folder mounted, which is root on the host. The read-only flag on the mount doesn't help, because it only stops changes to the socket file itself. Requests through it still work. So if that agent or its image were ever compromised, the whole server goes with it. I'd first ask what the agent needs. Usually it's listing containers and reading stats. A small socket proxy that allows only those read endpoints cuts the risk a lot. Some agents can read metrics from the cgroup files instead. If we truly must keep the socket, I'd pin the agent image by digest and write the risk down with an owner.”
Accepting that a read-only mount makes the Docker socket safe.
Why it matters: privileged mode grants every capability and every host device, and removes most confinement.
Find the real need: read the error without it; usually one capability or one device is missing.
Grant narrowly: add that one capability or device instead.
Deadline: if it truly can't wait, agree on a dated follow-up and keep the container away from sensitive hosts.
“I'd not just say no, because they have a deadline. I'd explain the cost in one line: privileged mode gives the container nearly everything the host's root user has, including every device, so a bug in it becomes a problem for the whole machine. Then I'd sit with them and run it without the flag to see what actually fails. Most of the time it's one thing. A VPN or network tool needs the NET_ADMIN capability, or something needs a single device like a USB serial port. I'd add just that capability or device and test again. That usually takes less than an hour. If it truly can't be narrowed today, I'd agree to ship on isolated hosts only, with a ticket and a date to fix it, rather than letting privileged mode quietly become normal.”
Either approving privileged mode without asking why, or blocking the release without helping find the real fix.
Find the digest: old deploy logs, CI logs, or docker images with digests on a host that still has it.
Roll back by digest: deploy image@sha256 instead of the tag.
Fallback: rebuild from the last good commit, knowing it may not be identical.
Prevent: unique tags per build, such as the commit hash, and tag immutability in the registry.
“The tag moved, but the old image may still exist under its digest, so I'd hunt for that first. Our deploy or CI logs often print the digest that was pushed. If not, a server that ran the old version may still have it locally, and docker images with the digests option shows it. Many registries also keep the untagged manifest until garbage collection runs. Once I have the digest, I deploy by digest, which points at exactly that image whatever happens to tags. If nothing turns up, I'd rebuild from the last good commit. That gets us close, but base images and packages may have moved, so I'd watch it closely. Afterwards I'd push for every build getting a unique tag like the commit hash, turn on tag immutability in the registry, and have deploys record the digest.”
Assuming a rebuild of the old commit gives exactly the same image, or searching for the old image by tag.
Cause: Docker Hub limits pulls, and anonymous pulls are counted by IP address.
Why busy hours: many runners behind one shared address use up the allowance together.
Quick fix: log in to Docker Hub in CI so pulls count against an account.
Lasting fix: a pull-through cache or copies of base images in your own registry.
“That error is Docker Hub's pull rate limit. Anonymous pulls are counted per IP address, and our CI runners probably share one outbound address, so in busy hours they use up the allowance together and the unlucky jobs fail. That explains why it looks random. The quick fix is to log in to Docker Hub in CI so pulls count against an account, which has a higher limit. The better fix is to stop pulling from the public hub on every job. I'd set up a pull-through cache or mirror, or copy the base images we use into our own registry and point our Dockerfiles at those. That also helps if the hub is down, and we control when a base image changes. I'd make the job log which registry each image came from, so it's clear in future.”
Adding automatic retries to CI and calling it fixed.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.