Scenario rounds hand you a symptom and watch how you think. Pods can't get an IP address, an event rule never fires, a cache restart takes the database down, an alert says a server is mining crypto. There is rarely one right answer. The interviewer wants the order you check things in, what you would look at first, and what would change your mind. This page is for anyone facing that round, from a first cloud job to a senior hire. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying the first three checks before the fix.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Why: resolvers and clients keep the old answer until its TTL runs out, and a few hold it longer.
Check: look the name up through several public resolvers and read the TTL they return.
Right now: keep the old servers serving, or forwarding to the new ones, until their traffic dies off.
Next time: lower the TTL well ahead of the change, switch, then raise it again.
“That's DNS caching doing its job. When a resolver looks up our name, it keeps the answer for as long as the record's TTL says. If the TTL was a day, some users won't see the change for up to a day, and a few clients with their own cache hold on even longer. I'd confirm it by querying the name through a few public resolvers with dig and looking at the TTL that's left. The important thing right now is not to switch the old servers off. I'd keep them serving, or have them forward to the new load balancer, and watch their traffic fall to zero before retiring them. For the next cutover, I'd lower the TTL to a minute or so at least one old TTL before the change, switch, then raise it once things are stable.”
Shutting the old servers down right after changing the record, as if a DNS change reached everyone instantly.
Measure: split time into DNS, connection, TLS, server time and download for those users.
Edge first: CloudFront for static files and as a front door for dynamic calls.
Network path: Global Accelerator for non-HTTP or latency-sensitive traffic.
Multi-Region: only if needed, with latency-based routing and a clear answer for where data is written.
“I'd measure before changing anything, using real user timing or synthetic checks from that continent. If most of the time is connection setup and downloading assets, distance is the problem, and CloudFront fixes a lot of it: static files come from a nearby edge, and even dynamic requests get faster because TLS ends close to the user and the edge keeps warm connections back to our origin. If server time is the slow part, that's our code or database, and geography isn't the fix. For non-HTTP traffic, Global Accelerator puts users onto the AWS network early. Only if that still isn't enough would I run the app in a second Region with latency-based routing, and then the hard question is where writes go, because the database can't simply be in two places.”
Copying the whole app to a second Region first without measuring where the delay actually is.
Suspect the volume: the Amazon-provided resolver accepts only so many packets per second from each network interface.
Confirm: count DNS queries per instance, and check the network driver's link-local allowance counter.
Fix: a local DNS cache on each instance, so repeated lookups never leave the box.
Fix the app too: reuse connections instead of resolving on every call.
“When DNS fails only under load and every service is healthy, I suspect the resolver limit rather than the services. The Amazon-provided DNS resolver in a VPC only accepts a certain number of packets per second from each network interface, and an app that looks up names on every single request can hit that as traffic climbs. Past the limit, queries get dropped, so lookups time out now and then. I'd confirm it by measuring how many DNS queries each instance sends, and by checking the network driver's link-local allowance counter, which counts packets dropped for exactly this reason. The fix is a local DNS cache on each instance, like systemd-resolved, dnsmasq or unbound, so repeated lookups are answered on the box. I'd also look at the app, because reusing HTTP connections instead of opening a new one per call cuts lookups a lot.”
Blaming the other services and raising timeouts without ever measuring how many lookups the instances send.
Cause: an AMI ID is Regional; the same image has a different ID elsewhere, and your own images exist only where you made them.
Public images: read the current ID from the public Systems Manager parameters AWS publishes.
Your own images: copy them to the new Region and store each Region's ID under the same parameter name.
Check the rest: key pair names, certificate ARNs and typed-in Availability Zone names don't travel either.
“AMI IDs are Regional. The same Amazon Linux image has a different ID in every Region, and our own custom images only exist where we built them. So a hard-coded ID works in one place and fails everywhere else. For AWS's public images, I'd stop hard-coding: CloudFormation can read the current image ID from the public Systems Manager parameters AWS publishes, using an SSM parameter type, so each Region resolves its own ID at deploy time. For our own AMIs, I'd copy them to the new Region as part of the image pipeline and keep each Region's ID in a parameter with the same name. While I'm in there, I'd check other things that don't travel: key pair names, certificate ARNs, and any Availability Zone names typed in instead of looked up.”
Parameters:
LatestAmiId:
Type: AWS::SSM::Parameter::Value<AWS::EC2::Image::Id>
Default: /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64
Resources:
WebServer:
Type: AWS::EC2::Instance
Properties:
ImageId: !Ref LatestAmiId
InstanceType: t3.micro
Keeping a hand-maintained list of image IDs per Region and hoping someone remembers to update it.
Which check failed: system check means the AWS host, fixed by stop and start; instance check means the OS.
Look at the boot: the system log, instance screenshot or serial console show where it stopped.
Rescue: stop it, attach the root volume to a helper instance, fix the file, attach it back.
“The instance status check failing, rather than the system check, tells me the AWS hardware is fine and the problem is inside the OS. So I'd look at how it booted: the system log from the console, the instance screenshot, or the serial console if it's enabled. After an update, common causes are a bad fstab entry for a disk that isn't attached, a kernel that won't boot, or broken network config. If I can fix it through the serial console, great. If not, I stop the instance, detach its root volume, attach it to a healthy helper instance in the same AZ, mount it, fix the file, like adding nofail to that fstab line, then move the volume back and start it. Before any of this I'd take a snapshot so I can't make it worse.”
Terminating the instance straight away without a snapshot when its disk may hold data that exists nowhere else.
Why: each block comes down from S3 the first time it's read, so first reads are slow.
Confirm: high read latency on the volume that improves as the data gets touched.
Warm it: read every block once with fio or dd before sending real traffic.
Skip it: Fast Snapshot Restore in the target Availability Zone gives full performance from the start, at a cost.
“The volume isn't really all there yet. When you create an EBS volume from a snapshot, you can use it straight away, but the blocks are fetched from S3 in the background, and any block you read before it has arrived is pulled on demand. So the first read of each block is slow, and a database that touches lots of data feels terrible until the volume has been read through. I'd confirm it from the volume metrics: high read latency that improves as the working set gets touched. To fix it now, I'd read the whole device once with fio or dd, which forces every block down. For next time, if restore speed matters, I'd turn on Fast Snapshot Restore for that snapshot in the Availability Zone we'd restore into, so new volumes are fast from the start. It costs money while it's on, so I'd keep it for the critical snapshots.”
# read every block once so later reads are fast
sudo fio --filename=/dev/nvme1n1 --rw=read --bs=1M --iodepth=32 \
--ioengine=libaio --direct=1 --name=volume-initialize
Blaming the instance size and upgrading it, when the delay comes from blocks still loading from the snapshot.
What it means: the host has a fault, and on that date the instance will be stopped, or terminated if it runs on instance store.
Check the storage: an EBS-backed instance moves to healthy hardware with a stop and start; instance store data is lost.
Your timing: stop and start in a quiet window before the deadline; a reboot is not enough.
Afterwards: check addresses, tell users, and ask why one instance matters this much.
“A retirement notice means the host underneath has a problem AWS can't fix in place, and on that date the instance will be stopped, or terminated if it's backed by instance store. I'd rather pick the moment myself than have it happen at three in the morning. First I check how it's built. If the root volume is EBS, a stop and start in a quiet window moves it to healthy hardware and the data stays. A reboot doesn't help, because the instance stays on the same host. If anything lives on instance store, like a cache or scratch files, it's gone after a stop, so I'd copy off anything that matters. I'd also check for a public IP without an Elastic IP, since that changes, and warn the service's users about the window. Then the bigger question: why is one instance critical enough that a hardware notice worries us?”
Ignoring the notice until the date, or rebooting and assuming that moved the instance to new hardware.
Data first: automated snapshots and one real test restore, because a lost database can end the company.
Why it went down: the two outages may have a cheap cause, like a full disk or memory running out.
Split the database: move it to RDS with Multi-AZ and automated backups.
Then the app: a launch template and two instances in different zones behind a load balancer.
“I'd rank by what could kill the business, and that's losing the data, not an hour of downtime. So on day one I set up automated EBS snapshots with a lifecycle policy, and I actually restore one to prove it works. Then I'd look at why it went down twice. If it's the disk filling or memory running out, that might be a quick fix that buys breathing room. Next I'd move the database to RDS with Multi-AZ and automated backups, because that removes the scariest single point of failure and takes patching and failover off their plate. If there's time left, I'd make the app server rebuildable from a launch template and run two instances in different zones behind a load balancer. I'd tell the founder plainly what's done, what's left and how the running cost changes, so they can decide.”
Proposing Kubernetes or a full redesign while there is still no backup that has been tested.
Find the weak point: usually the database, not the web tier.
Get ahead of it: scheduled scaling to raise the minimum before airtime, check service quotas.
Take load off: CloudFront caching, read replicas or a cache, queue non-urgent writes.
Plan B: a load test tonight, alarms, and a simple holding page if things tip over.
“Auto Scaling reacts in minutes, and a TV spike hits in seconds, so I'd scale ahead rather than wait. I'd add a scheduled action to raise the group's minimum before airtime, and check our EC2 service quotas so we can actually launch that many. But the web tier is rarely what breaks. The single database is. So I'd put CloudFront in front for static files and anything cacheable, like the landing page, and push heavy reads to a read replica or a cache. Writes that don't need to be instant, like signup emails, can go on a queue. Tonight I'd run a quick load test to see where it bends. And I'd prepare a light static holding page, so if it does tip over, people see something friendly instead of errors.”
Trusting Auto Scaling alone and ignoring the single database behind it.
Confirm: a cold start shows an Init Duration in the function's REPORT log line.
Shrink the start: smaller package, fewer heavy imports, set up clients once outside the handler.
Runtime choice: heavy runtimes start slower; some offer snapshot-based faster starts.
Pay to remove it: provisioned concurrency keeps environments ready.
“That pattern sounds like cold starts. When no environment is warm, Lambda has to create one, load the runtime and run my setup code before the request is handled. I'd confirm it in the logs: cold invocations show an Init Duration on the REPORT line, and I'd see if those match the slow requests. Then I'd shrink the startup. Trim the deployment package, avoid importing big libraries I don't need, and make sure SDK clients and connections are created once outside the handler. More memory also gives more CPU, which can speed up initialisation. If it's a Java function, I'd look at SnapStart. And if the endpoint really can't afford any delay, provisioned concurrency keeps a set number of environments warm, which costs money even when idle, so I'd use it only where users feel it.”
Blaming the network or the database without checking the logs for initialisation time.
Split the path: did the event arrive, did the rule match it, did the call to the function work?
Rule metrics: MatchedEvents and TriggeredRules show matching; FailedInvocations shows the target call failing.
Usual causes: a pattern that doesn't match the real event exactly, a rule in another Region or on the wrong bus, or no permission to invoke the function.
Make failures visible: test the pattern against a real sample event, and give the target a dead-letter queue.
“I'd split it into three hops: did the event reach EventBridge, did my rule match it, and did the call to the function work. The rule's metrics answer most of that. If nothing matched, the pattern is wrong or I'm looking in the wrong place. EC2 state change events go to the default bus in the Region where the instance runs, so a rule in another Region or on a custom bus never sees them. Patterns are also exact: a typo in the detail type, or Stopped with a capital S when the event says stopped, and nothing matches. I'd paste a real sample event into the pattern tester to check. If the rule matched but invocations failed, it's usually permissions: when the rule is created in code rather than the console, someone has to add the permission that lets EventBridge invoke the function. I'd also give the target a dead-letter queue so failures aren't silent.”
Rewriting the function's code first without checking whether the rule ever matched an event.
Tool: Systems Manager Session Manager gives a shell through the SSM agent, no inbound port.
Setup: agent installed, an instance role with the core SSM permissions, and a path to SSM through NAT or VPC endpoints.
Control: IAM decides who can start sessions, and sessions can be logged.
“I'd move everyone to Session Manager. The SSM agent on each instance makes an outbound connection to Systems Manager, so the engineer gets a shell in the browser or the CLI without any inbound port open and without SSH keys. To set it up, the instances need the agent, which many Amazon-provided images already include, an instance role with the core SSM managed policy, and a network path to the service, either through a NAT gateway or through VPC interface endpoints if the subnets are fully private. Access is controlled by IAM, so I can allow only certain teams or certain tagged instances, and session activity can be logged to CloudWatch or S3. It also supports port forwarding, which covers the case where someone needs to reach a private database from their laptop.”
Keeping a bastion host with a shared key as the answer, which is just moving port 22 somewhere else.
Federation: add the CI provider as an OIDC identity provider in IAM.
Role: each run assumes a deploy role and gets temporary credentials.
Tight trust: the trust policy checks the token's audience and its subject, such as the exact repository and branch.
Cut over: switch, watch a few deploys, then deactivate and delete the old key.
“I'd move the pipeline to federation. Most CI systems can issue a short-lived OIDC token for each run. In IAM I add that provider as an OIDC identity provider, then create a deploy role whose trust policy accepts tokens from it. The important part is the conditions: I check the audience, and I check the subject claim so only our repository, and only the main branch or a production environment, can assume the role. Without the subject check, other repositories on the same CI service could try. The pipeline then assumes the role with web identity and gets temporary credentials that expire soon after the run. I'd give the role only what the deploy needs. Once a few deploys pass, I deactivate the old key, watch for anything that breaks, and then delete it.”
Rotating the key every month and calling it solved, when the goal is to have no long-lived key at all.
Assume it is leaked: anyone with the image or repo has it, so it must change.
Move it: store it in Secrets Manager and have the app read it at startup through its role.
Rotate safely: two database users or a grace period so old and new passwords overlap.
Clean up: remove it from the image and config, and turn on scheduled rotation.
“I'd treat the password as leaked, since anyone who ever pulled that image or repo has it, so just moving it isn't enough; it has to change. Step one, I put the current password into Secrets Manager and change the app to fetch it at startup through its instance role, with a short cache and a re-fetch if a login fails. I deploy that while the old value still works. Step two, I rotate. To avoid downtime I'd use the alternating users approach: a second database user gets the new password, the secret switches to it, and once all instances have picked it up, the old user is disabled. Then I build a new image without the password, remove it from the config file, and turn on scheduled rotation so this stays fixed.”
Moving the same password into an environment variable and calling it done without ever changing it.
Act first: cancel the deletion now; the key comes back disabled, so enable it and check what was failing.
Why it's urgent: once the key is deleted, everything encrypted under it is unreadable, backups included.
Find out why: who scheduled it and why, from CloudTrail and a conversation.
Guard it: alert on ScheduleKeyDeletion, and deny it to everyone but a tightly held break-glass role.
“First I cancel the deletion, right now, before anything else. A key that's pending deletion also can't be used, so anything that needs to decrypt with it may already be failing. After cancelling, the key comes back disabled, so I enable it again and check the databases and backup jobs are healthy. The reason to move fast is that once a KMS key is actually deleted, all data encrypted under it, every backup included, can't be decrypted ever again. The waiting period, at least seven days, is the only safety net. Then I'd find out from CloudTrail who scheduled it and ask them why; it might be an honest cleanup mistake. Finally I'd make it hard to repeat: an EventBridge rule that alerts on any ScheduleKeyDeletion call, and a key policy or SCP that denies it to everyone except a break-glass role.”
Assuming AWS can bring back a deleted key, or finishing the investigation before cancelling the deletion.
Contain: swap to an isolation security group, detach from the load balancer and Auto Scaling group.
Keep evidence: snapshot the volumes before anything is terminated.
Cut credentials: revoke active sessions on the instance role and check what it did.
Find the way in: CloudTrail, flow logs and app logs, then rebuild from a clean image.
“I'd treat it as a real compromise. First, contain it without destroying it. I take it out of the load balancer and the Auto Scaling group, so it isn't replaced and terminated automatically, and swap its security group for an isolation group with no rules. Existing tracked connections can survive a security group change, so I might add a network ACL deny as well. Then I snapshot its volumes for evidence. The attacker may have used the instance role, so I revoke its active sessions and check CloudTrail for what those credentials did, like launching more instances in other Regions. Then I look for the way in: an unpatched app, an exposed port, a leaked key. The server itself gets rebuilt from a clean image, never cleaned in place. And I'd tell the security lead early.”
Killing the mining process, rebooting, and calling it fixed without checking how the attacker got in.
Two zones: uploads land in a quarantine bucket nothing can hand out links for.
Scan on arrival: an object event triggers the scanner, in Lambda for small files, a container for big ones, or a managed scanning feature.
Promote or block: clean files move to the serving bucket; infected ones go to a locked bucket and raise an alert.
Fail closed: a scan that errors or times out leaves the file blocked, and the user sees it as processing.
“I'd keep scanned and unscanned files physically apart. Uploads go straight from the browser into a quarantine bucket through presigned URLs, and nothing in the app can hand out download links for that bucket. A new object event triggers the scan. I could run an antivirus engine in a Lambda function for small files, or in a container task for big ones, or use GuardDuty's malware protection for S3 if I'd rather not maintain the engine and its signatures. Clean files are copied to the serving bucket and the database record flips to available. Infected files move to a locked-down bucket, and security gets an alert. Two details matter. A scan that errors or times out must leave the file blocked, never marked clean. And the user needs feedback, so the file shows as processing until the scan finishes instead of failing silently.”
Letting files be downloadable while the scan runs, or treating a failed scan as a pass.
Understand the goal: how much saving is needed and why now.
Show the risk plainly: what an AZ problem would do to customers and for how long.
Offer options: smaller instances in both AZs, savings elsewhere, or single AZ only for internal tools.
Record it: if overruled, write down the decision and the accepted risk.
“I'd start by asking what saving we're aiming for, because the goal might be reachable another way. Then I'd lay out the risk in plain terms: if that one zone has trouble, the whole app is down until it recovers, and customers see it. I'd bring options rather than just a no. We could run fewer or smaller instances but keep them spread across both zones, since most of the cost comes from how many instances we run and how big they are, not from how many zones they sit in. We could look for savings in idle dev environments, old snapshots or commitments for steady workloads. Or we could move internal tools to one zone and keep customer-facing parts in two. If my manager still decides on one zone, that's their call, and I'd support it, but I'd write the decision and the accepted risk down so nobody is surprised later.”
Either agreeing without mentioning the risk, or refusing outright with no alternatives.
See what exists: tag every resource with an owner and environment, find the untagged ones.
Schedule: stop instances and databases outside working hours, with an easy way to start early.
Trim: right-size oversized instances, delete unattached volumes and old snapshots.
Talk first: agree the schedule with the team so nothing breaks mid-test.
“I'd start by getting a clear picture: tag everything with an owner and an environment, and chase whatever isn't tagged, because unknown resources are usually the waste. Then the big win is scheduling. If people only work office hours, most dev and test instances and databases can stop in the evening and on weekends. I'd use a scheduler, something like EventBridge Scheduler triggering a small function, and give the team an easy way to start things early or skip a night for a long test. Next I'd right-size, since dev boxes are often copies of production size, and delete unattached EBS volumes and old snapshots. The key is agreeing the schedule with the developers first, so nobody loses a half-finished test run at seven in the evening.”
Shutting down environments without warning the team, or deleting untagged resources without finding their owner.
Users first: if the console change is the fastest safe fix, make it, with a second person watching if possible.
Leave a trail: say in the incident channel exactly what you changed and when.
Back into code: the next working day, the same change goes into the template, then check drift.
Fix the gap: ask why the pipeline couldn't ship the fix in time.
“I'd fix production. The infrastructure as code rule exists to keep things reliable, and it shouldn't keep users down longer. But I'd do it carefully: I post in the incident channel what I'm about to change, get a second pair of eyes if anyone's around, make the smallest change that works, and write down exactly what I did and when. The danger with console changes is that the code no longer matches reality, so the next deploy quietly undoes the fix or fails halfway. That makes the follow-up as important as the fix. The next working day I put the same change into the template, run drift detection to confirm they match, and note it in the incident review. And I'd ask why the pipeline couldn't get a fix out fast enough, because that's the real gap.”
Either leaving users down to follow the rule, or making the console change and never putting it back into code.
Read the reason: a stopped task's reason says whether it's a network timeout, an access error or a missing image.
Network path: private subnets need a NAT gateway, or VPC endpoints for the ECR API, the ECR registry and S3, where the layers live.
Right role: the task execution role pulls the image and writes logs, not the task role.
The image: the tag exists, and it was built for the CPU architecture the task uses.
“First I'd open a stopped task and read its stopped reason, because it tells me whether this is a network problem or a permissions problem. If it's a timeout, the task can't reach ECR at all. In private subnets, Fargate needs either a route to a NAT gateway or VPC endpoints. With endpoints, people often add the ECR API one and forget the rest: you also need the ECR Docker registry endpoint and an S3 gateway endpoint, because the image layers are actually stored in S3. If logs go to CloudWatch, a logs endpoint too, and the endpoints' security group has to allow HTTPS from the tasks. If the reason says access denied, I'd check the task execution role, not the task role, since that's the one that pulls the image. And if it says there's no matching manifest, the image was built for a different CPU architecture.”
Giving the task role admin rights, when the pull is done by the execution role and the real problem is often the network path.
How pods get IPs: the VPC CNI gives each pod an address from the node's subnet.
Two limits: the subnet running out of free addresses, and each instance type's cap on network interfaces and addresses.
Checks: free addresses in the subnets, the node's max pods, the CNI logs.
Fixes: prefix delegation for the per-node cap; a secondary CIDR with custom networking for a full subnet.
“On EKS with the default VPC CNI, every pod gets a real IP address from the subnet its node sits in, so pods can run out of addresses long before the nodes run out of CPU. There are two separate limits. One is the subnet: a small subnet with lots of pods simply has no free addresses left, and new pods get stuck creating with an IP assignment error. The other is per node: each instance type can hold only so many network interfaces and addresses, which caps how many pods it runs. I'd check free addresses in the subnets and the node's max pods, then read the CNI logs. If it's the node cap, prefix delegation hands each interface small blocks of addresses so a node fits many more pods. If the subnet is full, I'd add a secondary CIDR range to the VPC and use custom networking so pods draw from bigger subnets.”
Adding more nodes, which only takes even more addresses from the same full subnet.
Which queries: Performance Insights, CloudWatch Database Insights or the slow query log shows the top SQL by load.
What changed: data growth, a new report, a batch job, a traffic pattern, stale statistics.
Instance type: a burstable instance may have run out of CPU credits.
Short term vs lasting: scale up or kill a runaway query now, then fix the index or query.
“No deploy doesn't mean nothing changed, so I'd open Performance Insights or Database Insights and look at the top SQL by load for that hour. Usually one or two statements stand out. If it's a query that got slow because a table grew past the point where a missing index hurts, I'd check its plan and add the index. If it's a scheduled report someone started running at peak, I'd move it to a read replica or off hours. I'd also check the instance class: if it's a burstable type, it may simply have spent its CPU credits and dropped to baseline. And I'd look at connection count in case the app is retrying and piling on. If users are suffering right now, killing a runaway query or scaling up is fine as a stopgap, but I'd still find the query.”
Jumping straight to a bigger instance without ever looking at which queries are using the CPU.
What happened: every request missed at once and went to a database sized only for the misses.
Right now: shed or limit load, and warm the hottest keys before letting full traffic back.
Survive a node loss: a replica with automatic failover, so the cache doesn't restart cold.
Stop the herd: one request rebuilds a key while others wait or get the old value; add jitter to expiry times.
“That's a cache stampede. The database was quietly sized for the few requests that miss the cache, and when the node came back empty, every request missed at the same moment and hit the database together. First I'd protect the database: limit traffic or serve a lighter version of the busiest pages, and warm the hottest keys with a script before letting full traffic back. For the lasting fix, I'd run the cache with a replica and automatic failover, so losing one node promotes a replica that already holds the data instead of starting cold. In the code, I'd make sure only one request rebuilds a missing key, using a short lock, while the others wait briefly or get a slightly stale value. And I'd add random jitter to expiry times, because keys that all expire together cause a smaller version of the same stampede.”
Just making the cache node bigger, which does nothing for the moment it comes back empty.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.