Troubleshooting • Networking • Security Incidents • Databases • Serverless • Containers • Resilience • 2026

Scenario-Based AWS Interview Questions

Scenario rounds hand you a symptom and watch how you think. Pods can't get an IP address, an event rule never fires, a cache restart takes the database down, an alert says a server is mining crypto. There is rarely one right answer. The interviewer wants the order you check things in, what you would look at first, and what would change your mind. This page is for anyone facing that round, from a first cloud job to a senior hire. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying the first three checks before the fix.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Networking 3 questions

Easy Technical round Fresher, Mid-level Practice question

1. You pointed the Route 53 record for your site at a new load balancer an hour ago, but some users still reach the old servers. What's happening, and what do you do?

What the interviewer is really testing:
Whether you understand DNS caching and TTLs, and plan a cutover so the old side keeps working until caches run out.
Answer frame:

Why: resolvers and clients keep the old answer until its TTL runs out, and a few hold it longer.

Check: look the name up through several public resolvers and read the TTL they return.

Right now: keep the old servers serving, or forwarding to the new ones, until their traffic dies off.

Next time: lower the TTL well ahead of the change, switch, then raise it again.

Sample spoken answer:

“That's DNS caching doing its job. When a resolver looks up our name, it keeps the answer for as long as the record's TTL says. If the TTL was a day, some users won't see the change for up to a day, and a few clients with their own cache hold on even longer. I'd confirm it by querying the name through a few public resolvers with dig and looking at the TTL that's left. The important thing right now is not to switch the old servers off. I'd keep them serving, or have them forward to the new load balancer, and watch their traffic fall to zero before retiring them. For the next cutover, I'd lower the TTL to a minute or so at least one old TTL before the change, switch, then raise it once things are stable.”

Red flag to avoid:

Shutting the old servers down right after changing the record, as if a DNS change reached everyone instantly.

They may ask next:
  • Why can't you put a CNAME on the bare domain, and what does Route 53 offer instead?
  • How would you move traffic to the new load balancer gradually instead of all at once?
Say it in 60 seconds
Medium Case round Mid-level, Senior Practice question

2. The app is hosted in one Region. Users on another continent say it feels slow, while local users are fine. How do you investigate and improve it?

What the interviewer is really testing:
Whether you measure where the time goes before choosing between caching at the edge and running in more Regions.
Answer frame:

Measure: split time into DNS, connection, TLS, server time and download for those users.

Edge first: CloudFront for static files and as a front door for dynamic calls.

Network path: Global Accelerator for non-HTTP or latency-sensitive traffic.

Multi-Region: only if needed, with latency-based routing and a clear answer for where data is written.

Sample spoken answer:

“I'd measure before changing anything, using real user timing or synthetic checks from that continent. If most of the time is connection setup and downloading assets, distance is the problem, and CloudFront fixes a lot of it: static files come from a nearby edge, and even dynamic requests get faster because TLS ends close to the user and the edge keeps warm connections back to our origin. If server time is the slow part, that's our code or database, and geography isn't the fix. For non-HTTP traffic, Global Accelerator puts users onto the AWS network early. Only if that still isn't enough would I run the app in a second Region with latency-based routing, and then the hard question is where writes go, because the database can't simply be in two places.”

Red flag to avoid:

Copying the whole app to a second Region first without measuring where the delay actually is.

They may ask next:
  • What can CloudFront do for API calls that can't be cached?
  • How would you handle user data if you ran active in two Regions?
Say it in 60 seconds
Hard Technical round Senior Practice question

3. Under heavy load, a fleet of EC2 instances starts logging intermittent DNS lookup failures when calling other services, though nothing is actually down. What could cause that, and how do you fix it?

What the interviewer is really testing:
Whether you know the VPC resolver has a per-interface packet limit and that local caching takes the pressure off.
Answer frame:

Suspect the volume: the Amazon-provided resolver accepts only so many packets per second from each network interface.

Confirm: count DNS queries per instance, and check the network driver's link-local allowance counter.

Fix: a local DNS cache on each instance, so repeated lookups never leave the box.

Fix the app too: reuse connections instead of resolving on every call.

Sample spoken answer:

“When DNS fails only under load and every service is healthy, I suspect the resolver limit rather than the services. The Amazon-provided DNS resolver in a VPC only accepts a certain number of packets per second from each network interface, and an app that looks up names on every single request can hit that as traffic climbs. Past the limit, queries get dropped, so lookups time out now and then. I'd confirm it by measuring how many DNS queries each instance sends, and by checking the network driver's link-local allowance counter, which counts packets dropped for exactly this reason. The fix is a local DNS cache on each instance, like systemd-resolved, dnsmasq or unbound, so repeated lookups are answered on the box. I'd also look at the app, because reusing HTTP connections instead of opening a new one per call cuts lookups a lot.”

Red flag to avoid:

Blaming the other services and raising timeouts without ever measuring how many lookups the instances send.

They may ask next:
  • Which other services on an instance share that same link-local limit?
  • How would you handle this for containers, where every pod does its own lookups?
Say it in 60 seconds

Troubleshooting 3 questions

Easy Technical round Fresher, Mid-level Practice question

4. Your CloudFormation template works in one Region, but deploying it to a second Region fails saying the image ID doesn't exist. What's wrong, and how do you make the template portable?

What the interviewer is really testing:
Whether you know that AMI IDs belong to one Region, and how to look up the right image instead of hard-coding it.
Answer frame:

Cause: an AMI ID is Regional; the same image has a different ID elsewhere, and your own images exist only where you made them.

Public images: read the current ID from the public Systems Manager parameters AWS publishes.

Your own images: copy them to the new Region and store each Region's ID under the same parameter name.

Check the rest: key pair names, certificate ARNs and typed-in Availability Zone names don't travel either.

Sample spoken answer:

“AMI IDs are Regional. The same Amazon Linux image has a different ID in every Region, and our own custom images only exist where we built them. So a hard-coded ID works in one place and fails everywhere else. For AWS's public images, I'd stop hard-coding: CloudFormation can read the current image ID from the public Systems Manager parameters AWS publishes, using an SSM parameter type, so each Region resolves its own ID at deploy time. For our own AMIs, I'd copy them to the new Region as part of the image pipeline and keep each Region's ID in a parameter with the same name. While I'm in there, I'd check other things that don't travel: key pair names, certificate ARNs, and any Availability Zone names typed in instead of looked up.”

Code:
Parameters:
  LatestAmiId:
    Type: AWS::SSM::Parameter::Value<AWS::EC2::Image::Id>
    Default: /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64
Resources:
  WebServer:
    Type: AWS::EC2::Instance
    Properties:
      ImageId: !Ref LatestAmiId
      InstanceType: t3.micro
Red flag to avoid:

Keeping a hand-maintained list of image IDs per Region and hoping someone remembers to update it.

They may ask next:
  • Why might always taking the newest image be risky for production servers?
  • How would you share a custom AMI with other accounts as well as other Regions?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

5. After a routine OS update and reboot, an EC2 instance fails its instance status check and you can't connect at all. How do you get it back?

What the interviewer is really testing:
Whether you know the difference between system and instance status checks, and how to repair a server that will not boot.
Answer frame:

Which check failed: system check means the AWS host, fixed by stop and start; instance check means the OS.

Look at the boot: the system log, instance screenshot or serial console show where it stopped.

Rescue: stop it, attach the root volume to a helper instance, fix the file, attach it back.

Sample spoken answer:

“The instance status check failing, rather than the system check, tells me the AWS hardware is fine and the problem is inside the OS. So I'd look at how it booted: the system log from the console, the instance screenshot, or the serial console if it's enabled. After an update, common causes are a bad fstab entry for a disk that isn't attached, a kernel that won't boot, or broken network config. If I can fix it through the serial console, great. If not, I stop the instance, detach its root volume, attach it to a healthy helper instance in the same AZ, mount it, fix the file, like adding nofail to that fstab line, then move the volume back and start it. Before any of this I'd take a snapshot so I can't make it worse.”

Red flag to avoid:

Terminating the instance straight away without a snapshot when its disk may hold data that exists nowhere else.

They may ask next:
  • What would you do differently if the system status check had failed instead?
  • How would you design this server so a broken instance can just be replaced?
Say it in 60 seconds
Hard Technical round Senior Practice question

6. You restored a database server from an EBS snapshot during an incident. It boots fine, but queries are painfully slow for the first few hours. Why, and how would you avoid it next time?

What the interviewer is really testing:
Whether you know that a volume created from a snapshot loads its blocks lazily, and the ways to pay that cost before users feel it.
Answer frame:

Why: each block comes down from S3 the first time it's read, so first reads are slow.

Confirm: high read latency on the volume that improves as the data gets touched.

Warm it: read every block once with fio or dd before sending real traffic.

Skip it: Fast Snapshot Restore in the target Availability Zone gives full performance from the start, at a cost.

Sample spoken answer:

“The volume isn't really all there yet. When you create an EBS volume from a snapshot, you can use it straight away, but the blocks are fetched from S3 in the background, and any block you read before it has arrived is pulled on demand. So the first read of each block is slow, and a database that touches lots of data feels terrible until the volume has been read through. I'd confirm it from the volume metrics: high read latency that improves as the working set gets touched. To fix it now, I'd read the whole device once with fio or dd, which forces every block down. For next time, if restore speed matters, I'd turn on Fast Snapshot Restore for that snapshot in the Availability Zone we'd restore into, so new volumes are fast from the start. It costs money while it's on, so I'd keep it for the critical snapshots.”

Code:
# read every block once so later reads are fast
sudo fio --filename=/dev/nvme1n1 --rw=read --bs=1M --iodepth=32 \
  --ioengine=libaio --direct=1 --name=volume-initialize
Red flag to avoid:

Blaming the instance size and upgrading it, when the delay comes from blocks still loading from the snapshot.

They may ask next:
  • Does the same thing happen when you restore an RDS database from a snapshot?
  • How would you build this delay into the recovery time you promise the business?
Say it in 60 seconds

Resilience 3 questions

Easy Situational round Fresher, Mid-level Practice question

7. AWS emails that the EC2 instance running a critical internal service is scheduled for retirement in two weeks because of a hardware problem. What do you do?

What the interviewer is really testing:
Whether you know what a retirement notice means for your data, and act on your own schedule instead of waiting for the deadline.
Answer frame:

What it means: the host has a fault, and on that date the instance will be stopped, or terminated if it runs on instance store.

Check the storage: an EBS-backed instance moves to healthy hardware with a stop and start; instance store data is lost.

Your timing: stop and start in a quiet window before the deadline; a reboot is not enough.

Afterwards: check addresses, tell users, and ask why one instance matters this much.

Sample spoken answer:

“A retirement notice means the host underneath has a problem AWS can't fix in place, and on that date the instance will be stopped, or terminated if it's backed by instance store. I'd rather pick the moment myself than have it happen at three in the morning. First I check how it's built. If the root volume is EBS, a stop and start in a quiet window moves it to healthy hardware and the data stays. A reboot doesn't help, because the instance stays on the same host. If anything lives on instance store, like a cache or scratch files, it's gone after a stop, so I'd copy off anything that matters. I'd also check for a public IP without an Elastic IP, since that changes, and warn the service's users about the window. Then the bigger question: why is one instance critical enough that a hardware notice worries us?”

Red flag to avoid:

Ignoring the notice until the date, or rebooting and assuming that moved the instance to new hardware.

They may ask next:
  • What would you do differently if the instance were backed by instance store?
  • How would you set up alerts so notices like this don't sit unread in someone's inbox?
Say it in 60 seconds
Medium Case round Fresher, Mid-level Practice question

8. A small startup runs its whole product, app and database, on one EC2 instance. It went down twice this month, and the founder asks what to fix first with only a week of your time. What do you do?

What the interviewer is really testing:
Whether you rank fixes by risk to the business, protecting the data first and then removing single points of failure in a sensible order.
Answer frame:

Data first: automated snapshots and one real test restore, because a lost database can end the company.

Why it went down: the two outages may have a cheap cause, like a full disk or memory running out.

Split the database: move it to RDS with Multi-AZ and automated backups.

Then the app: a launch template and two instances in different zones behind a load balancer.

Sample spoken answer:

“I'd rank by what could kill the business, and that's losing the data, not an hour of downtime. So on day one I set up automated EBS snapshots with a lifecycle policy, and I actually restore one to prove it works. Then I'd look at why it went down twice. If it's the disk filling or memory running out, that might be a quick fix that buys breathing room. Next I'd move the database to RDS with Multi-AZ and automated backups, because that removes the scariest single point of failure and takes patching and failover off their plate. If there's time left, I'd make the app server rebuildable from a launch template and run two instances in different zones behind a load balancer. I'd tell the founder plainly what's done, what's left and how the running cost changes, so they can decide.”

Red flag to avoid:

Proposing Kubernetes or a full redesign while there is still no backup that has been tested.

They may ask next:
  • How would you move the database to RDS without a long outage for an app this small?
  • What would you leave out if you only had two days?
Say it in 60 seconds
Medium Case round Mid-level, Senior Practice question

9. Marketing tells you a TV spot airs tomorrow night and traffic could be many times normal for an hour. The app runs on EC2 behind a load balancer with one RDS database. What do you do today?

What the interviewer is really testing:
Whether you find the real bottleneck, prepare capacity ahead of a sudden spike, and have a plan if it still breaks.
Answer frame:

Find the weak point: usually the database, not the web tier.

Get ahead of it: scheduled scaling to raise the minimum before airtime, check service quotas.

Take load off: CloudFront caching, read replicas or a cache, queue non-urgent writes.

Plan B: a load test tonight, alarms, and a simple holding page if things tip over.

Sample spoken answer:

“Auto Scaling reacts in minutes, and a TV spike hits in seconds, so I'd scale ahead rather than wait. I'd add a scheduled action to raise the group's minimum before airtime, and check our EC2 service quotas so we can actually launch that many. But the web tier is rarely what breaks. The single database is. So I'd put CloudFront in front for static files and anything cacheable, like the landing page, and push heavy reads to a read replica or a cache. Writes that don't need to be instant, like signup emails, can go on a queue. Tonight I'd run a quick load test to see where it bends. And I'd prepare a light static holding page, so if it does tip over, people see something friendly instead of errors.”

Red flag to avoid:

Trusting Auto Scaling alone and ignoring the single database behind it.

They may ask next:
  • How would you load test this without hurting real users?
  • What alarms would you watch during that hour?
Say it in 60 seconds

Serverless 2 questions

Easy Technical round Fresher, Mid-level Practice question

10. An API on Lambda answers quickly most of the time, but the first request after a quiet period takes several seconds. Users notice. What do you check and change?

What the interviewer is really testing:
Whether you recognise cold starts, can confirm them from the logs, and know the options from cheap code fixes to paid warm capacity.
Answer frame:

Confirm: a cold start shows an Init Duration in the function's REPORT log line.

Shrink the start: smaller package, fewer heavy imports, set up clients once outside the handler.

Runtime choice: heavy runtimes start slower; some offer snapshot-based faster starts.

Pay to remove it: provisioned concurrency keeps environments ready.

Sample spoken answer:

“That pattern sounds like cold starts. When no environment is warm, Lambda has to create one, load the runtime and run my setup code before the request is handled. I'd confirm it in the logs: cold invocations show an Init Duration on the REPORT line, and I'd see if those match the slow requests. Then I'd shrink the startup. Trim the deployment package, avoid importing big libraries I don't need, and make sure SDK clients and connections are created once outside the handler. More memory also gives more CPU, which can speed up initialisation. If it's a Java function, I'd look at SnapStart. And if the endpoint really can't afford any delay, provisioned concurrency keeps a set number of environments warm, which costs money even when idle, so I'd use it only where users feel it.”

Red flag to avoid:

Blaming the network or the database without checking the logs for initialisation time.

They may ask next:
  • Why is scheduling a ping every few minutes a weak fix for cold starts?
  • How would you decide how much provisioned concurrency to buy?
Say it in 60 seconds
Easy Technical round Fresher, Mid-level Practice question

11. You added an EventBridge rule to run a Lambda function whenever an EC2 instance stops. You stopped a test instance and nothing happened. How do you find where it broke?

What the interviewer is really testing:
Whether you debug an event pipeline hop by hop, and know the usual reasons a rule never matches or never reaches its target.
Answer frame:

Split the path: did the event arrive, did the rule match it, did the call to the function work?

Rule metrics: MatchedEvents and TriggeredRules show matching; FailedInvocations shows the target call failing.

Usual causes: a pattern that doesn't match the real event exactly, a rule in another Region or on the wrong bus, or no permission to invoke the function.

Make failures visible: test the pattern against a real sample event, and give the target a dead-letter queue.

Sample spoken answer:

“I'd split it into three hops: did the event reach EventBridge, did my rule match it, and did the call to the function work. The rule's metrics answer most of that. If nothing matched, the pattern is wrong or I'm looking in the wrong place. EC2 state change events go to the default bus in the Region where the instance runs, so a rule in another Region or on a custom bus never sees them. Patterns are also exact: a typo in the detail type, or Stopped with a capital S when the event says stopped, and nothing matches. I'd paste a real sample event into the pattern tester to check. If the rule matched but invocations failed, it's usually permissions: when the rule is created in code rather than the console, someone has to add the permission that lets EventBridge invoke the function. I'd also give the target a dead-letter queue so failures aren't silent.”

Red flag to avoid:

Rewriting the function's code first without checking whether the rule ever matched an event.

They may ask next:
  • How would you check the function's side, to see whether it was invoked and failed?
  • How would you test a new rule safely without stopping real instances?
Say it in 60 seconds

Security 6 questions

Easy Technical round Fresher, Mid-level Practice question

12. Security wants port 22 closed on every server and all SSH keys removed. Engineers still need shell access to debug. How do you make that work?

What the interviewer is really testing:
Whether you know how to give audited shell access without open ports or shared keys.
Answer frame:

Tool: Systems Manager Session Manager gives a shell through the SSM agent, no inbound port.

Setup: agent installed, an instance role with the core SSM permissions, and a path to SSM through NAT or VPC endpoints.

Control: IAM decides who can start sessions, and sessions can be logged.

Sample spoken answer:

“I'd move everyone to Session Manager. The SSM agent on each instance makes an outbound connection to Systems Manager, so the engineer gets a shell in the browser or the CLI without any inbound port open and without SSH keys. To set it up, the instances need the agent, which many Amazon-provided images already include, an instance role with the core SSM managed policy, and a network path to the service, either through a NAT gateway or through VPC interface endpoints if the subnets are fully private. Access is controlled by IAM, so I can allow only certain teams or certain tagged instances, and session activity can be logged to CloudWatch or S3. It also supports port forwarding, which covers the case where someone needs to reach a private database from their laptop.”

Red flag to avoid:

Keeping a bastion host with a shared key as the answer, which is just moving port 22 somewhere else.

They may ask next:
  • How would you restrict one team to only their own instances?
  • What would you check if an instance doesn't show up in Session Manager?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

13. Your CI pipeline deploys to AWS with an access key stored as a secret in the CI system. Security wants every long-lived key gone. How do you keep deployments working?

What the interviewer is really testing:
Whether you know how a pipeline can get short-lived credentials through federation, and how to scope the trust so only the right pipeline gets in.
Answer frame:

Federation: add the CI provider as an OIDC identity provider in IAM.

Role: each run assumes a deploy role and gets temporary credentials.

Tight trust: the trust policy checks the token's audience and its subject, such as the exact repository and branch.

Cut over: switch, watch a few deploys, then deactivate and delete the old key.

Sample spoken answer:

“I'd move the pipeline to federation. Most CI systems can issue a short-lived OIDC token for each run. In IAM I add that provider as an OIDC identity provider, then create a deploy role whose trust policy accepts tokens from it. The important part is the conditions: I check the audience, and I check the subject claim so only our repository, and only the main branch or a production environment, can assume the role. Without the subject check, other repositories on the same CI service could try. The pipeline then assumes the role with web identity and gets temporary credentials that expire soon after the run. I'd give the role only what the deploy needs. Once a few deploys pass, I deactivate the old key, watch for anything that breaks, and then delete it.”

Red flag to avoid:

Rotating the key every month and calling it solved, when the goal is to have no long-lived key at all.

They may ask next:
  • What goes wrong if the trust policy checks the audience but not the subject?
  • How would the same pipeline deploy into several AWS accounts?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

14. You discover the production database password is baked into the server image and a config file in the repo. You're asked to fix it without taking the app down. How?

What the interviewer is really testing:
Whether you treat the password as already leaked, move it to a proper secret store, and rotate without an outage.
Answer frame:

Assume it is leaked: anyone with the image or repo has it, so it must change.

Move it: store it in Secrets Manager and have the app read it at startup through its role.

Rotate safely: two database users or a grace period so old and new passwords overlap.

Clean up: remove it from the image and config, and turn on scheduled rotation.

Sample spoken answer:

“I'd treat the password as leaked, since anyone who ever pulled that image or repo has it, so just moving it isn't enough; it has to change. Step one, I put the current password into Secrets Manager and change the app to fetch it at startup through its instance role, with a short cache and a re-fetch if a login fails. I deploy that while the old value still works. Step two, I rotate. To avoid downtime I'd use the alternating users approach: a second database user gets the new password, the secret switches to it, and once all instances have picked it up, the old user is disabled. Then I build a new image without the password, remove it from the config file, and turn on scheduled rotation so this stays fixed.”

Red flag to avoid:

Moving the same password into an environment variable and calling it done without ever changing it.

They may ask next:
  • Removing the password from the repo leaves it in the history. Does that matter?
  • How would the app notice the secret changed while it's running?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

15. You spot in CloudTrail that someone scheduled a customer managed KMS key for deletion three days ago. That key encrypts production databases and their backups. What do you do?

What the interviewer is really testing:
Whether you know that a deleted KMS key makes data unreadable for good, move fast inside the waiting period, and then stop it happening again.
Answer frame:

Act first: cancel the deletion now; the key comes back disabled, so enable it and check what was failing.

Why it's urgent: once the key is deleted, everything encrypted under it is unreadable, backups included.

Find out why: who scheduled it and why, from CloudTrail and a conversation.

Guard it: alert on ScheduleKeyDeletion, and deny it to everyone but a tightly held break-glass role.

Sample spoken answer:

“First I cancel the deletion, right now, before anything else. A key that's pending deletion also can't be used, so anything that needs to decrypt with it may already be failing. After cancelling, the key comes back disabled, so I enable it again and check the databases and backup jobs are healthy. The reason to move fast is that once a KMS key is actually deleted, all data encrypted under it, every backup included, can't be decrypted ever again. The waiting period, at least seven days, is the only safety net. Then I'd find out from CloudTrail who scheduled it and ask them why; it might be an honest cleanup mistake. Finally I'd make it hard to repeat: an EventBridge rule that alerts on any ScheduleKeyDeletion call, and a key policy or SCP that denies it to everyone except a break-glass role.”

Red flag to avoid:

Assuming AWS can bring back a deleted key, or finishing the investigation before cancelling the deletion.

They may ask next:
  • Why might you disable a key for a while instead of deleting it?
  • How would you have noticed sooner if nobody was watching CloudTrail?
Say it in 60 seconds
Hard Situational round Mid-level, Senior Practice question

16. GuardDuty raises a high-severity alert: one of your EC2 instances is talking to a known crypto-mining pool. Walk me through the next hour.

What the interviewer is really testing:
Whether you contain a compromised instance while keeping evidence, and look beyond the one server to how the attacker got in.
Answer frame:

Contain: swap to an isolation security group, detach from the load balancer and Auto Scaling group.

Keep evidence: snapshot the volumes before anything is terminated.

Cut credentials: revoke active sessions on the instance role and check what it did.

Find the way in: CloudTrail, flow logs and app logs, then rebuild from a clean image.

Sample spoken answer:

“I'd treat it as a real compromise. First, contain it without destroying it. I take it out of the load balancer and the Auto Scaling group, so it isn't replaced and terminated automatically, and swap its security group for an isolation group with no rules. Existing tracked connections can survive a security group change, so I might add a network ACL deny as well. Then I snapshot its volumes for evidence. The attacker may have used the instance role, so I revoke its active sessions and check CloudTrail for what those credentials did, like launching more instances in other Regions. Then I look for the way in: an unpatched app, an exposed port, a leaked key. The server itself gets rebuilt from a clean image, never cleaned in place. And I'd tell the security lead early.”

Red flag to avoid:

Killing the mining process, rebooting, and calling it fixed without checking how the attacker got in.

They may ask next:
  • Why not just terminate the instance straight away?
  • What would you look for in CloudTrail after a role's credentials were stolen?
Say it in 60 seconds
Hard System design round Mid-level, Senior Practice question

17. Users upload files that other users download. Security says nothing may be downloadable until it has been scanned for malware. Design this on AWS.

What the interviewer is really testing:
Whether you keep unscanned files physically apart, make the unsafe state the default, and think about failures and user feedback.
Answer frame:

Two zones: uploads land in a quarantine bucket nothing can hand out links for.

Scan on arrival: an object event triggers the scanner, in Lambda for small files, a container for big ones, or a managed scanning feature.

Promote or block: clean files move to the serving bucket; infected ones go to a locked bucket and raise an alert.

Fail closed: a scan that errors or times out leaves the file blocked, and the user sees it as processing.

Sample spoken answer:

“I'd keep scanned and unscanned files physically apart. Uploads go straight from the browser into a quarantine bucket through presigned URLs, and nothing in the app can hand out download links for that bucket. A new object event triggers the scan. I could run an antivirus engine in a Lambda function for small files, or in a container task for big ones, or use GuardDuty's malware protection for S3 if I'd rather not maintain the engine and its signatures. Clean files are copied to the serving bucket and the database record flips to available. Infected files move to a locked-down bucket, and security gets an alert. Two details matter. A scan that errors or times out must leave the file blocked, never marked clean. And the user needs feedback, so the file shows as processing until the scan finishes instead of failing silently.”

Red flag to avoid:

Letting files be downloadable while the scan runs, or treating a failed scan as a pass.

They may ask next:
  • How would you handle a file that's too large for the scanner to finish in time?
  • What would you do about files that were marked clean before a new signature showed they were malware?
Say it in 60 seconds

Cost and Operations 3 questions

Easy Situational round Fresher, Mid-level, Senior Practice question

18. To cut costs, your manager wants to move a customer-facing app from two Availability Zones to one. You think that's risky. How do you handle it?

What the interviewer is really testing:
Whether you can push back with facts and options instead of just saying no, and accept a decision that is not yours.
Answer frame:

Understand the goal: how much saving is needed and why now.

Show the risk plainly: what an AZ problem would do to customers and for how long.

Offer options: smaller instances in both AZs, savings elsewhere, or single AZ only for internal tools.

Record it: if overruled, write down the decision and the accepted risk.

Sample spoken answer:

“I'd start by asking what saving we're aiming for, because the goal might be reachable another way. Then I'd lay out the risk in plain terms: if that one zone has trouble, the whole app is down until it recovers, and customers see it. I'd bring options rather than just a no. We could run fewer or smaller instances but keep them spread across both zones, since most of the cost comes from how many instances we run and how big they are, not from how many zones they sit in. We could look for savings in idle dev environments, old snapshots or commitments for steady workloads. Or we could move internal tools to one zone and keep customer-facing parts in two. If my manager still decides on one zone, that's their call, and I'd support it, but I'd write the decision and the accepted risk down so nobody is surprised later.”

Red flag to avoid:

Either agreeing without mentioning the risk, or refusing outright with no alternatives.

They may ask next:
  • What would you say if your manager asked for the chance of an AZ outage?
  • Which parts of the app could safely run in one zone?
Say it in 60 seconds
Easy Situational round Fresher, Mid-level Practice question

19. Dev and test environments run all day and night, but the team only works office hours. You're asked to cut their cost without slowing anyone down. What do you do?

What the interviewer is really testing:
Whether you find waste methodically and involve the team, rather than switching things off and breaking someone's work.
Answer frame:

See what exists: tag every resource with an owner and environment, find the untagged ones.

Schedule: stop instances and databases outside working hours, with an easy way to start early.

Trim: right-size oversized instances, delete unattached volumes and old snapshots.

Talk first: agree the schedule with the team so nothing breaks mid-test.

Sample spoken answer:

“I'd start by getting a clear picture: tag everything with an owner and an environment, and chase whatever isn't tagged, because unknown resources are usually the waste. Then the big win is scheduling. If people only work office hours, most dev and test instances and databases can stop in the evening and on weekends. I'd use a scheduler, something like EventBridge Scheduler triggering a small function, and give the team an easy way to start things early or skip a night for a long test. Next I'd right-size, since dev boxes are often copies of production size, and delete unattached EBS volumes and old snapshots. The key is agreeing the schedule with the developers first, so nobody loses a half-finished test run at seven in the evening.”

Red flag to avoid:

Shutting down environments without warning the team, or deleting untagged resources without finding their owner.

They may ask next:
  • How would you handle a team that needs an environment running overnight for a test?
  • Which dev resources keep costing money even when stopped?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

20. Production is down, and the quickest fix is a change in the AWS console. Your team's rule is that every change goes through infrastructure as code. You're on call. What do you do?

What the interviewer is really testing:
Whether you put users first in an emergency while keeping a clear record, so the code and the real setup don't drift apart.
Answer frame:

Users first: if the console change is the fastest safe fix, make it, with a second person watching if possible.

Leave a trail: say in the incident channel exactly what you changed and when.

Back into code: the next working day, the same change goes into the template, then check drift.

Fix the gap: ask why the pipeline couldn't ship the fix in time.

Sample spoken answer:

“I'd fix production. The infrastructure as code rule exists to keep things reliable, and it shouldn't keep users down longer. But I'd do it carefully: I post in the incident channel what I'm about to change, get a second pair of eyes if anyone's around, make the smallest change that works, and write down exactly what I did and when. The danger with console changes is that the code no longer matches reality, so the next deploy quietly undoes the fix or fails halfway. That makes the follow-up as important as the fix. The next working day I put the same change into the template, run drift detection to confirm they match, and note it in the incident review. And I'd ask why the pipeline couldn't get a fix out fast enough, because that's the real gap.”

Red flag to avoid:

Either leaving users down to follow the rule, or making the console change and never putting it back into code.

They may ask next:
  • What if the change you made in the console turns out to be the wrong fix?
  • How would you make an emergency change through the pipeline fast enough next time?
Say it in 60 seconds

Containers 2 questions

Medium Technical round Mid-level Practice question

21. An ECS service on Fargate won't start: every task stops with an error saying it couldn't pull the container image from ECR. The tasks run in private subnets. What do you check?

What the interviewer is really testing:
Whether you can read the stopped reason and know every network path and role a Fargate task needs to pull its image.
Answer frame:

Read the reason: a stopped task's reason says whether it's a network timeout, an access error or a missing image.

Network path: private subnets need a NAT gateway, or VPC endpoints for the ECR API, the ECR registry and S3, where the layers live.

Right role: the task execution role pulls the image and writes logs, not the task role.

The image: the tag exists, and it was built for the CPU architecture the task uses.

Sample spoken answer:

“First I'd open a stopped task and read its stopped reason, because it tells me whether this is a network problem or a permissions problem. If it's a timeout, the task can't reach ECR at all. In private subnets, Fargate needs either a route to a NAT gateway or VPC endpoints. With endpoints, people often add the ECR API one and forget the rest: you also need the ECR Docker registry endpoint and an S3 gateway endpoint, because the image layers are actually stored in S3. If logs go to CloudWatch, a logs endpoint too, and the endpoints' security group has to allow HTTPS from the tasks. If the reason says access denied, I'd check the task execution role, not the task role, since that's the one that pulls the image. And if it says there's no matching manifest, the image was built for a different CPU architecture.”

Red flag to avoid:

Giving the task role admin rights, when the pull is done by the execution role and the real problem is often the network path.

They may ask next:
  • The same task starts in a public subnet only when a public IP is assigned. Why?
  • How would you tell whether the task role or the execution role is missing a permission?
Say it in 60 seconds
Hard Technical round Senior Practice question

22. New pods on your EKS cluster won't start, with errors about assigning an IP address, even though the nodes have plenty of spare CPU and memory. What's going on?

What the interviewer is really testing:
Whether you know that the default VPC CNI gives every pod a real VPC address, and can tell a full subnet from a per-node limit.
Answer frame:

How pods get IPs: the VPC CNI gives each pod an address from the node's subnet.

Two limits: the subnet running out of free addresses, and each instance type's cap on network interfaces and addresses.

Checks: free addresses in the subnets, the node's max pods, the CNI logs.

Fixes: prefix delegation for the per-node cap; a secondary CIDR with custom networking for a full subnet.

Sample spoken answer:

“On EKS with the default VPC CNI, every pod gets a real IP address from the subnet its node sits in, so pods can run out of addresses long before the nodes run out of CPU. There are two separate limits. One is the subnet: a small subnet with lots of pods simply has no free addresses left, and new pods get stuck creating with an IP assignment error. The other is per node: each instance type can hold only so many network interfaces and addresses, which caps how many pods it runs. I'd check free addresses in the subnets and the node's max pods, then read the CNI logs. If it's the node cap, prefix delegation hands each interface small blocks of addresses so a node fits many more pods. If the subnet is full, I'd add a secondary CIDR range to the VPC and use custom networking so pods draw from bigger subnets.”

Red flag to avoid:

Adding more nodes, which only takes even more addresses from the same full subnet.

They may ask next:
  • What's the catch with prefix delegation in a subnet that's already fragmented?
  • How would you size the subnets for a new cluster?
Say it in 60 seconds

Databases 2 questions

Medium Technical round Mid-level, Senior Practice question

23. The RDS database CPU has sat at the top of the graph for an hour and every page is slow. Nothing was deployed today. What do you check?

What the interviewer is really testing:
Whether you find the query or load causing the pressure before reaching for a bigger instance.
Answer frame:

Which queries: Performance Insights, CloudWatch Database Insights or the slow query log shows the top SQL by load.

What changed: data growth, a new report, a batch job, a traffic pattern, stale statistics.

Instance type: a burstable instance may have run out of CPU credits.

Short term vs lasting: scale up or kill a runaway query now, then fix the index or query.

Sample spoken answer:

“No deploy doesn't mean nothing changed, so I'd open Performance Insights or Database Insights and look at the top SQL by load for that hour. Usually one or two statements stand out. If it's a query that got slow because a table grew past the point where a missing index hurts, I'd check its plan and add the index. If it's a scheduled report someone started running at peak, I'd move it to a read replica or off hours. I'd also check the instance class: if it's a burstable type, it may simply have spent its CPU credits and dropped to baseline. And I'd look at connection count in case the app is retrying and piling on. If users are suffering right now, killing a runaway query or scaling up is fine as a stopgap, but I'd still find the query.”

Red flag to avoid:

Jumping straight to a bigger instance without ever looking at which queries are using the CPU.

They may ask next:
  • How would you add an index to a large production table without locking writes?
  • When is a read replica the wrong fix for a slow database?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

24. During a sale, the ElastiCache Redis node was replaced, the cache came back empty, and the database fell over within a minute. What happened, and how do you stop it happening again?

What the interviewer is really testing:
Whether you recognise a cache stampede and know how to protect the database when the cache is cold or many keys expire together.
Answer frame:

What happened: every request missed at once and went to a database sized only for the misses.

Right now: shed or limit load, and warm the hottest keys before letting full traffic back.

Survive a node loss: a replica with automatic failover, so the cache doesn't restart cold.

Stop the herd: one request rebuilds a key while others wait or get the old value; add jitter to expiry times.

Sample spoken answer:

“That's a cache stampede. The database was quietly sized for the few requests that miss the cache, and when the node came back empty, every request missed at the same moment and hit the database together. First I'd protect the database: limit traffic or serve a lighter version of the busiest pages, and warm the hottest keys with a script before letting full traffic back. For the lasting fix, I'd run the cache with a replica and automatic failover, so losing one node promotes a replica that already holds the data instead of starting cold. In the code, I'd make sure only one request rebuilds a missing key, using a short lock, while the others wait briefly or get a slightly stale value. And I'd add random jitter to expiry times, because keys that all expire together cause a smaller version of the same stampede.”

Red flag to avoid:

Just making the cache node bigger, which does nothing for the moment it comes back empty.

They may ask next:
  • How would you test what the database can take when the cache is completely empty?
  • When is serving a slightly stale value not acceptable?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card