AWS interviews for 5 years of experience skip definitions like a VPC and ask why you chose one service over another, what that choice cost, and how you proved a fix worked. Expect stories about failovers, throttling, runaway queues, risky stack updates, migrations and pushing back on a plan that was too big. This page is written for cloud and backend engineers with roughly five to seven years on AWS. You own a service or a slice of the platform, pick the pieces it runs on, carry the pager when it breaks and review other people's infrastructure changes. Most sample answers are first-person stories. Swap in your own project before you say it out loud.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Situation: what failed over, and how long users saw errors compared with the failover itself.
Root cause: why the app held on to the old primary, such as cached DNS or stale pooled connections.
Fix and proof: the client-side changes, then a planned failover test to show the gap shrank.
“We ran Postgres on RDS with Multi-AZ, and during a maintenance event the failover itself finished in about a minute, but our API kept throwing errors for much longer. When I dug in, our Java services had cached the endpoint's DNS answer for far too long, and the connection pool kept handing out connections to the old primary, which just hung until they timed out. I lowered the JVM's DNS cache time, set the pool to validate connections and drop ones that failed, and added retries with backoff on the few write paths that could safely retry. Then I forced a reboot with failover in staging and in a quiet production window to measure it. The error window went from many minutes to under a minute. The trade-off was a little extra connection churn, which I was happy to pay.”
Saying Multi-AZ makes failover invisible to the app, or never having tested a failover on purpose.
Split the time: load balancer target response time versus the load balancer itself.
Narrow down: per-instance CPU, credit balance, disk and database metrics against the slow window.
Fix and cost: the change you made and what it did to the bill.
“The load balancer's target response time rose in the same window every day, so the delay was in our instances, not the load balancer. CPU usage looked modest, which threw me at first. Then I noticed we were on burstable instances in standard mode, and the CPU credit balance graph slid to zero by early afternoon, when the instances got throttled to their baseline. Morning batch jobs had grown over a few months and were spending the credits. Short term I moved the batch work to its own small worker pool so it stopped starving the API. For the API, I compared the cost of unlimited mode against a fixed-performance instance family for our steady load and chose the latter, because we were busy most of the day anyway. Afternoon latency went flat, and I added an alarm on credit balance for any burstable instance we kept.”
Scaling out blindly or restarting instances without looking at per-instance metrics.
Cause: request rates are scaled per prefix, and every write was landing on one hot prefix.
Spread the keys: more prefixes so the load can be split across them.
Behave well: retries with exponential backoff and jitter, and fewer, larger objects where possible.
“The job wrote every output file under a single date prefix, and at peak it was sending tens of thousands of PUTs a second into that one prefix. S3 scales the request rate per prefix, and it adapts over time, but a sudden burst into one prefix hits the limit and you get Slow Down errors. Our SDK retries were set low and without jitter, so all the workers retried together and made it worse. I changed the key layout to add a short hash-based shard after the date, which spread writes across many prefixes, and turned on the SDK's adaptive retry mode. I also batched tiny records into larger files, which cut the request count and made the downstream readers faster too. The errors stopped, and the job finished sooner. The trade-off was that listing one day's data now meant reading several prefixes, which I wrapped in a small helper.”
Assuming S3 has unlimited throughput on any key layout, or retrying in a tight loop without backoff.
Cause: replicas apply changes asynchronously, so a read right after a write can land on a replica that hasn't caught up.
Evidence: replica lag metrics lining up with the complaints.
Fix: send a user's reads to the primary for a short window after they write, and alarm on lag.
“We'd pointed all reads at two read replicas to take load off the primary, and the complaints started the same week. The pattern was always the same: save a profile, the page reloads, and the old values show up. The replica lag metric told the story. Most of the time it was under a second, but during busy periods it jumped to several seconds, and the reload after a save hit a replica that hadn't applied the write yet. I didn't want to give up the replicas, so I added read-your-own-writes routing: after a user writes, we set a short-lived flag in their session, and for the next few seconds their reads go to the primary. Everyone else keeps reading from replicas. I also added an alarm on replica lag and a rule to pull a replica out of rotation if it falls far behind. The complaints stopped, and the primary kept most of the relief.”
Assuming replicas are always in sync, or fixing it by sending every read back to the primary.
Look wider: check throttle metrics on every index, not only the table.
Mechanism: each write that touches an indexed attribute also has to be written to the index, and a starved index throttles the base table.
Fix: give the index enough capacity or change modes, and trim what the index projects.
“The table's consumed write capacity was well under what we'd provisioned, so at first it looked like a bug in our client. When I opened the metrics per index, one global secondary index was throttling constantly. We'd added it a few weeks earlier for a new report and given it a small write capacity, assuming it was read-mostly. But every write to the table that touched the indexed attribute also had to be written into that index, and when the index couldn't keep up, DynamoDB throttled writes on the base table too. I raised the index's write capacity to match the table's write rate, cut the projection down to the few attributes the report needed so each index write was smaller, and added an alarm on index throttling specifically. The lesson I share now is that an index is part of the write path, not a free side table.”
Blaming DynamoDB or the SDK without checking index metrics, or treating an index as something that only affects reads.
Pain: what went wrong with functions calling functions, such as half-finished orders and no clear view of where one got stuck.
Choice: orchestration with retries, timeouts and compensating steps in one place.
Trade-off: cost per state change, a new tool to learn, and tighter coupling to one platform.
“Our checkout was four Lambdas, each invoking the next: reserve stock, take payment, create the shipment, send the email. When the shipping call failed, we had orders where payment was taken but no shipment existed, and finding where one got stuck meant searching four log groups. I moved it to a Step Functions state machine. Each step got its own retry policy and timeout, and if payment succeeded but shipping kept failing, a catch step released the stock and refunded the payment, so we never left an order half done. Support could open the execution and see exactly which step failed. The trade-offs were real: you pay per state transition, which mattered on our busiest days, the workflow definition is tied to one cloud, and the team had to learn a new way to test. For short, high-volume flows I'd still use a queue between two functions instead.”
Claiming chained functions are fine because each one retries, with no thought for partial failure or visibility.
Cause: all functions in an account and region share one concurrency pool.
Protect: reserve concurrency for critical functions and cap the noisy one.
Longer term: alarms on throttles and concurrency, and separate accounts for unrelated workloads.
“A marketing import dropped tens of thousands of files into S3 at once, and the function processing them scaled out fast enough to use up almost all of the account's concurrency in that region. Our checkout API ran on Lambda in the same account, and it started getting throttled, which customers noticed. Concurrency is one shared pool per account and region, so a single greedy function can starve everything else. I set reserved concurrency on the import function so it could never take more than a slice of the pool, and reserved a guaranteed amount for the checkout functions. I also put a queue in front of the import so bursts are smoothed out instead of hitting Lambda all at once. Later we moved batch workloads to their own account so they had their own pool. The import now runs a bit slower, which nobody minds.”
Not knowing concurrency is shared across functions, or fixing it only by asking for a higher limit.
Workload shape: traffic pattern, run time, latency needs and state.
Team fit: what the team could operate well at two in the morning.
Cost you accepted: the limit or bill you knowingly took on, and when you'd revisit it.
“Our order service was a long-running HTTP API with steady traffic during the day and a few background jobs that ran for up to half an hour. Lambda's time limit ruled it out for the jobs, and nobody on the team had run Kubernetes in production, so EKS would have meant learning a whole control plane just to host four containers. I picked ECS on Fargate: no servers to patch, a normal container image, and scaling on CPU and request count behind the load balancer. What it cost us was a higher price per unit of compute than well-packed EC2, less control over the host for debugging, and slower cold scale-out than Lambda during sudden bursts. I wrote those down in the design doc with a note that if our steady load grew a lot, we'd look at EC2 capacity for the cluster to cut the bill.”
Picking a platform because it is popular or on a CV, with no mention of the workload or what the choice gave up.
Requirements: one consumer or many, ordering, replay, throughput and message size.
Choice: why the picked service fit, and the next-best option you rejected.
Switch point: the change in requirements that would make you move.
“When an order is placed, three things need to happen: billing, email and analytics. Each is owned by a different team and can fail separately. I used SNS for fan-out with an SQS queue per consumer, so each team gets its own copy, its own retry and dead-letter queue, and a slow email service can't hold up billing. I looked at EventBridge, which would have given us content-based routing and a schema registry, but at the time we had only one event type and wanted the simplest thing. I ruled out Kinesis because we didn't need ordered replay of the whole stream and didn't want to manage shards. What would make me switch: if analytics needed to replay days of events in order, I'd add a stream for that consumer; if we grew to dozens of event types with rules, I'd move routing to EventBridge.”
Treating all four services as interchangeable, or picking one without mentioning retries, ordering or who consumes it.
The decision: what was chosen, and why it seemed fine at the time.
The bite: the concrete problem later, such as peering refused because ranges overlap.
The habit: the address plan or review step you use now.
“Early on, every team created VPCs with the default-style range from a tutorial, and I did the same for my service. Two years later we needed to connect my VPC to a data team's VPC and to the office network, and peering refused because the address ranges overlapped. We had two bad options: put a NAT layer in between, which made debugging painful, or rebuild one VPC with a new range. We rebuilt mine, which meant moving databases and load balancers over a few weekends. Since then I keep a simple shared address plan: every VPC in every account gets a non-overlapping block from it, sized for growth, and new VPCs are created from a module that takes its range from that plan. For bigger setups I'd connect them through a transit gateway, but none of that works if the ranges clash.”
Claiming no design decision ever went wrong, or blaming someone else without saying what you learned.
Read it literally: star means anyone, so everything rests on the condition.
Check the condition: is it a strong key, like the organization ID or a specific VPC endpoint, and is it on every statement?
Prove it: run an access analyzer check and a test from outside, and keep block public access on.
“I'd read it as if the condition wasn't there, because with Principal star the condition is the only thing standing between the bucket and the whole internet. Then I'd check what the condition actually uses. If it's something like the organization ID key, or a specific VPC endpoint ID, that's a real boundary, and it's a common pattern for sharing a bucket across many accounts. If it's a referer header or a source IP range that might change, I'd push back, because those are weak or easy to get wrong. I'd also check the condition sits on every statement that allows access, not just one. Finally I'd ask for proof: IAM Access Analyzer, with the organization as its zone of trust, should show no findings for the bucket, and a test from an account outside the organization should get AccessDenied. On my last team we made that analyzer check a required pipeline step for every bucket policy change.”
Approving because the description says it's private, or rejecting every Principal star policy without reading the condition.
Cause: snapshots encrypted with the AWS managed key cannot be shared with another account.
Fix: copy the snapshot using a customer managed key, then share both the snapshot and access to that key.
Target side: the other account copies it under its own key before restoring.
“The snapshot had been encrypted with the default AWS managed key for RDS, and snapshots encrypted with that key can't be shared with another account at all, because you can't change that key's policy. To fix it, I created a customer managed KMS key, copied the snapshot and re-encrypted it with that key, then shared the new snapshot with the target account. Sharing the snapshot isn't enough on its own, so I also added the target account to the key policy with permission to use the key. On their side, they copied the snapshot and re-encrypted it with their own key, so they didn't depend on ours later. After that, I changed our templates so new databases use a customer managed key from the start, which saved us the copy step during the next audit request.”
Suggesting you turn off encryption or make the snapshot public to get around it.
Cause: the app read the secret once at startup and kept using the old password after rotation.
Fix in the app: refetch the secret and retry when authentication fails, with sensible caching.
Fix in rotation: a strategy that keeps a valid credential during the switch, such as alternating users.
“We moved database credentials out of environment variables into Secrets Manager with rotation turned on. The first time it rotated, the password changed in the database, but our services had read the secret once at startup and kept trying the old one whenever they opened a new connection. Errors grew as old connections were recycled. I changed the services to use the caching client with a short refresh time and, on an authentication error, to refetch the secret and retry the connection once. I also switched rotation to the alternating-users strategy, where there are two database users and rotation updates the one not currently in use, so the old credential stays valid for a while after the switch. Then I triggered a rotation in staging during a load test to prove it. Now rotation is a non-event, and we rotate far more often than we could by hand.”
Turning rotation off to make the errors stop, or not knowing the app has to reload the secret.
Walk everything: paginate over users, then list each user's keys.
Useful output: key age, status and last-used date, so unused keys stand out.
Safe by default: report only; deactivate before deleting, after the owner confirms.
“I'd use boto3's paginator on list_users, and for each user call list_access_keys, work out the key's age from its create date, and skip anything younger than 90 days. For old keys I call get_access_key_last_used, because a key that hasn't been used in months is a much easier conversation than one used an hour ago. The script only reports. When I ran this at my last company, the plan was: share the list, deactivate unused keys first rather than deleting them, wait a couple of weeks for anyone to shout, then delete. For keys still in use, we moved those workloads to roles so they didn't need long-lived keys at all. The script needs only read permissions on IAM, and I ran it from a role, not from a user with a key of its own.”
import boto3
from datetime import datetime, timezone
iam = boto3.client("iam")
now = datetime.now(timezone.utc)
MAX_AGE_DAYS = 90
for page in iam.get_paginator("list_users").paginate():
for user in page["Users"]:
name = user["UserName"]
keys = iam.list_access_keys(UserName=name)["AccessKeyMetadata"]
for key in keys:
age = (now - key["CreateDate"]).days
if age < MAX_AGE_DAYS:
continue
info = iam.get_access_key_last_used(AccessKeyId=key["AccessKeyId"])
last = info["AccessKeyLastUsed"].get("LastUsedDate")
last_txt = last.date().isoformat() if last else "never"
print(f"{name} {key['AccessKeyId']} {key['Status']} {age} days last used {last_txt}")
Writing a script that deletes keys straight away, or skipping pagination so users are silently missed.
The change: what looked harmless, such as a rename or a property that forces replacement.
How it was caught: reading the change set or plan and spotting the removal or replacement.
Guardrails: retain policies on stateful resources, stack protection, and a pipeline check that blocks replacements.
“A teammate refactored our CloudFormation template and renamed the logical ID of the RDS instance to match a new naming scheme. To CloudFormation that looks like one resource removed and a new one added, so the change set showed the old database being removed and a new one added. I caught it because our pipeline posts the change set summary on the pull request and I read every line that removes or replaces a resource. We reverted the rename. Then I added guardrails: DeletionPolicy and UpdateReplacePolicy set to retain or snapshot on every stateful resource, termination protection on the stack, a stack policy that denies replacing or deleting the database, and a pipeline step that fails if a change set replaces or removes anything tagged as stateful unless someone approves it by hand. It made some legitimate changes slower, but nobody has come close to losing data since.”
Trusting that a refactor is safe because the template still validates, or relying on backups as the only guardrail.
Replicate: an initial load plus ongoing change capture until the target is caught up.
Verify: row counts, checksums on key tables, and the app tested against the target.
Cutover and rollback: a rehearsed, short write freeze, and a plan to go back if the first hour goes wrong.
“We had MySQL on two EC2 instances that we patched by hand, and I led moving it to Aurora MySQL. I used Database Migration Service for a full load followed by ongoing change capture from the binary log, so the target stayed a few seconds behind the source for weeks while we tested. I wrote checks comparing row counts and checksums on the biggest tables, and we ran the staging app against Aurora for a full release cycle. For the cutover we rehearsed twice: stop writes, wait for replication lag to hit zero, run the checks, point the app's config at the new endpoint and turn writes back on. Downtime was about ten minutes. For rollback, I kept the old primary and set up reverse replication so we could go back without losing new writes. We never needed it, but having it made the go decision easy.”
Describing a dump-and-restore over a weekend with no data checks or no way back.
Why: the price-performance case and which service went first.
Work: multi-architecture builds, native dependencies, and agents or tools that lacked ARM support.
Proof: a side-by-side test on real traffic before the full switch, and how you rolled it out.
“Our biggest spend was a fleet of containers running a Node API, and Graviton promised better price-performance, so I proposed trying it on that one service first. The code itself didn't change, but the build did: I switched our image builds to produce multi-architecture images, and two native modules plus an old monitoring agent had no ARM builds, so I upgraded them. Then we ran a small share of production traffic on ARM tasks next to the x86 ones for a week and compared latency, errors and cost per request. ARM was a bit faster and cheaper per request, so we moved the rest over in steps. The trade-off is that every new dependency now has to support both architectures, so I added an ARM build to CI to catch that early instead of at deploy time.”
Switching the whole fleet at once on the strength of a vendor claim, with no side-by-side measurement.
Fit: stateless or checkpointed work that can be retried, not the only copy of anything.
Handling: react to the two-minute interruption notice, drain work and checkpoint.
Availability: spread across many instance types and zones, with on-demand as a fallback.
“I used Spot for our nightly video transcoding workers, which pulled jobs from a queue. Each job could be retried, so losing a worker only cost time. The worker listened for the interruption notice, which gives about two minutes, stopped taking new jobs, and put its current job back on the queue if it couldn't finish. The bigger lesson was capacity: when we asked for one instance type, we sometimes couldn't get any Spot capacity at all during busy periods. I changed the Auto Scaling group to a mixed instances policy with several similar instance types across three zones, a price-capacity-optimized allocation, and a small on-demand base so the pipeline always made progress. Jobs occasionally took longer, but the batch still finished before morning at a fraction of the old compute cost, and the API servers stayed on on-demand.”
Putting critical stateful services on Spot, or not knowing there is an interruption notice to handle.
Audit: list every alarm that paged in the last month and whether anyone had to act.
Page on symptoms: errors and latency users feel, with sensible datapoints-to-alarm settings.
Demote the rest: causes like CPU go to tickets or dashboards, composite alarms for combined signals.
“I pulled a month of pages and marked each one: did someone have to do anything? Most were CPU spikes on single instances during deploys and one-minute error blips that recovered on their own. So I changed what pages us. Pages now fire on symptoms users feel, like the load balancer's error rate and p99 latency staying above a threshold for several datapoints in a row, not one bad minute. CPU and disk alarms became tickets or dashboard widgets. Where one alarm alone was noisy, I used a composite alarm so we only page when errors and latency are both bad. Every paging alarm got a short runbook link in its description. Night pages dropped to a handful a month, and when we did get paged, people trusted it and moved fast. The trade-off is we see slow creeping problems a bit later, so we review the dashboards weekly.”
Silencing alarms or raising thresholds until they stop firing, with no link to what users experience.
The drill: what you restored, where, and against what recovery target.
Surprises: the gaps found, such as missing keys, slow restores or forgotten config.
Follow-through: what you fixed, and how often the drill now runs.
“We had nightly snapshots and cross-region copies, and everyone assumed we were covered. I ran a drill: restore the production database and the file bucket into a clean account in the second region, bring the app up and check it works. Three things surprised us. The snapshot copies in the other region were encrypted with a key that the restore account couldn't use, so the first attempt failed outright. The restore of our largest database took far longer than the recovery time we'd promised. And the app needed secrets and parameters we'd never backed up at all. I fixed the key policies, added a cross-region read replica for the big database so recovery became a promotion instead of a long restore, and put secrets and config into code so they could be recreated. We now run the drill every quarter, and it takes an afternoon instead of a week.”
Saying backups are fine because the backup job reports success, without ever restoring one.
Shift gradually: canary or linear traffic shifting instead of all at once.
Watch real signals: alarms on errors and latency that gate each step.
Roll back on its own: the deploy tool reverts on an alarm, and changes stay backward compatible.
“For our ECS services I set up blue/green deployments with CodeDeploy behind the load balancer. A new version starts next to the old one, gets a small share of traffic for ten minutes, then the rest. The deployment is tied to CloudWatch alarms on the new version's error rate and latency, so if either alarm fires during the canary window, CodeDeploy shifts traffic back to the old tasks automatically. The part people miss is database changes: a rollback only works if the old code still runs against the new schema, so I set a rule that migrations are additive first and cleanup comes a release later. The costs are running two sets of tasks during each deploy and slower releases. The first month it caught two bad releases that would have been full outages, which ended the debate about speed.”
Relying on someone watching dashboards and redeploying the old version by hand, or ignoring schema changes.
Contain: close the rule, check for logins while it was open.
Coach: a private, blame-free talk about why it's risky and what to use instead.
Fix the system: a safer access path and a guardrail that catches the rule next time.
“Our config rules flagged a security group with port 22 open to everyone. I removed the rule straight away, then checked the instance's auth logs and our flow logs for connections from unknown addresses while it was open. Nothing had got in. Then I talked to the junior privately. He'd been stuck on an outage and couldn't reach the box any other way, which told me the real problem was ours: we had no good access path. So I set up Session Manager on our instances, which gives a shell through IAM with no open inbound port, set it to log every session, and walked the team through using it. I also added an automatic remediation that removes world-open SSH rules and posts in our channel. He later wrote the how-to page for the team himself, which I think did more than any lecture would have.”
Blaming the person publicly, or closing the rule and moving on without fixing why it happened.
Get the target: agree how long the service may be down and how much data it may lose.
Show the cost: data replication, conflicts, doubled infrastructure, testing and on-call load.
Offer a ladder: solid multi-AZ now, backups copied to a second region, a warm standby only if the targets demand it.
“I'd start by asking what problem we're solving. When I did this last year, the answer was a customer asking about disaster recovery, not a real need for zero downtime. So I got the manager and the product owner to agree on numbers: a few hours of downtime in a full regional outage was acceptable, and losing a few minutes of data was too. Then I laid out what active-active really means for our relational database: writes in two places, conflict handling, twice the infrastructure and a lot more testing. Instead I proposed hardening what we had across availability zones, copying backups and images to a second region, and keeping infrastructure as code that could stand up a copy there. We tested the restore and hit the target. It was a fraction of the effort, and the customer was satisfied with the documented plan.”
Saying yes to avoid conflict, or saying no flatly without offering numbers or an alternative.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.