Production Incidents • Design Trade-offs • Migrations • Reliability • Code Review • 2026

AWS Interview Questions for Experienced Candidates (5 Years)

AWS interviews for 5 years of experience skip definitions like a VPC and ask why you chose one service over another, what that choice cost, and how you proved a fix worked. Expect stories about failovers, throttling, runaway queues, risky stack updates, migrations and pushing back on a plan that was too big. This page is written for cloud and backend engineers with roughly five to seven years on AWS. You own a service or a slice of the platform, pick the pieces it runs on, carry the pager when it breaks and review other people's infrastructure changes. Most sample answers are first-person stories. Swap in your own project before you say it out loud.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Incidents 3 questions

Hard Behavioral round Mid-level, Senior Practice question

1. Tell me about a managed database failover that didn't go as smoothly as the docs promised. What broke in the application, and what did you change?

What the interviewer is really testing:
Whether you know that the database failing over is only half the story, and that the client side (DNS caching, connection pools, retries) decides how long users actually see errors.
Answer frame:

Situation: what failed over, and how long users saw errors compared with the failover itself.

Root cause: why the app held on to the old primary, such as cached DNS or stale pooled connections.

Fix and proof: the client-side changes, then a planned failover test to show the gap shrank.

Sample spoken answer:

“We ran Postgres on RDS with Multi-AZ, and during a maintenance event the failover itself finished in about a minute, but our API kept throwing errors for much longer. When I dug in, our Java services had cached the endpoint's DNS answer for far too long, and the connection pool kept handing out connections to the old primary, which just hung until they timed out. I lowered the JVM's DNS cache time, set the pool to validate connections and drop ones that failed, and added retries with backoff on the few write paths that could safely retry. Then I forced a reboot with failover in staging and in a quiet production window to measure it. The error window went from many minutes to under a minute. The trade-off was a little extra connection churn, which I was happy to pay.”

Red flag to avoid:

Saying Multi-AZ makes failover invisible to the app, or never having tested a failover on purpose.

They may ask next:
  • Which writes did you decide were not safe to retry, and how did you handle them?
  • Would a database proxy have changed this story, and what would it cost you?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

2. Response times behind your load balancer crept up every afternoon and nothing in the code had changed. How did you track it down?

What the interviewer is really testing:
Whether you debug from the load balancer inward with real metrics, and know the instance-level causes of slowdowns such as burstable CPU credits running out.
Answer frame:

Split the time: load balancer target response time versus the load balancer itself.

Narrow down: per-instance CPU, credit balance, disk and database metrics against the slow window.

Fix and cost: the change you made and what it did to the bill.

Sample spoken answer:

“The load balancer's target response time rose in the same window every day, so the delay was in our instances, not the load balancer. CPU usage looked modest, which threw me at first. Then I noticed we were on burstable instances in standard mode, and the CPU credit balance graph slid to zero by early afternoon, when the instances got throttled to their baseline. Morning batch jobs had grown over a few months and were spending the credits. Short term I moved the batch work to its own small worker pool so it stopped starving the API. For the API, I compared the cost of unlimited mode against a fixed-performance instance family for our steady load and chose the latter, because we were busy most of the day anyway. Afternoon latency went flat, and I added an alarm on credit balance for any burstable instance we kept.”

Red flag to avoid:

Scaling out blindly or restarting instances without looking at per-instance metrics.

They may ask next:
  • Why was CPU usage misleading here?
  • When is a burstable instance still the right choice?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

3. A batch job writing millions of small files to S3 started failing with 503 Slow Down errors. What was happening, and how did you fix it?

What the interviewer is really testing:
Whether you know S3 scales request rates per key prefix, how to spread load across prefixes, and that retries with backoff are expected rather than a hack.
Answer frame:

Cause: request rates are scaled per prefix, and every write was landing on one hot prefix.

Spread the keys: more prefixes so the load can be split across them.

Behave well: retries with exponential backoff and jitter, and fewer, larger objects where possible.

Sample spoken answer:

“The job wrote every output file under a single date prefix, and at peak it was sending tens of thousands of PUTs a second into that one prefix. S3 scales the request rate per prefix, and it adapts over time, but a sudden burst into one prefix hits the limit and you get Slow Down errors. Our SDK retries were set low and without jitter, so all the workers retried together and made it worse. I changed the key layout to add a short hash-based shard after the date, which spread writes across many prefixes, and turned on the SDK's adaptive retry mode. I also batched tiny records into larger files, which cut the request count and made the downstream readers faster too. The errors stopped, and the job finished sooner. The trade-off was that listing one day's data now meant reading several prefixes, which I wrapped in a small helper.”

Red flag to avoid:

Assuming S3 has unlimited throughput on any key layout, or retrying in a tight loop without backoff.

They may ask next:
  • Why does batching small files help both the writer and the readers?
  • How would you tell S3 throttling apart from a network problem on your side?
Say it in 60 seconds

Databases 2 questions

Hard Technical round Mid-level, Senior Practice question

4. After you moved read traffic to read replicas, users started saying their changes 'disappeared' right after saving. What was going on, and how did you fix it without giving up the replicas?

What the interviewer is really testing:
Whether you know read replicas are updated asynchronously, can spot replica lag as the cause, and can design read-your-own-writes routing instead of abandoning the scaling gain.
Answer frame:

Cause: replicas apply changes asynchronously, so a read right after a write can land on a replica that hasn't caught up.

Evidence: replica lag metrics lining up with the complaints.

Fix: send a user's reads to the primary for a short window after they write, and alarm on lag.

Sample spoken answer:

“We'd pointed all reads at two read replicas to take load off the primary, and the complaints started the same week. The pattern was always the same: save a profile, the page reloads, and the old values show up. The replica lag metric told the story. Most of the time it was under a second, but during busy periods it jumped to several seconds, and the reload after a save hit a replica that hadn't applied the write yet. I didn't want to give up the replicas, so I added read-your-own-writes routing: after a user writes, we set a short-lived flag in their session, and for the next few seconds their reads go to the primary. Everyone else keeps reading from replicas. I also added an alarm on replica lag and a rule to pull a replica out of rotation if it falls far behind. The complaints stopped, and the primary kept most of the relief.”

Red flag to avoid:

Assuming replicas are always in sync, or fixing it by sending every read back to the primary.

They may ask next:
  • Which reads in your app must always go to the primary, whatever the lag?
  • What would you do if replica lag kept growing and never recovered?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

5. Writes to a DynamoDB table were being throttled, yet the table's own write capacity looked mostly unused. What turned out to be the cause, and what did you do?

What the interviewer is really testing:
Whether you know that a global secondary index without enough write capacity pushes back on the base table, and can read the right metrics to prove it.
Answer frame:

Look wider: check throttle metrics on every index, not only the table.

Mechanism: each write that touches an indexed attribute also has to be written to the index, and a starved index throttles the base table.

Fix: give the index enough capacity or change modes, and trim what the index projects.

Sample spoken answer:

“The table's consumed write capacity was well under what we'd provisioned, so at first it looked like a bug in our client. When I opened the metrics per index, one global secondary index was throttling constantly. We'd added it a few weeks earlier for a new report and given it a small write capacity, assuming it was read-mostly. But every write to the table that touched the indexed attribute also had to be written into that index, and when the index couldn't keep up, DynamoDB throttled writes on the base table too. I raised the index's write capacity to match the table's write rate, cut the projection down to the few attributes the report needed so each index write was smaller, and added an alarm on index throttling specifically. The lesson I share now is that an index is part of the write path, not a free side table.”

Red flag to avoid:

Blaming DynamoDB or the SDK without checking index metrics, or treating an index as something that only affects reads.

They may ask next:
  • Would on-demand mode have avoided this, and what would you trade for it?
  • How would you backfill a new index on a very busy table without hurting live traffic?
Say it in 60 seconds

Serverless and Messaging 2 questions

Hard System design round Mid-level, Senior Practice question

6. You had a multi-step order workflow built as Lambdas calling each other. Why did you move it, or decide not to move it, to Step Functions, and what did that trade off?

What the interviewer is really testing:
Whether you can see the hidden cost of chained functions (lost state, retries, timeouts, no visibility) and weigh an orchestrator against its price and lock-in.
Answer frame:

Pain: what went wrong with functions calling functions, such as half-finished orders and no clear view of where one got stuck.

Choice: orchestration with retries, timeouts and compensating steps in one place.

Trade-off: cost per state change, a new tool to learn, and tighter coupling to one platform.

Sample spoken answer:

“Our checkout was four Lambdas, each invoking the next: reserve stock, take payment, create the shipment, send the email. When the shipping call failed, we had orders where payment was taken but no shipment existed, and finding where one got stuck meant searching four log groups. I moved it to a Step Functions state machine. Each step got its own retry policy and timeout, and if payment succeeded but shipping kept failing, a catch step released the stock and refunded the payment, so we never left an order half done. Support could open the execution and see exactly which step failed. The trade-offs were real: you pay per state transition, which mattered on our busiest days, the workflow definition is tied to one cloud, and the team had to learn a new way to test. For short, high-volume flows I'd still use a queue between two functions instead.”

Red flag to avoid:

Claiming chained functions are fine because each one retries, with no thought for partial failure or visibility.

They may ask next:
  • When would you pick the express type of workflow over the standard one?
  • How did you test the failure and refund path before production?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

7. A burst on one Lambda function caused throttling on other, more important functions in the same account. What happened, and how did you prevent it?

What the interviewer is really testing:
Whether you know that concurrency is shared per account and region, and how reserved concurrency and account separation protect critical functions.
Answer frame:

Cause: all functions in an account and region share one concurrency pool.

Protect: reserve concurrency for critical functions and cap the noisy one.

Longer term: alarms on throttles and concurrency, and separate accounts for unrelated workloads.

Sample spoken answer:

“A marketing import dropped tens of thousands of files into S3 at once, and the function processing them scaled out fast enough to use up almost all of the account's concurrency in that region. Our checkout API ran on Lambda in the same account, and it started getting throttled, which customers noticed. Concurrency is one shared pool per account and region, so a single greedy function can starve everything else. I set reserved concurrency on the import function so it could never take more than a slice of the pool, and reserved a guaranteed amount for the checkout functions. I also put a queue in front of the import so bursts are smoothed out instead of hitting Lambda all at once. Later we moved batch workloads to their own account so they had their own pool. The import now runs a bit slower, which nobody minds.”

Red flag to avoid:

Not knowing concurrency is shared across functions, or fixing it only by asking for a higher limit.

They may ask next:
  • What is the downside of reserving concurrency for a function?
  • How is reserved concurrency different from provisioned concurrency?
Say it in 60 seconds

Design Trade-offs 3 questions

Hard System design round Mid-level, Senior Practice question

8. For the last service you owned, why did you run it on the compute option you chose, such as Lambda, ECS on Fargate, EKS or plain EC2, and what did that choice cost you?

What the interviewer is really testing:
Whether you choose compute from the workload's shape and the team's skills, and can name the downside you accepted instead of claiming the choice was free.
Answer frame:

Workload shape: traffic pattern, run time, latency needs and state.

Team fit: what the team could operate well at two in the morning.

Cost you accepted: the limit or bill you knowingly took on, and when you'd revisit it.

Sample spoken answer:

“Our order service was a long-running HTTP API with steady traffic during the day and a few background jobs that ran for up to half an hour. Lambda's time limit ruled it out for the jobs, and nobody on the team had run Kubernetes in production, so EKS would have meant learning a whole control plane just to host four containers. I picked ECS on Fargate: no servers to patch, a normal container image, and scaling on CPU and request count behind the load balancer. What it cost us was a higher price per unit of compute than well-packed EC2, less control over the host for debugging, and slower cold scale-out than Lambda during sudden bursts. I wrote those down in the design doc with a note that if our steady load grew a lot, we'd look at EC2 capacity for the cluster to cut the bill.”

Red flag to avoid:

Picking a platform because it is popular or on a CV, with no mention of the workload or what the choice gave up.

They may ask next:
  • What would have pushed you toward EKS instead?
  • How did you size the tasks, and how did you know the sizing was right?
Say it in 60 seconds
Hard System design round Mid-level, Senior Practice question

9. Between two of your services, why did you pick the messaging option you used, such as SQS, SNS, EventBridge or Kinesis, and what would make you switch?

What the interviewer is really testing:
Whether you know what each messaging service is built for (work queues, fan-out, event routing, ordered streams) and chose from the requirements rather than habit.
Answer frame:

Requirements: one consumer or many, ordering, replay, throughput and message size.

Choice: why the picked service fit, and the next-best option you rejected.

Switch point: the change in requirements that would make you move.

Sample spoken answer:

“When an order is placed, three things need to happen: billing, email and analytics. Each is owned by a different team and can fail separately. I used SNS for fan-out with an SQS queue per consumer, so each team gets its own copy, its own retry and dead-letter queue, and a slow email service can't hold up billing. I looked at EventBridge, which would have given us content-based routing and a schema registry, but at the time we had only one event type and wanted the simplest thing. I ruled out Kinesis because we didn't need ordered replay of the whole stream and didn't want to manage shards. What would make me switch: if analytics needed to replay days of events in order, I'd add a stream for that consumer; if we grew to dozens of event types with rules, I'd move routing to EventBridge.”

Red flag to avoid:

Treating all four services as interchangeable, or picking one without mentioning retries, ordering or who consumes it.

They may ask next:
  • How do you stop a message schema change from breaking a consumer you don't own?
  • Where does ordering matter in this flow, if anywhere?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

10. Tell me about a network design decision on AWS that came back to bite you later, such as overlapping IP ranges. What did it cost, and what do you do differently now?

What the interviewer is really testing:
Whether you can own a past mistake honestly, explain its technical cost, and show the planning habit it taught you.
Answer frame:

The decision: what was chosen, and why it seemed fine at the time.

The bite: the concrete problem later, such as peering refused because ranges overlap.

The habit: the address plan or review step you use now.

Sample spoken answer:

“Early on, every team created VPCs with the default-style range from a tutorial, and I did the same for my service. Two years later we needed to connect my VPC to a data team's VPC and to the office network, and peering refused because the address ranges overlapped. We had two bad options: put a NAT layer in between, which made debugging painful, or rebuild one VPC with a new range. We rebuilt mine, which meant moving databases and load balancers over a few weekends. Since then I keep a simple shared address plan: every VPC in every account gets a non-overlapping block from it, sized for growth, and new VPCs are created from a module that takes its range from that plan. For bigger setups I'd connect them through a transit gateway, but none of that works if the ranges clash.”

Red flag to avoid:

Claiming no design decision ever went wrong, or blaming someone else without saying what you learned.

They may ask next:
  • How did you decide how big each VPC's range should be?
  • What would you do if you couldn't rebuild either VPC?
Say it in 60 seconds

Security and IAM 4 questions

Medium Situational round Mid-level, Senior Practice question

11. A teammate's pull request adds an S3 bucket policy with Principal set to star, plus a condition they say keeps it private. How do you review that?

What the interviewer is really testing:
Whether you can read a resource policy for what it really allows, check that the condition truly narrows access, and prove it with a tool instead of trusting the description.
Answer frame:

Read it literally: star means anyone, so everything rests on the condition.

Check the condition: is it a strong key, like the organization ID or a specific VPC endpoint, and is it on every statement?

Prove it: run an access analyzer check and a test from outside, and keep block public access on.

Sample spoken answer:

“I'd read it as if the condition wasn't there, because with Principal star the condition is the only thing standing between the bucket and the whole internet. Then I'd check what the condition actually uses. If it's something like the organization ID key, or a specific VPC endpoint ID, that's a real boundary, and it's a common pattern for sharing a bucket across many accounts. If it's a referer header or a source IP range that might change, I'd push back, because those are weak or easy to get wrong. I'd also check the condition sits on every statement that allows access, not just one. Finally I'd ask for proof: IAM Access Analyzer, with the organization as its zone of trust, should show no findings for the bucket, and a test from an account outside the organization should get AccessDenied. On my last team we made that analyzer check a required pipeline step for every bucket policy change.”

Red flag to avoid:

Approving because the description says it's private, or rejecting every Principal star policy without reading the condition.

They may ask next:
  • Why is a referer header condition a weak protection?
  • How does account-level block public access interact with a policy like this?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

12. You tried to share an encrypted database snapshot with another AWS account and it didn't work. Why not, and how did you fix it?

What the interviewer is really testing:
Whether you know how KMS key types limit cross-account sharing, and how key policies and grants fit together with the snapshot share.
Answer frame:

Cause: snapshots encrypted with the AWS managed key cannot be shared with another account.

Fix: copy the snapshot using a customer managed key, then share both the snapshot and access to that key.

Target side: the other account copies it under its own key before restoring.

Sample spoken answer:

“The snapshot had been encrypted with the default AWS managed key for RDS, and snapshots encrypted with that key can't be shared with another account at all, because you can't change that key's policy. To fix it, I created a customer managed KMS key, copied the snapshot and re-encrypted it with that key, then shared the new snapshot with the target account. Sharing the snapshot isn't enough on its own, so I also added the target account to the key policy with permission to use the key. On their side, they copied the snapshot and re-encrypted it with their own key, so they didn't depend on ours later. After that, I changed our templates so new databases use a customer managed key from the start, which saved us the copy step during the next audit request.”

Red flag to avoid:

Suggesting you turn off encryption or make the snapshot public to get around it.

They may ask next:
  • What permissions does the other account need on your key?
  • Why re-encrypt the copy with their own key instead of restoring directly?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

13. You moved database passwords into a secrets service with automatic rotation, and the first rotation broke the app. What went wrong, and how did you fix it?

What the interviewer is really testing:
Whether you understand how rotation interacts with cached credentials and long-lived connections, and know the patterns that make rotation safe.
Answer frame:

Cause: the app read the secret once at startup and kept using the old password after rotation.

Fix in the app: refetch the secret and retry when authentication fails, with sensible caching.

Fix in rotation: a strategy that keeps a valid credential during the switch, such as alternating users.

Sample spoken answer:

“We moved database credentials out of environment variables into Secrets Manager with rotation turned on. The first time it rotated, the password changed in the database, but our services had read the secret once at startup and kept trying the old one whenever they opened a new connection. Errors grew as old connections were recycled. I changed the services to use the caching client with a short refresh time and, on an authentication error, to refetch the secret and retry the connection once. I also switched rotation to the alternating-users strategy, where there are two database users and rotation updates the one not currently in use, so the old credential stays valid for a while after the switch. Then I triggered a rotation in staging during a load test to prove it. Now rotation is a non-event, and we rotate far more often than we could by hand.”

Red flag to avoid:

Turning rotation off to make the errors stop, or not knowing the app has to reload the secret.

They may ask next:
  • Why is the alternating-users strategy safer than changing one user's password?
  • How do you stop the app from calling the secrets service on every request?
Say it in 60 seconds
Medium Coding round Mid-level, Senior Practice question

14. Write a script that lists IAM users whose access keys are older than 90 days, and when each key was last used, so the team can rotate or remove them.

What the interviewer is really testing:
Whether you can write paginated, read-only automation against the IAM API and give owners the data they need to act safely.
Answer frame:

Walk everything: paginate over users, then list each user's keys.

Useful output: key age, status and last-used date, so unused keys stand out.

Safe by default: report only; deactivate before deleting, after the owner confirms.

Sample spoken answer:

“I'd use boto3's paginator on list_users, and for each user call list_access_keys, work out the key's age from its create date, and skip anything younger than 90 days. For old keys I call get_access_key_last_used, because a key that hasn't been used in months is a much easier conversation than one used an hour ago. The script only reports. When I ran this at my last company, the plan was: share the list, deactivate unused keys first rather than deleting them, wait a couple of weeks for anyone to shout, then delete. For keys still in use, we moved those workloads to roles so they didn't need long-lived keys at all. The script needs only read permissions on IAM, and I ran it from a role, not from a user with a key of its own.”

Code:
import boto3
from datetime import datetime, timezone

iam = boto3.client("iam")
now = datetime.now(timezone.utc)
MAX_AGE_DAYS = 90

for page in iam.get_paginator("list_users").paginate():
    for user in page["Users"]:
        name = user["UserName"]
        keys = iam.list_access_keys(UserName=name)["AccessKeyMetadata"]
        for key in keys:
            age = (now - key["CreateDate"]).days
            if age < MAX_AGE_DAYS:
                continue
            info = iam.get_access_key_last_used(AccessKeyId=key["AccessKeyId"])
            last = info["AccessKeyLastUsed"].get("LastUsedDate")
            last_txt = last.date().isoformat() if last else "never"
            print(f"{name}  {key['AccessKeyId']}  {key['Status']}  {age} days  last used {last_txt}")
Red flag to avoid:

Writing a script that deletes keys straight away, or skipping pagination so users are silently missed.

They may ask next:
  • Is there a faster way to get this data for every user in one call?
  • How would you stop new long-lived keys from being created in the first place?
Say it in 60 seconds

Changes and Migrations 2 questions

Hard Behavioral round Mid-level, Senior Practice question

15. An infrastructure change you were reviewing would have quietly replaced a production database. How was it caught, and what guardrails did you add so it can't happen again?

What the interviewer is really testing:
Whether you know which changes force resource replacement, read change sets or plans before applying, and protect stateful resources with policies rather than hope.
Answer frame:

The change: what looked harmless, such as a rename or a property that forces replacement.

How it was caught: reading the change set or plan and spotting the removal or replacement.

Guardrails: retain policies on stateful resources, stack protection, and a pipeline check that blocks replacements.

Sample spoken answer:

“A teammate refactored our CloudFormation template and renamed the logical ID of the RDS instance to match a new naming scheme. To CloudFormation that looks like one resource removed and a new one added, so the change set showed the old database being removed and a new one added. I caught it because our pipeline posts the change set summary on the pull request and I read every line that removes or replaces a resource. We reverted the rename. Then I added guardrails: DeletionPolicy and UpdateReplacePolicy set to retain or snapshot on every stateful resource, termination protection on the stack, a stack policy that denies replacing or deleting the database, and a pipeline step that fails if a change set replaces or removes anything tagged as stateful unless someone approves it by hand. It made some legitimate changes slower, but nobody has come close to losing data since.”

Red flag to avoid:

Trusting that a refactor is safe because the template still validates, or relying on backups as the only guardrail.

They may ask next:
  • How would you rename that resource properly if you really needed to?
  • What is the equivalent guardrail if the team used Terraform instead?
Say it in 60 seconds
Hard Behavioral round Mid-level, Senior Practice question

16. Walk me through a migration of a self-managed database onto a managed AWS database with little downtime. How did you plan the cutover and the way back?

What the interviewer is really testing:
Whether you can run a migration end to end: continuous replication, data checks, a short rehearsed cutover and a rollback path that actually works.
Answer frame:

Replicate: an initial load plus ongoing change capture until the target is caught up.

Verify: row counts, checksums on key tables, and the app tested against the target.

Cutover and rollback: a rehearsed, short write freeze, and a plan to go back if the first hour goes wrong.

Sample spoken answer:

“We had MySQL on two EC2 instances that we patched by hand, and I led moving it to Aurora MySQL. I used Database Migration Service for a full load followed by ongoing change capture from the binary log, so the target stayed a few seconds behind the source for weeks while we tested. I wrote checks comparing row counts and checksums on the biggest tables, and we ran the staging app against Aurora for a full release cycle. For the cutover we rehearsed twice: stop writes, wait for replication lag to hit zero, run the checks, point the app's config at the new endpoint and turn writes back on. Downtime was about ten minutes. For rollback, I kept the old primary and set up reverse replication so we could go back without losing new writes. We never needed it, but having it made the go decision easy.”

Red flag to avoid:

Describing a dump-and-restore over a weekend with no data checks or no way back.

They may ask next:
  • What did the data checks catch that you didn't expect?
  • Which database features did not come across cleanly, and how did you handle them?
Say it in 60 seconds

Cost and Capacity 2 questions

Medium Behavioral round Mid-level, Senior Practice question

17. Have you moved a workload to Graviton or another instance family? What did you have to change, and how did you decide it was worth it?

What the interviewer is really testing:
Whether you can run a cost or performance change as a measured experiment, including the unglamorous build and dependency work, rather than a blind switch.
Answer frame:

Why: the price-performance case and which service went first.

Work: multi-architecture builds, native dependencies, and agents or tools that lacked ARM support.

Proof: a side-by-side test on real traffic before the full switch, and how you rolled it out.

Sample spoken answer:

“Our biggest spend was a fleet of containers running a Node API, and Graviton promised better price-performance, so I proposed trying it on that one service first. The code itself didn't change, but the build did: I switched our image builds to produce multi-architecture images, and two native modules plus an old monitoring agent had no ARM builds, so I upgraded them. Then we ran a small share of production traffic on ARM tasks next to the x86 ones for a week and compared latency, errors and cost per request. ARM was a bit faster and cheaper per request, so we moved the rest over in steps. The trade-off is that every new dependency now has to support both architectures, so I added an ARM build to CI to catch that early instead of at deploy time.”

Red flag to avoid:

Switching the whole fleet at once on the strength of a vendor claim, with no side-by-side measurement.

They may ask next:
  • What would have made you stop the migration halfway?
  • How did you keep the option to roll back to x86 quickly?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

18. Where have you used Spot capacity for real work, and how did you make that workload survive interruptions?

What the interviewer is really testing:
Whether you match Spot to interruption-tolerant work and design for the interruption notice, capacity shortages and instance diversity instead of hoping it never happens.
Answer frame:

Fit: stateless or checkpointed work that can be retried, not the only copy of anything.

Handling: react to the two-minute interruption notice, drain work and checkpoint.

Availability: spread across many instance types and zones, with on-demand as a fallback.

Sample spoken answer:

“I used Spot for our nightly video transcoding workers, which pulled jobs from a queue. Each job could be retried, so losing a worker only cost time. The worker listened for the interruption notice, which gives about two minutes, stopped taking new jobs, and put its current job back on the queue if it couldn't finish. The bigger lesson was capacity: when we asked for one instance type, we sometimes couldn't get any Spot capacity at all during busy periods. I changed the Auto Scaling group to a mixed instances policy with several similar instance types across three zones, a price-capacity-optimized allocation, and a small on-demand base so the pipeline always made progress. Jobs occasionally took longer, but the batch still finished before morning at a fraction of the old compute cost, and the API servers stayed on on-demand.”

Red flag to avoid:

Putting critical stateful services on Spot, or not knowing there is an interruption notice to handle.

They may ask next:
  • Why would you never put a stateful database primary on Spot?
  • How did you check a job wasn't half-written when a worker disappeared?
Say it in 60 seconds

Reliability and DR 3 questions

Medium Situational round Mid-level, Senior Practice question

19. Your on-call rotation keeps getting paged at night by CloudWatch alarms that turn out to be nothing. How did you fix the alerting?

What the interviewer is really testing:
Whether you know the difference between a page and a dashboard signal, and can rebuild alerts around user impact without losing real warnings.
Answer frame:

Audit: list every alarm that paged in the last month and whether anyone had to act.

Page on symptoms: errors and latency users feel, with sensible datapoints-to-alarm settings.

Demote the rest: causes like CPU go to tickets or dashboards, composite alarms for combined signals.

Sample spoken answer:

“I pulled a month of pages and marked each one: did someone have to do anything? Most were CPU spikes on single instances during deploys and one-minute error blips that recovered on their own. So I changed what pages us. Pages now fire on symptoms users feel, like the load balancer's error rate and p99 latency staying above a threshold for several datapoints in a row, not one bad minute. CPU and disk alarms became tickets or dashboard widgets. Where one alarm alone was noisy, I used a composite alarm so we only page when errors and latency are both bad. Every paging alarm got a short runbook link in its description. Night pages dropped to a handful a month, and when we did get paged, people trusted it and moved fast. The trade-off is we see slow creeping problems a bit later, so we review the dashboards weekly.”

Red flag to avoid:

Silencing alarms or raising thresholds until they stop firing, with no link to what users experience.

They may ask next:
  • How do you treat missing data in an alarm, and why does it matter?
  • What alarm would you never demote?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

20. Tell me about a time you actually tested restoring from backups on AWS. What did the test show that you didn't expect?

What the interviewer is really testing:
Whether you treat a backup as unproven until restored, and can describe the specific gaps a real restore drill uncovers.
Answer frame:

The drill: what you restored, where, and against what recovery target.

Surprises: the gaps found, such as missing keys, slow restores or forgotten config.

Follow-through: what you fixed, and how often the drill now runs.

Sample spoken answer:

“We had nightly snapshots and cross-region copies, and everyone assumed we were covered. I ran a drill: restore the production database and the file bucket into a clean account in the second region, bring the app up and check it works. Three things surprised us. The snapshot copies in the other region were encrypted with a key that the restore account couldn't use, so the first attempt failed outright. The restore of our largest database took far longer than the recovery time we'd promised. And the app needed secrets and parameters we'd never backed up at all. I fixed the key policies, added a cross-region read replica for the big database so recovery became a promotion instead of a long restore, and put secrets and config into code so they could be recreated. We now run the drill every quarter, and it takes an afternoon instead of a week.”

Red flag to avoid:

Saying backups are fine because the backup job reports success, without ever restoring one.

They may ask next:
  • How did you measure how much data you would have lost?
  • Who else did you involve in the drill, and why?
Say it in 60 seconds
Hard System design round Mid-level, Senior Practice question

21. How have you set up deployments for a service on AWS so a bad release rolls itself back before most users notice?

What the interviewer is really testing:
Whether you can design a deployment path with gradual traffic shifting, health signals tied to alarms and an automatic rollback, and know what it costs to run.
Answer frame:

Shift gradually: canary or linear traffic shifting instead of all at once.

Watch real signals: alarms on errors and latency that gate each step.

Roll back on its own: the deploy tool reverts on an alarm, and changes stay backward compatible.

Sample spoken answer:

“For our ECS services I set up blue/green deployments with CodeDeploy behind the load balancer. A new version starts next to the old one, gets a small share of traffic for ten minutes, then the rest. The deployment is tied to CloudWatch alarms on the new version's error rate and latency, so if either alarm fires during the canary window, CodeDeploy shifts traffic back to the old tasks automatically. The part people miss is database changes: a rollback only works if the old code still runs against the new schema, so I set a rule that migrations are additive first and cleanup comes a release later. The costs are running two sets of tasks during each deploy and slower releases. The first month it caught two bad releases that would have been full outages, which ended the debate about speed.”

Red flag to avoid:

Relying on someone watching dashboards and redeploying the old version by hand, or ignoring schema changes.

They may ask next:
  • How do you pick the canary size and how long to wait?
  • What kinds of bad release would this setup not catch?
Say it in 60 seconds

Leadership 2 questions

Medium Behavioral round Mid-level, Senior Practice question

22. A junior engineer opened SSH to the whole internet on a production security group to debug something. How did you handle it, and what did you change afterwards?

What the interviewer is really testing:
Whether you can fix a security risk fast, coach without blame, and remove the reason people take the shortcut in the first place.
Answer frame:

Contain: close the rule, check for logins while it was open.

Coach: a private, blame-free talk about why it's risky and what to use instead.

Fix the system: a safer access path and a guardrail that catches the rule next time.

Sample spoken answer:

“Our config rules flagged a security group with port 22 open to everyone. I removed the rule straight away, then checked the instance's auth logs and our flow logs for connections from unknown addresses while it was open. Nothing had got in. Then I talked to the junior privately. He'd been stuck on an outage and couldn't reach the box any other way, which told me the real problem was ours: we had no good access path. So I set up Session Manager on our instances, which gives a shell through IAM with no open inbound port, set it to log every session, and walked the team through using it. I also added an automatic remediation that removes world-open SSH rules and posts in our channel. He later wrote the how-to page for the team himself, which I think did more than any lecture would have.”

Red flag to avoid:

Blaming the person publicly, or closing the rule and moving on without fixing why it happened.

They may ask next:
  • How would you check whether anyone actually got in while the port was open?
  • What would you do differently if it were a senior engineer who did it?
Say it in 60 seconds
Hard Situational round Mid-level, Senior Practice question

23. A manager wants the service running active-active in two regions by next quarter, but its uptime needs are modest. How do you push back?

What the interviewer is really testing:
Whether you can turn a vague wish into recovery targets, explain the real cost of multi-region, and offer a cheaper plan that meets the actual need.
Answer frame:

Get the target: agree how long the service may be down and how much data it may lose.

Show the cost: data replication, conflicts, doubled infrastructure, testing and on-call load.

Offer a ladder: solid multi-AZ now, backups copied to a second region, a warm standby only if the targets demand it.

Sample spoken answer:

“I'd start by asking what problem we're solving. When I did this last year, the answer was a customer asking about disaster recovery, not a real need for zero downtime. So I got the manager and the product owner to agree on numbers: a few hours of downtime in a full regional outage was acceptable, and losing a few minutes of data was too. Then I laid out what active-active really means for our relational database: writes in two places, conflict handling, twice the infrastructure and a lot more testing. Instead I proposed hardening what we had across availability zones, copying backups and images to a second region, and keeping infrastructure as code that could stand up a copy there. We tested the restore and hit the target. It was a fraction of the effort, and the customer was satisfied with the documented plan.”

Red flag to avoid:

Saying yes to avoid conflict, or saying no flatly without offering numbers or an alternative.

They may ask next:
  • What would the targets need to be before you agreed to active-active?
  • How do you keep a standby region from quietly going stale?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card