AWS interviews for 3 years of experience care less about definitions and more about what happened in your project, such as a Lambda that timed out, a queue that processed a message twice, an AccessDenied you traced, a stack update that failed, or an alarm that stayed quiet. This page is written for engineers with about two to four years of hands-on AWS work, the stage where you build and ship features on AWS yourself but someone else still designs the overall system. Each question shows what the interviewer is checking, the shape of a good answer and a short spoken answer. Swap the stories for your own before the interview.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Cause: in a VPC the function gets network interfaces in your subnets with no public IP, so it has no route to the internet.
Fix for the internet: run it in private subnets whose route table sends outbound traffic to a NAT gateway in a public subnet.
Fix for AWS services: use VPC endpoints for services like S3 or Secrets Manager so that traffic never needs the internet.
“We moved a Lambda into our VPC so it could talk to a private RDS database, and right after that every call it made to a payment provider just hung until the timeout. The reason is that a VPC-attached function gets network interfaces in your subnets, and those never get a public IP, even if you put them in a public subnet. So there was simply no way out. The fix was to place the function in private subnets and point their route table at a NAT gateway that sits in a public subnet. It was also reading a secret from Secrets Manager, so we added an interface endpoint for that, which kept the call inside the VPC. Since then, whenever I attach a function to a VPC, I list every outside thing it calls before I deploy.”
Raising the function timeout or memory, as if the slowness were a performance problem.
Cause: every concurrent execution environment opens its own connection, so a burst of concurrency outruns the database's connection limit.
Reuse: open the connection outside the handler so warm invocations reuse it instead of opening a new one each time.
Pool and cap: put RDS Proxy in front to share a pool of connections, and cap the function with reserved concurrency.
“We had an order API on Lambda talking to a Postgres database on RDS. When a promotion went out, concurrency jumped into the hundreds, and every execution environment opened its own connection, so we blew past the database's connection limit and requests started failing. The first thing I checked was the code, and sure enough the connection was being opened inside the handler, so even warm invocations opened a fresh one. I moved it outside the handler so each environment reuses its connection. That helped, but it can't stop a burst of new environments, so we put RDS Proxy in front, which keeps a shared pool and hands connections out as needed. We also set reserved concurrency on the function so it can never open more than the database can take.”
Only increasing the database instance size, without seeing that connections grow with concurrency.
What happened: the concrete mistake, in a sentence or two.
Stop it first: how you cut off the damage before looking for the cause.
Make it impossible: the design change, not just a promise to be careful.
“I wrote a Lambda that made thumbnails whenever an image landed in a bucket. To keep things simple, I wrote the thumbnails back into the same bucket. The trigger fired on every new object, including my thumbnails, so the function kept triggering itself. I noticed within about twenty minutes, because invocations on the dashboard shot straight up. The first thing I did was set the function's reserved concurrency to zero, which stops it instantly. Then I fixed the design: uploads go under one prefix, the trigger only listens to that prefix, and thumbnails go to a separate output bucket, so the function can't see its own output. I also added a CloudWatch alarm on invocations, and told the team what happened, since it cost us an unusual bill for the day.”
Telling a story where nothing went wrong, or blaming the service for doing what it was set up to do.
Thin handler: the handler parses the event and calls plain functions that you can unit test.
Fake AWS in unit tests: stub the SDK calls, or use a library that mocks S3 and DynamoDB in memory.
Real AWS once: invoke locally with a sample event, then run an integration test in a dev stage.
“I keep the handler thin. It pulls the bucket and key out of the event and passes them to plain functions, so most of the logic is ordinary Python I can unit test. For the parts that call AWS, I either stub the boto3 client with botocore's Stubber or use moto, which fakes S3 and DynamoDB in memory, so tests run fast and offline. I save real sample events as JSON files and use them with sam local invoke to check the wiring, like whether I'm reading the key right, since S3 URL-encodes it in the event. Then the pipeline deploys to a dev stage, where a small integration test uploads a real file and checks the item appears. Only after that does it go to production.”
Testing only by deploying straight to production and reading the logs.
Visibility timeout: a received message is only hidden for a while; if work takes longer, it reappears and another worker gets it.
At least once: a standard queue can deliver a message more than once anyway, so duplicates must be expected.
Fix: set the visibility timeout longer than the slowest processing, and make the handler idempotent with a record of what's done.
“When a worker receives a message from SQS, the message isn't removed, it's just hidden for the visibility timeout. Ours was 30 seconds, but sending an email through our provider sometimes took longer on a slow day. So the message became visible again, a second worker picked it up, and both sent the email before either deleted it. I raised the visibility timeout well above our slowest processing time. But a standard queue can also deliver the same message twice on its own, so that alone isn't enough. The real fix was making the handler idempotent: before sending, it does a conditional write to a DynamoDB table keyed by the email job id, and if that id is already there it skips the send. After that, duplicates became harmless.”
Believing SQS guarantees each message is delivered exactly once, so the duplicate must be a bug in SQS.
Dead-letter queue: a redrive policy with a max receive count moves a message aside after it fails that many times.
Partial batch: with batches, report only the failed message ids so the good ones in the batch are not retried.
Follow-up: alarm on the dead-letter queue depth, fix the cause, then redrive the messages back.
“We had a Lambda reading order events from SQS in batches of ten. One message had a malformed date, the function threw, and because the whole batch failed, all ten went back to the queue and came round again, over and over. First I added a dead-letter queue with a redrive policy, so after a few failed receives a message gets moved aside instead of retrying forever. Then I turned on partial batch responses, so the function returns just the ids of the messages that failed and the other nine get deleted normally. We also put a CloudWatch alarm on the dead-letter queue having any messages, because a DLQ nobody watches just hides failures. Once I'd fixed the date parsing, I used the redrive option to send those messages back to the main queue.”
Deleting the bad message by hand each time, or adding a DLQ with no alarm on it.
Who is calling: confirm the real principal, for example with get-caller-identity or the CloudTrail event.
Read the error: many services now say which kind of policy denied, and CloudTrail shows the exact action and resource.
Check each layer: identity policy, resource policy, permissions boundary, organization policies, and the key policy if KMS is involved.
“A deploy pipeline started failing with AccessDenied when it tried to write to a bucket in another account. My first step was to check which role was really making the call, and CloudTrail showed it was the pipeline role, not the deploy role I expected, because a step was missing its assume-role. After I fixed that, it still failed, but the error message now said no resource-based policy allowed the action. Cross-account access needs both sides, so our role policy allowed the write, but the bucket policy in the other account didn't mention our role. I asked that team to add a statement scoped to that one role and prefix. The habit I kept is to work through the layers in order rather than widening our own policy, which would never have fixed it.”
Fixing it by attaching a broad or admin policy without knowing which layer was denying.
Two permissions: reading an SSE-KMS object needs s3:GetObject and kms:Decrypt on that key; writing needs kms:GenerateDataKey.
Key policy: the key's own policy must allow the caller, or allow the account so IAM policies can grant it.
Cross-account: the AWS managed S3 key can't be shared across accounts; use a customer managed key and grant the other account.
“This one caught me once. S3 permissions were fine, but the objects were encrypted with a customer managed KMS key, and to read an object like that S3 has to ask KMS to decrypt it on your behalf, using your permissions. So the role also needs kms:Decrypt on that key, and for uploads it needs kms:GenerateDataKey. In our case the role's policy had nothing for KMS. There's a second layer too: the key has its own key policy, and unless it allows the caller or delegates to IAM in that account, an IAM policy alone won't help. I added kms:Decrypt on that key ARN only. When it's cross-account, you also can't use the default AWS managed key for S3, because you can't edit its policy, so it has to be a customer managed key.”
Checking only the bucket policy and IAM S3 actions, never looking at the object's encryption.
Store: Secrets Manager or a SecureString parameter in Parameter Store, both encrypted with KMS.
Access: the app's role is allowed to read only its own secret; the code fetches it at startup and caches it.
Rotation: Secrets Manager can rotate database credentials on a schedule, so the code must handle a changed password.
“We kept the database credentials in Secrets Manager. The service ran on ECS, its task role could read just that one secret, and the code fetched it at startup and cached it, so nothing sat in the repo, the image or the CloudFormation template. We picked Secrets Manager over a SecureString in Parameter Store mainly because it can rotate RDS passwords for you. That rotation bit us once: after the first rotation, the app kept its old cached password and started failing to connect. So I changed it to fetch the secret again when a login fails, then retry once. Before that project, I'd seen passwords sitting in plain Lambda environment variables in a template, and anyone who could read the stack could read the password.”
Keeping the password in a config file in the repo or in plain text in the infrastructure template.
Check: get-caller-identity shows the account and role; configure list shows where each setting came from.
Why it happens: leftover environment variables win over the default profile in your config file.
Habit: named profiles with short-lived SSO sessions, no long-lived keys, and the account shown in your prompt.
“The first thing I run is aws sts get-caller-identity, which tells me the account and the role I'm really using. aws configure list shows where each value came from, like an environment variable or a profile. In my case, I'd exported access keys for a test account in that terminal earlier, and environment variables win over the default profile in the config file, so my command went to the test account instead of the one I meant. Since then, I don't keep long-lived keys on my laptop at all. We use SSO, each account is a named profile, and I pass the profile on every command or set it per terminal. My shell prompt shows the current profile, and anything that deletes, I run get-caller-identity before it.”
Having no idea how to check which identity the CLI is using.
Unblock fast: read the exact denied action and resource from the error or CloudTrail with them.
Grant the gap: add just that permission, scoped, through the usual code change.
If it's truly urgent: time-boxed extra access with a clear owner, then remove it.
“I get why they're asking, they want to ship. But admin on a pipeline role means anything that can change the pipeline can do anything in the account, and those grants never seem to get removed. So I'd sit with them for ten minutes and look at the actual error. It usually names the action and the resource, and if not, CloudTrail does. Most of the time it's one or two missing actions, like permission to pass a role or to write to a new bucket. I'd add exactly that, scoped to the resource, as a small change to the role's template and get it reviewed quickly. That usually unblocks them faster than arguing about admin. If it really was urgent and complex, I'd ask our lead for a time-boxed exception, not a permanent one.”
Granting admin to get the deploy through, or refusing without helping them find the real gap.
Filter: pick the log group and time range, then filter to the error lines.
Group over time: count errors in small time bins to see when it started and whether it's still happening.
Drill in: look at the latest raw messages, then group by an error type or path field if the logs are structured.
“I'd open Logs Insights, pick the service's log group and the last hour, and first count the errors in five-minute bins, because the shape tells me a lot: did it start at a deploy, is it still climbing, or was it one burst. Then I'd run a second query to see the latest error messages themselves. In my last project our logs were JSON, so I could also group by the error code or the route, which quickly showed that nearly every error came from one endpoint. For Lambda I also query the REPORT lines, which carry the duration of each invocation, to see if timeouts are the real cause. Once I know the pattern, I save the query so the next person on call doesn't have to write it again.”
filter @message like /ERROR/
| stats count(*) as errors by bin(5m)
fields @timestamp, @message
| filter @message like /ERROR/
| sort @timestamp desc
| limit 50
filter @type = "REPORT"
| stats max(@duration) as slowest, count(*) as calls by bin(5m)
Only ever opening individual log streams in the console and reading them by eye.
No data: a job that never runs sends no error data points, so an errors alarm has nothing to breach.
Missing data setting: decide whether missing data counts as breaching, not breaching, or missing.
Alarm on success: a heartbeat metric or an alarm on too few invocations catches a job that silently stops.
“Our nightly export ran on Lambda from an EventBridge schedule, and I'd put an alarm on the function's Errors metric. Then someone disabled the rule during a change and forgot to turn it back on. The function didn't run, so it produced no errors, and the alarm just sat there, because no invocations means no data points to breach. The fix was to alarm on the thing we actually care about, that the job succeeded. I set an alarm on the Invocations metric being below one over a day, and set missing data to count as breaching, so silence itself raises the alarm. We also found the code caught some exceptions and only logged them, which kept the Errors metric at zero, so I added a metric filter on those log lines too.”
Assuming an errors alarm covers the case where the job never runs at all.
Find the drift: run drift detection on the stack to see which resources differ from the template.
Decide the truth: either put the manual change into the template or undo it by hand, then update again.
Prevent it: limit console write access in production and make fixes go through the template.
“During an incident, a teammate added an inbound rule to a security group in the console to let a new service in. A week later I added the same rule properly to our template, and the stack update failed because that rule already existed, so everything rolled back. I ran drift detection on the stack, and it showed the security group had a rule the template didn't know about. Since my change was going to add that exact rule, I removed the manual one in a quiet window and ran the update again, and CloudFormation put it back as part of the stack. After that we agreed that any change made by hand during an incident gets a ticket to move it into the template the next day.”
Deleting and recreating the whole stack in production to get rid of the error.
Versions and aliases: each publish makes a fixed version; callers use an alias like live, so a rollback means pointing the alias back.
Gradual shift: a canary or linear deployment sends a small slice of traffic to the new version first.
Automatic rollback: error alarms attached to the deployment shift traffic back by themselves if they fire.
“When it happened, API Gateway was calling the function's live alias, and every deploy published a new version, so the rollback was quick. I pointed the alias back at the previous version and the errors stopped within a minute, with no rebuild. The bug was new code expecting an environment variable that existed in dev but not in production. Afterwards I changed our SAM template so deploys go through CodeDeploy. With an auto-published alias and a canary deployment preference, a small slice of traffic goes to the new version for a few minutes first. We attached a CloudWatch alarm on the alias's errors, so if it fires during that window, traffic shifts back on its own. I also added a pre-traffic hook that invokes the new version once with a test event before any real traffic reaches it.”
OrderFunction:
Type: AWS::Serverless::Function
Properties:
Handler: app.handler
Runtime: python3.12
CodeUri: src/
AutoPublishAlias: live
DeploymentPreference:
Type: Canary10Percent5Minutes
Alarms:
- !Ref OrderErrorsAlarm
Hooks:
PreTraffic: !Ref PreTrafficCheck
Rebuilding and redeploying old code from source while errors keep coming, or deploying with no versions or aliases at all.
The feedback: what the reviewer pointed out, quoted plainly.
Why they were right: the risk or the cost you hadn't seen.
What you do now: the habit you kept, and how you pass it on.
“Early on, I wrote a CloudFormation template for a new service where the Lambda's role had s3:* on every resource, because I wasn't sure which actions it needed and I wanted it to just work. The reviewer asked me to list exactly what the function called. It was two actions on one bucket. She pointed out that with the wildcard, a bug or a stolen credential could delete every bucket in the account. That stuck with me. Now I start with no permissions, run the function, and add what the AccessDenied errors ask for, scoped to the exact bucket and prefix. The same reviewer also caught me hard-coding an image ID for one region. When I review others' templates, wildcards in IAM are the first thing I look for.”
Saying you've never had useful feedback, or describing feedback without any change in how you work.
Why no alarm: the built-in EC2 metrics cover CPU and network, not how full the file system is or memory use; that needs the CloudWatch agent.
Fix now: clear what filled the disk, or grow the EBS volume while it stays attached, then extend the partition and the file system.
Fix the cause: rotate logs or ship them off the box, and alarm on disk used through the agent.
“We heard about it from users, not from CloudWatch, and that was the first lesson: the built-in EC2 metrics show CPU and network, but not how full the disk is. I got in with Session Manager, ran df, and the root volume was full. du pointed at the app's log folder, where debug logging had been left on after a release and nothing rotated the files. I compressed the old logs and moved them to S3 to get the app back up. Since we wanted headroom, I also increased the EBS volume size while it stayed attached, then grew the partition with growpart and the file system with xfs_growfs, with no reboot. For the cause, we added logrotate, turned debug logging off, and installed the CloudWatch agent with an alarm on disk used.”
df -h # which file system is full
sudo du -xh /var --max-depth=2 | sort -h | tail
# after increasing the volume size in EBS:
lsblk # confirm the new size
sudo growpart /dev/nvme0n1 1 # grow partition 1
sudo xfs_growfs -d / # XFS; use resize2fs for ext4
Rebooting or replacing the instance and calling it fixed, without finding what filled the disk or adding a disk alarm.
502: the target closed the connection or sent a bad response; a common cause is an app keep-alive shorter than the balancer's idle timeout.
504: the target didn't answer within the idle timeout, or the balancer couldn't connect at all.
Where to look: balancer 5xx versus target 5xx metrics, and the access logs for target status and timings.
“First I split them using the metrics: errors generated by the load balancer and errors returned by the targets are counted separately. Ours were from the balancer. The 502s were random and low, and the cause was keep-alive. Our Node service closed idle connections after five seconds, which is Node's default, while the load balancer keeps them open much longer and reuses them. Now and then it sent a request just as the app closed the connection, and that came back as a 502. Setting the app's keep-alive timeout longer than the balancer's idle timeout stopped it. The 504s were different: a report endpoint ran longer than the idle timeout, so the balancer gave up waiting. We moved that report to a background job instead of raising the timeout.”
Treating every 5xx as an app bug without checking whether the balancer generated it.
Why it's replaced: with load balancer health checks on, an instance that fails them is marked unhealthy and replaced.
Usual causes: a grace period shorter than start-up, a health path that needs auth or a database, or the app failing on boot.
How to debug: read the scaling activity history, target health reasons, and the instance's start-up logs.
“The activity history showed every instance being terminated for failing load balancer health checks. The target group said the health checks were timing out. So either the app wasn't up yet, or it was broken. I suspended the ReplaceUnhealthy process so one instance would stay alive, and looked at its cloud-init output log. The new release had added a warm-up step, and the app now took longer to start than our health check grace period, so the group killed it before it was ready. On top of that, our health path queried the database, so a slow database could also fail it. We raised the grace period, and I changed the health endpoint to report only whether the app itself is up. Then I resumed the process.”
Switching off health checks altogether so the instances stop being replaced.
Scan cost: a Scan reads every item, and the filter is applied after the read, so you pay for everything.
Pages: each Scan call returns at most 1 MB, so a big table means many calls.
Fix: add a global secondary index keyed on what the page asks for, and Query it.
“The page listed a customer's open orders. The code scanned the orders table with a filter on customer id and status. That worked in testing with a small table, but a Scan reads every item and only then applies the filter, so we were paying to read the whole table to return five orders, and it took several paged calls to get through it. I added a global secondary index with customer id as the partition key and status plus created date as the sort key, then changed the code to Query that index with a key condition. The page went from seconds to milliseconds, and read usage dropped sharply. One thing I had to explain in review is that reads from a global index are eventually consistent, which was fine for that page.”
Thinking the filter expression makes a Scan cheaper, or fixing it by just raising capacity.
Cause: both requests read the item, changed it in code, and wrote it back, so the second write replaced the first.
Optimistic locking: keep a version number and only write if it still matches what you read.
Atomic updates: for counters, update in place with an update expression instead of read-modify-write.
“The bug was a read-modify-write. Two requests read the same order, each changed a field in code, and each wrote the whole item back with put_item, so whichever landed second wiped out the other one's change. I fixed it with optimistic locking. The item carries a version number, and the update has a condition that the version still equals the one I read. If someone else wrote in between, DynamoDB rejects it with a conditional check failure, and I read the item again and retry. I also switched to update_item so we only touch the fields we change. For a simple counter, you don't need a version at all, you just add to it in the update expression, which DynamoDB applies atomically.”
from botocore.exceptions import ClientError
try:
table.update_item(
Key={"order_id": order_id},
UpdateExpression="SET #s = :new, #v = #v + :one",
ConditionExpression="#v = :expected",
ExpressionAttributeNames={"#s": "status", "#v": "version"},
ExpressionAttributeValues={":new": new_status, ":one": 1, ":expected": version},
)
except ClientError as e:
if e.response["Error"]["Code"] == "ConditionalCheckFailedException":
retry_with_fresh_read(order_id)
else:
raise
Expecting DynamoDB to lock the item automatically, or wrapping the write in a retry that doesn't re-read.
Cause: list_objects_v2 returns at most 1000 keys per call, and the script read only the first page.
Fix: use a paginator, or loop on the continuation token until the response isn't truncated.
Habit: assume every list call in any AWS SDK is paginated, and test on real volumes.
“Almost certainly pagination. list_objects_v2 hands back at most 1000 keys per call, with a flag saying there's more and a continuation token for the next page. My script called it once, got the first thousand, and happily finished. It worked on the test bucket because that had a few hundred files, so the bug only showed up in production. The fix is to use boto3's paginator, which follows the tokens for you. Also, a page with no matches has no Contents key at all, so I read it with get and a default of an empty list. Since then I treat every list or describe call as paginated, and before running a script like that for real, I print a count and check it against what I expect.”
import boto3
s3 = boto3.client("s3")
paginator = s3.get_paginator("list_objects_v2")
keys = []
for page in paginator.paginate(Bucket=bucket, Prefix="exports/"):
for obj in page.get("Contents", []):
keys.append(obj["Key"])
print(len(keys), "objects found")
Assuming one list call returns everything in the bucket.
The task and estimate: what it was and what you said.
What you missed: the specific hidden dependencies, not a vague 'it was harder'.
What you do now: the checklist or spike you use before giving a number.
“I was asked to move a service's instances from public subnets into private ones, and I said two days, since it felt like a subnet change. It took almost two weeks. What I missed was everything that quietly depended on those public IPs. The instances pulled packages at boot, so we needed a NAT gateway. They read from S3 a lot, so we added a gateway endpoint to keep that off the NAT. People SSH'd in directly, so we set up Session Manager. And a partner allowlisted our old IPs, so they had to add the NAT's address, which took a week of emails. Now, before I estimate a network change, I spend half a day listing everything that goes in and out, and I give a range with those unknowns called out.”
Blaming the task or AWS for the miss, with nothing learned about estimating.
Find the owner: tags, CloudTrail for who created it, and asking in the team channel.
Check use: whether a volume is attached, when the bucket was last read or written, and what refers to it.
Delete safely: snapshot or back up, tag with a delete date, announce it, then delete after the wait.
“I wouldn't delete anything on day one. First I'd try to find an owner: tags if there are any, and CloudTrail to see who created it, if it's recent enough. Then I'd check use. An unattached EBS volume is a good candidate, but I'd look at its name and snapshots, because sometimes it's a detached data disk someone still needs. For the bucket, I'd look at its request metrics or access logs, and search our templates and code for its name. Then I'd post a list in the team channel with a delete date two weeks out, tag everything with that date, and snapshot the volumes before deleting them. The snapshots are cheap to keep for a while and they turn a wrong call into a small restore instead of an incident.”
Deleting everything without tags straight away because nobody answered fast enough.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.