Multi-Account Security • Resilience • Data at Scale • Serverless Corners • Leadership • 2026

AWS Interview Questions for 10+ Years Experience (Senior)

AWS interviews for 10+ years of experience skip basics like what EC2 is and go after what experience teaches, such as how guardrails work across many accounts, where encryption and scaling quietly break, why a health check can take down a fleet, how to keep one failure from reaching every customer, and how you lead migrations, set standards and grow a team. This page is written for cloud engineers and architects with around eight to fifteen years of experience, interviewing for senior, staff, principal or lead roles. Each question shows what the interviewer is checking, a shape for your answer and a sample to adapt. Say it in your own words, with your own stories.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Security and Governance 6 questions

Hard Technical round Senior Practice question

1. In a multi-account setup, how do service control policies interact with IAM policies, and what can an SCP never do?

What the interviewer is really testing:
Whether you know SCPs set the ceiling for a whole account and grant nothing, and whether you know the accounts and roles they do not touch.
Answer frame:

Ceiling, not grant: an SCP limits what any principal in a member account can do; access still needs an IAM allow.

Who it hits: every user and role in member accounts, including the account's root user.

Who it misses: the management account and service-linked roles are not restricted by SCPs.

Design: keep workloads out of the management account and test SCPs on a sandbox OU first.

Sample spoken answer:

“I think of an SCP as the fence around an account, not a key to any door. For a request to succeed, the SCPs on every level from the root down to the account have to allow it, and then IAM has to allow it too. So an SCP that allows everything grants nothing on its own. The catch people miss is scope. SCPs do restrict the root user of a member account, which is why they're the right place for things like 'nobody leaves the organization' or 'no resources outside our approved regions'. But they don't apply to the management account at all, and they don't restrict service-linked roles. That's why I keep nothing running in the management account and treat access to it as the most sensitive thing we have. I also roll any new SCP out to a sandbox OU first, because a bad deny can break every pipeline in the org at once.”

Red flag to avoid:

Saying an SCP grants permissions, or assuming it also protects the management account.

They may ask next:
  • How would you write a region restriction that doesn't break global services?
  • Why is a deny-list approach to SCPs usually easier to live with than an allow-list?
Say it in 60 seconds
Hard Technical round Senior, Mid-level Practice question

2. A third-party tool needs a role in your customers' AWS accounts. What is the confused deputy problem, and how do you design the trust policy to stop it?

What the interviewer is really testing:
Whether you understand why trusting another account's principal is not enough when that account acts for many customers.
Answer frame:

The risk: the vendor's account can assume roles for many customers, so a customer could trick it into acting on someone else's account.

The fix: require a unique external ID, issued by the vendor per customer, in the trust policy condition.

For AWS services: use aws:SourceArn or aws:SourceAccount conditions for the same reason.

Plus: least privilege on the role itself, and a narrow principal instead of the whole account where possible.

Sample spoken answer:

“The deputy here is the vendor's service. It holds permission to assume roles in lots of customer accounts. If the only check is 'trust the vendor's account', then customer A could type customer B's role ARN into the vendor's settings page, and the vendor would happily assume it, because from AWS's side the vendor is allowed. The fix is an external ID. The vendor generates a unique value per customer and always passes it when assuming the role, and the customer's trust policy only allows the assume call when that exact value is present. Customer A can't supply B's external ID because the vendor controls it. The same idea applies when an AWS service assumes a role for you: I add a source ARN or source account condition so the service can only act for my own resources. And the role itself still only gets the read actions the tool needs.”

Code:
{
  "Effect": "Allow",
  "Principal": { "AWS": "arn:aws:iam::111122223333:root" },
  "Action": "sts:AssumeRole",
  "Condition": {
    "StringEquals": { "sts:ExternalId": "customer-7f3a9c" }
  }
}
Red flag to avoid:

Treating the external ID as a secret password, or thinking an account-level trust alone is safe.

They may ask next:
  • Why should the vendor generate the external ID instead of letting the customer choose it?
  • How would you audit which third parties can assume roles into your accounts today?
Say it in 60 seconds
Hard System design round Senior Practice question

3. Product teams want to create their own IAM roles for their services without waiting on the platform team. How do you let them do that safely?

What the interviewer is really testing:
Whether you can delegate IAM without letting anyone escalate their own privileges, which is the core use of permission boundaries.
Answer frame:

Boundary: a managed policy that caps what any role the team creates can ever do.

Enforce it: allow iam:CreateRole only when the request attaches that boundary.

Close the gaps: deny editing or deleting the boundary policy and removing a boundary from a role.

Naming: restrict teams to roles under their own path or prefix so they can't touch others.

Sample spoken answer:

“The risk with letting teams create roles is simple escalation: create a role with admin rights, then assume it. Permission boundaries solve that. A boundary is a managed policy that sets the maximum a role can do; the role's effective permissions are the overlap between its own policies and the boundary. So I write one boundary per team, covering their services and nothing in IAM or billing. Then the team's deploy role may create roles only if the create call attaches that exact boundary, which I enforce with a condition on the permissions boundary key. The part people forget is the escape hatches: I also deny changing the boundary policy itself, deleting it, and removing the boundary from a role. And I scope their role names to a path, so they can't edit the platform's roles. With that, teams ship on their own and the worst they can build is a role no stronger than the boundary.”

Code:
{
  "Effect": "Allow",
  "Action": ["iam:CreateRole", "iam:PutRolePermissionsBoundary"],
  "Resource": "arn:aws:iam::111122223333:role/team-orders/*",
  "Condition": {
    "StringEquals": {
      "iam:PermissionsBoundary": "arn:aws:iam::111122223333:policy/team-orders-boundary"
    }
  }
}
Red flag to avoid:

Letting teams create roles with only a code review as protection, or forgetting to deny removing the boundary.

They may ask next:
  • What happens to effective permissions when the boundary allows an action but no identity policy does?
  • How would you catch a team that found a way around the boundary?
Say it in 60 seconds
Hard Technical round Senior, Mid-level Practice question

4. Why does IMDSv2 matter for security, and how would you enforce it across hundreds of accounts?

What the interviewer is really testing:
Whether you know how instance credentials leak through server-side request forgery and how the session token design blocks the common attack.
Answer frame:

The attack: a web app tricked into fetching a URL can fetch the metadata endpoint and return the instance role's credentials.

IMDSv2: needs a token from a PUT request with a TTL header first, which a simple forged GET can't do.

Hop limit: the token response has a hop limit; raise it only when containers truly need metadata.

Enforce: require tokens in launch templates, and block launches without it using an SCP condition, with Config to find old instances.

Sample spoken answer:

“The classic problem is server-side request forgery. If an app will fetch any URL a user gives it, an attacker points it at the metadata address and gets back the temporary credentials of the instance's role. IMDSv1 answered a plain GET, so that worked. IMDSv2 is session based: you first send a PUT with a TTL header to get a token, then send that token as a header on every read. Most forgery bugs can only make a GET with no custom headers, so they fail. There's also a hop limit on the token response, which stops it reaching further than it should, for example into containers that shouldn't have the host's role. To enforce it at scale, I set tokens to required in every launch template and AMI pipeline, add an SCP that denies launching instances unless the metadata tokens condition says required, and use a Config rule to list the old instances still allowing v1 so teams can fix them.”

Red flag to avoid:

Saying IMDSv2 is just a newer API version with no security difference.

They may ask next:
  • Why is a role on an instance still risky even with IMDSv2 on?
  • How would you find which applications still call the metadata service without a token before you enforce it?
Say it in 60 seconds
Hard Technical round Senior Practice question

5. An S3-heavy workload using SSE-KMS starts failing with KMS throttling errors. Explain how KMS fits into each request and how you fix it.

What the interviewer is really testing:
Whether you understand envelope encryption well enough to see that every object read or write can become a KMS call.
Answer frame:

Envelope encryption: S3 gets a data key from KMS to encrypt the object; the KMS key never leaves KMS.

Per request: writes ask KMS for a data key and reads ask KMS to decrypt it, so request volume maps to KMS volume.

Fix: turn on S3 Bucket Keys, which use a short-lived bucket-level key and cut calls to KMS sharply.

Also: KMS quotas are per account and region, so check who else shares them before asking for an increase.

Sample spoken answer:

“SSE-KMS uses envelope encryption. When I put an object, S3 asks KMS for a data key, encrypts the object with it, and stores the data key encrypted next to the object. When I read it, S3 sends that encrypted data key to KMS to decrypt. So at high request rates, every S3 call is also a KMS call, and KMS has request-rate quotas per account and region. That's where the throttling comes from. The first fix is S3 Bucket Keys: S3 gets a bucket-level key from KMS and uses it to make data keys locally for a while, so the KMS call count drops massively. It applies to new objects, so old ones still hit KMS until they're rewritten. Then I check what else in the account shares the quota, because one noisy job can starve the rest, and only then ask for a quota increase. I'd also make sure clients retry throttles with backoff and jitter.”

Red flag to avoid:

Suggesting you switch off encryption, or not knowing that each read also calls KMS.

They may ask next:
  • How would you find which application is using most of the account's KMS request quota?
  • What changes in CloudTrail once Bucket Keys are on?
Say it in 60 seconds
Medium Technical round Senior, Mid-level Practice question

6. An IAM admin has full KMS permissions in their policy but still gets access denied on a customer managed key. Why can that happen?

What the interviewer is really testing:
Whether you know that the key policy, not IAM, is the primary control on a KMS key.
Answer frame:

Key policy first: every KMS key has a key policy, and IAM policies only count if the key policy lets the account use them.

The default: the default key policy trusts the account, which switches IAM on for that key.

The trap: a custom key policy that names only certain roles makes IAM grants useless for everyone else.

Lockout: KMS refuses a key policy that would stop the caller changing it again, but that check can be bypassed and doesn't protect other admins.

Sample spoken answer:

“KMS is different from most services because the key policy is the main gate. IAM permissions only count for a key if its key policy lets the account delegate access through IAM. The default key policy does that by trusting the account, so people get used to IAM just working. But if someone wrote a custom key policy naming only one application role, an admin with full KMS rights in IAM still gets denied. So I check the key policy first, not the user. The bigger risk is lockout. KMS has a safety check that rejects a policy which would stop the person setting it from changing it again, but it can be bypassed, and it doesn't protect anyone else. If every admin is locked out, only AWS support can help, and that's slow. So key policies live in code, get reviewed like production changes, and always keep a break-glass admin role.”

Red flag to avoid:

Looking only at the user's IAM policy and never at the key policy.

They may ask next:
  • How do grants differ from key policies, and when would you use one?
  • How would you share a key with another account without trusting all of it?
Say it in 60 seconds

Resilience 4 questions

Hard System design round Senior Practice question

7. What is static stability, and why should your recovery plan not depend on launching new resources during a regional incident?

What the interviewer is really testing:
Whether you design for the fact that control planes are usually the first thing to struggle during a large event, while data planes keep running.
Answer frame:

Control vs data plane: the control plane creates and changes resources; the data plane serves the running ones.

The trap: plans that scale up, launch instances or change DNS records during the incident need the control plane at its worst moment.

Static stability: pre-provision enough so losing one AZ needs no action, for example each of three AZs sized for half the peak.

Failover: use mechanisms that act in the data plane, like health-checked DNS records, and rehearse them.

Sample spoken answer:

“Every AWS service has a control plane, which creates, changes and deletes things, and a data plane, which serves what's already running. Data planes are simpler and built to be more available. During a big incident the control plane is often what struggles, because everyone is trying to launch capacity at once. So if my recovery plan is 'when an AZ dies, Auto Scaling launches more instances elsewhere', I'm betting on the weakest part at the worst time. Static stability means the system keeps working through a failure without having to change anything. For AZ loss, I run in three AZs and give each enough capacity for half the peak, so losing one still leaves the full peak covered. That costs more, and I'd show leadership exactly what it buys. For failover between regions I prefer health-checked DNS records that flip on their own, and we rehearse it on a normal weekday.”

Red flag to avoid:

Answering that Auto Scaling will just replace the lost capacity automatically.

They may ask next:
  • What parts of your current architecture would need a control plane call to recover?
  • How would you explain the extra capacity cost to a finance lead?
Say it in 60 seconds
Hard System design round Senior, Mid-level Practice question

8. A downstream service slowed down for two minutes, but the whole platform stayed down for twenty. How can retries turn a small blip into a long outage, and how do you design against it?

What the interviewer is really testing:
Whether you know that retries multiply load at every layer and keep a recovering service pinned down, and whether you know the standard defences.
Answer frame:

Amplification: retries stack across layers; three layers each trying three times can mean twenty-seven calls at the bottom for one request.

Retry in one place: one layer retries, with capped exponential backoff, jitter and a retry budget.

Stop sending: circuit breakers and load shedding give a struggling service room to recover.

Timeouts that fit: each timeout shorter than its caller's, so nobody works on requests already given up on.

Sample spoken answer:

“The danger is that retries multiply. If the edge, the API and the service layer each try three times, one click can become twenty-seven calls on the database, right when it's slowest. The retries also arrive in waves, so when the service recovers, the backlog knocks it over again. That's how two minutes becomes twenty. My defences: retry at one layer only, usually the one that knows whether the work is safe to repeat, with capped exponential backoff and jitter so clients spread out. I add a retry budget, so retries can never be more than a small share of normal traffic. Circuit breakers stop calls to a dependency that's clearly failing, and the service sheds load it can't handle instead of queueing forever. Timeouts get shorter as you go down the stack. And since the AWS SDKs already retry with backoff, I check we aren't wrapping their retries in our own.”

Red flag to avoid:

Adding more retries or longer timeouts as the fix for a flaky dependency.

They may ask next:
  • Where would you put the retry in a chain of API Gateway, Lambda and DynamoDB?
  • How would you test that the system actually recovers once a dependency comes back?
Say it in 60 seconds
Hard System design round Senior Practice question

9. One large customer's traffic spike, or one bad deploy, takes down the service for every customer. How would you redesign it so a failure only hits a small slice of them?

What the interviewer is really testing:
Whether you can reason about blast radius with cells and staged rollout, and whether you see what that design costs.
Answer frame:

Cells: independent full copies of the service, each with its own compute and data, serving a fixed set of customers.

Thin router: a very simple layer maps each customer to a cell and does little else, so it rarely breaks.

Deploy by cell: changes reach one small cell first and widen only if its alarms stay quiet.

Noisy neighbours: per-customer limits, and a cell of its own for a customer that outgrows a shared one.

Sample spoken answer:

“Right now everyone shares one fleet and one database, so any failure is everyone's failure. I'd move to cells. A cell is a complete, independent copy of the service, with its own compute, queues and data, often in its own account, and sized to a known maximum. A thin routing layer looks up which cell a customer belongs to and forwards the request. It does almost nothing else, so it can be the most stable part of the system. Now a poison request, a runaway customer or a bad config hurts one cell, not the whole product. Deploys go cell by cell, starting with a small one, with automatic rollback if its alarms fire. Big customers get per-customer limits, or a cell of their own. The costs are real: more copies to run, cross-cell features like global search get harder, and moving a customer between cells needs a data migration plan. I'd present it to leadership as how many customers any one incident can touch.”

Red flag to avoid:

Answering only with more instances or a bigger database, which keeps every customer inside the same failure.

They may ask next:
  • How would you move a customer from one cell to another without downtime?
  • How is shuffle sharding different from plain cells, and when would you use it?
Say it in 60 seconds
Hard Technical round Senior, Mid-level Practice question

10. What does an ALB do when every target in a target group fails its health check, and how can that combine badly with a deep health check and an Auto Scaling group?

What the interviewer is really testing:
Whether you know how deep health checks turn one dependency blip into a fleet-wide outage, and how the load balancer behaves when every target is unhealthy.
Answer frame:

The chain: every target fails its check at once because they share the database.

ALB behaviour: when all targets are unhealthy, the ALB fails open and sends traffic to all of them anyway.

The real damage: an ASG using ELB health checks can start replacing every instance, so the fleet restarts cold.

Better design: shallow liveness for the load balancer, dependency checks in metrics and alarms.

Sample spoken answer:

“This is a nice trap. If the health check calls the database, a database blip makes every instance fail at the same moment, because they all share it. The load balancer itself is fairly forgiving here: when every target in every enabled zone is unhealthy, an ALB fails open and keeps sending traffic to all of them, since sending nowhere helps no one. The bigger danger is the Auto Scaling group. If it uses ELB health checks, it sees the whole fleet unhealthy and starts terminating and replacing instances. Now a thirty-second database blip turns into many minutes of cold starts, warming caches and a thundering herd on the database when it recovers. So I keep the load balancer check shallow, meaning the process is up and can serve, and I watch dependencies through metrics and alarms. If a deep check is needed, I make it degrade softly instead of failing the whole node.”

Red flag to avoid:

Saying a deep health check is always safer because it catches more problems.

They may ask next:
  • When would you still want a deep health check?
  • How does a health check grace period change this story?
Say it in 60 seconds

Data at Scale 3 questions

Hard Coding round Senior, Mid-level Practice question

11. Your DynamoDB users table is keyed on user id, and two sign-ups with the same email both succeeded. Why didn't a GSI on email stop it, and how do you enforce uniqueness? Show me the write.

What the interviewer is really testing:
Whether you know DynamoDB only enforces uniqueness on the primary key, that a GSI is updated asynchronously, and how a transaction can guard a second attribute.
Answer frame:

Why it failed: a GSI never enforces uniqueness and is updated asynchronously, so check-then-write races.

Guard item: write a second item whose key is the email, next to the user item.

Atomic: both puts in one transaction, each with an attribute_not_exists condition, so both succeed or neither does.

Upkeep: transactional writes use twice the write capacity, and an email change swaps the guard item in one transaction.

Sample spoken answer:

“DynamoDB only guarantees uniqueness on the primary key. A GSI on email will happily hold two items with the same email, and it's updated asynchronously, so 'query the index, then write' has a race: two requests both see no match and both write. The pattern I use is a guard item. Next to the user item, I write a second item in the same table whose key is EMAIL plus the address. Both puts go in one TransactWriteItems call, each with a condition that the key doesn't already exist. If the email is taken, the whole transaction is cancelled, and I return 'already registered'. The costs: transactional writes use twice the write capacity of normal ones, and changing an email means deleting the old guard and adding the new one in the same transaction. I also normalise the email first, or two spellings of one address slip through.”

Code:
import boto3
from botocore.exceptions import ClientError

ddb = boto3.client("dynamodb")

def create_user(user_id, email):
    email = email.strip().lower()
    guard = "attribute_not_exists(pk)"
    try:
        ddb.transact_write_items(TransactItems=[
            {"Put": {"TableName": "users", "ConditionExpression": guard,
                     "Item": {"pk": {"S": f"USER#{user_id}"}, "email": {"S": email}}}},
            {"Put": {"TableName": "users", "ConditionExpression": guard,
                     "Item": {"pk": {"S": f"EMAIL#{email}"}, "user_id": {"S": user_id}}}},
        ])
    except ClientError as e:
        reasons = e.response.get("CancellationReasons", [])
        if any(r.get("Code") == "ConditionalCheckFailed" for r in reasons):
            raise ValueError("user or email already registered")
        raise
Red flag to avoid:

Querying the index first and then writing, or saying a GSI can be marked unique.

They may ask next:
  • What happens if two transactions touch the same item at the same moment?
  • How would you add guard items for existing users when some duplicates already exist?
Say it in 60 seconds
Medium Technical round Senior, Mid-level Practice question

12. You set up S3 replication to another region for disaster recovery. What does it not copy, and what would you check before you trusted it?

What the interviewer is really testing:
Whether you know the gaps in S3 replication that tend to surprise teams during a real recovery.
Answer frame:

Only new writes: a rule copies objects written after it exists; older objects need S3 Batch Replication.

Deletes: delete markers are copied only if you turn that on, and deleting a specific version is never replicated.

Prerequisites: versioning on both buckets; KMS-encrypted objects replicate only when the rule opts in and can use a key in the destination.

Proof: check replication status and metrics, and read from the replica in a drill.

Sample spoken answer:

“The first surprise is that a new replication rule only covers objects written after it's created. Everything already in the bucket stays put unless I run S3 Batch Replication. Second, deletes. Delete markers are only replicated if I turn that on, and deleting a specific version is never replicated. That's useful protection if someone wipes the source, but it means the replica isn't a mirror. Third, prerequisites: both buckets need versioning, and objects encrypted with KMS keys are skipped unless the rule opts in and the replication role can use a key in the destination region. Before I'd trust it, I'd check the replication status on a sample of objects, watch the replication metrics for failures and lag, and in a drill point a real reader at the replica. Bucket policies, lifecycle rules and access points don't come along, and the app needs those there too.”

Red flag to avoid:

Assuming replication makes the second bucket a complete, exact mirror from day one.

They may ask next:
  • How would you keep the replica safe even if the source account is compromised?
  • What changes when the destination bucket is in a different account?
Say it in 60 seconds
Hard Technical round Senior Practice question

13. How is Aurora's storage different from standard RDS, and what does that change for replica lag, failover and cloning?

What the interviewer is really testing:
Whether you know Aurora's shared storage layout well enough to explain why its replicas, failover and clones behave differently from classic RDS.
Answer frame:

Storage: one shared cluster volume keeping six copies across three AZs; the write quorum is four of six, the read quorum three of six.

Replicas: readers attach to the same volume, so they don't replay changes into their own copy and usually lag far less than classic read replicas.

Failover: Aurora promotes a reader, chosen by the promotion tiers you set, and moves the cluster endpoint to it.

Clones: fast clones share pages copy-on-write, so a full-size copy is quick and only changed pages take new space.

Sample spoken answer:

“Standard RDS is an engine on an instance with its own disk, and Multi-AZ adds a standby with its own copy. Aurora separates compute from storage. One cluster volume keeps six copies of the data across three AZs, and a write is durable once four of the six acknowledge it. So it can lose an AZ and keep writing, and it repairs copies in the background. Readers attach to that same volume instead of replaying changes into their own storage, which is why their lag is usually far lower than a classic read replica's. It isn't zero, so read-after-write still goes to the writer. On failover, Aurora promotes a reader based on the promotion tiers I set, and moves the cluster endpoint to it, so I keep a reader sized like the writer in the top tier. And clones use copy-on-write on the shared volume, so a team can get a full-size copy of production data quickly, and only pages that change take new space.”

Red flag to avoid:

Describing Aurora as just a faster MySQL with no difference in storage or failover.

They may ask next:
  • How would you route read traffic so a failover doesn't overload the new writer?
  • When would Aurora be the wrong choice over standard RDS?
Say it in 60 seconds

Serverless at Scale 2 questions

Hard Technical round Senior, Mid-level Practice question

14. A Lambda consuming a Kinesis stream shows iterator age climbing on one shard while the others are fine. What's likely happening, and which settings change the behaviour?

What the interviewer is really testing:
Whether you know stream sources keep order per shard, so one failing record or one hot shard stalls everything behind it, and which event source mapping settings deal with that.
Answer frame:

Order blocks: Lambda reads each shard in order, so a batch that keeps failing is retried while nothing behind it moves.

Or a hot shard: one partition key with too much traffic overloads a single shard.

Failure settings: bisect the batch on error, cap retry attempts and record age, and add an on-failure destination.

Throughput: raise the parallelisation factor, and fix the partition key for good.

Sample spoken answer:

“Iterator age is how far behind the newest record the consumer is, so one shard falling behind points at that shard. There are two usual causes. The first is a record the function can't process. With streams, Lambda keeps order per shard, so when a batch fails it retries that same batch, and by default it keeps going until the records expire from the stream. Everything behind it waits. The fixes live on the event source mapping: bisect batch on error, so Lambda splits the batch to isolate the bad record, a maximum retry count and maximum record age, so it gives up, and an on-failure destination, so details of the failed batch land somewhere we can replay from. Partial batch responses work here too. The second cause is a hot shard, one partition key carrying too much traffic. The parallelisation factor lets several batches from one shard run at once while keeping order per key, but the lasting fix is a better key.”

Red flag to avoid:

Saying Lambda skips a failing record and moves on by default.

They may ask next:
  • Why does the on-failure destination get details of the batch rather than the records, and how would you replay them?
  • When would you give a consumer enhanced fan-out?
Say it in 60 seconds
Hard Technical round Senior, Mid-level Practice question

15. An asynchronously invoked Lambda, fed by EventBridge or S3, threw errors for an hour after a bad deploy. Which events were retried, which may be lost, and how do you stop events vanishing silently?

What the interviewer is really testing:
Whether you know how Lambda's asynchronous queue treats errors, throttles and event age, and how to keep what it gives up on.
Answer frame:

Internal queue: async events wait in a queue Lambda manages; function errors are retried twice by default.

Throttles differ: throttled events are retried for longer, up to the maximum event age, which can be as long as six hours.

Then gone: when retries or age run out, the event is dropped unless an on-failure destination or dead-letter queue is set.

Design: a destination with an alarm, idempotent handlers, and a way to replay.

Sample spoken answer:

“Asynchronous invokes go into a queue that Lambda runs for you, and the caller gets a success as soon as the event is queued. If the function throws, Lambda retries twice by default, with a wait between attempts. If the function was throttled instead, Lambda keeps retrying for longer, up to the maximum event age, which can be as long as six hours. Once retries or age run out, the event is dropped, and nothing tells you unless you set it up. So after an hour of errors, most of those events used up their retries and are gone. My design: an on-failure destination, usually an SQS queue, with an alarm on its depth, so each failure is kept with the original event and the error. Handlers are idempotent, because an async event can occasionally run more than once. And we keep a small tool that replays events from that queue after a fix. On EventBridge rules, I also set a dead-letter queue for events it couldn't deliver at all.”

Red flag to avoid:

Assuming Lambda keeps retrying failed async events until the code is fixed.

They may ask next:
  • When would you choose an on-failure destination over a dead-letter queue on the function?
  • Why might you lower the maximum event age on purpose?
Say it in 60 seconds

Networking and Compute 3 questions

Hard System design round Senior Practice question

16. With a few hundred accounts, how do you decide between Transit Gateway, shared VPCs and PrivateLink for connecting workloads?

What the interviewer is really testing:
Whether you know what each connectivity model is for, what it costs in isolation and operations, and can combine them rather than force one.
Answer frame:

Transit Gateway: a regional hub routing between many VPCs and on-premises; route tables segment it, but it needs ranges that don't overlap.

Shared VPCs: a network account owns the VPC and shares subnets through Resource Access Manager; teams launch into them but can't change the network.

PrivateLink: one-way access to a single service through an endpoint, and it works across overlapping ranges.

Combine: a hub for broad routing, PrivateLink for team-to-team services, and central IP planning.

Sample spoken answer:

“They solve different problems, so at that scale I use all three. Transit Gateway is a routing hub. VPCs attach to it, separate route tables keep production, non-production and shared services apart, and it's where on-premises links land. It gives broad network reach, which also means a broad blast radius, and it needs address ranges that don't overlap. Shared VPCs, through Resource Access Manager, let a central network account own the VPC and subnets while teams launch resources into them from their own accounts. That cuts the number of VPCs and keeps network changes with the network team, but teams share one network boundary. PrivateLink is the narrowest: a consumer gets an endpoint for one service behind a load balancer, connections only start from the consumer side, and it works across overlapping ranges. For team-to-team APIs I prefer it, because it exposes a service, not a network. And I'd put address planning in one place, with IPAM, so ranges stop colliding.”

Red flag to avoid:

Forcing one model onto everything, or not knowing overlapping ranges can't be routed through a hub.

They may ask next:
  • How would you inspect traffic between VPCs centrally without making the hub a bottleneck?
  • How would you move from a peering mesh to this design without downtime?
Say it in 60 seconds
Medium Technical round Senior, Mid-level Practice question

17. Your worker fleet runs jobs that take up to twenty minutes, and scale-in keeps killing instances mid-job. How do you fix it?

What the interviewer is really testing:
Whether you know the Auto Scaling hooks that let work finish, and whether you'd scale on the right signal.
Answer frame:

Lifecycle hook: a terminating hook holds the instance in a wait state so it can finish its current job.

Protection: the worker can turn on instance scale-in protection while busy and turn it off when idle.

Right metric: scale on queue backlog per instance, not CPU.

Job design: make jobs resumable or idempotent, since instances can still vanish.

Sample spoken answer:

“Auto Scaling doesn't know an instance is halfway through a job, so on scale-in it just picks one and ends it. There are two tools. A terminating lifecycle hook puts the chosen instance into a wait state; the worker stops taking new work, finishes the current job, then completes the lifecycle action so termination goes ahead. The other option is scale-in protection set by the worker itself: it protects itself when it picks up a job and removes protection when it's idle, so only idle instances are chosen. I'd also fix the scaling signal. CPU is a poor fit for a queue of jobs. Backlog per instance, meaning messages waiting divided by running workers, tracks the real need. Even with all that, instances can still disappear, for example on a Spot interruption, so the job itself should be safe to retry or able to resume from a checkpoint.”

Red flag to avoid:

Turning off scale-in altogether and paying for idle machines.

They may ask next:
  • How would Spot interruptions change this design?
  • What happens if the lifecycle hook's timeout is shorter than the longest job?
Say it in 60 seconds
Hard Technical round Senior, Mid-level Practice question

18. On EKS, every pod on a node can end up using the node's IAM role. Why is that a problem, and how do you give each workload only its own permissions?

What the interviewer is really testing:
Whether you know how pods pick up AWS credentials on EKS and the standard ways to scope them to one workload.
Answer frame:

The risk: pods can reach the instance metadata service and use the node role, so every pod gets what any pod needs.

Per workload: IAM roles for service accounts, through the cluster's OIDC provider, or EKS Pod Identity map a service account to an IAM role.

Trust policy: scope each role to one namespace and service account.

Close the fallback: IMDSv2 with a hop limit of one on nodes, and a minimal node role.

Sample spoken answer:

“By default a pod can call the instance metadata service just like the node, and get the node role's credentials. So if one app needs to write to a bucket and that permission goes on the node role, every other pod on that node can write there too. The fix is per-workload identity. With IAM roles for service accounts, the cluster has an OIDC provider, a Kubernetes service account is annotated with a role, and the SDK in the pod swaps its projected token for that role's credentials. EKS Pod Identity does the same job with an agent on each node and associations managed through the EKS API, which is simpler across many clusters. Either way, the role's trust policy names one namespace and service account, not anything in the cluster. Then I close the fallback: IMDSv2 required on nodes with a hop limit of one, so pods on the pod network can't reach the node's metadata, and a node role cut down to what the node itself needs.”

Red flag to avoid:

Putting every permission on the node role, or saying Kubernetes RBAC controls access to AWS.

They may ask next:
  • How would you audit which pods can reach which AWS resources today?
  • What changes for pods that use host networking?
Say it in 60 seconds

Platform Standards 2 questions

Hard System design round Senior Practice question

19. You're setting infrastructure-as-code standards for thirty teams. What rules do you put in place, and how do you enforce them without becoming a bottleneck?

What the interviewer is really testing:
Whether you can turn good habits into paved roads and automatic guardrails rather than manual review.
Answer frame:

Paved road: shared, versioned modules for common patterns, so the secure way is also the easy way.

Stack boundaries: split stacks by team and by how often things change, so a routine app deploy can't touch the network or the data stores.

Automatic checks: policy-as-code in the pipeline, SCPs as the hard floor, and drift detection on a schedule.

Ownership: teams own their stacks; the platform team owns the modules and the rules, and reviews only exceptions.

Sample spoken answer:

“I'd aim for rules that are enforced by the pipeline, so the platform team isn't reviewing every pull request. First, a paved road: versioned modules for the common shapes, like a service behind a load balancer or a queue with a dead-letter queue, with logging, encryption and tags built in. Teams using them get through faster, which is the real incentive. Second, stack boundaries. Networking, data stores and the app live in separate stacks, so a routine app deploy can't touch the database, and stateful resources keep a retain policy. Third, policy-as-code checks in the pipeline for things like public buckets and open security groups, with SCPs underneath as the floor nobody can get around. Fourth, nothing changes by hand in production, and drift detection runs on a schedule. Teams own their own stacks and on-call. My team reviews only exceptions, and when one keeps coming up, we turn it into a module.”

Red flag to avoid:

Answering with 'every change goes through my team for review'.

They may ask next:
  • How do you roll out a breaking change to a shared module used by thirty teams?
  • What would you do with a team that keeps making console changes in production?
Say it in 60 seconds
Hard System design round Senior Practice question

20. You're designing the AWS account structure for a company with thirty product teams. How do you lay out accounts and organizational units, and what does every new account get on day one?

What the interviewer is really testing:
Whether you use accounts as the main isolation boundary and can design the organization, the central accounts and the baseline so new accounts are safe by default.
Answer frame:

Accounts as boundaries: separate accounts per team and environment, so blast radius, quotas and costs stay apart.

OUs by policy: group accounts by the guardrails they need: security, infrastructure, production, non-production, sandbox.

Central accounts: a log archive for organization-wide CloudTrail and Config, a security tooling account as delegated admin, and a network account.

Vending: new accounts come from a pipeline with the baseline, single sign-on and a budget alert already in place.

Sample spoken answer:

“I treat the account as the strongest boundary AWS gives me, so each team gets separate production and non-production accounts, and developers get sandboxes. Organizational units follow policy, not the org chart: security, infrastructure, workloads split into production and non-production, a sandbox OU with looser rules, and a suspended OU for accounts being closed. Guardrails attach at the OU, so a new production account gets production rules automatically. A few central accounts do shared jobs: a log archive that receives an organization-wide CloudTrail and Config data, which almost nobody can touch, a security tooling account set as delegated admin for services like GuardDuty and Security Hub, and a network account that owns the hub and shared subnets. The management account runs nothing else. Accounts come from a vending pipeline, so each arrives with the baseline, single sign-on groups mapped to roles, and a budget alert. People sign in through single sign-on, never as IAM users.”

Red flag to avoid:

Putting every team in one account and separating them with IAM policies and tags alone.

They may ask next:
  • How would you move existing workloads out of one big shared account into this structure?
  • Why group accounts by policy rather than by the org chart?
Say it in 60 seconds

Leadership 4 questions

Hard Behavioral round Senior Practice question

21. Tell me about the largest migration to AWS you've led. How did you plan the waves, move the data and handle the cutover?

What the interviewer is really testing:
Whether you've led a migration end to end, including the unglamorous parts: discovery, data sync, cutover and rollback.
Answer frame:

Discovery: map apps and dependencies, then decide per app: rehost, replatform, refactor, retire or retain.

Waves: start with low-risk apps to build the method, then group tightly coupled apps together.

Data: full load plus ongoing change capture, so the cutover window only covers the last changes.

Cutover: a rehearsed runbook, a go or no-go check, and a tested rollback.

Sample spoken answer:

“At my last company I led moving about sixty applications and their databases out of a data centre whose lease was ending. We spent the first month on discovery, because the dependency map everyone believed in was wrong in several places; network flow data showed apps talking that nobody had listed. We sorted each app: most were rehosted as they were, a few databases moved to managed RDS, and around ten apps were retired outright. We ran it in waves, starting with internal tools so we could fix our runbook on low stakes. For databases we did a full load and then ongoing change capture, so the final cutover was only minutes of catch-up. Each cutover had a rehearsal, a go or no-go meeting and a rollback we'd actually tested. One wave did roll back because of a licensing check tied to hardware. We finished two weeks before the lease ended.”

Red flag to avoid:

Describing a migration as just copying servers across, with no rollback plan.

They may ask next:
  • What would you do differently if you ran that migration again?
  • How did you keep the business informed when a wave slipped?
Say it in 60 seconds
Medium Behavioral round Senior Practice question

22. Tell me about a cloud project you led that failed or had to be reversed. What went wrong and what did you change afterwards?

What the interviewer is really testing:
Whether you own a real failure, understand its root cause, and changed how the team works.
Answer frame:

The project: what you set out to do and why it seemed right.

What went wrong: the real cause, including your part in it.

The call: how and when you decided to stop or reverse.

After: the lasting change in how you plan or decide.

Sample spoken answer:

“I led a move of our main API from containers to Lambda, mainly to cut the time we spent on patching and scaling. On paper it fit. In practice, a big share of our traffic was long-running report requests and some endpoints kept warm in-memory caches, and I hadn't looked closely enough at the traffic mix. Latency got worse, the database got hammered by connections, and after two months we'd moved only a third of the endpoints. I called a stop, which was hard because I'd championed it. We kept the event-driven parts on Lambda, where they worked well, and moved the rest back. Afterwards I changed how we start projects like this: a one-page decision record, a small proof on the nastiest endpoint rather than the easiest, and exit criteria written down at the start, so stopping is a planned outcome, not a failure.”

Red flag to avoid:

Blaming the tool or another team, or choosing a failure that was really a success.

They may ask next:
  • How did you tell leadership you were reversing course?
  • How did the team take it, and what did you do about that?
Say it in 60 seconds
Medium Situational round Senior Practice question

23. You're leading the response to an outage that hit several services at once. Teams are arguing about the root cause, and a VP on the call keeps asking for updates. What do you do?

What the interviewer is really testing:
Whether you can run a major incident: clear roles, mitigation before diagnosis, and steady updates that keep leadership informed without letting them drive the call.
Answer frame:

Roles: name an incident lead, a communications owner and one lead per affected service.

Mitigate first: roll back recent changes, shift traffic or shed load before settling the root cause.

Updates on a clock: the communications owner updates leadership at a set interval, in a separate channel.

After: a blameless review with a timeline and actions that each have an owner.

Sample spoken answer:

“First I make the roles explicit, because arguing usually means nobody's in charge. I take incident lead, ask one person to own communications, and name one lead per affected service. Then I move the goal from 'why' to 'stop the bleeding': what changed recently, can we roll it back, can we move traffic away from a bad zone or turn off a non-essential feature to shed load? The root cause debate goes to a side thread while we mitigate. I also check the AWS Health Dashboard early, so we don't spend an hour debugging our code for a problem on the provider's side. For the VP, I'd say politely that the communications owner will post updates every fifteen minutes in the leadership channel, and we stick to that clock, even when the update is 'no change, trying this next'. Most leaders relax once updates are predictable. Afterwards we run a blameless review with a clear timeline and actions that each have an owner and a date.”

Red flag to avoid:

Diving into debugging yourself and leaving nobody running the call or talking to leadership.

They may ask next:
  • When would you call the incident over while the root cause is still unknown?
  • How do you make sure the actions from the review actually get done?
Say it in 60 seconds
Medium Culture fit round Senior, Mid-level Practice question

24. How do you hire and grow cloud engineers so the team can move fast without you approving every change?

What the interviewer is really testing:
Whether you build a team that scales past you, through hiring for judgement, clear guardrails and deliberate growth.
Answer frame:

Hiring: test judgement with a real design discussion and a debugging exercise, not trivia.

Guardrails: safe defaults and automatic checks so mistakes are small and caught early.

Growth: give people ownership of real systems, pair on incidents, and let them lead reviews.

Letting go: step back from approvals once the guardrails are in place.

Sample spoken answer:

“For hiring, I care more about judgement than memory. In interviews I give a small design problem with a twist, like a database that can't go down during a migration, and a broken setup to debug together, and I listen for how people reason and when they ask questions. Service trivia tells me very little. For speed, the trick is making mistakes cheap. We have sandbox accounts, guardrails from SCPs and pipeline checks, and nothing changes in production by hand, so a new engineer can ship in their first week without scaring anyone. For growth, I hand people ownership of a real system early, pair them with a senior on their first incidents, and rotate who leads design reviews. My own goal is to become unnecessary for approvals. When I find myself the bottleneck on something, that's a sign we're missing a guardrail or I haven't trusted someone enough yet.”

Red flag to avoid:

Saying you personally review every change to keep quality high.

They may ask next:
  • How do you give feedback to a strong engineer who ignores reviews?
  • What signal in an interview makes you say no even with a strong technical round?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card