AWS interviews for 10+ years of experience skip basics like what EC2 is and go after what experience teaches, such as how guardrails work across many accounts, where encryption and scaling quietly break, why a health check can take down a fleet, how to keep one failure from reaching every customer, and how you lead migrations, set standards and grow a team. This page is written for cloud engineers and architects with around eight to fifteen years of experience, interviewing for senior, staff, principal or lead roles. Each question shows what the interviewer is checking, a shape for your answer and a sample to adapt. Say it in your own words, with your own stories.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Ceiling, not grant: an SCP limits what any principal in a member account can do; access still needs an IAM allow.
Who it hits: every user and role in member accounts, including the account's root user.
Who it misses: the management account and service-linked roles are not restricted by SCPs.
Design: keep workloads out of the management account and test SCPs on a sandbox OU first.
“I think of an SCP as the fence around an account, not a key to any door. For a request to succeed, the SCPs on every level from the root down to the account have to allow it, and then IAM has to allow it too. So an SCP that allows everything grants nothing on its own. The catch people miss is scope. SCPs do restrict the root user of a member account, which is why they're the right place for things like 'nobody leaves the organization' or 'no resources outside our approved regions'. But they don't apply to the management account at all, and they don't restrict service-linked roles. That's why I keep nothing running in the management account and treat access to it as the most sensitive thing we have. I also roll any new SCP out to a sandbox OU first, because a bad deny can break every pipeline in the org at once.”
Saying an SCP grants permissions, or assuming it also protects the management account.
The risk: the vendor's account can assume roles for many customers, so a customer could trick it into acting on someone else's account.
The fix: require a unique external ID, issued by the vendor per customer, in the trust policy condition.
For AWS services: use aws:SourceArn or aws:SourceAccount conditions for the same reason.
Plus: least privilege on the role itself, and a narrow principal instead of the whole account where possible.
“The deputy here is the vendor's service. It holds permission to assume roles in lots of customer accounts. If the only check is 'trust the vendor's account', then customer A could type customer B's role ARN into the vendor's settings page, and the vendor would happily assume it, because from AWS's side the vendor is allowed. The fix is an external ID. The vendor generates a unique value per customer and always passes it when assuming the role, and the customer's trust policy only allows the assume call when that exact value is present. Customer A can't supply B's external ID because the vendor controls it. The same idea applies when an AWS service assumes a role for you: I add a source ARN or source account condition so the service can only act for my own resources. And the role itself still only gets the read actions the tool needs.”
{
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::111122223333:root" },
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": { "sts:ExternalId": "customer-7f3a9c" }
}
}
Treating the external ID as a secret password, or thinking an account-level trust alone is safe.
Boundary: a managed policy that caps what any role the team creates can ever do.
Enforce it: allow iam:CreateRole only when the request attaches that boundary.
Close the gaps: deny editing or deleting the boundary policy and removing a boundary from a role.
Naming: restrict teams to roles under their own path or prefix so they can't touch others.
“The risk with letting teams create roles is simple escalation: create a role with admin rights, then assume it. Permission boundaries solve that. A boundary is a managed policy that sets the maximum a role can do; the role's effective permissions are the overlap between its own policies and the boundary. So I write one boundary per team, covering their services and nothing in IAM or billing. Then the team's deploy role may create roles only if the create call attaches that exact boundary, which I enforce with a condition on the permissions boundary key. The part people forget is the escape hatches: I also deny changing the boundary policy itself, deleting it, and removing the boundary from a role. And I scope their role names to a path, so they can't edit the platform's roles. With that, teams ship on their own and the worst they can build is a role no stronger than the boundary.”
{
"Effect": "Allow",
"Action": ["iam:CreateRole", "iam:PutRolePermissionsBoundary"],
"Resource": "arn:aws:iam::111122223333:role/team-orders/*",
"Condition": {
"StringEquals": {
"iam:PermissionsBoundary": "arn:aws:iam::111122223333:policy/team-orders-boundary"
}
}
}
Letting teams create roles with only a code review as protection, or forgetting to deny removing the boundary.
The attack: a web app tricked into fetching a URL can fetch the metadata endpoint and return the instance role's credentials.
IMDSv2: needs a token from a PUT request with a TTL header first, which a simple forged GET can't do.
Hop limit: the token response has a hop limit; raise it only when containers truly need metadata.
Enforce: require tokens in launch templates, and block launches without it using an SCP condition, with Config to find old instances.
“The classic problem is server-side request forgery. If an app will fetch any URL a user gives it, an attacker points it at the metadata address and gets back the temporary credentials of the instance's role. IMDSv1 answered a plain GET, so that worked. IMDSv2 is session based: you first send a PUT with a TTL header to get a token, then send that token as a header on every read. Most forgery bugs can only make a GET with no custom headers, so they fail. There's also a hop limit on the token response, which stops it reaching further than it should, for example into containers that shouldn't have the host's role. To enforce it at scale, I set tokens to required in every launch template and AMI pipeline, add an SCP that denies launching instances unless the metadata tokens condition says required, and use a Config rule to list the old instances still allowing v1 so teams can fix them.”
Saying IMDSv2 is just a newer API version with no security difference.
Envelope encryption: S3 gets a data key from KMS to encrypt the object; the KMS key never leaves KMS.
Per request: writes ask KMS for a data key and reads ask KMS to decrypt it, so request volume maps to KMS volume.
Fix: turn on S3 Bucket Keys, which use a short-lived bucket-level key and cut calls to KMS sharply.
Also: KMS quotas are per account and region, so check who else shares them before asking for an increase.
“SSE-KMS uses envelope encryption. When I put an object, S3 asks KMS for a data key, encrypts the object with it, and stores the data key encrypted next to the object. When I read it, S3 sends that encrypted data key to KMS to decrypt. So at high request rates, every S3 call is also a KMS call, and KMS has request-rate quotas per account and region. That's where the throttling comes from. The first fix is S3 Bucket Keys: S3 gets a bucket-level key from KMS and uses it to make data keys locally for a while, so the KMS call count drops massively. It applies to new objects, so old ones still hit KMS until they're rewritten. Then I check what else in the account shares the quota, because one noisy job can starve the rest, and only then ask for a quota increase. I'd also make sure clients retry throttles with backoff and jitter.”
Suggesting you switch off encryption, or not knowing that each read also calls KMS.
Key policy first: every KMS key has a key policy, and IAM policies only count if the key policy lets the account use them.
The default: the default key policy trusts the account, which switches IAM on for that key.
The trap: a custom key policy that names only certain roles makes IAM grants useless for everyone else.
Lockout: KMS refuses a key policy that would stop the caller changing it again, but that check can be bypassed and doesn't protect other admins.
“KMS is different from most services because the key policy is the main gate. IAM permissions only count for a key if its key policy lets the account delegate access through IAM. The default key policy does that by trusting the account, so people get used to IAM just working. But if someone wrote a custom key policy naming only one application role, an admin with full KMS rights in IAM still gets denied. So I check the key policy first, not the user. The bigger risk is lockout. KMS has a safety check that rejects a policy which would stop the person setting it from changing it again, but it can be bypassed, and it doesn't protect anyone else. If every admin is locked out, only AWS support can help, and that's slow. So key policies live in code, get reviewed like production changes, and always keep a break-glass admin role.”
Looking only at the user's IAM policy and never at the key policy.
Control vs data plane: the control plane creates and changes resources; the data plane serves the running ones.
The trap: plans that scale up, launch instances or change DNS records during the incident need the control plane at its worst moment.
Static stability: pre-provision enough so losing one AZ needs no action, for example each of three AZs sized for half the peak.
Failover: use mechanisms that act in the data plane, like health-checked DNS records, and rehearse them.
“Every AWS service has a control plane, which creates, changes and deletes things, and a data plane, which serves what's already running. Data planes are simpler and built to be more available. During a big incident the control plane is often what struggles, because everyone is trying to launch capacity at once. So if my recovery plan is 'when an AZ dies, Auto Scaling launches more instances elsewhere', I'm betting on the weakest part at the worst time. Static stability means the system keeps working through a failure without having to change anything. For AZ loss, I run in three AZs and give each enough capacity for half the peak, so losing one still leaves the full peak covered. That costs more, and I'd show leadership exactly what it buys. For failover between regions I prefer health-checked DNS records that flip on their own, and we rehearse it on a normal weekday.”
Answering that Auto Scaling will just replace the lost capacity automatically.
Amplification: retries stack across layers; three layers each trying three times can mean twenty-seven calls at the bottom for one request.
Retry in one place: one layer retries, with capped exponential backoff, jitter and a retry budget.
Stop sending: circuit breakers and load shedding give a struggling service room to recover.
Timeouts that fit: each timeout shorter than its caller's, so nobody works on requests already given up on.
“The danger is that retries multiply. If the edge, the API and the service layer each try three times, one click can become twenty-seven calls on the database, right when it's slowest. The retries also arrive in waves, so when the service recovers, the backlog knocks it over again. That's how two minutes becomes twenty. My defences: retry at one layer only, usually the one that knows whether the work is safe to repeat, with capped exponential backoff and jitter so clients spread out. I add a retry budget, so retries can never be more than a small share of normal traffic. Circuit breakers stop calls to a dependency that's clearly failing, and the service sheds load it can't handle instead of queueing forever. Timeouts get shorter as you go down the stack. And since the AWS SDKs already retry with backoff, I check we aren't wrapping their retries in our own.”
Adding more retries or longer timeouts as the fix for a flaky dependency.
Cells: independent full copies of the service, each with its own compute and data, serving a fixed set of customers.
Thin router: a very simple layer maps each customer to a cell and does little else, so it rarely breaks.
Deploy by cell: changes reach one small cell first and widen only if its alarms stay quiet.
Noisy neighbours: per-customer limits, and a cell of its own for a customer that outgrows a shared one.
“Right now everyone shares one fleet and one database, so any failure is everyone's failure. I'd move to cells. A cell is a complete, independent copy of the service, with its own compute, queues and data, often in its own account, and sized to a known maximum. A thin routing layer looks up which cell a customer belongs to and forwards the request. It does almost nothing else, so it can be the most stable part of the system. Now a poison request, a runaway customer or a bad config hurts one cell, not the whole product. Deploys go cell by cell, starting with a small one, with automatic rollback if its alarms fire. Big customers get per-customer limits, or a cell of their own. The costs are real: more copies to run, cross-cell features like global search get harder, and moving a customer between cells needs a data migration plan. I'd present it to leadership as how many customers any one incident can touch.”
Answering only with more instances or a bigger database, which keeps every customer inside the same failure.
The chain: every target fails its check at once because they share the database.
ALB behaviour: when all targets are unhealthy, the ALB fails open and sends traffic to all of them anyway.
The real damage: an ASG using ELB health checks can start replacing every instance, so the fleet restarts cold.
Better design: shallow liveness for the load balancer, dependency checks in metrics and alarms.
“This is a nice trap. If the health check calls the database, a database blip makes every instance fail at the same moment, because they all share it. The load balancer itself is fairly forgiving here: when every target in every enabled zone is unhealthy, an ALB fails open and keeps sending traffic to all of them, since sending nowhere helps no one. The bigger danger is the Auto Scaling group. If it uses ELB health checks, it sees the whole fleet unhealthy and starts terminating and replacing instances. Now a thirty-second database blip turns into many minutes of cold starts, warming caches and a thundering herd on the database when it recovers. So I keep the load balancer check shallow, meaning the process is up and can serve, and I watch dependencies through metrics and alarms. If a deep check is needed, I make it degrade softly instead of failing the whole node.”
Saying a deep health check is always safer because it catches more problems.
Why it failed: a GSI never enforces uniqueness and is updated asynchronously, so check-then-write races.
Guard item: write a second item whose key is the email, next to the user item.
Atomic: both puts in one transaction, each with an attribute_not_exists condition, so both succeed or neither does.
Upkeep: transactional writes use twice the write capacity, and an email change swaps the guard item in one transaction.
“DynamoDB only guarantees uniqueness on the primary key. A GSI on email will happily hold two items with the same email, and it's updated asynchronously, so 'query the index, then write' has a race: two requests both see no match and both write. The pattern I use is a guard item. Next to the user item, I write a second item in the same table whose key is EMAIL plus the address. Both puts go in one TransactWriteItems call, each with a condition that the key doesn't already exist. If the email is taken, the whole transaction is cancelled, and I return 'already registered'. The costs: transactional writes use twice the write capacity of normal ones, and changing an email means deleting the old guard and adding the new one in the same transaction. I also normalise the email first, or two spellings of one address slip through.”
import boto3
from botocore.exceptions import ClientError
ddb = boto3.client("dynamodb")
def create_user(user_id, email):
email = email.strip().lower()
guard = "attribute_not_exists(pk)"
try:
ddb.transact_write_items(TransactItems=[
{"Put": {"TableName": "users", "ConditionExpression": guard,
"Item": {"pk": {"S": f"USER#{user_id}"}, "email": {"S": email}}}},
{"Put": {"TableName": "users", "ConditionExpression": guard,
"Item": {"pk": {"S": f"EMAIL#{email}"}, "user_id": {"S": user_id}}}},
])
except ClientError as e:
reasons = e.response.get("CancellationReasons", [])
if any(r.get("Code") == "ConditionalCheckFailed" for r in reasons):
raise ValueError("user or email already registered")
raise
Querying the index first and then writing, or saying a GSI can be marked unique.
Only new writes: a rule copies objects written after it exists; older objects need S3 Batch Replication.
Deletes: delete markers are copied only if you turn that on, and deleting a specific version is never replicated.
Prerequisites: versioning on both buckets; KMS-encrypted objects replicate only when the rule opts in and can use a key in the destination.
Proof: check replication status and metrics, and read from the replica in a drill.
“The first surprise is that a new replication rule only covers objects written after it's created. Everything already in the bucket stays put unless I run S3 Batch Replication. Second, deletes. Delete markers are only replicated if I turn that on, and deleting a specific version is never replicated. That's useful protection if someone wipes the source, but it means the replica isn't a mirror. Third, prerequisites: both buckets need versioning, and objects encrypted with KMS keys are skipped unless the rule opts in and the replication role can use a key in the destination region. Before I'd trust it, I'd check the replication status on a sample of objects, watch the replication metrics for failures and lag, and in a drill point a real reader at the replica. Bucket policies, lifecycle rules and access points don't come along, and the app needs those there too.”
Assuming replication makes the second bucket a complete, exact mirror from day one.
Storage: one shared cluster volume keeping six copies across three AZs; the write quorum is four of six, the read quorum three of six.
Replicas: readers attach to the same volume, so they don't replay changes into their own copy and usually lag far less than classic read replicas.
Failover: Aurora promotes a reader, chosen by the promotion tiers you set, and moves the cluster endpoint to it.
Clones: fast clones share pages copy-on-write, so a full-size copy is quick and only changed pages take new space.
“Standard RDS is an engine on an instance with its own disk, and Multi-AZ adds a standby with its own copy. Aurora separates compute from storage. One cluster volume keeps six copies of the data across three AZs, and a write is durable once four of the six acknowledge it. So it can lose an AZ and keep writing, and it repairs copies in the background. Readers attach to that same volume instead of replaying changes into their own storage, which is why their lag is usually far lower than a classic read replica's. It isn't zero, so read-after-write still goes to the writer. On failover, Aurora promotes a reader based on the promotion tiers I set, and moves the cluster endpoint to it, so I keep a reader sized like the writer in the top tier. And clones use copy-on-write on the shared volume, so a team can get a full-size copy of production data quickly, and only pages that change take new space.”
Describing Aurora as just a faster MySQL with no difference in storage or failover.
Order blocks: Lambda reads each shard in order, so a batch that keeps failing is retried while nothing behind it moves.
Or a hot shard: one partition key with too much traffic overloads a single shard.
Failure settings: bisect the batch on error, cap retry attempts and record age, and add an on-failure destination.
Throughput: raise the parallelisation factor, and fix the partition key for good.
“Iterator age is how far behind the newest record the consumer is, so one shard falling behind points at that shard. There are two usual causes. The first is a record the function can't process. With streams, Lambda keeps order per shard, so when a batch fails it retries that same batch, and by default it keeps going until the records expire from the stream. Everything behind it waits. The fixes live on the event source mapping: bisect batch on error, so Lambda splits the batch to isolate the bad record, a maximum retry count and maximum record age, so it gives up, and an on-failure destination, so details of the failed batch land somewhere we can replay from. Partial batch responses work here too. The second cause is a hot shard, one partition key carrying too much traffic. The parallelisation factor lets several batches from one shard run at once while keeping order per key, but the lasting fix is a better key.”
Saying Lambda skips a failing record and moves on by default.
Internal queue: async events wait in a queue Lambda manages; function errors are retried twice by default.
Throttles differ: throttled events are retried for longer, up to the maximum event age, which can be as long as six hours.
Then gone: when retries or age run out, the event is dropped unless an on-failure destination or dead-letter queue is set.
Design: a destination with an alarm, idempotent handlers, and a way to replay.
“Asynchronous invokes go into a queue that Lambda runs for you, and the caller gets a success as soon as the event is queued. If the function throws, Lambda retries twice by default, with a wait between attempts. If the function was throttled instead, Lambda keeps retrying for longer, up to the maximum event age, which can be as long as six hours. Once retries or age run out, the event is dropped, and nothing tells you unless you set it up. So after an hour of errors, most of those events used up their retries and are gone. My design: an on-failure destination, usually an SQS queue, with an alarm on its depth, so each failure is kept with the original event and the error. Handlers are idempotent, because an async event can occasionally run more than once. And we keep a small tool that replays events from that queue after a fix. On EventBridge rules, I also set a dead-letter queue for events it couldn't deliver at all.”
Assuming Lambda keeps retrying failed async events until the code is fixed.
Transit Gateway: a regional hub routing between many VPCs and on-premises; route tables segment it, but it needs ranges that don't overlap.
Shared VPCs: a network account owns the VPC and shares subnets through Resource Access Manager; teams launch into them but can't change the network.
PrivateLink: one-way access to a single service through an endpoint, and it works across overlapping ranges.
Combine: a hub for broad routing, PrivateLink for team-to-team services, and central IP planning.
“They solve different problems, so at that scale I use all three. Transit Gateway is a routing hub. VPCs attach to it, separate route tables keep production, non-production and shared services apart, and it's where on-premises links land. It gives broad network reach, which also means a broad blast radius, and it needs address ranges that don't overlap. Shared VPCs, through Resource Access Manager, let a central network account own the VPC and subnets while teams launch resources into them from their own accounts. That cuts the number of VPCs and keeps network changes with the network team, but teams share one network boundary. PrivateLink is the narrowest: a consumer gets an endpoint for one service behind a load balancer, connections only start from the consumer side, and it works across overlapping ranges. For team-to-team APIs I prefer it, because it exposes a service, not a network. And I'd put address planning in one place, with IPAM, so ranges stop colliding.”
Forcing one model onto everything, or not knowing overlapping ranges can't be routed through a hub.
Lifecycle hook: a terminating hook holds the instance in a wait state so it can finish its current job.
Protection: the worker can turn on instance scale-in protection while busy and turn it off when idle.
Right metric: scale on queue backlog per instance, not CPU.
Job design: make jobs resumable or idempotent, since instances can still vanish.
“Auto Scaling doesn't know an instance is halfway through a job, so on scale-in it just picks one and ends it. There are two tools. A terminating lifecycle hook puts the chosen instance into a wait state; the worker stops taking new work, finishes the current job, then completes the lifecycle action so termination goes ahead. The other option is scale-in protection set by the worker itself: it protects itself when it picks up a job and removes protection when it's idle, so only idle instances are chosen. I'd also fix the scaling signal. CPU is a poor fit for a queue of jobs. Backlog per instance, meaning messages waiting divided by running workers, tracks the real need. Even with all that, instances can still disappear, for example on a Spot interruption, so the job itself should be safe to retry or able to resume from a checkpoint.”
Turning off scale-in altogether and paying for idle machines.
The risk: pods can reach the instance metadata service and use the node role, so every pod gets what any pod needs.
Per workload: IAM roles for service accounts, through the cluster's OIDC provider, or EKS Pod Identity map a service account to an IAM role.
Trust policy: scope each role to one namespace and service account.
Close the fallback: IMDSv2 with a hop limit of one on nodes, and a minimal node role.
“By default a pod can call the instance metadata service just like the node, and get the node role's credentials. So if one app needs to write to a bucket and that permission goes on the node role, every other pod on that node can write there too. The fix is per-workload identity. With IAM roles for service accounts, the cluster has an OIDC provider, a Kubernetes service account is annotated with a role, and the SDK in the pod swaps its projected token for that role's credentials. EKS Pod Identity does the same job with an agent on each node and associations managed through the EKS API, which is simpler across many clusters. Either way, the role's trust policy names one namespace and service account, not anything in the cluster. Then I close the fallback: IMDSv2 required on nodes with a hop limit of one, so pods on the pod network can't reach the node's metadata, and a node role cut down to what the node itself needs.”
Putting every permission on the node role, or saying Kubernetes RBAC controls access to AWS.
Paved road: shared, versioned modules for common patterns, so the secure way is also the easy way.
Stack boundaries: split stacks by team and by how often things change, so a routine app deploy can't touch the network or the data stores.
Automatic checks: policy-as-code in the pipeline, SCPs as the hard floor, and drift detection on a schedule.
Ownership: teams own their stacks; the platform team owns the modules and the rules, and reviews only exceptions.
“I'd aim for rules that are enforced by the pipeline, so the platform team isn't reviewing every pull request. First, a paved road: versioned modules for the common shapes, like a service behind a load balancer or a queue with a dead-letter queue, with logging, encryption and tags built in. Teams using them get through faster, which is the real incentive. Second, stack boundaries. Networking, data stores and the app live in separate stacks, so a routine app deploy can't touch the database, and stateful resources keep a retain policy. Third, policy-as-code checks in the pipeline for things like public buckets and open security groups, with SCPs underneath as the floor nobody can get around. Fourth, nothing changes by hand in production, and drift detection runs on a schedule. Teams own their own stacks and on-call. My team reviews only exceptions, and when one keeps coming up, we turn it into a module.”
Answering with 'every change goes through my team for review'.
Accounts as boundaries: separate accounts per team and environment, so blast radius, quotas and costs stay apart.
OUs by policy: group accounts by the guardrails they need: security, infrastructure, production, non-production, sandbox.
Central accounts: a log archive for organization-wide CloudTrail and Config, a security tooling account as delegated admin, and a network account.
Vending: new accounts come from a pipeline with the baseline, single sign-on and a budget alert already in place.
“I treat the account as the strongest boundary AWS gives me, so each team gets separate production and non-production accounts, and developers get sandboxes. Organizational units follow policy, not the org chart: security, infrastructure, workloads split into production and non-production, a sandbox OU with looser rules, and a suspended OU for accounts being closed. Guardrails attach at the OU, so a new production account gets production rules automatically. A few central accounts do shared jobs: a log archive that receives an organization-wide CloudTrail and Config data, which almost nobody can touch, a security tooling account set as delegated admin for services like GuardDuty and Security Hub, and a network account that owns the hub and shared subnets. The management account runs nothing else. Accounts come from a vending pipeline, so each arrives with the baseline, single sign-on groups mapped to roles, and a budget alert. People sign in through single sign-on, never as IAM users.”
Putting every team in one account and separating them with IAM policies and tags alone.
Discovery: map apps and dependencies, then decide per app: rehost, replatform, refactor, retire or retain.
Waves: start with low-risk apps to build the method, then group tightly coupled apps together.
Data: full load plus ongoing change capture, so the cutover window only covers the last changes.
Cutover: a rehearsed runbook, a go or no-go check, and a tested rollback.
“At my last company I led moving about sixty applications and their databases out of a data centre whose lease was ending. We spent the first month on discovery, because the dependency map everyone believed in was wrong in several places; network flow data showed apps talking that nobody had listed. We sorted each app: most were rehosted as they were, a few databases moved to managed RDS, and around ten apps were retired outright. We ran it in waves, starting with internal tools so we could fix our runbook on low stakes. For databases we did a full load and then ongoing change capture, so the final cutover was only minutes of catch-up. Each cutover had a rehearsal, a go or no-go meeting and a rollback we'd actually tested. One wave did roll back because of a licensing check tied to hardware. We finished two weeks before the lease ended.”
Describing a migration as just copying servers across, with no rollback plan.
The project: what you set out to do and why it seemed right.
What went wrong: the real cause, including your part in it.
The call: how and when you decided to stop or reverse.
After: the lasting change in how you plan or decide.
“I led a move of our main API from containers to Lambda, mainly to cut the time we spent on patching and scaling. On paper it fit. In practice, a big share of our traffic was long-running report requests and some endpoints kept warm in-memory caches, and I hadn't looked closely enough at the traffic mix. Latency got worse, the database got hammered by connections, and after two months we'd moved only a third of the endpoints. I called a stop, which was hard because I'd championed it. We kept the event-driven parts on Lambda, where they worked well, and moved the rest back. Afterwards I changed how we start projects like this: a one-page decision record, a small proof on the nastiest endpoint rather than the easiest, and exit criteria written down at the start, so stopping is a planned outcome, not a failure.”
Blaming the tool or another team, or choosing a failure that was really a success.
Roles: name an incident lead, a communications owner and one lead per affected service.
Mitigate first: roll back recent changes, shift traffic or shed load before settling the root cause.
Updates on a clock: the communications owner updates leadership at a set interval, in a separate channel.
After: a blameless review with a timeline and actions that each have an owner.
“First I make the roles explicit, because arguing usually means nobody's in charge. I take incident lead, ask one person to own communications, and name one lead per affected service. Then I move the goal from 'why' to 'stop the bleeding': what changed recently, can we roll it back, can we move traffic away from a bad zone or turn off a non-essential feature to shed load? The root cause debate goes to a side thread while we mitigate. I also check the AWS Health Dashboard early, so we don't spend an hour debugging our code for a problem on the provider's side. For the VP, I'd say politely that the communications owner will post updates every fifteen minutes in the leadership channel, and we stick to that clock, even when the update is 'no change, trying this next'. Most leaders relax once updates are predictable. Afterwards we run a blameless review with a clear timeline and actions that each have an owner and a date.”
Diving into debugging yourself and leaving nobody running the call or talking to leadership.
Hiring: test judgement with a real design discussion and a debugging exercise, not trivia.
Guardrails: safe defaults and automatic checks so mistakes are small and caught early.
Growth: give people ownership of real systems, pair on incidents, and let them lead reviews.
Letting go: step back from approvals once the guardrails are in place.
“For hiring, I care more about judgement than memory. In interviews I give a small design problem with a twist, like a database that can't go down during a migration, and a broken setup to debug together, and I listen for how people reason and when they ask questions. Service trivia tells me very little. For speed, the trick is making mistakes cheap. We have sandbox accounts, guardrails from SCPs and pipeline checks, and nothing changes in production by hand, so a new engineer can ship in their first week without scaring anyone. For growth, I hand people ownership of a real system early, pair them with a senior on their first incidents, and rotate who leads design reviews. My own goal is to become unnecessary for approvals. When I find myself the bottleneck on something, that's a sign we're missing a guardrail or I haven't trusted someone enough yet.”
Saying you personally review every change to keep quality high.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.