Scenario rounds give you a broken system and watch how you think. An app can't reach its database after a private endpoint goes in, pods sit in Pending, a failover works but the app still errors, a bill jumps after a release. There is rarely one right answer. The interviewer wants the order you check things in, what you would look at first and what would change your mind. This page is for anyone facing that round, from a first cloud role to a senior hire. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying your first three checks before the fix.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Test the flow: Network Watcher IP flow verify names the exact rule that allows or denies the traffic.
Both layers: read the effective security rules; inbound traffic must pass the subnet NSG and the NIC NSG.
Get in safely: Bastion or the serial console if the fix takes time, never a wide-open rule.
Afterwards: put the rules in code so changes are reviewed.
“First I'd stop guessing and use Network Watcher. IP flow verify takes the VM, the port, and my source address, and tells me which rule allowed or denied it. I'd also open the effective security rules on the NIC, because there can be an NSG on the subnet and another on the NIC, and inbound traffic has to be allowed by both. Very often the tidy-up removed an allow rule in one of them, or added a deny with a lower priority number, which wins. If the fix needs a change approval, I'd reach the machine through Azure Bastion or the serial console in the meantime, rather than opening RDP to the whole internet. Once it's working, I'd get the NSG rules into Bicep so the next tidy-up goes through a review.”
Adding an allow-all rule from any source to fix it quickly and planning to tighten it later.
Read the error: it means the request still arrived over the public path.
Outbound path: the App Service needs VNet integration, or it never enters the VNet.
DNS: the server name must resolve to the private IP; the private DNS zone has to be linked to the VNet.
Prove it: resolve the name from the app's own console, then retest.
“That error is a useful clue. It means our connection still reached SQL over the public path, so the app isn't using the private endpoint at all. I'd check two things. First, the App Service has to be integrated with the VNet, otherwise its outbound traffic never enters the VNet. Second, DNS. The app still connects to the normal server name, and that name only points to the private IP if the privatelink database zone exists and is linked to the VNet the app resolves from. If the company runs its own DNS servers, they need to forward that zone to Azure's resolver. To prove it, I'd open the app's console and resolve the server name. If I get a public address back, it's DNS. Once it returns the private IP, the connection works, and public access can stay off.”
Turning public access back on to make the error go away without finding out why the private path wasn't used.
Theory: outbound calls go through a limited pool of SNAT ports; opening a new connection per request uses them up.
Test: the SNAT port exhaustion check in Diagnose and solve problems, plus how the code creates HTTP clients.
Fix in code: reuse connections through one shared or factory-managed client, with keep-alive.
Fix in platform: VNet integration with a NAT gateway for more ports; private endpoints for Azure services.
“When outbound calls time out under load but the app looks healthy and the other side sees nothing, I suspect SNAT port exhaustion. Outbound connections from App Service to the internet share a limited pool of ports, and if the code opens a new connection for every request, for example creating a new HTTP client each time, those ports run out and new calls wait until they time out. The payment team sees nothing because the requests never reach them. To test it, I'd run the SNAT port exhaustion check under Diagnose and solve problems for the peak window, and read how the code creates its HTTP clients. The main fix is in code: reuse connections through a client factory. If we truly need more outbound connections, I'd add VNet integration with a NAT gateway, and use private endpoints for Azure services so that traffic doesn't use SNAT at all.”
// Program.cs: one pooled client, reused across requests
builder.Services.AddHttpClient("payments", client =>
{
client.BaseAddress = new Uri("https://api.payments.example.com/");
client.Timeout = TimeSpan.FromSeconds(10);
});
// In a service: var http = httpClientFactory.CreateClient("payments");
Scaling up the plan to a bigger size because of timeouts, without checking outbound connections.
Peering settings: allow gateway transit on the hub side and use remote gateways on the spoke side.
Routes out: effective routes on a spoke VM's NIC show where on-premises traffic actually goes.
Routes back: on-premises must learn the spoke's address range, and a firewall in the path needs rules and matching return routes.
Prove it: Network Watcher next hop and connection troubleshoot from the spoke VM.
“Since the hub works, the tunnel is fine, so I'd focus on what's different for the spoke. First, the peering. The hub side needs allow gateway transit and the spoke side needs use remote gateways, or the spoke never learns the routes to on-premises. I'd then open the effective routes on a spoke VM's network card to see where traffic for the on-premises range actually goes. If there's a route table sending everything to a firewall in the hub, the firewall needs a rule allowing it. Then the return path. On-premises has to know the spoke's address range: with BGP it's advertised once gateway transit is set, but a static VPN needs it added on the on-premises side. If traffic goes through the firewall one way but not back, it gets dropped, so the gateway subnet may need its own route table. Next hop in Network Watcher confirms each step.”
Only checking the spoke side and never asking whether on-premises knows how to send traffic back.
Read it right: a script ignores CORS; the browser blocks the call because the storage service didn't allow the page's origin.
Fix: add a CORS rule on the Blob service for the site's origin, the methods and the headers the upload sends.
Rule out the SAS: a bad or expired token gives a 403 from storage, not a CORS message.
Keep it tight: your own origins only, short-lived SAS with write rights on one container.
“The fact that the script works tells me the SAS token is fine. CORS is a browser rule: before the real upload, the browser sends a preflight request, and if the storage account doesn't say this origin is allowed, the browser blocks it. Scripts don't do that check. So I'd go to the storage account and add a CORS rule on the Blob service with our site's origin, the PUT and OPTIONS methods, and the headers the upload sends, like the blob type header. If the error were a 403 with an authentication message instead, I'd look at the SAS: expiry, permissions, or a start time slightly in the future because of clock differences. I'd keep the allowed origins to our own domains, not a star, and hand out short-lived SAS tokens that can only write to one container.”
Regenerating the storage keys or making the container public to get past what is really a browser check.
The fact: an archived blob is offline; it has to be rehydrated to an online tier first.
Options: standard priority can take many hours; high priority is often under an hour for smaller files, but not promised.
How: copy it into a new blob in the Hot tier, which leaves the original in Archive.
Expectations and prevention: tell the manager straight away; revisit which data goes to Archive.
“First I'd set expectations, because an archived blob can't be read at all until it's rehydrated to an online tier. I'd say I can start now, but I can't promise an hour. Standard rehydration can take many hours. High priority often finishes in under an hour for smaller files, and a contract is small, but it isn't guaranteed. So I'd kick off high priority right away, and I'd do it as a copy into a new blob in the Hot tier rather than changing the tier of the original. That leaves the archived copy where it is and avoids an early deletion charge if it hasn't been in Archive long enough. I'd watch the rehydration status and send the file as soon as it lands. Afterwards, I'd look at the lifecycle rule: documents people may need at short notice probably belong in Cold, not Archive.”
Promising the file in minutes, or telling the manager it's lost because it was archived.
Confirm: the Usage and quotas page shows which limit you hit, regional total or VM family.
Now: request the increase; meanwhile delete idle test VMs (deallocated ones still count) or use a VM family with headroom.
Protect the app: caching and shedding non-essential work while capacity catches up.
Next time: check quotas as part of launch readiness, well before the date.
“A quota error like that means Azure isn't out of machines, our subscription is at its limit. There are two kinds: a total cores limit for the region and a limit per VM family, so first I'd open Usage and quotas and see which one we hit. Then I'd raise a quota increase request straight away; many go through quickly, but I can't count on it. Meanwhile I'd free cores. Deallocated VMs still count against the quota, so that means deleting idle test VMs in the same subscription and region, or scaling test scale sets to zero. Or I'd switch the scale set to a VM size from a family that still has room. I'd also ease the load with caching or by switching off heavy non-essential features. For the next launch, quota checks go on the readiness list weeks ahead, next to the load test.”
Assuming the region has run out of capacity and waiting for Azure to fix it.
Duplicates: a message becomes visible again if the run fails or times out after the email went out.
Poison queue: after the maximum number of tries, the message is moved aside instead of retried forever.
Find the cause: Application Insights failures and timeouts for those message IDs.
Fix: make processing idempotent, keyed on the order ID; handle poison messages on purpose.
“Queue triggers deliver messages at least once, not exactly once. The function picks up a message, and only when the run finishes successfully is it deleted. If the run throws or times out after the email already went out, the message comes back and runs again, so the customer gets a second email. If a message keeps failing, after the maximum number of tries, five by default, the runtime moves it to the poison queue so it stops blocking everything else. So I'd look in Application Insights for failures and timeouts on those message IDs, since it's probably one step failing after the email. The real fix is to make the function safe to run twice: record that confirmation was sent for that order ID and check before sending. And I'd alert on anything landing in the poison queue.”
Expecting the queue to guarantee exactly-once delivery, or silently clearing the poison queue.
Look first: boot diagnostics shows the screen or the serial log, so you see where boot stops.
Serial console: a text session that works without the network, often enough to roll back.
Repair VM: attach a copy of the OS disk to a rescue VM, fix it there, swap it back.
Protect the service: fail over or restore from backup if the fix will take too long.
“First I'd open boot diagnostics, which shows a screenshot of the console or the serial log, so I can see if it's stuck on a failed update, a disk check, or a kernel panic. If the service is down, I'd get it running elsewhere in parallel: fail over to the other instance, or restore last night's backup. Next I'd try the serial console, which works even when networking is broken. From there I can often roll back the update or fix a bad config file. If that doesn't work, I'd use the VM repair commands: they make a copy of the OS disk, attach it to a rescue VM, and I fix it there, for example removing the bad package or fixing the boot config. Then I swap the fixed disk back. Afterwards, I'd patch one machine first and wait before patching the rest.”
# Needs the vm-repair CLI extension; you'll be asked for a rescue admin password
az vm repair create -g rg-prod -n app-vm-01 --repair-username rescueadmin --verbose
# ...fix the attached OS disk copy on the rescue VM...
az vm repair restore -g rg-prod -n app-vm-01 --verbose
Deleting the VM and rebuilding it before looking at the boot log, losing the chance to learn the cause.
Cause: session or cart data held in memory on one instance; the next request lands on another.
Check: is the app using in-memory session or cache, and is ARR affinity switched off?
Quick relief: turning affinity on pins users, but restarts and scale-in still lose data.
Real fix: keep shared state in a distributed store such as Azure Cache for Redis or the database.
“When it breaks only after scaling out, I think state first. With one instance, keeping session data or carts in memory works. With three, the load balancer can send a user's next request to a different instance that has never heard of them, so they look logged out. I'd check how the app stores sessions and whether ARR affinity, the setting that pins a user to one instance with a cookie, has been switched off. Turning it back on would reduce the problem quickly, but it's only a patch: any restart, deploy or scale-in still wipes that memory, and load spreads unevenly. The proper fix is to make the instances stateless, keeping sessions and carts in Azure Cache for Redis or the database, so any instance can serve any request. After that change I'd test by restarting one instance mid-session.”
Scaling back to one instance and treating that as the fix.
Know the job: run time, memory, dependencies, what it reads and writes, what happens if it fails halfway.
Rule out: the Consumption plan for Functions caps a run at ten minutes, so an hour-long job doesn't fit there.
Good fits: a scheduled Container Apps job for a packaged task; Azure Batch for big parallel work.
Cheapest change: keep the VM but start and stop it on a schedule, if the job can't move yet.
“I'd start with the job itself: how long it really runs, how much memory it needs, what it installs, and what happens if it dies halfway. That rules things out quickly. Functions on the Consumption plan stop a run at ten minutes, so an hour-long job doesn't fit unless we split it into smaller steps. If the job can be packaged in a container, a scheduled Container Apps job is a good fit: it starts on a cron schedule, runs to completion, and we pay nothing between runs. If it's heavy work that splits into many parallel pieces, Azure Batch makes more sense. If it depends on software that's hard to containerise, the quickest win is keeping the VM but having automation start it before the job and deallocate it after. Whatever we pick, I'd add retries, an alert when it fails, and a log of each run.”
Moving it to a Consumption plan function without checking the run time limit.
Identity: invite them as a guest so they sign in with their own account; no shared logins.
Scope and role: assign on that resource group only, with the narrowest role that covers the work.
Time limit: an end date on the assignment, ideally through just-in-time access.
Guardrails: MFA for guests, activity log review, access removed at the end.
“I'd invite the contractor into Entra ID as a guest, so they sign in with their own company account and I never create or share a password for them. Then I'd ask what the work actually is. If they're building and changing resources, Contributor on that one resource group is usually enough; if it's only an App Service, there's often a narrower built-in role. I'd never give them anything at the subscription level. For the two weeks, I'd make the assignment time-bound, ideally through Privileged Identity Management with an end date, so it expires even if everyone forgets. I'd make sure a Conditional Access rule asks guests for MFA. At the end I'd confirm the access is gone, look over the activity log for what they changed, and make sure their work is in our code, not only in the portal.”
Creating a shared account in the tenant for the contractor, or making them Owner of the subscription to save time.
Permission model: is the vault using access policies or Azure RBAC? A grant in the other model does nothing.
Right role: reading secrets needs a data role such as Key Vault Secrets User; Reader doesn't cover it.
Right identity: system or user-assigned; a user-assigned one needs its client ID in the app's config.
Other causes: the vault firewall, and a new assignment that hasn't taken effect yet.
“First I'd check which permission model the vault uses. A vault uses either access policies or Azure RBAC, and if you add an RBAC role to a vault that runs on access policies, or the other way round, nothing changes. If it's RBAC, I'd check the role. Reader on the vault only covers management, so it can see the vault but not read secrets. The app needs something like Key Vault Secrets User. Then I'd check it's the right identity. If the app uses a user-assigned identity, the code has to be told its client ID, or it picks up nothing or the wrong one. I'd also look at the vault's firewall, because a blocked network also gives a 403, and the error text says which it is. And a role assigned a minute ago can take a few minutes to kick in.”
// User-assigned identity: tell the credential which one to use
var credential = new DefaultAzureCredential(new DefaultAzureCredentialOptions
{
ManagedIdentityClientId = builder.Configuration["AZURE_CLIENT_ID"]
});
var client = new SecretClient(new Uri("https://kv-orders-prod.vault.azure.net/"), credential);
KeyVaultSecret secret = await client.GetSecretAsync("SqlPassword");
Giving the identity Owner or Contributor on the vault, which grants far too much and, under RBAC, still doesn't let it read secrets.
Now: create a new short-lived secret on that app registration and update the service connection.
Check scope: list what the service principal can touch while you're in there, and trim it.
Remove the secret: move pipelines to workload identity federation so nothing expires or leaks.
Prevent: owners on every app registration and an alert before credentials expire.
“First I'd get deployments working again. I'd find the app registration behind the pipeline's service connection, add a new client secret with a short expiry, update the service connection, and rerun one pipeline to confirm. Anything urgent can deploy tonight. While I'm there, I'd check what that service principal can access, because old ones often have Owner on far more than they need, and I'd trim it. Then the real fix: move the pipelines to workload identity federation. The pipeline gets a short-lived token from Entra ID that trusts our pipeline directly, so there's no secret to expire, rotate or leak. Once that works, I'd delete the old secret. Finally, every app registration gets at least two owners who still work here, and a scheduled check that warns us weeks before any remaining credential expires.”
Creating a new secret that never expires, or one valid for years, so the same outage just moves further out.
Outside in: availability tests in Application Insights hitting the real URL from several locations.
Symptom alerts: server errors, response time and failed test locations, not only CPU.
Reach someone: action groups that page the on-call person, and a test of the whole chain.
Learn: a short review of why it went down and why nothing fired.
“The gap is that we were only watching the inside of the system, if anything. First, I'd add availability tests in Application Insights that request the home page and one key page every few minutes from several regions, and alert when more than one location fails, so one bad probe doesn't page anyone. Second, alerts on symptoms users feel: a spike in 5xx responses from the App Service and response time going up. CPU alerts alone miss most outages. Third, the alerts have to reach a person, so I'd set up an action group that pages whoever is on call, not a shared inbox nobody reads at night, and then trigger a test alert to prove the whole chain works. I'd also run a short blameless review of last night: what broke, and why nothing noticed.”
Adding a CPU alert and calling it done, or sending alerts to an email list nobody watches out of hours.
Resources: is the database hitting its CPU, data IO or log limits, or waiting on locks?
Query Store: compare this query's plans over time; a new plan often means a regression.
Why it flipped: changed statistics, data growth or a parameter value that suited one caller.
Fix: force the last good plan now, then fix the root cause with an index or query change.
“No deploy doesn't mean nothing changed: data grows, statistics update, and the optimiser can pick a new plan. First I'd check resource use for the database. If CPU or IO is pinned at the limit, the query is waiting for resources, not slow on its own. I'd also check for blocking. If resources look fine, I'd open Query Store, find the query, and look at its plans over time. Very often there's a new plan from a few days ago, maybe a scan where it used to seek, compiled for a parameter value that suited a different caller. The quick fix is to force the old plan from Query Store, or let automatic tuning do it. Then I'd fix the real cause, like a missing index or a query that behaves badly with skewed data, so we're not relying on a forced plan forever.”
Scaling the database up straight away without checking whether the plan changed.
Connection string: it must use the failover group listener, not a server name.
Logins: server logins must exist on the second server with the same SIDs, or use contained or Entra users.
Network rules: firewall rules and private endpoints are per server and must exist in both regions.
Make it routine: fix it in code and rerun the test until failover is a non-event.
“A failover group moves the data, not everything around it. First I'd check the connection string. If the app points at the primary server's own name, it's still talking to the old server. It should use the failover group's listener name, which always points at whichever side is primary. Next, logins. If the app uses a SQL login created on the server, that login lives in the primary's master database, and it has to exist on the second server with the same SID, or the database users end up orphaned. Contained users or Entra authentication avoid that. Then networking: firewall rules and private endpoints belong to each server, so the second one needs its own, plus DNS that works in that region. I'd write down every gap, fix it in code, and run the test again.”
Calling the DR test passed because the database came up, without checking the app could use it.
Why: provisioned throughput is spread evenly across physical partitions; one busy partition hits its share.
Prove it: normalized RU consumption by partition key range, and diagnostic logs showing the hot key values.
Also check: expensive queries that fan out across partitions, and heavy indexing on write-heavy data.
Fix: a better key means a new container and a migration; tune queries and indexing meanwhile.
“The total can look fine while one partition is overloaded. Cosmos DB spreads the container's throughput evenly across its physical partitions, so if most of the evening traffic lands on one key, that partition runs out of its share and returns 429s while the others sit idle. I'd check the normalized RU consumption metric split by partition key range. If one range is near the top and the rest are low, that's the hot partition. The diagnostic logs then show which key values are doing it. I'd also look for queries without the partition key in the filter, which fan out across partitions, and an indexing policy indexing fields nobody queries, which makes writes more expensive. The lasting fix is usually a better partition key, which means a new container and moving the data. Raising throughput alone mostly pays for idle partitions.”
Doubling the provisioned throughput and calling it fixed without checking how the load is spread.
Cause: Complete mode deletes anything in the group that the template doesn't list.
Now: stop the pipeline, try to recover the account, tell whoever owned it.
Prevent: Incremental by default, what-if output reviewed before apply, delete locks on production.
Longer term: decide who owns each resource; bring the stray one into code or move it out.
“That's almost certainly a deployment in Complete mode. The default, Incremental, only adds or updates what's in the template. Complete makes the resource group match the template exactly, so anything not listed gets deleted, including a storage account someone created by hand. First I'd pause the pipeline so it doesn't run again, check whether the account can be recovered, which is possible for a short time if the name hasn't been reused, and tell whoever owned it. Then I'd check the deployment history to confirm the mode. To prevent it, I'd switch to Incremental unless there's a strong reason, add a what-if step that shows every change before apply, and put delete locks on production resource groups. And the stray account needs an owner: either it goes into our template, or it moves to its own group.”
# Preview every create, change and delete before applying
az deployment group what-if --resource-group rg-app-prod --template-file main.bicep --mode Incremental
az deployment group create --resource-group rg-app-prod --template-file main.bicep --mode Incremental
Blaming whoever created the account by hand without seeing that the pipeline was set to delete anything it didn't know.
Find it: query the Usage table for billable volume by data type, before and after the release.
Source: usually debug logging left on, a chatty dependency, or a diagnostic setting sending every category.
Cut at the source: log level back down, sampling in Application Insights, only needed categories.
Guardrails: cheaper table plans for bulky logs, a budget alert, and a daily cap only as a last resort.
“A jump right after a release points at what the app is sending, not new resources. First I'd query the Usage table in the workspace to see billable volume by data type per day, which usually shows one table jumping on release day, often traces or dependencies from Application Insights. Then I'd find why: a log level left on debug, a new retry loop writing an error on every attempt, or a diagnostic setting someone switched to send every category. The fix is at the source: log level back to warning in production, sampling turned on for high-volume telemetry, and only the diagnostic categories we actually query. For bulky logs we keep but rarely search, a cheaper table plan helps. I'd add a budget alert on the workspace and add a volume check to the release checklist. A daily cap is a last resort, because it stops collecting right when you might need the logs.”
Usage
| where TimeGenerated > ago(30d)
| where IsBillable == true
| summarize IngestedGB = sum(Quantity) / 1000 by DataType, bin(TimeGenerated, 1d)
| sort by IngestedGB desc
Setting a tight daily cap straight away, so logs stop arriving in the middle of the next incident.
Triage: check sign-in logs on the exposed machines; any suspicious access becomes an incident.
Replace access: Bastion or just-in-time access, so people still get in without an open port.
Close it: remove the public IPs and the open rules, owner by owner, with a date.
Guardrail: policy at the management group that denies public IPs on VMs, with exemptions reviewed.
“RDP open to the internet gets hammered by password guessing all day, so I'd treat it as urgent but not panic. First I'd check the security logs on those machines for successful sign-ins from strange places, and if I find any, that VM becomes an incident. Then I'd give the owners a better way in before taking the old one away: Azure Bastion, so they connect through the portal over HTTPS without any public IP on the VM, or just-in-time access that opens the port to their own address for a few hours. With that ready, I'd remove the public IPs and the open rules subscription by subscription, with a date agreed with each owner. To stop it coming back, I'd assign a policy at the management group that denies public IPs on VM network cards, with a reviewed exemption process for real exceptions.”
Deleting the public IPs on day one without giving anyone another way in, so teams find a worse workaround.
Read the event: describe the pod; the message says not found, unauthorized or a timeout.
Not found: a wrong tag or registry name in the manifest.
Unauthorized: the kubelet identity needs AcrPull; attaching the registry sets it up.
Timeout: a registry firewall or private endpoint the nodes can't reach.
“Older pods working often just means their image was already cached on those nodes, so I wouldn't read too much into it. First I'd describe one of the stuck pods and read the event message, because it tells me which of three problems I have. If it says not found, it's the tag or the registry name, often a pipeline that pushed a tag nobody deployed. If it says unauthorized, the cluster's kubelet identity probably doesn't have AcrPull on the registry, maybe because the registry was recreated or the role assignment was removed. Attaching the registry to the cluster puts that role back, and the check command confirms the nodes can pull. If it's a timeout, I'd look at networking: the registry may now only allow private access, and the nodes can't reach its private endpoint or resolve its name.”
kubectl describe pod orders-api-7d9f8b6c4-x2k9p -n orders
az aks check-acr --name aks-prod --resource-group rg-prod --acr myregistry.azurecr.io
az aks update --name aks-prod --resource-group rg-prod --attach-acr myregistry
Turning on the registry's admin user and pasting its password into a secret as the fix.
Scheduler first: describe a pending pod; the event says insufficient CPU, memory, or no matching node.
Autoscaler limits: the node pool may be at its maximum count, or requests too large for any node size.
Azure limits: vCPU quota for the region, or no IPs left in the subnet with classic Azure CNI.
Fix and prevent: raise the limit that's hit; plan headroom, overlay networking or a bigger subnet.
“I'd start with a pending pod and read its events, because the scheduler says why: insufficient CPU or memory, or no node matching a selector or taint. If it's insufficient resources, the question becomes why the autoscaler didn't add nodes. I'd check its status and events. Common reasons: the node pool is already at its maximum count, or the pods request more than any single node of that size can offer, so a new node wouldn't help. Then the Azure-side limits. The subscription may have hit its vCPU quota for that VM family in the region, which shows up as failed scale operations in the activity log. And with classic Azure CNI, every node reserves IPs for its pods up front, so a small subnet runs out and new nodes can't join. Short term I'd raise whichever limit is hit; long term, headroom and overlay networking.”
Deleting and recreating the pods over and over, hoping they'll schedule the next time.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.