Troubleshooting • Identity • Networking • Databases • Containers • Resilience • Cost • 2026

Scenario-Based Azure Interview Questions

Scenario rounds give you a broken system and watch how you think. An app can't reach its database after a private endpoint goes in, pods sit in Pending, a failover works but the app still errors, a bill jumps after a release. There is rarely one right answer. The interviewer wants the order you check things in, what you would look at first and what would change your mind. This page is for anyone facing that round, from a first cloud role to a senior hire. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying your first three checks before the fix.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Networking 4 questions

Easy Technical round Fresher, Mid-level Practice question

1. Nobody can RDP into a Windows VM since a teammate tidied up the network security groups yesterday. How do you find the problem and get back in?

What the interviewer is really testing:
Whether you know NSGs can sit on both the subnet and the NIC, and that Azure gives you tools to test a flow instead of guessing.
Answer frame:

Test the flow: Network Watcher IP flow verify names the exact rule that allows or denies the traffic.

Both layers: read the effective security rules; inbound traffic must pass the subnet NSG and the NIC NSG.

Get in safely: Bastion or the serial console if the fix takes time, never a wide-open rule.

Afterwards: put the rules in code so changes are reviewed.

Sample spoken answer:

“First I'd stop guessing and use Network Watcher. IP flow verify takes the VM, the port, and my source address, and tells me which rule allowed or denied it. I'd also open the effective security rules on the NIC, because there can be an NSG on the subnet and another on the NIC, and inbound traffic has to be allowed by both. Very often the tidy-up removed an allow rule in one of them, or added a deny with a lower priority number, which wins. If the fix needs a change approval, I'd reach the machine through Azure Bastion or the serial console in the meantime, rather than opening RDP to the whole internet. Once it's working, I'd get the NSG rules into Bicep so the next tidy-up goes through a review.”

Red flag to avoid:

Adding an allow-all rule from any source to fix it quickly and planning to tighten it later.

They may ask next:
  • How does rule priority work when two rules match the same traffic?
  • What would you change so that RDP isn't open all the time?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

2. You added a private endpoint to Azure SQL and switched off public access. Now the App Service can't connect, and the error says public network access is denied. What do you check?

What the interviewer is really testing:
Whether you know a private endpoint only works when the app's traffic goes into the VNet and the name resolves to the private IP.
Answer frame:

Read the error: it means the request still arrived over the public path.

Outbound path: the App Service needs VNet integration, or it never enters the VNet.

DNS: the server name must resolve to the private IP; the private DNS zone has to be linked to the VNet.

Prove it: resolve the name from the app's own console, then retest.

Sample spoken answer:

“That error is a useful clue. It means our connection still reached SQL over the public path, so the app isn't using the private endpoint at all. I'd check two things. First, the App Service has to be integrated with the VNet, otherwise its outbound traffic never enters the VNet. Second, DNS. The app still connects to the normal server name, and that name only points to the private IP if the privatelink database zone exists and is linked to the VNet the app resolves from. If the company runs its own DNS servers, they need to forward that zone to Azure's resolver. To prove it, I'd open the app's console and resolve the server name. If I get a public address back, it's DNS. Once it returns the private IP, the connection works, and public access can stay off.”

Red flag to avoid:

Turning public access back on to make the error go away without finding out why the private path wasn't used.

They may ask next:
  • Why shouldn't the app connect to the private IP address directly instead of the name?
  • How would you set this up for on-premises clients that also need the database?
Say it in 60 seconds
Hard Technical round Senior Practice question

3. Under heavy load, your App Service gets intermittent timeouts calling an outside payment API. CPU and memory look fine, and the API's team says they see no errors. What's your theory and how do you test it?

What the interviewer is really testing:
Whether you think of SNAT port exhaustion on outbound connections, and know the fix is mostly connection reuse in code.
Answer frame:

Theory: outbound calls go through a limited pool of SNAT ports; opening a new connection per request uses them up.

Test: the SNAT port exhaustion check in Diagnose and solve problems, plus how the code creates HTTP clients.

Fix in code: reuse connections through one shared or factory-managed client, with keep-alive.

Fix in platform: VNet integration with a NAT gateway for more ports; private endpoints for Azure services.

Sample spoken answer:

“When outbound calls time out under load but the app looks healthy and the other side sees nothing, I suspect SNAT port exhaustion. Outbound connections from App Service to the internet share a limited pool of ports, and if the code opens a new connection for every request, for example creating a new HTTP client each time, those ports run out and new calls wait until they time out. The payment team sees nothing because the requests never reach them. To test it, I'd run the SNAT port exhaustion check under Diagnose and solve problems for the peak window, and read how the code creates its HTTP clients. The main fix is in code: reuse connections through a client factory. If we truly need more outbound connections, I'd add VNet integration with a NAT gateway, and use private endpoints for Azure services so that traffic doesn't use SNAT at all.”

Code:
// Program.cs: one pooled client, reused across requests
builder.Services.AddHttpClient("payments", client =>
{
    client.BaseAddress = new Uri("https://api.payments.example.com/");
    client.Timeout = TimeSpan.FromSeconds(10);
});

// In a service: var http = httpClientFactory.CreateClient("payments");
Red flag to avoid:

Scaling up the plan to a bigger size because of timeouts, without checking outbound connections.

They may ask next:
  • Why is wrapping a new HttpClient in a using block per request a problem here?
  • What else would you add so one slow outside API can't pile up requests in your app?
Say it in 60 seconds
Hard Technical round Senior Practice question

4. In a hub-and-spoke network, VMs in a new spoke can't reach on-premises servers, but the hub can. The VPN gateway is in the hub. What do you check?

What the interviewer is really testing:
Whether you know gateway transit on peering, effective routes and the return path, and debug routing in both directions.
Answer frame:

Peering settings: allow gateway transit on the hub side and use remote gateways on the spoke side.

Routes out: effective routes on a spoke VM's NIC show where on-premises traffic actually goes.

Routes back: on-premises must learn the spoke's address range, and a firewall in the path needs rules and matching return routes.

Prove it: Network Watcher next hop and connection troubleshoot from the spoke VM.

Sample spoken answer:

“Since the hub works, the tunnel is fine, so I'd focus on what's different for the spoke. First, the peering. The hub side needs allow gateway transit and the spoke side needs use remote gateways, or the spoke never learns the routes to on-premises. I'd then open the effective routes on a spoke VM's network card to see where traffic for the on-premises range actually goes. If there's a route table sending everything to a firewall in the hub, the firewall needs a rule allowing it. Then the return path. On-premises has to know the spoke's address range: with BGP it's advertised once gateway transit is set, but a static VPN needs it added on the on-premises side. If traffic goes through the firewall one way but not back, it gets dropped, so the gateway subnet may need its own route table. Next hop in Network Watcher confirms each step.”

Red flag to avoid:

Only checking the spoke side and never asking whether on-premises knows how to send traffic back.

They may ask next:
  • Why is it risky to send the gateway subnet's traffic through a firewall, and how would you do it safely?
  • How would Virtual WAN change the way you'd build this?
Say it in 60 seconds

Storage 2 questions

Easy Technical round Fresher, Mid-level Practice question

5. Your web page uploads files straight to Blob Storage with a SAS URL. It works from a script, but in the browser every upload fails with a CORS error. What's going on?

What the interviewer is really testing:
Whether you can tell a browser security check apart from a permission problem, and know CORS is set on the storage account.
Answer frame:

Read it right: a script ignores CORS; the browser blocks the call because the storage service didn't allow the page's origin.

Fix: add a CORS rule on the Blob service for the site's origin, the methods and the headers the upload sends.

Rule out the SAS: a bad or expired token gives a 403 from storage, not a CORS message.

Keep it tight: your own origins only, short-lived SAS with write rights on one container.

Sample spoken answer:

“The fact that the script works tells me the SAS token is fine. CORS is a browser rule: before the real upload, the browser sends a preflight request, and if the storage account doesn't say this origin is allowed, the browser blocks it. Scripts don't do that check. So I'd go to the storage account and add a CORS rule on the Blob service with our site's origin, the PUT and OPTIONS methods, and the headers the upload sends, like the blob type header. If the error were a 403 with an authentication message instead, I'd look at the SAS: expiry, permissions, or a start time slightly in the future because of clock differences. I'd keep the allowed origins to our own domains, not a star, and hand out short-lived SAS tokens that can only write to one container.”

Red flag to avoid:

Regenerating the storage keys or making the container public to get past what is really a browser check.

They may ask next:
  • Why would you generate the SAS on the server instead of in the browser?
  • What is a user delegation SAS, and why might you prefer it?
Say it in 60 seconds
Medium Situational round Fresher, Mid-level, Senior Practice question

6. An auditor needs a contract file today, and it's a blob a lifecycle rule moved to the Archive tier last year. The manager wants it within the hour. What do you tell them and do?

What the interviewer is really testing:
Whether you know archived blobs can't be read until rehydrated, the priority options, and set expectations honestly.
Answer frame:

The fact: an archived blob is offline; it has to be rehydrated to an online tier first.

Options: standard priority can take many hours; high priority is often under an hour for smaller files, but not promised.

How: copy it into a new blob in the Hot tier, which leaves the original in Archive.

Expectations and prevention: tell the manager straight away; revisit which data goes to Archive.

Sample spoken answer:

“First I'd set expectations, because an archived blob can't be read at all until it's rehydrated to an online tier. I'd say I can start now, but I can't promise an hour. Standard rehydration can take many hours. High priority often finishes in under an hour for smaller files, and a contract is small, but it isn't guaranteed. So I'd kick off high priority right away, and I'd do it as a copy into a new blob in the Hot tier rather than changing the tier of the original. That leaves the archived copy where it is and avoids an early deletion charge if it hasn't been in Archive long enough. I'd watch the rehydration status and send the file as soon as it lands. Afterwards, I'd look at the lifecycle rule: documents people may need at short notice probably belong in Cold, not Archive.”

Red flag to avoid:

Promising the file in minutes, or telling the manager it's lost because it was archived.

They may ask next:
  • How would you know when rehydration has finished without checking by hand?
  • What data really is a good fit for the Archive tier?
Say it in 60 seconds

Compute 5 questions

Easy Situational round Fresher, Mid-level, Senior Practice question

7. Launch day traffic is climbing, and your scale set won't add instances. The error mentions exceeding the approved regional cores. What do you do right now, and what do you change for next time?

What the interviewer is really testing:
Whether you know subscriptions have per-region and per-family vCPU quotas, and how to buy time while a quota request goes through.
Answer frame:

Confirm: the Usage and quotas page shows which limit you hit, regional total or VM family.

Now: request the increase; meanwhile delete idle test VMs (deallocated ones still count) or use a VM family with headroom.

Protect the app: caching and shedding non-essential work while capacity catches up.

Next time: check quotas as part of launch readiness, well before the date.

Sample spoken answer:

“A quota error like that means Azure isn't out of machines, our subscription is at its limit. There are two kinds: a total cores limit for the region and a limit per VM family, so first I'd open Usage and quotas and see which one we hit. Then I'd raise a quota increase request straight away; many go through quickly, but I can't count on it. Meanwhile I'd free cores. Deallocated VMs still count against the quota, so that means deleting idle test VMs in the same subscription and region, or scaling test scale sets to zero. Or I'd switch the scale set to a VM size from a family that still has room. I'd also ease the load with caching or by switching off heavy non-essential features. For the next launch, quota checks go on the readiness list weeks ahead, next to the load test.”

Red flag to avoid:

Assuming the region has run out of capacity and waiting for Azure to fix it.

They may ask next:
  • How would you split production and test so one can't eat the other's quota?
  • What else, besides quota, could stop a scale set from adding instances?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

8. A queue-triggered Azure Function processes orders. Some customers got two confirmation emails, and a few messages ended up in a queue ending in -poison. What's happening?

What the interviewer is really testing:
Whether you know queue triggers deliver at least once, how retries and the poison queue work, and that the fix is idempotent code.
Answer frame:

Duplicates: a message becomes visible again if the run fails or times out after the email went out.

Poison queue: after the maximum number of tries, the message is moved aside instead of retried forever.

Find the cause: Application Insights failures and timeouts for those message IDs.

Fix: make processing idempotent, keyed on the order ID; handle poison messages on purpose.

Sample spoken answer:

“Queue triggers deliver messages at least once, not exactly once. The function picks up a message, and only when the run finishes successfully is it deleted. If the run throws or times out after the email already went out, the message comes back and runs again, so the customer gets a second email. If a message keeps failing, after the maximum number of tries, five by default, the runtime moves it to the poison queue so it stops blocking everything else. So I'd look in Application Insights for failures and timeouts on those message IDs, since it's probably one step failing after the email. The real fix is to make the function safe to run twice: record that confirmation was sent for that order ID and check before sending. And I'd alert on anything landing in the poison queue.”

Red flag to avoid:

Expecting the queue to guarantee exactly-once delivery, or silently clearing the poison queue.

They may ask next:
  • How would you replay the poison messages once the bug is fixed?
  • When would Service Bus suit this better than a storage queue?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

9. A production VM didn't come back after last night's patching. You can't RDP or SSH in. How do you work out what's wrong and fix it?

What the interviewer is really testing:
Whether you know Azure's tools for a VM you can't log into: boot diagnostics, the serial console and a repair VM.
Answer frame:

Look first: boot diagnostics shows the screen or the serial log, so you see where boot stops.

Serial console: a text session that works without the network, often enough to roll back.

Repair VM: attach a copy of the OS disk to a rescue VM, fix it there, swap it back.

Protect the service: fail over or restore from backup if the fix will take too long.

Sample spoken answer:

“First I'd open boot diagnostics, which shows a screenshot of the console or the serial log, so I can see if it's stuck on a failed update, a disk check, or a kernel panic. If the service is down, I'd get it running elsewhere in parallel: fail over to the other instance, or restore last night's backup. Next I'd try the serial console, which works even when networking is broken. From there I can often roll back the update or fix a bad config file. If that doesn't work, I'd use the VM repair commands: they make a copy of the OS disk, attach it to a rescue VM, and I fix it there, for example removing the bad package or fixing the boot config. Then I swap the fixed disk back. Afterwards, I'd patch one machine first and wait before patching the rest.”

Code:
# Needs the vm-repair CLI extension; you'll be asked for a rescue admin password
az vm repair create -g rg-prod -n app-vm-01 --repair-username rescueadmin --verbose
# ...fix the attached OS disk copy on the rescue VM...
az vm repair restore -g rg-prod -n app-vm-01 --verbose
Red flag to avoid:

Deleting the VM and rebuilding it before looking at the boot log, losing the chance to learn the cause.

They may ask next:
  • What would you change about patching so one bad update can't take down every server?
  • How is the serial console protected so it isn't a way around your normal access controls?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

10. After you scaled an App Service from one instance to three, users started getting logged out at random and losing their carts. Nothing else changed. What's the likely cause?

What the interviewer is really testing:
Whether you spot state kept in one instance's memory, and fix it by moving state out rather than pinning users.
Answer frame:

Cause: session or cart data held in memory on one instance; the next request lands on another.

Check: is the app using in-memory session or cache, and is ARR affinity switched off?

Quick relief: turning affinity on pins users, but restarts and scale-in still lose data.

Real fix: keep shared state in a distributed store such as Azure Cache for Redis or the database.

Sample spoken answer:

“When it breaks only after scaling out, I think state first. With one instance, keeping session data or carts in memory works. With three, the load balancer can send a user's next request to a different instance that has never heard of them, so they look logged out. I'd check how the app stores sessions and whether ARR affinity, the setting that pins a user to one instance with a cookie, has been switched off. Turning it back on would reduce the problem quickly, but it's only a patch: any restart, deploy or scale-in still wipes that memory, and load spreads unevenly. The proper fix is to make the instances stateless, keeping sessions and carts in Azure Cache for Redis or the database, so any instance can serve any request. After that change I'd test by restarting one instance mid-session.”

Red flag to avoid:

Scaling back to one instance and treating that as the fix.

They may ask next:
  • What other things can differ between instances and cause odd bugs after scaling out?
  • How would you move existing users' carts without losing them during the change?
Say it in 60 seconds
Medium Case round Fresher, Mid-level, Senior Practice question

11. A nightly job runs for about an hour on a VM that stays on all day doing nothing else. The team wants to stop paying for idle time. How would you decide where it should run?

What the interviewer is really testing:
Whether you match the job's run time, dependencies and failure handling to a service, and know the limits that rule some out.
Answer frame:

Know the job: run time, memory, dependencies, what it reads and writes, what happens if it fails halfway.

Rule out: the Consumption plan for Functions caps a run at ten minutes, so an hour-long job doesn't fit there.

Good fits: a scheduled Container Apps job for a packaged task; Azure Batch for big parallel work.

Cheapest change: keep the VM but start and stop it on a schedule, if the job can't move yet.

Sample spoken answer:

“I'd start with the job itself: how long it really runs, how much memory it needs, what it installs, and what happens if it dies halfway. That rules things out quickly. Functions on the Consumption plan stop a run at ten minutes, so an hour-long job doesn't fit unless we split it into smaller steps. If the job can be packaged in a container, a scheduled Container Apps job is a good fit: it starts on a cron schedule, runs to completion, and we pay nothing between runs. If it's heavy work that splits into many parallel pieces, Azure Batch makes more sense. If it depends on software that's hard to containerise, the quickest win is keeping the VM but having automation start it before the job and deallocate it after. Whatever we pick, I'd add retries, an alert when it fails, and a log of each run.”

Red flag to avoid:

Moving it to a Consumption plan function without checking the run time limit.

They may ask next:
  • How would you make the job safe to rerun if it fails halfway through?
  • What changes if the job has to finish before the business opens every morning?
Say it in 60 seconds

Identity and Access 3 questions

Easy Situational round Fresher, Mid-level Practice question

12. An outside contractor needs to work on the resources in one resource group for two weeks. They don't have an account in your tenant. How do you set up their access?

What the interviewer is really testing:
Whether you use guest accounts, the narrowest role at the narrowest scope, and access that ends by itself.
Answer frame:

Identity: invite them as a guest so they sign in with their own account; no shared logins.

Scope and role: assign on that resource group only, with the narrowest role that covers the work.

Time limit: an end date on the assignment, ideally through just-in-time access.

Guardrails: MFA for guests, activity log review, access removed at the end.

Sample spoken answer:

“I'd invite the contractor into Entra ID as a guest, so they sign in with their own company account and I never create or share a password for them. Then I'd ask what the work actually is. If they're building and changing resources, Contributor on that one resource group is usually enough; if it's only an App Service, there's often a narrower built-in role. I'd never give them anything at the subscription level. For the two weeks, I'd make the assignment time-bound, ideally through Privileged Identity Management with an end date, so it expires even if everyone forgets. I'd make sure a Conditional Access rule asks guests for MFA. At the end I'd confirm the access is gone, look over the activity log for what they changed, and make sure their work is in our code, not only in the portal.”

Red flag to avoid:

Creating a shared account in the tenant for the contractor, or making them Owner of the subscription to save time.

They may ask next:
  • Why assign the role to a group rather than to the person directly?
  • What would you do if they need to read secrets from a Key Vault in that group?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

13. The app has a managed identity, and you added it to the Key Vault, but reading a secret still returns 403 Forbidden. Walk me through what you'd check.

What the interviewer is really testing:
Whether you know the two Key Vault permission models, the data-plane roles, and the other things that return a 403.
Answer frame:

Permission model: is the vault using access policies or Azure RBAC? A grant in the other model does nothing.

Right role: reading secrets needs a data role such as Key Vault Secrets User; Reader doesn't cover it.

Right identity: system or user-assigned; a user-assigned one needs its client ID in the app's config.

Other causes: the vault firewall, and a new assignment that hasn't taken effect yet.

Sample spoken answer:

“First I'd check which permission model the vault uses. A vault uses either access policies or Azure RBAC, and if you add an RBAC role to a vault that runs on access policies, or the other way round, nothing changes. If it's RBAC, I'd check the role. Reader on the vault only covers management, so it can see the vault but not read secrets. The app needs something like Key Vault Secrets User. Then I'd check it's the right identity. If the app uses a user-assigned identity, the code has to be told its client ID, or it picks up nothing or the wrong one. I'd also look at the vault's firewall, because a blocked network also gives a 403, and the error text says which it is. And a role assigned a minute ago can take a few minutes to kick in.”

Code:
// User-assigned identity: tell the credential which one to use
var credential = new DefaultAzureCredential(new DefaultAzureCredentialOptions
{
    ManagedIdentityClientId = builder.Configuration["AZURE_CLIENT_ID"]
});
var client = new SecretClient(new Uri("https://kv-orders-prod.vault.azure.net/"), credential);
KeyVaultSecret secret = await client.GetSecretAsync("SqlPassword");
Red flag to avoid:

Giving the identity Owner or Contributor on the vault, which grants far too much and, under RBAC, still doesn't let it read secrets.

They may ask next:
  • Would you give each app its own vault or share one, and why?
  • How do Key Vault references in App Service settings change this setup?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

14. At midnight every deployment pipeline started failing with an error that the client secret has expired. It was created two years ago by someone who has left. What do you do now and next?

What the interviewer is really testing:
Whether you can restore deployments quickly, then remove the secret entirely with workload identity federation.
Answer frame:

Now: create a new short-lived secret on that app registration and update the service connection.

Check scope: list what the service principal can touch while you're in there, and trim it.

Remove the secret: move pipelines to workload identity federation so nothing expires or leaks.

Prevent: owners on every app registration and an alert before credentials expire.

Sample spoken answer:

“First I'd get deployments working again. I'd find the app registration behind the pipeline's service connection, add a new client secret with a short expiry, update the service connection, and rerun one pipeline to confirm. Anything urgent can deploy tonight. While I'm there, I'd check what that service principal can access, because old ones often have Owner on far more than they need, and I'd trim it. Then the real fix: move the pipelines to workload identity federation. The pipeline gets a short-lived token from Entra ID that trusts our pipeline directly, so there's no secret to expire, rotate or leak. Once that works, I'd delete the old secret. Finally, every app registration gets at least two owners who still work here, and a scheduled check that warns us weeks before any remaining credential expires.”

Red flag to avoid:

Creating a new secret that never expires, or one valid for years, so the same outage just moves further out.

They may ask next:
  • How does the federated credential know it's your pipeline asking and not someone else's?
  • How would you find every other secret in the tenant that's about to expire?
Say it in 60 seconds

Monitoring 1 question

Easy Situational round Fresher, Mid-level, Senior Practice question

15. The site was down for forty minutes last night and a customer told you before any alert did. What do you put in place this week?

What the interviewer is really testing:
Whether you monitor from the user's side, alert on symptoms, and make sure an alert reaches a person who can act.
Answer frame:

Outside in: availability tests in Application Insights hitting the real URL from several locations.

Symptom alerts: server errors, response time and failed test locations, not only CPU.

Reach someone: action groups that page the on-call person, and a test of the whole chain.

Learn: a short review of why it went down and why nothing fired.

Sample spoken answer:

“The gap is that we were only watching the inside of the system, if anything. First, I'd add availability tests in Application Insights that request the home page and one key page every few minutes from several regions, and alert when more than one location fails, so one bad probe doesn't page anyone. Second, alerts on symptoms users feel: a spike in 5xx responses from the App Service and response time going up. CPU alerts alone miss most outages. Third, the alerts have to reach a person, so I'd set up an action group that pages whoever is on call, not a shared inbox nobody reads at night, and then trigger a test alert to prove the whole chain works. I'd also run a short blameless review of last night: what broke, and why nothing noticed.”

Red flag to avoid:

Adding a CPU alert and calling it done, or sending alerts to an email list nobody watches out of hours.

They may ask next:
  • How would you stop this from turning into so many alerts that people ignore them?
  • What would a Resource Health alert tell you that your own tests wouldn't?
Say it in 60 seconds

Databases 3 questions

Medium Technical round Mid-level, Senior Practice question

16. A report query on Azure SQL Database took two seconds last week and takes forty now. No code was deployed. Where do you start?

What the interviewer is really testing:
Whether you separate a plan change from a resource limit or blocking, and use Query Store to prove which one it is.
Answer frame:

Resources: is the database hitting its CPU, data IO or log limits, or waiting on locks?

Query Store: compare this query's plans over time; a new plan often means a regression.

Why it flipped: changed statistics, data growth or a parameter value that suited one caller.

Fix: force the last good plan now, then fix the root cause with an index or query change.

Sample spoken answer:

“No deploy doesn't mean nothing changed: data grows, statistics update, and the optimiser can pick a new plan. First I'd check resource use for the database. If CPU or IO is pinned at the limit, the query is waiting for resources, not slow on its own. I'd also check for blocking. If resources look fine, I'd open Query Store, find the query, and look at its plans over time. Very often there's a new plan from a few days ago, maybe a scan where it used to seek, compiled for a parameter value that suited a different caller. The quick fix is to force the old plan from Query Store, or let automatic tuning do it. Then I'd fix the real cause, like a missing index or a query that behaves badly with skewed data, so we're not relying on a forced plan forever.”

Red flag to avoid:

Scaling the database up straight away without checking whether the plan changed.

They may ask next:
  • What are the risks of leaving a forced plan in place for months?
  • How would you tell parameter sniffing apart from plain data growth?
Say it in 60 seconds
Medium Case round Mid-level, Senior Practice question

17. During a DR test you failed the Azure SQL failover group over to the second region. The database is healthy there, but the app is still throwing login and connection errors. What do you check?

What the interviewer is really testing:
Whether you know the pieces a failover group doesn't move for you: the listener name, logins, firewall rules and networking.
Answer frame:

Connection string: it must use the failover group listener, not a server name.

Logins: server logins must exist on the second server with the same SIDs, or use contained or Entra users.

Network rules: firewall rules and private endpoints are per server and must exist in both regions.

Make it routine: fix it in code and rerun the test until failover is a non-event.

Sample spoken answer:

“A failover group moves the data, not everything around it. First I'd check the connection string. If the app points at the primary server's own name, it's still talking to the old server. It should use the failover group's listener name, which always points at whichever side is primary. Next, logins. If the app uses a SQL login created on the server, that login lives in the primary's master database, and it has to exist on the second server with the same SID, or the database users end up orphaned. Contained users or Entra authentication avoid that. Then networking: firewall rules and private endpoints belong to each server, so the second one needs its own, plus DNS that works in that region. I'd write down every gap, fix it in code, and run the test again.”

Red flag to avoid:

Calling the DR test passed because the database came up, without checking the app could use it.

They may ask next:
  • What's the difference between the read-write and the read-only listener, and when would you use the second?
  • How would you decide between automatic and manual failover for this database?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

18. Every evening at peak, your Cosmos DB container returns lots of 429 errors, yet the portal says you're using well under the throughput you provisioned. What's going on?

What the interviewer is really testing:
Whether you know throughput is split across physical partitions, so one hot partition throttles while the total looks fine.
Answer frame:

Why: provisioned throughput is spread evenly across physical partitions; one busy partition hits its share.

Prove it: normalized RU consumption by partition key range, and diagnostic logs showing the hot key values.

Also check: expensive queries that fan out across partitions, and heavy indexing on write-heavy data.

Fix: a better key means a new container and a migration; tune queries and indexing meanwhile.

Sample spoken answer:

“The total can look fine while one partition is overloaded. Cosmos DB spreads the container's throughput evenly across its physical partitions, so if most of the evening traffic lands on one key, that partition runs out of its share and returns 429s while the others sit idle. I'd check the normalized RU consumption metric split by partition key range. If one range is near the top and the rest are low, that's the hot partition. The diagnostic logs then show which key values are doing it. I'd also look for queries without the partition key in the filter, which fan out across partitions, and an indexing policy indexing fields nobody queries, which makes writes more expensive. The lasting fix is usually a better partition key, which means a new container and moving the data. Raising throughput alone mostly pays for idle partitions.”

Red flag to avoid:

Doubling the provisioned throughput and calling it fixed without checking how the load is spread.

They may ask next:
  • How would you move the data to a new container without downtime?
  • Is a small number of 429s always a problem, and how does the SDK handle them?
Say it in 60 seconds

Governance and Cost 3 questions

Medium Technical round Mid-level, Senior Practice question

19. A pipeline ran your Bicep deployment and a storage account that isn't in the template got deleted from the resource group. How did that happen, and how do you stop it happening again?

What the interviewer is really testing:
Whether you know Incremental and Complete deployment modes, and use what-if and locks as safety nets.
Answer frame:

Cause: Complete mode deletes anything in the group that the template doesn't list.

Now: stop the pipeline, try to recover the account, tell whoever owned it.

Prevent: Incremental by default, what-if output reviewed before apply, delete locks on production.

Longer term: decide who owns each resource; bring the stray one into code or move it out.

Sample spoken answer:

“That's almost certainly a deployment in Complete mode. The default, Incremental, only adds or updates what's in the template. Complete makes the resource group match the template exactly, so anything not listed gets deleted, including a storage account someone created by hand. First I'd pause the pipeline so it doesn't run again, check whether the account can be recovered, which is possible for a short time if the name hasn't been reused, and tell whoever owned it. Then I'd check the deployment history to confirm the mode. To prevent it, I'd switch to Incremental unless there's a strong reason, add a what-if step that shows every change before apply, and put delete locks on production resource groups. And the stray account needs an owner: either it goes into our template, or it moves to its own group.”

Code:
# Preview every create, change and delete before applying
az deployment group what-if --resource-group rg-app-prod --template-file main.bicep --mode Incremental
az deployment group create --resource-group rg-app-prod --template-file main.bicep --mode Incremental
Red flag to avoid:

Blaming whoever created the account by hand without seeing that the pipeline was set to delete anything it didn't know.

They may ask next:
  • Is there ever a good reason to use Complete mode?
  • How would you handle someone changing a resource in the portal that your template also manages?
Say it in 60 seconds
Medium Case round Mid-level, Senior Practice question

20. The Log Analytics part of your monthly bill doubled right after a release. Nobody added new resources. How do you find the cause and bring it down?

What the interviewer is really testing:
Whether you can find which data is driving ingestion and cut it at the source without going blind.
Answer frame:

Find it: query the Usage table for billable volume by data type, before and after the release.

Source: usually debug logging left on, a chatty dependency, or a diagnostic setting sending every category.

Cut at the source: log level back down, sampling in Application Insights, only needed categories.

Guardrails: cheaper table plans for bulky logs, a budget alert, and a daily cap only as a last resort.

Sample spoken answer:

“A jump right after a release points at what the app is sending, not new resources. First I'd query the Usage table in the workspace to see billable volume by data type per day, which usually shows one table jumping on release day, often traces or dependencies from Application Insights. Then I'd find why: a log level left on debug, a new retry loop writing an error on every attempt, or a diagnostic setting someone switched to send every category. The fix is at the source: log level back to warning in production, sampling turned on for high-volume telemetry, and only the diagnostic categories we actually query. For bulky logs we keep but rarely search, a cheaper table plan helps. I'd add a budget alert on the workspace and add a volume check to the release checklist. A daily cap is a last resort, because it stops collecting right when you might need the logs.”

Code:
Usage
| where TimeGenerated > ago(30d)
| where IsBillable == true
| summarize IngestedGB = sum(Quantity) / 1000 by DataType, bin(TimeGenerated, 1d)
| sort by IngestedGB desc
Red flag to avoid:

Setting a tight daily cap straight away, so logs stop arriving in the middle of the next incident.

They may ask next:
  • What would you lose if you turned sampling on, and how would you limit that?
  • Who should be able to change log levels in production, and how?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

21. A security review finds thirty VMs across several subscriptions with public IPs and RDP open to the whole internet. The owners say they need to reach them. How do you handle it?

What the interviewer is really testing:
Whether you fix the risk without blocking people's work, then use platform guardrails so it can't come back.
Answer frame:

Triage: check sign-in logs on the exposed machines; any suspicious access becomes an incident.

Replace access: Bastion or just-in-time access, so people still get in without an open port.

Close it: remove the public IPs and the open rules, owner by owner, with a date.

Guardrail: policy at the management group that denies public IPs on VMs, with exemptions reviewed.

Sample spoken answer:

“RDP open to the internet gets hammered by password guessing all day, so I'd treat it as urgent but not panic. First I'd check the security logs on those machines for successful sign-ins from strange places, and if I find any, that VM becomes an incident. Then I'd give the owners a better way in before taking the old one away: Azure Bastion, so they connect through the portal over HTTPS without any public IP on the VM, or just-in-time access that opens the port to their own address for a few hours. With that ready, I'd remove the public IPs and the open rules subscription by subscription, with a date agreed with each owner. To stop it coming back, I'd assign a policy at the management group that denies public IPs on VM network cards, with a reviewed exemption process for real exceptions.”

Red flag to avoid:

Deleting the public IPs on day one without giving anyone another way in, so teams find a worse workaround.

They may ask next:
  • How would you handle a team that says they truly need a VM reachable from the internet?
  • Where would you see this kind of finding without waiting for the next manual review?
Say it in 60 seconds

Containers 2 questions

Medium Technical round Mid-level, Senior Practice question

22. New pods in your AKS cluster are stuck in ImagePullBackOff when pulling from your own Azure Container Registry. Older pods on the same image run fine. What do you check?

What the interviewer is really testing:
Whether you read the pod events and know AKS pulls with the kubelet identity, which needs AcrPull on the registry.
Answer frame:

Read the event: describe the pod; the message says not found, unauthorized or a timeout.

Not found: a wrong tag or registry name in the manifest.

Unauthorized: the kubelet identity needs AcrPull; attaching the registry sets it up.

Timeout: a registry firewall or private endpoint the nodes can't reach.

Sample spoken answer:

“Older pods working often just means their image was already cached on those nodes, so I wouldn't read too much into it. First I'd describe one of the stuck pods and read the event message, because it tells me which of three problems I have. If it says not found, it's the tag or the registry name, often a pipeline that pushed a tag nobody deployed. If it says unauthorized, the cluster's kubelet identity probably doesn't have AcrPull on the registry, maybe because the registry was recreated or the role assignment was removed. Attaching the registry to the cluster puts that role back, and the check command confirms the nodes can pull. If it's a timeout, I'd look at networking: the registry may now only allow private access, and the nodes can't reach its private endpoint or resolve its name.”

Code:
kubectl describe pod orders-api-7d9f8b6c4-x2k9p -n orders
az aks check-acr --name aks-prod --resource-group rg-prod --acr myregistry.azurecr.io
az aks update --name aks-prod --resource-group rg-prod --attach-acr myregistry
Red flag to avoid:

Turning on the registry's admin user and pasting its password into a secret as the fix.

They may ask next:
  • Why is pulling with the kubelet identity better than an image pull secret with a registry password?
  • How would you stop someone deploying an image with a mutable tag like latest?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

23. Traffic spiked, the HPA asked for more replicas, and the new pods have been Pending for twenty minutes. The cluster autoscaler is on. What could be stopping them?

What the interviewer is really testing:
Whether you can go from the scheduler's message to the real limit: node pool size, pod requests, subnet IPs or subscription quota.
Answer frame:

Scheduler first: describe a pending pod; the event says insufficient CPU, memory, or no matching node.

Autoscaler limits: the node pool may be at its maximum count, or requests too large for any node size.

Azure limits: vCPU quota for the region, or no IPs left in the subnet with classic Azure CNI.

Fix and prevent: raise the limit that's hit; plan headroom, overlay networking or a bigger subnet.

Sample spoken answer:

“I'd start with a pending pod and read its events, because the scheduler says why: insufficient CPU or memory, or no node matching a selector or taint. If it's insufficient resources, the question becomes why the autoscaler didn't add nodes. I'd check its status and events. Common reasons: the node pool is already at its maximum count, or the pods request more than any single node of that size can offer, so a new node wouldn't help. Then the Azure-side limits. The subscription may have hit its vCPU quota for that VM family in the region, which shows up as failed scale operations in the activity log. And with classic Azure CNI, every node reserves IPs for its pods up front, so a small subnet runs out and new nodes can't join. Short term I'd raise whichever limit is hit; long term, headroom and overlay networking.”

Red flag to avoid:

Deleting and recreating the pods over and over, hoping they'll schedule the next time.

They may ask next:
  • How would you keep critical pods running when the cluster is short of capacity?
  • What does it take to move an existing cluster to a bigger subnet or overlay networking?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card