Scenario rounds give you a situation and watch how you think. A customer named O'Brien can't save their profile, an app update wipes saved settings, the product owner lowers the priority of your bug, a manager compares your bug count with a colleague's. There is rarely one right answer. The interviewer wants to hear what you check first, what you would ask, and what would change your mind. This page is for testers at every level, from a first QA job to a test lead. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying your first two checks before the fix.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Suspect: the apostrophe is breaking how the name is saved, often because the input is glued straight into a database query.
Confirm: try other names with an apostrophe, then the same character in other fields such as address and search.
Widen: test hyphens, accents, spaces in a surname, very long names and names in other scripts.
Report: flag it as a possible security issue, since input that breaks a query may also be able to change it.
“The apostrophe is the obvious suspect, since that's what makes O'Brien different. In SQL an apostrophe marks the start or end of a piece of text, so if the app builds its query by gluing the name straight in, the apostrophe breaks the query and the save fails. To confirm, I'd try a couple of other names with an apostrophe, then put one into other fields, like the address, the search box and a comment field. If those break too, it's a pattern, not a one-off. Then I'd widen to other real names: hyphens, accents, a space in the surname, very long names, names in other scripts. In the report I'd say it may be more than a bug. If an apostrophe can break a query, someone might be able to change what the query does, so I'd ask for a security review too.”
Telling the customer to save the name without the apostrophe and closing it as a user error.
Suspect: the prices are being sorted as text, so it compares character by character and 1 comes before 2.
Confirm: try prices with different digit counts, such as 9, 20, 100 and 1000, and decimals.
Spread: check other numeric sorts and filters, like ratings, discounts and a price range filter.
Report: list the exact items, the order shown and the order expected.
“My guess would be that the price is being sorted as text rather than as a number. As text, 100 comes before 20 because it compares the first character, and 1 is smaller than 2. To check, I'd look for products priced at different lengths, say 9, 20, 150 and 1200, and see if the order follows the first digit. I'd also try prices with decimals and items on sale, because the sort might use the old price. Then I'd widen it: if price sorts wrongly, maybe rating, discount or newest do too, and the price range filter might have the same problem. In the bug report I'd list the items, the order I got and the order I expected, and mention my text-sort theory as a hint, not a fact.”
Logging it as a single cosmetic issue without checking whether other sorts and filters share the same cause.
What's different: probably the amount and age of their data; ask support how many orders and how old the oldest is.
Build a heavy account: create a test user with hundreds or thousands of orders, including very old ones.
Watch the request: in the browser's network tab, see whether the call times out, errors or returns everything at once.
Look at old data: an order from years ago may have an empty field or a retired product that breaks the page.
“My test accounts have maybe ten orders each, and a long-time customer might have hundreds, so I'd suspect the size or the age of their data. I'd ask support how many orders they have and how far back they go. Then I'd build something similar, probably by asking a developer for a script that creates a test user with a thousand orders, some dated years back. While that page loads, I'd watch the network tab in the browser tools. If the request runs for a minute and then fails, the page is probably loading everything at once instead of a page at a time. If it's fast with a thousand fresh orders, I'd look at old data instead. One ancient order with a product that no longer exists, or a field that was empty back then, can break the whole list. That would change the fix completely.”
Closing it as cannot reproduce because every test account loads fine.
Close the original: the cart fix passes, so close that bug; don't reopen it for a different symptom.
Log the regression: new bug, linked to the fix, with the build where it last worked.
Find the link: ask the developer what the fix changed; totals and cart often share price or tax logic.
Widen: test everything else that uses that logic, like invoices, emails and refunds.
“I'd treat them as two separate things. The cart bug is fixed, so I'd close it, because reopening it for a different symptom muddles the history. Then I'd log the order history problem as a new bug, mark it as a regression, and link it to the fix. I'd note the build where totals were last correct, since that narrows down what changed. Next I'd talk to the developer about what the fix touched. If they changed how totals, tax or discounts are calculated, that shared logic probably feeds order history too, and also invoices, confirmation emails and refunds. So I'd widen my testing to all of those before signing off. It also tells me our regression pack is missing something, and I'd add a test for order totals after this.”
Reopening the cart bug for an unrelated page, or logging the new bug with no link to the fix that likely caused it.
Find the pattern: which orders, which email providers, what time, which payment method.
Follow the email: was it created, was it sent, was it delivered, did it land in spam?
Check the rules: test environments often send to a catch-all inbox, so real delivery is never tested there.
Test the edges: long or unusual addresses, special characters, many orders at once.
“'Sometimes' usually means there's a pattern we haven't found yet. I'd ask support for a list of affected orders and look for what they share: one email provider, a certain payment method, orders at busy times, or guest checkout rather than logged in. Then I'd follow one email through its journey with the developers. Was the email created by the app? Was it handed to the sending service? Did that service deliver it, or did it bounce or land in spam? The sending service usually has a log that answers this. I'd also point out why test always works: test environments often send everything to one catch-all inbox, so real delivery never gets tested. Then I'd try the edges, like addresses with a plus sign or unusual domains, and many orders placed at once.”
Saying it works in test so it must be the customer's spam folder, without checking any logs.
Probe: after pressing back, refresh the page and click a link or button that needs the server.
Cached view: if every action sends you to login, the session is gone and only the old page was shown from the browser's memory.
Live session: if actions still work or fresh data loads, logout isn't really ending the session.
Report: either way raise it; a cached banking page on a shared computer is still a privacy problem.
“The first thing I'd find out is whether it's just a picture of the old page or a live session. So after pressing back, I'd refresh, and then click something that needs the server, like opening a statement. If everything throws me to the login page, the session really ended and the browser is only showing a cached copy. That's still worth a bug for a banking app, because someone on a shared computer could see a balance, and the usual fix is telling the browser not to store those pages. If refreshing loads fresh data, or I can still make a transfer, that's far worse: logout isn't killing the session on the server, and I'd raise it as critical. I'd also try the same thing after the session times out on its own.”
Calling it cosmetic without refreshing or clicking anything to see whether the session is still alive.
Stop: don't look through more records than needed; one example proves it.
Prove it cleanly: reproduce with two test accounts you own, so the evidence holds no real customer data.
Report: critical severity, straight to the lead and security contact, in a restricted ticket.
Widen: check other pages with IDs in the address or in requests, like invoices, addresses and returns.
“This is a serious security bug, so the first thing I'd do is stop. One example proves the problem, and I don't want to look at more real customer data than I already have. Then I'd reproduce it cleanly with two test accounts I control: place an order as user A, log in as user B, and open A's order by its number. That gives proof without any real customer in the screenshots. I'd report it as critical, tell my lead and whoever handles security straight away, and keep the ticket restricted rather than posting it in a public channel. After that I'd check the same pattern elsewhere: invoices, saved addresses, returns, anything with an ID in the address or the request. If this is live on production, I'd say so clearly, because that changes the urgency.”
Clicking through dozens of real customers' orders to see how far it goes, or posting the screenshots in a team chat.
Raise it: log the blocker as critical with steps, and tell the lead and developers right away.
Find what's reachable: list the areas that don't depend on the dashboard and test those.
Mark the rest: record blocked test cases as blocked, not failed or skipped.
Plan the retest: once a fixed build arrives, rerun the smoke test before anything else.
“I'd go along with it, but I'd make sure everyone sees what's blocked. First I'd log the dashboard failure as a critical blocker with the steps and a screenshot, and message the developers so the fix starts now rather than tomorrow. Then I'd work out what I can genuinely test without the dashboard. Maybe settings, profile pages or reports open from a direct link. I'd test those, but I'd tell the lead that anything touching the dashboard is blocked, and I'd mark those cases as blocked in the tracker so the numbers stay honest. When the fixed build comes, I'd run the smoke test again first, because the fix could break something else. If the blocker covered most of the app, I'd say clearly that the build isn't testable and suggest rejecting it.”
Quietly testing around the problem and reporting a normal pass rate at the end of the day.
Ask: browser and version, device, whether it happens every time, and a screenshot or short recording.
Match their setup: same browser and version, same screen size, and a normal browser with extensions on.
Look for a clue: open the developer tools console and click the button to see if a script error appears.
Common causes: an older browser missing something the page uses, an ad or privacy blocker, or a banner covering the button.
“Since it works for all of us, the difference is almost certainly on the customer's side, so I'd start with questions. Which browser and version, which device, does it happen every time, and can they send a screenshot or a short recording? Then I'd copy their setup as closely as I can: the same browser and version, the same screen size. Our team usually tests in a clean browser, but customers have extensions, so I'd also try with an ad blocker or privacy extension switched on, since those sometimes block scripts the button needs. Once I can reproduce it, I'd open the developer tools console and click the button, because a script error there usually tells the developer exactly where to look. I'd also check whether a cookie banner or chat widget is covering the button on small screens.”
Replying that it works on our side and asking the customer to try a different browser.
Find out why: is another tester using the same accounts, or is a job or data refresh running?
Own your data: create your own users and records, named so others know they're yours.
Agree rules: shared accounts listed, refresh times announced, destructive tests planned.
Suggest more: a data setup script, or separate environments if the conflicts keep costing time.
“First I'd find out what's actually changing my data. Usually it's one of two things: someone else is using the same test account, or there's a scheduled refresh or batch job nobody told us about. I'd ask the other testers and the developers. Then I'd fix my own part: create my own test users and orders, with names that make it obvious they're mine, like my initials in the name. After that I'd suggest a few simple rules for the team. A shared list of who uses which accounts, a heads-up in the team chat before anyone deletes data or runs a big import, and fixed times for environment refreshes. If it keeps wasting hours, I'd raise it with the lead with an estimate of the time lost, because that's what gets a separate environment approved.”
Blaming the other testers, or silently rerunning failed tests until they happen to pass.
Understand it: get the API description: what it accepts, what it returns and which errors it should give.
Send requests: use an API client, starting with a working example from the developer to prove your setup.
Same techniques: boundaries, missing fields, wrong types, and calls with no login or as the wrong user.
Check more than the status: read the response body and confirm the data really changed where it should.
“The good news is that the thinking doesn't change, only the tool. First I'd ask for the API description, so I know what a request should look like, what comes back, and which errors it should return. Then I'd use an API client to send requests, starting with a working example from the developer so I know my setup is right. After that, it's the same techniques I'd use on a form: values on the boundaries, missing required fields, text where a number should be, a very long value. I'd also call it without logging in, and as a user who shouldn't have access, because there's no screen hiding a button from them. For each call, I'd check the status code, read the response body, and then confirm the data really changed, through another call or a quick database check.”
Saying you'll wait until the screen is ready because manual testers don't test APIs.
Ask for a lever: a config setting in the test environment to shorten the expiry, say to two minutes.
Server time: the check runs on the server, so changing my own laptop clock proves nothing.
Boundaries: use the link just before and just after the limit, at least once with the real 30 minutes.
Other rules: link used twice, older link after a newer one is sent, and link used after the password changes.
“I'd start by asking the developers if the expiry time is a setting. In most apps it is, so in the test environment we can drop it to two minutes and run the checks quickly. I'd be careful about one thing: the expiry is checked on the server, so changing my own laptop's clock does nothing useful. With the short setting, I'd open the link just before it expires, then just after. I'd still run one test with the real 30 minutes before release, because the production setting could be wrong. Then there's the logic around it: can the same link be used twice, does an older link still work after I request a new one, and does the link die once the password is changed. Those are usually where the real bugs hide.”
Changing the laptop's clock and believing that tests a server-side expiry.
The gap: testers usually install fresh builds, so the update path with old saved data never got run.
Set it up: install the version users have now, create real settings and data, then update without uninstalling.
Check: settings, login, saved items and offline data all survive the update.
Cover the spread: update from more than one older version, since many users skip updates.
“My guess is that we only ever tested fresh installs. Testers tend to delete the app and install the new build, so everything starts clean, but real users update over the top with months of saved data. If the new version changed how settings are stored, the old ones can get lost on the way. To test it properly, I'd install the version that's live now, use it like a real person, change settings, log in and save a few things, then update to the new build without uninstalling. After that I'd check every setting, whether I'm still logged in, and any saved or offline data. I'd also try updating from an older version, not just the last one, because plenty of users skip a few updates. And I'd add an update test to every release checklist from now on.”
Saying new installs pass, so the users must have cleared their data themselves.
Before tonight: agree the checklist, the test accounts safe to use in production, and who decides on rollback.
First: confirm the right version is live, then verify the exact bug is fixed.
Second: the money path end to end: search, cart, checkout, payment, confirmation.
Last: a quick look at nearby features and error logs, then report go or rollback with what was and wasn't covered.
“I'd do most of the thinking before the hour starts. During the day I'd write a short checklist, agree which test accounts and payment methods are safe to use on production, and confirm who makes the rollback call. On the night, the first check is that the new version is actually live. Then I'd verify the bug itself using the exact steps from the ticket. Next comes the path that makes money: log in, search, add to cart, check out, pay, and get the confirmation email. If anything there breaks, it matters more than the original bug. With whatever time is left, I'd look at features near the change and ask the developers to watch the error logs. At the end I'd give a clear message: go or roll back, plus what I did and didn't cover.”
Only checking the fixed bug and ignoring the checkout path around it.
Counts and totals: records per type, and sums like balances or order values, must match between old and new.
Samples: compare a set of records field by field, chosen at random and on purpose.
Edge cases: special characters, empty fields, very long values, very old records, closed accounts.
Use it: run real journeys on migrated data, like an old customer logging in and viewing history.
“Nobody can eyeball millions of records, so I'd layer the checks. First, counts and totals: the number of customers, orders and accounts should match between old and new, and so should sums like total balances. If the totals differ, I know there's a problem before I look at a single record. Second, sampling. I'd pick records at random and compare every field, but I'd also pick awkward ones on purpose: names with accents or apostrophes, empty phone numbers, the longest addresses, the oldest accounts, closed customers. Third, I'd use the data the way customers will, logging in as migrated users, viewing old orders and making a change. I'd push to run all this on a rehearsal migration first, so problems turn up weeks before the real move, not on the night.”
-- Run on both systems and compare
SELECT COUNT(*) AS customers,
SUM(balance) AS total_balance
FROM customers;
Opening a few customers in the new system, seeing they look fine, and signing off.
Don't decide alone: the tester's job is to spot the conflict, not settle the business rule.
Take it to the owner: usually the product owner or business analyst, with the three versions side by side.
Record it: get the answer written in the story, and ask for the other document to be fixed.
Meanwhile: test everything that doesn't depend on the answer, and flag the story as at risk.
“Honestly, none of them yet. When three sources disagree, it isn't my call which is right, because that's a business rule with real money behind it. I'd put the three versions side by side, with a worked example showing what the customer pays under each, and take it to the product owner. A concrete example gets a quick answer much faster than a vague question. Once they decide, I'd make sure the decision is written into the story, and that the design or the requirement gets updated so the next person isn't confused. While I wait, I'd test everything around it that doesn't depend on the answer, like invalid codes, expired codes and removing a code. And I'd tell the team this story is at risk until we hear back.”
Testing against whatever the developer built because it's already there.
Make it measurable: agree a target, like the report opens within three seconds for a year of data.
Pin the conditions: how much data, how many users at once, which network and device.
Measure: time it consistently, repeat several runs, and test with the biggest realistic data set.
Escalate if needed: if it needs real load testing, say so and involve whoever does performance testing.
“I can't pass or fail 'fast', so the first job is turning it into a number. I'd ask the product owner something like: with a full year of sales data, how many seconds is acceptable? And under what conditions, one user or fifty at once, office network or mobile? Once we agree on a target, say three seconds for a year's data, I'd write it into the acceptance criteria. Then I'd test it properly: time it several times rather than once, with realistic data volumes, and with the heaviest filters people actually use. I'd note that a test environment may be slower or faster than production. If it turns out the real question is how it behaves with many users at once, I'd say that needs a proper load test and bring in whoever does performance testing.”
Opening the report once, deciding it felt quick, and marking the story done.
Pull out the serious ones: anything that blocks the user or loses data gets its own bug straight away.
Group the rest: one ticket for the form, with a numbered list and a screenshot for each item.
Check the source: compare with the design; some items may be design mistakes, not code.
Talk first: a quick word with the developer about the batch beats 15 notifications.
“First I'd sort them. If any one of them actually stops a user, like a field that won't accept valid input or an error message that points at the wrong field, that gets its own bug straight away. The truly small things, typos, spacing, alignment, I'd group into one ticket for that form, with a numbered list and a screenshot per item, so the developer can fix them in one sitting. Fifteen separate tickets would bury anything important and annoy everyone. Before logging, I'd check the design, because some of these might be built exactly as designed and the design is what's wrong, so those go to the designer instead. Then I'd have a quick word with the developer so the batch isn't a surprise, and ask the lead where it fits in the priorities.”
Logging 15 separate low-priority tickets in a row, or not logging the small ones at all because the developer is busy.
Understand why: ask what the decision is based on; they may know something you don't.
Add evidence: how often it happens, who it affects, and what it costs the customer, in plain words.
Accept and record: if the decision stands, note the known risk in the release notes or sign-off.
Follow up: suggest a check after release, like watching support tickets about addresses.
“First I'd ask why, because the product owner might know something I don't, like only a tiny number of users hit that path. If I still think it matters, I'd come back with evidence rather than opinion. How often did I reproduce it? Which users does it hit? What happens to them, maybe a failed delivery or a support call. I'd put that in plain words, not testing terms. If they still decide to ship, that's their call to make, and I'd respect it. But I'd make sure it's written down as a known issue in the sign-off, so nobody is surprised later. I'd also suggest a quick check after release, like watching support tickets about missing addresses, so we know early if it's worse than expected.”
Arguing until they give in, or quietly dropping it with nothing recorded.
Direction: Arabic reads right to left, so the layout, alignment and arrow icons should mirror.
Length: German words are long; look for cut-off buttons, overflowing labels and wrapped menus.
Missing text: find strings still in English, often in error messages, emails and pop-ups.
Formats: dates, numbers, decimal separators and sorting that differ by region.
“I'd split it by what usually breaks. Arabic is read right to left, so the whole layout should mirror: menus, alignment, form labels, and icons like a back arrow should point the other way. I'd also check screens that mix Arabic with numbers or English product names, because that's where text gets jumbled. German is the opposite problem: long words. I'd hunt for buttons where the text is cut off, labels that overflow their boxes and menus that wrap badly. For both, I'd look for leftover English, which usually hides in error messages, emails, pop-ups and anything added late. Then regional formats: dates, decimal commas in German, and alphabetical sorting. If there's time before real translations arrive, I'd ask for a test build with fake stretched text so layout problems show up early.”
Planning to just switch the language and click through the main screens once.
Ask: which screen reader, browser and device they use, and where exactly they get stuck.
Keyboard only: go through checkout with Tab, Enter and arrow keys; watch focus order and traps.
Screen reader: turn one on and listen: are fields, buttons and errors announced with clear names?
Common culprits: unlabelled fields, icon-only buttons, custom drop-downs, pop-ups that don't take focus, errors shown only in colour.
“I'd start by asking the customer which screen reader and browser they use and where they get stuck, because that saves hours. Then I'd try checkout myself using only the keyboard. Can I reach every field and button with Tab, is the order sensible, can I always see where focus is, and do I ever get trapped, say inside a pop-up? After that I'd turn on a screen reader, like the one built into my phone or computer, and go through it by listening. The usual problems are fields with no label, so it just says edit text, icon buttons with no name, custom drop-downs the keyboard can't open, and error messages that appear on screen but are never read out. I'd report each one with where it happens and what the screen reader actually said.”
Running one automatic scan, seeing few warnings, and replying that checkout is accessible.
Stay open: first check honestly whether you missed things, for example bugs found later in your area.
Give context: module size, maturity, risk and how early you were involved all change the count.
Show the value: severity of what you found, coverage, and bugs prevented in reviews.
Offer better measures: bugs that escaped to production, requirement coverage, critical bugs found early.
“I wouldn't get defensive, because it's a fair question. First I'd check myself honestly: were bugs found later in my area that I should have caught? If yes, I'd own that. If not, I'd give context. Maybe my colleague tested a brand new module and I tested a small change to a stable one. Maybe I reviewed the stories early and several problems were fixed before any code existed, so they never became tickets. I'd also point to what my five bugs were, if two of them were critical, and to my coverage. Then I'd gently suggest that count alone rewards logging lots of small issues. Better signals might be how many bugs escaped to production from each area, and how early the serious ones were found. I'd offer to pair with my colleague and compare approaches.”
Criticising the colleague's bugs as trivial, or promising to log more bugs next sprint.
Learn fast: read the stories, look at the designs, spend half an hour with the developer or analyst.
Break it down: list areas and test types: functional, integration, regression, data setup, retest time.
Compare: use a similar past module as a reality check.
Give a range: best and likely case, with assumptions and risks written next to the number.
“I'd use the morning to understand it rather than guess. I'd read the stories and designs and grab half an hour with the developer or analyst to ask what's new, what's risky and what it connects to. Then I'd break it down: the main flows, the integrations, the test data I'll need, regression on nearby features, and time for retesting fixes, which people always forget. I'd put a rough number on each piece and add them up. Next I'd sanity check it against a similar module we tested before. Finally I'd give my lead a range, say six to eight days, and list the assumptions: builds arrive on time, the environment is stable, requirements don't change. Writing the assumptions down matters, because when one breaks, the estimate moving is expected, not a failure.”
Throwing out a number from gut feeling with no breakdown or assumptions.
Don't rewrite first: three weeks isn't enough to fix 800 cases; protect the release.
Map the risk: agree with the team on the critical flows and what's changing this release.
Use what works: pick and update the cases covering those areas; use exploratory sessions for the gaps.
Clean as you go: tag each case as current, outdated or delete; retire duplicates after the release.
“My first instinct would be to not touch most of the 800 cases yet. Three weeks isn't enough to rewrite them, and the release matters more. So I'd sit with the product owner and developers and agree two lists: the flows that must never break, and what's changing in this release. Then I'd search the suite for cases covering just those areas, update them as I run them, and fill gaps with short exploratory sessions that I write notes on. Everything I touch gets tagged as current, and anything clearly dead gets marked to delete. After the release, I'd keep going area by area, merging duplicates and retiring cases for features that no longer exist. Within a few cycles the suite is smaller but trusted, which is worth more than a big number nobody runs.”
Spending the three weeks rewriting test cases and doing little actual testing before the release.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.