Manual testing interviews for 3 years of experience skip the textbook definitions and ask what you did: a bug you traced past the screen, how you used dev tools and SQL, how you estimated a story and where you got it wrong, and how you handled payments, time zones or double clicks. It is written for manual testers with about two to four years of real project work, the stage where you own the testing of a feature from story to release but someone else still sets the test strategy. Each question shows what the interviewer is checking, the shape of a strong answer and a short spoken answer. Swap the stories for your own before the interview.
Search all questions by round, difficulty and level, or save the ones you want to practice.
The symptom: what the user saw and why it wasn't enough to go on.
What you checked: network calls, console errors, logs or database rows, in the order you used them.
The handover: what you put in the bug and how much time it saved the developer.
“On my last project, users said their saved address sometimes vanished from checkout. On screen it just looked empty. I opened the network tab and saw the address call came back fine, with the data in it, so the back end wasn't losing it. Then I checked the console and found an error only when the address had an apartment number with a slash in it. The page was failing to render that address and showing nothing. I confirmed in the database that the row was saved correctly. So my bug said: data is saved, the API returns it, the front end fails to display addresses containing a slash, here's the console error and the exact test data. The developer fixed it the same day, because I'd already told him which layer to look at.”
A story where you only reported what the screen showed and left all the narrowing down to the developer.
Check the request: did the page send the right data?
Check the response: did the server return the right data and status?
Compare with the screen: right data but wrong display points to the front end.
“I look at the call behind the action in the network tab. First the request: did the page send what I entered? If the request is already wrong, say the date went in the wrong format, it's a front-end bug. If the request is right, I look at the response. A wrong value, a wrong status code or a server error points to the back end. If the response is correct but the screen shows something different, that's the front end again, usually rendering or state. It matters because on my teams those were different people, sometimes different sprints. A bug assigned to the wrong developer bounces around for a day or two. So I write the layer in the bug and attach the request and response, and it goes straight to the person who can fix it.”
Saying it's the developer's job to figure out the layer, or guessing from the screen alone.
Same build: confirm the version deployed to QA matches what the developer runs.
Differences: config, data, browser, user role, feature flags.
Reproduce together: a short call, same steps, both screens.
“First I check we're on the same code. Many times the QA environment simply hadn't got the latest build, so I look at the build number or the deployment note. If the build matches, I list what's different: the data, because QA data is messier than a developer's local data; the user role; config or feature flags; and the browser. Then I ask for ten minutes to run the steps together, him on his machine and me on QA. On one project it turned out a feature flag was on locally and off in QA. That wasn't a code bug, but it was a real release risk, because production had the flag off too. So I raised it as a config issue rather than closing it. Works on my machine is where the investigation starts, not where it ends.”
Closing the bug because the developer couldn't reproduce it, or escalating straight to the manager.
Network: status codes, request and response bodies, slow calls, throttling to a slow connection.
Console and storage: script errors, cookies, local storage and session values.
Device mode: quick checks of layouts at phone widths, backed up on a real device.
“The network tab is the one I use most. I check the status code and the response body when something looks wrong, and I throttle the connection to a slow network to see how the page behaves while it waits, which is where I've found double submits and missing loading states. The console I keep open all the time, because a red error there often explains a blank section before anyone reports it. In the application tab I look at cookies and local storage, for example to check that logging out really clears the session token. And I use device mode for a first pass at phone layouts, but I always confirm on a real phone, because device mode doesn't catch everything, like touch behaviour or the real keyboard covering a field.”
Only mentioning right-click and inspect, or treating device mode as a full replacement for real devices.
Find the order: filter by your test user and today's date.
Join the items: count them and compare with what you added to the cart.
Check the fields: status, totals and anything the UI never shows.
“I'd pick my test user's orders from today and join to the order items table, so I see the order status, the total and how many items were saved in one result. If I added three items in the cart, I expect three rows behind that order, and the total should match what the checkout page showed. I also look at the columns the UI doesn't show, like the payment status or the created-by field, because those are where quiet bugs hide. On my last project this caught a case where the confirmation page looked perfect, but one item with a discount code was never saved, because the insert for that item failed silently. The screen had simply shown what was in the cart, not what was in the database.”
SELECT o.order_id, o.status, o.total_amount,
COUNT(oi.item_id) AS item_count
FROM orders o
LEFT JOIN order_items oi ON oi.order_id = o.order_id
WHERE o.customer_email = 'qa.user01@example.com'
AND o.created_at >= CURRENT_DATE
GROUP BY o.order_id, o.status, o.total_amount
ORDER BY o.order_id DESC;
Trusting the success message, or writing a query that can't tell an order with no items from one that doesn't exist.
List the states: new, active, locked, expired, with and without history.
Build it: through the UI, an API, a script or seed files, fastest first.
Keep it reusable: named accounts, documented, reset after destructive tests.
“First I list the states I need. For a subscription feature it might be a brand new user, an active monthly one, one on a free trial, one whose payment failed, and one that's cancelled but still has days left. Then I pick the fastest way to create each. Some come quickly through the UI, some through an API call, and some need a developer's help with a script or a seed file, which I ask for early because it takes time. I don't copy production data. It holds real people's details, and even where the rules allow it, it's a privacy risk. I name the accounts clearly, like trial-expired-01, and keep them in a shared sheet so the team reuses them. And after tests that change data, like cancelling, I reset or recreate the account so the next run starts clean.”
Saying you'd just copy customer records from production, or waiting for data to appear by itself.
Break it down: test design, data and setup, execution, retesting, regression around it.
Look for risk: new area, outside services, unclear criteria, many devices.
Say it with assumptions: the number plus what it depends on.
“I break the story into the testing work, not just the running of tests. For a recent story adding a coupon field at checkout, I counted writing the cases, about half a day; setting up data, like valid, expired and single-use coupons, a few hours; running the cases across two browsers and the app; and time for at least one round of retesting, because a first build rarely passes clean. Then I added a slice for regression on checkout, since it touches payments. I also looked at risk: the coupon rules were partly unclear, so I said my estimate assumed the rules would be settled by day two. Saying the assumption out loud mattered. When the rules changed mid-sprint, nobody was surprised that testing took longer.”
Estimating only the time to execute test cases and forgetting setup, retesting and regression.
The estimate: what you said and what it was based on.
What blew it up: the specific thing you didn't count.
The change: the habit you use now, and proof it worked.
“I estimated two days to test a new report export, because the screen itself was simple. It took almost five. What I missed was data. The report had to be checked for customers with lots of records, empty records, special characters and different date ranges, and none of that data existed in QA. I spent two days just building it, partly by asking a developer for scripts. I also hadn't counted opening the files in different spreadsheet programs, where the formatting behaved differently. I told my lead as soon as I saw it slipping, on day two, not on the last day. Now, before I give any number, I ask myself what data I need and whether it exists yet. If it doesn't, data setup becomes its own line in the estimate.”
Blaming the developers or the requirements entirely, or saying your estimates are never wrong.
Before code: questions in refinement, scenarios shared with the developer.
During build: data ready, early builds checked, bugs raised and retested.
Release and after: regression, sign-off, a quick check in production.
“The best example is a feature that let users reschedule appointments. In refinement I asked what happens if the new slot is taken while you're choosing it, and that turned into a new acceptance criterion. Before coding finished, I wrote my scenarios and shared them with the developer, so he tested some of them himself. While it was being built I prepared users with past, future and cancelled appointments. When the first build came, I found four bugs, two around time zones, and retested each fix with the cases nearby. Before release I ran the booking regression pack and wrote a short sign-off with one known low bug. After release I did a quick check in production with a test account and watched the support queue for two days. Nothing came back, which felt good.”
A story that starts when the build arrived and ends when the test cases passed.
The feedback: who gave it and what exactly it said.
Your first reaction: honest, briefly.
The change: what you do now and how you know it helped.
“In my first year on my last team, my lead reviewed my test cases and said they were too tied to the screen. Every step said click this button, type in that box, so any small UI change broke half my cases, and someone new couldn't tell what each case was actually proving. I was a bit defensive at first, because they were detailed. But she was right. Now I write the purpose first, like checks that an expired coupon is rejected, keep the steps at the level of actions, and put exact test data in a separate column. When the checkout screen was redesigned later that year, I had to update very few cases, while an older suite in the same team needed a full rewrite. I still ask for a review on any new suite.”
Saying you've never had useful feedback, or describing feedback without any change that followed.
Automate first: stable, repeated, data-heavy checks on critical paths.
Keep manual: changing screens, look and feel, exploratory work, one-off checks.
Hand over well: clear steps, data and expected results the automation tester can use.
“I'd start with tests we run every release that rarely change: login, the main checkout path, key calculations with lots of data combinations. Those give the most back, because we run them again and again and a person gets bored and misses things by the fiftieth run. I'd keep manual the screens still being redesigned, since the scripts would break every sprint, plus anything about look and feel, usability and exploratory testing. On my last team I went through the regression suite with the automation engineer and marked each case. The big help from me was cleaning the cases first: exact data, one clear expected result each. Automating a vague case just gives you a vague script that nobody trusts when it fails.”
Saying everything should be automated, or seeing automation as a threat and avoiding the question.
Ask what changed: which function, and where else it is called.
Map the callers: screens, reports, jobs and APIs that use it.
Pick the tests: the fix, each caller's main path, and edge values the fix touched.
“I start by asking the developer exactly what changed and where else that code is used. Most developers can tell you in a minute, or show you in the code. Say the fix was in a function that rounds prices, found through a bug on the cart page. That same function might feed the invoice, the order history and a nightly report. So besides retesting the cart, I'd run the main path on each of those screens, and focus on the values the fix touched: amounts that round up, round down and sit exactly on the half. I also look at the pull request myself if I can read it, just the files list, because it often shows changes the bug report never mentioned. Then I note in the ticket what I covered, so if something slips, it's clear what was and wasn't checked.”
Retesting only the reported steps, or insisting on a full regression for every small fix.
Find the waste: cases for removed features, duplicates, cases nobody can follow.
Decide with data: past failures, risk of the area, whether automation already covers it.
Keep it clean: tags or tiers, an owner, and a review each release.
“I did this on my last project, where the suite had grown to around nine hundred cases and a full run took a week. I exported the cases with their last run results and went module by module with a developer and the product owner. Cases for features that no longer existed went straight out. Duplicates, often the same check written by two people, got merged. Then I tagged what was left into a small core set that runs every release, a wider set for the areas that changed, and the rest for major releases. I didn't delete a case just because it never failed, if it guarded something critical like payments. By the end it was just under five hundred cases, the core set ran in a day, and we added a rule that any story that changes behaviour updates its cases before it's closed.”
Deleting every case that never failed, or never removing anything because more tests always feel safer.
New users: the field is required, validated and saved.
Old users: login, profile view, profile edit, and every flow that reads the profile.
The decision: what should happen to old records, agreed before you test.
“New sign-ups are the easy part: the field is required, formats are validated, it saves. The risk is the old users. So I'd first ask what's supposed to happen to them. Are they asked for a number on next login, or only when they edit the profile? Is there a migration filling anything in? Then I'd test with accounts that have no phone number. Can they still log in? Can they place an order, or does checkout now fail because it reads the phone? If they edit only their address, does the save get blocked because the phone is empty, and is the message clear? I'd also check reports, exports and anything that sends messages, since those might now expect a number. I found exactly that kind of bug once: old users couldn't update anything on their profile.”
Testing only the sign-up form and forgetting the users who already exist.
Use the sandbox: the provider's test mode and test cards or accounts.
Force failures: declines, timeouts, closing the page, callbacks arriving late or twice.
Check both sides: your order status and the provider's record must agree.
“We used the payment provider's sandbox, which gives test card numbers for success, decline and extra verification. The happy path took an hour. Most of my time went on what goes wrong. I declined payments and checked the order stayed unpaid and the user could retry. I closed the browser right after paying, before the redirect came back, and checked the order still got marked paid when the provider's callback arrived. I asked a developer to send the same callback twice, to make sure we didn't create two orders or send two emails. And I compared our order statuses with the provider's dashboard for every test. That last check found a real bug: a payment that timed out on our side was successful on theirs, so the customer paid but the order showed failed.”
Testing only a successful payment, or assuming the sandbox behaves exactly like the live service.
Double submit: fast clicks, a slow network, refresh and back after submit.
Two editors: two browsers or users, open, edit, save in different orders.
Check the data: duplicates, lost updates, and what each user is told.
“For double submit, I throttle the network so the request is slow and click submit several times, then try refresh and the back button right after submitting. The screen often looks fine, so I check the database for duplicate rows, and I check emails and payments too. For two editors, I open the same record in two different browsers, or as two users. Both open it, the first saves a change, then the second saves a different change. The question is what happens to the first change. If it's silently overwritten, that's a lost update, and I raise it with the exact order of steps. A good system warns the second person that the record changed. I agree with the product owner which behaviour we want, because there's more than one acceptable answer.”
-- more than one order from the same cart = double submit got through
SELECT customer_id, cart_id, COUNT(*) AS orders_created
FROM orders
WHERE created_at >= CURRENT_DATE
GROUP BY customer_id, cart_id
HAVING COUNT(*) > 1;
Checking only the screen after a double click and never looking for duplicate records.
Time zones: users and servers in different zones, times shown and stored.
Edges of time: midnight, month end, leap day, clock changes.
Formats: day and month order, 12 and 24 hour, what the user sees versus what's saved.
“Time zones are the big one. I set my device or browser to a different zone and book something, then check what's shown to me, to the other person and what's saved. Often the server stores one zone and the screen shows another, and things are an hour or a day off. Then the edges: booking at eleven at night that lands on the next day in another zone, month end, the 29th of February, and the days the clocks change, where an hour can be skipped or happen twice, so a reminder fires twice or not at all. I also check formats, like whether 03/04 means March or April to this user. On my last project I found reminders going out an hour late for months after the clocks changed, because the job used a fixed offset instead of the user's time zone.”
Testing dates only with today's date in your own time zone.
Interruptions: calls, notifications, going to the background, locking the phone.
Network and device: switching or losing network, low battery, small and large screens.
Install life: permissions denied, updating from an old version, reinstalling.
“The flows are maybe half of it. I check interruptions: a call coming in during payment, switching to another app and coming back, locking the phone mid-form. Does it keep my data, or restart the screen? Then network: moving from wifi to mobile data during an upload, or going into airplane mode and coming back. I check permissions, especially what happens if someone says no to the camera or location, because a lot of apps just crash or show a blank screen. I also test upgrades, installing the old version, logging in, then updating, since that's how real users get the app, and I've seen saved data lost on upgrade. And I check a small screen with a large font setting, which is where text gets cut off.”
Treating a mobile app like a website on a smaller screen.
Access: change an ID in the URL, open another role's pages directly.
Session: back button after logout, old links after logout, session timeout.
Data shown: passwords and card numbers masked, no private data in errors or URLs.
“A few checks take minutes and catch real problems. When a page shows my order at a URL with an order number, I change the number to someone else's order. If I can see it, that's a serious bug. I also log in as a normal user and paste in the address of an admin page to see if it opens. After logging out I press back and try an old link, to be sure the pages don't come back. I check that passwords and card numbers are masked, and that error messages don't show technical details or other people's data. I also look for private details in the URL, because URLs end up in logs and browser history. Anything deeper I leave to the security team, but I raise these myself, because they come up in ordinary features all the time.”
Saying security is entirely someone else's job, or suggesting you'd attack production to test it.
What you reported: cases run and passed, open bugs by severity, bugs found after release.
The misleading one: a number that looked good but hid the real state.
What you do now: add context, not just counts.
“Each sprint I reported how many cases ran and passed, open bugs by severity, and how many bugs were found after release in my area. The misleading one was the pass count. One release, nearly everything passed, and the dashboard looked great. But the few that failed were all in payments, and one blocked refunds. Another time the number of bugs I raised was compared across testers, which pushed people to log small cosmetic bugs separately to look busy. So now I never send the numbers without a short note: what's risky, what's blocked, and what I didn't get to test. My lead said the note was the part she actually read. The numbers show the trend, but the note tells people whether we can release.”
Treating a high pass count as proof of quality, or saying you've never looked at metrics.
What was tested: scope, builds, environments, and what was left out.
Open bugs: each with its impact in user terms and any workaround.
Your view: a clear recommendation and the risks behind it.
“I keep it short enough to read in two minutes. First, what I tested: the stories, which build, which browsers and devices, and what regression ran. Then, clearly, what I didn't test and why, like a partner integration that was down in QA. Then open bugs. For each I write the impact the way a user would feel it, like customers with more than fifty orders see the last page load slowly, not just the ticket number, plus any workaround. Finally, my recommendation, for example: ready to release, with the partner integration checked in production first. I don't sign off as if everything's perfect when it isn't. The decision to release belongs to the product owner, but my job is to make sure they're deciding with the real picture in front of them.”
A sign-off that lists only passed test counts, or hides open bugs to avoid a delay.
Flag it early: tell the team and product owner that morning, not at the review.
Offer choices: test one fully, or both at risk, or carry one over.
Fix the pattern: raise it at the retro so stories arrive earlier.
“I'd tell the product owner and the team that morning, straight away, with a quick look at both stories: this one touches payments and needs a full day, this one is a small text change I can cover in an hour. Then I'd offer choices. I can test the small one fully and carry the payments one to next sprint, or test the payments one first and do only the main path of the other. What I won't do is rush both and mark them done, because then the board lies. If it happens once, that's life. If it keeps happening, I bring it to the retro. On my last team we agreed that stories had to reach testing two days before the sprint ended, and developers started handing over smaller pieces earlier.”
Silently testing both halfway and marking them done, or refusing to test either.
Check it: make sure it's real and not already logged.
Raise it: log it or hand it to the owner with steps.
Tell the owner: a quick message, so it isn't a surprise.
“I don't ignore it just because it's not my story. First I check it's real and not already in the tracker. If it's new, I either log it myself with full steps and link it to that feature, or send the details to the tester who owns it, depending on how the team works. Either way I message them directly, something like, found this in the order history while testing coupons, logged it here, have a look. That way they aren't caught off guard in stand-up. I also check whether my story caused it, because sometimes a bug in another area is actually a side effect of the change I'm testing. Once, that's exactly what it was, and catching it saved us from shipping it.”
Ignoring it because it's someone else's area, or raising it in public in a way that blames the other tester.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.