Manual testing interviews for 5 years of experience skip definitions like regression testing and ask why you tested something the way you did, what you chose to leave out, and what happened when a bug got past you anyway. Expect questions on estimates under pressure, regression next to automation and feature flags, data migrations, partner integrations, release decisions, quality numbers and mentoring. It is written for testers with roughly five to seven years of manual testing behind them, who own the testing for a module, write its strategy, run triage, make release calls and review what juniors write. Each answer below is a first-person story or a decision you can defend. Swap in your own project details.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Show the basis: what the estimate is made of, area by area, and where the time really goes.
Offer options: what one week can cover, ranked by risk, and what would need more people or less scope.
Make the gap visible: the untested areas written down and accepted by whoever owns the release.
“On a billing release I'd estimated three weeks, and my manager came back saying the business needed it in one. I didn't argue about the number itself. I showed him the breakdown: about a week on the new invoice rules, a week of regression around payments and tax, and several days on data setup and retests. Then I gave him three options. One week could cover the new rules and payment regression, leaving tax reports and older plan types lightly tested. Or two developers could help with data setup and we'd cover most of it in ten days. Or we could release the new rules to new customers first. He picked the first option, and I wrote the untested areas into the release note so the product owner accepted them knowingly. The trade-off was real risk in the tax reports, so we checked those in production the day after release.”
Agreeing to one week and quietly testing less, or refusing without offering any way to meet the date.
Risks first: list what would hurt users or the business most if it broke.
Choices: which levels, types and environments cover those risks, and who does what.
Left out: name what you won't test, why, and who accepted that.
“When I owned testing for a new returns module in an online store, I started with a risk session with the product owner and two developers. The top risks were refunds going to the wrong payment method, stock not coming back after a return, and warehouse staff scanning the wrong item. So my strategy put most effort there: detailed cases for refund paths, checks in the stock screens and the database, and exploratory sessions on the handheld scanner flow. I wrote down what I was leaving out: full browser coverage for the staff-only screens, since the warehouse used one managed browser, and deep performance testing, since volume was low at launch. The product owner signed off on both gaps. The cost was real risk if the warehouse ever changed browsers, so I added that as a trigger to revisit the strategy.”
Describing a copy of a standard template with every test type listed, and nothing left out on purpose.
Pairwise: pick a small set of combinations where every pair of setting values appears at least once.
Add risk: hand-add the combinations real customers use and the ones that broke before.
Limits: admit bugs that need three or more settings together can still slip through.
“We had a school management product where each school set things like grading scale, term count, attendance mode, language and fee schedule. All combinations ran into the thousands. I used pairwise testing: a generator gave me a set of about thirty combinations in which every pair of values showed up at least once. In my experience most config bugs came from two settings clashing, so that caught a lot. Then I added combinations by hand: the setups our five biggest customers actually used, and any mix that had caused a support ticket before. The trade-off I accepted is that a bug needing three specific settings together could still get through. We did get one of those once, and I added it to the fixed list so it never slipped again.”
Saying you test every combination, or picking combinations at random with no way to explain the coverage.
Work it out yourself: calculate expected figures from the raw data, without copying the report's own logic.
Small known data: a hand-built set where you know every answer in advance.
Sanity rules: parts add up to the total, and the old report agrees wherever the rules didn't change.
“We built a new monthly revenue report by region, and the business only knew roughly what the numbers should be. So I used three checks. First, I wrote my own SQL against the raw orders table to work out the totals, without looking at how the developers had built the report, and compared the two. Second, I made a small test month with about twenty orders I'd created by hand, including a cancelled order and one placed just before midnight at month end, so I knew every expected figure exactly. Third, simple rules: the regions had to add up to the grand total, and old months had to match the old report. The month-end order found a real bug: the report used the server's time zone, so late orders landed in the wrong month. The trade-off was that my SQL could share a wrong assumption, so I had the finance lead review my filters.”
-- My own expected totals, straight from raw orders
SELECT region,
SUM(amount) AS expected_revenue
FROM orders
WHERE status = 'PAID'
AND order_date >= DATE '2026-08-01'
AND order_date < DATE '2026-09-01'
GROUP BY region
ORDER BY region;
Checking that the report loads and the numbers look about right, or building the expected values by reusing the developer's own query.
Within reach: keyboard-only use, a screen reader on key flows, zoom and colour contrast.
Make it routine: a short checklist on every story, not a one-off audit before launch.
The line: a full audit and testing with disabled users need specialists, so raise it as a risk.
“A public sector client needed our portal to meet the accessibility guidelines at level AA, and nobody on the team had done it before. I took on the checks I could do reliably. I went through the key journeys using only the keyboard, checking I could reach every button and always see where the focus was. I used a free screen reader on the sign-up and payment forms, and found several fields with no labels, so the reader just said edit text. I zoomed pages to twice their size to see what broke, and ran our colours through a contrast checker. Then I added those checks to the story checklist so they happened every sprint. Where I drew the line was a full audit and testing with people who use assistive tools daily. I raised that as a risk, and the client booked an external audit before launch. It found problems with complex tables that I'd missed.”
Treating accessibility as optional, or claiming an automated scan alone proves the product is accessible.
New states: every live flag adds paths, so know which flags are on for real customers.
Choose: test both states for flags that touch risky areas, and only the common state for the rest.
Clean up: push to remove flags once a feature is on for everyone, and test the removal too.
“At my last company we moved to feature flags so new features could go to a few customers first. The first surprise was a bug that only happened with one new flag on and an older flag off, which was exactly the setup of our biggest customer. So I changed how we planned regression. I kept a list of every live flag and which customers had it on. Each release, I tested the flags that touched checkout or billing in both states, and for the rest I tested only the state most customers were in. I also pushed the team to remove a flag within a couple of sprints once a feature was on for everyone, and to treat that removal as a real change, since deleting an old path once broke a report. The trade-off was that rare flag combinations got no testing, and I wrote that down as a known risk.”
Testing only with every flag switched on, or treating flags as a developer detail that doesn't affect testing.
Line up the misses: for each escaped bug, find the automated test that should have caught it.
Read the checks: what each test really verifies, and with what data.
Close the gap: stronger checks, better data, and a clear split of what stays manual.
“We had three production bugs in two months in checkout, which the automation team said was fully covered. I took each bug and found the automated test that walked through that flow. Reading them with the automation engineer, the pattern was plain. Most tests checked that the next page loaded, not that the total, tax or discount on it was right. Two used the same test customer with one item in the cart. And one flaky test had been set to retry until it passed, which hid a real timing bug. We agreed the tests would check the actual values from my manual cases, use a few data sets including large carts, and that flaky tests got fixed or parked with a ticket, never silently retried. I kept a short manual pass on checkout totals too. The trade-off was a slower suite and a week of rework, but it was months before the next checkout bug escaped.”
Assuming a green suite means the area is tested, or blaming the automation team without reading what the tests check.
Reconcile: counts and totals by group must match between old and new.
Field mapping: sample records field by field, including the ugly ones.
Rehearse: run the full migration on a production-size copy for timing and rollback.
“We moved around a million insurance customers to a new platform. I started from the field mapping document and wrote checks for each rule, especially the transformed ones like merged address lines and old status codes mapped to new ones. For proof, I compared record counts and totals, like policy counts and premium totals, grouped by status and region, in both systems. Then I picked a sample on purpose: very old records, names with accents and apostrophes, missing phone numbers and customers with many policies. Those found three mapping bugs, one of which cut long names short. We did two full rehearsals on a masked production-size copy, which showed the migration took nine hours, longer than the planned window. The trade-off was that we only checked a sample field by field, so we relied on the reconciliation totals for everything else.”
-- Run on old and new systems, then compare the two results
SELECT status, region,
COUNT(*) AS customers,
SUM(policy_count) AS policies
FROM customers
GROUP BY status, region
ORDER BY status, region;
Checking a handful of records by eye and calling it done, or never rehearsing the migration at full size.
What happened: what the partner changed and why your tests couldn't see it.
Handle the unknown: test what your system does with values and fields it doesn't expect.
Watch the partner: regular checks against their test system and a way to hear about changes early.
“Our shipping partner added a new delivery status without telling anyone. Our code didn't recognise it and marked those orders as cancelled, so customers got cancellation emails for parcels that were on their way. The fix itself was quick. What I changed was the testing. We'd only ever tested the statuses listed in the partner's documents, so I added cases for an unknown status, a missing field and an empty response, and agreed with the developers that anything unknown should hold the order and alert us, never guess. I asked the automation engineer for a daily check that called the partner's test system and flagged any field or value we hadn't seen before. And our account manager got the partner to send us change notices. The trade-off is that held orders mean some manual work for support, but that's far better than wrong emails to customers.”
Blaming the partner and stopping there, or only adding a test for the one status that broke.
Find the pattern: why hotfixes caused new breaks.
Process: fixed steps sized for speed, with a short impact check and a core smoke set.
Cost: the time it added and how you kept it small.
“In one quarter, three of our hotfixes caused new production bugs. When I looked back, each fix was tested only on the exact bug, on a branch that was missing other recent changes. I set up a simple process with the developers. Every hotfix went on a branch cut from the current production version. The developer wrote two lines on what else the change could touch, and I tested the fix, those areas and a thirty-minute core smoke set of login, search, checkout and order history. If the change touched payments or data, the product owner had to approve a longer check. It added about an hour to each hotfix, and some developers grumbled at first. But we didn't have another hotfix break something for the rest of the year, and that ended the grumbling.”
Either testing only the reported bug, or insisting every hotfix wait for a full regression cycle.
Before: state the risk in plain words and get the decision recorded.
On the day: reduce the damage with monitoring, support notes and a fallback plan.
Afterwards: review what happened without blame, whichever way it went.
“On a loyalty app release, I recommended holding back because points weren't recalculated correctly when an order was partly refunded. The product owner had a marketing launch booked and decided to go. I made sure the risk was written in the release ticket in plain words, with roughly how many customers it could affect, and that she'd accepted it there. Then I switched to reducing the damage. I gave support a note on how to spot the problem and correct a balance by hand, asked the developers to log every partial refund, and checked those logs each morning. The bug hit a few dozen customers in the first week, support fixed them quickly, and the fix went out in the next release. In the review nobody played the blame game. We agreed that anything touching points balances gets tested before a marketing date is set. Releasing was her call, and I respected it.”
Refusing to support the release once it's decided, or going quiet and letting the risk go unrecorded.
Understand why: find the problem the plan is trying to solve.
Show the risk: late bug finding and weak independent checks, with examples from your own history.
Offer a middle path: more developer testing plus tester involvement on the risky stories.
“First I'd ask what's driving it. Usually it's that testing feels like a bottleneck at the end of the sprint. I'd agree that developers testing more is good, and say that finding problems only at release time makes each one costlier to fix and bunches all the work at the end. I'd bring examples from our own last few releases: which bugs came from misread requirements, which a developer's own checks would never catch because they built it with the same idea in their head. Then I'd suggest a middle path: developers own unit tests and basic checks on every story, I pair with them on test ideas during refinement, and testers still do hands-on testing on the stories we rate as risky. I'd offer to try it for two sprints and compare escaped bugs and cycle time.”
Refusing flatly because it threatens the tester's job, or agreeing without raising the risk.
Before the meeting: every bug has steps, impact and a proposed severity.
Rules: fix now, schedule, or close as won't fix, based on impact, frequency and workaround.
After: record decisions and reasons so nobody reopens the debate.
“I ran triage twice a week with the product owner and the lead developer, and kept it to thirty minutes. Before each meeting I made sure every new bug had clear steps, who it affected and how often, and a proposed severity, so we weren't investigating in the room. Our rule was simple: anything that lost data, blocked a main flow or hit many users was fixed in the current sprint. Things with an easy workaround went to the backlog with a target sprint. Old low-impact bugs on screens we planned to replace were closed as won't fix, with the reason written down. I also flagged patterns, like five bugs in the same screen, because that pointed to a design problem. The trade-off was that some testers felt their bugs were ignored, so I always posted the reason on each ticket.”
Letting whoever shouts loudest set priority, or never closing a bug as won't fix.
Group them: for each escaped bug, where it could have been caught and why it wasn't.
Pattern: the common cause across bugs, not one person's miss.
Change: one or two concrete process fixes, then check they worked.
“We had seven production bugs from the orders module in one quarter. I put them in a simple table: what broke, which environment it would have shown up in, and why we missed it. Four of the seven only happened with real-world data, like orders with more than a hundred lines or addresses in other scripts, and our test data was small and tidy. Two came from a config difference between staging and production. Only one was a plain miss in a test case. So I built a set of ugly test orders based on anonymised production patterns and got the team to add a config comparison step before each release. The next quarter we had two escaped bugs, neither from data or config. The trade-off was a week of my time on data and tooling instead of testing new stories.”
Blaming the developers or the deadlines for every escaped bug, or fixing each one with a single new test case and no wider look.
Useful numbers: escaped bugs, open bugs by severity over time, test progress against risk.
Bad rule: one you dropped and the behaviour it caused.
Replacement: what you did instead and why it was better.
“Each release I reported bugs found in production against bugs found before release, open bugs by severity and which way that trend was moving, and how much of the high-risk area we'd covered. The rule I got rid of was our release gate: no open high-severity bugs. It sounded strict, but triage turned into an argument about labels. A couple of days before each release, serious bugs kept getting downgraded to medium so the gate would pass, and once a real data bug went out labelled medium. I proposed dropping the automatic gate. Instead, my sign-off listed every open bug with its user impact in plain words, and the product owner decided. I also tracked how often a severity was lowered after triage, which made downgrading visible. The trade-off was losing a simple yes-or-no rule that managers liked, so every release needed a real conversation.”
Reporting only test case counts, or defending a number people were clearly gaming.
Narrow it: who is affected, since when, and what changed recently.
Reproduce safely: try the same steps in a test environment with similar data.
Feed back: hand clear findings to developers, then add the missing test.
“One Monday, support said some customers couldn't download their invoices. I joined the call and first asked three things: which customers, since when, and what was released last. It turned out only invoices created since Friday were affected, all for customers using one of our supported languages, and we'd released a change to the invoice template on Friday. I reproduced it in staging by creating an invoice for a test account set to that language and got the same error. The new template used a date format that failed for that language's month names. That gave the developer the exact trigger within the hour, and the fix went out the same day. Afterwards I added invoices in every supported language to our release checks. The trade-off was a slightly longer regression run, which I thought was worth it.”
Waiting for developers to solve it, or starting to test random areas with no attempt to narrow down who was affected.
Measure: log the hours lost and the causes for a couple of weeks.
Propose: deploy windows, owners, health checks or a separate environment.
Trade-off: what the fix costs and who pays for it.
“I'd start by keeping a simple log for two weeks: each time the environment was down, for how long, and why. When I did this on a past project, we'd lost about six working days across the test team in two weeks, mostly from unannounced redeployments and one team's test data scripts wiping shared tables. I took that log to the team leads and proposed three things: deploys only at fixed times twice a day, a message in the team channel before each one, and a basic health check page that showed whether key services were up. The data scripts team moved to their own schema. It meant developers sometimes waited a few hours to deploy, which they didn't love. But the lost time dropped to almost nothing, and the log made it easy to show why.”
Treating downtime as someone else's problem, or blaming another team without any record of what happened.
Risks: personal data exposure, privacy law, and tests reaching real customers.
Alternatives: masked copies, generated data built from real patterns, or a small seeded set.
Trade-off: what realism you lose and how you get it back.
“I agreed with the goal, because realistic data finds real bugs. But a raw copy puts real names, phone numbers and card details on an environment many people can reach, and privacy laws in most places treat that seriously. There's also the risk of a test sending a real email or SMS to a real customer, which I've seen happen at a previous job. So I pushed for a masked copy: names, emails, phone numbers and addresses replaced with fake but realistic values, keeping the shape of the data, like long names and odd characters. All outgoing email and SMS in that environment went to a catch-all inbox. The trade-off was a few days of setup and some lost realism, since masked addresses no longer matched real postcodes, so we kept a small hand-built set for address validation tests.”
Happily using raw production data with personal details, or rejecting realistic data with no alternative.
Coverage first: map cases against requirements and risks, find the gaps.
Quality: clear steps, one clear expected result, no duplicates.
Feedback: a few patterns explained, not forty separate comments.
“I don't start at case one. I first check the cases against the acceptance criteria and the risks I'd expect for that feature. With one junior's cases for a discount code feature, all the happy paths were there, but nothing on expired codes, codes used twice, or two codes on one cart, and nothing on what happens to the discount when an item is removed. Then I looked at the cases themselves: several had expected results like 'works correctly', and about ten were the same test with a different valid code. I sent back three points, not forty comments: the missing negative and state cases with one example each, the vague expected results, and merging the duplicates. Then we sat together for twenty minutes on the first rework. It took more of my time up front, but her next set needed very few changes.”
Only correcting formatting and wording, or rewriting the cases yourself without explaining why.
Look at examples: find what the bounced reports have in common.
Coach: show the gap using their own report, then practice together.
Follow up: watch the next reports and step back once they're solid.
“I pulled five of his bounced reports and read them next to the developers' comments. The pattern was clear: he wrote the steps from memory after testing, so he skipped the setup, like which user role he was logged in as and what data was already in the cart, and he never mentioned the build or browser. I sat with him and we tried to reproduce one of his own bugs using only his report. We couldn't, and that landed better than any lecture. We agreed he'd write steps while testing, start from a fresh login, and attach a short screen recording for anything with more than a few steps. For two weeks I read his reports before they went out. After that the bounces mostly stopped, and I stopped checking. It cost me a few hours, but it saved the developers far more.”
Just telling the junior to add more detail, or quietly rewriting their reports yourself.
Set up: scenarios in business language, test data ready, clear dates.
Sort feedback: defect against agreed requirement, or change request.
Keep goodwill: record every request and route it, never just reject it.
“For a new claims approval screen, I prepared about fifteen UAT scenarios written in the business's own terms, like approve a claim above your limit, and set up logins and sample claims so users didn't waste time on setup. In the first two days, half of what they raised wasn't a bug; it was things like wanting an extra column or a different button order. So I set up a simple rule: each item got labelled either defect, meaning it didn't match the agreed requirement, or change request. I went through them daily with the business lead and the product owner. Defects went to the team; change requests went to the backlog with the user's name on them, and a couple of small ones were squeezed into the release. Users felt heard, and we signed off UAT on time instead of it turning into a redesign.”
Rejecting every request that isn't a bug, or letting UAT turn into an open-ended redesign with no sign-off.
Situation: the story and what looked fine on paper.
Question: what you asked and why your experience made you ask it.
Result: what changed and what it saved.
“We had a story to let users change their email address on their profile. It had three acceptance criteria, all about the form. In refinement I asked what happens to the old email: do we confirm the new one before switching, and can someone who's taken over a session change the email and lock the real owner out? I'd seen a similar gap on a previous product become a support headache. Nobody had thought about it. We added a confirmation link to the new address, a notice to the old one, and a rule that a recent password entry was needed. That was an extra day of development, and the product owner wasn't thrilled at first. But finding it after release would have meant a security fix under pressure and possibly locked-out customers. Since then, the team adds a 'what could go wrong' minute to each story.”
Saying testers only get involved once the build is ready, or having no example of early influence.
Prepare: areas or charters, test accounts and a simple way to log findings.
Run: short time box, mixed groups, one person triaging live.
Judge: what it found that normal testing didn't, and what it cost.
“Before launching a new mobile app for a food delivery client, I ran a two-hour bug bash with developers, designers, support staff and two sales people. I prepared six charters, like order for a group and change the address mid-order, gave everyone test accounts and a shared form with fields for steps, device and a screenshot. People worked in pairs mixing roles, and I triaged findings live so duplicates didn't pile up. We got about sixty reports, which came down to twenty real bugs after merging. Four were serious, and two came from support staff who used the app the way real customers complain about it. The cost was two hours for fifteen people, plus a day of my prep and cleanup. I judged it worth it because those two serious bugs were outside anything our test cases covered.”
Letting people click around with no charters or logging, or counting raw reports as the result.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.