Manual testing interviews for 10+ years of experience skip definitions like regression testing and go after what only experience teaches: why some customers still see the old version after a release, how offline edits collide when they sync, how to call go or rollback in the middle of a cutover, and how you set strategy across many teams, hire testers and defend your team to leadership. It is written for testers with around eight to fifteen years behind them, interviewing for senior QA, test lead or QA manager roles. Each question shows what the interviewer is really checking, a shape for your answer and a sample you can adapt. Say your own version out loud.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Sources: rounding each line versus rounding the total, tax per line versus on the total, discounts split across lines.
Traps: floating-point maths, currencies with different decimal places, partial refunds.
Tests: amounts that land exactly on a half unit, many small lines, and parts that must add up to the total everywhere.
“A difference of one unit almost always means two parts of the system round at different moments. The invoice might round tax on every line and add them up, while the report calculates tax on the total and rounds once. Both look right alone and disagree together. Discounts spread across lines leave a remainder that has to land somewhere, and if the code uses floating-point numbers instead of exact decimals, values like one tenth aren't stored exactly. I'd first ask finance which rule is correct, because that's a business and sometimes legal decision that varies by country. Then I'd design amounts that land exactly on a half unit, invoices with many tiny lines, currencies with no decimals or three, and partial refunds. The check is that the lines, the invoice, the payment and the refund all agree exactly.”
Three lines of 3.335 each
Round per line: 3.34 + 3.34 + 3.34 = 10.02
Round the total: 10.005 -> 10.01 (half up)
Same order, two answers, one unit apart
Dismissing a one-unit difference as too small to matter in a financial system.
Same letter, two forms: an accented letter can be one character or a plain letter plus a separate accent mark, so a strict match fails.
Lost characters: question marks mean a system along the way couldn't represent those characters and replaced them.
Length: a limit counted in bytes cuts text differently from one counted in characters.
Tests: accents, apostrophes, non-Latin scripts and emoji, followed through every system the text reaches.
“Both point at text changing as it moves between systems. The login failure smells like two forms of the same letter. An accented e can be stored as one character, or as a plain e followed by a separate accent mark. They look identical on screen, but if sign-up saved one form and the login request carries the other, a strict comparison says they don't match unless both sides normalise the text first. The question marks usually mean the name passed through something, often an older invoice or export tool, that couldn't represent those characters and swapped them out. Length is the third trap. A form may allow thirty characters while the database column counts bytes, and accented letters and emoji take more than one byte, so text gets cut or rejected. So I test accents, apostrophes, non-Latin scripts and emoji, and follow each one through the database, search, emails, PDFs and exports.”
Treating it as one customer's typing mistake, or testing only plain English letters in name and username fields.
Test permissions, not roles: get the permission list and check each one with a user who has it and one who doesn't.
Blocked, not hidden: a missing button proves nothing; send the action directly as a user who shouldn't have it.
Corners: users with two roles, custom roles with odd mixes, and access removed while someone is logged in.
Risk order: exports, payments, admin settings and other customers' data first.
“Twelve roles times hundreds of screens is a plan nobody finishes, so I test permissions rather than roles. I get the actual list of permissions and what each one should allow, then check each one with a user who has it and a user who doesn't. Custom roles are just bundles of permissions, so that covers them better than any list of roles. Two things matter more than raw coverage. First, hiding isn't blocking. If a viewer can't see the delete button, I still send the delete request directly, through the browser's developer tools or the API, because that's where the rule has to hold. Second, the corners: someone with two roles where one allows and one denies, a custom role with a strange mix, and access removed while the person is logged in. Does their open session lose it straight away or only at next login? I start with exports, payments, admin settings and anything touching another customer's data.”
Checking only that buttons are hidden for each role, without trying the actions directly.
Ask first: which data may lag, for how long, and who agreed it.
Always a bug: the user's own change seeming to revert, data that never catches up, or stale data driving a wrong decision.
Tests: measure the delay normally and under load, check the screen tells the user, and look for decisions made on old data.
“First I'd find out whether it's by design. When services copy data to each other through messages, some delay is normal, so I ask the architect which data is allowed to lag, for how long, and whether product agreed to it. If nobody can answer, that's the first finding. Some things are bugs regardless. If the person who made the change sees their old address come back, that feels like data loss and users will retype it. If the other screen never catches up, a message was lost. And if stale data drives a decision, like shipping to the old address because the order service read it before the update arrived, that's serious. So I test the lag under normal traffic and while the queue is busy, check whether the screen shows that a change is still processing, and walk the flows where old data could cause real harm.”
Raising every delay as a bug, or accepting any delay because the system is distributed.
Know the spread: which app versions are still in use, and by how many people.
Old against new: keep old builds on test devices and run core flows against the new backend before release.
Usual breaks: a removed or renamed field, a newly required field, a value the old app doesn't recognise.
Policy: a minimum supported version and a tested forced or suggested update screen.
“On the web everyone gets the new version, but on mobile people keep old apps for years. So first I ask for the version spread from analytics: which versions are still in use and by how many people. Then I keep a few old builds installed on test devices, usually the oldest one we support plus the most popular ones, and before a backend release I run the core flows on them against the new backend. What breaks is fairly predictable. A field the old app expects was removed or renamed, the backend now requires something the old app never sends, or a new status value arrives that the old app doesn't know, so it crashes or shows a blank screen. I also push for a written minimum supported version, and a forced or suggested update screen we've actually tested, because the day you need it is the wrong day to find out it doesn't work.”
Testing only the latest app build against the new backend and assuming everyone updates.
Conflicts: the same job edited offline on the phone and online in the office.
Interrupted sync: signal lost halfway, the app closed mid-sync, and whether a retry uploads things twice.
Time: a phone with the wrong clock, and the order changes get applied in.
Server moved on: a job cancelled or reassigned while the engineer was offline.
“Working offline is the easy part. The risk is the moment it syncs. First, conflicts: an engineer updates a job on the phone while the office edits the same job online. Which change wins, and does anyone find out? Then interrupted syncs. I cut the signal halfway through a sync or force-close the app, and check nothing is lost and a retry doesn't upload the same photos or notes twice. Time is a trap too. If the app orders changes by the phone's clock and someone's phone is an hour wrong, older changes can overwrite newer ones. And the server moves on while they're offline. A job gets cancelled or reassigned, so what happens when they complete it and sync? I also test a full day's backlog, with lots of photos on a weak connection, because that's what real engineers bring back. For every case I check the office system, not just the phone.”
Testing that screens load in airplane mode and never testing what happens when the queued changes sync.
Cheap checks first: configuration, feature flags and whether every service really runs the same version.
Data: old records created by earlier versions, volume, roles and permissions of real users.
Infrastructure: several servers behind a load balancer, caches, live third parties, server time zone.
Evidence: logs for one failing request, then rebuild that record's shape in staging.
“Same build rarely means same system, so I compare layer by layer, cheapest first. Configuration and feature flags come first, because a flag switched on only in production explains a lot. Then versions: is every service really on the same release, or is one a version behind? Next, data. Production has records created years ago by older versions, with fields that staging data never has, plus real users with odd permission mixes. Then infrastructure. Production might run many servers where staging runs one, so a bug that needs the next request to land on a different server never shows in staging. Caches, live third-party services instead of sandboxes, and the server's time zone come next. I ask for logs of one failing request, find the exact record and user involved, and rebuild that shape in staging. That usually turns the mystery into a normal bug.”
Closing it as cannot reproduce because it works in staging.
Layers: the browser cache, a content delivery network, a service worker, and caches on the server.
The broken mix: an old cached page loading a new script, or new code reading a cached value in the old format.
Test the change-over: keep a session open from before the deploy, and test as a returning visitor, not only a fresh one.
Ask for: version numbers in file names, a cache clear step in the release, and a visible build number.
“Testers usually open the app fresh after a deploy, so they see exactly what we shipped. Customers don't. Their browser may hold old scripts, a content delivery network may keep serving old files until they expire, and a service worker can keep running the old version until every tab of the app is closed. On the server, cached responses can outlive the change too. The broken mix comes from layers disagreeing, like an old cached page loading a new script whose file name never changed, or new code reading a cached value in the old format. So I test the change-over itself. I keep a logged-in session open from before the deploy and carry on using it afterwards, and I test as a returning visitor as well as a new one. I also ask for version numbers in file names and a build number shown somewhere, so support can tell which version a customer has.”
Telling customers to clear their cache and closing the ticket without finding which layer served the old version.
Before the weekend: written go and rollback criteria per check, who decides, and the last hour a rollback still fits.
At 2 a.m.: find the records behind the mismatch and whether it's an agreed exclusion or real loss.
Decide: unexplained money or customer data differences mean rollback, unless a fix fits the window.
Changes during the move: confirm the old system was frozen, or that its changes were captured.
“Most of that decision should be made before the weekend. I'd have agreed written criteria with the business: which checks must match exactly, like money and customer counts, which can carry a known, explained difference, who makes the call, and the latest hour a rollback still fits. We'd have rehearsed the rollback too, not just the migration. So at 2 a.m. I don't debate, I investigate against the clock. I break the totals down by day to find where the gap sits, then pull those records. Is it a group we agreed to leave out, like old test orders, or a rounding rule we already knew about? Or are real orders missing or wrong? If it's money nobody can explain and the fix doesn't fit the window, we roll back. That hurts, but a second attempt costs less than wrong balances for years. I'd also check the old system was really frozen during the move.”
-- run on old and new systems, compare day by day
SELECT CAST(created_at AS DATE) AS order_day,
COUNT(*) AS orders,
SUM(total_amount) AS amount
FROM orders
GROUP BY CAST(created_at AS DATE)
ORDER BY order_day;
Deciding on the night with no agreed criteria, or going live because the mismatch looks small.
Log it as a difference: not a bug in either system, a change in behaviour.
Find the impact: reports, integrations, customer habits and scripts that rely on the old way.
Get a decision: the product owner chooses keep, change with notice, or change behind an option, and it's written down.
“I'd record it as a behaviour difference, not a bug in either system. On one migration the old system sorted search results in a way that was technically wrong, but a downstream team's nightly export depended on that order, and some long-time users had learned to rely on it. Correct behaviour can still break people. So I found everyone who consumed that output: reports, integrations, support scripts and customers. Then I took it to the product owner with three options and their cost: copy the old behaviour, change it with notice to the affected teams and customers, or change it with a setting for a while. They chose to change with notice. My job afterwards was to test the new behaviour, confirm the downstream team had updated its export, and check support had a line ready for users who noticed.”
Quietly accepting the new behaviour as correct without checking who relied on the old one.
Listen: map teams, releases and tools, and study where recent production incidents came from.
Few standards: shared severity definitions, testing in the definition of done, one place for bugs.
Strategy by risk: what each test level covers and who owns it, by product area.
Pilot and measure: try it with one team, track escaped defects, then roll out.
“The first month is mostly listening. I'd sit with each team, watch a release go out, and go through the last few months of production incidents to see where problems really came from. That usually shows a pattern, like most escapes coming from the API changes breaking mobile. In the second month I'd set a small number of shared rules: one set of severity definitions with examples, testing written into the definition of done, and one place where bugs live. Then a written strategy by risk, saying which areas get deep testing, what each level of testing covers and who owns it, including what developers check before handing over. By the third month I'd pilot it with one team, track escaped defects and release delays, fix what didn't work, and only then roll it out to the other four.”
Arriving with a heavy template process and rolling it out to every team in week one.
What doesn't change: a small group of real users is still real users, and some damage can't be rolled back.
New things to test: the flag on and off, switching it off mid-flow, and data created while it was on.
Watch the rollout: agree the signals and the levels that pause it before it starts.
Flag hygiene: an owner and a removal date for every flag.
“I welcome it, but I'd correct the idea behind it. Rolling out to a small group limits how many people a bug hits. It doesn't make the bug fine. And some damage can't be undone by switching a flag off: corrupted data, wrong charges, one customer seeing another's details. Those still need testing before anyone sees them. What changes is that there's more to test. The feature with the flag on and off, the switch-off itself while someone is halfway through the flow, and whether data created under the new feature still works once it's off. Before a rollout starts, I agree with the team which signals we watch, like error rates, support tickets and the key business numbers, and what level pauses it. And I push for flag hygiene. Every flag gets an owner and a removal date, because old flags multiply the combinations nobody tests.”
Agreeing that a small rollout means testing can be light, or ignoring the flag-off path entirely.
Every version works: the losing version still takes real orders while the test runs.
Assignment: a customer keeps the same version across visits and after login, if that's the agreed rule.
Tracking: the events that decide the winner fire once, at the right moment, with the right version attached.
Ending it: switching the experiment off moves everyone cleanly to the chosen version.
“I test three things beyond the screens. First, every version fully works, because the losing version still takes real orders while the experiment runs. Second, assignment. A customer should keep the same version when they come back and after they log in, and ideally on another device, if that's the rule we agreed. If people flip between versions, the result is noise and the experience is confusing. I also check who's meant to be left out, like staff accounts. Third, the tracking. The decision will be made on events like added to cart or order paid, so I watch those events in the network tab or the analytics tool's debug view and check each one fires once, at the right moment, with the right version attached. A missing or doubled event can pick the wrong winner. Finally, I test switching the experiment off, so everyone lands cleanly on the chosen version.”
Testing each version's screens and never checking assignment or the events the decision rests on.
Traceability: requirement to risk to test to result to defect to sign-off, in both directions.
Evidence: who ran it, when, on which build, with screenshots or logs captured at the time.
Control: plans approved before execution, test changes versioned, deviations written up and approved.
“In a regulated setting, a test that wasn't recorded properly might as well not have happened. The rules differ by industry and country, so I work with the compliance team from day one, but the core stays the same. Every requirement, especially safety or money-related ones, traces to a risk, to tests, to results, to any defects and to a sign-off, and an auditor should be able to walk that chain both ways. Evidence is captured while the test runs: who executed it, the date, the build version, the environment, and screenshots or logs for the key steps. The test plan is approved before we start, and if a test case changes mid-cycle, the change is versioned and reviewed. When we skip something or accept a known defect, that's a written deviation with a reason and an approver. Nothing gets reconstructed after the fact.”
Planning to gather screenshots and fill in the traceability matrix after testing is finished.
Before the sprint: product, developer and tester go through each story with concrete examples.
Questions: what if, what about, and what happens when it fails, written into the story.
Testability: ask for logs, test data and feature flags at design time.
Show it worked: track where bugs come from, release after release.
“First I'd prove the pattern with data. I tag the last few releases' serious bugs by where they started, so the conversation is about facts, not blame. Then I'd put testers into the room before the sprint starts. For each story, the product owner, a developer and a tester spend a short session turning the acceptance criteria into concrete examples: this customer, this cart, this result. The tester's job there is to ask what if. What if the coupon expires mid-checkout? What if the user has two accounts? The answers go straight into the story. I'd also ask for testability at design time: logs we can read, a way to set up test data, a feature flag to switch things off. Then I'd keep tagging bugs by origin, and show teams the requirement bugs shrinking. That trend is what keeps the habit alive.”
Blaming business analysts and asking only for longer requirement documents.
Open the estimate: show what the six weeks covers, area by area, ranked by risk.
Offer options: what three weeks covers fully, lightly and not at all.
Change the levers: scope, phased rollout, feature flags, earlier builds, more people.
Record it: the decision and the accepted risks, with the owner's name.
“I wouldn't argue about the number, I'd open it up. I'd show the six weeks broken down by area and ranked by risk: payments and data migration at the top, settings pages near the bottom. Then I'd lay out what three weeks buys. These areas fully tested, these lightly, these not at all, and here's what could go wrong in the ones we skip. Next I'd look for other levers. Can we release the risky part a week later behind a feature flag? Roll out to a small group of customers first? Get stable builds earlier so testing starts sooner? Extra people help only if they arrive early enough to learn the product. Whatever we choose, I write it down with the risks accepted and who accepted them. Usually the programme manager finds a middle option once the risks have names.”
Quietly agreeing to three weeks and cutting corners without telling anyone what won't be tested.
Early: business users write scenarios from their real working day, weeks before testing.
Ready kit: stable environment, prepared data, short training and simple scripts.
Protected time: agreed with their managers, since they have day jobs.
Clear sign-off: criteria agreed upfront, coverage tracked by business process.
“The last round failed for predictable reasons. Users were squeezed in around their day jobs, the environment was shaky, and sign-off meant nothing specific. This time I'd start weeks earlier by asking a few experienced users to describe their real working week, including month-end and the awkward cases, and we'd turn that into their scenarios. I'd give them a ready kit: a stable environment, their kind of data already loaded, a short walkthrough and simple steps. I'd agree protected time with their managers, because an hour squeezed between calls isn't testing. We'd run a short triage every day so their issues get answers fast and they see the point. Progress is tracked by business process covered, not by scripts ticked. And sign-off criteria are agreed before we begin, so signing means something.”
Handing business users a long script library and waiting for the sign-off email.
Agree first: fewer escapes was the goal, and asking about cost is fair.
Show what's caught: serious bugs found before release, by area, with what each would have meant in production.
Give options: what a smaller team stops covering, and which risks leadership then owns.
Offer a path: move effort to prevention or automation and measure it for a quarter.
“I'd start by agreeing that fewer bugs reaching customers is exactly what we wanted, and that asking what it costs is fair. Then I'd point out the number is down partly because the team catches problems earlier. So I'd bring what we find before release: over the last few releases, the serious bugs caught in testing, grouped by area, with a couple of concrete examples of what each would have meant in production, like wrong invoices or a broken sign-up. Then options, not a flat no. Here's what a team one third smaller would stop covering, and here are the risks leadership would be accepting. Often there's a better middle path, like moving some people into automation or earlier requirement reviews, and tracking escapes for a quarter before deciding. If they still choose the cut, I make sure the dropped coverage is written down and owned by name.”
Getting defensive and saying quality will collapse, with no evidence and no options.
Look beneath green: trace recent escapes to their test cases, test data and evidence.
Find the cause: thin product knowledge, happy-path cases, clean data, or a contract that rewards cases run.
Change the measures: escaped defects by severity and coverage of risky areas.
Close the gap: their leads in refinement and triage, and one named contact on your side.
“Green reports with production escapes means the reports measure the wrong thing. I'd take the last few serious escapes and trace each one back. Did the vendor have a test for it? Did they run it, on what data, and what evidence did they keep? I'd also sit in on a few of their sessions. Usually the cause is a mix of thin product knowledge, test cases that only follow the happy path, clean data that looks nothing like production, and a contract that pays for cases executed, which quietly rewards finishing over finding. So I'd change what we measure to escaped defects by severity and coverage of the risky areas, and bring their leads into our refinement and bug triage so they understand the product. I'd name one person on our side to answer their questions fast. Then I'd give it a couple of release cycles with clear targets before deciding whether to keep them.”
Demanding more test cases and longer reports from the vendor without finding why the escapes happened.
Exercise: a small real app or a spec with gaps, explored live.
Signals: the questions they ask, how they choose what to test first, one written bug report.
Fairness: a structured scorecard and more than one interviewer.
“I stopped hiring on definitions years ago, because anyone can learn the difference between smoke and sanity testing from a list. Now every candidate gets a practical round. I give them a small web app with a few planted problems, or a short spec with gaps, and let them explore while talking. I watch for three things. Do they ask questions before clicking, like who uses this and what matters most? Do they choose what to test first based on risk? And can they write one bug report a developer could act on without a follow-up chat? For senior hires I add a conversation about a disagreement with a developer or a release they pushed back on. Everyone gets the same exercise and a scorecard, and two of us interview, so I'm not just hiring people who remind me of myself.”
Describing hiring as a quiz of testing definitions, or relying only on gut feel.
Starting point: their strengths and the specific gap.
What you did: pairing, stretch ownership, feedback on real work.
Result: what they could do afterwards and how you knew.
“At my last company there was a tester who was thorough but wrote hundreds of shallow test cases and never questioned a requirement. She wanted to become a lead. I started by pairing with her on exploratory sessions, thinking out loud about risk, and then swapped so she drove and I watched. I gave her feedback on her bug reports, mainly on explaining impact in business terms. After a couple of months I gave her one feature end to end: the test approach, the estimate and the release recommendation, with me reviewing but not deciding. She got the estimate badly wrong the first time, and we went through why together. By the second quarter she ran bug triage for her team, and the next major release she led on her own. I knew it had worked when developers started going to her first instead of me.”
An answer that's only about sending someone on a course or a certification.
Embedded: faster feedback and deeper product knowledge, but standards drift and testers can feel alone.
Central: shared practice and flexible staffing, but testing becomes a hand-off at the end.
What you chose: the model, the reason, and the trade-off you named.
Evidence: what you tracked to see whether it worked.
“I've run both, and each fails in its own way. With a central team, standards were consistent and I could move people to wherever the crunch was, but work got thrown over the wall at the end and testers never learned the product deeply. When we moved testers into the delivery teams, feedback got much faster and they joined refinement, but within a year every team tested differently, and a few testers felt alone with nobody to learn from. What I recommended to leadership was a mix. Testers belong to their delivery team day to day, and also to a testing practice I lead, with shared standards, peer review, a regular meetup and one career path. The trade-off I named honestly was losing the ability to move people quickly between teams. I tracked cycle time and escaped defects per team to check it was working.”
Declaring one model always right without naming what it costs.
What happened: the release, the failure and its impact on customers.
Your part: the specific decision or silence that was yours.
What changed: the lasting process change, and evidence it held.
“A few years ago we rebuilt the checkout flow against a fixed launch date. Late in the project, the integration with the warehouse system slipped, and I agreed to cut its testing to a quick check so we could keep the date. After launch, some orders with mixed stock never reached the warehouse, and for two days customers didn't get what they'd paid for. We rolled that part back and fixed it. My part was clear: I raised the risk in a meeting but never wrote it down, and I didn't make the product owner decide on it explicitly. So the risk quietly became nobody's. Since then, every testing cut goes into a release risk list that the product owner signs, with what we skipped and what could happen. I've used that on every release since, and twice it's moved a date.”
Choosing a story where the failure was entirely someone else's fault.
Keep it small: a few required bug fields, shared severity with examples, a short case style guide.
Show, don't tell: a handful of great real examples beats a long document.
Make it stick: peer review, a regular testers' meetup across teams, spot checks.
“I'd start by collecting real bug reports from each team and asking developers which ones they liked and why. That gives us standards people already believe in. Then I keep the rules small: a bug needs steps, expected and actual results, environment and build, evidence and impact, and severity follows one shared table with examples from our own product. For test cases, a one-page style guide, mainly telling people to write the intent, not every click. I pin a few excellent real examples where everyone can see them, because people copy examples far more than they read rules. To make it stick, testers review a couple of each other's reports every week across teams, and we hold a short monthly meetup at a time that works for every zone. Every quarter I sample reports and share what's improved rather than who broke the rules.”
Writing a long QA policy document and expecting every team to follow it.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.