Framework Decisions • Grid and CI • Test Data • Selenium 4 • Suite Health • Mentoring • 2026

Selenium Interview Questions for Experienced Candidates (5 Years)

Selenium interviews for experienced candidates with around five years never ask you to define an explicit wait; they ask why you built something the way you did, what it cost, and how you handled the day it broke. Expect questions on test data, Grid and CI, Selenium 4 features, flaky suites, pushback and mentoring. It is written for automation engineers with roughly five to seven years of Selenium behind them, who own the framework or a large part of it, decide where the tests run, get the message when the pipeline goes red and review other testers' code. Each answer below is a first-person story or a decision you can defend. Swap in your own project details.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Suite Health 4 questions

Hard Technical round Mid-level, Senior Practice question

1. How did you measure flakiness across your whole Selenium suite, and what rule did you set for quarantining a test?

What the interviewer is really testing:
Whether you manage flakiness with data and a clear policy instead of gut feel, and whether you know quarantine has a cost.
Answer frame:

Data: store every result with the test id and the commit it ran on.

Definition: a test that passed and failed on the same unchanged code is flaky.

Rule and cost: quarantine with an owner and a deadline, and admit what is unguarded meanwhile.

Sample spoken answer:

“We used to argue about which tests were flaky, so I started saving every run's results into a small table: test id, pass or fail, commit, browser and Grid node. Any test that both passed and failed on the same commit counted as flaky, because the code hadn't changed. Once a test failed that way more than twice in fifty runs, a job moved it into a quarantine group. Quarantined tests still ran, but they couldn't block a merge, and each one got a ticket with an owner and a two-sprint deadline. If nobody fixed it by then, we deleted it and wrote down the coverage we'd lost. The trade-off I accepted is that a quarantined test isn't protecting anything, so I kept that list short and visible in the team channel every week.”

Red flag to avoid:

Calling a test flaky from memory, or quarantining tests with no owner, no deadline and no record of what stopped being covered.

They may ask next:
  • How did the node column help you find causes that weren't in the test code?
  • What stops quarantine from turning into a place where tests go to be forgotten?
Say it in 60 seconds
Hard Situational round Mid-level, Senior Practice question

2. Developers have started ignoring red builds from your UI suite. How do you win back their trust?

What the interviewer is really testing:
Whether you see a noisy suite as your problem to fix and can rebuild trust with evidence rather than reminders.
Answer frame:

Own it: a suite people ignore is worth nothing, so noise is the automation team's bug.

Triage: sort every failure into product bug, test bug or environment for a few weeks.

Show it: fix the worst offenders, make failures readable, share the real bugs caught.

Sample spoken answer:

“When that happened on my team, I didn't send a reminder to check the builds. I treated it as my bug. For two weeks I looked at every red run myself and sorted each failure into product bug, test bug or environment. Most were test bugs and a shaky test database, which told me exactly where to start. I paused new automation for two sprints, fixed or quarantined the worst tests, and changed failure messages so they say what the user couldn't do, like cannot apply coupon on checkout, instead of a timeout on some locator. Then every week I posted the real product bugs the suite had caught. After about a month, developers were opening failures on their own again. The cost was two sprints with no new coverage, and I explained that trade to my manager up front.”

Red flag to avoid:

Blaming developers for not looking, or asking them to rerun until green.

They may ask next:
  • How would you convince your manager to pause new automation for two sprints?
  • What would you measure to know trust was actually back?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

3. Tell me about UI tests you deleted or moved down to the API or unit level. How did you decide which ones, and how did you convince the team?

What the interviewer is really testing:
Whether you think of the suite as a portfolio with costs, and can remove tests without losing real coverage.
Answer frame:

Pick: tests that check logic the UI only displays, slow ones and ones that never catch bugs.

Replace: cover the same rule lower down, keep one UI test that proves it shows.

Convince: run time, failure history and a coverage map, not opinion.

Sample spoken answer:

“We had about forty UI tests on a signup form, each typing a bad value and checking the error message. The rules lived on the server, so I proposed moving them to API tests and keeping two UI tests: one that shows an error appears in the right place, and one happy path. To convince the team, I pulled the history. Those forty tests had failed dozens of times over six months, and every failure was flakiness, not a bug. I also mapped each rule to its new API test so nobody felt coverage had quietly vanished. The team agreed, the API tests ran in seconds, and the UI suite got shorter and calmer. The trade-off is that a bug where the page shows the wrong message for one rule could slip, and we accepted that because the message text comes from one place.”

Red flag to avoid:

Deleting tests to speed up the build without replacing what they covered or telling anyone.

They may ask next:
  • Which kinds of UI tests would you never move down?
  • How do you keep the suite from growing back to the same size?
Say it in 60 seconds
Hard Behavioral round Mid-level, Senior Practice question

4. You inherited a Selenium suite of about fifteen hundred tests from a team that left, and nobody knows what half of them cover. What did you do in the first month?

What the interviewer is really testing:
Whether you can take over an unknown suite calmly: learn it, make a trusted core reliable and decide what to keep before adding anything new.
Answer frame:

Learn: run it several times, map tests to features and pull the failure history.

Stabilise: get a trusted core reliably green before touching the rest.

Decide: keep, fix or delete, agreed with the people who own each feature.

Sample spoken answer:

“When that happened to me, I didn't rewrite anything for the first few weeks. I ran the full suite several times and recorded what passed, what failed and how long each test took. Then I mapped tests to features from their names and the screens they touched, and sat with each feature's product owner for half an hour. That showed a lot of duplication: several tests checked the same login and search paths in slightly different ways. I picked a core of a couple of hundred tests on the journeys people cared about and got those reliably green first, so the pipeline meant something again. The rest I sorted into keep, fix or delete, and deleted a good chunk with the owners' agreement. The trade-off was a month with almost no new automation, and I agreed that with my manager on day one so it didn't look like I was doing nothing.”

Red flag to avoid:

Announcing a full rewrite in week one, before knowing what the old suite actually protects.

They may ask next:
  • How did you decide a test was safe to delete?
  • What would you have done if the framework code itself was the real problem?
Say it in 60 seconds

Grid and CI 4 questions

Hard Technical round Mid-level, Senior Practice question

5. You had to choose where the tests run: your own Grid in containers or a hosted browser cloud. What did you pick, and what did it cost you?

What the interviewer is really testing:
Whether you made the infrastructure call from speed, coverage and upkeep, and can say what you took on by choosing it.
Answer frame:

Needs: which browsers, how many parallel sessions, and whether the app is reachable from outside.

Split: often your own Grid for the common browsers and a hosted service for the rest.

Cost: image upgrades, capacity and on-call work you now own.

Sample spoken answer:

“We ended up with a split. Our staging app sat inside a private network, and most runs were Chrome and Firefox, so I set up our own Selenium Grid using the official Docker images, with nodes that scale up during working hours. Tests ran close to the app, which cut network lag, and we controlled browser versions exactly. For Safari and real mobile devices we kept a hosted cloud, but only for the nightly run, because its concurrency limit would have queued our pull request builds. The cost of owning a Grid was real: I had to rebuild images when browsers updated, watch node memory, and I became the person paged when the Grid itself was down. I'd make the same call again, but I'd write the runbook on day one instead of month three.”

Red flag to avoid:

Picking an option because it was already there, with no view of browser coverage, network access or who maintains it.

They may ask next:
  • How did you decide how many sessions each node could safely run?
  • What would push you to move everything to the hosted option?
Say it in 60 seconds
Hard Behavioral round Mid-level, Senior Practice question

6. You doubled the parallel threads on your suite and it got slower and flakier. What was the real bottleneck, and how did you find it?

What the interviewer is really testing:
Whether you look past the test code to machines and the app under test when scaling parallel runs.
Answer frame:

Measure: CPU and memory on Grid nodes, and response times of the test environment.

Find the limit: browsers starved of CPU, or a backend that cannot take the load.

Fix and cost: cap sessions per node, add capacity, accept a practical ceiling.

Sample spoken answer:

“We went from eight to sixteen threads expecting half the run time, and instead it got longer and timeouts shot up. The test code was fine, since it already had one driver per thread. So I looked at the machines. Our Grid nodes allowed eight sessions each, but each node only had four CPUs, and a busy Chrome needs most of a core, so pages were rendering slowly and waits ran out. At the same time, the staging database was hitting its connection limit because sixteen tests were creating data at once. I capped sessions per node to match the CPU, added two nodes, and asked the platform team to raise the staging pool. We landed at twelve threads, which was the fastest stable point. The lesson: the Grid and the app under test are part of the suite's performance.”

Red flag to avoid:

Adding more threads until things break and then blaming the tests, without ever looking at node or backend load.

They may ask next:
  • How did you decide twelve was the right number and not fourteen?
  • What would you monitor so the next person does not repeat this?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

7. Which browsers did you run on every pull request, which only nightly, and how did you make that call?

What the interviewer is really testing:
Whether your browser coverage is based on real usage and past bugs rather than testing everything everywhere.
Answer frame:

Evidence: product analytics for real browser usage, plus past browser-specific bugs.

Split: the main browser on every pull request, the rest nightly or before release.

Cost: a bug in a less-used browser is found a day later, not in the pull request.

Sample spoken answer:

“I asked product for browser usage from analytics and went through a year of bug tickets tagged by browser. Chrome was most of our traffic and caught nearly every functional bug, so pull requests ran the smoke group on Chrome only, which kept feedback fast. Firefox and Edge ran the full regression every night, and Safari ran nightly on a hosted service because our own Grid had no Macs. The browser-specific bugs we'd had were mostly layout and date inputs on Safari, so I made sure the nightly Safari run covered those screens. The trade-off came up once: a Firefox-only bug slipped past a pull request and was caught the next morning instead. We agreed that a day's delay on a smaller browser was fine, and we'd revisit the split if usage shifted.”

Red flag to avoid:

Running every browser on every commit without thinking about cost, or testing one browser only with no evidence it is enough.

They may ask next:
  • What would make you add a second browser to the pull request run?
  • How do you handle a test that only fails on one browser?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

8. Which Selenium tests block a merge, which block a release, and how is that split set up in your framework?

What the interviewer is really testing:
Whether you have designed where UI tests sit in the pipeline so they give fast feedback without slowing every change.
Answer frame:

Merge gate: a short smoke group on the critical user journeys.

Release gate: the full regression, nightly and before a release.

Setup: TestNG groups chosen at run time, so the same tests serve both.

Sample spoken answer:

“On my last project the merge gate was a smoke group of about thirty tests covering the journeys that make money or lock people out: sign up, log in, search, checkout and password reset. It ran in parallel in about twelve minutes on every pull request. The full regression of several hundred tests ran nightly and again on the release candidate, and a release couldn't go out with a red regression unless the product owner signed off on each failure. In the framework, every test carries TestNG groups, so the pipeline just asks for smoke or regression and the same code serves both. The trade-off is that a bug outside the smoke paths is found the next morning, not in the pull request. I review the smoke list every quarter so it stays short and actually matches what matters.”

Code:
@Test(groups = {"smoke", "regression"})
public void userCanCheckOutWithSavedCard() {
    // critical journey: runs on every pull request
}

@Test(groups = {"regression"})
public void userCanEditGiftMessageAfterOrder() {
    // nightly and before release only
}
Red flag to avoid:

Running the full multi-hour regression on every pull request, or having no UI tests in the merge gate at all.

They may ask next:
  • How do you decide what earns a place in the smoke group?
  • What happens when the smoke group creeps past your time budget?
Say it in 60 seconds

Selenium 4 2 questions

Medium Technical round Mid-level, Senior Practice question

9. Selenium Manager can fetch drivers, and even browsers, on its own. Did you let it do that in CI or pin versions yourself, and why?

What the interviewer is really testing:
Whether you set a deliberate browser and driver version policy for CI, instead of leaving it to whatever gets downloaded on the day.
Answer frame:

Local: let Selenium Manager match the driver to the installed browser, so setup is painless.

CI: fix the browser and driver versions in the image, so Monday's run matches Friday's.

Cost: you test a browser a little behind users, so upgrades become a planned task with an early-warning run.

Sample spoken answer:

“Locally, I let Selenium Manager do its job. New joiners install Chrome, run the suite, and it finds or downloads the matching driver, so we deleted our old driver setup page. In CI I made the opposite call. Our agents sat behind a proxy, and a run that downloads things at start-up can fail for reasons that have nothing to do with the app. Worse, if the browser moves on its own, a red build on Tuesday might mean a new browser, not a new bug. So CI runs in images with the browser and driver versions fixed, and nothing downloads during the run. The cost is that we sometimes test a browser a few weeks behind what users have. To cover that, a weekly job runs the smoke tests on the newest stable browser, and upgrading the image is a planned monthly task, not a surprise.”

Red flag to avoid:

Letting CI pick up whatever browser version happens to be installed that day, then chasing the same version break every few weeks.

They may ask next:
  • What would you check first if only the Firefox runs broke right after an image upgrade?
  • How do you balance a pinned browser against testing what users actually run?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

10. When did you use the Chrome DevTools features in Selenium 4, like network interception, and what did that choice cost you?

What the interviewer is really testing:
Whether you used the newer Selenium 4 abilities for a real reason and understand their browser and version limits.
Answer frame:

Use case: stub a slow or flaky third-party call, or catch console errors.

How: a DevTools session from the Chromium driver, scoped to specific tests.

Cost: Chromium only, and the DevTools classes are tied to browser versions.

Sample spoken answer:

“Our checkout called a third-party address lookup that was slow and sometimes down in staging, which made a dozen tests fail for reasons outside our control. I used the Selenium 4 DevTools support on Chrome to intercept that request and return a fixed response, but only in the tests that weren't about the lookup itself. I kept one test that hits the real service, so we'd still notice if the integration broke. I also used DevTools to collect console errors and fail a test if the page threw a JavaScript error. The cost was clear to me going in. It only works on Chromium browsers, so Firefox runs couldn't use it, and the DevTools classes are tied to browser versions, so every browser upgrade meant bumping that dependency too. I'm watching WebDriver BiDi as the cross-browser way to do this.”

Red flag to avoid:

Stubbing every backend call so the UI tests no longer test the real system, or not knowing the feature is Chromium only.

They may ask next:
  • How would you stub that call for the Firefox runs?
  • Where is the line between stubbing a dependency and hiding a real problem?
Say it in 60 seconds

Test Data 1 question

Hard Technical round Mid-level, Senior Practice question

11. Your tests need accounts with realistic history, like years of orders. Did you copy production data, generate it, or seed a fixed set, and what did that choice cost?

What the interviewer is really testing:
Whether you weigh realism against privacy, repeatability and upkeep when you decide where test data comes from.
Answer frame:

Options: a masked copy of production, data generated per test, or a fixed seed loaded before the run.

Choice: how realistic the tests really need it, and who is allowed to hold real customer data.

Cost: a seed goes stale as the schema changes, so it needs an owner.

Sample spoken answer:

“Our order history and reports screens needed accounts with hundreds of orders across several years, and building that per test through the API took far too long. Someone suggested copying production. I said no, because even masked, customer data in a test environment is a risk, and masking tends to miss things like free-text notes. Instead I built a seed: a script that generates a few dozen accounts with realistic patterns, like refunds, cancelled orders and old addresses, and loads them into a fresh database before the nightly run. Tests that only read that history use the seeded accounts. Tests that change data still create their own, so they can't damage the shared seed. The cost is upkeep. Any schema change can break the seed script, so it also runs in the backend team's pipeline, and I'm listed as its owner.”

Red flag to avoid:

Copying real customer data into a test environment for convenience, or building long histories through the UI on every run.

They may ask next:
  • How do you stop a test from quietly changing the seeded accounts other tests rely on?
  • When would a masked copy of production be the right call?
Say it in 60 seconds

Framework Design 5 questions

Medium Coding round Mid-level, Senior Practice question

12. Your app fills each screen through many background requests, and element waits keep passing too early. Show me the page-ready wait you built, and why you chose it.

What the interviewer is really testing:
Whether you can design one shared wait from what the app really does, and know the cost of relying on a hook inside the app.
Answer frame:

Why: document.readyState only covers the first load, not the requests that follow it.

Signal: a counter of open requests the app exposes in test builds, read with JavascriptExecutor.

Cost: it depends on the frontend keeping that hook, and it doesn't replace waiting for the exact element a test checks.

Sample spoken answer:

“Our dashboards load in pieces. The page reports complete almost straight away, then eight or nine requests fill in the widgets. Element waits kept passing on a half-loaded screen, and then a number changed under the assertion. I asked the frontend team to add a tiny counter in test builds: it goes up when a request starts and down when it finishes. Then I wrote one helper that waits until the document is complete and that counter is zero. Page objects call it after navigation and after actions that trigger loading. I chose it over chaining ten element waits because it describes when the page is actually done. The cost is a dependency on that hook, so the helper fails with a clear message if the counter is missing instead of silently waiting. For the exact thing a test checks, I still wait for that element or text.”

Code:
public static void waitForPageReady(WebDriver driver) {
    JavascriptExecutor js = (JavascriptExecutor) driver;
    new WebDriverWait(driver, Duration.ofSeconds(20)).until(d -> {
        Object state = js.executeScript("return document.readyState;");
        if (!"complete".equals(state)) {
            return false;
        }
        Object pending = js.executeScript("return window.__pendingRequests;");
        if (pending == null) {
            throw new IllegalStateException("App has no __pendingRequests hook");
        }
        return ((Number) pending).longValue() == 0;
    });
}
Red flag to avoid:

Adding a fixed sleep after every navigation, or trusting document.readyState to mean the data has arrived.

They may ask next:
  • What happens with requests that never finish, like a polling call or a websocket?
  • Why keep that counter out of production builds?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

13. Your checkout page object has grown to two thousand lines. How did you split it, and what trade-off came with the new structure?

What the interviewer is really testing:
Whether you can refactor a framework around reusable components and are honest about the extra structure it adds.
Answer frame:

Find the seams: parts of the page that repeat or change on their own, like forms and panels.

Components: small classes scoped to a root element, composed into the page.

Trade-off: more classes to navigate, and a naming convention the team has to follow.

Sample spoken answer:

“Our checkout page object had every locator and action for five steps in one class, and two people editing it meant merge conflicts every week. I noticed the same address form appeared in checkout, the account page and returns, so I pulled out components: address form, cart summary, payment panel. Each component finds elements inside its own root, so a locator can't accidentally match something elsewhere on the page. The checkout page became a thin class that hands out those components. Duplicate code dropped, and a change to the address form now happens in one place. The trade-off was more files, and new testers had to learn where things live, so I wrote a one-page guide and a template. I also look up the root each time rather than storing it, so a re-rendered form doesn't go stale.”

Code:
public class AddressForm {
    private final WebDriver driver;
    private final By rootBy = By.cssSelector("[data-testid='address-form']");

    public AddressForm(WebDriver driver) { this.driver = driver; }

    private WebElement root() { return driver.findElement(rootBy); }

    public AddressForm postcode(String code) {
        root().findElement(By.name("postcode")).sendKeys(code);
        return this;
    }
}
Red flag to avoid:

Splitting the class by line count alone, or keeping a stored WebElement root that goes stale after the page re-renders.

They may ask next:
  • Why look up the root element each time instead of keeping it in a field?
  • How do you stop components from depending on each other?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

14. How does your framework switch between test environments and keep credentials out of the code?

What the interviewer is really testing:
Whether your framework is configurable without code edits, and whether you treat test credentials as real secrets.
Answer frame:

One switch: an environment name passed in at run time picks a config file.

Secrets: passwords and tokens come from the CI secret store as environment variables.

Safety: fail fast with a clear message if something is missing.

Sample spoken answer:

“When I joined, URLs were hard-coded and a test password sat in a Java class in the repo. I changed it so the run picks its environment from one system property, defaulting to staging, and that loads a properties file with the URLs and timeouts for that environment. Anything secret, like passwords and API tokens, never lives in the repo: CI injects it from its secret store as environment variables, and locally people keep it in a file that git ignores. At startup the framework checks every required value and fails in the first second with a message naming what's missing, instead of a confusing login failure ten minutes in. We also rotated that old password, since it was in the git history. The cost is a bit more setup for new joiners, so I put it in the readme.”

Code:
String env = System.getProperty("env", "staging");
Properties props = new Properties();
try (InputStream in = getClass().getResourceAsStream("/config/" + env + ".properties")) {
    if (in == null) {
        throw new IllegalStateException("No config for environment: " + env);
    }
    props.load(in);
}
String password = System.getenv("TEST_USER_PASSWORD"); // injected by CI
if (password == null) {
    throw new IllegalStateException("TEST_USER_PASSWORD is not set");
}
Red flag to avoid:

Keeping credentials in the repo because they are only test accounts, or editing code to switch environments.

They may ask next:
  • Why rotate a password that was only ever used in tests?
  • How do new joiners get the secrets they need locally without anyone pasting them in chat?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

15. The product runs feature flags and A/B experiments, so the same page looks different from one run to the next. How did you make your suite handle that?

What the interviewer is really testing:
Whether you control variation on purpose in tests, instead of writing locators and checks that try to cope with every version.
Answer frame:

Control it: each test sets its flag or experiment variant, through an override the flag service honours for test accounts.

Cover both: run the key journeys on each live variant, not only the default.

Clean up: delete the old variant's tests when the flag is removed.

Sample spoken answer:

“We had tests failing at random because a new checkout layout was on for some sessions and off for others. My first rule was that no test leaves a variant to chance. We agreed with the platform team that the flag service would honour an override for test accounts, so each test says which variant it wants during setup, and the framework applies it before the first page loads. By default tests get whatever the release would ship, and the checkout journey runs once per live variant. I turned down the idea of page objects with if-else branches that guess which layout appeared, because those hide real bugs. The cost is bookkeeping: every flag adds tests, so each flag-specific test is tagged with the flag name, and when a flag is removed its old tests go in the same sprint. Without that, the suite fills up with dead branches.”

Red flag to avoid:

Writing checks that accept either layout, or letting tests land in random variants and calling the failures flaky.

They may ask next:
  • How would you test a variant that is only switched on for a small slice of real users?
  • What do you do when someone flips a flag in production that your suite never ran with?
Say it in 60 seconds
Hard System design round Mid-level, Senior Practice question

16. Did you build your own layer on top of Selenium or adopt an existing open-source framework? What did you choose, and what did it cost?

What the interviewer is really testing:
Whether you made a build-or-adopt choice from team skills and needs, and understand the long-term upkeep either way.
Answer frame:

Needs: team language skills, Grid control, reporting and how custom the app is.

Choice: a thin in-house layer or an existing framework, with the reason.

Cost: who maintains waits, reporting and upgrades from then on.

Sample spoken answer:

“My team was five testers who knew Java well, and we needed tight control over the Grid, test data and reporting. I looked at a couple of open-source frameworks that wrap Selenium with built-in waits and reports. They'd have given us a fast start, but they hid the driver in ways that made our custom Grid setup awkward, and upgrades would depend on their release cycle. So I built a thin layer ourselves: a driver factory, a small set of wait helpers, base components for page objects, and a failure listener. Everything else stayed plain Selenium and TestNG, so new joiners could use normal documentation. The cost is that we own that code. When Selenium 4 arrived, the upgrade was ours to do. I kept the layer small on purpose, because every helper we add is one more thing to maintain.”

Red flag to avoid:

Building a large custom framework nobody else understands, or adopting one without checking it fits how you run tests.

They may ask next:
  • What would make you choose the ready-made framework next time?
  • How do you stop a thin layer from growing into a heavy one?
Say it in 60 seconds

Code Review and Mentoring 2 questions

Medium Behavioral round Mid-level, Senior Practice question

17. A junior's pull request adds five new Selenium tests. Walk me through what you look for in review and what you'd send back.

What the interviewer is really testing:
Whether you review test code with the same care as product code and can teach through review instead of just rejecting.
Answer frame:

Level: should each test be a UI test at all, or would an API test do?

Code: sleeps, locators outside page objects, asserts inside them, shared data.

Tone: explain the why, fix one example together, approve the rest.

Sample spoken answer:

“First I ask whether each test belongs at the UI level. In one review, two of the five were checking server validation messages that an API test could cover in a second, so I suggested moving them. Then I read the code. That PR had a Thread.sleep of three seconds, which I asked to replace with a wait for the actual condition. It had assertions inside the page object, which hides what the test checks, so I asked to move them into the test. And every test used the same hard-coded user, which would clash once we ran in parallel. I don't leave twenty comments. I paired with him for half an hour on the first test, then he fixed the rest himself. Next time his PR came in clean, and later he started catching the same things in others' reviews.”

Red flag to avoid:

Approving test code without reading it because it is only tests, or rewriting it yourself without explaining anything.

They may ask next:
  • What would you do if the junior pushed back and said the sleep made the test pass?
  • Which review checks did you turn into automated lint rules?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

18. A tester you mentor keeps writing long absolute XPaths copied from the browser. How did you coach them out of it?

What the interviewer is really testing:
Whether you mentor by building understanding and habits, not by rewriting their locators for them.
Answer frame:

Show the break: let them see one of their locators fail after a harmless layout change.

Give an order: test id, id, name, attribute-based CSS, then short relative XPath.

Make it stick: a review check, then let them review others.

Sample spoken answer:

“Instead of rewriting their locators, I waited for the next small UI change, when a designer wrapped a section in an extra div, and sat with them while their tests broke. That made the point better than any rule. Then I gave them a short order to follow: a test id if there is one, then id or name, then a CSS selector on stable attributes, and only then a short relative XPath, for example one that finds a row by its text. We rewrote one page object together, and they did the next two alone. I also added a simple check in review that flags any locator starting from html, so they got feedback without waiting for me. A couple of months later they were the one leaving that comment on other people's pull requests, which is when I knew it had stuck.”

Red flag to avoid:

Quietly fixing their locators yourself so they never learn, or banning XPath entirely without explaining why.

They may ask next:
  • What would you do if the app has no test ids and developers won't add them?
  • How do you mentor someone who is senior to you in years but not in automation?
Say it in 60 seconds

Production Incidents 2 questions

Hard Behavioral round Mid-level, Senior Practice question

19. Tell me about a time your automated suite caused real damage outside the test, like sending real emails or wiping data someone else needed. What did you change?

What the interviewer is really testing:
Whether you own your suite's side effects in shared systems and fix them with guards in the framework, not a promise to be careful.
Answer frame:

What happened: the side effect, how you found out and how far it spread.

Contain: pause the job, tell the people affected and help clean up.

Guard: checks in the framework that make the mistake impossible, not just unlikely.

Sample spoken answer:

“At my last company a nightly run on a newly built staging environment sent a few hundred order emails to real addresses. Someone had loaded that staging with production users, and its email service wasn't pointed at a sandbox. We found out from a customer's reply. I paused the nightly job, told support and my lead within the hour, and helped list who got what so support could write back. Then I added guards to the framework itself. At start-up it checks the target environment against an allow list and refuses to run anywhere unknown. Every test account now uses an address on a domain we own, and a smoke check confirms the email service is in sandbox mode before any order test runs. The trade-off is that a run fails fast whenever someone stands up a new environment, until it's added to the list, and I'm fine with that.”

Red flag to avoid:

Blaming the environment team and changing nothing in the framework.

They may ask next:
  • Who should own the rule that test environments never hold real customer data?
  • How would you explain this incident to a manager outside engineering?
Say it in 60 seconds
Hard Behavioral round Mid-level, Senior Practice question

20. CI agents kept running out of memory because of leftover browser processes. How did you trace it and fix it?

What the interviewer is really testing:
Whether you can trace a resource leak to the framework's lifecycle code and fix it at every exit path.
Answer frame:

Evidence: process lists on the agent, matched to which jobs left them.

Causes: teardown skipped after a setup failure, close instead of quit, jobs killed on timeout.

Fix: teardown that always runs, plus a cleanup step the framework cannot skip.

Sample spoken answer:

“Our CI agents slowed down over a day and eventually failed builds with out-of-memory errors. I listed processes on one agent and found dozens of Chrome and ChromeDriver processes from finished jobs. Matching them to job logs showed three causes. When our before-method setup failed after the browser had started, TestNG skipped the after-method, so quit never ran. A few old tests called close instead of quit, which leaves the driver process running. And jobs killed on timeout never reached teardown at all. I set alwaysRun to true on the teardown and made it null-safe, replaced the stray close calls, and added a final pipeline step that kills any browser process the job started. Longer term we moved runs into short-lived containers, so nothing can survive the job. Agents stayed flat for weeks after.”

Code:
@AfterMethod(alwaysRun = true) // runs even if setup failed
public void tearDown() {
    WebDriver driver = DriverHolder.get();
    if (driver != null) {
        driver.quit();
        DriverHolder.remove();
    }
}
Red flag to avoid:

Rebooting the agents on a schedule and never finding out why processes were left behind.

They may ask next:
  • Why did short-lived containers fix what the teardown changes alone could not?
  • How would you spot this leak before the agents fall over?
Say it in 60 seconds

Debugging Failures 1 question

Medium Technical round Mid-level, Senior Practice question

21. A batch of tests fails only in the last few days of each month, then passes again. How did you track that down, and what did you change in the framework?

What the interviewer is really testing:
Whether you spot time-based causes behind failures and fix them with controlled dates, instead of rerunning until the calendar moves on.
Answer frame:

Pattern: line failures up against the date and time, not just the commit.

Cause: dates picked relative to today that cross into the next month, or a time zone gap between CI and the app.

Fix: one date helper that works in the app's time zone and handles month changes.

Sample spoken answer:

“Our booking tests went red near the end of every month, and people just reran them. When I charted failures by date the pattern was obvious. The tests picked a delivery date as today plus three days and clicked that day in a calendar widget. Near month end that date fell in the next month, so the widget was showing the wrong page and the click missed. A second bug hid behind it: the CI agents ran in UTC while the app worked in the store's local time, so late-evening runs disagreed about what today was. I replaced every hand-made date with one helper that builds dates in the app's time zone and moves the calendar to the right month before clicking. It has its own unit tests for the last day of a month and of a year. The cost is one more shared helper, so review now flags any LocalDate.now in test code.”

Red flag to avoid:

Calling it flaky and rerunning, or hard-coding one date that will be in the past next year.

They may ask next:
  • How would you test the last day of the year without waiting for December?
  • What else in a UI test depends on the clock besides date pickers?
Say it in 60 seconds

Team Decisions 2 questions

Medium Situational round Mid-level, Senior Practice question

22. Your manager wants a new Selenium test added for every bug fixed from now on. How do you respond?

What the interviewer is really testing:
Whether you can push back on a well-meant rule with a better one, instead of either refusing or quietly growing a slow suite.
Answer frame:

Agree with the goal: every fixed bug should have a test that would catch it again.

Better rule: the test goes at the lowest level that catches that bug.

Show the cost: what the rule would do to run time and upkeep.

Sample spoken answer:

“I'd agree with the goal first, because a fixed bug without a regression test can come back. Then I'd suggest a sharper rule: every bug fix gets a test at the lowest level that would have caught it. A wrong tax calculation needs a unit test, a bad API response needs an API test, and only bugs that live in the UI flow, like a button that stops working after going back a step, get a Selenium test. I'd bring numbers from our tracker: most recent bugs were logic bugs that a Selenium test would catch slowly and flakily. I'd also show how long the suite would get if every bug added a UI test. When I proposed this before, the manager agreed, and we added a field to the bug template asking where the regression test lives.”

Red flag to avoid:

Saying yes and letting the UI suite grow without limit, or refusing without offering anything in its place.

They may ask next:
  • Who writes the test when the bug is in backend logic?
  • What would you do if the manager still insisted on UI tests for everything?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

23. The frontend team won't add test ids, and your locators break every sprint. How did you get that changed?

What the interviewer is really testing:
Whether you can influence another team with evidence and a low-cost offer, rather than complaining or working around it forever.
Answer frame:

Evidence: how many hours and failures the broken locators cost last sprint.

Make it cheap: open the first pull request yourself on one screen.

Agree a convention: naming, where ids go, and a review check.

Sample spoken answer:

“I stopped asking and started measuring. For one sprint I logged every test broken by markup changes that didn't change behaviour, and it was most of our failures and a couple of days of my time. I brought that to the frontend lead, along with a pull request I'd written myself adding data-testid attributes to the login and checkout screens. It showed the change was small and didn't affect users. We agreed a naming convention, that new components get ids on the elements tests need, and that removing one needs a heads-up to us. The frontend team also liked that the ids made their own component tests easier. Until each screen was covered, I kept using stable attributes like names and labels. Broken locator failures dropped sharply within two sprints. The cost was my time writing those first pull requests, and it was worth it.”

Red flag to avoid:

Accepting constant breakage as normal, or escalating to management before trying to make the change easy.

They may ask next:
  • What would you do if the frontend lead still said no?
  • Should test ids be stripped from production builds?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card