Selenium interviews for experienced candidates with around five years never ask you to define an explicit wait; they ask why you built something the way you did, what it cost, and how you handled the day it broke. Expect questions on test data, Grid and CI, Selenium 4 features, flaky suites, pushback and mentoring. It is written for automation engineers with roughly five to seven years of Selenium behind them, who own the framework or a large part of it, decide where the tests run, get the message when the pipeline goes red and review other testers' code. Each answer below is a first-person story or a decision you can defend. Swap in your own project details.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Data: store every result with the test id and the commit it ran on.
Definition: a test that passed and failed on the same unchanged code is flaky.
Rule and cost: quarantine with an owner and a deadline, and admit what is unguarded meanwhile.
“We used to argue about which tests were flaky, so I started saving every run's results into a small table: test id, pass or fail, commit, browser and Grid node. Any test that both passed and failed on the same commit counted as flaky, because the code hadn't changed. Once a test failed that way more than twice in fifty runs, a job moved it into a quarantine group. Quarantined tests still ran, but they couldn't block a merge, and each one got a ticket with an owner and a two-sprint deadline. If nobody fixed it by then, we deleted it and wrote down the coverage we'd lost. The trade-off I accepted is that a quarantined test isn't protecting anything, so I kept that list short and visible in the team channel every week.”
Calling a test flaky from memory, or quarantining tests with no owner, no deadline and no record of what stopped being covered.
Own it: a suite people ignore is worth nothing, so noise is the automation team's bug.
Triage: sort every failure into product bug, test bug or environment for a few weeks.
Show it: fix the worst offenders, make failures readable, share the real bugs caught.
“When that happened on my team, I didn't send a reminder to check the builds. I treated it as my bug. For two weeks I looked at every red run myself and sorted each failure into product bug, test bug or environment. Most were test bugs and a shaky test database, which told me exactly where to start. I paused new automation for two sprints, fixed or quarantined the worst tests, and changed failure messages so they say what the user couldn't do, like cannot apply coupon on checkout, instead of a timeout on some locator. Then every week I posted the real product bugs the suite had caught. After about a month, developers were opening failures on their own again. The cost was two sprints with no new coverage, and I explained that trade to my manager up front.”
Blaming developers for not looking, or asking them to rerun until green.
Pick: tests that check logic the UI only displays, slow ones and ones that never catch bugs.
Replace: cover the same rule lower down, keep one UI test that proves it shows.
Convince: run time, failure history and a coverage map, not opinion.
“We had about forty UI tests on a signup form, each typing a bad value and checking the error message. The rules lived on the server, so I proposed moving them to API tests and keeping two UI tests: one that shows an error appears in the right place, and one happy path. To convince the team, I pulled the history. Those forty tests had failed dozens of times over six months, and every failure was flakiness, not a bug. I also mapped each rule to its new API test so nobody felt coverage had quietly vanished. The team agreed, the API tests ran in seconds, and the UI suite got shorter and calmer. The trade-off is that a bug where the page shows the wrong message for one rule could slip, and we accepted that because the message text comes from one place.”
Deleting tests to speed up the build without replacing what they covered or telling anyone.
Learn: run it several times, map tests to features and pull the failure history.
Stabilise: get a trusted core reliably green before touching the rest.
Decide: keep, fix or delete, agreed with the people who own each feature.
“When that happened to me, I didn't rewrite anything for the first few weeks. I ran the full suite several times and recorded what passed, what failed and how long each test took. Then I mapped tests to features from their names and the screens they touched, and sat with each feature's product owner for half an hour. That showed a lot of duplication: several tests checked the same login and search paths in slightly different ways. I picked a core of a couple of hundred tests on the journeys people cared about and got those reliably green first, so the pipeline meant something again. The rest I sorted into keep, fix or delete, and deleted a good chunk with the owners' agreement. The trade-off was a month with almost no new automation, and I agreed that with my manager on day one so it didn't look like I was doing nothing.”
Announcing a full rewrite in week one, before knowing what the old suite actually protects.
Needs: which browsers, how many parallel sessions, and whether the app is reachable from outside.
Split: often your own Grid for the common browsers and a hosted service for the rest.
Cost: image upgrades, capacity and on-call work you now own.
“We ended up with a split. Our staging app sat inside a private network, and most runs were Chrome and Firefox, so I set up our own Selenium Grid using the official Docker images, with nodes that scale up during working hours. Tests ran close to the app, which cut network lag, and we controlled browser versions exactly. For Safari and real mobile devices we kept a hosted cloud, but only for the nightly run, because its concurrency limit would have queued our pull request builds. The cost of owning a Grid was real: I had to rebuild images when browsers updated, watch node memory, and I became the person paged when the Grid itself was down. I'd make the same call again, but I'd write the runbook on day one instead of month three.”
Picking an option because it was already there, with no view of browser coverage, network access or who maintains it.
Measure: CPU and memory on Grid nodes, and response times of the test environment.
Find the limit: browsers starved of CPU, or a backend that cannot take the load.
Fix and cost: cap sessions per node, add capacity, accept a practical ceiling.
“We went from eight to sixteen threads expecting half the run time, and instead it got longer and timeouts shot up. The test code was fine, since it already had one driver per thread. So I looked at the machines. Our Grid nodes allowed eight sessions each, but each node only had four CPUs, and a busy Chrome needs most of a core, so pages were rendering slowly and waits ran out. At the same time, the staging database was hitting its connection limit because sixteen tests were creating data at once. I capped sessions per node to match the CPU, added two nodes, and asked the platform team to raise the staging pool. We landed at twelve threads, which was the fastest stable point. The lesson: the Grid and the app under test are part of the suite's performance.”
Adding more threads until things break and then blaming the tests, without ever looking at node or backend load.
Evidence: product analytics for real browser usage, plus past browser-specific bugs.
Split: the main browser on every pull request, the rest nightly or before release.
Cost: a bug in a less-used browser is found a day later, not in the pull request.
“I asked product for browser usage from analytics and went through a year of bug tickets tagged by browser. Chrome was most of our traffic and caught nearly every functional bug, so pull requests ran the smoke group on Chrome only, which kept feedback fast. Firefox and Edge ran the full regression every night, and Safari ran nightly on a hosted service because our own Grid had no Macs. The browser-specific bugs we'd had were mostly layout and date inputs on Safari, so I made sure the nightly Safari run covered those screens. The trade-off came up once: a Firefox-only bug slipped past a pull request and was caught the next morning instead. We agreed that a day's delay on a smaller browser was fine, and we'd revisit the split if usage shifted.”
Running every browser on every commit without thinking about cost, or testing one browser only with no evidence it is enough.
Merge gate: a short smoke group on the critical user journeys.
Release gate: the full regression, nightly and before a release.
Setup: TestNG groups chosen at run time, so the same tests serve both.
“On my last project the merge gate was a smoke group of about thirty tests covering the journeys that make money or lock people out: sign up, log in, search, checkout and password reset. It ran in parallel in about twelve minutes on every pull request. The full regression of several hundred tests ran nightly and again on the release candidate, and a release couldn't go out with a red regression unless the product owner signed off on each failure. In the framework, every test carries TestNG groups, so the pipeline just asks for smoke or regression and the same code serves both. The trade-off is that a bug outside the smoke paths is found the next morning, not in the pull request. I review the smoke list every quarter so it stays short and actually matches what matters.”
@Test(groups = {"smoke", "regression"})
public void userCanCheckOutWithSavedCard() {
// critical journey: runs on every pull request
}
@Test(groups = {"regression"})
public void userCanEditGiftMessageAfterOrder() {
// nightly and before release only
}
Running the full multi-hour regression on every pull request, or having no UI tests in the merge gate at all.
Local: let Selenium Manager match the driver to the installed browser, so setup is painless.
CI: fix the browser and driver versions in the image, so Monday's run matches Friday's.
Cost: you test a browser a little behind users, so upgrades become a planned task with an early-warning run.
“Locally, I let Selenium Manager do its job. New joiners install Chrome, run the suite, and it finds or downloads the matching driver, so we deleted our old driver setup page. In CI I made the opposite call. Our agents sat behind a proxy, and a run that downloads things at start-up can fail for reasons that have nothing to do with the app. Worse, if the browser moves on its own, a red build on Tuesday might mean a new browser, not a new bug. So CI runs in images with the browser and driver versions fixed, and nothing downloads during the run. The cost is that we sometimes test a browser a few weeks behind what users have. To cover that, a weekly job runs the smoke tests on the newest stable browser, and upgrading the image is a planned monthly task, not a surprise.”
Letting CI pick up whatever browser version happens to be installed that day, then chasing the same version break every few weeks.
Use case: stub a slow or flaky third-party call, or catch console errors.
How: a DevTools session from the Chromium driver, scoped to specific tests.
Cost: Chromium only, and the DevTools classes are tied to browser versions.
“Our checkout called a third-party address lookup that was slow and sometimes down in staging, which made a dozen tests fail for reasons outside our control. I used the Selenium 4 DevTools support on Chrome to intercept that request and return a fixed response, but only in the tests that weren't about the lookup itself. I kept one test that hits the real service, so we'd still notice if the integration broke. I also used DevTools to collect console errors and fail a test if the page threw a JavaScript error. The cost was clear to me going in. It only works on Chromium browsers, so Firefox runs couldn't use it, and the DevTools classes are tied to browser versions, so every browser upgrade meant bumping that dependency too. I'm watching WebDriver BiDi as the cross-browser way to do this.”
Stubbing every backend call so the UI tests no longer test the real system, or not knowing the feature is Chromium only.
Options: a masked copy of production, data generated per test, or a fixed seed loaded before the run.
Choice: how realistic the tests really need it, and who is allowed to hold real customer data.
Cost: a seed goes stale as the schema changes, so it needs an owner.
“Our order history and reports screens needed accounts with hundreds of orders across several years, and building that per test through the API took far too long. Someone suggested copying production. I said no, because even masked, customer data in a test environment is a risk, and masking tends to miss things like free-text notes. Instead I built a seed: a script that generates a few dozen accounts with realistic patterns, like refunds, cancelled orders and old addresses, and loads them into a fresh database before the nightly run. Tests that only read that history use the seeded accounts. Tests that change data still create their own, so they can't damage the shared seed. The cost is upkeep. Any schema change can break the seed script, so it also runs in the backend team's pipeline, and I'm listed as its owner.”
Copying real customer data into a test environment for convenience, or building long histories through the UI on every run.
Why: document.readyState only covers the first load, not the requests that follow it.
Signal: a counter of open requests the app exposes in test builds, read with JavascriptExecutor.
Cost: it depends on the frontend keeping that hook, and it doesn't replace waiting for the exact element a test checks.
“Our dashboards load in pieces. The page reports complete almost straight away, then eight or nine requests fill in the widgets. Element waits kept passing on a half-loaded screen, and then a number changed under the assertion. I asked the frontend team to add a tiny counter in test builds: it goes up when a request starts and down when it finishes. Then I wrote one helper that waits until the document is complete and that counter is zero. Page objects call it after navigation and after actions that trigger loading. I chose it over chaining ten element waits because it describes when the page is actually done. The cost is a dependency on that hook, so the helper fails with a clear message if the counter is missing instead of silently waiting. For the exact thing a test checks, I still wait for that element or text.”
public static void waitForPageReady(WebDriver driver) {
JavascriptExecutor js = (JavascriptExecutor) driver;
new WebDriverWait(driver, Duration.ofSeconds(20)).until(d -> {
Object state = js.executeScript("return document.readyState;");
if (!"complete".equals(state)) {
return false;
}
Object pending = js.executeScript("return window.__pendingRequests;");
if (pending == null) {
throw new IllegalStateException("App has no __pendingRequests hook");
}
return ((Number) pending).longValue() == 0;
});
}
Adding a fixed sleep after every navigation, or trusting document.readyState to mean the data has arrived.
Find the seams: parts of the page that repeat or change on their own, like forms and panels.
Components: small classes scoped to a root element, composed into the page.
Trade-off: more classes to navigate, and a naming convention the team has to follow.
“Our checkout page object had every locator and action for five steps in one class, and two people editing it meant merge conflicts every week. I noticed the same address form appeared in checkout, the account page and returns, so I pulled out components: address form, cart summary, payment panel. Each component finds elements inside its own root, so a locator can't accidentally match something elsewhere on the page. The checkout page became a thin class that hands out those components. Duplicate code dropped, and a change to the address form now happens in one place. The trade-off was more files, and new testers had to learn where things live, so I wrote a one-page guide and a template. I also look up the root each time rather than storing it, so a re-rendered form doesn't go stale.”
public class AddressForm {
private final WebDriver driver;
private final By rootBy = By.cssSelector("[data-testid='address-form']");
public AddressForm(WebDriver driver) { this.driver = driver; }
private WebElement root() { return driver.findElement(rootBy); }
public AddressForm postcode(String code) {
root().findElement(By.name("postcode")).sendKeys(code);
return this;
}
}
Splitting the class by line count alone, or keeping a stored WebElement root that goes stale after the page re-renders.
One switch: an environment name passed in at run time picks a config file.
Secrets: passwords and tokens come from the CI secret store as environment variables.
Safety: fail fast with a clear message if something is missing.
“When I joined, URLs were hard-coded and a test password sat in a Java class in the repo. I changed it so the run picks its environment from one system property, defaulting to staging, and that loads a properties file with the URLs and timeouts for that environment. Anything secret, like passwords and API tokens, never lives in the repo: CI injects it from its secret store as environment variables, and locally people keep it in a file that git ignores. At startup the framework checks every required value and fails in the first second with a message naming what's missing, instead of a confusing login failure ten minutes in. We also rotated that old password, since it was in the git history. The cost is a bit more setup for new joiners, so I put it in the readme.”
String env = System.getProperty("env", "staging");
Properties props = new Properties();
try (InputStream in = getClass().getResourceAsStream("/config/" + env + ".properties")) {
if (in == null) {
throw new IllegalStateException("No config for environment: " + env);
}
props.load(in);
}
String password = System.getenv("TEST_USER_PASSWORD"); // injected by CI
if (password == null) {
throw new IllegalStateException("TEST_USER_PASSWORD is not set");
}
Keeping credentials in the repo because they are only test accounts, or editing code to switch environments.
Control it: each test sets its flag or experiment variant, through an override the flag service honours for test accounts.
Cover both: run the key journeys on each live variant, not only the default.
Clean up: delete the old variant's tests when the flag is removed.
“We had tests failing at random because a new checkout layout was on for some sessions and off for others. My first rule was that no test leaves a variant to chance. We agreed with the platform team that the flag service would honour an override for test accounts, so each test says which variant it wants during setup, and the framework applies it before the first page loads. By default tests get whatever the release would ship, and the checkout journey runs once per live variant. I turned down the idea of page objects with if-else branches that guess which layout appeared, because those hide real bugs. The cost is bookkeeping: every flag adds tests, so each flag-specific test is tagged with the flag name, and when a flag is removed its old tests go in the same sprint. Without that, the suite fills up with dead branches.”
Writing checks that accept either layout, or letting tests land in random variants and calling the failures flaky.
Needs: team language skills, Grid control, reporting and how custom the app is.
Choice: a thin in-house layer or an existing framework, with the reason.
Cost: who maintains waits, reporting and upgrades from then on.
“My team was five testers who knew Java well, and we needed tight control over the Grid, test data and reporting. I looked at a couple of open-source frameworks that wrap Selenium with built-in waits and reports. They'd have given us a fast start, but they hid the driver in ways that made our custom Grid setup awkward, and upgrades would depend on their release cycle. So I built a thin layer ourselves: a driver factory, a small set of wait helpers, base components for page objects, and a failure listener. Everything else stayed plain Selenium and TestNG, so new joiners could use normal documentation. The cost is that we own that code. When Selenium 4 arrived, the upgrade was ours to do. I kept the layer small on purpose, because every helper we add is one more thing to maintain.”
Building a large custom framework nobody else understands, or adopting one without checking it fits how you run tests.
Level: should each test be a UI test at all, or would an API test do?
Code: sleeps, locators outside page objects, asserts inside them, shared data.
Tone: explain the why, fix one example together, approve the rest.
“First I ask whether each test belongs at the UI level. In one review, two of the five were checking server validation messages that an API test could cover in a second, so I suggested moving them. Then I read the code. That PR had a Thread.sleep of three seconds, which I asked to replace with a wait for the actual condition. It had assertions inside the page object, which hides what the test checks, so I asked to move them into the test. And every test used the same hard-coded user, which would clash once we ran in parallel. I don't leave twenty comments. I paired with him for half an hour on the first test, then he fixed the rest himself. Next time his PR came in clean, and later he started catching the same things in others' reviews.”
Approving test code without reading it because it is only tests, or rewriting it yourself without explaining anything.
Show the break: let them see one of their locators fail after a harmless layout change.
Give an order: test id, id, name, attribute-based CSS, then short relative XPath.
Make it stick: a review check, then let them review others.
“Instead of rewriting their locators, I waited for the next small UI change, when a designer wrapped a section in an extra div, and sat with them while their tests broke. That made the point better than any rule. Then I gave them a short order to follow: a test id if there is one, then id or name, then a CSS selector on stable attributes, and only then a short relative XPath, for example one that finds a row by its text. We rewrote one page object together, and they did the next two alone. I also added a simple check in review that flags any locator starting from html, so they got feedback without waiting for me. A couple of months later they were the one leaving that comment on other people's pull requests, which is when I knew it had stuck.”
Quietly fixing their locators yourself so they never learn, or banning XPath entirely without explaining why.
What happened: the side effect, how you found out and how far it spread.
Contain: pause the job, tell the people affected and help clean up.
Guard: checks in the framework that make the mistake impossible, not just unlikely.
“At my last company a nightly run on a newly built staging environment sent a few hundred order emails to real addresses. Someone had loaded that staging with production users, and its email service wasn't pointed at a sandbox. We found out from a customer's reply. I paused the nightly job, told support and my lead within the hour, and helped list who got what so support could write back. Then I added guards to the framework itself. At start-up it checks the target environment against an allow list and refuses to run anywhere unknown. Every test account now uses an address on a domain we own, and a smoke check confirms the email service is in sandbox mode before any order test runs. The trade-off is that a run fails fast whenever someone stands up a new environment, until it's added to the list, and I'm fine with that.”
Blaming the environment team and changing nothing in the framework.
Evidence: process lists on the agent, matched to which jobs left them.
Causes: teardown skipped after a setup failure, close instead of quit, jobs killed on timeout.
Fix: teardown that always runs, plus a cleanup step the framework cannot skip.
“Our CI agents slowed down over a day and eventually failed builds with out-of-memory errors. I listed processes on one agent and found dozens of Chrome and ChromeDriver processes from finished jobs. Matching them to job logs showed three causes. When our before-method setup failed after the browser had started, TestNG skipped the after-method, so quit never ran. A few old tests called close instead of quit, which leaves the driver process running. And jobs killed on timeout never reached teardown at all. I set alwaysRun to true on the teardown and made it null-safe, replaced the stray close calls, and added a final pipeline step that kills any browser process the job started. Longer term we moved runs into short-lived containers, so nothing can survive the job. Agents stayed flat for weeks after.”
@AfterMethod(alwaysRun = true) // runs even if setup failed
public void tearDown() {
WebDriver driver = DriverHolder.get();
if (driver != null) {
driver.quit();
DriverHolder.remove();
}
}
Rebooting the agents on a schedule and never finding out why processes were left behind.
Pattern: line failures up against the date and time, not just the commit.
Cause: dates picked relative to today that cross into the next month, or a time zone gap between CI and the app.
Fix: one date helper that works in the app's time zone and handles month changes.
“Our booking tests went red near the end of every month, and people just reran them. When I charted failures by date the pattern was obvious. The tests picked a delivery date as today plus three days and clicked that day in a calendar widget. Near month end that date fell in the next month, so the widget was showing the wrong page and the click missed. A second bug hid behind it: the CI agents ran in UTC while the app worked in the store's local time, so late-evening runs disagreed about what today was. I replaced every hand-made date with one helper that builds dates in the app's time zone and moves the calendar to the right month before clicking. It has its own unit tests for the last day of a month and of a year. The cost is one more shared helper, so review now flags any LocalDate.now in test code.”
Calling it flaky and rerunning, or hard-coding one date that will be in the past next year.
Agree with the goal: every fixed bug should have a test that would catch it again.
Better rule: the test goes at the lowest level that catches that bug.
Show the cost: what the rule would do to run time and upkeep.
“I'd agree with the goal first, because a fixed bug without a regression test can come back. Then I'd suggest a sharper rule: every bug fix gets a test at the lowest level that would have caught it. A wrong tax calculation needs a unit test, a bad API response needs an API test, and only bugs that live in the UI flow, like a button that stops working after going back a step, get a Selenium test. I'd bring numbers from our tracker: most recent bugs were logic bugs that a Selenium test would catch slowly and flakily. I'd also show how long the suite would get if every bug added a UI test. When I proposed this before, the manager agreed, and we added a field to the bug template asking where the regression test lives.”
Saying yes and letting the UI suite grow without limit, or refusing without offering anything in its place.
Evidence: how many hours and failures the broken locators cost last sprint.
Make it cheap: open the first pull request yourself on one screen.
Agree a convention: naming, where ids go, and a review check.
“I stopped asking and started measuring. For one sprint I logged every test broken by markup changes that didn't change behaviour, and it was most of our failures and a couple of days of my time. I brought that to the frontend lead, along with a pull request I'd written myself adding data-testid attributes to the login and checkout screens. It showed the change was small and didn't affect users. We agreed a naming convention, that new components get ids on the elements tests need, and that removing one needs a heads-up to us. The frontend team also liked that the ids made their own component tests easier. Until each screen was covered, I kept using stable attributes like names and labels. Broken locator failures dropped sharply within two sprints. The cost was my time writing those first pull requests, and it was worth it.”
Accepting constant breakage as normal, or escalating to management before trying to make the change easy.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.