Selenium interviews for 10+ years of experience skip waits and locators and go after what only experience teaches: how WebDriver behaves underneath, why the obvious fix is wrong, how browser infrastructure holds up at company scale, how you move a large suite without breaking trust, and how you set standards and grow people. It is written for test automation engineers with around eight to fifteen years behind them, interviewing for senior, lead or architect roles. Each question shows what the interviewer is really checking, a shape for your answer and a sample you can adapt. Say your own version out loud with a story from your own work.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Protocol change: Selenium 4 dropped the old JSON Wire protocol and speaks only W3C WebDriver.
The rule: a capability that is not one of the standard ones must carry a vendor prefix with a colon, like goog:chromeOptions or myteam:build.
The fix: move browser settings into the Options classes, rename custom keys with a prefix, and update whatever reads them on the Grid or reporting side.
“Selenium 3 sent capabilities in both the old JSON Wire shape and the W3C shape, so a loose key could ride along in the legacy part and nobody noticed. Selenium 4 speaks only W3C, and the spec says any capability that isn't a standard one, like browserName or platformName, must have a vendor prefix with a colon. Depending on the binding and driver, a bare custom key is either refused before the request is sent or rejected by the driver as an invalid argument, so the session fails before a single test runs. In our case keys like build and testName, which our dashboards read, had to become something like myteam:build. Browser settings belong in ChromeOptions or FirefoxOptions, which write goog:chromeOptions and moz:firefoxOptions for you. I'd move everything off DesiredCapabilities onto the Options classes, rename the custom keys, and change whatever reads them in the same pull request.”
ChromeOptions options = new ChromeOptions();
options.addArguments("--window-size=1920,1080");
options.setCapability("myteam:build", buildId); // custom keys need a prefix
WebDriver driver = new RemoteWebDriver(gridUrl, options);
Blaming the driver version and downgrading instead of reading the capability rule in the error.
get: normal waits for readyState complete, eager for interactive, and none returns soon after the navigation starts.
Clicks: drivers try to wait for a navigation the click started, using the same strategy, but detection depends on timing.
Single-page routes: a client-side route change isn't a document load, so nothing waits for it.
“get follows the session's page load strategy. Normal, the default, waits until the document's readyState is complete, so the HTML, scripts and images of that first load are done. Eager returns at interactive, once the HTML is parsed. None returns almost as soon as the navigation starts. None of them know about fetch calls the app makes afterwards. A click is different. The spec asks the driver to try to wait if the click started a navigation, using the same strategy, but whether a navigation is seen depends on timing, especially when script starts it a moment later. In a single-page app, clicking a menu item usually changes the route through the history API, which isn't a page load at all, so nothing waits. That's why I never treat get or click returning as the page being ready, and why every step waits for its own element.”
Saying get waits for the whole app to be ready, including its own API calls.
Page load timeout: bounds how long get waits for the load strategy's condition, then throws a timeout even if the page is usable.
The others: the script timeout bounds async JavaScript, and implicit wait only affects finding elements.
Remove the cause: drop or block outside domains in test environments, through a build flag, the browser's network controls or a proxy.
“A session has three timeouts. Implicit wait only affects finding elements, the script timeout bounds async JavaScript, and the page load timeout bounds how long get waits for the load condition. A tag that holds the load event open keeps get waiting until that page load timeout fires and throws. Lowering it makes the failure faster, but you still lose the test, and catching the exception and carrying on is fragile. The better fix is taking outside scripts out of the test environment. Where I could, I got a config flag that drops analytics and chat tags in test builds. Where I couldn't, on Chromium I blocked those domains through the DevTools network commands, and for other browsers we sent traffic through a proxy with a block list. Pages that still carry heavy assets then get the eager strategy.”
ChromeDriver chrome = new ChromeDriver(options);
chrome.executeCdpCommand("Network.enable", Map.of());
chrome.executeCdpCommand("Network.setBlockedURLs",
Map.of("urls", List.of("*analytics.example.com*", "*chat.example.net*")));
chrome.manage().timeouts().pageLoadTimeout(Duration.ofSeconds(30));
Raising the page load timeout until the stall stops showing up in the logs.
The contract: Selenium passes a callback as the last argument, and the command returns only when the script calls it.
The bound: the session's script timeout; if the callback is never called, you get a script timeout error.
Common traps: calling the callback only on the success path, and errors thrown later in the page that never reach Selenium.
“executeScript runs a function body and comes back with whatever it returns. executeAsyncScript is for work that finishes later, like a fetch or a timer in the page. Selenium adds a callback as the last argument, so inside the script it's arguments[arguments.length - 1], and the command returns only when the script calls it. Whatever you pass to that callback comes back to Java. If it's never called, the command waits for the session's script timeout and throws. The trap I've seen most is a script that calls the callback only when a request succeeds. When the request fails, or something throws inside a later then handler, Selenium never hears about it, so a quick error turns into a slow timeout with no clue why. I use it rarely, for things like asking the app whether it's idle, and I always call the callback on both paths.”
driver.manage().timeouts().scriptTimeout(Duration.ofSeconds(10));
Object status = ((JavascriptExecutor) driver).executeAsyncScript(
"const done = arguments[arguments.length - 1];" +
"fetch(arguments[0]).then(r => done(r.status), e => done('error: ' + e));",
healthUrl);
Thinking executeAsyncScript runs in the background while the test carries on.
Input sources: W3C actions model a keyboard and a pointer, and the session remembers which keys and buttons are down.
The leak: a keyDown with no matching keyUp, or an exception before it, leaves the key held for later action chains.
The fix: always pair downs with ups, and reset input state in any teardown that keeps the browser.
“Selenium 4 speaks W3C only, and there Actions builds sequences for input sources, a keyboard and a pointer, whose state the remote end keeps for the whole session. If a test does keyDown on Shift and then fails before its keyUp, or the author simply forgot the keyUp, the session still thinks Shift is held. Later pointer actions in that session are sent with Shift down, so a plain click becomes a range selection, and the failure appears in a test that did nothing wrong. It only bites when a browser is reused across tests, which is why it looks random and passes when run alone. The fix is to always pair each down with an up, and in any teardown that keeps the browser, call the release actions command. In Java that's resetInputState on the driver, which tells the remote end to release anything still pressed.”
new Actions(driver)
.keyDown(Keys.SHIFT)
.click(lastRow)
.keyUp(Keys.SHIFT) // never leave a key held
.perform();
// in teardown when the session is reused
((RemoteWebDriver) driver).resetInputState();
Adding a retry because the failing test changes from run to run.
Opt in: the session asks for the webSocketUrl capability, and the new session reply carries a WebSocket address.
Events: the client subscribes to modules like log and network, and the browser pushes events over that socket as they happen.
On the Grid: the router proxies the WebSocket to the node, so every proxy or load balancer in between must allow WebSocket upgrades.
“Classic WebDriver is one request, one response, so the browser can never tell you something happened. BiDi adds a second channel. When the session is created with the webSocketUrl capability set to true, the reply includes a WebSocket address, and the client connects to it. From then on the client subscribes to events, like log entries or network requests and responses, and the browser pushes them as they happen, on Chrome and Firefox alike. On a Grid, that socket goes to the router, which proxies it through to the node running the browser. So classic commands can work perfectly while BiDi fails, because an ingress or load balancer in front drops WebSocket upgrades or cuts idle connections. Before building listeners on it, I'd prove the path with one test per browser that subscribes to log events through the same Grid address the pipelines use.”
FirefoxOptions options = new FirefoxOptions();
options.setCapability("webSocketUrl", true); // ask for the BiDi channel
WebDriver driver = new RemoteWebDriver(gridUrl, options);
Calling BiDi the DevTools protocol under a new name, or assuming it passes through any proxy.
The parts: a router in front, a new session queue, a distributor that matches requests to node stereotypes, a session map, and an event bus joining them.
Free but unmatched: the request asks for a browser, version or platform no node advertises, so it waits in the queue until it times out.
Free but unreachable: nodes look registered, but the distributor cannot actually reach them, often because of an address that isn't routable.
“In Grid 4 the router takes every request. A new session request goes into the new session queue, the distributor pulls from that queue and matches the requested capabilities against each node's stereotypes, then asks a node to start the session and records it in the session map. The event bus carries registrations and status between the parts. Free slots with a growing queue usually means one of two things. Either the request doesn't match any stereotype, like a browserVersion or platformName nobody runs, so it sits in the queue until the request timeout. Or the nodes look registered but the distributor can't reach them, often because a node advertises a hostname that isn't routable inside the container network. I'd compare the queued request's capabilities with the node stereotypes from the status endpoint, then read the distributor and node logs for connection errors.”
Restarting the hub and adding nodes without looking at what the queued requests actually asked for.
Isolation: one session per short-lived container, so cookies, downloads and crashed browsers never leak into the next test.
Fairness: a concurrency limit per team, capacity that grows with the queue, and a separate pool for merge-gating runs.
Blame the right layer: per-session video and logs linked to the pipeline, plus platform error rates tracked on their own.
“I'd run Grid 4 with the router, queue and distributor as long-lived services, and browsers as short-lived containers running one session each, so nothing leaks between tests. Capacity grows with the new session queue. Fairness needs rules, because one pipeline with a huge parallel setting can starve everyone. The Grid has no idea of teams, so each team gets a concurrency limit, enforced at our own gateway or through separate node pools, and merge-gating runs get their own pool so a nightly job never delays a pull request. The part people skip is attribution. Every session gets video, browser logs and a link to its pipeline and commit. Separately, I track the platform's own errors: failed session starts, time in the queue, sessions lost mid-run and node crashes, per pool. When those rise, it's our problem, and we tell teams before they open a ticket.”
One big shared pool with no isolation, no limits and no way to tell platform failures from test failures.
The cause: a scale-down or node rotation removes a pod that is still running a session, so the next command finds no session.
Drain, don't kill: a drained Grid node stops taking new sessions and finishes its current ones before it exits.
One session per pod: pods that exit after their session make scaling down mostly a matter of not starting new ones.
“The cluster picks pods to remove on its own terms. It doesn't know one of them is halfway through a checkout test, so that browser dies and the test's next command fails because the session is gone. It shows up at night because that's when load drops and the cluster shrinks. The fix is making sure only idle browsers go away. A Grid 4 node can be drained: it stops accepting new sessions, finishes the ones it has, then shuts down. So I'd have the pod's shutdown hook call drain, and set the termination grace period longer than our slowest test. Better still, run each pod for a single session, using the node's option to drain after a set number of sessions, and let an autoscaler that reads the Grid's queue start fresh pods for new requests. Then pods mostly leave when their own session ends, not in the middle of one.”
Adding retries for lost sessions and leaving the scale-down rules alone.
What it does: a script that returns the live property when one exists, falls back to the attribute, and treats boolean attributes specially.
Why it surprises: a property is the current state, like a resolved URL or typed text, while an attribute is what the markup says.
Be explicit: getDomAttribute for the attribute as written, getDomProperty for the live state, both direct W3C commands.
“getAttribute is older than the W3C spec, and it tries to guess what you meant. Under W3C the binding sends a small JavaScript routine that looks for a property with that name first and uses the attribute only if there isn't one, with special handling for boolean attributes, which come back as true or null. The href property is the resolved absolute URL, so that's what you get, even though the markup has a relative path. Same with value: you get what's in the box now, not the default in the HTML. Selenium 4 added getDomAttribute, which returns the attribute as written, and getDomProperty, which returns the live property, and both are plain W3C commands rather than a script. In a framework I'd use those two and treat getAttribute as legacy, so each check says what it means: a link target is an attribute, typed text is a property.”
WebElement link = driver.findElement(By.id("help"));
link.getDomAttribute("href"); // as written, e.g. /help
link.getDomProperty("href"); // resolved absolute URL
link.getAttribute("href"); // guesses: property first, here the full URL
Saying attributes and properties are the same thing, so getAttribute always returns what's in the HTML.
How it's decided: a JavaScript routine checks CSS like display and visibility, size and hidden overflow, not the pixels on screen.
What it misses: an element under an overlay, or with pointer events turned off, still counts as displayed.
Check what matters: wait for the covering element to go, let the click's own hit check fail loudly, and use layout checks for what users see.
“Displayedness isn't something browsers compute natively. The W3C spec only describes it in an appendix, and drivers use a JavaScript routine from the Selenium project that looks at display, visibility, the element's size and whether overflow hides it. It never checks what's actually painted on top. So a button under a cookie banner, a modal backdrop or a sticky header still reports displayed, and so does a button with pointer events turned off. That's why isDisplayed passes and the click then throws an intercepted exception. Even elementToBeClickable only means displayed and enabled, so it won't catch an overlay. For actions, I wait for whatever covers it to disappear, like the spinner or banner, and let the click's own hit check fail loudly. For checks about what a user sees, I add layout or visual assertions rather than trusting isDisplayed alone.”
Treating isDisplayed as proof that a user can see and use the element.
Proxies: initElements fills each field with a proxy, not an element, and nothing is looked up yet.
Every call looks up: each click or getText finds the element again, so re-rendered elements are picked up, and a bad locator fails only at first use.
CacheLookup: keeps the first element found, which saves a lookup but goes stale as soon as that part of the page re-renders.
“initElements doesn't find anything. It fills each annotated field with a proxy, and every time you call a method on it, like click or getText, the proxy runs the lookup again and passes the call to whatever it finds. That's why a field declared once keeps working after the page re-renders, and also why a wrong locator only blows up when the field is first used, not when the page object is built. Lists are proxies too: each call on the list finds the whole list again, but the elements you take out of it are real ones that can go stale. @CacheLookup tells the proxy to keep the first element it found. That saves a round trip on things that never change, like a static header, but on anything the front end re-renders, the cached reference points at a removed node. I'd allow it only on truly static elements, or skip PageFactory and use plain By locators.”
Believing initElements finds every element up front, then blaming timing when a field fails.
Own your data: each test creates what it needs through APIs before the browser opens and never touches another test's records.
Lease the scarce things: accounts or cards you can't create on demand come from a pool with time-limited leases.
Read-only and tagged: heavy reference data is never changed, and created data carries a run id for a sweeper.
“Shared accounts are fine at one thread and fall apart at hundreds, because one test changes the cart or password another test is using. The rule is that a test owns its data: it creates the user or order it needs through APIs before the browser opens. The harder part at scale is data you can't just create, like a small set of single sign-on users, payment test cards or partner sandbox accounts. For those I've built a lease service. A test checks out an account for a few minutes and returns it in teardown, and an expired lease goes back to the pool on its own, so a killed job doesn't lock it forever. The pool size then caps how many of those tests run at once, and the scheduler has to know that. Heavy reference data, like a big catalogue, is built once and read-only. Everything created carries a run id, and a sweeper cleans up by tag.”
Forcing tests into a fixed order so they don't collide on the same accounts.
One library: component objects for each design-system widget, published as a versioned package every suite depends on.
Test it upstream: the library's own tests run against the design system's component showcase, inside the design-system pipeline.
Release together: a library version per design-system release, a changelog, and a way to extend a component instead of forking it.
“Inside one app, component objects are the obvious move. Across twelve teams, the real problem is distribution. If each team copies a date picker class, one markup change still means twelve fixes. So I'd publish the component objects, like the date picker, data table and modal, as a versioned package every suite pulls in, and each app's page objects are built from them. The library's own tests run against the design system's component showcase inside the design-system pipeline, so a change that breaks the date picker's automation fails there, before any app upgrades. Each design-system release ships with a matching library version and a short changelog, and apps move up on their own schedule. For apps that customised a component, the library lets them extend the class instead of forking it. And it has a clear owner, usually my platform team, with design-system engineers as reviewers.”
Letting every team copy the same widget code into its own suite and calling that reuse.
Threads are reused: TestNG runs many tests on the same pool threads, so a ThreadLocal value outlives one test.
The leak: teardown calls quit but never remove, and a lazy getter only creates a driver when the slot is empty.
The fix: quit then remove in teardown, a fresh driver in setup, and driver lifetime matched to the test scope.
“The clue is that it only happens in long runs and always at the first step. TestNG reuses a pool of threads, so a ThreadLocal keeps its value across many tests on the same thread. A common pattern is a getter that creates a driver only if the slot is empty. If teardown calls quit but never remove, the slot still holds the old driver object, pointing at a session that no longer exists. The next test on that thread gets it, and its first command fails with NoSuchSessionException. A similar thing happens when the Grid ends an idle session between a class-level setup and a later method. I fix it by making teardown quit and then remove in a finally, and making setup always create a new driver rather than reusing whatever is in the slot. I also match scopes: a driver created per method is quit per method.”
Wrapping every first step in a retry instead of finding out why the session was dead.
The pressure: what leadership wanted, and why it made sense from their side.
The case: your evidence, the risk in their words, and an option they could say yes to.
The outcome: what happened, including what slipped or what you gave up.
“Leadership wanted my team to start automating a newly acquired product straight away, to show coverage within the quarter. Our own suite was in poor shape: long runs, a flaky Grid, and developers overriding red builds. Adding a second product on top would have doubled the noise. I pulled three months of data: how often builds were overridden, how many failures were infrastructure, and the hours spent on reruns. I put it to our VP in her terms. Right now a red build doesn't stop a bad release, so more tests buy us nothing. Then I offered a deal: six weeks on stability with clear targets for overrides and run time, plus a small smoke set for the new product so it wasn't left bare. She agreed, though not happily. We hit the targets in eight weeks, not six, which I told her early. After that, automating the new product went much faster.”
A story where leadership was simply wrong and you won without any evidence or compromise.
Exercise: a real flaky test with its logs to debug, not a trivia quiz.
Design talk: how they'd structure automation for several teams, pushed on every trade-off.
The senior signal: they talk about causes, data and people, and say what shouldn't be automated.
“I'd skip locator trivia. The core is a working session: a small framework with one flaky test plus its logs and screenshots, and I watch how they find the cause. A mid-level engineer usually adds a wait and reruns. A senior one asks what changed, reproduces it, reads the timing, and can say why the wait they add is the right one. Then a design conversation: how would you structure automation for four teams sharing one app, where do API tests fit, how do you keep the suite fast. I push on each choice to see whether they can defend it or change their mind for a good reason. Last, a talk about influence, like a time they pushed back on what to automate. The senior signal is that they talk about causes, test data and people, and they're comfortable saying some checks shouldn't be UI tests at all.”
Hiring on how many Selenium methods someone can name from memory.
Starting point: what they were strong at, often product knowledge and test design, and what was missing.
What you did: real tasks of growing size, pairing, and code review that taught rather than corrected.
Result: what they owned at the end, shown by what they did without you.
“At my last company one of our manual testers knew the billing product better than anyone but had never written code. I didn't send her on a course and wait. We agreed a plan: first she'd add rows to existing data-driven tests, then fix broken locators, then write new tests on page objects I'd set up, and finally design the page objects for a new screen herself. We paired twice a week for the first two months. In code review I explained why, not just what to change, and I asked her to review my pull requests too, which built her confidence fast. After about nine months she owned the billing suite, and she found a timing bug in our wait helper that I'd missed. The real proof was that she started mentoring the next manual tester using the same plan.”
Describing someone you sent on a course, with nothing you did yourself.
The promise: what you committed to and why it mattered to the business.
Own the cause: the decisions of yours that led there, not only outside factors.
The change: how you recovered, and the habit you keep now.
“I once promised that within a quarter every product team would move onto one shared automation framework I'd designed, so we could retire five home-grown ones. At the end of the quarter two teams had moved, and the rest were still on their old code. The causes were mine. I built it with my own team and showed it to others only when it was finished, so it didn't fit how they worked, like their reporting and their data setup. I set the date without asking what moving would cost each team. And I measured success by teams migrated, not by whether their runs got better. I stopped pushing, sat with the three biggest holdouts and added what they needed. Then we moved the team with the worst flakiness first and shared their before and after numbers, which did more than any mandate. Now I bring teams in before design starts.”
A failure story where every cause was someone else.
What breaks: removed findElementBy helper methods, waits built from plain numbers, DesiredCapabilities habits and unprefixed custom capabilities.
Separate steps: bindings in one change, Grid in another, so any failure points at one cause.
Prove it: run old and new on the same commits and compare results test by test before switching.
“Most code changes are mechanical. The findElementById style helpers are gone, so everything goes through By. Waits and timeouts take a Duration instead of a plain number. DesiredCapabilities gives way to the browser Options classes, and custom capabilities need a vendor prefix. The Grid is the bigger change, because Grid 4 is a different design with different config, so I'd run it as its own project. I'd do the bindings first on a branch, lean on the compiler and a few search-and-replace scripts, then run old and new against the same commits for a week or two and compare results test by test. Only when they match do we switch. Then the Grid the same way: a new Grid 4 beside the old one, pipelines moved over one team at a time, and the old Grid kept until the last team is across.”
Upgrading bindings, drivers and the Grid in one big change on the same day.
Name the problem: what hurts today, like flakiness, speed or debugging time, measured, and whether a new tool or better practice fixes it.
Price it: people and months for the rewrite, the period with two suites, and what goes untested meanwhile.
Prove it small: a time-boxed trial on the worst part of the suite, judged on the same numbers.
“I'd start by asking what problem the rewrite is meant to solve. If it's flaky tests, I'd check how many flakes come from Selenium itself versus test data, environments or waits we wrote badly, because those follow you into any tool. If it's speed or debugging, a newer framework may genuinely help. Then I'd price it: how many tests, how many people for how long, and the period where we maintain two suites. I'd propose a time-boxed trial, porting the thirty or so flakiest tests and measuring flake rate, run time and time to debug a failure against the Selenium versions. To leadership I'd bring one page with the problem, the trial numbers, the cost and a recommendation. That could well be writing only new tests in the new tool and letting the old suite shrink as features change.”
Recommending a full rewrite because the new tool is popular, with no numbers on today's pain.
Report: flake rate, time from commit to UI result, failures that found real bugs, and escaped defects in covered areas.
Refuse to lead with: test count and raw pass rate, since both are easy to game.
Make it act: split everything by team and area so each number has an owner and a next step.
“I'd report four things. The flake rate, meaning failures that passed on rerun with no code change, per team. The time from a commit to a UI result, because slow feedback is why people stop looking. How many failures that month were real bugs, since that's the value we're paying for. And escaped defects, bugs that reached production in areas we claim to cover. I'd push back on test count and raw pass rate as headline numbers. Test count rewards shallow tests, and pass rate goes up nicely when people quietly disable or retry failing ones. Every number would be split by team and feature area, with an owner, so the report leads to a decision, like quarantining a flaky area or moving slow checks down to API tests.”
Leading with the number of automated tests as proof of quality.
What goes wrong: real intermittent bugs get hidden, run time grows, and nobody fixes flaky tests.
Policy: retry only classified infrastructure errors, and report a pass after retry as flaky, not passed.
Close the loop: a flaky test gets an owner and a deadline, then it's fixed or quarantined.
“Blanket retries feel great for a month. Then three things happen. Intermittent product bugs, like a race in checkout that fails one time in five, now pass on the second try, so the suite stops catching exactly the bugs that are hardest to find by hand. Run time creeps up because slow failing tests run several times. And since everything ends green, nobody fixes flaky tests, so their number grows. My policy would be to retry only errors we've classified as infrastructure, like a node dropping the session or a browser that stops answering, never assertion failures. Every retry gets recorded, and a test that passed after a retry shows as flaky in the report. Any test that goes flaky more than a set number of times a week goes to its owner with a deadline, and if it isn't fixed, it leaves the gating run.”
public class InfraRetry implements IRetryAnalyzer {
private int attempts = 0;
@Override
public boolean retry(ITestResult result) {
Throwable t = result.getThrowable();
boolean infra = t instanceof UnreachableBrowserException
|| t instanceof NoSuchSessionException;
return infra && attempts++ < 1; // one retry, infrastructure only
}
}
Saying retries are fine because the build ends green.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.