Scenario rounds for API testers skip the definitions. You get a symptom and a clock: tests go red after an hour, a retry charges a customer twice, the browser shows a CORS error while Postman is green, a report times out only for the biggest accounts. The interviewer wants the order you check things in, the evidence you would collect, and when you would push back. This page is for testers and QA engineers at any level. Each question shows what is being tested, the shape of a good answer and a sample that thinks out loud. Practice saying your first two checks before the fix.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Pattern: failures that start at a steady point in time usually mean something expired, not something broke.
Check: read the token's expiry (the exp claim in a JWT, or expires_in from the login call) and when the suite fetched it.
Fix the suite: fetch or refresh the token when it is close to expiry, instead of once at the start or pasted by hand.
Keep the lesson: add a test that an expired token gets a clean 401, not a 500.
“A clean pass for an hour and then a wall of 401s smells like an expired token, not a bug in the API. First I'd check how the suite gets its token. Often it's fetched once at the start, or someone pasted one into the environment. Then I'd decode it and look at the exp claim, or check expires_in from the login response. If it lives for sixty minutes, that's the whole story. The fix is in the suite: a small step before each request that checks the expiry and gets a new token when there's less than a minute left. I wouldn't just ask for longer-lived tokens in test, because then we stop exercising expiry at all. And I'd keep one deliberate test that sends an expired token and expects a clean 401 with no stack trace.”
Rerunning the suite until it passes, or asking for tokens that never expire in test.
Check the claim: can it be reached from outside, through the gateway, a mis-routed path or a public host name?
Check the design: if other services must call it with credentials, test that calls without them are refused.
Why it matters: internal networks get breached, and one wrong route makes internal public.
If they still say no: write it down as a risk and let the person who owns security accept it.
“I wouldn't argue about it, I'd check it. Internal only is a claim about the network, so I'd try reaching the endpoint from outside: through the public gateway, with small path tricks like a trailing slash or different letter case, and on any public host name that routes to the same service. If I can reach it, the conversation is over. If I can't, I'd still ask how other services are meant to call it. If the design says they use a service token, I'd test that a call without one is refused. I'd explain the reason calmly: networks get breached, and one routing mistake turns an internal admin endpoint into a public one. If the team still decides no auth, that's fine, but I'd write it down as a risk for whoever owns security to accept, not me.”
Agreeing and skipping the tests, or refusing to sign off without checking whether the claim is true.
Reproduce: run the exact customer search and read the full response, not just the status.
Read the tests: most likely they check status and response time only, and a 200 with an empty list passes.
Assert on content: search for a known seeded record and expect it in the results, plus a schema check.
Report both: the search bug, and the gap in the monitor.
“First I'd reproduce it with a real customer query and read the whole response. My guess is we're getting a 200 with an empty results array, or a body that says success false. Then I'd open the monitor's tests. If they only check the status code and response time, an empty list is a pass, so the monitor was green while search was broken. I'd fix that by seeding one product we control and searching for it by name, then asserting it's in the results. For a broad query I'd also assert the list isn't empty, and validate the body against the schema. I'd log two things: the search bug itself, and the monitor gap, because the gap is why customers found it before we did.”
pm.test('search finds the seeded product', () => {
const body = pm.response.json();
pm.expect(body.results).to.be.an('array').that.is.not.empty;
const names = body.results.map(r => r.name);
pm.expect(names).to.include(pm.environment.get('seededProductName'));
});
Blaming the customers' search terms without looking at what the monitor actually asserts.
Why it passed: the developer updated the tests with the code, so the tests described the new API, not what clients rely on.
Schema guard: mark the fields clients read as required in the response schema, owned by QA, not edited casually.
Consumer contracts: the mobile app publishes what it reads; the backend pipeline verifies against it before merge.
Review: flag any rename or removal in the API spec diff as breaking by default.
“The tests passed because they moved with the code. The same pull request that renamed the field also updated the assertion, so our suite described the new API, not what the app depends on. First, I'd help confirm which app versions are affected, because the quick fix is usually to send both names for a while. Then I'd change two things. I'd make the response schema list the fields clients actually read as required, and treat edits to it as a review point, not a side effect. And I'd push for consumer-driven contract tests: the mobile team records what they read, and the backend build fails if a change breaks that. I'd also have the spec diff checked on every pull request, so a rename shows up as a breaking change before anyone merges it.”
Saying the fix is to be more careful next time, with no check that would actually fail.
Freeze v1: its regression suite runs unchanged on every release for the whole six months.
Cross-version data: create in v2 and read in v1, and the other way round, since they share one database.
New values: a new status or field value from v2 must not break a v1 response or client.
Retirement: deprecation notices on v1 and usage tracking now, and a planned end response later.
“Running the v1 suite once isn't enough, because both versions write to the same data. First, I'd freeze the v1 regression suite and run it unchanged on every release for the whole six months, so nobody quietly updates it to match new behaviour. Then the interesting part: create an order through v2 and read it through v1, and the other way round. If v2 adds a new order status, what does v1 return for those orders? An old app might crash on a value it's never seen. If v2 makes a field required, what happens to orders created through v1 without it? I'd also check v1 sends a deprecation notice, and that we can see how much traffic still uses it. And I'd agree now what v1 returns once it's switched off, so that day isn't a surprise.”
Testing v1 and v2 separately and never checking data created by one version and read by the other.
Sort order: check the sort has a unique tie-breaker; rows with equal timestamps can come back in any order.
Moving data: with offset paging, a new or deleted order between requests shifts every row by one.
Test: seed orders with identical timestamps, walk every page, and assert no ID repeats and none is missing.
Suggest: sort by a unique key as well, or move to cursor-based paging.
“There are two usual suspects, and I'd check both. First, the sort. If the list is ordered by created date and several orders share the same timestamp, the database can return those ties in a different order on each query, so one lands on both pages and another on neither. I'd seed twenty-five orders with the same timestamp, walk every page, and collect the IDs. Second, data changing between requests. With page and offset, a new order arriving while you're on page one pushes everything down one place, so you see a repeat. A delete does the opposite and you skip one. I'd test that by inserting an order between page calls. The fixes I'd suggest are adding the order ID as a tie-breaker in the sort, and cursor-based paging for lists that change often.”
seen, page = [], 1
while True:
r = requests.get(f"{BASE}/orders", params={"page": page, "size": 20}, headers=AUTH)
assert r.status_code == 200
items = r.json()["items"]
if not items:
break
seen.extend(o["id"] for o in items)
page += 1
assert len(seen) == len(set(seen)), "an order appeared twice"
assert set(seen) == set(expected_ids), "an order was skipped"
Testing paging with a handful of records that all have different timestamps and calling it done.
Pattern: 'some customers, near midnight' points to time zones, not random data.
Follow the value: what is stored, what the API returns (with an offset or Z, or a bare date), and who converts it.
Test the edges: create orders just before and after midnight in zones ahead of and behind the server.
Settle the contract: timestamps in UTC with a zone marker; date-only fields defined in a named zone.
“Late at night and only some customers tells me it's time zones. I'd follow one order through the system. What's in the database: UTC, or the server's local time? What does the API send back: a full timestamp with a Z or an offset, or a bare date the server already formatted in its own zone? Then who turns it into a date on screen. If the server formats dates in its zone, a customer eight hours behind sees tomorrow. To test it, I'd create orders at 11:30 pm and 12:30 am from clients set to zones ahead of and behind the server, and check the response and any 'orders today' filters. I'd also try the days clocks change. The contract I'd push for is timestamps in UTC with the zone marked, and date-only fields that say which zone they mean.”
Testing only at midday from one machine, where the bug can never show.
Clue: raw text is right but the parsed value is wrong, so it happens during JSON parsing.
Cause: JavaScript numbers are doubles; integers are exact only up to 2 to the power 53, minus 1, which is 16 digits.
Fix: send IDs as strings in the API; the contract says so.
Test: seed IDs above that limit and assert the type is string in the schema.
“If the raw bytes from curl are right but the browser shows something else, the value is changing when it's parsed. JavaScript stores every number as a double, which can only hold whole numbers exactly up to about nine quadrillion, that's 2 to the 53 minus one, sixteen digits. A 19-digit ID gets rounded to the nearest value it can hold, which is why they end in zeros. Then the app requests that rounded ID and gets not found. I'd prove it by parsing the raw response in a browser console and comparing. The fix is on the API: send IDs as strings, since nobody does maths on an ID. For testing, I'd make sure our test data includes IDs above that limit, and have the schema say the ID is a string so a regression fails straight away.”
const raw = '{"id": 1234567890123456789}';
console.log(JSON.parse(raw).id); // 1234567890123456800
console.log(Number.MAX_SAFE_INTEGER); // 9007199254740991
Blaming the database or the front-end display code without comparing the raw response to the parsed value.
Mechanism: confirm the fix, usually an idempotency key the client sends and reuses on every retry.
Core cases: same key twice gives one payment and the same response; same key with a different body is rejected.
The real bug: make the first call slow so the retry arrives while it is still running, and send both at once.
Proof: count charges in the database and at the provider's sandbox, not just the responses.
“I'd start by asking how the fix works, because the usual answer is an idempotency key: the app makes a unique key per payment and sends the same one on every retry. Then I'd test that from the outside. Same key, same body, twice in a row: I expect one payment and the same response both times. Same key, different amount: I expect an agreed error, not a silent second charge. But the customer bug happened while the first request was still in flight, so I'd recreate that. I'd slow the payment provider's mock down past the timeout, fire the retry while the first call is still running, and also send two identical requests at the same moment. Then I'd check the proof that matters: one row in our payments table and one charge in the provider's sandbox.”
Testing only two calls one after the other, which never recreates the overlap that caused the double charge.
Find the layer: response headers like Cache-Control, Age or a cache hit marker; call the service directly, past any CDN or gateway.
Other sources: an app cache not cleared on write, or reads going to a replica that lags.
Judge it: a stale price on a product page may be fine; a stale price at checkout is a real bug.
Test the rule: agree the allowed delay, then test that checkout always charges the new price.
“Whether it's a bug depends on what the product needs, but first I'd find where the old value lives. I'd look at the GET response headers: Cache-Control, Age, or a header saying it was a cache hit. Then I'd call the service directly, skipping the CDN and gateway. If the direct call shows the new price, it's the edge cache. If it's still old, it's inside: maybe the service caches products and doesn't clear the entry on update, or reads go to a database replica that's behind. Then the judgement. A product page showing the old price for a minute might be acceptable. Charging the old price at checkout isn't. So I'd get the allowed delay written down, and add a test that updates a price and immediately places an order, checking the new price is charged.”
Adding a sleep to the test so the new price shows up, without asking where the stale value came from.
Recreate: both GET the product, A saves a new price, B saves the full object with the old price.
Cause: PUT replaces the whole record, so B's stale copy wins.
Expected fix: GET returns an ETag or version; the save sends If-Match; a stale save gets 412.
Test cases: matching tag succeeds with a new tag, stale tag gets 412, missing tag gets whatever the design says.
“I can recreate this with four calls, no clicking. User A and user B both GET the product. A sends a PUT with the new price. B then sends a PUT with the fixed description, but the rest of B's body is the old copy, including the old price. PUT replaces the whole record, so the price change is lost. The fix I'd expect is optimistic locking. The GET returns an ETag or a version number, and the save sends it back in an If-Match header. If the product changed in the meantime, the server refuses with 412, Precondition Failed, and the client reloads. So my tests would be: matching tag saves and returns a new one, stale tag gets 412, and a missing tag gets 428 if the server requires it. Using PATCH helps for different fields, but not when both edit the same one.”
Saying it can only be tested with two people clicking at the same time.
Meaning: 429 is Too Many Requests, so a rate limit is being hit.
Read the headers: Retry-After and any rate-limit headers show the limit and when it resets.
What changed: the suite grew, jobs now run in parallel on one key or IP, or a new limit was deployed.
Fix: a separate test client with its own limit, or pace the run; keep one test that checks the limit works.
“429 means Too Many Requests, so we're hitting a rate limit. I'd open one failing response and read the headers: Retry-After, and any rate-limit headers that show the limit and how many calls are left. Then I'd work out what changed since last month. Maybe the suite doubled in size, maybe we started running jobs in parallel with the same API key, or maybe CI runners share one outbound IP with other pipelines. It's also possible the backend team just switched the limiter on in this environment. The fix I'd want is a dedicated test client with its own sensible limit, or pacing our run. I wouldn't just add retries everywhere, because that hides real failures. And I'd keep one test that hits the limit on purpose and checks for a 429 with a Retry-After header.”
Wrapping every request in automatic retries so the suite goes green without anyone knowing why.
Why both are right: browsers enforce CORS; tools like Postman don't, so a green suite proves nothing here.
Preflight: a JSON POST with an Authorization header makes the browser send an OPTIONS request first.
Reproduce: send that OPTIONS with an Origin header and read the Access-Control headers that come back.
Usual causes: OPTIONS blocked by auth, the origin not allowed, or a header missing from the allowed list.
“We're probably both right. CORS is a rule the browser enforces. Postman doesn't apply it, so my green tests say nothing about this. For a POST with a JSON body and an Authorization header, the browser first sends an OPTIONS request, the preflight, asking if this origin may use that method and those headers. If the answer is missing or wrong, the browser blocks the real call. So I'd reproduce the preflight with curl, sending the front end's origin. Then I'd read the response: it should allow that origin, POST, and the authorization and content-type headers. Common causes are the OPTIONS call hitting auth and getting a 401, the origin not on the allowed list, or a wildcard origin used with credentials, which browsers refuse. Then I'd add that check to the suite.”
curl -i -X OPTIONS "https://api.example.test/orders" \
-H "Origin: https://app.example.test" \
-H "Access-Control-Request-Method: POST" \
-H "Access-Control-Request-Headers: authorization, content-type"
Closing it as a front-end problem because Postman works.
Size the gap: how many records do the big accounts have against the test data?
Reproduce: seed an account at production scale in a test environment; small data can't show this.
Measure growth: time the call at several sizes; a steep curve points to a query per item or no paging.
Pass on evidence: timings per size, response size, and the timeout that fires, for the fix discussion.
“My first question is how big those accounts really are. If our test account has fifty records and the big customers have two hundred thousand, the test environment can't show this bug. So I'd seed an account at that scale in a performance or staging environment. Then, rather than one timing, I'd measure the call at a few sizes, maybe a thousand, ten thousand and a hundred thousand records. If time grows in a straight line, it's probably just too much work in one request. If it climbs much faster, there's often a separate query for every item hiding in there, and I'd ask the developer to check the query count in the logs. I'd also note which timeout fires, the app's or a gateway's in front of it. Then the fix could be paging, an index, or making it a background export.”
Rerunning the call on the small test account and closing the bug as not reproducible.
Both can be right: an average hides a slow tail; look at the 95th and 99th percentiles and the maximum.
Find the pattern: break slow calls down by input, user, instance, time of day and first call after idle.
Reproduce: repeat the suspect request many times and record every timing, not just the mean.
Suspects: a slow dependency on one path, cold starts, one bad instance, retries or locking.
“Both can be right. If most requests take 150 milliseconds and a few take eight seconds, the average still looks healthy. So first I'd ask for the 95th and 99th percentiles and the maximum, or work them out from the logs. Then I'd look for what the slow ones have in common. Is it certain carts, like ones with many items or a coupon? One server instance? The first request after a quiet period? A certain time of day when a batch job runs? Once I have a suspect, I'd send that request a few hundred times and record every timing, so I can show the spread rather than one number. Common causes are one path calling a slow dependency, cold starts, one unhealthy instance, or retries stacking up behind a timeout.”
Trusting the average and telling users their network must be slow.
Poll, don't sleep: check the job status every couple of seconds, with an overall deadline that fails the test.
Check the journey: valid status changes, then a result link that works and holds the right data.
Failure path: bad input ends in a failed status with a reason, never a job stuck forever.
Edges: unknown job ID gives 404; another user's job isn't visible.
“A fixed sleep is either too long, which wastes time on every run, or too short on a busy day, which makes the test flaky. I'd poll instead. Call the status endpoint every two seconds until it says done or failed, with a hard deadline, say two minutes, after which the test fails with a clear message. Once it's done, I'd download the file and check the contents match what I asked for, not just that a link exists. Then the paths people skip. An export with bad filters should end as failed with a reason, not sit in running forever. An unknown job ID should give 404. And I'd log in as a different user and make sure they can't read my job or download my file, because export links often leak data.”
def wait_for_job(job_id, timeout=120, every=2):
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
r = requests.get(f"{BASE}/exports/{job_id}", headers=AUTH)
assert r.status_code == 200
body = r.json()
if body["status"] in ("done", "failed"):
return body
time.sleep(every)
raise AssertionError(f"job {job_id} not finished after {timeout} seconds")
Keeping the fixed sleep and just making it longer when the test flakes.
Likely cause: events arriving out of order, so an older 'pending' event overwrites a newer 'paid' one.
Order and duplicates: replay events in the wrong order and twice; state must never move backwards or double-process.
Security: missing, wrong or old signatures are rejected.
Timing: a slow handler makes the provider retry; the endpoint should acknowledge quickly and process after.
“Paid then back to pending sounds like events arriving out of order. Webhooks aren't guaranteed to come in sequence, and if our handler just writes the status from each event, a late pending event overwrites paid. I'd capture real payloads from the provider's sandbox and replay them against our endpoint, signed with the test secret. First the out-of-order case: paid, then pending. The order must stay paid. Then duplicates: the same event ID twice should be processed once, but still get a success response so the provider stops retrying. Then security: no signature, a wrong one, and a correctly signed but old event should all be rejected. Finally timing: if our handler takes too long, the provider retries, so I'd check we reply fast and do the heavy work afterwards.”
Testing only one well-formed event delivered once, which is the one case that never causes trouble.
Size: just under, at and over the limit; over should give 413 quickly.
Early rejection: a huge file should be refused without the server reading it all; watch memory while testing.
Content: wrong type renamed to .jpg, empty file, corrupt image, strange file names.
Volume: several uploads at once, since one large file was enough to take it down.
“The first missing test is size. I'd upload files just under the limit, exactly at it, and well over it. Over should come back fast with 413, which means the body is too large, and a clear message. But speed matters too. If the server reads the whole file into memory before checking, a big upload can still take it down even though the answer is right. So I'd watch the service's memory while sending a very large file, and check the rejection happens early. Then content: a video renamed to photo.jpg should be rejected by checking the actual file, not the extension. An empty file, a corrupt image, and names with slashes, dots or very long names too. Finally I'd send several uploads at the same time, because that's where memory problems show.”
Only testing one normal photo and one file slightly over the limit.
Reproduce: send the exact name and read the error message.
400 causes: a validation rule that only allows plain A to Z letters, or an apostrophe breaking something.
Garbled causes: an encoding mismatch in the request charset, the database column or the response header.
Keep a name set: accents, apostrophes, hyphens, spaces, non-Latin scripts and emoji in every run.
“I'd reproduce with the exact name first and read the error body. For the 400, the likely cause is a validation rule that only allows plain A to Z letters, so the accents or the hyphen fail. If it's the apostrophe that breaks it, and especially if I get a 500, that's worth a closer look, because it can mean the input is going into a database query unsafely. For the garbled name, I'd check the request is sent as UTF-8, then look at what's stored, then the response's content type. Somewhere one layer is using a different encoding. Emoji are a good extra test because they need four bytes in UTF-8 and some databases are set up to store only three. After that I'd keep a set of real-world names in the regression data so this stays fixed.”
Testing names only with plain English letters, as if every customer's name fit that pattern.
Check the spec: does it say what invalid filter values should do?
The risk: a client with a typo thinks it filtered and acts on the wrong data.
Worst case: the same habit on a bulk update or delete would hit every record.
Expect: a 400 naming the bad value; then try the same on every filter.
“Yes, I'd log it, and first I'd check what the spec says. The problem isn't the typo, it's what the client believes. If a report or a script sends a wrong status and gets a normal 200 with data, it assumes the filter worked. Someone could email a list of active products that includes discontinued ones. And if the same code pattern is used on a bulk update or delete endpoint, ignoring a bad filter means the action hits everything, which is serious. I'd separate two cases: unknown parameter names, which some APIs ignore by design, and known parameters with invalid values, which should give a 400 naming the value. I'd rate severity by where it's used, and I'd check every other filter for the same behaviour.”
Dropping it because the developer said it's fine, without checking the spec or thinking about who uses the filter.
Repeat-safe: DELETE is idempotent; calling it again must not break anything.
Expected code: usually 404 for the second call, or 204 if the team agreed that; never 500.
Why it matters: clients retry deletes after timeouts, and a 500 sets off alerts and error handling.
Check around it: GET after delete, lists, and related records.
“I'd say it does matter. DELETE is meant to be idempotent: calling it twice leaves the server in the same state as calling it once. The second call can answer 404, or 204 if the team agreed on that, but a 500 means the server hit something it didn't handle, probably code trying to delete a record that isn't there. It matters in practice because clients retry a delete when the first response times out. That retry will now look like a server failure, fire alerts and maybe show the user an error for something that worked. While I'm there, I'd check a GET on the item returns 404, it's gone from lists and search, and anything linked to it was handled the way the spec says.”
Accepting a 500 because the end state looks right.
Rank by risk: money, login, personal data, changed code and the endpoints clients call most.
Wide first: a quick smoke pass over all 40, generated from the spec where possible: auth works, happy path, right shape.
Deep second: negative, boundary and access tests on the top handful.
Say what's left: a short list of untested areas and their risk, shared before sign-off.
“I can't test forty endpoints deeply in two days, so the plan is about choosing. First hour: I'd rank them. Anything touching payments, login or personal data goes top, then whatever changed most, then whatever the app calls most. I'd ask the developers and product owner to check my ranking, because they know things I don't. Then I'd go wide: a quick smoke pass on all forty, ideally generated from the API spec, checking each one needs auth, works on the happy path and returns the right shape. That catches the embarrassing breaks. Then I'd spend most of the time deep on the top eight or so: bad input, boundaries, one user reaching another's data. Before sign-off, I'd share a short list of what I didn't test and why, so the go decision is made knowing the risk.”
Starting at endpoint one and working down the list until time runs out.
Collect: the exact request, timestamps with time zone, and any request or correlation ID from the response headers.
Measure: run it many times and count the failures, so 'sometimes' becomes a number.
Look for a pattern: certain data, parallel calls, one server instance, a time of day.
Reopen together: share the evidence and look at the server logs for those IDs with the developer.
“Cannot reproduce usually means I didn't give them enough to go on, so I'd fix that rather than argue. I'd export the exact failing request as curl, with headers and body. I'd note the time of each failure with the time zone, and grab the request ID or correlation ID from the response headers, because with that the developer can find the exact log line. Then I'd turn sometimes into a number: run the same request a couple of hundred times and count. If it's seven failures, that's real. I'd look for a pattern: particular carts, parallel calls, one server instance if a header shows it, or a time when a batch job runs. Then I'd reopen the bug with all of that and ask for fifteen minutes to look at the logs together.”
Reopening the bug with the same description and a complaint, or dropping it because it's rare.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.