Python interviews for experienced candidates with around five years move from how you use a feature to why you chose it and what it cost: incidents you traced in production, design decisions you would defend, performance and reliability fixes, migrations, reviewing other people's code and mentoring. Expect follow-ups that push on the trade-off you accepted. It is written for Python developers with around five to seven years of work who own a module or a service. Each question has the answer's shape and a spoken sample told as a story. Replace the stories with your own; interviewers can tell a real incident from a rehearsed one.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Inventory: dependencies without support for the new version, deprecated APIs, compiled extensions.
Run both: CI on the old and new versions until the new one is green.
Roll out: one worker or one environment first, with a quick way back.
Result: what broke, what it cost, what you would do again.
“At my last company I moved a reconciliation service up three minor versions of Python. First I listed every dependency and checked which had builds for the new version; two didn't, so upgrading those came first, as separate releases. Then I ran CI on both versions side by side and turned deprecation warnings into errors in the tests, which surfaced a few standard library calls that were going away. The surprise was a date parsing library that behaved slightly differently, and it only showed up because we compared outputs on a week of real files. We rolled it out to one worker first and kept the old image ready. It took about three weeks instead of one, but there were no customer-facing issues.”
Describing a version bump with no dependency check, no side-by-side testing and no way to roll back.
Measure first: is the time spent waiting on I/O, and is concurrency the real limit?
Costs: coroutines spread up the call stack, you need async clients, and one blocking call freezes the loop.
Stages: new async entry points, blocking calls pushed to threads with asyncio.to_thread, then hot paths converted one by one.
Decision: what you chose and what it cost.
“I had this exact debate on a gateway service I owned. First I measured: most requests spent their time waiting on three downstream calls, and we ran lots of worker processes just to hold connections open, so async genuinely fit. I was honest about the cost, though. Async spreads, so every function on the path becomes a coroutine, and our database driver and one vendor SDK were blocking. We migrated in stages: new async entry points, with the old blocking calls wrapped in asyncio.to_thread so they didn't freeze the event loop, then we swapped in async clients one at a time. We also turned on logging for slow loop callbacks. We ended up needing far fewer workers. For a CPU-heavy service elsewhere, I argued against the same move, because async wouldn't have helped at all.”
Saying async makes Python faster in general, or migrating without checking that the service really waits on I/O.
Cause: an exception in a submitted task is stored on its future; nothing shows it until result() or exception() is called.
Find it: compare records in with records done; clean logs are the clue, not the all-clear.
Fix: keep every future, call result() inside as_completed, and log each failure with its record ID.
Decide: retry, skip or fail the run, and exit non-zero when anything failed.
“On a nightly sync I owned, we pushed each customer record to a partner API through a thread pool. The job always logged 'done', but the partner kept finding missing customers. The logs were clean, and that was the clue. We called submit for each record and never looked at the futures, so when a call raised, the exception sat on its future and nobody ever saw it. I changed it to keep every future, loop over them with as_completed, and call result() inside a try block, logging the record ID on failure. Failed records got one retry, and anything still failing went to a file and made the job exit non-zero, so the scheduler flagged it. The trade-off was noisier mornings at first: about one record in a thousand had been failing for months, and we had to fix the data behind them.”
from concurrent.futures import ThreadPoolExecutor, as_completed
failed = []
with ThreadPoolExecutor(max_workers=8) as pool:
futures = {pool.submit(push_record, r): r for r in records}
for fut in as_completed(futures):
record = futures[fut]
try:
fut.result() # re-raises the task's exception here
except Exception:
logger.exception(f"push failed for record {record.id}")
failed.append(record)
if failed:
raise SystemExit(1)
Assuming an exception in a pool thread will crash the job or show up in the logs by itself.
executor.map behave when one of the calls raises?Symptom: the error, when it appeared and what it affected.
Trace: pool or open-handle counts over time, then the code path that opens without closing.
Fix: context managers so release always happens, plus sensible timeouts.
Guard: a metric or a rule so it can't creep back.
“On an order service I owned, requests started failing with pool timeouts a few hours after each deploy. A restart fixed it, which pointed to a leak. I added the pool's checked-out count to our metrics and it climbed steadily and never came back down. Then I logged where connections were taken, and nearly all the stragglers came from one report endpoint: an early return for empty results skipped the line that gave the session back. The fix was making every database session a with block, so it's released on errors and early returns too, and I added a lint rule against opening sessions outside our helper. I pushed back on just raising the pool size, because that only delays the crash and puts more load on the database.”
Fixing it by raising the connection limit or scheduling restarts, with no root cause.
Problem: logs from many requests interleaved, with no way to group them.
Mechanism: set an ID at the entry point, keep it in a ContextVar, add it to every record with a logging filter.
Why contextvars: each asyncio task runs in its own context; a thread-local is shared by all tasks on that thread.
Propagate: send the ID on outgoing calls and into background jobs.
“After an outage on an async API I owned, we had thousands of interleaved log lines and couldn't tell which belonged to the failing requests. I added middleware that reads an incoming request ID header or makes a new one, stores it in a ContextVar, and a logging filter stamps it on every record. I chose contextvars over thread-locals on purpose: in asyncio many requests share one thread, so a thread-local would be overwritten by whichever task ran last, while each task gets its own context. We also sent the ID on outgoing calls and into queued jobs, so one search showed a request across three services. One catch: run_in_executor doesn't carry the context over the way asyncio.to_thread does, so we had to copy it there ourselves.”
import contextvars
import logging
request_id = contextvars.ContextVar("request_id", default="-")
class RequestIdFilter(logging.Filter):
def filter(self, record):
record.request_id = request_id.get()
return True
handler = logging.StreamHandler()
handler.addFilter(RequestIdFilter())
handler.setFormatter(logging.Formatter("{asctime} {request_id} {message}", style="{"))
# In middleware, at the start of each request:
# request_id.set(incoming_id or new_id())
Using a global or a thread-local for per-request data in async code without seeing why it breaks.
Cause: the parent opened connections before forking, so the children share the same sockets.
Symptom: replies get mixed up between processes, showing as garbled responses and protocol or SSL errors.
Fix: open connections inside each child, for example in a pool initializer, and never reuse inherited ones.
Guard: create clients lazily, never at import time.
“I hit this on a reporting job. We created the database engine at import time, then used a process pool with the fork start method. Fork copies the parent's memory, including open sockets, so four children were talking over the same connection. Their replies got mixed up, which showed as odd protocol errors and the occasional SSL failure, and only under load, so tests never caught it. The fix was to stop creating clients at import time and open them inside each worker using the pool's initializer. I also made sure no child ever reused a pooled connection it inherited from the parent. The trade-off was more connections open at once, so I checked our database's connection limit before shipping.”
from multiprocessing import Pool
_conn = None
def init_worker():
global _conn
_conn = make_db_connection() # one connection per child
def handle(report_id):
return build_report(_conn, report_id)
with Pool(processes=4, initializer=init_worker) as pool:
results = pool.map(handle, report_ids)
Blaming the database or the network without asking what the child processes inherited from the parent.
Cause: asyncio.gather over thousands of coroutines schedules them all at once, and each holds a connection.
Cap: an asyncio.Semaphore around the call, or a fixed number of worker tasks reading from an asyncio.Queue.
Pick the limit: from the partner's limit and your connection pool size, not a guess.
Trade-off: a slower, steady run instead of fast failures and retry storms.
“A backfill job I owned fetched about twenty thousand records from a partner API. The first version made a coroutine per record and passed them all to gather, which schedules every one of them at once. The partner answered with rate-limit errors, our retries made it worse, and then the process ran out of file descriptors because every request held a socket open. I wrapped the request in an asyncio.Semaphore of twenty and shared one HTTP client so connections were reused. Twenty came from the partner's documented limit and our client's pool size. For a much bigger run I'd switch to a fixed number of worker tasks pulling IDs from an asyncio.Queue, because creating twenty thousand tasks up front still costs memory. The run took longer, but it finished first time with no errors, which is what that job needed.”
import asyncio
async def fetch_all(client, ids, limit=20):
sem = asyncio.Semaphore(limit)
async def fetch_one(record_id):
async with sem: # at most `limit` requests in flight
return await client.get_record(record_id)
return await asyncio.gather(*(fetch_one(i) for i in ids))
Starting every request at once and blaming the partner, or picking a limit with no link to what the downstream can handle.
Processes: several worker processes means one dict each, so the limit multiplies, and a restart wipes it.
Threads: inside one process, counts[key] += 1 is a read, an add and a write; the GIL doesn't make it atomic.
Fix: a shared store with an atomic increment and an expiry, or limiting at the proxy.
Trade-off: a network call per request, and a new dependency that can fail.
“This was a per-client limit I added to an API I owned. The tests ran one process with one thread, so it looked perfect. In production the server ran four worker processes, and each had its own copy of the module and its own dict, so a client could get close to four times the limit, and every restart or deploy wiped the counts. There was a smaller bug too: inside a process, requests ran on several threads, and counts[key] += 1 is a read, an add and a write, so two threads could read the same value and one increment would vanish. The GIL doesn't prevent that. The real fix was moving the counter to a shared store with an atomic increment and an expiry per time window. That added a network call per request and a new dependency, so we agreed that if the store was down, requests would be let through rather than blocked.”
Saying the GIL makes the increment safe, or missing that separate worker processes never share the dict.
Incident: what broke and why an unlocked indirect dependency caused it.
Lock: direct dependencies as ranges, plus a lock file with exact versions and hashes for every build.
Upgrade on purpose: scheduled, automated update pull requests that run the full test suite.
Trade-off: fewer surprises, a little more maintenance.
“At my last company we listed only our direct dependencies, with loose ranges. One Friday a sub-dependency of our HTTP client released a version that changed a default, and a fresh build broke staging. Production only survived because it was still on an older image. After that I moved the service to a lock file: our requirements stay as ranges, but every build installs exact versions with hashes from the lock, so the install is identical on my laptop, in CI and in the container. Then I set up a weekly automated pull request that refreshes the lock and runs the whole suite, so upgrades happen when we choose, in small steps. The cost is a short review every week, and now and then a bump we have to hold back and investigate.”
Saying they pin everything by hand and never upgrade, or that they just rebuild and hope.
Cause: the platform sends SIGTERM, and by default Python exits at once without running finally blocks.
Handle it: a signal handler that only sets a flag; stop taking new jobs, finish the current one.
Fit the window: keep jobs shorter than the grace period, or make them safe to redo.
Acknowledge late: mark a job done only after it succeeds.
“On a worker fleet I owned, every deploy left a handful of half-processed jobs. The orchestrator sends SIGTERM, waits a grace period, then kills the process. Python's default for SIGTERM is to exit immediately, so our cleanup and finally blocks never ran. I added a handler that only sets a stop flag; the main loop checks it between jobs, finishes the current one and exits cleanly. Python only runs signal handlers in the main thread, so I kept the loop there. We also moved the acknowledgment to after the job succeeds, so anything cut off by a hard kill goes back on the queue. The trade-off was slower deploys, since we raised the grace period to cover our slowest job.”
import signal
import threading
stop = threading.Event()
def on_sigterm(signum, frame):
stop.set() # only set a flag; no real work in the handler
signal.signal(signal.SIGTERM, on_sigterm)
while not stop.is_set():
job = fetch_next_job(timeout=1)
if job is not None:
process(job)
job.ack() # only after it succeeded
Assuming Python cleans up on SIGTERM by itself, or doing heavy work inside the signal handler.
Idempotent writes: upserts keyed on a natural ID, not blind inserts.
Checkpoints: record progress per batch so a rerun skips finished work.
Side effects: emails or payments guarded by a record of what was already sent.
Trade-off: extra bookkeeping and slightly slower runs.
“The first version of a billing export I inherited used plain inserts, so when it crashed at batch forty, rerunning it created duplicates and someone spent a morning deleting rows by hand. When I rebuilt it, I made every write an upsert keyed on the invoice ID and date, so running it twice gives the same result. It works in batches and records each finished batch in a small progress table, so a rerun picks up where it stopped. The part people forget is side effects: it also emailed a summary, so it checks a sent record first and writes one right after sending. A crash between those two could send the summary twice, which is harmless for a summary. Now on-call just reruns it.”
Relying on someone to clean up by hand after a failure, or forgetting side effects like emails.
Keep both: accept the old form for a while and map it onto the new one.
Warn: warnings.warn with DeprecationWarning and stacklevel=2 so it points at the caller's line.
Version and tell: the warning in a minor release, the removal in the next major, a changelog entry and a message to each team.
Find callers: search the other repositories and help the biggest users move.
“I did this with a retry helper whose delay argument we wanted to replace with a backoff object. I kept the old keyword working, converted it inside the function, and issued a DeprecationWarning with stacklevel=2, so the warning pointed at the caller's line, not mine. Those warnings are hidden by default outside tests and the main script, so I also asked teams to run their tests with warnings from our package treated as errors. It shipped as a minor version with a changelog entry. I searched the other repositories, found about thirty call sites and opened pull requests for the two biggest teams myself. Two releases later we removed the old argument in a major version. The cost was carrying some ugly compatibility code for a couple of months.”
import warnings
def retry(func, *, backoff=None, delay=None):
if delay is not None:
warnings.warn(
"retry(delay=...) is deprecated; pass backoff=Backoff(delay)",
DeprecationWarning,
stacklevel=2,
)
backoff = Backoff(delay)
...
Changing the signature in place and expecting other teams to find out from a failed build.
Abstract base class: plugins must inherit; a missing method fails when the class is instantiated; shared code can live in the base.
Protocol: any class with the right methods fits; the type checker enforces it, not the runtime by default.
Choice: who writes plugins, whether you need shared behaviour, and when you want the error.
Cost: what you gave up.
“On an export module I owned, other teams wrote their own exporters. I first used an abstract base class, which was nice because a missing method failed as soon as someone instantiated the class, and I could put shared helpers in the base. But it forced every exporter to import and inherit from our package, and one team wanted to plug in a class from their own library. So I switched to a typing.Protocol: anything with an export(records) method and a name fits, and the type checker in CI catches mismatches. The cost was losing the runtime check, so I added a small check at registration that the method exists and is callable. For an interface with lots of shared code, I'd still pick the base class.”
Treating the two as interchangeable, or not knowing that a Protocol isn't enforced at runtime by default.
@runtime_checkable give you, and what does it not check?Normal or broken: if not finding it is normal, return None; if it means a bug or broken data, raise.
Collections: return an empty list, never None, so callers' loops just work.
Make it visible: the type hint and the name say which one it is.
Own exception: raise a named exception from your module, not a bare KeyError from its insides.
“In a customer module four teams used, lookups were inconsistent: some returned None, some leaked a KeyError from deep inside, one returned an empty dict. Callers kept getting it wrong, and the worst bug was a None passed along until something failed three layers away. So I set one rule. If not finding it is normal, like checking an email at signup, the function is find_customer and returns Customer | None, so the type checker makes callers deal with None. If the ID came from our own database and must exist, it's get_customer and raises our own CustomerNotFound, so the error points at the real problem. Anything returning many items returns an empty list, never None. The cost was renaming functions across four teams' code over a few releases, with deprecation warnings on the old names.”
class CustomerNotFound(LookupError):
pass
def find_customer(email: str) -> "Customer | None":
return _by_email.get(email.lower())
def get_customer(customer_id: int) -> "Customer":
try:
return _by_id[customer_id]
except KeyError:
raise CustomerNotFound(customer_id) from None
Saying they always return None, or letting a KeyError from inside the module leak out to callers.
LookupError for your own not-found exception?The need: untrusted input at the edge versus trusted data inside.
Runtime validation: parses and checks types at the boundary, with clear error messages.
Plain classes: light, fast and dependency-free, but they check nothing.
Cost: speed, a dependency to keep upgrading, or bugs you let through.
“On an ingestion API I owned, I used pydantic models at the edge, because the input came from partners and was often wrong. It gave us type coercion and error messages clients could act on, instead of a KeyError deep inside. Inside the service I used plain dataclasses, since that data was already validated and I didn't want to pay the checking cost on every internal object. The split had costs. A major version of the validation library changed some behaviour and took a sprint to work through, and I had to keep an eye out for people quietly passing raw dicts past the edge. For a small internal tool I'd use dataclasses alone. The rule I'd defend is: validate once, at the boundary.”
Picking a library because it's popular, with no idea where validation belongs or what it costs.
Coupling: a pickle refers to classes by module and name, so moving or renaming a class breaks messages already waiting.
Safety: loading a pickle can run arbitrary code, so only load data you fully trust.
Replacement: a versioned JSON message with explicit plain fields, often just IDs.
Cost: conversion code to write, and datetimes and decimals handled by hand.
“On a job system I owned, producers pickled whole objects onto the queue because it was one line of code. It bit us during a deploy: we moved a class to a new module, and every message already waiting failed to load, because a pickle finds the class by its module and name. Workers on the new code couldn't read what the old code had written, and we drained the queue by hand that night. Around the same time a security review pointed out that loading a pickle can run arbitrary code, and that queue was shared with another team's service. So I moved the payloads to JSON with a version field and only plain values: IDs, strings, numbers, timestamps as ISO strings. Workers load the full record from the database by ID. The cost was writing conversion code and handling datetimes and decimals by hand, but deploys stopped caring what was in the queue.”
import json
def to_message(order_id, requested_at):
return json.dumps({
"version": 2,
"order_id": order_id,
"requested_at": requested_at.isoformat(),
})
Not knowing that loading a pickle from an untrusted source can run code, or that pickled messages break when classes move.
Suspect: a big heap of long-lived objects makes full collections slow.
Prove: time each collection with gc.callbacks and line the slow ones up with the spikes.
Fix: fewer reference cycles, gc.freeze() after startup, or tuned thresholds.
Trade-off: what freezing or tuning risks.
“On a pricing service I owned, p99 latency jumped every few minutes while CPU and traffic looked flat. Nothing external lined up, so I suspected the cycle collector, because we kept a huge in-memory lookup table of long-lived objects that a full collection had to walk. I added a gc.callbacks hook that logged how long each collection took and which generation, and the slow ones were the oldest generation, right on the spikes. Two fixes helped. We loaded the table at startup and called gc.freeze() afterwards, so the collector stopped rescanning it, and we fixed a code path that created reference cycles on every request. The trade-off is that frozen objects are never checked for cycles again, so anything frozen that later ends up in a cycle leaks. That's why I froze right after loading, before serving requests.”
import gc
import logging
import time
_started = {}
def gc_timer(phase, info):
if phase == "start":
_started["t"] = time.perf_counter()
elif phase == "stop":
ms = (time.perf_counter() - _started["t"]) * 1000
if ms > 20:
logging.warning(f"gc generation {info['generation']} took {ms:.1f} ms")
gc.callbacks.append(gc_timer)
Tuning or disabling the garbage collector on a hunch, without measuring collection times first.
Ask what matters: how bad is a lost or doubled email for this case?
Name the risks: the thread dies when a worker restarts or a deploy happens, there's no retry, and failures go unseen.
Middle path: write the email to an outbox table in the same transaction, and let one small worker send it.
Draw the line: when the shortcut is fine and when it isn't.
“I had this exact conversation about receipt emails. I didn't just say no; I asked how bad a lost email would be, and for receipts the answer was support tickets. Then I explained what would really happen: our web workers get recycled after a set number of requests and on every deploy, and a thread that's halfway through sending just dies with them. There's no retry, and an exception in that thread only gets printed to stderr, so failures would be invisible. A full queue setup was more than the deadline allowed, so I suggested a middle path: write the email as a row in an outbox table in the same transaction as the order, and have one small worker process send pending rows and mark them sent. It took two extra days. For a low-stakes 'someone viewed your profile' notice, I agreed the thread was fine.”
Blocking the shortcut on principle with no cheaper option, or agreeing without asking what a lost email costs.
Starting point: where the person was and what held them back.
What you did: pairing, smaller tasks, clearer reviews.
What you changed: where your first approach didn't work.
Result: what they could do on their own afterwards.
“A new graduate joined my team and was slow to ship; their pull requests sat for days. My first move was detailed reviews, and that backfired: long comment lists made them more anxious and slower. So I changed approach. We paired for an hour twice a week on their own tickets, with them typing. I split tasks into pieces that could merge within a day, and in reviews I marked each comment as must-fix or optional so they knew what mattered. I also gave them a small module to own, the CSV importer, alerts included. Within three months they were reviewing other people's code in that area and handled an incident there without me. The cost was some of my own output for a while, which I agreed with my manager up front.”
Describing mentoring as answering questions when asked, or quietly rewriting the junior's code for them.
Why: the pain that justified it, in terms the team and manager cared about.
Plan: small steps behind a stable interface, each one mergeable on its own.
Safety: characterisation tests, flags, and old and new code run side by side.
Result and cost: what improved and what it took.
“I led a refactor of our pricing code, which had grown into a two-thousand-line module where every change broke something. I sold it to my manager on lead time: small pricing changes were taking a week. We never had a big branch. First I wrote characterisation tests that captured current prices for a few thousand real orders. Then I put a new interface in front of the old code and moved one family of rules at a time behind it, each move a small pull request merged the same day. For the riskiest part, the new code ran alongside the old in production and logged any difference, and we only switched after a full week with none. It took two months with feature work continuing, and pricing changes now take a day or two.”
A months-long branch merged in one go, or no way to show the behaviour stayed the same.
Shape: many fast unit tests on logic, fewer integration tests against real dependencies, a thin end-to-end layer.
Real dependencies: where mocks lied, and where a real database in a container paid off.
Left out: what you don't test and why.
Speed: how long CI takes and what you did about it.
“For an order service I owned, most tests are fast unit tests on pure functions: pricing, discounts, state changes. Then there are about forty integration tests against a real database in a container, because mocked queries had lied to us twice about constraint errors. On top sits a handful of end-to-end checks of the main flows. What I chose not to test: thin glue code and the framework's own behaviour, like whether the router routes. I also don't chase a coverage number; we dropped our target because people were writing tests that asserted nothing. Integration tests are slower, so CI runs the unit tests on every push and the full set before merge. The whole suite takes about six minutes.”
Quoting a coverage target as the whole strategy, or mocking everything including the database.
Acknowledge: it works, and the instinct behind it is sound.
Name the cost: harder to read, debug, type-check and search; surprises for the next person.
Offer a simpler path: a plain class, a dict of handlers or a decorator-based registry.
Deliver it well: a short call rather than a wall of comments.
“I had almost this exact pull request: a metaclass that auto-registered handler classes and generated attributes through __getattr__. It worked, and I said so first, because the instinct to remove repetition was right. Then I explained the cost in concrete terms: our type checker couldn't see those attributes, editors couldn't jump to definitions, and a typo returned something odd instead of raising an error. I suggested a plain dict registry filled by a small decorator, which kept the auto-registration and dropped the magic. Rather than leave twenty comments, I had a fifteen-minute call and we rewrote one handler together; they did the rest. The trade-off I accepted was a little more code to type. A metaclass is fine when nothing simpler works, but that bar is high.”
Approving it because it works, or rejecting it as too clever with no explanation and no alternative.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.