Production Incidents • Design Trade-offs • Performance • Migrations • Leadership • 2026

Python Interview Questions for Experienced Candidates (5 Years)

Python interviews for experienced candidates with around five years move from how you use a feature to why you chose it and what it cost: incidents you traced in production, design decisions you would defend, performance and reliability fixes, migrations, reviewing other people's code and mentoring. Expect follow-ups that push on the trade-off you accepted. It is written for Python developers with around five to seven years of work who own a module or a service. Each question has the answer's shape and a spoken sample told as a story. Replace the stories with your own; interviewers can tell a real incident from a rehearsed one.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Migrations 2 questions

Hard Behavioral round Mid-level, Senior Practice question

1. Tell me about upgrading a service you owned to a newer Python version. How did you plan it, and what broke along the way?

What the interviewer is really testing:
Whether they can run a migration safely: list the risks, test on both versions and roll out gradually, rather than bumping the version and hoping.
Answer frame:

Inventory: dependencies without support for the new version, deprecated APIs, compiled extensions.

Run both: CI on the old and new versions until the new one is green.

Roll out: one worker or one environment first, with a quick way back.

Result: what broke, what it cost, what you would do again.

Sample spoken answer:

“At my last company I moved a reconciliation service up three minor versions of Python. First I listed every dependency and checked which had builds for the new version; two didn't, so upgrading those came first, as separate releases. Then I ran CI on both versions side by side and turned deprecation warnings into errors in the tests, which surfaced a few standard library calls that were going away. The surprise was a date parsing library that behaved slightly differently, and it only showed up because we compared outputs on a week of real files. We rolled it out to one worker first and kept the old image ready. It took about three weeks instead of one, but there were no customer-facing issues.”

Red flag to avoid:

Describing a version bump with no dependency check, no side-by-side testing and no way to roll back.

They may ask next:
  • How did you justify the upgrade against feature work?
  • What would you automate so the next upgrade is cheaper?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

2. Your team wants to move a synchronous Python service to asyncio. How do you decide whether it's worth it, and how would you migrate in stages?

What the interviewer is really testing:
Whether they tie the choice to where the service actually waits, and know the real costs: async spreads through the call stack and one blocking call stalls everything.
Answer frame:

Measure first: is the time spent waiting on I/O, and is concurrency the real limit?

Costs: coroutines spread up the call stack, you need async clients, and one blocking call freezes the loop.

Stages: new async entry points, blocking calls pushed to threads with asyncio.to_thread, then hot paths converted one by one.

Decision: what you chose and what it cost.

Sample spoken answer:

“I had this exact debate on a gateway service I owned. First I measured: most requests spent their time waiting on three downstream calls, and we ran lots of worker processes just to hold connections open, so async genuinely fit. I was honest about the cost, though. Async spreads, so every function on the path becomes a coroutine, and our database driver and one vendor SDK were blocking. We migrated in stages: new async entry points, with the old blocking calls wrapped in asyncio.to_thread so they didn't freeze the event loop, then we swapped in async clients one at a time. We also turned on logging for slow loop callbacks. We ended up needing far fewer workers. For a CPU-heavy service elsewhere, I argued against the same move, because async wouldn't have helped at all.”

Red flag to avoid:

Saying async makes Python faster in general, or migrating without checking that the service really waits on I/O.

They may ask next:
  • How would you find a blocking call hiding inside an async code path?
  • What happens to your existing tests when the code turns async?
Say it in 60 seconds

Production Incidents 3 questions

Medium Technical round Mid-level, Senior Practice question

3. A sync job you own uses ThreadPoolExecutor to call an API for each record. It reports success, but some records were never processed. What happened, and how did you fix it?

What the interviewer is really testing:
Whether they know an exception inside a submitted task is stored on its future and stays silent until someone asks for the result, and build jobs that account for every item.
Answer frame:

Cause: an exception in a submitted task is stored on its future; nothing shows it until result() or exception() is called.

Find it: compare records in with records done; clean logs are the clue, not the all-clear.

Fix: keep every future, call result() inside as_completed, and log each failure with its record ID.

Decide: retry, skip or fail the run, and exit non-zero when anything failed.

Sample spoken answer:

“On a nightly sync I owned, we pushed each customer record to a partner API through a thread pool. The job always logged 'done', but the partner kept finding missing customers. The logs were clean, and that was the clue. We called submit for each record and never looked at the futures, so when a call raised, the exception sat on its future and nobody ever saw it. I changed it to keep every future, loop over them with as_completed, and call result() inside a try block, logging the record ID on failure. Failed records got one retry, and anything still failing went to a file and made the job exit non-zero, so the scheduler flagged it. The trade-off was noisier mornings at first: about one record in a thousand had been failing for months, and we had to fix the data behind them.”

Code:
from concurrent.futures import ThreadPoolExecutor, as_completed

failed = []
with ThreadPoolExecutor(max_workers=8) as pool:
    futures = {pool.submit(push_record, r): r for r in records}
    for fut in as_completed(futures):
        record = futures[fut]
        try:
            fut.result()  # re-raises the task's exception here
        except Exception:
            logger.exception(f"push failed for record {record.id}")
            failed.append(record)

if failed:
    raise SystemExit(1)
Red flag to avoid:

Assuming an exception in a pool thread will crash the job or show up in the logs by itself.

They may ask next:
  • How does executor.map behave when one of the calls raises?
  • How would you choose the number of threads for an API with a rate limit?
Say it in 60 seconds
Hard Behavioral round Mid-level, Senior Practice question

4. Tell me about a time a Python service ran out of database connections or file handles. How did you find the leak?

What the interviewer is really testing:
Whether they can trace resource exhaustion to the code path that forgets to release, and fix it in the structure of the code rather than by raising the limit.
Answer frame:

Symptom: the error, when it appeared and what it affected.

Trace: pool or open-handle counts over time, then the code path that opens without closing.

Fix: context managers so release always happens, plus sensible timeouts.

Guard: a metric or a rule so it can't creep back.

Sample spoken answer:

“On an order service I owned, requests started failing with pool timeouts a few hours after each deploy. A restart fixed it, which pointed to a leak. I added the pool's checked-out count to our metrics and it climbed steadily and never came back down. Then I logged where connections were taken, and nearly all the stragglers came from one report endpoint: an early return for empty results skipped the line that gave the session back. The fix was making every database session a with block, so it's released on errors and early returns too, and I added a lint rule against opening sessions outside our helper. I pushed back on just raising the pool size, because that only delays the crash and puts more load on the database.”

Red flag to avoid:

Fixing it by raising the connection limit or scheduling restarts, with no root cause.

They may ask next:
  • How would you find a file handle leak on a Linux box without changing any code?
  • What pool timeout would you set, and what should happen when it's hit?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

5. After an incident you couldn't follow one request through the logs. How did you add request IDs, and why use contextvars instead of thread-locals?

What the interviewer is really testing:
Whether they can build tracing into logs and understand that thread-local state is shared by every asyncio task on a thread, while context variables follow each task.
Answer frame:

Problem: logs from many requests interleaved, with no way to group them.

Mechanism: set an ID at the entry point, keep it in a ContextVar, add it to every record with a logging filter.

Why contextvars: each asyncio task runs in its own context; a thread-local is shared by all tasks on that thread.

Propagate: send the ID on outgoing calls and into background jobs.

Sample spoken answer:

“After an outage on an async API I owned, we had thousands of interleaved log lines and couldn't tell which belonged to the failing requests. I added middleware that reads an incoming request ID header or makes a new one, stores it in a ContextVar, and a logging filter stamps it on every record. I chose contextvars over thread-locals on purpose: in asyncio many requests share one thread, so a thread-local would be overwritten by whichever task ran last, while each task gets its own context. We also sent the ID on outgoing calls and into queued jobs, so one search showed a request across three services. One catch: run_in_executor doesn't carry the context over the way asyncio.to_thread does, so we had to copy it there ourselves.”

Code:
import contextvars
import logging

request_id = contextvars.ContextVar("request_id", default="-")

class RequestIdFilter(logging.Filter):
    def filter(self, record):
        record.request_id = request_id.get()
        return True

handler = logging.StreamHandler()
handler.addFilter(RequestIdFilter())
handler.setFormatter(logging.Formatter("{asctime} {request_id} {message}", style="{"))

# In middleware, at the start of each request:
# request_id.set(incoming_id or new_id())
Red flag to avoid:

Using a global or a thread-local for per-request data in async code without seeing why it breaks.

They may ask next:
  • How would you connect these IDs with traces across several services?
  • What would you log with the ID, and what must never appear in those logs?
Say it in 60 seconds

Concurrency 3 questions

Hard Technical round Mid-level, Senior Practice question

6. After you moved a batch job to multiprocessing, it started throwing random database and SSL errors. What's the likely cause, and how do you fix it?

What the interviewer is really testing:
Whether they understand what fork copies into child processes, and have seen shared sockets or connection pools go wrong in practice.
Answer frame:

Cause: the parent opened connections before forking, so the children share the same sockets.

Symptom: replies get mixed up between processes, showing as garbled responses and protocol or SSL errors.

Fix: open connections inside each child, for example in a pool initializer, and never reuse inherited ones.

Guard: create clients lazily, never at import time.

Sample spoken answer:

“I hit this on a reporting job. We created the database engine at import time, then used a process pool with the fork start method. Fork copies the parent's memory, including open sockets, so four children were talking over the same connection. Their replies got mixed up, which showed as odd protocol errors and the occasional SSL failure, and only under load, so tests never caught it. The fix was to stop creating clients at import time and open them inside each worker using the pool's initializer. I also made sure no child ever reused a pooled connection it inherited from the parent. The trade-off was more connections open at once, so I checked our database's connection limit before shipping.”

Code:
from multiprocessing import Pool

_conn = None

def init_worker():
    global _conn
    _conn = make_db_connection()  # one connection per child

def handle(report_id):
    return build_report(_conn, report_id)

with Pool(processes=4, initializer=init_worker) as pool:
    results = pool.map(handle, report_ids)
Red flag to avoid:

Blaming the database or the network without asking what the child processes inherited from the parent.

They may ask next:
  • Why might the same code work fine with the spawn start method?
  • How would you size the number of workers against the database's connection limit?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

7. An asyncio job you owned fired thousands of requests at once, got rate-limited, then hit 'too many open files'. How did you limit how much ran at the same time?

What the interviewer is really testing:
Whether they know that gathering thousands of coroutines starts them all together, and can cap in-flight work with a semaphore or a fixed set of worker tasks, choosing the limit from what the downstream can take.
Answer frame:

Cause: asyncio.gather over thousands of coroutines schedules them all at once, and each holds a connection.

Cap: an asyncio.Semaphore around the call, or a fixed number of worker tasks reading from an asyncio.Queue.

Pick the limit: from the partner's limit and your connection pool size, not a guess.

Trade-off: a slower, steady run instead of fast failures and retry storms.

Sample spoken answer:

“A backfill job I owned fetched about twenty thousand records from a partner API. The first version made a coroutine per record and passed them all to gather, which schedules every one of them at once. The partner answered with rate-limit errors, our retries made it worse, and then the process ran out of file descriptors because every request held a socket open. I wrapped the request in an asyncio.Semaphore of twenty and shared one HTTP client so connections were reused. Twenty came from the partner's documented limit and our client's pool size. For a much bigger run I'd switch to a fixed number of worker tasks pulling IDs from an asyncio.Queue, because creating twenty thousand tasks up front still costs memory. The run took longer, but it finished first time with no errors, which is what that job needed.”

Code:
import asyncio

async def fetch_all(client, ids, limit=20):
    sem = asyncio.Semaphore(limit)

    async def fetch_one(record_id):
        async with sem:  # at most `limit` requests in flight
            return await client.get_record(record_id)

    return await asyncio.gather(*(fetch_one(i) for i in ids))
Red flag to avoid:

Starting every request at once and blaming the partner, or picking a limit with no link to what the downstream can handle.

They may ask next:
  • What changes if the partner limits requests per second rather than requests in flight?
  • How would you notice in production that the limit was set too low?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

8. You built a rate limiter on a module-level dict. It passed every test, but in production clients get far more requests through than the limit. What went wrong?

What the interviewer is really testing:
Whether they spot that each worker process holds its own copy of module state, that threads in one process still race on read-modify-write, and can decide where shared state should live.
Answer frame:

Processes: several worker processes means one dict each, so the limit multiplies, and a restart wipes it.

Threads: inside one process, counts[key] += 1 is a read, an add and a write; the GIL doesn't make it atomic.

Fix: a shared store with an atomic increment and an expiry, or limiting at the proxy.

Trade-off: a network call per request, and a new dependency that can fail.

Sample spoken answer:

“This was a per-client limit I added to an API I owned. The tests ran one process with one thread, so it looked perfect. In production the server ran four worker processes, and each had its own copy of the module and its own dict, so a client could get close to four times the limit, and every restart or deploy wiped the counts. There was a smaller bug too: inside a process, requests ran on several threads, and counts[key] += 1 is a read, an add and a write, so two threads could read the same value and one increment would vanish. The GIL doesn't prevent that. The real fix was moving the counter to a shared store with an atomic increment and an expiry per time window. That added a network call per request and a new dependency, so we agreed that if the store was down, requests would be let through rather than blocked.”

Red flag to avoid:

Saying the GIL makes the increment safe, or missing that separate worker processes never share the dict.

They may ask next:
  • Would a free-threaded build of Python change anything here?
  • How would you write a test that catches the multi-process problem before production?
Say it in 60 seconds

Reliability 3 questions

Medium Behavioral round Mid-level, Senior Practice question

9. A deploy broke because a library you don't even import directly shipped a new release. How do you manage dependencies on a service you own so that can't happen?

What the interviewer is really testing:
Whether they separate declared version ranges from locked, reproducible installs, and upgrade on purpose instead of by accident.
Answer frame:

Incident: what broke and why an unlocked indirect dependency caused it.

Lock: direct dependencies as ranges, plus a lock file with exact versions and hashes for every build.

Upgrade on purpose: scheduled, automated update pull requests that run the full test suite.

Trade-off: fewer surprises, a little more maintenance.

Sample spoken answer:

“At my last company we listed only our direct dependencies, with loose ranges. One Friday a sub-dependency of our HTTP client released a version that changed a default, and a fresh build broke staging. Production only survived because it was still on an older image. After that I moved the service to a lock file: our requirements stay as ranges, but every build installs exact versions with hashes from the lock, so the install is identical on my laptop, in CI and in the container. Then I set up a weekly automated pull request that refreshes the lock and runs the whole suite, so upgrades happen when we choose, in small steps. The cost is a short review every week, and now and then a bump we have to hold back and investigate.”

Red flag to avoid:

Saying they pin everything by hand and never upgrade, or that they just rebuild and hope.

They may ask next:
  • How would you handle a security fix that needs an upgrade today?
  • Would you lock versions the same way for a library other people install?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

10. Every deploy, your Python queue workers drop a few jobs halfway through. How do you make them shut down gracefully?

What the interviewer is really testing:
Whether they know how orchestrators stop processes, what Python does on SIGTERM by default, and how to finish in-flight work safely.
Answer frame:

Cause: the platform sends SIGTERM, and by default Python exits at once without running finally blocks.

Handle it: a signal handler that only sets a flag; stop taking new jobs, finish the current one.

Fit the window: keep jobs shorter than the grace period, or make them safe to redo.

Acknowledge late: mark a job done only after it succeeds.

Sample spoken answer:

“On a worker fleet I owned, every deploy left a handful of half-processed jobs. The orchestrator sends SIGTERM, waits a grace period, then kills the process. Python's default for SIGTERM is to exit immediately, so our cleanup and finally blocks never ran. I added a handler that only sets a stop flag; the main loop checks it between jobs, finishes the current one and exits cleanly. Python only runs signal handlers in the main thread, so I kept the loop there. We also moved the acknowledgment to after the job succeeds, so anything cut off by a hard kill goes back on the queue. The trade-off was slower deploys, since we raised the grace period to cover our slowest job.”

Code:
import signal
import threading

stop = threading.Event()

def on_sigterm(signum, frame):
    stop.set()  # only set a flag; no real work in the handler

signal.signal(signal.SIGTERM, on_sigterm)

while not stop.is_set():
    job = fetch_next_job(timeout=1)
    if job is not None:
        process(job)
        job.ack()  # only after it succeeded
Red flag to avoid:

Assuming Python cleans up on SIGTERM by itself, or doing heavy work inside the signal handler.

They may ask next:
  • What happens if a single job takes longer than the grace period?
  • How would you test shutdown behaviour before a real deploy?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

11. A nightly job you own died halfway through. How did you design it so the fix is simply to run it again?

What the interviewer is really testing:
Whether they build jobs to be idempotent and resumable, so a failure means a rerun, not a manual clean-up.
Answer frame:

Idempotent writes: upserts keyed on a natural ID, not blind inserts.

Checkpoints: record progress per batch so a rerun skips finished work.

Side effects: emails or payments guarded by a record of what was already sent.

Trade-off: extra bookkeeping and slightly slower runs.

Sample spoken answer:

“The first version of a billing export I inherited used plain inserts, so when it crashed at batch forty, rerunning it created duplicates and someone spent a morning deleting rows by hand. When I rebuilt it, I made every write an upsert keyed on the invoice ID and date, so running it twice gives the same result. It works in batches and records each finished batch in a small progress table, so a rerun picks up where it stopped. The part people forget is side effects: it also emailed a summary, so it checks a sent record first and writes one right after sending. A crash between those two could send the summary twice, which is harmless for a summary. Now on-call just reruns it.”

Red flag to avoid:

Relying on someone to clean up by hand after a failure, or forgetting side effects like emails.

They may ask next:
  • What would you do if the side effect were a payment, where sending twice is not harmless?
  • How do you test that a job really is safe to rerun?
Say it in 60 seconds

Design Choices 5 questions

Hard Technical round Mid-level, Senior Practice question

12. You own an internal Python package that five teams import. How do you change a function's signature without breaking them?

What the interviewer is really testing:
Whether they can evolve a shared API safely: a deprecation period, warnings that point at the caller, clear versioning and actually finding the callers.
Answer frame:

Keep both: accept the old form for a while and map it onto the new one.

Warn: warnings.warn with DeprecationWarning and stacklevel=2 so it points at the caller's line.

Version and tell: the warning in a minor release, the removal in the next major, a changelog entry and a message to each team.

Find callers: search the other repositories and help the biggest users move.

Sample spoken answer:

“I did this with a retry helper whose delay argument we wanted to replace with a backoff object. I kept the old keyword working, converted it inside the function, and issued a DeprecationWarning with stacklevel=2, so the warning pointed at the caller's line, not mine. Those warnings are hidden by default outside tests and the main script, so I also asked teams to run their tests with warnings from our package treated as errors. It shipped as a minor version with a changelog entry. I searched the other repositories, found about thirty call sites and opened pull requests for the two biggest teams myself. Two releases later we removed the old argument in a major version. The cost was carrying some ugly compatibility code for a couple of months.”

Code:
import warnings

def retry(func, *, backoff=None, delay=None):
    if delay is not None:
        warnings.warn(
            "retry(delay=...) is deprecated; pass backoff=Backoff(delay)",
            DeprecationWarning,
            stacklevel=2,
        )
        backoff = Backoff(delay)
    ...
Red flag to avoid:

Changing the signature in place and expecting other teams to find out from a failed build.

They may ask next:
  • What would you do about a team that ignores the warning until the removal date?
  • In Python almost everything is reachable. How do you decide what counts as a breaking change?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

13. You're designing a plugin interface for a module you own. Would you use an abstract base class or a typing Protocol, and what did that choice cost?

What the interviewer is really testing:
Whether they understand inheritance-based versus structural typing and choose based on who writes the plugins and when errors should surface.
Answer frame:

Abstract base class: plugins must inherit; a missing method fails when the class is instantiated; shared code can live in the base.

Protocol: any class with the right methods fits; the type checker enforces it, not the runtime by default.

Choice: who writes plugins, whether you need shared behaviour, and when you want the error.

Cost: what you gave up.

Sample spoken answer:

“On an export module I owned, other teams wrote their own exporters. I first used an abstract base class, which was nice because a missing method failed as soon as someone instantiated the class, and I could put shared helpers in the base. But it forced every exporter to import and inherit from our package, and one team wanted to plug in a class from their own library. So I switched to a typing.Protocol: anything with an export(records) method and a name fits, and the type checker in CI catches mismatches. The cost was losing the runtime check, so I added a small check at registration that the method exists and is callable. For an interface with lots of shared code, I'd still pick the base class.”

Red flag to avoid:

Treating the two as interchangeable, or not knowing that a Protocol isn't enforced at runtime by default.

They may ask next:
  • What does @runtime_checkable give you, and what does it not check?
  • How would you version this interface once many plugins depend on it?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

14. In a module other teams call, how did you decide whether a lookup returns None, raises an exception or returns an empty result when nothing is found?

What the interviewer is really testing:
Whether they design the error contract on purpose, from how callers use the function, and make the choice visible in names and type hints.
Answer frame:

Normal or broken: if not finding it is normal, return None; if it means a bug or broken data, raise.

Collections: return an empty list, never None, so callers' loops just work.

Make it visible: the type hint and the name say which one it is.

Own exception: raise a named exception from your module, not a bare KeyError from its insides.

Sample spoken answer:

“In a customer module four teams used, lookups were inconsistent: some returned None, some leaked a KeyError from deep inside, one returned an empty dict. Callers kept getting it wrong, and the worst bug was a None passed along until something failed three layers away. So I set one rule. If not finding it is normal, like checking an email at signup, the function is find_customer and returns Customer | None, so the type checker makes callers deal with None. If the ID came from our own database and must exist, it's get_customer and raises our own CustomerNotFound, so the error points at the real problem. Anything returning many items returns an empty list, never None. The cost was renaming functions across four teams' code over a few releases, with deprecation warnings on the old names.”

Code:
class CustomerNotFound(LookupError):
    pass

def find_customer(email: str) -> "Customer | None":
    return _by_email.get(email.lower())

def get_customer(customer_id: int) -> "Customer":
    try:
        return _by_id[customer_id]
    except KeyError:
        raise CustomerNotFound(customer_id) from None
Red flag to avoid:

Saying they always return None, or letting a KeyError from inside the module leak out to callers.

They may ask next:
  • Why subclass LookupError for your own not-found exception?
  • When would you return a result object instead of raising?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

15. For data coming into a service you owned, did you pick plain dataclasses, attrs or pydantic? What did that choice cost you?

What the interviewer is really testing:
Whether they chose deliberately between lightweight containers and runtime validation, and know where each one hurts.
Answer frame:

The need: untrusted input at the edge versus trusted data inside.

Runtime validation: parses and checks types at the boundary, with clear error messages.

Plain classes: light, fast and dependency-free, but they check nothing.

Cost: speed, a dependency to keep upgrading, or bugs you let through.

Sample spoken answer:

“On an ingestion API I owned, I used pydantic models at the edge, because the input came from partners and was often wrong. It gave us type coercion and error messages clients could act on, instead of a KeyError deep inside. Inside the service I used plain dataclasses, since that data was already validated and I didn't want to pay the checking cost on every internal object. The split had costs. A major version of the validation library changed some behaviour and took a sprint to work through, and I had to keep an eye out for people quietly passing raw dicts past the edge. For a small internal tool I'd use dataclasses alone. The rule I'd defend is: validate once, at the boundary.”

Red flag to avoid:

Picking a library because it's popular, with no idea where validation belongs or what it costs.

They may ask next:
  • Where exactly would you draw the boundary in a service fed by both a queue and an API?
  • How would you stop unvalidated data sneaking past the edge in a later change?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

16. Your service put pickled Python objects on a job queue. Why would you move the messages to a plain format like JSON, and what does that cost?

What the interviewer is really testing:
Whether they know a pickle ties each message to the exact code that wrote it and can run code when loaded, and can weigh that against the convenience.
Answer frame:

Coupling: a pickle refers to classes by module and name, so moving or renaming a class breaks messages already waiting.

Safety: loading a pickle can run arbitrary code, so only load data you fully trust.

Replacement: a versioned JSON message with explicit plain fields, often just IDs.

Cost: conversion code to write, and datetimes and decimals handled by hand.

Sample spoken answer:

“On a job system I owned, producers pickled whole objects onto the queue because it was one line of code. It bit us during a deploy: we moved a class to a new module, and every message already waiting failed to load, because a pickle finds the class by its module and name. Workers on the new code couldn't read what the old code had written, and we drained the queue by hand that night. Around the same time a security review pointed out that loading a pickle can run arbitrary code, and that queue was shared with another team's service. So I moved the payloads to JSON with a version field and only plain values: IDs, strings, numbers, timestamps as ISO strings. Workers load the full record from the database by ID. The cost was writing conversion code and handling datetimes and decimals by hand, but deploys stopped caring what was in the queue.”

Code:
import json

def to_message(order_id, requested_at):
    return json.dumps({
        "version": 2,
        "order_id": order_id,
        "requested_at": requested_at.isoformat(),
    })
Red flag to avoid:

Not knowing that loading a pickle from an untrusted source can run code, or that pickled messages break when classes move.

They may ask next:
  • How would you roll out a new message version while old messages are still in the queue?
  • When is pickle still a reasonable choice?
Say it in 60 seconds

Performance 1 question

Hard Technical round Senior Practice question

17. Your service's p99 latency spikes every few minutes, but CPU and traffic look flat. How would you check whether Python's garbage collector is behind it?

What the interviewer is really testing:
Whether they know the cycle collector can pause a process when it walks many long-lived objects, and prove it with measurements before tuning anything.
Answer frame:

Suspect: a big heap of long-lived objects makes full collections slow.

Prove: time each collection with gc.callbacks and line the slow ones up with the spikes.

Fix: fewer reference cycles, gc.freeze() after startup, or tuned thresholds.

Trade-off: what freezing or tuning risks.

Sample spoken answer:

“On a pricing service I owned, p99 latency jumped every few minutes while CPU and traffic looked flat. Nothing external lined up, so I suspected the cycle collector, because we kept a huge in-memory lookup table of long-lived objects that a full collection had to walk. I added a gc.callbacks hook that logged how long each collection took and which generation, and the slow ones were the oldest generation, right on the spikes. Two fixes helped. We loaded the table at startup and called gc.freeze() afterwards, so the collector stopped rescanning it, and we fixed a code path that created reference cycles on every request. The trade-off is that frozen objects are never checked for cycles again, so anything frozen that later ends up in a cycle leaks. That's why I froze right after loading, before serving requests.”

Code:
import gc
import logging
import time

_started = {}

def gc_timer(phase, info):
    if phase == "start":
        _started["t"] = time.perf_counter()
    elif phase == "stop":
        ms = (time.perf_counter() - _started["t"]) * 1000
        if ms > 20:
            logging.warning(f"gc generation {info['generation']} took {ms:.1f} ms")

gc.callbacks.append(gc_timer)
Red flag to avoid:

Tuning or disabling the garbage collector on a hunch, without measuring collection times first.

They may ask next:
  • Why can't reference counting alone free objects that point at each other?
  • What would make you switch automatic collection off entirely, and what would you do instead?
Say it in 60 seconds

Leadership 3 questions

Medium Situational round Mid-level, Senior Practice question

18. To ship faster, a teammate wants to send emails from a background thread inside your web workers instead of adding a job queue. Do you agree?

What the interviewer is really testing:
Whether they can push back on a shortcut by naming what would actually go wrong, and still find a middle path that fits the deadline.
Answer frame:

Ask what matters: how bad is a lost or doubled email for this case?

Name the risks: the thread dies when a worker restarts or a deploy happens, there's no retry, and failures go unseen.

Middle path: write the email to an outbox table in the same transaction, and let one small worker send it.

Draw the line: when the shortcut is fine and when it isn't.

Sample spoken answer:

“I had this exact conversation about receipt emails. I didn't just say no; I asked how bad a lost email would be, and for receipts the answer was support tickets. Then I explained what would really happen: our web workers get recycled after a set number of requests and on every deploy, and a thread that's halfway through sending just dies with them. There's no retry, and an exception in that thread only gets printed to stderr, so failures would be invisible. A full queue setup was more than the deadline allowed, so I suggested a middle path: write the email as a row in an outbox table in the same transaction as the order, and have one small worker process send pending rows and mark them sent. It took two extra days. For a low-stakes 'someone viewed your profile' notice, I agreed the thread was fine.”

Red flag to avoid:

Blocking the shortcut on principle with no cheaper option, or agreeing without asking what a lost email costs.

They may ask next:
  • What happens if the worker crashes after sending an email but before marking it sent?
  • How would you raise this if the teammate were more senior than you?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

19. Tell me about a junior developer you helped grow on a Python team. What did you change about how you helped them along the way?

What the interviewer is really testing:
Whether they mentor deliberately, adapt to the person, and let the junior own work instead of fixing everything for them.
Answer frame:

Starting point: where the person was and what held them back.

What you did: pairing, smaller tasks, clearer reviews.

What you changed: where your first approach didn't work.

Result: what they could do on their own afterwards.

Sample spoken answer:

“A new graduate joined my team and was slow to ship; their pull requests sat for days. My first move was detailed reviews, and that backfired: long comment lists made them more anxious and slower. So I changed approach. We paired for an hour twice a week on their own tickets, with them typing. I split tasks into pieces that could merge within a day, and in reviews I marked each comment as must-fix or optional so they knew what mattered. I also gave them a small module to own, the CSV importer, alerts included. Within three months they were reviewing other people's code in that area and handled an incident there without me. The cost was some of my own output for a while, which I agreed with my manager up front.”

Red flag to avoid:

Describing mentoring as answering questions when asked, or quietly rewriting the junior's code for them.

They may ask next:
  • How did you know when to step back?
  • What would you do with a junior who doesn't take feedback well?
Say it in 60 seconds
Hard Behavioral round Mid-level, Senior Practice question

20. Tell me about a large refactor you led in a Python codebase. How did you keep it shippable while it was in progress?

What the interviewer is really testing:
Whether they can lead a multi-week change without a long-lived branch, keep the team unblocked and prove behaviour didn't change.
Answer frame:

Why: the pain that justified it, in terms the team and manager cared about.

Plan: small steps behind a stable interface, each one mergeable on its own.

Safety: characterisation tests, flags, and old and new code run side by side.

Result and cost: what improved and what it took.

Sample spoken answer:

“I led a refactor of our pricing code, which had grown into a two-thousand-line module where every change broke something. I sold it to my manager on lead time: small pricing changes were taking a week. We never had a big branch. First I wrote characterisation tests that captured current prices for a few thousand real orders. Then I put a new interface in front of the old code and moved one family of rules at a time behind it, each move a small pull request merged the same day. For the riskiest part, the new code ran alongside the old in production and logged any difference, and we only switched after a full week with none. It took two months with feature work continuing, and pricing changes now take a day or two.”

Red flag to avoid:

A months-long branch merged in one go, or no way to show the behaviour stayed the same.

They may ask next:
  • What did you do when the side-by-side run found a difference?
  • How did you split the work so others could help?
Say it in 60 seconds

Code Quality 2 questions

Medium Technical round Mid-level, Senior Practice question

21. What does the test suite for a service you own look like, and what did you deliberately decide not to test?

What the interviewer is really testing:
Whether they design testing for confidence per minute of CI rather than a coverage number, and can defend what they left out.
Answer frame:

Shape: many fast unit tests on logic, fewer integration tests against real dependencies, a thin end-to-end layer.

Real dependencies: where mocks lied, and where a real database in a container paid off.

Left out: what you don't test and why.

Speed: how long CI takes and what you did about it.

Sample spoken answer:

“For an order service I owned, most tests are fast unit tests on pure functions: pricing, discounts, state changes. Then there are about forty integration tests against a real database in a container, because mocked queries had lied to us twice about constraint errors. On top sits a handful of end-to-end checks of the main flows. What I chose not to test: thin glue code and the framework's own behaviour, like whether the router routes. I also don't chase a coverage number; we dropped our target because people were writing tests that asserted nothing. Integration tests are slower, so CI runs the unit tests on every push and the full set before merge. The whole suite takes about six minutes.”

Red flag to avoid:

Quoting a coverage target as the whole strategy, or mocking everything including the database.

They may ask next:
  • What made you trust a real database more than mocks here?
  • How do you decide a bug deserves a new test?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

22. A junior's pull request solves the problem with a metaclass and dynamic attribute tricks. It works. How do you review it?

What the interviewer is really testing:
Whether they review for the team's ability to maintain the code, explain the reasoning, and teach rather than just reject.
Answer frame:

Acknowledge: it works, and the instinct behind it is sound.

Name the cost: harder to read, debug, type-check and search; surprises for the next person.

Offer a simpler path: a plain class, a dict of handlers or a decorator-based registry.

Deliver it well: a short call rather than a wall of comments.

Sample spoken answer:

“I had almost this exact pull request: a metaclass that auto-registered handler classes and generated attributes through __getattr__. It worked, and I said so first, because the instinct to remove repetition was right. Then I explained the cost in concrete terms: our type checker couldn't see those attributes, editors couldn't jump to definitions, and a typo returned something odd instead of raising an error. I suggested a plain dict registry filled by a small decorator, which kept the auto-registration and dropped the magic. Rather than leave twenty comments, I had a fifteen-minute call and we rewrote one handler together; they did the rest. The trade-off I accepted was a little more code to type. A metaclass is fine when nothing simpler works, but that bar is high.”

Red flag to avoid:

Approving it because it works, or rejecting it as too clever with no explanation and no alternative.

They may ask next:
  • When would you actually approve a metaclass in production code?
  • How do you keep reviews from feeling like personal criticism?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card