Testing • Debugging • Standard Library • Asyncio • Real Work • 2026

Python Interview Questions for 3 Years Experience (2 to 4 Years)

Python interviews for three years of experience (two to four years) stop asking what a list is and start asking what you built and fixed yourself: how you test your code, how you tracked down a strange bug, what went wrong with dates, encodings or caching, and what review feedback changed your habits. Concepts come up as 'how did this behave in your project'. It is written for Python developers with about two to four years of work behind them. Each question has the answer's shape and a spoken sample. Swap in your own stories wherever you can; a real example beats a perfect definition.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Testing 3 questions

Medium Technical round Mid-level Practice question

1. How do you test your own Python code day to day? Walk me through how you use pytest fixtures and parametrize.

What the interviewer is really testing:
Whether testing is a habit in their daily work, and whether they know the two pytest features that keep a test file short and readable.
Answer frame:

Fixtures: shared setup lives in one function that tests ask for by name; conftest.py shares it across files.

Parametrize: one test body runs over a table of inputs and expected results.

Habit: run the relevant tests on every change, the full suite before pushing.

Sample spoken answer:

“I write the test next to the change, usually before I open the pull request. For setup I use fixtures: a function marked with @pytest.fixture that builds, say, a sample cart, and any test that names it as an argument gets a fresh copy. Fixtures used across files go in conftest.py. When I'm checking the same logic against many inputs, I use @pytest.mark.parametrize so each case shows up as its own test in the report, and a failure tells me exactly which input broke. While working I run just the file I'm touching with -k or --lf for last-failed, and the whole suite runs in CI. For floats I compare with pytest.approx rather than equals.”

Code:
import pytest
from pricing import apply_discount

@pytest.fixture
def cart():
    return {"items": [10.0, 20.0], "coupon": None}

@pytest.mark.parametrize("coupon, expected", [
    (None, 30.0),
    ("HALF", 15.0),
])
def test_apply_discount(cart, coupon, expected):
    cart["coupon"] = coupon
    assert apply_discount(cart) == pytest.approx(expected)
Red flag to avoid:

Saying they test by running the script and eyeballing the output, or sharing one mutable object across tests so they depend on each other's order.

They may ask next:
  • What does fixture scope do, and when have you used a scope wider than function?
  • How do you clean up after a fixture, for example a temp file or a database row?
Say it in 60 seconds
Hard Technical round Mid-level Practice question

2. You patched a function with unittest.mock.patch, but the real one still ran in your test. What was most likely wrong?

What the interviewer is really testing:
Whether they understand that patch replaces a name in one namespace, which is the most common reason a mock silently does nothing.
Answer frame:

Cause: the module under test imported the function with from x import f, so it holds its own reference.

Rule: patch where the name is looked up, not where it is defined.

Check: assert the mock was called, so a wrong target fails loudly.

Sample spoken answer:

“Nine times out of ten it's the patch target. patch swaps a name inside one module. If report.py does from app.clock import today, then report has its own name today pointing at the real function. Patching app.clock.today changes the clock module, but report never looks there again, so the real one runs. The fix is to patch where it's used: app.report.today. If the module had done import app.clock and called app.clock.today(), then patching app.clock.today would work, because the lookup happens at call time. I learned this the hard way on a test that passed only on certain days. Now I always add assert_called_once() on the mock so a wrong target shows up as a failure.”

Code:
# app/report.py
from app.clock import today      # report now holds its own reference

# tests/test_report.py
from datetime import date
from unittest.mock import patch

with patch("app.clock.today"):   # wrong: report never looks here again
    ...

with patch("app.report.today", return_value=date(2026, 1, 5)) as fake:
    build_report()
    fake.assert_called_once()
Red flag to avoid:

Blaming the mock library or adding more patches until it works, without being able to say which namespace the code reads the name from.

They may ask next:
  • When would you use patch.object instead of a string target?
  • What is autospec, and what bug does it catch?
Say it in 60 seconds
Hard Situational round Mid-level Practice question

3. A test passes every time on your laptop but fails in CI about once a week. How do you track it down?

What the interviewer is really testing:
Whether they treat a flaky test as a real bug with a small set of usual causes, instead of re-running the pipeline until it goes green.
Answer frame:

Reproduce: run it many times, in random order, and alone versus with the full suite.

Usual causes: shared state between tests, real time or dates, randomness, network calls, dict or set ordering assumptions.

Fix the cause: isolate state, freeze time, seed randomness, fake the network; never just add a retry.

Sample spoken answer:

“First I stop treating it as noise. I read the CI failure closely, the time it ran and what else ran before it. Then I try to make it fail locally: run the test in a loop, run it alone and then with the full suite, and shuffle the test order. Most of the ones I've hit came from shared state, like a module-level list one test appends to and another reads, or a fixture with too wide a scope. Another classic was a test comparing against today's date that broke near midnight in UTC on the CI machine. Once I know the cause I fix it properly: fresh state per test, a fixed clock, seeded randomness, no real network. Marking it skip or adding retries just hides a bug that might be in the product too.”

Red flag to avoid:

Saying they would just re-run the job or add a retry decorator, with no plan to find why it fails.

They may ask next:
  • How would you find which earlier test is leaking state into this one?
  • When, if ever, is quarantining a flaky test acceptable?
Say it in 60 seconds

Debugging 3 questions

Medium Technical round Mid-level Practice question

4. A function returns the wrong value only for some inputs. Walk me through how you debug it. Print, logging or a debugger?

What the interviewer is really testing:
Whether they have a calm, repeatable way to narrow a bug down, and whether they actually use the debugger that ships with Python.
Answer frame:

Reproduce: capture a failing input and turn it into a small test first.

Narrow: breakpoint() or a post-mortem session to inspect state at the moment it goes wrong.

Keep: the test stays, so the bug can't come back.

Sample spoken answer:

“I start by getting one failing input and writing a test that shows the wrong result, so I can rerun it in a second. Then I drop breakpoint() just before the suspect line, which opens pdb. I step with n, go into calls with s, print variables with p, and use w to see how I got there. If it crashes rather than returning a bad value, I run it under python -m pdb -c continue and it stops right at the exception with the frame intact. A quick print is fine for a one-minute check, but I remove it before committing. Logging is for things I want to see in production later, not for this hunt. Once fixed, the test stays in the suite.”

Code:
def split_invoice(total, parts):
    breakpoint()            # opens pdb here; remove before committing
    share = total // parts
    return [share] * parts

# crash hunting without editing code:
#   python -m pdb -c continue job.py
Red flag to avoid:

Only ever adding print statements and committing them, or changing code at random until the output looks right.

They may ask next:
  • How do you debug something that only goes wrong in production, where you can't attach a debugger?
  • How would you set a breakpoint that only triggers for one specific input?
Say it in 60 seconds
Easy Technical round Mid-level Practice question

5. Your job crashed with 'RuntimeError: dictionary changed size during iteration'. What causes it, and how did you fix it?

What the interviewer is really testing:
Whether they recognise a very common runtime error from real code and know more than one clean fix, with the trade-off of each.
Answer frame:

Cause: adding or deleting keys while a loop is walking the same dict.

Fix one: loop over a snapshot, list(d), and delete from the original.

Fix two: build a new dict with a comprehension, keeping in mind others may hold the old one.

Sample spoken answer:

“It happens when you add or remove keys from a dict while you're looping over it. In my case it was a cleanup loop deleting expired sessions from a cache dict. Python notices the size changed under the iterator and raises instead of giving you unpredictable results. There are two clean fixes. I can loop over a copy of the keys with list(sessions) and delete from the real dict, which keeps the same object, so anything else holding a reference sees the change. Or I build a new dict with a comprehension that keeps only the live entries. That's tidier, but it rebinds the name, so if another module holds the old dict, it won't see the cleanup. In a shared cache I used the first one.”

Code:
# fails: RuntimeError
for key, s in sessions.items():
    if s.expired:
        del sessions[key]

# fix 1: loop over a snapshot, change the same dict
for key in list(sessions):
    if sessions[key].expired:
        del sessions[key]

# fix 2: new dict (other references keep the old one)
sessions = {k: s for k, s in sessions.items() if not s.expired}
Red flag to avoid:

Wrapping the loop in try and except to swallow the error, or not knowing why a new dict and an in-place change behave differently.

They may ask next:
  • Is changing a value for an existing key inside the loop also a problem?
  • What about removing items from a list while looping over it?
Say it in 60 seconds
Hard Technical round Mid-level Practice question

6. A sort in your code crashed with a TypeError saying '<' isn't supported between NoneType and str. Why, and how do you sort by last name, missing ones last, then newest signup first?

What the interviewer is really testing:
Whether they know Python 3 refuses to order unlike types, and can build a sort key that handles missing values and mixed directions by leaning on sort stability.
Answer frame:

Cause: Python 3 won't compare None with a string, so one missing value breaks the whole sort.

Key: a tuple like (name is None, name or "") puts missing values last on purpose.

Mixed directions: sort by the secondary key first, then the primary; the sort is stable, so ties keep the earlier order.

Sample spoken answer:

“Python 3 won't order values of different types, so the moment one record had None as its last name, comparing it with a string raised a TypeError and the whole sort failed. Our tests only had complete records. The fix is a key that never compares None directly: a tuple of name is None and name or "". False sorts before True, so missing names go to the end, and the empty string stands in so the second slot is always a string. For newest signup first within the same name, the tidy trick is two sorts. Python's sort is stable, so I sort by signup date descending first, then by name. Records with the same name keep the date order from the first pass. With a numeric field I could negate it inside one key instead.”

Code:
rows.sort(key=lambda r: r["signed_up"], reverse=True)   # secondary key first
rows.sort(key=lambda r: (r["last_name"] is None,          # missing names last
                         r["last_name"] or ""))
Red flag to avoid:

Filtering out the rows with None to make the error go away, or not knowing that Python's sort is stable.

They may ask next:
  • How would you make the name order ignore upper and lower case?
  • How would you get just the ten newest rows out of a million without sorting all of them?
Say it in 60 seconds

Tooling & Setup 3 questions

Medium Technical round Mid-level Practice question

7. How did you set up logging in your last Python project, and why not just use print?

What the interviewer is really testing:
Whether they use the logging module the way it is designed: named loggers per module, configuration once at the entry point, and exceptions logged with tracebacks.
Answer frame:

Per module: logging.getLogger(__name__) in each file.

Configure once: handlers, level and format set at the entry point, never inside library code.

Use it well: levels that mean something, logger.exception in except blocks, pass arguments instead of formatting early.

Sample spoken answer:

“Each module gets its own logger with logging.getLogger(__name__), so every line says where it came from, and I can turn one noisy module down without touching the rest. The configuration, level, format and handlers, happens once, in the entry point or the app's startup, never in the modules themselves. Print can't do any of that: no levels, no timestamps, no way to send it to a file or a log service, and you can't switch it off. Inside an except block I use logger.exception, which records the full traceback. And I pass values as arguments to the log call rather than building an f-string, so the message is only formatted if that level is actually on, and log tools can group lines with the same template.”

Code:
import logging

logger = logging.getLogger(__name__)

def sync_orders(batch):
    logger.info("syncing %d orders", len(batch))
    try:
        push(batch)
    except TimeoutError:
        logger.exception("push failed, first id %s", batch[0].id)
        raise

if __name__ == "__main__":
    logging.basicConfig(level=logging.INFO,
                        format="%(asctime)s %(levelname)s %(name)s %(message)s")
Red flag to avoid:

Calling logging.basicConfig inside library modules, or logging an exception with only its message and losing the traceback.

They may ask next:
  • How would you add a request id to every log line without passing it into every function?
  • What have you had to be careful not to log?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

8. Running a file inside your package with python app/jobs/sync.py fails on an import, but python -m app.jobs.sync works. Why?

What the interviewer is really testing:
Whether they understand how Python finds modules when a script starts, which explains a whole family of confusing import errors in real projects.
Answer frame:

Script mode: the file runs as __main__ with no parent package, and its own folder goes on the path.

Module mode: -m puts the current folder on the path and keeps the package, so both import styles work.

Lasting fix: run with -m, install the project in editable mode, or add an entry point; don't edit sys.path.

Sample spoken answer:

“When you run a file by its path, Python treats it as a standalone script. It's named __main__, it has no parent package, and the folder it lives in, app/jobs, goes first on the import path instead of the project root. So a relative import like from .utils import x fails with 'attempted relative import with no known parent package', and from app.utils import x fails because app can't be found. With python -m app.jobs.sync run from the project root, the current folder goes on the path and Python knows the package the module belongs to, so both work. On my team we fixed it for good by installing the project with pip install -e . and adding a console entry point, and we stopped people patching sys.path at the top of files.”

Red flag to avoid:

Fixing it by inserting paths into sys.path at the top of each file without understanding why the import failed.

They may ask next:
  • How would you make this job runnable as a plain command after the project is installed?
  • What does an editable install do behind the scenes?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

9. How did you handle settings like database URLs and feature flags across dev, staging and production? Did an environment variable ever bite you?

What the interviewer is really testing:
Whether they keep config out of the code, read it once at startup with proper types, and know the classic trap that every environment variable arrives as a string.
Answer frame:

Out of code: settings and secrets come from environment variables or a secrets store, never from the repo.

Read once: one settings module parses and checks everything at startup and fails fast if a required value is missing.

Types: every value is a string, so "false" is truthy; convert on purpose.

Sample spoken answer:

“We kept settings out of the code. Each environment set its own environment variables, and secrets came from the platform's secret store, never from the repo. One settings module read them all once at startup, turned them into the right types, and stopped the app with a clear message if a required one was missing, instead of crashing an hour later on the first database call. The bite: someone wrote DEBUG = os.environ.get("DEBUG", False). Every environment variable is a string, so when staging set it to "false", that non-empty string was truthy and debug mode came on. We fixed it by parsing booleans against a short list of accepted words. For local work we used a .env file that git ignores, plus an example file with dummy values.”

Code:
import os

def env_bool(name, default=False):
    raw = os.environ.get(name)
    if raw is None:
        return default
    return raw.strip().lower() in {"1", "true", "yes", "on"}

DEBUG = env_bool("DEBUG")
DATABASE_URL = os.environ["DATABASE_URL"]      # KeyError at startup if missing
TIMEOUT_SECONDS = float(os.environ.get("TIMEOUT_SECONDS", "10"))
Red flag to avoid:

Hard-coding credentials in the source, or reading environment variables all over the codebase with no type conversion.

They may ask next:
  • How do you keep secrets out of your logs and error reports?
  • How do you override a setting in one test without it leaking into the others?
Say it in 60 seconds

OOP & Data Classes 2 questions

Medium Technical round Mid-level Practice question

10. Have you replaced dicts passed around your code with dataclasses? What did you gain, and what caught you out?

What the interviewer is really testing:
Whether they have refactored real code towards clearer data shapes and hit the practical dataclass gotchas, rather than reciting what the decorator generates.
Answer frame:

Gain: named fields, type hints, a readable repr and value equality for free; typos become errors.

Gotcha: a list or dict default needs field(default_factory=...).

Frozen: blocks reassignment, but a list inside can still change and makes hashing fail.

Sample spoken answer:

“Yes. We had order data flowing through a pipeline as dicts, and a typo in a key would just give a KeyError three functions later. I moved it to a dataclass: named, typed fields, a clear repr in logs, equality by value, and my editor autocompletes the fields. Two things caught me. First, I wrote tags: list[str] = [] and Python refused with a ValueError, because every instance would share one list. The fix is field(default_factory=list). Second, I made it frozen=True expecting it to be fully read-only and hashable, but the list inside could still be appended to, and calling hash on it raised a TypeError because lists aren't hashable. For truly fixed data I switched that field to a tuple.”

Code:
from dataclasses import dataclass, field

@dataclass(frozen=True)
class Shipment:
    order_id: str
    weight_kg: float
    tags: tuple[str, ...] = ()          # tuple keeps it truly fixed

@dataclass
class Batch:
    shipments: list[Shipment] = field(default_factory=list)
Red flag to avoid:

Thinking frozen=True makes everything inside immutable, or not knowing why a mutable default is refused.

They may ask next:
  • When would you choose a NamedTuple or a TypedDict instead of a dataclass?
  • How do you add validation when a dataclass is created?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

11. When have you used @property in a class, and when did it cause trouble?

What the interviewer is really testing:
Whether they know properties are for cheap, attribute-like values, and can spot the hidden costs a property can bring into a codebase.
Answer frame:

Good use: a derived value or a validated setter, while callers keep attribute syntax.

Trouble: slow work or I/O hidden behind what looks like a plain attribute.

Alternative: functools.cached_property for costly values that won't change, a method when it does real work.

Sample spoken answer:

“I use it for values derived from other fields, like a full name from first and last name, and for setters that validate, so a negative quantity raises straight away. Callers still write item.quantity, so I can add the check without changing any call sites. It caused trouble once when a teammate's property ran a database query. It looked like a plain attribute, so someone used it inside a loop over a thousand rows and the page got very slow. We turned it into a method called load_history(), because a method name tells the reader it does work. For expensive values that don't change after creation, cached_property computes once per instance, but it goes stale if the underlying data changes.”

Code:
from functools import cached_property

class Report:
    def __init__(self, rows):
        self.rows = rows

    @property
    def row_count(self):          # cheap, always fresh
        return len(self.rows)

    @cached_property
    def summary(self):            # costly, computed once per instance
        return build_summary(self.rows)
Red flag to avoid:

Hiding network or database calls behind properties, or using getters and setters for every field out of habit from other languages.

They may ask next:
  • How does cached_property store its value, and how do you clear it?
  • Why can't you use cached_property on a class that defines __slots__ without a __dict__?
Say it in 60 seconds

I/O & Concurrency 3 questions

Medium Technical round Mid-level Practice question

12. Your code calls a third-party HTTP API. What do you put around that call so it can't hang or sink the whole job?

What the interviewer is really testing:
Whether they have been burned by network calls in real work and know the defaults that make them dangerous, starting with the missing timeout.
Answer frame:

Timeout: always set one; the popular client waits forever without it.

Retries: only on safe calls and temporary errors, with backoff.

Fail well: check the status, log the context, decide whether one failure stops the job.

Sample spoken answer:

“The first thing is a timeout, because by default the requests library will wait forever, and I've seen a nightly job sit for hours on one stuck connection. I pass a connect and a read timeout. Then I use a session, which reuses connections, and mount a retry policy for temporary failures like 502, 503 or a rate-limit 429, with a growing delay between tries. I only retry calls that are safe to repeat, like GETs, unless the API supports idempotency keys. After the call I use raise_for_status() so an error page doesn't get parsed as data. And in a batch job, one bad record logs its id and moves on, instead of killing the whole run.”

Code:
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

session = requests.Session()
retry = Retry(total=3, backoff_factor=0.5,
              status_forcelist=[429, 502, 503, 504],
              allowed_methods=["GET"])
session.mount("https://", HTTPAdapter(max_retries=retry))

resp = session.get(url, timeout=(3.05, 10))   # connect, read
resp.raise_for_status()
Red flag to avoid:

Making network calls with no timeout, or retrying every error in a tight loop with no delay.

They may ask next:
  • How would you test this code without calling the real API?
  • How would you respect a rate limit the API tells you about in its response headers?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

13. A CSV import you wrote crashes with a UnicodeDecodeError on one customer's file. How do you handle it?

What the interviewer is really testing:
Whether they understand that text files carry an encoding the code must choose, and handle messy real-world files without silently corrupting data.
Answer frame:

Cause: the file isn't in the encoding open assumed; without encoding= the default depends on the machine.

Be explicit: open with a stated encoding, utf-8-sig to drop a byte order mark, and newline="" for csv.

Bad files: tell the user clearly, or try a known fallback; replacing characters is the last resort.

Sample spoken answer:

“It means the bytes aren't valid in the encoding I opened the file with. The first thing I check is whether my code even named one, because without encoding= the Python versions we ran used the machine's locale default, which is why it can work on my laptop and fail on a server. I always open with an explicit encoding now. For CSVs I use utf-8-sig, which also strips the invisible byte order mark some spreadsheet tools add, otherwise the first column header comes out with junk at the front. The csv module also wants newline="". When a customer really sends a different encoding, I either reject it with a clear message saying which row failed, or try one agreed fallback. I avoid errors="ignore", because it quietly drops characters from names.”

Code:
import csv

def read_rows(path):
    with open(path, newline="", encoding="utf-8-sig") as f:
        for line_no, row in enumerate(csv.DictReader(f), start=2):
            yield line_no, row
Red flag to avoid:

Adding errors="ignore" to make the error go away without realising it silently deletes customer data.

They may ask next:
  • How would you process a two-gigabyte CSV without loading it into memory?
  • What is the difference between bytes and str in Python 3, and where do you convert?
Say it in 60 seconds
Hard Technical round Mid-level Practice question

14. You used asyncio.gather to call many endpoints at once, and one failure lost all the results. How do you handle partial failures?

What the interviewer is really testing:
Whether they know how gather behaves when one task raises, and can collect partial results while keeping concurrency under control.
Answer frame:

Default: the first exception is raised from gather, and you lose the list of results.

Keep results: return_exceptions=True puts errors in the list next to successes.

Control load: a semaphore caps how many requests run at once.

Sample spoken answer:

“By default, if one coroutine in gather raises, that exception comes straight out of the await, so I never get the result list, even though the other calls may have finished fine. They aren't cancelled either, so they keep running in the background. When I want whatever succeeded, I pass return_exceptions=True. Then the list has a result or an exception in each slot, in the same order as the inputs, and I split them into successes and failures and retry or log the failures by URL. I also wrap each call in a semaphore, because firing two thousand requests at once got us rate-limited. If one failure should cancel everything, a TaskGroup is the better fit on newer Python versions.”

Code:
import asyncio

async def fetch_all(client, urls, limit=20):
    sem = asyncio.Semaphore(limit)

    async def one(url):
        async with sem:
            return await client.get(url)

    results = await asyncio.gather(*(one(u) for u in urls),
                                   return_exceptions=True)
    ok = [r for r in results if not isinstance(r, BaseException)]
    failed = [(u, r) for u, r in zip(urls, results)
              if isinstance(r, BaseException)]
    return ok, failed
Red flag to avoid:

Not knowing that the other tasks keep running after the first error, or launching unlimited concurrent requests.

They may ask next:
  • What is the difference between gather and a TaskGroup when one task fails?
  • How would you put an overall time limit on the whole batch?
Say it in 60 seconds

Standard Library 5 questions

Medium Technical round Mid-level Practice question

15. Tell me about a date or timezone bug you've hit in Python. What is the difference between a naive and an aware datetime?

What the interviewer is really testing:
Whether they have handled time correctly in real code: store in UTC, keep datetimes aware, and convert to local time only for display.
Answer frame:

Naive vs aware: a naive datetime has no tzinfo; an aware one knows its offset.

The bug: mixing them raises a TypeError, or naive local times get stored as if they were UTC.

Habit: datetime.now(timezone.utc), store UTC, convert with zoneinfo only for display.

Sample spoken answer:

“A naive datetime has no timezone attached, so it's just a clock reading; an aware one carries tzinfo and pins down an exact moment. My bug was a report job that used datetime.now() on a server set to local time, while the database stored UTC. Records near midnight landed in the wrong day. Another time, comparing a naive and an aware datetime just raised a TypeError. Since then I create times with datetime.now(timezone.utc), keep everything aware, store UTC, and convert to a user's zone with zoneinfo only when showing it. I also stopped using datetime.utcnow(): it returns a naive value that looks like UTC but doesn't say so, and recent Python versions deprecate it.”

Code:
from datetime import datetime, timezone
from zoneinfo import ZoneInfo

created = datetime.now(timezone.utc)                 # aware, store this
shown = created.astimezone(ZoneInfo("Europe/Berlin"))  # display only
Red flag to avoid:

Storing local server time in the database, or saying timezones don't matter because all users are in one place.

They may ask next:
  • How would you find all records created on a given local date for a user in another timezone?
  • What goes wrong when you schedule a daily job at 02:30 local time on a daylight-saving change day?
Say it in 60 seconds
Easy Technical round Mid-level Practice question

16. Your endpoint fails with 'Object of type datetime is not JSON serializable'. How do you fix it cleanly?

What the interviewer is really testing:
Whether they know how to extend the standard json module properly, instead of converting values by hand all over the codebase.
Answer frame:

Why: json only knows dicts, lists, strings, numbers, booleans and None.

Fix: pass a default function (or subclass JSONEncoder) that converts known types.

Be strict: raise TypeError for anything unexpected rather than calling str on everything.

Sample spoken answer:

“The json module only knows the basic types: dicts, lists, strings, numbers, true, false and null. A datetime or a Decimal isn't one of them, so it gives up. Rather than converting fields by hand everywhere, I give json.dumps a default function. It's called only for objects json doesn't understand, and I return an ISO string for datetimes and a string for Decimals, so precision isn't lost by going through float. For anything else I raise TypeError on purpose, because turning every unknown object into str hides bugs, like a whole model object leaking into an API response. If the project uses a web framework, I plug the same function into its encoder so there's one place for the rule.”

Code:
import json
from datetime import datetime
from decimal import Decimal

def to_json(obj):
    if isinstance(obj, datetime):
        return obj.isoformat()
    if isinstance(obj, Decimal):
        return str(obj)
    raise TypeError(f"Cannot serialize {type(obj).__name__}")

body = json.dumps(payload, default=to_json)
Red flag to avoid:

Using default=str for everything, which hides bugs, or scattering manual .isoformat() calls across the codebase.

They may ask next:
  • How would you turn those ISO strings back into datetimes when reading the JSON?
  • Why is sending a Decimal as a JSON number risky for the client?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

17. Your report's totals disagreed with the finance team's spreadsheet in the last digit on a few rows. What was going on, and how did you fix it?

What the interviewer is really testing:
Whether they understand binary floating point well enough to choose the right type for exact values, and know Python's rounding surprises.
Answer frame:

Why: floats are binary, so values like 0.1 can't be stored exactly and small errors add up.

Decimal: exact base-ten arithmetic; build it from strings, round with quantize.

Rounding: built-in round rounds halves to even, which surprises people.

Sample spoken answer:

“Two things were going on. We summed amounts as floats, and floats are binary, so most decimal fractions can't be stored exactly and tiny errors add up over thousands of rows. On that invoicing feature I switched to Decimal for anything that had to match exactly. The key detail is building it from strings, Decimal("0.1"), because Decimal(0.1) copies the float's error. The second thing was rounding. The built-in round(2.5) gives 2, because it rounds halves to the even number, while the spreadsheet rounded halves up. So I round with quantize and an explicit rounding mode the finance team agreed on, in one helper everyone calls. For measurements and science code, floats are fine.”

Code:
from decimal import Decimal, ROUND_HALF_UP

Decimal(0.1)     # carries the float's error: Decimal('0.1000000000000000055511...')
Decimal("0.1")   # Decimal('0.1'), exact

Decimal("2.675").quantize(Decimal("0.01"), rounding=ROUND_HALF_UP)  # Decimal('2.68')
round(2.5)       # 2, halves go to the even number
round(2.675, 2)  # 2.67, the float is really a hair below 2.675
Red flag to avoid:

Fixing the drift with round() calls sprinkled everywhere, or creating Decimals from floats.

They may ask next:
  • How do you compare two floats safely in a test?
  • Where would you store these values in the database, and as what type?
Say it in 60 seconds
Medium Technical round Mid-level Practice question

18. How do you run a shell command from Python safely, and get its output and errors back?

What the interviewer is really testing:
Whether they use subprocess.run correctly in scripts and know why shell=True with outside input is dangerous.
Answer frame:

Call: subprocess.run with the command as a list, not one string.

Capture: capture_output=True, text=True for strings, check=True so failures raise, a timeout.

Safety: avoid shell=True, especially with any value a user or file supplied.

Sample spoken answer:

“I use subprocess.run and pass the command as a list of arguments, like ["git", "rev-parse", "HEAD"]. With a list there's no shell in between, so a filename with spaces or a semicolon can't turn into a second command. That's the big risk with shell=True: if any part of the string comes from a user or a file, it's an injection hole. I add capture_output=True and text=True so I get stdout and stderr as strings, check=True so a non-zero exit raises CalledProcessError, which carries the stderr I log, and a timeout so a stuck tool can't hang my script. Before shelling out at all, I check whether Python can do it directly, like shutil for copying files.”

Code:
import subprocess

result = subprocess.run(
    ["git", "rev-parse", "--short", "HEAD"],
    capture_output=True, text=True, check=True, timeout=10,
)
commit = result.stdout.strip()
Red flag to avoid:

Building a command string with an f-string and shell=True, or using os.system and ignoring the exit code.

They may ask next:
  • How would you stream the output of a long-running command line by line?
  • What does shlex.quote do, and when would you still need it?
Say it in 60 seconds
Hard Technical round Mid-level Practice question

19. You added functools.lru_cache to speed up a function. What can go wrong with that?

What the interviewer is really testing:
Whether they know the real-world pitfalls of memoisation, not just that it makes repeat calls fast.
Answer frame:

Arguments: must be hashable; a list argument raises TypeError.

Shared results: the cached object is returned to every caller, so mutating it changes it for everyone.

Staleness and memory: no expiry, one cache per process, and on methods it keeps instances alive.

Sample spoken answer:

“I cached a function that loaded tenant settings from the database and it worked, but a few things bit us. The cache never expires, so when someone changed settings, the old ones stayed until restart; we added cache_clear() on the update path. The function returned a dict, and one caller modified it, which quietly changed the cached copy for every later caller, so now I return something read-only or copy it. Arguments must be hashable, so passing a list raises TypeError. Each worker process has its own cache, so two servers can disagree. And putting it on an instance method caches self as part of the key, which keeps those objects alive. cache_info() helped me check it was actually getting hits.”

Code:
from functools import lru_cache

@lru_cache(maxsize=1024)
def tenant_settings(tenant_id: str) -> dict:
    return db.fetch_settings(tenant_id)

tenant_settings.cache_info()    # hits, misses, currsize
tenant_settings.cache_clear()   # call after settings change
Red flag to avoid:

Treating caching as free speed with no thought about stale data, shared mutable results or unbounded memory.

They may ask next:
  • How would you add a time limit so cached entries expire?
  • When would you move this to a shared cache instead of an in-process one?
Say it in 60 seconds

Errors & Exceptions 1 question

Medium Technical round Mid-level Practice question

20. In a module you owned, how did you turn low-level errors like a KeyError into your own exceptions, and what does raise ... from add?

What the interviewer is really testing:
Whether they design errors so callers can handle them sensibly, and keep the original cause visible for whoever debugs it.
Answer frame:

Hierarchy: one base exception per module or package, specific subclasses under it.

Translate: turn low-level errors into your own at the boundary so callers don't depend on internals.

Chain: raise X from exc keeps the original traceback; from None hides it on purpose.

Sample spoken answer:

“In the inventory module I owned, I made one base class, InventoryError, and a few specific ones under it, like ItemNotFound and OutOfStock. Callers can catch the specific one they know how to handle, or the base class to catch anything from my module, without also catching unrelated bugs. At the edges I translate: if the database layer raises a KeyError, I raise ItemNotFound instead, so callers don't depend on how I store things. I use raise ItemNotFound(item_id) from exc, which sets the original as the cause, so the traceback shows both, with the line 'The above exception was the direct cause'. When the inner error is just noise, from None suppresses it.”

Code:
class InventoryError(Exception):
    """Base for everything this module raises."""

class ItemNotFound(InventoryError):
    pass

def load_item(item_id):
    try:
        return store[item_id]
    except KeyError as exc:
        raise ItemNotFound(item_id) from exc
Red flag to avoid:

Raising bare Exception everywhere, or catching and re-raising in a way that throws away the original traceback.

They may ask next:
  • What does the traceback look like if you raise a new exception inside except without from?
  • Should custom exceptions carry extra fields, and how do you add them?
Say it in 60 seconds

Real Work 4 questions

Medium Behavioral round Mid-level Practice question

21. Walk me through a Python feature you built end to end, from picking up the ticket to seeing it work in production.

What the interviewer is really testing:
Whether they genuinely own work from start to finish at this stage: clarifying the ask, testing, shipping safely and checking it afterwards.
Answer frame:

Clarify: what the ticket really needed and what you asked before coding.

Build and test: how you split the work and what you tested.

Ship and check: how it went out, what you watched, what you fixed after.

Sample spoken answer:

“At my last company I built a nightly job that synced stock levels from a supplier API into our database. The ticket said 'sync inventory', so first I asked what should happen to items the supplier dropped, and we agreed to mark them inactive, not delete them. I wrote the parsing as small pure functions with tests built from real sample responses, then the API client with timeouts and retries. I put it behind a flag and ran it against staging for two nights, comparing counts with the supplier's export. After launch I watched the logs for a week. One supplier sent empty strings for missing quantities, which I'd assumed were nulls, so I added that case and a test. It's run every night since.”

Red flag to avoid:

Describing only the coding, with no clarifying questions, no testing story and no check after release.

They may ask next:
  • What would you do differently if you built it again?
  • How did you know it was working, beyond no errors in the logs?
Say it in 60 seconds
Medium Behavioral round Mid-level Practice question

22. Which comment from a reviewer on one of your Python pull requests do you still think about when you write code?

What the interviewer is really testing:
Whether they take feedback well and can name a concrete habit that changed, which is how people grow in their second and third year.
Answer frame:

The feedback: what the reviewer pointed at, in plain terms.

Your reaction: whether you understood it, pushed back, or asked why.

The change: the habit you kept afterwards, with an example.

Sample spoken answer:

“Early in my second year a senior reviewer left one comment on a pull request: 'this function fetches, transforms and saves; I can't test the middle part.' She was right. My function called the API, reshaped the data and wrote to the database all in one go, so every test needed mocks for both ends. I split it so the transform was a plain function taking a dict and returning a list, with the I/O around it. The tests for the tricky part became three lines each, no mocks. She also pointed out I was changing a dict the caller had passed in, which had caused an odd bug before. Now I keep logic pure where I can and return new objects instead of mutating arguments.”

Red flag to avoid:

Saying reviews never taught them anything, or describing feedback only as style nitpicks they grudgingly accepted.

They may ask next:
  • Tell me about a time you disagreed with a review comment. What happened?
  • What do you look for now when you review someone else's Python code?
Say it in 60 seconds
Medium Behavioral round Mid-level Practice question

23. Tell me about a Python task you estimated badly. What did you miss, and what do you do differently now?

What the interviewer is really testing:
Whether they are honest about misses and have learned to find the unknowns in a task before committing to a number.
Answer frame:

The task: what it was and what you estimated.

What you missed: the specific unknowns, not just 'it was harder'.

What changed: how you estimate and communicate now.

Sample spoken answer:

“I estimated two days for a CSV export of user activity. It took over a week. I'd pictured one query and a csv.writer. What I missed: some accounts had hundreds of thousands of rows, so building the file in memory crashed the worker, and I had to stream it. Managers were only allowed to see their own team's rows, which meant a permission check I hadn't planned for. And the product manager expected it emailed, not downloaded, which nobody had written down. I told my lead on day two that it was slipping, which helped. Now for anything unfamiliar I spend a few hours spiking the riskiest part first, list what I don't know, and give a range instead of one number.”

Red flag to avoid:

Blaming others for the miss, or claiming they always estimate accurately.

They may ask next:
  • How do you tell your lead a task will slip, and when?
  • How do you estimate work in a part of the codebase you've never touched?
Say it in 60 seconds
Medium Situational round Mid-level Practice question

24. While fixing a bug, you find the real cause is in a shared utility that five other modules call. What do you do?

What the interviewer is really testing:
Whether they act carefully on code they don't fully own: checking who depends on the current behaviour before changing it.
Answer frame:

Scope: find every caller and check whether any of them depends on the buggy behaviour.

Prove it: write a failing test for the utility itself.

Change safely: fix with tests around callers, flag it in the pull request, tell the owners.

Sample spoken answer:

“I wouldn't just patch it and move on, because five callers means five places that might rely on the current behaviour, even if it's wrong. First I'd search for every call site and read how each uses the result. Then I'd write a failing test for the utility that shows the bug. If all callers want the fixed behaviour, I fix it, run the full suite, and call out the change clearly in the pull request so reviewers from those areas look at it. If one caller depends on the old behaviour, I'd talk to whoever owns it; sometimes the safer move is a new function or a parameter, and moving callers over one at a time. Either way I'd tell my lead it's bigger than the original ticket.”

Red flag to avoid:

Quietly changing shared code to fix their own ticket without checking the other callers.

They may ask next:
  • What if the caller that depends on the old behaviour belongs to another team that is busy?
  • How would you make sure this bug never comes back?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card