Production Incidents • Data and Performance • Safe Rollouts • Resilience • Code Review • 2026

Spring Boot Interview Questions for Experienced Candidates (5 Years)

Spring Boot interviews for experienced candidates with around five years rarely ask what an annotation does; they ask why you chose an approach and what it cost you: deadlocks and bulk loads, safe rollouts and health checks, caching and rate limits, poison messages, security holes you closed, and how you mentor and push back. It is written for Spring Boot developers with roughly five to seven years of experience, who own a service, make design calls inside it, get paged when it breaks and review other people's pull requests. Each answer below is a first-person story with a trade-off. Swap in your own project details before you say it out loud.

Search all questions by round, difficulty and level, or save the ones you want to practice.

Data and Transactions 3 questions

Hard Technical round Mid-level, Senior Practice question

1. Your logs showed occasional database deadlock errors between two endpoints that both update accounts. How did you trace the cause, and what did you change?

What the interviewer is really testing:
Whether you can read a deadlock report, fix the order locks are taken in and how long transactions run, and add a safe retry rather than just catching the error.
Answer frame:

Trace: the database's deadlock report names the two statements and the rows each one held.

Fix the order: lock rows in one fixed order, such as lowest id first, and keep transactions short.

Safety net: a small retry around the whole transaction, never inside it.

Sample spoken answer:

“The database's deadlock report showed the pattern. A transfer endpoint updated the sender's account and then the receiver's, while a refund went the other way round, so two opposite operations could each lock one row and wait for the other. The database killed one, and Spring surfaced it as a PessimisticLockingFailureException subclass. The fix was to lock both accounts in a fixed order, lowest id first, with one locking select before any change. I also moved a call to our audit service out of that transaction, because the longer locks are held, the more chances to collide. Deadlocks can still happen rarely, so I added a retry of up to three attempts around the service method, outside the transaction, so each attempt started clean. The cost was a little more locking up front, and the errors disappeared from the logs.”

Code:
@Lock(LockModeType.PESSIMISTIC_WRITE)
@Query("select a from Account a where a.id in :ids order by a.id")
List<Account> lockInIdOrder(@Param("ids") Collection<Long> ids);
Red flag to avoid:

Catching the deadlock error and logging it, or raising the lock timeout, without fixing the order the locks are taken in.

They may ask next:
  • Why must the retry sit outside the transaction rather than inside it?
  • When would a version check be a better fit here than a locking select?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

2. You needed to rename a column that a live Spring Boot service reads and writes, with no downtime during a rolling deploy. How did you do it?

What the interviewer is really testing:
Whether you know old and new versions run side by side during a deploy, and can split a schema change into safe, reversible steps.
Answer frame:

Constraint: during a rolling deploy, old and new code use the same schema at the same time.

Expand: add the new column, write to both, backfill in batches.

Contract: switch reads, then drop the old column in a later release.

Sample spoken answer:

“The easy way is one Flyway migration that renames the column, but during a rolling deploy the old pods are still running and would fail the moment the column disappears. So I spread it over three releases. The first added the new column, and the code wrote to both columns while still reading the old one. A backfill then copied existing rows in small batches so we never locked the table for long. The second release switched reads to the new column while still writing both, so a rollback stayed safe. The third stopped writing the old column, and a later migration dropped it. It took about two weeks instead of an afternoon, and the team grumbled, but there was no outage and every step could be rolled back on its own. I wrote the pattern into our team wiki so the next rename followed it.”

Red flag to avoid:

Shipping the rename and the code change in one release and hoping the deploy is quick.

They may ask next:
  • Why is a big backfill in one transaction a problem on a busy table?
  • How would you add a NOT NULL column to a large table in the same spirit?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

3. Your database was overloaded by reads, so the team added a read replica. How did you route read-only work to it in Spring Boot, and what problems came with it?

What the interviewer is really testing:
Whether you can wire the routing correctly and have thought about replication lag and reading your own writes.
Answer frame:

Routing: an AbstractRoutingDataSource that picks the replica when the current transaction is read-only.

Gotcha: wrap it in LazyConnectionDataSourceProxy so the real connection is fetched after the read-only flag is set.

Lag: a read right after a write may miss it, so those paths stay on the primary.

Sample spoken answer:

“I built a routing data source that checks whether the current transaction is marked read-only and picks the replica or the primary. At first everything still went to the primary. The connection was being taken as the transaction began, before the read-only flag was visible to the router. Wrapping the router in LazyConnectionDataSourceProxy fixed it, because the real connection is only fetched at the first statement. Then replication lag bit us: a user saved their profile, was redirected, and saw the old data because the page read from the replica. So any read that follows a user's own write stayed on the primary, and only heavy reporting and list endpoints moved to the replica. The trade-off is that developers must think about which transactions are safe to mark read-only, so I added that to our review checklist.”

Red flag to avoid:

Sending every read to the replica without thinking about lag, then treating stale screens as a front-end bug.

They may ask next:
  • How would you monitor replication lag and stop using a replica that falls too far behind?
  • What does a read-only transaction change even when there's no replica?
Say it in 60 seconds

Performance 3 questions

Hard Technical round Mid-level, Senior Practice question

4. A nightly import calls saveAll on a few hundred thousand new rows. You set a JDBC batch size, but the SQL log still shows one insert per row and it takes hours. Why?

What the interviewer is really testing:
Whether you know what quietly stops Hibernate from batching inserts, and can judge when to step down from JPA to plain JDBC.
Answer frame:

IDENTITY ids: Hibernate must run each insert at once to read back the generated id, so it cannot batch them.

Assigned ids: save treats an entity that already has an id as existing and merges it, which adds a select per row.

Fix: sequence ids with a pooled allocation, or plain JDBC batch inserts for a pure load job.

Sample spoken answer:

“Two things were fighting the batch setting. The entities used IDENTITY ids, and Hibernate has to run each insert straight away to read back the generated id, so it quietly sends them one at a time. And some rows came with ids from the source file, so save thought they already existed and merged them, which means a select before every insert. We were on PostgreSQL, so I switched the ids to a sequence with an allocation size of fifty: Hibernate hands out ids from memory and can batch, and I turned on ordered inserts too. For the rows with source ids I used a plain JdbcTemplate batch insert, since that job needed no entity behaviour at all. It went from hours to minutes. The trade-off was a schema change for the sequence and two ways of writing data in one service, which I wrote down in the module's readme.”

Code:
@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "price_row_seq")
@SequenceGenerator(name = "price_row_seq", sequenceName = "price_row_seq",
                   allocationSize = 50) // sequence must also INCREMENT BY 50
private Long id;

// application.yml: spring.jpa.properties.hibernate.jdbc.batch_size: 50
//                  spring.jpa.properties.hibernate.order_inserts: true
Red flag to avoid:

Raising the batch size again and again without checking whether Hibernate is able to batch at all.

They may ask next:
  • Why must the database sequence's increment match the allocation size?
  • What would you do on a database that has no sequences?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

5. A hot lookup was hammering your database. How did you decide between an in-process cache like Caffeine and a shared one like Redis, and what did each cost you?

What the interviewer is really testing:
Whether you weigh staleness across instances, memory, network hops and failure modes before choosing a cache, and whether you always bound it.
Answer frame:

In-process: fastest and nothing new to run, but every instance holds its own copy, and those copies can disagree.

Shared: one copy for all instances, but a network hop, serialisation and one more thing that can go down.

Either way: a size limit and an expiry chosen from how stale the data is allowed to be.

Sample spoken answer:

“The lookup was product categories, read on almost every request and changed a few times a day. With Spring's cache abstraction the code looks the same either way, so it was really an operations choice. I picked Caffeine: no network hop and nothing new to run. With six instances each holding a copy, the worst case was one instance showing an old category for a few minutes, which the business accepted. I set a maximum size and a ten-minute expiry, because Boot's fallback cache, the one you get with no provider on the classpath, is a plain map with no limit and no expiry. Later a pricing lookup needed every instance to agree right after an update, so there I used Redis and evicted on write. That cost us a network hop per read, some serialisation bugs, and a decision about Redis being down: we fell back to the database.”

Code:
spring:
  cache:
    type: caffeine
    cache-names: categories
    caffeine:
      spec: maximumSize=5000,expireAfterWrite=10m
Red flag to avoid:

Adding a cache with no size limit or expiry and calling the performance problem solved.

They may ask next:
  • How do you stop many requests reloading the same expired entry at once?
  • What would you check before putting user-specific data in a shared cache?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

6. An admin screen pages through millions of rows. Early pages are fast, but page five thousand takes seconds, and every page also runs a slow count. What did you change?

What the interviewer is really testing:
Whether you understand why offset paging slows down with depth, and what returning a Page costs compared with a Slice.
Answer frame:

Offset cost: the database still reads and throws away every skipped row.

Count cost: returning a Page runs a count query on every request; a Slice doesn't.

Fix: keyset paging on an indexed, unique sort order, giving up jumps to arbitrary pages.

Sample spoken answer:

“The repository returned a Page, so every request ran the data query plus a count over millions of rows, and the offset meant the database scanned and discarded everything before the requested page. First I switched to a Slice, which fetches one extra row to know whether a next page exists, and changed the screen to show next and previous instead of page numbers. Then I moved to keyset paging: the client sends the created time and id of the last row it saw, and the query asks for rows after that, ordered by the same columns, backed by a matching index. Each page now costs the same, first or ten-thousandth. We lost jump-to-page, which the product owner agreed nobody used. For full exports we streamed rows in chunks instead of paging at all.”

Code:
@Query("""
    select e from AuditEvent e
    where e.createdAt < :createdAt
       or (e.createdAt = :createdAt and e.id < :id)
    order by e.createdAt desc, e.id desc
    """)
List<AuditEvent> nextPage(@Param("createdAt") Instant createdAt,
                          @Param("id") Long id,
                          Pageable page); // PageRequest.of(0, 50), no count query
Red flag to avoid:

Asking for a bigger database instead of seeing that deep offsets scan every skipped row.

They may ask next:
  • Why must the sort order be unique for keyset paging to be correct?
  • How would you show users an approximate total cheaply?
Say it in 60 seconds

API Design 3 questions

Hard Technical round Mid-level, Senior Practice question

7. The mobile app retries on timeouts, and you start seeing duplicate orders from a single tap. How did you make the create endpoint safe to retry?

What the interviewer is really testing:
Whether you can design an idempotency key properly, including the race when two retries arrive at the same moment.
Answer frame:

Key: the client sends one unique key per logical action in a request header.

Store: save the key and the result under a unique constraint, in the same transaction as the order.

Replay: a repeated key returns the stored response instead of creating a new order.

Sample spoken answer:

“The order was actually created; the response just never reached the phone, so the app retried. I asked the app team to generate a unique idempotency key per checkout and send it in a header. On the server, I stored that key with the new order id in a table with a unique constraint, inside the same transaction that created the order. When a request arrived with a key we'd seen, we returned the original response instead of creating another order. The tricky case was two retries landing at once. The unique constraint made one of them fail, and we caught that and returned the first result. Keys were kept for a day, then cleaned up. The cost was an extra table and one more write per request, which was nothing next to cancelling duplicate orders by hand.”

Red flag to avoid:

Checking for an existing record and inserting afterwards with no unique constraint, which still races under concurrent retries.

They may ask next:
  • What should happen if the same key arrives with a different request body?
  • Why not just look for an identical order placed in the last minute?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

8. Another team's service consumes your REST API. How do you add, rename or remove JSON fields without breaking them, and what did you learn doing it?

What the interviewer is really testing:
Whether you know which changes are safe on each side of a JSON contract, and how Spring Boot's Jackson defaults play into that.
Answer frame:

Usually safe: adding optional fields, because Spring Boot's Jackson setup ignores unknown properties.

Renames and removals: add the new field, keep the old one, deprecate it, remove it once consumers have moved.

Proof: contract tests, or at least logs showing who still uses the old field.

Sample spoken answer:

“Adding a field is usually safe, because Spring Boot sets up Jackson not to fail on unknown properties, so consumers built on Spring just ignore new fields. I learned not to assume every consumer works that way: one team used a strict client in another language that rejected unknown fields, and our harmless addition broke them. Since then I find out who consumes an API before changing it. For a rename, I add the new field, keep sending the old one, mark it deprecated in the docs and tell consumers directly. We logged which clients still sent the old field in requests, so we knew when removing it was safe. Only for a truly breaking change did we add a second version of the endpoint, because running two versions doubles the testing. The payload carries some duplicate fields for a while, which is cheap next to a broken integration.”

Red flag to avoid:

Renaming a field in place and promising to tell the other team, with no overlap period.

They may ask next:
  • How would consumer-driven contract tests have caught that break earlier?
  • When is a new version in the URL better than evolving the same endpoint?
Say it in 60 seconds
Medium System design round Mid-level, Senior Practice question

9. A new downstream team needs to know when orders in your service ship. Did you call their API directly or publish an event, and what did that choice cost you?

What the interviewer is really testing:
Whether you choose between a direct call and messaging on coupling, failure and ownership, not on fashion.
Answer frame:

First question: does my flow need their answer to finish, or are they only being told something happened?

Choice: an event, because shipping should not depend on their uptime.

Costs accepted: data that lags a little, a schema to version, and harder debugging across the broker.

Sample spoken answer:

“My first question was whether shipping an order needed anything back from them. It didn't; they only had to be told. A direct call would have tied our shipping flow to their uptime, and every new team wanting the same news would mean another call from our code. So we published an 'order shipped' event and they subscribed. The costs were real. Their data runs a few seconds behind ours, which they agreed was fine. The event became a public contract, so we wrote a schema, only added optional fields, and planned a new version for anything breaking. And debugging got harder, so we put a correlation id in the message headers and traced it across both services. Where we needed an answer straight away, like checking stock before accepting an order, I kept a normal REST call with a timeout.”

Red flag to avoid:

Choosing messaging for everything because it sounds scalable, with no plan for duplicates, ordering or schema changes.

They may ask next:
  • How should the downstream team handle getting the same event twice?
  • What would make you switch that integration back to a direct call?
Say it in 60 seconds

Operations 4 questions

Hard Technical round Mid-level, Senior Practice question

10. A short database outage made Kubernetes restart every pod of your service at once, which made recovery slower. What was wrong with your health checks?

What the interviewer is really testing:
Whether you know the difference between liveness and readiness, and what each one should and shouldn't depend on.
Answer frame:

Cause: the liveness probe used the full health endpoint, which included the database check.

Fix: Actuator's separate liveness and readiness groups, with outside dependencies kept out of liveness.

Trade-off: decide per dependency whether losing it should take the pod out of traffic.

Sample spoken answer:

“Our liveness probe pointed at the main Actuator health endpoint, and that included the database health check. When the database blipped, every pod reported down, Kubernetes decided they were all broken and restarted them together, and they all came back hammering the database with connection attempts. The pods weren't broken, and a restart can't fix a database outage. I moved the probes to Actuator's liveness and readiness endpoints, which by default only reflect the app's own state. Liveness now answers one thing: is this process stuck. Readiness took a real debate. If every pod goes unready when the database is down, the load balancer has nowhere to send traffic and clients get a vaguer error. We left the database out of readiness too and let requests fail fast with a clear error, which gave better messages and a faster recovery.”

Red flag to avoid:

Putting every downstream check into liveness because health should cover everything.

They may ask next:
  • When would you add a dependency to the readiness group on purpose?
  • How would you stop every pod from reconnecting to the database at the same moment after an outage?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

11. Pods of your Spring Boot service kept getting OOMKilled, yet the heap graphs never went near the limit. What was going on, and how did you size memory after that?

What the interviewer is really testing:
Whether you know a JVM uses far more than its heap, and how the container limit, the heap size and the kernel's out-of-memory killer relate.
Answer frame:

Cause: the heap was set almost as large as the container limit, leaving no room for anything else.

Outside the heap: metaspace, thread stacks, the code cache, direct buffers and the garbage collector all count against the limit.

Fix: size the heap as a share of the limit, measure the rest with Native Memory Tracking, and leave headroom.

Sample spoken answer:

“The container limit was two gigabytes, and someone had set the maximum heap to nearly the same. The heap looked healthy, but the kernel kills the container when the whole process passes the limit, and a JVM uses plenty outside the heap: metaspace for all of Spring's classes, a stack for each of our two hundred Tomcat threads, the code cache, and direct buffers from the HTTP client. None of that shows on a heap graph. I turned on Native Memory Tracking in a test pod to see the real breakdown. Then I replaced the fixed heap size with MaxRAMPercentage, so the heap follows the container limit and leaves room for the rest; we settled on about three quarters. The kills stopped, and when we later raised the thread count, we knew to check memory too. The trade-off was a smaller heap than people wanted, so we watched garbage collection time after the change.”

Code:
# heap follows the container limit, leaving room outside the heap
JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75.0 -XX:NativeMemoryTracking=summary"  # tracking in a test pod only

# inside the pod: where the memory really goes
jcmd 1 VM.native_memory summary
Red flag to avoid:

Raising the container limit and the heap together, which only moves the kill to later.

They may ask next:
  • What would you look at if memory kept growing slowly outside the heap?
  • Why can adding threads push a pod over its memory limit even when the heap is fine?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

12. You had to change the pricing logic in a service you own, where a bug would charge customers the wrong amount. How did you roll it out safely?

What the interviewer is really testing:
Whether you separate deploying code from releasing behaviour, and can prove a risky change is right before customers depend on it.
Answer frame:

Deploy is not release: ship the new code behind a feature flag that starts switched off.

Prove it: run old and new logic side by side and log every difference before switching.

Release slowly: internal accounts first, then a small share of customers, with a switch-off that needs no deploy.

Sample spoken answer:

“I shipped the new calculation behind a feature flag, off by default, so deploying it changed nothing. Then for two weeks the service ran both versions on every quote: customers got the old result, and we logged every case where the new one differed. That found two rounding bugs and a discount rule the spec had missed, all before a customer saw anything. Once the only differences were the intended ones, we turned the flag on for internal accounts, then a small share of customers, then everyone, watching refunds and support tickets at each step. Switching it off needed a config change, not a deploy. The cost was running both calculations for a while and some messy code, so removing the old path and the flag became a ticket with a date, not a someday.”

Red flag to avoid:

Relying on staging tests alone and switching every customer over in one deploy.

They may ask next:
  • When is running old and new logic side by side unsafe?
  • How do you stop old feature flags piling up in the codebase?
Say it in 60 seconds
Medium Behavioral round Mid-level, Senior Practice question

13. Walk me through the worst production incident in a Spring Boot service you owned, from the first alert to what you changed afterwards.

What the interviewer is really testing:
Whether you restore service first and find the root cause second, keep people informed while it happens, and turn the incident into lasting fixes without blame.
Answer frame:

Restore first: roll back, switch off or work around before digging for the root cause.

Communicate: one channel, regular updates, clear owners.

Afterwards: a blameless review whose actions have owners and dates.

Sample spoken answer:

“Early one morning our alerts showed every call to the payment provider failing. The first ten minutes went on confirming it was us, not them: the logs showed a TLS handshake error, because the client certificate in our service's keystore had expired overnight. Rolling back wouldn't help, so the fix was a new certificate. While a colleague got it issued, I opened an incident channel, posted updates every fifteen minutes, and had support tell customers payments were delayed, not lost, since we queued the orders. We were back in about an hour. The blameless review found the real cause: the expiry date lived in one person's calendar, and that person had left. The actions were an alert well before any certificate expires, moving certificates out of the image into our secrets store so they could be rotated without a rebuild, and a runbook. Each had an owner and a date.”

Red flag to avoid:

Hunting for the root cause while customers stay broken, or naming the person who forgot instead of fixing the process.

They may ask next:
  • What did you tell customers and stakeholders while it was still broken?
  • What would you do differently in the first ten minutes next time?
Say it in 60 seconds

Reliability 3 questions

Hard Technical round Mid-level, Senior Practice question

14. You put a circuit breaker around a flaky downstream service. How did you choose the thresholds and the fallback, and what did getting them wrong look like?

What the interviewer is really testing:
Whether you set resilience settings from real traffic and failure data, and whether your fallback is honest about what the user gets.
Answer frame:

Timeouts first: a breaker without timeouts only trips after threads are already stuck.

Thresholds: failure rate, slow-call rate and a minimum number of calls, sized from real traffic.

Fallback: something true, like cached data or a clear degraded response, never a fake success.

Sample spoken answer:

“The downstream was a recommendations service that sometimes went slow for minutes. I used Resilience4j through its Spring Boot starter. First I set proper connect and read timeouts, because a breaker only counts failures it can see. Our first thresholds came from an example and tripped on a handful of errors during quiet night hours, because the minimum number of calls was far too low for that traffic. I raised the minimum, added a slow-call threshold since slowness was the real problem, and kept the breaker open long enough for the service to recover. For the fallback we showed popular items from a cached list instead of personal picks, which was honest and still useful. The trade-off is that while the breaker is open, some calls that would have worked get the fallback, so we put the breaker state on a dashboard and alerted on it.”

Code:
resilience4j:
  circuitbreaker:
    instances:
      recommendations:
        slidingWindowSize: 50
        minimumNumberOfCalls: 20
        failureRateThreshold: 50
        slowCallDurationThreshold: 2s
        slowCallRateThreshold: 50
        waitDurationInOpenState: 30s
Red flag to avoid:

A fallback that returns an empty success, so failures look like no data and nobody notices.

They may ask next:
  • How does a bulkhead differ from a circuit breaker, and when would you want both?
  • Should retries sit inside or outside the circuit breaker, and why?
Say it in 60 seconds
Hard Technical round Mid-level, Senior Practice question

15. One malformed message landed on a Kafka topic your Spring Boot service consumes. The consumer got stuck on it and lag kept growing. What did you change?

What the interviewer is really testing:
Whether you can stop one bad record from blocking a partition, and you know which errors are worth retrying and which never will succeed.
Answer frame:

Why it stuck: the record failed every time, and a partition is read in order, so nothing behind it moved.

Fix: a deserializer that turns bad payloads into handled errors, a few retries with backoff, then a dead-letter topic.

Classify: errors that can never succeed, like bad data, skip the retries.

Trade-off: dead-lettered records need an alert, an owner and a way to replay them.

Sample spoken answer:

“Another team produced a payload with a field of the wrong type, so it failed to deserialise every time. A partition is read in order, so the consumer kept failing on that record, everything behind it waited, and lag climbed. I made three changes. I wrapped our deserializer in ErrorHandlingDeserializer, so a bad payload reached our error handler instead of failing inside the poll. I configured a DefaultErrorHandler with a short backoff and a DeadLetterPublishingRecoverer, so after a few tries the record went to a dead-letter topic and the partition moved on. Deserialisation errors already skip retries in that handler, and I added our validation exception to that list, because retrying bad data never helps. The risk is that dead-lettered records get forgotten, so we alerted on that topic and wrote a small tool to replay them once the producer fixed their bug.”

Code:
@Bean
DefaultErrorHandler errorHandler(KafkaTemplate<Object, Object> template) {
    DefaultErrorHandler handler = new DefaultErrorHandler(
        new DeadLetterPublishingRecoverer(template),
        new FixedBackOff(1000L, 3)); // 3 retries, 1 second apart
    handler.addNotRetryableExceptions(ValidationException.class);
    return handler;
}
Red flag to avoid:

Catching every exception in the listener and just logging it, so bad data disappears without a trace.

They may ask next:
  • How would you replay records from the dead-letter topic safely once the bug is fixed?
  • What changes if order matters for one customer's messages and one of them goes to the dead-letter topic?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

16. One client's batch script started hammering your API and slowed it down for everyone else. How did you protect the service, and where did you put the limit?

What the interviewer is really testing:
Whether you can decide where a limit belongs, key it per client, and tell clients clearly when they are being limited.
Answer frame:

Right now: find the client in the logs and throttle it at the gateway, before it reaches the app.

Design: limits keyed on the client's API key, not the IP, answered with a 429 and a Retry-After header.

Where: the gateway for simple limits; the app only for limits that need business data.

Sample spoken answer:

“The logs showed one API key calling a search endpoint hundreds of times a second, which used up our database connections. The quick fix was a limit for that key at our API gateway, so the traffic never reached the app. Then I did it properly. Limits were keyed on the client's API key, not the IP address, because several customers sat behind one shared office network. Over the limit, clients got a 429 with a Retry-After header, so a well-written script could slow down instead of failing. Most limits stayed at the gateway, where blocking is cheap and every instance sees the same counts. One limit, on expensive exports per account, needed business data, so it lived in the app with a counter in Redis that all instances shared. I also called the customer. They needed a bulk endpoint, and building one was the real fix.”

Red flag to avoid:

Blocking the client's IP address and moving on, with no per-client limit and no clear 429 response.

They may ask next:
  • Why is a limit counted separately on each instance misleading behind a load balancer?
  • How would you give one customer a higher limit without a code change?
Say it in 60 seconds

Security 2 questions

Hard Technical round Mid-level, Senior Practice question

17. A security review found that any logged-in user could read another user's invoice by changing the id in the URL. How did you fix it, and where did you put the check?

What the interviewer is really testing:
Whether you know authentication is not authorisation, and can place ownership checks so the next endpoint can't forget them.
Answer frame:

Cause: the endpoint checked the user was logged in, never that they owned the record.

Fix: enforce ownership in the query or with method security such as @PreAuthorize, not in scattered if statements.

Stop it recurring: tests that call each endpoint as the wrong user, plus a review checklist item.

Sample spoken answer:

“The filter chain only checked that the token was valid, and the controller loaded the invoice by id. Authentication was fine; authorisation was missing. The first fix was to scope the repository call to the current user, finding the invoice by id and owner, and returning 404 when nothing matched, so we didn't even confirm the id existed. For support staff who could view any invoice, I used @PreAuthorize with a small bean that checks the role or ownership, so the rule lived in one place. Then I searched for the same pattern and found two more endpoints with it. To stop it coming back, I added integration tests that call every id-based endpoint as a second user and expect a 404 or 403. Repository methods got a little more specific, which I think is a fair price.”

Code:
@PreAuthorize("hasRole('SUPPORT') or @invoiceAccess.isOwner(#id, authentication)")
@GetMapping("/invoices/{id}")
public InvoiceResponse get(@PathVariable Long id) {
    return invoiceService.find(id);
}
Red flag to avoid:

Treating hard-to-guess ids, such as switching to UUIDs, as the fix.

They may ask next:
  • Why return 404 rather than 403 when a user asks for someone else's record?
  • How would you enforce the same rule on list and search endpoints?
Say it in 60 seconds
Medium Technical round Mid-level, Senior Practice question

18. You found customer emails, phone numbers and even auth tokens in your service's production logs. How did they get there, and how did you fix it for good?

What the interviewer is really testing:
Whether you treat logs as data that can leak, and fix what writes the data instead of scrubbing it afterwards.
Answer frame:

Sources: full request logging, toString on DTOs, and error messages that echo the input.

Fix at the source: no bodies or auth headers in logs, sensitive fields left out of toString, messages that name the field, not the value.

Clean up and guard: revoke exposed tokens, shorten retention, and add a test so it cannot creep back.

Sample spoken answer:

“A support engineer spotted a bearer token in our log search. I traced three sources. A debug filter that logged full request bodies and headers had been left on in production. Several DTOs had generated toString methods, and one error log printed the whole object. And a validation message echoed the rejected value back, phone numbers included. First I had the exposed tokens revoked, since anyone with log access could have used them. Then I removed body and header logging, left sensitive fields out of toString, and changed messages to name the field, not the value. We shortened log retention too. To keep it fixed, I added a test that sends known fake values through the main endpoints and fails if any appear in the captured logs. Debugging got a little harder without bodies, so we log record ids and people look up the record with proper access.”

Red flag to avoid:

Deleting the old log files and stopping there, without revoking the tokens or fixing what wrote them.

They may ask next:
  • Would masking values in the log pattern be enough on its own? What would it miss?
  • Who should be able to search production logs, and why does that matter here?
Say it in 60 seconds

Testing 1 question

Medium Technical round Mid-level, Senior Practice question

19. Your Spring Boot integration tests grew past twenty minutes, and the build log shows the application context starting again and again. What caused that, and how did you fix it?

What the interviewer is really testing:
Whether you know how Spring caches test contexts, and which habits quietly create a new context for almost every test class.
Answer frame:

How caching works: Spring reuses a test context only when the test configuration is exactly the same.

What broke it: each class mocking a different set of beans, different properties, and @DirtiesContext used just in case.

Fix: a few shared test setups, and @DirtiesContext only where state truly cannot be reset.

Sample spoken answer:

“Spring caches the application context between test classes, but only when their configuration is exactly the same, and ours almost never was. Each class mocked its own mix of beans, set a few different properties, and several used @DirtiesContext just in case, which throws the context away. Counting startup banners in the build log, I found over forty context starts, each paying for startup and a fresh database container. I created three shared base classes, one per kind of test, each with the same mocks and properties, and moved per-test differences into setup code. I removed @DirtiesContext wherever cleaning the tables between tests did the job. We got down to a handful of contexts and the suite ran in about six minutes. The trade-off is that base classes mock a few beans some tests don't need, which looked odd at first, so I wrote a short note explaining why.”

Red flag to avoid:

Buying bigger build machines or running tests in parallel without asking why each class starts its own context.

They may ask next:
  • How would you stop a test that changes a singleton bean's state from affecting the next test class?
  • When is a sliced test a better answer than a faster full-context test?
Say it in 60 seconds

Observability 1 question

Medium Technical round Mid-level, Senior Practice question

20. When you took over on-call for a Spring Boot service, alerts fired all night and the team ignored them. How did you decide what should wake someone up?

What the interviewer is really testing:
Whether you alert on what users feel rather than on every internal signal, and can cut noise without going blind.
Answer frame:

Symptoms, not causes: page on error rate and latency of the endpoints users depend on.

Over a window: judge a few minutes of traffic, so one bad request never pages anyone.

Everything else: dashboards or daytime tickets, and every page links to a short runbook.

Sample spoken answer:

“We had alerts on CPU, heap, every single 500 and a dozen log patterns. The team had muted the channel, and a real outage went unnoticed for twenty minutes. I rebuilt the paging alerts around what users feel: error rate and latency for our main endpoints, from the http.server.requests metrics Boot already records through Micrometer, judged over a few minutes so one bad request doesn't wake anyone. Consumer lag got a page too, because a stuck consumer is invisible to users until it's very late. CPU and heap moved to dashboards and daytime warnings. Every page linked to a runbook saying what to check first. Pages went from several a night to a couple a week, and the next real incident paged within minutes. The trade-off is that slow-burn problems like a filling disk no longer page, so we kept a few long-lead alerts for those.”

Red flag to avoid:

Adding more alerts after every incident until nobody reads any of them.

They may ask next:
  • How would you choose the error-rate threshold for a page?
  • What do you do with an alert that fired but needed no action?
Say it in 60 seconds

Leadership 3 questions

Medium Behavioral round Mid-level, Senior Practice question

21. What do you look for when you review a junior developer's pull request on a Spring Boot service, and how do you give feedback so they actually improve?

What the interviewer is really testing:
Whether your reviews catch the Spring mistakes that hurt later, and whether you teach instead of rewriting.
Answer frame:

What I check: layering, where the transaction starts and ends, error handling, config and secrets, and tests of real behaviour.

How: separate must-fix from nice-to-have, explain why, and point to a good example.

Growth: pair on a pattern once, then let them apply it.

Sample spoken answer:

“I check a few things that are expensive to fix later. Is business logic creeping into the controller? Where does the transaction start and end, and is anything slow, like an HTTP call, sitting inside it? Are errors going through our shared handler, or being swallowed? Is anything hard-coded that should be config, or worse, a secret in the properties file? And do the tests exercise real behaviour, not mocks returning mocks? For feedback, I label comments as blocking or optional so a junior knows what matters, and I explain why instead of pasting the fix. One developer put @Transactional on every method; instead of twenty comments, I spent half an hour walking through one of their classes with them, and their next pull requests changed. Teaching is slower than fixing it myself, but it scales and I stop being the bottleneck.”

Red flag to avoid:

Rewriting the code yourself in review comments, or approving everything to keep the peace.

They may ask next:
  • How do you handle a review where you and the author disagree on the design?
  • How do you keep reviews quick when the team sends many pull requests a day?
Say it in 60 seconds
Medium Situational round Mid-level, Senior Practice question

22. A bug left bad data in production, and your manager asks you to run an UPDATE straight on the live database right now to fix it for customers. What do you do?

What the interviewer is really testing:
Whether you can push back on an unsafe shortcut without blocking the fix, and turn the urgency into a safe, quick process.
Answer frame:

Say yes to the goal: customers need the fix today, so the answer is how, not whether.

Make it safe: count the rows first, copy them aside, run in a transaction, and have a second person read the script.

Stop the source: fix the code that wrote the bad data, or the same UPDATE will be needed again.

Sample spoken answer:

“I'd agree the data had to be fixed today, but not by typing an UPDATE into a production console. When this happened to me, a bad release had set the wrong status on a few thousand orders. I wrote the fix as a script: a select that counted the affected rows, a copy of those rows into a backup table, then the update inside a transaction that checked the row count matched before committing. A teammate reviewed it in ten minutes, and it ran within the hour, so the manager got speed and we kept the safety. Then we fixed the bug that wrote the bad data, because otherwise we'd be back the next week. Afterwards I proposed that any production data change goes through a reviewed script, which the team adopted. It adds a few minutes, and one mistyped WHERE clause costs far more.”

Red flag to avoid:

Running the update alone with no count, backup or review because the manager asked, or refusing with no faster safe option.

They may ask next:
  • What if the fix really could not wait for a reviewer?
  • When would you build an admin endpoint instead of running scripts?
Say it in 60 seconds
Hard Behavioral round Mid-level, Senior Practice question

23. Tell me about a design decision in a Spring Boot service that you made and later reversed. Why did you make it, what did it cost, and what changed your mind?

What the interviewer is really testing:
Whether you can own a call that didn't hold up, explain your reasoning without excuses, and show what you weigh differently now.
Answer frame:

The decision: what you chose and why it looked right with what you knew then.

The cost: the concrete signal that it wasn't working.

The reversal: how you undid it safely and what you'd check earlier next time.

Sample spoken answer:

“In a service I owned, I decoupled modules with Spring application events. Creating an order published an event, and loyalty points, emails and analytics each listened for it. It looked clean, because the order code knew about none of them. A year later it cost us. Nobody could see the whole flow by reading the code, and a new listener threw an exception. Plain @EventListener methods run in the publisher's thread and transaction, so a failing loyalty update rolled back the order itself. I reversed most of it. Steps that must happen with the order became explicit calls in the order service, so the flow reads top to bottom. Things that can happen afterwards, like emails, moved to @TransactionalEventListener, which by default runs after the commit. What I weigh now is that indirection has a reading cost, and I only pay it when there really are many independent listeners.”

Red flag to avoid:

Picking a story where the decision was really someone else's, or one with no cost and nothing learned.

They may ask next:
  • What catches people out when an after-commit listener needs to write to the database?
  • How did you tell the team you had got that call wrong?
Say it in 60 seconds
Were you asked something else? Share it A person checks every question before it goes on the site. No name is shown.
For the call itself

You practiced these. On the real call, ClapAssist helps with the rest.

ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.

Download with 10 free minutes
Mac and Windows · Stays out of screen share · No card