Spring Boot interviews for experienced candidates with around five years rarely ask what an annotation does; they ask why you chose an approach and what it cost you: deadlocks and bulk loads, safe rollouts and health checks, caching and rate limits, poison messages, security holes you closed, and how you mentor and push back. It is written for Spring Boot developers with roughly five to seven years of experience, who own a service, make design calls inside it, get paged when it breaks and review other people's pull requests. Each answer below is a first-person story with a trade-off. Swap in your own project details before you say it out loud.
Search all questions by round, difficulty and level, or save the ones you want to practice.
Trace: the database's deadlock report names the two statements and the rows each one held.
Fix the order: lock rows in one fixed order, such as lowest id first, and keep transactions short.
Safety net: a small retry around the whole transaction, never inside it.
“The database's deadlock report showed the pattern. A transfer endpoint updated the sender's account and then the receiver's, while a refund went the other way round, so two opposite operations could each lock one row and wait for the other. The database killed one, and Spring surfaced it as a PessimisticLockingFailureException subclass. The fix was to lock both accounts in a fixed order, lowest id first, with one locking select before any change. I also moved a call to our audit service out of that transaction, because the longer locks are held, the more chances to collide. Deadlocks can still happen rarely, so I added a retry of up to three attempts around the service method, outside the transaction, so each attempt started clean. The cost was a little more locking up front, and the errors disappeared from the logs.”
@Lock(LockModeType.PESSIMISTIC_WRITE)
@Query("select a from Account a where a.id in :ids order by a.id")
List<Account> lockInIdOrder(@Param("ids") Collection<Long> ids);
Catching the deadlock error and logging it, or raising the lock timeout, without fixing the order the locks are taken in.
Constraint: during a rolling deploy, old and new code use the same schema at the same time.
Expand: add the new column, write to both, backfill in batches.
Contract: switch reads, then drop the old column in a later release.
“The easy way is one Flyway migration that renames the column, but during a rolling deploy the old pods are still running and would fail the moment the column disappears. So I spread it over three releases. The first added the new column, and the code wrote to both columns while still reading the old one. A backfill then copied existing rows in small batches so we never locked the table for long. The second release switched reads to the new column while still writing both, so a rollback stayed safe. The third stopped writing the old column, and a later migration dropped it. It took about two weeks instead of an afternoon, and the team grumbled, but there was no outage and every step could be rolled back on its own. I wrote the pattern into our team wiki so the next rename followed it.”
Shipping the rename and the code change in one release and hoping the deploy is quick.
Routing: an AbstractRoutingDataSource that picks the replica when the current transaction is read-only.
Gotcha: wrap it in LazyConnectionDataSourceProxy so the real connection is fetched after the read-only flag is set.
Lag: a read right after a write may miss it, so those paths stay on the primary.
“I built a routing data source that checks whether the current transaction is marked read-only and picks the replica or the primary. At first everything still went to the primary. The connection was being taken as the transaction began, before the read-only flag was visible to the router. Wrapping the router in LazyConnectionDataSourceProxy fixed it, because the real connection is only fetched at the first statement. Then replication lag bit us: a user saved their profile, was redirected, and saw the old data because the page read from the replica. So any read that follows a user's own write stayed on the primary, and only heavy reporting and list endpoints moved to the replica. The trade-off is that developers must think about which transactions are safe to mark read-only, so I added that to our review checklist.”
Sending every read to the replica without thinking about lag, then treating stale screens as a front-end bug.
IDENTITY ids: Hibernate must run each insert at once to read back the generated id, so it cannot batch them.
Assigned ids: save treats an entity that already has an id as existing and merges it, which adds a select per row.
Fix: sequence ids with a pooled allocation, or plain JDBC batch inserts for a pure load job.
“Two things were fighting the batch setting. The entities used IDENTITY ids, and Hibernate has to run each insert straight away to read back the generated id, so it quietly sends them one at a time. And some rows came with ids from the source file, so save thought they already existed and merged them, which means a select before every insert. We were on PostgreSQL, so I switched the ids to a sequence with an allocation size of fifty: Hibernate hands out ids from memory and can batch, and I turned on ordered inserts too. For the rows with source ids I used a plain JdbcTemplate batch insert, since that job needed no entity behaviour at all. It went from hours to minutes. The trade-off was a schema change for the sequence and two ways of writing data in one service, which I wrote down in the module's readme.”
@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "price_row_seq")
@SequenceGenerator(name = "price_row_seq", sequenceName = "price_row_seq",
allocationSize = 50) // sequence must also INCREMENT BY 50
private Long id;
// application.yml: spring.jpa.properties.hibernate.jdbc.batch_size: 50
// spring.jpa.properties.hibernate.order_inserts: true
Raising the batch size again and again without checking whether Hibernate is able to batch at all.
In-process: fastest and nothing new to run, but every instance holds its own copy, and those copies can disagree.
Shared: one copy for all instances, but a network hop, serialisation and one more thing that can go down.
Either way: a size limit and an expiry chosen from how stale the data is allowed to be.
“The lookup was product categories, read on almost every request and changed a few times a day. With Spring's cache abstraction the code looks the same either way, so it was really an operations choice. I picked Caffeine: no network hop and nothing new to run. With six instances each holding a copy, the worst case was one instance showing an old category for a few minutes, which the business accepted. I set a maximum size and a ten-minute expiry, because Boot's fallback cache, the one you get with no provider on the classpath, is a plain map with no limit and no expiry. Later a pricing lookup needed every instance to agree right after an update, so there I used Redis and evicted on write. That cost us a network hop per read, some serialisation bugs, and a decision about Redis being down: we fell back to the database.”
spring:
cache:
type: caffeine
cache-names: categories
caffeine:
spec: maximumSize=5000,expireAfterWrite=10m
Adding a cache with no size limit or expiry and calling the performance problem solved.
Offset cost: the database still reads and throws away every skipped row.
Count cost: returning a Page runs a count query on every request; a Slice doesn't.
Fix: keyset paging on an indexed, unique sort order, giving up jumps to arbitrary pages.
“The repository returned a Page, so every request ran the data query plus a count over millions of rows, and the offset meant the database scanned and discarded everything before the requested page. First I switched to a Slice, which fetches one extra row to know whether a next page exists, and changed the screen to show next and previous instead of page numbers. Then I moved to keyset paging: the client sends the created time and id of the last row it saw, and the query asks for rows after that, ordered by the same columns, backed by a matching index. Each page now costs the same, first or ten-thousandth. We lost jump-to-page, which the product owner agreed nobody used. For full exports we streamed rows in chunks instead of paging at all.”
@Query("""
select e from AuditEvent e
where e.createdAt < :createdAt
or (e.createdAt = :createdAt and e.id < :id)
order by e.createdAt desc, e.id desc
""")
List<AuditEvent> nextPage(@Param("createdAt") Instant createdAt,
@Param("id") Long id,
Pageable page); // PageRequest.of(0, 50), no count query
Asking for a bigger database instead of seeing that deep offsets scan every skipped row.
Key: the client sends one unique key per logical action in a request header.
Store: save the key and the result under a unique constraint, in the same transaction as the order.
Replay: a repeated key returns the stored response instead of creating a new order.
“The order was actually created; the response just never reached the phone, so the app retried. I asked the app team to generate a unique idempotency key per checkout and send it in a header. On the server, I stored that key with the new order id in a table with a unique constraint, inside the same transaction that created the order. When a request arrived with a key we'd seen, we returned the original response instead of creating another order. The tricky case was two retries landing at once. The unique constraint made one of them fail, and we caught that and returned the first result. Keys were kept for a day, then cleaned up. The cost was an extra table and one more write per request, which was nothing next to cancelling duplicate orders by hand.”
Checking for an existing record and inserting afterwards with no unique constraint, which still races under concurrent retries.
Usually safe: adding optional fields, because Spring Boot's Jackson setup ignores unknown properties.
Renames and removals: add the new field, keep the old one, deprecate it, remove it once consumers have moved.
Proof: contract tests, or at least logs showing who still uses the old field.
“Adding a field is usually safe, because Spring Boot sets up Jackson not to fail on unknown properties, so consumers built on Spring just ignore new fields. I learned not to assume every consumer works that way: one team used a strict client in another language that rejected unknown fields, and our harmless addition broke them. Since then I find out who consumes an API before changing it. For a rename, I add the new field, keep sending the old one, mark it deprecated in the docs and tell consumers directly. We logged which clients still sent the old field in requests, so we knew when removing it was safe. Only for a truly breaking change did we add a second version of the endpoint, because running two versions doubles the testing. The payload carries some duplicate fields for a while, which is cheap next to a broken integration.”
Renaming a field in place and promising to tell the other team, with no overlap period.
First question: does my flow need their answer to finish, or are they only being told something happened?
Choice: an event, because shipping should not depend on their uptime.
Costs accepted: data that lags a little, a schema to version, and harder debugging across the broker.
“My first question was whether shipping an order needed anything back from them. It didn't; they only had to be told. A direct call would have tied our shipping flow to their uptime, and every new team wanting the same news would mean another call from our code. So we published an 'order shipped' event and they subscribed. The costs were real. Their data runs a few seconds behind ours, which they agreed was fine. The event became a public contract, so we wrote a schema, only added optional fields, and planned a new version for anything breaking. And debugging got harder, so we put a correlation id in the message headers and traced it across both services. Where we needed an answer straight away, like checking stock before accepting an order, I kept a normal REST call with a timeout.”
Choosing messaging for everything because it sounds scalable, with no plan for duplicates, ordering or schema changes.
Cause: the liveness probe used the full health endpoint, which included the database check.
Fix: Actuator's separate liveness and readiness groups, with outside dependencies kept out of liveness.
Trade-off: decide per dependency whether losing it should take the pod out of traffic.
“Our liveness probe pointed at the main Actuator health endpoint, and that included the database health check. When the database blipped, every pod reported down, Kubernetes decided they were all broken and restarted them together, and they all came back hammering the database with connection attempts. The pods weren't broken, and a restart can't fix a database outage. I moved the probes to Actuator's liveness and readiness endpoints, which by default only reflect the app's own state. Liveness now answers one thing: is this process stuck. Readiness took a real debate. If every pod goes unready when the database is down, the load balancer has nowhere to send traffic and clients get a vaguer error. We left the database out of readiness too and let requests fail fast with a clear error, which gave better messages and a faster recovery.”
Putting every downstream check into liveness because health should cover everything.
Cause: the heap was set almost as large as the container limit, leaving no room for anything else.
Outside the heap: metaspace, thread stacks, the code cache, direct buffers and the garbage collector all count against the limit.
Fix: size the heap as a share of the limit, measure the rest with Native Memory Tracking, and leave headroom.
“The container limit was two gigabytes, and someone had set the maximum heap to nearly the same. The heap looked healthy, but the kernel kills the container when the whole process passes the limit, and a JVM uses plenty outside the heap: metaspace for all of Spring's classes, a stack for each of our two hundred Tomcat threads, the code cache, and direct buffers from the HTTP client. None of that shows on a heap graph. I turned on Native Memory Tracking in a test pod to see the real breakdown. Then I replaced the fixed heap size with MaxRAMPercentage, so the heap follows the container limit and leaves room for the rest; we settled on about three quarters. The kills stopped, and when we later raised the thread count, we knew to check memory too. The trade-off was a smaller heap than people wanted, so we watched garbage collection time after the change.”
# heap follows the container limit, leaving room outside the heap
JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75.0 -XX:NativeMemoryTracking=summary" # tracking in a test pod only
# inside the pod: where the memory really goes
jcmd 1 VM.native_memory summary
Raising the container limit and the heap together, which only moves the kill to later.
Deploy is not release: ship the new code behind a feature flag that starts switched off.
Prove it: run old and new logic side by side and log every difference before switching.
Release slowly: internal accounts first, then a small share of customers, with a switch-off that needs no deploy.
“I shipped the new calculation behind a feature flag, off by default, so deploying it changed nothing. Then for two weeks the service ran both versions on every quote: customers got the old result, and we logged every case where the new one differed. That found two rounding bugs and a discount rule the spec had missed, all before a customer saw anything. Once the only differences were the intended ones, we turned the flag on for internal accounts, then a small share of customers, then everyone, watching refunds and support tickets at each step. Switching it off needed a config change, not a deploy. The cost was running both calculations for a while and some messy code, so removing the old path and the flag became a ticket with a date, not a someday.”
Relying on staging tests alone and switching every customer over in one deploy.
Restore first: roll back, switch off or work around before digging for the root cause.
Communicate: one channel, regular updates, clear owners.
Afterwards: a blameless review whose actions have owners and dates.
“Early one morning our alerts showed every call to the payment provider failing. The first ten minutes went on confirming it was us, not them: the logs showed a TLS handshake error, because the client certificate in our service's keystore had expired overnight. Rolling back wouldn't help, so the fix was a new certificate. While a colleague got it issued, I opened an incident channel, posted updates every fifteen minutes, and had support tell customers payments were delayed, not lost, since we queued the orders. We were back in about an hour. The blameless review found the real cause: the expiry date lived in one person's calendar, and that person had left. The actions were an alert well before any certificate expires, moving certificates out of the image into our secrets store so they could be rotated without a rebuild, and a runbook. Each had an owner and a date.”
Hunting for the root cause while customers stay broken, or naming the person who forgot instead of fixing the process.
Timeouts first: a breaker without timeouts only trips after threads are already stuck.
Thresholds: failure rate, slow-call rate and a minimum number of calls, sized from real traffic.
Fallback: something true, like cached data or a clear degraded response, never a fake success.
“The downstream was a recommendations service that sometimes went slow for minutes. I used Resilience4j through its Spring Boot starter. First I set proper connect and read timeouts, because a breaker only counts failures it can see. Our first thresholds came from an example and tripped on a handful of errors during quiet night hours, because the minimum number of calls was far too low for that traffic. I raised the minimum, added a slow-call threshold since slowness was the real problem, and kept the breaker open long enough for the service to recover. For the fallback we showed popular items from a cached list instead of personal picks, which was honest and still useful. The trade-off is that while the breaker is open, some calls that would have worked get the fallback, so we put the breaker state on a dashboard and alerted on it.”
resilience4j:
circuitbreaker:
instances:
recommendations:
slidingWindowSize: 50
minimumNumberOfCalls: 20
failureRateThreshold: 50
slowCallDurationThreshold: 2s
slowCallRateThreshold: 50
waitDurationInOpenState: 30s
A fallback that returns an empty success, so failures look like no data and nobody notices.
Why it stuck: the record failed every time, and a partition is read in order, so nothing behind it moved.
Fix: a deserializer that turns bad payloads into handled errors, a few retries with backoff, then a dead-letter topic.
Classify: errors that can never succeed, like bad data, skip the retries.
Trade-off: dead-lettered records need an alert, an owner and a way to replay them.
“Another team produced a payload with a field of the wrong type, so it failed to deserialise every time. A partition is read in order, so the consumer kept failing on that record, everything behind it waited, and lag climbed. I made three changes. I wrapped our deserializer in ErrorHandlingDeserializer, so a bad payload reached our error handler instead of failing inside the poll. I configured a DefaultErrorHandler with a short backoff and a DeadLetterPublishingRecoverer, so after a few tries the record went to a dead-letter topic and the partition moved on. Deserialisation errors already skip retries in that handler, and I added our validation exception to that list, because retrying bad data never helps. The risk is that dead-lettered records get forgotten, so we alerted on that topic and wrote a small tool to replay them once the producer fixed their bug.”
@Bean
DefaultErrorHandler errorHandler(KafkaTemplate<Object, Object> template) {
DefaultErrorHandler handler = new DefaultErrorHandler(
new DeadLetterPublishingRecoverer(template),
new FixedBackOff(1000L, 3)); // 3 retries, 1 second apart
handler.addNotRetryableExceptions(ValidationException.class);
return handler;
}
Catching every exception in the listener and just logging it, so bad data disappears without a trace.
Right now: find the client in the logs and throttle it at the gateway, before it reaches the app.
Design: limits keyed on the client's API key, not the IP, answered with a 429 and a Retry-After header.
Where: the gateway for simple limits; the app only for limits that need business data.
“The logs showed one API key calling a search endpoint hundreds of times a second, which used up our database connections. The quick fix was a limit for that key at our API gateway, so the traffic never reached the app. Then I did it properly. Limits were keyed on the client's API key, not the IP address, because several customers sat behind one shared office network. Over the limit, clients got a 429 with a Retry-After header, so a well-written script could slow down instead of failing. Most limits stayed at the gateway, where blocking is cheap and every instance sees the same counts. One limit, on expensive exports per account, needed business data, so it lived in the app with a counter in Redis that all instances shared. I also called the customer. They needed a bulk endpoint, and building one was the real fix.”
Blocking the client's IP address and moving on, with no per-client limit and no clear 429 response.
Cause: the endpoint checked the user was logged in, never that they owned the record.
Fix: enforce ownership in the query or with method security such as @PreAuthorize, not in scattered if statements.
Stop it recurring: tests that call each endpoint as the wrong user, plus a review checklist item.
“The filter chain only checked that the token was valid, and the controller loaded the invoice by id. Authentication was fine; authorisation was missing. The first fix was to scope the repository call to the current user, finding the invoice by id and owner, and returning 404 when nothing matched, so we didn't even confirm the id existed. For support staff who could view any invoice, I used @PreAuthorize with a small bean that checks the role or ownership, so the rule lived in one place. Then I searched for the same pattern and found two more endpoints with it. To stop it coming back, I added integration tests that call every id-based endpoint as a second user and expect a 404 or 403. Repository methods got a little more specific, which I think is a fair price.”
@PreAuthorize("hasRole('SUPPORT') or @invoiceAccess.isOwner(#id, authentication)")
@GetMapping("/invoices/{id}")
public InvoiceResponse get(@PathVariable Long id) {
return invoiceService.find(id);
}
Treating hard-to-guess ids, such as switching to UUIDs, as the fix.
Sources: full request logging, toString on DTOs, and error messages that echo the input.
Fix at the source: no bodies or auth headers in logs, sensitive fields left out of toString, messages that name the field, not the value.
Clean up and guard: revoke exposed tokens, shorten retention, and add a test so it cannot creep back.
“A support engineer spotted a bearer token in our log search. I traced three sources. A debug filter that logged full request bodies and headers had been left on in production. Several DTOs had generated toString methods, and one error log printed the whole object. And a validation message echoed the rejected value back, phone numbers included. First I had the exposed tokens revoked, since anyone with log access could have used them. Then I removed body and header logging, left sensitive fields out of toString, and changed messages to name the field, not the value. We shortened log retention too. To keep it fixed, I added a test that sends known fake values through the main endpoints and fails if any appear in the captured logs. Debugging got a little harder without bodies, so we log record ids and people look up the record with proper access.”
Deleting the old log files and stopping there, without revoking the tokens or fixing what wrote them.
How caching works: Spring reuses a test context only when the test configuration is exactly the same.
What broke it: each class mocking a different set of beans, different properties, and @DirtiesContext used just in case.
Fix: a few shared test setups, and @DirtiesContext only where state truly cannot be reset.
“Spring caches the application context between test classes, but only when their configuration is exactly the same, and ours almost never was. Each class mocked its own mix of beans, set a few different properties, and several used @DirtiesContext just in case, which throws the context away. Counting startup banners in the build log, I found over forty context starts, each paying for startup and a fresh database container. I created three shared base classes, one per kind of test, each with the same mocks and properties, and moved per-test differences into setup code. I removed @DirtiesContext wherever cleaning the tables between tests did the job. We got down to a handful of contexts and the suite ran in about six minutes. The trade-off is that base classes mock a few beans some tests don't need, which looked odd at first, so I wrote a short note explaining why.”
Buying bigger build machines or running tests in parallel without asking why each class starts its own context.
Symptoms, not causes: page on error rate and latency of the endpoints users depend on.
Over a window: judge a few minutes of traffic, so one bad request never pages anyone.
Everything else: dashboards or daytime tickets, and every page links to a short runbook.
“We had alerts on CPU, heap, every single 500 and a dozen log patterns. The team had muted the channel, and a real outage went unnoticed for twenty minutes. I rebuilt the paging alerts around what users feel: error rate and latency for our main endpoints, from the http.server.requests metrics Boot already records through Micrometer, judged over a few minutes so one bad request doesn't wake anyone. Consumer lag got a page too, because a stuck consumer is invisible to users until it's very late. CPU and heap moved to dashboards and daytime warnings. Every page linked to a runbook saying what to check first. Pages went from several a night to a couple a week, and the next real incident paged within minutes. The trade-off is that slow-burn problems like a filling disk no longer page, so we kept a few long-lead alerts for those.”
Adding more alerts after every incident until nobody reads any of them.
What I check: layering, where the transaction starts and ends, error handling, config and secrets, and tests of real behaviour.
How: separate must-fix from nice-to-have, explain why, and point to a good example.
Growth: pair on a pattern once, then let them apply it.
“I check a few things that are expensive to fix later. Is business logic creeping into the controller? Where does the transaction start and end, and is anything slow, like an HTTP call, sitting inside it? Are errors going through our shared handler, or being swallowed? Is anything hard-coded that should be config, or worse, a secret in the properties file? And do the tests exercise real behaviour, not mocks returning mocks? For feedback, I label comments as blocking or optional so a junior knows what matters, and I explain why instead of pasting the fix. One developer put @Transactional on every method; instead of twenty comments, I spent half an hour walking through one of their classes with them, and their next pull requests changed. Teaching is slower than fixing it myself, but it scales and I stop being the bottleneck.”
Rewriting the code yourself in review comments, or approving everything to keep the peace.
Say yes to the goal: customers need the fix today, so the answer is how, not whether.
Make it safe: count the rows first, copy them aside, run in a transaction, and have a second person read the script.
Stop the source: fix the code that wrote the bad data, or the same UPDATE will be needed again.
“I'd agree the data had to be fixed today, but not by typing an UPDATE into a production console. When this happened to me, a bad release had set the wrong status on a few thousand orders. I wrote the fix as a script: a select that counted the affected rows, a copy of those rows into a backup table, then the update inside a transaction that checked the row count matched before committing. A teammate reviewed it in ten minutes, and it ran within the hour, so the manager got speed and we kept the safety. Then we fixed the bug that wrote the bad data, because otherwise we'd be back the next week. Afterwards I proposed that any production data change goes through a reviewed script, which the team adopted. It adds a few minutes, and one mistyped WHERE clause costs far more.”
Running the update alone with no count, backup or review because the manager asked, or refusing with no faster safe option.
The decision: what you chose and why it looked right with what you knew then.
The cost: the concrete signal that it wasn't working.
The reversal: how you undid it safely and what you'd check earlier next time.
“In a service I owned, I decoupled modules with Spring application events. Creating an order published an event, and loyalty points, emails and analytics each listened for it. It looked clean, because the order code knew about none of them. A year later it cost us. Nobody could see the whole flow by reading the code, and a new listener threw an exception. Plain @EventListener methods run in the publisher's thread and transaction, so a failing loyalty update rolled back the order itself. I reversed most of it. Steps that must happen with the order became explicit calls in the order service, so the flow reads top to bottom. Things that can happen afterwards, like emails, moved to @TransactionalEventListener, which by default runs after the commit. What I weigh now is that indirection has a reading cost, and I only pay it when there really are many independent listeners.”
Picking a story where the decision was really someone else's, or one with no cost and nothing learned.
ClapAssist is an AI interview assistant for Mac and Windows. It listens to the interview on your computer and shows you what to say, in short lines you can read while you talk. Your live interview audio and screen are never stored. Your resume and notes are saved to your account so the app fills them in on any computer. It stays out of screen share on every plan, including Free; only you can see it.