skip to content

A web service with Open Session in View enabled degrades badly under load, with requests queueing on the connection pool. Explain what that pattern does to connection usage and query count during response rendering, and how you would confirm it in production.

level: seniorimportance: must knowfreq 40%

answer

  1. pool used during rendering, the slowest phase
  2. query count scales with rendered rows
  3. delayed acquisition = churn, eager = pinned
  4. acquisition wait clustered after commit
  5. fix fetch strategy before pool size

basics

~20 s

Rendering triggers ad-hoc lazy loads, so the pool is being used at the end of the request, when slow serialization stretches the window. Query count grows with the rendered graph. Confirm with per-request query counts and connection-acquisition timings.

solid answer

~50 s

Two effects compound. **Query count.** Every unfetched association the serializer or template walks becomes a separate `SELECT`. A list endpoint that renders one field of a lazy association produces one query per row — invisible in the service code, so it never shows up in review. **Connection timing.** Those statements need a connection *during rendering*, the slowest and least predictable phase. Older setups pinned one connection for the whole request; modern Hibernate acquires lazily and releases between statements outside a transaction, so you instead get repeated acquire/release churn while rendering. Either way, throughput is now bounded by pool size across the full request duration rather than just the service transaction. To confirm: count statements per request (a statement-count assertion or query-log interceptor), measure connection acquisition wait time from the pool's metrics, and compare latency of an endpoint fetched eagerly versus lazily. The tell is a query count that scales with result-set size and pool wait time concentrated after the transaction has committed.

code

java · 6 lines
java
Statistics stats = em.getEntityManagerFactory()
                     .unwrap(SessionFactory.class)
                     .getStatistics();
long before = stats.getPrepareStatementCount();
// ... handle request, including rendering ...
long perRequest = stats.getPrepareStatementCount() - before;

go deeper

for a junior

Recall that keeping the session open for the request lets rendering trigger extra queries, which costs database round trips.

for a middle

Explain how query count scales with the rendered graph and why those statements land in the rendering phase.

for a senior

Own the diagnosis: statement counts per request, pool acquisition wait after commit, connection handling mode, and fixing fetch strategy before pool sizing.

for a principal

Reason about capacity — pool size as throughput times hold time, the risk of coupling database concurrency to response size, and the policy of forbidding data access in the presentation layer.

## Why it degrades specifically under load A request with a request-scoped session has two phases that touch the database: the service transaction, and rendering. Without the pattern only the first exists; the connection is borrowed and returned inside a tight, code-reviewed window. With it, rendering can also hit the database — and rendering is where the unpredictable work lives: serializing a large graph, template evaluation, content negotiation, sometimes writing to a slow client socket. Pool capacity is throughput times hold time. Adding database work to the longest phase inflates hold time and, worse, correlates it with response size. Under load the pool saturates, requests queue for a connection, and latency goes non-linear. ## The query-count side This is the larger problem in practice. A service method returns 200 entities with a lazily mapped `customer`. Nothing in the service triggers it, so the code looks clean. The serializer then walks each element and touches `customer.name`, producing 200 extra `SELECT`s. Without a request-scoped session those accesses would have thrown immediately in development — the pattern converts a loud, early failure into a silent per-row query that only hurts in production. That is the deepest objection to it: it does not cause N+1 selects, but it removes the feedback that would have caught them. The cost is unbounded in the wrong dimension too: the number of statements depends on how much of the graph the *view* decides to render, which is not the service author's decision. ## The connection side, precisely The folklore claim is "the connection is held for the whole request". Be careful — modern Hibernate uses delayed connection acquisition: a connection is obtained on first statement need, and for non-transactional work it can be released between statements, so a session that is open but idle is not necessarily holding a connection. The accurate statement is: - with eager acquisition or a JTA setup, the connection genuinely can be pinned for the request; - with delayed acquisition, you get repeated borrow/return cycles during rendering, adding pool contention and acquisition latency per lazy load rather than one long hold. Either way the pool is a participant in the rendering phase, which is the structural problem. Check `hibernate.connection.handling_mode` before asserting which of the two your system does. ## Confirming it in production 1. **Count statements per request.** Enable statement counting (Hibernate's statistics, or a proxying DataSource that logs and counts) and correlate the count with response size. Query count that scales linearly with rendered rows is the signature. 2. **Split the timeline.** Instrument transaction-commit time versus response-complete time. If pool-acquisition waits cluster *after* commit, the database work is coming from rendering. 3. **Read pool metrics.** Connection acquisition wait time, pending threads, and active-connection duration percentiles. Rising acquisition wait with unchanged query latency means contention, not a slow database. 4. **A/B one endpoint.** Rewrite one hot endpoint to fetch what it needs explicitly (join fetch or a projection) and return data the view cannot extend. Compare query count and p99. This both proves the diagnosis and is the fix. 5. **Check the startup log.** Frameworks that enable this by default typically warn about it at startup; its presence tells you the pattern is on without an explicit decision. ## Mitigation ordering Raising the pool size treats the symptom and moves the bottleneck to the database. Fix the query count first: fetch the graph the endpoint actually needs in the service transaction, or return a projection so rendering has nothing left to load. Only then reconsider whether the request-scoped session is still earning its keep — by that point it usually is not, and turning it off makes any regression fail loudly in tests rather than silently in production.

  • Is it accurate to say the JDBC connection is held for the entire request?
    Only for eager acquisition or JTA setups. Modern Hibernate acquires connections lazily and, for statements outside a transaction, can release them between statements, so an idle open session need not hold one. The accurate framing is that the pool is exercised during rendering — as a long hold or as repeated churn depending on the connection handling mode.
  • Why not just increase the connection pool size?
    It hides the real defect and moves the bottleneck. The query count is still proportional to rendered rows, so the database now absorbs the multiplied load, and a larger pool means more concurrent sessions competing there. Fix the fetching so rendering issues no queries, then re-measure; pool sizing should follow the workload, not compensate for accidental queries.

saying these in an interview costs you the question

  • Blaming the database for slowness when query count per request is the actual variable
  • Asserting flatly that one connection is pinned per request without checking the connection handling mode
  • Raising pool size as the primary fix
  • Claiming the pattern causes N+1 — it hides N+1 by letting it succeed instead of throwing
  • Measuring only average latency, missing that the cost scales with response size

context