skip to content

How would you set unit-of-work and connection-lifetime policy for a service mixing short requests, long batch jobs and multi-step flows?

level: principalimportance: should knowfreq 38%

answer

  1. one invariant for the whole service
  2. open only while issuing statements
  3. chunk the batch, checkpoint the chunk
  4. detached state between flow steps
  5. alert on borrow wait, not query latency

basics

~20 s

Set policy per workload against one invariant: a unit of work is open only while it is issuing statements. Requests get a thin boundary; batches one unit per chunk with the tracked set discarded between chunks; flows keep detached state.

solid answer

~50 s

Write one rule the whole service is judged against - **a unit of work stays open only while it is issuing statements, never while waiting on a remote call, a lock, rendering or a human** - then specialise it. Request handlers get a thin transactional boundary with everything the response needs loaded before it ends. Batch jobs must never run under a single unit of work: chunk the work, take one unit and one transaction per chunk, discard the tracked set between chunks so memory does not grow, and checkpoint so a re-run resumes. Multi-step flows keep their state detached between steps and re-load at each one, with a version column supplying the concurrency check the long transaction would have given. Then make the policy observable: hold time per boundary, and alerts on time spent waiting to borrow rather than only query latency.

go deeper

for a junior

Take away the one rule that generalises: a unit of work should be open only while it is actually running statements, and long-running work should be split into pieces that each commit.

for a middle

Be able to give the shape per workload — thin boundary for a request, one unit and transaction per chunk for a batch, detached state between the steps of a flow — and say what each avoids.

for a senior

Show how you would enforce it: hold time per boundary, statements attributed to code paths, borrow-wait alerts, and a review rule against calls to other systems inside a boundary.

for a principal

Name the price you are choosing to pay — lost whole-job atomicity, version conflicts surfaced to users, read paths that must declare what they need — and why a policy that depends on a particular pool size is not a policy.

## Start from one invariant Every workable policy reduces to a single sentence that a reviewer can apply without knowing the domain: > **A unit of work is open only while it is issuing statements.** Everything else is a consequence. A connection slot is a service-wide allocation; concurrency is slots divided by average hold time. Any span during which a boundary is open but not querying is capacity spent on nothing, and the four classic offenders are the same everywhere: a call to another system inside the boundary, waiting on a lock, rendering a response, and waiting for a human. ## Policy per workload class **Short request handlers.** A thin transactional boundary around the statements, with validation, authorisation checks against other systems, and response formatting outside it. Everything the response needs is loaded before the boundary ends, so nothing downstream can fetch. This is also where the default acquisition policy matters least, because the boundary contains little but statements. **Batch and background jobs.** The tempting shape — one unit of work for the whole run — fails on three axes at once: memory grows with everything ever loaded, one connection is held for hours, and a failure at 90% loses everything. The workable shape is chunked: 1. Read a bounded slice of work, ordered so slices are stable and resumable. 2. Open one unit of work and one transaction for that slice; process and commit it. 3. Discard the tracked set so the next slice starts empty, and let the connection go back to the pool between slices. 4. Record a checkpoint with the commit, so a re-run resumes at the boundary of the last committed slice. 5. Size the slice so a chunk's transaction is short; then a retry is cheap and a failure costs one slice. **Multi-step flows.** State survives across requests as detached objects held in the client or in a store with an expiry; each step opens its own short unit of work. Concurrency is enforced with a version column checked in the final update rather than by a transaction held open across the user's think-time. Abandoned flows expire on their own. **Streaming and export endpoints.** Treat them as batches with a socket attached. Either materialise before responding, or process in committed chunks; never hold one boundary open for the consumer's read speed. ## The budget, stated in numbers | Workload | Unit of work | Connection held for | Bounded by | |---|---|---|---| | Request handler | One per request | The statement phase | A latency budget per endpoint | | Batch job | One per chunk | One chunk | A maximum chunk duration | | Multi-step flow | One per step | One step | Nothing held between steps | | Streamed export | One per chunk | One chunk | Never the client's read speed | A concrete way to express the budget: publish a **maximum open time** per class — for example, request boundaries in the low tens of milliseconds, chunk transactions in the low seconds — and treat a breach as a defect rather than as tuning. Note the deliberate omission: none of this is about how many slots the pool has. Sizing is a separate decision, and a policy that only works at a particular pool size is not a policy. ## Making it hold - **Measure hold time, not query time.** The gap between them is the waste, and it is invisible on a query-latency dashboard. - **Alert on time spent waiting to borrow.** Rising borrow wait with an unremarkable server is the signature of scope problems; it fires long before user-visible timeouts. - **Attribute statements to code paths** and assert their counts in tests for the busiest paths, so a new per-row loop fails a build. - **Make "no calls to other systems inside the boundary" a review rule**, because it is checkable by reading a diff and it is where the worst holds come from. - **Cap what a single unit of work may accumulate** in a long-running process, so an unbounded loop fails loudly instead of degrading. ## The tradeoffs a lead actually owns The policy costs something and the honest version says so. Chunked batches give up whole-job atomicity, which means every job must be re-runnable and every chunk idempotent — real design work pushed onto job authors. Detached multi-step flows give up the guarantee that the data has not moved, replacing it with a version check and a conflict the user must resolve on the last screen. Loading everything at the boundary means read paths must state what they need, so a screen change becomes a change in two places. These are worth paying because the alternative — long-held scopes — fails as a *system* property: it degrades every workload at once, at peak, and it is diagnosed last, after the team has looked at the database, the queries and the pool size.

  • What do you give up by refusing to run a batch job under one transaction?
    Whole-job atomicity. Partial results become visible, so every job must be idempotent and re-runnable from a checkpoint, and consumers must tolerate a partially applied run. That is real design work, accepted because a job-long transaction holds a connection for hours, grows memory with everything loaded, and loses all progress on a late failure.
  • Why alert on borrow wait rather than on query latency?
    Because the failure being hunted is occupancy, not slowness. Long-held scopes leave query latency looking normal while callers queue in the application for a slot. Borrow wait rises first and points at the right layer; query latency rises later, or never, and sends the investigation to the database.
  • How do you keep the policy from depending on the pool being a particular size?
    State it as a bound on open time per workload class and verify each class independently, so correctness does not rest on capacity. Sizing then answers a separate question — how much concurrency to buy — and a change to it never silently invalidates the scope rules.

saying these in an interview costs you the question

  • Runs a whole batch job inside one unit of work and transaction
  • Sets one global scope default for every workload in the service
  • Answers pool starvation by raising the pool size first
  • Keeps a transaction open across a call to another system
  • Measures query latency but never how long a boundary stays open
  • Assumes chunked jobs need no idempotency or checkpointing