skip to content

For a bulk-write API endpoint, would you promise all-or-nothing atomicity or best-effort per-item processing? Argue the tradeoff and say how you would express the choice in the contract.

level: principalimportance: should knowfreq 30%

answer

  1. best-effort = scales, one bad item costs one item
  2. atomic = one transaction, caps size, holds locks, often impossible across shards/side effects
  3. poison item blocks whole batch forever under atomic
  4. explicit atomic flag + tighter limit; never silently degrade
  5. atomicity != idempotency - retry still duplicates

basics

~20 s

Default to best-effort with per-item results: it scales, degrades gracefully and matches how callers actually recover. Offer atomicity only for small batches within one datastore, as an explicit opt-in flag, and reject batches too large to hold in one transaction.

solid answer

~60 s

**Best-effort** processes each item independently and reports per-item outcomes. It parallelises, keeps batches large, and a bad item costs only that item. The price is that the caller must own reconciliation - track which items landed and retry precisely the rest. **Atomic** means the whole batch commits or none of it does. Callers love it because failure handling collapses to "fix and resend", but it needs one transaction spanning every item, so it caps batch size, holds locks longer, raises deadlock and timeout risk, and is often simply impossible - across shards, across services, or when items trigger side effects like emails or payments. My default is best-effort with rigorous per-item reporting, plus an explicit `atomic: true` opt-in bounded to a small item limit and a single datastore. What I refuse to do is leave it unstated: silence means callers assume atomicity, discover otherwise during an incident, and build recovery on a guarantee that never existed. And atomicity is not idempotency - an atomic batch retried after a timeout still duplicates.

go deeper

for a junior

State the difference plainly - all-or-nothing versus each item independent - and that the API must document which one it does.

for a middle

Give concrete consequences: one bad item rejecting the whole batch, transaction size limits, and the need for per-item results in best-effort mode.

for a senior

Bring in operational cost - lock duration, timeouts, poison items, cross-shard impossibility - and design the opt-in flag with mode-specific status semantics.

for a principal

Frame it as where partial-state complexity should live, decide per endpoint on whether partial application is incorrect or merely inconvenient, and separate atomicity from idempotency and from ordering guarantees.

## The two contracts **Atomic (all-or-nothing)**: every item commits or none does. The client's mental model is a single transaction: on failure, fix the payload and resend the whole thing, with no reconciliation logic. **Best-effort (per-item)**: each item is processed independently; the response reports each outcome. The client retries only what failed. ## Why best-effort is usually the right default **It scales.** Independent items can be processed in parallel or in chunks, so batch size is bounded only by timeouts, not by how long a transaction can be held open. **It degrades gracefully.** One item referencing a deleted product costs one item. Under an atomic contract that single bad row rejects 999 good ones - and if the caller's data source keeps producing that row, the caller is permanently blocked and must implement bisection to find the poison item. **It fits real recovery.** Bulk callers - importers, sync jobs, mobile flush queues - already track per-record state. Per-item results feed that machinery directly. **It is often the only thing you can implement honestly.** Atomicity requires a transaction spanning all items. That is unavailable across shards or partitions, across microservice boundaries without a saga or two-phase commit, and whenever an item causes an external side effect - a charge, an email, a webhook - that cannot be rolled back. Promising atomicity you cannot deliver is worse than not offering it, because clients will build on it. ## When atomicity genuinely earns its place - **Financial or ledger writes** where a half-applied batch is a correctness violation, not an inconvenience. - **Referentially coupled items** - a parent and its children in one payload, where partial application leaves orphans. - **Configuration or policy updates** that must take effect as a set, or the system enters a state no one designed. In those cases pay the price deliberately: cap the item count hard (tens, not thousands), keep the batch inside a single datastore, and say so. ## Expressing the choice in the contract Make it explicit and, where you can, selectable: - A request field (`"atomic": true`) with a documented default. Google's API guidance and several database APIs use exactly this shape. - **Different status semantics per mode.** Atomic mode has a single outcome, so a failure is one 4xx naming the offending item - a 207 Multi-Status would be incoherent. Best-effort mode returns per-item results under 207 or 200. - **A tighter size limit for atomic mode**, and a documented rejection when a caller asks for atomicity beyond it, rather than silently degrading to best-effort. Silent degradation is the worst outcome available: the caller believes a guarantee that is not being honoured. - **A written statement about ordering and side effects** - whether items may be processed in parallel, and whether a rolled-back atomic batch can still have emitted events or notifications. ## The trap: atomicity is not idempotency These get conflated constantly. Atomicity says the batch does not apply halfway. It says nothing about what happens when a client times out and resends: an atomic batch applied twice is two complete applications. If callers retry - and any client behind a network will - you still need per-item deduplication keys or a request-level idempotency mechanism. Answering "we made it atomic" to a duplicate-records question is the classic misdiagnosis. ## The principal framing Decide by asking who absorbs partial-state complexity. Best-effort pushes it to callers, where they usually already have per-record tracking. Atomic absorbs it into the server, at the cost of batch size, lock duration and implementability. Pick per endpoint according to whether partial application is merely inconvenient or actually incorrect - and never leave the answer unwritten.

  • A caller says an atomic batch keeps failing because one record is always bad. What do you tell them?
    That is the structural cost of atomicity: one poison item blocks the set indefinitely, and the caller ends up bisecting the batch to find it. Either switch that endpoint to best-effort so the failure isolates to one item, or add a pre-validation endpoint that reports which items would fail without committing anything.
  • Does making a batch atomic solve duplicate records from client retries?
    No. Atomicity governs partial application within one execution; it says nothing about repeated executions. A timed-out atomic batch that actually committed will commit again on retry, producing a full duplicate set. Duplicates are solved by deduplication on a caller-supplied key or a request-level idempotency mechanism.

Best-effort is a delivery van that drops what it can and brings back the failures; atomic is a van that returns fully loaded if a single address is wrong. The second is simpler to reason about and useless at scale.

saying these in an interview costs you the question

  • Promising atomicity across shards, services or externally visible side effects
  • Conflating atomicity with idempotency when asked about duplicate records
  • Silently falling back to best-effort when an atomic batch exceeds the size limit
  • Leaving the semantics undocumented so callers assume all-or-nothing
  • Choosing atomic by default at large batch sizes, ignoring lock duration and timeout risk

context