skip to content

An in-process broker stand-in never redelivers a message. Which behaviours go untested?

level: seniorimportance: nice to knowfreq 22%

answer

  1. Exactly once is not what production offers
  2. The guard is never on the exercised path
  3. Acknowledgement, lease expiry, attempt count
  4. A reference is passed, not bytes
  5. Make the stand-in harsher than production

basics

~20 s

Everything that depends on delivery not being exactly once: consumer idempotency, acknowledgement and retry, backoff, dead-lettering after repeated failure, ordering after a redelivery, and reprocessing when a consumer restarts. Handing the object straight to the handler also hides serialization and shared-mutation defects.

solid answer

~50 s

A stand-in that calls the handler once, in-process, with the same object reference models a delivery guarantee no real broker offers. What disappears is the whole redelivery surface: duplicates from an acknowledgement lost after the work was done, a visibility or lease timeout expiring on a slow handler, retry with backoff, the attempt count that routes a repeatedly failing message aside, and reprocessing after a consumer restarts mid-batch. Because nothing is encoded, serialization defects and handlers that mutate a message another consumer still holds also vanish. The practical consequence is that the idempotency guard — the thing that makes at-least-once delivery safe — is written but never exercised. If you keep the stand-in, make it deliberately hostile: deliver twice, deliver out of order, fail an acknowledgement, and encode and decode the payload on the way through.

code

pseudocode · 19 lines
pseudocode
deliver(batch):
    doubled = []
    for m in batch:
        doubled.add(m)
        doubled.add(m)
    shuffle(doubled)

    for m in doubled:
        payload = decode(encode(m))
        try:
            handler(payload)
            if dropAckAfterHandle:
                requeue(m)
        except:
            m.attempts += 1
            if m.attempts >= maxAttempts:
                routeAside(m)
            else:
                requeue(m)

go deeper

for a junior

Know the headline: real messaging delivers at least once, so duplicates happen, and a stand-in that hands each message to the handler exactly once is modelling a guarantee production does not give.

for a middle

Explain the mechanisms that create a duplicate — a lost acknowledgement, an expired lease on a slow handler, a crash between doing work and acknowledging — and what a consumer must do to be safe under them.

for a senior

Show that you would redesign the instrument: deliver twice by default, shuffle within a batch, encode and decode, model acknowledgement and an attempt count, so the guard is exercised by every case rather than by none.

for a principal

Own where the line sits between a deliberately hostile in-process stand-in and a smaller tier against real middleware, and be able to justify the run-time cost of that tier against the class of defect it is the only thing that catches.

## What the substitute is modelling An in-process broker stand-in replaces the messaging middleware with a direct call: publish puts the message on a list, the test drains the list, and the handler runs on the same thread with the same object. It is fast, needs nothing running, and makes an asynchronous flow behave synchronously so an assertion can follow the publish immediately. That last property is the reason it is popular and also the reason it misleads — it converts an at-least-once, eventually-delivered, possibly-reordered channel into an exactly-once, immediate, in-order function call. ## The behaviours that disappear **Redelivery and duplicates.** Real brokers deliver at least once. A duplicate arrives when an acknowledgement is lost after the handler has already committed its work, when a lease or visibility timeout expires because the handler was slow, when a consumer crashes between doing the work and acknowledging, or when the client library retries a publish it never saw confirmed. Every one of these produces a second execution of a handler that already succeeded. The guard against it — an idempotency key, a processed-message record, a conditional update — is therefore the single most important thing about the consumer, and a stand-in that delivers once never runs the guard's interesting branch. **Acknowledgement semantics.** Whether the handler acknowledges before or after doing its work decides whether a crash loses a message or duplicates it. A stand-in with no acknowledgement at all cannot express the difference, so the choice is invisible in review and untested in the suite. **Retry, backoff and giving up.** A failing handler on a real broker is retried, often with growing delay, and after a threshold the message is routed aside for inspection rather than retried forever. The threshold logic, the delay, and the behaviour of everything queued behind a stuck message are all absent from a stand-in that simply propagates the exception into the test. **Ordering.** Real ordering guarantees are narrow — commonly per key or per partition, and broken by a redelivery, a rebalance or a parallel consumer. A stand-in that drains a list preserves publish order globally, which is a stronger guarantee than production offers and lets order-dependent handlers pass. **Restart and rebalance.** Consumers restart, and work partly done is repeated from the last committed position. Group membership changes move partitions between consumers mid-flight. Neither exists in-process. **Encoding.** The most under-appreciated gap: a real broker carries bytes. A stand-in passes a reference. So an unencodable field, a schema change the decoder rejects, a lost precision on a number, a dropped timezone, and a handler that mutates a message the publisher still holds are all invisible. Encoding and decoding on the way through the stand-in costs almost nothing and recovers most of this. ## The worked case A subscription renewal job publishes one renewal instruction per due subscription and a consumer charges the card. In the suite each instruction is handled exactly once, so the consumer's guard — skip if a renewal already exists for this subscription and this period — is only ever hit on its miss path. In production, one batch takes long enough for the lease to expire and 118 instructions are redelivered. The guard compares the period boundary with a strict inequality where it needed a non-strict one, so an instruction whose renewal landed exactly on the boundary is not recognised as already processed and the customer is charged twice. The defect is in a line written specifically to prevent duplicates, and the suite could not reach it because the suite had no duplicates. ## Making the substitute hostile instead of polite If the stand-in stays, stop letting it model the happy path. Useful properties to build in, each cheap: - deliver every message twice by default, so the idempotency guard is on the main path of every case rather than an afterthought; - shuffle delivery order within a batch, so an accidental ordering assumption fails; - encode and decode the payload with the real serializer; - model acknowledgement explicitly, and offer a mode where an acknowledgement is dropped after the handler returns; - expose an attempt counter and a route-aside destination so retry limits and their consequences are assertable; - let a test make a handler slow enough to lose its lease. A stand-in with these properties is a genuinely useful test instrument, because it is now *harsher* than production and a passing consumer is safe under a weaker guarantee than it will actually get. Complementing it with a smaller tier of cases against a real broker is what covers what remains — the client library's own behaviour, and the broker's actual redelivery and rebalance timing. Worth stating plainly: none of this argues the stand-in is worthless. It argues that its default configuration encodes an assumption nobody would defend out loud, and that changing the default is cheaper than any other fix available here.

  • Why does a stand-in that passes the message object by reference hide defects a real broker exposes?
    Because a real broker transports bytes. Passing a reference skips encoding and decoding entirely, so an unencodable field, a decoder that rejects a changed shape, lost numeric precision, a dropped timezone and a handler that mutates an object the publisher still holds all go unseen. Encoding and decoding inside the stand-in recovers most of that for almost no cost.
  • If you make the stand-in deliver every message twice, do you still need any cases against a real broker?
    Yes, a small tier. Duplicate delivery covers your consumer's idempotency, but not the client library's configuration, the acknowledgement mode actually in force, lease and heartbeat timing, partition assignment during a membership change, or the broker's own ordering guarantee. Those are properties of the real system and only the real system demonstrates them.
  • How would you prove a consumer is genuinely idempotent rather than accidentally so?
    Assert on observable effect counts, not just on the absence of an exception: run the same message twice and require exactly one charge, one outbound notification and one state transition. Then vary the interleaving — a duplicate arriving before the first attempt commits — because a guard that reads then writes without a conditional update passes the sequential case and fails the concurrent one.

saying these in an interview costs you the question

  • Assumes messaging delivers each message exactly once
  • Writes an idempotency guard and never tests a duplicate
  • Relies on global publish order the stand-in invents
  • Ignores that a real broker carries bytes, not references
  • Treats a repeatedly failing message as retried forever

context