Walk through the claim check pattern step by step for a pipeline where a producer emits a large report and a consumer needs to process it: what exactly happens, in what order, and what does each side own?
answer
- write-then-publish ordering
- reference + metadata, not payload, in the message
- consumer makes a second call to storage
- checksum in the claim check for validation
- who owns cleanup is a separate decision
basics
~20 sProducer saves the big report to storage first, gets back a location, then sends a small message with that location through the queue. Consumer reads the message, fetches the report from storage using the location, then processes it.
solid answer
~50 sFour steps. One, the producer writes the payload to external storage (object store, blob store, or cache) and waits for confirmation that the write succeeded and is durable. Two, the producer builds a claim check -- a message containing the storage reference plus light metadata like size, content type, and maybe a checksum -- and publishes it to the broker. Three, the consumer receives the small message from the broker through its normal subscription and extracts the reference. Four, the consumer calls the storage system directly to fetch the payload, validates it if a checksum was included, and processes it; it may also be responsible for deleting or marking the payload for cleanup once done, depending on who owns the payload's lifecycle. The broker never sees the payload bytes at any point -- it only ever carries the reference.
go deeper
Should be able to list the four steps in the right order at a basic level: write payload, publish reference, receive reference, fetch payload.
Should explain why the write-then-publish ordering is required and what metadata belongs in the reference message versus what stays out of it.
Should discuss what happens on partial failure at each step (crash after write but before publish, crash after fetch but before processing) and how idempotency and retries interact with each step.
Should reason about ownership of the payload lifecycle across multiple consumers in a fan-out topology and design the reference schema to be extensible/versioned as the pipeline evolves.
## The cast and the one hard constraint The claim check pattern has exactly two participants, a producer and a consumer, plus two infrastructure pieces, a message broker and an external storage system, and the mechanism is best understood as a strict **four-step handoff** with one hard ordering constraint. ## The producer side **Step one: write the payload.** The producer generates or receives the large payload -- a report, an image, a video, a batch extract -- and writes it to the storage system. This is a synchronous call from the producer's point of view: it must wait for the storage system to acknowledge that the write is durable before moving on, because everything downstream depends on the payload already existing by the time anyone tries to read it. The storage system returns, or the producer deterministically constructs, an identifier for what it just wrote: - an object key, - a full URL, - a composite of container name plus blob name, or similar. **Step two: publish the claim check.** The producer constructs the claim check itself and publishes it to the broker. The claim check is not the payload -- it is a small envelope containing the reference from step one plus whatever metadata the consumer will need to safely and efficiently retrieve and validate the payload: - a **content type**, so the consumer knows how to parse it; - a **size**, so it can pre-allocate buffers or reject unexpectedly large fetches; - often a **checksum** or hash, so the consumer can detect corruption or a mismatched object after fetching. This publish happens only after step one's write is confirmed durable -- publishing the reference before the write completes is the single most common bug in claim-check implementations, because a consumer could receive and act on the reference before the payload actually exists in storage, producing a spurious not-found error that looks like a transient failure but is actually a logic bug in the producer. ## The consumer side **Step three: receive the reference.** The consumer receives the small reference message through whatever subscription mechanism the broker offers -- a queue poll, a push subscription, a `Kafka` partition read -- exactly as it would for any other message, since from the broker's perspective this is an ordinary small message with no special handling. The consumer deserializes it and extracts the reference and metadata. **Step four: fetch the payload.** The consumer makes a separate, direct call to the storage system, not the broker, to fetch the actual payload using the reference. This is a genuinely new network hop that doesn't exist in a broker-only design, and it introduces its own failure surface: - the storage system might be unavailable; - the object might have been deleted by a retention policy before the consumer got to it; - the fetch might return corrupted bytes that fail the checksum check. Once fetched, the consumer processes the payload as if it had arrived directly in the message. ## Who owns cleanup A design decision that doesn't always get made explicitly is who owns cleanup: 1. does the producer delete the payload after some retention window regardless of consumers, 2. does the consumer delete it once processing succeeds, 3. or does a separate lifecycle policy on the storage bucket handle expiry independently of both? Each choice has different failure implications, covered in more depth by the pattern's lifecycle and cleanup concerns. ## A worked example A concrete worked example makes the ordering concrete: a video-upload service accepts a 200 MB file from a user, streams it to an `S3` bucket, and only after S3 confirms the multipart upload is complete does it publish a message like `{'videoId': 'abc123', 'bucket': 'raw-uploads', 'key': 'abc123.mp4', 'sizeBytes': 209715200, 'sha256': '...'}` to an `SQS` queue. A transcoding worker polls that queue, receives the message, calls S3's `GetObject` with the bucket and key from the message, streams the video into its transcoder, and on success publishes its own downstream claim check pointing at the transcoded output. Note that the transcoding worker never once receives video bytes through SQS -- SQS carries only the roughly 150-byte JSON reference, while S3 does all the heavy lifting of storing and serving the actual 200 MB of data. ## The reference schema is the contract This division of labor is exactly why the pattern exists: each system does the job it's built for, and the coupling between them is a small, well-defined, easily-versioned reference schema rather than the raw bytes themselves. Because that schema is the only contract between producer and consumer, it is worth treating it with the same care as any other API: - adding new optional fields over time is safe; - but renaming or repurposing an existing field can silently break every consumer still expecting the old shape, which is a maintenance concern that a purely inline payload design would not have surfaced in the same way.
- What happens if the consumer crashes after fetching the payload from storage but before finishing processing it?This depends on the broker's acknowledgment model, same as any other message. If the consumer hasn't acknowledged the reference message yet, most brokers will redeliver it, and the consumer simply re-fetches the payload from storage and retries -- storage being immutable and independently addressable makes this safe and idempotent as long as processing itself is idempotent.
- Should the claim check message include the full payload's checksum, and why does that matter operationally?Yes, it's good practice: it lets the consumer detect silent corruption or a stale/overwritten object at the storage layer without having to parse or fully process a bad payload first. Without it, corruption surfaces as a confusing downstream processing failure instead of a clear, fast validation error.
- Who should own deleting the payload from storage once processing is done -- producer or consumer?There's no universal answer: single-consumer pipelines often let the consumer delete on success, but multi-consumer fan-out pipelines can't, because deleting after the first consumer finishes would break the others still processing. In fan-out scenarios a time-based lifecycle policy on the storage bucket, independent of any consumer, is usually safer.
Like a warehouse shipping process: the warehouse (storage) physically receives and shelves the pallet first and confirms it's stored; only then does dispatch send the delivery slip (the message) with the shelf location to the driver (consumer), who goes and picks up the actual pallet from the warehouse rather than the slip carrying the goods.
saying these in an interview costs you the question
- Has the producer publish the reference before confirming the storage write succeeded
- Thinks the metadata/checksum in the claim check is optional decoration rather than a validation tool
- Believes the broker is involved in the payload fetch step
- Can't explain that the consumer's fetch is a separate network call to a separate system
- Assumes cleanup ownership is automatically obvious without stating who's responsible