skip to content

questions

6

A message queue advertises a 256 KB maximum message size, but a service needs to move a 50 MB video file through an event-driven pipeline. What is the standard fix, and why don't teams just send the file directly in the message body?

level: juniorimportance: must knowfreq 72%

answer

  1. coat-check ticket
  2. broker = small messages only
  3. payload in blob store, pointer on bus
  4. S3 + SQS combo
  5. write payload before publishing reference

basics

~10 s

Save the big file to storage like S3, then send only a small pointer (its location or ID) through the queue. The receiver downloads the real file from storage when it needs it.

solid answer

~50 s

This is the Claim Check pattern. Brokers are built for high-throughput delivery of many small messages, not bulk bytes -- most enforce hard size caps (SQS is 256 KB) and even uncapped brokers like Kafka degrade once records get large, because size cost scales across every replica and every fanned-out consumer. Instead of pushing the payload through the bus, the producer writes it to external storage (an object store, blob store, or cache) and publishes a tiny message containing just a reference -- a claim check -- such as an object key or URL. Consumers read the reference and fetch the payload directly from storage. The broker stays fast and cheap doing what it is good at; storage absorbs the bulk bytes. The cost is an extra network hop and a second dependency whose availability and consistency with the broker now matter.

go deeper

for a junior

Should recognize the basic shape of the fix -- store the big thing externally, send a small reference through the queue -- and be able to name a concrete real pairing such as S3 plus a queue.

for a middle

Should articulate specifically why brokers suffer with large payloads (memory pressure, replication cost, per-consumer fan-out cost) and clearly separate the producer's write-then-publish role from the consumer's read-then-fetch role.

for a senior

Should flag the ordering/consistency hazard (payload must be durably written before the reference is visible) and discuss what happens operationally when storage and broker fall out of sync.

for a principal

Should reason about when the second dependency and extra hop are not worth it, and compare the pattern against alternatives such as compression, chunking, or simply picking a broker whose size limits fit the workload.

## What a broker is built for Message brokers -- the queues or pub/sub topics that sit between producers and consumers in an event-driven system -- are engineered for one job: moving a very large number of small messages quickly, reliably, and with predictable delivery guarantees. To do that well, brokers typically: - keep messages in fast, often memory-backed storage with disk as a durability backstop; - **replicate** every message across multiple nodes for fault tolerance; - frequently deliver the same message to several independent consumers through **fan-out**. All three of those mechanisms scale with message size: a message ten times larger costs roughly ten times the memory footprint, ten times the replication bandwidth between broker nodes, and ten times the network cost per additional subscriber it is delivered to. ## Why brokers cap message size Because of this, almost every managed broker enforces a hard size cap -- Amazon `SQS` tops out at **256 KB** per message, for instance -- and brokers without a hard cap, like self-managed `Kafka` or `RabbitMQ`, still degrade badly once messages routinely run into the megabytes, because large records: - defeat batching; - blow through network and disk buffers; - slow down replication and log compaction for every other message sharing that broker, not just the large one. ## Splitting the payload from its notification The **Claim Check** pattern resolves the mismatch between 'I have a 50 MB file to move' and 'my broker is only good at moving small events' by splitting the payload from its notification. 1. The producer first writes the large payload to a storage system built for bulk bytes rather than message throughput -- commonly an object store such as Amazon `S3`, Azure Blob Storage, or Google Cloud Storage, though a distributed cache or shared filesystem can play the same role. 2. That write returns, or lets the producer construct, a **reference**: an object key, a URL, or some other identifier the payload can later be located by. 3. The producer then publishes a small message onto the broker carrying that reference instead of the payload itself -- the claim check. 4. A downstream consumer reads the reference off the bus and only then makes a separate call to the storage system to fetch the real payload, using it to complete the unit of work. ## Where the name comes from The pattern earns its name from the coat-check analogy: you do not carry your winter coat to your seat in a theater; you hand it to an attendant and receive a small paper ticket. The ticket is cheap to carry, verify, and pass around; the heavy coat sits in a room built to hold coats. Redeeming the ticket later retrieves the coat. Here, the ticket is the reference message on the bus, and the coat room is the object store. ## The core trade-off The core trade-off is that a single dependency (the broker) becomes two, and a single hop becomes two hops. Where a broker-only design has one moving part whose availability governs the whole pipeline, a claim-check design now depends on: - the broker delivering the reference, - the storage system serving the payload, - and the two staying consistent with each other -- the reference must never become visible to consumers before the payload it points to is durably written, or a consumer will fetch and get a not-found error. That ordering rule, **write payload then publish reference and never the reverse**, is the single most common correctness bug in claim-check implementations, and it follows directly from the fact that the storage write and the broker publish are two independent, non-transactional operations that nothing forces to happen atomically. ## Where it shows up in production In production this shows up constantly: - an image or video pipeline uploads a raw file to S3 and then emits a small 'uploaded, key=abc123' event that fans out to transcoding, thumbnailing, and moderation consumers, each fetching the file from S3 directly rather than each receiving a multi-hundred-megabyte copy through the bus; - a document-processing system drops a PDF into blob storage and triggers OCR off a small event referencing it; - an ETL pipeline moves large batch extracts between stages the same way. The payoff is that the broker stays small, fast, and cheap to operate at the throughput profile it was designed for, while an object store -- cheap per gigabyte and built for bulk storage rather than low-latency fan-out -- absorbs the bytes, and the two systems can be scaled and priced independently of each other.

  • Why not just raise the broker's max message size instead of adding external storage?
    Some brokers let you raise the cap, but the underlying cost doesn't go away: larger messages still consume more broker memory, slow down replication between broker nodes, and multiply network cost on every fan-out to a subscriber. A managed broker like SQS gives you no dial to turn at all -- 256 KB is a hard limit -- so for genuinely large payloads external storage is the only option, not just the cheaper one.
  • What should actually go inside the claim check message itself?
    Keep it minimal: the object key or URL, and enough metadata to fetch and validate it safely -- content type, size, and a checksum are common additions. Avoid duplicating business data that belongs in the payload; the reference's only job is to locate and verify the real object.
  • Does using a claim check change what delivery guarantee the consumer gets?
    The broker still guarantees delivery of the small reference message under its normal semantics, such as at-least-once delivery. But it says nothing about the payload's durability or availability -- that responsibility moves entirely to the storage system, so the consumer now has to handle the case where the reference arrives but the payload isn't fetchable yet or has already been removed.

Like checking your coat at a theater: you don't carry the heavy coat to your seat, you get a small ticket, and you hand back the ticket later to retrieve the coat from the coat room.

saying these in an interview costs you the question

  • Suggests just raising the broker's message-size limit and stopping there
  • Doesn't mention that storage becomes a second dependency with its own availability
  • Assumes the queue durably stores the payload bytes itself
  • Can't explain why large messages specifically hurt broker replication and fan-out throughput
  • Doesn't realize the payload must be written before the reference is published

context

open as a page

Walk through the claim check pattern step by step for a pipeline where a producer emits a large report and a consumer needs to process it: what exactly happens, in what order, and what does each side own?

level: middleimportance: must knowfreq 65%

basics

~20 s

Producer saves the big report to storage first, gets back a location, then sends a small message with that location through the queue. Consumer reads the message, fetches the report from storage using the location, then processes it.

open as a page

A claim-check pipeline has been running for six months and the object storage bucket backing it has grown to 40 TB, far more than the traffic volume would suggest. What's the likely root cause, and how would you diagnose and fix it?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Old files in storage are probably never getting deleted after being processed -- orphaned blobs pile up forever. Fix by adding a cleanup step or an automatic expiry rule on the storage bucket so old, unused files get removed.

open as a page

What does a team actually give up by adopting the claim check pattern for a pipeline that previously sent every payload inline on the message bus?

level: middleimportance: should knowfreq 55%

basics

~20 s

It gets slower per message (extra round trip to fetch the payload) and more complex (a second system to run, pay for, and secure). In exchange, the queue stays fast and cheap and can handle way more messages per second.

open as a page

A claim-check pipeline stores customer documents in a shared object storage bucket and passes references through a queue. What access-control mistakes commonly show up here, and how should read access to the payload actually be scoped?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The common mistake is making the storage bucket wide-open or giving every consumer permanent access to everything in it. Better: give each consumer only the narrow permission it needs, ideally a short-lived link scoped to just that one file.

open as a page

You're reviewing a proposed architecture where every message on an internal event bus, regardless of size, gets routed through claim check by default: producers always write to object storage first and always publish just a reference, even for a 500-byte status update. What's wrong with this default, and when does claim check genuinely not pay for itself?

level: principalimportance: should knowfreq 38%

basics

~20 s

Using claim check for every message, even tiny ones, adds an extra network call and a second system for no benefit -- small payloads fit on the bus fine. It's worth it only when payloads are large.

open as a page