A social app is built as several microservices — posts, comments, notifications — each backed by its own eventually consistent store. A user reads a post, then writes a comment on it, expecting that anyone who can see their comment can also see the post it refers to. What has to happen at the infrastructure level to actually guarantee this writes-follow-reads property across service boundaries, and where does it typically break?
answer
- causal token must survive every service hop
- vector clock entry per read value
- async/event bus often drops custom metadata
- CDN/cache has no version awareness
- scope enforcement narrowly, not everywhere
basics
~20 sEvery service in the chain needs to pass along a marker, like a version or vector clock entry, of what the user has already read, and the comment service must attach that marker to the write so downstream readers of the comment are forced to also be caught up on the post before they see it.
solid answer
~50 sWrites-follow-reads across service boundaries requires propagating causal metadata — typically a vector clock entry, a monotonic version or log sequence number for the read post, or an opaque session token — from the read response all the way through to the subsequent write request, with every service in the chain forwarding it rather than dropping it at a service-to-service hop. The comment service then stores that dependency alongside the comment, and any reader of the comment must be served from a replica or service state that is at least as fresh as that dependency, or the comment must be withheld until the dependency is satisfiable. In practice this breaks most often at service boundaries where request context isn't propagated through an internal RPC call, an async queue or event bus strips custom metadata, or a caching layer serves the comment from a CDN edge with no way to check the dependency at all.
go deeper
Not expected to reason about microservice propagation; can restate the goal in plain language, that a comment shouldn't outrun its post.
Should recognize this requires passing some marker along and can name at least one place it might get dropped, such as an internal service call.
Should walk through the propagation chain end to end, from client to gateway to service to async bus to downstream consumer, and name multiple concrete failure seams.
Should discuss the cost and scope trade-off of enforcing this guarantee narrowly versus everywhere, and connect it to adjacent solved problems like distributed tracing context propagation.
## What the guarantee says Writes-follow-reads, or session causality, says that if a client reads value `X` and then performs a write `Y`, that write is causally dependent on `X`, and any observer of `Y` must also be able to observe `X` or a version at least as new. Within a single service backed by a single data store, this is a comparatively contained problem — the same version-vector or LSN scheme covers all data. In a microservice architecture, though, the read and the dependent write cross service and often data-store boundaries entirely, and a guarantee each individual service offers internally for its own data doesn't automatically compose across those boundaries. ## The chain that has to hold end to end Concretely, for the post, comment, and notification scenario: 1. The user's client calls the posts service, which returns the post along with a **causal marker** — a version number, an LSN, or an entry in a vector clock keyed by post ID. 2. For writes-follow-reads to hold end to end, the client, or more robustly an API gateway or backend-for-frontend layer acting on the client's behalf, must attach that marker to the subsequent write request to the comments service. 3. The comments service must then persist that marker as a causal dependency alongside the new comment record — not just the comment text, but a record that this comment depends on the post being at or beyond that version. 4. Any service or client that later reads that comment must enforce the dependency before exposing it: by checking the posts service or its own cached copy of post state is caught up, by proactively fetching the post fresh before returning the comment, or, in an event-driven model, by attaching the dependency to any downstream event, such as a comment-created event for the notifications service, so every consumer that eventually surfaces the comment can re-validate or wait on the dependency. ## The seams where it breaks This is genuinely expensive to implement correctly, which is precisely why it breaks down at several common seams. - **First, request-context propagation:** many internal RPC frameworks and HTTP clients don't automatically forward custom headers or metadata across service-to-service hops unless every team explicitly wires it in; a causal token attached to the client's original request can silently get dropped the moment it passes through an internal call the comments service makes back to the posts service, or through a shared library unaware of this convention. - **Second, async messaging:** if the comments service publishes an event onto a message bus for the notifications service to consume, the event schema has to explicitly carry the causal dependency; if it's added as an afterthought, older event versions or unrelated consumers silently have no way to honor it. - **Third, caching and CDNs:** if comments are served through an edge cache or read-replica fan-out with no concept of per-item causal dependencies, there is no hook at all for enforcing the dependency at serve time — the cache just returns whatever bytes it has, with no version awareness. - **Fourth, cost and blocking:** even where the mechanism is wired correctly, enforcing block-or-fetch-fresh on every read adds latency, and under partition or heavy replication lag can turn into unbounded waiting, which teams often quietly relax rather than enforce everywhere. ## Why teams scope it narrowly Because of this cost, most production systems don't implement full writes-follow-reads end to end across every service boundary; instead they scope it narrowly to interactions that would be visibly broken if violated, such as a comment thread rendering with a missing parent post, and accept looser guarantees elsewhere, such as push notifications arriving slightly before the referenced content is globally visible being treated as an acceptable, low-stakes artifact. ## Plumbing borrowed from tracing A concrete real-world analog: distributed tracing systems, such as OpenTelemetry's context propagation via trace headers, solve a structurally similar propagation problem — carrying causal or contextual metadata across every service hop, including through async queues — and teams building writes-follow-reads guarantees across microservices often reuse the same context-propagation plumbing, headers through synchronous calls and message attributes through async buses, rather than inventing a separate mechanism, precisely because ensuring propagation survives every hop, including async ones, is already a solved and tested problem in that adjacent space.
- Would using a single shared database across the posts and comments services eliminate this problem?It would eliminate the cross-service propagation problem specifically, since a single store's own internal version or vector-clock scheme could cover both post and comment reads and writes uniformly. It doesn't eliminate the broader architectural trade-off, though — that essentially reverts to a shared-database, monolith-style design specifically to avoid solving cross-service causal propagation, sacrificing the independent scaling and deployment benefits microservices were adopted for.
- Why is dropping the causal token in an async message queue a particularly dangerous failure mode compared to dropping it in a synchronous call?A synchronous call failing to propagate the token typically fails fast and visibly, since the immediate response is wrong and easy to notice in testing, whereas an async event silently missing the dependency can propagate a causally-broken state to many downstream consumers, such as search indexes or notification services, before anyone notices, and by the time it's caught the incorrect state may already be cached or surfaced to end users at scale.
- If enforcing the dependency at read time is too expensive, what's a cheaper alternative to still avoid obviously broken UX?A common pragmatic compromise is to enforce the dependency only at write or publish time — refusing to accept or index the comment write until the comments service itself can confirm the post is visible in its own read path — rather than re-checking on every subsequent read, shifting the cost to a one-time check at the expense of a slightly weaker end-to-end guarantee if the post later becomes unavailable through some other path.
It's like passing a chain-of-custody tag through every hand a package goes through — if even one courier in the relay forgets to re-attach the tag, the final recipient has no way to know the package causally depended on an earlier delivery, and the whole guarantee silently evaporates at that one dropped handoff.
saying these in an interview costs you the question
- Assumes each microservice offering strong consistency internally is sufficient for cross-service writes-follow-reads
- Doesn't mention that async or event-bus hops are a common place the guarantee breaks
- Thinks this can be solved with retries or timeouts alone rather than explicit causal metadata
- No mention of caching or CDN layers as a place where causal checks can't be enforced
- Proposes enforcing this guarantee everywhere with no discussion of latency or cost trade-off