skip to content

Companion Service Dependencies

The services a cluster is useless without, such as the store holding its message contracts, and what becomes of them in a restore or a switch. Asked because half a plan restores nothing.

part ofBroker & streaming operationsoverview, primer and where to startread it →
on this pageshow

questions

3

A cluster's brokers are all healthy yet every producer's publish fails - which service beside the cluster do you suspect, and how do you confirm it?

level: middleimportance: must knowfreq 50%

answer

  1. green dashboards measure only the brokers
  2. the failure is on the write side
  3. something else sits on the publish path
  4. the contract store is the usual culprit

basics

~20 s

Suspect a companion service on the publish path - most often the contract store a producer must reach before it may publish. Broker metrics cannot see it, so the cluster reads green while every write is refused.

solid answer

~50 s

Producers depend on more than the brokers. The usual extra is a **contract store** - the service a producer reaches to obtain or register the contract its payload must satisfy - and nothing about it is measured on the brokers, so the cluster reads green. The signature is distinctive: every stream fails at once, including unrelated teams'; the error is raised client-side, often before any record is handed to a broker; and readers often keep working for a while, because they already hold what they need and consult the service less often than a writer does. A producer that has just been restarted fails first, since a restart discards what it held. Confirm by checking the blast radius, where the error is raised, and what the client is configured to reach besides the cluster. Not every platform has such a service - say which model you mean.

go deeper

for a junior

Know that publishing can fail for reasons outside the brokers. If every producer fails at once while the cluster looks healthy, say so and start asking what else the client has to reach.

for a middle

Explain the mechanics: which companion services sit on the publish path, why broker metrics cannot see them, and why the read side often keeps working while every write is refused.

for a senior

Show the diagnosis: blast radius across unrelated streams, where the error is raised, the read/write asymmetry, and the restarted producer that fails while its untouched neighbour keeps going.

for a principal

Talk about ownership. A companion service typically has fewer copies, a quieter alert and no place on the rota, yet it gates every write to the cluster - that imbalance is a choice somebody should be asked to justify.

## What "beside the cluster" means Publishing a record looks like one hop - client to broker - and on some estates it is exactly that. On many it is not. A producer may have to reach one or more **companion services** first: separate deployments, with their own addresses, their own storage and their own availability, that stand beside the cluster and are not part of it. The brokers do not monitor them, do not report on them and do not fail when they fail. The most common one is the **contract store** (often called a schema registry in lower-case ordinary usage): the service that holds the message contracts a producer must obtain or register before it may hand a payload to the cluster. That is why this failure has such a distinctive shape. Almost everything the operations team looks at is a broker measurement, and every broker measurement is fine. ## The signature - every broker is up, accepting connections and serving reads; - copy sets are complete; nothing is short of its copies; - disk, network and request latency on the cluster are ordinary; - publish failures appear on **every** stream at once, including streams owned by teams that share nothing but the platform; - the error is raised on the client side, often before a record is ever handed to a broker; - restarting a producer tends to make things worse rather than better, because a restart discards whatever the client had already obtained. That last point is the tell that separates this from a cluster problem: a process that had been publishing happily for hours starts failing the moment it is restarted, while its untouched neighbours keep going a little longer. ## Confirming it during an incident 1. **Check the blast radius.** Unrelated streams and unrelated teams failing at the same moment, with healthy brokers, points outside the cluster to something shared. 2. **Check where the error is raised.** A failure that never produced a connection attempt to a broker is not a broker failure, whatever the ticket says. 3. **Check the read side.** Readers commonly keep working for a while, because they often already hold what they need and consult the companion service less often than a writer does. An asymmetry between a dead write path and a live read path is strong evidence. 4. **Check what the client is configured to reach.** A producer's configuration usually names more than one address. Anything in it that is not a broker is a candidate. ## The companions that produce this shape | dependency | what a client needs it for | what its loss looks like | |---|---|---| | the contract store | obtaining or registering the contract a payload must satisfy | publishes refused at the client while the cluster sits idle and healthy | | an external coordination service, on designs that keep cluster metadata outside the brokers | the cluster's own bookkeeping | broader than publishing: administrative work and sometimes leadership stall too | | the service that answers whether a client may connect at all | admission | connections refused rather than publishes refused - a different shape, worth telling apart | ## What varies between platforms This is the part a candidate has to get right, because it decides whether the answer travels: - **Some platforms ship a contract store alongside the cluster; others have none.** Where none exists, the contract travels with the payload or lives in a build-time artifact shared by both sides, and this dependency is simply not present. - **Some designs keep cluster metadata in a separate coordination service; others keep it inside the cluster itself.** The first has a companion whose loss stops far more than publishing; the second does not have that companion at all. - **Clients differ in what they do when the companion is unreachable.** Some hold what they last obtained and keep publishing until they are restarted; others refuse immediately. The same outage therefore looks gradual on one estate and instantaneous on another. Say which model you are describing and the answer is right everywhere; assert one model flatly and it is wrong somewhere. ## Why an operator cares beyond the incident A companion service is the part of a messaging estate that nobody owns on the operations rota, because it is not the cluster. It usually has fewer copies than the cluster, a smaller machine, a quieter alert, and no line in the continuity plan. The incident above is the cheap version of that neglect: an hour of refused writes, a confusing dashboard, and a post-incident action to monitor the thing. The expensive version arrives when the cluster has to be restored or switched to a standby, and the plan covers the brokers only - at which point the restore finishes, every broker is healthy, and not one producer can publish.

  • Why might readers keep working while every producer is blocked?
    Because the two paths consult a companion service at different rates. A reader often already holds what it needs and can run on that for a long time, while a producer's next publish needs the lookup now. Some clients also hold what they obtained until they are restarted, which is why a restarted producer fails first.
  • How would you make this failure visible before a producer reports it?
    Monitor the companion service as a first-class dependency: its availability and latency measured from where a client sits rather than from the service itself, plus a synthetic publish that exercises the whole path end to end. Broker health will not move when it fails, so it cannot be the signal.

saying these in an interview costs you the question

  • Claims green broker dashboards prove the whole messaging path is working.
  • Assumes only the brokers have to be reachable for a producer to publish.
  • Blames the network to the brokers without checking the read side is fine.
  • Insists readers must fail whenever producers do.
  • Treats a companion service as unimportant because it holds no records.
open as a page

Where does the contract store beside a cluster keep its own state, and why does that decide whether a restore of the brokers is usable?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A contract store keeps state of its own, and where it lives decides the restore: on the cluster it serves, in another team's database, or in a checked-in definition. Only the last comes back for free.

open as a page

Which services standing beside your clusters belong in the continuity plan, and what do you do about the ones you exclude?

level: principalimportance: should knowfreq 32%

basics

~20 s

Inventory everything a publish and a read touch besides the brokers, then give each companion a posture: one shared instance, a copy per site, or rebuild on demand. Anything excluded should be excluded on purpose, in writing.

open as a page