skip to content

When several services each hold part of one generated entity, what must agree across their datasets?

level: seniorimportance: should knowfreq 38%

answer

  1. One entity, several stores
  2. The parts must describe the same thing
  3. Same join value, same state, same copies
  4. Project one description into per-service views
  5. Reconcile a sample as part of the build

basics

~10 s

The value used to join the parts, presence on every side, the lifecycle state, and every attribute more than one service copies. Generate all views from one description, then reconcile a sample after loading.

solid answer

~50 s

Four things. The **join value** - the identifier one service stores to refer to an entity another owns - must be identical, not merely similar in format. **Existence** must be symmetric: no service may hold a reference to an entity another never produced, in either direction. **Lifecycle state** must correspond to one point in the entity's life, so one side is not holding an active customer the other has closed. And the **copied attributes** each service keeps for its own use - the display name, the tier, the currency - must carry the same values, because a flow reads whichever copy is nearest and will otherwise behave differently depending on the route taken. The way to get all four is to generate the entity once as a single description and project a per-service view from it, then reconcile a sample after loading and fail the environment build.

code

pseudocode · 19 lines
pseudocode
// one authoritative description, generated once
entity = {
    key:      "CUST-000417",
    name:     "Aurelia Vance",
    tier:     "GOLD",
    state:    "ACTIVE",
    currency: "EUR"
}

// every service gets a projection of the same description
accountsView = { id: entity.key, fullName: entity.name, status: entity.state }
billingView  = { customerRef: entity.key, tier: entity.tier, currency: entity.currency }
supportView  = { customerRef: entity.key, displayName: entity.name, status: entity.state }

// reconcile a sample after loading, before any test runs
for e in sample(entities, 200):
    failBuildUnless(billing.read(e.key) exists, "missing in billing: " + e.key)
    failBuildUnless(accounts.read(e.key).status == support.read(e.key).status, e.key)
    failBuildUnless(accounts.read(e.key).fullName == support.read(e.key).displayName, e.key)

go deeper

for a junior

Know that one business entity can be split across several services, each storing its own part, and that generated test data has to describe the same entity on every side or any flow crossing those services will not find it.

for a middle

Explain what must match: the value used to join the parts, presence on both sides, the lifecycle state, and any attribute more than one service copies. Be able to say why generating each side independently does not guarantee any of it.

for a senior

An interviewer expects the discipline: generate one description, project each service's view from it, then reconcile a sample after loading and fail the environment build. Expect to describe how the mismatch presents - as a spurious defect against the product.

for a principal

Own the boundary: one shared generated dataset for all services, or each team manufacturing its own with only agreed entities aligned. Both are defensible; the choice turns on how many flows cross services and on who maintains the alignment when a view changes.

## One entity, several owners Split a system into services and business entities get split with them. A customer exists as an account in one service, a billing party in another, a contact in a third. Each service holds its own slice, and each slice carries a value that ties it to the others plus a few attributes it copied so it does not have to ask. Manufacturing test data for that arrangement is not the same as manufacturing it for one store. Each service can be filled correctly and independently, and the estate — the collection of environments and datasets the team tests against — can still be wrong, because *correct on its own* and *consistent with the others* are different properties. ## What has to agree | What must agree | The failure when it does not | Where it surfaces | |---|---|---| | The join value | One side stores a padded, trimmed or differently-cased variant | A lookup returns nothing; a combining step silently drops the row | | Existence, in both directions | One service holds a reference to an entity another never produced | A not-found at the first hop, or an empty result that reads as "no data yet" | | Lifecycle state | One side active, the other closed or removed | Behaviour depends on which service is asked; failures are intermittent and entity-specific | | Copied attributes | Two spellings of a name, two tiers, two currencies | A figure reconciles differently depending on the route; a rule fires on one path only | | Stored totals | An aggregate one service holds does not match the rows another holds | An assertion on a total passes at the wrong number | The copied-attribute row is the sly one. Denormalised copies exist precisely so a service can answer without a call, which means a flow reads whichever copy is nearest — and a set where the two copies disagree produces behaviour that changes with the route taken rather than with the input. ## Generate once, project per service The reliable construction is to manufacture the entity **once**, as a single description, and derive each service's view from it: 1. Build one in-memory description per entity: its key, its lifecycle state, and every attribute more than one service will hold. 2. Derive each service's view from that description by projection. Never run each service's generator separately and rely on matching configuration. 3. Derive identifiers from the entity's own key by a fixed rule, so a rebuild produces the same values and a partial reload cannot invent new ones. 4. Write each service's slice in dependency order, and record which entities were written where. 5. Reconcile a sample after loading, before any test runs. Step 2 deserves its reason. Two generators configured identically produce identical values only while both consume their random source in exactly the same order. Add a field to one service's view, reorder two lines, or fix a bug in one of them, and every value after that point shifts — silently, and only on that side. The coupling is invisible in review and it breaks on an unrelated change months later. A projection has no such coupling: there is one description, and the views are functions of it. ## What an inconsistency looks like when it surfaces It arrives as a defect filed against the product, which is what makes it expensive. A flow fails with a not-found; an engineer reads the service, reads its caller, adds logging, and eventually discovers that the dataset disagreed with itself. That is a day, and it recurs. The worse variant is silent. A combining step drops the rows whose other side is missing, a total comes out lower than it should, and the test asserting that total passes — because the expected figure was computed from the same inconsistent set. Nothing fails; the suite is simply measuring a smaller world than it thinks it is. ## Reconcile as part of building the environment The check is small and it belongs to the build, not to a test: - For a sample of entities, read the view back from every service that should hold one. - Assert presence on every side, the same join value character for character, the same lifecycle state, and equality of each attribute that is deliberately duplicated. - Fail the environment build with the entity's key in the message. Sampling is enough — inconsistency produced by a generator is systematic rather than sporadic, so a couple of hundred entities will find it. What this check does not cover is whether the services can still parse each other's messages after a change; that is a separate concern about the shape of what is exchanged, not about whether the data agrees. One more habit worth adopting: print the count of entities written per service in the build output. A slice that came out with a different number of entities than its siblings is the cheapest possible signal that a projection silently dropped something.

  • Each service already has its own generator. Why not just run all of them with identical configuration?
    Because identical configuration produces identical values only while every generator consumes its random source in exactly the same order. Add a field to one service's view, reorder two lines, or fix a bug in one of them, and every value after that point shifts - silently, and only on that side. The coupling is invisible in review and breaks months later on an unrelated change. One description projected into several views has no such coupling.
  • How does a cross-service data inconsistency usually reach the team?
    As a defect filed against the product. A flow fails with a not-found, or a total comes out low, and somebody spends a day inside the service before suspecting the dataset. The worse variant is silent: a combining step drops the rows whose other side is missing and the assertion still passes, because the expected figure came from the same inconsistent set.

saying these in an interview costs you the question

  • Runs each service's generator separately and trusts matching configuration
  • Assumes only the identifier has to match across services
  • Ignores lifecycle state, so one side holds a closed entity
  • Treats a cross-service mismatch as a product defect
  • Checks consistency only after a test has already failed