skip to content

questions

4

Why does the order a connected test dataset is generated in decide whether its references resolve?

level: juniorimportance: must knowfreq 55%

answer

  1. Rows point at other rows
  2. Something must exist before it is referenced
  3. Sort the dependency graph, generate level by level
  4. Who assigns the identifier forces the order
  5. Mutual references need a second pass

basics

~20 s

A child row can only point at a parent that already exists and whose identifier is known. Generating in dependency order - parents first, then the rows referencing them - makes every reference resolve to a real row.

solid answer

~50 s

Treat the dataset as a graph of dependencies and produce it in **topological order**: an entity is generated only after everything it points at. The other half is identifier allocation. If the generator invents identifiers itself, it can hand a parent's identifier to its children immediately; if the store assigns them at write time, the generator has to write the parent, read the assigned value back, and carry it into the children - so the allocation scheme, not just the schema, forces the order. Mutual references cannot be sorted topologically at all: write one side with the reference empty and fill it in a second pass, or defer constraint checking to the end of the load. Get it wrong and the load either stops on a rejected write or, worse, succeeds with references pointing at rows nobody generated.

code

pseudocode · 18 lines
pseudocode
plan = newGraph()

// 1. build the whole set in memory, parents before children
for i in 1..200:
    customer = plan.add("customer", { id: nextId("cust") })
    for j in 1..fanOut():
        order = plan.add("order", { id: nextId("ord"), customerId: customer.id })
        for k in 1..lineCount():
            plan.add("orderLine", { id: nextId("line"), orderId: order.id })

// 2. write in dependency order
for entity in topologicalOrder(plan):
    store.writeAll(plan.rowsOf(entity))

// 3. prove every reference resolved, as part of the build
for ref in plan.references():
    dangling = countRows(ref.child, where: ref.value not in plan.idsOf(ref.parent))
    failBuildUnless(dangling == 0, "dangling " + ref.name)

go deeper

for a junior

Be ready to say why a parent has to be generated before anything that references it, and what happens when it is not: either the write is rejected, or the dataset ends up holding references that point at nothing.

for a middle

Explain the mechanics - sorting the dependency graph and generating a level at a time - and how the allocation scheme changes the job when the store assigns identifiers only at write time, including the read-back and the batching that keeps it fast.

for a senior

An interviewer at this level expects the failure you engineer against: the silent one, where nothing rejects the write and every later query quietly sees fewer rows. Describe the post-load check that catches it during the build rather than in a test.

for a principal

Own the policy: whether shared generated datasets use identifiers the generator invents, so builds are reproducible and a row can be named in a defect report, or identifiers the stores assign, and what each choice costs the loader and the teams that share the set.

## What "resolving" means in a manufactured dataset A connected dataset is a set of rows that point at each other: an `order` carries a `customerId`, an `orderLine` carries an `orderId`, a `shipment` carries both. A reference **resolves** when the value it holds identifies a row that actually exists in the set. Manufacturing a connected set is a different job from manufacturing one row at a time, because the value a child needs is produced by generating its parent, and until the parent has been produced that value does not exist. Two things decide whether every reference resolves: the **order entities are generated in**, and **who allocates the identifier**. ## Generating in dependency order Draw the entities as a directed graph with one edge per reference, pointing from the referring entity to the thing it depends on. Generating in **topological order** means producing each entity only after everything it points at. In a plain sales set that is customers, then orders, then order lines, then shipments — a level at a time rather than a row at a time, so the generator can hold a pool of parent identifiers and draw from it. Two structures defeat a plain sort: 1. **Cycles.** Two entities reference each other — an account naming its primary contact while the contact names its account. No single order satisfies both. Write one side with the reference left empty, write the other side, then return and fill the first in a second pass. If the store refuses the empty value, defer constraint checking until the whole load has been written and check once at the end. 2. **Self-references.** A row that points at another row of its own kind — a category's parent category, an employee's manager. The sort then has to run *inside* the entity: roots first, descendants after, with the depth chosen deliberately rather than left to recursion. ## Who allocates the identifier | Allocation | Who invents the value | What the generator must do | |---|---|---| | Generator-allocated | The generator, before anything is written | Nothing extra: a parent's identifier is known before the parent is written, so children can be built in the same pass | | Store-allocated | The store, at write time | Write the parent, read the assigned value back, hold it in a lookup, and only then build the children | | Derived from fields | A fixed rule over the entity's own values | Nothing extra, but the rule has to be collision-free across the whole set | The middle row is where most loaders both slow down and go wrong. Reading assigned identifiers back one row at a time turns a bulk load into a round trip per row; writing in batches and reading back a block of assigned values is the usual repair. Generator-allocated identifiers avoid the read-back entirely and make the whole set reproducible — the same inputs produce the same identifiers, so a failing case can name a specific row by value instead of describing it. ## What going wrong looks like - **Loud failure.** The store rejects the child because the row it references is not there. Annoying but cheap: the load stops at the mistake, and the message names it. - **Quiet failure.** Nothing enforces the constraint — the reference is a plain column, or the two sides live in different stores — so the load succeeds and the set now holds rows pointing at nothing. Every later query that combines the two sides silently sees fewer rows than it should, and a suite can go green against a set that is missing a third of its relationships. - **Reproducibility failure.** Identifiers come out different on every build, so a test that names one works today and fails after the next rebuild, with no change to blame. The quiet failure is the one worth engineering against, because it never announces itself. The cheap defence is a **post-load check that belongs to building the dataset, not to any test**: for every reference in the set, count the rows whose target is absent, and fail the build unless every count is zero. It is a handful of counts, it runs once per build, and it converts a whole class of silent wrongness into an immediate build error that names the offending entity. ## Practical shape Keep planning separate from writing. Build the entire object graph in memory first — entities, their identifiers, their references — and only then write it out in dependency order. Building in memory lets the generator sort, count and validate before anything is committed, and it makes the dependency order an explicit, reviewable property of the generator instead of an accident of the order somebody happened to write the code in. It also makes the two-pass repair for cycles trivial, because both sides already exist as objects before either is written. A last practical note: the order that matters is the order of *writes*, not the order of the code. A generator that builds children eagerly inside a loop over parents is still correct, as long as nothing is committed until the graph is complete and sorted.

  • The store assigns identifiers only at write time. How does that change how the generator builds a connected set?
    The parent's identifier does not exist until the parent is written, so children cannot be built in the same pass. The generator writes one level, reads the assigned values back, holds them in a lookup keyed by whatever it used internally to identify each parent, and builds the next level from that lookup. Read the values back in batches rather than one row at a time, or a bulk load becomes a round trip per row.
  • Two entities reference each other, so no generation order satisfies both. What do you do?
    Break the cycle across two passes: write one side with the reference left empty, write the other side, then update the first with the value that now exists. If the store will not accept the empty value, defer constraint checking until the whole batch is written and validate once at the end. Either way, record which reference was filled late, so a failed load can be diagnosed rather than guessed at.
  • How do you prove a generated dataset holds no reference pointing at a row that was never produced?
    Make it part of building the dataset rather than part of a test. After the load, count for each reference the rows whose target is absent, and fail the build unless every count is zero. It is a handful of counts, it runs once per build, and it turns a silent shortage of rows in every later query into an immediate build failure that names the offender.

It is flat-pack furniture: you cannot fix the shelf until the frame is standing, and the frame is what tells you where the holes are.

saying these in an interview costs you the question

  • Assumes the store's constraints will catch every dangling reference
  • Generates children before parents and retries until it works
  • Expects identifiers to be identical on every rebuild without allocating them
  • Calls mutual references impossible instead of writing them in two passes
  • Reads each assigned identifier back one row at a time
open as a page

What breaks in a generated dataset whose records all carry the same creation timestamp?

level: middleimportance: should knowfreq 52%

basics

~20 s

Anything that filters, sorts or groups by time. Every range query returns all rows or none, ageing paths never fire, ordering ties are decided by the store, and the causal orderings the product depends on are neither exercised nor violated.

open as a page

When several services each hold part of one generated entity, what must agree across their datasets?

level: seniorimportance: should knowfreq 38%

basics

~10 s

The value used to join the parts, presence on every side, the lifecycle state, and every attribute more than one service copies. Generate all views from one description, then reconcile a sample after loading.

open as a page

What does a generated dataset with exactly two children per parent fail to exercise?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Everything at the ends of the distribution: parents with no children at all, and the rare parent with thousands. A flat two-per-parent set never reaches empty-list handling, paging past the first page, chunking, truncation or tie-heavy ordering.

open as a page