How do you decide whether to enforce document shape in the database or in application code?
answer
- who else can write to this collection?
- not every write goes through your service
- one layer catches shape, the other meaning
- coarse contract below, semantics above
- the database check is a backstop, not a spec
basics
~20 sEnforce the coarse contract in the database — required fields, types, allowed discriminator values — because it catches every writer including scripts and migrations. Enforce semantics and cross-field business rules in application code. Use both layers; they cover different failure modes.
solid answer
~50 sThey protect against different things, so it is not either-or. A **database-side validator** is the only check that applies to *every* writer: a second service, a one-off maintenance script, a data-fix run from a shell, an import job. It should carry the coarse contract — the fields that must exist, their types, the closed set of values a discriminator may take — because those are exactly the invariants that must hold no matter who wrote the document. **Application code** is where expressive rules belong: conditional requirements per variant, cross-field consistency, anything needing another document or external state, and anything that must produce a good user-facing error message. It is also versioned, testable and reviewable with the feature. The failure mode of skipping the database layer is silent corruption from an unexpected writer; the failure mode of pushing everything into the database is unmaintainable rules that only fail at write time with an opaque error.
code
javascript · 9 lines// application layer: conditional, contextual, good errors
function validatePayment(doc, previousStatus) {
if (doc.type === "card" && !doc.details?.last4)
throw new ValidationError("details.last4", "required for card payments");
if (doc.status === "PAID" && previousStatus === "REFUNDED")
throw new ValidationError("status", "illegal transition");
}
// database layer holds only: type, amount, currency, createdAt required;
// amount numeric; type one of card | transfer | vouchergo deeper
Know that a document store can optionally check documents at write time, and that your application also validates input. Be able to say the two are not the same check.
Explain what each layer catches: the database sees every writer including scripts, while the application expresses conditional and cross-field rules and produces usable error messages.
Show the decision procedure — who can write here, is the rule pure shape, will it hold for every future variant — and describe how you keep the two layers from drifting: rules as reviewed code plus tests that assert rejection.
Own the framing that the database check is a backstop rather than a specification, and be able to justify how narrow that backstop should be given the number of independent writers and the cost of unwinding a rule later.
## Two layers, two different failure modes The question sounds like a choice and is really a layering decision. Ask what each layer catches that the other cannot. **The database-side check catches writers you did not write.** Application validation only runs when the write goes through the application. In any system older than a few months, that is not every write: there are maintenance scripts, backfills, an admin console, an import from a partner, a second service that grew its own write path, and an engineer fixing production data by hand at 2 a.m. A collection-level check is the only thing standing in front of all of those. **The application check catches meaning.** Whether an order's status transition is legal, whether the discount is allowed for this customer tier, whether the shipping address is required *because* the order is physical — these are conditional, contextual, and often depend on data outside the document. They also need to fail with a message a human can act on. A weak candidate picks one and defends it. A strong one describes the split and the reason for it. ## What belongs in the database layer Keep it coarse, stable and cheap: - **Required core fields.** The small set every document must carry — identifier, timestamps, owner, the discriminator for a polymorphic collection. - **Types.** A field that must be a number must not accept the string "12". Type drift is the most damaging kind of schema drift because it breaks sorting and comparison silently. - **Closed value sets.** Status and type fields whose values must come from a known list. - **Bounds where they are genuinely invariant.** A quantity that can never be negative. Why so narrow? Because a collection-level rule applies to the entire collection, including documents written years ago and by variants you have not designed yet. Every rule you add is a rule you must keep true forever, or unwind. A small required core has almost no maintenance cost and blocks the worst outcomes; an elaborate rule set becomes a second schema you must evolve in lockstep with the code, in a place your tests do not naturally reach. ## What the database layer cannot do Be explicit about the limits, because interviewers probe here: - **Anything involving another document.** "The order total must equal the sum of its line items stored elsewhere" is not a shape rule. - **Anything involving prior state.** "Status may go from PENDING to PAID but not back" requires knowing what the document said before. - **Anything involving external state.** Rates, entitlements, feature flags, another service's answer. - **Good error reporting.** A rejected write surfaces as a driver error; mapping it back to "the postcode is missing" for an end user is application work regardless. There are also practical costs: a check runs on every write and consumes some CPU; the rules live outside your application repository unless you deliberately manage them as code; and tightening a rule on a collection that already holds non-conforming documents is a project of its own rather than a config change. ## What belongs in the application layer Everything expressive, plus the user-facing contract: - Per-variant conditional requirements — a card payment needs its masked digits, a voucher needs its code. - Cross-field and cross-document consistency. - Normalisation: trimming, case-folding, unit conversion, and the absent-versus-null convention, applied in **one** place on the write path. - Rich, localised, field-level error messages. Because it is code, it is versioned with the feature, covered by unit tests, and reviewable in the same pull request as the change it protects. ## Making the two layers agree The classic operational trap is divergence: the application requires a field the database does not, and eventually a script writes documents the application will later choke on; or the database requires a field the application stopped sending, and a deploy starts failing writes. Two habits prevent it: 1. **Treat the database-side rules as code.** Keep them in the repository, apply them through the same reviewed, ordered mechanism as any other schema change, and never edit them ad hoc in a console. 2. **Test against the real check.** Integration tests that write a deliberately malformed document and assert it is rejected keep both layers honest and catch the day someone loosens one of them. ## Choosing in an interview scenario A good verbal decision procedure: *Who can write to this collection?* If the answer is more than one code path, you want a database-side check. *Is the rule expressible as "this field exists and has this type or one of these values"?* If yes, it belongs in the database. If it needs another document, prior state, or a good error message, it belongs in the application. *Will this rule still be true for every future variant?* If not, do not put it in a collection-wide check. The honest summary is that the database layer is a **backstop**, not a specification: it exists to make the worst corruption impossible, while the application layer expresses what the data actually means.
- Name a rule a collection-level shape check can never enforce, and say why.Anything depending on another document, on the document's prior state, or on external state. "The order total equals the sum of its line items stored elsewhere" and "status may not go from PAID back to PENDING" both need information the check does not have at write time — the first needs a second read, the second needs the old version of the document.
- Your application already validates everything. Why add a database-side check at all?Because application validation only runs on writes that go through the application. Maintenance scripts, backfills, imports, an admin console, a second service and a hand-typed fix in a shell all bypass it. The database check is the only layer that sees every writer, which is exactly the population that produces the corrupt documents nobody expected.
- How do you stop the two layers from drifting apart?Treat the database-side rules as code in the repository, applied through the same reviewed, ordered mechanism as any other schema change rather than edited by hand. Then add integration tests that write a deliberately malformed document and assert the database rejects it, so loosening either layer fails a test rather than going unnoticed until production.
- Why keep the database-side rule set deliberately small?A collection-level rule applies to every document, including old ones and variants not yet designed, and must stay true forever or be unwound. A small required core costs almost nothing to maintain and blocks the worst corruption; an elaborate rule set becomes a second schema evolving in lockstep with the code, in a place your unit tests do not reach.
saying these in an interview costs you the question
- Assumes every write goes through the application
- Puts business rules and state transitions into a collection-level check
- Treats database-side rules as a full specification of the model
- Edits the database's rules by hand outside version control
- Argues it is strictly either database or application, never both