In a document store, when should a field be absent versus stored as an explicit null?
answer
- there are more than two empty states
- one means the fact was never recorded
- the other means recorded as deliberately empty
- queries can tell the two apart
- consistency matters more than the choice
basics
~20 sAbsent usually means the fact was never recorded or does not apply; explicit null means it was recorded as deliberately empty. Document stores distinguish the two, so pick one convention per field and apply it everywhere rather than letting both appear.
solid answer
~50 sThe two carry different meanings and you should use that deliberately. **Absent** is the natural encoding for "not applicable to this kind of document" or "never captured" — and it costs nothing to store, which is why sparse optional fields are cheap in a document model. **Explicit null** is the encoding for "we asked and the answer is nothing", which is a real fact worth recording. The danger is not choosing: once a collection holds documents where a field is sometimes missing, sometimes `null` and sometimes an empty string, every reader grows a three-way fallback and every query has to be written twice. Document stores generally treat missing and null as distinguishable in matching, so the inconsistency is not cosmetic — it changes results. Decide the convention when you design the field, document it, and make the write path the only place that decision is made.
code
json · 3 lines{ "_id": 1, "name": "Ada" }
{ "_id": 2, "name": "Grace", "middleName": null }
{ "_id": 3, "name": "Alan", "middleName": "" }go deeper
Know that a document can simply omit a key, and that omitting it is not the same as storing null. Be able to give one example of each meaning.
Explain the three empty states and what each should mean, and note that queries distinguish presence from a null value so the choice changes results, not just aesthetics.
Show how you make a convention stick: one decision point on the write path, normalisation at the read boundary, and periodic sampling of stored documents to catch the drift early.
Own the rule that presence may vary but type must not, and be able to justify the cost of cleaning up a mixed-encoding field against leaving readers to branch forever.
## Three states, not two A field in a document can be in more than one "empty" state, and they are not interchangeable: - **Absent** — the key is not present in the document at all. - **Explicit null** — the key is present and its value is the null literal. - **A sentinel** — the key is present with an empty string, a zero, an empty array, or a magic value like `"UNKNOWN"`. Relational schemas collapse the first two: a column exists on every row, so the only way to say "nothing here" is NULL. A document model has genuinely more expressive room, and with it the obligation to decide what each state means. ## What each state should mean A workable convention, and the one most teams converge on: **Absent = the fact is not applicable or was never captured.** A `cancelledAt` field on an order that has not been cancelled should simply not be there. A `vatNumber` on a consumer account should not be there. This reads naturally: the document describes what is true about the entity, and says nothing about what is not. **Explicit null = the fact was captured, and the answer is nothing.** If a form asked for a middle name and the user actively answered "none", storing `middleName: null` records that the question was asked. If the same field is merely absent, you cannot tell whether the user declined or was never asked. That distinction matters for anything that later needs to backfill, re-prompt or audit. **Sentinels are usually a mistake.** An empty string that means "unknown" collides with an empty string that means "the user really entered nothing", and a zero that means "not measured" corrupts every average computed over that field. If you must use one, it should be a value that cannot occur naturally, and it should be documented as loudly as any other part of the contract. ## Why the distinction is not cosmetic Document stores generally let a query distinguish "key not present" from "key present with a null value", and equality matching against null does not behave identically to an existence test. That means a collection where the same logical state is encoded three different ways will return three different result sets depending on which query a developer happened to write. Reports disagree with the UI; a cleanup job misses a third of the documents it was meant to fix. There is also a storage and shape consequence. Absent fields consume no space per document, which is precisely why a document model handles sparse attributes well: a field present on a small fraction of documents costs roughly nothing on the rest. Materialising it as null on every document throws that advantage away for no informational gain — unless the null genuinely carries the "asked and answered" meaning above. ## Type stability matters more than presence A related and more damaging inconsistency is a field whose *type* varies: sometimes a string, sometimes an array of strings, sometimes a number that happens to be stored as text. Readers can cope with a field that is sometimes missing — that is a single well-defined branch. A field whose type varies forces type inspection at every use, breaks comparisons and sorts, and quietly changes how ordering behaves. If you take one rule away: presence may vary, type must not. ## Making the convention stick Conventions decay unless something holds them: - Put the decision in one place on the write path — a factory, a mapper, a constructor — so no caller gets to improvise. - Have the read path normalise once at the boundary, so the rest of the code sees one representation instead of three. - Make it part of the shape check if you enforce one: a required field with a nullable type says "always present, may be null"; leaving it out of the required set says "may be absent". - Periodically sample stored documents and count how each state actually appears. Drift shows up as a small but nonzero third variant, and it is far cheaper to fix while it is rare than after it is common. ## What interviewers are checking They want to see that you noticed there is a decision here at all. The weak answer is "they're the same thing". The strong answer distinguishes the meanings, notes that queries can tell them apart so the choice affects results, mentions that absent fields are free in a document model, and finishes on the point that consistency of the convention matters more than which convention you picked.
- Why are sparse optional fields cheap in a document model but expensive as columns in a fixed schema?An absent key occupies no space in the document that omits it, so a field present on a small fraction of documents costs nothing on the rest. A fixed row layout reserves space or a null marker on every row, so a mostly-empty column is paid for by every record whether or not it carries a value.
- Which is worse for readers: a field whose presence varies, or one whose type varies?Type variation is far worse. Varying presence is a single defined branch — use a default when it is missing. Varying type forces type inspection at every use, breaks comparison and sorting, and produces results that differ depending on which documents a query happened to touch.
- How would you find out which convention a legacy collection actually follows?Sample stored documents and count the states per field: present with a value, present and null, and absent. The distribution usually shows one dominant convention plus a small contaminated tail from a particular writer or era, which tells you both what the intended rule was and where to fix it.
saying these in an interview costs you the question
- Says missing and null are equivalent in a document store
- Materialises null on every document out of relational habit
- Uses empty string or zero as an unknown marker
- Lets each service pick its own convention per field
- Ignores that a field's type can drift as well as its presence