skip to content

When a teammate proposes a specialised datastore for a new service, what access-pattern signals justify it over the relational database you already run?

level: middleimportance: should knowfreq 55%

answer

  1. default is what you already operate
  2. measured need, not popularity
  3. search, time-series, blobs, write volume
  4. JSON columns and replicas cover much
  5. every store adds ops and sync work

basics

~10 s

Justified signals are measured ones: write volume beyond one node on key-partitionable data, relevance-ranked text search, high-rate time-stamped appends with retention, or large blobs. Popularity, vague scale worries or one query are not enough.

solid answer

~40 s

I treat the relational database as the default because it is flexible, transactional, well understood and already operated. A specialised store earns its place when the access pattern clearly outgrows it: a sustained write rate or data volume one node cannot absorb, on data that is only ever accessed by a partition key; **relevance-ranked full-text search** with typo tolerance and facets; **high-rate timestamped appends** with time-range aggregates and retention; **large binary objects**; or deep **multi-hop relationship traversal**. I also ask for evidence, such as a measured bottleneck or a query the relational engine cannot serve well, and I count the price: another system to back up, monitor and staff, plus the work of keeping it consistent with the source of truth. 'It scales better' with no measurement is not a reason.

go deeper

for a junior

Recall that the relational database is a sensible default and name a few needs, like text search or large files, that point elsewhere.

for a middle

Explain which access-pattern signals justify each store class and why weak signals like variable fields or read volume usually have relational answers first.

for a senior

Demonstrate that you demand measurements and count operational and consistency costs before approving a new store, and that 'not yet' is a valid outcome.

for a principal

Frame the decision as reversibility and total cost of ownership, distinguishing derived stores from systems of record and setting thresholds to revisit.

## Why the relational database is the default In most teams the relational database is already running, backed up, monitored and understood. It offers several things that are expensive to rebuild elsewhere: - **Query flexibility**: SQL lets you answer questions nobody planned for, which matters when the product is still changing. - **Transactions and constraints**: multi-row atomic changes and enforced integrity. - **Maturity**: well-known tooling for migrations, backups, point-in-time restore and access control. - **Breadth**: many relational engines also offer JSON columns, basic full-text search, partitioning and replicas, which covers a surprising range of 'specialised' needs at modest scale. So the question in a design review is rarely 'which store is best?' but '**what does this workload need that the store we already run cannot give us?**' ## Signals that genuinely justify a different store class | Signal | What it looks like | Store class it points to | |---|---|---| | Write volume beyond one node | Sustained writes one primary cannot absorb, data always accessed by a partition key | Wide-column or distributed key-value store | | Relevance-ranked text search | Typo tolerance, stemming, ranking, facet counts | Search engine's index | | High-rate time-stamped appends | Metrics or events, range aggregates, downsampling, expiry | Time-series store | | Large binary objects | Images, video, backups, exports | Object store | | Extreme low-latency key reads | Per-request lookups where a network round trip to the database is too slow | In-memory key-value store | | Deep relationship traversal | Many-hop 'friends of friends' queries | Graph store | Each row is a **measured** property of the access pattern. Notice that 'variable attributes' alone is weak, because a JSON column often handles it, and that 'high read volume' alone is weak, because replicas and caching usually handle it before a new store class is needed. ## Signals that do not justify it 1. **Popularity or familiarity**: 'everyone uses it' or 'I used it at my last job'. 2. **Unmeasured scale fears**: 'it might not scale' without a load estimate or a benchmark. 3. **One awkward query**: a single report that a view, an index or a nightly job could serve. 4. **Schema friction**: a desire to skip migrations, which usually moves schema enforcement into application code rather than removing it. 5. **Future-proofing**: choosing a store for a scale the product may never reach. ## Counting the price of saying yes Every new store class adds work that does not show up in a feature estimate: - **Operations**: provisioning, upgrades, backups, restore drills, monitoring, alerting and on-call knowledge. - **Consistency**: if data lives in two stores, something must copy it and handle failures and ordering. Dual writes from request code are a classic source of silent divergence. - **Query limits**: many specialised stores require you to know the queries up front. A wide-column or key-value store designed for one access path cannot easily answer a new one. - **Skills**: the team must learn a new data model and new failure modes. - **Exit cost**: moving off a store later means backfilling, dual running and verifying data. ## A practical review checklist - What exact queries and write rates does the feature have, with numbers? - Can the existing relational database serve them with indexes, partitioning, replicas, a JSON column or a materialized view? - If not, is the new store a **derived view** that can be rebuilt, or would it become a **system of record**? - How will data reach it, and what happens when that pipeline fails? - Who operates it at 3 a.m.? A proposal that answers these with measurements is a good one. A proposal that answers them with enthusiasm is usually premature. The outcome is often 'not yet': keep the data relational, record the threshold at which the decision should be revisited, and design the data-access layer so that a later move is possible. ## Interview framing When asked 'why not just use the relational database?', a strong answer names the specific access pattern, the measured limit it hits, the store class that removes that limit, and the new cost the team accepts in exchange. If any of those four parts is missing, the honest answer is to keep the relational database for now.

  • What evidence would you ask for before approving a new store class?
    Concrete numbers: current and projected write and read rates, data volume, and latency targets, plus a demonstration that the relational database cannot meet them with indexes, partitioning, replicas or caching. I would also want the queries written down, a plan for how data reaches the new store, and a named owner for operating it.
  • When is a specialised store a lower-risk addition than it first appears?
    When it is a derived view rather than a system of record. A search index or cache fed from the relational database can be rebuilt from the source if it misbehaves or the choice proves wrong, so the decision is largely reversible. Replacing the authoritative store is far harder to undo.

saying these in an interview costs you the question

  • Relational databases do not scale, so new services should avoid them
  • Varying fields per record always require a document store
  • Adding a new store only costs the time to write the integration
  • High read volume by itself justifies a different store class
  • Schemaless stores remove the need to manage schema changes