skip to content

In a product with an order ledger, sessions, a catalog, device metrics, a search box and media files, which store class fits each part?

level: middleimportance: must knowfreq 62%

answer

  1. interrogate each data shape separately
  2. transactions, expiry, variable attributes
  3. append-only timestamps and retention
  4. derived index versus system of record
  5. bytes elsewhere, metadata in database

basics

~20 s

Order ledger: relational. Sessions: in-memory key-value with expiry. Catalog: document or relational. Device metrics: time-series. Search box: a search engine's index fed from the source of truth. Media files: an object store, with metadata in the database.

solid answer

~40 s

I match each data shape to its access pattern. The **order ledger** needs multi-row transactions and constraints, so it goes in a **relational database** as the system of record. **Sessions** are read by ID on every request and expire, so an **in-memory key-value store** with per-key expiry fits. The **catalog** has attributes that vary per category and is read as a whole, so a **document store** fits, although a relational table with a JSON column is often enough. **Device metrics** are append-only, timestamped and queried by time range, so a **time-series store**. The **search box** needs tokenised, relevance-ranked text matching, so a **search engine's index**, fed asynchronously and rebuildable from the source. **Media files** are large immutable bytes, so an **object store** holds them while the database keeps metadata and the object key.

go deeper

for a junior

Recall the store classes by name and one data shape each is good for: ledger, sessions, catalog, metrics, search and media.

for a middle

Explain the access-pattern questions behind each mapping, especially which store is the source of truth and which is derived and rebuildable.

for a senior

Show you weigh the operational cost of each added store and would keep shapes in the relational database until a measured need justifies moving them.

for a principal

Discuss how the mapping evolves with scale and team size, and which specialisations are reversible derived views versus hard-to-change systems of record.

## Start from the access pattern, not the product list A single product rarely has one kind of data. The mistake is to pick one store out of habit and bend every feature to it, or to pick a trendy store and discover it cannot answer the questions the product asks. The disciplined approach is to interrogate each data shape first: - **Read/write ratio**: is it mostly reads, mostly writes, or both? - **Query shape**: lookup by key, range by time, ad-hoc filters, joins, or free-text relevance? - **Transactional needs**: must several records change together, with constraints enforced? - **Size and growth**: small records, huge blobs, or an endless append-only stream? - **Lifetime**: kept forever, expired after minutes, or downsampled after weeks? - **Source of truth**: is this data authoritative, or derived from something else and rebuildable? ## Mapping each shape to a store class | Data shape | Access pattern | Store class | Why | |---|---|---|---| | Order ledger | Mixed read/write, joins, money must add up | **Relational database** | Multi-row transactions, constraints, flexible reporting | | Sessions | Read on every request by ID, expire on idle | **In-memory key-value store** | Sub-millisecond key lookup, per-key expiry | | Product catalog | Read-heavy, whole-item reads, varied attributes | **Document store** (or relational with a JSON column) | Nested, per-category fields read as one unit | | Device metrics | Write-heavy append, time-range aggregates | **Time-series store** (or wide-column) | Time-chunked storage, compression, retention | | Search box | Tokenised text, typo tolerance, ranking, facets | **Search engine's index** | Inverted index and relevance scoring | | Media files | Large immutable bytes, streamed or downloaded | **Object (blob) store** | Cheap durable bytes; database keeps metadata | ## Walking through the choices 1. **Order ledger.** Placing an order debits stock, records payment state and writes line items. Those changes must succeed or fail together, and finance will ask questions nobody has thought of yet. A relational database is the natural **system of record** here. 2. **Sessions.** Every request loads the session by its ID, and abandoned sessions must disappear. An in-memory key-value store with a time-to-live per key does exactly this. Decide consciously whether losing sessions on a node failure is acceptable, because volatile stores trade durability for speed. 3. **Product catalog.** A laptop has RAM and screen size; a shirt has size and fabric. Forcing both into fixed columns leads to sparse tables or awkward attribute tables. A document store keeps each product as one nested document. Many relational databases can store and index JSON too, so this is a judgment call, not a rule. 4. **Device metrics.** Thousands of devices each send a reading every few seconds, queries ask 'average temperature per hour for device X last week', and old raw data is thrown away. Time-series stores chunk data by time, compress it heavily and drop old chunks cheaply. A wide-column store keyed by device and time bucket is a common alternative. 5. **Search box.** Users type partial, misspelled words and expect the best match first. A search engine's **inverted index** maps terms to documents and scores relevance. It is a **derived store**: the catalog stays authoritative, changes flow to the index asynchronously, and the index can be rebuilt from scratch. 6. **Media files.** Photos and videos are large, immutable and read by URL. Storing them as database blobs bloats backups and replication. An **object store** holds the bytes; the database row holds the object key, size, content type and owner. A CDN typically serves them to users. ## The cost of the resulting polyglot architecture Using several store classes is called **polyglot persistence**. It is justified only when each store earns its place, because every addition brings: - another system to back up, restore-test, monitor, patch and staff on call; - another failure mode when data must be copied between stores; - consistency work, since a derived store lags behind its source. A small product might reasonably keep the catalog, the metrics and even basic search in the relational database until measurements show a bottleneck. The mapping above is where each shape naturally lands at scale, not a checklist to deploy on day one.

  • Why would you still keep the catalog in the relational database at first, even though a document store fits its shape?
    Many relational databases can store and index JSON, so varied attributes do not force a new system. Keeping one store avoids a second backup, monitoring and on-call surface, and avoids copying data between stores. I would move the catalog only when a measured need appears, such as read volume or document size that the relational setup handles poorly.
  • Why should media files not be stored as binary columns in the order database?
    Large blobs inflate table size, backups, restore time and replication traffic, and they compete with transactional queries for memory and I/O. An object store is built for durable, cheap storage of large immutable bytes and pairs naturally with a CDN. The database keeps only the object key and metadata, which stays small and queryable.

saying these in an interview costs you the question

  • One database product should hold every kind of data in the system
  • The search index can be the only copy of the product catalog
  • Store images as binary columns so one backup covers everything
  • Every data shape needs its own store from day one
  • Session data must always live in the relational database for durability