skip to content

An in-memory store holds a one-byte flag under a ninety-character key, forty million times over. What is that ratio costing?

level: middleimportance: should knowfreq 48%

answer

  1. compare the two numbers, not just one
  2. the key is stored, not free metadata
  3. the key travels on every call too
  4. constant segments distinguish nothing
  5. share times count decides if it matters

basics

~20 s

The addressing has become the data: about 3.6 gigabytes of key text against 40 megabytes of payload. Keys are stored bytes, they travel on every call, and at high counts a key can far outweigh what it addresses.

solid answer

~50 s

The key is not free metadata — it is stored, and at this ratio it is roughly ninety times the value it addresses. Forty million entries at ninety bytes of key is about 3.6 gigabytes of key text against about 40 megabytes of payload, before the fixed bookkeeping each entry carries on top of both. The key also travels: it is sent on every request, held by whatever the application logs, shipped to any replica and contained in any copy the tier writes out. The remedies are to cut segments that are constant across every key and therefore carry no information, or to mint fewer keys so the key text is amortised over more value — which, where the server understands the value as a structure, means one key addressing several fields rather than one key per field.

go deeper

for a junior

Recall that the key string is stored data, not a free label. Multiplying key length by the number of entries gives a real number, and with tiny values that number can exceed the payload.

for a middle

Explain the ratio and both levers: cut segments that are identical on every key and carry no information, or mint fewer keys so one key addresses more value where the server understands the value.

for a senior

Demonstrate the judgment about when not to act: compute the share, weigh it against legibility on a shared tier, and refuse micro-optimisation where the value dominates the entry.

for a principal

The trade is bytes against the ability to attribute a keyspace years later. Decide where the organisation sits on that line once, so teams are not each inventing an answer under pressure.

The habit this question tests is comparing two numbers that most designs never put side by side: the size of the key against the size of the value it addresses. ## Do the subtraction Ninety bytes of key, one byte of value, forty million entries: - key text: forty million times ninety bytes, about **3.6 gigabytes**; - payload: forty million times one byte, about **40 megabytes**; - ratio: the address is about ninety times the thing it addresses. On top of both sits the fixed bookkeeping every entry carries, which is a separate subject and is measured rather than assumed. Even ignoring it, the shape of this design is unmistakable: the tier is a warehouse of addresses that happen to have a bit of data attached. ## Why the key is not free It is tempting to treat the key as a label — something the store uses to find the entry and does not really keep. It keeps it. Concretely the key is paid for in several places: - **In memory.** The key string is stored alongside the value, for every entry, for as long as the entry lives. - **On the wire.** It is sent on every request that touches the entry, and by any call that names several keys at once. At small values the key can be the majority of the bytes on the network too. - **In the copies.** Anything the tier ships to a replica or writes out as a copy contains the key text as well as the value. - **In everything around the store.** Application logs, traces, monitoring labels and the rows of any application-maintained index hold the key again. ## What is actually worth cutting | Part of the key | Information it carries | Worth keeping? | |---|---|---| | a segment identical on every key | none, by definition | first thing to cut | | a segment with a handful of values | a little | shorten rather than remove | | the identifier that distinguishes this entry | all of it | must stay in some form | | a segment a human needs on call | none technically, a lot operationally | keep, and pay for it deliberately | A constant segment repeated on forty million keys is the purest waste in the design: it distinguishes nothing, and every byte of it is multiplied by the count. Removing it is free of risk because it cannot cause two keys to become one. ## Minting fewer keys, and the premise that decides it The other lever is to amortise the key across more value. One key addressing four fields carries one copy of the key text instead of four. This is worth doing **where the server understands the value as a structure and can read or change part of it**. Where a value is opaque bytes the server only hands back, the same move makes every change a full read, modify and write and every read pull the whole value across the network — the key bytes you saved come back as traffic. What such structures can actually do is another subject; here it is only the premise that decides whether the amortisation is real. ## When the ratio does not matter The same ninety-byte key against a four-kilobyte value is about two percent overhead, and no effort spent shortening it will ever be visible. Three things decide whether the ratio is worth acting on: 1. **The share.** Key bytes over total chosen bytes. Below a few percent, stop. 2. **The count.** A high share on ten thousand entries is a rounding error in absolute terms. 3. **What you give up.** A readable key is what lets whoever is on call in two years attribute an entry to a team, a feature or a tenant. That is worth real bytes, and the trade should be made deliberately rather than by default in either direction. One caution about numbers: how long a key may be at all differs between stores by a wide margin, and any specific ceiling is a property of one store stated as a premise, never a general fact. Design away from any ceiling rather than up to it, and find out what yours is before the scheme is carrying traffic.

  • Does the key cost anything beyond the memory it occupies?
    Yes, in four more places. It is sent on every request that names the entry, echoed by calls that address many entries at once, carried to any replica and into any copy the tier writes out, and repeated in application logs, traces and the rows of any application-maintained index. At small values the key can dominate network bytes as decisively as it dominates stored bytes.
  • When is a long, descriptive key the right call regardless of the arithmetic?
    When the value dominates, so the key is a small share of each entry, and when the tier is shared by teams who need to know at a glance who owns an entry. Legibility has real operational value: an unreadable keyspace cannot be attributed, audited or cleaned up. Pay the bytes on purpose when the share is small; revisit only when the count makes the share expensive.

A warehouse of forty million envelopes, each holding a single postage stamp. The stamps would fit in one drawer; what fills the building is the addressing.

saying these in an interview costs you the question

  • Treats the key as metadata the store does not really store
  • Prices entries by value size alone
  • Assumes shortening keys helps even when values are kilobytes
  • Quotes one store's key-length ceiling as a general rule
  • Forgets the key is sent on every request and every copy
  • Cuts the segment that identifies the entry rather than the constant one