skip to content

The Access Cost Model

What one call to this tier actually costs the caller: server time against network time, how many round trips a design makes, and what a connection is holding while it waits.

on this pageshow

questions

21

An in-memory store answers a lookup in microseconds, but the caller opens a fresh connection for every call - what does each call pay?

level: juniorimportance: must knowfreq 70%

answer

  1. setup, not service
  2. count the round trips first
  3. the handshake precedes any request
  4. credentials and encryption add trips
  5. microseconds of work, milliseconds of setup

basics

~20 s

Connection setup, paid before any work begins: at least one network round trip for the handshake, plus further trips where an encrypted transport or a credential exchange is required. Microseconds of server work end up buried under milliseconds of setup.

solid answer

~50 s

Split the call into the server's service time, which is what the store spends executing, and the caller's wall time, which is what the caller measures end to end. The service time is tens of microseconds. Opening a connection first costs a handshake round trip, then whatever the store requires on top - an encrypted transport and a credential exchange each add trips - and only then is the request sent and its own round trip paid. On a same-zone tier that turns a call of roughly 0.35 ms into something nearer a millisecond, with the store doing no extra work at all. A pool amortises the setup across every later call on that connection. It does not remove the request round trip, and an idle pooled connection still occupies a slot against the store's connection ceiling.

go deeper

for a junior

Recall that a call to this tier is a tiny amount of server work plus a network trip, and that connections are reused from a pool rather than opened per call. Being able to say setup is paid before any work is enough here.

for a middle

Explain the setup sequence in order - handshake, any encrypted transport, any credential exchange, then the request - and do the arithmetic that shows a per-call connection costing several times a pooled one on the same tier.

for a senior

Show what a pool does not fix: the request round trip per call, the standing slot each connection holds, and a liveness check on borrow that silently reinstates the cost the pool was installed to remove.

for a principal

Frame connect cost as a placement and contract decision: where the tier sits relative to callers, whether an encrypted transport or credentials are required at all, and what standing connection footprint the fleet is permitted to hold.

## The two clocks in one call An in-memory store is asked for a value and answers from memory. **The server's service time** - what the store spends executing the operation - is typically tens of microseconds, because the work is a lookup in a hash table and a copy of some bytes. **The caller's wall time** - what the caller measures end to end - is a different number, and on a healthy system almost none of it belongs to the store. For a pooled connection to a tier in the same zone, the caller's wall time is roughly the sum of: - **one round-trip time** across the network, which on a same-zone hop is a few hundred microseconds; - the **serialisation** of the request and the parsing of the reply, at both ends; - the store's service time, the smallest term in the sum; - whatever the caller's own runtime adds between issuing the call and being scheduled to read the reply. Assume a same-zone round-trip time of 0.3 ms and a service time of 0.05 ms. A pooled call costs about **0.35 ms**, and the store owns roughly a seventh of it. ## What opening a connection adds Opening a connection is not one more term in that sum. It is a sequence that must run to completion **before the first byte of the request is sent**: 1. **Transport establishment** - the handshake that creates the connection costs at least one round trip, and nothing useful can be sent until it finishes. 2. **Whatever the store requires on top** - an encrypted transport adds further round trips, and a store that expects a credential adds an exchange of its own. 3. **Server-side accept work** - the store allocates per-connection state, including **the reply buffer at each end**, and consumes one slot against **the connection ceiling**, the limit the component itself sets on concurrent connections. 4. **Teardown afterwards** - closing leaves socket state on the caller's host for a while, so a high rate of connect-and-close can exhaust resources on the caller long before the store notices anything. ## The arithmetic | Cost, same-zone tier | Pooled connection | New connection per call | |---|---|---| | Transport establishment | none | about 0.3 ms, one round trip | | Encrypted transport, where used | none | about 0.3 to 0.6 ms | | Credential exchange, where required | none | about 0.3 ms | | Request round trip | about 0.3 ms | about 0.3 ms | | Server service time | about 0.05 ms | about 0.05 ms | | **Caller's wall time** | **about 0.35 ms** | **about 0.95 to 1.25 ms** | Those figures illustrate the shape; they are not any product's published numbers. Connecting per call multiplies the caller's wall time by roughly three to four **while the store does not a microsecond of extra work**. That is the entire point of the category: the cost of a design against this tier is trips, and connecting is trips. ## What a pool does not remove A pool is often described as though it made calls free. It removes one term and leaves the rest standing. - **The request round trip is still paid on every call.** A pool of a hundred connections does not make a loop of a hundred single-key reads cheaper than a hundred round trips. - **An idle pooled connection still holds a slot** against the connection ceiling and still carries per-connection buffers on the server, so a pool converts a per-call cost into a standing one. - **A liveness check on every borrow gives the saving back** - a pool that makes a call to validate a connection each time it is handed out has re-added exactly one round trip per use. Validate on idle or on a schedule instead. - **Nothing about the store's own work changes.** If a call is expensive because of what it asks for, no pool affects it. ## Where stores in this class differ The connect cost is not a single number across the class, and quoting one means quoting a product: - **Credentials vary.** Some stores of this class ship with no authentication at all and are protected only by the network they sit on; others require a credential exchange, which costs a trip at connect and nothing afterwards. - **Encryption varies.** Some tiers are reached over an encrypted transport as a matter of course, some cannot do it at all, and some sit behind an intermediary that terminates the connection, in which case the caller's connect cost is the cost of reaching the intermediary. - **The accept costs the server different amounts** depending on how large its per-connection buffers are, which is why the same connection count is comfortable on one tier and expensive on another. The claim that holds for every store in the class is the ratio: **setup is measured in round trips and the work is measured in microseconds**. Any design that pays setup per call has made the store's speed irrelevant to its own latency.

  • What does a connection that is sitting idle in the caller's pool still occupy?
    A socket and its buffers on the caller, and on the store a slot against its connection ceiling plus the per-connection state that comes with an accepted connection, including the reply buffer at that end. Idle is cheap per call and not free in total, which is why an oversized pool is a real cost rather than a harmless margin.
  • Why does validating a pooled connection on every borrow undo part of the benefit?
    Because a validation call is itself a round trip, added before the real request goes out. On a tier whose calls are microseconds of service time and one trip of network, that roughly doubles the caller's wall time for every call. Validate connections when they have been idle, or on a background schedule, so the check is amortised rather than per use.

Placing a phone call to say one word. The dialling, the ringing and the greeting are the cost; the word itself is free. Keeping the line open is the pool - it removes the dialling, never the speaking.

saying these in an interview costs you the question

  • Calls connection setup free because the store holds everything in memory.
  • Says a pool exists to save client memory rather than round trips.
  • Quotes the store's service time as what the caller experiences.
  • Assumes an idle pooled connection costs the store nothing at all.
  • Believes every store of this class demands a credential exchange at connect.
open as a page

A request makes 500 single-key reads, each answered in microseconds, yet takes 300 ms overall — where did the time go?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Almost all of it went into the network. Each read pays one full crossing, so 500 sequential reads pay 500 crossings. The store's microseconds are far too small to explain the delay; the number of crossings is the cost.

open as a page

One write in a run of two hundred sent to a volatile tier without waiting for replies fails — what happened to the others?

level: middleimportance: must knowfreq 62%

basics

~20 s

The other one hundred and ninety-nine were executed and their effects stand. Sending a run of operations without waiting for replies buys latency and nothing else: no atomicity, no isolation, no rollback — each operation succeeds or fails on its own.

open as a page

A request needs forty values from a volatile tier one zone away — how does a pipelined run differ from one operation taking all forty keys?

level: middleimportance: must knowfreq 68%

basics

~20 s

Both collapse forty round trips into one, and that is all they share. A pipelined run stays forty independent operations with forty replies matched in order; one operation taking many keys is a single operation with a single reply and a single outcome.

open as a page

While a caller is parked on a wait-with-deadline call, what does it consume on the store and what does it not?

level: middleimportance: must knowfreq 58%

basics

~20 s

A parked caller consumes a connection slot for the entire wait and essentially no server processing time. The store registers the waiter and goes on serving everyone else, so every store-side signal stays flat while the caller's connection is unusable.

open as a page

An in-memory store takes 40,000 calls per second, each holding a connection 0.4 ms of wall time - how many connections does the pool need?

level: middleimportance: must knowfreq 62%

basics

~20 s

About sixteen, plus a margin. Pool size follows concurrent calls in flight, which is arrival rate times hold time - 40,000 per second times 0.0004 seconds - not the request rate, so microsecond-scale calls need few connections.

open as a page

A page render needs 60 values from a tier in another zone within a 50 ms budget; how do you check it fits?

level: middleimportance: must knowfreq 64%

basics

~20 s

Measure one round trip from the real caller to the real tier at a high percentile, multiply by the crossings the path makes, and compare with the share of the 50 ms the tier gets. Sixty sequential crossings will not fit.

open as a page

In an in-memory store where most operations cost microseconds of server service time, what makes one call cost hundreds of milliseconds?

level: middleimportance: must knowfreq 62%

basics

~20 s

A call's server service time scales with how much stored data it must touch, not with how many bytes the caller sent. Returning every member of one entry, or walking the whole keyspace, does work proportional to what is stored.

open as a page

Eighteen workers park on wait-with-deadline calls against a shared pool of twenty connections, request-path lookups stall, and the store reports no load — why?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The scarce resource is the pool, not the store. Eighteen parked callers hold eighteen slots for their whole wait while spending no server processing time, so the store is genuinely idle and the request path queues for two connections.

open as a page

One call keeps an in-memory store busy for 200 milliseconds: who is delayed if the server executes one operation at a time, and who if it serves requests from a thread pool?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Where the server executes one operation at a time, every other caller waits the full 200 ms and then the backlog drain. Where a thread pool serves requests, one worker is lost and only callers needing the same lock block.

open as a page

Why would a caller ask a store to hold its request until an entry appears instead of re-asking every hundred milliseconds?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Parking trades network trips for a held connection: one call instead of one per ask, and the caller learns of the entry a round trip after it is written rather than an interval later. Not every store offers it.

open as a page

A caller requests a thirty-second server-side wait while its client library caps every operation at five seconds — what happens?

level: middleimportance: should knowfreq 44%

basics

~20 s

The inner deadline wins. Every wait longer than five seconds is abandoned by the caller, which usually discards the connection and retries, so the design degrades into polling at five-second granularity plus a connect cost per attempt.

open as a page

A caller pipelines a million writes to a volatile tier in one run — where does the memory go, and what bounds a safe run size?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Unread replies pile up in the reply buffer at each end — the server's for what it has produced, the caller's for what it has received and not read. Run size is bounded by memory, not by the protocol, so size the chunk in bytes of expected reply.

open as a page

A parked caller's connection is lost and re-established on another node after a failover — why might it wake with nothing?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A wait is state on the connection that registered it. When that connection dies the registration dies too, nothing on the new node knows the caller was waiting, and a value written across the gap may be gone or already taken.

open as a page

An in-memory store has hit its ceiling on concurrent connections: one caller sees instant connect errors, another only gets slower - why does one limit produce two unlike symptoms?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The connection ceiling has two surfaces. A store that refuses above it returns a fast, loud error naming the dependency; a store that simply stops accepting leaves the caller's connect attempt unanswered, which presents as general application slowness.

open as a page

Your tier moves from the same zone to a second region; what happens to a path making 12 sequential calls?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Each of the twelve crossings now pays a cross-region round trip, so a path that cost a few milliseconds costs hundreds. Distance is set by physics and topology, not configuration, so only fewer crossings or a closer tier help.

open as a page

An in-memory store stalls for 200 ms a few times a minute, yet its mean server service time and operations-per-second counter look normal; why?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Averaging buries a rare stall: one 200 ms sample among millions of microsecond samples barely moves the mean, and the delayed operations are still counted, only later. The waiting shows in the callers' wall time, at a high percentile.

open as a page

Two teams collapse round trips to a shared volatile tier in different ways — what must a platform-wide guideline fix rather than leave to each team?

level: principalimportance: should knowfreq 38%

basics

~20 s

Fix what a single team cannot see: a bound on unread reply bytes per connection, a trip budget per request path, and the rule that collapsing trips buys latency only. Fix nothing that depends on a capability the stores may not share.

open as a page

One shared in-memory store serves thirty autoscaling services, each instance holding its own pool - how do you budget its connection ceiling across them?

level: principalimportance: should knowfreq 42%

basics

~20 s

Treat the ceiling as a budget to allocate, not a limit to discover. Demand is per-instance pool maximum times each service's maximum instance count, plus operators and jobs - so cap pools from real concurrency and reserve incident headroom.

open as a page

How would you set and enforce a per-request trip budget for a shared in-memory tier across many teams?

level: principalimportance: should knowfreq 38%

basics

~20 s

State, per request path, how many crossings it may make and at what placement, derived from the user-facing deadline and a measured per-crossing cost. Then measure trips per path, not calls per second, and treat a removed crossing as better than a tuned one.

open as a page

A service makes several calls to a volatile tier whose cost grows with stored data; which bounded forms do you require, and what does each cost?

level: principalimportance: should knowfreq 36%

basics

~20 s

Require that no call's cost be decided by stored data: ask for a bounded range, walk incrementally with a cursor, cap what one request may return, or precompute the answer into its own entry. Each bound costs trips, freshness or a refusal.

open as a page