skip to content

A shared routing table's reference count is the hottest address in a service whose throughput drops as workers are added. How do you fix it?

level: seniorimportance: should knowfreq 48%

answer

  1. make the update rarer, not faster
  2. who already keeps it alive
  3. own once, borrow below
  4. process-lifetime means never counted
  5. pin it and re-measure to confirm

basics

~20 s

Remove ownership changes rather than making them cheaper: hold one long-lived handle per worker and borrow below it, never count a process-lifetime object, or defer the updates. Making a contended atomic update faster is not an option.

solid answer

~50 s

The tax is one atomic update per ownership change on one address, and no tuning makes that address scale, so every real fix reduces the number of ownership changes. In order of how much they cost you: **borrow instead of own** — take the handle once per worker or per request and pass a non-owning reference into everything below it, which deletes whole chains of pairs; **stop counting the object at all** if it lives for the process, by publishing it as an immortal object whose count is never touched; **defer or batch** the updates so they leave the fast path; and only as a last resort **change the object's reclamation discipline**. Confirm before you cut: re-run with a handle deliberately held forever, so no count updates occur, and see whether throughput recovers. If it does not, the count was not your ceiling.

go deeper

for a junior

The takeaway to keep: a count is written on every share, so the way to make counting cheap is to share ownership less often, not to speed up the write.

for a middle

Be able to name the ladder — borrow below an existing owner, one handle per worker, never count a process-lifetime object, defer the updates — and say what each one gives up.

for a senior

Demonstrate the confirmation step before the fix: pin the object so no count updates occur, re-measure, and only then restructure. Then show you know borrowing trades an atomic update for a lifetime you now assert by hand.

for a principal

The durable version of this is a convention about where ownership may be taken at all — outermost frames own, inner frames borrow, published data is immortal — so the hot-count failure cannot re-enter the codebase one helper at a time.

## First, prove it is the count A profile that names the count's address is suggestive, not conclusive, because the count usually shares an object header with other fields that the request path also touches. Two cheap experiments settle it: 1. **Pin the object.** Take one extra handle at start-up and never drop it, then change the request path to use a non-owning reference. No ownership changes occur, the object can never be reclaimed, and the only thing that changed is the counting traffic. If throughput recovers and scales with threads again, the count was the ceiling. 2. **Vary the thread count deliberately.** Plot throughput against workers. Counting contention has a distinctive shape: it climbs, flattens early, and then goes *down* as more threads are added, because each extra thread adds ownership transfers to the same word without adding capacity. Both are reversible and neither needs a specialised tool, which is why they are the right first move. ## The fixes, in the order you should reach for them ### Borrow instead of own Most ownership changes in a request path are unnecessary. A handle is copied into a helper, into a request context, into a closure — each pair is two atomic updates, and none of them is keeping the object alive for any interval that is not already covered by an outer handle. If a call's lifetime is provably nested inside an existing owner's, the call needs a **non-owning reference**, not a handle. - Take the handle once, at the outermost point that owns the work. - Pass a non-owning reference downward through the call chain. - Only take a new handle where ownership genuinely escapes — stored beyond the call, queued, or handed to another thread. This routinely removes an order of magnitude of counting traffic and changes nothing about when the object dies. ### Hold one handle per worker If the table is looked up many times per request, or the request path is too fragmented to thread a borrow through, give each worker thread one long-lived handle taken when the worker starts. The remaining updates still land on the same single word, but that word is now touched a few times per process lifetime instead of millions of times per second, which is the whole difference. ### Stop counting it An object published for the lifetime of the process has a count that will never reach zero. The accounting is pure loss. Many schemes allow such an object to be marked **immortal** — a sentinel or saturated count that update paths recognise and skip — turning every share into a plain read. This is the strongest fix available and it is exactly right for configuration, routing tables, interned values and other publish-once data. Its cost is that the object is now unreclaimable by construction, so it must genuinely be process-lifetime, and replacing it requires a separate mechanism. ### Defer or batch the updates If the object must stay reclaimable, the updates can leave the fast path instead of disappearing: buffer count deltas per thread and apply them in batches at chosen points, or leave the most frequent class of references out of the count entirely and re-establish the truth periodically. The throughput win is real and the price is **promptness** — reclamation no longer happens at the instant of the last drop. ### Change the discipline If a hot, widely shared object must be reclaimable and none of the above fits, the honest conclusion is that counting is the wrong mechanism for this object, and a reachability-based scheme — which puts no work on the sharing path at all — suits it better. This is a large change and belongs last. ## What does not work | Attempted fix | Why it fails | |---|---| | Making the count field wider or narrower | The cost is exclusive ownership of the line, not the value's size | | Taking the handle earlier in the request | The same two updates still happen, just closer together | | Splitting the table into many objects | Each request then counts several objects instead of one | | Adding more worker threads | More threads means more transfers of the same line | | Caching lookup results | Helps if lookups were the cost, but the count is written per share regardless | ## The shape of the judgment The useful generalisation is that **counting cost is a function of ownership churn, and ownership churn is a property of your code's structure, not of the object**. Every fix above is a restructuring: own less often, own for longer, or do not own at all. That is why the first question to ask about a hot count is not 'how do I make this update faster' but 'why does this thread need to own the table at all, when something above it already does'.

  • What is the risk of replacing handles with non-owning references down the call chain?
    You are now asserting a lifetime by hand: the borrow is only valid while some outer handle lives. If a borrowed reference is stored beyond the call, captured by a callback that outlives it, or queued to another thread, it outlives its guarantee and becomes a dangling reference. The rule is that a borrow may travel down the stack, never out of it.
  • Why is marking the table immortal so much cheaper than counting it correctly?
    An immortal object's count is never written, so sharing it touches only unwritten lines, which every core may cache at once. The contention disappears rather than being reduced. The price is that the object can never be reclaimed, so this is only honest for data that genuinely lives as long as the process.

saying these in an interview costs you the question

  • Tries to make the atomic update faster instead of rarer
  • Adds threads to recover throughput and makes the curve worse
  • Splits the object into pieces, multiplying the ownership changes
  • Assumes the profile naming the count proves it without pinning the object
  • Stores a borrowed non-owning reference beyond the call that borrowed it
  • Marks a replaceable object immortal and then needs to replace it