skip to content

You are setting the accessor policy for an in-memory dataset library: when should a call hand back a borrowed view rather than an owned copy?

level: principalimportance: nice to knowfreq 26%

answer

  1. bytes against exported constraint
  2. a view bans mutation while held
  3. retention means a key, not an address
  4. public boundary defaults to copies
  5. view to copy migrates, copy to view does not

basics

~20 s

Return a view only for large results read immediately inside your own build, where copying would dominate. Default to an owned copy otherwise: a view bans mutation of the dataset for as long as any caller holds it.

solid answer

~50 s

The trade is not bytes versus bytes; it is bytes versus **constraint**. A view costs no allocation, but while a caller holds one, the dataset cannot be mutated, and that ban propagates through the caller's code and into anything it hands the view to. A copy costs bytes and can drift from the source, but it constrains nobody. My default policy: copy for small scalar results and for anything crossing a public boundary; view for large results consumed inside a tight read loop; key or index for anything the caller may store. Two second-order points decide the edge cases. A view leaks your internal representation into your signature, so you cannot change the layout later without breaking callers. And the migrations are asymmetric: narrowing a published view to a copy is usually source-compatible, while widening a copy to a view is not.

go deeper

for a junior

Know the two options exist: a view points at the library's own data, a copy is yours to keep. If you plan to hold the result, you want a copy or a key.

for a middle

Explain why a held view blocks mutation of the dataset, and why a copy can go stale. Those two sentences are the whole trade at this level.

for a senior

Argue it with the caller's code in view: which shape forces a restructure, what the copy costs in the real access pattern, and how a scoped cursor bounds the constraint.

for a principal

Own the policy and its migration path. Say where the constraint should live, what evidence would move your default, and why narrowing a view to a copy later is cheap while the reverse is not.

## What each option really costs | Accessor returns | Runtime cost | Constraint on the caller | Staleness | |---|---|---|---| | A borrowed view | None beyond an address | The dataset cannot be mutated while it is held | Impossible: it is the live data | | An owned copy | Allocation plus the bytes | None | Yes, from the moment it is taken | | A key or index | One lookup on each use | None | Detectable: the lookup can fail | The column that decides most arguments is the third, not the second. Copying is a cost you pay once and can measure. A view is a **constraint you export**, and its cost shows up later, in a place that has no idea where it came from: a caller who cannot append to the dataset inside their loop, a team who cannot hold your result across an await point, a refactor that stalls because your return type pins their control flow. ## The policy I would write down 1. **Small scalar results are always copies.** A number, a flag, a short identifier. There is no performance argument for a view, and the constraint is pure loss. 2. **Bulk reads in a tight loop get a view**, ideally through an explicitly scoped cursor so the window is visible at the call site and obviously short. 3. **Anything a caller may retain gets a key**, not an address. Retention is the case where a view is both most tempting and most dangerous, and a key turns "is this still valid?" into "does this still exist?", which the dataset can answer. 4. **A public boundary defaults to copies.** You control the compile of every caller inside your own build; you do not control theirs, and a view makes their lifetimes your business. 5. **Mutating accessors take an exclusive borrow of the dataset and return nothing borrowed.** Handing out a view from a mutating call re-creates the aliasing problem inside your own API. ## Second-order effects worth raising unprompted - **A view is a representation leak.** Its type says the caller is looking at *your storage*. Change from contiguous rows to a column-per-field layout and every view signature changes with it; the copy-returning version of the same API would not have noticed. - **Migration is asymmetric.** Turning a published view into a copy usually keeps callers compiling, because a copy satisfies every use a read-only view did. Turning a copy into a view rarely does, because callers hold the result across mutations that used to be fine. That asymmetry argues for starting conservative. - **Views compose badly with buffering.** The moment a caller wants to batch, queue or defer what they read, a view forces either a copy at the boundary or a lifetime constraint on the queue. You may as well have made the copy where it was cheap and obvious. - **The benchmark that justifies a view must include the caller's restructuring.** A view often wins the microbenchmark and loses the program, because the caller has to reorganise around the mutation ban. ## How to decide, in order 1. How big is the result compared with an address? Under a few dozen bytes, copy and stop. 2. Will the caller plausibly hold it past the call? If yes, key. 3. Does the dataset mutate during a typical session? If yes, the mutation ban is expensive, so bias to copy. 4. Is the caller outside your build? If yes, copy, unless you have measured that it matters. 5. Only if a large result, consumed immediately, inside your own build, in a measured hot path — return a view, and scope it explicitly. ## What a strong answer sounds like It does not say "views are faster". It says the library is choosing **where a constraint lives**: inside the library as copy cost, or outside it as a mutation ban on every caller, or split as a lookup on each use. It states a default, names the measurement that would override it, and admits the case the default handles badly — large results read in a loop, where a copy per row is a real regression and a scoped cursor is the honest answer.

  • Why does returning a view make your internal layout harder to change?
    Because the view's type is a statement about how the data is stored. Reorganising the layout changes what a view can be, so every caller that holds one is affected. An accessor that returns owned values keeps that decision inside the library.
  • Your hot path really does need per-row views. How do you limit the damage?
    Hand out a scoped cursor rather than loose views: the caller receives it inside a call you control, uses it there, and cannot store it. The mutation ban is then visibly bounded to that block, and callers who want to retain anything are pushed to the key-based accessor.
  • What measurement would override a copy-by-default policy?
    A profile showing the copies dominating a path that matters, taken with the caller's real access pattern and including whatever restructuring a view would force on them. A microbenchmark of the accessor alone does not qualify, because it omits the cost the policy actually exports.

saying these in an interview costs you the question

  • Argues views are simply faster, with no mention of the constraint
  • Returns views across a public boundary by default
  • Lets callers store views and calls the resulting rule a convention
  • Thinks an owned copy stays in step with the underlying row
  • Ignores that a view pins the library's internal layout
  • Optimises the accessor in isolation from the caller's loop