The same selection expression shares storage in one tool and allocates in another - what should portable code commit to instead?
answer
- the text cannot say where bytes live
- the shape permits, the design decides
- state the requirement, do not detect it
- materialise where the source must go
basics
~20 sCommit to the requirement, not to the observed behaviour. Whether a result shares bytes is a design decision that the text of an expression cannot express, so state whether the source must be releasable and materialise when it must be, instead of inheriting whatever the tool did.
solid answer
~50 sOwnership is not a property of the expression; it is a policy of the design underneath. The same selection can be a **window onto the same bytes** in one tool, **storage of its own** in a second that duplicates eagerly, and a promised duplicate in a third that allocates only when something is modified - with identical text and identical values. Code that reads correctly under one of those and holds a large source resident under another is not a porting bug, it is an assumption that was never written down. The portable move is to say what the code needs: if a returned selection must not keep its source alive, materialise it into storage of its own before returning it. Detecting which kind of result you were handed and branching on it encodes the very per-tool detail you were trying to escape.
go deeper
Take away that two tools can return the same values from the same expression and differ in whether the result shares memory with its source.
Name the policies - duplicate eagerly, share where the shape allows, share untouched buffers, defer the duplicate until a write - and explain why the source text cannot distinguish them.
Show the engineering move: identify the boundary where the source must become releasable, materialise there, and treat detect-and-branch as re-importing the dependency you were removing.
Own the rule for the codebase - what a function returning a selection promises about retention - and defend it by the class of silent difference it removes when the tool underneath changes.
## Why identical text answers differently A selection expression says which rows and columns you want. It has no way to say where the result's bytes should live, because that is not part of what is being asked. The answer therefore comes from the design, and the designs in this family have genuinely different policies: - **Eager duplication.** Every result is materialised into freshly allocated storage. Callers never reason about sharing; every step costs bytes. - **Sharing where the shape allows it.** A selection describable as a start, a count and a step is handed back as a window on the original's storage; anything irregular is gathered and owns its bytes. - **Immutable by default, sharing untouched buffers.** The step returns a new handle that shares every buffer it did not need to change, so what is shared is decided per buffer rather than per selection. - **Deferring the duplicate until something is written.** The result is a logical duplicate immediately, and bytes are allocated only when one side is modified. None of these is wrong, and each was chosen for a reason - predictability, cheapness, safety, or a late and smaller cost. What matters is that the choice is invisible in the source text, so an expression carries no information about ownership from one tool to another. ## What actually breaks when the assumption travels The values do not change, which is why this is a slow failure rather than a loud one. What changes is memory behaviour: | the code assumed | under a design that shares | under a design that duplicates | |---|---|---| | the result is cheap to take in a loop | cheap, nothing allocated | an allocation per iteration | | the source is released after the loop | held alive by the kept result | released as expected | | the small result is small | small to look at, large in retention | small in both senses | So the same routine can be memory-clean in one setting and pin a source in another, or be allocation-free in one and allocate once per iteration in another. Neither shows up in a test that checks the numbers. ## Commit to the requirement The fix is to write down what the code needs rather than what the tool currently does. There are only a few requirements worth stating: 1. **The source must be releasable after this point.** Then materialise the kept rows into storage of their own at that point. Under a design that already duplicates, this costs a redundant allocation that you can accept for the predictability; under one that shares, it is the thing that makes the release happen. 2. **The result is transient and consumed immediately.** Then state nothing and pay nothing. Sharing is a benefit here, and a blanket rule to copy everything throws it away. 3. **The result crosses an interface** - it is returned, cached, or stored in something long-lived. Then the requirement belongs to the interface: a function that read something large and returns a small piece of it should return a result that does not keep the large thing alive, because its caller cannot see the source and cannot reason about it. ## Why detect-and-branch is the wrong instinct The tempting alternative is to establish what kind of result you were handed and act accordingly. It is the wrong instinct for portable code for three reasons: - It puts a per-design detail into the code, which is what made the routine unportable in the first place. - The answer can be genuinely in between. Under a design that defers the duplicate until something is written, the result is neither a plain window nor fully allocated storage until a modification forces the issue. - It reasons about the situation you are in, rather than removing the dependency on it. Materialising is unconditional and its cost is known. ## The two sentences not to carry into an interview Both of these are commonly said and both state one design as the class: - *Selecting rows always copies, so a selection never shares the original's memory.* True of every gather. False of a stride under the designs that share. - *A contiguous run of rows is always a window.* False wherever the design duplicates eagerly, and only partly true where the duplicate is promised and deferred. The repair is not to drop the claim but to say what it varies with: the shape of the selection decides whether sharing is **possible**, and the design decides whether it **happens**. ## How to answer this in an interview Say that ownership is a design policy that the expression cannot express, give two of the policies concretely, and then give the engineering answer: state the requirement at the boundary and materialise where the requirement is that the source can go away. That answer works in every tool in the family, which is the point of it.
- Why is materialising unconditionally acceptable even where the tool already duplicates?Because the cost is a known, bounded allocation proportional to the rows you are keeping, and you are keeping few - that is the premise. In exchange the routine has one behaviour everywhere instead of two, and the caller gets a promise about retention that does not depend on which tool is underneath.
- Where is sharing worth keeping rather than materialising away?Wherever the result is consumed and dropped without outliving its source: a step inside a loop iteration, an intermediate in a chain, a value read once. Sharing exists to avoid allocations on exactly those, and retention cannot bite when nothing is retained. Materialise at boundaries, not at every step.
saying these in an interview costs you the question
- Calls the difference between tools a bug in one of them
- Assumes memory behaviour ports along with the expression
- Branches on which kind of result the tool returned
- Says one policy is simply safer and should always be preferred
- Copies every result everywhere to avoid thinking about it