Extract 100 projected components or select 100 of 1,000 original columns for a latency budget?
answer
- ask where the latency actually is
- combinations of all versus a subset
- extraction still needs all 1,000 inputs
- named columns survive monitoring
- fit the reduction inside the folds
basics
~20 sSelection cuts what you must collect and keeps every input explainable; extraction keeps signal spread thinly across all 1,000 columns but still requires computing every one of them at serving time. If the latency is upstream, only selection helps.
solid answer
~40 sFirst establish where the latency actually goes. Extraction builds 100 new columns as combinations of all 1,000 originals, so at serving time you still fetch, clean and compute every original column and then pay a matrix multiply on top — it narrows the model, not the pipeline. Selection keeps 100 real columns and lets you stop producing the other 900 entirely, which is the only option that removes upstream cost. Against that, on wide sparse data the signal is often diffuse, so dropping 900 columns can discard power a projection would have retained. Selection also keeps every input nameable, which matters for monitoring, drift attribution and any explanation you owe a regulator. My default is to select first, measure the loss, and reach for extraction only if that loss is material.
go deeper
Be ready to state the difference cleanly: selection keeps some of the original columns untouched, extraction builds new columns out of all of them, and only the first reduces what data you have to collect.
Explain the accuracy side — diffuse signal across many weak columns favours extraction — and the discipline point that the reduction must be fitted inside the training folds or the comparison is inflated.
Show that you decompose the latency budget before choosing, and that you can name the operational costs you would inherit: drift you cannot attribute, incidents you cannot trace, and models coupled to a projection version.
Own the decision as a written tradeoff — which of latency, metric, explainability and drift attribution you optimised and which you accepted — and be willing to defend a hybrid or a fixed random map as the deliberate choice.
## The two things being compared **Feature extraction** creates new features as functions — usually linear combinations — of the originals. Supervised discriminant axes, variance-maximising components and random projections all belong here. The output columns are new objects that do not correspond to anything in the source data. **Feature selection** keeps a subset of the original columns unchanged and discards the rest. The output columns are the same objects that went in, with the same names, units and provenance. Both can take 1,000 columns down to 100. They differ in almost every consequence that follows. ## Where the latency actually lives This is the first question to ask, and the one candidates skip. Decompose the serving budget: - **Upstream cost** — fetching each feature from its store, joining, cleaning, computing derived values. Proportional to how many *original* features you still need. - **Transform cost** — the projection multiply, if you extract. - **Model cost** — inference, proportional to the model's input width. Extraction reduces only the third. It leaves the first untouched — every one of the 1,000 originals must still be produced to compute the 100 components — and it *adds* the second. If your budget is being eaten by feature fetches and joins, extraction can make the pipeline slower while making the model narrower. Selection is the only one of the two that removes upstream work, because the 900 dropped columns need never be computed, stored or fetched again. ## What each costs in accuracy On wide, sparse data the signal is often spread thinly across many weak columns. Keeping 100 of 1,000 discards whatever the other 900 collectively carried, and no amount of care in choosing which 100 recovers a signal that only exists in aggregate. A projection distributes the contribution of all 1,000 columns across its outputs and can preserve that diffuse signal, which is why extraction usually wins on pure held-out metric for this data shape. The honest way to settle it is to measure both under identical cross-validation with the reduction fitted inside each training fold — fitting a projection or running a selection procedure on the full dataset before splitting leaks information and inflates both estimates, with the selection side typically inflating more dramatically. ## What each costs operationally - **Interpretability.** A selected column has a name, a unit and an owner. A component is a weighted blend of a thousand things and explains nothing. If you owe a customer or a regulator a reason for a decision, working from components means bolting on a post-hoc explanation method rather than reading the model. - **Monitoring and drift.** When a selected feature drifts you know which upstream source broke and who to call. When a component drifts you know only that something among a thousand inputs moved. Attribution is the operational cost people underestimate most. - **Refit coupling.** Refitting a data-dependent projection changes what every component means, so the downstream model must be retrained in lockstep and old models cannot score new data. Selection has a milder version of the same problem — the chosen subset changes — but each surviving column keeps its meaning. A random projection avoids the coupling entirely because its matrix never changes. - **Data acquisition.** If some of the 1,000 columns are bought, rate-limited, or expensive to compute, selection lets you stop paying for the ones you dropped. Extraction commits you to all of them forever. - **Debuggability.** A bad prediction traced to a named feature with an implausible value is a fixable incident. The same prediction traced to component 37 is not. ## A defensible answer Start with selection, because it is the only choice that pays back the upstream budget, keeps the system explainable and keeps incidents debuggable. Measure the held-out loss against the full 1,000-column model. If the loss is small, you are done and the system is materially simpler. If the loss is material, examine whether the constraint is genuinely upstream: if the bottleneck is model width or downstream compute, extraction becomes viable, and if the fitting and refit coupling is what worries you, a fixed random projection buys the width reduction without a data-dependent map to version. A hybrid is often the real answer — drop the columns that are genuinely dead or unaffordable, then project what remains — and it should be presented as an explicit choice, not as a compromise nobody decided on. Whatever you choose, write down which of the four costs (latency, metric, explainability, drift attribution) you optimised and which you accepted, because that is the artefact the next engineer needs.
- Which choice reduces the cost of the feature pipeline itself, and why only that one?Selection. The 900 dropped columns never have to be fetched, joined, cleaned or computed again, so the saving is upstream and permanent. Extraction computes its components from all 1,000 originals, so every one must still be produced at serving time, and the projection multiply is added on top. Extraction narrows the model; only selection narrows the pipeline.
- How would you compare the two options without fooling yourself?Cross-validate with the reduction fitted inside each training fold, never on the full dataset — a projection fitted or a subset chosen before splitting has seen the held-out labels and inflates the estimate. Compare both against the full 1,000-column model on the same folds, and report the latency measured end to end, including feature fetch, not just model inference time.
- What happens to monitoring if you ship the projected version?Drift alerts lose their address. A component that shifts tells you something among a thousand inputs moved but not which one, so triage means recomputing per-source statistics you did not plan to keep. If you ship extraction, keep monitoring the original inputs independently of the model's inputs, and budget for the extra storage that implies.
- When is a fixed random projection preferable to a fitted one here?When the coupling of refits is the real problem. A fitted projection changes the meaning of every component when it is refit, so old models cannot score new data and everything downstream retrains together. A randomly drawn matrix is regenerated from a seed, never changes, and works identically on every worker — you trade some fidelity per dimension for a stable, versionable contract.
Selection is cancelling 900 subscriptions. Extraction is still paying for all 1,000 magazines and hiring someone to write you a 100-page digest.
saying these in an interview costs you the question
- Says extraction always wins because it keeps more information
- Forgets extraction still requires computing every original column at serving
- Fits the projection or picks the subset before splitting the data
- Treats loss of interpretability as a free cost
- Never asks where the latency budget is actually spent
- Ignores that projected components make drift impossible to attribute