What does FeatureUnion do in scikit-learn, and when is it the wrong tool?
answer
- parallel branches, not a chain
- every branch sees the whole input
- outputs are stacked side by side
- different columns means a different composer
- branch names prefix the output feature names
basics
~20 sFeatureUnion fits several transformers on the same input in parallel and horizontally concatenates their outputs into one feature matrix. It is the wrong tool when each transformer should see a different subset of columns — that is ColumnTransformer's job.
solid answer
~40 sWhere a `Pipeline` composes transformers in series, `FeatureUnion` composes them in parallel: it takes a `transformer_list` of `(name, transformer)` pairs, fits every one on the **whole** input, and horizontally stacks their outputs into a single matrix. `make_union` is the shorthand with auto-generated names. It supports `n_jobs` to fit branches concurrently, `transformer_weights` to scale each block's contribution, and `"drop"` to disable a branch. A branch is itself often a `Pipeline`, so you can build "raw counts plus an SVD projection" in one object. It is the wrong tool whenever the branches are really about different **columns** — text here, numerics there — because every branch receives the full input; `ColumnTransformer` is the composer that routes column subsets. Inside a pipeline you tune a branch through nested names, e.g. `features__pca__n_components`.
code
python · 18 linesfrom sklearn.decomposition import PCA
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import FeatureUnion, Pipeline
union = FeatureUnion(
transformer_list=[
("pca", PCA(n_components=5)),
("kbest", SelectKBest(f_classif, k=3)),
],
transformer_weights={"pca": 1.0, "kbest": 2.0},
n_jobs=1,
)
pipe = Pipeline([("features", union), ("clf", LogisticRegression(max_iter=1000))])
# Tune inside a branch, or drop a branch entirely:
grid = {"features__pca__n_components": [2, 5], "features__kbest": ["drop"]}go deeper
Remember the one-liner: a pipeline chains transformers in series, a union runs them in parallel on the same input and glues the outputs together side by side.
Explain the API — transformer_list, n_jobs, transformer_weights, "drop" — and be clear that every branch sees the whole input, which is why column-subset work belongs to a different composer.
Show judgment about cost: the feature count is additive across branches, sparse and dense branches interact badly, and nested names let you search over whether a whole feature family earns its place.
Own the feature-composition architecture — which feature families are worth maintaining as separate branches, how their cost and interpretability scale, and when a union should be split into independent models instead of one wide matrix.
## Series versus parallel Scikit-learn has two composition primitives with different shapes. `Pipeline` is **series**: step one's output is step two's input. `FeatureUnion` is **parallel**: every branch gets the same input, and their outputs are glued side by side. ``` FeatureUnion([("pca", PCA(n_components=5)), ("kbest", SelectKBest(k=3))]) ``` Both branches are fitted on the full matrix; the result has 5 + 3 = 8 columns. `fit_transform` on the union calls `fit_transform` on each branch and `numpy.hstack`s (or sparse-hstacks) the pieces; `transform` calls each branch's `transform`. Because a union is itself a transformer, it slots into a pipeline as one step, and because a pipeline is a transformer when its last step is one, a branch can itself be a pipeline. That nesting is where the expressiveness comes from. ## The arguments worth knowing - **`transformer_list`** — the ordered `(name, transformer)` pairs. Names must be unique and are used both for the output column-name prefixes and for parameter addressing. - **`n_jobs`** — fits and transforms branches in parallel processes. It helps when branches are genuinely heavy and independent; for small branches the process overhead dominates. - **`transformer_weights`** — a dict mapping branch name to a multiplier applied to that branch's output block. It is a crude way to control the relative influence of blocks in a distance-based or regularised model, and it is a hyperparameter you can search over. - **`"drop"`** — assign this string in place of a branch's transformer to disable that branch, which makes "is this feature family worth including?" a searchable question. - **`make_union(t1, t2)`** — builds the union with lowercased class names as branch names. Output naming: `get_feature_names_out()` on the union prefixes each branch's names with the branch name, so you get `pca__pca0`-style names rather than a nameless block — worth knowing when you need to explain a model's coefficients. Fitting passes `y` through: `union.fit(X, y)` gives each branch's `fit` the labels, which is what allows a supervised branch such as feature selection to work inside a union. ## When it is right The honest use case is **multiple views of the same input**. Text is the classic: word-level TF-IDF and character-level TF-IDF over the same documents, concatenated. Or a numeric block kept raw next to a low-rank projection of it, so the model gets both. Or a domain-specific extractor next to a generic one. In all of these, "the same input" is literally true — each branch legitimately consumes everything it is given. ## When it is wrong The frequent misuse is column routing. People reach for `FeatureUnion` to say "scale the numeric columns and encode the categorical ones", then discover that both branches receive the whole frame and each one has to select its own columns first. That works — historically people wrote a small column-selecting transformer as the first step of each branch — but it is exactly the job `ColumnTransformer` was added to do, with column selection declared once, by name or dtype, and remainder handling built in. If your branches begin by slicing columns, you want the column-routing composer, not the union. The second smell is dimensionality: unions concatenate, so a union of several expansive transformers multiplies the feature count, and every downstream cost — memory, fit time, regularisation strength — scales with it. Sparse branches keep the result sparse, which is why text unions stay tractable; mixing a dense branch into a sparse union densifies everything and can blow up memory. ## Tuning through a union Because the union exposes its branches as parameters, nested addressing composes. In a pipeline `Pipeline([("features", union), ("clf", LogisticRegression())])`, a grid key reaching a parameter inside a branch reads `features__pca__n_components`, and `features__kbest: ["drop"]` removes that branch for a set of candidates. Weights are addressable too, as `features__transformer_weights`. ## The take-away line for an interview "`Pipeline` is series, `FeatureUnion` is parallel over the same input, `ColumnTransformer` is parallel over different columns." Say that, then give one real example of each, and the follow-ups tend to be about tuning names or about sparsity — both of which you now have answers for.
- How do you tune a parameter that lives inside one branch of a union?Compose the names: if the union is a pipeline step called `features` and the branch is called `pca`, the key is `features__pca__n_components`. Each double underscore descends one level, and the same addressing lets you set a branch to `"drop"` to search over whether that feature family helps at all.
- Does FeatureUnion pass the labels to its branches?Yes — `fit(X, y)` forwards `y` to every branch's `fit`, which is why a supervised branch such as feature selection works inside a union. Each branch still sees the same `X`; the union only fans out the input and concatenates results, it does not partition anything.
- What happens to sparsity when branches produce different output types?If every branch returns sparse output the union stays sparse, which is what keeps text unions affordable. Mixing a dense branch in forces the concatenation to be dense, and a large sparse text block densified that way can exhaust memory. Keep sparse and dense families in separate pipelines, or densify deliberately.
- What do the output feature names look like after a union?`get_feature_names_out()` prefixes each branch's names with that branch's name, so you can tell which block a coefficient belongs to. That prefixing is a good reason to give branches short, meaningful names — the names travel all the way into model interpretation output.
saying these in an interview costs you the question
- Each branch receives a different slice of the columns
- FeatureUnion runs branches in sequence like a pipeline
- It is interchangeable with ColumnTransformer
- Concatenating extra views is always free
- Branch parameters cannot be tuned in a search