In scikit-learn, how does ColumnTransformer apply transforms per column, and what happens to unlisted columns?
answer
- different transforms, different columns
- one container, one fit
- a triple: name, transformer, columns
- columns you forgot to list
- remainder decides their fate
basics
~20 sColumnTransformer fits each listed transformer on its own column subset and concatenates the results side by side in the order the transformers are declared. Columns you did not list are dropped, because remainder defaults to 'drop'.
solid answer
~40 s`ColumnTransformer` takes a list of `(name, transformer, columns)` triples. Each transformer is fitted only on its own slice of the input and the transformed blocks are concatenated horizontally, in the order the triples are listed — so the output column order follows the transformer list, not the input DataFrame. The parameter that surprises people is `remainder`: it defaults to `'drop'`, so any column you did not mention is silently discarded from the output. Set `remainder='passthrough'` to append the untouched columns at the end, or pass an estimator to transform them. Columns can be selected by name, by position, by boolean mask, or by dtype via `make_column_selector`. Afterwards `get_feature_names_out()` gives the output names, prefixed with the transformer name (`num__age`) because `verbose_feature_names_out` defaults to `True`, and `set_output(transform='pandas')` returns a DataFrame instead of an array.
code
python · 14 linesimport pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
df = pd.DataFrame({"age": [30, 40], "city": ["A", "B"], "note": ["x", "y"]})
ct = ColumnTransformer(
transformers=[
("cat", OneHotEncoder(handle_unknown="ignore"), ["city"]),
("num", StandardScaler(), ["age"]),
],
remainder="drop",
)
ct.fit(df)
print(ct.get_feature_names_out()) # cat block first, 'note' is gonego deeper
Know that ColumnTransformer lets you scale numeric columns and encode categorical ones in a single object, and that you list the columns each transformer applies to.
Be able to state the (name, transformer, columns) triple, that remainder defaults to 'drop', and that output order follows the transformer list rather than the input frame.
Show that you debug these: get_feature_names_out and set_output(transform='pandas') for interpretability, sparse_threshold flipping the container type, and named_transformers_ for inspecting a fitted encoder's vocabulary.
Own the schema contract between the feature source and the model: which columns are declared, how a new column is admitted, and whether a silent drop is acceptable or should fail loudly in your training pipeline.
## The problem it solves Real tabular data is heterogeneous: numeric columns want scaling, categorical columns want encoding, a text column wants vectorising, and some columns want nothing at all. A single transformer cannot express that, and doing it by hand in pandas before the split is the most common source of leakage in a project. `ColumnTransformer` is the container that applies different transformers to different column subsets as one fitted object, so the whole thing can sit inside a pipeline and be fitted once on training data. ## How it is constructed The first argument is a list of `(name, transformer, columns)` triples. - `name` is a string used for parameter addressing and output-name prefixes; it must be unique. - `transformer` is any transformer object, or the string `'drop'` to discard those columns explicitly, or `'passthrough'` to keep them untouched. - `columns` may be a list of column names, a list of integer positions, a slice, a boolean mask, a single string, or a callable. `make_column_selector(dtype_include=...)` is the callable scikit-learn ships for selecting by dtype, which is how you say 'all numeric columns' without enumerating them. `make_column_transformer(...)` is the shorthand constructor that generates the names for you from the transformer class names. ## The two behaviours that catch people **Unlisted columns vanish.** `remainder='drop'` is the default. Add a column to your data, forget to add it to the transformer list, and it is simply gone from the model's input — no error, no warning, just a quieter model. `remainder='passthrough'` keeps them, appended after all the transformed blocks. You can also pass an estimator (for example `SimpleImputer()`) as `remainder` to apply one transform to everything not otherwise claimed. **Output column order changes.** The blocks are concatenated in the order the transformers are declared, each block internally ordered as that transformer emits. So a DataFrame with `age, city, income` transformed by a list that puts the `OneHotEncoder` on `city` first produces the city indicators before the scaled numeric columns. If any downstream code indexes the output by position — plotting coefficients, slicing a feature block — it must derive positions from `get_feature_names_out()` rather than from the input frame. ## Names, output types and inspection `get_feature_names_out()` returns the full output names. `verbose_feature_names_out=True` (the default) prefixes each with its transformer name and a double underscore, giving `num__age`, `cat__city_Berlin`; set it to `False` for bare names, which raises an error if that produces duplicates. `set_output(transform='pandas')` makes the container and its children return DataFrames, which is the single biggest quality-of-life improvement for debugging a preprocessing stack; `'polars'` is also supported. The output container type is decided by `sparse_threshold` (default `0.3`): if the fraction of sparse columns in the concatenated result exceeds it, you get a sparse matrix, otherwise everything is densified. That is why adding one wide `OneHotEncoder` block can flip your whole output from a NumPy array to a CSR matrix, and code downstream that calls array-only methods then breaks. After fitting, `named_transformers_` gives dictionary access to the fitted children (`ct.named_transformers_['cat'].categories_`), and `transformers_` is the resolved list including the remainder entry. ## Fitting semantics Each transformer sees only its own columns, and they are independent — `n_jobs` can fit them in parallel. Fitting happens on whatever data is passed to `ColumnTransformer.fit()`, which is why the container belongs inside a pipeline: then each cross-validation fold fits the encoders and scalers on that fold's training part only, and a category or a mean from the validation part cannot reach the model. ## How to answer Describe the triple, state plainly that unlisted columns are dropped by default, and mention that output order follows the transformer list. If you add `make_column_selector` for dtype-based selection and `get_feature_names_out()`/`set_output` for keeping the result interpretable, you sound like someone who has maintained one of these rather than copied it from a tutorial.
- You add a new column to your training data and the model gets slightly worse. How could ColumnTransformer be involved?If the new column is not listed in any triple, `remainder='drop'` discards it silently, so the model never sees it. More subtly, if it *is* picked up by a shared selector, the output width changes and any downstream code indexing by position now points at different features. Check `get_feature_names_out()`.
- How do you select all numeric columns without enumerating them?Pass `make_column_selector(dtype_include='number')` as the columns entry; it is a callable evaluated against the input DataFrame at fit time. `dtype_exclude` works the same way. Be aware it re-evaluates on whatever frame it is given, so a dtype change between train and serving silently changes the selection.
- Why does your ColumnTransformer sometimes return a sparse matrix and sometimes a dense array?`sparse_threshold` (default 0.3) decides: if the sparse fraction of the concatenated output exceeds it, the result stays sparse, otherwise it is densified. Adding a wide `OneHotEncoder` block can flip the whole output. Set `sparse_threshold=0` to force dense, or use `set_output(transform='pandas')`.
saying these in an interview costs you the question
- Assuming unlisted columns are passed through by default
- Expecting output columns in the original input order
- Indexing the output by position instead of feature names
- Fitting the ColumnTransformer on the full dataset before splitting
- Thinking each transformer sees the whole feature matrix