How do you choose among XGBoost, LightGBM and CatBoost for a given table?
answer
- same algorithm, different engineering priorities
- constraints first, leaderboard second
- row count against the retrain window
- tens of thousands of categorical levels
- after tuning, gaps sit inside fold variance
basics
~20 sChoose on constraints, not on accuracy claims. LightGBM for training throughput on very large, wide tables; CatBoost for high-cardinality categorical columns and fast uniform-depth inference; XGBoost as a strongly regularized general default. Tuned properly, their accuracies usually land within noise.
solid answer
~40 sAll three are gradient boosted decision trees, so I decide on engineering constraints first. If a nightly retrain window has to absorb a 60-million-row, 2,000-column table, LightGBM's design is built for exactly that throughput. If a marketplace table is dominated by three categorical columns with tens of thousands of levels each, CatBoost handles them natively instead of forcing me to invent an encoding that leaks target information across folds. XGBoost is my default when neither pressure is dominant: a regularized objective with second-order gradient information and mature tooling. Then I run a like-for-like bake-off — same folds, same metric, same early-stopping rule, comparable search budgets — because after tuning the accuracy gap is usually inside fold-to-fold variance. At that point throughput, inference latency and what the team can operate decide it.
go deeper
Know that all three are gradient boosted trees rather than three different algorithms, and that one is known for speed on large data and one for handling categorical columns without manual encoding.
Be able to name what each design buys you — throughput, native categorical handling, a regularized general-purpose objective — and explain why accuracy differences shrink once each is tuned with early stopping.
Show the decision procedure: identify the binding constraint, run a like-for-like bake-off with equal budgets, and report the gap against fold-to-fold variance rather than asserting a winner.
Frame it as a standardisation call: one boosted stack the whole organisation can tune, monitor and reproduce usually beats a per-project pick, and you should be able to say what evidence would justify adding a second.
## They are variants of one algorithm XGBoost, LightGBM and CatBoost are all gradient boosted decision trees: an additive ensemble where each new tree is fitted to the negative gradient of the loss with respect to the current predictions and added with a shrinkage factor. They differ in how they find splits, how they grow trees, how they treat categorical columns and how their objective is regularized. Those differences change *cost and convenience* far more often than they change *accuracy*. That framing matters in an interview, because the weak answer is a ranking ("LightGBM is the best") and the strong answer is a decision procedure. ## What each one is actually optimised for **XGBoost** — a regularized objective that uses second-order (curvature) information about the loss when scoring candidate splits, with explicit L1 and L2 penalties on leaf weights and a learned default direction for missing values at each split. Grows trees level by level by default. It is the conservative general-purpose choice: predictable, heavily regularized, widely operated. **LightGBM** — designed above all for throughput and memory on large data: continuous features are bucketed into a small number of bins so split search scans histograms rather than sorted values, and trees are grown best-first rather than level by level. The practical consequence is that it trains dramatically faster on wide and tall tables. The practical caution is that its growth strategy will happily build very unbalanced trees, so on small data you must cap tree size and require a minimum number of rows per leaf, or it overfits. **CatBoost** — designed around categorical data: it converts categorical levels into numeric statistics computed in a way that avoids letting a row's own target leak into its own feature value, so you do not have to hand-build an encoding. Its trees are symmetric — every node at a given depth uses the same split — which makes prediction unusually fast and uniform in cost. It also tends to be strong at or near its defaults. Deliberately shallow treatment of the internals here: what matters for the choice is what each design *buys you*. ## The decision procedure **1. Start from the constraints, not the leaderboard.** - *Scale and retrain window.* A **60-million-row, 2,000-column table with a nightly retrain window** is a throughput problem before it is an accuracy problem. Binning-based split search and fast parallel histogram building are what make that window achievable, which points at LightGBM. If the model cannot be rebuilt before the batch job that consumes it, its error rate is irrelevant. - *Categorical cardinality.* A **marketplace table dominated by three categorical columns with tens of thousands of levels each** — seller id, product category, postcode — is the case for CatBoost. One-hot encoding explodes the width; hand-rolled target encoding is fine in principle but leaks unless it is computed strictly inside each training fold, and that is a bug people ship constantly. - *Inference budget.* Symmetric trees evaluate at uniform depth with very predictable cost, which matters when you serve per-request rather than in batch. - *Neither dominant.* Take XGBoost as the default and spend the saved time on features and validation. **2. Run a fair bake-off.** Same folds, same metric, same early-stopping rule against the same held-out data, and *comparable* tuning budgets. The single most common way this comparison goes wrong is spending a day tuning your favourite and running the others at defaults — that measures your effort, not the library. **3. Expect a small gap.** On most business tables the tuned results sit within fold-to-fold variance of each other. When that happens, say so and decide on the things that are not within noise: training time, memory, inference latency, reproducibility across versions, and whether the team can debug it at 3am. ## What the choice does not fix None of the three rescues a bad target definition, leakage from a feature computed after the event you are predicting, or a validation split that ignores time or grouping. Candidates who reach for a different library when a model underperforms are usually solving the wrong problem: the leakage audit and the validation protocol come first, and only then does it matter which booster you fit. ## The answer to give "They are the same algorithm with different engineering priorities. I pick on the binding constraint — throughput, categorical cardinality or inference latency — default to the regularized general-purpose option when nothing binds, and then prove the choice with a like-for-like comparison instead of asserting a ranking."
- How do you make the comparison across the three genuinely fair?Fix the folds, the metric and the early-stopping rule, then give each an equal search budget measured in trials or wall-clock, not in whichever one you enjoy tuning. Report mean and spread across folds, not a single best score, and repeat with a couple of seeds. Without that, a reported win is usually a tuning-effort artefact.
- Why is a hand-rolled target encoding for tens of thousands of categorical levels risky?Because the encoded value for a row is computed from targets that include that row, or from rows in the same fold, so the feature carries label information the model will not have at prediction time. Validation then looks great and production does not. Rare levels are also estimated from a handful of rows, so they need smoothing toward the global mean.
- When does inference latency, rather than training cost, decide the choice?When you serve per-request under a tight budget rather than scoring a nightly batch. Then the number of trees, their depth and how uniform the evaluation path is dominate; symmetric trees of fixed depth are attractive because per-row cost is predictable. Measure at the ensemble size you would actually ship, not at defaults.
saying these in an interview costs you the question
- Ranking one library as universally most accurate
- Tuning your favourite hard and running the others at defaults
- Assuming every booster handles high-cardinality categoricals natively
- Ignoring the retrain window and the inference budget
- Switching library instead of auditing leakage and validation
- Treating them as different algorithms rather than one algorithm's variants