skip to content

A colleague says boosted trees are always the best model on tabular data — how do you respond?

level: principalimportance: nice to knowfreq 30%

answer

  1. is it a theorem or a habit?
  2. a claim about which problems
  3. the benchmark sample is not your sample
  4. tuning budget and baselines unstated
  5. name the conditions that void the default

basics

~10 s

No algorithm is best on every problem, so the claim is empirical — a statement about the problems that team has seen. Ask what evidence supports the default and which observable conditions void it.

solid answer

~50 s

I would agree it is a good default and disagree that it is a law. No-free-lunch forbids any algorithm being best on all problems, so the statement can only be a claim about the *distribution* of problems people call tabular — and it is a decent one, because the inductive bias of boosted trees (axis-aligned piecewise-constant splits, insensitivity to monotone feature rescaling, tolerance of irrelevant features) matches how many business datasets are shaped. But the evidence behind it is a sample of public benchmarks, made under a particular tuning budget and against particular baselines, and none of that is guaranteed to resemble our problems. So I would push the conversation from `which algorithm wins` to `which problem distribution are we defaulting for, what evidence do we have from our own past projects, and which observable conditions void the default`.

go deeper

for a junior

Know that popular defaults exist for good empirical reasons, but that 'always best' is never literally true. Being able to ask 'best compared with what, and on which datasets?' is enough at this level.

for a middle

Be ready to explain why the boosted-tree prior fits many business datasets — threshold-like structure, insensitivity to feature rescaling, tolerance of irrelevant columns — and to name a situation where that prior does not apply.

for a senior

Show how you would test the claim on your own data: comparable tuning effort for each contender, a split that respects grouping or time, and a simple baseline that makes the win meaningful rather than assumed.

for a principal

Own the organisational tradeoff between a standing default and a per-project bake-off. Be able to state what evidence licenses the default, which observable conditions void it, and how often that evidence gets refreshed.

## What kind of claim is being made 'Boosted trees always win on tabular data' is not a theorem and cannot be one — the no-free-lunch result says that averaged uniformly over all target functions no algorithm outperforms any other, so a universal winner is ruled out by construction. What the claim can legitimately be is an *empirical generalization about a distribution of problems*: among the datasets people label 'tabular' and care to benchmark, this family wins often. That is a much weaker and much more useful statement, and treating it as the weaker statement is the whole skill here. ## Why the empirical version is largely true The claim has real support, and dismissing it with a slogan is as bad as over-claiming it. The boosted-tree family (gradient boosting and its widely used variants such as XGBoost and LightGBM) carries a prior that matches many business datasets well: - The learned function is piecewise constant over axis-aligned regions, which handles sharp thresholds and non-smooth interactions that a linear form would have to be told about in advance. - Splits depend only on the order of a feature's values, so monotone rescaling of a feature changes nothing. - Uninformative features tend to be ignored, because they rarely produce the best split. - Boosting fits stages sequentially to the residual signal of the current ensemble, which lets a modest number of shallow trees compose fairly intricate structure. When the problems you get look like that, a default from this family is a rational bet and it saves real time. ## Where the claim is quietly weak Four things are usually left unstated: 1. **The sample of problems.** Benchmarks over-represent datasets that were interesting enough to publish. Your organisation's problems are a different sample, and the claim transfers only as far as the resemblance goes. 2. **The tuning budget.** Comparative wins depend on how much effort each contender received. A conclusion drawn where one family was carefully tuned and the alternatives were left at their usual settings is a statement about effort as much as about algorithms. 3. **The baselines.** 'Best' is relative to whatever was in the comparison. A well-specified simple model that nobody ran cannot lose. 4. **The metric and the split.** A win on one metric under a random split can vanish under a split that respects grouping or time ordering, because the family's advantage may partly be the ability to exploit structure that leaks. ## The inductive-bias conditions that void it The honest form of the default names its own escape hatches. The clearest is extrapolation: a piecewise-constant model cannot continue a trend past the range of feature values it saw in training — its prediction saturates at the value of the outermost region. If the target rises with a feature and production inputs run beyond the training range, the prior is simply wrong for the job, no matter what any benchmark says. Similar reasoning applies whenever the problem has structure the family's prior cannot express, or whenever a hard requirement exists that the family's form cannot satisfy. ## What I would actually do At the level of a team rather than a project, the response is a policy, not a rebuttal: - **Make the default explicit and evidence-backed.** Collect the projects you have already done, and check that the default really did win *on those*, tuned comparably. - **Write down the voiding conditions.** Extrapolation beyond the training range, strongly grouped or time-ordered rows, a hard constraint on the model's form — each of these turns the default into a hypothesis to be tested rather than an answer. - **Keep a mandatory cheap baseline.** If every project compares the default against one simple, well-specified alternative, the default can never quietly become an unchallenged rule, and you accumulate exactly the evidence that would tell you when it stops being right. - **Re-run the comparison periodically.** The claim is about a problem distribution; distributions move as the business changes. ## The failure modes on both sides Over-claiming looks like 'we do not need to try anything else'. That is a prior masquerading as a proof, and it goes wrong silently — nothing in the pipeline reports 'a different family would have been better here'. Under-claiming is the mirror error: invoking no-free-lunch to justify benchmarking everything on every project. The theorem says there is no *universal* best; it does not say you have no information. Refusing to hold a well-evidenced prior wastes the most expensive resource on the team. The mature position is a default you can defend, with written conditions under which you stop trusting it.

  • What evidence would justify making one model family your team's standing default?
    A set of past projects that genuinely resembles the incoming work — similar feature types, sizes, label noise and horizons — with each contender tuned under a comparable budget and evaluated on splits that respect the data's structure. That is evidence about your problem distribution rather than someone else's, and it should be regenerated periodically rather than treated as settled.
  • How do you stop a default from hardening into an unquestioned rule?
    Write down what the default assumes and which observable conditions void it, then make one cheap, well-specified baseline mandatory in every project so the default is always compared with something. When a project trips a voiding condition — extrapolation beyond the training range, strongly grouped rows, or a requirement the family's form cannot meet — a real comparison becomes required rather than optional.
  • Someone cites no-free-lunch to argue every project must benchmark every algorithm. What is wrong with that?
    It reads 'no universally best algorithm' as 'no usable prior', which does not follow. The theorem removes a guarantee, not your accumulated evidence about the problems you actually face. Exhaustive benchmarking spends your scarcest resource to rediscover what past projects already showed; the right response is a defensible default plus stated conditions for abandoning it.

saying these in an interview costs you the question

  • Treats a benchmark result as a theorem about all tabular data
  • Cites no-free-lunch to refuse any default and bake off everything
  • Ignores that comparisons depend on tuning budget and baselines
  • Assumes a piecewise-constant model can extrapolate a trend
  • Never asks whether the benchmark problems resemble ours

context