skip to content

Why hand-build a debt-to-income ratio when the model already has both raw columns?

level: middleimportance: must knowfreq 62%

answer

  1. coefficients add columns, they never divide them
  2. which functions can the model reach?
  3. a diagonal boundary versus axis-aligned splits
  4. logs turn a ratio into a difference

basics

~20 s

A quotient is not a linear combination of its two columns, so no coefficients on debt and income reproduce debt divided by income. A tree can only approximate it with a staircase of axis-aligned splits, which costs depth and data.

solid answer

~50 s

A linear model on debt and income can only produce `w0 + w1*debt + w2*income` — a plane. The quantity that actually drives default risk, `debt / income`, is not in that family, so no coefficient setting recovers it. Trees are not much better off: constant-ratio contours are straight lines through the origin, and axis-aligned splits can only approximate a diagonal boundary with a staircase, spending a split and a chunk of data on every step. Supplying the ratio directly turns that boundary into one split. The same argument gives price per square metre in a property table. Two cautions: guard the denominator, because near-zero income sends the ratio to enormous values that dominate a squared-error fit, and keep the raw columns as well, since the ratio alone throws away absolute scale — a debt-to-income of 0.3 means something different at 20,000 and at 2,000,000.

go deeper

for a junior

Be ready to say that a linear model can only add up its columns after multiplying each by a number, so it can never divide one by another — that is why the ratio has to be built by hand.

for a middle

Explain the mechanics both ways: the function family a linear model can reach, and why a tree needs a staircase of axis-aligned splits to approximate a diagonal ratio boundary.

for a senior

Demonstrate the production care — denominator guards fitted on training data and reapplied identically at scoring, whether to keep the parent columns, and what collinearity does to a linear model's coefficients.

for a principal

Own the discipline question: a small set of domain-meaningful ratios that stakeholders already reason with, rather than a generated sweep that widens the table and buys nothing you can explain.

## What the model can and cannot reach Every model has a family of functions it can express. A linear model in two columns `d` (debt) and `i` (income) can produce anything of the form `w0 + w1*d + w2*i`. Geometrically that is a plane over the (d, i) space, and its contours — the sets of points that get the same prediction — are parallel straight lines. The function `d / i` is not in that family. Its contours are straight lines through the *origin*, fanning out at different angles: every point with debt equal to a third of income gets the same value, whether that is 10,000 on 30,000 or 200,000 on 600,000. No plane has fanning contours. So the honest statement is not "the model would find it slowly" — it is **cannot, at any sample size, with any coefficients**. That is the whole case for hand-built ratios. A ratio is a *normalisation*: it says the meaningful quantity is one column relative to another, not either alone. ## Why trees do not rescue you A tree can, in principle, approximate any function, so the argument is quantitative rather than absolute. A tree splits on one column at a time, `d < t` or `i < t`, so its decision regions are axis-aligned rectangles. The boundary `d = 0.4 * i` is a diagonal line, and a rectangle-based partition can only approximate a diagonal with a staircase. Each step of that staircase costs a split, and each split needs enough rows beneath it to be estimated reliably. With the ratio supplied as a column, the same boundary is a single split at `ratio < 0.4`. That is the difference between spending depth and data on rediscovering arithmetic and spending it on the parts of the problem that are genuinely hard. ## Ratios versus differences — not the same argument It is worth being precise here, because interviewers probe it. A **difference** such as `ship_date - order_date` *is* a linear combination: coefficients of +1 and -1 on the two columns produce it exactly. So a linear model can already represent fulfilment days from the two raw dates. The reasons to build the column anyway are different in kind: - **Trees cannot form it.** Axis-aligned splits on two date columns cannot express "the gap exceeded five days" any more than they can express a ratio. - **Regularisation and limited data.** A penalised linear model shrinks coefficients toward zero; the exact +1/-1 pairing that encodes a difference is one point in a large space, and with few rows the fit may never land near it. - **The raw columns carry a trend that the difference does not.** Order date drifts forward forever; fulfilment days do not. The derived column is the stable, extrapolable one. A useful bridge between the two cases: on the log scale a ratio *becomes* a difference, since `log(d) - log(i)` equals `log(d / i)`. So if both columns are already log-transformed, a linear model can reach the log-ratio on its own. ## Doing it safely **Guard the denominator.** Near-zero income, or a property listing with an area of 0.5 square metres, produces a value orders of magnitude above the rest of the column. Under squared-error loss, one such row can pull the fit visibly. Choose a rule and apply it identically at training and at scoring time: floor the denominator, add a small constant, clip the ratio at a training-set percentile, or emit a separate indicator column for the degenerate rows and let the model handle them apart. Zero denominators must never reach the model as infinities. **Decide whether to keep the parents.** The ratio deliberately discards absolute scale. Price per square metre says nothing about whether a listing is a 30-square-metre studio or a 300-square-metre house, and those are different markets. Keeping price, area *and* the ratio gives the model both views. For a tree ensemble this costs little beyond spreading importance across correlated columns. For a linear model, the ratio is correlated with its parents, which inflates coefficient variance and makes individual coefficients unstable — regularise, or validate which subset actually helps. **Keep it domain-driven.** Good ratios are the ones practitioners already reason with: debt-to-income, price per square metre, revenue per employee, cost per acquisition. They are predictive *and* explainable, which makes the model easier to defend. Generating every pairwise ratio in a wide table is the opposite: it multiplies columns without adding understanding, and most of the new columns can only fit noise. ## The interview summary Ratios encode a relationship no linear model can reach and no tree can reach cheaply. Build the handful the domain justifies, protect the denominator, keep the parents unless validation says otherwise, and be able to state precisely why a difference is a different case from a quotient.

  • Does the same argument hold for a difference like ship date minus order date?
    Not exactly. A difference is a linear combination — coefficients +1 and -1 — so a linear model can already represent it from the two raw columns. The gains are elsewhere: a tree cannot form it from axis-aligned splits, regularisation and small samples make the exact pairing hard to land on, and the raw dates drift forward while the gap does not.
  • How do you handle rows where income is zero?
    Pick an explicit rule and fit it on training data only: floor the denominator, add a small constant, clip the ratio at a training percentile, or add an indicator column marking the degenerate rows. Apply the identical rule at scoring time. What must never happen is an infinity or a value orders of magnitude above the rest reaching a squared-error fit.
  • Should you keep debt and income once the ratio is in the table?
    Usually yes, because the ratio deliberately drops absolute scale, and 30% of a small income is not the same risk as 30% of a large one. For tree ensembles the extra correlated columns cost little. For a linear model the collinearity inflates coefficient variance, so regularise or let a validated comparison decide which subset to keep.
  • Why not just generate every pairwise ratio in the table and let the model sort it out?
    The count grows roughly with the square of the number of columns, so a modest table becomes a very wide one where most columns can only fit noise. It also destroys interpretability and creates many degenerate denominators to guard. A handful of ratios the domain already reasons with beats an exhaustive sweep on almost any real table.

Giving a model debt and income and hoping it invents the ratio is like giving someone the numerator and denominator on separate cards and expecting them to feel the fraction — the value only exists once you actually divide.

saying these in an interview costs you the question

  • The model will discover the ratio itself if it matters
  • Dividing without checking for zero or near-zero denominators
  • Treating a ratio and its two parent columns as independent
  • Generating every pairwise ratio in the table
  • Claiming a quotient is just a linear combination of its columns

context