skip to content

What does an OLS regression coefficient mean when it is described as holding the other predictors fixed?

level: middleimportance: must knowfreq 78%

answer

  1. compare like with like
  2. other predictors sit at the same values
  3. outcome units per predictor unit
  4. residualise the predictor, then regress

basics

~20 s

An OLS slope is the average difference in the outcome per one-unit increase in that predictor, among observations that share the same values of the other predictors in the model. It is a comparison, not an intervention.

solid answer

~50 s

In a fitted model `y_hat = b0 + b1*x1 + ... + bk*xk`, each `bj` is a partial effect: the average change in the outcome for a one-unit increase in `xj` while the other predictors in the model stay at whatever values they already have. Its units are outcome units per predictor unit, so a price-on-square-footage slope is dollars per square foot. In a house-price model with square footage and bedrooms, the bedrooms coefficient compares two houses of the same square footage that differ by one bedroom, which means smaller rooms rather than a bigger house, so it can come out near zero or negative. The coefficient is defined relative to the model's predictor list: change which predictors are included and the number changes meaning. And "held fixed" is arithmetic within the fit, not something anyone performed.

go deeper

for a junior

Be ready to state the one-liner: a slope is the average change in the outcome per one unit of that predictor with the others unchanged, and to say its units out loud as outcome units per predictor unit.

for a middle

An interviewer expects you to explain that the comparison is like-for-like within the model's predictor set, and to work a concrete case such as bedrooms with square footage held fixed without stumbling.

for a senior

Show that you check support before quoting a partial effect, that you report the held-fixed set and an interval alongside the point estimate, and that you can explain why the number moved when the specification changed.

for a principal

Own the framing risk: decide what the team is allowed to say publicly about a coefficient, set a house standard for how partial effects are written up, and push back when a comparison is presented as an action.

## The partial-effect reading Ordinary least squares fits `y_hat = b0 + b1*x1 + b2*x2 + ... + bk*xk` by choosing the coefficients that minimise the sum of squared differences between observed and fitted outcomes. Each slope `bj` is the estimated average change in the outcome associated with a one-unit increase in `xj` **when the other predictors in the model do not change**. Formally it is the partial derivative of the fitted surface with respect to `xj`: everything else in the equation is a constant while you move `xj` by one. This is why a multiple-regression slope is called a *partial* effect and why a simple one-predictor slope is usually a different number. A simple slope lets everything correlated with the predictor drift along with it; a multiple-regression slope does not. ## Always say the units The units of `bj` are **outcome units per predictor unit**. Price in dollars regressed on floor area in square feet gives dollars per square foot. Test score in points regressed on study time in hours gives points per hour. A coefficient with no units attached is not an interpretation, it is a number. Speaking the units aloud also catches nonsense fast: "each extra square foot is worth 210 dollars" is obviously wrong for a coefficient of 210 if the outcome is in thousands of dollars. ## The bedrooms-and-square-footage example Regress house price on square footage and number of bedrooms. People read the bedrooms coefficient as "what an extra bedroom is worth". It is not. It is the average price gap between two houses **of the same total square footage** that differ by one bedroom. Total space is fixed by construction, so the extra bedroom is carved out of the existing space: the same house with more, smaller rooms. That comparison is genuinely different from the one people imagine (bolting an extension onto the house, which would change square footage too), and it is why such a coefficient is often small, sometimes negative, and always a surprise to a stakeholder who has not been told which comparison is on offer. The general lesson: a partial effect answers a question about **like-for-like comparisons within the model's predictor set**, and you have to state that set out loud for the number to mean anything. ## What the arithmetic actually does The Frisch-Waugh-Lovell theorem makes this concrete. Regress `xj` on all the other predictors and keep the residuals: that is the part of `xj` the other predictors cannot explain. Now regress the outcome on those residuals. The slope you get is exactly `bj` from the full model. So a multiple-regression coefficient is estimated from the variation in that predictor that remains after the other predictors have taken their share. Two consequences follow immediately. First, if the other predictors explain most of `xj`, very little variation is left and the estimate is imprecise, with a wide standard error. Second, the coefficient is a property of *this* model specification: the same data with a different set of controls produces a different residualised predictor and therefore a different number. "The coefficient on tenure" is not a fact about the data; it is a fact about a model. ## Holding fixed is not intervening The fit compares observations that already happen to share values on the other predictors. Nobody set anything. Whether the resulting comparison also answers a causal question - what would happen if you *changed* the predictor - depends on assumptions about the data-generating process that live entirely outside the arithmetic of the fit, and it is a separate topic from reading the number. There is also a support question. "Holding square footage fixed while adding a bedroom" describes a comparison that must actually exist in the data. If no 600-square-foot house in the sample has four bedrooms, the model's prediction for that combination is extrapolation driven by the assumed linear form, not evidence. Before quoting a partial effect, check that the combination of predictor values you are describing is populated. ## Reporting it well A defensible sentence has four parts: the size of the change you are describing, the units, the set of predictors held fixed, and the uncertainty. "Among houses of the same square footage, each additional bedroom is associated with about 4,000 dollars less in price, with a 95% interval from 9,000 less to 1,000 more." That is honest, checkable, and immune to the usual misreadings. One caveat on scope: this single-coefficient reading assumes the predictor enters the model in exactly one term. If a predictor appears in several terms at once, the effect of moving it is spread across those terms and no single coefficient carries it.

  • In a price model containing square footage and bedrooms, what exactly does the bedrooms coefficient compare?
    Two houses with the same total square footage that differ by one bedroom. Because the footage is held fixed, the extra bedroom is carved out of existing space, so the comparison is more and smaller rooms rather than a larger house. That is why the coefficient is often small or negative, and why it does not answer "what is an extra room worth".
  • How does the Frisch-Waugh-Lovell theorem describe the same coefficient?
    Regress the predictor on all the other predictors and keep the residuals, then regress the outcome on those residuals. The slope equals the multiple-regression coefficient. So the estimate uses only the variation in that predictor the others cannot explain, which is why heavily overlapping predictors give imprecise coefficients.
  • When does "holding the others fixed" describe a comparison the data cannot support?
    When that combination of predictor values does not occur in the sample. If every large house has many bedrooms, the model still reports a bedrooms slope, but the comparison at a small footage with many bedrooms is extrapolation from the assumed linear form. Check the joint support before quoting the number.
  • Why can the same predictor's coefficient change when you add another predictor to the model?
    Because the coefficient is defined relative to the predictor set. Adding a predictor changes what is being held fixed and removes some of the first predictor's variation from the estimate. The two numbers answer different comparison questions, so neither is wrong on its own terms - they must be reported with their model.

It is like comparing two runners' times only among races run on the same course in the same weather: the number answers a same-conditions comparison, not what happens if you change the weather.

saying these in an interview costs you the question

  • Says the coefficient is what happens if you change that predictor for everyone
  • Quotes a coefficient with no units attached
  • Reads a coefficient without naming which predictors are in the model
  • Assumes every combination of held-fixed values exists in the data
  • Treats a multiple-regression slope as identical to the simple one-predictor slope

context