skip to content

Is an R-squared of 0.04 ever good enough to ship a regression model?

level: principalimportance: should knowfreq 41%

answer

  1. no universal threshold exists
  2. explanation and prediction carry different bars
  3. the outcome's noise sets the ceiling
  4. a very high value deserves suspicion
  5. the decision defines the acceptance criterion

basics

~20 s

Yes, when the model's job is to identify and size drivers in a noisy outcome rather than to predict individuals. The acceptance bar comes from the decision the model supports, not from any fixed R-squared threshold.

solid answer

~50 s

R-squared has no universal pass mark; the bar is set by what the model is for. A churn-driver model at 0.04 can be genuinely valuable, because individual churn is close to unpredictable while a stable, well-sized coefficient on price or tenure can still redirect a retention budget. The same 0.04 is unacceptable if the output is a per-account score someone acts on one case at a time. Symmetrically, 0.95 on business data should trigger an investigation rather than a celebration - the usual cause is a predictor that repackages the outcome, such as an amount computed from the thing being predicted or recorded after it. So my framing to stakeholders is: settle first whether we need explanation or prediction, agree what fit level would actually change the decision, and treat the fit statistic as one input alongside coefficient stability and whether the inputs even exist at decision time.

go deeper

for a junior

Do not quote a magic cutoff. Say that acceptable fit depends on how noisy the outcome is and on what the model will be used for.

for a middle

Explain concretely why a small share of explained variation can coexist with a useful, precisely estimated coefficient, and why a near-perfect fit is a warning sign in business data.

for a senior

Show the investigation you run at both extremes: what is driving a suspiciously high value, and what a low one limits you to claiming.

for a principal

Own the acceptance criterion itself - define it from the decision and its costs before modelling starts, and defend it when stakeholders anchor on a number they read somewhere.

## There is no threshold, and saying so is the answer The question is a test of whether a candidate imports a remembered cutoff or reasons from purpose. Any number quoted as a universal minimum - 0.5, 0.7, 0.8 - is arbitrary. R-squared is the share of the outcome's variation the model accounts for, and how much variation is available to explain is a property of the outcome, not of the modeller's skill. ## The ceiling belongs to the outcome Some outcomes are intrinsically almost unpredictable at the level of the individual unit. Whether a particular customer cancels next month, whether a particular visitor converts, whether a particular employee resigns - these are dominated by circumstances that are not in any dataset. In such domains no model reaches a high R-squared, and demanding one is demanding that reality be different. Meanwhile in engineered or physical settings, where the outcome is a near-deterministic function of measurable inputs, 0.95 is unremarkable and 0.6 would suggest something is wrong. The same number carries opposite verdicts in the two settings. ## Explanation and prediction have different bars Separating the two uses does most of the work: - **Explanation / driver analysis.** The deliverable is a set of coefficients with their uncertainty: which levers move the outcome, in which direction, and by roughly how much. What must be defensible here is the coefficient - its sign, its magnitude, its stability under reasonable changes to the specification, and whether confounding could plausibly explain it. A model can explain 4 percent of individual variation and still say, credibly, that a 5 percent price increase costs about so many points of retention. That statement can redirect real money. - **Prediction / scoring.** The deliverable is a number attached to a specific unit that triggers an action. Here the outcome-level variation the model fails to capture is exactly the error someone will act on, and 0.04 means the score is close to noise at the individual level. Shipping it would mean intervening on essentially randomly chosen accounts. So 'is 0.04 enough?' cannot be answered without knowing which of those two the model is. ## The other side: suspiciously high fit A leadership answer is symmetric. When a business regression jumps to 0.95, the first move is not to ship it but to find out what the strongest predictor is and when its value becomes known. The usual explanations are unglamorous: the predictor is derived from the outcome (an invoiced amount used to predict revenue), it is recorded after the outcome occurs and would not exist at decision time, an identifier effectively indexes the outcome, or the model carries nearly as many parameters as rows. Each of these produces an impressive number and a useless model. Treating a high fit statistic as self-validating is the more expensive failure mode of the two, because low-fit models get scrutinised while high-fit ones get deployed. ## Setting the bar before modelling The way to avoid arguing about R-squared after the fact is to define acceptance in advance, working backwards from the decision: 1. **Name the action** the output triggers and who takes it. 2. **Name the status quo** the model must beat - usually a simple rule or a human judgement, not a blank page. 3. **Price the errors** in each direction, since asymmetric costs change what accuracy is needed. 4. **State what would change the decision.** If no plausible model output changes anyone's behaviour, the project is not a modelling project. 5. **Add the non-fit requirements**: the predictors must be available at decision time, the coefficients must be stable enough to defend, and the model must be explainable to whoever is accountable for the action. With those written down, the fit statistic becomes one input to a decision that was already framed, rather than a scoreboard that stakeholders anchor on. ## Communicating a low value When a stakeholder reads 0.04 as failure, reframe it as a statement about the outcome rather than about the model: most of what drives individual behaviour is idiosyncratic and no model will claim it. Then show what the model does deliver - direction and magnitude of the drivers with their uncertainty - and connect that to the decision it supports. Refusing to inflate the number, and refusing to hide it, is the credibility-preserving move.

  • How would you set the acceptance bar before the model is built?
    Work backwards from the decision. Name the action the output triggers, the status quo it must beat, and the cost of being wrong in each direction. That turns an abstract fit target into a concrete question: at what quality does anyone's behaviour actually change? If no answer exists, it is not a modelling project.
  • A model's R-squared jumps from 0.30 to 0.95 when one feature is added. What is your first move?
    Find out what the feature is and when its value becomes known. The usual cause is that it is derived from the outcome or recorded after it, so it could never be available at decision time. Trace its definition back to the source system before touching the model itself.
  • How do you explain a low R-squared to a stakeholder who reads it as failure?
    Reframe it as a fact about the outcome, not the model: individual behaviour is largely idiosyncratic, so no model will explain most of it. Then show what the model does deliver - the direction and size of each driver with its uncertainty - and tie that directly to the decision it supports.

Predicting which individual raindrop hits your window is hopeless; predicting how much rain falls this month is routine. A driver model is doing the second job, so judging it by the first job's accuracy is a category error.

saying these in an interview costs you the question

  • Quotes a fixed R-squared threshold every model must clear
  • Treats the highest R-squared model as automatically the best
  • Never questions a suspiciously high fit statistic
  • Ignores whether predictors exist at decision time
  • Judges a driver-analysis model by per-unit prediction accuracy

context