With day-of-week dummies and Monday as the reference level, what does the Saturday coefficient mean?
answer
- a contrast, not a level
- one day is missing from the columns
- the intercept holds the omitted day
- Saturday minus Monday, other predictors fixed
basics
~20 sThe Saturday coefficient is the estimated difference in the outcome between Saturday and Monday, with the model's other predictors held fixed. It is a contrast against the omitted baseline day, not Saturday's own average level.
solid answer
~40 sMonday is the omitted level, so there is no Monday column in the model. Every remaining day gets a 0/1 indicator, and its coefficient answers one question: how much does the expected outcome move when you go from Monday to that day, holding the other predictors constant. So a Saturday coefficient of `-120` means Saturday runs 120 units below Monday, in the outcome's own units. The intercept carries the baseline: it is the expected outcome on Monday when the other predictors are zero. Contrasts between two non-reference days come from differencing their coefficients — Saturday versus Sunday is `b_Sat - b_Sun`. And a t-test on the Saturday coefficient tests only Saturday versus Monday; it says nothing about Saturday versus Friday, or about whether day-of-week matters at all.
go deeper
Be ready to say in one sentence that a dummy coefficient is the gap between that level and the omitted level, and to point to the intercept as where the baseline lives.
Expect to derive it: write the fitted equation for a baseline row and a Saturday row and subtract, then explain why the difference between two non-baseline levels is a difference of coefficients.
Show you report these safely — always naming the baseline, resisting the reading that one dummy's p-value settles whether the factor matters, and getting a contrast's standard error right rather than eyeballing two printed ones.
Own the convention: which baseline the team uses, whether coefficients reach stakeholders as raw contrasts or as level estimates, and how model summaries are written so nobody downstream reads a contrast as an absolute level.
## What a dummy variable is A categorical predictor such as day of week has no numeric meaning — coding Monday as 1 and Tuesday as 2 would tell the model that Tuesday is 'one more' than Monday and that the step from Monday to Tuesday is the same size as the step from Thursday to Friday. Neither claim is intended. Instead each category is represented by a **dummy** (also called an indicator): a column that is 1 when the row belongs to that category and 0 otherwise. With seven days you create **six** dummies, not seven. One level is left out; it is called the **reference level**, the **baseline**, or the omitted category. Here that level is Monday. ## Where Monday went Monday has no column, but it has not left the model. A Monday row has 0 in all six day columns, so the only thing describing it is the intercept. That is what makes the intercept the baseline: with an intercept plus dummies for Tuesday through Sunday, the intercept is the expected outcome for a Monday row whose other predictors are all zero. Every other day's coefficient is then measured *from that anchor*. Write the fitted equation for a row with no other predictors: ``` yhat = b0 + b_Tue*Tue + b_Wed*Wed + ... + b_Sat*Sat + b_Sun*Sun ``` A Monday row gives `yhat = b0`. A Saturday row gives `yhat = b0 + b_Sat`. Subtracting, `b_Sat = (Saturday mean) - (Monday mean)`. That is the whole interpretation. ## Units and direction The coefficient is in the outcome's units. If the outcome is orders per day and `b_Sat = -120`, Saturday is estimated to see 120 fewer orders than Monday. The sign is relative to the baseline only — a negative dummy coefficient does not mean the level is bad or small in absolute terms, it means it sits below the omitted level. The usual ceteris paribus clause applies: the comparison holds every other predictor in the model at the same value. If the model also contains, say, a promotion indicator, `b_Sat` is the Saturday-minus-Monday gap *among days with the same promotion status*, not the raw gap in the data. ## Comparing two non-reference levels Because each coefficient is measured from Monday, the difference between two other days is the difference of their coefficients: ``` Saturday - Sunday = b_Sat - b_Sun ``` The point estimate is easy. The **standard error** is not: `SE(b_Sat - b_Sun)` needs the variances of both estimates *and* their covariance, which is why you cannot eyeball it from the two printed standard errors. The practical routes are to compute the standard error of that linear contrast properly, or simply to refit the model with Sunday as the reference level, which makes the Saturday coefficient *be* the Saturday-minus-Sunday contrast with its correct standard error attached. ## What a significance test on one dummy does and does not say The t-statistic printed beside `b_Sat` tests the null hypothesis that Saturday and Monday have the same expected outcome. Three consequences follow, and interviewers probe all three. First, a non-significant Saturday coefficient does **not** mean Saturday is unremarkable — it means you could not distinguish it *from Monday*. Saturday might sit far from Wednesday. Second, significance depends on which level you omitted. Choose a different baseline and the same data produce a different set of individual t-tests, because they are tests of different comparisons. Third, the question 'does day of week matter at all?' is a joint question about all six coefficients together, not a scan of six individual t-tests. That joint test is unchanged by the choice of baseline. ## Choosing a reference level Nothing in the mathematics prefers one level. The fit is identical whichever you drop. But the *readability* of the output is not identical, so pick a baseline that makes the contrasts you care about the ones printed: a level with plenty of observations (so the anchor is precisely estimated, since every contrast inherits its noise), a level that is a natural status quo or control, and — where there is an order — one end of that order rather than the middle. ## Common mistakes Reading `b_Sat` as Saturday's average outcome is the most common error; that quantity is `b0 + b_Sat`, and only when the other predictors are zero. Reporting the coefficient as a raw fact about Saturday without naming the baseline is the second; a coefficient with no stated reference level is uninterpretable, which is why written model summaries should always say which level was omitted.
- What does the intercept represent in that day-of-week model?It is the expected outcome for a Monday row when every other predictor in the model equals zero. Monday has no column of its own, so the baseline day is folded into the intercept. If the other predictors have no meaningful zero, the intercept is still the correct anchor for the dummy contrasts even though its own value may not describe any real day.
- How would you get the estimated Saturday-versus-Sunday difference from that same model?The point estimate is `b_Sat - b_Sun`, since both are measured from Monday. The standard error needs the variances of both coefficients and their covariance, so you cannot combine the two printed standard errors. Either compute the standard error of that linear contrast directly, or refit with Sunday as the reference level so the Saturday coefficient becomes the contrast you want.
- If the Saturday coefficient is not statistically significant, what exactly has failed to be shown?Only that Saturday differs from Monday, the omitted level. It says nothing about Saturday versus Friday, and nothing about whether day of week matters overall — that is a joint test on all the day coefficients at once. Change the baseline and the individual significance pattern can change even though the fit is identical.
- How would you choose which day to make the reference level?Any choice fits the data identically, so choose for readability. Prefer a level with many observations, because every contrast is measured against it and inherits its noise, and prefer a natural status quo so the printed coefficients are the comparisons stakeholders actually ask about. Then state the baseline explicitly wherever the coefficients are reported.
The reference level is sea level on a map of elevations. Each dummy coefficient is a height above or below sea level, not an absolute altitude, and moving sea level renumbers every peak without moving any mountain.
saying these in an interview costs you the question
- Says the coefficient is Saturday's average outcome
- Reports a dummy coefficient without naming the reference level
- Reads a negative dummy coefficient as a small absolute value
- Compares two dummies' p-values instead of testing their difference
- Thinks a non-significant dummy proves day of week is irrelevant