Why prefer the OLS slope over one computed from only the first and last observations?
answer
- check the class first, then the variance
- a weighted sum, two non-zero weights
- the intercept cancels in the difference
- sigma-squared over Sxx, versus the endpoints
- unbiased is a low bar to clear
basics
~20 sA slope from two endpoints is linear in the data and unbiased, so it is a fair competitor, but its sampling variance is larger. Gauss-Markov says OLS wins that class, so unbiasedness alone does not make an estimator good.
solid answer
~50 sDefine `b1tilde = (y_n - y_1) / (x_n - x_1)`. It is a linear function of the outcomes with fixed weights, and under the usual conditions its expectation is the true slope, so it is linear and unbiased — a fair entrant in the Gauss-Markov comparison class. Its variance is `2*sigma^2 / (x_n - x_1)^2`, while the OLS slope's is `sigma^2 / Sxx` with `Sxx = sum (x_i - xbar)^2`. For `x = 0, 1, ..., 10` that is `sigma^2/50` versus `sigma^2/110` — the two-point estimator is 2.2 times noisier, because it throws away the nine interior observations. The lesson generalises: infinitely many unbiased estimators exist and they differ enormously in precision, so unbiasedness is a weak criterion on its own. The endpoint estimator is also fragile — one mismeasured endpoint moves the whole answer.
go deeper
Recall that many different unbiased estimators of the same slope exist, and that OLS is preferred because it is the most precise among the linear unbiased ones, not because the others are wrong on average.
Be ready to show the endpoint estimator is a weighted sum with two non-zero weights, take its expectation to confirm unbiasedness, and say why it therefore falls inside the Gauss-Markov comparison class.
Carry the variance comparison through with numbers and explain the tie cases, then add the practical point that concentrating an estimate on two observations makes it fragile to a single bad measurement.
Turn it into a design argument. Precision in a slope comes from spread in the predictor, but a design chosen purely for spread cannot detect curvature — decide explicitly which of the two you are buying.
## The competitor Suppose the data are ordered by the predictor and you estimate the slope by connecting the two extreme observations: ``` b1tilde = (y_n - y_1) / (x_n - x_1) ``` This feels crude, and the instinct is to call it biased. That instinct is wrong, and understanding why is the whole point of the exercise. ## It is linear Write it as a weighted sum of the outcomes: ``` b1tilde = w_1*y_1 + w_2*y_2 + ... + w_n*y_n with w_1 = -1/(x_n - x_1), w_n = +1/(x_n - x_1), and all other w_i = 0 ``` The weights depend only on the predictor values, not on the outcomes, so this is a linear estimator in exactly the sense Gauss-Markov requires. ## It is unbiased Under the model `y_i = b0 + b1*x_i + e_i` with `E[e_i | X] = 0`, ``` E[b1tilde | X] = (E[y_n] - E[y_1]) / (x_n - x_1) = ((b0 + b1*x_n) - (b0 + b1*x_1)) / (x_n - x_1) = b1*(x_n - x_1) / (x_n - x_1) = b1 ``` The intercept cancels in the difference, and the slope comes back exactly. Discarding observations does not create bias — it destroys precision. ## The variance comparison With uncorrelated errors of common variance `sigma^2`, the variance of any weighted sum is `sigma^2 * sum w_i^2`. For the endpoint estimator: ``` Var(b1tilde) = sigma^2 * (w_1^2 + w_n^2) = 2*sigma^2 / (x_n - x_1)^2 ``` For the OLS slope: ``` Var(b1hat) = sigma^2 / Sxx, Sxx = sum (x_i - xbar)^2 ``` Gauss-Markov guarantees `Var(b1hat) <= Var(b1tilde)`, which is equivalent to the purely arithmetic fact `Sxx >= (x_n - x_1)^2 / 2`. **Concrete case.** Take eleven equally spaced predictor values `x = 0, 1, 2, ..., 10`. Then `xbar = 5` and ``` Sxx = 2*(25 + 16 + 9 + 4 + 1) + 0 = 110 Var(b1hat) = sigma^2 / 110 ~= 0.0091 * sigma^2 Var(b1tilde) = 2*sigma^2 / 100 = 0.0200 * sigma^2 ``` The endpoint estimator has 2.2 times the variance — its standard deviation is about 1.5 times larger, so it needs materially more data to reach the same precision. Note that the ratio does not depend on `sigma^2` at all: the common factor cancels, so the comparison is a property of the design of predictor values, not of the noise level. ## When the two tie Gauss-Markov promises *no larger* variance, not strictly smaller, and there are designs where the promise is tight. - With `n = 2`, the two estimators are literally the same thing: the least-squares line through two points is the line through those two points. - With three equally spaced values, say `x = 0, 5, 10`, the middle observation sits exactly at `xbar` and contributes nothing to `Sxx`. Then `Sxx = 50 = (10 - 0)^2 / 2`, and the two variances are identical. In fact the OLS slope reduces algebraically to `(y_3 - y_1)/(x_3 - x_1)` — the middle point influences the intercept but not the slope. That second case is instructive in its own right: an observation placed at the mean of the predictor carries no information about the slope. ## The wider lesson **Unbiasedness is a weak criterion.** For any linear model there are infinitely many linear unbiased estimators of the slope — take any two distinct points, or any weighted average of such pairwise slopes. They all get the right answer on average and differ wildly in how far a single sample can land from it. Choosing among them requires a second criterion, and variance is the natural one. That is exactly the gap Gauss-Markov fills. **Robustness cuts the other way here.** Beyond variance, the endpoint estimator concentrates all of its dependence on two observations, so a single mismeasured endpoint moves the estimate arbitrarily far, and no amount of good data in the middle offsets it. It also uses none of the interior data to reveal whether the relationship is even straight — a strongly curved pattern between the endpoints leaves it untouched. OLS spreads influence across all observations, though not equally: points far from `xbar` still carry more weight in the slope. **Reading it as a design statement.** Since `Var(b1hat) = sigma^2 / Sxx`, precision comes from spread in the predictor. That is why a deliberately chosen design that places observations at the extremes is efficient for estimating a slope — and also why it gives you no ability to detect curvature. The endpoint estimator is the degenerate limit of that tradeoff: maximally spread, completely blind to shape.
- Is there a design where the endpoint slope is exactly as good as OLS?Yes. With only two observations they are the same estimator. And with three equally spaced predictor values the middle point sits at the mean of the predictor and contributes nothing to the slope, so OLS reduces algebraically to the endpoint formula and the variances coincide. Gauss-Markov promises no larger variance, not strictly smaller, and these are the cases where the bound is tight.
- Beyond variance, what makes the endpoint estimator dangerous in practice?All of its dependence sits on two observations, so one mismeasured or unusual endpoint moves the estimate arbitrarily far, and clean interior data cannot compensate. It also ignores the interior entirely, so a strongly curved relationship between the endpoints leaves it completely unwarned. OLS spreads influence across every observation, which both stabilises the estimate and lets residual patterns expose a wrong functional form.
- Does the comparison depend on knowing the error variance?No. Both variances carry the same `sigma^2` factor, so it cancels in the ratio and the comparison depends only on the configuration of predictor values. For `x = 0, 1, ..., 10` the ratio is 110 to 50 whatever the noise level. That also means you can rank candidate designs for slope precision before collecting a single observation.
Two witnesses standing at opposite ends of a street versus eleven spread along it. All are honest on average, but the estimate built from just two of them swings much more from one incident to the next.
saying these in an interview costs you the question
- Assumes any estimator using fewer observations must be biased
- Claims unbiasedness alone makes an estimator acceptable
- Says Gauss-Markov guarantees strictly smaller variance in every design
- Thinks the comparison requires knowing the error variance
- Cannot state why the estimator counts as linear in the outcomes