skip to content

Flight instructors report that praise after a great landing precedes a worse one — what explains this?

level: middleimportance: nice to knowfreq 30%

answer

  1. look at when the feedback was triggered
  2. the extremes were selected, not assigned
  3. great landings contain favourable luck
  4. randomise feedback within equally extreme landings

basics

~20 s

Regression to the mean, not the praise. Exceptional landings are partly luck, so the next attempt is usually closer to the trainee's normal standard whatever the instructor says. The same logic makes criticism after a terrible landing look effective.

solid answer

~50 s

The instructors are describing an artifact of when they choose to speak. A landing quality is part stable skill and part occasion-specific noise — gusts, approach angle, timing. Exceptional landings are enriched with favourable noise, so the next attempt averages closer to the trainee's normal level; awful landings are enriched with unfavourable noise, so the next attempt averages better. Praise is triggered by the first case and criticism by the second, which means praise reliably precedes a decline and criticism reliably precedes an improvement even if feedback does nothing at all. The experience is systematically miseducating: it teaches that punishment works and encouragement backfires. To measure the real effect you must hold the trigger constant — among landings of similar quality, give praise to a randomly chosen subset and withhold it from the rest, then compare the following attempts.

go deeper

for a junior

Recognise that the landings which triggered praise were unusual ones, and that unusual results tend not to repeat. Naming regression to the mean here is most of the answer.

for a middle

Explain the decomposition into stable skill and occasion-specific noise, and show why triggering feedback on the outcome guarantees the pattern even when feedback does nothing.

for a senior

Turn it into a measurable design: randomise feedback within a band of similarly extreme performances, or use a less noisy outcome, and say what you would conclude from each result.

for a principal

Own the organisational version. Decide how performance interventions are triggered and evaluated so that leaders stop learning, from consistent lived experience, that punishment works and encouragement fails.

## The pattern and the trap Instructors who give feedback based on how a landing went will report, honestly and consistently, that praise seems to hurt and criticism seems to help. The observation is real; the causal reading is wrong. Decompose a landing's quality into a stable part — the trainee's actual skill at that stage of training — and a transient part covering wind, traffic, approach setup, fatigue and luck. Because the transient part is redrawn each time, the very best landings in a session are disproportionately ones where the transient part was favourable, and the very worst are ones where it was unfavourable. Whatever happens next, the transient part has expectation zero again, so the following landing sits nearer the trainee's stable level. Better landings tend to follow bad ones; worse landings tend to follow great ones. That is regression to the mean. Now overlay the feedback rule. Praise fires on the extreme highs; criticism fires on the extreme lows. The feedback is not applied to a representative sample of landings — it is applied to a *selected* sample chosen on the outcome variable itself. The apparent effect of each kind of feedback is therefore the expected regression of the group it was applied to, plus whatever the feedback actually does. With no effect at all, the observed pattern is exactly as reported. ## Why this misleading experience is so persuasive Three features make it unusually convincing: 1. **It replicates.** Every instructor sees it, every cohort, year after year. Reliability feels like validity, but a consistent selection rule produces a consistent artifact. 2. **It has a ready mechanism.** 'Praise breeds complacency' is a plausible story, so the mind supplies a cause and stops looking. 3. **The counterfactual is invisible.** Nobody observes what the trainee would have done after a great landing with no praise, because the instructor always praises. ## What it does *not* prove It does not show that praise is useless or harmful. The observation carries essentially no information about the sign of feedback's real effect, because the comparison is between two differently-selected groups rather than between treated and untreated cases from the same pool. Praise could be strongly beneficial and the pattern would still appear, merely a little attenuated. The correct conclusion is 'this data cannot answer the question', not 'encouragement backfires'. ## How to actually measure it Hold the trigger constant and vary the feedback. Among landings of comparable quality — say, all landings graded in the top band — praise a randomly chosen subset and stay quiet for the rest, then compare the next landing across those subsets. Because both subsets were selected by the same extreme rule, both regress by the same amount, and any residual difference is attributable to the praise. The same design works at the bottom band for criticism. Lower-effort alternatives exist when withholding feedback is unacceptable: use a stable outcome rather than a single noisy landing (an average over the next five attempts regresses far less), or predict how far each landing should regress from the historical attempt-to-attempt correlation and compare the actual next landing to that prediction rather than to the extreme one. ## The general shape The structure recurs anywhere feedback, help or scrutiny is triggered by an extreme outcome. A performance-improvement plan opened after someone's worst quarter will appear to work; a bonus paid after a stellar quarter will appear to backfire; extra tutoring given to the lowest-scoring pupils will appear effective; an audit launched after a spike will appear to have fixed things. In every case the trigger is an extreme observation of a noisy measure, and the rebound arrives on schedule. The practical rule to carry out of this is short: **whenever the rule for who gets an intervention is 'they were extreme on the outcome we will measure again', the naive before-and-after comparison is not evidence.** Either randomise within the extreme group, or measure the outcome on an independent, less noisy instrument, or state up front how much rebound to expect so the result can be judged against it.

  • Does this mean feedback has no effect on performance?
    No. It means this comparison cannot measure the effect, because when feedback is given is entangled with how extreme the performance was. Praise could be helpful, harmful or neutral and the same pattern would appear, because pure regression to the mean produces it on its own. The honest conclusion is that the observation carries no information about the sign of the real effect.
  • How would you design a test of whether praise actually helps?
    Hold the trigger constant and vary the feedback. Among landings of similar quality, praise a randomly chosen subset and withhold praise from the rest, then compare the following attempt across the subsets. Both were selected by the same extreme rule so both regress equally, and any remaining difference is attributable to the praise itself.
  • If withholding feedback is unacceptable, what weaker approach still helps?
    Measure the outcome on something less noisy than the very next attempt — an average over the next several landings regresses far less — or compute from the historical attempt-to-attempt correlation how much rebound to expect, and judge the observed next landing against that prediction rather than against the extreme one that triggered the feedback.

It is like praising the weather on an unusually warm October day and concluding that praise cools things down. The next day was always going to be closer to normal.

saying these in an interview costs you the question

  • Concludes that praise damages performance
  • Ignores that the landings were selected for being extreme
  • Treats consistent instructor experience as controlled evidence
  • Assumes the next single attempt measures the feedback's effect
  • Invents a psychological mechanism before checking the selection rule

context