skip to content

Two analysts overlay a fitted smoothing curve on the same scatter; one curve dips mid-year and the other rises steadily - how do you settle it?

level: seniorimportance: should knowfreq 40%

answer

  1. the curve is a fit, not a reading
  2. one knob decides the wiggliness
  3. sweep the span, then the rows
  4. the ends have neighbours on one side

basics

~20 s

Neither curve is a finding. A fitted smoothing curve averages nearby points, and how wide that neighbourhood is - the span - is a setting: wide erases a real dip, narrow invents one. Fix one span rule, then go to the rows.

solid answer

~50 s

A smoothing curve is not a measurement of the data; it is a fit whose wiggliness is a parameter. Each part of the curve is built from the points in a neighbourhood around it, and the width of that neighbourhood - the span - decides everything: wide enough and a genuine mid-year dip is averaged away, narrow enough and ordinary scatter becomes a feature. So the two curves are not in conflict about the data, they are two settings. Settle it by applying one span rule to both, checking whether the dip survives a range of spans, and then going to the rows in that window - how many are there, and does one account dominate them. Also establish where any shaded band came from: some interfaces compute an interval as part of the same smoothing default, others give you a bare line you fitted yourself. Publish the span.

go deeper

for a junior

Recall that a curve drawn over a scatter was fitted, not measured, and that how smooth it looks is a setting. Do not quote a shape from a curve without knowing who chose that setting.

for a middle

Explain the trade the span makes: wide averages a real feature away, narrow turns scatter into bumps, and the ends are supported by points on one side only. Say what you would redraw to test a shape.

for a senior

Run the process: recover both spans, redraw under one rule, sweep the span, count the rows under the feature, and check whether any band was computed by the smoother or added by hand. Then publish the span with the chart.

for a principal

The call is what your team is allowed to publish. A fitted curve on a decision chart without its span and its points is an assertion with its evidence removed; deciding that is a standing rule costs review time and prevents a whole class of retracted claims.

## What the curve actually is A smoothing curve drawn over a scatter is **a statistic the chart computed, not a measurement it read**. Each point along the curve is built from the observations in a neighbourhood around that horizontal position - roughly, a local average, usually weighted so nearer points count for more. The chart then joins those local results into a line and draws it on top of the marks. The single parameter that decides what the curve says is **the span**: how wide that neighbourhood is. Everything else about the fit is downstream of it. | span | what happens to a genuine feature | what happens to noise | what happens at the ends | |---|---|---|---| | wide | a real dip is averaged away with its neighbours | suppressed, correctly | stable, but heavily pulled toward the interior | | middling | survives if enough rows support it | mostly suppressed | reasonable, still the weakest part | | narrow | preserved, and so is everything else | promoted into visible bumps | swings, because few points support the end | That table is the whole of the disagreement between the two analysts. **Neither curve is a claim about the data that the other contradicts**; they are two settings of the same knob, and the chart drew each faithfully. ## Settling the disagreement 1. **Recover both spans.** If one or both are the tool's default, say so explicitly - a default is still a choice, just not one a person made. 2. **Apply one rule to both.** Redraw the two charts with a single, stated span so the comparison is about the data. 3. **Sweep the span.** Draw the same scatter at several spans from clearly wide to clearly narrow. A dip that persists across all of them is robust to the setting, which is worth reporting in exactly those words. A dip that appears at one span is a property of that span until something else supports it. 4. **Go to the rows in the window.** Count the observations under the dip. Ask whether a handful of records, one large customer or one bad import produced it. A feature that only the curve shows is still the curve's. 5. **Look at the scatter without the curve.** If nobody can see the dip in the points, the curve is doing the asserting, and it should not be the thing a decision rests on. ## The band, and where it came from Designs differ here, and the difference matters for how much weight a reader gives the curve: - **Where the smoothing is a statistical transform belonging to the mark**, the interface fits the curve and computes an interval around it as part of the same default. The band appears without being asked for. - **Where you fit the curve yourself and draw it as another line**, there is no band unless you computed one, and nothing on the page says the curve is uncertain at all. So the question to ask of a shaded band is not whether it is there but **what produced it**. A band that arrived by default still rests on the same span and on assumptions the chart made on your behalf, and a reader who treats it as confirmation of the shape has been misled by a default rather than by a claim. ## The ends of the curve are the least supported part Every point of the curve averages its neighbours, but the first and last points have neighbours on one side only. Fewer observations support them, and the fit leans on whatever happens to sit nearest the edge - which is exactly where readers extend the line with their eyes and talk about where things are heading. Treat the ends as the weakest region of the chart, say so, and consider trimming the curve short of the data's edge rather than inviting the extrapolation. ## What to publish - **The span and the fitting rule**, in the caption, in the same breath as the axis labels. - **The row count in any window you make a claim about**, because a dip over eleven observations and a dip over eleven thousand are different assertions drawn identically. - **What the band is**, if one is drawn - computed by the smoother, or added by you. - **The scatter itself**, at a density the reader can see. A curve published without its points is an assertion with its evidence removed. The habit to carry away is the same one this whole surface teaches: a curve is smooth because a setting made it smooth. Ask who chose the setting, check whether the shape survives another one, and put the answer on the chart.

  • The dip survives every span you try - is it real then?
    It is robust to the setting, which is genuinely worth reporting, but robustness is not evidence by itself. Look at the observations in that window: how many are there, do a few records dominate them, and can you see the dip in the raw points with no curve drawn. A feature only the curve shows remains the curve's.
  • Why is the curve least trustworthy at its left and right ends?
    Each part of the curve averages the points around it, and at the ends there are points on one side only. Fewer observations support the fit and it leans on whatever sits nearest the edge - which is exactly where readers extrapolate. Trim the curve short of the data's edge, or say plainly that the ends are weak.
  • A reviewer asks for the curve without the scatter, to keep the slide clean - what do you say?
    Publishing the curve alone removes the evidence and leaves only the assertion, and the smoothness reads as certainty the fit does not have. If the marks are too dense to show as they are, summarise them honestly and say what the summary did, rather than deleting them from the picture.

It is like judging your weight from a bathroom scale that jumps half a kilo a day. Average over a month and a real week-long change is invisible; average over two readings and every glass of water looks like a trend. The averaging window is your choice, and whichever shape you show, you chose it.

saying these in an interview costs you the question

  • Treats the drawn curve as a measurement rather than a fit with a setting.
  • Chooses the span that shows the expected dip and never reports it.
  • Assumes a shaded band around the curve confirms the shape.
  • Reads the curve at its ends as confidently as in the middle.
  • Believes every smoothing curve arrives with a band computed for it.
  • Publishes the curve without the points it was fitted to.