Why does a click-prediction model need the ad server's impression log and not only its click log?
answer
- a rate needs a denominator
- positives only is not a dataset
- the label is a join result
- negatives are unclicked exposures
- no impression line, no training row
basics
~20 sA click log holds positives only. The negative examples are the ads that were shown and not clicked, so the impression log is what decides which training rows exist at all and what the base click rate is.
solid answer
~40 sA click log is a log of positives. One training row is a join, not a record: take each impression the ad server actually rendered, look for a matching click inside the attribution window, label the row `1` if one exists and `0` if none does. That makes the impression log — one line per ad actually shown, carrying the impression identifier, the serve timestamp and the context known at serve time — the thing that defines the row set. With clicks alone you can count clicks but you cannot estimate the probability of one, because you have no denominator: every row would be labelled `1`, and the cheapest model that fits is the one predicting `1` everywhere.
go deeper
Remember the shape: clicks are the numerator, impressions the denominator, and a training row is one rendered ad plus whether a click followed it. Positives-only data cannot teach a rate.
Explain the join: which key matches a click to an impression, why the exposure record must carry the serve timestamp and the serve-time context, and why a rendered ad rather than an ad request is the unit.
Show that you treat the exposure log as a product surface with its own reliability: a dropped beacon on one placement is a silent, slice-shaped hole in the dataset that no downstream check will error on.
The trade-off is what the platform commits to logging forever versus what it can reconstruct. Exposure is the one thing that cannot be recovered later, which is why its retention and completeness budget outranks convenience.
## What one training row actually is A training row for a click model is **not** a record anyone writes down. It is the result of a join between two logs that the serving path emits independently: - the **exposure record** — one line written when an ad was actually rendered for a user, carrying an impression identifier, the serve timestamp, and the context that was known at that moment (slot, page category, audience segment, the candidate that won the auction); - the **click event** — one line per click, carrying the same impression identifier so it can be matched back. The label is *derived*: `1` where a click event matches the impression inside the attribution window, `0` where none does. This inversion is what trips people up. Learners expect the log to contain the answer. It contains half of it — the click log records the things that happened, and a click model has to learn about the far more numerous things that did not. ## Why a click log alone cannot train a probability model - A probability needs a **denominator**. Clicks are the numerator; impressions are the denominator. Without exposures there is no rate to estimate. - Every row would carry label `1`. Any loss on that set is minimised by a constant prediction of one, which is exactly what such a model learns. - You cannot measure the quantity you intend to predict. The observed click rate — clicks divided by impressions — is not computable from the numerator. - Slice coverage disappears. "Which placements never get clicked" is answerable only when the unclicked exposures are on disk. | log | one line per | can answer alone | cannot answer alone | |---|---|---|---| | click log | click | how many clicks, on what, when | how often an ad shown was clicked | | impression log | ad actually rendered | how much was shown, where, to whom | which of those were clicked | | joined rows | rendered ad + its outcome | the click rate and what drives it | outcomes past the attribution window | ## What the exposure record has to carry 1. **An impression identifier** that the click beacon echoes back — the join key. If the click event carries only a creative identifier and a user identifier, the join becomes ambiguous the moment the same creative is shown twice to the same user. 2. **The serve timestamp**, which anchors both the label window and every as-of feature computed for that row. 3. **The context as it stood at serve time** — either the feature values themselves, or a key plus a version that lets them be reconstructed. Re-deriving them later from current state is how a leak gets in. 4. **Enough placement detail to distinguish a rendered ad from an ad that was merely returned.** An ad response that never painted is not an exposure, and counting it as a negative teaches the model that a perfectly good ad was declined. ## Where the row set silently goes wrong - **Dropped beacons.** If one placement's impression beacon fails for a client type, those exposures never land, and the assembled set quietly under-represents that slice while showing no error anywhere. - **Orphan clicks.** A click whose impression line is missing must not be promoted into a row of its own — that fabricates a positive with no features and no denominator partner. Count orphans as a data-quality metric instead. - **Multi-slot pages.** Several ads render in one page view; each rendered ad is its own row, and collapsing the page into one row destroys the per-ad label. - **Deduplication.** A retried beacon can write the same impression twice; the assembly job needs the identifier to be the unit of deduplication, or the base rate shifts. ## What this decides downstream Everything about the assembled set follows from the row definition. The base click rate comes out of it, and that rate sets the sampling rates, the per-row weights, and the calibration the model has to reproduce. Coverage of slices comes out of it, and a slice that lost its exposure lines is a slice the model will be confidently bad at. This is why the exposure log is treated as a first-class product surface rather than as telemetry: it is the only place the negative class exists, and a gap in it is unrecoverable after the fact, because nothing else on the platform records what a user was shown and ignored.
- A single placement's impression beacon starts failing for one client type. What does that look like in the assembled dataset?Not an error — a hole. Rows for that slice simply stop appearing, so the model is trained on less of it and its clicks may show up as orphans with no impression partner. The detectable signals are a drop in row counts for that slice against its serving volume, and a rise in orphan clicks. Both belong in the assembly job's own data-quality checks, because nothing downstream distinguishes "few exposures" from "few rows".
- A click event arrives with no matching impression in the log. What should the assembly job do with it?Drop it from the training rows and count it. Creating a row from the click alone fabricates a positive with no serve-time context and inflates the base rate; silently ignoring it hides a logging defect. The orphan rate is a standing quality metric of the join — a few are normal from retention boundaries and late beacons, and a jump in it means the exposure path is broken.
saying these in an interview costs you the question
- Says the click log alone is enough to train a click model.
- Treats the label as a column the ad server writes at serve time.
- Promotes a click with no matching impression line into a training row.
- Counts every ad request as a row even when no ad was rendered.
- Assumes a missing click event means the ad was never served.