How do you choose an embedding similarity threshold for auto-closing incidents?
answer
- no universal number exists
- a property of model, corpus and task
- false positives cost more here
- margin over runner-up beats absolute cutoff
- threshold is versioned with the model
basics
~20 sDerive it from labelled pairs on that exact model and corpus, choose the operating point from the cost of a wrong auto-close, and prefer a margin over the runner-up to a bare number. Re-validate on every model or corpus change.
solid answer
~50 sThere is no universal cutoff, because a similarity score is a property of one model applied to one corpus, not of cosine itself. Different spaces put their unrelated-pair baseline in different places, so a number quoted in a blog post or borrowed from another team means nothing here. The method is empirical: assemble labelled pairs of genuinely-duplicate and genuinely-distinct incidents from your own history, plot the two score distributions, and read off the precision and recall at each candidate cutoff. Then let the cost asymmetry pick the point — auto-closing a live incident that was real is far more expensive than leaving a duplicate open, so you want a high-precision operating point, not the F1 optimum. Strengthen it with relative signals, requiring the top match to beat the runner-up by a margin, and add an abstain band routed to a human. Finally, treat the threshold as a versioned artifact pinned to the model and re-derived whenever either side changes.
go deeper
Know that similarity thresholds are found by testing on your own labelled examples, and that a number taken from a tutorial or another model does not carry over.
Explain how to build positive and negative pairs, read precision and recall off the two score distributions, and why realistic near-miss negatives matter more than random ones.
Show the production mechanics: an abstain band with human review, a margin requirement over the runner-up, a calibration mapping fitted on held-out pairs, and sampled audits of automated actions to measure real precision.
Own the tradeoff itself — force an explicit cost ratio from the people who bear a wrong auto-close, decide whether autonomous action is justified at the achievable precision, and make the threshold a versioned artifact with drift monitoring and a reversal path.
## Why the number cannot be borrowed Similarity scores have no absolute meaning. Every embedding space has its own baseline: the level unrelated pairs reach because the vectors crowd into a narrow region and share a large common direction. In one model that baseline may sit near 0.35, in another near 0.80. A cutoff of 0.85 is aggressive in the first space and barely above noise in the second. The corpus matters too — a homogeneous corpus of incident tickets, all written in the same register with the same product nouns, compresses scores further than a diverse corpus would. So the first thing to say is that a threshold is a property of the triple (model, corpus, task). Copying one across any of those three is guessing. This is the single most common mistake in production semantic systems, and it is why an interviewer asks the question at all. ## Turning it into a measurement The procedure is unglamorous and works. Take historical incidents where the duplicate relationship is already known — linked tickets, merged records, post-incident reviews — and build two sets of pairs: true duplicates and true distinct pairs, the latter sampled to look like what the system will actually see rather than uniformly at random, because uniformly-random negatives are far too easy. Score every pair with the exact production pipeline, including the query and document roles the model expects. You now have two distributions. Read the operating characteristics off them: for each candidate cutoff, what fraction of things above it are real duplicates (precision) and what fraction of real duplicates are above it (recall). If the two distributions overlap so heavily that no cutoff gives acceptable precision, the answer is not a cleverer threshold — it is a better representation or a second-stage check. ## Letting the costs choose the point This is where judgment enters and where there is no single right answer. The default habit is to maximize F1, which implicitly declares a false positive and a false negative equally bad. For auto-closing incidents they are not remotely equal. Wrongly closing a live incident suppresses a real signal, delays response, and destroys trust in the automation; wrongly leaving a duplicate open costs an engineer a few minutes of triage. The asymmetry might be a hundred to one, which puts the operating point far out on the high-precision end, accepting low recall deliberately. The honest framing is to state the cost ratio explicitly with the people who own the consequence, then pick the point that minimizes expected cost under that ratio. If nobody will state the ratio, that itself is the finding — the automation is not ready to be autonomous. ## Designs that beat a bare threshold **Margin over the runner-up.** Require that the best match exceeds the second-best by some gap. This is far more robust than an absolute score, because it is a within-query comparison: it survives a shift in the whole score distribution, which an absolute cutoff does not. **An abstain band.** Above a high line, act automatically; below a low line, do nothing; in between, surface a suggestion to a human. This converts a binary threshold problem into a coverage problem, and coverage can be tuned over time as confidence grows. **Calibration.** Fit a monotonic mapping from raw score to estimated probability of a true duplicate on held-out labelled pairs, using an approach such as isotonic regression or logistic (Platt) scaling. Then the decision rule is stated in probability terms, which is the only form in which a cost ratio can be applied honestly. The mapping is as model-and-corpus specific as the threshold it replaces. **A cheap second check.** A high-precision confirmation step — a stricter comparison, a structural check on the affected service and time window, or a model judgment on the pair — lets you run a lower similarity gate without the precision loss. ## Operating it over time Pin the threshold, or the calibration mapping, to the model version and the pipeline configuration, and store them together. A model upgrade invalidates the number completely; so does a change to how the two sides are encoded, and so, more gradually, does a corpus that shifts as your product changes. Monitor the inputs, not just the outputs: track the score distribution of the daily match candidates, and alert on drift in its mean or spread. Track the auto-close rate and, most importantly, sample outcomes — audit a fraction of automated closes for correctness, because that is your only direct measurement of production precision. Keep a reversal path so that a wrong auto-close is recoverable, which is what makes a slightly-too-aggressive threshold survivable rather than catastrophic. ## The short answer to give No universal number; derive it on your own labelled pairs with your own model; let an explicitly stated cost asymmetry choose the operating point; prefer margins and abstain bands to a bare cutoff; version the threshold with the model and audit the outcomes.
- Why is maximizing F1 the wrong default for this decision?F1 weights false positives and false negatives equally, which is a statement about costs, not a neutral choice. Auto-closing a real incident is far more damaging than leaving a duplicate open, so the correct operating point sits well into the high-precision region and deliberately accepts lower recall. Make the cost ratio explicit with the owners of the consequence, then choose the point that minimizes expected cost under it.
- What makes a margin over the runner-up more robust than an absolute cutoff?A margin is a comparison within a single query, so it is insensitive to shifts that move the whole score distribution — a corpus that grows more homogeneous, a slightly different encoder, a change of domain. An absolute cutoff silently changes meaning under any of those. A margin also encodes the thing you actually care about: that this candidate is distinctly better than the alternatives, not merely above a line.
- What triggers a re-derivation of the threshold after launch?Any change to the model, to the encoding roles or prefixes, to chunking, or to the comparison pipeline invalidates it outright — those change the score scale. Corpus drift invalidates it gradually, so monitor the candidate-score distribution and re-derive on meaningful drift. Sampled audits of automated actions are the ground-truth signal: a falling audited precision means re-derive now, regardless of what changed.
- How do you decide whether the system is ready to act without a human at all?By whether anyone will state the cost ratio and accept the residual error rate. Estimate the automated action's precision at the chosen point with a confidence interval from the labelled set, translate it into expected wrong actions per month at your volume, and put that number in front of the owner. If the answer is unacceptable, run in suggest-only mode and use the accumulating human decisions as the labelled data that earns autonomy later.
saying these in an interview costs you the question
- Reusing a cutoff quoted for a different embedding model
- Picking the F1 optimum regardless of the cost asymmetry
- Treating cosine similarity as a probability of being a duplicate
- Setting the threshold once and never revisiting it after a model upgrade
- Sampling easy random negatives instead of realistic near-miss pairs