What confidence threshold lets a self-healing locator repair silently, and who owns that number?
answer
- There is no correct number to copy
- Two ways to be wrong, traded off
- Label real repairs, then place the cut
- Bands, not a single line
- A team setting with evidence behind it
basics
~20 sNo universal number exists. Calibrate the threshold on a labelled sample of the mechanism's own past repairs, trading wrong substitutions against needless failures. It is a team-owned setting with evidence behind it, not a shipped default left untouched.
solid answer
~50 sThe score is a similarity number, not the probability that a substitution is right, so the threshold is a policy choice about which error you would rather make. Set it low and the suite absorbs wrong substitutions and still reports green; set it high and it fails on changes a person would call an obvious match. Calibrate instead of guessing: record every repair the mechanism *would* make for a few weeks, have someone label each one same-control or different-control, then place the cut using the overlap between those two score distributions. Prefer bands to a single line - a high band substitutes and logs, a middle band substitutes but marks the run as needing review, everything below fails - and require a margin over the runner-up as well. The number belongs to the group that owns the suite, recorded with its calibration and revisited when the screens change.
code
yaml · 16 lineslocator_repair:
enabled: true
bands:
- min_score: 0.90 # confident: substitute and log
action: repair
run_status: pass
- min_score: 0.70 # uncertain: substitute, but surface it
action: repair
run_status: repaired_needs_review
- min_score: 0.00 # not identified
action: fail
required_margin: 0.15 # gap to the runner-up candidate
calibrated_on: 412-labelled-proposals
calibrated_at: 2026-03-18
owner: suite-maintainers
review_after: 2 releasesgo deeper
Know that repair is gated by a number: below a certain similarity the mechanism gives up and the step fails, and above it the run quietly continues against a substitute element. Ask what that number is set to.
Explain what the score actually measures and why no threshold is universal. Be ready to describe both failure modes - a wrong element accepted, and a good match rejected - and how moving the line trades one for the other.
Demonstrate calibration from real data rather than intuition: a period of recorded proposals, human labels, a cut chosen from the overlap, a required margin over the runner-up, and bands that separate a silent repair from one flagged for review.
Own the threshold as a policy the organisation can defend: who sets it, what evidence sits behind it, when it is revisited, and how much silent change the business will let a green run hide. Watch for it being lowered under release pressure.
A self-healing mechanism produces a **similarity score** for its best substitute. The **confidence threshold** is the line that turns that score into a decision: above it the run continues on the substitute, below it the step fails. Choosing that line is not a tuning detail. It decides how much unreviewed change a suite is willing to absorb and still report green. ## What the number is, and what it is not The score is a weighted sum of per-feature similarities. It is **not** a probability that the substitution is correct, and it is not comparable across mechanisms, across applications, or even across two screens of the same application whose elements carry different amounts of identifying information. A screen where every control has a distinct identifier produces confident, well-separated scores; a screen of near-identical rows produces a cluster of high scores that mean very little. The same 0.82 therefore means different things in those two places, which is the first reason a number copied from somebody else is indefensible. ## The two errors you are trading Every threshold sits between two failure modes, and moving it trades one for the other. | Setting | What happens | What it costs | |---|---|---| | **Too low** | A wrong element is accepted and the case proceeds against it | The run reports green while the case no longer tests what it claims; the failure surfaces later, from a user | | **Too high** | A substitution a person would call obvious is rejected | The suite fails on changes that broke nothing, and the team learns to re-run rather than read | The two are not symmetric. A false repair is silent and durable; a false alarm is loud and self-correcting. That asymmetry is the argument for erring high — but *err high* is still not a number. ## Calibrating rather than guessing The only defensible way to place the line is to measure your own application. 1. Run for a period — a few weeks, or a few hundred proposals — with the mechanism recording every repair it *would* make: its score, the runner-up's score, the features that agreed, and a captured image of the screen. 2. Have someone who knows the product label each proposal *same control* or *different control*. This is the expensive step and nothing substitutes for it. 3. Plot the two labels against score. You get two overlapping distributions, and the overlap is the part of the problem that no threshold solves. 4. Choose the cut from the shape of that overlap and the cost asymmetry above, not from a round number. 5. Write down the sample size, the date and the application version you calibrated on. A threshold with no provenance cannot be argued about later. Recalibrate when the screens change materially. A redesign, or a change in how identifiers are assigned, shifts the whole distribution, and last year's line is describing a product that no longer exists. ## Bands beat a single line A single cut forces every uncertain match into one of two outcomes, and the uncertain ones are exactly the interesting ones. Most mature setups use three bands, each with its own `min_score`: - **High** — substitute and log. The run passes normally and the repair sits in the record. - **Middle** — substitute, but mark the run repaired-and-unreviewed so it reads differently from a clean pass, and put the repair in front of a person before the next release. - **Low** — fail the step. The mechanism has not identified the element and should say so plainly. Add a `required_margin` over the runner-up as a second condition on the high band: a 0.93 that beats a 0.91 is not a confident identification, however high it looks. ## Whose number it is The threshold does not belong to whoever wrote the failing case, and it is not the mechanism's to decide. It belongs to the group that owns the suite and acts on its results, for the same reason a coverage gate or a release checklist does: it encodes a risk decision on behalf of everyone who reads a green run. In practice that means it lives in version control beside the suite rather than in someone's local settings, it carries a note naming the calibration behind it, changes to it are reviewed like code, and a named person is asked when the healing rate moves. The failure mode worth naming out loud is **threshold drift under pressure**. The suite goes red the day before a release, someone lowers the number until it goes green, and nobody records why. That single edit converts a safety net into a silencer, and because it is one line in one file it is invisible in every report the team looks at afterwards. A threshold that can be changed without evidence is not really a threshold; it is a slider that always ends up at whichever setting makes the pipeline quiet.
- How would you gather the labelled sample you need before choosing a threshold at all?Run the mechanism in a mode that records every repair it would make - score, runner-up score, features agreed, a captured screen image - without that score being what decides the case, for long enough to collect a few hundred proposals across real product changes. Have someone who knows the product mark each substitution same-control or different-control. The two score distributions over those labels are what the cut is chosen from.
- Why is a threshold calibrated a year ago suspect today?The score distribution is a property of the application's screens, not of the mechanism. A redesign, a new component set, or a change in how identifiers are assigned shifts what a given score means, and the labelled sample it was chosen from no longer describes the product. Treat it like any tuned setting: date it, keep the evidence, and recalibrate whenever the screens change materially.
saying these in an interview costs you the question
- Quotes one specific threshold as the industry-standard number
- Treats the similarity score as the probability the repair is correct
- Leaves the threshold at whatever the mechanism came configured with
- Lowers the threshold whenever the suite goes red, with no evidence
- Cannot name the two errors the setting is trading against each other
- Assumes one threshold suits every screen the suite ever touches