skip to content

How would you validate edit distance as a quality proxy for AI-drafted agent replies?

level: seniorimportance: should knowfreq 44%

answer

  1. a proxy is a hypothesis
  2. normalise before you compare
  3. strip the signature first
  4. check it against human labels
  5. measured editing stops happening

basics

~20 s

Treat it as a hypothesis, not a metric. Normalise the distance for length, then check on a few hundred human-labelled drafts whether high distance actually co-occurs with bad answers, per segment. If the correlation is weak or driven by boilerplate, the proxy is measuring style, not quality.

solid answer

~50 s

Edit distance between what the model drafted and what the human agent actually sent is attractive because it is continuous and emitted on every assisted reply. But it is a proxy, so it needs validating. First normalise: divide the character-level distance by the longer string's length, otherwise long replies dominate. Then take a few hundred drafts that already carry human quality labels and check whether the normalised distance separates good from bad — a rank correlation or a simple decile plot is enough. Then hunt confounds: appended signatures, mandatory disclaimers, agents who rewrite everything by habit, and localisation all inflate distance with no quality meaning. Segment by agent and by intent before comparing. Finally, guard against Goodhart's law: the moment agents are measured on how much they edit, they stop editing, and the signal dies.

code

python · 15 lines
python
def norm_edit_distance(a: str, b: str) -> float:
    if not a and not b:
        return 0.0
    prev = list(range(len(b) + 1))
    for i, ca in enumerate(a, 1):
        cur = [i]
        for j, cb in enumerate(b, 1):
            cur.append(min(prev[j] + 1, cur[j - 1] + 1, prev[j - 1] + (ca != cb)))
        prev = cur
    return prev[-1] / max(len(a), len(b))


draft = "Your refund of 42 EUR will arrive in 3-5 business days."
sent = "Your refund of 42 EUR should arrive within 5 business days."
print(round(norm_edit_distance(draft, sent), 3))

go deeper

for a junior

Understand what the signal is: the gap between what the model drafted and what the human actually sent, available on every draft. Say that it needs to be scaled for reply length before drafts can be compared.

for a middle

Explain the mechanics — normalisation, boilerplate stripping, cohorting by intent and agent — and why per-draft readings are noise while aggregates are informative. Be ready to name two confounds that inflate the number with no quality meaning.

for a senior

Demonstrate the validation instinct: state that a proxy is a hypothesis, describe how you would test it against labelled drafts, and say what result would make you drop it. Also cover how you use it as a sampler for review rather than a verdict.

for a principal

Own the incentive design. Argue why this stays a system metric and never a per-agent one, what independent signal guards against the metric being optimised into uselessness, and how proxies get retired when they stop earning their dashboard space.

## Why a continuous proxy is worth the trouble Most quality signals are sparse or binary. In an assist flow — where a model drafts a reply and a human sends it — you get something rarer: a continuous, universally emitted signal. Every single draft produces a pair (what the model wrote, what the human sent), and the gap between them is a numeric read on how much repair the human had to do. Unlike a thumbs rating it exists for 100% of drafts, and unlike abandonment it has magnitude, not just presence. That makes it a strong candidate proxy — and exactly the kind of metric that quietly measures the wrong thing if nobody validates it. ## Step 1: define the measure precisely Raw Levenshtein distance is a character count, so a 900-character reply with a small fix scores higher than a 40-character reply that was rewritten wholesale. Normalise by the longer string's length so the value sits in 0..1 and is comparable across replies. Alternatives worth considering: token-level distance (less sensitive to punctuation churn), or a semantic similarity between draft and sent (catches paraphrase, but adds a model dependency and its own drift). Decide, write it down, and version it — changing the formula silently makes every historical comparison meaningless. ## Step 2: test it against labels Validation means: does this number move with quality as a human would judge it? Take a sample of drafts that carry independent quality labels — from explicit ratings, from a review queue, or from a small deliberate labelling exercise. Then look at the relationship: - Bucket normalised distance into deciles and plot the share of human-labelled-bad drafts in each. A usable proxy shows a monotone rise. - Or compute a rank correlation between distance and label. You are not looking for a large coefficient; you are looking for a stable, direction-correct one that survives slicing. - Check both tails. A proxy that flags catastrophic rewrites but cannot distinguish a good draft from a mediocre one is still useful, provided you only use it that way. If the relationship vanishes once you control for reply length or intent, the proxy was measuring length or intent. ## Step 3: enumerate the confounds This is where most edit-distance programmes fail. Systematic edits with no quality meaning include: - **Boilerplate.** Signatures, greeting personalisation, legal disclaimers, ticket numbers appended by policy. Strip these before measuring, or the floor is never zero. - **Agent habit.** Some agents rewrite everything to their own voice. Distance then measures the agent, not the draft. Segment per agent, or compare each agent against their own baseline. - **Intent mix.** A refund explanation gets edited more than a store-hours answer, for reasons of risk rather than quality. Compare like with like. - **Locale and language.** Different average lengths and different editing conventions shift the distribution. - **Formatting.** A draft pasted into a channel that strips markdown will show distance from formatting alone. A validated proxy is really "normalised distance, boilerplate-stripped, within intent and agent cohort". ## Step 4: place it in the signal stack Edit distance is one member of a family. Pair it with discrete repair signals that carry different information — most usefully, whether the user rephrased the same question within about thirty seconds of a reply, which is often a stronger negative than an unclicked thumbs-down because it is evidence the answer failed for someone who was still engaged. Distance tells you *how much* repair; a rephrase tells you *that* the answer missed. When they agree, confidence is high. When they disagree, you have a candidate for human review. Also decide the unit of use. Per-draft, the signal is too noisy for a verdict; nobody should be judged on one number. Aggregated per intent per day, it is a serviceable trend line. As a sampler, it is excellent: pull the worst decile into a review queue and you will find real defects fast. ## Step 5: protect it from Goodhart's law The moment editing volume appears on an agent's scorecard, agents stop editing, and the metric improves while quality falls — the proxy has been optimised into uselessness and, worse, into a customer-facing harm. Keep it a system metric, not a person metric; state that explicitly when you socialise it; and keep at least one independent signal (sampled judging, downstream reopen rate) that does not depend on human editing behaviour, so you can detect the collapse if it happens. ## When to abandon it If, after normalisation and cohorting, the proxy still fails to separate labelled good from labelled bad, delete it rather than keeping it on a dashboard. A validated weak signal is fine; an unvalidated signal that people trust is a liability, because it will eventually be used to justify a decision it cannot support.

  • Why normalise by length rather than reporting the raw distance?
    Because raw distance scales with reply size, so long replies with trivial fixes outrank short replies that were rewritten completely. Dividing by the longer string's length puts every draft on a 0-to-1 scale where 0 means sent verbatim and values near 1 mean effectively rewritten. Without that, your worst-decile review queue fills with your longest answers and you learn nothing about quality.
  • What would make you prefer a semantic similarity score over character-level distance here?
    When agents routinely paraphrase without changing meaning — different house voice, different register — character distance flags every reply while the content was fine. A semantic measure sees those as near-identical and reserves low scores for genuine content changes. The cost is a model dependency: the similarity model must be pinned and versioned, or your quality trend will move when it changes.
  • How do you use this signal without turning it into a performance metric for agents?
    Publish it aggregated by intent and by model version, never by individual, and say so when you introduce it. Keep an independent check — sampled judged quality or ticket reopen rate — that does not depend on human editing behaviour, so if agents ever stop editing to look good, the independent signal exposes the drop instead of confirming the illusion.

It is like judging a translator by how much a proofreader changes their text: informative in aggregate, but worthless until you subtract the house style guide and account for which proofreader you got.

saying these in an interview costs you the question

  • Reporting raw edit distance without normalising for reply length
  • Treating a single high-distance draft as proof of a bad answer
  • Never checking whether the proxy correlates with human quality labels
  • Comparing distance across agents and intents without cohorting
  • Putting edit volume on an individual agent's scorecard

context