A vendor memo ranks models for a permit-review assistant using SWE-bench Verified and Terminal-Bench — what is wrong, and what would you measure instead?
answer
- measuring a different capability entirely
- construct validity, not accuracy
- the product writes no code
- score citations, fabrication, abstention
- transfer is a hypothesis to test
basics
~20 sThose benchmarks score agentic coding and command-line work, and a construction-permit assistant writes no code — the memo ranks models on a construct the product does not use. Replace it with a suite of real permit questions scored on citation correctness and safe abstention.
solid answer
~50 sThe defect is **construct validity**: the memo measures agentic software engineering and terminal task completion, while the product retrieves municipal permit code, answers an applicant's question, cites the governing section and declines when the code does not cover the case. Those are different capabilities, and no published correlation licenses the inference from one to the other. Public scores can still build a shortlist, but the decision needs a suite of real permit questions with known-correct answers, run through the full assistant — retrieval included — and scored on whether the cited section is the right one, whether any clause was fabricated, and whether unanswerable cases are escalated rather than guessed. In practice a modest internal suite of a few dozen well-chosen permit cases separates models that the public leaderboards rank identically, because it asks about the one narrow capability the product depends on. Transfer from any benchmark is an empirical claim: test it, do not assume it.
go deeper
Recognise that SWE-bench Verified and Terminal-Bench score coding and command-line agents, so they say nothing about a product that answers permit questions from documents. Ask what the assistant is actually supposed to get right.
Name construct validity and explain the mismatch concretely, then list scoreable criteria for the real task — correct governing section, no invented clauses, escalation when the corpus is silent — evaluated over the full retrieval pipeline.
Diagnose why the memo's evidence is invalid before proposing a replacement, concede the partial transfer honestly, and design a compact suite around real failure modes that can actually settle a procurement decision.
Own the procurement standard: no model decision lands on external scores alone, every product surface has a suite that outlives model versions, and the cost of building it is weighed against the compliance exposure of a fabricated citation.
## Naming the defect **Construct validity** asks whether a measurement captures the thing you intend to measure. SWE-bench Verified measures whether an agent can resolve real issues in real code repositories, verified by running the repository's tests. Terminal-Bench measures whether an agent can complete multi-step tasks in a command-line environment. Both are good benchmarks. Neither measures: retrieving the correct section of a municipal building code, reading a zoning table, distinguishing a rule that applies to a detached garage from one that applies to an accessory dwelling, citing the section that governs, or refusing when the code is silent. So the memo's ranking is not wrong-in-detail, it is wrong-in-kind. It answers a question nobody asked. This is the single most common eval mistake in a procurement decision, and it survives because agentic benchmark scores *feel* more product-relevant than knowledge quizzes do. ## Is there any transfer at all? Be honest here rather than absolutist, because a good interviewer will push. Strong agentic coding scores do correlate loosely with instruction-following, long-context handling and multi-step reliability, all of which the permit assistant uses. Coding benchmarks are also among the few with objective end-state verification, so they are less noisy than most. But *loosely correlated* is not *sufficient for a ranking decision*, and the correlation is weakest exactly where these two models differ — that is what makes the top of any leaderboard a near-tie. The defensible framing: benchmark-to-product transfer is an empirical question. If you believe SWE-bench Verified predicts permit-answer quality, that is a hypothesis you can test by running both models on your own suite and checking whether the ranking survives. Usually it does not. ## What to measure instead Start from the failure modes the product actually has, then write items that expose them: - **Citation correctness.** Does the response point at the section that genuinely governs the applicant's situation? This is scoreable objectively against a reference section number. - **Fabrication.** Does the model invent a clause, a setback distance or a permit class that does not exist in the corpus? A single hallucinated code reference in a permit context is a compliance incident, not a quality ding. - **Safe abstention.** For questions the corpus does not answer — a jurisdiction you do not hold, a case requiring a variance — does the assistant escalate to a human reviewer or does it produce a confident guess? - **Case discrimination.** Near-miss pairs where two similar structures fall under different rules. These are where models separate. - **Grounded reasoning over tables.** Zoning and setback data live in tables; getting the right row and column is a distinct, testable skill. Run candidates through the **whole assistant** — the retrieval layer, the prompt, the citation formatter — not the bare model. A weaker model with better retrieval routinely beats a stronger model with worse retrieval, and a model-only comparison would attribute that entirely to the model. ## Why a small suite beats a big leaderboard here A few dozen carefully chosen permit cases carry far more decision value than any public benchmark, for one reason: every item is on the axis you care about. Public suites spend nearly all their resolution on capabilities you do not exercise. When two models sit within noise of each other on MMLU-Pro, GPQA-Diamond and LMArena — which top models routinely do — a targeted internal suite is the only instrument that can still tell them apart on the thing you are buying. Expect to find that one model reliably abstains and the other confabulates a plausible section number, a distinction that appears nowhere on any leaderboard and completely determines which one you can ship. One caveat to state out loud: a small suite resolves large differences, not marginal ones. If two models come out two items apart, treat that as a tie on quality and decide on cost, latency or operational factors instead of over-reading it. ## Where public benchmarks still belong in this decision Use them upstream. They cheaply eliminate models that are far off the frontier, and they surface newly released candidates worth testing. Occupational suites like GDPval, which score deliverables on real professional tasks against human-produced work, sit closer in spirit to a permit assistant than a coding benchmark does — but still not on your corpus, your jurisdiction or your citation rules. The funnel is: leaderboards narrow the field, your suite decides, and the suite is the artifact you re-run when the vendor ships the next model. ## How to say it in an interview Name the concept (construct validity), say concretely what the cited benchmarks measure and what the product needs, concede the partial transfer honestly, then describe the replacement measurement in terms of scoreable failure modes. Candidates who jump straight to "just build an internal eval" without diagnosing *why* the memo's evidence is invalid give a much weaker answer.
- The vendor argues that strong agentic coding scores prove general reliability. How do you respond?Concede the partial point and then make it testable. Agentic benchmarks do correlate loosely with instruction-following and multi-step reliability, which the assistant uses. But the claim being made is a ranking claim between two near-tied models, and that is exactly where a loose correlation carries no information. Offer the resolution: run both on our permit suite and see whether the ranking survives.
- Your permit suite ranks the two models two items apart out of a few dozen. What do you conclude?That it is a tie on quality. A small suite resolves large gaps, not marginal ones, and a two-item difference is comfortably inside what item wording and sampling variation produce. Decide on the other axes — cost per request, latency, provider reliability, deprecation history — and if the quality question genuinely matters, add harder discriminating items rather than re-reading the noise.
- Which single failure mode would you weight most heavily in a permit-review assistant, and why?Fabricated code references. A wrong-but-cited section is worse than a refusal, because an applicant or reviewer may act on it and the error carries an official-looking citation that invites trust. I would score fabrication as a hard fail rather than a partial-credit deduction, and require the system to abstain and escalate whenever the retrieved passages do not support an answer.
saying these in an interview costs you the question
- Accepts an agentic coding score as evidence for a document-retrieval product
- Says the top-ranked model is best because it leads several benchmarks
- Compares bare models instead of the full retrieval pipeline
- Treats a hallucinated code citation as a minor quality issue
- Reads a two-item difference on a small suite as a real ranking