Why is queries-per-user a dangerous OEC for a search product?
answer
- which way is better?
- two opposite stories, same movement
- counting attempts, not accomplishments
- reformulation loop inflates the number
- fine to explain, unfit to decide
basics
~20 sQueries per user has no fixed good direction. More queries can mean users are engaged, or that results were bad and they had to keep reformulating. A criterion whose direction is ambiguous cannot decide a launch on its own.
solid answer
~50 sAn OEC needs a declared direction: up is better, or down is better. Queries per user fails that test, because two opposite stories produce the same movement. A rise can mean the product got more useful and people search more; it can equally mean results degraded and users had to reformulate three times to find one answer. Since the metric cannot distinguish success from failure, no decision rule written on it alone is sound. The way out is to make the criterion resolve the ambiguity itself — count successful sessions rather than queries, use task-level success such as a query answered without reformulation, or define the criterion at the session level where a shorter path to an answer is unambiguously better. Any metric where you cannot state, before launch, which direction counts as an improvement and why, is not yet a criterion.
go deeper
Remember that a criterion needs a direction stated before the test, and that more activity is not automatically better — repeated attempts can mean the product failed the user.
Explain the general pattern: metrics that count attempts rather than accomplishments rise under both success and failure, and describe how redefining toward successful tasks removes the ambiguity.
Show you catch this in design review rather than in the results meeting, and that you can keep the ambiguous metric as a diagnostic while moving the decision onto a success-shaped criterion.
Be ready to argue how an org stops ambiguous activity counts becoming default success measures across teams, and what review standard forces a stated direction and rationale before any experiment launches.
## The requirement being violated An Overall Evaluation Criterion is not just a metric; it is a metric *plus a direction* plus a decision rule. If you cannot write down before launch which way the number has to move for the change to count as an improvement, you do not have a criterion — you have a number that will be interpreted after the fact to suit whatever happened. Queries per user is the classic metric that fails at the direction step. ## The two stories **Engagement story.** The change made search more useful. Users trust it with more of their questions, so they issue more queries per user. Up is good. **Failure story.** The change made results worse. A user who would have found the answer with one query now searches, scans, fails, rephrases, and searches again. Each failed attempt is another query. Up is bad. Both stories move the same metric in the same direction, and the metric contains no information that separates them. This is not a subtle statistical point — the estimate can be perfectly precise and the interpretation is still undetermined. The reverse ambiguity exists too. A *fall* in queries per user can mean users abandoned the product, or that each query now works so well that fewer are needed. Neither direction is safe to declare in advance, which is precisely the problem. ## What makes a metric directionally ambiguous The general pattern: the metric counts an *activity* that users perform as a means to an end, rather than the *end* itself. Activity counts go up both when the means is more attractive and when the means is less effective and has to be repeated. Any metric of the form *number of attempts* inherits this. Recognising the pattern matters more than memorising the example. When a candidate proposes a criterion, a good question to ask is: can I tell a plausible story where this number rises and users are worse off? If yes, either the metric is not the criterion, or the criterion needs to be redefined so the bad story can no longer produce a rise. ## Fixing it Three standard repairs, in increasing order of how much they cost you: **Move up a level, from action to task.** Count queries that were *answered* rather than queries issued — for example, a query where the user acted on a result and did not immediately reformulate. Now a reformulation loop, the signature of the failure story, no longer inflates the number; it deflates it. **Redefine the unit as the task, not the action.** Make the criterion successful tasks per user, where a task is a coherent run of activity toward one need. The failure story now collapses several queries into one task with no success, so the two stories separate cleanly and the direction is unambiguous: more successful tasks is better. **Move to an outcome only a success can produce.** Use something downstream that a frustrated user cannot generate by trying harder. This is the least ambiguous and usually the least sensitive, since such events are rarer. All three share a principle: make the metric count *ends*, not *attempts*. That is what restores a declarable direction. ## Where the ambiguous metric still belongs Queries per user is not useless. As a driver or diagnostic metric it is genuinely informative — read alongside a criterion that already tells you whether the change helped, a rise in queries means something quite different depending on the criterion's verdict, and the pair together locate the mechanism. What it cannot do is carry the decision by itself. This distinction is the point of the question. Interviewers use it to check that a candidate understands that the metric hierarchy is about *roles*, and that a metric can be excellent in the explaining role and disqualified from the deciding role. ## The wrong answers - *Just look at whether it went up and use judgment.* This is post-hoc interpretation, exactly what declaring a criterion is meant to prevent. Whichever story fits the team's hopes will be the one told. - *Pair it with a second metric and read them together.* Better, but if there is no rule for what to do when they conflict, the decision is still unstructured. If two metrics jointly define success, they need to be combined into one criterion with an explicit rule. - *It is fine because in practice more search means more engagement.* This asserts one of the two stories as a fact. If it were reliably true, the metric would have a direction — and the whole reason the question is asked is that it is not. ## Compact summary A criterion must have a direction you can state before you see the data. Metrics that count attempts rather than accomplishments usually cannot, because effort rises both when the product is more attractive and when it is less effective. Redefine toward success-shaped counts, keep the ambiguous metric as a diagnostic, and never let a number with two opposite readings decide a launch.
- How would you redefine the criterion so its direction becomes unambiguous?Count accomplishments rather than attempts. Successful tasks per user is the usual repair: a run of activity toward one need counts once, and counts only if it ended in the user acting on a result without immediately reformulating. The failure story then collapses several queries into a single unsuccessful task and pushes the metric down, so more is unambiguously better.
- Is queries per user worth reporting at all, then?Yes, as a diagnostic. Once a well-directed criterion has told you whether the change helped, movement in queries per user helps explain how — the same rise means engagement if the criterion improved and struggle if it did not. That is the driver-metric role. The restriction is only that it must not carry the decision by itself.
- What quick test flags a directionally ambiguous metric before you commit to it?Try to tell a plausible story in which the number rises and users are worse off. If you can, the metric cannot carry a pre-declared direction. Run the same test downward. Doing this in the design review is cheap; discovering it while arguing about a finished experiment is not.
Counting queries is like judging a library by how many aisles a visitor walks down. It goes up for someone happily browsing and for someone who cannot find the one book they came for.
saying these in an interview costs you the question
- Assuming more activity always means more engagement
- Deciding the metric's good direction after seeing results
- Pairing two metrics with no rule for conflicts
- Confusing an explaining metric with a deciding one
- Ignoring that failed attempts inflate activity counts