skip to content

Sessions completed per week is proposed as a tester productivity target — how do you respond?

level: principalimportance: should knowfreq 36%

answer

  1. Good as capacity, poisonous as a target
  2. The measured party writes the measurement
  3. Splitting a mission costs nothing
  4. Separate capacity, effectiveness, individual work
  5. Keep one signal nobody is graded on

basics

~20 s

Push back, and offer something better. A session count is a good planning input and a terrible individual target: it is trivially inflated by shorter, shallower, easier sessions, and the moment it drives appraisal the session sheets stop being honest.

solid answer

~50 s

The count is fine as a capacity figure — roughly how much chartered testing a team delivers in a week — and destructive as a per-person target, because every way of raising it is cheaper than testing well. A tester can split one mission into three, pick familiar areas, skip setup-heavy work, cut the debrief and stop pairing; each raises the number and lowers the value, and the artefacts that would expose it are written by the person being measured. So the counter-offer is: keep sessions as an aggregate planning number, report progress as sessions and areas with the obstacles attached, and judge individuals through debriefs and the readability of their sheets, which a lead can assess directly. If the real question is whether the team is delivering enough testing, the honest answers are the areas grid and the obstacle list.

go deeper

for a junior

Understand that a session count describes how much chartered testing happened, not how good it was, and that the same work can be reported as three sessions or one depending only on how the missions were split.

for a middle

Be able to name the concrete gaming levers — splitting charters, favouring easy areas, cutting setup and debriefs — and explain why the self-reported sheet cannot detect them.

for a senior

Show you can protect the reporting artefacts under pressure: keeping the time split unattributed, keeping debriefs alive, and reporting lost sessions with their obstacles rather than absorbing them quietly.

for a principal

Own the negotiation. Separate capacity from effectiveness from individual performance, offer the aggregate figure with its stated limits, name the trade between a number that rises and sheets that stay honest, and keep an unmeasured signal to detect drift.

### Why the proposal is tempting Sessions are the first thing exploratory testing has ever offered that looks countable. A lead who has spent two years unable to answer "how much testing did we do?" now has a unit, and the pull toward turning that unit into a target is immediate. The reasoning even looks sound: sessions are roughly uniform by construction, so more sessions should mean more testing. ### Why it collapses The uniformity is a convention, not a property of the world, and the person producing the number is the person being measured by it. Every lever that raises a session count is cheaper than testing well: - **Split the mission.** One coherent charter becomes three narrow ones, and the week's count goes from four to twelve without a minute of extra testing. Charter granularity is a judgement call, so this is undetectable from the outside. - **Choose the easy area.** Familiar, well-seeded, quick-to-set-up areas yield clean sessions. The nasty area with no test data — the one carrying the risk — costs a session and a half and yields a messy sheet, so it drifts to the bottom of the list. - **Skip the expensive parts.** Setup, careful notes and the debrief all consume the block. Under a count target, the block shrinks toward pure execution and the reporting quality dies, which removes exactly the artefacts that would reveal the problem. - **Stop pairing.** Two testers on one area produce one session; separately they produce two. The count rewards splitting up precisely where paired work is most valuable. - **Corrupt the time split.** The estimated setup / design / investigation breakdown only works because nobody is judged on it. Attach consequences and every sheet converges on a socially safe distribution, and the one number that reliably exposed a broken environment stops working. This is the ordinary pattern of a measure becoming a target and ceasing to be a good measure. It bites unusually hard here because the measurement instrument is a self-report from the measured party, and because testing has no accepted individual productivity metric at all — the evidence base for any of the candidates is thin and genuinely contested, which is worth saying out loud rather than replacing one confident number with another. ### What to offer instead Separate the two questions hiding in the proposal. *Capacity.* Sessions answer this well, in aggregate and unattributed. "This team sustains about eleven to fourteen sessions a week" is a real forecasting input for how much chartered testing a release can get, and it is safe because no individual is graded on it. *Effectiveness.* Report it as areas and obstacles rather than volume: which areas have had attention, which have had none, what is currently making sessions expensive. A team whose session count fell from fourteen to eight because the payment sandbox was down for a day and a half has not become less productive, and the obstacle list says so. *Individual performance.* Judge it where a lead has direct evidence: the debrief. Does this tester's reasoning hold up when questioned, are their notes usable by someone else six weeks later, do they take the ugly areas or avoid them, do their charters improve after a finding? That is qualitative, it does not aggregate into a dashboard, and it is the honest answer. ### Holding the line in practice Agreeing to publish the count as a team-level capacity figure, with the areas grid attached and a stated rule that it is never broken down per person, usually satisfies the real need behind the request. Where the request survives that, it is worth naming the trade explicitly to the person asking: you can have a number that goes up, or you can have session sheets you trust, and within about two quarters you will not have both. A lead who insists after hearing that has made an informed choice, and the useful move then is to keep an unmeasured signal — sampled debriefs, or simply reading sheets — so you can tell how far the reported number has drifted from the work.

  • Is there any version of a session count you would agree to publish?
    Yes — an aggregate, team-level figure used for capacity planning, published with the areas grid and the obstacle list beside it, and never broken down per tester. That form answers the legitimate question of how much chartered testing a release can expect, while removing the individual incentive that does the damage. The condition worth insisting on is that it is reported as effort applied, never as a completeness claim.
  • How would you detect that a session count has already started drifting from the work?
    Look for charter granularity shrinking over time, average session length creeping down, the on-charter share moving suspiciously toward 100%, debriefs shortening, and the estimated time split converging on the same figures across every sheet. The strongest single check is qualitative: pick two sheets at random and see whether a colleague could reconstruct what was covered from them.
  • What would you measure instead if you genuinely need to compare two teams?
    Prefer comparing the situations rather than the people: which areas each team has covered, what obstacles each is carrying, how long obstacles survive before someone clears them. Those describe the system the teams work in, which is usually where the difference actually lives. Any individual-level testing productivity metric should be offered with the caveat that no well-evidenced one exists.

Counting sessions to rate testers is like rating surgeons by operations per week: the number moves fastest by taking the easy cases and skipping the preparation.

saying these in an interview costs you the question

  • Accepting the target because sessions are uniform by design
  • Assuming testers would not game a metric they self-report
  • Ignoring that shorter charters raise the count for free
  • Turning the estimated time split into a timesheet
  • Offering defect counts as the alternative target
  • Refusing the request without proposing anything usable

context