skip to content

How do you stop a bug-bash capture channel from filling with duplicate reports?

level: seniorimportance: should knowfreq 39%

answer

  1. Duplicates are arithmetic, not carelessness
  2. Cut them before, during and at merge
  3. A visible running list scans itself
  4. One watcher who does not hunt
  5. Merge while the reporter is still there

basics

~10 s

Post known issues before the start, fix one short title shape with an area tag, keep the growing list visible so people scan before filing, and put one person on duty merging duplicates live.

solid answer

~50 s

Duplicates are the tax on the format: many people looking at one build will find the same loud things. Cut them at three points. Before the event, publish what is already known so nobody refiles it. During it, require a one-line title in an agreed shape with an area tag, and keep the running list where everyone can see and scan it — a visible list is the cheapest deduper there is. And staff a duplicate-watcher who does not hunt at all: they merge as reports land, ask for the one missing detail while the reporter is still at the keyboard, and call out in the channel when an area has gone quiet. Merging live is far cheaper than merging cold, because the reporter's memory is still available. Accept that some duplicates survive — chasing zero slows the hunt, which is the thing you are paying for.

code

pseudocode · 16 lines
pseudocode
STOPWORDS = {"the", "a", "on", "when", "is", "not", "of", "in"}

function group_key(title):
    tag, rest = split_once(title, "]")
    words = lowercase(strip(rest)).split()
    words = [w for w in words if w not in STOPWORDS]
    return strip(tag, "[") + ":" + join(sort(first(words, 4)), "-")

buckets = empty_map()
for report in incoming_reports:
    key = group_key(report.title)
    buckets[key] = append(buckets.get(key, []), report)

for key, group in buckets:
    if count(group) > 1:
        flag_for_watcher(key, group)

go deeper

for a junior

Know the participant-side habits: use the agreed one-line title shape with the area tag, glance at the recent channel before filing on a busy screen, and file rather than agonise when unsure.

for a middle

Explain the mechanics — a published known-issues list, a fixed title shape that makes scanning and grouping work, a visible running list, and a named duplicate-watcher who merges as reports land.

for a senior

Demonstrate running it live: keeping the capture bar low on purpose, spotting an environment-wide fault from the stream before a dozen copies arrive, rebalancing participants when one area floods, and refusing to work findings up during the window.

for a principal

Own the accounting. Be able to say which duplicates were preventable and which are the format's irreducible tax, what you report as the outcome instead of the raw count, and how you keep the aftermath owned so attendance survives to the next event.

## Why duplicates are structural, not sloppiness Put twenty-three people on one build and they will not distribute themselves evenly over the defect population. Defects differ enormously in how easy they are to trip: a broken date format on the busiest screen will be found by nine people in the first ten minutes, while a permission edge on a rarely used reconciliation report is found by nobody unless someone is sent there. So the arrival curve of a bug bash is front-loaded and lumpy, and duplicates are the arithmetic consequence of the format rather than evidence that participants were careless. That framing matters because it tells you where to spend effort: reduce the *cost* of duplicates, and prevent the avoidable ones, rather than lecturing participants. ## A worked shape Take a warehouse stock ledger, a **94-minute** window, twenty-three participants across nine areas. A realistic pile looks like this: 133 raw reports arrive; after merging, 71 are distinct; 22 of those were already open before the event; 12 turn out to be one intermittent save timeout in the shared environment reported under a dozen different titles; the rest are new. Read the numbers and the costs are obvious. The 22 already-open ones were fully preventable with a published known-issues list. The 12 timeout reports were preventable with a smoke pass and one announcement. The remaining duplicates — the same date-format defect from five different areas — are the irreducible tax. Sorting the pile cold, days later, means reading 133 items with no reporter to ask; sorting it live means most merges happen in seconds while the two reporters are both present. ## The three points of intervention **Before.** Publish known issues in the channel and in the briefing, including anything the smoke pass turned up. State the areas and who has each, so a participant seeing something outside their area knows whose it is. **During.** Fix the title shape — area tag, screen or object, what went wrong, in that order — because a consistent prefix is what makes scanning and text search work. Keep the running list visible: a channel where each report is one message, or a board that the watcher updates. Ask participants to scan the last few minutes of the channel before filing anything on a busy screen. Keep the capture bar low: the point during the window is to record the observation, not to work it up. **Live merging.** Staff one person as duplicate-watcher who does not hunt. Their job is to read every incoming report, merge the obvious repeats immediately, ask the one clarifying question that makes a vague report usable, and watch the area distribution — when six reports come from receiving and none from transfers, they say so in the channel and someone moves. The watcher role is also the natural home for spotting an environment-wide fault early: they are the only person who sees the whole stream. ## Titles and normalisation A consistent title shape produces a cheap grouping key. If reports look like `[adjustments] save spinner never ends on quantity change`, then lowercasing, stripping the tag and matching on a few distinctive words groups most repeats mechanically, leaving the watcher to judge the ambiguous ones. This is a helper, not a replacement for a human: two reports of the same underlying fault often describe two different symptoms, and only a person notices that the frozen spinner and the duplicated ledger row are the same timeout. ## What not to do Do not make the filing form long in the hope of getting cleaner reports. A heavy form during a time-boxed hunt suppresses reporting — participants who are visitors will simply not bother — and you lose observations, which is a worse trade than absorbing some duplicates. Do not ask participants to search an issue tracker before filing, either; searching is slow and they do not know the vocabulary the tracker uses. Do not rank participants by report count: it directly incentivises volume and re-reporting. Also resist finishing the sort during the window. The temptation is to have the watcher fully work up each finding as it arrives; that turns one person into a bottleneck and stalls the stream. Capture broadly now, work items up afterwards, with a named owner and booked time. ## Judging afterwards The honest measures of a bug bash are the distinct count after merging, how many of those were in areas the designed suite claims to cover, and how many turn out to matter once someone looks properly. A pile of 133 reports that collapses to 71 distinct, of which 22 were already known, is a normal and healthy outcome — reporting the 133 as the result of the event is the mistake, because it flatters the event and misleads whoever reads the summary.

  • Why merge duplicates live rather than sorting the pile afterwards?
    Because the reporters are present. A live merge is a question and an answer in seconds; a cold merge is guessing from two thin descriptions with nobody to ask. Live merging also surfaces environment-wide faults early enough to announce them, which stops a dozen more arriving, and it keeps the visible list short enough for participants to scan.
  • Would you make the report form richer to raise report quality during the window?
    No. A heavy form during a time-boxed hunt suppresses reporting, especially from visiting participants, and lost observations cost more than duplicates do. Keep capture to a one-line title, an area tag and enough to find it again; the working-up happens afterwards with a named owner and booked time.
  • What do you report as the outcome of a bug bash, and why not the raw count?
    Report the distinct count after merging, how many were already known, how many fell in areas the designed suite claims to cover, and what the aftermath found to matter. The raw count mostly reflects attendance and the filing bar, so quoting it flatters the event and misleads anyone comparing one bug bash with another.

saying these in an interview costs you the question

  • Blames participants for duplicates instead of the format
  • Adds a long filing form and loses reports
  • Leaves all merging until days after the event
  • Reports the raw submission count as the result
  • Rewards whoever filed the most reports
  • Lets one person work up findings live and stall the stream

context