skip to content

When is paying an external crowd for device and locale breadth worth it?

level: principalimportance: should knowfreq 31%

answer

  1. You are buying breadth, not understanding
  2. External testers have no domain oracle
  3. Count reviewer hours, not just the invoice
  4. Acceptance ratio decides the economics
  5. Pilot it, then judge on your own numbers

basics

~10 s

Buy a crowd when the real-device, network and locale breadth you need exceeds what you can staff, the build is safe to expose, and you have reviewer capacity to filter a low acceptance ratio.

solid answer

~50 s

A paid crowd sells one thing well: many real people on real devices, networks and locales you do not own, at a time you choose. It sells other things badly — anything needing domain judgement, long setup or knowledge of what the product should do, because an external tester has no oracle for that. So the decision turns on whether your unmet risk is breadth-shaped. Then weigh three underestimated costs: the acceptance ratio, since most submissions are duplicates, known issues or non-defects; your own reviewers' time filtering them, which no invoice shows; and the exposure of a pre-release build and account access to people outside the company. Run a paid pilot on one release, measure accepted-unique per reviewer-hour and cost per accepted finding, and decide on your own numbers — published yield figures vary wildly.

code

pseudocode · 12 lines
pseudocode
submissions      = 312
accepted_unique  = 61
vendor_cost      = 4_180
reviewer_minutes = 6.4 * submissions
reviewer_rate    = 0.95   # cost per minute of internal reviewer time

internal_cost = reviewer_minutes * reviewer_rate
total_cost    = vendor_cost + internal_cost

print("acceptance ratio  ", accepted_unique / submissions)
print("cost per accepted ", total_cost / accepted_unique)
print("internal share    ", internal_cost / total_cost)

go deeper

for a junior

Know what a paid crowd is and the one thing it is good at: many real people on real devices, networks and locales you do not own. Knowing that most submissions are filtered out is enough at this level.

for a middle

Explain the mechanics of an engagement — a scoped brief, a safe build, credentials, a known-issues list — and why an external tester finds locale and device defects but not domain-correctness ones.

for a senior

Show you have filtered the output: reviewer time per hundred submissions, how you separate duplicates and non-defects, and how you brief so the pool works the areas you actually care about.

for a principal

Own the decision and its economics: whether the risk is breadth-shaped at all, contract shape and the behaviour it incentivises, data and confidentiality exposure, the pilot you would run, and the numbers that would end the engagement.

## What a crowd actually sells A paid crowd is a marketplace: you describe a target, a build and a window, and a pool of external testers works it for a fee — per accepted report, per hour, or on a subscription. What you are buying is **breadth you cannot staff**: real handsets and older operating-system versions nobody in the office owns, genuinely slow or lossy mobile networks, real locales with the keyboards, input methods, calendars, number formats and text lengths that come with them, and a time-zone spread that lets a window run overnight. What you are not buying is understanding. An external tester has no oracle for your domain. Given a warehouse stock ledger, they cannot tell whether a put-away that splits across two locations is correct, because correctness there lives in the operational rules and a receiving manager's expectations. They can tell you that the quantity field rejects a comma decimal separator in a locale that uses one, that the reason-code dropdown clips its longest translated label, and that the adjustment screen is unusable one-handed on a small display. Those are real defects with real customer impact and they are exactly the ones an internal team, all on the same recent hardware in one language, systematically fails to see. ## The three costs that decide it **Acceptance ratio.** Crowd output is dominated by duplicates, already-known issues, misunderstandings of intended behaviour and environment problems. Any commercial claim about yield should be treated as unverified; the only number worth planning on is the one your own pilot produces. As an illustrative shape from a single pilot on the stock ledger — not a benchmark — 312 submissions produced 61 accepted distinct findings, of which 14 were locale or text-length defects and 9 were device-specific rendering or input faults across 18 device models. That is roughly one accepted finding in five submissions, and it was judged worth repeating for the mobile client only. **Your reviewers' time.** Every submission is read by someone who knows the product. That internal cost never appears on the vendor invoice and is routinely left out of the business case. Measure it: reviewer-hours per hundred submissions, and therefore reviewer-hours per accepted finding. If filtering the crowd consumes more skilled time than the findings save, the answer is no regardless of the per-report price. **Exposure.** External people get a build, credentials and a view of your data shapes before release. That means contractual confidentiality, disposable accounts, synthetic data, no production connectivity, and a clear position on what an external tester may keep, screenshot or publish. In regulated or competitively sensitive contexts this can be disqualifying on its own, and it is the argument a security or legal reviewer will make. ## Choosing the contract shape Pay-per-accepted-report puts the filtering risk on the vendor and caps your spend per finding, but it incentivises volume, re-reporting and shallow surface defects, and it drags you into arguments about what counts as accepted. Pay-per-hour buys directed effort — you can ask for one specific flow on one device family — but you pay for time whatever comes back, so it only works with a trusted, briefed pool. Many teams use per-hour for a narrow, well-specified sweep and per-report for a broad one-off before a launch. Whichever shape, expect to write the scope, the in-scope build, the areas that are off-limits and the known-issues list — a crowd briefed badly returns noise, and the brief is your work, not the vendor's. ## Where it fits against the internal bug bash The two formats answer different questions and are often confused. An internal bug bash buys **domain-aware fresh eyes** cheaply, in one room, on one build, with people who know what the product is for. A paid crowd buys **environment diversity** at a price, from people who do not. If the outstanding risk is 'we do not know how this behaves in fourteen locales on hardware we do not own', a crowd is the honest answer. If it is 'we do not know whether this workflow makes sense to someone who was not in the design meetings', the crowd will not tell you and colleagues will. ## Deciding, and deciding again Treat the first engagement as an experiment with a stated hypothesis: which risk it addresses, what acceptance ratio and reviewer cost would make it repeatable, and what result would end it. Run it on one release, count accepted-unique per reviewer-hour, cost per accepted finding, and how many of those findings a designed suite could plausibly have caught. Then decide with your own numbers. The common failure is buying a crowd as a substitute for test capacity in general. It is not a general capacity purchase; it is a narrow instrument for a breadth-shaped gap, and used as a general one it produces a large invoice, a large pile, and a team that spends its week reading submissions instead of testing.

  • How do the pay-per-report and pay-per-hour contract shapes change crowd behaviour?
    Pay-per-accepted-report shifts filtering risk to the vendor and caps cost per finding, but rewards volume and shallow surface defects and invites disputes over acceptance. Pay-per-hour buys directed effort on a flow or device family you specify, but you carry the risk of an empty result, so it needs a briefed, trusted pool. Narrow sweeps suit hourly; broad pre-launch passes suit per-report.
  • What would make you refuse a crowd engagement outright?
    An unreleased build or data whose exposure is unacceptable contractually or under regulation; no internal reviewer capacity to filter submissions; or a risk that is not breadth-shaped — a domain-judgement or long-workflow gap that external testers have no oracle for. A crowd bought as generic test capacity produces an invoice and a pile, not information.
  • What would you measure in a first paid pilot before committing?
    Accepted distinct findings per reviewer-hour, cost per accepted finding including internal review time, the mix of findings by class — locale, device, input, general — and how many the existing designed suite could plausibly have caught. State up front what result ends the experiment, so the decision is not made by momentum.

saying these in an interview costs you the question

  • Buys a crowd as general test capacity
  • Ignores internal reviewer time in the business case
  • Expects external testers to judge domain correctness
  • Quotes vendor yield figures as if verified
  • Exposes a pre-release build with no confidentiality terms
  • Judges the engagement on raw submission volume

context