skip to content

When does planting more documents in a review corpus stop being worth what it costs?

level: principalimportance: nice to knowfreq 26%

answer

  1. rising cost against an uneven payoff
  2. thin neighbourhoods fall first
  3. attributable writes accumulate
  4. held ground decays as content arrives
  5. the count is what coverage is worth

basics

~20 s

When the next document costs more than the question it would move is worth. Cheap, thinly covered questions go first; each further one is more crowded, more attributable, and needs upkeep as genuine content keeps arriving.

solid answer

~50 s

Treat it as a rising marginal cost against an uneven value curve. The thinly covered questions fall first and cheaply; after those, each additional neighbourhood is one the genuine corpus already answers well, so it costs several aimed documents rather than one. Meanwhile the value of moving a question is very uneven — a handful of buying-decision questions carry almost all of it, and the long tail carries none. Two more costs sit on top: every post is attributable to a listing, so the accumulated write history is itself an exposure, and ground held in a live corpus decays as genuine content keeps arriving, making this a standing bill rather than a one-off. The honest stopping line is the point where those curves cross, and in a sanctioned exercise the count you spent to reach it is the result — it is what a corpus's existing coverage is demonstrably worth.

go deeper

for a junior

Know that this is an economic question: more documents cost more, and the questions worth moving are far fewer than the questions available to move.

for a middle

Be able to explain why the cost per additional question rises — cheap thin neighbourhoods are consumed first — and why identical posts do not lower that cost.

for a senior

Show how you would set a stopping line in practice: decision-adjacent questions first, measured per neighbourhood, with upkeep and attribution counted rather than assumed away.

for a principal

Own the tradeoff and the claim. Report the shape of the cost curve rather than a number, say what the count does and does not establish, and be willing to conclude that the method itself is the wrong purchase.

## Why this is a judgment and not a calculation The mechanics are settled by the time this question arrives: one passage owns one query neighbourhood, copies buy nothing, breadth costs distinct documents. What remains is a decision somebody has to own — how far to push, and what to claim from where they stopped. It has no single right answer because the inputs are a cost curve, a value distribution and a tolerance for exposure, and reasonable people weigh them differently. ## The cost curve rises, and it rises for a structural reason The order in which questions fall is not random. Thin neighbourhoods go first: a question the corpus barely answers can be taken with one aimed document. Once those are gone, what remains is the well-covered material — exactly the questions that many genuine buyers already answered in detail — and each of those needs a plant that clears a stack of established passages, often several documents per phrasing. So the count is not linear in the questions moved. The first five neighbourhoods might cost seven documents and the next five might cost forty. Anyone quoting an average is hiding the shape that matters. ## The value curve is uneven the other way Moved questions are not worth the same. A handful sit immediately before a purchase decision — fit, compatibility, safety, whether the cheaper option is equivalent — and those carry nearly all the payoff, because a false answer there changes what somebody does. The rest of the family is decoration. That asymmetry is what makes the stopping decision tractable at all: the right move is usually to stop far earlier than "we covered the family", once the decision-adjacent questions are held. ## Three costs that are not document counts 1. **Attribution accumulates and cannot be withdrawn.** Every post is tied to a listing and an account. A campaign that reaches thirty documents has left thirty attributable writes, and the pattern of a single controlled set of listings producing consistent, similar-purpose content is more visible than any single post. Later deletion does not undo the fact that the write happened. 2. **Held ground decays.** A live marketplace keeps producing genuine reviews and answers. A neighbourhood held today can be outranked next quarter by nothing more than ordinary buyer activity, so coverage is a subscription, not a purchase. Budget for upkeep or accept a shrinking result. 3. **The pipeline can move under you.** Chunking, encoder, or ranking changes shift every position at once, and none of that is observable from the write side. The dependency is on a configuration nobody involved controls. ## The alternative-use test The strongest form of the stopping question is not "is one more document worth it" but "is this construction still the cheapest route to this payoff". Corpus poisoning buys influence over answers to questions people ask. If what is actually wanted is influence over one specific decision, and there is a shorter path to it, the marginal document is the wrong purchase however cheap it is. A lead who cannot say when their favourite method stops being the right method will over-invest in it. ## What the stopping point is worth as a result In a sanctioned exercise the count is the deliverable's substance. "Eleven of forty enumerated buyer questions moved, for thirty-one distinct posts, of which the first six questions cost seven posts and the rest cost twenty-four" is a statement about how much protection genuine coverage actually provides — and it is falsifiable, because someone can re-run the same phrasings later. Be careful about what is being claimed on either side. A count is not an assurance that a cheaper route does not exist; someone better at aiming passages may move the same questions for fewer documents, and a corpus with different coverage would price differently. Equally, a small count of moved questions is not "the assistant is fine" if the moved ones were the decision-adjacent ones. Report the shape of the curve, not a single number, and let the person who owns the product weigh it. ## What a weak answer looks like It treats the budget as the only constraint and pushes until the money runs out; it quotes an average cost per question and hides the rising curve; it ignores that the writes are attributable and permanent; and it assumes ground taken stays taken in a corpus that people are still adding to every day.

  • Why does the cost per additional moved question rise rather than stay flat?
    Because the cheap neighbourhoods are consumed first. Once the thinly covered questions are taken, what is left is the material many genuine buyers already answered in detail, and clearing a tight band of established passages needs several aimed documents per phrasing instead of one. An average cost per question hides exactly that.
  • Why is coverage of a live corpus a recurring cost rather than a one-off purchase?
    Because genuine content keeps arriving. New reviews and answers land in the same neighbourhoods and can outrank a plant without anyone intervening, and pipeline changes can reshuffle positions wholesale. Coverage held today is a lease with an unknown term, so any claim about it should carry a date.

saying these in an interview costs you the question

  • Pushes until the budget runs out with no value test
  • Quotes an average cost per question and hides the rising curve
  • Ignores that every post is attributable and permanent
  • Assumes ground taken in a live corpus stays taken
  • Treats a small count of moved questions as reassurance regardless of which ones

context