How would you estimate how many distinct planted posts a family of buyer questions requires?
answer
- the genuine corpus is the competition
- one intent, several phrasings
- you probe the product, not the index
- read how tight the score band is
- sum per neighbourhood, allow little overlap
basics
~20 sEnumerate the phrasings buyers use, then gauge genuine competition for each by probing the assistant itself. Well-covered questions need several distinct passages each; thinly covered ones need one. The count is the sum, not a constant.
solid answer
~50 sSizing is a competition estimate, not a formula. Start by enumerating the question family the way buyers phrase it, because one phrasing is not the family. For each phrasing, probe the assistant from the outside — you cannot query the index — and see how much genuine material it draws on: how consistent the answers are, how many distinct sources they reflect, whether the top of the ranking looks tight or thin. A phrasing that already has several detailed genuine passages sitting in a narrow score band needs a plant that outranks a stack of them, and often more than one. A phrasing barely covered needs one. Sum across the family, add overlap where one passage happens to cover two phrasings, and you get a count of distinct documents — each an attributable post. That number is the cost line, and it is the thing worth reporting.
code
json · 7 lines[
{"chunk_id": "c-4471", "score": 0.86, "source": "buyer_review", "listing_id": "L-1180"},
{"chunk_id": "c-9032", "score": 0.84, "source": "buyer_review", "listing_id": "L-0442"},
{"chunk_id": "c-2185", "score": 0.81, "source": "seller_qa", "listing_id": "L-0442"},
{"chunk_id": "c-7710", "score": 0.79, "source": "seller_qa", "listing_id": "L-6631"},
{"chunk_id": "c-1096", "score": 0.62, "source": "buyer_review", "listing_id": "L-6631"}
]go deeper
Know that the number depends on how much genuine content already answers the question, and that a well-covered question is expensive while a barely-covered one is cheap.
Explain how you would enumerate phrasings and infer competition from the assistant's own answers, and why a tight band of similar genuine passages raises the count for that phrasing.
Demonstrate the full estimate and its caveats: per-phrasing measurement, little assumed overlap, and honest separation between what you observed in the product and what you inferred about ranking.
Own the interpretation. A per-question-family document count is an economic statement about how much genuine coverage is worth, and it is the number an owner can actually weigh.
## What the number is actually a measure of "How many documents do I need?" has no fixed answer, because it is not a property of the pipeline. It is a property of **how well the existing corpus already answers the target questions**. The genuine documents are the competition; the count you need is whatever it takes to outrank them, question by question, phrasing by phrasing. That reframing is most of the skill. Someone who answers "about five" has not understood the question; someone who answers "depends on how crowded each phrasing already is, and here is how I would find out" has. ## Step one: enumerate the family, not the question Buyers do not ask a question one way. "Does it fit a standard shelf", "what are the dimensions", "will it sit under a 30cm gap" may be one intent and three neighbourhoods. Before counting documents, write out the phrasings that actually matter — the ones a buyer would type before spending money, since those are the answers with a payoff attached. A family of five real phrasings and a family of forty are different projects. ## Step two: measure competition from outside The write channel is user content under listings; there is no read channel into the index. So competition has to be inferred from the product's own behaviour: - Ask each phrasing and see what the assistant draws on. Consistent, specific, detail-rich answers imply several strong genuine passages. Vague, hedged or "I could not find" answers imply thin coverage — and thin coverage is the cheap target. - Vary the phrasing slightly and watch whether the answer's substance changes. If it does, the neighbourhoods are distinct and each will have to be paid for separately. - Where the product shows which content it used, read the spread of sources rather than the content: many distinct sources means many established competitors. Every one of these is an inference about ranking, not an observation of it. Say so when reporting. ## Step three: read the shape of the competition Given a candidate set for one phrasing, the useful question is not "did my passage appear" but "what would it have to beat". A tight band of similar scores across several genuine passages means a plant has to clear all of them to change what the generator sees, and one plant that lands mid-band changes nothing. A single strong passage above a long drop means one well-aimed document can plausibly take the top. ## Step four: sum, with overlap Count is per neighbourhood, minus whatever a single passage happens to cover twice. Overlap is real but modest: a passage broad enough to cover several phrasings is by construction less specific, which usually costs it rank against the specific genuine passages in each of them. Assume less overlap than you would like. ## What the count is worth The resulting number — say, "thirty-one distinct posts to move eleven of the forty phrasings we enumerated" — is the substance of the exercise. It converts an alarming demonstration into an economic statement about how much genuine coverage protects a question family and how much it costs to overcome. It also gives whoever inherits the corpus something falsifiable: re-run the same phrasings later and see whether the count still holds. ## What the estimate does not prove - Probing the assistant measures **the product's answers**, not the index. A change in chunking, retrieval configuration or the encoder can move every one of these estimates without anyone touching the corpus. - A phrasing that moved once may not move again. A probabilistic generator can decline to use a retrieved passage on a later run, and retrieval itself can shift as the corpus grows. - Documents planted and questions moved are different numbers, and only the second is the effect. Near-copies inflate the first and not the second. - A high score on a planted passage proves it matched the phrasing, never that a reader will believe it or act on it. Payoff is a separate question from rank.
- You cannot query the index. What is the strongest signal available from outside?The assistant's own answers across deliberately varied phrasings. Specific, consistent, detail-rich answers imply several strong genuine passages competing; hedged or empty answers imply thin coverage. Where the product shows which content it used, the spread of distinct sources is the closest thing to a competition count you can get — all of it inference about ranking, never observation of it.
- Why does one passage rarely cover a whole question family?Because a passage broad enough to sit near several phrasings is less specific, and specificity is what wins rank against the genuine passages in each neighbourhood. Breadth inside one document trades against rank in every neighbourhood it straddles, so real coverage usually costs one aimed document per phrasing.
saying these in an interview costs you the question
- Gives a fixed number without reference to existing coverage
- Tests one phrasing and calls it the question family
- Claims to know the ranking rather than inferring it from answers
- Counts documents planted instead of questions moved
- Assumes broad overlap between neighbourhoods to cut the count