Asked which of 40M warehouse rows ever carried an invisible-codepoint span, what can you honestly claim?
answer
- refuse the word ever
- separate measurable from unmeasurable
- a census is population, not a sample
- eyeballing has zero detection power here
- the copies left the warehouse
basics
~20 sClaim only what an instrument supports: a codepoint census shows which rows carry such a span now, across the whole population. It cannot show what overwritten values held, or where copies already travelled. Say where the boundary is.
solid answer
~50 sSplit the question into what is measurable, what is unmeasurable, and what the organisation is going to be told. Measurable and cheap: a codepoint census over current stored values - every row, not a sample - identifying which fields carry non-rendering codepoints today. Unmeasurable with what you have: values overwritten or aged out before the census, and the copies. In this setting the assistant's answers quote the field, and those answers get saved as dashboard annotations and shared notebooks that other analysts open and re-prompt from, so the span has left the warehouse into stores nobody keeps an inventory of. And the thing to say out loud: a human spot-check of a few hundred rows buys nothing here, because reviewers read a rendering and the carrier was chosen to defeat exactly that. Offering it as partial assurance is worse than offering none, because it is believed.
go deeper
Know that a person reading rows cannot detect this class at all, so any assurance about a corpus has to come from a machine check over stored values.
Be able to describe the census as a full-population scan over current values and to say what it does not cover: overwritten history, and copies that left the warehouse.
Show that you scope a statement precisely - what was measured, when, over what - and that you separate naturally-occurring format characters from spans that read as directives.
Own the refusal. Decline the zero-power spot-check even when it is the politically easy deliverable, and convert 'we cannot know' into a priced menu of what could be known and what stays unanswerable.
## Why this is a judgement call and not a technical one The technical part is small. The hard part is that somebody - a data owner, a platform lead, a customer - is going to receive a sentence from you, act on it, and repeat it. The job is to make sure that sentence is true and that its boundaries are visible in it. Start by refusing the framing. 'Which rows ever carried it' contains a word - *ever* - that no instrument in the building can address. Answering the question as asked is how people end up making claims about history they never had access to. ## Three buckets **1. Knowable, cheaply, and over the whole population.** A codepoint census over current stored values. For each free-text column, count codepoints, bucket by Unicode general category and block, and flag rows containing format characters or Tags-block codepoints. This is a full scan, not a sample, and that distinction matters when you write it down: you are stating a population fact about the present, which is a strong claim and an honest one. It costs a pass over the warehouse and some review of the hits, since format characters occur naturally in pasted text. **2. Knowable only partially.** History. If rows are overwritten in place and no change history exists for the column, values that carried a span and were later edited are simply gone. Where an append-only history, a change-data-capture stream or dated snapshots exist, the census can run against those too, and your claim extends exactly as far back as they do - and no further. Say which it is. **3. Not knowable with what you have.** Where the span went after the assistant read it. The assistant's answers quote field values, and in this setting those answers are saved as dashboard annotations and pasted into shared notebooks that other analysts open and re-prompt from. The codepoints ride along through every one of those copies, invisible at each stop. Exports, ticket comments, screenshots-turned-transcripts and downstream marts extend the surface further. There is no inventory of these, so any statement about them is a guess. The honest move is to name the class of destination and decline to bound it. ## The spot-check question, which is the one you will actually be asked Someone will propose that a team eyeballs a few hundred rows so there is *something* by Friday. Refuse it clearly, and say why: reviewers read a rendering, and the span is a carrier selected precisely so that no rendering shows it. The detection power of that exercise is not low, it is approximately zero - and unlike a genuinely low-power sample, it will be reported upward as 'we reviewed the data'. A control with zero power that manufactures confidence is worse than nothing, and being the person who says so is the job. The same argument disposes of the softer variants: a second reviewer, a different tool, more time. All of them are the disqualified instrument, applied harder. ## What you actually hand over A defensible statement has four parts, and its usefulness comes from the last one: - **The measurement:** what was scanned, over what columns, at what time, by what test. - **The result:** the population-level count of rows currently carrying non-rendering codepoints, with the naturally-occurring hits separated from the ones that read as directives. - **The boundary:** the history you could not see, and the derived stores you could not enumerate. - **The decision it supports:** what the owner may now reasonably do, and what they may not say. ## The funding argument, since it is the real disagreement Extending coverage is not free, and it is somebody's budget. Re-running the census across snapshots, derived marts, annotation stores and shared notebooks is a project, not an afternoon; the owner will reasonably ask whether it is worth it. That is their call, not yours, and your contribution is to make the trade legible: what each additional store costs to cover, what class of question it would let them answer, and - crucially - what they will still not be able to say afterwards. A lead who converts 'we cannot know' into a priced menu of what could be known has done the principal-level work. A lead who lets 'we checked the data' stand because it is easier has not. ## The sentence to avoid 'The corpus is clean.' It has no instrument behind it, it will be quoted for years, and it is the exact claim the carrier was designed to make available to anyone who only looks.
- The owner wants one number by Friday. What do you give them?The population count of rows currently carrying non-rendering codepoints, with the date and the columns scanned attached, and a separate line for the hits that plausibly read as directives. One number without those qualifiers becomes 'the corpus is clean' within a week, which is the claim you cannot support.
- Why not report a sampled estimate from a manual review?Because the sample's detection power is not merely low, it is nil: reviewers see a rendering and the carrier draws nothing. A sampled estimate implies an instrument that detects at some rate, and here that rate is zero, so the estimate is not conservative - it is fabricated.
- How do you talk about the copies in annotations and shared notebooks?Name the class and decline to bound it. Say that any answer quoting the field carried the codepoints, that those answers were saved into annotation and notebook stores with no inventory, and that you can price a scan of the ones you can enumerate. Do not offer a number for stores you have not enumerated.
- Who owns the decision about how far the scan goes?The data owner funds it, so they decide. Your job is to make the menu legible: cost per additional store, the question each one would answer, and what remains unanswerable after all of it. Presenting an unbounded remediation as mandatory usually gets the whole thing declined.
saying these in an interview costs you the question
- Answers ever without noticing no instrument covers history
- Offers a manual review of sampled rows as partial assurance
- Reports a clean corpus after scanning only current values
- Ignores that the assistant's own answers carried the span onward
- Presents unbounded remediation instead of a priced menu