skip to content

A team wants one classifier class per customer recipe; internal users can query it freely. What's your call?

level: principalimportance: should knowfreq 31%

answer

  1. the label space is the control
  2. a shared console, many customers
  3. trusted staff is not the question
  4. a membership floor, recorded and enforced
  5. the thin tail grows every quarter

basics

~20 s

Per-customer classes with a handful of examples each turn a shared console into a disclosure channel: the best-scoring input for such a class is that customer's signature. Own the class cardinality decision, not the attack.

solid answer

~50 s

Say clearly what the design does: any user who can query the console for per-class scores can search for the input that class scores highest, and for a class built from a few of one customer's records that prototype **is** that customer's signature. Then treat it as a trade, because the granularity is real product value. The levers are a minimum-membership rule before a narrow class may exist, folding the thin tail into coarser labels, or scoping per-customer class scores to that customer's own tenancy. What I would not accept is "internal users are trusted" as the answer — on a shared multi-customer console, the confidentiality obligation runs between customers, and trust in staff does not discharge it. I would put in writing which classes may be per-subject, who signed that, the membership floor, and a review each time new classes are added, because the thin tail grows with every triage improvement.

go deeper

for a junior

Recognise the shape of the risk: a class built from one customer's few records means the model can be asked what that customer's data looks like, on a console other customers' staff also use.

for a middle

Explain why the exposure comes from the label space rather than the attack, and why hiding scores raises cost without removing a property of the trained function.

for a senior

Be ready to run the review: bucket the classes by membership, name the ones effectively equal to one subject, and propose the coarser label or tenancy-scoped scores with the granularity cost stated.

for a principal

Own the decision and its durability — who signs which classes may be per-subject, the membership floor and why it exists, the recurring histogram review, and the framing that keeps this a product call rather than a veto.

## What is actually being proposed A foundry's inspection console classifies wafer maps by defect signature and shows per-class scores to internal users across customers. The dashboard team wants finer triage: instead of forty generic signatures, a class per customer process recipe. Several of those classes will start life with a handful of examples. This is a good product idea. It is also a privacy decision, and somebody has to own it as one. ## The mechanism, in one paragraph An adversary here needs nothing exotic: query access and the returned per-class scores. Searching for the input a class scores highest recovers that class's prototype — the model's idea of the class. For a large generic class that prototype is an average matching nobody. For a class built from twelve maps belonging to one customer, the average and its members are nearly the same object, so the prototype is that customer's proprietary signature. The label space, not the attack, is what makes those two cases different. ## Why the usual answers do not settle it **"Internal users are trusted."** This changes who, not whether. On a shared console the confidentiality obligation runs between customers; an employee legitimately viewing customer A's dashboard is not thereby authorised to reconstruct customer B's recipe signature. Staff trust does not discharge a contractual duty to a third party. **"We will meter or log the queries."** Useful, and it does not address a design where the disclosive object is a property of the trained function. Metering raises the cost of a search that should not be possible for that user at all. **"We will only return the top class, not scores."** It genuinely raises the bill — the fine-grained signal a search climbs is what you removed — but it is a cost control, not a boundary, and it takes away most of what makes the console useful. Offer it as a trade, never as a fix. **"Nobody would bother."** Perhaps. That is a statement about motivation among competitors sharing a foundry, which is exactly the population with a motive. ## The decision I would actually make Separate the classes into three buckets and decide once, in writing: 1. **Generic signatures with broad membership** — ship, no constraint. Their prototypes are averages. 2. **Narrow classes above a stated membership floor** — ship, with the floor recorded and enforced at class-creation time, and with a note that the floor is a privacy parameter, not a modelling one. 3. **Classes that are effectively one customer** — either fold into a coarser label, or scope their scores to that customer's own tenancy so the console shows them only where the data already belongs. The cost is honest and should be stated: option 3 loses triage granularity precisely where the team wanted it most, and tenancy-scoped scores mean a cross-customer analytics view the product team may have promised somebody. Naming that cost is the point of the decision, not an argument against it. ## What to put in writing - Which class definitions are permitted to be per-subject, and who approved that. - The membership floor for creating a new narrow class, and the fact that it exists for disclosure reasons so a future engineer does not tune it away as a modelling threshold. - A recurring review of the class-membership histogram, because the thin tail grows every time somebody adds a class for better triage — a launch-time check that is never repeated will be wrong within two quarters. - What the console returns, and to whom, as a product contract rather than a configuration detail. ## The framing that wins the room Do not argue that the model is insecure; argue that the label space chooses what the model can be asked to describe, and that this proposal asks it to describe individual customers. That reframes the conversation from "security is blocking a feature" to "which of these classes may exist", which is a decision the product owner can actually make and can be held to.

  • The team says all console users are employees under NDA. Does that close it?
    No. It changes who could do it, not whether the design exposes it. On a shared multi-customer console the obligation runs between customers, and an employee entitled to view one customer's dashboard is not entitled to reconstruct another's signature. Staff trust does not discharge a duty owed to a third party.
  • They offer to return only the top class instead of all scores. Do you take it?
    As a trade, not a fix. Removing per-class scores strips the fine-grained signal a search climbs and raises the cost a lot, but the prototype remains a property of the trained function, and the scores are usually why the console is worth using. Say both halves out loud.
  • What do you commit to reviewing after launch?
    The class-membership histogram, on a schedule. Every triage improvement adds narrow classes, so the thin tail grows continuously; a one-off launch check is stale within a couple of quarters. I would also record the membership floor's rationale so it is not later retuned as a modelling knob.
  • How do you frame this so it is not read as security blocking a feature?
    Argue about the label space rather than the model. The classes decide what the model can be asked to describe, and this proposal asks it to describe individual customers. That turns it into a question the product owner can decide and be accountable for, with a stated cost in triage granularity.

saying these in an interview costs you the question

  • Accepts trusted internal users as the answer
  • Treats query logging as addressing the design
  • Presents hiding scores as a fix rather than a cost
  • Checks class membership once at launch only
  • Frames it as insecure model rather than label-space choice

context