skip to content

A reconstruction from a classifier's scores resembles a real wafer in one run of five — what can you claim?

level: seniorimportance: nice to knowfreq 24%

answer

  1. one run of five is a number
  2. compare against the class average
  3. the measure comes before the look
  4. resemblance judged by a hopeful reviewer
  5. state the strongest claim you can support

basics

~10 s

Only that one search run produced something a reviewer judged similar. Without a stated similarity measure and a comparison against the class average, the resemblance may be the class prototype plus reviewer expectation.

solid answer

~50 s

Report it as what it is: five restarts, one output a reviewer called similar, similarity assessed by eye. That is a lead, not a finding. Three things turn it into one. First, a **similarity measure fixed before you look**, applied to every restart, not the one you liked. Second, a **baseline**: the class average and a held-out record that is not a class member. If the output is no closer to the specific wafer than to the class average, you have recovered a prototype and nothing more — which is the expected result. Third, the **class membership count** stated beside the number, because a three-member class makes its average look like a member by construction. Also report the four failed restarts; a one-in-five reproduction rate is a number the reader needs, not an embarrassment to hide.

code

text · 13 lines
text
class: "recipe-C tear signature"   members in class: 12
access: per-class scores from the inspection console, no weights

restart  steps   final class score   reviewer: looks like a real wafer?
1        2000    0.97                no
2        2000    0.98                no
3        2000    0.96                yes
4        2000    0.97                no
5        2000    0.98                no
...
# no similarity metric recorded
# no comparison against the class average
# no comparison against a non-member record

go deeper

for a junior

Know that a single striking result is not evidence yet, and that the class average is the thing any reconstruction must be compared against before anyone calls it somebody's record.

for a middle

Be able to name the baselines: output versus specific record, output versus class average, output versus a non-member. Explain why the middle one usually explains the result.

for a senior

Demonstrate write-up discipline under pressure — fixed measure, all restarts reported, reproduction rate stated, and the strongest claim the evidence actually supports rather than the most quotable one.

for a principal

Own the credibility cost. A team that ships the top claim on weak evidence gets the underlying real issue dismissed with it, so the reporting standard is an organisational asset you set, not a personal habit.

## The situation You run a search against an inspection console that returns a score for every defect class. You pick a class, run five restarts of the optimisation, and one of the five outputs makes a reviewer say "that looks like a real wafer we have seen". The other four look like nothing. Now you have to write it up, and the write-up is the actual deliverable. ## Why this is a lead, not a finding Three distinct things could explain the observation, and the report must distinguish them. 1. **You recovered the class prototype and it happens to look plausible.** This is the default outcome and the least interesting one. A prototype assembled from what class members share will naturally resemble a member; that resemblance is the attack working exactly as expected and is not a disclosure of any record. 2. **The reviewer expected a match.** Judging visual similarity by eye, knowing what you are hoping to see, is the weakest evidence in the building. It is not a reason to distrust the reviewer; it is a reason not to let their impression be the measurement. 3. **You genuinely recovered something specific to one record.** Possible, and the only version worth escalating — and it needs evidence the other two explanations do not produce. ## What makes it a finding **A similarity measure chosen before you look, applied to every run.** Picking the best of five restarts and then describing it is selection on the outcome. Report the distribution across restarts and say the measure was fixed in advance. **Baselines, and the right ones.** The comparison that matters is not "output versus real wafer" in isolation. It is output-versus-specific-record against output-versus-class-average, and against output-versus-a-record-from-a-different-class as a floor. If the output is about as close to the class average as to the specific record, the honest conclusion is that you recovered the average. **The class membership count in the same sentence as the result.** A class built from a handful of records has an average that looks like a member by construction, so a strong resemblance there tells you about the label space, not about extraction. **The reproduction rate as a reported number.** One in five is information: it tells a reader what re-running the assessment will cost and how stable the effect is. Suppressing the four failures is the defect; describing them is the finding's credibility. ## The claim ladder Write the strongest claim your evidence supports and no higher: - "A search against returned class scores produces a plausible-looking input for this class." — supported by one run. - "The recovered input is the model's class prototype." — supported once you compare against the class average. - "The recovered input is closer to one specific training record than to the class average, and this class has twelve members." — this is the escalation, and it needs the measure, the baselines and the count. - "The model leaks a customer's proprietary wafer." — needs everything above plus somebody who can confirm the record, and it is the sentence a reader will check hardest. ## The failure mode this avoids The worst outcome for a security team is not missing a finding; it is shipping the top claim on the bottom claim's evidence. It gets corrected in public by whoever reproduces it, it costs the next report its credibility, and it usually gets the underlying real issue — a thin per-customer class — dismissed along with it.

  • What single baseline would most change your conclusion here?
    Distance from the output to the class average, alongside distance to the specific record. If they are comparable, you recovered a prototype and the finding closes. If the output is markedly closer to one record, you have something worth escalating and a concrete number to escalate with.
  • Should the four restarts that produced nothing appear in the report?
    Yes, and prominently. The reproduction rate is part of the result: it tells the reader how stable the effect is and what re-testing costs. Reporting only the successful restart is selection on the outcome and is the first thing a careful reader will look for.
  • The class has twelve members. How does that change the write-up?
    It has to be stated next to the result. With twelve members the class average is close to any member by construction, so resemblance is expected and is evidence about the label space rather than about extraction. It also points at the real recommendation: the thin class.

saying these in an interview costs you the question

  • Reports the best restart and drops the rest
  • Calls visual resemblance a match without a measure
  • Omits the class membership count from the claim
  • Compares only to a real record, never to the class average
  • Escalates a prototype as a recovered training record

context