skip to content

Your engagement covers a chat product whose backend calls a hosted content-moderation service. Do you send your probe corpus straight to the vendor's moderation endpoint, or through the product? What does each choice actually measure?

level: seniorimportance: must knowfreq 52%

answer

  1. vendor endpoint = filter only
  2. through the app = the deployment
  3. unattributable without backend logs
  4. correlation marker joins the two runs
  5. fail-open on timeout is a finding

basics

~20 s

Direct calls measure the vendor's classifier alone. Going through the product measures the deployment: which categories the team acts on, the cut-offs they chose, what text the service actually receives after the app transforms it, and what happens when the call fails. Only the second yields findings the customer can fix, so do both and report them separately.

solid answer

~50 s

Direct probing is cheap, attributable and useless on its own: the customer cannot change the vendor's model, so every finding lands as "switch vendors" or nothing. Its real job is **calibration** — what the filter does with a payload in isolation is the baseline you need to read the other run. Probing through the product measures what is deployed. Between your text and the service sit normalisation, truncation, whatever the app chose to send, and the app's behaviour when the call errors or times out. Those are the customer's decisions and every one is fixable, so this is where the report's value is. The cost is attribution. A payload that gets through the product could have been scored clean by the service, scored above threshold and ignored by the app, or never sent at all. Resolve that by correlating with the backend's own logs of the moderation call — and if those logs do not exist, that gap is itself a finding.

go deeper

for a junior

Knows you can call the moderation service directly and can also test through the product, and that the results are not the same thing.

for a middle

Explains that the app's transformation, cut-off and category selection sit between the user text and the service, so end-to-end results include decisions the vendor never made.

for a senior

Runs both, plants correlation markers, joins to backend logs to attribute each miss to a layer, and tests failure and timeout behaviour explicitly.

for a principal

Allocates the metered budget between the two runs by fixability, and sets the reporting shape so a filter-level number is never read as the product's exposure.

## Two probe points ask two different questions "What does this filter do with input X?" and "what does this product do when a user sends X?" are not the same question, and a report that mixes their results is one the customer cannot act on. Both runs are legitimate; they buy different things, cost differently, and their numbers must never share a denominator. ## What the direct run buys Calling the vendor's moderation endpoint yourself gives an isolated, reproducible reading per payload, at a known price per call, with no application in the way. Its jobs are narrow and real: - build and sanity-check the corpus; - establish the **control set** of known positives that validates every later session; - and hold a **reference value** for anything the end-to-end run surprises you with. What it does *not* buy is a fixable finding. The customer cannot retrain a black-box classifier, so a direct-run miss lands as "switch vendors" or as nothing — and it ages fast, because the thing it describes can change server-side without notice. ## What the end-to-end run buys Everything the deployment decided, and every one of those decisions is the customer's to change this week. Sitting between the text a user typed and the text the service scores are: - normalisation and encoding; - truncation of long inputs; - the choice of what to send at all (the latest turn, the whole thread, the user's message only); - whether the model's *output* is screened as well as the input; - which categories the app treats as blocking versus merely logged; - and the cut-off, which the vendor's numbers only advise. Then the **error path**: what the product does when the moderation call throttles, errors or times out. If it proceeds unscreened — **fail-open** — then vendor availability has quietly been promoted into a security control, and that is a finding no direct-endpoint run can ever produce, because the behaviour under test belongs to the application, not the classifier. ## The cost of the end-to-end run is attribution A payload that gets through the product could have been: - (a) scored clean by the service, - (b) scored above the line and ignored by the app, - (c) sent in a mangled or truncated form that no longer carried the harm, - or (d) never sent to the service at all. Four different fixes, in four different places. End-to-end traffic alone cannot tell them apart, and a finding that names the wrong layer wastes the remediation as surely as no finding at all. ## How you resolve it Ask for the backend's own logged request and response for the moderation call, and join on a distinctive, benign **correlation marker** you plant in each probe so the two records can be matched without guessing. With that join every end-to-end miss resolves cleanly into a filter miss, a threshold decision, an app-logic decision, or a call that never happened. Without it, you are guessing — and the honest write-up says "layer unattributed" rather than naming one. If those logs do not exist, stop and write *that* up: a security control whose invocations are not logged cannot be operated, tuned or incident-reviewed, and this is usually a higher-severity finding than whatever payload you were chasing. ## Sequencing under a metered budget 1. **Direct first**, small and stratified, to prune payload families that are uninteresting even in isolation and to fix the baseline. 2. **Then spend the larger share end-to-end** on the survivors, where the actionable findings live. The end-to-end calls are also the slower ones — a product round trip includes the model, not just the filter — so wall-clock, not just billing, argues for pruning first. ## How the misleading number appears here A direct-run miss rate reads, to any executive skimming the report, as the product's exposure. It is not: the app may block on a lower cut-off, screen both directions, and catch downstream what the filter missed — or it may fail open and be far worse. Report the two runs as **separate sections**, each with its own named denominator, and never compute a combined figure. If the customer wants one number, the only defensible one comes from the end-to-end run, and it is a rate against your corpus during your window, not a property of the product.

  • The customer cannot give you backend logs for the moderation call. How do you proceed?
    Write that up as a finding first, then compensate: plant distinctive correlation markers in probes, use direct-run baselines as the reference, and mark every end-to-end miss as layer-unattributed rather than guessing.
  • Which single end-to-end behaviour would you test first, and why?
    What the product does when the moderation call fails or times out. Fail-open turns vendor availability into a security control, and it is invisible to any direct-endpoint run.
  • You see a payload flagged in the direct run but accepted by the product. What are the candidate explanations?
    The app transformed or truncated the text before sending, sent only part of the exchange, used a higher cut-off, treated that category as log-only, or did not call the service on that path.

saying these in an interview costs you the question

  • Delivering only a direct-endpoint run and calling it an assessment of the product.
  • Reporting an end-to-end miss as a vendor classifier defect with no log evidence that the call was even made.
  • Never testing what the product does when the moderation call errors or times out.
  • Merging the two runs into one miss rate with a single unnamed denominator.
  • Assuming the text the app sends to the service is the text the user typed.

context