skip to content

You lead a fixed-hours red-team engagement on a product that relies on a hosted content-moderation service. How much of the engagement do you spend probing a layer the customer cannot change, and what should the deliverable about it say?

level: principalimportance: should knowfreq 33%

answer

  1. sort work by fixability
  2. vendor model = bounded slice
  3. hours go to thresholds, categories, failure behaviour
  4. named residual risk with an owner
  5. hand over the corpus, not just the report

basics

~20 s

Spend little on proving a vendor classifier imperfect; the customer cannot fix it. Spend the hours on the decisions they own: which categories they act on, the cut-offs, what text is screened, and what happens on error or timeout. The deliverable should recommend configuration, fallbacks and a residual-risk statement, not a vendor swap.

solid answer

~60 s

Sort candidate work by **fixability**. Anything whose only remedy is "the vendor should train a better model" earns a small, bounded slice — enough to characterise the coverage boundary and support a residual-risk statement, and no more. Everything downstream of the vendor is the customer's, is cheap to change, and is where the hours belong. The exception is procurement. If the engagement's real question is *which* service to adopt, or whether a specific harm class is covered at all, then deep probing of the hosted layer is the deliverable, and it should be scoped and priced that way explicitly rather than smuggled in. The deliverable's shape follows. State what the filter is measured to cover and, plainly, what it is not — the out-of-taxonomy classes are a named residual risk with an owner, not a silent gap. Then the actionable list: category and threshold choices, failure behaviour, what is screened and what is not, monitoring, and a regression corpus the customer runs themselves against a control they cannot version.

go deeper

for a junior

Understands the vendor's classifier is not something the customer can change, so findings about it are less actionable.

for a middle

Directs effort at the configuration the team owns — categories, cut-offs, what is screened — and reports failure behaviour.

for a senior

Bounds the vendor-layer work explicitly, tests fail-open and screening scope, and writes a residual-risk statement for out-of-taxonomy harms.

for a principal

Allocates a fixed budget by fixability, names the procurement exception, hands over a reusable regression corpus, and shapes the recommendation as defence in depth with owners rather than a vendor swap.

**Start from what an engagement is bought to do.** A red team is paid to change something. That makes **fixability**, not severity, the right axis for allocating fixed hours. A finding against a black-box hosted classifier changes nothing the customer controls: they cannot retrain it, cannot inspect its policy, cannot pin its version. It also decays, because the behaviour it describes can move server-side without notice. A report full of such findings reads as diligence and delivers no remediation, and six weeks later half of it will not reproduce. **So bound that work explicitly, at the start.** A stratified corpus, one direct baseline run against the vendor's endpoint, a control set, and a written coverage statement. That is enough to say what the filter is measured to cover and what it demonstrably is not, which is all the residual-risk paragraph needs. Then defend the boundary, because curiosity pulls hard here: generating misses against a hosted filter is the most immediately gratifying activity available on the engagement and the least useful one. **Where the hours actually go.** Everything downstream of the vendor, all of it same-week fixable: - **Cut-off selection**, with the false-positive cost on the customer's real traffic made visible rather than assumed. Tightening a threshold is free in a test corpus and expensive in production. - **Category handling** — which categories block, which merely log, and whether anyone reads the log. - **Screening scope** — is the model's output screened as well as the user's input; is the text sent for screening the text the user actually sees; are long inputs truncated before they are screened. - **Failure behaviour** — what happens on a throttle, an error or a timeout. Fail-open promotes the vendor's availability into a security control, and it is invisible to any direct-endpoint run. - **What is behind it** — whether anything downstream catches what the filter misses, and whether a human is ever on the escalation path. **The residual-risk paragraph, written down.** A hosted filter has a fixed taxonomy. Harms outside it are not measured by this control, and no cut-off change reaches them. That sentence must appear in the deliverable in language a non-specialist can carry into a risk meeting, with a **named owner**. Teams end up treating a purchased filter as their entire safety story mostly because nobody ever wrote that sentence down — the coverage boundary was implicit, and implicit boundaries do not survive the trip from the report to the risk register. **Where the numbers mislead the reader of your report.** A miss rate against the vendor's endpoint reads as the product's exposure and it is not — it excludes every decision the deployment made, some of which are worse and some far better. A high finding *count* against the vendor layer looks like thoroughness, but count is a measure of where you spent time, not of where the risk is; balancing findings evenly across layers so the report looks symmetrical optimises the document's appearance at the expense of the customer's remediation queue. And a coverage percentage means nothing without its denominator: probes sent, harm families represented, or vendor categories exercised are three different numbers from the same run. **The one exception worth naming.** If the engagement's real question is **procurement** — which service to adopt, or whether a specific harm class is covered at all — then deep probing of the hosted layer *is* the deliverable. Scope and price it that way explicitly rather than letting it arrive as scope creep, and expect the corpus, not the report, to be the artefact the customer keeps. **Shape the recommendation as depth, not replacement.** The honest conclusion is rarely "change vendors" and usually "stop asking one purchased layer to be the whole control": a second, differently-failing layer for the harm classes that matter most, human review retained on the escalation path, and monitoring on the join. Advise on the shape of that stack from what you measured; the detailed engineering of the customer's own moderation controls belongs to their team, and your distinctive value is having measured what the bought layer actually does. **Finally, hand over the harness.** The corpus, the runner, the control set and the storage format — not just a PDF. A regression corpus the customer runs continuously converts an ageing one-off finding into a standing control, and it is the only mechanism that catches silent drift in a layer nobody can version.

  • When is deep probing of the hosted layer the right place for the budget?
    When the question is procurement or coverage evidence — choosing between services, or demonstrating whether a specific harm class is covered at all. Then it is the deliverable and should be scoped as such.
  • What single sentence do you insist appears in the report?
    That harms outside the service's fixed category set are not measured by this control, with a named owner for that residual risk.
  • How do you keep the engagement's value from expiring when the service changes?
    Hand over the corpus and harness so the customer runs it continuously, and word findings as behaviour classes with observation windows rather than single instances.

saying these in an interview costs you the question

  • Spending most of the engagement generating misses against a black-box vendor model.
  • A deliverable whose main recommendation is "switch vendors" with no configuration findings.
  • Never writing down the out-of-taxonomy residual risk, leaving the filter to read as the whole safety story.
  • Ignoring failure and timeout behaviour because it is not a content finding.
  • Delivering a static report and no reusable corpus for a control that can drift silently.

context