How do you price the per-category scores in a moderation API's response at a design review?
answer
- arrive with a cost, not a verdict
- what would a vendor have charged
- name what you cannot claim
- the field exists for a reason
- ten contracts or open self-serve
basics
~20 sReport what the field buys an adversary in money and time: the annotation budget a competitor no longer funds, and how fast a stand-in reaches usable quality. Leave the keep-or-cut call to the product owner.
solid answer
~50 sArrive with a number and a boundary, not a verdict. The number: at published pricing, what a campaign of ordinary-looking calls costs, against what the equivalent annotated corpus would have cost from a vendor with a policy team behind it. That ratio is the finding, and it is denominated in something the owner already funds. The boundary is what earns you the second meeting — you cannot claim the copy will match us, cannot claim any schema change would stop a determined buyer, and cannot claim we would notice. Then hand back the decisions that are genuinely not yours: the field exists because integrators threshold on it, so whether it stays is a product call; whether the subscription price reflects what a buyer is really acquiring is a commercial one; and whether we face this adversary at all differs between ten contracted enterprises and an open self-serve tier.
go deeper
Know that a response field can be discussed as a cost to the business, and that the person raising the concern is not the person who decides the schema.
Be able to turn the mechanism into an estimate: calls needed, published price per call, and the annotation spend a buyer avoids by making those calls instead.
Show that you bound your own finding out loud — no parity, no claim of detection, no claim that a schema change ends it — because the unbounded version is the one that gets discounted.
Own the framing that this is a pricing, terms and channel question as much as an engineering one, and that the write-up should say which buyers it is about rather than averaging them.
## The chair you are sitting in This is not the engineer improving the model and not the person building a guardrail. It is the red-teamer invited to an API design review, where a response schema is being agreed, and asked what each field costs. The deliverable is a **price**, not a prescription — and getting that distinction right is most of what separates a finding that changes a decision from one that gets filed. ## What to bring **Two figures, both in money.** 1. *What the campaign costs the buyer.* Published price per call, multiplied by the calls needed for an annotated corpus of a size that would train a usable stand-in. Every input to this is public; that is the point, and it is worth saying out loud that a competitor can compute it from the documentation without ever contacting us. 2. *What it replaces.* What that same annotated corpus would have cost as a human programme — policy authorship, annotator training, per-judgment billing, adjudication, quality sampling. On a policy task this is normally the larger figure by a wide margin. The ratio between them is the finding. It says, in the owner's own units, what the response body is giving away per dollar of revenue it earns. **A time estimate, roughly.** How long a funded competitor takes from first call to a system good enough to ship. "Weeks, not quarters" changes roadmap conversations in a way that a severity label never does. ## What you must refuse to claim An unbounded claim is a discounted claim, and this area invites three overstatements: - **You cannot claim parity.** The copy is coverage-limited and typically weaker on the ambiguous tail our annotators argued hardest about. Say that before somebody else does. - **You cannot claim a schema change would stop anyone.** Reducing what the body returns changes the buyer's bill; it does not make the campaign impossible, and pretending otherwise sets up a defence that fails in front of an audience. - **You cannot claim detection.** If we do not have a demonstrated way to tell this traffic from a large legitimate integrator's, saying we would notice is an assurance we cannot honour. A reviewer who states these three unprompted is believed on everything else in the memo. ## The decisions that are not yours Three calls belong to other people, and naming their owners is part of the deliverable: - **Product owns the schema.** The per-category scores exist for a reason: integrators threshold per policy, route to their own human review queues, and tune their own precision. Opening with "remove the field" tells the room you did not learn why it is there. - **Commercial owns the price and the terms.** If a buyer's realistic alternative is a large annotation programme, then the subscription price and the contract are levers that already exist and are cheaper to pull than any engineering change. This is the reframing that most often lands: the exposure is a **pricing** problem wearing an engineering costume. - **The business owns the risk appetite.** Some products should absolutely publish rich replies and compete on something other than the labels. ## Say for whom The same schema carries different exposure through different channels, and a write-up that gives one number for all of them is weaker than it needs to be. With ten contracted enterprise buyers there is a named counterparty, an agreement, and a legal surface. With open self-serve sign-up there is an anonymous buyer, no counterparty, and a payment method. Nothing about the response body differs; everything about who can obtain it does. Separate the two in the memo. ## If they want a verdict anyway Sometimes the room wants a single rating. Give the numbers first and let them rate it if their process needs a rating — a compressed severity discards the only content that lets an owner trade this against the other twelve things competing for the same quarter. The judgment being asked of you is not "is this dangerous"; it is "what is it worth, to whom, and what would I be lying about if I said more". That is a principal-level answer, and it is one an interviewer can distinguish from a senior one within about thirty seconds.
- The owner asks for a single severity rating instead. What do you give them?The two figures first — what a campaign costs a buyer, and the annotation spend it replaces — plus the note that exposure differs by sales channel. If their process needs a rating they can apply it to that; a rating handed over on its own discards the only content that lets them weigh this against everything else in the quarter.
- Why does pricing and contract belong in an engineering design review at all?Because the exposure is commercial. If a competitor's alternative is a large annotation programme, then price, terms and channel are levers that already exist and cost nothing to pull, while a schema change costs every honest integrator something. Raising it invites the people who own revenue into a decision that is partly theirs.
- How does an open self-serve tier change the write-up?It changes who the buyer can be and how cheaply they stay anonymous. Contracted enterprises come with a named counterparty and an agreement; open sign-up comes with neither. The response body is identical, so state the exposure per channel rather than averaging them into one misleading number.
- What do you say if asked whether we should just stop returning the scores?That it is a product decision with a real cost to honest integrators, and that it changes the buyer's bill rather than their feasibility. Give the room the price on both sides and decline to make the call — a reviewer who quietly annexes the schema decision does not get invited to the next review.
saying these in an interview costs you the question
- Opens with a recommendation to strip the field
- Quotes exposure in adjectives rather than money and time
- Claims a schema change would stop a determined buyer
- Ignores why integrators need the per-category scores
- Gives one exposure number for enterprise and self-serve alike
- Promises the traffic would be noticed without evidence