skip to content

What determines the query bill for cloning a paid speech-to-text API?

level: middleimportance: should knowfreq 46%

answer

  1. two factors multiplied, not one number
  2. fidelity is a curve with an expensive tail
  3. they only buy the slice they sell
  4. both columns priced per audio hour
  5. the copy cannot beat its teacher

basics

~20 s

Two numbers multiplied: the price per call and the number of calls the wanted fidelity needs. Matching a target across its whole input range costs far more than reaching sellable accuracy on the narrow slice the adversary intends to serve.

solid answer

~50 s

The bill is price per call times calls needed, and the second term is set almost entirely by the fidelity target. Fidelity is a curve, not a threshold: close agreement with the endpoint across everything it handles gets steeply more expensive out in the tail, while adequate accuracy on one language, one audio channel and one domain is cheaper by orders of magnitude. An adversary reselling transcription buys the cheapest point on that curve that clears their business need, so they never pay to match the vendor. Two other terms move the bill: how much each reply reveals, which sets how much fidelity a unit of spend buys, and how well the queried inputs match the traffic they plan to serve, since coverage bought off-distribution is wasted money. The ceiling is the teacher: a copy trained on the endpoint's answers inherits its errors as ground truth.

go deeper

for a junior

Know that copying costs money per call, and that the total is a call count times a price. Be able to say that a cheap copy and a faithful copy are not the same purchase.

for a middle

Explain why agreement with the target gets steeply more expensive out in the tail, and why an adversary buys only the slice of behaviour they intend to sell rather than matching the whole endpoint.

for a senior

Put units on an estimate. Tie a call count to a named fidelity target and an input distribution, price it, and state the ceilings the copier accepts: the teacher's errors, off-distribution gaps and staleness.

for a principal

Own the consequence of the fidelity curve for pricing and product: the part of your quality a copier declines to pay for is your tail, so ask whether the durable asset is the tail and the refresh rather than the snapshot.

## The bill has two factors, and only one is interesting **Query spend = price per call x number of calls.** The price is published by the victim. The call count is the adversary's decision, and it is set by how good the copy has to be. Everything worth saying about clone economics lives in that second factor. ## Fidelity is a curve, not a number "Fidelity" here means how closely the copy agrees with the endpoint it was fitted to. It is not one target, and treating it as one is the usual mistake. Agreement rises fast for the first tranche of purchased labels — the common cases, the everyday accents and clean recordings — and then flattens. Squeezing out the last few points of agreement means buying coverage of the tail: rare vocabulary, difficult acoustics, the unusual speakers whose handling is precisely what the vendor's five years of purchased transcription paid for. So the spend-versus-fidelity relationship has a cheap early stretch and an expensive tail, and the adversary is free to stop wherever their business stops caring. A reseller who intends to serve one domain — say, one language over a single call-centre audio channel — does not need the tail at all. They stop at the point where their own customers stop complaining. That is why a clone can be viable at a small fraction of what the vendor believes their capability is worth: the copier is not buying the vendor's product, they are buying the part of it that they can sell. ## Coverage has to match the traffic they intend to serve The copy is only dependable where supervision was purchased. Queries concentrated on the distribution they will actually serve buy far more usable fidelity per unit of spend than queries spread evenly over everything the endpoint can do. This is why the input corpus the adversary already holds matters as much as their budget: raw audio in the target language is cheap, so a copier who already has hours of representative recordings is holding most of the non-monetary cost of the attack before they start. ## What each reply is worth The other lever on the exchange rate between money and fidelity is how much information comes back per call. A richer reply teaches more per unit of spend than a bare answer, so the same budget lands at a different point on the curve depending on what the endpoint chooses to return. This is a term in the ledger, not a defence discussion; what controls actually bind a determined copier is its own subject. ## Denominating both columns in the same unit A transcription service sold per audio hour is the cleanest possible instance of this arithmetic, because the honest alternative is priced in the same unit. The comparison is literally: - machine transcripts: vendor's price per audio hour x hours needed - human transcripts: going transcription rate per audio hour x hours needed, plus curation, plus one training run Skilled transcription of a low-resource language is expensive per hour and machine inference is cheap per hour, so the ratio between the columns can be very large. And crucially, the number of *hours* needed is roughly comparable on both sides, so it cancels out of the comparison and the price ratio dominates. ## The ceilings on the steal route Three things the copier accepts by choosing this route: 1. **They cannot exceed the teacher.** Labels are the endpoint's outputs, so the endpoint's errors become the copy's ground truth. The copy converges toward the target's behaviour, mistakes included. 2. **They degrade off the queried distribution.** Anything they did not buy coverage for is a gap, and it is a gap they may not know they have. 3. **They inherit staleness.** The copy is a snapshot. A vendor whose corpus keeps growing keeps moving; a copy taken once does not. For a reseller undercutting on price these are acceptable. For anyone who needs to *lead* the market rather than trail it, they are not, and that is a real limit on what the clone route can buy. ## What a good answer sounds like Do not quote a query count. Quote a query count **attached to a fidelity target and an input distribution**, note that the price per call turns it into money, and say which point on the curve the adversary's business actually needs. A number without those attachments is not an estimate of anything.

  • Why does a copy trained on an endpoint's replies rarely beat that endpoint?
    Because its supervision is the endpoint's output, so the endpoint's errors arrive labelled as truth. The copy converges toward the target's behaviour rather than toward the task, and it degrades wherever coverage was not purchased. A reseller undercutting on price accepts that ceiling; anyone who needs to lead the market cannot.
  • Does a low price per call always make an endpoint attractive to clone?
    No. Price moves one column only. If a good labelled corpus for the same task is already public, the honest-build column is also small and nobody bothers regardless of price. Attractiveness is the gap between the two columns, not the level of either one.
  • How does the input corpus the adversary already holds change their bill?
    It sets how efficiently the spend converts to usable fidelity. Someone holding hours of representative audio in the target language buys coverage exactly where they will sell it; someone querying with unrepresentative inputs pays for agreement they cannot use. Cheap unlabelled data is most of the non-monetary cost of the route.

saying these in an interview costs you the question

  • Quotes a query count with no fidelity target attached
  • Assumes the copy must match the target everywhere
  • Leaves the honest labelling alternative out of the comparison
  • Believes a copy can be more accurate than what it copied
  • Ignores which input distribution the coverage was bought over

context