skip to content

You lead a red team whose targets are mostly third-party hosted chat endpoints, with one or two open-weights models deployed in house. A senior engineer proposes buying GPU capacity and standing up a weight-handling process so the team can run gradient-guided prompt searches. How do you decide, and what recurring costs sit behind the hardware line item?

level: principalimportance: should knowfreq 30%

answer

  1. coverage set by weight-holding share
  2. custody, deletion, licence review
  3. compute never amortises
  4. fidelity maintenance is the hidden hour sink
  5. rent, or run in the client environment

basics

~20 s

Decide from the target mix: the capability only applies where you hold weights, which here is one or two systems. Beyond hardware you take on weight custody and deletion duties, licence review, per-target GPU hours that never amortise, re-runs after every checkpoint update, and staff time to keep the rig matching each serving stack.

solid answer

~50 s

Frame it as capability-versus-coverage, not as a hardware purchase. **Coverage.** Gradient search applies to the fraction of your scope where weights are obtainable. If that is two systems out of twenty, the investment buys deeper results on a tenth of the portfolio while the other nine tenths still need query-driven methods that the team must be good at regardless. **Recurring costs.** GPU hours per target that do not amortise; a re-run whenever a checkpoint moves; secure storage, access control and documented deletion for client weights; licence and contract review for holding them; and engineer time keeping each local rig faithful to the corresponding serving stack, which is where the real hours go. **Alternatives.** Rent capacity per engagement instead of owning it; or ask clients to run the job inside their own environment, which removes custody entirely. **Decide yes** when weight-holding engagements are a repeating line of business, or when in-house models are consequential enough to justify the deepest testing available. Otherwise buy it by the hour.

go deeper

for a junior

Not expected to lead this; should at least know the capability requires local weights and therefore does not apply to hosted targets.

for a middle

Can list the direct costs — hardware, GPU hours per target, re-runs after updates — and note that most of the portfolio is unaffected.

for a senior

Adds custody, licensing and rig-fidelity maintenance, and proposes renting or running in the client environment as alternatives.

for a principal

Decides from the portfolio share, sets a scope and utilisation review, caps hours per engagement, and controls how the resulting depth asymmetry is described in reports.

**Start from the portfolio, not from the technique.** The ceiling on this capability is set by how often the team is handed loadable weights, and nothing about the hardware raises that ceiling. Count the last year of engagements and ask how many delivered a checkpoint you could load and were licensed to run. That fraction is the maximum share of work the investment can ever touch. In the situation described — mostly third-party hosted endpoints, one or two in-house open-weights models — the fraction is small, and that alone argues for renting capacity per engagement rather than owning it. **Costs that are not the accelerator.** - *Weight custody.* A client's model weights are among the most sensitive artefacts a consultancy will ever hold: they embody the training investment and, with a fine-tune, sometimes the training data's fingerprints. Holding them means encrypted storage, controlled and logged access, a documented retention window, evidenced deletion, and an answer to "what happens if that machine is compromised". Some clients refuse on this basis alone, which makes custody a sales constraint as well as a line item. - *Legal review.* Open-weights licences carry use restrictions; client weights carry contract terms. Someone reviews both, per engagement, before a byte lands on disk. - *Non-amortising compute.* Every target is a fresh search, because the gradient belongs to one parameter tensor. Every checkpoint update invalidates the previous result. A client on a quarterly release cadence turns a one-off purchase into a standing re-run obligation. - *Rig-fidelity maintenance.* This is the hidden hour sink and the one most often absent from proposals. Keeping a local rig faithful to each serving stack — quantised build, chat template, system preamble, filters — is recurring engineer time per target, and it is the work that decides whether the run's output means anything. - *Utilisation and ageing.* Owned accelerators idle between engagements are pure loss, and they age against model sizes that keep growing; a card that comfortably held last year's targets shards this year's. **Where the number misleads.** Two figures typically decide this review and both can be constructed to flatter the rig. The first is *cost per confirmed finding computed over weight-holding engagements only*: it drops the idle months and the fixed cost carried while the rest of the portfolio ran query-only, so the rig looks efficient on a slice chosen after the fact. Compute it over the whole portfolio and the same rig usually looks marginal. The second is *depth asymmetry in the reporting*: two in-house models receive the deepest testing available while eighteen hosted targets receive query-driven testing, and unless every report says so plainly, a clean result on a hosted endpoint gets read as though it had survived a search that was never run against it. That reading is the real institutional risk of owning the capability, and it is a writing discipline, not a hardware problem. **Alternatives to price against.** | Option | What it removes | What it introduces | |---|---|---| | Rent capacity per engagement | Fixed cost, idle time, ageing hardware | Still full custody obligations; procurement latency per job | | Run the job inside the client's environment | Custody, storage and deletion entirely | Scheduling dependence, restricted debugging, slower iteration | | Skip the class; invest in query-driven capability | Custody and compute both | No white-box depth on the in-house models | **How to say yes well.** If the answer is yes, scope it: name the specific high-consequence in-house models it exists for, publish the fidelity standard the local rig must meet before any hours are spent, cap hours per engagement with an agreed stopping rule, require calibration before full allocation, and review utilisation across the *whole* portfolio after two quarters rather than across the engagements the rig happened to serve. And write the standing rule that every report states which targets received this depth of testing and which did not.

  • What single metric would you review after two quarters to judge the decision?
    Utilisation against weight-holding engagements — GPU hours spent on searches that produced confirmed, replayed findings, over hours available. Low utilisation means it should have been rented.
  • A client is willing to run your search job in their own environment. What changes in your cost case?
    Custody, storage and deletion obligations disappear, and their hardware absorbs the compute. You trade that for scheduling dependence, restricted debugging access, and a slower iteration loop.

saying these in an interview costs you the question

  • Justifying the purchase on the technique's sophistication rather than on the share of scope it can touch.
  • Omitting weight custody, deletion and licence obligations from the cost case.
  • Assuming one purchase covers all targets, with no allowance for re-runs after checkpoint updates.
  • No comparison against renting capacity or running the job inside the client's environment.
  • Letting a deep result on one in-house model set the reporting tone for hosted targets that were never searched this way.

context