skip to content

When does renting GPUs to self-host an open model beat paying per token?

level: principalimportance: should knowfreq 44%

answer

  1. Fixed cost versus variable cost
  2. the floor is replicas, not one GPU
  3. utilization decides the bet
  4. some reasons are not about price
  5. re-run it when either side moves

basics

~20 s

When sustained volume is high enough to keep GPUs genuinely busy, and when something other than price — data control, a custom checkpoint, predictable latency — is also on the table. Below that, per-token pricing wins because you are not paying for idle hardware.

solid answer

~60 s

Frame it as a fixed-versus-variable cost decision. Self-hosting has a **floor**: enough GPUs to hold the model, times enough replicas for availability, billed continuously. Per-token pricing has no floor and scales linearly. So there is a crossover volume, and it is usually higher than people expect, because the floor is not one GPU — it is two or three replicas of however many GPUs the model needs. Below the crossover, per-token wins outright. Above it, self-hosting wins only if you can hold **utilization** up; spiky interactive traffic with a 10x daily peak wastes most of what you rent, while steady or batch traffic does not. Then weigh the costs that never appear in a spreadsheet: on-call, upgrades, capacity planning, quota risk. And weigh the reasons that override cost entirely — data residency, a fine-tuned checkpoint nobody else hosts, guaranteed capacity, or a per-token price that a vendor can change. Prove it with a measured cost per million tokens against the current price sheet, and re-run it when either side moves.

go deeper

for a junior

Know that rented GPUs bill continuously while per-token APIs bill only for use, so low or bursty volume is almost always cheaper on an API.

for a middle

Build the comparison: monthly GPU floor including replicas versus volume times per-token price, using a measured throughput figure rather than a benchmark peak.

for a senior

Bring utilization and operational cost into it, and name the traffic shapes that break the self-hosting case. Show the load test the throughput number came from.

for a principal

Own the decision and its shelf life: set the volume and utilization thresholds, weigh residency, capacity guarantees and vendor pricing risk, propose a hybrid split, and name the events that trigger re-evaluation.

## The shape of the decision This is fixed cost versus variable cost, dressed in GPUs. - **Per-token APIs** are pure variable cost. Zero traffic, zero bill. Your unit cost is constant at any volume, and your capacity is somebody else's problem. - **Self-hosting** is mostly fixed cost. You rent whole GPUs by the hour; the bill is the same at 3 a.m. as at peak. Your marginal token is nearly free once the fleet exists, and your unit cost falls as volume rises. Two cost curves, one crossing point. The engineering content of the answer is where that crossing sits and how confident you are in the inputs. ## The floor is higher than one GPU A serious self-hosted endpoint is never one card. Count: - **Model footprint.** However many GPUs the checkpoint needs to hold weights plus workable KV cache. - **Replicas for availability.** At least two, so a node failure or a rolling upgrade is not an outage. - **Headroom for peak.** Weights take minutes to load, so you cannot scale from cold in response to a spike; some capacity is paid for while idle by design. - **Environments.** A staging replica for validating checkpoint and engine upgrades. Multiply GPUs by hourly rate by 730 hours and you get a monthly floor that exists whether you serve a million tokens or a billion. That floor, divided by your realistic monthly volume, is your true unit cost — and it is the number to compare against the price sheet. ## Utilization is the whole argument The self-hosting case is entirely a bet that you can keep the fleet busy. Traffic shape decides whether you win it: - **Steady, high-volume, batch-tolerant** work — bulk classification, embedding, offline scoring, overnight enrichment — can saturate GPUs and drives cost per token down dramatically. This is the strongest case for self-hosting. - **Interactive traffic with a 10x daily peak** is the weakest. You size for peak and pay for the trough, so effective utilization lands at 20-30% and your unit cost is three to five times the saturated figure you put in the business case. A useful test: if you cannot describe a workload that will keep the cluster busy at night, you are probably buying idle GPU hours. ## The costs that are not on the invoice - **People.** Someone owns capacity planning, engine upgrades, checkpoint validation, OOM incidents and the pager. That is a meaningful fraction of an engineer, indefinitely. - **Quality drift.** Hosted frontier models improve without you doing anything; your self-hosted checkpoint improves when you do the work of upgrading and revalidating it. - **Quota and supply risk.** The instance type you priced may not be available in your region when you need to scale. - **Time to first token of the *project*.** Weeks of platform work before the first production request, against an API key that works this afternoon. ## The reasons that beat cost Sometimes the arithmetic is not the decision: - **Data control.** Regulatory or contractual requirements that inference never leaves your boundary. - **A checkpoint nobody hosts.** Your own fine-tune, or many per-tenant adapters, is a capability argument rather than a price argument. - **Predictable capacity and latency.** No shared-tenancy variance, no rate limits imposed from outside, no surprise deprecation of the model version you built on. - **Pricing risk.** A vendor can change a price sheet or retire a model; owning the stack converts that exposure into hardware you control. Conversely, the strongest case *against* self-hosting is uncertainty. Early products whose volume, model choice and prompt shape are all still moving should stay variable-cost until at least one of those stabilizes. ## How to answer this as a lead Do not give a universal threshold. Give a method: measure cost per million tokens on a representative load test, build the monthly floor from replica count and hourly rate, project volume with an honest utilization assumption, compare against the current price sheet at your input/output token mix, and state the non-cost factors that could override the result either way. Then say when you would revisit it — a price cut, a new model generation, or a volume milestone — because this decision has a shelf life measured in months, not years. A hybrid answer is often the right one: batch and high-volume paths self-hosted, low-volume and frontier-quality paths on an API.

  • What traffic shape makes self-hosting most attractive, and which makes it worst?
    Steady or batch-tolerant volume is best: it saturates the GPUs you already pay for, driving cost per token toward the marginal rate. Spiky interactive traffic is worst — you must size for peak, weights load too slowly to scale from cold, and you pay full price through the trough. A 10x peak-to-trough ratio can leave effective utilization near 20%, tripling or quadrupling your real unit cost.
  • Someone proposes scale-to-zero to fix the idle-cost problem. What do you tell them?
    That weights are tens of gigabytes and a cold replica needs minutes to pull, load and warm up, so scale-to-zero converts an idle bill into a first-request latency disaster. It is viable for batch or internal tooling that tolerates a multi-minute wait, and not for interactive traffic. The usual compromise is a warm floor of one replica with elastic capacity above it.
  • Which non-financial factors would make you self-host even below the break-even volume?
    Data residency or contractual requirements that inference stay inside your boundary; a fine-tuned or per-tenant checkpoint no vendor hosts; a need for guaranteed capacity and stable latency without externally imposed rate limits; and protection against a model version being deprecated or repriced underneath a product you have already built.
  • How would you structure a hybrid so the decision is not all-or-nothing?
    Route by path. Send high-volume, latency-tolerant or privacy-constrained work to the self-hosted open model where the fleet is already paid for, and keep low-volume or frontier-quality paths on a per-token API. Keep one request interface in front of both so switching a path is configuration. That preserves optionality while the volume and quality assumptions are still moving.

saying these in an interview costs you the question

  • Comparing one GPU's hourly price to a per-token rate
  • Assuming 100% utilization in the business case
  • Ignoring the engineering and on-call cost of a fleet
  • Treating scale-to-zero as a solved problem for large weights
  • Deciding on price alone when data residency is the real driver

context