In OpenRouter, what do the :nitro and :floor model-slug suffixes do?
answer
- variant suffixes on the slug, same model
- one picks speed, one picks price
- shorthand for a provider sort setting
- they switch off the default balancing
- the other axis is what you pay
basics
~20 sThey are shorthand for a provider-sorting rule on the model slug. Appending :nitro ranks the providers hosting that model by throughput, so the request goes to the fastest; :floor ranks them by price, so it goes to the cheapest.
solid answer
~50 sFor a model served by several upstream providers, OpenRouter normally balances between them rather than always picking one extreme. The slug suffixes let you override that inline without sending a `provider` object: `vendor/model:nitro` sorts candidates by throughput and routes to the fastest, while `vendor/model:floor` sorts by price and routes to the cheapest. They are equivalent to setting `provider.sort` to `throughput` or `price` respectively, and like any explicit sort they replace the default load balancing for that request. The practical consequence is that you are opting into one dimension and accepting the other: `:nitro` can land you on a materially more expensive upstream, `:floor` can land you on a slow or heavily loaded one. Use `:nitro` for interactive, latency-visible paths and `:floor` for bulk offline work, and log which provider served so you can see what the choice actually bought.
code
json · 4 lines{
"model": "vendor-a/model-name:floor",
"messages": [{ "role": "user", "content": "Summarise this row for the nightly report." }]
}go deeper
Recall that the suffix follows a colon on the model slug, that :nitro means fastest provider and :floor means cheapest, and that the model itself is unchanged.
Explain that they are shorthand for a provider sort, that specifying a sort replaces the default load balancing, and what each one gives up on the opposite axis.
Tie the choice to workload class and defend it with data — serving provider, cost per request, and tail latency measured against the unsorted default.
Decide where speed-first routing is worth its premium across a portfolio of services, and keep that decision reviewable as the provider mix and price sheet change.
## What the suffixes are An OpenRouter model slug looks like `vendor/model-name`. A **variant suffix** after a colon modifies how that slug is handled. Two of them concern routing: - **`:nitro`** — rank the providers hosting this model by **throughput** and prefer the fastest. - **`:floor`** — rank them by **price** and prefer the cheapest. They are conveniences, not separate models. Each is equivalent to sending a `provider` object with `sort` set to `throughput` or `price`. The suffix form exists because it travels through anything that accepts a model string — a config file, an environment variable, a third-party client that only exposes `model` and never lets you add custom body fields. ## What they replace Without any sort, OpenRouter **load-balances** across the providers serving a model, weighing price against how reliably each has been performing. That default is deliberately a compromise: it avoids hammering the single cheapest provider until it degrades, and it avoids paying premium rates when you did not ask to. Specifying a sort — by suffix or by `provider.sort` — turns that compromise off. The gateway now ranks candidates on the one dimension you named and works down the ranking. That is the whole trade: you get predictability on the axis you care about and give up the balancing that protected the other axis. ## Choosing between them `:nitro` is for **latency-visible work**: interactive chat, autocomplete, anything where a user is watching tokens appear, and agent loops where many sequential calls compound into wall-clock time. The cost is real — the fastest upstream for a popular model is frequently not the cheapest, and if you route all traffic this way your bill reflects a speed preference you may not have priced. `:floor` is for **throughput-insensitive work**: batch enrichment, offline evaluation, nightly summarisation, backfills. The cost is queueing and tail latency, because the cheapest provider is often the most contended. On an interactive path that shows up as a bad p95 rather than an error, which makes it easy to miss. A useful discipline is to make the choice a property of the *workload*, not of the codebase: the same service can send `:nitro` for its user-facing endpoint and `:floor` for its background job, using the same prompts and the same parser. ## What they do not do They do not change the model. Same slug, same weights, same capabilities — only the upstream differs. They do not guarantee an SLA; they express a ranking preference at selection time, based on observed characteristics, not a contractual latency or a fixed price. They do not fence anything: if you also need to stay inside an approved set of providers, that is `only`/`order` with `allow_fallbacks: false`, and a sort suffix is no substitute. And they do not filter by capability — that is `require_parameters`, and combining a cheap-first sort with structured output requests without it is a good way to discover a provider that quietly ignores your schema. Because they are sorts, they also compose oddly with an explicit `order`: an order is already a sequence, so adding a sort on top is contradictory intent. Pick one mechanism per request. ## Verifying the choice The honest way to evaluate a suffix is measurement, not belief. Log the serving provider and the per-request cost and latency, then compare the suffix against the default over real traffic. Teams frequently find that `:nitro` bought less latency than expected for their prompt shape — because the dominant term was output length, not provider speed — or that `:floor` bought a big cost saving with an acceptable p95. Neither answer is knowable a priori, which is exactly why the suffixes are cheap to try and worth re-checking as the provider mix changes. ## Common mistakes Treating `:nitro` as a different, faster *model*; assuming `:floor` is always the cheapest possible outcome even when a fallback carries you elsewhere; and using either as if it were a compliance or capability control. They are one-dimensional ranking hints, and every other routing concern still needs its own field.
- What is the request-body equivalent of appending :nitro to the slug?Sending the plain slug together with a provider object whose sort is set to throughput. The suffix exists mainly so the preference can travel inside a model string — through config, environment variables, or a client library that only exposes the model field. Behaviourally the two forms express the same ranking, and both replace the default load balancing for that request.
- Why might :floor hurt an interactive endpoint even though every request still succeeds?The cheapest upstream is usually the most contended, so requests queue. That does not show up as errors; it shows up as a stretched p95 and p99 while the median looks fine. On a user-facing path that reads as an app that feels sluggish under load. Reserve price-first routing for batch and offline work, and measure tail latency rather than averages when you evaluate it.
- Do these suffixes offer any guarantee about which provider serves the request?No. They express a ranking preference evaluated at selection time from observed characteristics, not a contract. If the top-ranked provider cannot serve, routing continues down the ranking, and unless you also set allow_fallbacks to false it can continue past your preferences entirely. For guarantees about where traffic lands you need the explicit provider allowlist fields, not a sort.
saying these in an interview costs you the question
- Thinks :nitro is a separate, faster model rather than a routing hint
- Assumes :floor guarantees the lowest possible cost regardless of fallback
- Believes the suffixes pin one specific provider
- Uses a sort suffix as if it were a compliance or capability control
- Applies :floor to interactive traffic and only watches average latency