How do you serve one LLM for interactive chat and an overnight 900K-listing scoring job?
answer
- Opposite ends of the same curve
- Nobody watches the batch job
- Different metric, different deployment
- Idle peak capacity is the cheap lever
- Enforced caps, not good intentions
basics
~20 sTreat them as two deployments of the same model tuned to opposite ends of the latency-throughput frontier: the chat path optimised for tail first-token latency at modest concurrency, the scoring job optimised for tokens per accelerator-hour at maximum concurrency, with the bulk work scheduled into off-peak capacity.
solid answer
~50 sThese are the two extremes of one curve. Interactive chat is judged on p95 first-token latency and per-token rate, so it runs at modest concurrency with headroom for bursts and pays for idle capacity. The 900K-listing scoring job has no first-token requirement at all — nobody watches it — so its only real metric is cost per unit of work, which means maximum concurrency, tight output caps, retries, and running when interactive demand is low. Serving both from one shared configuration forces a compromise that is wrong for both: the bulk queue inflates chat's first-token tail, and the chat objective caps the concurrency the bulk job would happily use. The usual answer is separate pools, or one pool with strict admission control and a hard cap on bulk concurrency. Whether isolation is worth the idle capacity is a genuine judgment call at small scale.
go deeper
Know that a background job nobody is waiting on should be tuned for total cost and completion, while a live chat feature is tuned for how quickly the first words appear.
Explain that the two workloads sit at opposite ends of the throughput-versus-latency curve and that a shared configuration must compromise, degrading the interactive tail or wasting bulk capacity.
Show the operational design: separate deployments with separate objectives, bulk work scheduled into off-peak troughs, enforced concurrency caps, cost estimated and outputs sampled before committing the full run.
Own the trade honestly — isolation buys a protected tail at the price of idle headroom, sharing is defensible at small scale with enforced admission control, and the decision turns on scale, owned versus rented capacity, and what the interactive tail is worth.
## Two workloads, one model, opposite objectives A consumer chat surface and an overnight job that scores 900,000 product listings can run the identical model and still have almost nothing in common operationally. The chat path is judged on **tail first-token latency** and **sustained per-token rate**. Both are user-perception metrics, both must hold during arrival spikes, and both are violated by queueing. Meeting them requires deliberately leaving capacity idle, because a pool sized to average demand will queue during peaks and blow the tail. The scoring job is judged on **tokens per accelerator-hour** and, equivalently, **cost per listing**. No human waits on any individual response, so first-token latency is meaningless and even a multi-minute queue wait per item costs nothing. What matters is finishing within the window at the lowest cost, which argues for the highest concurrency the memory will support, sitting far past the knee that would be unacceptable for chat. ## Why one shared configuration is the wrong answer Run them together and each pays the other's bill. A large backlog of bulk items in the queue puts every interactive request behind a wall of work, so p99 first-token latency for chat degrades exactly when the batch job is most productive. Conversely, holding the shared pool to the chat objective caps concurrency far below what the bulk job could exploit, so the batch runs slower and costs more than necessary. There is no single setting that is right for both; that is precisely what "a frontier" means. The standard structure is therefore two deployments of the same weights, each tuned to its own end, with independent objectives and independent dashboards. The chat deployment: modest concurrency, burst headroom, alerting on p95 and p99 first-token latency and on token rate. The bulk deployment: concurrency pushed to the memory limit, alerting on throughput and completion-by-deadline, with no latency objective at all. ## Scheduling as a cost lever Interactive demand is diurnal; the fleet you size for peak chat is largely idle at 03:00. Running the scoring job in that trough is the single cheapest optimisation available, because it converts already-paid-for capacity into work rather than buying new capacity. This is why the job is described as overnight in the first place. Where capacity is elastic rather than owned, the same logic becomes a scheduling decision against price rather than against idle hardware. ## Sizing and budgeting the bulk job Estimate before you run. Total tokens is roughly 900,000 x (input tokens per listing + expected output tokens). Divide by the measured tokens-per-accelerator-hour of the bulk configuration to get accelerator-hours, then add margin for failures and re-runs — at this scale a small percentage of retries is thousands of items. Trim the prompt: a 200-token reduction in a shared prefix is 180 million tokens across the run, which is a real line item. And cap output length aggressively, because output tokens are both the slowest to produce and the ones that multiply straight into total time. Checkpoint progress so an interrupted run resumes rather than restarts, and sample-evaluate the first few thousand outputs before committing the remaining 895,000 — the expensive failure at this scale is discovering after eight hours that the prompt was subtly wrong. ## When sharing is defensible Isolation is not free. Two pools means two sets of idle headroom and two sets of operational surface. At small volume — a bulk job that is a few percent of daily tokens, or an organisation with a handful of accelerators — a single pool with strict admission control can be the right call: cap bulk concurrency to a fixed slice, give interactive traffic strict priority in admission, and accept that the bulk job runs slower. What makes this work is enforcement, not intent; "we'll run the batch when it's quiet" without an enforced cap reliably becomes an incident. This is genuinely contested ground and depends on scale, on whether capacity is owned or rented, and on how much the interactive tail is worth. What is not contested is that the two workloads have different objectives, and that a design which does not name them separately is a design that will fail one of them silently. ## What an interviewer is listening for That you separate the objectives before proposing any mechanism; that you name cost per unit of work as the bulk metric rather than reaching for latency; that you understand idle interactive capacity as a resource the batch can consume; that you size and sample before committing 900,000 items; and that you can argue honestly for sharing at small scale instead of reciting isolation as dogma.
- When is a single shared pool actually the right call?At small scale, where bulk traffic is a few percent of daily tokens or the fleet is a handful of accelerators. Two pools means two sets of idle headroom and twice the operational surface, which can cost more than the contention it prevents. It only works with an enforced cap on bulk concurrency and strict admission priority for interactive traffic.
- How do you estimate the cost of the scoring job before running it?Multiply 900,000 by input plus expected output tokens per listing to get total tokens, divide by the measured tokens per accelerator-hour of the bulk configuration, and add margin for retries and re-runs. Then look for prompt savings: trimming a shared prefix by 200 tokens removes 180 million tokens from the run, which is a material line item.
- What is the risk of running the bulk job on the interactive endpoint at low priority?Priority in admission does not undo work already in flight, so long-running bulk sequences still occupy memory and share generation steps with interactive traffic. The result is a degraded first-token tail during exactly the periods the batch is most productive. Without an enforced concurrency cap and a way to shed bulk work, low priority is a label rather than a control.
saying these in an interview costs you the question
- Applies the same latency SLO to interactive and offline traffic
- Runs the bulk job on the interactive pool during peak hours
- Measures the batch job in requests per second instead of cost per item
- Assumes one tuning can serve both ends of the frontier
- Commits all 900,000 items before sampling the first few thousand outputs