skip to content

What is the difference between an open and a closed workload model in a load test?

level: middleimportance: must knowfreq 54%

answer

  1. What decides when the next request goes
  2. Rate as input versus population as input
  3. One model throttles itself, one does not
  4. Waiting users cannot also be arriving users
  5. Bounded seats versus an unbounded audience

basics

~20 s

In an open model the generator sends requests at a chosen arrival rate regardless of how the system responds. In a closed model a fixed population of simulated users each waits for a response, then thinks, then sends again, so slowness throttles the load.

solid answer

~50 s

An **open** workload is driven by an arrival rate: the generator issues requests on a schedule — often randomised around a target rate — and keeps doing so whether or not earlier requests have returned. Concurrency is an outcome, and if the system slows, in-flight work piles up. A **closed** workload is driven by a population: a fixed number of simulated users each sends a request, waits for the response, waits a think time, and sends the next. Concurrency is the input and throughput is the outcome, so a system that slows down automatically receives less load — the workload throttles itself. The choice changes what the test can show. Traffic arriving from an unbounded internet audience is open-like; sessions behind a fixed pool of terminals, agents or connections are closed-like. Picking the wrong one either invents backpressure the real system does not have, or hides the collapse the real system would suffer.

code

pseudocode · 10 lines
pseudocode
// closed: population is the input, throughput is the result
for each of N simulated_users in parallel:
    loop:
        response = send(request)     // blocks until it returns
        wait(think_time)

// open: arrival rate is the input, concurrency is the result
every interval_drawn_from(mean: 1 / target_rate):
    spawn:
        response = send(request)     // no one waits for it before the next arrival

go deeper

for a junior

Be ready to state the difference in one sentence each: a rate that keeps arriving no matter what, versus a fixed set of simulated users who each wait for a response before sending again. Knowing which quantity you set and which you read back is the core.

for a middle

Explain the mechanics: concurrency as an outcome in one model and a setting in the other, think time as the throttle inside a closed loop, and the relationship linking population, think time and achieved rate. Show why a population alone is not a workload.

for a senior

Demonstrate judgement about which model matches the real arrival process, and describe the failure each one hides — the queue pile-up a closed run cannot produce, and the phantom load an open run invents for a genuinely bounded population.

for a principal

Own the modelling decision across a system with several channels: mixed open and closed workloads running together, how results are reported so releases stay comparable, and how the organisation avoids optimising for a failure mode its real traffic cannot cause.

### Two ways to decide when the next request is sent Every load generator has to answer one question repeatedly: when do I send the next request? The two families of answers define the workload model. In an **open model**, arrivals are governed by a rate. The generator has a schedule — 320 requests per second, perhaps with randomised spacing so arrivals are bursty rather than metronomic — and it follows that schedule independent of what the system under test is doing. The number of requests in flight at any moment is a *result* of the run, not a setting. If the system slows, arrivals keep coming, in-flight work accumulates, and the concurrency the system experiences rises. In a **closed model**, arrivals are governed by a population. You configure a number of simulated users. Each one repeats a cycle: send a request, block until the response arrives, wait a think time, send the next. The concurrency is therefore capped at the population size, and throughput is the *result*. If the system slows, each user completes fewer cycles per minute, so the offered load falls automatically. ### Why the difference is not cosmetic The two models produce different failures on the same system. A closed model has built-in negative feedback. Doubling response time roughly halves the request rate the population produces, which relieves pressure. Real systems that behave this way exist — a fixed pool of internal terminals, an agent desktop where each agent can only do one thing at a time, an upstream service with a bounded connection pool. For those, the closed model is the honest one, and it will faithfully show demand backing off as latency rises. An open model has no such feedback. If arrivals continue at a rate the system cannot sustain, queues grow without bound, memory follows them, timeouts fire, and the failure mode is the queue-collapse behaviour that public-facing systems actually exhibit under a surge. If your real traffic comes from a large audience that does not coordinate with your service's health, a closed model can hide this entirely: you never observe the pile-up, because your own generator quietly stopped sending. The reverse error is just as real. Modelling a genuinely closed population as an open stream invents load that could not exist — arrivals from users who, in reality, are still sitting and waiting for a response — and produces a pessimistic picture that leads to over-provisioning or to chasing a failure mode nobody can trigger. ### The relationship between the two settings A useful identity connects them: at steady state, average concurrency equals throughput multiplied by average residence time (Little's Law). In a closed run with think time, the population is split between users waiting on a response and users in think time, so the request rate produced by N users is roughly N divided by (response time + think time). That is why raising think time lowers the rate a population generates, and why a closed run's throughput is not something you set — it is something you read back and report. Practically, this means a closed configuration cannot be quoted as a load level on its own. Two hundred simulated users is not a workload; two hundred users with a stated think time, on a system with a stated response time, produced a measured 137 requests per second — that is a workload. Reporting only the population makes runs incomparable across releases, because the same population produces different load as the system gets faster or slower. ### A worked example An insurance quote engine is exercised two ways. The closed run configures 180 simulated brokers, each pausing 4.5 seconds between quotes; it settles at 31 quotes per second, and when a rating dependency slows down, the achieved rate drops to 22 per second while the reported latency stays modest. The open run offers a steady 31 quotes per second; when the same dependency slows, arrivals keep coming, in-flight quotes climb from about 40 to several hundred, the thread pool saturates and the run ends in timeouts and a rising error rate. Both runs are correct — for different production realities. Broker desktops really are a bounded population, so the closed run predicts the broker channel well. The public quote widget on the website is not bounded, so only the open run predicts what a marketing burst will do to it. A team that runs only the closed shape will be surprised the first time the widget is featured somewhere. ### Choosing, and saying which you chose Ask what actually limits arrivals in production. If it is a countable population with a device or seat each, model closed and set think times from observed session behaviour. If arrivals come from an audience whose size and impatience are independent of your latency, model open and set the rate from observed traffic. Mixed systems deserve mixed workloads — an open stream for the public path and a closed population for the back-office path, running together. Whatever you choose, state it in the report along with the rate or population, because a result without its workload model cannot be interpreted or reproduced.

  • Your open-model run is configured for 400 requests per second but only achieves 260. What does that tell you?
    That the generator could not keep to its schedule — either the system under test is holding responses long enough to exhaust the generator's own resources, or the generator is undersized. An open run that silently degrades into a closed one is the classic way a test loses its unbounded-arrival property, so achieved rate against offered rate should be a first-class chart, not something you discover afterwards.
  • How would you choose the think time for a closed workload?
    From observed behaviour rather than intuition: the distribution of gaps between consecutive actions in the same real session. Use a distribution with spread rather than a single constant, because identical constant pauses make a population march in lockstep and produce artificial waves of load. Then sanity-check the result — the achieved rate the population produces should resemble the traffic rate you are trying to reproduce.
  • Can a closed model ever show the queue pile-up an open model shows?
    Only by raising the population until the concurrency you want is reached, and that is fragile: the population needed depends on the response time you are trying to measure, so the setting drifts as the system changes. If the question is about behaviour under arrivals the system cannot absorb, the open model states it directly and reproducibly.

A closed workload is a barbershop with a fixed set of regulars who each come back only after their last haircut; an open workload is the street outside, where people keep walking in at whatever rate they please no matter how long the queue already is.

saying these in an interview costs you the question

  • Treating a count of simulated users as a load level on its own
  • Assuming a closed run can reveal unbounded queue growth
  • Setting think time to zero to get more load
  • Reporting a closed run's population but never its achieved rate
  • Choosing the model by generator default rather than by real arrivals
  • Believing an open run guarantees the requested rate was actually sent

context