When would you use Triton BLS in the Python backend instead of an ensemble?
answer
- ensembles are static, BLS is imperative
- branch, loop, retry, dynamic model choice
- pb_utils.InferenceRequest then exec()
- blocking exec holds the instance
- deadlock risk when instances all wait
basics
~20 sUse Business Logic Scripting when the pipeline needs data-dependent control flow. Inside a Python-backend model you build a pb_utils.InferenceRequest and call exec() to invoke other loaded models, so you can branch, loop, retry or choose a model at runtime — none of which a static ensemble DAG can express.
solid answer
~50 sAn ensemble is a fixed graph: steps, tensor wiring, done. BLS is ordinary Python running in the Python backend's `execute()` that calls other models in the same server. You construct `pb_utils.InferenceRequest(model_name=..., requested_output_names=[...], inputs=[...])` and call `.exec()`, or `await .async_exec()` from an `async def execute()`, then check `response.has_error()` and pull tensors with `pb_utils.get_output_tensor_by_name`. That unlocks conditionals (route cheap requests to a small model), loops (call a model until a stop condition), retries, and aggregation logic. The costs are real: a blocking `exec()` occupies that Python instance for the whole nested call, so concurrency is bounded by instance count and you can deadlock if a model calls itself and exhausts its instances. Tensors may round-trip through CPU unless you use DLPack (`to_dlpack`/`from_dlpack`) or request GPU-preferred memory. The usual production shape is an ensemble for the static spine with one BLS step for the branch.
code
python · 23 linesimport triton_python_backend_utils as pb_utils
class TritonPythonModel:
def execute(self, requests):
responses = []
for request in requests:
image = pb_utils.get_input_tensor_by_name(request, "IMAGE")
infer_request = pb_utils.InferenceRequest(
model_name="classifier",
requested_output_names=["LOGITS"],
inputs=[pb_utils.Tensor("INPUT", image.as_numpy())],
)
infer_response = infer_request.exec()
if infer_response.has_error():
raise pb_utils.TritonModelException(
infer_response.error().message()
)
logits = pb_utils.get_output_tensor_by_name(infer_response, "LOGITS")
responses.append(
pb_utils.InferenceResponse(output_tensors=[logits])
)
return responsesgo deeper
Know that an ensemble is a fixed pipeline declared in config, and that BLS is Python code in the Python backend that calls other models when the flow depends on the data.
Name the actual API — pb_utils.InferenceRequest with exec() or async_exec() — and give a concrete case an ensemble cannot express, such as a confidence-threshold cascade.
Own the operational cost: instance occupancy under blocking calls, deadlock from exhausted instances, DLPack to avoid device-host copies, and how nested time is attributed across metrics.
Decide where orchestration should live at all — in-server BLS versus an external service — weighing latency saved against putting business logic inside a GPU server that must be versioned and deployed with the models.
## The dividing line Triton gives you two ways to compose models in-server, and they differ on exactly one axis: **can the graph depend on the data?** - **Ensemble** — a static DAG declared in `config.pbtxt`. Triton knows the whole shape before any request arrives. Cheap, declarative, no user code in the path. - **BLS (Business Logic Scripting)** — a Python-backend model that calls other models imperatively at request time. The control flow is your code, so it can look at a tensor and decide what to do next. If you can draw the pipeline on a whiteboard without an arrow labelled "if" or "repeat", use an ensemble. ## How a BLS call is written Inside `TritonPythonModel.execute(self, requests)`, for each request you build an inference request against another model in the repository: - `pb_utils.InferenceRequest(model_name="classifier", requested_output_names=["LOGITS"], inputs=[pb_utils.Tensor("INPUT", array)])` - `response = infer_request.exec()` for a blocking call, or `response = await infer_request.async_exec()` inside an `async def execute()`. - `response.has_error()` then `response.error().message()` to propagate failures, usually by raising `pb_utils.TritonModelException`. - `pb_utils.get_output_tensor_by_name(response, "LOGITS")` to read results, then wrap them in `pb_utils.InferenceResponse(output_tensors=[...])`. The called model goes through its own scheduler, so it still batches and still uses its own instances. Nothing about BLS bypasses the target model's configuration. ## What it lets you build - **Cascades.** Run a small cheap model; if its confidence clears a threshold, return; otherwise call the large model. The routing rule is a Python `if`. - **Iteration.** Call a model repeatedly, feeding its output back in, until a stopping condition holds. - **Dynamic dispatch.** Pick `model_name` from a field in the request — per-tenant models, per-language models, A/B splits. - **Aggregation and fallback.** Call two models, compare, retry one on failure, merge outputs. ## The costs interviewers want you to name **Occupancy.** A blocking `exec()` holds the calling Python instance for the entire duration of the nested inference. With `count: 1`, your BLS model handles exactly one request at a time no matter how idle the GPU is. Two fixes: raise the Python model's `instance_group` count, or make `execute` async and use `async_exec()` so one instance can have many calls in flight. **Deadlock.** If a BLS model calls itself, or two BLS models call each other, and every instance is blocked waiting on a nested call, nothing can make progress. Instance counts must exceed the maximum call depth, and self-recursion should be avoided outright. **Tensor copies.** Naive BLS converts tensors to NumPy, which means device-to-host and host-to-device copies around every hop. For large tensors that can erase the benefit of staying in-process. Use `pb_utils.Tensor.from_dlpack` and `to_dlpack` to pass GPU tensors without copying, and set `preferred_memory` on the request when you want the response left on the GPU. **Python itself.** The Python backend runs a separate process per instance and is subject to Python's execution model. Heavy per-request Python work between model calls is CPU-serial inside that instance. **Observability.** The nested call is attributed to the called model's metrics, while the wall-clock is charged to the BLS model. When a BLS model looks slow, split the time yourself with your own timing or by comparing the inner model's compute duration against the outer model's total. ## Decoupled mode A BLS model can also be **decoupled** — declared with `model_transaction_policy { decoupled: true }` — so it sends zero, one or many responses per request using the request's response sender. That is how streaming-shaped pipelines are built in the Python backend, and it pairs with BLS when each loop iteration should emit a partial result rather than waiting for the whole loop to finish. ## The pragmatic answer Most real pipelines are mostly static with one decision in the middle. The best structure is usually an ensemble whose spine is declarative and whose one branching step is a small Python-backend model doing BLS. You keep the declarative wiring, its per-step tuning and its metrics, and you pay for Python only where you actually need a decision.
- Your BLS model shows one request in flight at a time despite an idle GPU. What are the two fixes?Either raise the Python model's instance_group count so several processes can each hold a blocking call, or convert execute to an async coroutine and use await infer_request.async_exec(), which lets one instance keep many nested calls outstanding. Async scales better because it does not multiply process and weight memory, but it requires the surrounding Python to be genuinely non-blocking.
- How do you avoid copying a large GPU tensor when passing it between BLS calls?Move tensors through DLPack rather than NumPy: pb_utils.Tensor.from_dlpack builds a tensor from a device buffer and to_dlpack hands one out without a host copy. You can also set preferred_memory on the InferenceRequest so the response is placed in GPU memory. Without this, every as_numpy() call round-trips through host memory and can dominate the pipeline's latency.
- When is a BLS deadlock possible, and how do you prevent it structurally?When all instances of a model are blocked inside exec() waiting on a model that ultimately needs one of those same instances — self-recursion, or a cycle between two BLS models. Prevent it by forbidding cycles in the call graph and sizing instance counts above the maximum nesting depth. Async execution helps because a waiting coroutine does not hold the instance's only execution slot.
saying these in an interview costs you the question
- Claiming ensembles support if/else or loops
- Leaving instance count at 1 with blocking exec calls
- Ignoring host copies from as_numpy on GPU tensors
- Letting a BLS model call itself recursively
- Thinking BLS bypasses the called model's batcher