skip to content

A cgo call into a C scoring engine takes 50-400 ms; how do you decide whether it stays on the RPC request path?

level: principalimportance: nice to knowfreq 24%

answer

  1. concurrency times duration is threads
  2. publish a ceiling, enforce on entry
  3. your deadline cannot stop C
  4. one fault kills every in-flight request
  5. isolation costs serialisation, buys blast radius

basics

~20 s

Turn it into a capacity and blast-radius decision. Concurrency times call duration is the OS-thread budget, so publish a cap and enforce it on entry; then judge whether a library that can kill the whole process, and that ignores your deadlines, belongs in the request path at all.

solid answer

~50 s

Three questions decide it. First, capacity: in-flight calls times duration is the thread count, so I pick a ceiling from something real — cores, or the engine's own parallelism — enforce it with a semaphore at the boundary, and let overload queue in Go where a deadline can shed it. Second, honesty about deadlines: a `context.Context` cancellation cannot stop a running C call, so unless the engine takes its own timeout I must not promise callers a bound I cannot enforce, and I must make sure retries do not stack calls. Third, blast radius: a fault in that library takes down every in-flight request on the instance, so if it is unmaintained or memory-unsafe I argue for an out-of-process worker and pay serialisation for isolation. What I commit to is a published thread ceiling, an alert on gate wait time rather than on raw thread count, an owner for engine upgrades, and a documented reason the performance or reliability lead can overrule.

go deeper

for a junior

Take away the two facts the decision rests on: a call into C occupies a real operating-system thread until it finishes, and a crash inside that library ends the whole process rather than one request.

for a middle

Be able to describe the mechanics you would put in place — a semaphore capping concurrent calls, shedding at the gate with the request deadline — and to explain why a context deadline alone does not stop the C call.

for a senior

Show the diagnosis and the fix together: measure the engine's throughput knee, size the ceiling from it, prove thread count tracks in-flight calls, and identify whether the failure mode is containable in process.

for a principal

Own the tradeoff and its record: publish the concurrency ceiling and its reasoning, state the latency you can actually enforce, name the owner for engine upgrades and the rollback path, and say in advance on what grounds a reliability lead should overrule you.

## Why this is a decision and not a tuning exercise A 50-400 ms C call on the request path is not a slow function. It is a commitment of an operating-system thread per concurrent request, a failure mode that is not contained to one request, and a latency promise you cannot actually keep. Each of those is somebody's budget, which is what makes this an ownership question rather than a code change. ## Capacity: the thread budget is the real number A goroutine inside C holds a thread for the duration and cannot be preempted, so in-flight calls translate almost one-for-one into threads. At 300 ms and a few hundred requests per second you are provisioning hundreds of threads, each with a stack and scheduling cost, and the runtime does not give them back afterwards — the count is a high-water mark for the process's life. So the first decision is a **number you publish**: the maximum number of concurrent calls into the engine, derived from cores, memory, or the engine's own internal parallelism and any licence limit — never from incoming request concurrency, which is not yours to control. Enforce it on entry with a semaphore. The important consequence is that overload becomes queueing in Go instead of thread growth, and queueing is measurable and sheddable: a request whose deadline would expire waiting at the gate never enters C at all. That also changes what you page on. Raw thread count is a lagging, noisy signal; **wait time at the gate** tells you the engine is undersized while there is still time to act. ## Honesty: your deadline is not a deadline Cancellation in Go is cooperative. When the request context expires, the handler can stop waiting, but the C function runs to completion holding its thread. Two things follow. You must not advertise a p99 the engine can violate at will, and you must make sure a client retry does not put a second call in flight while the first is still running — otherwise a latency blip becomes a thread avalanche. If the engine exposes its own timeout or cancel handle, use it and your deadline becomes real; if it does not, that absence is a legitimate reason to push the work off the request path. ## Blast radius: what one bad pointer costs An in-process C dependency shares your address space. A memory fault inside it is not a recoverable panic; the process dies and takes every in-flight request with it. A panic that reaches the boundary from an exported callback does the same. That risk is acceptable for a library your team maintains and fuzzes, and much less acceptable for a vendored numeric engine nobody has touched in three years. The alternative is running the engine as a separate process and talking to it over a socket. You pay serialisation of the feature vector and a copy, and you buy: a blast radius of one worker, the ability to restart it, an operating-system-level memory and CPU limit on it, and a real timeout because you can abandon or kill a process. On a call already costing hundreds of milliseconds, a serialisation cost measured in microseconds is not the deciding factor — the deciding factor is whether you trust the code. ## Off the path entirely Before any of that, ask whether the call must be synchronous. If the score can be precomputed for common inputs, cached with a short lifetime, or computed asynchronously with the response returning a pending state, the cheapest engineering answer is not to make the call per request. If it must be synchronous, batching several requests into one crossing amortises both the boundary cost and the thread. ## What you actually commit to A decision here is worth writing down, because it will be revisited by someone who was not in the room: - the concurrency ceiling and the reasoning behind the number; - the alert on gate wait time, and the thread-count reading that confirms the model; - the latency the service promises, stated as what it can enforce rather than what the engine usually does; - who owns engine upgrades, and that a behaviour change on upgrade is their cost to carry; - the rollback path, which for an in-process engine usually means a flag that routes to the off-path implementation. ## Who overrules you The service owner makes the call; the performance or reliability lead can overrule it, and the strongest grounds for doing so are the ones that are not about this service: a thread budget that threatens co-tenants on the node, an unpatchable dependency, or a stated latency objective the boundary cannot honour. Naming those grounds in advance is what turns a preference into a decision.

  • What would convince you to move the engine out of process despite the extra cost?
    Evidence that its failures are not containable: crashes in production, a memory-unsafe library nobody maintains, or a need for a hard timeout the C API cannot provide. Against a 50-400 ms call, serialisation is noise, so the trade is bought with risk, not with latency.
  • What do you alert on, given the thread count itself is a lagging signal?
    Wait time at the admission gate, plus the number of requests shed there. Both move before saturation and both point directly at the engine being undersized. Thread count and the threadcreate profile stay as confirmation that reality matches the model, not as the paging signal.
  • How do you size the concurrency ceiling without simply guessing?
    Measure the engine's throughput against concurrency on the target hardware until latency rises without throughput improving; that knee is the useful ceiling. Cross-check it against the engine's own internal threading and any licence limit, then leave headroom and re-measure after upgrades.
  • The engine's maintainers ship an upgrade that changes scores slightly. Whose problem is that?
    Whoever owns the dependency, and that owner must be named before adoption. Practically it means a shadow comparison of old and new scores on real traffic, an agreed tolerance, and a rollback flag — an upgrade that silently changes output is a product change, not a maintenance task.

saying these in an interview costs you the question

  • Sizes the ceiling from request concurrency instead of hardware
  • Promises a latency bound the C engine can ignore
  • Treats a crash in the library as a per-request failure
  • Dismisses out-of-process isolation purely on serialisation cost
  • Leaves engine upgrades unowned because it builds fine