skip to content

Your Go service needs a blocking C codec — how do you decide between keeping it in-process with a thread cap and moving it out?

level: principalimportance: nice to knowfreq 26%

answer

  1. make it a number before making it an argument
  2. concurrency times duration
  3. the peak becomes the permanent floor
  4. the cap is a tripwire, not a control
  5. who accepts a crash loop instead of degradation

basics

~20 s

Decide from measured numbers: concurrent calls times call duration is the OS thread demand, and each thread costs a stack. Keep the codec in-process only behind an owned concurrency bound; move it out when the memory or blast radius is unacceptable.

solid answer

~50 s

Turn it into a number first. Concurrent calls multiplied by time spent inside C is your steady-state OS thread demand, each thread carries a system stack, and the runtime never gives those threads back, so a codec that blocks for seconds prices itself out. Keeping it in-process is defensible when calls are short, concurrency is bounded at a single choke point your team owns, and the memory at peak fits the budget. Move it out when a fault in C would take down unrelated traffic, or when bursty load makes the high-water mark the permanent footprint. `debug.SetMaxThreads` belongs in the conversation but is not the answer: it converts a slow leak into a crash, which the on-call team may want and the availability owner may refuse. Write down the ceiling, the bound, and who accepted the trade.

go deeper

for a junior

You are not expected to make this call, but know that a blocking call into C occupies a real OS thread for its whole duration, so the number of them running at once is a resource decision, not a detail.

for a middle

Be able to derive the thread demand — concurrent calls times how long each one blocks — and to explain why an unbounded call site has no capacity limit at all until memory runs out.

for a senior

Show that you would bound the call at a single owned choke point, size the bound from measurement, and set an explicit thread ceiling above it as a tripwire rather than as the control.

for a principal

Own the trade in the open: name whose availability budget a low ceiling spends, what evidence would move the codec out of process, and what standard applies to any blocking foreign call in the estate.

## The decision, stated honestly A transcoding service must run a C codec. Two shapes are on the table: call it from Go through cgo, or run it in a separate worker process the Go service talks to. This is not a taste question and it is not a performance micro-benchmark. It is a capacity and blast-radius decision with an owner, and it is normally settled by four numbers and two risks. ## The four numbers **1. Time inside the call.** Measure it, at the percentiles that matter. A call that takes 50 microseconds and a call that takes 8 seconds are different products. While a goroutine is inside C, the operating-system thread carrying it is unavailable to the runtime. **2. Concurrency at peak.** How many of those calls are in flight simultaneously at the traffic you actually plan for, not the average. **3. Thread demand = 1 × 2.** Arrival rate times duration. This is the number of OS threads the runtime will be pushed to create, and it appears nowhere in the source code, which is why it surprises teams in production. **4. Memory per thread.** Kernel bookkeeping plus a system stack; threads created for calls into C get the platform's default thread stack reservation, commonly 8 MB of address space each on Linux, resident as it is touched. Multiply by number 3 and compare against the container's memory limit. Remember that the runtime does not reap idle threads, so the peak is the permanent floor. If those four numbers put you comfortably inside the memory budget with room for a bad day, in-process is a reasonable answer. If they do not, no amount of tuning rescues it, and the argument is over. ## The two risks **Blast radius.** Everything in one process shares one fate. A fault or a runaway allocation inside the C library takes down the API traffic that had nothing to do with transcoding. A separate worker turns that into a failed job and a restarted worker. **Debuggability.** Thread growth caused by C calls is visible from Go only indirectly — a thread count against a flat goroutine count, a threadcreate profile that under-reports, stacks that stop at the boundary. Out of process, the same problem is a worker's resource usage that your existing operating tools already understand. ## What each option actually costs **In-process, bounded.** You pay engineering discipline: exactly one choke point in one package, with a counting semaphore in front of the C call, sized from measurement, with queue depth and wait time exported. You pay a permanent memory floor set by the bound. You gain no serialisation of the payload, no extra hop, and one binary to deploy. The bound is load-bearing — without it the design has no capacity limit at all — so it must be owned, tested and impossible to bypass, not a convention. **Out of process.** You pay latency for the hand-off, a serialisation cost on a large blob, a supervision and restart story, and a second artefact to build, ship and version. You gain isolation, an independent memory limit, the ability to scale codec capacity separately from request capacity, and an out-of-memory kill that lands on a worker rather than on the API. The blob size matters here: passing a large payload across a process boundary is a real cost that a shared-memory or file hand-off can reduce, and it should be measured rather than assumed either way. ## Where the thread ceiling fits `debug.SetMaxThreads` will come up, usually from whoever carries the pager, in the form "cap it at 500 so we find out immediately". Be clear about what a cap is and is not. It is **not** a throttle. Crossing it is a fatal runtime error: the process dies, `recover` cannot intervene, and in-flight requests unrelated to transcoding die too. It **is** an excellent tripwire: an early crash with full stacks, close in time to the cause, is a far better artifact than a slow slide into an out-of-memory kill with nothing to read. So the two positions are both rational. On call wants a fuse because an unbounded leak at 3am with no stacks is the worst outcome for them. The service owner resists because a fuse converts degradation into a crash loop and spends the availability budget. The resolution is not to pick a side but to order the controls: the **concurrency bound** is the control, the **ceiling** is the tripwire set well above it, and the crash is treated as a bug report rather than as the mitigation. Then write the number down, with the reasoning and the name of whoever accepted it, so the next incident does not relitigate it from scratch. ## The policy layer At organisation scale the useful rule is not "cgo is banned" or "cgo is fine". It is: a blocking foreign call is allowed only behind a single owned interface with an enforced concurrency bound, a documented thread and memory budget derived from the four numbers above, an explicit process thread ceiling, and an alert on thread count against goroutine count. A team that cannot produce those four numbers has not finished the design, and that — not the language of the codec — is what a reviewer should push back on. ## What would change your mind later Call duration regressing after a codec upgrade; peak concurrency growing past the bound often enough that queue wait becomes the latency story; a memory limit cut; or a second team wanting the same codec, at which point a shared worker is cheaper than two bounded copies. Name those triggers when you make the call, so the decision has an expiry condition rather than becoming folklore.

  • The on-call team wants debug.SetMaxThreads lowered to 500. What do you tell them?
    That I agree with the instinct and want it set above the concurrency bound, not below it. A ceiling is a tripwire: crossing it is a fatal error that kills in-flight requests unrelated to transcoding, so it must never be reachable in normal operation. If 500 is below our measured peak we are choosing a crash loop; if it is comfortably above it, it buys an early crash with stacks and I will take that.
  • What evidence would move you from in-process to an out-of-process worker?
    Measured call duration or peak concurrency pushing thread demand past what the memory limit tolerates; a fault in the C library taking down unrelated traffic even once; bursty load whose high-water mark becomes the permanent footprint; or a second consumer wanting the same codec. I would name those triggers when making the original call so the decision has an expiry condition.
  • How do you stop the concurrency bound from being bypassed as the codebase grows?
    Put the C call behind one package that owns the semaphore and exposes only the bounded entry point, keep the raw binding unexported, and add a test that fails if a new call path appears. Export queue depth and wait time so saturation is visible. A bound that lives in a convention rather than in the type boundary will be routed around within a quarter.
  • Does moving the codec out of process eliminate the thread problem?
    No, it relocates it and bounds the damage. The worker still has whatever thread behaviour the codec induces, but its memory limit, restart policy and scaling are independent, so an incident kills a worker instead of the API. You have traded a shared-fate memory risk for a hand-off cost and a supervision story you now own.

saying these in an interview costs you the question

  • Argues from preference without measuring call duration
  • Treats a thread ceiling as the fix rather than a tripwire
  • Ignores that the thread peak becomes the permanent floor
  • Assumes a crash is always safer than degrading
  • Relies on a convention instead of an enforced bound
  • Bans cgo outright without a capacity argument