skip to content

A multithreaded service shows noticeably worse throughput on Alpine Linux than on a glibc-based distribution, and a deeply recursive worker thread crashes with a segmentation fault that never happened before. Which musl defaults explain the two symptoms, and what would you do?

level: seniorimportance: should knowfreq 32%

answer

  1. same kernel, different libc defaults
  2. allocator design versus thread scaling
  3. per-thread arenas versus a shared structure
  4. spawned thread stacks are tiny
  5. main thread gets its stack from the kernel

basics

~20 s

Two musl defaults: its allocator is designed for small size and low fragmentation rather than multithreaded scalability, so allocation-heavy threads contend; and its default thread stack is far smaller than glibc's 8 MiB, so deep recursion in a spawned thread overruns it.

solid answer

~50 s

Both symptoms come from musl's defaults rather than from anything Alpine-specific in the kernel. glibc's allocator keeps per-thread arenas so concurrent allocation mostly avoids a shared lock; musl's is built for small footprint and low fragmentation, and allocation-heavy multithreaded workloads contend on it, which shows up as throughput that does not scale with cores. The crash is separate: musl gives new threads a default stack orders of magnitude smaller than glibc's 8 MiB, so recursive or large-frame code that never came close to the limit before now runs off the end and hits the guard page. Note that only spawned threads are affected — the main thread's stack comes from the kernel and `RLIMIT_STACK`, which is the same on any distribution. The fixes are to set thread stack size explicitly, for example with the runtime's own option, and to link or preload a different allocator — or to accept that this workload wants a glibc base.

code

bash · 3 lines
bash
# Only spawned threads use musl's small default stack:
ulimit -s                    # main-thread limit, set by the kernel, same on any distro
java -Xss1m -jar app.jar     # state the thread stack size instead of inheriting it

go deeper

for a junior

Know that the C library affects how a program runs, not just whether it compiles, and that musl on Alpine has different defaults from glibc for memory allocation and thread stacks.

for a middle

Explain both mechanisms: glibc's per-thread allocator arenas versus musl's small-footprint allocator, and musl's much smaller default thread stack against glibc's 8 MiB, plus why the main thread is exempt.

for a senior

Show the diagnosis and the remedies: prove the allocator is the bottleneck by measuring scaling and preloading an alternative, set thread stack size explicitly rather than inheriting a default, and check for a musl-targeted runtime build before anything else.

for a principal

Own the base-system decision on evidence. Weigh the footprint and attack-surface gain against measured throughput loss, source builds of native dependencies and the runtime-availability constraint, and be willing to reverse a fleet-wide standard when the numbers say so.

## Same kernel, different libc Nothing here is about Alpine's kernel — it is the same Linux everyone else runs. It is about musl, and specifically about two defaults musl chose differently from glibc because musl optimises for small, correct and predictable rather than for maximum throughput on a large machine. ## Symptom one: allocation does not scale glibc's allocator maintains multiple arenas and hands threads their own, so concurrent `malloc`/`free` from many threads mostly proceeds without touching a shared structure. It costs memory — arenas reserve address space — and it can look wasteful in RSS, but it scales. musl's allocator (the mallocng design, since musl 1.2.1) is written for a small footprint, strong fragmentation behaviour and hardening. It is not designed around per-thread arenas, so a workload whose hot path is allocating and freeing small objects across many threads serialises far more than it did on glibc. The signature is characteristic: single-threaded benchmarks look fine or even good, and throughput flattens as you add cores while CPU time goes into the allocator rather than your code. Managed runtimes and interpreted languages are hit hardest, because their execution model allocates constantly. The mitigations, in order of preference: reduce allocation in the hot path, which helps everywhere; link or `LD_PRELOAD` an allocator built for concurrency, such as jemalloc, which is the usual pragmatic fix; or move the workload to a glibc base and be honest that the image-size saving was not worth the throughput. ## Symptom two: the stack is much smaller glibc gives each new thread an 8 MiB stack by default. musl's default is smaller by orders of magnitude — a design choice that suits small systems and thousands of threads, and that quietly breaks code written against the glibc default. The failure mode is a clean segmentation fault with no application-level error: the thread walks past the end of its stack into the guard page and the kernel kills it. Anything with deep recursion, large stack-allocated buffers, or a framework that recurses per request will hit it, and it will hit only in threads the program spawned. That last point is the detail worth knowing: the **main** thread's stack is not allocated by libc at all. The kernel sets it up and grows it up to `RLIMIT_STACK`, which `ulimit -s` shows and which is typically 8 MiB on any distribution. So the same recursive function succeeds on the main thread and dies on a worker, on the same Alpine host — a genuinely confusing bug report if you do not know why. ``` ulimit -s # main-thread limit, from the kernel, not from musl java -Xss1m -jar app.jar # give every JVM thread an explicit stack ``` The fix is to stop relying on the default. Code that creates threads directly should set the stack size explicitly through the thread attribute API rather than inheriting whatever the libc chose; managed runtimes expose their own option for it. Setting it explicitly is better practice on every platform, because it turns an invisible platform-dependent default into a stated requirement. ## Related traps in the same family **Runtimes need a musl build.** A JVM, or any large runtime, must be a build targeting musl; a glibc build will not start, and it fails at the dynamic-loader level rather than with a helpful message. Several vendors publish musl-targeted JDK builds precisely for this. **Native extensions.** Anything loading compiled native code — a language's C extensions, a database driver, an ML library — needs a musl build of that code too. Where upstream publishes prebuilt artifacts only for glibc, the fallback is compiling from source, which turns a fast install into a slow one and needs a toolchain present. **Diagnosis discipline.** Before attributing a difference to musl, confirm the two systems are otherwise comparable: same CPU count, same memory limits, same runtime version and flags. musl explains a great deal on Alpine, and it is also a convenient scapegoat for a configuration difference nobody checked. ## The judgment being tested The interviewer is not looking for allocator internals. They want to see that you know a libc is a runtime with performance and resource defaults, not just a set of function names; that you can separate two co-occurring symptoms with different causes; and that you can say out loud that the right resolution may be to abandon the smaller image, because the footprint was never the point — the service was.

  • Why does the same recursive function survive on the main thread but crash on a worker thread?
    Because the two stacks come from different places. The main thread's stack is created by the kernel at exec and grows on demand up to RLIMIT_STACK — commonly 8 MiB, and identical across distributions. A spawned thread's stack is allocated by the C library at creation time using its own default, and musl's default is far smaller than glibc's. Set the size explicitly when creating threads and the discrepancy disappears.
  • How would you confirm the allocator is the bottleneck rather than guessing?
    Measure rather than assume. Compare single-threaded against multi-threaded throughput on the same host: an allocator bottleneck shows scaling that flattens as threads increase while CPU time concentrates in allocation paths. Then test the hypothesis directly by preloading a concurrency-oriented allocator such as jemalloc and re-running the identical benchmark. If throughput jumps, you have your answer with evidence.
  • Is switching to a glibc base image an admission of failure?
    No — it is the trade being priced correctly. Alpine buys a smaller footprint and a smaller attack surface; if a service pays for that with throughput, crashes and source builds of every native dependency, the saving was never real. The engineering answer is to state what each option costs and pick deliberately, rather than treating the base system as a fixed constraint nobody may revisit.

saying these in an interview costs you the question

  • Blames Alpine's kernel rather than the C library
  • Says musl limits recursion depth directly
  • Thinks ulimit -s controls spawned thread stacks
  • Assumes any JDK build runs on a musl system
  • Treats the base image choice as unchangeable

context