skip to content

Which confinement wall explains a thumbnail worker that fails with a permission error inside its decoding library, on files that are readable?

level: seniorimportance: should knowfreq 46%

answer

  1. the error names the wrong thing
  2. no file is involved at all
  3. a refused call, not a refused path
  4. reproduce once with confinement relaxed
  5. permit one call, not the whole tail

basics

~20 s

Most often the syscall allowlist. A refused kernel call is returned to the caller as a plain permission error, so it surfaces as a library failure on an operation that has nothing to do with file permissions.

solid answer

~50 s

Two walls can produce that symptom, and the failing operation tells you which. If the operation touches no path or device at all — initialising an accelerated decoder, starting a helper, inspecting another process — a refused kernel call is the likely cause, because the allowlist returns its refusal as a plain error and the library reports it in its own vocabulary. If the operation is opening something that demonstrably exists and is readable, suspect the profile governing which files and devices are reachable. Confirm by reproducing in a throwaway environment with the same image and the same input, then running once with confinement relaxed: if the failure disappears you have the wall, though not yet the call. Then narrow to the single call and permit that one. Switching the filter off is not a smaller version of the fix — it is removing the wall.

go deeper

for a junior

Remember the shape of the clue: a permission error on something that is not a file usually means a refused kernel call, not a broken file permission.

for a middle

Explain why the error is misleading — the refusal is returned to the caller as an ordinary error, so it is reported in the library's vocabulary and nowhere else.

for a senior

Walk the narrowing: reproduce with the same input, relax confinement once to identify the wall, isolate the single call, then permit only that call in confinement scoped to this workload.

for a principal

Make the point that the default is doing its job and the organisation's answer must be a narrow, owned, expiring exception — otherwise every diagnosis of this kind ends with a wall quietly removed.

## Why the error points at the wrong thing When a kernel call is refused by an allowlist, the refusal comes back to the caller as an ordinary error return. The library that made the call takes the failure branch it already had, and reports what it was trying to do: it could not initialise the accelerated path, it could not open the decoder, the operation is not permitted. Nothing in that message mentions confinement, because nothing inside the process knows a filter exists. The result is a **permission error on an operation that involves no file permissions at all**, which is the single most recognisable signature in this material. The reason it is so reliably misdiagnosed is that it looks like two much more familiar things: a file-ownership problem in the image, or a bug in a library version. Engineers therefore rebuild the image, change ownership, and pin a different library — none of which touches the wall that is actually refusing. ## Narrowing it to the right wall | what the failing operation touches | the wall to suspect | how to check it | |---|---|---| | nothing on disk — a decoder handle, a helper process, a timer, another process | the syscall allowlist | run once with confinement relaxed | | a device or path that exists and looks readable | the profile governing file and device reachability | list what the workload can actually see from inside | | a path that does not exist inside the workload at all | what was mounted into it, which is a different wall | inspect the filesystem the workload was given | | a plain read of a file owned by another account | ordinary ownership on that file | cheapest to rule out, so rule it out first | The useful discriminator is not the error text — it is identical in several of these cases — but the **object of the failing operation**. A refusal with no object is a refused call. ## Confirming it without ripping the wall out 1. **Reproduce outside production**, in a throwaway environment, with the same image *and the same input*. This failure is usually input-dependent: only some media reaches the code path that makes the unusual call, which is exactly why it passed through testing. 2. **Run once with confinement relaxed.** If the failure disappears, you have identified the wall. You still do not know the call, and that distinction matters — this is a bisection step, not a fix. 3. **Get the refusal reported if the platform can do it.** Some filters can be configured to record a call rather than refuse it; where that is available it names the call directly and ends the investigation. 4. **Otherwise bisect.** Permit the denied tail back in groups until the failure stops, then split the group. Slow, but it converges on one call. 5. **Ask whether the workload needs the call at all.** Many libraries take a fast path only when it is available and have a configuration switch to stay on the portable one. Then the fix is one setting in the workload, and the profile never changes. ## The fix, and the thing that looks like a fix The fix is to **permit the one call**, in confinement scoped to this workload, with a note saying which call and why. The thing that looks like a fix is switching the filter off for the workload — and it is not a smaller version of the same action. It restores the entire denied tail for the life of that workload, including every entry point the default was there to remove, and it does so permanently: the relaxed setting outlives the release it unblocked, survives in the spec, and gets copied by the next team that starts from this one. The same reasoning applies to the other wall. If a device is genuinely needed, name that device in the workload's reachability profile rather than making everything reachable. ## What a senior answer adds A middle answer names the filter. A senior answer adds four things: that the misleading error is a **property of the design**, not bad luck; that a relaxed run is a bisection step rather than a resolution; that the exception must be scoped to one workload and carry an expiry; and that this class of failure is frequently input-dependent, which is why it reaches production having passed every test. The last point is what tells an interviewer you have actually chased one.

  • How would you tell a refused kernel call from a profile refusing a device?
    By what the failing operation was reaching for. An operation with no object — a handle, a helper, an inspection — points at the call filter. An attempt to open a device or path that exists and is readable points at the reachability profile. Listing what the workload can actually see from inside separates the two in one step.
  • The failure disappears when you switch confinement off entirely. Is that a diagnosis?
    It narrows the cause to that wall, but it does not name the call, and a fix built on it removes the wall rather than the obstacle. Keep narrowing — have the platform report refusals if it can, or permit the denied set back in groups — and ship confinement that adds exactly one call.
  • Why did this pass every test and fail in production?
    Because the unusual call sits on a code path only some inputs reach. A media worker takes the accelerated path for particular formats or sizes, so a test corpus that never contains one exercises only permitted calls. Reproducing it means reproducing the input, not just the image.

saying these in an interview costs you the question

  • Says the fix is to run the workload as the superuser.
  • Blames file ownership in the image and rebuilds it.
  • Switches the whole filter off permanently to unblock the release.
  • Expects a distinctive error that names the confinement filter.
  • Treats a relaxed-confinement run as the fix rather than a bisection step.
  • Assumes the failure is deterministic across all inputs.