A packaged Linux binary exits immediately with a vague message and writes no log. How do you use strace to find out which file or permission it is actually failing on?
answer
- ask the kernel, not the error message
- filter to path-taking calls
- read the tail, not the head
- -1 plus a symbolic errno
- ENOENT missing, EACCES forbidden
basics
~20 sRun strace -f -e trace=file on the program and read the failing calls near the end: openat returning -1 ENOENT names a missing path, EACCES names a permission problem. The failed system call is the real error.
solid answer
~40 sstrace prints every system call a process makes, so it tells you what the program actually asked the kernel for rather than what its error message claims. I would run `strace -f -e trace=file -o /tmp/app.trace /usr/local/bin/myapp`, then read the trace backwards from the end. A line like `openat(AT_FDCWD, "/etc/myapp/config.yml", O_RDONLY) = -1 ENOENT` names the missing file outright; `EACCES` means the path exists but the process cannot traverse or read it, which is usually a mode or ownership problem on the file or one of its parent directories. `-f` matters because launcher scripts and daemons that re-exec themselves would otherwise give a trace that stops at `execve`. The one skill is filtering noise: the dynamic loader and libc probe several candidate paths at startup, so a burst of harmless ENOENT is normal.
code
bash · 2 linesstrace -f -y -s 4096 -o /tmp/app.trace -e trace=file /usr/local/bin/myapp
grep -E 'ENOENT|EACCES|EPERM' /tmp/app.trace | tail -20go deeper
Know that strace prints the system calls a program makes, and be able to say that a call ending in -1 ENOENT means a missing path while EACCES means a permission problem. Naming -f and -e trace=file is enough at this level.
Be ready to explain the shape of a trace line — decoded arguments, return value, symbolic errno — and why the loader's harmless ENOENT probes at startup are not the failure. Mention -o, -y and -s and say what each buys you.
Show you can triage quickly: filter the trace, read backwards from the exit, and separate a kernel-level failure from a user-space rejection where every syscall succeeded. Mention that SELinux or AppArmor denials can appear as EACCES with correct mode bits.
Frame when reaching for strace is the right call at all versus fixing the diagnosability gap: a service that fails with an unusable message is a defect in its error reporting, and repeatedly strace-ing production binaries is a symptom worth engineering out.
## What a system-call trace is Every time a program needs something it cannot do on its own — open a file, read a socket, map memory, start another process — it asks the kernel through a system call. `strace` attaches to the process with the kernel's `ptrace` interface and prints each of those requests with its decoded arguments and its result. That makes a trace the ground truth about what the program tried to do, independent of the error message its authors happened to write. Most "it just fails with no useful error" bugs are a missing file, a wrong path, or a permission the process does not have, and all three are visible in one screen of trace output. ## Anatomy of a trace line ``` openat(AT_FDCWD, "/etc/myapp/config.yml", O_RDONLY|O_CLOEXEC) = -1 ENOENT (No such file or directory) ``` Left to right: the syscall name, its arguments with constants decoded (`AT_FDCWD` means "resolve relative to the current working directory"), then `=` and the return value. A failure is printed as `-1` followed by the symbolic errno and the message the C library would render for it. A success shows the real value — a file descriptor number for `openat`, a byte count for `read`. Note that on current glibc a call written as `open()` in C source shows up as `openat()` in the trace, because that is the syscall the library actually issues. ## The options that matter for this job - `-f` follows children created by `fork`/`clone`, and threads. A shell wrapper, a supervisor, or a daemon that re-execs itself will otherwise leave you a trace that ends at `execve`. - `-e trace=file` limits output to the path-taking syscalls. Current strace spells the same class `%file`; `-e trace=network` and `-e trace=openat,connect` are the other two forms you use daily. - `-o /tmp/app.trace` sends the trace to a file. Without it strace writes to stderr and gets tangled with the program's own output. - `-y` annotates each file descriptor with the path or socket it refers to, so `read(7, ...)` becomes readable without cross-checking `/proc/<pid>/fd`. - `-s 4096` raises the string print limit; the default truncates strings at 32 bytes, which will cut off the interesting half of a long path or payload. - `-tt` adds timestamps and `-T` shows how long each call took, which is how a trace turns from "what failed" into "what was slow". ## Read it backwards The useful part of a startup failure is almost always the last few dozen lines: the final failing syscall, then `write(2, "...")` of the program's own message, then `exit_group`. Working from the end avoids drowning in initialization. `grep -E 'ENOENT|EACCES|EPERM' trace.log | tail -20` is the crude version of the same move and is usually enough. ## The errno vocabulary - **ENOENT** — no such path. Check for a typo, a relative path resolved against an unexpected working directory, or a package that never installed the file. - **EACCES** — permission denied by file mode or ACL. Remember that this also fires when a *parent directory* lacks the execute bit, so the named file may be perfectly readable. - **EPERM** — the operation itself is not permitted for this process, typically a missing capability rather than a file mode. On SELinux or AppArmor systems a denial can surface as EACCES with nothing wrong in the mode bits, so check the audit log too. - **ENOTDIR / EISDIR** — a path component is not the kind of object the program assumed. - **EMFILE / ENFILE** — the per-process or system-wide open-file limit is exhausted; a leak, not a path bug. ## Noise you must learn to skip A normal startup contains many failed calls. The dynamic loader reads `/etc/ld.so.cache` and then probes several directories for each shared library; libc probes locale files, `nsswitch` databases, and timezone data. Dozens of ENOENT lines before the program does anything are expected behaviour, not the bug. The failure you want is the one immediately before the program gives up. ## Where strace will not help If the program read its config successfully and then rejected its *contents*, no syscall failed and the trace shows a clean read followed by an exit — that decision was made in user space. Interpreted and JIT'd runtimes bury the interesting logic behind their own abstractions. And a setuid binary traced by an unprivileged user runs without its elevated privileges, so the failure you observe under strace may not be the failure users report.
- You ran strace on the service's launcher script and the trace ends at execve with nothing after it. What went wrong?The trace followed only the script's own process. `execve` succeeded and the real work happened in a child, which strace never attached to. Re-run with `-f` so clones and forks are followed, and consider `-o` with a file so the child's output does not interleave with the parent's on the terminal.
- A trace shows openat succeeding on the config file, yet the program still exits with an error. What does that tell you?That the failure is not a syscall failure. The kernel handed over the bytes, so the rejection happened in user space — a parse error, a validation rule, or a value the program refuses. strace has answered its question; the next step is the program's own debug or verbose flag, or reading the config against its schema.
- The trace prints read(9, ...) and you have no idea what descriptor 9 is. How do you find out?Add `-y`, which annotates every descriptor argument with the path or socket it points at, and `-yy` to expand socket descriptors with protocol details such as the local and remote address. For a process you are attached to rather than launching, `ls -l /proc/<pid>/fd` gives you the same mapping directly.
saying these in an interview costs you the question
- Treating every ENOENT in the trace as the bug
- Tracing a wrapper script without -f and concluding nothing happens
- Believing strace shows internal function calls, not just syscalls
- Reading EACCES as a file mode problem without checking parent directories
- Assuming a clean trace means the program did not fail