What does it mean to join a thread you started, and what problem does joining solve that simply starting the thread does not?
answer
- start = fire-and-forget, join = wait for death
- join gives happens-before: read results safely after
- join ≠ cancel; passive waiting only
- exception escaping child still looks like normal end
- always prefer a timed join in request paths
basics
~20 sJoining means one thread blocks until another finishes. Starting a thread only launches it; the starter keeps running and has no idea when the work is done or whether it succeeded. Join gives you completion timing and a safe point to read results.
solid answer
~50 sStarting a thread is fire-and-forget: the caller continues immediately and the new thread runs independently. **Joining** blocks the caller until the target thread terminates (normally or by throwing out of its entry function). Join buys three things: 1. **Ordering** — after a successful join, everything the joined thread did is guaranteed visible to the joiner, so you can read results it produced without extra synchronization. 2. **Lifetime control** — you know the thread is gone, so you can release resources it used or exit the program safely. 3. **Structure** — join turns a fan-out into a scoped operation: split work across N threads, join all N, combine. Caveats: join tells you the thread ended, not that it succeeded — an exception that escaped the thread still ends it "normally" from the joiner's point of view, so results/errors must be passed out explicitly. Prefer a timed join, otherwise a hung worker hangs the joiner too.
code
text · 10 linesresults = new array[N]
threads = []
for i in 0..N-1:
threads.add(start(() -> results[i] = compute(i)))
// without the joins below, reading results[] is a data race
for t in threads:
t.joinWithTimeout(30s) // timed: a stuck worker must not hang us forever
total = sum(results)go deeper
Say clearly that start returns immediately and join blocks until the other thread finishes, and give the fan-out/collect-results example.
Add the visibility guarantee — after join you can read the child's writes without extra synchronization — and note that join does not carry errors.
Stress timed joins in production paths, cancel-then-join as the correct stop sequence, and join as a wait-for edge that can deadlock.
Frame join as the low-level primitive whose scoped use becomes structured concurrency, and argue for futures/scopes so lifetime and error propagation are enforced rather than left to convention.
## What starting a thread actually does When you start a thread, you hand the runtime a function to execute and it schedules a new independent flow of control. The call returns immediately. From that instant there are two threads: the *starter*, which continues at the next statement, and the *child*, which begins executing its entry function whenever the scheduler picks it. Nothing coordinates them. The child may finish before the next line of the starter runs, or minutes later, or never. That is exactly the problem join solves. ## Join = wait for termination `join` (also called *await*, *wait for*, or in some runtimes simply waiting on the thread handle) blocks the calling thread until the target thread has terminated. While blocked, the joiner consumes no CPU; the scheduler wakes it when the target ends. Termination here means the target's entry function returned *or* threw an exception that escaped it. Either way, the thread is dead. Join does not tell you which happened — that is a common trap. If you need the outcome, the child must publish it: write a result into a shared slot, put it on a queue, or complete a future. ## The three guarantees join gives you **1. A visibility/ordering guarantee.** Under any realistic memory model, the actions of a thread that has terminated *happen-before* the return of a successful join on that thread. Concretely: the child can write to a plain, unsynchronized variable, and after join the parent may read it and is guaranteed to see the final value — no lock, no atomic needed. Without the join, reading that variable is a data race and the parent may observe a stale or partially-constructed value. This is the single most under-appreciated property of join in interviews. **2. Lifetime control.** Resources the child touches — buffers, file handles, memory it borrowed from the parent's stack in languages that permit that — can only be reclaimed once you know the child cannot touch them again. Join is the proof. It is also how a program guarantees that all work is complete before it exits, since in most runtimes the process may terminate while ordinary background threads are still running. **3. Structure.** Join is what turns an unbounded fan-out into a bounded, scoped operation: ``` results = array of N slots threads = for i in 0..N-1: start(worker(i, results)) for t in threads: t.join() combine(results) ``` After the loop, every worker is finished and every slot is safely readable. The lifetime of the child threads is nested inside the lifetime of the block that created them. This nesting is the seed of *structured concurrency*, where the language or library enforces the pattern instead of leaving it to your discipline. ## What join does not do - **It does not stop the thread.** Join is passive waiting. If the child loops forever, join waits forever. To end a child early you must cancel cooperatively (a flag, an interrupt-style signal, closing the input it reads) and *then* join to confirm it stopped. - **It does not propagate failures.** If the child throws, most runtimes log or hand the exception to a per-thread failure hook and the thread simply ends; the joiner sees a normal return. Errors must be carried out of the thread explicitly, which is precisely what futures/promises add on top of raw threads. - **It does not compose safely without timeouts.** An unbounded join in a request path converts a stuck worker into a stuck request, then a stuck thread pool, then an outage. Use a timed join: wait up to T, and if it expires, take a deliberate decision (abandon and report, retry, escalate) rather than blocking indefinitely. ## Self-join and cycles A thread joining itself deadlocks by definition — it would have to terminate before it can proceed. More subtly, two threads that join each other deadlock the same way. Join participates in deadlock cycles exactly like a lock does, because it is a wait-for edge in the same dependency graph. ## When you would not join Long-lived background threads — a metrics flusher, a log shipper, a scheduler tick — are not joined at every use; they are joined once, at shutdown, after being asked to stop. And in async models you rarely join threads at all: you await a future, which carries both the value and the error and does not tie up an OS thread while waiting. Join is the primitive; futures and structured-concurrency scopes are the ergonomic layers built over it.
- After a successful join, do you still need a lock to read a variable the child thread wrote?No. A successful join establishes an ordering edge: everything the joined thread did happens-before the join returns, so the joiner is guaranteed to see the child's final writes. A plain read is correct and race-free. You would still need synchronization if any *other* thread can also write that variable, since join only orders you against the one thread you joined.
- How do you get an exception thrown inside a child thread back to the thread that started it?Not through join — join reports termination, not outcome. The child must capture the failure and publish it: store it in a shared result slot, push it onto a queue, or complete a future/promise exceptionally so the waiter can rethrow it. That is exactly the value futures add over raw threads: a single handle carrying value-or-error plus completion. Runtimes also offer a per-thread uncaught-handler hook, but that is a last-resort log-and-notify path, not a way to return the error to a specific caller.
- What happens if a thread joins itself, and why?It deadlocks permanently. Join waits for the target thread to terminate, and the target is the waiter, which cannot terminate while it is blocked waiting. Some runtimes detect the trivial self-join and throw, but the general case — A joins B while B joins A — is an ordinary wait-for cycle and no runtime detects it for you.
Starting a thread is mailing a task to a colleague; joining is standing at their desk until they hand it back. Standing there doesn't make them work faster, and it doesn't tell you whether the answer is correct — only that they are done.
saying these in an interview costs you the question
- Believing join stops or kills the other thread rather than merely waiting for it
- Assuming join tells you whether the work succeeded, so failures vanish silently
- Reading data written by the child without joining or otherwise synchronizing, then calling it 'fine because it's just an int'
- Using unbounded joins in a request path, so one stuck worker hangs the caller forever
- Thinking join makes execution sequential — the threads still ran concurrently; only the waiting point is ordered