The Java Memory Model explicitly forbids out-of-thin-air values. What is an out-of-thin-air value, why must a language memory model rule them out, and why is that the hardest part of such a specification to write?
answer
- Self-justifying value: read causes the write that supplies it
- Reference forging -> type safety and sandbox
- Java defines racy behaviour; C++ leaves it undefined
- Causality: commit actions in stages, each independently justified
- Formulation known to be imperfect, still unreplaced
basics
~20 sAn out-of-thin-air value is one a read returns that no write ever justified except circularly - the read causes the write that supplies it. Java forbids them so racy code stays type-safe; it is hard because the rule must block circular causality without banning real optimizations.
solid answer
~60 sIn a racy program, imagine two threads where each copies a value from one field to the other. If a read could speculate a value, and that speculation makes the write that produces it happen, the value 42 could appear having never been written by anyone. That is an out-of-thin-air (OoTA) value: it is self-justifying. Java must forbid this for security, not just cleanliness. If a racy read of a reference field could return a fabricated value, untrusted bytecode could forge a pointer and break type safety and the sandbox. Java therefore defines behaviour even for racy programs, unlike C++ where a data race is undefined. It is the hardest part of the model because the obvious fix - forbid speculation - would outlaw optimizations compilers legitimately perform on unshared data. The JMM instead uses causality rules built on committing actions in successive stages, where each committed action must be justified by an execution not relying on it. Later analysis showed this formulation also rejects some transformations real compilers do, so the problem is not considered fully solved.
go deeper
Not expected. Knowing that a racy read can only return a value someone actually wrote is enough.
Be able to state that Java forbids fabricated values so racy code stays type-safe, and contrast that with undefined behaviour in C++.
Give the two-thread self-justifying example, explain the type-safety motivation, and note that the difficulty is separating illegitimate speculation from legitimate optimization.
Discuss the specification-design trade-off - defining behaviour for all programs costs a non-compositional model with known imperfections - and what that implies for language and compiler evolution.
## The shape of the problem Consider a racy program where `x` and `y` start at 0: ``` Thread 1: r1 = x; y = r1; Thread 2: r2 = y; x = r2; ``` Under any sensible reading, both reads return 0. But suppose a compiler or processor speculates that `r1` will be 42, writes 42 to `y` on that speculation, thread 2 then reads 42 and writes it to `x`, and thread 1's read of `x` returns 42 - retroactively confirming the guess. Every step is locally justified, and the value 42 appeared from nowhere. That is an out-of-thin-air value: it exists only because it was assumed to exist. ## Why a language must forbid it With `int` fields this is merely bizarre. Change the type to a reference and it becomes a security hole. If a racy read could return a fabricated reference, code could obtain a pointer to memory it was never given, and the JVM's type safety - the basis of its sandbox and of every security assumption above it - would fall apart. Java's designers therefore committed to defining behaviour for *all* programs, including badly synchronized ones. C++ chose the opposite: a data race there is undefined behaviour, so its model needs no such rule, but a racy C++ program has no guarantees whatsoever. This is the reason Java's model is materially harder to specify than C++'s. Java pays specification complexity in exchange for a hard floor under every program. ## Why forbidding it is difficult The naive fix is "a read may only return a value written by a write that actually happened, and the write must not depend on the read." The trouble is that legitimate optimizations look exactly like the illegitimate ones from the model's viewpoint. - Compilers reorder and speculate constantly, and doing so on data that turns out to be shared is not something the compiler can detect. - Some transformations are provably safe yet produce executions whose formal justification looks circular - for instance when redundant-read elimination or branch-condition analysis proves a value is the same on all paths and then hoists the store. So the rule cannot simply outlaw speculation; it must separate self-justifying results from results that could have been produced without assuming themselves. ## How the specification does it The JMM defines legality by **committing actions iteratively**. Starting from an empty set, actions are committed in successive stages; every committed action must be justified by a well-formed execution that does not itself depend on that action having occurred - intuitively, a witness execution in which the value already arises for independent reasons. A value that only appears if you first assume it can never be committed at any stage, so out-of-thin-air executions are excluded. Sequentially consistent executions are trivially justifiable, which is why race-free programs get SC for free. ## Why it is still an open area Subsequent work found the causality formulation is not tight. It rejects some transformations that real compilers perform and that are believed safe, which means a strictly conforming compiler would have to give up optimizations that in practice everyone keeps performing. It is also not compositional: the legality of a program fragment cannot be decided locally. Research on the Java and C++ models has produced candidate replacements built on notions like promising semantics or dependency tracking, but no revision has been adopted into the specification. The pragmatic consequence is that OoTA is banned in principle, no real JVM produces such values, and the formal machinery that forbids them is acknowledged to be imperfect. ## What this is worth in an interview This is a differentiator question. What signals depth is being able to state *why the rule exists* - type safety and the sandbox, hence the deliberate divergence from C++'s undefined behaviour - and *why it is hard* - that the model cannot distinguish an illegitimate speculation from a legitimate optimization by looking at the resulting execution alone. Candidates who can add that the specification's answer is known to be imperfect, without turning it into a practical worry, are demonstrating genuine familiarity with the specification rather than with a blog summary.
- Why does C++ not need an out-of-thin-air rule of this kind?Because a data race in C++ is undefined behaviour, so the standard owes racy programs nothing and never has to describe what values they may see. Java cannot take that route: untrusted bytecode must remain type-safe and memory-safe no matter how badly it is synchronized, so the model must bound the behaviour of every program, including racy ones.
- Does the imperfection of the causality rules cause practical problems for Java developers today?Not for application code. No production JVM manufactures out-of-thin-air values, and race-free programs are covered by the sequential-consistency guarantee, which the causality rules do handle cleanly. The imperfection matters to compiler writers and specification researchers, because certain optimizations are formally questionable under the current rules while being universally implemented.
saying these in an interview costs you the question
- Describing out-of-thin-air as a bug that actually happens on real JVMs
- Claiming Java handles racy programs the way C++ does, as undefined behaviour
- Thinking the rule is simply 'a read returns some previously written value' with nothing subtle about it
- Presenting the causality rules as a settled, fully satisfactory solution