skip to content

JVM Architecture & Execution

The path from a compiled class file to running native code: the class loader subsystem, the runtime data areas, and the execution engine that interprets first and compiles later. Interviewers use it to check you can describe the machine your code runs on instead of falling back on "the JVM handles it".

on this pageshow

explore

questions

14

What is the bytecode that a Java compiler emits into .class files, and why does the JVM execute that form instead of executing source text directly or compiling straight to native machine code?

level: juniorimportance: must knowfreq 62%

answer

  1. one-byte opcodes → "bytecode"
  2. abstract machine defined on paper, not a CPU
  3. constant pool = symbolic references
  4. verifiable + language-neutral
  5. portability now, translation cost later

basics

~20 s

Bytecode is a compact, platform-neutral instruction set for an abstract machine, stored in .class files. The compiler targets that abstract machine instead of a real CPU, so one compiled artifact runs on any JVM, and the JVM can verify it before running it.

solid answer

~60 s

A Java compiler does not produce machine code for x86 or ARM. It produces **bytecode**: a stream of one-byte opcodes plus operands, one method at a time, stored inside `.class` files together with a constant pool, field/method tables and attributes. Bytecode targets an *abstract* machine defined by the JVM specification — it has an operand stack, local variable slots, and typed instructions (`iadd`, `dadd`, `aload`, `invokevirtual`, `getfield`). Because the target is a specification rather than a chip, the same `.class` or jar runs unchanged on any conforming JVM on any OS/CPU. That is the concrete meaning of "compile once, run anywhere": portability lives at the bytecode level, not the source level. The intermediate form buys two more things. It is **verifiable** — the JVM can statically check types, stack depth and control flow before executing untrusted code. And it is **language-neutral**: anything that can emit a valid classfile runs on the JVM. The cost is that bytecode still has to be turned into real instructions at run time: first by interpretation, later by JIT compilation for hot code.

code

text · 6 lines
text
public int twice(int);
  Code:
     0: iload_1        // push local slot 1 (the parameter)
     1: iconst_2       // push int constant 2
     2: imul           // pop two ints, push product
     3: ireturn        // return int on top of stack

go deeper

for a junior

Be able to say: javac emits .class files containing platform-neutral bytecode for an abstract stack machine, and the JVM turns it into native instructions at run time. Mention portability.

for a middle

Add the structure — constant pool, symbolic references, per-method Code attribute — and explain verification and language neutrality as deliberate benefits of the intermediate form.

for a senior

Frame the tradeoff explicitly: late symbolic linking and runtime profiles buy optimization and dynamism that static native compilation cannot match, at the cost of startup translation work.

for a principal

Discuss when the bytecode contract is the right deployment boundary at all — plugin ecosystems and multi-language platforms versus workloads where startup latency or footprint push toward ahead-of-time images.

## What a .class file actually contains When you compile `Foo.java`, the output is `Foo.class` — a binary file in a format fixed by the JVM specification. It is not "compressed Java source" and not machine code. Its main parts are: - a **magic number** (`0xCAFEBABE`) and a class-file version; - a **constant pool**: a table of every string literal, class name, field name, method name and descriptor the class refers to, so instructions can refer to these by index instead of repeating text; - access flags, the superclass and interfaces; - **field and method tables**; each non-abstract method carries a `Code` attribute; - inside `Code`: the actual **bytecode array**, the maximum operand stack depth, the number of local-variable slots, exception-handler ranges, and optional debug attributes (line numbers, local variable names). The bytecode array is a sequence of instructions. Each instruction begins with a single-byte **opcode** (hence "bytecode"), optionally followed by operand bytes. There are around 200 defined opcodes, which is why one byte suffices. Examples: `iconst_1` (push int 1), `iload_2` (push local slot 2 as an int), `iadd` (pop two ints, push their sum), `getfield #7` (pop an object reference, push a field named by constant-pool entry 7), `invokevirtual #12`, `areturn`. ## Why an intermediate form at all **Portability.** A native compiler must commit to an instruction set, register file, calling convention and OS ABI. Bytecode commits to none of them: it targets an abstract machine defined on paper. Shipping bytecode means shipping one artifact for every platform that has a JVM. The famous slogan "write once, run anywhere" is really "compile once to bytecode, and let each platform's JVM deal with its own CPU". **Verifiability and safety.** Because bytecode is typed and structured, the JVM can run a **verifier** over a class before executing it: it checks that the operand stack never underflows or overflows the declared maximum, that types match the instructions applied to them, that jumps land on instruction boundaries inside the method, and that `final` classes are not subclassed. Native machine code cannot be checked that way in any practical sense. This is what made downloading and running untrusted code plausible in the first place. **Language neutrality.** The JVM knows nothing about Java syntax. Its contract is the classfile format. Anything that emits a valid classfile — Kotlin, Scala, Groovy, Clojure, a compiler you write this afternoon — runs on the JVM and interoperates with Java libraries, because at run time everything is just classes, methods and descriptors. **Compactness and late binding.** Bytecode is small and, crucially, symbolic: a call site names the target as `com/acme/Service.run:()V` rather than as an address. The JVM resolves those symbolic references lazily at run time, which is what makes dynamic linking, class loading and hot-swapping-style tooling possible at all. A statically linked native binary has no equivalent flexibility. ## What it costs Bytecode is not executable by the CPU. Something must bridge the gap: 1. The **interpreter** reads one instruction at a time and performs its effect. This starts instantly — no compile pause — but every instruction costs decoding and dispatch overhead on top of the work it represents. 2. The **JIT compiler** translates methods that prove to be hot into optimized native code, using profile data gathered while interpreting. So the JVM pays a translation cost that an ahead-of-time-compiled C binary does not. In exchange it gets portability, verification, and profile-guided optimization that can beat static compilation on real workloads because it sees actual behavior (which branches are taken, which receiver types actually appear). ## Common misreadings - "Bytecode is interpreted, so Java is slow." Bytecode is *initially* interpreted; hot code ends up as native machine code with inlining and register allocation. - "Bytecode is machine-independent, so it hides everything about the machine." It abstracts the instruction set, not the machine model — data races, memory visibility and 64-bit value handling are governed by the Java memory model, not erased by bytecode. - "The .class file is obfuscated/compiled beyond recovery." It retains names, descriptors and often line numbers; `javap -c` prints it, and decompilers reconstruct near-source. Bytecode is a *portability* format, not a protection mechanism.

  • If bytecode is portable, why do some libraries still ship platform-specific artifacts?
    Because bytecode only covers code that runs inside the JVM. Anything reaching outside — native libraries loaded through JNI or the FFI API, memory-mapped drivers, OS-specific system calls — is real machine code compiled per platform. Such libraries ship a portable jar plus per-platform native binaries selected at run time.
  • Does bytecode being verified mean it is safe to run untrusted classes?
    Verification only guarantees structural and type safety of the instruction stream: no stack underflow, no treating an int as a reference, no jumps into the middle of an instruction. It says nothing about what the code does with the APIs it is allowed to call — file access, network, reflection. Sandboxing is a separate concern from verification.

Bytecode is like shipping IKEA parts plus an assembly diagram instead of a fully built cabinet: the diagram is the same everywhere, and each local workshop (JVM) assembles it with its own tools.

saying these in an interview costs you the question

  • Saying the JVM executes Java source code, or that .class files contain source
  • Claiming bytecode is machine code for a specific CPU
  • Claiming Java is 'always interpreted' and therefore inherently slow
  • Believing bytecode hides implementation details from decompilation
  • Thinking the JVM knows which source language produced a class

context

open as a page

Compiling a Java application ahead of time into a standalone native executable is an alternative to running it on the JVM with just-in-time compilation. What do you gain and what do you give up?

level: middleimportance: must knowfreq 45%

basics

~20 s

You gain millisecond startup, immediate peak-ish performance, and a much smaller memory footprint with no compiler or class loading at runtime. You give up profile-guided peak throughput, runtime dynamism such as unconstrained reflection, and fast builds.

open as a page

Trace, in order, what a Java Virtual Machine does from the moment running code first references a type until that type's methods are executing as JIT-compiled native code — and name which JVM subsystem owns each step.

level: middleimportance: must knowfreq 52%

basics

~20 s

Order: load, verify, prepare, resolve, initialize, interpret, profile, compile. The class loader subsystem owns the first five (loading through running the static initializer); the execution engine owns the last three (interpreting bytecode, gathering profile counters, and JIT-compiling hot methods into native code).

open as a page

Describe the three top-level subsystems that make up a Java Virtual Machine implementation and what each one is responsible for.

level: middleimportance: must knowfreq 62%

basics

~20 s

Class loader subsystem: finds class files and loads, links (verify, prepare, resolve) and initializes types. Runtime data areas: heap, per-thread stacks and program counters, class-metadata area, code cache. Execution engine: interpreter, JIT compiler and garbage collector, which actually run and maintain the code.

open as a page

Interpreting bytecode instruction by instruction is portable but slow. Mechanically, where does the slowness come from, and why does a runtime that interprets first still keep an interpreter around after it can compile code to native instructions?

level: middleimportance: must knowfreq 55%

basics

~20 s

Each bytecode costs decode plus an indirect dispatch on top of its actual work, values shuttle through the operand stack instead of registers, and nothing is optimized across instruction boundaries. The interpreter is kept because it starts instantly, costs no compile time for cold code, and is the fallback when compiled code must be abandoned.

open as a page

Ahead-of-time compilation of a Java application to a native binary assumes a closed world. What does that assumption mean concretely, and how do you make reflection, dynamic proxies, and resource loading work under it?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Closed world means every class and method reachable at runtime must be known at build time; unreachable code is not in the binary. Anything named by string — reflection, proxies, resources, serialization — is invisible to the analysis and must be declared in build-time metadata, generated by a framework, or recorded by a tracing agent.

open as a page

Java, Kotlin, Scala, Groovy and Clojure all run on the JVM. Concretely, what does it mean for a language implementation to 'target the JVM', and how much does the running JVM know about which source language produced a class?

level: middleimportance: should knowfreq 38%

basics

~20 s

Targeting the JVM means emitting valid classfiles — bytecode plus constant pool and metadata — that pass verification. At run time the JVM sees only classes, methods and descriptors; source-language features that bytecode lacks must be desugared into ordinary code or encoded in attributes.

open as a page

JVM bytecode is a stack-based instruction set: most opcodes carry no operand fields and act on an implicit working stack. Real CPUs — and some other virtual machines, such as Android's Dalvik — use register-based instruction sets whose instructions name their operands explicitly. Why did the JVM's designers pick the stack-based form, and what does that choice cost when the code actually runs?

level: middleimportance: should knowfreq 40%

basics

~20 s

Implicit operands make most opcodes one byte: bytecode stays compact, a compiler emits it as a postfix tree walk with no register allocator, and typed opcodes plus StackMapTable allow single-pass verification. Cost: more instructions and stack traffic, which the JIT removes.

open as a page

When a Java application is compiled ahead of time into a native binary, some class initialization can run at build time and be snapshotted into the image. What does that buy, and what can go wrong?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Running static initializers during the build and storing the resulting objects in the image means the binary starts with that state already built, removing startup work. The hazard is baking in state that must be per-run or per-environment — seeds, hostnames, file handles, timestamps, secrets.

open as a page

Why does a Java application compiled ahead of time into a native binary usually reach lower peak throughput than the same application on a warmed-up JVM, and what narrows that gap?

level: seniorimportance: should knowfreq 30%

basics

~20 s

A build-time compiler has no runtime profile, so it cannot speculate on the branches and receiver types actually taken, and it has no way to recover if a guess is wrong. Profile-guided optimization from an instrumented run recovers much of the gap.

open as a page

In the JVM, the heap and class-metadata area are shared by all threads while each thread gets its own stack and program-counter register. Why is the split drawn there, and what practical consequences follow from it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Method frames and the instruction pointer are private to one execution, so they need no coordination — locals are unsynchronized and cheap. Objects outlive the frame that created them and must be reachable from anywhere, so the heap is shared, which forces garbage collection, safepoints and a memory model. Static state is per-type, hence global.

open as a page

You own a fleet of Java services and are asked whether to move them to ahead-of-time-compiled native binaries. How would you decide, service by service?

level: principalimportance: should knowfreq 28%

basics

~20 s

Decide by process lifetime and constraint. Short-lived, scale-from-zero, high-instance-count, or memory-capped workloads favour native. Long-running throughput services favour the JVM. Then gate on framework support, dynamic-feature usage, and whether the pipeline can test the native artifact.

open as a page

HotSpot does not execute bytecode with a plain C switch statement. Describe how its template interpreter works and what the interpreter does besides producing the instruction's result.

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

At startup HotSpot generates a small machine-code stub for each opcode into a dispatch table, and each stub ends by jumping straight to the next opcode's stub. It also keeps the top of stack in a register and, while executing, records invocation counts, back-edge counts, branch and receiver-type profiles.

open as a page

A team is puzzled that the same compiled application behaves identically everywhere but performs very differently across JVM builds and flags. How would you explain what the Java Virtual Machine specification actually mandates versus what each implementation is free to decide?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

The specification fixes the observable contract: class file format, instruction semantics, verification, class-initialization rules, exception behaviour and the memory model. It leaves the mechanism open — object layout, garbage collection algorithm, whether and how code is compiled, stack representation, internal caching. Portability is of semantics, not of performance.

open as a page