How do you reuse a MessageDigest instance, and what do reset(), update(), and the two digest() overloads do to its internal state?
answer
- update = accumulate, digest = finalize + auto-reset
- digest(input) == update(input); digest()
- reset() = abandon without a result
- Many updates == one update of the concatenation
- Not thread-safe → per-thread or per-call instance
basics
~20 sA MessageDigest holds running state. update() adds bytes to it; digest() finishes the hash and then auto-resets, so you can hash another message right away. reset() throws away any buffered bytes without producing a hash.
solid answer
~40 sMessageDigest is a stateful engine. update(byte[]) accumulates input into an internal buffer/state, and you may call it many times to feed data incrementally. digest() finalizes the computation, returns the byte[], and then automatically resets the engine to its initial state — so the very next update/digest starts a brand-new message with no leftover bytes. The overload digest(byte[] input) is just update(input) followed by digest(). If you started feeding data with update but want to abandon it without finishing, call reset() to clear the state. Because all this state is mutable and unsynchronized, MessageDigest is not thread-safe: never share one instance across threads — use a fresh instance per computation, a ThreadLocal, or pool one per thread. The auto-reset on digest() is what makes single-threaded reuse safe and cheap, avoiding a getInstance call per message.
go deeper
Knows update feeds bytes and digest produces the hash, and that you can hash a new message after digest.
Explains the auto-reset on digest(), reset() semantics, the digest(input) convenience overload, and incremental update for streaming.
Adds thread-safety reasoning (per-thread/per-call instances) and the equivalence of chunked updates to a single update.
Considers reuse strategies under load (ThreadLocal vs fresh instance vs clone for shared-prefix hashing) and codifies thread-safety conventions for shared utility code.
## MessageDigest is a stateful machine Unlike a pure function, a `MessageDigest` object carries **mutable internal state** — a running computation plus any not-yet-processed input. Understanding the three operations that touch that state is the whole topic. ## update(...) — accumulate input ```java md.update(chunk1); md.update(chunk2); ``` `update` folds the given bytes into the running hash state. The key property is that **many small updates produce the same result as one big update of the concatenation**: `update(a); update(b); digest()` equals `digest(a ++ b)`. That is exactly what lets you hash data you cannot (or don't want to) hold entirely in memory — a large file read block by block, or a network stream. Overloads accept a single byte, a `byte[]`, a slice `(byte[], offset, len)`, or a `ByteBuffer`. ## digest() — finalize and auto-reset ```java byte[] h = md.digest(); ``` `digest()` applies the final padding/length step, returns the fixed-length result, **and then resets the engine to its initial state**. This auto-reset is the single most important fact for reuse: after `digest()`, the instance is *as good as new*, so the next `update`/`digest` cleanly begins a new message. You do **not** need to call `getInstance` again, and you must **not** assume leftover bytes carry over. There are two finalizing overloads: - `digest()` — finalize whatever was fed via prior `update` calls. - `digest(byte[] input)` — convenience equal to `update(input); digest();` for the one-shot case. - `digest(byte[] buf, int offset, int len)` — write the result into a caller-provided buffer (returns the number of bytes written). ## reset() — abandon without finishing ```java md.reset(); ``` `reset()` discards any buffered input and running state, returning the engine to initial state **without** producing a digest. You use it when you've started `update`-ing a message and decide to throw it away (e.g. an error mid-stream) and start over. Since `digest()` already resets, you rarely need `reset()` in the happy path. ## Reuse pattern ```java MessageDigest md = MessageDigest.getInstance("SHA-256"); for (String msg : messages) { byte[] h = md.digest(msg.getBytes(StandardCharsets.UTF_8)); // auto-resets each time store(h); } ``` This is correct *only because* `digest()` auto-resets. Reusing one instance avoids the (small) cost of repeated `getInstance` provider lookups in a single-threaded loop. ## Thread-safety — the big caveat All that mutable state is **unsynchronized**, so `MessageDigest` is **not thread-safe**. If two threads call `update`/`digest` on the same instance, their bytes interleave and corrupt both results. Options: - **Fresh instance per call** — simplest and usually fast enough; `getInstance` is cheap relative to most workloads. - **`ThreadLocal<MessageDigest>`** — one instance per thread, reused across calls; remember it's reset by `digest()`. - A small **per-thread pool**. Never cache a single shared static `MessageDigest`. ## Bonus: clone() Some providers' digests support `clone()`, letting you snapshot the running state (e.g. hash a common prefix once, then clone to finish several different suffixes). It throws `CloneNotSupportedException` if the provider doesn't support it — so guard for it.
- If you call update three times then digest once, how does the result relate to hashing the concatenated bytes?They are identical. update is incremental: update(a); update(b); update(c); digest() produces the same digest as digest(a ++ b ++ c). This is what enables streaming a large input in chunks.
- What's the difference between reset() and the auto-reset that digest() performs?Both return the engine to its initial state, but digest() first finalizes and returns the hash, whereas reset() discards the buffered input without producing any digest. Use reset() to abandon a partially-fed message.
saying these in an interview costs you the question
- Calling getInstance again for every message because you think digest() doesn't reset
- Sharing one MessageDigest across threads
- Assuming leftover bytes survive a digest() call
- Thinking reset() returns or produces the hash