What is the difference between sha256.Sum256 and sha256.New, and when do you use each?
answer
- one call versus a running hash
- how do the bytes arrive?
- one returns an array, one a slice
- you cannot io.Copy into a value
- Sum(nil) finishes the streaming form
basics
~20 ssha256.Sum256 hashes a byte slice you already hold and returns a [32]byte array. sha256.New returns a streaming hash.Hash you write data into and finish with Sum(nil) - use that when the input is a file or stream too large to buffer.
solid answer
~40 s`sha256.Sum256(b)` is the one-shot form: you hand it every byte at once and it returns a `[32]byte` **array**, so you usually write `sum[:]` to get a `[]byte` for `hex.EncodeToString` or for storage. `sha256.New()` returns a `hash.Hash`, which is an `io.Writer` with `Sum`, `Reset`, `Size` and `BlockSize`. You feed it incrementally - typically `io.Copy(h, f)` for a file - and call `h.Sum(nil)` at the end to get the same 32-byte digest as a `[]byte`. The streaming form keeps memory constant regardless of input size, which is why a tool that walks a directory hashing files uses it rather than `os.ReadFile` plus `Sum256`. Both compute identical digests for identical bytes; the choice is purely about how the bytes arrive.
code
go · 2 linessum := sha256.Sum256([]byte("hello")) // sum is [32]byte
hexDigest := hex.EncodeToString(sum[:])go deeper
Be able to write both forms from memory: sha256.Sum256(b) for data in hand, sha256.New() plus io.Copy for a file. Remember the one-shot returns a [32]byte array and needs sum[:] before hex encoding.
Explain why the streaming form exists: the hash carries a small fixed running state, so memory does not grow with the input. Be ready to name the hash.Hash methods and say that it is an io.Writer.
Show where the choice bites in production: whole-file reads that blow up under concurrency, and the fact that a hash.Hash value is not safe for concurrent use. Say how you would hash while simultaneously copying, using io.MultiWriter.
Own the consequence of a digest becoming an identifier other systems store. Its length, its encoding and whether the algorithm can ever change are decisions you cannot quietly reverse once other teams persist those strings.
## Two entry points, one algorithm The `crypto/sha256` package exposes SHA-256 twice, and the difference is about **plumbing**, not cryptography. Fed the same bytes, both produce the same 32 bytes. ### The one-shot: `sha256.Sum256` ```go sum := sha256.Sum256([]byte("hello")) ``` The signature is `func Sum256(data []byte) [Size]byte`, where `sha256.Size` is 32. Two things trip people up: 1. **It returns an array, not a slice.** `[32]byte` is a value type: assigning it copies all 32 bytes, and you cannot pass it where a `[]byte` is wanted. That is why you constantly see `sum[:]` - slicing the array yields a `[]byte` backed by it, which is what `hex.EncodeToString`, `base64`, a database driver or a `Write` call expects. 2. **Every byte must already be in memory.** Reading a multi-gigabyte file with `os.ReadFile` just to hash it works on your laptop and falls over when a handful of goroutines do it at once. The array return is actually a small gift: because it is a value, it is comparable with `==`, usable as a map key, and it needs no allocation. ### The streaming form: `sha256.New` ```go h := sha256.New() if _, err := io.Copy(h, f); err != nil { return err } digest := h.Sum(nil) ``` `sha256.New()` returns a `hash.Hash`, the standard-library interface every hash in Go implements: ```go type Hash interface { io.Writer Sum(b []byte) []byte Reset() Size() int BlockSize() int } ``` Because it embeds `io.Writer`, a hash is a legal destination for `io.Copy`, `fmt.Fprintf`, `io.MultiWriter` (hash while you also copy somewhere else) and `io.TeeReader`. That composability is the real reason the interface exists. The hash keeps a fixed-size running state - a few dozen bytes - and consumes the input in 64-byte blocks (`sha256.BlockSize`). Memory does not grow with the input. `Sum(nil)` returns the digest as a `[]byte` of length `Size()`, here 32. ### Which to reach for - The data is a small in-memory `[]byte` or `string` you already have: **`Sum256`**. It is one line and allocates nothing. - The data comes from an `io.Reader` - a file, a request body, a network stream, a tar entry - or is assembled from several pieces: **`New` plus `io.Copy` or repeated `Write`**. A directory indexer that names each blob by its digest is the canonical streaming case: you do not know how large the next file is, and you must not find out the hard way. ### Things worth knowing early - A `hash.Hash` value is **not safe for concurrent use**. One per goroutine, or one per file; never a package-level shared instance. - `Write` on a `hash.Hash` is documented never to return an error, so the only error `io.Copy(h, f)` can produce is the **reader's**. Handle it - dropping it leaves you a well-formed digest of a prefix. - `hash/fnv` and `hash/crc32` also implement `hash.Hash` (via `hash.Hash64` and `hash.Hash32`). The interface will happily let you substitute one for the other, and the compiler will not object. They are checksums for detecting accidental corruption and for keying in-memory tables - not for naming content that anyone else can influence. - Digest bytes are raw binary. Anywhere they meet a log line, a filename or JSON, encode them - `hex.EncodeToString` for the familiar 64-character form. ### The mental model Think of `Sum256` as `strings.ToUpper`: a pure function over a value you hold. Think of `New` as opening a pipe you push bytes through, then asking it what it saw. Same answer either way; the second one never has to hold the whole message.
- Why do you so often see sum[:] right after a call to sha256.Sum256?Because `Sum256` returns `[32]byte`, an array value, and almost everything downstream wants a `[]byte`: `hex.EncodeToString`, `base64` encoders, `io.Writer.Write`, database drivers. Slicing the array with `sum[:]` produces a slice backed by it, with no copy. Passing the array itself simply will not compile where a slice is required.
- Someone swaps sha256.New for fnv.New64a in a content-addressed indexer to make it faster. Does it compile, and what has been given up?It compiles: `fnv.New64a()` returns a `hash.Hash64`, which embeds `hash.Hash`, so it fits anywhere the interface is expected. What is gone is any resistance to deliberately constructed collisions - `hash/fnv` and `hash/crc32` exist to catch accidental corruption and to key in-memory tables. The digest also shrinks to 8 bytes, so even accidental collisions become plausible at scale.
- Is it worth reusing one hash.Hash across files instead of calling sha256.New per file?Rarely. `sha256.New()` is a small allocation; the hashing itself dominates. If you do reuse one, you must call `Reset()` between inputs or the second digest covers both files, and you must keep it to a single goroutine - a `hash.Hash` is not safe for concurrent use. The reset bug is far more expensive than the allocation.
Sum256 is weighing a parcel you are already holding. sha256.New is a turnstile: people walk through one at a time and it tells you the count at the end, never needing room for the whole crowd.
saying these in an interview costs you the question
- Thinks sha256.Sum256 returns a []byte
- Reads a whole multi-gigabyte file into memory to hash it
- Hashes each chunk separately and concatenates or XORs the digests
- Passes the [32]byte array where a []byte is required
- Treats hash/crc32 and crypto/sha256 as interchangeable because both fit hash.Hash
- Shares one hash.Hash across goroutines