skip to content

What is the difference between sha256.Sum256 and sha256.New, and when do you use each?

level: juniorimportance: must knowfreq 52%

answer

  1. one call versus a running hash
  2. how do the bytes arrive?
  3. one returns an array, one a slice
  4. you cannot io.Copy into a value
  5. Sum(nil) finishes the streaming form

basics

~20 s

sha256.Sum256 hashes a byte slice you already hold and returns a [32]byte array. sha256.New returns a streaming hash.Hash you write data into and finish with Sum(nil) - use that when the input is a file or stream too large to buffer.

solid answer

~40 s

`sha256.Sum256(b)` is the one-shot form: you hand it every byte at once and it returns a `[32]byte` **array**, so you usually write `sum[:]` to get a `[]byte` for `hex.EncodeToString` or for storage. `sha256.New()` returns a `hash.Hash`, which is an `io.Writer` with `Sum`, `Reset`, `Size` and `BlockSize`. You feed it incrementally - typically `io.Copy(h, f)` for a file - and call `h.Sum(nil)` at the end to get the same 32-byte digest as a `[]byte`. The streaming form keeps memory constant regardless of input size, which is why a tool that walks a directory hashing files uses it rather than `os.ReadFile` plus `Sum256`. Both compute identical digests for identical bytes; the choice is purely about how the bytes arrive.

code

go · 2 lines
go
sum := sha256.Sum256([]byte("hello")) // sum is [32]byte
hexDigest := hex.EncodeToString(sum[:])

go deeper

for a junior

Be able to write both forms from memory: sha256.Sum256(b) for data in hand, sha256.New() plus io.Copy for a file. Remember the one-shot returns a [32]byte array and needs sum[:] before hex encoding.

for a middle

Explain why the streaming form exists: the hash carries a small fixed running state, so memory does not grow with the input. Be ready to name the hash.Hash methods and say that it is an io.Writer.

for a senior

Show where the choice bites in production: whole-file reads that blow up under concurrency, and the fact that a hash.Hash value is not safe for concurrent use. Say how you would hash while simultaneously copying, using io.MultiWriter.

for a principal

Own the consequence of a digest becoming an identifier other systems store. Its length, its encoding and whether the algorithm can ever change are decisions you cannot quietly reverse once other teams persist those strings.

## Two entry points, one algorithm The `crypto/sha256` package exposes SHA-256 twice, and the difference is about **plumbing**, not cryptography. Fed the same bytes, both produce the same 32 bytes. ### The one-shot: `sha256.Sum256` ```go sum := sha256.Sum256([]byte("hello")) ``` The signature is `func Sum256(data []byte) [Size]byte`, where `sha256.Size` is 32. Two things trip people up: 1. **It returns an array, not a slice.** `[32]byte` is a value type: assigning it copies all 32 bytes, and you cannot pass it where a `[]byte` is wanted. That is why you constantly see `sum[:]` - slicing the array yields a `[]byte` backed by it, which is what `hex.EncodeToString`, `base64`, a database driver or a `Write` call expects. 2. **Every byte must already be in memory.** Reading a multi-gigabyte file with `os.ReadFile` just to hash it works on your laptop and falls over when a handful of goroutines do it at once. The array return is actually a small gift: because it is a value, it is comparable with `==`, usable as a map key, and it needs no allocation. ### The streaming form: `sha256.New` ```go h := sha256.New() if _, err := io.Copy(h, f); err != nil { return err } digest := h.Sum(nil) ``` `sha256.New()` returns a `hash.Hash`, the standard-library interface every hash in Go implements: ```go type Hash interface { io.Writer Sum(b []byte) []byte Reset() Size() int BlockSize() int } ``` Because it embeds `io.Writer`, a hash is a legal destination for `io.Copy`, `fmt.Fprintf`, `io.MultiWriter` (hash while you also copy somewhere else) and `io.TeeReader`. That composability is the real reason the interface exists. The hash keeps a fixed-size running state - a few dozen bytes - and consumes the input in 64-byte blocks (`sha256.BlockSize`). Memory does not grow with the input. `Sum(nil)` returns the digest as a `[]byte` of length `Size()`, here 32. ### Which to reach for - The data is a small in-memory `[]byte` or `string` you already have: **`Sum256`**. It is one line and allocates nothing. - The data comes from an `io.Reader` - a file, a request body, a network stream, a tar entry - or is assembled from several pieces: **`New` plus `io.Copy` or repeated `Write`**. A directory indexer that names each blob by its digest is the canonical streaming case: you do not know how large the next file is, and you must not find out the hard way. ### Things worth knowing early - A `hash.Hash` value is **not safe for concurrent use**. One per goroutine, or one per file; never a package-level shared instance. - `Write` on a `hash.Hash` is documented never to return an error, so the only error `io.Copy(h, f)` can produce is the **reader's**. Handle it - dropping it leaves you a well-formed digest of a prefix. - `hash/fnv` and `hash/crc32` also implement `hash.Hash` (via `hash.Hash64` and `hash.Hash32`). The interface will happily let you substitute one for the other, and the compiler will not object. They are checksums for detecting accidental corruption and for keying in-memory tables - not for naming content that anyone else can influence. - Digest bytes are raw binary. Anywhere they meet a log line, a filename or JSON, encode them - `hex.EncodeToString` for the familiar 64-character form. ### The mental model Think of `Sum256` as `strings.ToUpper`: a pure function over a value you hold. Think of `New` as opening a pipe you push bytes through, then asking it what it saw. Same answer either way; the second one never has to hold the whole message.

  • Why do you so often see sum[:] right after a call to sha256.Sum256?
    Because `Sum256` returns `[32]byte`, an array value, and almost everything downstream wants a `[]byte`: `hex.EncodeToString`, `base64` encoders, `io.Writer.Write`, database drivers. Slicing the array with `sum[:]` produces a slice backed by it, with no copy. Passing the array itself simply will not compile where a slice is required.
  • Someone swaps sha256.New for fnv.New64a in a content-addressed indexer to make it faster. Does it compile, and what has been given up?
    It compiles: `fnv.New64a()` returns a `hash.Hash64`, which embeds `hash.Hash`, so it fits anywhere the interface is expected. What is gone is any resistance to deliberately constructed collisions - `hash/fnv` and `hash/crc32` exist to catch accidental corruption and to key in-memory tables. The digest also shrinks to 8 bytes, so even accidental collisions become plausible at scale.
  • Is it worth reusing one hash.Hash across files instead of calling sha256.New per file?
    Rarely. `sha256.New()` is a small allocation; the hashing itself dominates. If you do reuse one, you must call `Reset()` between inputs or the second digest covers both files, and you must keep it to a single goroutine - a `hash.Hash` is not safe for concurrent use. The reset bug is far more expensive than the allocation.

Sum256 is weighing a parcel you are already holding. sha256.New is a turnstile: people walk through one at a time and it tells you the count at the end, never needing room for the whole crowd.

saying these in an interview costs you the question

  • Thinks sha256.Sum256 returns a []byte
  • Reads a whole multi-gigabyte file into memory to hash it
  • Hashes each chunk separately and concatenates or XORs the digests
  • Passes the [32]byte array where a []byte is required
  • Treats hash/crc32 and crypto/sha256 as interchangeable because both fit hash.Hash
  • Shares one hash.Hash across goroutines