skip to content

What precision do you lose choosing float32 over float64 in Go, and when is float32 still worth it?

level: middleimportance: should knowfreq 46%

answer

  1. twenty-four bits against fifty-three
  2. about seven decimal digits
  3. Go never widens a float for you
  4. the math package speaks one width only
  5. memory bandwidth is the reason to pick it

basics

~20 s

A float32 carries a 24-bit significand, about seven decimal digits, against float64's 53 bits and about fifteen. Go never converts between the two implicitly, and the math package is float64-only, so float32 pays off only when memory or bandwidth dominates.

solid answer

~50 s

`float32` gives you a 24-bit significand — about seven decimal digits — and a maximum around `math.MaxFloat32`, roughly 3.4e38; `float64` gives 53 bits, about fifteen to seventeen digits, and a range to roughly 1.8e308. Converting a float64 to float32 rounds, so `float64(float32(x)) == x` is usually false, and the error compounds: summing thousands of ratios in float32 drifts long before float64 does, because each add re-rounds to 24 bits and a large running total starts absorbing small addends entirely. Go makes the cost visible — there is no implicit numeric conversion, so mixing widths is a compile error until you write `float64(x)`, and the `math` package is float64-only, so float32 code round-trips through float64 for anything beyond `+ - * /`. Choose float32 only when you hold millions of values and memory or wire size is the constraint; use float64 for anything you accumulate or compare.

code

go · 4 lines
go
var ratio float64 = 0.1
f := float32(ratio)
fmt.Println(float64(f) == ratio, math.Float32bits(f))
// prints: false 1036831949

go deeper

for a junior

Know that float64 is the default working width in Go, that float32 holds far fewer digits, and that converting between them must be written out explicitly as float64(x) or float32(x).

for a middle

Explain the numbers: 24-bit versus 53-bit significands, roughly seven versus fifteen decimal digits, rounding on every conversion and every arithmetic step, and the fact that the math package offers float64 signatures only.

for a senior

Show the judgment: accumulate in float64 even when you store in float32, convert once at the boundary, and reach for exact integers when a total has to reconcile with another system rather than merely look right.

for a principal

Be able to defend the storage-versus-compute split across a system — where the width is fixed by a wire format or a dataset, what it costs to change later, and how you keep the accumulate-in-float64 rule from eroding as code spreads.

### What the two widths actually hold Both of Go's float types are IEEE 754 binary formats: | | sign | exponent | stored fraction | significand | decimal digits | largest finite | |---|---|---|---|---|---|---| | `float32` | 1 | 8 | 23 | 24 bits | ~7 | `math.MaxFloat32` ≈ 3.4e38 | | `float64` | 1 | 11 | 52 | 53 bits | ~15–17 | `math.MaxFloat64` ≈ 1.8e308 | The significand is what matters for accuracy. Twenty-four bits means a float32 can distinguish about one part in sixteen million; a running total of 100,000 and a change of 0.001 cannot both be represented, so the change is simply lost. Exponent range is rarely the reason to pick float64 — precision is. ### Conversion is lossy and explicit Go has no implicit numeric conversion at all. `a + b` with `a float32` and `b float64` does not compile: `invalid operation: a + b (mismatched types float32 and float64)`. You must write `float64(a) + b`, which makes every widening and narrowing a visible act in the source rather than a silent promotion the reader has to infer. That is a deliberate design choice and it is the main reason precision bugs in Go tend to be locatable. Narrowing rounds to the nearest representable float32, so a round trip usually does not return the original: `float64(float32(0.1)) != 0.1`. Print `math.Float32bits` and you can see exactly which pattern you landed on. Two edge cases are worth knowing: a *constant* that does not fit is a compile error (`constant 1e39 overflows float32`), while a *non-constant* conversion of a too-large float64 does not fail at run time — in practice you get an infinity — so a defensive `math.IsInf(f32Result, 0)` is the check, not an error return. ### The math package is float64-only `math.Sqrt`, `math.Log`, `math.Pow` and the rest all take and return `float64`. There are no float32 variants, so float32 code that needs any of them converts up, computes, and converts back down — two extra roundings per call, plus code noise. In the same spirit, `strconv.ParseFloat(s, 32)` and `strconv.FormatFloat(f, 'g', -1, 32)` return and consume `float64` values while promising the result round-trips as a float32; the `bitSize` argument is how you tell the conversion which width you actually mean. ### How error accumulates Every arithmetic operation rounds its result to the nearest representable value of the type. Two effects follow: - **Growth.** Adding n values accumulates up to n roundings. In float64 the per-step error is around 1e-16 relative; in float32 it is around 1e-7. Sum ten thousand percentages and the float32 total can be visibly wrong in the second decimal place while the float64 total is still exact to more digits than you will ever display. - **Absorption.** Once the running total is much larger than the next addend, the add rounds back to the total and the addend vanishes. The threshold arrives roughly 2^29 times sooner in float32 than in float64. If a total has to be exact — money, counts, anything reconciled against another system — neither width is the answer. Use integers in minor units and keep floats for derived ratios only. ### When float32 is the right call It halves memory and halves the bytes you move, and for large homogeneous arrays that is often the entire performance story: cache lines hold twice as many values, and the same is true of network payloads and files. Signal samples, pixel and geometry data, model weights and large telemetry buffers are all legitimate float32 territory, because the inputs themselves carry only a few significant digits and the values are consumed rather than accumulated. The rule of thumb: **store in float32 if the volume demands it, compute and accumulate in float64.** Convert at the boundary, once, explicitly — which is the only way Go lets you do it anyway.

  • Why does summing ten thousand small ratios drift further in float32 than in float64?
    Every add re-rounds to the type's significand, so error grows with the number of operations, and once the running total dwarfs the next addend the add rounds straight back to the total and the addend disappears. With 24 bits that absorption point arrives early; with 53 bits it is far out of reach for realistic sums. Accumulate in float64, or use exact integers.
  • Can you add a float32 and a float64 directly in Go?
    No. Go has no implicit numeric conversion, so the expression is a compile error until you write `float64(a) + b` or `a + float32(b)`. The upside is that every width change is visible at the point it happens, which makes precision loss reviewable instead of invisible.
  • What happens when you convert a float64 larger than math.MaxFloat32 to float32?
    If the value is a constant, the compiler rejects it outright — `constant 1e39 overflows float32`. If it is a variable, the conversion does not fail at run time; in practice you get an infinity, so the way to detect it is `math.IsInf(v, 0)` on the converted value rather than an error check.

saying these in an interview costs you the question

  • Assumes Go promotes float32 to float64 automatically
  • Says float32 is precise to about fifteen decimal digits
  • Claims float32 is fine for money because the amounts are small
  • Expects math.Sqrt to accept a float32 argument
  • Picks float32 for speed without any memory-bound measurement