What is sign extension, and what goes wrong when a negative 16-bit value is widened to 32 bits by zero-filling?
answer
- widening must fill the new high bits with something
- what did the old top bit's weight mean?
- negative values need ones, not zeros, up top
- run of ones plus new sign bit telescopes to -2^15
- unsigned reading minus 2^n recovers the signed value
basics
~20 sSign extension copies the sign bit into every new high-order bit when widening, which preserves the two's-complement value. Zero-filling instead reinterprets a negative as a large positive: the 16-bit value -200 arrives in 32 bits as 65336.
solid answer
~40 sWidening a two's-complement value to more bits must fill the new high-order bits with **copies of the sign bit**: zeros for a non-negative value, ones for a negative one. That replication preserves the value because the old top bit's negative weight is exactly reproduced by the new sign bit plus the run of ones beneath it. Zero-filling is only correct for unsigned values; applied to a negative it produces the unsigned reading of the old pattern — 16-bit `0xFF38` (-200) becomes `0x0000FF38` = 65336. The bug classically appears when assembling signed samples from raw payload bytes, because bytes come off the wire as 0..255 and the assembled number never carries its sign into the wider variable. The manual fix: after assembling an n-bit value u, if u >= 2^(n-1), subtract 2^n.
code
pseudocode · 13 lines// two raw bytes off the wire, each read as 0..255
high = payload[k]
low = payload[k + 1]
// assembled into a 32-bit signed variable
sample = high * 256 + low
// wire bytes 0xFF, 0x38 encode -200,
// but sample now holds 65336 — no sign was extended
// fix: interpret the top bit of the 16-bit pattern
if sample >= 32768:
sample = sample - 65536go deeper
Be ready to define sign extension in one sentence — copy the sign bit into the new high bits — and to show 0xFF38 becoming 0xFFFFFF38 rather than 0x0000FF38.
An interviewer expects the why: the telescoping argument for why a run of ones reproduces the old sign bit's weight, the u >= 2^(n-1) subtraction fix, and where the bug hides in byte-assembly code.
Demonstrate diagnosis from symptoms: negatives arriving as values just under 2^n is the fingerprint of a missing sign extension at a serialization boundary; be ready to say where in the pipeline you'd assert ranges to catch it.
Own the boundary policy: binary format specs must state width and signedness for every field, and decoding layers should normalize to a single wide signed representation at the edge so the rest of the codebase never re-litigates fill rules.
## Widening and the two fill rules **Widening** stores a narrow integer in a wider one — 16 bits into 32, 8 into 64. The low bits copy across unchanged; the question is what fills the new high-order bits. - **Zero extension**: fill with zeros. Correct for **unsigned** values, whose value is just the sum of positive place weights — leading zeros change nothing. - **Sign extension**: fill with **copies of the sign bit**. Correct for **signed** two's-complement values. For a non-negative signed value the two rules coincide (the sign bit is 0). They diverge exactly when the value is negative. ## Why replicating the sign bit preserves the value In 16 bits, the top bit carries weight -2^15. After widening to 32 bits, that position's weight becomes an ordinary +2^15, and the new top bit carries -2^31. Fill bits 16..31 with ones (for a negative value) and their combined contribution is: ``` -2^31 + (2^30 + 2^29 + ... + 2^16 + 2^15) = -2^15 ``` — exactly the weight the old sign bit used to carry. The geometric run of ones plus the new negative top weight telescopes back to the original contribution, so the value is unchanged. This is also why a negative number's wide form is a long run of leading ones: -200 is `0xFF38` in 16 bits and `0xFFFFFF38` in 32. ## The classic bug: assembling samples from raw bytes A firmware payload delivers a signed 16-bit temperature sample as two bytes. Each byte reads as 0..255, and the assembly looks innocent: ``` high = payload[k] // 0..255 low = payload[k + 1] // 0..255 sample = high * 256 + low // held in a 32-bit variable ``` For a reading of -20.0 degrees (stored as -200 in tenths), the wire bytes are `0xFF` and `0x38`. The arithmetic yields 255*256 + 56 = **65336** — a huge positive — because nothing ever told the 32-bit variable that bit 15 of the assembled pattern was a *sign* bit. Every reading below zero arrives as a value slightly under 65536, which is precisely the symptom to recognize: negatives showing up as `65536 + value`. ## The fix, portably After assembling an n-bit signed value into an unsigned reading `u`: ``` if u >= 2^(n-1): value = u - 2^n else: value = u ``` For n = 16: 65336 >= 32768, so value = 65336 - 65536 = **-200**. This subtraction is the arithmetic form of sign extension — it converts the unsigned reading of a pattern into its two's-complement reading. Equivalently, one can replicate bit 15 into bits 16..31 with masks. Statically typed languages sign-extend automatically when widening a *signed* type — a 16-bit signed value assigned into a 32-bit one in Go or Java carries its sign — which is exactly why the bug hides in *byte-level* code: the bytes are small non-negative numbers, so the type system never knows a sign bit exists. Environments where wire bytes surface as plain numbers, such as decoding binary payloads in Python or JavaScript, must always apply the subtraction manually. ## The mirror image: narrowing Going the other way — 32 bits into 16 — is **truncation**: the high bits are dropped. That is not a fill-rule question but a range question (the value may simply not fit), and its failure modes belong with overflow. The asymmetry is worth stating in an interview: widening is always value-preserving *if you pick the right fill*, narrowing is only safe when the value already fits the narrow range.
- Why is zero-filling the correct rule when widening an unsigned value?An unsigned value is a pure sum of positive place weights, so prepending zeros adds nothing and the value is unchanged. There is no negatively weighted bit to reproduce. The fill rule is determined entirely by the signedness of the *source* interpretation, which is why the same bit pattern can demand different fills.
- You only have unsigned byte values available — how do you sign-extend an assembled 16-bit sample by hand?Assemble the unsigned reading u = high*256 + low, then test the sign bit: if u >= 32768, the pattern is negative, so subtract 65536. That subtraction converts the unsigned reading into the two's-complement reading, because the two interpretations of any pattern differ by exactly 2^n when the top bit is set.
- Does narrowing from 32 bits to 16 have the same problem?No — narrowing truncates the high bits rather than choosing a fill, so there is no extension rule to get wrong. Its hazard is different: the value may not fit the narrow range at all, and the kept low bits then encode a different number. Widening is lossless with the right fill; narrowing is only lossless when the value already fits.
Writing -5 as a five-digit number requires repeating the minus context, not padding with zeros: "-00005" keeps the sign in force, while blindly writing "00-5" style digits loses it. Sign extension is the binary form of carrying the minus sign across the new columns.
saying these in an interview costs you the question
- Believes copying the low bits across is enough and widening never changes meaning
- Zero-fills everything and treats the resulting huge positives as sensor glitches
- Cannot say why replicating the sign bit preserves the value
- Applies sign extension when widening unsigned values too
- Confuses widening's fill problem with narrowing's range problem