Why does extracting a 12-bit length field from a packed 32-bit header need both a right shift and a mask?
answer
- two neighbours share one word
- the field has an offset and a width
- one operator moves, one operator clears
- AND with as many ones as the field is wide
- shift down by the field's offset first
basics
~20 sThe shift moves the field down to bit 0 so it reads as a number; the mask clears the neighbouring fields still sitting above it. Shift alone leaves garbage on top, mask alone leaves the value scaled.
solid answer
~50 sA packed header stores several fields side by side in one word, so a field has both an offset and a width. The right shift by the offset slides the field's lowest bit down to position 0, which fixes its place value; everything that lived above the field is now sitting in the high bits of the result. The AND with a width-shaped mask of 12 ones then clears those higher bits, because AND keeps a bit only where both operands have a 1. Do only the shift and you get the field plus every field above it; do only the mask and you get the right bits at the wrong place value, i.e. the field scaled by 2 to the offset. Masking first with a *placed* mask and shifting afterwards is equivalent — the shift-then-mask order is just the one where the mask literal reads as the field's width.
code
pseudocode · 10 lines// header layout, high bits first:
// bits 31..28 : version (4 bits)
// bits 27..16 : length (12 bits)
// bits 15..0 : flags
version = (header >> 28) & 0xF
length = (header >> 16) & 0xFFF
...
// shift alone would leave version riding above length
// mask alone would leave length scaled by 65536go deeper
Recall that AND keeps bits, OR merges them, XOR flips them and NOT inverts everything, and that a packed field needs a shift for its position plus a mask for its width. Be ready to write both halves out loud.
Explain why each half alone fails, how (1 << w) - 1 builds a width-shaped mask, and how to write a field back with clear-then-set without disturbing the fields around it.
Show the judgment that keeps the layout maintainable: named shift/mask constants derived from one layout definition, masking values on the way in, and tests where neighbouring fields are non-zero so a missing mask actually fails.
Own the call between hand-rolled shift/mask accessors and a generated or declarative codec: hand-rolling is fast and obvious in a hot path, but every open-coded offset is a place the next wire-format revision can be missed.
### The four operators, one bit position at a time Bitwise operators combine two values position by position; bit 7 of the result depends only on bit 7 of each operand. That independence is the whole reason masks work. - **AND** yields 1 only where *both* operands have 1. Against a mask it means *keep the selected bits, clear everything else*. - **OR** yields 1 where *either* has 1. It merges and never clears. - **XOR** yields 1 where the operands *differ*. It flips exactly the positions the mask marks, and applying the same mask twice restores the original — XOR is its own inverse. - **NOT** flips every bit, which is how a "keep these" mask becomes a "clear these" mask. ### Packing: a field has an offset and a width Say a 32-bit header word holds a 4-bit version in bits 31..28, a 12-bit length in bits 27..16, and smaller flags below. Two numbers describe each field: its **width** (how many bits) and its **offset** (how far up the word its lowest bit sits). Writing the header puts each value in place with a **left shift**: `value << offset`. Left shift by k moves every bit k positions up and feeds zeros in at the bottom, which is exactly multiplication by 2^k as long as nothing meaningful falls off the top. The placed fields are then merged with OR. Because the fields occupy disjoint positions, OR and XOR would give the same answer here, but OR is the operator that states the intent: combine, never flip. ### Unpacking: shift, then mask `length = (header >> 16) & 0xFFF` The right shift by 16 brings the field's lowest bit to position 0, so the digits now carry their true place values — 1, 2, 4, 8 and so on. But the version field that lived above the length rode down with it and is now occupying bits 12..15 of the intermediate result. The mask `0xFFF` — twelve 1 bits, which is `(1 << 12) - 1` — clears them. **The mask's width must match the field's width**: an 8-bit mask silently truncates a 12-bit field, and a 16-bit mask lets four foreign bits through. ### Why neither half works alone - **Shift only.** `header >> 16` is the length *plus* the version scaled up by 4096. Every reading is wrong whenever a neighbour is non-zero, which is exactly the case your unit test with a zeroed header will not catch. - **Mask only.** `header & 0x0FFF0000` selects the correct bits but leaves them where they were. The value you get is the field multiplied by 65536 — right pattern, wrong magnitude. The mask-first ordering, `(header & 0x0FFF0000) >> 16`, is equally correct; it just requires the mask literal to encode the field's position as well as its width, so the constant no longer reads as "twelve bits". ### Writing a field back without disturbing neighbours The symmetric operation is clear-then-set: build the placed mask, invert it to clear the old field, then OR in the new value: `word = (word & ~(0xFFF << 16)) | ((value & 0xFFF) << 16)` Note the mask on the incoming value. If `value` exceeds 12 bits, an unmasked write spills straight into the version field above it — a corruption that shows up far from the line that caused it. ### Things that bite beginners here - **Mask/width mismatch** is the most common defect, and it is asymptomatically invisible: it only misbehaves once the neighbouring field becomes non-zero in production traffic. - **Signed types.** If the field sits at the very top of a signed word, the right shift may drag copies of the sign bit down with it. The mask removes them, which is another reason never to skip it. - **Byte order is a different concern.** How a multi-byte value is laid out in memory is a serialization question; once the bytes are assembled into one integer value, bit offsets within that value behave the same everywhere. - **Named constants beat literals.** `(header >> LENGTH_SHIFT) & LENGTH_MASK` documents the layout once; scattering `>> 16` through a codec guarantees that one of them will not be updated when the wire format changes.
- Does it matter whether you mask first and then shift?Both orders work, but the mask differs. Shift-then-mask uses a width-shaped mask of as many ones as the field is wide; mask-then-shift needs that mask already moved to the field's position. The results are identical, so pick the one whose constant reads as the field's width, since that is the property a reviewer can check against the layout table.
- How do you build the mask for a field of width w?One shifted into position w, minus one: `(1 << w) - 1` gives w low ones. It reads directly as the width, so it stays correct when the layout changes. The one boundary to respect is w equal to the whole integer's width — a shift count that large is not a normal shift, so a full-width mask must be built differently.
- How do you write a new length back without touching the version above it?Clear then set: AND the word with the inverted placed mask to zero the field, mask the incoming value to its width, shift it into position, and OR it in. Masking the incoming value is not optional — an over-range value spills into the neighbouring field and corrupts a header that still parses.
It is like pulling one column out of a long paper receipt: first you slide the paper sideways until that column starts at the edge, then you cut away everything past its width.
saying these in an interview costs you the question
- Shifting is enough; the mask is defensive noise
- Masking alone gives the field's numeric value
- OR and XOR are interchangeable for merging fields
- The mask width need not match the field width
- Writing a field back needs no clearing step first