What surprising results arise from integer promotion and sign extension when applying bitwise and shift operators to byte, short, and char, and how do you handle them safely?
answer
- byte/short/char promote to int first
- byte/short sign-extend; char zero-extends
- 0xFF byte becomes int -1, not 255
- mask with & 0xFF to read unsigned
- Byte.toUnsignedInt for clarity
basics
~20 sbyte, short, and char are widened to int before any bitwise or shift operation. A negative byte gets its high bits filled with 1s (sign extension), so masking with & 0xFF is often needed to get the value you expect.
solid answer
~50 sJava has no bitwise/shift arithmetic on types narrower than int. A `byte`, `short`, or `char` operand is first promoted to `int`. For `byte` and `short` this is *sign-extending*: a byte value like 0x80 (-128) becomes the int 0xFFFFFF80, so reading a 'raw' byte 0xFF gives the int -1, not 255. `char` is unsigned, so it zero-extends. The classic bug is `int hi = b >> 4` on a negative byte, which leaks sign bits; or building an int from bytes with `(b0 << 24) | (b1 << 16) | ...` where each negative byte pollutes higher bytes. The fix is to mask each byte to its low 8 bits: `b & 0xFF` (which promotes, then keeps only the low byte as a non-negative int). Also note the result of any byte/short bitwise op is an int, so assigning back to a byte needs an explicit cast.
code
java · 15 linesbyte b = (byte) 0xFF;
System.out.println((int) b); // -1 (sign-extended)
System.out.println(b & 0xFF); // 255 (masked -> unsigned)
System.out.println(Byte.toUnsignedInt(b)); // 255 (clear intent)
// packing 4 bytes into an int safely
byte b0 = (byte)0x12, b1 = (byte)0xFF, b2 = (byte)0x34, b3 = (byte)0x56;
int packed = ((b0 & 0xFF) << 24)
| ((b1 & 0xFF) << 16)
| ((b2 & 0xFF) << 8)
| (b3 & 0xFF);
// result type is int:
// byte c = b0 & b1; // does NOT compile
byte c = (byte) (b0 & b1); // explicit cast requiredgo deeper
Aware that byte/short are treated as int in arithmetic; may not yet know sign extension details.
Knows to use & 0xFF to read a byte as unsigned and that bitwise results on bytes are ints.
Explains sign vs zero extension, the byte-packing corruption bug, the compound-assignment narrowing cast, and uses Byte.toUnsignedInt.
Designs binary parsing/serialization code that is correct across signed types, codifies masking conventions, and reviews for these subtle promotion bugs.
## The root cause: no sub-int bitwise math The JVM's bitwise and shift operations work on `int` and `long` only. So whenever you apply `& | ^ ~ << >> >>>` to a `byte`, `short`, or `char`, Java first performs **unary numeric promotion**: the operand is widened to `int`. The result is therefore an `int` too. ## Sign extension vs zero extension How the value widens depends on whether the source type is signed: - `byte` and `short` are **signed**. Widening copies the sign bit into all the new high bits — **sign extension**. A `byte` holding `0xFF` is the value -1; promoted to int it becomes `0xFFFFFFFF` = -1, *not* 255. - `char` is **unsigned** (16-bit). Widening fills the new high bits with 0 — **zero extension**. ## Concrete bugs this causes 1. **Reading bytes as unsigned 0..255.** `byte b = (byte) 0xFF; int v = b;` gives -1. To get 255: `int v = b & 0xFF;` — the `& 0xFF` keeps only the low 8 bits after promotion, discarding the sign-extended 1s. 2. **Packing bytes into an int/long.** `int x = (b0 << 24) | (b1 << 16) | (b2 << 8) | b3;` — if any `bi` is negative, its sign bits overwrite the higher bytes via OR. Correct: mask each: `((b0 & 0xFF) << 24) | ((b1 & 0xFF) << 16) | ((b2 & 0xFF) << 8) | (b3 & 0xFF)`. 3. **Shifting a byte right.** `b >> 4` sign-extends first, so a negative byte yields 1s in the high nibble. Use `(b & 0xFF) >> 4` to treat it as unsigned. 4. **Assigning the result back.** Because the result is `int`, `byte c = b1 & b2;` does not compile without `(byte)`. Compound assignment `b1 &= b2;` *does* compile because it has an implicit narrowing cast — which can silently truncate. ## Why `& 0xFF` works `0xFF` is the int `255` = `0000...0011111111`. After the byte is sign-extended to int, ANDing with `0xFF` forces all bits except the low 8 to 0, recovering the original 8-bit pattern as a non-negative int in 0..255. Use `& 0xFFFF` for shorts/chars and `& 0xFFFFFFFFL` to treat an int as unsigned in a long. ## char specifics Since `char` zero-extends, `char` bitwise ops are already 'unsigned' over 0..65535. But mixing `char` and `byte` still requires care because the byte path sign-extends. ## Best practices - Always mask bytes you intend as unsigned: `& 0xFF`. - Prefer `Byte.toUnsignedInt(b)`, `Short.toUnsignedInt(s)`, `Integer.toUnsignedLong(i)` (Java 8+) for readability. - Be wary of compound assignment's hidden narrowing cast.
- Why does (b0 << 24) | b1 corrupt the result when b1 is a negative byte?b1 is sign-extended to a negative int (high 24 bits all 1), and OR sets those high bits, overwriting b0's byte. Masking b1 & 0xFF removes the sign bits first.
- Why does b1 &= b2 compile but byte c = b1 & b2 does not?Compound assignment includes an implicit narrowing cast back to byte; the plain expression yields an int that won't auto-narrow without an explicit (byte) cast.
saying these in an interview costs you the question
- Assuming byte 0xFF reads as 255 without masking
- Packing bytes with << / | but forgetting & 0xFF on each
- Thinking char sign-extends (it zero-extends)
- Believing b1 & b2 returns a byte (it returns int)