Why does the single byte 0xC8 decode as 200 when read as unsigned but as -56 when read as signed two's complement?
answer
- the byte never changes — the weighting does
- sum the place values twice, two ways
- top bit: +128 in one reading, -128 in the other
- the two readings differ by exactly 2^n
- u >= 128 means signed value is u - 256
basics
~20 sThe bits never change — the interpretation does. Unsigned reading weights the top bit +128, so 11001000 is 200. Two's complement weights the same bit -128, giving -128 + 64 + 8 = -56. Signedness is a property of the reading, not the byte.
solid answer
~50 sA byte is just eight bits; a number only appears once you choose a weighting. `0xC8` is `11001000`. Read unsigned, the top bit is worth +128, so the value is 128 + 64 + 8 = 200. Read as two's complement, the top bit is worth **-128**, so the value is -128 + 64 + 8 = -56. Note the two readings differ by exactly 256 = 2^8 — that is always true when the top bit is set, which gives the quick conversion rule: if the unsigned reading u >= 128, the signed reading is u - 256. Crucially, the top bit is not a detachable sign *flag* over an unchanged magnitude — that is sign-magnitude, a different encoding under which this byte would mean -72. In two's complement the top bit participates in the arithmetic with negative weight, so a format spec that omits a field's signedness leaves the decoder guessing between two legitimate numbers.
go deeper
Be ready to decode one concrete byte both ways by summing place values, and to state clearly that signedness is chosen by the reader, not stored in the bits.
An interviewer expects you to contrast two's complement with sign-magnitude on the same byte, produce the u - 2^n rule, and explain why 11111111 is -1 and not -127.
Show boundary judgment: recognize a signed/unsigned reinterpretation from symptoms like a negative version number, and locate the fix in the format contract and decode layer rather than in downstream patches.
Own the contract question: every field in a wire or file format needs explicit width and signedness, and cross-team decoding disputes should be settled by tightening the spec, not by matching one implementation's habit.
## Bits are not numbers Storage holds patterns; numbers arise from an agreed weighting of those patterns. The single byte `0xC8` — `11001000` — is a perfectly ordinary pattern with at least three historically real numeric readings: | Encoding | How it reads 11001000 | Value | |----------|----------------------|-------| | Unsigned | 128 + 64 + 8 | **200** | | Two's complement | -128 + 64 + 8 | **-56** | | Sign-magnitude | flag set, magnitude 1001000 = 72 | **-72** | | One's complement | flip to 00110111 = 55, negate | **-55** | Only the first two survive in modern practice, but the table makes the point an interviewer is probing: *the same bits, four defensible numbers*. A binary file format that documents a version field as "one byte" without saying *signed or unsigned* has not defined the field. ## Why the top bit is not a sign flag The most common misconception is sign-magnitude thinking: "the first bit says minus, the rest is the number." Under that model `11001000` would be -(1001000) = -72 — and the wrong answer is *confidently* wrong because the model feels intuitive. Two's complement instead makes the top bit a full participant in the sum with weight **-2^(n-1)**. Consequences that follow only from the weighted model: - `11111111` is -1 (not "-127", which the flag model suggests): -128 + 127 = -1. - Negative numbers *grow toward* the all-ones pattern as they approach -1, rather than mirroring the positives. - Ordinary binary addition works across the sign boundary with no special cases, because the encoding is arithmetic modulo 2^n. ## The 2^n offset rule When the top bit is clear, the signed and unsigned readings agree. When it is set, they differ by exactly 2^n: ``` signed(u) = u if u < 2^(n-1) signed(u) = u - 2^n if u >= 2^(n-1) ``` For the byte: 200 >= 128, so signed reading = 200 - 256 = -56. This rule is the workhorse for hand-decoding hex dumps and for converting in environments that only hand you non-negative byte values. It is also sign extension in arithmetic clothing — the same subtraction that fixes a widened negative. ## Where the ambiguity bites Reinterpretation bugs are boundary bugs: a header byte, a checksum field, an offset in a binary protocol. One team writes 200 meaning "version 200, unsigned"; another decodes the byte into a signed type and logs version -56; a range check `version >= 0` then rejects a valid file. Neither side corrupted a bit. Mainstream languages even disagree about the *default* reading of a byte — Java's byte type is signed (-128..127) while Go's byte is an alias for an unsigned 8-bit type (0..255) — which is precisely why serialization specs, not language habits, must carry the signedness decision. ## What to say at the whiteboard Decode it live: write `11001000`, sum the unsigned weights to 200, then replace +128 with -128 and sum to -56, then state the u - 2^n shortcut and the sign-magnitude contrast (-72) to show you know *which* wrong model the interviewer is fishing for. That covers value, mechanism, and misconception in under a minute.
- Under sign-magnitude, what would 0xC8 mean, and why did that encoding lose out?Sign-magnitude reads the top bit as a pure flag over the 7-bit magnitude 1001000, giving -72. It lost because it has two zeros (+0 and -0), complicates equality, and forces adders to special-case operand signs. Two's complement has one zero and lets a single adder handle signed and unsigned arithmetic identically.
- Give me the general rule for converting an n-bit unsigned reading into its signed reading.If the unsigned reading u is below 2^(n-1) the readings agree; otherwise the signed value is u - 2^n. The two interpretations of a pattern with the top bit set always differ by exactly 2^n, so one subtraction converts between them — for a byte, 200 becomes 200 - 256 = -56.
- Two decoders disagree about a header byte's value — what is the first question you ask?What does the format spec say the field's signedness and width are? If the spec is silent, that is the defect: both decoders are internally consistent, and no amount of inspecting bits resolves which number was meant. The fix is a spec clarification plus a normalization at the decode boundary, not a code patch on one side.
The digits "12" mean twelve in decimal and eighteen in hexadecimal — same marks on paper, different agreed weighting. Signed vs unsigned is the same ambiguity inside one byte, and the file format spec plays the role of saying which base you're in.
saying these in an interview costs you the question
- Explains the top bit as a sign flag over an unchanged magnitude
- Says the byte itself is signed or unsigned, rather than the reading
- Cannot produce the u - 2^n conversion rule
- Believes all-ones is the minimum value instead of -1
- Treats a decoder disagreement as data corruption rather than a spec gap