skip to content

questions

6

Why must a binary wire format pin the byte order of its fixed-width integer fields instead of leaving it to each writer?

level: middleimportance: must knowfreq 66%

answer

  1. several bytes, two possible orders
  2. the bytes alone cannot say which
  3. most significant first or last
  4. network byte order means big-endian
  5. zero reads the same under both

basics

~20 s

Because a multi-byte integer can be laid out most-significant byte first or least-significant byte first, and the two readings disagree: the bytes 00 00 00 01 mean 1 one way and 16,777,216 the other. Only the format can settle which.

solid answer

~40 s

A fixed-width integer occupies several bytes, and nothing about the bytes themselves says in which order they carry significance. **Big-endian** puts the most significant byte first and is what 'network byte order' means; **little-endian** puts the least significant byte first and is the in-memory layout of most mainstream processor families. If the format does not pin one, a writer will naturally emit its own memory layout and a reader on the other convention will read a byte-reversed value - a count of 1 arriving as 16,777,216 in a 32-bit field. The choice only touches **multi-byte fixed-width numeric fields**: a byte string and UTF-8 text already have a defined order, and a variable-length integer's group order is part of its own definition.

go deeper

for a junior

Recall that a multi-byte number can be written most significant byte first or last, that both are valid, and that the two sides must agree because the bytes themselves do not say which was meant.

for a middle

Explain which fields are affected and which are not - fixed-width integers yes, single bytes and UTF-8 text no - and work a concrete reversal such as 1 reading back as 16,777,216.

for a senior

Demonstrate the diagnosis: hex dump the field, read it both ways, and recognise the byte-reversed value. Note why zero-heavy fixtures let the bug reach production undetected.

for a principal

Own the contract term. Decide whether the format pins one order or carries a marker, and be able to defend the cost of a rarely exercised reader branch against the writer-side conversion it saves.

## What byte order actually means A 32-bit integer is four bytes. Writing it to a flat stream forces a decision no single byte can express: which of those four is the most significant. The two answers have names. | Convention | Bytes on the wire for the value 0x0A0B0C0D | |---|---| | **Big-endian**, also called network byte order | `0A 0B 0C 0D` | | **Little-endian** | `0D 0C 0B 0A` | Both are complete, correct, lossless representations. Neither is discoverable from the bytes: `0D 0C 0B 0A` is a perfectly good big-endian number too, it just means something else. Byte order is therefore not a property of the data, it is a **term in the contract**, and a contract that omits it is incomplete. ## Which fields it touches This is where most confusion lives. The choice applies only to fields made of several bytes whose numeric significance is ordered: - **Fixed-width integers** - 16-, 32-, 64-bit fields: yes, this is the whole problem. - **A single byte** - a flag, a small enumerated value, a tag: no. There is nothing to order, and bit order inside a byte is not what endianness means. - **Byte strings and UTF-8 text**: no. The sequence of bytes *is* the value, in the order given. (The contrast is instructive: UTF-16 does have two byte orders for the same reason integers do, which is why it needs a byte-order mark and UTF-8 does not.) - **A base-128 variable-length integer**: no. Its groups run least-significant first by definition of the encoding, independent of what the format chose for its fixed-width fields. ## What a mismatch looks like The failure is quiet in a specific and dangerous way. 1. A 32-bit field holding **1** is written big-endian as `00 00 00 01` and read little-endian as **16,777,216**. 2. The same field holding **0** reads as **0** under both conventions. 3. Any value whose byte pattern happens to be a palindrome also survives. So a test suite whose fixtures are mostly zeros and empty records passes, and the bug surfaces later as absurd numbers in production: a count in the millions where a handful was expected, a size field that asks for gigabytes. The diagnosis is mechanical - take a hex dump of the field, interpret it both ways, and check whether the value you expected is the byte-reversed reading of the value you got. If it is, the byte order is the cause and nothing else needs investigating. ## How formats settle it There are two workable designs and one that is not. 1. **Pin one convention for the whole format.** Every reader and writer converts between its own memory layout and the wire convention at the boundary. This is the standard answer, and it is why so many wire formats simply declare big-endian: the convention needs to be stated once, not negotiated. 2. **Carry the order in the data.** The stream opens with a known constant - a magic word or a byte-order mark whose correct value is fixed - and the reader inspects it to learn the convention, then branches. This buys writers the ability to emit their native layout without conversion, and it costs every reader a branch on every multi-byte field plus a code path that is rarely exercised and therefore rarely tested. 3. **Leave it unstated.** Not a design. It works exactly as long as both sides happen to share a processor family, and breaks the first time one does not. The practical rule when defining a layout is to state byte order once, in the same breath as the integer widths, and apply it uniformly. Mixing conventions between fields of one record - which happens when a layout grows by accretion and someone copies a field definition from elsewhere - makes the record impossible to read by hand from a hex dump, which is precisely the capability the whole primitive layer exists to preserve.

  • Does the format's byte order affect a UTF-8 text field or a base-128 variable-length integer?
    No, neither. The bytes of UTF-8 text are the value in the order given, which is exactly why it needs no byte-order mark while UTF-16 does. A variable-length integer's groups run least-significant first because its own encoding says so, not because the format chose a convention for its fixed-width fields.
  • Two services share a format, but one reads an implausible value from a 32-bit count. How do you confirm byte order is the cause?
    Take a hex dump of those four bytes and interpret them both ways. If the value you expected is the byte-reversed reading of the value you got - 1 against 16,777,216, or 256 against 65,536 - the diagnosis is settled. The confirming detail is that records whose counts are zero look perfectly healthy.
  • What does a format buy and lose by carrying a byte-order marker instead of pinning one convention?
    It buys writers the right to emit their native layout with no conversion, which can matter for a high-volume producer. It costs every reader a branch per multi-byte field and a second code path that is almost never exercised, so the rarely-used branch is the one most likely to be wrong when it finally runs.

saying these in an interview costs you the question

  • Thinks byte order also reverses the bytes of a UTF-8 string
  • Says network byte order means least significant byte first
  • Assumes every machine is little-endian now, so it cannot matter
  • Believes a byte-order bug always shows up in tests
  • Treats endianness as a property of a single-byte field
open as a page

How does a base-128 varint use the top bit of each byte to encode an integer whose width the reader does not know in advance?

level: middleimportance: must knowfreq 62%

basics

~20 s

A base-128 varint carries seven payload bits per byte and spends the eighth as a continuation flag: set means another byte belongs to this number, clear means this is the last one. The reader stops at the first byte whose top bit is clear.

open as a page

For a variable-length field inside a binary record, what changes when its length is written as a prefix rather than marked by a terminator byte?

level: middleimportance: should knowfreq 52%

basics

~20 s

A length prefix lets the reader take exactly that many bytes and step over the field without inspecting it, and the payload may hold any byte value. A terminator forces a scan and forbids that byte inside the payload unless it is escaped.

open as a page

A service trims a UTF-8 text field so it fits a byte-length budget on the wire; what can go wrong, and why?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A length counts bytes, but a UTF-8 character takes one to four of them, so a cut at an arbitrary byte position can land inside a character. The field then carries an incomplete sequence that a strict decoder rejects outright.

open as a page

Your teams are defining a shared in-house binary record layout; which byte-level conventions do you pin first, and what does each cost?

level: principalimportance: should knowfreq 30%

basics

~20 s

Pin four things: how integers are represented, the byte order of fixed-width fields, how every variable-length field declares its end, and the text encoding. Each one trades payload size against random access, writer complexity and how readable a hex dump stays.

open as a page

Why does a binary format apply ZigZag encoding to a signed integer before writing it as a varint?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Two's complement puts every negative number at the top of the unsigned range, so -1 would fill ten varint bytes. ZigZag interleaves the signs, mapping 0, -1, 1, -2 onto 0, 1, 2, 3, so small magnitudes stay short.

open as a page