Your teams are defining a shared in-house binary record layout; which byte-level conventions do you pin first, and what does each cost?
answer
- decide the primitives before the fields
- integers, byte order, field ends, text
- size against random access
- uniformity beats any single choice
- some reversals silently reinterpret old bytes
basics
~20 sPin four things: how integers are represented, the byte order of fixed-width fields, how every variable-length field declares its end, and the text encoding. Each one trades payload size against random access, writer complexity and how readable a hex dump stays.
solid answer
~50 sFour conventions carry almost all the weight. **Integer representation**: variable-length shrinks small values but forbids offset arithmetic and costs more than fixed width above 2^56, so a mixed rule - variable-length for counts and lengths, fixed width for identifiers and digests - is usually right. **Byte order**: pin one for the whole format rather than negotiating per field. **Self-delimiting fields**: decide whether every field carries a tag and a length, which lets any reader walk any record, or whether small scalars stay positional to save the overhead. **Text**: one encoding, stated once, with the unit of every length limit named. The meta-decision matters more than any single choice: apply the conventions uniformly, because the ones that reinterpret rather than reject old bytes - byte order and integer representation - are the ones you cannot quietly reverse once archives exist.
go deeper
Know that a binary layout is more than its field list: integer representation, byte order, how a field ends and the text encoding are decided once and apply everywhere.
Be able to argue each choice on mechanism - what variable-length integers save and what they forbid, what a tag and a length on every field buy a reader that lacks the schema.
Bring the operational angle: which conventions fail loudly and which reinterpret existing bytes in silence, and what it actually takes to change one after data has been archived.
Own the meta-decision. Defend uniformity over per-field optimisation, name the capability you are protecting - a record decodable by hand from a hex dump - and say which of these you would never reopen without a new format and a discriminator.
## The four decisions A layout is more than a list of fields. Before any field is defined, four conventions have to be settled, and they are worth settling explicitly because each one is invisible in the bytes and expensive to change later. 1. **How an integer is represented** - variable-length or fixed width, and which widths exist. 2. **Byte order** for every multi-byte fixed-width field. 3. **How a variable-length field declares its end** - a length prefix, and whether every field also carries a tag. 4. **The text encoding**, and the unit in which any length limit on text is counted. ## Integer representation | Property | Variable-length integer | Fixed-width integer | |---|---|---| | Small values | one or two bytes | the full declared width | | Large 64-bit values | nine or ten bytes | eight | | Reaching a field by offset | not possible, decode from the start | arithmetic on known widths | | Decode cost | a loop with a branch per byte | one aligned read | | Reading a hex dump by hand | groups must be reassembled | the value is visible directly | The honest rule is that variable-length encoding is a bet on the distribution. Counts, lengths, small keys and offsets near zero win; identifiers, digests and anything high-entropy lose outright, since they use their full width and then pay a byte or two of overhead on top. A layout that uses both deliberately is a considered design, not a failure of nerve - but write down which rule applies where, because inconsistency is what makes a format unreadable. ## Byte order Pin one convention for the whole format and convert at the boundary on both sides. The alternative - a marker in the stream that tells readers which order this producer used - buys writers their native layout and costs every reader a branch that is rarely exercised and therefore rarely correct when it finally runs. Whichever you choose, the decision belongs in the format's opening paragraph, not in each field's description. ## Self-delimiting fields The choice is between **tag-length-value everywhere** and **positional fields with lengths only where needed**. - Tag-length-value everywhere means any reader can walk any record without a schema: a field it does not recognise still announces its size, so the walk continues. That buys generic tooling - a dump utility, a byte-level diff, a debugging view - written once and useful for every record type you ever define. - The cost is a tag and a length on every field, which is heavy when the payload is a single small number. A record of twenty small scalars can easily double. Positional layouts invert both the benefit and the cost. The common compromise is tag-length-value for anything variable-length or optional, and a fixed positional prefix for the handful of fields every record has. ## The decision behind the decisions The conventions are worth less than their uniformity. The capability you are protecting is that an engineer with a hex dump and the spec can decode a record by hand at three in the morning. A layout that changes convention between fields destroys that capability even though every individual field is well defined. Rank the choices by how they fail if reversed: - **Byte order and integer representation reinterpret old bytes.** Change either and an existing archive still parses, producing wrong numbers with no error anywhere. That is the worst failure mode there is, and it is why these two are the decisions to get right first and never revisit casually. - **Text encoding mostly fails loudly**, as invalid sequences or visibly wrong characters, though a lenient reader can turn that into quiet corruption too. - **Field-delimiting conventions fail structurally** - the walk desynchronises and the record is rejected - which is unpleasant but detectable. If one of these must change after bytes exist in archives, the honest move is a new format with a discriminator in front, not a silent redefinition of what the old bytes meant. ## What is deliberately not settled here These are primitive-level conventions only. How a stream of records is cut into individual messages, what bounds a reader enforces on input it does not trust, and how the schema is allowed to change over time are three further decisions with their own answers. Settling the primitives first is what lets those conversations be about policy rather than about bytes.
- Which of these conventions is hardest to reverse once archives exist, and why?Byte order and integer representation, because old bytes still parse under the new rule and yield wrong numbers with no error raised anywhere. A text-encoding change at least tends to produce invalid sequences, and a change to how fields are delimited desynchronises the walk and gets rejected. Silent reinterpretation is worse than loud failure, so those two decisions carry the most weight.
- When would you accept fixed-width integers everywhere despite the size cost?When readers reach fields by offset instead of decoding from the start, when the values are high-entropy identifiers or digests that a variable-length encoding cannot shrink anyway, or when the per-byte branch in the decode loop shows up in a profile. Uniform fixed width also keeps a hex dump readable by hand, which is worth real bytes on a format people will debug for years.
- What do you gain by applying one self-delimiting layout to every field rather than only where it is needed?A reader can walk any record without a schema, so dump, diff and inspection tooling is written once and works for every record type you define later. The cost is a tag and a length on every field, which dominates for small scalars - hence the common compromise of self-delimiting variable-length fields over a fixed positional prefix.
saying these in an interview costs you the question
- Chooses variable-length integers everywhere without checking value magnitudes
- Leaves byte order to whichever side happens to write the field
- Assumes a layout decision stays revisable once archives exist
- Mixes self-delimiting and positional fields in one record for convenience
- Treats hand-readability of the bytes as worth nothing