How do you decide which numeric and temporal scalars in a cross-team wire contract must cross as strings rather than native numbers?
answer
- decide per field, not per contract
- identity or quantity is the first question
- measure against the weakest reader
- strings move checks into every reader
- boundary vectors make the policy real
basics
~20 sDecide per field, not per contract: compare the value's required exactness and magnitude against the weakest consumer's numeric type, and let identities and money cross as text while ordinary quantities stay native. Then enforce it with boundary vectors.
solid answer
~50 sThe decision has two inputs: what the value **is**, and what the **weakest reader** can hold exactly. A value that is an identity, or whose exact digits carry legal or financial meaning, gets text — you never compute on it, so the ergonomic loss is nil and the exactness gain is total. A quantity that is aggregated, compared or filtered stays a native number, because a text field pushes range checking out of the schema and into every reader and costs real money in an analytics layout. Between those, the test is magnitude: if the field can exceed the exact-integer range of the weakest consumer's numeric type, it cannot cross as a bare number. The policy is worthless without enforcement, so publish the allowed scalar shapes, put boundary vectors in the contract's conformance suite, and re-run the decision whenever a consumer with a weaker reader joins.
go deeper
Take away the shape of the decision: identities and money need exactness on the wire, ordinary quantities do not, and the limit that matters is what the receiving system can hold.
Be able to run the mechanical test — the field's widest possible value against the weakest consumer's exact range — and name what a text field costs in validation and payload size.
Show enforcement, not intent: declared scalar shapes, boundary vectors in the conformance suite, and the schema review as the moment a widened range is re-examined.
Frame the asymmetry explicitly. Text moves cost into readers and queries; native typing moves undetectable risk onto the consumer with the narrowest numeric type, who will not notice. Decide per field and re-decide when the consumer set changes.
## Why this is a judgment call and not a rule Both blanket policies are defensible and both are wrong. *Everything native* assumes every consumer's numeric type can hold every value exactly, which is false the moment a reader whose only number is a binary float joins the contract. *Everything a string* buys exactness by giving up the things a typed schema does for free. The work is deciding field by field, and then making the decision stick. ## Four questions to ask about each scalar 1. **Is it an identity or a quantity?** An identity — an entry reference, an account number, a tax registration — is a name that happens to be made of digits. Nothing is ever added to it, so representing it as text costs nothing and removes an entire class of silent corruption. 2. **Do the exact digits carry meaning?** A monetary amount, a tax figure, a quantity that must reconcile to a control total: exactness beats ergonomics, and the field gets a representation that preserves the decimal scale. 3. **What is the widest value the field can ever hold, against the weakest consumer's exact range?** This is the mechanical test. A field bounded well below a reader's exact-integer limit is safe as a native number; one that can exceed it is not, regardless of how the schema declares it. 4. **What is done with it downstream?** A value that is filtered, ranged, aggregated or sorted in a columnar analytics layout pays for being text on every query, and the schema can no longer express its bounds. That is a real cost to weigh against the risk. ## What each blanket policy actually costs | Policy | What you gain | What you give up | |---|---|---| | Every number native | Compact payloads, schema-level range validation, cheap aggregation and ordering | Silent rounding at any reader whose numeric type is narrower than the value | | Every number a string | Exactness regardless of the reader's type system | Range and type checks move from the schema into every reader's code; larger payloads; ordering and comparison hazards; analytics cost | | Per-field decision with declared shapes | Exactness where it matters, native typing where it does not | Ongoing governance: the policy must be written down, tested and re-applied | The middle row hides the most dangerous failure mode: a **false sense of safety**. A quoted value that the reader parses straight back into a binary float has moved the rounding one step later and fixed nothing, and because the field is now visibly a string, everyone assumes it is handled. A string policy is only real if the contract also states the type the reader must parse into, and the conformance suite proves it. ## Making the decision stick - **Publish the allowed scalar shapes.** A short list — identity as text, money as a minor-unit integer or decimal text with a currency, instants as a unit-named epoch integer or a date-time string, everything else native with declared bounds — is far more enforceable than case-by-case review. - **Ship boundary vectors with the contract.** A value just above the exact-integer limit, an amount with a significant trailing zero, a far-dated instant, a name in both canonical text forms. A new consumer's conformance run either round-trips them or fails on day one instead of in production. - **Review the decision at the schema, not at the code.** A pull request that widens a field's range is the moment to re-ask question three, and it is cheap there. - **Re-run it when the consumer set changes.** The weakest reader is a property of the *current* consumer set, not of the contract. Onboarding a consumer with a narrower numeric type is a contract event, not just an integration task. - **Record the reason in the contract.** A field that is a string for a reason nobody can recall gets "simplified" back to a number two years later. ## The trade-off to say out loud Every exactness choice moves work somewhere. Text moves validation from the schema into readers and cost into the query layer. Native typing moves risk onto whichever consumer has the narrowest numeric type — and, crucially, that consumer often does not know it is at risk, because nothing fails. That asymmetry is the argument for spending the bytes on the small number of fields where a wrong digit is a defect nobody will detect, and for refusing to spend them everywhere else.
- A team proposes making every number in the contract a string. What is your counter-argument?It trades one silent failure for several loud costs: the schema can no longer express bounds, so every reader re-implements range checking; payloads grow; ordering and filtering in an analytics layout get more expensive. It also invites a false sense of safety, because a string parsed back into a binary float rounds exactly as before. Spend exactness where a wrong digit is undetectable, not everywhere.
- How does onboarding a new consumer change the decision?The test is against the weakest reader in the current consumer set, so a new consumer with a narrower numeric type can invalidate a field that was safe yesterday. Treat onboarding as a contract event: run the boundary vectors against the new reader, and if a native field cannot round-trip them, the contract changes rather than the consumer quietly rounding.
- What belongs in the conformance suite that ships with the contract?Values on the edges of every fidelity decision: an identifier just past the exact-integer limit of a binary float, an amount with a significant trailing zero and one in a currency with a non-standard exponent, a far-dated instant and one near the epoch, and a text value in both canonical forms. Each must round-trip unchanged through a consumer before that consumer is considered conformant.
saying these in an interview costs you the question
- Applies one representation rule to every field in the contract
- Measures fidelity against the strongest reader rather than the weakest
- Assumes a string field is safe without stating the parsed type
- Treats onboarding a new consumer as purely an integration task
- Leaves the reason for a field's shape undocumented
- Ignores what a text field costs in an aggregation workload