skip to content

A GraphQL field typed Int fails only for the oldest vehicles in a fleet — why, and what do you change?

level: seniorimportance: should knowfreq 41%

answer

  1. Only some rows fail, and only old ones
  2. Staging fixtures are always small
  3. The specification fixes a width, not a meaning
  4. Two point one billion is the wall
  5. Exactness can be lost after the response leaves

basics

~20 s

Int is specified as a signed 32-bit value, so it tops out at 2,147,483,647. High-mileage vehicles have crossed that boundary and their values cannot be coerced. Fix it by changing the field's type, not by clamping the data.

solid answer

~50 s

The tell is that the failure is **data-dependent**: it appears only for rows whose value has grown past 2,147,483,647, which is the maximum of `Int`'s specified signed 32-bit range. The backend column is a 64-bit integer, so nothing broke below the boundary and nothing broke in staging, where the seeded odometers are small. Result coercion cannot represent the value, so it fails rather than truncating — which is the correct behaviour, because a silently wrapped odometer would be worse than an error. The fix is a **type change on the field**, and the honest options are: a custom scalar for large integers, serialized as a string so it survives clients whose numbers lose exactness above 2^53; or a unit change so the value fits, which is lossy; or `Float`, exact only to 2^53 and quietly wrong beyond it. The type change is breaking, so it follows the usual additive path.

code

graphql · 11 lines
graphql
# before: breaks once a vehicle passes 2,147,483,647 metres
type Vehicle {
  odometerMetres: Int!
}

# after: a custom scalar carrying a large integer as a string
scalar Int64

type Vehicle {
  odometerMetres: Int64!
}

go deeper

for a junior

Remember the number: a GraphQL Int stops at 2,147,483,647. If a value can grow past that, Int is the wrong type for it, whatever the database column says.

for a middle

Explain that the failure happens during result coercion and that refusing to serialize is deliberate, then compare the unit change, the Float change and a custom scalar rather than naming one fix.

for a senior

Diagnose from the shape of the symptom — data-dependent, time-dependent, invisible in staging — and argue for the string-serialized custom scalar on the grounds that exactness can be lost in the client after the response leaves you.

for a principal

Turn the incident into a rule: audit every Int for its five-year maximum, decide centrally which large-number scalar the organisation uses so clients configure it once, and treat monotonic counters as scheduled failures rather than hypotheticals.

## The symptom, and why it looks like a ghost A fleet telematics graph exposes `Vehicle.odometerMetres: Int!`. It has worked for two years. Then a support ticket says the vehicle detail page is blank for a handful of vehicles, and separately a telemetry subscription that had been streaming happily for weeks silently stopped delivering payloads to one operator's dashboard — no reconnect, no visible failure, just nothing new on screen for the vehicles that operator watches. The two reports look unrelated and neither reproduces in staging, where the seeded fixtures have odometers in the tens of thousands. They are the same bug. Every affected vehicle is an old one — a long-haul unit whose cumulative metres has crept past 2,147,483,647. ## The cause `Int` is specified as a **signed 32-bit non-fractional value**. That range — -2,147,483,648 to 2,147,483,647 — is part of the type's definition, not a quirk of one server. A vehicle at 3,147,000,000 metres has an odometer the type cannot represent. Crucially, the value did not arrive corrupted. The database column is a 64-bit integer, the service holds it correctly, the resolver returns it correctly. The failure is at **result coercion**: the scalar is asked to turn a value into an `Int` and it cannot, so it raises a coercion failure rather than emitting a number. That refusal is the right design. A truncated or wrapped odometer would be a plausible-looking lie propagated into billing, maintenance scheduling and compliance reports — an error is strictly better. And because `odometerMetres` was declared non-null, the failure does not stay local; it removes the containing object from the response. That is why the page is blank rather than missing one number, and why the subscription's per-event payloads stopped carrying anything useful. ## Why it survived review, tests and staging Three properties make this class of bug slippery: * **It is data-dependent.** The schema is not wrong for 99% of rows. No amount of schema linting catches a boundary that only some values cross. * **It is time-dependent.** A counter that only increases converts a latent modelling mistake into a scheduled outage. The question is not whether the field will break, but which week. * **Fixtures are small.** Seeded test data almost never contains a value near a 32-bit boundary, because whoever wrote the fixtures picked a readable number. The general lesson worth stating in an interview: when you type a field `Int`, you are asserting that the value can never exceed roughly 2.1 billion — for all time, for every row. Counters, byte totals, millisecond durations, cumulative distances and anything denominated in a small unit routinely break that assertion. ## The options, honestly compared **Change the unit.** `odometerKilometres: Int` puts the value back in range for the lifetime of any real vehicle. Cheapest change, and it loses precision permanently. Acceptable when the extra precision was never used; unacceptable if some consumer bills per metre. **Change to `Float`.** Tempting, and wrong in a specific, quiet way: a double is exact for integers only up to 2^53. Values below that are fine, values above are silently rounded. You have replaced an error with an incorrect number, which is a downgrade. It also changes the field's meaning from a count to a measurement, and typed clients will start treating it as fractional. **Change to a custom scalar for large integers.** The conventional fix. Declare a scalar — the ecosystem tends to call it something like `Long` or `BigInt`, none of which is specified — and serialize it **as a string**. The string is the point: many clients decode JSON numbers into a type that itself loses exactness above 2^53, so a large integer sent as a JSON number can be corrupted after it leaves your server, no matter how correct your coercion was. A string survives the trip intact, and the client's mapping for that scalar decides how to parse it. The cost is real: every client generator now needs a configuration entry for that scalar, and callers doing arithmetic must parse first. **Change to `String`.** Works on the wire, and throws away every signal that the value is a number — no ordering, no arithmetic, no generated numeric type. Prefer a custom scalar, which carries the same string on the wire but keeps the intent visible in the schema and mappable in clients. ## Two more things a senior answer includes First, changing a field's type is not a compatible change for existing consumers, so it goes through the same additive path as any other contract change rather than being swapped in place. Second, this is worth a sweep rather than a one-off fix. Grep the schema for every `Int` and ask what its maximum is in five years. Cumulative counters, epoch values in milliseconds, byte counts and any quantity in a small unit are the usual suspects, and finding the other three before they page you is worth more than the fix you just shipped.

  • Why serialize a large integer as a string rather than as a JSON number?
    Because exactness can be lost after your server is done. Plenty of clients decode JSON numbers into a double-based numeric type that is exact only up to 2^53, so a correct 64-bit value sent as a number can be silently rounded on arrival. A string crosses the wire intact and the client's mapping for that scalar decides how to parse it.
  • Should the server clamp the value to Int's maximum instead of failing?
    No. Clamping or wrapping produces a plausible number that is wrong, and it will be trusted — fed into billing, maintenance intervals and compliance reporting. A coercion failure is loud, localised and diagnosable; a corrupted odometer is none of those. If the value cannot be represented, the type is wrong and the type is what should change.
  • How would you find the other fields with this problem before they fail?
    Sweep every `Int` in the schema and ask what its maximum plausible value is over the service's lifetime. The usual offenders are cumulative counters, byte totals, durations in milliseconds, epoch timestamps in milliseconds and any quantity denominated in a small unit. Anything monotonically increasing is a scheduled failure rather than a hypothetical one.

saying these in an interview costs you the question

  • Assumes Int follows the backend's 64-bit column
  • Proposes clamping or wrapping the value
  • Switches to Float and calls it fixed
  • Blames the client because staging passes
  • Thinks a JSON number always survives the client intact
  • Treats the type change as non-breaking for consumers

context