skip to content

Two strings hold identical characters — why can an identity comparison still report them as different?

level: middleimportance: must knowfreq 68%

answer

  1. two different questions, not two ways to ask one
  2. one implication holds, the reverse does not
  3. where do run-time strings come from
  4. why the unit test passed anyway
  5. identity is safe only after canonicalization

basics

~20 s

Identity comparison asks whether two references point at one object; content comparison asks whether the characters match. Equal contents can live in two separately allocated strings, so identity is false while content is true. Compare contents unless every value was canonicalized.

solid answer

~50 s

They are two different questions. Identity asks "is this one object under two names?" and is a single reference comparison, O(1). Content equality asks "do these two objects hold the same characters?" and is O(n) in the worst case, though it usually short-circuits on length or on a cached hash. Nothing about immutability forces equal values to share one object: a string assembled at run time from input has the same characters as one written in the source, but it was allocated separately, so identity is false while content equality is true. Identity comparison *appears* to work in development because test values are typically fixed literals that the runtime canonicalizes into one shared instance — pure survivorship bias, which collapses the first time a value arrives over the network. Identity is a legitimate fast path only when every value on both sides provably came from one canonicalization table.

code

pseudocode · 12 lines
pseudocode
a = "service-alpha"                 // fixed in the program text
b = concat("service-", "alpha")     // assembled at run time
c = a                               // a second name for one object

identity(a, c)                      // true: one object, two names
identity(a, b)                      // may be false: two objects
same_characters(a, b)               // true: 13 identical characters

// a content comparison may use identity only as a fast path:
//   if identity(x, y) then return true
//   if length(x) != length(y) then return false
//   for i in 0..length(x)-1 ... compare x[i] == y[i]

go deeper

for a junior

Be ready to say that one comparison asks whether two names point at a single object and the other asks whether the characters match, and that two separately created strings can hold identical characters.

for a middle

Explain where the second object comes from — decoded input, a slice, a trimmed or rebuilt value — and why literals in the source are the canonicalized exception that makes tests pass while production fails.

for a senior

Diagnose the intermittent shape of this bug: comparisons that succeed only for values that happened to pass through a cache, so failures correlate with cache warmth and disappear under a reproducer built from literals.

for a principal

Decide whether an identity fast path is worth its enforcement cost at all. If you allow it, the canonical value needs its own type so no future caller can route around the table; if you cannot guarantee that, standardise on content comparison.

## Two different questions | Comparison | The question it answers | Cost | |---|---|---| | Identity | Do these two references denote **one object**? | O(1), one reference comparison | | Content equality | Do these two objects hold the **same characters**? | O(n) worst case; short-circuits on length, and on a cached hash mismatch | The two answers coincide in one direction only. **Identical implies equal** — one object trivially has the same characters as itself. **Equal does not imply identical** — two separately allocated strings may hold exactly the same characters and remain two objects. Immutability does not change this. Immutability guarantees the characters of a value cannot drift; it makes no promise that the system stores each distinct value exactly once. Storing each distinct value once is a *separate* mechanism — canonicalization, also called interning — and it applies only to the values you actually route through the canonical table. ## Where the extra objects come from ``` a = "service-alpha" // fixed in the program text b = concat("service-", "alpha") // assembled at run time c = a // a second name for one object ``` `a` and `c` are one object; identity is true. `b` holds thirteen identical characters but was allocated by a run-time operation, so identity against `a` is false while content equality is true. Every source of run-time strings behaves like `b`: characters decoded from a socket or a file, a slice taken out of a larger buffer, the result of trimming or lowercasing, a value deserialized from a request body, a key rebuilt from two fields. Fixed literals in the program text are the *exception* that gets canonicalized, not the rule. ## Why the bug survives testing This is the trap the question exists to set. In a unit test you write the expected value as a literal and often construct the input as a literal too. Both sides are canonicalized to one shared object, identity returns true, and the test passes. In production one side arrives over the wire and identity returns false — for values that are, character for character, equal. Worse, this often only breaks *partly*: the values that happen to have gone through a deduplication cache still compare identical, so the failure is intermittent and correlates with cache warmth, which is about the most expensive shape a bug can have. The right lesson is not "identity comparison is flaky"; it is that identity was answering a different question all along and happened to give the answer you wanted on your sample. Mainstream ecosystems differ enough here to make the point vividly: Java and C# automatically canonicalize compile-time string literals so identity often succeeds on them, Python canonicalizes short identifier-like strings by an unspecified rule that has changed between versions, and Go performs no such canonicalization at all. Any code whose correctness depends on which of those rules is in force is depending on an implementation detail. ## When identity is legitimate Identity comparison of strings is a valid and genuinely fast optimisation under one condition: **every value being compared, on both sides, is guaranteed to have come from the same canonicalization table.** That is achievable in a closed system — a parser that maps every token to a canonical instance on the way in, an event pipeline that canonicalizes the field it later groups by — and it is worth real money when the comparison sits in a hot loop, because you trade an O(n) character walk for one reference comparison. It is unachievable, and therefore a bug, whenever an un-canonicalized value can reach the comparison from any path: a new caller, a new deserialization route, a cached value that expired. The discipline that keeps this safe is to make the canonicalization the *only* constructor for that type of value — wrap the canonical instance in its own type, so a raw string cannot be passed where a canonical one is expected. If you cannot enforce that, compare contents. ## A useful pattern from inside content comparison A well-implemented content comparison usually starts with an identity check as a fast path: if the two references are the same object, return true immediately without touching a single character. That is identity used *correctly* — as an optimisation inside an operation whose contract is content equality, where a false answer costs nothing but a character walk. Contrast that with identity used *as* the equality test, where a false answer is simply wrong. The difference between those two uses is what an interviewer is really probing.

  • Why is an identity check often the first line of a content-equality routine?
    Because it is a free short-circuit with no false positives: if the two references are one object, the characters are trivially equal and the routine returns without touching them. Used this way a false result costs nothing — the routine falls through to the character walk. The defect is using identity *instead of* the routine, where a false result is simply wrong.
  • A colleague says immutability guarantees equal strings are stored only once. What's wrong with that?
    It conflates two independent mechanisms. Immutability says an existing value's characters cannot change, which is what makes sharing *safe*; it never says the system will find and reuse an existing equal value. Storing each distinct value once requires a canonicalization table plus a lookup on every construction, which nothing does automatically for run-time-built strings.
  • How would you make an identity-based fast path safe to rely on?
    Give the canonical value its own type whose only constructor goes through the canonicalization table, and accept that type — not a raw string — everywhere the fast path is used. Then a value that skipped the table cannot reach the comparison. Without that enforcement, one new deserialization path silently reintroduces false negatives.

saying these in an interview costs you the question

  • Says equal contents means the same object
  • Believes immutability guarantees one instance per distinct value
  • Treats identity comparison as a faster spelling of equality
  • Concludes identity works because the tests pass
  • Thinks identity comparison falls back to comparing characters
  • Claims content comparison is constant time

context