skip to content

Your schema-to-Go generator maps field names to identifiers: how do you apply initialism casing and catch two fields collapsing to one name?

level: seniorimportance: nice to knowfreq 26%

answer

  1. canonicalise first, then detect
  2. the mapping is many-to-one
  3. one initialism table, versioned
  4. fail before writing the file
  5. the golden diff shows the rename

basics

~20 s

Split the schema name into words, canonicalise each through one explicit initialism list, then join in MixedCaps. Before writing any file, fail the run when two source fields produce the same Go identifier, naming both.

solid answer

~50 s

Generation happens in two passes. First, canonicalise: split the schema name on its own separator, look each word up in a single explicit initialism table — `id` to `ID`, `url` to `URL`, `http` to `HTTP` — and title-case the rest, so `user_id` becomes `UserID` and never `UserId`. Second, detect collisions: build a map from the produced Go identifier back to the source field, and abort the run when a second field maps to an identifier already taken, with an error naming both source names. Without that pass, `user_id` and `userId` both produce `UserID`, and the failure surfaces as a compile error pointing at a generated line rather than at the schema that caused it. Keep the initialism table in one versioned place and review generated output through a golden-file diff, because widening the table renames identifiers across the whole generated API at once.

code

go · 21 lines
go
var initialisms = map[string]string{
	"id":   "ID",
	"url":  "URL",
	"http": "HTTP",
	"api":  "API",
}

func goName(field string) string {
	var b strings.Builder
	for _, word := range strings.Split(field, "_") {
		if word == "" {
			continue
		}
		if up, ok := initialisms[strings.ToLower(word)]; ok {
			b.WriteString(up)
			continue
		}
		b.WriteString(strings.ToUpper(word[:1]) + strings.ToLower(word[1:]))
	}
	return b.String()
}

go deeper

for a junior

Know what the transform must produce: user_id becomes UserID, not UserId. Recognise that a naive title-case of every word breaks Go's initialism rule.

for a middle

Explain that the mapping from schema names to Go identifiers is many-to-one, so two differently spelled fields can produce one identifier, and that a Go struct cannot declare the same field name twice.

for a senior

Design both passes and defend where the check lives: fail in the generator with both source names before writing output, never silently disambiguate, and use a golden-file diff so a rename is reviewed as one decision.

for a principal

Own the initialism table as shared infrastructure. Widening it renames the generated API for every consumer at once, so decide who approves that, how it is staged, and which boundaries are allowed to mint Go names at all.

## The problem A generator that turns an interface-definition schema into Go types has to mint Go identifiers from names that were written for a different naming culture — usually `snake_case`, sometimes `camelCase`, occasionally both in the same file because different teams contributed different messages. Two things go wrong, and only one of them is obvious. ## Pass one: canonicalisation The naive transform is "split on underscore, upper-case the first letter of each word, join". That produces `UserId`, `HttpUrl`, `ApiKey` — every one of which violates Go's rule that an initialism keeps uniform case. Generated code is read by humans and, worse, it is *depended on* by humans, so `UserId` propagates into every hand-written adapter around it. The fix is a single explicit table of initialisms — `id`, `url`, `http`, `api`, `json`, `xml`, `sql`, `tls`, `uuid`, `db`, `os`, `ip` — consulted per word after splitting. Three properties matter: - **One table, one place.** If the templates, the field-name transform and the method-name transform each carry their own list, they drift, and you get `UserID` fields on a type with a `GetUserId` method. - **Versioned and reviewed.** Adding `api` to the table renames every `Api...` identifier in the generated package at once. That is a deliberate, reviewable event, not a silent Tuesday. - **Position-aware.** If the generator also emits unexported names, a leading initialism lowercases entirely (`urlPool`, not `uRLPool`) while an interior one stays uppercase. ## Pass two: collision detection Canonicalisation is many-to-one. `user_id`, `userId`, `userID` and `USER_ID` all collapse to `UserID`. Schemas evolve independently of your generator, so the day someone adds a field whose spelling differs only in separators or case from an existing one, your transform silently produces two struct fields with the same name. The generated file then fails to compile, because a Go struct cannot declare the same field name twice. That is a *safe* failure — nothing incorrect ships — but it is a bad one, because the error points at a line in a file nobody wrote and the engineer has to reverse-engineer the transform to work out which two schema fields fought. So the generator should detect it itself: ``` seen := map[string]string{} for _, f := range fields { name := goName(f) if prev, dup := seen[name]; dup { return fmt.Errorf("fields %q and %q both generate %s", prev, f, name) } seen[name] = f } ``` The error names both sources and the identifier they fought over, and it fires before a byte is written. That turns a ten-minute archaeology session into a one-line fix in the schema. Do not silently disambiguate — appending a suffix to the loser produces `UserID` and `UserID2`, which is worse than failing, because it compiles and ships a field name nobody can guess from the schema. Fail, and make the schema author choose. Also collide-check against everything else the generator emits into the same namespace: a field named `Name` and a generated method `Name()` cannot coexist on one Go type, so the check has to cover method names, not just fields. ## Reviewing the output Generated packages are reviewed through a **golden-file diff**: the committed generated output is regenerated in CI and the run fails if the result differs from what is checked in. This is what makes naming changes visible. When the initialism table gains an entry, the diff is a large, obvious rename set that a reviewer can approve as one decision. Without the golden file, the same change lands as an unnoticed API break in whichever package regenerates first. That is also the argument for making the table hard to edit casually: the diff is the safety net, and it only helps if somebody is required to read it. ## What a strong answer includes - The two passes, distinctly: canonicalise, then detect. - That the detection must precede file output, because the compiler's version of this error is correct but useless. - That silent disambiguation is worse than failing. - That the initialism table is a shared, versioned artefact, and widening it is a rename of the generated API. - That generated identifiers are somebody's public API the moment the package is imported, so the naming rule cannot be revisited cheaply later. ## The wider point Every boundary that mints Go names from foreign ones has this shape: a database schema, a wire format, an environment-variable mapping. The naming convention is not enforced by the toolchain — `gofmt` will format `UserId` as happily as `UserID` — so the enforcement point is whatever code creates the name. Putting the rule there, once, is the difference between a codebase with one spelling and a codebase that argues about it in review forever.

  • Why not let the Go compiler catch the duplicate instead of checking in the generator?
    It does catch it — a struct cannot declare the same field twice — but the message points at a generated line rather than at the two schema fields that collided. The engineer then has to reverse-engineer the transform. Detecting it in the generator turns the same failure into an error naming both source fields, before any file is written.
  • Should the generator disambiguate a collision automatically?
    No. Emitting `UserID` and `UserID2` compiles and ships, which is worse than failing: the second name is unguessable from the schema and becomes permanent API the moment somebody imports it. Fail the run and let the schema author rename, so the decision is made by the person who knows what the field means.
  • What breaks when you add a new entry to the initialism table?
    Every generated identifier containing that word is renamed at once — `ApiKey` becomes `APIKey` across the package. Since generated identifiers are API for anyone importing it, that is a breaking change disguised as a config tweak. The golden-file diff is what makes it visible as one reviewable decision rather than a surprise.

saying these in an interview costs you the question

  • Title-cases every word, producing UserId and ApiKey
  • Keeps a separate initialism list in each template
  • Silently appends a suffix to disambiguate a collision
  • Relies on the compiler to report duplicate generated fields
  • Treats an initialism table change as a harmless config edit