What are "disinformative" names and "noise words" in identifiers, and what does the rule "one word per concept" require?
answer
- accountList that isn't a list = disinformation
- noise words: Data / Info / Object / the
- test: delete the word — did meaning change?
- one verb per operation: get vs fetch vs retrieve
- don't pun: `add` must mean one thing
basics
~20 sA disinformative name implies something untrue — accountList for a set, getUser that also creates one. Noise words add characters but no meaning — Data, Info, the, Object. One word per concept means picking a single verb per operation across the codebase (get, not get/fetch/retrieve).
solid answer
~50 s**Disinformation** is a name that makes a false promise: `accountList` when the value is a set or map; `hp` when it means hypotenuse but reads as horsepower; `getUser()` that lazily creates and persists a user; `price` holding cents while a sibling holds dollars. It is worse than vagueness, because readers act on it without checking. **Noise words** are morphemes that add no distinguishing information: `Data`, `Info`, `Object`, `Manager`, `the`, `a`, `variable`. Their tell is that removing them changes nothing — if `Product`, `ProductData` and `ProductInfo` coexist, no one can choose between them. **One word per concept** means the codebase uses exactly one verb per operation (`fetchUser`/`getUser`/`retrieveUser` collapse to one) and one noun per domain entity, so names become predictable and greppable. The dual rule — don't pun — says the same word must not mean two different things: `add` must not sometimes mean list-append and sometimes mean arithmetic sum.
code
pseudocode · 19 lines// disinformation
accountList: Set<Account> // not a list
getUser(id) // secretly creates + persists on miss
timeout = 30 // seconds? milliseconds?
// noise words: which one holds the price?
class Product { ... }
class ProductData { ... }
class ProductInfo { ... }
// vocabulary drift: three names, one operation
fetchUser(id); getUser(id); retrieveUser(id)
// repaired
accounts: Set<Account>
findUser(id) -> User? // find = may be absent
getUser(id) -> User // get = throws if absent (documented convention)
timeoutMillis = 30_000
class Product; class ProductDto // suffix carries a real distinctiongo deeper
Give one concrete example of each: a name that lies (accountList on a set), a noise word (ProductInfo next to Product), and a synonym drift (get/fetch/retrieve).
Add the operational tests — delete the word and see if meaning is lost; ask whether a caller could pick between two names — plus units-in-names and side-effect-hiding verbs.
Cover the don't-pun dual rule, deliberate vocabularies like find (optional) vs get (throws), and why a half-completed vocabulary migration is worse than either endpoint.
Institutionalise it: a written glossary aligned with domain language, lint and architecture rules for suffix conventions and banned noise words, and vocabulary migrations planned as complete, fenced changes rather than opportunistic edits.
## Three related failures ### 1. Disinformation — the name makes a false promise A vague name costs the reader time. A *disinformative* name costs them correctness, because they will act on the implication without verifying it. Sources of disinformation: - **Container-type lies.** `accountList` when the value is a `Set`, a `Map`, or an array. "List" carries real semantics — ordering, duplicates allowed, index access. Prefer `accounts`, or name the actual structure if it matters (`accountsById`). - **Domain-term collisions.** `hp` may be hypotenuse to you and horsepower to the reader; `aix` and `sco` are OS names. Avoid abbreviations that already mean something else in the room. - **Verbs that lie about effects.** `getUser()` that lazily creates, caches, and persists is not a getter. `validate()` that also mutates. `isValid()` that throws. Anything named like a cheap query but doing expensive or state-changing work will be called in loops, in logs, and in debugger watch expressions — where the side effect is a genuine bug source. - **Unit and currency lies.** `timeout = 30` (seconds? milliseconds?), `price` in cents next to `total` in dollars, `date` holding a local time when everything else is UTC. Put the unit in the name (`timeoutMillis`, `priceCents`, `createdAtUtc`) or, better, in a type (`Duration`, `Money`, `Instant`). - **Near-identical names.** `getActiveAccount`, `getActiveAccounts`, `getActiveAccountInfo` in one class — the differences are invisible when scanning, and autocomplete will hand you the wrong one. - **Characters that look alike.** Lowercase `l` versus `1`, uppercase `O` versus `0`. ### 2. Noise words — the name is longer but no more informative A noise word is a token you could delete without losing meaning. Common offenders: `Data`, `Info`, `Object`, `Item`, `Value`, `Manager`, `Processor`, `Helper`, `Util`, `the`, `a`, `an`, `my`, `variable`, `table`, `String`. The operational test: **if two names differ only by a noise word, no caller can tell which to use.** `Product`, `ProductData`, `ProductInfo` — which one holds the price? `getAccount()`, `getAccountInfo()`, `getAccountData()` — you have to open all three. That ambiguity is the harm; it is not an aesthetic complaint. Note the necessary exception: some of these words are meaningful in some contexts. `Customer` (entity) vs `CustomerDto` (wire shape) vs `CustomerEntity` (persistence mapping) is a *real* distinction and a legitimate, conventional suffix. The rule is not "never use a suffix"; it is "every word in the name must do work." `variable`, `the`, and `Object` never do. ### 3. One word per concept, and its dual: don't pun **One word per concept (consistent vocabulary).** Pick one verb for one operation and use it everywhere: `fetch`, `get`, `retrieve`, `load`, and `find` should not all mean the same lookup. Same for nouns: `Controller`, `Manager`, and `Driver` should not all denote the same layer. The benefit is *predictability* — a developer who has never seen your `InvoiceRepository` can guess `findById` correctly — and *searchability*, because one search finds every occurrence of a concept. A useful refinement: reserve *different* words for *genuinely different* semantics, and document the distinction. Many teams settle on `find…` returning an optional/nullable result and `get…` throwing when absent. That is not inconsistency; that is a vocabulary, as long as it is applied uniformly. **Don't pun (the dual rule).** Never use the same word for two different concepts. If `add` means arithmetic addition on `Money` and list-append on `Collection`, a reader who learned one meaning will misread the other. When the semantics differ, choose a different verb — `append`, `insert`, `plus`. ## How this shows up in review and tooling - **Review heuristic:** if you have to open the implementation to know what a name means or whether it lies, that is a finding. - **Tooling:** many linters ban a configurable list of noise words in type names; spell-checkers catch unpronounceable abbreviations; architecture tests can enforce suffix conventions per layer (`*Controller`, `*Repository`). - **Glossary:** the durable fix for vocabulary drift is a written glossary of the team's canonical verbs and domain nouns, reviewed when new terms appear. Without it, every new joiner adds one more synonym. ## Trade-off Consistency has a cost: enforcing one vocabulary across a large or legacy codebase means mass renames, and half-finished renames are *worse* than either consistent state, because now there are two vocabularies and no signal about which is current. Do vocabulary migrations as a deliberate, complete change (or fence them per module), not opportunistically file by file.
- Why is a disinformative name considered worse than a merely vague one?A vague name makes the reader stop and check; a disinformative one makes them confidently proceed on a false assumption. `accountList` invites index access and duplicate assumptions; a `getX()` that mutates invites calls from loops, logs, and debugger watches where side effects become real bugs.
- Is `Dto` or `Entity` a noise word?No, when the codebase genuinely distinguishes a wire/serialization shape from a persistence mapping — the suffix carries information the reader needs. The test is whether deleting the word loses meaning. `Data`, `Info`, `Object`, `the`, and `variable` fail that test; `Dto` in a layered codebase passes it.
- How do you actually enforce one word per concept in a large codebase?Write the canonical verbs and domain nouns into a short glossary, encode what you can as lint or architecture-test rules (suffix conventions per layer, banned noise words), and do vocabulary migrations as complete, module-scoped changes — a half-done rename leaves two vocabularies and no signal about which is current.
Disinformation is a road sign pointing the wrong way — worse than no sign, because drivers obey it. Noise words are signs reading "SIGN AHEAD": technically present, entirely uninformative.
saying these in an interview costs you the question
- "The name is fine, the type is right there" — ignoring that readers scan names, not signatures.
- Adding `Data`/`Info`/`Manager` to distinguish two classes instead of naming the actual distinction.
- Treating all suffixes as noise, including meaningful ones like `Dto`/`Entity`/`Repository`.
- Keeping several synonyms for one operation because "renaming is churn", leaving a half-migrated vocabulary in place.
- Naming a value `timeout` or `price` with no unit and relying on a comment or convention.