skip to content

A legacy codebase has table and variable names like cust_tbl, proc_flg, and acct_stat_cd that no longer match how the business actually talks about the domain today. Is renaming these to match current domain vocabulary worth doing, and how would you approach it without a risky big-bang rewrite?

level: middleimportance: nice to knowfreq 22%

answer

  1. code renames cheap, schema renames risky
  2. gap exists because old constraints froze names in amber
  3. strangler pattern: translate at boundary before touching schema
  4. additive migration: new column, backfill, dual-write, cutover, drop old
  5. weigh compounding confusion tax vs one-time migration risk

basics

~20 s

Yes, it's usually worth doing gradually — renaming things to match how people actually talk saves confusion for years to come. Do it in small, safe steps (rename in code first behind the old database names, or use views/aliases) rather than one giant risky change all at once.

solid answer

~40 s

It's worth doing when the mismatch is actively causing costly confusion — bugs from misread abbreviations, slow onboarding, mistranslation in requirements discussions — but it should be scoped and incremental, not a big-bang rewrite. A practical approach mirrors the strangler pattern: introduce a translation layer (a repository or mapping class) that exposes clean, domain-accurate names in the application code while the underlying legacy schema keeps its cryptic column names for now, decoupling the risky, slow database migration from the immediate, low-risk win of a readable domain model in code. Database renames, if pursued, happen later and separately, via additive migrations (new column, backfill, dual-write, cut over, drop old column) rather than an in-place rename that risks breaking every consumer of that table at once.

go deeper

for a junior

Can recognize that cryptic legacy names like proc_flg are confusing and suggest renaming would help; not expected to know migration safety techniques like dual-write or the strangler pattern.

for a middle

Proposes an incremental approach (translation layer first, schema later) and understands why an in-place schema rename is riskier than a code rename.

for a senior

Designs the actual migration plan (additive column, backfill, dual-write, cutover, drop) and can weigh when the migration isn't worth the risk for a given system.

for a principal

Sets policy for how and when legacy renaming efforts get prioritized across many systems, balancing long-term maintainability against migration risk and competing roadmap priorities at a portfolio level.

## Code renames and schema renames are different problems The decision to rename legacy artifacts like `cust_tbl` or `proc_flg` to match current domain vocabulary is fundamentally a cost-benefit judgment, not an automatic 'always do it' rule, and the mechanism for doing it safely is different depending on which layer — code or schema — you're touching. - **In application code** (class names, variable names, method names), renaming is comparatively cheap: modern IDEs perform safe, compiler-verified renames across an entire codebase in seconds, and the risk is low because the change is purely textual with no runtime behavior difference. - **In a database schema, however**, a column or table rename can break every application, report, ETL job, and ad hoc query that references the old name directly, some of which the renaming team may not even know exist — so schema renames need to be treated as a distinct, higher-risk migration, not folded into the same effort as a code cleanup. ## Why the gap opens in the first place The reason this problem exists at all is that legacy names are usually artifacts of old constraints that no longer apply: - decades-old naming conventions that capped identifiers at 8 or 30 characters, forcing abbreviations like `proc_flg` for 'processing flag' or `acct_stat_cd` for 'account status code'; - or names chosen by an early team before the current domain vocabulary had even stabilized. Over years, the business's actual language moves on (analysts stop saying 'account status code' and start saying 'membership tier', say), but the schema is frozen in amber because changing it is expensive and risky, so a widening gap opens between what the business calls something and what the database calls it — exactly the ubiquitous-language drift problem, except calcified in a place (a production schema with years of dependents) where fixing it is much more expensive than fixing a class name. ## The trade-off The trade-off is real engineering time and migration risk versus an ongoing, compounding **translation tax**. Every engineer who touches this code pays a small toll indefinitely: they have to remember that `proc_flg` means 'processing flag' (not, say, 'processed' or 'procurement'), that acct_stat_cd's cryptic values map to specific business states, and new hires take longer to become productive because the schema actively obscures the domain rather than expressing it. That toll, paid by every engineer for years, can easily exceed the one-time cost of a careful migration — but only if the migration is actually careful; a rushed, big-bang rename that breaks a downstream reporting job during a busy period can cost far more in incident response and trust than the years of minor confusion it was meant to fix. ## The safe approach The safe approach borrows from the **strangler fig pattern** used for legacy modernization generally: introduce translation at a layer boundary rather than renaming the thing everyone already depends on. Concretely, an application-level repository or mapping class can expose clean, domain-accurate methods and types — `findCustomerAccountStatus()` returning a `MembershipTier` enum — while internally still querying `cust_tbl.acct_stat_cd`, so every new piece of application code gets to work against clear, ubiquitous-language-aligned names immediately, with zero schema risk. If and when the schema itself is worth migrating (because, say, direct SQL access by other teams or BI tools makes the cryptic names a persistent problem beyond the application layer), that's done as its own additive, reversible migration: 1. add a new column with the clear name, 2. backfill it from the old column, 3. dual-write to both during a transition window, 4. migrate all known consumers over, 5. verify, 6. and only then drop the old column — never an in-place rename that has no rollback path if an unknown consumer breaks. ## Failure modes on both sides Failure modes show up on both sides of this decision. - **Doing nothing indefinitely** means the confusion tax keeps compounding — bugs from misreading acct_stat_cd's encoded values, onboarding friction, and a codebase that increasingly requires tribal knowledge to navigate, which is a real, if slow-burning, cost. - **Doing it recklessly** — an in-place `ALTER TABLE RENAME COLUMN` run against a production schema without first finding every consumer — is the more dramatic failure: an unknown nightly ETL job or a BI dashboard's raw SQL query breaks silently until someone notices stale numbers days later, which is a strong argument for the application-layer-first, additive-migration-later approach. A concrete real-world pattern this maps to is any large, long-lived system with a mainframe-era or early-2000s schema underneath a modern application layer (common in banking and insurance systems) — teams there routinely build a clean domain layer in code on top of a schema they can't easily touch, precisely because the schema-migration risk and the code-cleanup benefit are decoupled problems best solved on different timelines.

  • When is it NOT worth renaming legacy names at all, even incrementally?
    When the schema is stable, rarely touched by new development, and well understood by the small team that maintains it — the ongoing confusion cost is low and localized, so the migration risk and effort of even an application-layer translation may not be worth it. This is common for a legacy system in maintenance mode that isn't accumulating new features.
  • How do you find all the unknown consumers of a table before attempting even an additive migration?
    Check database access logs or query logs over a representative time window (ideally including month-end or quarter-end jobs that run infrequently), grep application and ETL codebases for the table and column names, and check any BI tool's saved queries or dashboards; even after this, treat the dual-write/verification window as the real safety net, since some consumer will likely be missed by static search alone.
  • Does introducing a translation layer in application code count as fully solving the ubiquitous language problem, or is it a partial fix?
    It's a real, meaningful fix for anyone working in the application code going forward, since they now read and write domain-accurate names — but it's partial in that anyone querying the raw schema directly (BI tools, DBAs, ad hoc SQL) still sees the old cryptic names, so the underlying drift isn't fully resolved until the schema itself is eventually migrated.

Like renovating an old building's interior room labels before touching its load-bearing plumbing — you can repaint the signage and give staff a clear map of what each room is actually used for today without yet ripping out the pipes behind the walls, and you only reroute the plumbing later, carefully, once you've confirmed nothing unknown depends on the old layout.

saying these in an interview costs you the question

  • Proposes an in-place schema rename without first identifying all consumers
  • Says legacy names never need fixing since 'the team already knows what they mean'
  • Doesn't distinguish the risk profile of a code rename from a schema rename
  • Treats this as an all-or-nothing decision rather than a cost-benefit judgment
  • Has no incremental plan and jumps straight to a full rewrite

context