skip to content

Why do EA frameworks separate data (information) architecture from application architecture as two distinct BDAT layers instead of treating "the data model" as just part of each application's own design?

level: middleimportance: should knowfreq 55%

answer

  1. canonical vs local schema
  2. MDM = master data management, golden record
  3. one Customer definition, many apps
  4. data outlives any single application

basics

~20 s

Because the meaning of information, like what a "Customer" is, needs to stay consistent across many different apps that use it. If each app defines it its own way, you end up with conflicting, duplicated data that nobody can fully trust.

solid answer

~50 s

Data architecture defines the canonical, system-independent meaning, structure, ownership, and lifecycle of enterprise information - such as a single agreed definition of "Customer" with agreed attributes and identifiers - while application architecture defines the software that implements, stores, and exposes that information for specific use cases. Keeping them separate lets multiple applications share and interoperate around the same information without each inventing its own definition; this is the basis of master data management and canonical data models. When data architecture collapses into "whatever the application's database schema says," organizations end up with several incompatible definitions of the same business entity across several systems, no single source of truth, and expensive reconciliation and ETL work to make systems agree. The trade-off is governance cost: someone has to own and arbitrate the canonical model across application teams who would otherwise move faster with their own local schema.

go deeper

for a junior

Understands that data architecture and application architecture are different concepts, not the same thing under two names.

for a middle

Can explain the canonical-versus-local-model distinction and give a master-data-management example.

for a senior

Discusses governance mechanisms - stewardship, mapping/translation at integration boundaries - and when local divergence from the canonical model is acceptable.

for a principal

Weighs organizational design for data governance (federated versus centralized stewardship), the cost of enterprise-wide MDM programs, and how to prioritize which entities most need canonical treatment.

## Two different sets of artifacts **Data architecture's artifacts** are conceptual and logical: - canonical entity definitions (what "Customer," "Order," or "Policy" mean and which attributes they carry) - data ownership and stewardship assignments - data flow and lineage diagrams showing how information moves between systems - lifecycle rules covering creation, retention, and deletion These are all deliberately independent of which specific database stores the data. **Application architecture's artifacts** are physical and local: the actual services and their database schemas, the APIs they expose, and the specific technology used to persist data for that application's own needs. The two are related but distinct — an application's local schema is one implementation of (a slice of) the canonical model, not the model itself. ## Why the separation exists This separation exists because information in a real enterprise is almost never used by only one application. A customer's identity, contact details, and status are needed by sales, billing, support, and marketing systems alike. If each application team independently designs its own notion of "Customer" — different fields, different identifiers, different rules for what "active" means — there is no way for those systems to agree on who a given customer is or what state they're in without expensive point-to-point reconciliation. **Master data management (MDM)** exists specifically to solve this: a canonical, authoritative "golden record" for core entities that other applications synchronize against or query, so meaning is established once, at the data layer, rather than reinvented per application. ## The trade-off: governance overhead versus local speed The trade-off is governance overhead versus local speed. Maintaining a canonical model requires ongoing stewardship — someone has to: - arbitrate disagreements (does "active customer" mean logged in within 30 days, or holding a paid subscription?) - keep the model current as the business evolves - enforce that new applications map to it rather than inventing parallel definitions Application teams often experience this as friction: their delivery is faster if they can design a schema that fits only their immediate use case without waiting for cross-team agreement. A common, pragmatic compromise is to let applications keep local schemas optimized for their own performance and bounded context, provided there's an explicit, maintained mapping — a translation layer at integration points — back to the canonical model, so meaning isn't silently lost or diverged when data crosses system boundaries. ## The failure mode when the separation is skipped The failure mode when this separation is skipped shows up as data drift and reconciliation sprawl: several systems each holding their own "Customer" table with different identifiers for the same real person, requiring nightly batch jobs or manual spreadsheets to reconcile counts and statuses that should simply agree. It surfaces painfully during events that require a unified view — a regulator asking for total exposure per individual across all products, or a company wanting a single customer 360 view for marketing — when nobody can produce one without a lengthy, error-prone data-matching project. It's also visible in day-to-day work as: - duplicate customer records - inconsistent reporting numbers between departments - staff quietly trusting one system's numbers over another's without anyone having formally decided which is authoritative ## What it looks like in practice A concrete scenario: a retail bank's savings-account application, mortgage application, and call-center CRM each independently built a "Customer" table over the years, each with its own customer ID and its own definition of fields like "address" or "employment status." When a merger or a new regulatory reporting requirement demands a single view of total credit exposure per individual across all three product lines, the bank discovers there is no reliable way to match the same person across the three systems — different spellings, different ID schemes, no shared key. Untangling this typically becomes a multi-year master data management initiative: 1. defining a canonical Customer model at the data-architecture layer 2. assigning a single golden identifier 3. then updating each application to either adopt that identifier directly or maintain an explicit, governed mapping to it That is precisely the work that treating data architecture as "just each application's database" had skipped from the start.

  • What is master data management and how does it relate to the data architecture layer?
    MDM is the discipline and tooling for creating and maintaining a single authoritative "golden record" for core business entities across all systems; it's the practical, operational implementation of data architecture's canonical model, typically via a hub or service that other applications synchronize against or query rather than each maintaining their own version of the truth.
  • If two application teams disagree on an attribute's definition - say, whether "active customer" means logged in recently or holding a paid subscription - whose job is it to resolve that, and at which layer?
    It's resolved at the data (information) architecture layer, typically by a data steward or governance body, who establishes one canonical definition that both applications must then align to. Leaving it unresolved means each application silently encodes a different answer, which is exactly the kind of drift the data layer exists to prevent.
  • Is it ever acceptable for an application to keep a local data model that differs from the canonical enterprise model?
    Yes - especially for performance or bounded-context reasons - as long as there's an explicit, maintained mapping or translation at integration points back to the canonical model, so meaning isn't lost when the data crosses a system boundary into another application or a shared report.

It's like a company's official style guide and dictionary (data architecture) versus each department's own memo templates (applications) - the templates can look different, but they all need to use the same defined terms so nobody in another department misreads a memo.

saying these in an interview costs you the question

  • equates data architecture with "the database schema of the main application"
  • doesn't know what master data or a canonical model means
  • considers five applications each with their own incompatible "Customer" table acceptable as long as each application works on its own
  • cannot say who resolves conflicting attribute definitions or at what layer

context