At the scale of hundreds of concurrent projects, how would you design principle enforcement so it doesn't collapse into either governance theater or an unstaffable bottleneck?
answer
- fitness functions = continuous automated architecture checks
- policy-as-code in CI/CD + cloud posture scanning
- risk-tiered human review only above threshold
- self-attestation for the long tail
- feedback loop from waivers back into principle revision - or you just automate stale rules faster
basics
~10 sAutomate the easy, objective rule checks so they run constantly without needing people, save human review time for the genuinely hard judgment calls, and keep revisiting the rules themselves so they don't go stale.
solid answer
~40 sAt portfolio scale, manual review can't be the enforcement mechanism for every principle - there simply aren't enough architect-hours. The design splits principles into two tiers: objectively checkable ones (approved tech stack, mandatory encryption config, network segmentation rules) get encoded as automated policy-as-code / fitness functions running continuously in CI/CD and cloud config scanning, so violations are caught in minutes, not at a quarterly review. Genuinely judgment-heavy principles stay with human review, but only for projects above a risk/cost threshold, with everything else self-attested. The system stays healthy only if there's a feedback loop from waiver and violation data back into revising the principles themselves - otherwise automation just enforces stale rules faster and at greater scale.
go deeper
Should understand that not every rule can be checked by a person one by one and some checks can be automated.
Should know the basic split between automated checks and human review, roughly which kinds of principles fit each.
Should be able to design risk-tiered review thresholds and identify which specific principles are good automation candidates versus which need judgment.
Should be able to design the full layered system at portfolio scale, including the platform investment trade-off and the mandatory feedback loop that keeps automated enforcement from ossifying around stale principles.
## The capacity problem Enforcing architecture principles across a portfolio of hundreds of concurrent projects is fundamentally a **capacity problem**: a fixed, small number of expert architects cannot manually review every design decision made by every team without becoming the binding constraint on the entire portfolio's delivery speed. The design response is to stop treating 'enforcement' as synonymous with 'human review' and instead build a **layered system** where different kinds of principles are enforced by different mechanisms matched to how much judgment they actually require. ## The first layer: automated, continuous enforcement The first layer is automated, continuous enforcement for principles that are **objectively checkable** - meaning a machine can evaluate compliance without needing to understand business context or make a trade-off call. Examples: - mandatory encryption configuration - an approved list of cloud services or dependencies - network segmentation rules - required tagging for cost allocation - disallowed data flows between regions These get encoded as **policy-as-code** - infrastructure-as-code linting rules, admission-control policies, CI pipeline gates that fail a build on a disallowed dependency, cloud security posture management scanning production configuration continuously. This is often called **'fitness functions'** in the architecture literature: automated, repeatable checks that continuously verify a system still fits its intended architectural characteristics, the same idea behind automated test suites but pointed at architectural properties instead of functional correctness. The payoff is enormous: checks that would take a human reviewer an hour to manually verify run in seconds, on every commit, for every project, with zero marginal architect time - which is the only way enforcement scales **sublinearly** with portfolio size instead of linearly with it. ## The second layer: human judgment, risk-tiered The second layer stays as human review, reserved for principles that genuinely require judgment a machine can't replicate: - whether a particular integration pattern is the right choice for a specific business problem - whether a novel technology choice is a reasonable bet given the team's skills and the vendor's maturity - whether a proposed data model correctly balances normalization against query performance for this workload Applying human review to only these questions, and only for projects above a defined risk or cost threshold, is what keeps the review board's workload proportional to genuinely hard decisions rather than the whole portfolio. Everything below that threshold, and everything the automated layer already covers, is handled by **self-attestation** - the team declares compliance, spot-audited rather than gate-reviewed, trading some detection latency for a large reduction in review volume. ## The trade-off The trade-off in this design is **upfront investment versus ongoing review cost**. Building and maintaining policy-as-code checks - writing the rules, integrating them into every relevant pipeline and platform, keeping them from producing false positives that erode developer trust - is real, ongoing engineering work, often requiring a platform or DevEx team, not just architects. Skipping that investment and relying purely on human review is cheaper upfront but doesn't scale: as portfolio size grows, either review quality degrades or coverage silently narrows, which is governance theater by omission rather than by explicit rubber-stamping. The right amount of automation investment is itself a judgment call tied to portfolio size and principle stability - automating a principle that changes every quarter costs more in maintenance than it saves in review time. ## The failure mode at this scale The failure mode specific to this scale problem is **'enforcement without revision'** - automation makes it cheap to enforce a rule at massive scale, but nothing about automation makes the rule itself correct or current. If a fitness function encodes a principle that's become outdated, automating it just means the organization enforces a stale rule faster and more comprehensively than it ever could manually, generating friction and shadow-IT workarounds at scale. The fix has to be structural: a **feedback loop** where waiver requests, automated-check override requests, and exception patterns are aggregated and periodically fed back into a principle-review cycle, with clear ownership for retiring or updating rules - not just adding new ones. Without that loop, the enforcement system optimizes for compliance with an increasingly disconnected rulebook. ## A concrete illustration A concrete illustration: a large financial institution running several hundred concurrent digital projects might encode roughly 30 objectively-checkable principles as automated policy checks running in every CI pipeline and continuously scanning live cloud configuration, catching the large majority of violations within minutes of introduction with no architect involvement. A much smaller set of maybe 8-10 principles that require genuine judgment route to full board review, but only for projects above a defined cost or risk tier - perhaps the top 15% of the portfolio by risk score - with everything else self-attesting. Quarterly, the governance team reviews aggregated waiver and override data specifically looking for principles generating disproportionate exception volume, and retires or rewrites the ones that no longer fit, treating the automated rule set as a living asset rather than a fixed constraint.
- Why can't every architecture principle be automated?Some principles require weighing business context, team capability, and trade-offs that a rule engine can't evaluate - like whether a specific integration pattern fits a specific business problem, which depends on judgment rather than a checkable fact. Automation works for objectively verifiable properties (is encryption enabled) but not for evaluative ones (is this the right architectural choice here).
- What's the risk of investing heavily in automated policy-as-code checks without a process to revise the underlying principles?The organization ends up enforcing outdated rules faster and more comprehensively than it ever could manually, since automation removes the natural friction that might otherwise force a rule to be reconsidered. This tends to generate shadow IT and workarounds at scale, because the fastest-growing gap is between what's automatically enforced and what's actually still a good idea.
- How would you decide which principles get full human review versus self-attestation in a risk-tiered model?Tier by a combination of the project's cost/risk score and whether the specific principle in question is objectively checkable versus judgment-heavy - high-risk projects touching judgment-heavy principles get full review, while low-risk projects or objectively-checkable principles route to automated checks or self-attestation. The threshold should be revisited periodically as portfolio composition and automation coverage change.
Like airport security: metal detectors and X-ray machines (automated, scale to every passenger) handle the objective checks instantly, while a human agent only gets pulled in for the ambiguous case the machine flags - you'd never staff enough agents to hand-search every bag, and you'd never trust a machine alone with a genuinely judgment-heavy call.
saying these in an interview costs you the question
- treats human review as the only real enforcement mechanism at any scale
- assumes automating a check makes the underlying principle correct or permanent
- proposes automating everything including genuinely judgment-heavy decisions
- no risk-tiering - treats every project as needing the same review depth
- no feedback loop from exception/override data back into revising principles