How is architecture compliance scoring typically constructed for a project or a portfolio, and what makes a compliance score misleading if it's built poorly?
answer
- weight by risk not headcount of principles
- gate principles can block regardless of %
- waived != non-compliant in scoring
- portfolio average can hide concentrated risk
- score should trend across gates, not one snapshot
basics
~20 sIt's a percentage showing how well a project follows the architecture rules, based on how many principles it meets versus breaks - but it can lie if it treats a minor rule and a critical security rule as equally important.
solid answer
~40 sCompliance scoring rolls a project's status against each applicable architecture principle (met / not met / waived / not applicable) into a single number or category, usually weighted so higher-risk principles (security, data protection) count more than cosmetic ones. It's built from ARB review outcomes and the waiver register, refreshed at each gate, and rolled up to portfolio level for governance reporting. The core failure is treating it as a single unweighted percentage: a project that's 95% compliant but missing the one control that actually matters looks healthier than it is, so a good scoring model weights by risk, separates 'waived' from 'non-compliant' (waived is known risk, non-compliant is unmanaged risk), and is read alongside trend rather than as an absolute pass/fail gate on its own.
go deeper
Should understand a compliance score is a summary of how well a project follows the rules, and that a waiver is different from just breaking a rule.
Should know that scores are usually weighted by risk rather than treating every principle equally, and where the score inputs come from (review outcomes, waiver register).
Should be able to design a scoring model - gate vs weighted principles, waived vs non-compliant treatment - and explain how it's reported at portfolio level.
Should be able to identify gaming and false-comfort failure modes from portfolio data, and redesign incentive structures so scoring drives real behavior change, not just reporting optics.
## How a score is built **Compliance scoring** is the mechanism that turns a set of individual architecture-principle checks into a single number, grade, or status (e.g., 'green/amber/red', or a percentage) that can be reported up through governance forums and rolled up across an entire project portfolio. Mechanically, it starts from a defined set of applicable principles for a given project - not every principle applies to every project, so the first step is scoping which ones are in play. For each applicable principle, the project is assessed into a status: - **compliant** - **non-compliant** - **waived** (a tracked, approved exception) - **not applicable** Those statuses are typically weighted, because principles are not equally important - a violation of a data-residency or authentication principle should move the score far more than a violation of a documentation-naming convention - and rolled into a single score or banding per project. At the portfolio level, individual project scores are aggregated so governance forums can see where compliance risk is concentrated rather than drowning in project-by-project detail. ## Why it exists This exists because architecture principles are only useful if adherence is visible and comparable across many concurrent projects; without a scoring mechanism, governance has no way to answer 'which of our forty active projects are the ones actually at risk' except by manually re-reading every review. A score also creates an **accountability artifact** - it can be tied to program funding health checks or executive dashboards, which is exactly what gives it teeth to actually change behavior instead of being an FYI nobody acts on. ## The trade-off The central trade-off is **simplicity versus accuracy**. | Model | What it buys | What it costs | |---|---|---| | A simple unweighted percentage (e.g., '18 of 20 principles met = 90% compliant') | easy to compute and compare across projects | silently assumes every principle carries equal risk, which is essentially never true | | Risk-weighted scoring | more accurate | harder to build and maintain: someone has to assign and periodically revisit weights, and a weighted score is harder to explain quickly to an audience that wants a simple number | A project that fails two low-stakes principles looks identical, at 90%, to a project that fails one principle protecting customer PII, even though the second carries order-of-magnitude more actual risk. Most mature governance models land on a **hybrid**: a small set of non-negotiable, heavily weighted 'gate' principles that can single-handedly block a project regardless of overall percentage, plus a broader weighted score for trend and portfolio reporting. ## How a waived item is scored A second design decision with real consequences is how waived items are scored. If a tracked, approved, time-boxed waiver counts identically to an undocumented, silent violation, the scoring model destroys the incentive to ever request a waiver - teams learn that deviating quietly costs the same score as deviating transparently. Better models score waived items separately from the raw compliance percentage, which preserves the incentive to go through the formal exception process. ## Failure modes Failure modes are common and mostly stem from over-trusting the number. - **'Gaming the score'** happens when project teams learn exactly which principles are weighted lightly and deliberately concentrate non-compliance there, or push hard for waivers specifically because waivers are scored more leniently than violations - the score goes up while real risk doesn't change. - **'Point-in-time blindness'** happens when the score is measured once at a gate and never refreshed, so a project that was 95% compliant at design review can drift substantially by go-live with nobody noticing until the next scheduled checkpoint; better models track compliance as a trend across gates, not a single snapshot. - **'False portfolio comfort'** happens at the aggregate level: an 88% average portfolio compliance score sounds healthy, but if it's an average across projects of wildly different risk, it can mask a small number of high-risk projects with genuinely dangerous gaps. A well-designed portfolio report segments by risk tier rather than presenting one blended number. ## A concrete example A concrete example: an insurer runs compliance scoring where five gate-level principles (data encryption, identity federation, PII data-residency, DR/RTO commitment, approved-tech-stack) are pass/fail and block go-live regardless of overall score, while roughly thirty additional principles feed a weighted percentage used for trend reporting and portfolio heat-maps; waived items are shown as a separate 'managed risk' band rather than folded into the raw percentage, so a governance forum can immediately tell the difference between a project that's 85% clean-compliant versus one that's 60% compliant plus 25% actively-managed-waiver, which are very different risk profiles even though a naive unweighted score might make the second look worse.
- Why is it dangerous to score a waived principle the same as an unaddressed violation?It removes the incentive to use the formal exception process, since going through review, getting sign-off, and accepting an expiry buys no scoring advantage over just quietly not complying. Teams under deadline pressure will rationally choose the path with equal cost and less friction, which is silent non-compliance.
- A portfolio dashboard shows 88% average compliance across 40 projects. What follow-up question should a governance lead ask before treating that as healthy?Whether that average is masking a small number of high-risk projects with severe gaps, since a blended average across projects of very different risk profiles can look fine in aggregate while a payments or PII-handling system sits dangerously low. They should ask for the score segmented by risk tier rather than trusting one number.
- How would you design scoring so a project can't game it by concentrating violations on lightly-weighted principles?Make a small set of high-risk principles gate-level and pass/fail - unable to be offset by scoring well elsewhere - so no amount of compliance on low-stakes items compensates for failing something critical. This caps the maximum benefit of gaming the weighted portion, because the gate can block progress regardless of overall percentage.
Like a credit score built only from 'number of accounts in good standing' with no weighting - it would rate someone who's never missed a $10 subscription payment above someone who's missed one mortgage payment, because it can't tell a trivial miss from a critical one.
saying these in an interview costs you the question
- treats compliance score as a single unweighted percentage
- can't explain the difference between a waived item and a silent violation in scoring
- assumes a high portfolio average means no serious risk exists
- scores compliance once and never revisits it as the project evolves
- no concept of gate-level/non-negotiable principles that can block regardless of overall score