You own architecture governance for dozens of teams on AWS. How would you run Well-Architected reviews so they change what gets built, instead of becoming an annual paperwork exercise?
answer
- Calendar cadence breeds ritual
- Trigger on launches, incidents, re-architecture
- Encode local standards as your own lens
- Findings need an owner and a date
- Measure change between snapshots, not scores
basics
~20 sTie reviews to lifecycle events rather than a calendar, encode organisation-specific standards as custom lenses, save milestones so progress is measurable, and require that high risk issues become funded, owned backlog items. A review nobody is resourced to act on produces documents, not change.
solid answer
~60 sThree decisions matter. **When**: trigger reviews on real events — before a first production launch, after a significant incident, before a major re-architecture — rather than on a quarterly calendar that guarantees ritual. **What is asked**: use the AWS lens as the base but add a custom lens encoding your own non-negotiables, and use review templates so organisation-wide guardrails are answered once instead of by forty teams. **What happens after**: every high risk issue leaves the review with an owner, a decision — fix, accept with written justification, or defer with a date — and a place on a real backlog. Save a milestone at the end so the next review measures movement rather than asserting it. Share workloads with the central group so you can see the estate without owning every account, and feed Trusted Advisor findings in as evidence. The failure mode to design against is a review whose output nobody is funded to act on; at that point you are paying senior engineers to fill in a form.
go deeper
Understand that a review is only useful if the risks it finds turn into work someone owns, and that a snapshot at the end is what makes the next review measurable.
Be ready to describe how findings leave a review: an owner, a decision to fix or accept or defer, a date, and an entry on the backlog the team actually plans from.
Show judgment about triggers and scope — which workloads get a deep review, which events justify one, and how you keep the improvement plan alive after the meeting ends.
Own the programme's incentive design: what you measure, why absolute risk counts corrupt the data, where custom lenses and templates encode your standards, and where you accept imperfect coverage of the long tail.
## The problem with a review programme Every governance instrument decays the same way: it becomes a form, the form becomes a gate, and the gate gets satisfied rather than used. Well-Architected reviews decay especially fast because the questionnaire is long, the answers are self-reported, and nothing in AWS verifies them. A programme design has to fight that specifically. ## Trigger on events, not a calendar An annual review is a ritual. Reviews land hardest at moments when the design is genuinely in question: - **Before first production launch.** The team is already thinking about the design and has not yet accumulated the sunk cost that makes findings unwelcome. - **After a significant incident.** The failure is concrete, and the reliability and operational excellence questions stop being hypothetical. - **Before a major re-architecture or a large scale-up.** Changing the design is on the table anyway. - **On ownership change.** A team inheriting a workload has to build a model of it regardless; a review is a structured way to do that. A light annual refresh for the most critical workloads on top of that is defensible. Making the annual pass the *only* trigger is not. ## Decide what "all workloads" means You cannot review everything at depth and should not try. Tier the estate: workloads whose failure is a business event get full reviews with the operators in the room; the long tail gets a lightweight subset, or nothing beyond automated checks. Being explicit about the tiers is better than a nominal universal policy that quietly produces rubber-stamped reviews at the bottom. ## Encode your own standards The AWS framework lens asks generic questions. Your organisation has specific non-negotiables — an identity pattern, a tagging scheme, a deployment mechanism, a data classification rule — and the way to make them part of the same conversation is a **custom lens**: your questions, your best practices, your risk rules, uploaded as JSON, versioned, and shared to the accounts that need it. Versioning matters, because the standard will change and you want to know which version a given review answered. **Review templates** handle the other half: answers identical across every workload because they follow from a platform guardrail. Answering "we enforce this at the organisation level" once, centrally, is both more accurate and less tedious than forty teams answering separately and three getting it wrong. **Profiles** let a workload record its business context so the tool prioritises the questions and risks that matter for its current goals. ## Make the output into work This is where programmes live or die. A review produces high and medium risk issues; if they stay in the tool they are documentation. Each high risk issue should leave the session with three things: an owner, a decision, and a date. The decision is one of *fix*, *accept* (with a written justification signed by someone who can accept the risk), or *defer* (with a date and a trigger). Accepted risk is a legitimate outcome — a programme that only permits "fix" trains teams to answer optimistically so nothing is found. Then mirror the items into whatever backlog the team actually works from. A finding that does not appear where engineers plan work does not exist. ## Measure movement, not scores Save a milestone at the end of every review. The metric worth reporting upward is *change between milestones* — high risk issues closed, risks explicitly accepted with justification — not a raw count that punishes honest teams. If you report absolute risk counts, teams will optimise them, and the cheapest way to lower a risk count is to mark questions not applicable. Design the metric so honesty is not penalised, or you will get dishonest reviews and never know it. ## Feed the automated evidence in Trusted Advisor findings, Config rules and security tooling give account-level facts. The Well-Architected Tool surfaces relevant Trusted Advisor checks next to design questions where your Support plan allows, and that pairing — the claim next to the evidence — is worth more than either alone. Keep the boundary clear: automation covers what is machine-checkable, so the human review spends its time on the parts that are not. ## Run it as a service, not an audit The practical difference between a programme engineers welcome and one they route around is whether the central group facilitates or judges. Facilitators bring the lens, run the session, take the notes, and help write the improvement plan. Auditors send a spreadsheet and a deadline. Share workload records with the central group with contribute access so you participate in reviews rather than demanding reports about them, and accept that the estate view you get will be imperfect.
- How do you stop teams from gaming a review by marking questions not applicable?Two ways. Report movement between milestones rather than absolute risk counts, so an honest team with many findings is not punished. And have the central group facilitate rather than receive — a reviewer in the room asking why a question was marked not applicable catches most of it. Also make accepted risk a legitimate, low-friction outcome, so honesty has somewhere to go.
- What belongs in a custom lens versus in automated policy enforcement?Anything machine-checkable belongs in automation — organisation policies, config rules, CI checks — because a questionnaire is a bad place to assert facts a computer can verify. A custom lens is for judgment: whether the recovery objective matches the business need, whether the team can operate what it built, whether a design decision was deliberate. Ask humans only what humans can answer.
- How would you justify the cost of this programme to a sceptical executive?By tying it to events with a price. Reviews before launch and after incidents catch design faults while they are still cheap to change, and the milestone history shows exactly which risks were closed and which were consciously accepted and by whom. If the programme cannot show closed risks between milestones, the sceptic is right and it should be cut back.
saying these in an interview costs you the question
- Mandates a review for every workload every quarter
- Reports absolute risk counts as a team scorecard
- Leaves findings in the tool with no owner
- Treats accepted risk as a failed review
- Uses the questionnaire for facts automation could verify