skip to content

You're comparing three candidate solutions using a weighted decision matrix with criteria like cost, time-to-market, and scalability. Walk through how you'd build and use it, and name one way it commonly gets misused.

level: middleimportance: must knowfreq 75%

answer

  1. weights sum to 100 (or 1.0)
  2. score x weight per cell
  3. set weights before scoring
  4. matrix informs, doesn't decide
  5. hard constraints filter before scoring

basics

~20 s

List the things that matter (like cost and speed), give each a weight for importance, score each option on each thing, multiply and add up the scores, and the highest total is the suggested winner -- but you still sanity-check it, because the weights themselves were a judgment call.

solid answer

~50 s

A weighted decision matrix lists evaluation criteria as rows and candidate options as columns. Each criterion gets a weight (say, out of 100 total) reflecting its relative importance, agreed with stakeholders before scoring to reduce bias. Each option is then scored per criterion (e.g., 1-5) against a defined scale, and the weighted score per cell is weight times score, summed per option. The highest total is a strong signal, not an automatic verdict -- you sanity-check it against qualitative factors the matrix can't capture (team morale, political risk, one-off vendor concerns) and against constraints that should have excluded an option outright rather than just penalized it. Common misuse: picking weights or scores after already favoring an outcome, so the matrix becomes retrofitted justification instead of genuine analysis -- reviewers should be able to see the weights were set before scoring, ideally by a cross-functional group.

go deeper

for a junior

Can fill in a decision matrix template given the criteria, weights, and scoring scale, and correctly compute the weighted totals.

for a middle

Can propose the criteria and weights for a moderately complex comparison, defend the weighting rationale, and knows to separate hard constraints from scored criteria.

for a senior

Facilitates the weighting session with stakeholders to avoid post-hoc bias, and knows when the matrix's output should be overridden by qualitative judgment -- and can articulate why in writing.

for a principal

Establishes the org-wide convention for how these matrices get used in architecture decisions (e.g., as an ADR appendix, not the whole decision), and prevents the tool from being used as a rubber stamp for decisions already made politically.

## What the tool is A weighted decision matrix is a simple scoring tool for comparing a small number of options against a shared set of criteria in a way that makes the comparison's logic **visible rather than implicit**. ## The mechanics Mechanically: 1. **List the rows and columns.** You list the evaluation criteria as rows -- for example cost, time-to-market, scalability, team familiarity, operational burden -- and the candidate options as columns. 2. **Weight the criteria.** Each criterion is given a weight reflecting how much it matters relative to the others, commonly normalized so the weights sum to 100 or to 1.0 (e.g., `cost=40`, `time-to-market=30`, `scalability=30`). 3. **Score each option.** Then each option is scored against each criterion on a fixed scale, typically 1-5, ideally against a written description of what each score point means (a '5' for 'operational burden' might mean 'no new on-call load,' a '1' might mean 'requires a new 24/7 on-call rotation'), rather than left to gut feel. 4. **Total the weighted cells.** The weighted score for each cell is weight multiplied by score, and the option's total is the sum of its weighted cells across all criteria. The option with the highest total is a data point favoring that choice, **not an automatic verdict**. ## Why the technique exists The reason this technique exists is that unstructured, purely qualitative comparisons ('Option A feels more scalable, Option B feels cheaper') let the loudest voice or the architect's gut carry the decision without exposing which factors actually drove it, or whether people even agree on what matters most. By forcing an explicit weighting exercise before scoring, the matrix separates two questions that often get conflated in a debate: - **the weights** -- 'what matters, and how much' - **the scores** -- 'how does each option perform' This separation is valuable because disagreements about a decision are frequently actually disagreements about priorities -- one stakeholder implicitly weighting cost heavily, another implicitly weighting scalability heavily -- and the matrix surfaces that disagreement explicitly instead of letting it play out as a vague argument about which option is 'better.' ## The trade-off The trade-off is that the matrix's apparent objectivity is a construction, not a measurement: the weights and the scores are still human judgments, just made explicit and quantified rather than left implicit. This is valuable for transparency but risks manufacturing **false precision** -- a computed total of '87' out of 100 sounds authoritative, but it's built from several subjective 1-5 judgments multiplied together, and small differences in scoring choices can swing the total more than the underlying reality actually differs between options. The matrix also can't capture everything: qualitative factors like team morale, political feasibility, or strategic optionality either get force-fit into a scored criterion poorly, or left out of the process entirely, and then get raised anyway, informally, undermining the matrix's authority. ## Failure modes 1. **Post-hoc weighting** is the most damaging failure mode in practice: someone favors an option before the matrix exists, and either the weights or the scores get quietly adjusted, consciously or not, until the numbers land where that person already wanted them. This is why setting and locking weights with the relevant stakeholders before anyone sees preliminary scores is important -- if weights get revisited after a result is visible, the matrix has stopped being an analysis tool and become a justification exercise. 2. **Conflating a hard constraint with a scored criterion** is a related failure: an option that violates a genuine pass/fail requirement, such as a data-residency law, should be eliminated before scoring, not merely given a low score on some related criterion, because a low score still lets a disqualified option appear competitive in the final ranking, and a busy reader may miss the disqualification buried in the numbers. 3. **Vague scoring scales** are a third common failure -- 'scalability: 1-5' without a defined meaning per point invites two people to score the identical option a '2' and a '5' for entirely different, unstated reasons, and averaging those scores hides a real disagreement instead of resolving it. ## A worked example A concrete example: a team choosing a message queue technology sets weights of cost=20, operational maturity=35, ecosystem/community support=25, and latency=20. They score Kafka, RabbitMQ, and a cloud-managed queue against each, using pre-agreed scoring definitions (e.g., 'operational maturity 5' means 'team has run it in production for 2+ years'). - **Kafka** scores highest on ecosystem and latency but lowest on operational maturity for this particular team, since nobody has run it before. - **RabbitMQ** is a strong all-rounder. - **The cloud-managed queue** wins on operational maturity by outsourcing it entirely. The totals come out close between RabbitMQ and the managed queue, which is itself useful information -- it tells the team the decision hinges on factors the matrix didn't weight, like vendor lock-in risk, prompting a focused conversation on exactly that point rather than a debate about the whole comparison from scratch.

  • What's the difference between a hard constraint and a scored criterion in this kind of comparison?
    A hard constraint is a pass/fail gate -- e.g., 'must run in our existing AWS region for data residency' -- that eliminates an option entirely regardless of how well it scores elsewhere. A scored criterion, like cost or time-to-market, is a matter of degree that trades off against other criteria via weights. Conflating the two is a common mistake: teams give a low score to an option that actually violates a hard constraint, letting a disqualifying option still appear competitive in the totals.
  • How do you stop stakeholders from reverse-engineering weights to guarantee their preferred option wins?
    Set and lock the weights (and ideally the scoring scale) with the group before anyone scores the options, and write the rationale for each weight down next to it. If someone proposes changing a weight after seeing preliminary totals, treat that as a signal worth discussing openly rather than silently applying it. Some teams also have a neutral facilitator run the weighting session separately from the scoring session.
  • Why might two people scoring the same option on the same criterion disagree by a lot, and what do you do about it?
    It usually means the criterion or scoring scale wasn't defined concretely enough -- 'scalability: 1-5' is far too vague to score consistently, whereas 'can sustain 10x current peak load without re-architecture: 1-5' is scorable. When scores diverge significantly, that's valuable signal to discuss rather than average away, since it often surfaces a hidden assumption one person is making that the other isn't.

It's like grading job candidates with a rubric: you decide upfront that 'system design' is worth more than 'whiteboard trivia,' score each candidate against the same rubric, and the totals guide -- but don't replace -- the hiring committee's judgment.

saying these in an interview costs you the question

  • Sets weights after seeing preliminary scores
  • Treats the total score as an automatic, unquestionable decision
  • Scores a criterion that should have been a hard pass/fail constraint
  • Uses vague, undefined scoring scales like 'good/bad'
  • No named stakeholders behind the weights

context