skip to content

Solution Architecture

Turning a set of requirements into one concrete, buildable design: analysing needs, selecting components, defining integrations and NFRs, modelling cost, comparing options and getting stakeholders to agree. It sits between enterprise architecture and the code.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

When deciding whether to build a piece of software in-house or buy an off-the-shelf/SaaS product, what is the single most useful question to ask about the capability itself, and why does it matter more than a raw cost comparison?

level: juniorimportance: must knowfreq 78%

answer

  1. core vs context
  2. differentiating vs commodity
  3. does the vendor's average serve your edge case
  4. revisit, don't set and forget

basics

~20 s

Ask whether this makes your company special, or is something every company needs the same way, like payroll. If special, building may be worth it. If it's the same everywhere, buying is usually cheaper and faster.

solid answer

~40 s

The key question is whether the capability is 'differentiating' (tied to why customers choose you over competitors) or 'commodity' (a supporting capability every company in the industry needs, implemented roughly the same way everywhere). Commodity capabilities -- auth, payroll, ticketing, CI runners -- should almost always be bought, because a vendor amortizes R&D across many customers and reaches maturity faster than any single internal team can justify. Differentiating capabilities are worth building in-house because owning them lets you iterate on the exact behavior that creates advantage, and no vendor tunes their product to your specific edge case. This classification isn't permanent: capabilities drift from differentiating to commodity as a market matures, so the answer must be revisited over time.

go deeper

for a junior

Should be able to explain the basic idea in plain language and give one example each of a commodity and a differentiating capability from a product they know.

for a middle

Should apply the framework to a real decision at their company, articulating which side a specific capability falls on and why, not just reciting the definitions.

for a senior

Should push back on stakeholders who mislabel a commodity capability as differentiating out of attachment or habit, and connect the decision to opportunity cost of engineering time.

for a principal

Should treat the classification as a living, organization-wide input to portfolio and staffing decisions, revisited on a cadence, not a one-time answer baked into an architecture doc.

## Classify the capability, not the price tag The build-vs-buy decision starts with classifying the capability, not comparing price tags. | Class | When it applies | Examples | |---|---|---| | **commodity** (or **context**, in Wardley Mapping terminology) | it is necessary for the business to function but does not distinguish it from competitors | authentication, expense reporting, a CI pipeline, a support ticketing system | | **differentiating** (or **core**) | it directly encodes the thing customers pay for | a recommendation algorithm for a media company, a matching engine for a marketplace, the physics engine for a games studio | Every build-vs-buy conversation should open by placing the capability on this spectrum, because the spectrum, not the dollar figure, determines which factors matter. ## The two questions that classify it Mechanistically, classify a capability by asking two things: 1. **First**, if this capability vanished tomorrow, would a customer notice a difference in why they picked this company over a rival? 2. **Second**, does a mature market of vendors already sell this exact capability well? If the answer to the first is no and the second is yes, buy -- there is no advantage to protect and someone else has already paid down the maturity curve. If the answer to the first is yes, building deserves serious consideration, because a vendor's roadmap is driven by the average need across their whole customer base, and average is precisely what a differentiator cannot afford to be. ## Why the classification outranks the dollar figure This distinction exists because engineering time is the scarcest resource in most organizations, and every hour spent reinventing a commodity capability is an hour not spent on the thing that actually grows the business. This is **opportunity cost** expressed at the capability level: the true cost of building an internal ticketing system is not just the dev-months it consumes, it is the roadmap feature that did not ship because those dev-months went elsewhere. ## The trade-offs, in both directions The trade-offs run in opposite directions depending on which side of the line a capability sits. - **Buying a commodity capability** gets a team of specialists' full-time attention, a much faster time-to-market (days or weeks of integration instead of months or years of building and hardening), and the benefit of edge cases the vendor has already hit and fixed on someone else's dime -- a payments vendor has already solved the currency-rounding bug this team has not encountered yet. The cost is inheriting the vendor's roadmap, pricing power, and outages, plus the integration and data-migration cost of ever leaving. - **Building a differentiating capability** gets full control over behavior, roadmap, and intellectual property, and the ability to move at the business's pace rather than a vendor's release cadence. The cost is permanently owning the operational burden: security patching, on-call, and the multi-year cost of a team a vendor would otherwise have spread across hundreds of customers. ## Failure modes - **The most common failure mode** is misclassification in the 'build' direction: a team builds a capability because 'our workflow is special,' when in reality eighty percent of the requirement is standard and only a thin slice is genuinely unique. In production this shows up years later as a home-grown authentication system missing security features a mature identity provider shipped long ago, because the team never had spare bandwidth to catch up while also serving the business. - **The mirror failure** is buying a genuinely differentiating capability off the shelf and discovering the vendor cannot or will not build the one feature that would have been the moat, so the product quietly converges toward parity with every other customer of that vendor. - **A subtler failure** is treating the decision as permanent: a build call that was correct in year one can become a maintenance tax by year five once the market catches up and the capability commoditizes underneath the team that built it. ## A worked example **A concrete worked example.** a fintech startup building a lending product needs identity verification and KYC, a core underwriting risk model, and internal support ticketing. KYC and ticketing are commodity -- mature vendors like Persona and Zendesk sell exactly this, so buying gets the team to market in weeks and the vendor's compliance certifications come for free. The underwriting risk model is the actual product; it is what makes the startup's loan approvals faster, cheaper, or more accurate than a bank's, so it gets built in-house, with engineering time protected specifically for it. If a generic underwriting-as-a-service vendor later emerges that is genuinely as good, the calculus might flip -- but only once the capability has actually commoditized, not merely because a vendor's marketing claims it has.

  • Give an example of a capability that used to be differentiating and became commodity over time.
    Running your own data centers and provisioning physical servers was a meaningful competitive advantage for early web companies in the mid-2000s, since it required real infrastructure expertise. By the mid-2010s, cloud providers like AWS had commoditized elastic compute so thoroughly that owning data centers became a cost and distraction rather than an edge, and most companies migrated to buying compute as a utility instead.
  • Can a capability be commodity for one company and differentiating for a competitor in the same industry?
    Yes -- classification depends on the business model, not the industry label. A generic e-commerce checkout flow is commodity for most retailers, but for a company whose entire value proposition is a novel one-click or embedded checkout experience, that same capability is the differentiator and is worth building and owning.
  • How does team size or stage of the company affect this decision even for a differentiating capability?
    A very early-stage startup may still choose to buy or heavily leverage an existing platform even for a differentiating capability, because it lacks the engineering capacity to build and operate anything reliably; the build decision often becomes affordable only once the company has proven the differentiator matters and can staff it properly.

It's like deciding whether to bake your own bread or buy it from a bakery: if you run a sandwich shop and the bread is what people line up for, bake it yourself and control every detail; if bread is just something the sandwich sits on, buy a decent loaf and spend your energy on the filling that actually sells the sandwich.

saying these in an interview costs you the question

  • Treats build-vs-buy purely as a price comparison without discussing what the business actually gets paid for
  • Assumes 'we're unique' without checking whether a mature vendor market already exists
  • Never mentions that the classification can change over time
  • Recommends building every capability 'for control' regardless of whether it is differentiating
  • Cannot name a real example of a commodity vs. differentiating capability

context

open as a page

When evaluating a third-party library or framework for a new feature, what does checking its 'capability fit' involve, and why is comparing feature checklists on vendor or project websites not enough on its own?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Capability fit means testing whether a tool actually solves your specific problem in practice, not just matching feature names on a list. A checklist can say 'yes' to a feature while the real behavior, limits, or edge cases don't match what you need.

open as a page

In cost modeling for a system, what is the difference between capital expenditure (capex) and operating expenditure (opex), and why does moving from on-premises data centers to cloud computing typically shift spending from capex to opex?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Capex is money spent upfront to buy something you own long-term, like servers. Opex is ongoing spending to run things day to day, like a monthly cloud bill. Cloud turns a big upfront hardware purchase into a recurring subscription-like cost.

open as a page

When integrating two systems, what is the core difference between a synchronous request-reply call and an asynchronous message-based call, and what does each cost you?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Sync: caller waits for an instant answer, like a phone call. Async: caller sends a message and moves on, checking back later, like a letter. Sync is simple but ties your uptime to the other side; async is more resilient but harder to reason about.

open as a page

What is the difference between an SLA, an SLO, and an SLI, and how do the three fit together when you design a system's reliability target?

level: juniorimportance: must knowfreq 85%

basics

~20 s

SLI is what you actually measure (e.g., % of successful requests). SLO is the internal goal for that measurement (e.g., 99.9%). SLA is the external, often contractual, promise to customers — usually set looser than the SLO so you have margin.

open as a page

In requirements analysis, what is the difference between a functional requirement and a non-functional requirement, and why does an architect need to treat them differently?

level: juniorimportance: must knowfreq 85%

basics

~20 s

Functional requirements describe what the system must do, e.g. 'let a user reset their password.' Non-functional requirements describe how well it must do it, e.g. 'password reset must respond in under 2 seconds for 99% of requests.'

open as a page

When architecting a solution for a business requirement, why is it considered bad practice to design and propose only a single technical option?

level: juniorimportance: must knowfreq 65%

basics

~10 s

Because comparing choices lets you see costs and risks you'd miss with just one idea, and it lets stakeholders make an informed pick instead of trusting a single opinion.

open as a page

Why would a solution architect create a high-level context diagram for a business sponsor and a separate, more detailed container/component diagram for the engineering team, instead of just handing everyone the same architecture diagram?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Different people care about different things. A sponsor wants to know cost, risk, and business value; an engineer wants to know how the system's pieces talk to each other in detail. One diagram can't do both jobs well.

open as a page

In vendor selection, what is the difference between an RFI (Request for Information) and an RFP (Request for Proposal), and why do organizations typically issue an RFI before an RFP?

level: juniorimportance: must knowfreq 70%

basics

~20 s

RFI asks vendors 'tell us about yourselves and your product' to narrow a long list. RFP asks the shortlisted vendors 'here's exactly what we need, give us a formal, priced proposal.' RFI happens first so you don't waste a detailed RFP on unqualified vendors.

open as a page

A company buys a COTS CRM platform and heavily customizes it to match its exact sales process by directly modifying core workflow logic. Eighteen months later, upgrading to the vendor's new major version breaks most of the customizations. What went wrong in how the customization was done, and how should it have been managed instead?

level: middleimportance: must knowfreq 65%

basics

~20 s

They edited the software's core parts directly instead of using the vendor's safe, supported ways to customize it, like plugins or settings. The next update assumed nobody touched the core, so it broke. Use only supported extension points instead.

open as a page

A team estimates that building an in-house feature-flagging system will take six engineer-weeks, versus paying a SaaS feature-flag vendor a few thousand dollars a month. Why is comparing 'six engineer-weeks' to the monthly subscription price the wrong comparison, and what should be compared instead?

level: middleimportance: must knowfreq 72%

basics

~20 s

Six weeks isn't the real cost. After launch, someone must keep fixing and running it forever, and those engineers could've built something else instead. Compare the full years-long cost, including that lost opportunity, not just build time.

open as a page

What signals should you check to judge whether an open-source library is mature and well-supported enough to depend on in production, beyond its star count on a code-hosting platform?

level: middleimportance: must knowfreq 70%

basics

~20 s

Look past star counts at things like how often it's updated, how fast bugs get fixed, how many people maintain it, and whether real companies use it in production. A popular but abandoned project is riskier than a smaller, active one.

open as a page

When comparing two architecture options — for example, building a custom service in-house versus buying a vendor SaaS product — how would you construct a total cost of ownership (TCO) model to make a fair comparison?

level: middleimportance: must knowfreq 80%

basics

~20 s

List every cost each option will cause over the same time period — not just the price tag, but also setup, running, and people costs — then add them up and compare like for like over the same number of years.

open as a page

When two teams integrate over an API, what makes a good API contract, and how do you evolve it without breaking existing consumers?

level: middleimportance: must knowfreq 75%

basics

~20 s

A contract is the agreed shape of requests and responses between two systems - like a form both sides fill out the same way. To change it safely, add new optional fields instead of removing or renaming old ones, so old callers keep working.

open as a page

You're told a service must meet a 99.95% availability NFR. Walk through how you'd translate that number into concrete architectural decisions.

level: middleimportance: must knowfreq 80%

basics

~20 s

99.95% means about 4.4 hours of downtime allowed per year. To hit that you typically need redundancy (multiple instances/zones), automatic failover, health checks, no single point of failure, and enough capacity headroom to survive losing a node without falling over.

open as a page

You're gathering requirements for a new claims-processing system from customer support, compliance, and engineering leads, and each group hands you a different, sometimes contradictory, list of needs. What elicitation techniques would you use, and how do you resolve the conflicts before they reach the architecture?

level: middleimportance: must knowfreq 78%

basics

~20 s

Talk to each group separately first (interviews, workshops) to understand their real needs, then bring the conflicting groups together with a shared list of options and trade-offs so they negotiate and agree, usually prioritizing by business impact, instead of the architect silently picking a side.

open as a page

You're comparing three candidate solutions using a weighted decision matrix with criteria like cost, time-to-market, and scalability. Walk through how you'd build and use it, and name one way it commonly gets misused.

level: middleimportance: must knowfreq 75%

basics

~20 s

List the things that matter (like cost and speed), give each a weight for importance, score each option on each thing, multiply and add up the scores, and the highest total is the suggested winner -- but you still sanity-check it, because the weights themselves were a judgment call.

open as a page

You need a non-technical steering committee to approve a choice between two architecture options - for example, building a custom integration in-house versus buying a third-party platform. How do you present the trade-off so they can actually make an informed decision and sign off?

level: middleimportance: must knowfreq 70%

basics

~20 s

Translate the technical trade-off into things they care about: cost, time, risk, and what the business gets or gives up with each option. Give a clear recommendation, not just a list of pros and cons.

open as a page

How does a weighted scoring matrix work for comparing vendor proposals, and what's the most common way teams misuse it to justify a decision they'd already made?

level: middleimportance: must knowfreq 65%

basics

~20 s

You list the things that matter (price, features, support...), give each a weight based on importance, score every vendor on each item, multiply and add up. The misuse: picking the weights AFTER seeing the favorite vendor score well, to make the math match a decision already made.

open as a page

What are the distinct mechanisms by which an organization becomes 'locked in' to a vendor after a buy decision, and what concrete steps can be taken at contract-signing time to reduce the switching cost later?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Lock-in has several forms: data trapped in the vendor's format, systems wired to their specific APIs, staff knowing only that vendor's tools, or a contract penalizing exit. Reduce it by negotiating data export and standard interfaces before signing.

open as a page

How does 'vendor lock-in' accumulate when you adopt a third-party component or managed service, and what concrete architectural techniques reduce your exposure without giving up the component's benefits entirely?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Lock-in happens when a product uses so many of a vendor's unique features that switching away becomes expensive or impossible. You reduce it by using standard interfaces where you can, isolating vendor-specific code behind your own abstraction layer, and keeping your data in portable formats.

open as a page

As a solution architect practicing FinOps, how do you estimate and control cloud infrastructure run costs for a new architecture, and what mechanisms (tagging, showback/chargeback, reserved capacity) do you rely on?

level: seniorimportance: must knowfreq 75%

basics

~20 s

FinOps means treating cloud spend like a thing you actively manage, not a surprise bill. You estimate cost per unit of usage before building, label every resource so you know which team or feature caused a cost, show teams their own spend so they feel responsible for it, and commit to steady baseline capacity in advance to get a discount.

open as a page

A team is integrating a pricing service with three downstream consumers: a checkout UI that needs the current price before rendering, an analytics warehouse that aggregates prices nightly, and a recommendation engine that reacts to price drops. How should the choice between request-reply and event-driven integration differ across these three consumers, and what do you give up by picking event-driven for all three anyway?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Pick request-reply when a consumer needs the answer right now to keep working, like a checkout screen. Pick events when a consumer just needs to know 'something changed' and can react later, like analytics or recommendations. Using events everywhere adds delay and complexity where you didn't need it.

open as a page

A business NFR states the system must scale to support 10x current traffic within 12 months without a redesign. What concrete architectural decisions does that requirement drive, and how do you validate the system will actually get there?

level: seniorimportance: must knowfreq 75%

basics

~20 s

It pushes you toward stateless services that can scale horizontally, autoscaling on real metrics, partitioning/sharding data stores so no single DB instance is the ceiling, and caching/async processing to cut load on the bottleneck. You validate it with load testing at the target scale, not by assuming the design works.

open as a page

What problem does a requirements traceability matrix solve on a solution architecture project, and what specifically tends to go wrong on a multi-year system when a team stops maintaining one?

level: seniorimportance: must knowfreq 65%

basics

~20 s

It's a table linking every requirement to the design, code, and tests that satisfy it, so you can prove nothing important was missed, silently dropped, or built without ever being verified. Skip it and requirements quietly disappear or get marked 'done' without proof they actually work.

open as a page

You recommended Option B over Option A using a documented trade-off analysis, but six months later a director asks 'why didn't we just do Option A, it looked cheaper?' What should your original rationale document have contained so this question is easy to answer, and what's commonly missing from real-world rationale write-ups?

level: seniorimportance: must knowfreq 60%

basics

~20 s

A good rationale write-up explains what you compared, what mattered most and why, and what you'd have to see change to reconsider -- so months later anyone can read it and understand the decision without you having to remember or re-explain it from memory.

open as a page

A security lead insists on mutual TLS across every internal service call before launch, while the product owner needs to ship in two weeks and says that timeline is non-negotiable for a committed customer date. Both have legitimate authority over their domain. How do you drive this to an actual decision instead of a stalemate?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Get both people talking about the actual risk and the actual cost of delay, not just their positions. Find a middle option, like phasing in the strongest protection first, and get someone with the authority to accept the trade-off to make the call.

open as a page

What concrete forms does vendor lock-in take beyond 'it's hard to switch,' and what should an exit strategy documented during vendor selection actually contain?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Lock-in isn't just a feeling - it's specific things like your data being stuck in a proprietary format, custom code that only works with that vendor's API, or contract penalties for leaving. An exit strategy written down before you sign lists exactly how you'd get your data out, how long it would take, and what it would cost, so you're not stuck figuring it out under pressure later.

open as a page

When running a proof-of-concept 'bake-off' between two or three shortlisted vendors' products, what design choices keep the comparison fair, and what commonly makes POC results look good in the trial but misleading once the product is actually in production?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Give every vendor the same test, same data, same amount of help, and same success criteria decided beforehand - otherwise whoever gets more attention or an easier task wins unfairly. POCs often look great because vendors hand-hold and use small clean data, which doesn't reflect real production scale, messy data, or the support you'll get after you've already paid.

open as a page

When adding an open-source library to a commercial product, why can a permissive license (like MIT or Apache 2.0) and a copyleft license (like GPL or AGPL) lead to very different legal outcomes for your codebase, even if the two libraries do the exact same technical job?

level: middleimportance: should knowfreq 45%

basics

~20 s

Some licenses (MIT, Apache) let you use code freely in a closed-source product. Others (GPL, AGPL) can require you to release your own product's source code, or at least the parts that link to that library, under the same license if you distribute or run it as a service.

open as a page

showing 1–30 of 52