skip to content

Solution Architecture

Turning a set of requirements into one concrete, buildable design: analysing needs, selecting components, defining integrations and NFRs, modelling cost, comparing options and getting stakeholders to agree. It sits between enterprise architecture and the code.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

What are the main enterprise software licensing models — per-core, per-user/seat, and consumption-based — and what pitfalls do they each create when you try to fold them into a TCO model?

level: middleimportance: should knowfreq 55%

basics

~20 s

Some software charges by how much hardware you run it on (per-core), some by how many people use it (per-user), and some by how much you actually use it (consumption-based). Each one can surprise you: scaling your servers, hiring more staff, or growing usage can all silently blow up your bill.

open as a page

When choosing a data exchange format for an integration - say, JSON versus a binary format like Protobuf or Avro - what are you actually trading off, and how does each handle schema evolution over time?

level: middleimportance: should knowfreq 55%

basics

~20 s

JSON is text you can read with your eyes and is easy to debug, but it's bigger and slower to parse. Formats like Protobuf or Avro are compact and fast but need a shared schema file and special tools to read.

open as a page

A product NFR states an API must respond in under 300ms at p99. The request fans out to a database call, an auth check, and two downstream microservices. How do you turn that single latency number into a design constraint across the call chain?

level: middleimportance: should knowfreq 65%

basics

~20 s

Split the 300ms budget across every hop in the request's path (network, auth, each downstream call, your own processing), leaving slack for tail variance, not just averages. If the sum of the pieces plus slack doesn't fit, redesign — e.g., call downstreams in parallel instead of one after another, cache, or drop something from the critical path.

open as a page

When modeling a 'place order' use case for an e-commerce checkout, what belongs in the main success scenario versus the alternate and exception flows, and why does keeping them separate matter for the resulting architecture?

level: middleimportance: should knowfreq 60%

basics

~20 s

The main flow is the happy path: everything goes right and the order completes. Alternate and exception flows are what happens when something differs or goes wrong, like a declined card or an out-of-stock item. Keeping them separate matters because each exception often needs its own piece of architecture to handle it.

open as a page

When documenting solution options for a technical decision, you list both 'constraints' and 'assumptions' separately. What's the practical difference between the two, and why does conflating them cause problems later?

level: middleimportance: should knowfreq 55%

basics

~20 s

A constraint is a hard rule you can't change (like a budget cap or a law), while an assumption is a guess you're making that could turn out wrong (like expecting traffic to stay under a certain level). Mixing them up means nobody knows what's truly fixed versus what might need to be revisited.

open as a page

Before you even get to a decision meeting, how do you figure out which stakeholders' priorities you actually need to reconcile on an architecture decision, out of everyone who has an opinion?

level: middleimportance: should knowfreq 55%

basics

~20 s

Sort people by how much power they have over the decision and how much they care about it. Spend your real effort on the ones high in both; keep others informed without dragging them into every debate.

open as a page

When you call the reference customers a vendor gives you, why is that alone not enough to trust reference checks, and what should you actually do to get more honest signal?

level: middleimportance: should knowfreq 50%

basics

~20 s

The vendor obviously hands you their happiest customers, so those calls will sound great almost by design. To get honest signal, find some references yourself, ask specific pointed questions about problems and support response times instead of 'are you happy,' and talk to people who actually use the product day to day.

open as a page

Before committing to a candidate component, why is a short technical spike or proof-of-concept often more reliable than a written comparison for estimating real integration effort, and what hidden costs does it tend to surface?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Reading docs and comparison tables can't show you the real friction of hooking a component into your actual system - things like data format mismatches, auth setup, or deployment quirks. Building a small working version for a day or two surfaces those surprises before you're committed.

open as a page

When you're asked to justify an architecture decision — say, migrating a monolith to microservices — with a cost-benefit or ROI analysis, what should that analysis actually contain, and what hidden costs and benefits are easy to leave out?

level: seniorimportance: should knowfreq 60%

basics

~20 s

You compare what the change costs (money, time, risk of things breaking) against what it saves or earns (faster releases, lower running costs, fewer outages) over the same time period, and show the payback point. The easy mistake is only counting the obvious costs and only counting the hoped-for benefits, without pricing the migration pain or the chance the payoff never fully arrives.

open as a page

Why does an integration contract for a 'create order' operation typically need to define an idempotency key, and what breaks in production if it doesn't?

level: seniorimportance: should knowfreq 60%

basics

~20 s

An idempotency key lets a caller safely retry a request without accidentally doing it twice - like writing a unique order number on a form so resubmitting it doesn't create a second order. Without it, network retries can cause duplicate orders, charges, or emails.

open as a page

A solution architecture must satisfy a security NFR such as 'all sensitive data encrypted at rest and in transit, with defense in depth against a compromised application server.' What architectural decisions does this drive, and what does it cost elsewhere in the design?

level: seniorimportance: should knowfreq 60%

basics

~20 s

It means encrypting data both while stored (at rest) and while moving over the network (in transit, e.g. TLS), and not relying on just one layer of defense — network segmentation, least-privilege access, and secrets management so that if one layer (like the app server) is breached, the attacker still can't reach everything. The cost is extra latency, operational complexity, and key-management overhead.

open as a page

Beyond functional and non-functional requirements, what counts as a 'constraint' during solution architecture requirements analysis, for example a rule that all customer data must stay within a specific country for regulatory reasons, and how should an architect handle constraints differently from ordinary requirements?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A constraint is something you're not allowed to choose, like a law, an existing system you must integrate with, or a fixed budget or deadline. Requirements describe what to build; constraints narrow down which solutions are even legal or possible before you start designing.

open as a page

You need formal sign-off on an architecture decision from a stakeholder who keeps deferring - not rejecting it outright, just never quite approving it - and the deadline to start implementation is now days away. What do you do?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Find out what's really holding them back - maybe they don't understand it, don't trust it, or have a concern they haven't said out loud. Address that directly, and if they still won't decide, take it to someone who can make the call.

open as a page

How should a principal-level architect structure a build-vs-buy decision so the organization doesn't get stuck with a stale answer as the surrounding vendor market and the company's own strategy evolve over several years?

level: principalimportance: should knowfreq 45%

basics

~20 s

Don't treat build-vs-buy as permanent. Re-check major decisions yearly, watching whether the market or your own strategy has shifted, and design systems so switching direction later -- building what you bought, or vice versa -- doesn't mean starting over.

open as a page

What are the most common, systemic ways that TCO and cost-benefit models for cloud architecture decisions turn out to be wrong in practice, and how would you design a modeling process to catch these failures before the decision is locked in?

level: principalimportance: should knowfreq 40%

basics

~20 s

Cost models are usually wrong in predictable ways: people forget ongoing running costs, forget data-transfer/egress fees, assume the migration will go smoothly and on time, and only check the model once instead of updating it as reality unfolds. Fixing this means building in checkpoints, naming every assumption, and revisiting the numbers after the decision, not just before.

open as a page

When integrating a modern service with a legacy system (or an external system your team doesn't control) whose domain model is a poor fit for yours, what does an anti-corruption layer do, and what does it cost to maintain?

level: principalimportance: should knowfreq 45%

basics

~20 s

An anti-corruption layer is a translation wall between your system and someone else's messy or outdated one - it converts their concepts into yours at the boundary, so their quirks don't leak into your code. It costs extra code to build and keep updated.

open as a page

A solution architecture must satisfy a compliance NFR requiring EU customer data to stay within EU data centers and be retained for a fixed audit period, while the product also has a global latency target and a cost ceiling. How do you architect for a compliance NFR that actively conflicts with other NFRs, and how do you decide the trade-off?

level: principalimportance: should knowfreq 45%

basics

~20 s

You typically split data by residency requirement (data-partitioning by region), keeping EU customer data physically in EU infrastructure while allowing non-regulated data or metadata to flow globally. That usually means non-EU users hitting EU-hosted data pay a latency penalty, or you replicate read-only, non-sensitive views elsewhere — the trade-off is decided by ranking compliance as a hard constraint (non-negotiable, legal risk) versus everything else as tunable.

open as a page

A weighted decision matrix comparing two solution options for a platform migration comes out nearly tied, and you suspect the losing side is quietly lobbying to change the weights to flip the result. As the architect accountable for the final recommendation, how do you handle both the near-tie and the political pressure?

level: principalimportance: should knowfreq 35%

basics

~20 s

Don't let people quietly change the scoring after the fact to get the answer they want. Instead, be upfront that the numbers are close, bring the real disagreement into the open, and make the final call based on judgment and the most important unscored risks -- then write down clearly why.

open as a page

You're the lead architect on a program with a dozen or more stakeholder groups across several teams, running for a year or more. How do you keep everyone aligned over that time without becoming a bottleneck that every decision has to pass through personally?

level: principalimportance: should knowfreq 40%

basics

~20 s

Don't try to be in every conversation. Set up clear rules for who decides what, write decisions down so people can find them without asking you, and check in on a regular schedule instead of only when something breaks.

open as a page

What signals indicate a vendor might not be viable long-term, financially or strategically, and how should that risk change how you structure the contract and architecture around their product?

level: principalimportance: should knowfreq 45%

basics

~20 s

Watch for things like the vendor burning cash with no clear path to profit, losing key customers or executives, being a tiny player in a market big companies are entering, or getting acquired by a competitor of yours. If a vendor looks risky, you build in ways to leave fast, like source-code escrow, shorter contracts, and less deeply integrated architecture, rather than betting everything on them surviving.

open as a page

Beyond checking a candidate component's own functionality and maintenance health, how should an architect evaluate its supply-chain security risk before adopting it, and what specific attack patterns make this a distinct concern from ordinary maturity checks?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Beyond checking if a component is well-made, you also need to check if it, or its own dependencies, could be a way for an attacker to sneak malicious code into your system - through a hijacked update, a fake similarly-named package, or a compromised maintainer account.

open as a page

A stakeholder asks for a complete requirements specification with a full traceability matrix before any design work starts, on a product that's expected to pivot significantly based on early user feedback. As the architect, when do you push back on that level of upfront requirements analysis, and what do you propose instead?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Full upfront specs make sense for fixed-scope, regulated, or high-stakes systems, not for something expected to change fast based on user feedback. Push back by proposing to lock down only the non-negotiables, such as security and compliance, up front and let the detailed feature requirements emerge iteratively as you learn from real usage.

open as a page

showing 31–52 of 52