How does team cognitive load act as a constraint on architectural boundaries, and what do you do when a team is overloaded?
answer
- Intrinsic / extraneous / germane load
- Boundary must fit one team's head
- Kill extraneous first: platform, tooling, decommission
- Enabling team lowers intrinsic load temporarily
- Split the domain last — it's the expensive move
basics
~20 sA team can only hold so much in its head. If it owns more systems and domains than it can understand, quality and speed drop. So you size each team's responsibilities to fit, and shrink their load by simplifying, delegating to a platform, or splitting the domain.
solid answer
~50 sTeam Topologies treats cognitive load — the sum of what a team must know to build and run its systems — as the real limit on how much a team can own, and therefore as a constraint on where architectural boundaries belong. Cognitive load has three parts: intrinsic (fundamental skills: the language, testing, the domain), extraneous (accidental burden: deployment ceremony, brittle environments, undocumented tooling), and germane (the valuable domain thinking you actually want them spending on). Overload shows up as slow lead time, rising change-failure rate, shallow on-call knowledge, unfixed toil, and single points of person-failure. Remedies in order of preference: eliminate extraneous load (self-service platform, better tooling, delete dead systems); reduce intrinsic load with an enabling team's coaching; move rare deep-expertise components to a complicated-subsystem team; and only then split the domain and hand a bounded context to a new stream-aligned team. Crucially, this means a boundary that looks clean on a diagram is wrong if the owning team can't hold it.
go deeper
Say a team can only understand so much; if it owns too many systems, quality and speed drop.
Distinguish intrinsic/extraneous/germane load, list overload symptoms, and name platform + smaller scope as remedies.
Explain that cognitive load bounds how much a boundary can enclose, order the remedies from cheap to expensive, and cover the standardization-vs-autonomy trade-off.
Treat load as an org-wide budget: fracture-plane selection, paved-road strategy, legacy decommission programs, platform-as-product funding, and how to detect overload from delivery and incident data before it shows up as attrition.
## What cognitive load means here Cognitive load theory (John Sweller) distinguishes three kinds of mental effort. *Team Topologies* applies it at team scale: - **Intrinsic load** — effort inherent to the task and required skills: knowing the language, the framework, how to test, how the domain works. Reduced by training, coaching, and hiring, not by tooling. - **Extraneous load** — accidental effort caused by how work is organized: hand-rolled deploy scripts, flaky environments, undocumented internal tools, ticket queues, five different CI systems. This is pure waste and the first thing to attack. - **Germane load** — effort spent on the valuable thinking: understanding the business domain, designing the model. This is what you want maximized. A team's total capacity is roughly fixed. Every subsystem it owns consumes some of that budget in all three categories, permanently — not just while building it, but for running, patching, and answering questions about it. ## Why it constrains architecture Conway's Law says team boundaries become architectural boundaries. Cognitive load says how *much* a team boundary can enclose. Put together: > An architecture is only viable if each boundary fits inside one team's cognitive-load budget. A "beautiful" decomposition where one team owns nine services across three unrelated domains will degrade: the team will context-switch, learn each system shallowly, take shortcuts across the boundaries they own (because internal shortcuts are cheap), and the services will drift into a tangle. Conversely, boundaries so small that a team owns one trivial service leave capacity unused and multiply cross-team coordination. This inverts a common habit: instead of "design the system, then staff it", you check each candidate boundary against "can one long-lived team realistically own this — build, run, and be on call for it?" If not, the boundary is wrong, or the team's other responsibilities must go. ## Symptoms of overload - Lead time for changes rises; work sits in progress waiting for the one person who knows. - Change failure rate and MTTR climb; incidents in the less-loved systems take much longer. - On-call is dreaded; runbooks are stale; the team can't explain half of what it owns. - Toil and upgrades are perpetually deferred; dependency versions rot. - "Bus factor of one" per subsystem — knowledge is siloed inside the team. - The team refuses or slow-walks new work not because of capacity in hours but because of unfamiliarity. A useful low-tech measurement: ask the team to list every domain and subsystem it owns and self-rate its comfort with each; the count of low-comfort items is your signal. Cognitive load is not measurable in story points; it is measured in *distinct things that must be understood*. ## Remedies, in order 1. **Delete or consolidate.** The cheapest load reduction is decommissioning a system nobody needs, or merging two near-duplicate services. 2. **Remove extraneous load with a platform.** A self-service deployment pipeline, standard runtime, and turnkey observability strip the same accidental burden from every stream team at once. This is the *purpose* of a platform team — not centralization for its own sake. 3. **Standardize.** Fewer languages, frameworks, and deployment shapes lowers intrinsic load across the org, at the cost of local optimization. A deliberate trade. 4. **Bring in an enabling team.** Temporary coaching (test automation, security, performance) raises capability so the same scope costs less effort. Time-boxed — a permanent helper becomes a dependency. 5. **Extract a complicated-subsystem team.** If part of the scope needs genuinely rare expertise (an optimization solver, a codec, an ML pipeline), concentrate it in one specialist team so stream teams consume it as a service. 6. **Split the domain.** Last and most expensive: identify a fracture plane (business subdomain, change cadence, compliance, user persona), carve out a bounded context, and give it to a new long-lived stream-aligned team — with its own data and pipeline, or you'll create a distributed monolith. Adding people to the existing team is *not* on the list as a first move: beyond roughly 8-9 people the internal communication cost grows superlinearly, so a bigger team does not buy proportionally more cognitive capacity; it usually means the team informally splits anyway. ## Trade-offs and edge cases - **Standardization vs. autonomy.** Lower intrinsic load through uniformity conflicts with teams choosing the best local tool. Most orgs pick a "paved road" that is optional-but-excellent. - **Platform load is real too.** The platform team itself has a cognitive-load budget; "thinnest viable platform" exists so the platform doesn't grow into its own unmaintainable monolith. - **Legacy nobody wants.** Load from legacy systems is often invisible until an incident. Explicitly account for it, and prefer decommission plans over quiet ownership. - **Small organizations.** With 15 engineers, everyone communicates cheaply; formal load management is overhead. The constraint bites as headcount and system count grow. - **Domain vs. technology load.** A team can absorb a wide domain with familiar technology, or narrow domain with unfamiliar technology, but rarely both at once — useful when planning a migration. ## The one-line takeaway Cognitive load turns "where should the boundaries be?" from a purely technical question into a socio-technical one: a boundary is correct when one stable team can own everything inside it, end to end, without exceeding what it can hold in its head.
- Why isn't 'add more people to the team' the first remedy for cognitive overload?Communication paths grow roughly quadratically with team size, so past about 8-9 people the coordination overhead cancels the added capacity, trust degrades, and the team informally fragments. You also haven't reduced the number of distinct things that must be understood — you've just spread shallower knowledge across more people.
- How would you actually measure a team's cognitive load?There's no clean metric. Practical proxies: have the team enumerate every domain and subsystem it owns and self-rate familiarity; count distinct runtimes/languages/data stores; track on-call pages per system and time-to-diagnose; watch lead time and change failure rate per subsystem. The signal is the count of things owned but poorly understood, not hours of work.
- How does cognitive load interact with the decision to build an internal platform?The platform's justification is precisely extraneous-load removal for many stream teams at once. That implies it must be self-service, well-documented, and optional (teams adopting it because it's easier), and kept 'thinnest viable' so the platform team itself doesn't become overloaded or a ticket-based bottleneck.
A team's cognitive load is like a backpack with a fixed weight limit. You can repack it (standardize), take out gear someone else can carry (platform, specialist team), or take a shorter route (smaller domain) — but you cannot just keep adding rocks and expect the same pace.