skip to content

Codebase & Dependency Organization

How code is laid out and how its dependencies are managed: one repo or many, how package managers resolve versions, and how packages and modules are arranged inside a project. It is the day-to-day structure everyone on the team lives in.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

explore

questions

99 · 3 sections

In a monorepo build system, what does it mean for a build task to get a 'cache hit', and why does that make rebuilds faster?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A cache hit means the build system already ran this exact task before with the exact same inputs, so it reuses the saved output instead of redoing the work — like reheating leftovers instead of cooking again.

open as a page

In a monorepo containing hundreds of independently deployable projects, why would a CI pipeline choose to only build and test the projects 'affected' by a given code change, rather than running the full build and test suite on every commit?

level: juniorimportance: must knowfreq 75%
basics
~20 s

Because rebuilding and testing everything every time gets too slow as the codebase grows. Affected detection figures out which projects a change could actually break and only runs those, so a small change gets fast feedback instead of a multi-hour full run.

open as a page

In a shared monorepo hosted on GitHub or GitLab, what does a CODEOWNERS file do, and how does it change what happens when someone opens a pull request?

level: juniorimportance: must knowfreq 65%
basics
~10 s

A CODEOWNERS file maps folders/files to people or teams. When a PR touches those files, the platform automatically asks the listed owners to review, and can require their approval before merge.

open as a page

In a monorepo containing many internal projects, what does it mean to have a single source of truth for a shared dependency's version, and why do teams enforce it instead of letting each project pin its own version independently?

level: juniorimportance: must knowfreq 65%
basics
~10 s

One place in the repo says "we use version X of this library" for everyone, so all teams building together use the same code and can't accidentally clash with incompatible versions.

open as a page

In a monorepo containing dozens of separate packages, why do teams adopt a dedicated build orchestrator (like Nx or Turborepo) instead of just writing a shell script that runs `build` and `test` in every package folder?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A shell script rebuilds and tests every package every run, which gets slow fast. A monorepo orchestrator tracks changes and dependencies between packages, so it only reruns what's actually affected and reuses cached results for the rest.

open as a page

Imagine your application directly depends on library A and library B. Library A internally requires a shared library C at version 1.0, while library B internally requires that same library C at version 2.0. Only one version of C typically ends up available at runtime. What is this situation called, and why can't both versions simply be used side by side?

level: juniorimportance: must knowfreq 70%
basics
~20 s

It's called a 'diamond dependency conflict.' Two things your app needs each secretly need a different version of the same third thing, but most systems can only load one version of a library at once, so something has to give.

open as a page

In a Node.js project using npm, package.json lists dependency versions as semver ranges like ^18.2.0. What extra file does npm generate, and what specific problem does it solve that the ranges alone can't?

level: juniorimportance: must knowfreq 85%
basics
~20 s

A lockfile records the exact version of every package (including indirect ones) that got installed, so everyone who installs later gets the identical set of files instead of whatever the version ranges happen to resolve to that day.

open as a page

In Semantic Versioning (SemVer), a library bumps its published version from 2.3.5 to 2.4.0, and later from 2.4.0 to 3.0.0. What does each of those two jumps tell a consumer about what changed, and why does that difference matter when deciding whether to upgrade?

level: juniorimportance: must knowfreq 80%
basics
~10 s

SemVer is MAJOR.MINOR.PATCH. 2.3.5→2.4.0 (MINOR) means new features were added, safe to upgrade. 2.4.0→3.0.0 (MAJOR) means something incompatible changed - your code might break, so check the changelog first.

open as a page

A company publishes an internal package called `acme-auth-utils` on a private registry, but a developer's laptop still has the default public registry configured. One day a build silently pulls in a different `acme-auth-utils` package that was never written by anyone at the company. What attack is this, and how does the resolution logic actually let it happen?

level: juniorimportance: must knowfreq 60%
basics
~20 s

An attacker publishes a public package with the exact same name as a company's private internal package. If a build tool checks the public registry and grabs whichever version looks newest/highest, it installs the attacker's fake package instead of the real internal one — no phishing or malware download needed, just a naming collision.

open as a page

In a build tool's dependency manifest (like package.json, pom.xml, or build.gradle), what is the difference between a 'direct' dependency and a 'transitive' dependency, and where do transitive dependencies come from?

level: juniorimportance: must knowfreq 75%
basics
~10 s

A direct dependency is a library you add yourself. A transitive dependency is a library that gets pulled in automatically because one of your direct dependencies needs it to work.

open as a page

In a layered application, what does it mean for source-level dependencies to 'point inward', and why does this matter for the codebase's design?

level: juniorimportance: must knowfreq 65%
basics
~20 s

Code in the core business logic should not need to know about outer layers like databases or web frameworks. Only the outer layers should import and depend on the inner ones, never the other way around.

open as a page

In Java, what's the difference between declaring a class `public` versus leaving it with default (package-private) access, and why would a team choose package-private for an internal helper class?

level: juniorimportance: must knowfreq 65%
basics
~20 s

Public means any code anywhere in the program can use the class. Default/package-private means only code living in the same folder (package) can use it. Teams hide helper classes as package-private so other parts of the codebase can't accidentally depend on internal details that might change.

open as a page

In a multi-module project, what does it mean to separate a module into a public 'API' part and a hidden 'implementation' part, and why would a team bother doing that instead of just putting everything in one module?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Split the module into two: a small public part other code depends on (the API), and a bigger hidden part with the real logic (the implementation). This lets you change the hidden part freely without breaking dependents.

open as a page

What is the difference between organizing a codebase's packages by technical layer (e.g. separate `controller`, `service`, and `repository` packages) versus organizing them by feature (e.g. a package per business capability like `orders` or `users`)?

level: juniorimportance: must knowfreq 85%
basics
~20 s

Layer-based grouping sorts files by their technical role (all controllers together, all services together). Feature-based grouping sorts files by what business thing they do (everything about orders lives in one folder, everything about users in another).

open as a page

In a classic three-tier layout where a `controller` package depends on a `service` package, which in turn directly imports concrete classes from a `repository` package that wraps the database, the code compiles and runs fine. Why is this still considered a dependency-direction problem, and how would you restructure the packages to fix it?

level: middleimportance: must knowfreq 80%
basics
~20 s

It compiles, but the business-logic layer is now stuck to one specific way of talking to the database — you can't swap it or test the logic without the real database. Fix: put an interface next to the business logic and make the database code implement that interface instead of the other way around.

open as a page