skip to content

Writing Clean Code

Code-in-the-small: how a single name, function, comment or error path is written so intent is visible at a glance. This is the material reviewers comment on daily and the level a take-home assignment is judged at.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

In Robert Martin's Clean Code terminology, what is the difference between an "object" and a "data structure"?

level: juniorimportance: must knowfreq 62%

answer

  1. Object = hide data, expose behaviour
  2. Data structure = expose data, no behaviour
  3. Getters+setters everywhere = data structure in disguise
  4. Hybrid = worst of both
  5. Objects add types easily, data adds operations easily

basics

~20 s

An object hides its data and exposes behaviour - you tell it to do something. A data structure exposes its data and has almost no behaviour - other code reads its fields and does the work.

solid answer

~50 s

The line is drawn by what is public. An **object** keeps its representation private and publishes *behaviour*: callers say shape.area() without knowing whether a circle stores a radius or a bounding box. A **data structure** publishes its *data* (public fields, or getters/setters that are just field access with ceremony) and has little meaningful behaviour: a Point{x,y}, a JSON payload, a DTO, a database row. Neither is wrong; they solve opposite problems. Objects let you add new *types* without touching existing callers, because behaviour travels with the type. Data structures let you add new *operations* in one place without touching the types. The mistake is the hybrid: a class with public getters/setters for everything *and* significant business rules. It gets the drawbacks of both - callers reach into its data, yet it hides enough that you cannot treat it as plain data. Pick one side per type, deliberately.

code

pseudocode · 9 lines
pseudocode
// Data structure: exposes data, no behaviour
struct Rectangle { width, height }
function area(r) { return r.width * r.height }   // logic lives outside

// Object: hides data, exposes behaviour
class Rectangle {
  private width, height
  function area() { return width * height }      // logic lives inside
}

go deeper

for a junior

State the one-line contrast (hide data + expose behaviour vs expose data + no behaviour) and give one example of each, e.g. an Account object vs a Point data structure.

for a middle

Add that trivial getters/setters do not create encapsulation, and name the hybrid as the thing to avoid.

for a senior

Frame the choice by expected axis of change - new types favour objects, new operations favour data structures - and describe where in a system each belongs (edges vs core domain).

for a principal

Connect it to the Expression Problem and to architectural boundaries: dumb data at the seams for versioning and serialization, behaviour-rich objects inside; and set team conventions so hybrids do not accumulate.

## The distinction The vocabulary comes from the *Objects and Data Structures* chapter of Robert C. Martin's *Clean Code*, but the idea is older and language-independent. **Object** = hides its representation, exposes behaviour. - Fields are private/internal; nobody outside can see how state is stored. - Public methods are domain verbs: `withdraw(amount)`, `area()`, `render()`, `isEligible()`. - Callers *tell* it what to do rather than pulling out its parts and deciding for it. **Data structure** = exposes its representation, has (almost) no behaviour. - Fields are public, or accessed via trivial getters/setters that add nothing. - Behaviour lives in *other* code: functions, services, procedures that read the fields. - Examples: geometric `Point{x, y}`, a request/response DTO, a row mapped from a database, a parsed JSON document, a struct in a C-style language. ## Getters and setters do not make it an object The single most common misconception: "my fields are private and I wrote getX/setX, therefore it is encapsulated." It is not. If every field has a public getter and setter, the representation is fully visible and fully mutable - you have written a data structure with extra typing. Real encapsulation means an outside caller *cannot* learn how the data is stored. `account.getBalance()` may still be legitimate if balance is part of the abstraction the class promises; `account.getInternalLedgerEntries()` almost certainly is not. ## Why the distinction matters Because it determines *where new code goes* when requirements change: - With objects, adding a new **type** (a new shape, a new payment method) means writing one new class. Existing callers do not change - they already call the abstract method. - With data structures plus procedures, adding a new **operation** (a new report, a new export format) means writing one new function. Existing types do not change. Those are exactly opposite strengths, and the trade-off is symmetric (this is the *data/object anti-symmetry*, and it is the same force behind the Expression Problem in language design and the Visitor pattern). ## Choosing - Domain concepts with invariants to protect (money, an order, a policy) - make them objects: hide the fields, publish verbs. - Transport / edge shapes (API payloads, config, query results, coordinates, events) - make them data structures: plain, public, dumb, easy to serialize. - Boundary rule of thumb: data structures are fine *at* the boundary of the system; convert them into real objects once inside. ## The hybrid anti-pattern A class with public accessors for all its state *and* important business methods is a hybrid. It invites callers to bypass the behaviour (`if (order.getStatus() == PAID && order.getItems().isEmpty())`) while also being too opinionated to serialize or diff cleanly. Hybrids are the usual habitat of train-wreck chains and Law-of-Demeter violations, because callers get used to reaching in.

  • If a class has private fields but a public getter and setter for every one of them, is it an object or a data structure?
    Effectively a data structure. The representation is fully exposed and mutable, so callers can (and will) build logic on its shape. The accessors add ceremony, not encapsulation.
  • Is a DTO carrying data across a service boundary a bad design because it has no behaviour?
    No. A DTO is a deliberate data structure: dumb, serializable, versionable. The mistake would be letting business rules grow on it, turning it into a hybrid.

An object is a vending machine: you press a button and get a drink; the mechanism is sealed. A data structure is a supermarket shelf: everything is visible and you assemble the meal yourself.

saying these in an interview costs you the question

  • Claiming private fields plus getters/setters automatically means encapsulation
  • Saying data structures are always bad / everything must be an object
  • Confusing the term 'data structure' here with algorithmic data structures like hash maps or trees
  • Treating DTOs, config records, and API payloads as design smells
  • Believing objects are always the safer default regardless of how the code will change

context

open as a page

What is the difference between cyclomatic complexity and cognitive complexity as code metrics, and why was cognitive complexity introduced?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Cyclomatic complexity counts the independent execution paths through code, which estimates how many tests you need. Cognitive complexity estimates how hard code is for a human to read: it penalises nesting and ignores structures people find easy.

open as a page

How do guard clauses and early returns reduce the cognitive load of a function, and when is the 'single exit point' rule still justified?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A guard clause handles an invalid or special case immediately and returns, instead of wrapping the real work in an if. That flattens nesting, so the reader stops carrying conditions in their head and the main path stays at the left margin.

open as a page

In code review, why is a comment that explains WHY the code does something usually more valuable than a comment that restates WHAT the code does?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The code already shows what it does; a reader can see that. It cannot show the reason, constraint, or bug that forced this approach. "Why" comments add information; "what" comments repeat it and can go stale.

open as a page

Why does Clean Code advise signalling failures with exceptions rather than returning error codes, and when might error codes still be the better choice?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Error codes force every caller to check the returned value right at the call site, mixing error checks into the happy path. Exceptions move failure handling out of the main flow, and they cannot be silently ignored, so the code reads cleaner and is safer.

open as a page

Clean Code says "don't return null" and "don't pass null". What problems does null cause, and what should you return or accept instead?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Returning null forces every caller to add a null check; one forgotten check crashes at runtime, far from the cause. Instead return an empty collection, a default "do-nothing" object (Null Object), or an explicit optional type. Never pass null as an argument.

open as a page

In Robert C. Martin's Clean Code, what is the "newspaper metaphor" for formatting a source file, and how should it shape the order of functions in that file?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Read a source file like a newspaper article: the name at the top tells you what it is, the first code gives the high-level story, and details get lower-level as you scroll down. So put high-level functions first and their helpers below.

open as a page

In function design, what does the guideline "a function should do one thing" actually mean, and how can you tell that a given function violates it?

level: juniorimportance: must knowfreq 78%

basics

~20 s

It means the function performs a single, clearly nameable job. Signs of violation: you can extract a meaningful chunk into another function that isn't just restating the original, the name needs "and"/"or", or the body mixes unrelated steps.

open as a page

Why are boolean "flag" arguments considered a function-design smell, and what refactorings remove them? Also: what is the practical guidance on how many parameters a function should take?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A boolean parameter means the function does two things — one per branch — and the call site save(user, true) is unreadable. Fix by splitting into two clearly named functions. Prefer 0-2 parameters; 3 is suspicious; 4+ usually means a missing object.

open as a page

What makes a variable, function, or class name "intention-revealing", and why is renaming usually a better fix than adding an explanatory comment?

level: juniorimportance: must knowfreq 80%

basics

~20 s

An intention-revealing name says what the thing is, why it exists, and how it is used — so no comment is needed. Comments drift out of date silently; a name is re-read every time the code is read.

open as a page

What naming conventions distinguish classes, methods, and boolean-returning members, and why does following them consistently matter?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Classes and modules get noun phrases (things): Invoice, PaymentGateway. Methods get verb phrases (actions): send, calculateTotal. Boolean members read as a true/false statement: isEmpty, hasPermission. Readers then parse code as sentences instead of decoding it.

open as a page

What does the Law of Demeter state, and why is a chain like order.getCustomer().getAddress().getCity().toUpperCase() considered a problem?

level: middleimportance: must knowfreq 70%

basics

~20 s

The Law of Demeter says a method should only talk to its immediate neighbours: itself, its own fields, its parameters, and objects it created. Long chains like that one couple your code to the internal structure of three other classes, so any of them changing breaks you.

open as a page

Clean-code guidance splits comments into a small set of justified categories and a longer list of harmful ones. Name several of each and state the single criterion that decides which side a comment falls on.

level: middleimportance: must knowfreq 58%

basics

~20 s

Good: legal/licence headers, intent, clarification of code you can't change, warnings of consequences, TODOs, and public API docs. Bad: redundant restatement, journal/changelog entries, commented-out code, noise, banners, attributions. Criterion: does it tell the reader something the code cannot?

open as a page

What does "define exception classes in terms of the caller's needs" mean, and how do you use wrapping to apply it to third-party libraries?

level: middleimportance: must knowfreq 66%

basics

~20 s

Design exception types around how callers will handle failures, not around where they came from. If a caller treats ten library exceptions identically, wrap them in one of your own exception types thrown from a thin adapter, so the calling code has one catch block and no dependency on the library.

open as a page

What are "disinformative" names and "noise words" in identifiers, and what does the rule "one word per concept" require?

level: middleimportance: must knowfreq 55%

basics

~20 s

A disinformative name implies something untrue — accountList for a set, getUser that also creates one. Noise words add characters but no meaning — Data, Info, the, Object. One word per concept means picking a single verb per operation across the codebase (get, not get/fetch/retrieve).

open as a page

How should you manage the boundary between your code and a third-party library or external API, and what are "learning tests" in that context?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Do not let a third-party type spread through your codebase. Wrap it behind an interface you own, expressed in your domain's terms, and keep the vendor types inside that wrapper. Learning tests are small tests you write against the library to verify how it actually behaves - and they warn you when a version upgrade changes that behaviour.

open as a page

What does "self-documenting code" mean in practice, which specific refactorings let you replace a comment with code, and where does that technique stop working?

level: seniorimportance: must knowfreq 55%

basics

~20 s

It means naming and structuring code so it explains itself: rename vague identifiers, extract a commented block into a well-named function, and put a complex condition into a named variable. It fails for reasons outside the code — external constraints, hazards, and decisions.

open as a page

Why should code formatting be enforced by tooling rather than by code review, and how do you introduce an auto-formatter into a large existing codebase without destroying version-control blame history?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Formatting arguments waste review time and have no effect on behavior, so a formatter should decide and CI should check it. To adopt one on an old codebase, reformat everything in one dedicated commit and tell the blame tool to ignore that commit.

open as a page

Explain the "one level of abstraction per function" rule and the step-down (newspaper) ordering rule. How do you detect a mixed-abstraction function and fix it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Every statement in a function should sit at roughly the same conceptual distance from the domain — don't mix high-level policy with low-level string or byte fiddling. Step-down means each function is followed by the ones it calls, so the file reads top to bottom, general to detailed.

open as a page

What is a "hybrid" class in the objects-versus-data-structures sense, and how do Tell-Don't-Ask and feature envy relate to it?

level: middleimportance: should knowfreq 40%

basics

~20 s

A hybrid is a class that both exposes all its data through getters/setters and carries important business behaviour. It gets the downsides of both styles. Feature envy is code that uses another object's data more than its own; Tell-Don't-Ask fixes it by moving the behaviour to the data.

open as a page

Explain the scoring rules of SonarSource's Cognitive Complexity metric: what adds points, what adds nothing, and how the nesting penalty works.

level: middleimportance: should knowfreq 45%

basics

~20 s

Each structure that breaks straight-line reading — if, else, loops, catch, switch, jumps — adds one point, plus one more for each level of nesting it sits inside. Shorthand that condenses code (a whole switch, the method declaration itself) adds nothing extra.

open as a page

A teammate leaves a 40-line block of commented-out code in a pull request, arguing "we might need it back". What concrete harms does that cause, and is there any situation where keeping it is defensible?

level: middleimportance: should knowfreq 48%

basics

~20 s

It is dead text nothing checks: no compiler, no tests, no refactoring tool touches it, so it silently goes stale. Readers won't delete it because they assume it matters. Version control already stores it — delete it and recover from history if needed.

open as a page

Clean Code advises writing the try-catch-finally block first and treating error handling as "one thing". What does that mean in practice, and how does it relate to fail-fast and resource cleanup?

level: middleimportance: should knowfreq 52%

basics

~20 s

Write the failing case first — start with a test that expects the exception, then a try/catch/finally skeleton, then fill in the happy path. A function that handles errors should do only that: try is the first statement, and there's nothing after the catch/finally block.

open as a page

Why is column-aligning consecutive assignments or declarations ("horizontal alignment") generally discouraged, and what is the modern rationale behind line-length limits?

level: middleimportance: should knowfreq 45%

basics

~20 s

Aligning values into a neat column looks tidy but breaks the moment a name changes, producing large diffs and constant re-alignment. Line limits exist mainly so code fits side-by-side diffs and review windows without horizontal scrolling.

open as a page

What do "vertical openness" and "vertical density" mean in code formatting, and how should blank lines and vertical distance between related declarations be used?

level: middleimportance: should knowfreq 48%

basics

~10 s

Vertical openness means putting a blank line between separate thoughts so each stands out. Vertical density means keeping tightly-related lines packed together with no blank lines, so they read as one unit.

open as a page

What is Command-Query Separation (CQS) in function design, what concrete problems does it prevent, and where is it deliberately violated?

level: middleimportance: should knowfreq 58%

basics

~20 s

CQS says a function should either change state (a command, returning nothing) or return a value (a query, changing nothing) — never both. It keeps queries safe to call and read, so if (set(x)) style confusion disappears.

open as a page

What is Hungarian notation, and why do most modern style guides discourage encoding type, scope, or membership into identifier names?

level: middleimportance: should knowfreq 45%

basics

~20 s

Hungarian notation prefixes a name with its type — strName, iCount, lpszBuffer. Modern languages and IDEs already show types, so the prefix adds noise and becomes a lie the moment the type changes. Name the meaning instead.

open as a page

What is the data/object anti-symmetry, and how does it guide whether you should model something as polymorphic objects or as plain data plus procedures?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Objects make it easy to add new types (write one class; existing code is untouched) but hard to add new operations (every class must change). Plain data plus procedures is the exact opposite. So pick based on which change you expect more often.

open as a page

A teammate lowers a function's cognitive complexity from 24 to 8 by extracting six private helpers. How do you judge whether the code actually got easier to read?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Check whether each helper's name lets you skip its body. If you must open all six to understand the original function, the complexity just moved and reading now costs six jumps. Good extraction removes detail; bad extraction only relocates it.

open as a page

Beyond nesting, what specific code properties increase the reader's short-term memory load, and what techniques reduce it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Anything the reader must hold in their head while reading on: long-lived mutable variables, boolean flag parameters, unnamed intermediate conditions, hidden side effects, and jumping between distant files. Fix by naming intermediates, shrinking variable lifetimes, and making each named function a single chunk.

open as a page

showing 1–30 of 42