skip to content

Peer-to-Peer Style

Nodes that are all both client and server, organised by overlay networks, distributed hash tables such as Chord and Kademlia, and gossip. The interview value is in the discovery and trust problems that a central server would otherwise solve for you.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a peer-to-peer (P2P) network architecture, what does it mean for a node to act as both a client and a server at the same time?

level: juniorimportance: must knowfreq 80%

answer

  1. servent = server+client
  2. capacity scales with peers
  3. no single point of failure
  4. discovery is the hard part
  5. churn = constant membership change

basics

~20 s

Every computer in the network can both ask for things and give things to other computers -- there's no single 'boss' machine. Each one shares data directly with others, like friends swapping items instead of going through a store.

solid answer

~40 s

In P2P, every node (sometimes called a 'servent' -- server + client) exposes the same capabilities: it can originate requests for resources and it can service requests from other peers. There's no dedicated, privileged server that all traffic must pass through. This symmetry lets the network's capacity scale with its population -- each new peer brings its own upload bandwidth, storage, and CPU rather than adding pure load to a fixed set of servers. It also removes the single point of failure a central server represents: if one peer goes offline, others can still serve the same content (assuming replication) and requests get rerouted. The cost is added complexity: peers must discover each other, agree on protocols for locating resources, and handle the fact that any peer can vanish at any moment without warning.

go deeper

for a junior

Can state that peers both request and serve, and give a rough example like file sharing.

for a middle

Explains why this improves capacity/resilience versus a fixed server, and names discovery as a needed extra piece.

for a senior

Discusses churn, free-riding, and hybrid designs (trackers/supernodes) that combine coordination with symmetric data exchange.

for a principal

Weighs when the symmetric-peer model is the wrong choice -- e.g., when strong access control, audit trails, or guaranteed availability outweigh the capacity/resilience gains -- and articulates the operational cost of running software on machines you don't control.

## Both roles at once Peer-to-peer (P2P) architecture inverts the basic assumption behind most networked systems: instead of dividing participants into a small set of servers that provide a service and a large set of clients that consume it, every node in a P2P network plays both roles simultaneously. Practitioners sometimes coin the word `servent` (server + client) to name this dual capability. A peer can issue a request -- 'who has file X?', 'what's the value for key K?' -- and it can also receive and answer the same kind of request from any other peer. There is no node that is structurally privileged to always answer and never ask, and no node that is structurally forced to always ask and never answer. ## Capacity that grows with the crowd Why does this matter? The core motivation is **capacity that scales with population**. In a system where a fixed set of servers answers requests from an arbitrarily large set of clients, growth in the user base is pure load -- more clients means more demand against the same supply, and the operator must keep buying bigger or more servers. In a P2P system, each new participant brings its own resources: - **upload bandwidth**, disk space, CPU cycles; - in filesharing or content-distribution use cases, a fresh copy of whatever content it downloaded that it can now redistribute to others. Popular content becomes cheaper to serve precisely because more people have copies of it -- the opposite of the situation on a centralized server, where popularity causes overload. **BitTorrent** is the canonical illustration: a peer downloading a large file simultaneously uploads the pieces it already has to other downloaders, so a single 'seed' can bootstrap distribution to thousands of peers without the seed's own bandwidth being the bottleneck. ## Resilience The second big motivation is **resilience**. A centralized server is a single point of failure -- if it goes down, is partitioned by a network fault, or is legally forced offline, every client loses access simultaneously. In a P2P network, no single peer's departure has that effect, assuming the resource in question is available from more than one peer, which the application has to arrange rather than getting it for free from the architecture. **Skype's** original architecture (before it moved to centralized infrastructure) used ordinary user machines as 'supernodes' to route calls precisely to avoid depending on Skype-owned servers for the actual call traffic. ## Where the complexity lands The trade-offs are real and land on the complexity side of the ledger. 1. **First, discovery.** In a client-server system, a client trivially knows where to send its request -- a fixed, well-known address. In P2P, a peer usually doesn't know in advance which of potentially millions of other peers holds the resource it wants, so the network needs a discovery mechanism, whether that's a distributed hash table doing structured routing, a flooding/gossip-based unstructured search, or a hybrid that uses lightweight coordination nodes (like a BitTorrent tracker or a DHT bootstrap node) purely to help peers find each other before they start talking directly. 2. **Second, churn.** Peers are typically end-user machines -- laptops that get closed, phones that lose signal, home connections that drop -- and they join and leave far more often and far less predictably than a data-center server does. Any protocol built on the symmetric-peer assumption has to tolerate a routing table, a replica set, or a neighbor list changing constantly, with no central authority to notify everyone when a member disappears. 3. **Third, trust and heterogeneity.** A client-server deployment lets the operator control every server's software version, uptime, and behavior. A P2P network is made of independently owned, independently administered machines of wildly varying bandwidth, uptime, and honesty. This opens the door to failure modes that don't exist in a controlled server fleet -- free-riding peers that consume but never contribute (documented in early Gnutella measurements), and malicious peers that lie about what they have or corrupt data (mitigated with techniques like piece hashing in BitTorrent). ## Being precise about symmetry Finally, it's worth being precise about what 'symmetric' does and does not mean. Nodes have symmetric **capability** -- the software each peer runs can serve and can request. It does not mean every peer has identical resources or plays an identical actual role at every moment: some peers have more bandwidth and end up serving far more requests than they issue (BitTorrent seeds, or long-uptime DHT nodes that others route through more heavily), and hybrid designs deliberately elect certain peers to short-lived coordinating roles (Skype's old supernodes, BitTorrent trackers) without breaking the underlying peer symmetry of the protocol itself.

  • If every peer can be a server, why do many real P2P systems still use a small number of dedicated coordination nodes, like BitTorrent trackers?
    Pure discovery among symmetric peers with no fixed rendezvous point is hard to bootstrap -- a brand-new peer has no way to find even one other peer to start talking to. A tracker or bootstrap node solves the 'first contact' problem cheaply without becoming a bottleneck for the actual data transfer, which still happens directly peer-to-peer. This is a hybrid design, not a contradiction of the P2P model -- the data plane stays symmetric even though the discovery plane isn't.
  • How does free-riding undermine the capacity argument for P2P?
    The capacity-scales-with-population argument assumes peers contribute resources roughly in proportion to what they consume. Free-riders -- peers that download but never upload -- break that assumption, and if enough of the population free-rides, the network degenerates toward a small set of altruistic peers effectively acting as unpaid servers. BitTorrent responds to this with tit-for-tat choking, preferentially uploading to peers who upload back.

Like a potluck dinner versus a restaurant: at a restaurant (client-server) one kitchen serves everyone; at a potluck (P2P) every guest brings a dish and also eats from what others brought -- the more guests show up, the more food there is to go around.

saying these in an interview costs you the question

  • Describes P2P as 'no servers at all' rather than every node running both roles
  • Thinks discovery is solved automatically just because there's no central server
  • Assumes all peers are equally reliable/available
  • Can't name why churn matters for a P2P design
  • Conflates P2P with 'decentralized consensus' or replication guarantees

context

open as a page

In a peer-to-peer system, how does a gossip (epidemic) protocol propagate information like membership or liveness updates, and what does it trade off against a structured DHT lookup for that job?

level: middleimportance: must knowfreq 60%

basics

~20 s

Gossip is how peers spread news the way rumors spread in a crowd: each peer periodically tells a few random other peers what it knows, and they tell a few more, until eventually almost everyone has heard it -- no one broadcasts to everyone at once.

open as a page

What is an 'overlay network' in a peer-to-peer system, and what's the practical difference between a structured overlay and an unstructured overlay when a peer needs to find a resource?

level: middleimportance: must knowfreq 70%

basics

~20 s

An overlay network is a map of which peers talk to which, layered on top of the real internet. 'Structured' means peers are organized in a strict pattern so you can find things quickly and predictably; 'unstructured' means connections are more random, so finding things means asking around and hoping.

open as a page

In a distributed hash table (DHT) like Chord or Kademlia, how does a peer route a lookup for a given key to the peer responsible for it, and why does this typically take O(log n) hops for a network of n peers?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Every peer and every piece of data gets a number. Each peer only directly knows a few other peers, chosen so at each step you jump about halfway closer to the number you're looking for -- like narrowing down a phonebook by splitting it in half each time, so it only takes a handful of hops even in a huge network.

open as a page

Peer-to-peer networks experience high 'churn' -- peers constantly joining and leaving without warning. How does this affect a structured overlay's routing tables and stored data's availability, and what mechanisms mitigate it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Peers come and go all the time in a P2P network, like people wandering in and out of a crowd. If the network doesn't keep updating its 'who's near whom' info and doesn't keep spare copies of data on multiple peers, searches start failing and data can vanish when the one peer holding it leaves.

open as a page

What is a Sybil attack against a peer-to-peer network's discovery/routing layer, and why is it particularly damaging to a structured overlay like Kademlia compared to simply degrading service?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A Sybil attack is when one person pretends to be hundreds or thousands of different peers using fake identities. If enough fake peers surround the part of the network responsible for a specific piece of data, the attacker can quietly control access to that data or hide it, since new peers only trust who the network tells them to trust.

open as a page