What is asyncio's transport and protocol layer, and when would you use it instead of streams?
answer
- The high-level API sits on this one
- One owns the socket, one owns logic
- The loop calls you, not the reverse
- Nothing in those calls can suspend
- Pays off per connection, at scale
basics
~20 sTransports own the socket and its write buffer; protocols are your callback object, receiving connection_made, data_received, eof_received and connection_lost. Streams are a coroutine-friendly façade built on that pair. Drop to protocols for lower per-connection overhead and zero-copy reads.
solid answer
~40 sThe transport/protocol pair is asyncio's low-level networking layer. A **transport** wraps the socket and exposes `write`, `close`, `abort`, `pause_reading`/`resume_reading` and `set_write_buffer_limits`. A **protocol** is an object you write, and the loop calls back into it: `connection_made`, `data_received`, `eof_received`, `connection_lost`, plus `pause_writing`/`resume_writing` for flow control. `asyncio.open_connection` and `asyncio.start_server` build exactly this pair with a stock protocol that feeds a `StreamReader`, which is why streams and protocols are not competing APIs but two floors of one building. You go down a floor when per-connection cost matters — no task and no reader buffer per connection — or when you want `asyncio.BufferedProtocol`, which hands the transport your own buffer and avoids a copy per read. The cost is that callbacks are synchronous and cannot await, so sequential logic becomes a state machine.
code
python · 25 linesimport asyncio
class EchoProtocol(asyncio.Protocol):
def connection_made(self, transport):
self.transport = transport
def data_received(self, data):
self.transport.write(data.upper())
self.transport.close()
def connection_lost(self, exc):
print("connection_lost:", exc)
async def main():
loop = asyncio.get_running_loop()
server = await loop.create_server(EchoProtocol, "127.0.0.1", 0)
port = server.sockets[0].getsockname()[1]
async with server:
reader, writer = await asyncio.open_connection("127.0.0.1", port)
writer.write(b"scrape\n")
print(await reader.read())
writer.close()
await writer.wait_closed()
asyncio.run(main())go deeper
Know that the stream API you use is built on a lower layer of transports and callback protocols, and that you are not expected to write one for ordinary work.
Name the callbacks — connection_made, data_received, eof_received, connection_lost — and explain that they are synchronous, so anything needing to wait must be handed to a task.
Justify dropping down with a measurement: per-connection memory and scheduling cost at high connection counts, or an allocation per read that BufferedProtocol removes, and state what you give up in readability and automatic backpressure.
Decide whether owning a protocol implementation is worth the maintenance and on-call cost against adopting an existing framed protocol, and set the bar of evidence a team must clear before writing one.
## Two floors of one building asyncio's networking has a low level and a high level, and the high level is implemented in terms of the low one. * **Transport** — owns the socket, the write buffer and the connection lifetime. Its interface is synchronous: `write(data)`, `writelines`, `close()`, `abort()`, `pause_reading()`, `resume_reading()`, `set_write_buffer_limits(high, low)`, `get_write_buffer_size()`, `get_extra_info(name)`. You never construct one; the loop hands you one. * **Protocol** — the object *you* write, subclassing `asyncio.Protocol`. The loop calls: `connection_made(transport)` once at the start, `data_received(data)` for every chunk that arrives, `eof_received()` when the peer half-closes, `connection_lost(exc)` once at the end, and `pause_writing()`/`resume_writing()` when the write buffer crosses its water marks. Streams sit on top: `asyncio.open_connection` creates a transport plus a stock protocol whose `data_received` pushes bytes into a `StreamReader` and wakes whichever coroutine is awaiting a read, and whose `pause_writing`/`resume_writing` are what make `await writer.drain()` suspend and resume. Knowing that is what makes `drain()` stop being magic. ## The essential difference: callbacks cannot await `data_received` is an ordinary synchronous method. It cannot `await` anything, so any logic that needs to wait — query a store, call a downstream, apply a timeout — cannot live inline. You either keep an explicit parsing state machine in the protocol instance and hand completed messages to a task, or you push them into an `asyncio.Queue` a worker task drains. That is real complexity, and it is why the standard advice is: use streams unless you have a reason not to. The mirror is also true and is the reason `data_received` is fast: nothing suspends inside it, so there is no task creation, no scheduling round trip, and no per-connection coroutine frame. ## When going down a floor pays **Per-connection overhead at high connection counts.** A stream-based server creates a task per connection, a `StreamReader` with its own buffer and a `StreamWriter`. Across tens of thousands of mostly-idle connections — a metrics scraper holding one connection per agent in a large fleet — that fixed cost is measurable in both memory and the loop's scheduling work. A protocol keeps one small object per connection and no task at all. **Zero-copy reads.** `asyncio.BufferedProtocol` inverts the read path: instead of the transport allocating a `bytes` object per chunk and calling `data_received`, it asks your protocol for a writable buffer via `get_buffer(sizehint)`, reads straight into it, and calls `buffer_updated(nbytes)`. For a protocol that parses in place — a fixed-size header followed by a body — that removes one allocation and one copy per read. It is the reason several high- throughput protocol implementations are written as protocols rather than over streams. **Precise lifecycle control.** `connection_lost(exc)` fires exactly once with the reason, which makes cleanup accounting exact. `transport.abort()` drops the connection immediately without flushing, which is the right hammer for a misbehaving peer. And read-side backpressure is yours explicitly: `pause_reading()` stops the loop from pulling from the socket, so the kernel receive buffer fills and TCP's window closes — the same effect that a slow `StreamReader` consumer gets automatically. ## Flow control, both directions Write side: the transport calls `pause_writing()` when its buffer passes the high-water mark and `resume_writing()` when it drops below the low mark. Your protocol is responsible for actually stopping — nothing prevents you from ignoring both and buffering until you die. The streams layer is exactly the code that turns those callbacks into a future so a coroutine can await it. Read side: `pause_reading()` and `resume_reading()` on the transport. ## Choosing, in practice Start with streams. Everything reads sequentially, timeouts compose with `asyncio.timeout`, errors propagate as exceptions where they happened. Move to a protocol when you have measured a cost you cannot pay: per-connection memory at a connection count in the tens of thousands, or an allocation and copy per read in a hot parsing path. Treat it as an optimisation with a real readability price, not as the "proper" way — and note that when you switch, backpressure stops being automatic and becomes something you must implement in both directions. ## What interviewers listen for That streams are built on transports and protocols rather than parallel to them; that the protocol callbacks are synchronous and why that constrains the design; that flow control exists at both layers with `pause_writing`/`resume_writing` underneath `drain()`; and that the honest default is streams.
- How does backpressure work when you write a protocol directly instead of using streams?The transport calls your protocol's `pause_writing()` when its write buffer crosses the high-water mark and `resume_writing()` when it falls under the low mark, and you must actually stop producing in between — nothing enforces it. On the read side you call `pause_reading()` and `resume_reading()` on the transport yourself. Streams exist precisely to turn that callback pair into an awaitable `drain()`.
- What does asyncio.BufferedProtocol change about the read path?Instead of the transport allocating a bytes object per chunk and passing it to `data_received`, it asks your protocol for a writable buffer with `get_buffer(sizehint)`, reads into it directly, and then calls `buffer_updated(nbytes)`. That removes an allocation and a copy per read, which matters in a hot parsing loop; the price is that you own buffer sizing and the parse position.
- Why can't a protocol's data_received call an async function directly?It is a plain synchronous method invoked by the event loop while it is dispatching I/O readiness; suspending there would mean suspending the loop. You keep a parsing state machine on the protocol instance and hand completed messages to a task or an `asyncio.Queue`, which is exactly the complexity that the streams layer hides for you.
saying these in an interview costs you the question
- Thinks streams and protocols are unrelated alternatives
- Tries to await inside data_received
- Says protocols are always the faster right choice
- Cannot name a single protocol callback
- Assumes backpressure is automatic when writing a protocol