Ch. 21 · Operating Systems

Operating Systems: I/O Models

Compare blocking, non-blocking, multiplexed and asynchronous I/O, and why the choice drives server scalability.

~2 min readadvancedupdated Oct 5, 2026

How a program waits for I/O determines how many connections it can serve. Blocking I/O ties up a thread per operation, while non-blocking and asynchronous models let one thread manage many connections. The models differ in who waits and when.

Before you start

You should understand file descriptors and threads. This article compares the models conceptually.

Step-by-step walkthrough

Step 1: Blocking I/O is simple but costly per connection

A blocking read returns only when data is available, so the calling thread sleeps. This is easy to reason about but needs one thread per in-flight operation, which does not scale to tens of thousands of connections.

Step 2: Non-blocking plus readiness avoids the wait

A non-blocking read returns immediately, either with data or a “would block” signal. The program then asks the OS which descriptors are ready and handles only those. This is I/O multiplexing, and one thread can manage many connections. The cost is more complex control flow.

Step 3: Asynchronous I/O removes the readiness step

In true async I/O the program submits an operation and is notified on completion, so it never polls readiness. This can be more efficient but depends on OS support, which varies. Many runtimes emulate async with a readiness loop, which is why the distinction matters.

Worked scenario

The readiness loop handles only sockets that are ready.

loop:
  ready = poll(fds)            # block until at least one is ready
  for fd in ready:
    data = read(fd)            # will not block; data is available
    handle(fd, data)
Text

Walk through the example

poll blocks once, waking when at least one descriptor is ready, and the program reads only those. A single thread handles many connections because it never blocks on any individual read. This is the structure behind event-loop servers.

Common mistake

Busy-waiting by looping on a non-blocking read without a readiness call, which burns CPU. Another is spawning a thread per connection for blocking I/O and hitting a thread and memory ceiling under load.

Verify the behavior

Compare memory and thread counts for a thread-per-connection server and an event-loop server at the same connection count. Confirm the event-loop server stays responsive with few threads. Observe the readiness call returning only ready descriptors.

Interview exercise

Why does a thread-per-connection model fail at high concurrency?

Answer and reasoning

Each thread needs a stack and kernel bookkeeping, so memory grows roughly linearly with connections, and context switching between many runnable threads adds overhead. At tens of thousands of idle connections, most threads are blocked on I/O yet still consume resources. An event-loop or async model handles the same connections with far fewer threads by waiting once for readiness.

Continue learning

Compare the mechanism in epoll and select and interrupts in Interrupts and I/O. Read the Linux epoll documentation and try the Operating systems interview questions.

More in Operating Systems

read ✓Operating Systems · hard

Operating Systems: CPU Cache Locality

Improve performance with spatial and temporal locality, avoid pointer chasing, and understand false sharing between threads.

~2 min readread →
read ✓Operating Systems · hard

Copy-on-Write Memory and fork

How copy-on-write lets fork share pages until a write, why RSS can be misleading, and the implications for memory and latency.

~2 min readread →
read ✓Operating Systems · hard

Operating Systems: epoll and select

How select, poll and epoll report I/O readiness, and the difference between level- and edge-triggered notifications.

~2 min readread →
esc