I/O multiplexing needs a way to ask the OS which descriptors are ready. select and poll scan a set of descriptors each call, while epoll keeps a kernel list and returns only the ready ones, which scales better. The trigger mode then decides what “ready” means.
Before you start
You should understand file descriptors and non-blocking I/O. This article covers the readiness interfaces and trigger modes.
Step-by-step walkthrough
Step 1: select and poll scan on every call
select and poll take a set of descriptors and return which ones are ready, but they examine the whole set each time, so cost grows with the number of descriptors. They are portable and fine for small sets, less so for thousands of connections.
Step 2: epoll keeps a ready list
epoll registers descriptors once and maintains a kernel-side list of ready events, so each call returns only those. The per-call cost is proportional to the number of ready descriptors, not the total, which is why it scales to large connection counts.
Step 3: Choose level- or edge-triggered
Level-triggered reports a descriptor as ready as long as data remains, so you may read it in several calls. Edge-triggered reports only on a change, so you must drain the descriptor until it would block; missing that leaves data unread and the event never re-fires.
Worked scenario
epoll reports only the descriptors that are ready.
epfd = epoll_create1(0)
epoll_ctl(epfd, ADD, fd)
loop:
n = epoll_wait(epfd, events, timeout)
for i in 0..n: handle(events[i]) # only ready descriptorsWalk through the example
After registration, epoll_wait blocks until at least one descriptor is ready and returns just that subset, so the loop never scans every descriptor. With level-triggered mode, a descriptor with remaining data is reported again next call; with edge-triggered, the handler must read until the socket would block.
Common mistake
Using edge-triggered mode without draining the descriptor, so the remaining data is never processed. Another is using select with a large descriptor set and paying the scan cost on every call.
Verify the behavior
Register many descriptors and compare per-call overhead between poll and epoll. In edge-triggered mode, read partially and confirm the event does not re-fire, then drain fully and confirm it does. Confirm level-triggered mode re-reports lingering data.
Interview exercise
Why is edge-triggered epoll easy to get wrong?
Answer and reasoning
It notifies only when the state changes, so if the handler reads part of the available data and stops, no new edge occurs and the rest is never processed. The handler must loop until a read would block, at which point the descriptor is drained. Level-triggered mode avoids this by re-reporting, at the cost of redundant wakeups.
Continue learning
Compare the models in I/O models and interrupts in Interrupts and I/O. Read the Linux epoll documentation and try the Operating systems interview questions.