nginx handles thousands of concurrent connections with a handful of worker processes — typically one per CPU core. Mechanically, how does one worker do that without one OS thread per connection, and why does a thread-per-connection server run out of resources before nginx does at the same connection count?
A thread-per-connection server blocks one OS thread on each connection's socket, whether or not that connection currently has anything to do — an idle keep-alive connection still owns a thread's stack (hundreds of kilobytes to a few megabytes) and the kernel scheduler has to consider that thread every time it decides what to run, even while it's just waiting. That cost scales with the number of open connections, active or not, which is why such a server runs out of threads, then memory, long before it runs out of CPU. Each nginx worker is instead a single-threaded event loop built on the kernel's epoll facility: it registers every socket it's watching and makes one call that returns only the sockets that actually have something ready right now — a request arrived, a write buffer drained, a connection closed — and services exactly those in a tight loop. An idle connection between requests costs the worker almost nothing: a file descriptor and a small amount of state, not a blocked thread, so the cost scales with the number of events, not the number of open connections — which is why doubling the idle-but-open connection count barely moves nginx's CPU usage while it roughly doubles the threads a thread-per-connection server is juggling.