A Java service has been running fine for weeks and suddenly starts throwing `java.io.IOException: Too many open files` and refuses new connections, with no code deploy that day. What's actually happening, and where do you look first?
Every open file, socket, pipe and epoll instance counts against a per-process limit on file descriptors (ulimit -n, or LimitNOFILE= in a systemd unit), and the default on some systems (1024) is a number a busy Java service reaches within minutes, not weeks — every accepted connection is a descriptor, every outbound HTTP call, every open log file. Once the table fills, the process can't open anything new, and because accepting a connection also needs a descriptor, the service stops accepting new connections entirely, not just failing the operation that happened to hit the limit. "No deploy that day" doesn't mean nothing changed — a slow leak (a socket or file handle that's opened and never closed) accumulates gradually and only crosses the limit after enough time or traffic has passed, which is exactly why it can look sudden. The first places to look: ls /proc/<pid>/fd | wc -l for the count, and lsof -p <pid> for what they actually are — a leak shows up as the same file or the same remote address repeated hundreds of times.