← back to section

A service restarts without a single log line, Kubernetes shows "exit code 137", under load you get "Too many open files", and the disk is full even though the logs were just deleted. None of these stories is about code; they are about how Linux manages a process. The network side (ports, connections, DNS) is covered in network diagnostics; this article covers the rest.

SIGTERM"please shut down" wait 10–30 sprocess finishes work SIGKILLkilled, code 137 made it: code 0 docker stop: 10 s, Kubernetes: 30 s cannot be caught

Stopping always takes two steps: a polite request, then a forced kill if the process does not make it in time.

Signals: how a process is stopped

A process cannot be "switched off" from outside directly; it is sent a signal, a short notification from the kernel. Two signals stop it. SIGTERM is a request to exit: the process can catch it, finish responses, close database connections and exit on its own. SIGKILL cannot be caught: the kernel kills the process instantly and anything unwritten is lost.

docker stop and Kubernetes behave the same way: first SIGTERM, then they wait (Docker 10 seconds, Kubernetes 30 seconds by default), and if the process is still alive they send SIGKILL. The exit code tells you what happened: a process killed by a signal exits with 128 plus the signal number. SIGTERM is 15, hence 143; SIGKILL is 9, hence 137.

The example is in Python because it is the easiest to run; Java, Go and Node work the same way, and writing a handler in your language is covered in the section on graceful shutdown:

live example

import signal
import subprocess
import sys

graceful = """
import signal, sys, time
def stop(signum, frame):
    print("  got SIGTERM, closing connections", flush=True)
    sys.exit(0)
signal.signal(signal.SIGTERM, stop)
print("  working", flush=True)
time.sleep(30)
"""
careless = """
import time
print("  working", flush=True)
time.sleep(30)
"""

for name, code, sig in [
    ("with handler, SIGTERM", graceful, signal.SIGTERM),
    ("without handler, SIGTERM", careless, signal.SIGTERM),
    ("with handler, SIGKILL", graceful, signal.SIGKILL),
]:
    print(name)
    proc = subprocess.Popen([sys.executable, "-c", code], stdout=subprocess.PIPE, text=True)
    print(proc.stdout.readline().rstrip())
    proc.send_signal(sig)
    out, _ = proc.communicate()
    if out:
        print(out.rstrip())
    shell_code = 128 - proc.returncode if proc.returncode < 0 else proc.returncode
    print(f"  shell exit code: {shell_code}")
Run

Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →

A container pitfall: the kernel protects the process with PID 1, and a signal without a handler does not affect it. An application without its own handler, started as the first process, ignores SIGTERM and still gets SIGKILL with code 137 after the timeout. Why this happens and how to fix it is covered in the article on running containers.

File descriptors and Too many open files

To the kernel, an open file, a network connection and a pipe between processes are the same thing: a file descriptor, a number in the process table. The number of descriptors is limited; the soft limit is often 1024. Every database connection, every incoming request, every open log uses one number. When they run out, any attempt to open a file or accept a connection fails with EMFILE, and the logs say "Too many open files".

There are two causes. Either the limit is honestly too low for the load, and it is raised with ulimit -n or the service settings. Or descriptors leak: a file or connection was opened and never closed, and their count grows until it hits the limit. A leak shows up as ls /proc/<pid>/fd | wc -l that only grows, and lsof -p <pid> shows what exactly is open. This is how it looks from inside, with a limit of 64:

live example

import errno
import resource
import tempfile

soft, hard = resource.getrlimit(resource.RLIMIT_NOFILE)
resource.setrlimit(resource.RLIMIT_NOFILE, (64, hard))

leaked = []
try:
    while True:
        leaked.append(tempfile.TemporaryFile())
except OSError as error:
    print(f"files opened: {len(leaked)}, then error {errno.errorcode[error.errno]}")

for f in leaked:
    f.close()

with tempfile.TemporaryFile() as f:
    f.write("with closes the file itself".encode())
print("after closing, files open again")
Run

Running examples is part of paid access. There the same code runs inside the article: editor, run and check next to the paragraph. Three free days →

Fewer than 64 get opened: the first numbers are already taken by standard input, output and error. A leak is fixed not by raising the limit but by closing the resource where it was opened: with in Python, try-with-resources in Java, defer in Go, finally in Node.

Disk: space, inodes and deleted files

"No space left on device" can happen with gigabytes free. The file system keeps each file's metadata in a separate record, an inode, and their number is fixed when the disk is created. Millions of small files, such as a cache or sessions, exhaust inodes before space. So check both commands: df -h shows space, df -i shows inodes.

The second common story: a log was deleted but the space was not freed. While a process keeps the file open, the kernel does not free the data; only the name is gone. lsof +L1 shows such files marked (deleted). The space comes back when the process closes the file or restarts. That is why large logs are not deleted but truncated with truncate -s 0, or rotated with a signal telling the process to reopen the file.

Memory and the OOM killer

When the whole system, or a container with a limit, runs out of memory, the kernel does not wait: the OOM killer picks a process, usually the largest, and kills it with SIGKILL. The application logs will show nothing; the process had no chance to write anything. The signs: code 137, the reason OOMKilled on a Kubernetes pod, and lines like Out of memory: Killed process in dmesg.

How much a process really uses is shown by RSS, the resident memory in ps and top. A common cause of OOM in containers is an application sizing its memory by the whole machine rather than the container limit. The JVM since version 10 bases its heap on the container limit. Go does not: the garbage collector is told the limit via GOMEMLIMIT. Node sets its heap limit with --max-old-space-size. Everywhere, thread stacks and buffers come on top of the heap, so the limit is set with headroom below the pod limit.

Summary

  • SIGTERM is a request to exit and can be caught; SIGKILL kills immediately with no handler.
  • Exit code is 128 + signal number: 143 is SIGTERM, 137 is SIGKILL (often the OOM killer).
  • docker stop waits 10 s, Kubernetes 30 s, then SIGKILL.
  • Files and sockets are file descriptors; the limit is ulimit -n, a leak shows in ls /proc/<pid>/fd and lsof -p.
  • df -h shows space, df -i shows inodes; a deleted but open file keeps its space, lsof +L1 shows it.
  • The OOM killer kills silently: look for OOMKilled, dmesg, and compare RSS with the limit.

Further reading