Skip to content
Riadh Mnasri
← Back to blog
7 min read

OOMKilled or OutOfMemoryError: where a JVM's memory goes in a pod

OOMKilled or OutOfMemoryError: where a JVM's memory goes in a pod

The scenario is a classic. A Spring Boot service runs in a pod limited to 1 Gi, with -XX:MaxRAMPercentage=75, as recommended in the article on day-to-day Kubernetes. Metrics show a heap that tops out at 500 MB. And yet, twice a day, the pod restarts. kubectl describe pod shows OOMKilled, exit code 137. No OutOfMemoryError in the logs, no heap dump.

It isn't a contradiction: the JVM and Kubernetes don't measure the same memory. This article explains the difference between the two errors, where the memory that isn't in the heap goes, how to measure it, and how to turn that into consistent requests and limits.

Two errors, two mechanisms#

OutOfMemoryErrorOOMKilled
Who triggers itThe JVMThe Linux kernel (the container's cgroup)
What is fullThe heap (or metaspace, or direct memory)The process's total memory exceeded the container limit
What you getAn exception, a stack trace, a heap dump if configuredA process killed outright, exit code 137 (128 + SIGKILL)
Where to lookApplication logskubectl describe pod, Last State field

The key difference is in the second row. -Xmx or MaxRAMPercentage cap the heap. The Kubernetes limit applies to the whole process. Everything the JVM uses on top of the heap counts for Kubernetes and not for the JVM.

Practical consequence: -XX:+HeapDumpOnOutOfMemoryError is useless against an OOMKilled. The kernel kills the process without warning, it has no time to write anything.

What the JVM uses outside the heap#

┌─────────────────── container limit: 1 Gi ────────────────────┐
│ Heap                           │ Off-heap                     │
│ MaxRAMPercentage=75 → 768 Mi   │ metaspace, code cache,       │
│                                │ thread stacks, GC,           │
│                                │ direct buffers, native libs  │
└────────────────────────────────┴──────────────────────────────┘
                                   ▲
                    256 Mi left for all of that

The main items, with orders of magnitude for a typical Spring Boot service:

  • Metaspace: metadata of loaded classes. A Spring Boot service with Hibernate and a few starters easily loads 20,000 classes, or 100 to 200 MB. Uncapped by default.
  • Code cache: code compiled by the JIT. Up to 240 MB reserved, often 50 to 100 MB actually used.
  • Thread stacks: about 1 MB reserved per thread by default (-Xss). 200 Tomcat threads plus connection pools, Kafka threads and GC threads, and you quickly go past 200 MB reserved, even if only part of it is actually used.
  • GC structures: G1 uses memory for its own bookkeeping, on the order of a few percent of the heap.
  • Direct buffers: Netty, reactive HTTP clients and some drivers allocate off-heap (ByteBuffer.allocateDirect). By default, the limit equals the maximum heap size.
  • Native libraries and the allocator: compression, TLS, and fragmentation in the libc memory allocator, which alone can account for tens of megabytes in a heavily multithreaded process.

Added up, the off-heap footprint of an ordinary Spring Boot service is often between 250 and 400 MB. With a 1 Gi limit and 75% for the heap, 256 MB remain: exactly the zone where the pod survives most of the time, then dies when metaspace, threads and buffers grow at the same time under load.

Measure instead of guessing#

The JVM can break down its own usage, as long as you ask at startup:

-XX:NativeMemoryTracking=summary

Then, on the running pod:

bash
kubectl exec -it api-7d9f-xk2p -- jcmd 1 VM.native_memory summary
Total: reserved=2143MB, committed=931MB
-                 Java Heap (reserved=768MB, committed=512MB)
-                     Class (reserved=1105MB, committed=182MB)
-                    Thread (reserved=231MB, committed=231MB)
-                      Code (reserved=247MB, committed=74MB)
-                        GC (reserved=96MB, committed=71MB)
-                  Internal (reserved=5MB, committed=5MB)
-                     Other (reserved=34MB, committed=34MB)

The column that matters is committed: the memory actually requested from the system. Here, 931 MB for a container limited to 1 Gi, with a heap at only 512 MB. The problem jumps out: 230 MB of threads and 180 MB of classes.

Two caveats. NMT doesn't see everything: allocations made by native libraries outside the JVM and allocator fragmentation don't show up. The Kubernetes metric container_memory_working_set_bytes, the one the kernel compares to the limit, will therefore be a bit higher. And NMT has a small cost: I turn it on to diagnose, not permanently.

Fixing it: shrink off-heap or leave more room#

Once you've measured, the levers are well known.

Leave more room for off-heap. 75% suits containers of 2 Gi and above. For 512 Mi or 1 Gi, 60 to 65% is often more realistic, or a fixed heap size computed from the NMT measurement.

Cap what can be capped, so growth produces an explicit JVM error rather than a silent OOMKilled:

-XX:MaxRAMPercentage=65
-XX:MaxMetaspaceSize=256m
-XX:ReservedCodeCacheSize=128m
-XX:MaxDirectMemorySize=128m
-Xss512k

Lowering -Xss isn't risk-free: deeply recursive code can trigger StackOverflowError. 512 KB is enough for most web services, but check it with load tests.

Reduce the thread count. A 200-thread Tomcat pool in a pod limited to one CPU makes no sense. Bringing it down to 50, or switching to virtual threads (Java 21, spring.threads.virtual.enabled=true), whose stacks live in the heap and grow on demand, directly shrinks the "Thread" line.

Requests and limits: what each value changes#

Kubernetes uses two values per resource, and they don't play the same role.

  • The request is for placement: the scheduler only puts the pod on a node that has that amount available. It limits nothing at runtime.
  • The limit is a ceiling. For memory, exceeding it means OOMKilled. For CPU, exceeding it means being slowed down (throttling).

For a JVM's memory, I set request = limit. A JVM almost never gives heap memory back to the system once it has taken it. If the request is lower than the limit, the pod can land on a node that can't really absorb its actual usage, and it will be among the first killed when the node runs short of memory.

For CPU, the question is more debated. Two effects are worth knowing:

  • the JVM sizes GC, JIT and common ForkJoinPool threads from the number of CPUs it sees, derived from the CPU limit (and not the request, since Java 19);
  • with fewer than 2 visible CPUs, or less than 1,792 MB of memory, the JVM defaults to the Serial GC, poorly suited to a web service. A 1 Gi pod is therefore affected, whatever its CPU count.

A container with a 500m CPU limit therefore sees a single processor, runs the Serial GC, and gets throttled on the slightest spike, for instance during Spring startup, which compiles a lot of code. Two reasonable options: no CPU limit, relying on the request (the pod uses the node's spare CPU), or a limit of at least 2 CPUs. Either way, you can force the value the JVM sees with -XX:ActiveProcessorCount=2 and the GC with -XX:+UseG1GC.

yaml
resources:
  requests:
    memory: "1Gi"
    cpu: "500m"
  limits:
    memory: "1Gi"
    # no CPU limit: the pod can exceed its request when the node has spare CPU
env:
  - name: JAVA_TOOL_OPTIONS
    value: >-
      -XX:MaxRAMPercentage=65
      -XX:MaxMetaspaceSize=256m
      -XX:ActiveProcessorCount=2
      -XX:+UseG1GC
      -XX:+ExitOnOutOfMemoryError

-XX:+ExitOnOutOfMemoryError completes the setup: a JVM that hit an OutOfMemoryError is often in an inconsistent state. Letting it exit cleanly lets Kubernetes restart it, instead of leaving it serving errors for hours.

The approach in short#

  1. Identify the error: OOMKilled (code 137 in kubectl describe pod) or OutOfMemoryError (in the logs). They don't have the same causes.
  2. For an OutOfMemoryError: heap dump and heap analysis, as described in the heap dump article.
  3. For an OOMKilled: enable NMT, measure off-heap under real load, compare with the limit.
  4. Adjust: a lower heap percentage, caps on metaspace, code cache and direct memory, fewer threads.
  5. Set request = limit for memory, and handle CPU with what the JVM derives from it in mind.

The "let's raise the limit to 2 Gi" reflex often fixes the symptom, but without measuring, you don't know whether you gave room to the heap, which had plenty, or to off-heap, which was short. And you pay for twice the memory per replica to fix a 200 MB problem.