Free sample · Kubernetes Postmortem

Pod Lifecycle & OOMKilled

Part I · Chapter 1

Foundation · Week 1 · Depends on: —

Understanding how pods live, die, and exhaust memory — and why your first instinct is usually wrong.

A pod is restarting every few hours. The status flickers to OOMKilled, exit code 137, then back to Running as if nothing happened. You check the deployment. The memory limit is already generous — in fact, someone raised it last week. You check the metrics. The pod is nowhere near its limit. Everything you can see says this pod should be fine, and yet roughly every five hours, the kernel reaches in and kills it.

This is a true story, and over the course of this chapter we'll solve it. But here's the part worth bracing for now: you are going to be wrong about it three times before you're right. The first fix will be obvious and useless. The second fix will appear to work and won't. The third realisation is that you were solving the wrong problem from the very first instinct — and that the actual fix involves changing no memory setting at all.

That three-layer mystery is the spine of this chapter. We'll build up the concepts you need to solve it, and then we'll come back and solve it properly in the first postmortem. Along the way, this single failure — a container exceeding a memory ceiling — will turn out to be the doorway to nearly everything: cgroups, the QoS hierarchy, admission control, the difference between a pod being alive and a workload being healthy, and the brutal honesty of running your own hardware.

1. Definition & Core Concept

Before we can talk about how a pod dies, we have to be precise about how it lives.

The lifecycle, mechanically

A pod moves through a small set of phases — Pending while it waits to be scheduled and have its images pulled, Running once at least one container has started, and finally Succeeded or Failed. But the phase is the coarse view. The interesting state lives one level down, in the container statuses, where each container is Waiting, Running, or Terminated, and carries a history of how it got there.

On every node, the kubelet is the local supervisor. It watches the containers it's responsible for, and when one dies, it consults the pod's restartPolicy (which defaults to Always) and decides whether to restart it. This is the loop at the heart of this chapter: a container dies, the kubelet restarts it, and if it keeps dying quickly, that loop becomes the CrashLoopBackOff we'll meet at the end of this chapter and dissect properly in the next.

The crucial mental shift for a newcomer is this: the kubelet does not run your container. The kernel does. The kubelet asks the container runtime to start a process inside a set of Linux primitives — namespaces for isolation, and cgroups for resource accounting and limits. When we talk about a memory limit, we are talking about a number the kubelet wrote into a cgroup, which the kernel then enforces with no knowledge of, or sympathy for, Kubernetes.

What a memory limit actually is

When you write this in a manifest:

resources:
requests:
memory: "2Gi"
limits:
memory: "6Gi"

the limit becomes a hard ceiling on a cgroup. On a cgroup v2 system (which is what you'll find on any modern node) that ceiling is written to a file called memory.max. The kernel tracks the cgroup's current usage in memory.current, and when usage tries to exceed memory.max and cannot be reclaimed, the kernel's cgroup-level OOM killer fires. It picks a process inside that cgroup — almost always your main process, PID 1 in the container — and sends it SIGKILL.

SIGKILL is signal 9. A process killed by signal N exits with code 128 + N. So:

128 + 9 = 137

That is the entire meaning of exit code 137. It is not a Kubernetes error code. It is the Unix convention for "killed by signal 9," and when you see it next to Reason: OOMKilled, the kernel is telling you it ran a process out of its memory cgroup and terminated it without ceremony.

The triad you must keep separate

Here is the single most important distinction in this chapter, and one that newcomers conflate constantly. There are three different ways memory pressure kills your pods, they have different causes, and treating one as another will send you debugging in exactly the wrong direction.

  1. cgroup OOM kill. The container exceeded its own limit. The kernel kills a process inside that one cgroup. Every other container on the node is unaffected. This is the classic OOMKilled, exit 137, single-pod event. The cause is inside that workload.
  2. Host (global) OOM kill. The node itself ran out of memory. The kernel's global OOM killer fires and chooses a victim across the entire machine, ranked by a score. It can — and does — kill pods that are comfortably under their own limits, because the problem was never that pod; the problem was the node. This is the co-located-overcommit scenario, and on self-hosted infrastructure it is a live risk, not a theoretical one. We'll meet it in Postmortem 2.
  3. Kubelet node-pressure eviction. Before the kernel's global killer fires, the kubelet is also watching node memory. When available memory drops below a configured threshold, the kubelet proactively evicts pods to reclaim memory — these show up with status Evicted, not OOMKilled. This is the comparatively graceful path, driven by Kubernetes rather than the raw kernel, and it has its own victim-selection rules based on QoS class and how far a pod has exceeded its requests.

Same symptom family — a pod stops running because something ran out of memory — three entirely different mechanisms. The first question in any memory investigation is which of these am I looking at, because they have nothing in common except the word "memory."

Why the metrics can say "fine" while the kernel says "dead"

One more concept, because it explains the opening mystery's most maddening feature: the pod that OOMs while the dashboard shows it using less than its limit.

What counts toward the cgroup limit is, roughly, the working set — anonymous memory (heap, stacks) plus page cache that cannot be reclaimed under pressure. Reclaimable page cache doesn't ultimately keep you from staying under the ceiling, because the kernel can drop it. But two things routinely fool the naive observer. First, a metrics scrape every 30 or 60 seconds will completely miss a fast allocation spike that crosses the limit and triggers the kill between samples — the average looks calm because the kill happened in the gap. Second, and more insidiously, memory the kernel counts but your runtime's garbage collector does not — native allocations, off-heap buffers, C-extension memory — is fully charged to the cgroup while being entirely invisible to a heap dashboard. The GC graph looks healthy. The RSS climbs anyway. The kernel goes by RSS.

Hold onto that last point. It is the mechanism behind two of this chapter's examples, and it is the single most common reason engineers stare at a "healthy" memory graph next to a pod that keeps dying.

2. Why It Happens

If you remember one sentence from this chapter, make it this one:

The limit is not the leak.

OOMKilled is a symptom, and the number on the limit is very nearly the least informative thing about which underlying mechanism produced it. Raising the limit is the universal first instinct precisely because it's the only lever that's obviously connected to the word "memory" — and it's wrong far more often than it's right. Here are the mechanisms that actually drive the symptom, grouped so you can recognise them.

Genuinely under-provisioned for real concurrency. Sometimes the boring answer is correct: the limit really is too low for the workload's honest peak. A service sized against average load meets a burst of concurrent requests, each holding a buffer or a connection or a partial response, and the sum briefly exceeds the ceiling. This is the case where raising the limit is the fix — but it's a minority of cases, and you only earn the right to conclude it after ruling out the others.

Native / off-heap memory the runtime never sees. This is the big one. A Node.js service that processes images through sharp is calling into libvips, a native C library that allocates memory entirely outside the V8 heap. The Node garbage collector neither sees nor manages it. A JVM service using direct byte buffers or memory-mapped files allocates off-heap, outside the space -Xmx governs. A Python service calling a C extension — NumPy, a database driver, a serialisation library — allocates in the C runtime, invisible to Python's own accounting. In every case the heap dashboard stays flat and green while RSS climbs until the cgroup ceiling is hit. The limit "hides" the leak because the tool you'd naturally reach for to measure memory is measuring the wrong memory.

Unbounded growth from an error or retry loop. This is subtler than a classic leak, and it's the root cause of our opening mystery. A client that cannot complete an operation — connect to a broker, find a topic, reach a dependency — may enter a tight retry loop, and if each iteration accretes state (buffered metadata, retry bookkeeping, log lines held in memory, growing backoff structures), the heap grows not because the program is wrong about memory but because it is stuck, and being stuck has a memory cost that compounds. There is no leak to find in the allocation sense. There is a hot loop that should not be running at all.

Runtime heap misaligned with the container limit. A runtime that isn't container-aware sizes itself to the machine, not the cgroup. A JVM without container support enabled sees the host's total RAM and sets a heap far larger than your limit; the first time it tries to grow into that heap, the cgroup kills it. Or the reverse: someone sets -Xmx or NODE_OPTIONS=--max-old-space-size larger than the container limit, and the runtime cheerfully tries to use more than the kernel will allow. The runtime and the cgroup disagree about how much memory exists, and the cgroup always wins.

QoS and eviction interplay. Finally, your pod may die not because it misbehaved but because of where it sits in the eviction hierarchy when the node is squeezed. Under node pressure, BestEffort pods die first, then Burstable pods that have exceeded their requests, while Guaranteed pods are the last to be touched. A perfectly well-behaved Burstable service can be evicted or OOM-scored into the line of fire because something else on the node ate the memory. This is the mechanism behind co-located cascades, and it's why your QoS class is a survival trait, not a billing detail.

The meta-point ties the list together: the symptom is downstream of at least five distinct mechanisms, and four of the five are made worse, not better, by raising the limit. The discipline this chapter is trying to build is the reflex to ask which mechanism before touching the only knob that's obvious.

3. Real-World Examples

Concepts don't stick; stories do. Here are the failure classes from the wild — some genuine production incidents, some reproducible classes I'll flag honestly as lab material. Each one maps to a mechanism above.

The native-heap OOM at 512 MiB

(A representative scenario and a lab you can run — this is a real, well-documented class, presented as a reproduction rather than claimed as one of our outages.)

Picture an image-upload service: a small Node.js application that accepts a photo, resizes it into a few thumbnail variants, and stores the results. It runs happily for months inside a 512 MiB limit, because on an ordinary day it handles a trickle of uploads and each resize completes in milliseconds. The memory graph is a flat, boring line. Everyone forgets it exists.

Then a promotion lands — a sale, a launch, the kind of day where upload volume jumps five-fold and arrives in concurrent bursts. Suddenly a dozen resize operations are in flight at once. Each one hands a buffer to libvips, which allocates native memory outside the V8 heap to decode and transform the image. The Node heap graph barely moves. RSS, however, climbs past 512 MiB in seconds, the cgroup OOM killer fires, the pod dies with exit 137, the kubelet restarts it, and three minutes later it happens again. Restart, die, restart, die — and every dashboard you own that's measuring the heap swears the service is healthy.

The wrong instinct is to raise the limit and move on. The right fix has three parts, and it's instructive because each part addresses a different facet:

# 1. Give native-heavy work a home with real headroom.
# The heap is small; the *native* allocations need the room.
resources:
requests:
memory: "3Gi"
limits:
memory: "4Gi"

# 2. Align the runtime heap *below* the container limit, leaving
# deliberate headroom for libvips' native allocations.
env:
- name: NODE_OPTIONS
value: "--max-old-space-size=1536" # ~1.5Gi V8 heap inside a 4Gi cgroup

# 3. Cap concurrency so native memory can't be multiplied without bound.
- name: IMAGE_MAX_CONCURRENCY
value: "3" # at most 3 resizes in flight

Better still, this kind of bursty, native-heavy work doesn't belong in your latency-sensitive request path at all. Move it to a dedicated Job or a worker pool on its own node group, throttle it to a fixed concurrency, and let the rest of your services run undisturbed. The lesson that generalises: native memory is invisible to the GC, so you size for RSS and you bound concurrency, because every concurrent native operation multiplies the part of memory your dashboards can't see.

The limit that was simply too low

The honest, boring case. A service sized for average load, no native trickery involved, meets a genuine concurrency burst and exhausts a ceiling that was always too tight. Here the metrics do tell the truth — RSS rises to meet the limit under load — and raising the limit (or, better, fixing the unbounded concurrency that let load translate directly into memory) is the correct fix. It's listed here mostly so you remember it exists, because after enough native-heap mysteries you start to assume every OOM is exotic. Sometimes the limit really is too low. You're allowed to conclude that — after you've checked it isn't one of the others.

The JVM that thought it owned the machine

A Java service deployed with a 2 GiB limit, OOMKilled within seconds of starting under any real load. The cause: the runtime wasn't reading the cgroup. An older JVM, or one with container support disabled, inspects the host and sees — on one of our hypervisor-backed nodes — 32 GiB of "available" RAM, and sets a default heap sized to a fraction of that, far beyond the 2 GiB the cgroup will permit. The first time the heap tries to grow into the space the JVM believes it has, the cgroup kills it.

The fix is to make the runtime container-aware and size the heap as a fraction of the limit, not the host:

env:
- name: JAVA_TOOL_OPTIONS
# Modern, container-aware: heap is 75% of the cgroup limit, not the host.
value: "-XX:MaxRAMPercentage=75.0"
resources:
requests:
memory: "2Gi"
limits:
memory: "2Gi" # requests == limits → Guaranteed; see below

The mirror-image of this bug is setting -Xmx (or Node's --max-old-space-size) larger than the limit by hand — telling the runtime it may use more memory than the kernel will ever allow. Either way the root cause is the same: the runtime and the cgroup disagree about how much memory exists, and the cgroup is the one holding the SIGKILL.

The retry loop that grew a heap

The opening mystery, in its essential shape, belongs here — and we'll give it the full postmortem treatment shortly. A client that can't find what it's looking for enters a hot loop, each iteration retains a little more state, and the heap grows until the ceiling is hit. No allocation bug. No leak in the textbook sense. A process that is stuck, paying for being stuck in memory. The tell, when you finally look, isn't in the metrics at all — it's in the logs, the same warning scrolling past dozens of times a second. We'll come back to it.

QoS classes and the order of death

Every pod is assigned one of three Quality-of-Service classes, and the kubelet derives that class purely from the relationship between requests and limits:

# Guaranteed: requests == limits for every resource.
resources:
requests: { memory: "4Gi", cpu: "2" }
limits: { memory: "4Gi", cpu: "2" } # Burstable: requests set, but less than limits (the common case).
resources:
requests: { memory: "2Gi", cpu: "500m" }
limits: { memory: "6Gi", cpu: "2" } # BestEffort: no requests or limits at all.
resources: {}

This classification is not bookkeeping. The kubelet translates it directly into the kernel's victim-selection score, oom_score_adj, which the global OOM killer consults when the node is out of memory:

  • Guaranteed pods are set to oom_score_adj = -997 — the most protected workload tier, last to be killed.
  • BestEffort pods are set to oom_score_adj = 1000 — the most expendable, first to be killed.
  • Burstable pods land somewhere between 2 and 999, scaled so that a pod requesting more of the node's memory gets a lower (safer) score.

In practice, most platform services are written with requests below limits — a service requesting 2 GiB and limited to 6 GiB is Burstable — which means almost everything you run is Burstable, and therefore in line ahead of your Guaranteed workloads when a node is squeezed. That's usually the right trade for density, but it has a consequence: under host-level memory pressure, your ordinary services are the kernel's preferred victims, and they can be killed while sitting well under their own limits. Meanwhile the control-plane's static pods — the apiserver, etcd — sit in the protected tier, which is exactly why, in Postmortem 2, a memory cascade takes the control plane down one member at a time instead of all at once. The QoS hierarchy is the difference between an incident and an outage.

Killed → restart → CrashLoopBackOff

Finally, the pattern that bridges into the next chapter. When a container dies and the kubelet restarts it, and it dies again quickly, the kubelet stops restarting it immediately and starts backing off. The delay grows exponentially — 10 seconds, then 20, 40, 80, 160 — capped at 5 minutes, and the pod sits in CrashLoopBackOff between attempts:

NAME                          READY   STATUS             RESTARTS         AGE
api-7d9f8c5b4-2xk9p 0/1 CrashLoopBackOff 8 (2m11s ago) 34m

CrashLoopBackOff is not itself an error — it's the kubelet protecting your node from a pod that's failing fast, by refusing to spin it in a tight loop. The error is whatever keeps killing the container. An OOM that recurs is one way to land here; a process that exits on a bad config or a missing dependency is another. We'll take CrashLoopBackOff apart properly — the back-off curve, how to read it, and how liveness and readiness probes both rescue you and, misconfigured, trap you — in → Chapter 2. For now, recognise it as the downstream consequence of a container that won't stay alive, and note that the number in parentheses (RESTARTS) is one of your most important early signals.

4. How to Debug

There is an order to this, and following the order is most of the skill. The instinct under pressure is to jump to the fix; the discipline is to characterise first. Here is the sequence.

Step 1 — Confirm it's actually an OOM, and which kind

Do not trust a dashboard that says "memory looks fine." Go to the source of truth, the container's last terminated state:

$ kubectl describe pod api-7d9f8c5b4-2xk9p
...
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Started: Sat, 14 Jun 2026 02:09:41 +0300
Finished: Sat, 14 Jun 2026 02:14:55 +0300
Restart Count: 8

Reason: OOMKilled and Exit Code: 137 confirm the mechanism. Now determine which OOM it was — cgroup or host — by reading the kernel log on the node:

# A cgroup OOM (the container exceeded its own limit):
$ dmesg -T | grep -i 'killed process'
[Sat Jun 14 02:14:55 2026] Memory cgroup out of memory: Killed process 12894 (python) ...

# A host OOM (the whole node ran out) reads differently:
[Sat Jun 14 02:14:55 2026] Out of memory: Killed process 12894 (python) ...

"Memory cgroup out of memory" means the workload hit its own ceiling — the cause is inside the pod. A plain "Out of memory" means the node ran out, and the cause may be an entirely different pod. This single line redirects your entire investigation. Don't skip it.

Step 2 — Verify your last fix actually took effect

This step saves more hours than any other, and it's the one everyone forgets. If someone "already raised the limit," confirm that the pod running right now is the one with the new limit. The deployment will happily report the new value in its template while an old pod keeps serving — a trap we'll see in full detail in Postmortem 1. Check whether the rollout actually completed:

$ kubectl get deploy events-analytics -n analytics
NAME READY UP-TO-DATE AVAILABLE AGE
events-analytics 1/1 0 1 41d

UP-TO-DATE: 0 is the alarm. It means zero pods are running the current template — the new ReplicaSet has not produced a healthy pod, and the "1 available" is the old one. When you see this, ask the deployment why:

$ kubectl describe deploy events-analytics -n analytics
...
Conditions:
Type Status Reason
---- ------ ------
Available True MinimumReplicasAvailable
ReplicaFailure True FailedCreate

ReplicaFailure: True. The new pods are failing to be created at all. The reason is one level down, in the ReplicaSet's events:

$ kubectl describe rs events-analytics-6c8b94f7d9 -n analytics
...
Events:
Warning FailedCreate ... Error creating: pods "events-analytics-6c8b94f7d9-..." is forbidden:
maximum memory usage per Container is 4Gi, but limit is 6Gi

There it is — your change was silently rejected by admission control (a LimitRange, which we'll meet properly in the next two steps), the old pod is still serving and still dying, and kubectl get deploy was showing you a limit that no running pod actually has. The dashboard lied because the change never landed. Always confirm the running pod is the changed pod before you conclude anything about whether the change worked.

Step 3 — Find the growth driver before you touch the limit

Once you know the kill is real and your view of it is honest, find what is consuming the memory — and crucially, do this before reaching for the limit, because for four of the five mechanisms the limit is the wrong lever. Tail the logs and watch for a hot loop:

$ kubectl logs -f events-analytics-... -n analytics | head
WARN Topic events.rsvp not found in cluster metadata; refreshing...
WARN Topic analytics.attendance.updated not found in cluster metadata; refreshing...
WARN Topic events.rsvp not found in cluster metadata; refreshing...
WARN Topic analytics.attendance.updated not found in cluster metadata; refreshing...
...

The same warnings, dozens of times a second. That is not a memory-sizing problem wearing an OOM costume — that's a stuck client in a hot retry loop, and the heap is growing because the process is spinning, not because the workload is large. No limit you set will outrun a loop that grows the heap faster than you can provision.

For the native-memory cases, watch RSS diverge from heap directly. From inside the container on a cgroup v2 node:

# Current cgroup memory usage (what the kernel charges you):
$ cat /sys/fs/cgroup/memory.current
4187593728

# OOM-kill bookkeeping — proof the kernel has fired here before:
$ cat /sys/fs/cgroup/memory.events
...
oom 5
oom_kill 5

If memory.current is climbing toward the limit while your application's heap metric is flat, you are looking at native or off-heap memory — libvips, JVM direct buffers, a C extension — and your fix is headroom and concurrency limits, not a bigger heap. The mantra for the whole section: characterise the kill, verify the change took, find the driver. Only then size.

5. How to Prevent

Debugging is recovery. Prevention is the work that means you sleep through the night. Five concrete practices, each tied to a failure we've now seen.

Make limit changes LimitRange-aware. The trap in Step 2 — a limit bump silently rejected because a namespace LimitRange caps the container lower than your new value — is entirely preventable by checking the constraint before you change the limit:

$ kubectl get limitrange -n analytics -o yaml
...
limits:
- type: Container
max:
memory: 4Gi # <-- any container limit above this is rejected
default:
memory: 1Gi
defaultRequest:
memory: 512Mi

If you're raising a container limit past the LimitRange max, you must raise the LimitRange in the same change, or your "fix" never runs. Treat the two as a single atomic edit.

Alert on restart rate and OOM count — not just readiness. Readiness tells you whether a pod is currently serving; it does not tell you that a pod has been quietly OOMing and restarting every five hours, or — as we'll see in this chapter's sidebar — that a workload has been blackholing writes for two days while reporting itself Running. Add alerts on the things readiness can't see:

# Restarts climbing for any container — catches recurring OOMs
# that self-recover before a human notices.
increase(kube_pod_container_status_restarts_total[15m]) > 3

# OOM kills happening at all, anywhere.
increase(container_oom_events_total[15m]) > 0

A pod that recovers on its own is not a pod that's healthy — it's a pod that's failing on a schedule. Restart-rate alerting is how you find the failures that hide inside successful recoveries.

When the same thing keeps OOMing after a bump, find the offender — don't re-bump. This is the apiserver lesson of Postmortem 2, stated as a rule: a recurring OOM that survives a limit increase is telling you the limit was never the constraint. Each additional bump buys a little time and obscures the cause further. The fix is to identify the thing — the loop, the watcher, the runaway client — that keeps filling whatever ceiling you give it, and address that. If your reflex after the second bump is a third bump, stop; the bump is the wrong tool and the failure is somewhere you haven't looked.

Size for RSS, with deliberate native headroom, and bound concurrency. For any workload that touches native memory — image processing, anything calling C libraries, JVM off-heap — set the runtime heap explicitly below the cgroup limit (MaxRAMPercentage, --max-old-space-size), leave room for the native allocations the GC can't see, and cap concurrency so native usage can't be multiplied without bound. Push bursty native-heavy work onto dedicated node pools or Jobs so it can't take latency-sensitive services down with it.

Use QoS deliberately. Set requests honestly so that Burstable scoring protects your more important pods, and make genuinely critical singletons Guaranteed (requests equal to limits) so they sit in the protected tier when a node is under pressure. On a dense self-hosted cluster, your QoS assignments are the policy that decides who survives a node-level squeeze — don't leave them to accident.

6. CI/CD, Hardware & Proxmox Factors

This is the section the managed-cloud books don't have, because the managed cloud hides exactly these failures from you. On self-hosted infrastructure they're yours to own.

The overcommit you signed up for

Your Kubernetes nodes are virtual machines, and those VMs live on a handful of physical hypervisors. It is entirely normal — often economically necessary — to allocate more memory to the VMs than the host physically has, betting that they won't all peak at once. Layer Kubernetes on top, where the sum of your pods' memory limits can itself exceed a node's RAM (because limits are ceilings, not reservations, and Burstable pods are expected not to all burst simultaneously), and you have two independent layers of overcommit stacked on each other:

   Physical hypervisor host  (e.g. 64 GiB physical RAM)
├── VM: worker-04 (allocated 32 GiB) ─┐
├── VM: worker-05 (allocated 32 GiB) ├─ allocated > physical: layer 1
└── VM: worker-06 (allocated 24 GiB) ─┘

Inside worker-05 (32 GiB):
├── pod A limit 6Gi ─┐
├── pod B limit 8Gi │ sum of limits > node RAM: layer 2
├── pod C limit 6Gi │ (fine *if* they don't all peak together)
└── pod D limit 16Gi ─┘

Figure 1.1 — Two stacked layers of overcommit. The host promises the VMs more RAM than it has; the node promises its pods more RAM than it has. Both bets pay off until the day they don't, and when the host loses, the kernel's global OOM killer picks a victim across the whole machine — by oom_score, which means a well-behaved Burstable pod under its own limit can die so that the host can survive.

This is the precise mechanism by which a pod is killed without breaching its own limit: the host ran out, not the pod. It is also why, on this kind of topology, we've had to grow a control-plane VM's allocation from 8 GiB to 16 GiB — not because the control plane was leaking, but because co-located pressure left it no room to breathe at peak. When you run your own hardware, "the pod is under its limit, so the pod is fine" is no longer a safe inference. The node can kill it for reasons that have nothing to do with the pod.

A few hardware-layer subtleties compound this. Hypervisor features like memory ballooning and same-page merging change how much RAM a VM actually has available moment to moment, which means the node's view of its own memory and the host's view can disagree — yet another layer that can lie to you. Three layers now: the cgroup limit, the node's available memory, and the hypervisor's real allocation. Each can report something the layer below it will contradict.

The pipeline that shipped a change that never ran

There's a CI/CD failure mode hiding in the LimitRange trap, and it's worth making explicit because it's invisible to most pipelines. A change to a Helm values file — raising a memory limit — passes review, passes CI, and kubectl apply (or your GitOps controller) reports success. The pipeline is green. The deployment's template shows the new value. And yet, because admission control silently rejected the new pods, production is still running the old pod with the old limit. Green pipeline, unchanged production.

The defence is to gate on the outcome, not the submission. A successful kubectl apply only means the API server accepted your object; it says nothing about whether the resulting pods ever became healthy. In a CD pipeline, follow the apply with a rollout check:

$ kubectl rollout status deploy/events-analytics -n analytics --timeout=120s
Waiting for deployment "events-analytics" rollout to finish: 0 of 1 updated replicas are available...
error: deadline exceeded # <-- the pipeline should FAIL here, not pass

If you run GitOps, the same discipline applies: a reconciler will happily report Synced (the desired state was applied) while the workload is Degraded (the applied state never became healthy). Watch health, not just sync. The distinction between "the change was accepted" and "the change is running" is exactly the gap a LimitRange rejection lives in, and a pipeline that only checks acceptance will ship silent failures all day.

Lab first, especially for the host-level kill

The host-RAM OOM — the kernel killing an innocent pod because the node ran dry — is genuinely dangerous to learn about in production, and it's one I'll show you how to reproduce deliberately rather than wait to be ambushed by. We'll do exactly that in the lab reproduction inside Postmortem 2. The rule the whole section serves: on self-hosted infrastructure, reproduce the failure class on a scratch node until your mental model matches what the kernel actually does, before you're standing in the incident.

7. Production Postmortems

Two cases. The first is a genuine production incident, told in full, with the three-strikes structure this chapter has been building toward. The second is a real cascade paired with an explicitly labelled lab reproduction for the textbook variant — because, as promised in the preface, I won't dress a reproduction up as history.

Postmortem 1 — The OOM that was never a memory problem

Status: genuine production incident, fully reproduced here, anonymised.

The symptom. A Python ML/analytics service began OOMing — exit 137, OOMKilled, recovering on its own roughly every five to six hours, accumulating about four restarts a day. Classic self-healing-on-a-schedule: bad enough to alert on, benign enough to be ignored, which is the most dangerous combination there is.

Strike one — "raise the limit." The first instinct, and the obvious one. A pull request (call it PR #181) bumped the service's memory limit from 4 GiB to 6 GiB in its Helm values. Merged, applied, pipeline green. The deployment now showed a 6 GiB limit. We considered it handled and moved on.

It was not handled. It hadn't even changed anything.

Strike two — "the bump took, problem solved." It hadn't taken. The platform's base chart renders a per-namespace LimitRange that caps any single container at 4 GiB. The new 6 GiB pod spec violated that cap, so admission control rejected it outright — maximum memory usage per Container is 4Gi, but limit is 6Gi — and the new ReplicaSet never produced a pod. The old 4 GiB pod kept right on serving, and kept right on OOMing, while kubectl get deploy cheerfully displayed 6Gi. The rollout had silently stalled with ReplicaFailure: True, and nobody had checked UP-TO-DATE. For some hours, every artefact we looked at told us the limit was 6 GiB, and the process dying every five hours had 4.

The fix for this layer was a second pull request (PR #203) that raised the LimitRange max to 8 GiB alongside the container limit, followed by a kubectl rollout restart to un-stick the wedged rollout and let a new pod actually schedule. Now a genuine 6 GiB pod was running.

It kept OOMing.

Strike three — "it's a memory-sizing problem at all." This was the real lesson, and it arrived only after the first two instincts had been thoroughly disproven: the problem was never the size of the memory. The service consumed eight Kafka topics — events.rsvp, analytics.attendance.updated, and six others — whose producers ran elsewhere and which, on this cluster, had never been created. The Kafka client, unable to find them, spun a tight metadata-refresh loop: topic not found, refresh metadata, topic still not found, refresh again, dozens of times a second, accreting state on every pass until the heap reached whatever ceiling we'd given it. Six gigabytes just bought the loop a little more runway before the same kill.

The actual fix changed no memory setting at all:

$ kafka-topics.sh --bootstrap-server kafka:9092 \
--create --topic events.rsvp --partitions 6 --replication-factor 3
# ...and the same for the other seven topics.

Within a single metadata refresh, the loop found its topics, stopped spinning, and the heap stabilised. The OOMs ended. No memory change was the fix. The limit was never the problem; the first two "fixes" were treating a symptom that wasn't even the right symptom — and the only reason we ever got to the real cause was that, eventually, someone tailed the logs instead of reaching for the limit a third time.

The three-strikes arc of this incident is the reason it opens the chapter. It is "the limit is not the leak," "verify your change actually ran," and "find the driver before you size," all lived in sequence, on one pod, in one afternoon.

Postmortem 2 — A cascade on co-located members (and the host-OOM lab)

Status: the cascade is a genuine production incident (timeline 15–17 May 2026), anonymised. The pure host-RAM kill at the end is an explicitly labelled lab reproduction, not a historical event.

The incident. One control-plane API server began climbing — 1.5 GiB to 5.2 GiB — while its two peers sat steady at around 1.5 GiB. The cause was a watch storm: a long-running, expensive watch pinned to whichever API server happened to be answering it, driven by a controller relisting far more than it needed (the usual suspects in our stack are the GitOps application controller, the secrets operator, and the storage manager). That one API server climbed alone until it hit its 8 GiB limit, was OOMKilled with exit 137, and dropped into CrashLoopBackOff.

Then the cascade. With that member down, its clients reconnected and the watch traffic redistributed to the next API server — which began climbing, hit the same ceiling, and OOMed in turn. For about ten minutes the cluster ran on a single Ready API server while its siblings cycled through CrashLoopBackOff:

$ kubectl get pods -n kube-system | grep apiserver
kube-apiserver-cp-1 0/1 CrashLoopBackOff 3 (40s ago) ...
kube-apiserver-cp-2 1/1 Running 0 ...
kube-apiserver-cp-3 0/1 CrashLoopBackOff 4 (28s ago) ...

Crucially, the control plane went down one member at a time, not all at once. That's the QoS hierarchy from Section 3 doing its job: control-plane static pods sit in the protected tier, the redistribution was sequential rather than simultaneous, and so there was always one survivor. On a less carefully prioritised topology this is the difference between a degraded control plane and a dead one.

The bump trap, again. The memory limit on these API servers had already been raised once — 4 GiB to 8 GiB — in a previous incident. It made no difference here, because a sustained watch storm simply refills whatever ceiling you give it; 8 GiB just delayed each OOM by a few minutes. The real fix was not a third bump but finding the offending watcher and constraining it — throttling the misbehaving client through API Priority and Fairness, and fixing the controller's relist behaviour at the source. This is the "find the offender, don't re-bump" rule from Section 5, learned the hard way at control-plane scale.

A second front (→ Chapter 12). Concurrently, a worker node crossed a disk-pressure threshold and entered a node-pressure eviction loop: the kubelet evicted DaemonSet pods to reclaim disk, the DaemonSet controller immediately recreated them, the node re-evicted them, and within minutes there were hundreds of Evicted and Pending pods churning. Note the shape — node-level pressure evicting pods that never breached their own limits, the exact dynamic as the host-OOM cascade, but driven by disk, not RAM. Same lesson, different resource. Node pressure in all its forms is the subject of → Chapter 12, and this is your first sight of it.

Lab reproduction — the textbook host-RAM OOM. Here is the variant I haven't lived through in production and won't pretend I have: the kernel's global OOM killer terminating a perfectly innocent, under-limit pod because the node ran out of RAM. The cascade above is co-location pressure on the control plane; this is the cleanest possible demonstration of "the host ran out, not the pod," and you should reproduce it on a scratch node rather than trust my description.

The recipe is deterministic. On a small scratch node, schedule pods whose memory limits sum to more than the node's RAM — the overcommit of Figure 1.1 — then drive one of them to actually consume its share with a stress workload. Alongside them, run two innocent pods doing nothing: one Guaranteed (requests equal to limits) and one BestEffort (no requests or limits at all). As the node's free memory collapses, watch the kernel's global OOM killer choose its victim by oom_score:

# On the node, watch the kernel make its choice:
$ dmesg -T -w | grep -i 'out of memory'
[...] Out of memory: Killed process 20114 (pause-besteffort) ...

The BestEffort pod — oom_score_adj = 1000 — dies first, while the Guaranteed pod — oom_score_adj = -997 — survives untouched, neither of them having done anything wrong. The innocent BestEffort pod was killed not for its own behaviour but for its QoS class, so that the node could survive the memory the stress pod consumed. Run this until it's in your hands and not just on this page. It is the mechanism every dense self-hosted cluster carries, and the day you meet it in production you'll want it to feel familiar.

Sidebar — A pod can be Running and still be dead to its users.

One more incident, not an OOM, but the sharpest illustration of why "the pod is Running" is a claim you must never fully trust — and a direct bridge into the next chapter.

An operator-managed Postgres primary filled its 30 GiB write-ahead-log volume (the usual cause: a stale replication slot or a failing archive pinning WAL so it can't be recycled). With its WAL volume full, Postgres refused to start — not enough WAL disk space — and the container exited in seconds, landing in CrashLoopBackOff with the back-off climbing toward five minutes. So far, an ordinary lifecycle failure.

The dangerous part: the operator never failed over. Its instance-manager subprocess stayed alive inside the pod and kept answering status checks as "Postgres is still booting" — and the operator's failover trigger is pod gone, not process dead. So the pod reported itself as present and being managed, the failover that should have promoted a replica never fired, and writes were blackholed for two days behind a pod that looked, by every coarse signal, alive.

This is the cleanest possible statement of a principle that runs through this entire book: a pod being alive is not the same as a workload being healthy. A subprocess can keep a pod looking Running while the thing the pod exists to do is completely dead. It's also exactly why Section 5 insists on alerting on restart rate and real workload health rather than readiness alone — and it's the question that → Chapter 2 is built to answer, with liveness and readiness probes that ask whether the workload, not just the process, is actually doing its job.

8. Lessons Learned

Distilled to the principles worth carrying out of this chapter and into the rest of the book:

  1. OOMKilled is a symptom with at least five different mechanisms. Exit 137 tells you a process was killed by SIGKILL for memory reasons; it tells you nothing about why. Characterise the kill before you change anything.
  2. The limit is not the leak. Raising the limit is the universal first instinct and the wrong fix for four of the five mechanisms. It buys time and hides the cause — the worst combination. Find the driver first.
  3. cgroup OOM, host OOM, and eviction are three different failures. One is the pod's fault, one is the node's, one is the kubelet protecting the node. The kernel log line — "Memory cgroup out of memory" versus "Out of memory" — tells you which, and redirects your entire investigation.
  4. Verify your change actually runs. Admission control (a LimitRange) can silently reject a change while every dashboard shows it applied. UP-TO-DATE: 0 and ReplicaFailure: True are the tells. The deployment's template is the desired state, not the running one. Confirm the running pod is the changed pod.
  5. Native and off-heap memory is invisible to the GC. The heap graph stays flat while RSS climbs to the ceiling. Watch the cgroup (memory.current) and RSS, not just heap, and bound concurrency for native-heavy work because every concurrent native operation multiplies the part you can't see.
  6. A recurring OOM that survives a bump means the bump is the wrong tool. Find the offending loop, watcher, or client and constrain it. If your reflex after the second bump is a third, stop.
  7. Alert on restart rate and OOM count, not just readiness. A pod that self-recovers is failing on a schedule, not healthy. And a pod can report Running while its workload has been dead for days. Readiness is necessary and nowhere near sufficient.
  8. On self-hosted infrastructure, the node can kill an innocent pod. Two stacked layers of overcommit — hypervisor and cgroup — mean "the pod is under its limit" is no longer a safe inference. Use QoS deliberately; it's the policy that decides who survives a squeeze.
  9. Lab first. Reproduce the class on a scratch node until your model matches the kernel's behaviour, before you're standing in the incident.

What's next

We've established how a pod lives and dies, and we've seen the failure that bridges everything: a container that won't stay alive, restarting until the kubelet backs it off into CrashLoopBackOff. We've also seen, twice — in the Postgres sidebar and the silent rollout stall — that a pod being alive is not the same as a workload being healthy.

That gap is Chapter 2's entire subject. We'll take CrashLoopBackOff apart: the exponential back-off curve and how to read it, why a pod loops, and how liveness and readiness probes are supposed to close the alive-versus-healthy gap — and how, misconfigured, they pry it open instead, killing healthy pods or hiding dead ones. The eviction loop we glimpsed in Postmortem 2, and the overcommit of Figure 1.1, open the door to Chapter 12 on node pressure and hardware. We'll be back at both doors.

Field notes — Chapter 1

  • Exit 137 = 128 + 9 (SIGKILL). It's a Unix convention, not a Kubernetes code.
  • First question, always: cgroup OOM, host OOM, or eviction? Read the kernel log.
  • UP-TO-DATE: 0 / ReplicaFailure: True = your change never ran. Check before you theorise.
  • Heap flat + RSS climbing = native/off-heap memory. Size for RSS, cap concurrency.
  • Second bump didn't work? Don't do a third. Find the offender.
  • Running ≠ healthy. Alert on restart rate, not just readiness.
  • Two layers of overcommit on self-hosted: the host can kill an innocent pod. QoS decides who.

Want to continue reading?

Get the complete Kubernetes Postmortem

$15 or KES 1,950 via M-PESA