Production failures on infrastructure you own.
Every Kubernetes tutorial ends at the same place: a green Running status, a service that responds to curl, and a sense that you now understand the thing. You don't. Nobody does at that point, because the tutorial deliberately removed everything that makes Kubernetes hard. It gave you one pod on one node with no memory pressure, no admission controllers fighting your changes, no controller quietly relisting ten thousand objects, and no kernel deciding — at 02:14 on a Saturday — which of your processes gets to die so the others can live.
This book is about that second part. The part that starts the moment something that was working stops working, the dashboard insists everything is fine, and your first three ideas about what's wrong all turn out to be the same wrong idea wearing different clothes.
Why self-hosted, and why that changes everything
Most Kubernetes writing quietly assumes a managed control plane — EKS, GKE, AKS — where a cloud provider absorbs the failures you'd otherwise have to understand. Your etcd is somebody else's problem. Your node sizing is smoothed over by an autoscaler. Your control-plane memory leak is a line item on a status page you'll never see.
This book assumes none of that. It is written from inside a self-hosted, effectively bare-metal cluster: Kubernetes running on virtual machines on Proxmox hypervisors, with the full production stack — a CNI carrying real traffic, distributed block storage, an operator-managed Postgres, a Kafka cluster, GitOps reconciliation, secret management, policy enforcement, autoscaling — all of it owned, operated, and broken by the same small team.
When you run your own platform, there is no provider to page. The etcd is yours. The kernel OOM killer is yours. The overcommitted hypervisor where the sum of your pods' memory limits quietly exceeds the physical RAM of the host — that is very much yours. The failures in this book are not exotic. They are the ordinary consequences of operating real infrastructure, and they are invisible to anyone who has only ever rented a control plane.
Why postmortems
You do not learn distributed systems from the happy path. You learn them from the moment the happy path breaks and the abstraction leaks all over your terminal. So every chapter in this book is organised around real failures: what broke, what we believed at the time, what was actually true, how we fixed it, and — most importantly — how we'd catch it earlier next time.
There is one idea underneath all of it, and it is worth stating plainly because it will recur in every single chapter:
Your first instinct is usually wrong.
The obvious fix — raise the limit, restart the pod, add a node — almost always treats a symptom and hides a cause. It makes the red thing turn green for a few hours, which is the worst possible outcome, because now the real failure is both still present and harder to see. This book is an extended argument for the second and third instinct: the discipline to characterise a failure before you change anything, to verify that your change actually took effect, and to find the mechanism rather than mute the alarm.
A note on honesty
The incidents in this book are real. Names, namespaces, node identifiers, and anything that could identify a specific system have been changed. The mechanics have not — the error messages, the timelines, the cgroup behaviour, and the fixes are reproduced as they happened.
Where a failure class is real and well-known but I have not personally hit it in production, I will tell you so and show you how to reproduce it in a lab, rather than dressing up a lab exercise as a war story. You will see this most clearly in Chapter 1, where one of the two postmortems is a genuine production incident with pull-request numbers attached, and the other is a real incident paired with an explicitly labelled lab reproduction for the textbook variant I haven't been unlucky enough to live through. Treating a reproduction as history would teach you the wrong lessons about how often these things actually happen and what they actually look like. I won't do it.
How the book is built
There are thirty-two chapters, ordered from foundation to advanced. Each one follows the same spine:
- Definition & Core Concept — what the thing actually is, mechanically.
- Why It Happens — the causes, grouped by mechanism rather than by symptom.
- Real-World Examples — the failure class, told as stories, because dry enumeration doesn't stick.
- How to Debug — the exact commands and the order to run them in.
- How to Prevent — actionable, not aspirational.
- CI/CD, Hardware & Proxmox Factors — the self-hosted reality the managed-cloud books skip.
- Production Postmortems — at least two, with real timelines.
- Lessons Learned — the durable principles, distilled.
The chapters build on each other. When a concept here depends on one we haven't reached yet, I'll flag it with a forward arrow (→ Chapter 12). When a later chapter leans on something we've already established, it'll point back. By the end you should be carrying a single connected model of how this system fails, not thirty-two disconnected facts.
How to read it
Lab first. Every example in every chapter can be reproduced in kind, minikube, or a scratch namespace, and you should reproduce it before you read the fix. The failure will teach you more in five minutes of watching it happen than the explanation will in an hour. The explanation is here to confirm what you saw, not to substitute for seeing it.
If you operate your own cluster, this book is for you. If you're about to, it's a map of the territory ahead. And if you've only ever rented a control plane, consider it a preview of what your provider has been quietly doing on your behalf — and what you'd be signing up for the day you decide to bring it in-house.
Let's start where every cluster's troubles start: with a single pod, a memory limit, and a kernel that has run out of patience.