
By Mike Murango
Kubernetes Postmortem
Real production failures from a self-hosted cluster — and how to survive them
About the book
Practical lessons from Kubernetes incidents, failures and production recovery on self-hosted, bare-metal clusters: OOMKilled pods, NotReady nodes, etcd NOSPACE, Kafka lag, Vault sync errors and the runbooks that fixed them.
Who this book is for
- SREs and platform engineers running Kubernetes in production
- Teams self-hosting clusters on bare metal or Proxmox
- Engineers who learn best from real incidents rather than toy examples
What you'll learn
- Diagnose OOMKilled, CrashLoopBackOff and Pending pods from first principles
- Debug Services, DNS, ingress, TLS and Cilium networking
- Recover NotReady nodes, the control plane and etcd (including NOSPACE)
- Operate Longhorn, PostgreSQL/pgBouncer, Redis, MongoDB and Kafka on Kubernetes
- Run GitOps with ArgoCD, secrets with Vault/ESO, and policy with Gatekeeper
- Build runbooks, GameDays and an SRE culture that prevents repeats
Table of contents
- —PrefaceFree sample
- Part I · Chapter 1Pod Lifecycle & OOMKilledFree sample
- Part I · Chapter 2CrashLoopBackOff & Probe Failures
- Part I · Chapter 3Container Images, Build Args & Registry Pulls
- Part I · Chapter 4Pod Scheduling, Pending & Resource Requests
- Part II · Chapter 5Services, Endpoints, targetPort & DNS
- Part II · Chapter 6Ingress Controllers, Kong & HTTPRoutes
- Part II · Chapter 7TLS, Certificates & cert-manager
- Part II · Chapter 8Authentication, JWT & Identity Drift
- Part III · Chapter 10Cilium CNI: Pod & Node Networking, NetworkPolicy & CNPs
- Part III · Chapter 11The "NotReady" Node: kubelet, containerd & CNI Bring-up
- Part III · Chapter 12Node-Level OOM, Memory Pressure & the Proxmox Over-Commit Cascade
- Part III · Chapter 13Kured, Node Drains & OS Patching
- Part III · Chapter 14Proxmox HA, VM Placement & the Bare-Metal Layer
- Part IV · Chapter 15API Server Performance & Slow Responses
- Part IV · Chapter 16Control Plane Repair: etcd, Scheduler & Controller-Manager Leaders
- Part IV · Chapter 17etcd Operations: Backup, Defrag, NOSPACE & Disaster Recovery
- Part V · Chapter 18Persistent Storage & Longhorn
- Part V · Chapter 19PostgreSQL, Connection Pooling & pgBouncer
- Part V · Chapter 20Redis: Caching, Sessions & Pub/Sub
- Part V · Chapter 21MongoDB: ReplicaSets & Operators
- Part V · Chapter 22Kafka: Consumer Lag, Poison Pills, Rebalances & ISR
- Part V · Chapter 23Backup & Disaster Recovery: Velero, WAL Archiving & Off-Site Replication
- Part VI · Chapter 24GitOps with ArgoCD: Sync, Drift & ApplicationSets
- Part VI · Chapter 25Secrets, Vault & ESO: Sealed Vault, Sync Errors & Token Expiry
- Part VI · Chapter 26Admission Control & Policy: OPA Gatekeeper & Namespace Isolation
- Part VI · Chapter 27CI/CD & GitHub Actions ARC Runners
- Part VI · Chapter 29Security, RBAC, Falco & Supply Chain
- Part VII · Chapter 30Cluster Migration & Multi-Cluster Communication
- Part VII · Chapter 31Observability with Qentra: Identity-Aware Causality & Incident Management
- Part VII · Chapter 32Managing a Full Production Cluster: Runbooks, GameDays & SRE Culture
- Part VII · AppendixThe Field Guide
FAQ
Is this beginner material?
It assumes you know Kubernetes basics. Each chapter starts from the symptom and works down to the root cause.
How do I read it after paying?
Payment unlocks the book in your library instantly. Read in your browser, with progress synced across devices.
Can I pay with M-PESA?
Yes. Checkout is handled securely by Paystack. Pay by card in US dollars, or with M-PESA in Kenyan Shillings.