Kubernetes Postmortem cover

By Mike Murango

Kubernetes Postmortem

Real production failures from a self-hosted cluster — and how to survive them

$15

or KES 1,950 with M-PESA · digital edition · read online & listen

Read a free preview →
Already purchased? Sign in to read

About the book

Practical lessons from Kubernetes incidents, failures and production recovery on self-hosted, bare-metal clusters: OOMKilled pods, NotReady nodes, etcd NOSPACE, Kafka lag, Vault sync errors and the runbooks that fixed them.

Who this book is for

  • SREs and platform engineers running Kubernetes in production
  • Teams self-hosting clusters on bare metal or Proxmox
  • Engineers who learn best from real incidents rather than toy examples

What you'll learn

  • Diagnose OOMKilled, CrashLoopBackOff and Pending pods from first principles
  • Debug Services, DNS, ingress, TLS and Cilium networking
  • Recover NotReady nodes, the control plane and etcd (including NOSPACE)
  • Operate Longhorn, PostgreSQL/pgBouncer, Redis, MongoDB and Kafka on Kubernetes
  • Run GitOps with ArgoCD, secrets with Vault/ESO, and policy with Gatekeeper
  • Build runbooks, GameDays and an SRE culture that prevents repeats

Table of contents

  1. —PrefaceFree sample
  2. Part I · Chapter 1Pod Lifecycle & OOMKilledFree sample
  3. Part I · Chapter 2CrashLoopBackOff & Probe Failures
  4. Part I · Chapter 3Container Images, Build Args & Registry Pulls
  5. Part I · Chapter 4Pod Scheduling, Pending & Resource Requests
  6. Part II · Chapter 5Services, Endpoints, targetPort & DNS
  7. Part II · Chapter 6Ingress Controllers, Kong & HTTPRoutes
  8. Part II · Chapter 7TLS, Certificates & cert-manager
  9. Part II · Chapter 8Authentication, JWT & Identity Drift
  10. Part III · Chapter 10Cilium CNI: Pod & Node Networking, NetworkPolicy & CNPs
  11. Part III · Chapter 11The "NotReady" Node: kubelet, containerd & CNI Bring-up
  12. Part III · Chapter 12Node-Level OOM, Memory Pressure & the Proxmox Over-Commit Cascade
  13. Part III · Chapter 13Kured, Node Drains & OS Patching
  14. Part III · Chapter 14Proxmox HA, VM Placement & the Bare-Metal Layer
  15. Part IV · Chapter 15API Server Performance & Slow Responses
  16. Part IV · Chapter 16Control Plane Repair: etcd, Scheduler & Controller-Manager Leaders
  17. Part IV · Chapter 17etcd Operations: Backup, Defrag, NOSPACE & Disaster Recovery
  18. Part V · Chapter 18Persistent Storage & Longhorn
  19. Part V · Chapter 19PostgreSQL, Connection Pooling & pgBouncer
  20. Part V · Chapter 20Redis: Caching, Sessions & Pub/Sub
  21. Part V · Chapter 21MongoDB: ReplicaSets & Operators
  22. Part V · Chapter 22Kafka: Consumer Lag, Poison Pills, Rebalances & ISR
  23. Part V · Chapter 23Backup & Disaster Recovery: Velero, WAL Archiving & Off-Site Replication
  24. Part VI · Chapter 24GitOps with ArgoCD: Sync, Drift & ApplicationSets
  25. Part VI · Chapter 25Secrets, Vault & ESO: Sealed Vault, Sync Errors & Token Expiry
  26. Part VI · Chapter 26Admission Control & Policy: OPA Gatekeeper & Namespace Isolation
  27. Part VI · Chapter 27CI/CD & GitHub Actions ARC Runners
  28. Part VI · Chapter 29Security, RBAC, Falco & Supply Chain
  29. Part VII · Chapter 30Cluster Migration & Multi-Cluster Communication
  30. Part VII · Chapter 31Observability with Qentra: Identity-Aware Causality & Incident Management
  31. Part VII · Chapter 32Managing a Full Production Cluster: Runbooks, GameDays & SRE Culture
  32. Part VII · AppendixThe Field Guide

FAQ

Is this beginner material?

It assumes you know Kubernetes basics. Each chapter starts from the symptom and works down to the root cause.

How do I read it after paying?

Payment unlocks the book in your library instantly. Read in your browser, with progress synced across devices.

Can I pay with M-PESA?

Yes. Checkout is handled securely by Paystack. Pay by card in US dollars, or with M-PESA in Kenyan Shillings.