How Production Actually Breaks
The Universal Failure Map
(Any System. Any Scale. Any Company.) Why This Exists
Most engineers think outages are:
-
random
-
complex
-
unique
They are not. All production failures fall into a small number of repeatable shapes. If you don’t know the shapes, every incident feels new. If you know them, incidents feel… boring. And boring is good.
The First Brutal Truth
Systems don’t fail where you look.
They fail where you didn’t draw the boundary.
The Universal Failure Stack (Memorize This)
Every production system: monolith, microservices, Kubernetes, mainframe - collapses along the same vertical stack:
User Traffic
↓
Entry Point (LB / Gateway / Edge)
↓
Service Logic
↓
Runtime (VM / Container / Process)
↓
Operating System
↓
Network
↓
Storage
↓
Control Plane / State
↓
Physical Reality
Every outage lives at exactly one layer first. The rest are symptoms.
The 5 Universal Failure Modes
Forget vendor-specific nonsense. Everything maps to one of these:
1. State Is Wrong
-
etcd corruption
-
stale caches
-
split-brain
-
config drift
-
“It worked yesterday”
Symptoms:
-
Weird behavior
-
Inconsistent results
-
Restarts don’t help
2. Capacity Is Exhausted
-
CPU lies
-
Memory leaks
-
conntrack full
-
file descriptors gone
-
IPs exhausted
Symptoms:
-
Latency before errors
-
Random failures
-
“But metrics look fine”
3. Dependency Is Slow or Dead
-
DNS
-
Auth
-
Metadata APIs
-
Downstream services
Symptoms:
-
Timeouts
-
Retries explode
-
Cascading failures
4. Control Plane Is Impaired
-
API servers
-
Schedulers
-
Controllers
-
Cloud control planes
Symptoms:
-
You can’t change anything
-
Running things still run
-
Humans panic first
5. Reality Disagrees with Your Model
-
Pod says Running, process is dead
-
Node says Ready, kernel is dying
-
“Green dashboards, red users”
Symptoms:
-
Ghosts
-
Illusions
-
Debugging hell
The Pattern Most Engineers Miss
Incidents rarely start where alerts fire.
Alerts trigger at the surface. Failures originate deep.
That gap is where bad engineers thrash and good engineers navigate.
Real Incident Translation Table
Symptom || Likely Reality
-
Random 500s || Resource pressure or dependency slowness
-
Everything slow || DNS, network, or storage
-
Can’t deploy || Control plane issue
-
Works after restart || State or memory corruption
-
Only prod breaks || Scale exposed a hidden limit
The Rule of Containment
Before debugging anything, answer this: Is this failure contained or spreading?
-
Contained → slow down, investigate
-
Spreading → stop the bleed first
Most outages escalate because engineers debug before containment.
Senior Engineer Principle #1
You don’t fix outages.
You collapse the search space.
The Universal Failure Map is how you do that.
PLAYBOOK 0.2 (Next)
Control Plane vs Data Plane
“Stop Mixing Them or You’ll Always Be Late”
We will cover:
-
Why engineers misdiagnose healthy systems as dead
-
Why scaling doesn’t fix control-plane outages
-
Why “kubectl works” means almost nothing
-
How to tell in 60 seconds which plane is broken
PLAYBOOK 0.3 (Then)
Why Healthy Metrics Mean Nothing During Incidents
We will destroy:
-
Blind faith in dashboards
-
CPU obsession
-
“Green = Good” thinking
And replace it with:
-
Leading indicators
-
Negative signals
-
Absence-of-evidence detection
Series 0 Exit Criteria (Non‑Negotiable)
If someone finishes Series 0, they should be able to say:
-
“I know where to look first.”
-
“I know what not to touch.”
-
“I don’t panic when data disagrees.”
-
“I can explain this outage to leadership calmly.”
If they can’t; they’re not ready for Series 1.
Discussion
to read and post comments.