The Night SSL Quietly Turned Against Everyone 3:00 AM.
# The Night SSL Quietly Turned Against Everyone
3:00 AM.
No deployments.
No infrastructure changes.
No database maintenance.
Yet dashboards suddenly lit up with a flood of HTTPS failures.
Customers weren't reporting slow applications.
They were seeing something far worse.
"Your connection is not private."
Browsers refused to trust production.
The war room immediately checked the ingress controllers. Healthy.
Pods? Healthy.
Load balancers? Healthy.
Applications? Running perfectly.
Everything looked alive.
Except nobody could securely reach it.
Someone opened the certificate dashboard.
Dozens...
Then hundreds...
Certificates stuck in Pending.
The team watched cert-manager repeatedly retry renewals, only to be rejected again.
It wasn't Kubernetes.
It wasn't networking.
It wasn't DNS.
The enemy was hiding outside the cluster.
Every certificate had been issued on the very same day months ago. As they approached expiration together, nearly 180 renewals launched simultaneously. Let's Encrypt saw the avalanche and did exactly what it was designed to do—protect itself with rate limits.
Production had accidentally DDoSed its own Certificate Authority.
Only a fraction of certificates renewed successfully.
The rest expired one by one while perfectly healthy applications became unreachable.
The fix wasn't "increase retries."
The real solution started much earlier: stagger certificate issuance, tune renewBefore windows, introduce meaningful renewal distribution instead of relying on default jitter, and where appropriate, use DNS-01 validation to avoid HTTP-01 bottlenecks. Sometimes the biggest outage is caused by everything working exactly as configured.
Production rarely collapses because one component fails.
It collapses when hundreds of healthy components make the exact same decision at the exact same moment.
At InfraThrone, we recreate production incidents like these—not to memorize commands, but to understand why seemingly harmless defaults can evolve into platform-wide outages. Because certificates, just like infrastructure, don't fail in isolation—they fail as part of a story.
Discussion
to read and post comments.