The 9 AM Push That Froze Every Deployment
The 9 AM Push That Froze Every Deployment
9:00 AM.
The coffee was still hot.
Developers across the organization started pushing their feature branches before the morning stand-up. New APIs, bug fixes, UI tweaks—nearly 50 pull requests landed within minutes.
Normally, CI pipelines would spring to life immediately.
Today...
Nothing happened.
Pull requests sat in "Queued".
One minute became ten.
Ten became forty-five.
Slack channels filled with the same question:
"Is GitHub Actions down?"
It wasn't.
The problem was much closer to home.
The engineering team relied on 10 self-hosted GitHub Actions runners running on EC2 instances. Under normal traffic, that capacity was more than enough. But the daily 9 AM push hour created a sudden burst of 50 concurrent jobs. The runners were saturated almost instantly, and every new workflow joined an ever-growing queue.
The autoscaler did react—but only after the queue had already exploded. New EC2 instances needed nearly three minutes to boot, register as runners, and become available. By then, dozens of builds were already waiting.
The bottleneck wasn't compilation.
It wasn't testing.
It wasn't GitHub.
It was infrastructure reacting to demand after the demand had arrived.
The postmortem uncovered a familiar anti-pattern: reactive scaling. Instead of provisioning runners only after queues formed, the platform team adopted ephemeral GitHub Actions runners backed by Karpenter and AWS Auto Scaling Groups, allowing runners to be created and destroyed automatically. More importantly, they introduced pre-warmed runner capacity before predictable traffic spikes like the morning push window.
The result wasn't just shorter queue times—it was uninterrupted developer flow. Engineers stayed focused, features shipped on schedule, and the CI platform stopped being the slowest member of the team.
Production incidents aren't always caused by broken applications.
Sometimes they're caused by the systems that build them.
At InfraThrone, we recreate incidents like this from real production environments. You don't just learn GitHub Actions, Kubernetes, or cloud autoscaling independently—you experience how they interact under real-world pressure, investigate the hidden bottlenecks, and build the mindset needed to solve outages before they impact an entire engineering organization.
Discussion
to read and post comments.