The Node That Lived for One Minute
# The Node That Lived for One Minute
Wednesday.
5:45 PM.
The traffic graph was climbing faster than expected.
Autoscaling kicked in.
Everything looked exactly as designed.
Karpenter detected pending pods, requested a new EC2 Spot instance, and AWS responded almost instantly. Another node was on its way, ready to absorb the incoming workload while keeping infrastructure costs low.
The operations dashboard looked healthy.
"Perfect timing," someone said.
Two minutes later, the node joined the Kubernetes cluster.
The team waited for workloads to land.
They never did.
Before the scheduler could place a single pod, another notification appeared.
EC2 Spot Interruption Notice.
The brand-new node had been marked for termination.
Sixty seconds after joining the cluster.
No applications had started.
No requests had been served.
No business value had been delivered.
Karpenter gracefully drained the node, removed it from the cluster, and immediately began provisioning another replacement.
The cycle repeated.
Nodes appeared.
Nodes disappeared.
Cloud costs quietly increased while application capacity barely moved.
At first glance, it looked like Kubernetes scheduling was broken.
It wasn't.
The scheduler was waiting for stable capacity.
The real culprit was hiding beneath it.
Karpenter had treated every Spot capacity pool as equally reliable. It provisioned instances from highly volatile Spot families without considering how likely they were to be reclaimed moments later. The platform was winning the race to launch nodes, only to lose them before they became useful.
The fix wasn't abandoning Spot instances.
It was making provisioning decisions more intelligent.
Teams diversified across multiple instance families and Availability Zones instead of depending on a single volatile pool. Karpenter disruption budgets prevented excessive node churn during scaling events, while node TTLs ensured underutilized nodes were recycled in a controlled manner. Historical EC2 Spot interruption patterns became part of provisioning decisions, steering workloads toward more stable capacity pools instead of simply choosing the cheapest available instance.
The next scale-up looked uneventful.
Nodes joined.
Pods scheduled.
Traffic stabilized.
Exactly as autoscaling should.
Cloud-native systems don't fail only because applications crash.
Sometimes the infrastructure is healthy, Kubernetes is healthy, and yet the platform struggles because capacity disappears before it can do any work.
At InfraThrone, we recreate production incidents like these—not to teach another Kubernetes feature, but to expose the hidden interactions between cloud infrastructure, autoscalers, and real-world workloads. Because in production, success isn't just about launching resources quickly—it's about ensuring they're still there when your applications need them most.
Discussion
to read and post comments.