The Registry Was Clean... Until the Finance Team Looked at the Bill
# The Registry Was Clean... Until the Finance Team Looked at the Bill
Quarter-end infrastructure review.
Everything looked healthy.
No production incidents. No deployment failures. Developers were shipping containers every day without a second thought. The self-hosted Harbor registry had become just another reliable piece of infrastructure.
Then someone opened the monthly AWS invoice.
S3 Storage: $5,000.
Silence.
"How do container images cost this much?"
The registry UI showed only a few thousand active images. Old releases had been cleaned up regularly. Teams proudly followed image retention policies.
So where was the storage coming from?
The investigation led into the object storage bucket backing the registry.
Layer after layer.
Blob after blob.
Millions of them.
Deleted images were gone from the UI—but their underlying layers still occupied nearly 50 TB of storage.
The cleanup had never really happened.
Months earlier, someone had suggested running the registry's garbage collector.
Another engineer had quickly replied, "Not during business hours. It blocks the registry."
The discussion ended there.
Next sprint.
Next quarter.
Next maintenance window.
It kept getting postponed because nobody wanted to risk slowing down image pulls during deployments.
Meanwhile, every CI pipeline kept pushing new image versions.
Every feature branch.
Every hotfix.
Every rollback.
Every forgotten experiment.
The zombies quietly multiplied.
Eventually someone finally asked the uncomfortable question:
"Do we actually know how our registry deletes images?"
That's when the real lesson surfaced.
Deleting an image removes its manifest, not necessarily the blobs (layers) stored underneath. As long as orphaned blobs remain in object storage, S3 continues billing for every byte. Only the registry's garbage collection (GC) process walks the OCI distribution graph, identifies unreferenced layers, and permanently removes them.
Running GC isn't just clicking a button.
Depending on the registry, it may require offline mode, temporarily blocking writes, while newer implementations support online GC with fewer disruptions. Knowing which mode your platform supports—and planning maintenance accordingly—is what separates routine operations from costly surprises.
The registry never failed.
Deployments never stopped.
Monitoring never complained.
Yet thousands of dollars quietly disappeared every month because of one maintenance task everyone was afraid to schedule.
Production engineering isn't only about fixing outages.
Sometimes the biggest incident is the one nobody notices.
At InfraThrone, we recreate these overlooked production scenarios—the ones hidden behind successful deployments and green dashboards. Because real DevOps expertise isn't just learning Kubernetes or Docker; it's understanding the operational decisions, hidden dependencies, and infrastructure behaviors that quietly shape production every single day.
Discussion
to read and post comments.