The Merger That Looked Complete... Until Two Clouds Refused to Trust Each Other
The Merger That Looked Complete... Until Two Clouds Refused to Trust Each Other
It was the first Monday after the acquisition.
Executives announced that both platforms were officially connected.
AWS and GCP.
One company.
One product.
One future.
The dashboards looked healthy. VPN tunnels were up. DNS resolution worked. Kubernetes clusters could finally see each other.
Everyone expected the integration to be uneventful.
Then the first customer request crossed the Atlantic.
It failed.
Not because the network was down.
Not because the service crashed.
Because one service simply said:
"I don't know who you are."
Inside the war room, engineers verified firewalls, Transit Gateway routes, cloud VPNs, and Kubernetes Services. Everything looked perfectly connected.
Packets reached their destination.
The destination just refused to trust them.
Hours later, someone noticed something subtle.
The workloads in AWS carried one SPIFFE identity.
The workloads in GCP carried another.
Each Istio service mesh trusted only its own Certificate Authority.
Both applications were speaking mutual TLS.
Neither believed the other's identity.
A bridge of trust had to be built manually between two different trust domains before a single API call could succeed.
Finally…
Traffic started flowing.
Relief lasted exactly three months.
The quarterly root CA rotation arrived.
Nearly 200 Kubernetes Secret resources across two clouds needed updating before certificates expired. One forgotten Secret caused workloads to silently lose trust again. Applications stayed healthy. Pods stayed Ready.
Only service-to-service communication began disappearing.
One request at a time.
Production didn't fail loudly.
It slowly became strangers talking across a perfectly healthy network.
Cross-cloud networking isn't just about VPNs and routing tables.
It's about federating identities, synchronizing certificate authorities, combining AWS IRSA with GCP Workload Identity, managing Istio multi-primary service meshes, handling SPIFFE identities, and ensuring security policies remain consistent across platforms.
These are the problems documentation explains in isolation.
Production introduces them all at once.
At InfraThrone, we recreate incidents like this as hands-on labs where every layer matters—from networking and identity federation to certificate rotation and service mesh trust. Because modern outages rarely happen because a single component fails. They happen when dozens of perfectly healthy systems quietly stop trusting each other.
Discussion
to read and post comments.