Start with the failure path, not the cloud provider
At 02:17, a payment API started timing out for customers in Gauteng. The application dashboard showed healthy pods in AWS Cape Town. The database dashboard showed normal CPU. The incident channel blamed the carrier, then the DNS provider,…