Start with a causal graph, not a pile of alerts
At 18:07 on a Thursday, a payment API in Johannesburg started timing out. The first alert blamed CPU saturation. Three minutes later, alerts arrived for database connections, queue depth, checkout failures and elevated latency to an upstream service…