Start with the failure path, not the dashboard
At 02:17 on a Tuesday morning, a payment service in Johannesburg started timing out. CPU was ordinary, error rates were climbing slowly, and the application logs showed little more than retries. The useful clue was elsewhere: traces showed…