Start with a telemetry budget, not a bigger cluster
At 18:07 on a Tuesday evening, a load-shedding event moved a customer-facing workload from a Johannesburg data centre to a European region. The application recovered, but the telemetry did not: Prometheus remote writes queued up, Loki received a…