Real-Time Infrastructure Health Analytics starts with user impact
At 02:17, a payment API in Johannesburg started returning intermittent 502s. CPU stayed below 40%, memory looked normal, and the Kubernetes nodes were “healthy”. The first alert pointed at the application. The second pointed at the database. By…
Real-Time Infrastructure Health Analytics: Finding the Failure Behind the Red Dashboard
At 02:17, a payment API in Johannesburg started returning intermittent 502s. CPU stayed below 40%, memory looked normal, and the Kubernetes nodes were “healthy”. The first alert pointed at the application. The second pointed at the database. By the time someone correlated the traces with packet loss on the office-to-cloud link, the incident had already become a customer-facing outage.
This is the problem Real-Time Infrastructure Health Analytics is meant to solve: not displaying more graphs, but connecting weak signals quickly enough to explain what users are experiencing. In a hybrid South African estate, that means correlating infrastructure metrics with logs, traces, network conditions, power events, regional latency and cloud spend.
Real-Time Infrastructure Health Analytics starts with user impact
A node can be healthy while a service is effectively unavailable. A dashboard can show green CPU, memory and disk panels while requests time out because the path to an upstream dependency is degraded. Infrastructure health therefore needs to be measured at several layers:
- Service health: request rate, error rate, latency and saturation.
- Platform health: node pressure, container restarts, control-plane latency and storage performance.
- Dependency health: database connections, queue depth, DNS resolution and third-party API response times.
- Network health: packet loss, route changes, transit latency and cross-region performance.
- Operational context: deployments, maintenance, load-shedding schedules, backup windows and scaling activity.
The practical starting point is the familiar service-level trio: errors, latency and traffic. Prometheus records these signals; Grafana turns them into dashboards and alerts; Loki adds searchable logs; Tempo connects a slow or failed request to the services it traversed. For longer retention and multi-cluster scale, Mimir provides a Prometheus-compatible metrics backend.
Recent observability practice is moving towards combining Prometheus with OpenTelemetry rather than treating them as competing choices. Grafana’s 2025 observability survey reported that 71% of organisations use both in some capacity, while 67% use Prometheus in production. That direction matters because metrics identify the symptom, traces show the path and logs often contain the reason.
Build a signal path that survives a hybrid estate
A workable architecture does not require every workload to move into one platform. It requires consistent labels and reliable transport. A typical flow looks like this:
- Prometheus scrapes Kubernetes, virtual machines, databases and network exporters.
- OpenTelemetry agents or Grafana Alloy collect application telemetry and forward metrics, logs and traces.
- Mimir stores high-volume metrics across clusters and regions.
- Loki stores structured logs with carefully controlled labels.
- Tempo stores traces and exposes trace identifiers for correlation.
- Grafana presents service, infrastructure and incident views from the same operating context.
For a Johannesburg workload calling a service in eu-west, label every signal with fields such as environment, region, cluster, service and team. Without that metadata, a query for “latency” quickly becomes a hunt through unrelated workloads.
Do not put highly variable values such as request IDs, user IDs or full URLs into Loki labels. That creates excessive stream cardinality and increases cost. Keep those fields inside the log body, then extract them during investigation. The same discipline applies to Prometheus labels: a metric with an unbounded customer or session label can become an expensive incident of its own.
sum by (service, region) (
rate(http_server_request_duration_seconds_count{
environment="production",
status_code=~"5.."
}[5m])
)
/
sum by (service, region) (
rate(http_server_request_duration_seconds_count{
environment="production"
}[5m])
)
> 0.02
This query identifies services whose five-minute error ratio exceeds two percent, split by region. The threshold is only useful if it reflects the service’s SLO. A batch worker, payment endpoint and public search API should not necessarily share the same alert policy.
Make load-shedding and connectivity visible, not anecdotal
Power events are operational data. During load-shedding, a site may switch to batteries or generators, cooling capacity may change, and network equipment may restart even when the primary application remains available. Treating these events as unlabelled background noise leads to false escalations and missed capacity risks.
South Africa’s power situation improved materially in 2024 and the first half of 2025, but the system remained vulnerable to renewed interruptions. The CSIR reported 749 GWh of load-shedding in the first half of 2025, down 82% from the comparable 2024 figure, while interruptions returned during parts of early 2025. The operational lesson is not to remove power-related monitoring when conditions improve.
Export schedule information into a small time series or event stream. Then use it in dashboards and alert routing. During a planned Stage 2 window, a brief UPS transfer should not page the entire on-call team. A battery discharge outside the expected window should.
Last-mile connectivity deserves the same treatment. Probe critical endpoints from more than one vantage point: a cloud region, the data centre, a branch network and, where relevant, an ISP-facing probe. A successful synthetic check from AWS Cape Town does not prove that customers using a mobile network in Limpopo can reach the service.
Use traces and logs to explain the red metric
Metrics are excellent at detecting change. They are weaker at explaining causality. That is where exemplars, logs and traces earn their place.
Suppose the latency alert fires for a checkout service. A Grafana panel should make it possible to move from the latency histogram to an exemplar trace, then from the trace to the relevant Loki records. The trace may show that the checkout request spent 900 milliseconds waiting on an inventory call, while the application log identifies connection-pool exhaustion. The infrastructure view can then reveal that the inventory database is healthy but the private link is dropping packets.
Structure logs for investigation rather than for human reading alone:
{
"timestamp": "2026-10-09T02:17:43.120Z",
"level": "error",
"service": "checkout-api",
"environment": "production",
"region": "johannesburg",
"trace_id": "7f3a1d9c",
"dependency": "inventory-api",
"error": "upstream timeout",
"duration_ms": 1842
}
Trace sampling should reflect business value. Head-based sampling may discard the one failed payment that matters. Tail-based sampling can retain slow, failed or unusually expensive traces, although it requires more processing and storage. Keep the policy explicit, and review it when traffic patterns change.
Alert on degradation, not every uncomfortable number
An alert is useful when it changes an operator’s next action. “CPU above 80%” often does not. “Checkout success rate below the SLO for ten minutes, with database connection saturation above 90%” does.
Use separate warning and page paths. A warning can create a ticket for rising disk utilisation. A page should be reserved for customer impact or an imminent loss of service. Add runbook links, ownership and the affected region to the alert annotations.
For cross-region services, compare latency against a local baseline rather than applying one global threshold. A Johannesburg-to-eu-west request will naturally have a different round-trip time from a Johannesburg-to-Cape-Town request. Alert on unexpected change, not on geography.
Grafana helps here by combining alert rules, annotations and dashboards into an incident view rather than forcing an engineer to open separate tools. It is particularly useful when deployment markers, power events and trace exemplars are displayed on the same time axis.
Control ZAR cost and protect sensitive telemetry
Real-time analytics can become a cost problem if every log line, trace span and high-resolution metric is retained indefinitely. Measure telemetry spend like any other platform workload. Mimir retention tiers, recording rules, metric relabelling, log sampling and trace policies can reduce volume without making investigations impossible.
Keep high-cardinality, high-value data for a shorter period and retain aggregated service indicators for longer. A payment trace may need detailed retention for days; an hourly error-rate series may be useful for months. Align retention with incident response, compliance and capacity-planning requirements.
POPIA also changes the design. Do not send customer names, identity numbers, payment details or session contents into a central observability system by default. Redact at the source or collection layer, restrict access by team and region, and document where telemetry is stored. Data sovereignty is not solved by adding a “South Africa” label after sensitive data has already crossed a border.
The strongest implementation is deliberately boring: consistent labels, sensible cardinality, clear SLOs, tested alert routes and a short path from symptom to evidence. Real-Time Infrastructure Health Analytics is successful when the on-call engineer can answer three questions quickly: what is failing, who is affected and which change or dependency explains it.