Predictive Failure Detection for Hybrid Clouds starts with failure paths
At 02:17, the cloud-facing checkout service was still returning 200s. The useful warning was elsewhere: packet loss on the Johannesburg-to-Dublin link had climbed slowly, one on-premise queue was growing, and the UPS runtime had fallen below the time…
Predictive Failure Detection for Hybrid Clouds: Catching the Fault Before the Pager
At 02:17, the cloud-facing checkout service was still returning 200s. The useful warning was elsewhere: packet loss on the Johannesburg-to-Dublin link had climbed slowly, one on-premise queue was growing, and the UPS runtime had fallen below the time required to shut down its storage controllers cleanly. Nothing had breached the usual alert thresholds yet. Forty minutes later, the failover run began under pressure.
This is the practical case for Predictive Failure Detection for Hybrid Clouds: not predicting every failure with machine learning, but combining weak signals early enough to make a controlled decision. In a hybrid estate, the failure may start in a rack, a carrier network, a cloud region, a certificate store or a cost-control policy. The telemetry must follow the dependency chain, not the organisational boundary.
Predictive Failure Detection for Hybrid Clouds starts with failure paths
A useful design begins with a small number of failure paths rather than a large dashboard catalogue. For each service, document what it depends on and what “degraded” looks like before the customer sees an error.
- On-premise to cloud: VPN or direct-connect packet loss, BGP changes, rising round-trip time and retransmits.
- Power and facilities: UPS battery health, generator state, cooling alarms and the remaining runtime during load-shedding events.
- Application capacity: queue age, database connection saturation, garbage-collection pauses and the rate at which work is entering versus leaving a system.
- Cloud dependencies: API throttling, quota consumption, spot-instance interruption notices and regional error-rate differences.
- Data movement: replication lag, object-store upload failures and the age of the newest successful backup.
The important question is not “What is the CPU percentage?” It is “Which combination of signals says this service will miss its recovery objective within the next hour?” A rising queue with flat CPU can indicate a downstream dependency problem. A falling UPS runtime with normal server metrics can be the most urgent signal in the room.
Build a time-aligned signal layer with the LGTM stack
Prometheus is well suited to short-range infrastructure and service metrics. For a hybrid estate, remote-write selected series into Mimir so that Johannesburg, Cape Town and cloud workloads can be queried over a common retention period without treating one Prometheus server as the historical system of record. Keep labels disciplined: site, region, environment, service and dependency are generally more useful than labels containing request IDs or pod names.
Loki adds operational context, provided logs are structured and sensitive fields are removed before centralisation. Tempo supplies the missing path through distributed services: a slow checkout trace can show that the apparent application problem is actually an on-premise inventory call traversing a congested link. Grafana ties those signals together in an investigation view and makes alert annotations visible alongside deployments, maintenance windows and power schedules.
That architecture also supports a sensible sovereignty boundary. POPIA does not create a blanket ban on international processing, but cross-border transfers require appropriate safeguards. For government workloads, South Africa’s 2024 National Policy on Data and Cloud states that certain national-security and sovereignty-related data must be stored on infrastructure within the country. Keep raw personal-data-bearing logs and verbose traces in a controlled local tenant where required; send aggregated metrics, redacted events and carefully sampled traces to the wider analytical plane.
Measure the leading indicators, not just the symptoms
Predictive alerts are usually derived from rates, slopes and relationships. For example, queue age is more informative when paired with throughput and error rate. The following PromQL expression identifies a queue whose age is increasing while consumers are failing:
(
deriv(work_queue_oldest_age_seconds{service="orders"}[15m]) > 0.02
)
and
(
rate(order_consumer_failures_total{service="orders"}[5m]) > 0
)
for 10mThe threshold is deliberately modest and must be calibrated against the workload. A ten-minute warning is valuable only if the team can drain the queue, add capacity or route traffic elsewhere within ten minutes.
For links to Europe, track both absolute latency and deviation from a rolling baseline. A Johannesburg workload talking to eu-west-1 may remain technically available while timeouts accumulate at the application layer. Alerting on a fixed latency number alone creates noise during normal peak periods; alerting on a sustained deviation from the service’s own baseline is more useful.
Turn forecasts into decisions, not decorative dashboards
There are three practical levels of prediction.
- Threshold with persistence: alert when a leading indicator remains abnormal for a defined interval. This is the most explainable option and should cover most production cases.
- Trend-based detection: estimate whether a resource will cross a safe limit soon. Disk exhaustion, certificate expiry, queue growth and replication lag are good candidates.
- Multi-signal correlation: combine independent evidence, such as packet loss, trace latency and retry volume, before escalating. This reduces the chance that a single noisy metric pages the team.
Prometheus recording rules can make these calculations cheap enough for continuous evaluation. For capacity forecasts, use a conservative window and include operational headroom:
- record: node_filesystem:free_bytes:predict_6h
expr: predict_linear(
node_filesystem_avail_bytes{mountpoint="/var/lib"}[6h],
6 * 3600
)
- alert: FilesystemLikelyFull
expr: node_filesystem:free_bytes:predict_6h < 20 * 1024 * 1024 * 1024
for: 15m
labels:
severity: warning
annotations:
summary: "Filesystem projected to breach 20 GiB within six hours"A forecast is not a fact. Compaction, a planned data export or a retention change can invalidate it. The alert should therefore include the observed slope, the forecast horizon and a link to the runbook, rather than claiming that failure is certain.
Make load-shedding and connectivity part of the model
Power events should not be treated as an unrelated facilities concern. Record the scheduled stage, UPS runtime, generator state and the availability of network equipment as metrics. During an outage, suppress alerts for hosts intentionally shut down, but keep alerts for the control plane, core routing, storage integrity and services that are expected to remain available.
A useful alert annotation can state: “UPS runtime is 28 minutes; graceful storage shutdown requires 20 minutes; cloud replication lag is 14 minutes.” That is an operational decision, not merely a red panel.
Last-mile instability deserves similar treatment. A branch office may lose connectivity without the application being unhealthy in the cloud. Use black-box probes from representative South African networks, not only probes launched inside the same cloud region. Loki can retain carrier and tunnel state changes, while Tempo shows whether retries are amplifying the outage. The objective is to distinguish a local access problem from a genuine service failure before teams start unnecessary failovers.
Control ZAR cost before observability becomes the failure
Hybrid telemetry can become expensive quickly, particularly when high-cardinality labels and unbounded traces are sent across regions. Mimir helps with long-term metrics storage, but it does not make poor metric design free. Drop unused series at collection time, sample traces according to service value, and retain full-fidelity data only for security or contractual requirements.
Attach cost metadata to telemetry pipelines: tenant, environment, retention class and estimated ingest volume. A sudden increase in trace volume should produce an operational warning before the monthly bill becomes the incident. In a ZAR-sensitive environment, a cheaper local buffer that survives a link outage may be preferable to streaming every log immediately to a foreign region.
Grafana can present reliability and spend together: projected capacity exhaustion beside ingestion volume, or replication lag beside cross-region transfer. That pairing discourages false economies, such as reducing retention while an unresolved data-loss risk remains.
Test the prediction during daylight hours
Predictive detection earns trust through rehearsal. Inject controlled packet loss, pause a consumer, fill a test filesystem, expire a non-production certificate and simulate reduced UPS runtime. Record three timings:
- When the first leading indicator moved.
- When the alert became actionable.
- When the team could complete the mitigation.
Then review false positives. An alert that fires during every month-end batch will be ignored, even if its mathematics is sound. Adjust the baseline, add maintenance context or change the action threshold. Keep the final notification short: affected service, expected time to breach, evidence, owner and next action.
The strongest implementation is often deliberately boring. Prometheus rules identify a worsening trend; Mimir preserves the history; Loki explains the surrounding events; Tempo confirms the dependency path; Grafana puts the evidence in front of the person who must choose between draining traffic, starting capacity, isolating a site or waiting. Prediction is useful only when it buys enough time to make that choice deliberately.
For teams standardising their platform, the Grafana ecosystem provides the visual and alerting layer, while the data-retention and sovereignty decisions remain yours. The engineering work is not choosing a fashionable algorithm. It is defining failure paths, collecting trustworthy signals and proving that an early warning changes the outcome.