Reliability starts with a transaction, not a server
At 18:07, the payment API did not fail. Its pods were healthy, Prometheus was scraping, and the dashboard was mostly green. Customers still could not complete transactions because the application was waiting 11 seconds for a dependency hosted…
Distributed System Reliability Engineering: Keeping SA Services Alive When the Network Does Not Co-operate
At 18:07, the payment API did not fail. Its pods were healthy, Prometheus was scraping, and the dashboard was mostly green. Customers still could not complete transactions because the application was waiting 11 seconds for a dependency hosted in eu-west. The alert fired on request errors, but not on the growing queue, the stretched latency, or the fact that the retry storm was consuming the last available connections.
That is the practical problem Distributed System Reliability Engineering must solve: not merely detecting broken containers, but proving whether a user can complete a meaningful transaction across unreliable networks, power interruptions, regions and third-party systems.
Reliability starts with a transaction, not a server
A distributed service is only as reliable as its slowest required hop. In a South African deployment, that may include a Kubernetes workload in Johannesburg, an on-premise database, a mobile user on a congested last-mile connection, and an identity provider reached through eu-west. A CPU dashboard cannot describe that customer journey.
Start by writing a service-level objective around an observable outcome. For example:
- 99.9% of checkout attempts return a valid response within 1.5 seconds over a rolling 30-day window.
- 99.95% of authorised payments are either confirmed or safely reconciled, rather than left ambiguous.
- Background settlement jobs recover within 30 minutes after a regional or power-related interruption.
The second objective matters because availability is not the same as correctness. A payment endpoint returning HTTP 200 while losing the confirmation event is operationally worse than a clean failure with a durable retry path.
Use a consistent correlation ID across HTTP requests, queue messages and database operations. Instrument the critical path with OpenTelemetry where practical, while keeping Prometheus-compatible metrics for established infrastructure and application exporters. Grafana’s 2025 observability survey reported that 67% of respondents used Prometheus in production and 41% used OpenTelemetry in production, showing why teams increasingly need both rather than treating them as competing standards.Grafana
Distributed System Reliability Engineering needs four telemetry views
Metrics tell you that a symptom is spreading. Logs explain an individual decision. Traces show where time disappeared. Profiles or runtime signals can explain why a process consumed its budget. The useful design is not four disconnected tools, but one navigable evidence chain.
- Prometheus and Mimir: Store request rate, error rate, latency histograms, queue depth, saturation, dependency outcomes and SLO burn rates. Mimir is useful when several clusters or regions need horizontally scalable, long-term Prometheus storage.
- Loki: Keep structured logs with labels such as service, environment, region and severity. Do not label every request ID; that creates expensive high cardinality. Put correlation IDs in the log body and extract them during investigation.
- Tempo: Retain traces that connect the public API to the database, message broker and remote dependency. Tail-based sampling is especially valuable for retaining slow, failed and unusual transactions without storing every successful request.
- Grafana: Provide a shared investigation surface. A panel showing a latency burn rate should link to the relevant traces and logs, not leave the on-call engineer searching three systems by hand.
Keep metric labels bounded. region="jhb", region="cpt" and dependency="identity" are useful. A label containing a full URL, user ID or exception message is a cardinality incident waiting to happen.
sum by (service, region) (
rate(http_request_duration_seconds_count{
environment="production",
route="/checkout",
status=~"2..|5.."
}[5m])
)
The query above is only a traffic view. Pair it with a histogram-based latency query and an error-budget calculation. A dashboard that displays averages hides the long tail that users experience when a cross-region link degrades.
Design alerts for failure domains, including load-shedding
Load-shedding schedules create an unusual alerting trap. If a site loses power, alerts generated by that site may arrive late or not at all. Alerting must therefore have an external perspective: probe the public endpoint from another region, monitor the state of the power-aware infrastructure, and alert when telemetry itself goes silent.
Separate symptoms from causes:
- Alert on a rising SLO burn rate for customer impact.
- Alert on queue age before workers are fully exhausted.
- Alert on scrape failures and remote-write backlog so missing data is not mistaken for health.
- Alert on dependency latency and timeout ratios, not only dependency HTTP errors.
- Use inhibition carefully during a known maintenance or power event; suppress duplicate noise, not the primary customer-impact alert.
A short interruption should not wake someone for every pod restart if workloads recover within their availability budget. It should wake someone when retries, queue age or lost redundancy indicate that recovery is no longer automatic.
Example: a burn-rate alert
groups:
- name: checkout-slo
rules:
- alert: CheckoutFastBurn
expr: |
(
1 - (
sum(rate(http_request_duration_seconds_bucket{
service="checkout",
route="/checkout",
le="1.5",
status=~"2.."
}[5m]))
/
sum(rate(http_request_duration_seconds_count{
service="checkout",
route="/checkout"
}[5m]))
)
) > 0.05
for: 10m
labels:
severity: page
annotations:
summary: Checkout latency or success SLO is burning rapidly
In production, validate the denominator and bucket semantics carefully. A copied alert with the wrong status labels can page a team for healthy traffic or silently ignore failures.
Trace the eu-west dependency before adding another retry
Retries are not resilience by themselves. When the network between Johannesburg and eu-west becomes slow, every retry increases concurrency, extends request lifetimes and consumes connection pools. The system may then fail locally even though the original problem is remote.
Use traces to measure each dependency’s contribution to the end-to-end budget. Set explicit deadlines, cap retries, add jitter, and use circuit breakers for dependencies that cannot meet the deadline. Prefer asynchronous workflows for operations that do not need an immediate answer.
- Give the customer-facing request a total deadline.
- Reserve part of that budget for response handling and persistence.
- Pass the remaining deadline downstream rather than allowing each service to invent one.
- Return a useful degraded response where business rules permit it.
- Make messages idempotent so replay after a network partition is safe.
Tempo traces make this visible when spans carry the same trace context across service boundaries. Loki then answers the next question: did the remote call time out, fail DNS resolution, hit a circuit breaker, or return a syntactically valid but unusable response?
Control telemetry cost and sovereignty deliberately
ZAR cloud-cost pressure changes the reliability calculation. Retaining every debug log and full trace indefinitely is not an engineering virtue if the observability bill forces teams to disable telemetry during an incident. Use retention tiers, structured logging, metric relabelling and targeted sampling.
Keep sensitive data out of telemetry by default. POPIA considerations apply to log fields, trace attributes, support exports and dashboard access, not only to the primary database. Mask tokens, payment details and identity attributes at the instrumentation boundary. Decide which telemetry must remain in South Africa and which aggregate data may cross borders before a vendor or incident workflow makes that decision for you.
A practical retention model might keep high-value SLO metrics for months, ordinary logs for days, and failed or slow traces longer than routine successful traces. The exact periods depend on contractual, regulatory and forensic requirements; the important point is to document them and test deletion.
Prove recovery with failure drills
Reliability claims are hypotheses until tested. Run controlled experiments against the failure modes that are normal for the estate:
- Drop or delay traffic to the eu-west dependency and verify deadline propagation, retry limits and customer messaging.
- Stop telemetry export from one cluster and confirm that monitoring reports the blind spot.
- Drain a region while queues are active; measure duplicate processing and reconciliation time.
- Simulate a load-shedding event at a non-production site and check whether external probes, paging and failover behave as intended.
- Restore from backup and compare recovery time with the stated objective, including DNS, secrets, certificates and dashboards.
After each drill, record the evidence: the first detectable signal, the first actionable alert, the decision that reduced impact, and the telemetry that was missing. That record is more valuable than a dashboard screenshot because it shows whether the system remains understandable when several components fail together.
Distributed System Reliability Engineering is therefore a discipline of budgets and evidence. Define reliability around completed user outcomes, connect metrics to logs and traces, treat regional latency and power interruptions as design inputs, and make telemetry affordable enough to survive the incident it is meant to explain.