Start with the failure path, not the dashboard

At 02:17 on a Tuesday morning, a payment service in Johannesburg started timing out. CPU was ordinary, error rates were climbing slowly, and the application logs showed little more than retries. The useful clue was elsewhere: traces showed…

Start with the failure path, not the dashboard

High-Scale Performance Analytics Ecosystems: Designing Observability That Survives Real Traffic

At 02:17 on a Tuesday morning, a payment service in Johannesburg started timing out. CPU was ordinary, error rates were climbing slowly, and the application logs showed little more than retries. The useful clue was elsewhere: traces showed requests taking an unfamiliar path through a Cape Town dependency, while a Prometheus query had become so slow that the on-call engineer stopped trusting the dashboard.

That is the operational test for High-Scale Performance Analytics Ecosystems. The challenge is not collecting more telemetry. It is keeping metrics, logs and traces queryable, correlated and affordable when traffic, cardinality and failure conditions all increase at once.

Start with the failure path, not the dashboard

A scalable observability design begins with the questions an incident requires answering:

  • Is the customer-facing request slow, or is a downstream service slow?
  • Did latency increase in Johannesburg, Cape Town, Nairobi or across all regions?
  • Is the problem application capacity, a database, network reachability or an infrastructure event?
  • Which release, tenant, route or dependency changed?

These questions map naturally to the three core signal types. Prometheus-compatible metrics expose saturation and service-level indicators. Loki preserves the event detail needed to understand what a request did. Tempo connects the request across service boundaries and makes a latency budget visible rather than theoretical.

Grafana provides the investigation surface, but the important design decision happens underneath: every signal must carry a small, deliberate set of dimensions that can be joined. Useful dimensions include service, environment, region, cluster and route. A random request ID does not belong in a metric label. It belongs in logs and traces.

High-Scale Performance Analytics Ecosystems need separated data paths

A single Prometheus server is an excellent starting point and a poor long-term architecture for a large hybrid estate. It couples scraping, rule evaluation, storage and querying. A busy incident can therefore affect both the system being observed and the system used to observe it.

A more resilient pattern uses Prometheus or Grafana Alloy at the edge, with regional collection and remote writing into Grafana Mimir. Mimir provides horizontally scalable, Prometheus-compatible storage, while object storage provides durable retention. Query capacity can then be expanded independently from ingestion capacity.

The same principle applies to logs and traces. Loki should receive structured logs with low-cardinality index labels, while the log body carries fields such as request ID, customer segment and upstream status. Tempo should receive sampled traces, with tail sampling used to retain slow, failed or unusual requests at a higher rate than ordinary traffic.

In practice, the topology might look like this:

  • South African Kubernetes clusters scrape locally and buffer during connectivity interruptions.
  • A regional gateway applies authentication, tenant limits and metric relabelling.
  • Mimir stores metrics centrally, with retention and replication selected by business importance.
  • Loki stores operational logs locally or regionally where POPIA and contractual requirements demand it.
  • Tempo receives traces from South Africa and selected European dependencies, including eu-west latency.
  • Grafana presents a shared view while permissions restrict sensitive tenants and data sets.

This is not merely an availability pattern. It is a cost-control pattern. Recent observability surveys have continued to show strong adoption of Prometheus and increasing investment in OpenTelemetry, while cost remains a major selection criterion. The practical response is to scale the useful signals, not to retain every event forever.

Keep Prometheus labels boring and queries useful

Most high-scale metric failures begin with label design. The following query is operationally useful because it aggregates by a small, intentional set of dimensions:

histogram_quantile(
  0.99,
  sum by (le, service, region) (
    rate(http_request_duration_seconds_bucket{
      environment="production",
      route=~"/payments|/checkout"
    }[5m])
  )
)

By contrast, labels such as user_id, order_id, full URL, exception text or trace ID create a new time series for every value. That increases memory usage, remote-write volume and query cost without improving an alert about regional service latency.

Use recording rules for queries that appear on many dashboards or alerts. A rule can calculate the service-level indicator once, rather than forcing every dashboard panel to repeat an expensive range query across months of data.

groups:
  - name: payments-sli
    interval: 30s
    rules:
      - record: service_region:http_requests_error_ratio5m
        expr: |
          sum by (service, region) (
            rate(http_requests_total{service="payments",status=~"5.."}[5m])
          )
          /
          sum by (service, region) (
            rate(http_requests_total{service="payments"}[5m])
          )

Alert on symptoms that matter to users: error-budget burn, sustained latency, unavailable capacity and failed dependencies. Do not page because a single node briefly lost power during load-shedding. Route that event to an infrastructure channel unless it materially affects a service objective.

Design for load-shedding and last-mile uncertainty

South African infrastructure introduces failure modes that a generic global template often misses. A site may lose utility power while generators and batteries keep workloads running. A branch may remain reachable from inside the network while its last-mile connection cannot reach cloud endpoints. A regional link to eu-west may degrade without being completely unavailable.

Alerting should distinguish those conditions. A node-down alert can be paired with quorum, workload availability and power telemetry. For remote sites, use longer evaluation windows and an explicit connectivity state. Avoid converting every scrape failure into a customer-facing incident.

For cross-region services, measure both application latency and network latency. A trace span from Johannesburg to eu-west may explain why an API is slow, but a synthetic probe can show whether the problem affects all traffic or only one service path. Include region in dashboards and alerts, but avoid a separate alert for every low-volume combination.

Buffering matters as well. Collectors should queue telemetry briefly when a remote destination is unreachable, with limits that protect application resources. When the queue fills, prioritise errors, exemplars and traces selected by policy over routine debug logs.

Use traces and logs to explain the metric

Metrics identify the shape of an incident. Traces explain the path. Logs provide the local evidence.

Exemplars are particularly effective for latency investigations: a latency time series can link directly to a representative trace. From there, the engineer can inspect the slow database span, retry loop or remote dependency. Loki can then be queried using the trace ID or service metadata without indexing every field in every log line.

Structured logging should be deliberate:

  • Index stable fields such as service, environment, cluster and severity.
  • Keep request IDs and trace IDs in the log body for correlation.
  • Redact identity, payment and health information before export.
  • Set retention by operational value, not by habit.
  • Sample successful, repetitive access logs while retaining failures and audit records according to policy.

POPIA and data-sovereignty requirements make placement part of observability architecture. Telemetry can contain email addresses, IP addresses, account identifiers and request payload fragments. Keep sensitive fields out of labels, apply redaction at collection time, and document which data leaves South Africa. A dashboard that is technically available but legally overexposed is not a successful platform.

Make ZAR cost visible before scale makes it painful

Cloud bills rarely increase because one dashboard is expensive. They increase through many small decisions: verbose logs, unnecessary label dimensions, long trace retention, duplicate scraping and unbounded remote writes.

Build cost views around telemetry volume and ownership. Track samples received, bytes written, active series, log ingestion, query duration and object-storage growth by team or environment. Add those signals to the same operational review as CPU and availability.

A practical retention policy might keep high-resolution metrics for a short period, downsampled metrics for longer trend analysis, recent logs for investigation, and traces according to error and latency value. The exact periods depend on compliance and incident history, but the principle is consistent: retain what changes a decision.

Recent platform work has also made profiling more relevant. Continuous profiling can reveal CPU and memory hot spots that metrics alone cannot explain, particularly when a service appears healthy but spends excessive time in a library or garbage collector. It should be introduced carefully,