Start with a telemetry budget, not a bigger cluster

At 18:07 on a Tuesday evening, a load-shedding event moved a customer-facing workload from a Johannesburg data centre to a European region. The application recovered, but the telemetry did not: Prometheus remote writes queued up, Loki received a…

Start with a telemetry budget, not a bigger cluster

Enterprise Telemetry Optimisation Strategies for a Cost-Conscious Hybrid Estate

At 18:07 on a Tuesday evening, a load-shedding event moved a customer-facing workload from a Johannesburg data centre to a European region. The application recovered, but the telemetry did not: Prometheus remote writes queued up, Loki received a flood of duplicate reconnect errors, and trace volume tripled as retries multiplied across the WAN. The incident was visible everywhere except in the one dashboard the on-call engineer needed.

This is why Enterprise Telemetry Optimisation Strategies must address more than retention and storage. In a hybrid South African estate, telemetry has to remain useful when power is unstable, links to eu-west are slow, cloud costs are billed in foreign currency, and personal data cannot casually cross borders.

Start with a telemetry budget, not a bigger cluster

Telemetry is a production dependency. Every metric series, log line and span consumes CPU, network capacity, storage and engineering attention. Treating those resources as unlimited produces an observability bill that grows faster than the platform it describes.

Set separate budgets for each signal:

  • Metrics: active series, samples per second and remote-write bandwidth.
  • Logs: gigabytes per day, indexing labels and retention periods.
  • Traces: spans per request, sampling rate and object-storage growth.
  • Alerts: notification rate, evaluation cost and the number of actionable pages.

Prometheus should retain the metrics required for fast operational decisions locally, while long-term history can be sent to Mimir. Mimir provides scalable long-term storage for Prometheus data, making it appropriate for capacity analysis and quarterly reporting rather than every short-lived debugging query.[12]

Measure telemetry cost per service or team. A simple allocation model can combine active series, log volume and trace ingestion. The exact weighting matters less than exposing ownership. When a new deployment adds 40,000 series or a debug logger emits request bodies, the responsible team should see the impact in the same way it sees compute consumption.

Enterprise Telemetry Optimisation Strategies for Prometheus and Mimir

High cardinality is usually a naming and labelling problem, not a Prometheus problem. Labels such as pod, container_id, request path and customer reference can create a separate time series for every changing value. Keep dimensions that support a decision: service, namespace, cluster, region, status class and a bounded route name.

scrape_configs:
  - job_name: application
    kubernetes_sd_configs:
      - role: pod
    metric_relabel_configs:
      - source_labels: [__name__]
        regex: 'http_request_duration_seconds_bucket'
        action: keep
      - source_labels: [customer_id]
        action: labeldrop
      - source_labels: [request_id]
        action: labeldrop

Do not blindly drop every histogram. Latency distributions are often essential for SLOs, but bucket boundaries should reflect the service’s user experience. A payment API may need buckets around 100 ms and 1 second; a batch job does not need the same granularity.

Use recording rules for expensive, repeatedly queried calculations. A dashboard should not recalculate a complex rate over billions of samples every time an executive opens it.

groups:
  - name: checkout-sli
    interval: 30s
    rules:
      - record: service:http_requests:rate5m
        expr: sum by (service, status_class) (
          rate(http_requests_total[5m])
        )
      - record: service:http_error_ratio:rate5m
        expr: |
          sum by (service) (rate(http_requests_total{status_class="5xx"}[5m]))
          /
          sum by (service) (rate(http_requests_total[5m]))

Federation, remote write and long-term storage should be deliberate choices. Keep an operational Prometheus close to the workloads it scrapes, then send selected metrics to Mimir. During a cross-region outage, local dashboards and alert evaluation should continue even if the path to cloud storage is impaired.

Make logs searchable without making every field a Loki label

Loki’s cost advantage depends on restraint. Its index is built around labels, so putting high-cardinality fields into labels recreates the same problem seen in metrics. Use stable labels such as cluster, namespace, app and environment. Keep request IDs, usernames and transaction references inside structured log content.

Structured JSON makes this practical:

{
  "ts": "2026-10-07T18:08:12.442Z",
  "level": "WARN",
  "service": "checkout",
  "event": "payment_retry",
  "region": "johannesburg",
  "attempt": 2,
  "request_id": "7f3c..."
}

In LogQL, the stable labels narrow the search before content filters do the expensive work:

{namespace="payments", app="checkout"} 
  |= "payment_retry"
  | json
  | attempt > 1

Sample noisy logs at the collector, not after they have crossed an expensive international link. Health checks, successful polling messages and repeated connection failures often have low diagnostic value. Preserve counts and exemplars where possible, and retain the complete event only for errors or selected percentage samples.

Grafana is useful here because a single investigation can move from a Loki log line to its trace and then to the Prometheus or Mimir metric that shows the wider impact. The value is correlation, not another screen full of charts.

Trace the failure path, not every request

Tempo is designed for high-volume distributed tracing and can use object storage as its primary persistence layer.[1] That makes it a sensible home for traces, but it does not make unlimited tracing free. A retry storm can generate more spans precisely when storage and network capacity are under pressure.

Use head sampling for predictable baseline control and tail sampling for errors, slow requests and unusual latency. Retain complete traces for failed payments, authentication errors and requests exceeding the service’s SLO threshold. Sample healthy, fast traffic more aggressively.

Trace propagation must survive asynchronous boundaries. HTTP headers alone will not connect a Kafka consumer, queue worker or scheduled job to the originating request. Define propagation conventions and attach a trace ID to operational logs. Tempo can then provide the causal path while Loki supplies the detailed event context.[1]

Span-derived metrics are particularly valuable when teams need RED signals without retaining every trace. Tempo’s metrics generator can derive rate, error and duration metrics and remote-write them to Prometheus or Mimir.[8] This creates a useful feedback loop: traces explain individual failures, while metrics support cheap, persistent alerting.

Design for South African failure modes

Load-shedding-aware alerting requires separating infrastructure loss from application failure. A node disappearing during a planned power event should not page every service owner as if each service independently failed. Use maintenance windows, inhibition rules and a site-health signal. Continue paging when an active region loses redundancy or when recovery exceeds the declared window.

For workloads failing over from Johannesburg or Cape Town to eu-west, measure the full path: DNS time, TLS negotiation, application latency and queue age. A synthetic check from the same network conditions as customers is more useful than a healthy internal service metric.

Hybrid estates also need local buffering. Collectors should queue telemetry briefly during WAN interruptions and apply back-pressure rather than allowing memory growth to take down the application. Prioritise alerts, error logs and exemplars over verbose debug streams.

Data sovereignty and POPIA add another design constraint. Classify telemetry before exporting it. Remove names, phone numbers, identity numbers, access tokens and request bodies at the source. Keep sensitive logs in the approved jurisdiction, and send only aggregated metrics or redacted events to a central platform. A trace attribute is still data; calling it “telemetry” does not make it exempt from governance.

Turn optimisation into an operating practice

Telemetry controls decay unless they are tested like application controls. Add a review to the service onboarding process:

  1. Define the service’s golden signals and SLOs.
  2. List required metric labels and reject unbounded dimensions.
  3. Specify log fields, redaction rules and Loki labels.
  4. Set baseline and exceptional trace-sampling policies.
  5. Declare local and long-term retention.
  6. Assign a team responsible for telemetry volume and alert quality.

Review the top ten metric families by series count and the top ten log streams by daily volume each month. Look for sudden changes after releases, especially new Kubernetes labels and debug statements. Grafana dashboards can expose these trends alongside service health, while Mimir, Loki and Tempo retain the evidence needed to investigate them.

Recent observability practice is moving towards vendor-neutral instrumentation and cost-aware collection. Grafana’s 2025 survey reported that 71% of organisations used Prometheus and OpenTelemetry in some capacity, while cost remained a major selection concern.[7] OpenTelemetry is useful as a collection and propagation standard, but it does not remove the need for sampling, cardinality controls or data classification.

The practical target is not maximum telemetry. It is maximum useful evidence per rand, byte and CPU cycle. When the next regional failover occurs, the winning platform will be the one that still answers three questions quickly: what broke, who is affected, and which dependency caused it.

Grafana can provide the investigation layer across these signals, provided the underlying telemetry has been designed with clear budgets, ownership and failure conditions.