Start with a telemetry bill of materials

At 02:17 on a Tuesday, a Kubernetes deployment in Johannesburg began emitting unusually verbose debug logs. The application was healthy, but Loki ingestion climbed rapidly, the cloud bill moved in the wrong direction, and the on-call engineer spent…

Start with a telemetry bill of materials

Observability Cost Governance Strategies: Stop Paying for Telemetry Nobody Uses

At 02:17 on a Tuesday, a Kubernetes deployment in Johannesburg began emitting unusually verbose debug logs. The application was healthy, but Loki ingestion climbed rapidly, the cloud bill moved in the wrong direction, and the on-call engineer spent the next hour deciding whether the spike was an incident or merely expensive noise. That is the operational failure that Observability Cost Governance Strategies must prevent: not the existence of telemetry, but the absence of ownership, limits and useful defaults.

The pressure is measurable. Grafana Labs’ 2024 observability survey reported that 61% of respondents regarded cost or unexpected bills as one of their biggest observability concerns.[1] Its 2025 survey put average observability spend at 17% of compute infrastructure spend, while noting that the ratio varies widely between organisations.[2] The answer is not to switch off monitoring. It is to make every metric, log and trace justify its storage, transmission and query cost.

Start with a telemetry bill of materials

Before changing retention or sampling, build an inventory that resembles a cloud resource register. For each Prometheus scrape job, Loki stream and Tempo instrumentation library, record:

  • The owning team and service.
  • The environment, region and data classification.
  • Ingestion volume and cardinality.
  • Retention and storage tier.
  • Primary operational questions answered.
  • Monthly cost in rand, including egress where applicable.

This exposes common surprises: a staging cluster scraping production-style exporters, a pod UID embedded in a metric label, or an access log copied into both Loki and a separate object store. Grafana dashboards are useful here because cost panels can sit beside service-level indicators rather than in a finance spreadsheet nobody checks during an incident.

Use a naming convention that makes ownership queryable. For example, labels such as team, service, environment and region belong in a controlled vocabulary. Do not add arbitrary labels such as request IDs, email addresses or full URLs to Prometheus metrics. Those values create time series faster than most teams realise.

groups:
  - name: telemetry-governance
    rules:
      - record: service:http_requests:rate5m
        expr: sum by (service, environment, region) (
          rate(http_requests_total[5m])
        )

The recording rule above keeps a useful service-level view without preserving every high-cardinality dimension in every dashboard query. It does not replace raw metrics; it gives engineers a stable, cheaper interface for common questions.

Make Prometheus cardinality a production control

Prometheus cost is driven less by the number of dashboards than by active time series, scrape frequency and retention. A metric with labels for HTTP method and status may be reasonable. Add route, customer, container ID, pod hash and an unbounded exception message, and the same metric becomes a liability.

Use relabelling to drop labels that have no alerting or diagnostic purpose. Reduce scrape intervals for slow-changing infrastructure, but keep short intervals where alert detection depends on them. A node exporter rarely needs the same cadence as a payment API.

scrape_configs:
  - job_name: node-exporter
    scrape_interval: 30s
    metric_relabel_configs:
      - source_labels: [__name__]
        regex: node_network_(carrier|info)
        action: drop

For Mimir, apply per-tenant limits and monitor rejected samples, active series and query load. A limit is not a substitute for engineering discipline, but it prevents one team’s broken instrumentation from consuming the entire shared platform. Alert on approaching limits before ingestion fails; a blocked telemetry pipeline during a major release is a poor time to discover the budget.

OpenTelemetry adoption has also increased the importance of consistent metric semantics. Grafana Labs reported in 2024 that 85% of surveyed organisations used OpenTelemetry and 89% invested in Prometheus.[3] Standardisation can reduce duplicate collectors, but only if teams agree which system is authoritative for each signal.

Put Loki logs on a retention ladder

Logs are where cost governance often becomes visible first. Retaining every application log for 30 days in hot storage is convenient, but it is rarely necessary. Separate logs by operational value:

  • Security and audit records: retain according to legal, contractual and incident-response requirements, with restricted access.
  • Application errors and structured events: keep longer when they support reliability investigations.
  • Routine request logs: retain briefly in Loki and archive only when there is a defined need.
  • Debug output: enable temporarily, with automatic expiry.

Use structured logging and avoid indexing every field as a Loki label. Labels should identify streams, not become a second database. Keep request IDs, user-agent values and URLs in the log body or structured metadata unless there is a clear query requirement.

{app="checkout", environment="prod"} |= "timeout"
| json
| line_format "{{.timestamp}} {{.service}} {{.request_id}} {{.message}}"

Sampling routine access logs at the agent can materially reduce ingestion, but sample by decision rather than randomly. Preserve all errors, slow requests and requests associated with a failed deployment; sample successful 200 responses. This makes a post-incident investigation more useful than a uniform reduction in volume.

For South African estates, include connectivity in the calculation. Shipping verbose logs from an on-premises data centre or a remote African site to a central cloud region costs bandwidth as well as storage. A local agent or regional Loki gateway can filter and batch data before forwarding it. During an unreliable last-mile connection, buffering and backpressure are preferable to allowing telemetry traffic to compete with customer traffic.

Sample Tempo traces according to risk

Tracing is valuable precisely because it follows a request across boundaries, but full-fidelity tracing can become expensive in high-volume services. Tempo should receive more traces from transactions that are slow, failed, novel or linked to a release, and fewer routine successful requests.

Tail-based sampling is well suited to this policy because the decision can be made after observing the complete trace. Retain traces containing errors, high latency or selected business-critical routes. Use head sampling for low-risk traffic where tail sampling infrastructure would cost more than the data it saves.

  • Keep 100% of failed traces.
  • Keep 100% of traces exceeding the service latency objective.
  • Keep a controlled percentage of successful requests.
  • Increase sampling temporarily during a release, migration or incident.

Do not use trace IDs as metric labels. Instead, connect Grafana panels to Tempo through exemplars. That preserves a path from an aggregated latency metric to a representative trace without creating a separate metric series for every request.

Build governance around failure modes, not invoices

A monthly bill arrives too late to control a telemetry explosion. Governance needs operational signals and clear escalation paths. Create dashboards for ingestion rate, active series, log bytes, trace spans, retention, query latency and rejected data. Set budgets per team or platform tenant, expressed in both volume and rand.

Useful alerts include:

  • A sudden increase in Loki bytes per service.
  • Prometheus or Mimir active series exceeding an agreed baseline.
  • Tempo span volume increasing without a corresponding traffic increase.
  • Cross-region egress rising after a configuration change.
  • Telemetry cost per successful transaction exceeding its budget.

Load-shedding-aware alerting matters in South Africa. During scheduled or unscheduled power interruptions, an on-premises cluster may reconnect in bursts, replay buffered logs and trigger false alerts. Suppress duplicate notifications during a known maintenance window, but keep alerts for data loss, battery failure and prolonged service unavailability. Alerting should become quieter during a power event, not blind.

Review telemetry in the same change process as infrastructure. A new exporter, label, log field or instrumentation library should include an estimated volume and owner. When a team cannot explain what operational decision a signal supports, that signal should not receive indefinite retention.

Keep POPIA and sovereignty in the cost discussion

Cheaper storage is not automatically acceptable storage. Logs and traces can contain names, account numbers, IP addresses, headers or payload fragments. Apply redaction before data leaves the workload boundary, and classify telemetry before choosing a regional backend. For workloads subject to POPIA obligations or contractual data-location requirements, retaining sensitive telemetry outside the required jurisdiction may create compliance exposure as well