Why a service map is not enough

At 02:17, the checkout error rate rose in Johannesburg while every Kubernetes health check remained green. The application team saw a healthy API, the database team saw normal query latency, and the network team saw no packet loss.…

Why a service map is not enough

End-to-End Service Dependency Visibility Models: Finding the Real Failure Path Across Hybrid African Estates

At 02:17, the checkout error rate rose in Johannesburg while every Kubernetes health check remained green. The application team saw a healthy API, the database team saw normal query latency, and the network team saw no packet loss. The missing detail was a dependency three hops away: a payment token call travelling from an on-premises cluster through a cloud firewall and across the Atlantic to eu-west.

That incident exposed the practical value of End-to-End Service Dependency Visibility Models. A service map is not merely a diagram of applications. It is a working model that connects user symptoms to requests, infrastructure, queues, databases, third-party APIs and regional network paths. Without that model, teams investigate dashboards in isolation and mistake “healthy” components for a healthy service.

Why a service map is not enough

Most dependency views begin with topology: service A calls service B, which reads database C. Useful, but incomplete. Operators also need to know whether the relationship is synchronous or asynchronous, which region handles the traffic, how often the dependency fails, and whether the dependency is on the critical path.

For a South African estate, geography is part of the dependency. A request from Cape Town to a Johannesburg workload may encounter a different failure mode from one routed to a European control plane. Load-shedding can alter power, connectivity and failover behaviour at the same time. A backup link may keep packets moving while adding enough latency to breach an API timeout.

A useful model therefore joins four views:

  • Logical dependency: the service-to-service call graph, including queues and scheduled jobs.
  • Runtime dependency: pods, nodes, virtual machines, load balancers, databases and gateways executing the work.
  • Evidence dependency: metrics, logs and traces that prove a relationship exists and show its health.
  • Business dependency: the customer journey, revenue path or operational process affected by failure.

Grafana is most useful when these views can be explored from the same incident context rather than displayed as unrelated dashboards. A trace should lead to the relevant service metrics, logs and infrastructure panels without forcing an engineer to reconstruct labels manually.

Three models, three different questions

1. The declared model: what should depend on what?

Start with service ownership and architecture definitions. Kubernetes labels, Helm values, infrastructure-as-code, API specifications and deployment metadata can describe intended relationships. This model is stable and valuable for access control, ownership and change review.

It is also the least trustworthy model during an incident. Configuration frequently lags reality. A forgotten feature flag may activate a legacy endpoint; a shared Redis cluster may serve six teams despite appearing in only one application repository. Declared topology tells the team what was designed, not necessarily what handled the request at 02:17.

2. The observed model: what is actually communicating?

Distributed traces provide the strongest evidence for request-driven dependencies. Instrument services with OpenTelemetry and propagate trace context across HTTP, gRPC, messaging and background workers. In Grafana Tempo, spans can expose a path such as:

web-checkout
  -> order-api
     -> fraud-service
        -> payment-adapter
           -> eu-west token endpoint

Metrics complete the picture at scale. Prometheus can measure request rate, errors and latency by service, route, region and dependency. Mimir provides a horizontally scalable long-term metrics store when local Prometheus retention is not sufficient, while Loki adds the operational narrative: timeout messages, circuit-breaker transitions and provider response codes.

Observed topology has its own blind spots. Sampling may omit rare failures. Uninstrumented libraries create broken traces. High-cardinality labels can make a seemingly precise model expensive in ZAR-denominated cloud bills. A dependency map should show confidence, not pretend that every edge is equally proven.

3. The impact model: what matters when something fails?

The impact model ranks dependencies by user and business consequence. A slow image CDN is not equivalent to an unavailable payment authorisation endpoint. A failed analytics export may be tolerated for hours; a failed identity provider may stop every customer journey.

Combine service-level objectives with dependency criticality. For example, an order API may have a 99.9% availability objective, but its payment dependency may require a tighter timeout budget and a documented degraded mode. This prevents teams from paging on every downstream warning while missing the dependency that consumes the entire user-facing error budget.

Building the evidence chain with Grafana, Prometheus, Loki, Tempo and Mimir

The practical design is less about drawing edges and more about agreeing on identity. Use consistent resource attributes and metric labels such as service.name, service.namespace, deployment.environment, cloud.region and network.zone. Keep route labels controlled; raw URLs and customer identifiers do not belong in metric labels.

For a service-level view, record the expensive PromQL expressions once instead of recalculating them in every dashboard. Prometheus documentation recommends recording rules for frequently needed or computationally expensive expressions.[1] A ratio should aggregate numerator and denominator separately:

groups:
- name: service-dependencies
  interval: 30s
  rules:
  - record: service:http_requests:rate5m
    expr: sum by (service, region) (
            rate(http_server_requests_total[5m])
          )

  - record: service:http_errors:rate5m
    expr: sum by (service, region) (
            rate(http_server_requests_total{
              status_class="5xx"
            }[5m])
          )

  - record: service:http_error_ratio:rate5m
    expr: |
      service:http_errors:rate5m
      /
      service:http_requests:rate5m

Tempo traces can supply exemplars from a latency or error panel, allowing an engineer to move from an aggregate metric to a representative request. Loki then answers questions that traces often cannot: Was the timeout caused by DNS, a TLS handshake, an exhausted connection pool or an upstream 429?

For cross-region dependency visibility, include region and availability-zone context in span attributes and logs. A panel showing “payment latency” is weak; panels separating Johannesburg-to-Johannesburg, Johannesburg-to-Cape-Town and Johannesburg-to-eu-west traffic are actionable. Mimir is particularly useful when those comparisons need weeks of history across multiple Prometheus instances.

Designing for load-shedding and last-mile failure

Load-shedding should not trigger a storm of misleading alerts. If an on-premises cluster loses a power feed, node availability, packet loss, DNS resolution and application latency may all change together. Alerting on every symptom creates noise precisely when the incident requires focus.

Prometheus recommends keeping alerts simple, alerting on symptoms associated with user pain, and linking alerts to consoles that help identify the cause.[2] A better pattern is to page on a sustained customer-facing SLO breach, then annotate the dependency view with power, connectivity and failover signals.

  • Track request success and latency by site, region and customer journey.
  • Track UPS, generator and node power telemetry where it is available.
  • Track WAN packet loss, DNS latency and BGP or tunnel state separately.
  • Label failover traffic so a backup path is visible rather than mistaken for normal capacity.
  • Use longer alert windows for infrastructure symptoms that commonly blip during failover.

Last-mile connectivity deserves the same treatment. A mobile customer in a low-bandwidth area may experience failure while the application’s internal latency remains excellent. Synthetic probes from multiple South African networks, combined with real-user measurements, prevent the internal service map from becoming an operations-only view.

POPIA, data sovereignty and the cost of excessive visibility

More telemetry is not automatically better telemetry. Traces and logs can contain email addresses, identity tokens, payment references or free-text payloads. For estates subject to POPIA obligations, define what leaves the country, what is retained, and who can query it. Service dependency visibility should expose relationships and failure evidence without copying sensitive payloads into every backend.

Useful controls include attribute redaction in collectors, restricted log fields, short retention for high-volume debug data and separate tenant access for production teams. Keep low-cardinality service metadata broadly queryable, while protecting request content and personally identifiable information.

Cost also shapes the model. Retaining every span from every service in a European region may be technically simple but financially indefensible when exchange-rate pressure is high. Tail-based sampling can retain slow,