The real problem is correlation, not a lack of dashboards
At 02:17, a payment API in Johannesburg started timing out. Prometheus showed rising latency, Loki showed a burst of connection errors, and Tempo showed traces stretching across a dependency in eu-west. The first alert named the symptom correctly,…
AI-Augmented Root Cause Analysis Systems: Finding the Fault Before the War Room Gets Expensive
At 02:17, a payment API in Johannesburg started timing out. Prometheus showed rising latency, Loki showed a burst of connection errors, and Tempo showed traces stretching across a dependency in eu-west. The first alert named the symptom correctly, but not the cause. An engineer restarted two pods, another checked the database, and the incident bridge filled with plausible guesses.
This is where AI-Augmented Root Cause Analysis Systems can earn their keep. Not by declaring “the answer” from a single dashboard, but by joining metrics, logs, traces, deployments and topology into a ranked set of hypotheses that engineers can test quickly.
The real problem is correlation, not a lack of dashboards
Most production estates already collect enough telemetry to explain an incident. The difficulty is that the evidence is distributed. A Prometheus alert may identify an SLO breach; Loki may contain the first useful error; Tempo may show that all slow requests pass through one remote service; and deployment records may reveal a configuration change eight minutes earlier.
Grafana provides the practical investigation surface for these signals. Metrics can be stored in Prometheus or Mimir, logs in Loki, and traces in Tempo, with links between related exemplars, trace IDs and log labels. The LGTM pattern—Loki, Grafana, Tempo and Mimir—is particularly useful when teams need one investigative workflow across Kubernetes, virtual machines and on-premise services.
An AI layer should sit above this evidence, not replace it. Its job is to:
- Group alerts that are probably symptoms of the same event.
- Identify unusual changes in service behaviour and dependency latency.
- Search logs and traces using the incident’s time window and affected dimensions.
- Compare current behaviour with a known-good baseline.
- Present hypotheses with supporting evidence and confidence, rather than inventing certainty.
The distinction matters. A language model that summarises an alert is useful. A system that can query the actual time series, inspect trace exemplars and show its evidence is substantially more useful.
What the Johannesburg payment incident should have shown
The initial alert was based on API latency:
histogram_quantile(
0.95,
sum by (le, route) (
rate(http_request_duration_seconds_bucket{
job="payments-api",
status=~"2..|5.."
}[5m])
)
) > 1.5
That alert was valid, but incomplete. It did not distinguish between an application regression, a saturated database pool, packet loss on the local network, or a slow cross-region dependency. An augmented investigation would add several related questions:
- Did the error rate increase for all routes or only payment authorisation?
- Did the trace duration accumulate in the application, database, or eu-west call?
- Did the problem begin after a deployment, certificate rotation or network change?
- Were retries amplifying load?
- Did the same pattern occur in Cape Town, Johannesburg and the secondary region?
Tempo can expose the slow span, while Loki can find the corresponding retry and timeout messages. Mimir can retain the longer baseline needed to compare today’s p95 with the same service’s normal weekday pattern. The AI component then turns these separate queries into an investigation timeline.
A useful finding might read: “Payment authorisation latency increased at 02:16:40. Seventy-eight percent of affected traces spend more than 900 ms in the eu-west fraud-check span. Loki records increased upstream timeout messages, while local database latency remains within its seven-day range. The likely cause is dependency degradation or network latency, not database saturation.”
That statement is actionable because every assertion can be opened and checked.
Build the evidence path before adding the model
The quality of an AI-Augmented Root Cause Analysis System is constrained by telemetry quality. Start with consistent resource attributes across all services:
service.name,service.versionanddeployment.environment.- Region, availability zone, cluster and namespace.
- Trace ID and span ID in application logs.
- Stable route, dependency and outcome labels.
- Deployment, feature-flag and infrastructure-change events.
A typical telemetry path might send OpenTelemetry data through a collector, then route metrics to Mimir, logs to Loki and traces to Tempo. Keep high-cardinality data out of metric labels. A customer ID belongs in a trace or structured log field, not in a Prometheus label that can multiply time-series costs.
For a hybrid estate, record network boundaries explicitly. A service calling eu-west should expose a dependency or region dimension; otherwise an AI system sees “latency increased” without knowing whether the increase is local or transcontinental. This is especially important where last-mile connectivity, submarine cable incidents or carrier routing changes can affect African users without affecting workloads in Europe.
Use [Grafana](https://grafana.co.za) to create investigation views that preserve the links between signals. The dashboard should not be a decorative wall of panels; it should let an engineer move from an alert to a trace, from a trace to logs, and from a log event to the deployment or infrastructure change that preceded it.
Design the AI investigator as a cautious analyst
The safest design is an evidence-retrieval loop rather than an unrestricted chatbot. The model receives the alert, service ownership, time window and permitted tools. It then executes bounded queries, records the results and proposes the next query.
- Normalise the alert into an incident scope: service, region, symptoms and start time.
- Measure the change against a baseline, not merely against a static threshold.
- Inspect correlated metrics for saturation, errors, retries and dependency latency.
- Follow trace relationships to the slowest or most frequently failing spans.
- Search Loki for matching error patterns, trace IDs and configuration changes.
- Check deployment and infrastructure events.
- Rank hypotheses and show the evidence that supports or weakens each one.
Every generated claim should carry a query, time range and source panel. If the model cannot retrieve evidence, it should say so. “No evidence found” is safer than a confident explanation based on an incomplete scrape.
Guardrails are equally important. Begin with read-only access. Restrict queries by tenant and time range. Redact personal information before logs reach an external model. Require human approval for remediation, especially for actions such as scaling, failover, deleting workloads or changing routing.
South African constraints change the design
Cost and sovereignty are not afterthoughts. Sending every log line and trace payload to a foreign AI service can create POPIA concerns, increase egress charges and place sensitive operational data outside an approved boundary. Keep raw telemetry and primary investigation data in the required South African environment where policy demands it. Send the model a minimised evidence set: aggregated metrics, redacted log excerpts, service names and selected trace attributes.
For ZAR-sensitive cloud budgets, sampling needs to be deliberate. Retain complete traces for errors and high-latency requests, but sample ordinary successful traffic. Store long-term metrics in Mimir while applying sensible retention to verbose logs. An AI investigator that requires every event at maximum resolution is not operationally efficient.
Load-shedding also deserves explicit treatment. A power event may cause hosts to disappear, battery-backed network equipment to switch state, or workloads to fail over. Alerting on every node outage can create a false root-cause trail. Model the event as an operational condition, suppress dependent alerts during an approved maintenance or power window, and continue alerting on customer-facing SLOs. The question is not whether a server went down; it is whether the service remained available and whether recovery behaved as designed.
Measure whether the system improves incident work
Do not evaluate an AI investigator by how fluent its incident summary sounds. Measure:
- Time from first alert to a validated hypothesis.
- Percentage of incidents where the cited evidence was correct.
- Number of irrelevant alerts grouped into one incident.
- False-cause rate and engineer overrides.
- Query cost, model cost and telemetry egress.
- Reduction in repeated manual investigation steps.
Run it in shadow mode first. Let the system produce hypotheses while engineers follow the existing process. Compare its ranked causes with the final post-incident findings, including cases where the correct answer was “ins