When the trace goes dark
The alert is real and several explanations are plausible. Missing spans remove the evidence needed to distinguish them, so the investigation stops short of a fabricated answer.

Several explanations fit. None can be verified.
At 21:08 UTC, the synthetic API gateway begins returning intermittent 502 responses for account requests. The alert is narrow—roughly 4% of traffic—and the affected requests share no region, tenant, or deployment version.
Metrics prove that failures exist. Logs offer two plausible stories: DNS resolution errors on the profile path and token-validation timeouts on the auth path. Distributed traces should separate them. Instead, trace coverage for failed requests falls from 96% to 18% at the same moment.
“Most likely” is a hypothesis. A verified RCA requires the evidence that distinguishes it from the alternatives.
Signal packet
gateway.route with no linked downstream span.Map what the evidence cannot connect.
1. Confirm the symptom without averaging it away
Overall gateway success remains above 95%, so a fleet-wide view makes the system look mostly healthy. Splitting by response code confirms a real failure cohort, but the usual dimensions do not isolate it. This establishes the symptom and little else.
2. Inspect the last reliable span
Every surviving failed trace stops at the gateway boundary. A missing child span could mean the downstream request never started, the instrumentation failed, or the collector dropped the span later. The trace alone cannot distinguish those states.
3. Test both log-derived hypotheses
DNS failures and token-key refresh timeouts both occur during the alert window. Their counts are high enough to explain the observed 502 volume, but neither contains a shared request ID. Time proximity is correlation, not request-level evidence. Choosing either one would require assuming that the other is unrelated noise.
4. Check the evidence system itself
Collector metrics reveal a full export queue and dropped spans. This explains the telemetry gap, not the gateway failures. It also removes the one signal that could have assigned each failure to auth or profile. The investigation has discovered an observability incident nested inside an application incident.
The available signals support two competing causal chains and provide no request-level link that falsifies either one. A verified root cause cannot be assembled from this packet.
The correct conclusion is incomplete.
This investigation should not emit a normal Causal Chain RCA. It can report the confirmed symptom, the trace-collection failure, the two live hypotheses, and the next evidence required. Anything stronger would convert missing telemetry into invented certainty.
{
"status": "evidence_insufficient",
"confirmed": [
"gateway 502 cohort",
"trace export queue saturation"
],
"hypotheses": [
"profile DNS failure",
"auth key-refresh timeout"
],
"missing_evidence": "request-level downstream span or correlated log id",
"next_action": "restore collector capacity and preserve failed-request correlation"
}The record above is an editorial representation of how to communicate uncertainty. It is not presented as Tiravan’s current production response schema.
What an on-call engineer should carry forward
- Alert on telemetry-pipeline health with the same seriousness as application health.
- Propagate one request identifier into traces and structured logs so either signal can repair gaps in the other.
- Record competing hypotheses and the evidence required to falsify each one.
- Make “insufficient evidence” an acceptable operational outcome. Confidence theatre is not reliability.