Three simulated investigations. Explicit assumptions. No customer telemetry.Read the method

Incident Lab

Failure, reconstructed.

Synthetic production incidents, investigated without hindsight. Each file follows the signals from the first alert to a verified causal chain—or the point where the evidence runs out.

All incidents are simulated
Abstract service topology compressing through an amber bottleneck before breaking into rose error pathsLead investigation · 8 min read
01Simulated incident

Checkout p99 after a harmless-looking deploy

pool cap reducedconnection queue growspayment latencycheckout p99 alert

A routine payment-service release changes one connection-pool default. Checkout owns the alert; a dependency two hops away owns the failure.

  • Prometheus
  • OpenTelemetry
  • Loki
  • GitHub
8 min read
Verified causal chain
02Simulated incident

A retry storm disguised as database latency

deadline mismatchunbounded retriesdatabase saturationorder latency

  • Prometheus
  • OpenTelemetry
  • Loki
  • Deploy diff
9 min read
Root cause with contributing factor
03Simulated incident

When the trace goes dark

gateway errorstelemetry gapcompeting hypothesesno verified RCA

  • Prometheus
  • OpenTelemetry
  • Loki
  • Collector health
7 min read
Evidence insufficient

Evidence before opinion. Causality before correlation.

These are not customer case studies and they are not Tiravan performance benchmarks. We author plausible services, alerts, metrics, spans, logs, and deploy events to examine one question: what can the available evidence actually support?

Every file distinguishes the initiating trigger from downstream effects, records rejected hypotheses, and stops when a conclusion would require invented evidence.

Synthetic
Services, traffic, telemetry, timestamps
Representative
Failure modes and investigation sequence
Not claimed
Customer outcomes, benchmarks, production proof

One new failure at a time.

Follow the evidence here. Bring your own telemetry when you are ready to evaluate Tiravan.