A retry storm disguised as database latency
Database load is real, but it is not the initiating fault. A timeout mismatch turns one slow call into six concurrent attempts and hides the upstream trigger.

The database really was saturated.
At 10:42 UTC, the synthetic order API’s p95 latency rose from 280 ms to 1.9 s. Database CPU climbed to 86%, connection utilization reached 100%, and the slow-query log filled with ordinary reads taking four times longer than baseline.
Every dashboard pointed at the database. Scaling it would have reduced the immediate pressure. It would not have explained why request volume at the database increased sixfold while customer traffic remained flat.
A downstream system can be both overloaded and innocent of initiating the incident.
Signal packet
inventory.reserve and orders.persist spans with the same request key.deadline exceeded at 250ms; retrying attempt=2 jitter=0msCount attempts per request.
1. Normalize by customer traffic
Absolute database activity makes the database look like the source. Dividing database calls by completed order requests changes the picture. The ratio moves from roughly 1.1 calls per order to 6.6 while incoming traffic remains stable. Something inside the request path is multiplying work.
2. Find repeated spans
The trace tree contains nearly identical child spans starting 250 ms apart. Each timed-out inventory reservation remains alive downstream while the order API starts another attempt. Those attempts all continue into the persistence path, so one customer request can hold several concurrent database connections.
3. Reject three attractive explanations
- Slow query regression: query plans are unchanged; latency increases only after concurrency rises.
- Cache failure: hit rate remains at 94%, and repeated calls carry the same already-warm keys.
- Traffic spike: edge request volume is flat; only internal attempts increase.
4. Identify the initiating mismatch
A client-library update reduces the order API’s inventory timeout from 800 ms to 250 ms but leaves two immediate retries enabled with no jitter. Inventory legitimately takes 300–420 ms during a routine reconciliation job. The client abandons healthy work early, retries it, and multiplies the load that eventually slows the database.
Root cause and contributing factor
- A client update sets a 250 ms deadline below inventory’s normal reconciliation latency.
- Two immediate retries start while earlier attempts continue downstream.
- Duplicated persistence work consumes database connections and CPU.
- Database saturation slows every attempt, producing still more timeouts and retries.
- The feedback loop raises order latency despite unchanged customer traffic.
{
"summary": "internal retries multiplied database work 6x",
"trigger": "inventory client deadline reduced to 250ms",
"effect": "overlapping attempts saturated database connections",
"root_cause": "deadline mismatch combined with immediate retries",
"evidence_cell_ids": [
"E-0204",
"E-0211",
"E-0216",
"E-0229"
]
}The reconciliation job is a contributing condition, not the root cause. It raises legitimate inventory latency enough to expose the unsafe client policy. Without the deadline mismatch and retry behavior, the same job completes without multiplying database work.
What an on-call engineer should carry forward
- Graph attempts per originating request, not only aggregate throughput.
- Budget nested deadlines from the outside in; a client should not abandon work that remains active downstream.
- Add jitter, a retry budget, and idempotency protection before enabling retries in multiple layers.
- Treat scaling as mitigation. Continue investigating why the workload changed.