Three Retries Per Layer Is 27 Calls to the Service That's Already Down
Stacked retry logic multiplies instead of adding: three layers with three attempts each means 27 calls to a dependency that is already saturated. How retry budgets, deadline propagation and jittered backoff replace the hardcoded retry count, and what each one costs you.

Checkout started returning 504s at about 2am. Traffic was flat, dead flat, same shape as the previous Tuesday. The Postgres pool was pinned at 100% anyway and the pricing service was answering roughly one request in twenty.
The trigger was boring: a long GC pause on pricing pushed its p99 from 40ms to around 900ms. That alone should have been a blip. What turned it into forty minutes of downtime was three separate pieces of retry code, each written by someone reasonable, each reviewed and approved, none of them aware of the others.
The multiplier nobody writes down
The mobile client retried a failed checkout up to three times. The gateway's HTTP client had retries: 3 in its config, set during a flaky-network incident two years earlier. And the checkout service used an axios interceptor that also retried three times, because of course it did.
Nobody typed the number 27. But that's what pricing saw for every one checkout a user attempted.
Retry layers multiply. When you review a one line PR that adds a retry, your brain models it as addition: one extra request, worst case. It isn't. The amplification factor is the product of every layer in the call path, and that product sits dormant until the day something gets slow.
Retries assume a failure mode you probably don't have
A retry is the right tool for an independent, transient fault. A dropped packet. A pod that got rescheduled mid request. A leader election blip in your database. The second attempt lands somewhere healthy, the user never notices, everyone goes home.
Overload is a different animal. When a dependency is failing because it is saturated, your retry is perfectly correlated with the cause of the failure. You are adding load to the thing that broke from load. That feedback loop is why these outages outlive their trigger: the GC pause ended in 900 milliseconds, but the retry traffic it generated kept pricing pinned for half an hour. The literature calls this a metastable failure, and the useful part of that framing is that removing the trigger does not fix it. You have to drain the retries.
Worth saying out loud: none of this is safe unless the work is idempotent. If POST /charge gets retried by three layers, you need an idempotency key that the server stores and dedupes on, or you have just invented a way to bill someone 27 times.
Deadlines travel, timeouts don't
Here's the companion bug, and it's the one I see more often. Every service picks its own timeout independently. The client gives up after 2 seconds. The gateway's timeout is 5 seconds. The service's query timeout is 10.
So three seconds after the user's phone has stopped listening, your gateway is still holding a socket, your service is still holding a connection from a pool of 20, and Postgres is still executing a query whose result will be thrown into a void. Under load, a big slice of your capacity goes to requests nobody is waiting for anymore.
The fix is to pass the remaining budget down the call chain and have every hop refuse work it cannot finish in time. gRPC does this for you. Over HTTP you do it by hand:
// inbound: trust the caller's remaining budget, clamped to our own ceiling
const claimed = Number(req.headers.get("x-timeout-ms"));
const budgetMs = Math.min(Number.isFinite(claimed) ? claimed : 5000, 5000);
const deadline = Date.now() + budgetMs;
async function callPricing(deadline: number) {
const left = deadline - Date.now() - 50; // reserve time to send our own response
if (left < 100) {
throw new DeadlineExceeded("not enough budget left to start");
}
return fetch(PRICING_URL, {
signal: AbortSignal.timeout(left),
headers: { "x-timeout-ms": String(left) },
});
}Send a remaining duration, not an absolute timestamp. Absolute deadlines are cleaner right up until you discover one host's clock is 400ms off, and then they are a mystery to debug.
The counterintuitive win is that rejecting a request early is cheap capacity. A request you refuse in 2 microseconds frees a connection for a request that can still succeed.
Spend a retry budget, not a retry count
The actual replacement for retries: 3 is a token bucket, fed by successes.
Every success deposits a fraction of a token, every retry spends a whole one:
class RetryBudget {
private tokens = 0;
constructor(private ratio = 0.1, private max = 100) {}
recordSuccess() {
this.tokens = Math.min(this.max, this.tokens + this.ratio);
}
tryRetry() {
if (this.tokens < 1) return false;
this.tokens -= 1;
return true;
}
}With ratio = 0.1, retries can never exceed 10% of your successful traffic. A healthy dependency funds plenty of retries for the odd dropped packet. A dying dependency produces no successes, so the bucket empties, retry traffic goes to zero on its own, and the dependency gets a chance to recover. Envoy ships this as retry_budget, Finagle has had it for years, and gRPC's retryThrottling is the same idea with a different shape. Real implementations add decay so the bucket doesn't hold stale credit from an hour ago.
Then do the unglamorous parts: retry at exactly one layer (the one closest to the user usually knows best whether anyone still cares), use full jitter rather than plain exponential backoff so your clients don't resynchronise into waves, and honour Retry-After when a server bothers to send it.
What this costs you
A budget means some requests that would have succeeded on attempt two now fail. You are trading a slightly worse p99.9 error rate for not having a forty minute outage, which I think is obviously the right trade, but it is a trade and someone will file a ticket about it.
The budget is per process unless you share state, and that's fine here in a way it isn't for rate limiting. The bucket is funded by local successes, so it scales with however much traffic that pod is handling. No Redis, no coordination.
Deadline propagation is weaker: one service in the path that ignores the header burns everyone's budget. And idempotency keys cost you a write per mutating request plus a decision about how long to keep them.
What to watch
Attempts per logical request, not request count. That ratio is your amplification factor. Mine should sit near 1.02 and alert above 1.2.
Budget exhaustion events, as a counter per dependency. A spike here is the early warning that used to be an outage.
Requests rejected for insufficient deadline. Rising numbers mean the mechanism is working.
Connection pool wait time, which is where this always shows up first.
Go open your HTTP client config this week and find the retry number. It is probably 3. The interesting question is what it's being multiplied by.