All writing
Distributed Systems5 min read

Exponential backoff alone will not save you

Backoff spreads retries out in time. It does not spread clients apart from each other. Without jitter a synchronised herd stays synchronised — it just gets quieter in between the spikes.

Every engineer learns to add exponential backoff to a retry loop. Far fewer learn the second half of the lesson, which is that exponential backoff on its own barely helps when the clients are synchronised — and clients are almost always synchronised, because they were all knocked over by the same event.

The canonical demonstration is Marc Brooker's Exponential Backoff And Jitter on the AWS Architecture Blog, published in 2015 and still the clearest treatment of the problem. It is worth walking through what the simulation actually shows, because the conclusion is more nuanced than the version that gets repeated second-hand.

The setup

Brooker models optimistic concurrency control: many clients trying to update the same row, each reading a value, computing a new one, and writing it back conditionally. The simulated network has a mean delay of 10ms with a variance of 4ms.

OCC guarantees progress — exactly one client wins each round — which produces two unpleasant properties. Completion time grows linearly with the number of contending clients, because it takes N rounds for N clients to succeed. Total work grows with the square of contention, because N clients compete in the first round, N−1 in the second, and so on.

Backoff helps less than you would expect

The standard fix is capped exponential backoff: after each failure, sleep for an exponentially growing interval up to some ceiling.

// Capped exponential backoff — the version most people ship
sleep = Math.min(cap, base * 2 ** attempt)

Running the simulation with this in place reduces client work only slightly. Plotting when the calls happen shows why: the calls do get less frequent, but they still arrive in tight clusters. Every client computed the same sleep interval from the same starting point, so they all wake up together.

Instead of reducing the number of clients competing in every round, we've just introduced times when no client is competing.
Marc Brooker, Exponential Backoff And Jitter

That sentence is the whole point of the post. Backoff without randomness does not decorrelate the herd; it just adds silence between stampedes.

Three ways to add randomness

The fix is a small change to the sleep calculation. Brooker names three variants:

// Full Jitter — sleep anywhere between zero and the backoff ceiling
sleep = random(0, Math.min(cap, base * 2 ** attempt))

// Equal Jitter — keep half the backoff, randomise the other half
temp  = Math.min(cap, base * 2 ** attempt)
sleep = temp / 2 + random(0, temp / 2)

// Decorrelated Jitter — grow the ceiling from the previous sleep
sleep = Math.min(cap, random(base, sleep * 3))

Equal Jitter exists to avoid very short sleeps, keeping some guaranteed slowdown. Decorrelated Jitter is stateful — each delay is derived from the last one, so clients drift into independent rhythms rather than all sampling from one shared distribution.

What the simulation actually found

With 100 contending clients, adding full jitter cut the call count by more than half and significantly improved time to completion compared with un-jittered backoff. Beyond that headline, the rankings are closer than most summaries admit:

  • No-jitter exponential backoff is the clear loser on both work and time — by so much that it had to be left off the comparison graph.
  • Among the jittered approaches, Equal Jitter comes last: slightly more work than Full Jitter, and it takes much longer.
  • Full Jitter versus Decorrelated Jitter is genuinely unresolved. Full Jitter uses less work; Decorrelated Jitter is slightly faster. Brooker calls the decision "less clear" and does not pick a winner.

One caveat Brooker states plainly and almost everyone drops: none of these approaches change the N² nature of the work. Jitter reduces the constant substantially at reasonable contention levels. It does not fix an algorithm that scales quadratically.

What jitter does not fix

Jitter solves synchronisation. It does nothing about three other failure modes that show up in layered systems, and in my experience these cause more production damage than the herd itself.

Retry amplification

Retries compose multiplicatively through a call stack. If every layer independently retries three times, the load reaching the bottom is not 3× — it is 3 raised to the number of layers.

Layers retryingRequests at the bottom
13
29
327
481
Requests arriving at the deepest service from a single user request, when every layer retries 3 times.

A struggling database at the bottom of a four-layer stack sees an 81× load multiplier at exactly the moment it can least afford one. This is why retrying at every layer is an anti-pattern: pick the layer with the best information about whether a retry is worthwhile, and make the others fail fast.

No retry budget

Per-request retry limits do not bound aggregate retry load. The standard fix is a token bucket at the client: retries consume tokens, successful requests replenish them, and when the bucket empties the client stops retrying entirely. This caps retries as a percentage of successful traffic, so a total outage produces roughly zero retries instead of infinite ones. The AWS SDKs implement this in their adaptive retry mode.

Retries that are not safe to repeat

A timeout tells you nothing about whether the work happened. The request may have been lost on the way out, or completed successfully with the response lost on the way back. Retrying a non-idempotent operation on timeout is how duplicate charges happen. If an operation can be retried, it needs an idempotency key that the server deduplicates on — and that key has to be generated by the caller before the first attempt, not per attempt.

Politeness is not a control

Client-side backoff is voluntary. It protects you from well-behaved clients you control, which is not the population that takes services down. Server-side admission control — load shedding, concurrency limits, returning 429 with Retry-After — is the only mechanism that works against clients you do not control, including the badly configured internal one somebody deploys next quarter.

Reasonable defaults

  • Use full jitter unless you have measured a reason not to. It is one line, stateless, and close enough to optimal that the remaining difference is noise.
  • Always cap the backoff. Unbounded exponential growth turns a brief blip into a client that reconnects tomorrow.
  • Add a retry budget as a token bucket, and emit a metric when it empties — that metric is an excellent early warning.
  • Retry only what is safe: timeouts, 429, 503, connection failures. Never retry a 400 or a 422; the answer will not change.
  • Put a circuit breaker in front of dependencies that can fail wholesale, so you stop paying timeout latency on every call.
  • Shed load at the server. Assume some client will ignore every rule above.

The 2015 post closes by noting that the return on implementation complexity for jittered backoff is huge. That remains true. It is also, still, only the first of the four things a retry loop needs.

Sources

  1. 01Marc Brooker. Exponential Backoff And Jitter. AWS Architecture Blog, 2015.
  2. 02Marc Brooker. Timeouts, retries, and backoff with jitter. Amazon Builders' Library.
  3. 03aws-arch-backoff-simulator. GitHub.