Back to BlogCloud

Your Worker 404'd a Row You Just Wrote, and the Replica Is Only 40ms Behind

A commit, an enqueue, and a read pool that is 40ms stale is enough to kill a background job and leave nothing useful in the logs. Why sleep() and retries only hide it, how to tie reads to a replication position instead of a stopwatch, and what that plumbing actually costs you.

replication lagread-your-writespostgresqueuesconsistency
Your Worker 404'd a Row You Just Wrote, and the Replica Is Only 40ms Behind

A user changes their display name, we redirect them to their profile, and the page shows the old name. Somebody files it as a caching bug. We add a cache buster, it goes quiet, and it comes back three weeks later.

It was never the cache. The write went to the primary, the read went to a replica that was 40ms behind, and the redirect won.

The profile version of this is cosmetic. The version that wakes you up is uglier.

The race nobody sized

An upload endpoint inserts an asset row, commits, pushes a thumbnail job onto the queue, returns 201. The worker picks the job up about 12ms later. It selects the asset by id from the read pool, because at some point somebody sensibly decided background jobs shouldn't hammer the primary. It gets zero rows.

The job throws, retries three times over the next 30 seconds while a bulk import is holding the standby back, then lands in the dead letter queue. No thumbnail. Nothing in the logs mentions replication, because from the worker's point of view the row genuinely does not exist.

So it gets triaged as a null check. It isn't one. The app asked two different machines about the same fact at two different points in logical time and believed both of them.

sleep(250) and other crimes

The usual first fix is a delay. await sleep(250) before the read, or a 5 second visibility delay on the job, or exponential backoff that "works" because by attempt three the replica has caught up.

Replication lag is not a constant you can budget against. It moves with write volume, with vacuum, with a long analytics transaction on the standby, with a cross-region network hiccup, with the 4 million row backfill someone scheduled at 2am. Your 250ms holds until the day it needs nine seconds, and that is precisely the day you already have a different incident open.

The retry flavour bothers me more, because it mostly works, so nobody learns anything. Then a teammate adds a read that can't be retried, inside a synchronous HTTP response or an email send, and the same bug shows up wearing a new hat.

There's a worse shape too. If the worker does read-modify-write instead of just read, a stale read becomes a stale write: it loads the old row from the replica, computes from it, and writes the result back to the primary, quietly clobbering a field that was already correct. That one doesn't fail. It corrupts, and you find out in a support ticket.

Tie the read to a position, not a clock

What you actually want are two of the guarantees DDIA spends a chapter on. Read your writes: once I've written something, my next read sees at least that. Monotonic reads: once I've seen a value, I never see an older one, which is the thing that breaks when a client bounces between two replicas with different lag and time appears to run backwards on refresh.

Both get cheap once reads carry a replication position instead of a hope. On the write path, grab the primary's WAL position and hand it back to the caller:

await tx.commit();
const { rows } = await primary.query('select pg_current_wal_lsn() as lsn');
res.cookie('rlsn', rows[0].lsn, { httpOnly: true, maxAge: 30_000 });

On the read path, pick the pool based on whether the replica has replayed that far:

async function readPool(minLsn?: string) {
  if (!minLsn) return replica;
  const { rows } = await replica.query(
    'select pg_last_wal_replay_lsn() >= $1::pg_lsn as ok', [minLsn]);
  return rows[0].ok === true ? replica : primary;
}

Note the === true. On a primary, pg_last_wal_replay_lsn() returns NULL, so the comparison is NULL rather than false, and a lazy truthiness check sends you somewhere surprising.

Postgres 18 added pg_wal_replay_wait(), which lets the standby block until it reaches a target LSN rather than making you poll and route. MySQL has had the equivalent for years with @@GLOBAL.gtid_executed on the writer and WAIT_FOR_EXECUTED_GTID_SET() on the replica. Different syntax, same idea: the client names a point, the server tells you whether it's past it.

For the worker case the token goes in the job payload, which is the part people skip:

await queue.add('thumbnail', { assetId, minLsn: rows[0].lsn });

Now the job carries its own causality. The worker can wait, or route to the primary, and the decision is made with a fact instead of a guess.

What it costs

Some primary read traffic comes back. Not much in practice, since tokens expire in tens of seconds and most reads aren't preceded by a write from the same session, but you should measure the split rather than assume it.

The plumbing is the real tax. Every boundary has to forward the token: cookie, internal header, gRPC metadata, job payload. A service that silently drops it doesn't error, it just quietly downgrades to "stale is fine," which is why I'd emit a counter for requests that arrive with no token and a counter for ones where the replica check failed. Those two numbers tell you whether the mechanism is still alive six months from now.

You also need a ceiling and a breaker. Wait at most a couple hundred milliseconds, then read the primary. And if the caught-up check is failing for most requests, stop checking and route everything to the primary, otherwise you pay an extra round trip to learn what you already knew while the standby is falling further behind.

If none of that sounds worth it, the honest alternative is to not have the problem. Classify your reads once: anything a session reads right after writing it, and anything a freshly enqueued job needs, goes to the primary. Replicas serve search, listings, dashboards, analytics. That's a worse architecture on paper and it survives contact with a team of twelve much better than a token that one service forgets to forward.

What I'd push back on is the middle ground, where the read pool is used everywhere because it's the default in the config, and the gaps are papered over with sleeps nobody can justify. Pick a guarantee, write down which reads get it, and make the code say so.