One Tenant's Spreadsheet Took Every Worker for Forty Minutes
A single FIFO job queue is a scheduling policy, and the policy is that whoever enqueues the most wins. Here is how one 50,000 row import starved every other tenant for forty minutes without firing a single alarm, what per-tenant lanes with concurrency caps and leases look like instead, and what that architecture actually costs you.

The queue was healthy. The customer was not.
A billing service I worked on had one Redis list named jobs and twenty workers pulling off it. Median job: about 280ms. Boring, and boring was the point.
Then a customer uploaded a spreadsheet with 50,000 line items, and the importer enqueued one job per row. For the next forty minutes, every other tenant's invoices sat behind that spreadsheet. Nothing paged. Job duration was still 280ms. Error rate was zero. Worker CPU was fine. The queue depth alarm was set at 100,000, so it stayed quiet too.
Every dashboard was right about what it measured. None of them measured the thing that broke: how long a job waits before a worker picks it up, split by who enqueued it.
A single FIFO queue looks like a buffer. It is really a scheduling policy, and the policy says whoever enqueues the most gets the most capacity. Totally fine when one team owns all the work. It turns into a liability the moment the work belongs to different customers who each think they bought a share of the system.
Adding workers does not fix this
The instinct is to scale the pool. Go from 20 workers to 40 and the forty minute wait drops to twenty. Still bad, and now you are hammering Postgres twice as hard on behalf of jobs nobody is waiting for. We found that out the slow way: autoscaling on queue depth turned a latency problem into a lock contention problem, and the small tenant's invoice was still late.
Head-of-line blocking is a position problem, not a throughput problem. If your job is 50,001st in line, the only real fix is changing the line.
One lane per tenant, plus something that chooses
Replace the single list with a list per tenant and a small scheduler in the claim path. Workers stop asking "what is next" and start asking "who is owed time".
The claim logic is maybe fifteen lines:
// tenants:ready is a sorted set scored by the last time we served that tenant
async function claimNext(workerId: string) {
const tenants = await redis.zrange("tenants:ready", 0, 50);
for (const tenant of tenants) {
const inflight = await countLeases(tenant);
if (inflight >= capFor(tenant)) continue; // this tenant is at its ceiling
const job = await redis.lpop(`jobs:${tenant}`);
if (!job) {
await redis.zrem("tenants:ready", tenant);
continue;
}
await takeLease(job, workerId, { ttlMs: 60_000 });
await redis.zadd("tenants:ready", Date.now(), tenant); // back of the line
return { tenant, job };
}
return null; // genuinely idle, not blocked
}Three things are doing work in that loop.
Least recently served ordering means a tenant that just got a slot goes to the back, so a new arrival waits for one round of active tenants rather than for the whole backlog. The per-tenant cap means no single tenant can hold all twenty slots even when everyone else happens to be idle for a second. And the loop keeps going when lanes are empty: if tenant A is capped and tenant B has nothing, the worker still takes tenant C's job instead of sleeping. Fair schedulers that are not work conserving are just slow schedulers.
Do not trust that inflight counter either. Count live leases with an expiry. Workers get OOM killed mid-job, and a counter you only decrement on success drifts upward until a tenant is capped at zero forever. We had one tenant stuck at zero concurrency for about a day before anyone connected it to a crash loop from the week before.
If you have 20,000 tenants rather than 200, scanning a ready set gets silly. That is where shuffle sharding earns its keep: hash each tenant onto a small random subset of lanes so a noisy tenant only collides with a fraction of the others. Marc Brooker's piece on fairness in multi-tenant systems is the best thing written on when to pick which.
What this costs you
The bulk import gets slower. A 50,000 row job capped at four concurrent slots now takes three hours instead of forty minutes. That is usually the right call, but it is as much a product decision as an infra one, so somebody has to tell the customer and give them a real progress indicator instead of a spinner that lies.
Fairness is per tenant, not per second of work. A tenant whose jobs take five minutes each still occupies a slot five minutes at a time, so a tenant sending 300ms jobs sees worse latency even under perfect round robin. The practical answer is chunking: make jobs small enough that a slot turns over quickly. Preemption sounds nicer and is miserable to implement.
You also just made your claim path stateful. That loop needs to be atomic, which in Redis means a Lua script, and in Postgres means SELECT ... FOR UPDATE SKIP LOCKED with the cap folded into the query. A racy scheduler will happily hand one tenant twenty-five slots under load, which is the exact bug you were trying to prevent.
The dashboard we should have had
Job duration is a worker health metric. It says nothing about fairness. The numbers that would have caught this:
p99 queue wait per tenant, meaning enqueue time to claim time. This is the one that would have paged us at minute two.
Share of worker seconds consumed per tenant over a five minute window, compared against their share of enqueued jobs. Drift is your fairness bug.
Longest consecutive run of dequeues handed to one tenant. If that number is in the thousands, your scheduler is not scheduling anything.
We shipped per-tenant lanes in an afternoon and spent two weeks on lease reconciliation. The scheduler is the easy part. Knowing a worker died while holding tenant C's only slot is the part that keeps the thing honest.