Your Durable Workflow Replays Its Own History, and Your Deploy Just Changed the Function
A one-line change to a signup workflow froze ten thousand in-flight trials, because durable execution re-runs your function from the top against a recorded history. What deterministic replay actually demands, how to split decisions from effects, and the two ways to version a workflow while old runs are still in the air.

Durable execution sells itself on one line: write an ordinary async function, and the platform makes it survive crashes. Temporal, Cloudflare Workflows, Azure Durable Functions, Restate, they all pitch some flavor of it, and the pitch is basically honest. What gets left off the slide is that your function does not run once. It runs again from the top every time a worker resumes it, with the results of the steps it already finished replayed back in from an event history. That detail is fine right up until you deploy.
The deploy that stranded 10,000 signups
The workflow was a signup: create the account, charge the card, provision the tenant, then wait out the 14 day trial and record whether it converted. Because of that wait, there were always something like ten thousand runs parked mid-flight.
Someone added subdomain reservation. Sensible change, one line, in the obvious place:
// v2
export async function signup(ctx: Ctx, input: Signup) {
const account = await ctx.run("create-account", () => createAccount(input));
await ctx.run("reserve-subdomain", () => reserve(input.slug)); // new
await ctx.run("charge", () => chargeCard(account, input.plan));
await ctx.sleep("trial", DAYS_14);
// ...
}Deploy goes out. New workers start picking up runs that v1 started, and the engine does what it always does: feed the recorded history back into the code and check that the code asks for the same things in the same order. v1's history says step two was charge. v2's code asks for reserve-subdomain. Mismatch, so the workflow task fails, and it keeps failing on every retry, because the code is never going to change its mind.
Nobody got double charged, which is the good news and also why it took two days to notice. The runs were not erroring in any way a customer could see. They were frozen, and the alerting everyone had was on activity failures, not on workflow tasks that fail and retry forever. Ten thousand trials quietly stopped aging.
The worse version of this bug is the hand rolled one. If you built your own resume logic on a DB row with a last_completed_step integer, there is no history to compare against, so inserting a step renumbers everything after it and your code cheerfully re-runs a charge it already made. The nondeterminism error is the system working.
Keep the decisions boring
The fix is structural: the function that makes decisions gets to do nothing else. No clock, no random, no network, no reading config that shifts between deploys. Anything that touches the world moves behind a recorded step, so its result lands in history once and replays as a constant from then on.
This is the rule people break by accident, and it never looks like a mistake while you are writing it:
// replayed six days later, this is a different answer
if (Date.now() - account.createdAt > TRIAL_MS) await downgrade(account);
// recorded once, same answer on every replay
const now = await ctx.run("now", async () => Date.now());
if (now - account.createdAt > TRIAL_MS) await ctx.run("downgrade", () => downgrade(account));Steps also have to be idempotent, because delivery is at least once and a worker can die between completing a side effect and writing it down. Derive the idempotency key from the run id and the step name instead of generating one inside the step. A UUID minted inside a retried charge is just a second charge with good intentions.
The version problem does not go away
Deterministic replay does not solve deploys. It converts a silent corruption into a loud stall, and you still have to ship changes while old runs are in the air. Two answers work.
Pin runs to the code that started them. Temporal does this with worker build IDs, routing each run to a compatible worker so v1 runs finish on v1 workers. Clean, and it means you keep shipping the old code until the last v1 run drains. With a 14 day sleep that is two weeks. With an annual renewal workflow it is a year, and teams tend to learn that right after deleting the old build.
Or gate the change inside the code. if (await ctx.patched("add-subdomain")) takes the new branch for new runs and the old branch for runs that predate the patch. Cheaper to operate, worse to read, and the gates stay until the cohort drains. They never get removed on their own, so put the cleanup on a calendar with the drain date written on it.
Either way, the number you need is open executions grouped by code version. That query is most of the operability story here and almost nobody has it on a dashboard. Add an alert on workflow task failures kept separate from activity failures, since the first one is what a stuck run looks like.
What you are paying for
History is append only, so a workflow that loops for a long time grows until the engine refuses it. You cut it with continue-as-new, starting a fresh run from a summarized input and dropping the old history. Replay also makes debugging strange. A stack trace can come from code that ran four days ago on a machine that no longer exists, and your breakpoint fires five times for one logical execution.
The alternative is a state machine in Postgres with a cron poking it. No replay rules, no versioning dance, no engine to operate. You write every timeout, every retry, and the crash recovery path by hand, and sleep-for-14-days becomes a scheduled job plus a status column plus the edge cases you find in production. For anything that waits in units of days, I would still take the engine and do the versioning homework. For a three step pipeline that finishes in a second, the Postgres row is fine and probably better.
None of this is really about architecture. Your workflow code is a pure function of its history, and a deploy changes the function while the history stays exactly where it was. Once you see it that way, the versioning rules stop feeling like ceremony.