Idempotent Stripe Webhooks in Node.js TypeScript: Patterns That Survive Replay
Why Stripe webhook duplicates are a documented certainty, not a bug, and the three-part pattern that makes billing survive replay.
Stripe will deliver the same webhook event more than once. That's not an edge case, it's documented behavior: Stripe's own docs say your endpoint may receive the same event more than once and that delivery order isn't guaranteed (Stripe webhooks documentation). The same event.id can land on your endpoint twice, sometimes minutes apart, sometimes back to back.
If your webhook sends an email, swaps a subscription, or closes a checkout session, that second delivery is where things go wrong. I found this out at Loom, the first time around it wasn't a stack trace, it was a support message asking why a customer had gotten two "you're now on the Pro plan" emails ten minutes apart. Nothing had crashed. Nothing had errored. Two separate webhooks, customer.subscription.updated and invoice.paid, had both fired for the same upgrade close enough together that each one thought it was the only thing finalizing the change. The handler looked correct in isolation and broke under replay. That's the bug that doesn't show up in a demo, it shows up three weeks later in your support queue.
The fix that fixed it for me is simple to state and easy to get wrong: dedupe the event so it never runs twice, and dedupe the business outcome so two different events can't both finalize the same checkout. The examples below come from a real codebase, Refract, the TypeScript SaaS foundation I built, file paths included, because I think showing the actual implementation is worth more than describing it abstractly. But if you're not on Node, Postgres, or Redis, ignore the file names and keep the shape: a fast lock, a durable ledger, one owner per side effect. That part travels.
What replay actually breaks
Before the patterns, it's worth being concrete about what goes wrong, because none of it shows up in testing:
- Duplicate invoices or credits. An upsert without a dedupe key writes the row twice.
- Duplicate emails. A plan-changed notification fires once per delivery instead of once per business event.
- Race conditions across event types.
payment_intent.succeededandinvoice.paidboth try to finalize the same checkout. - Partial transactions. A worker dies mid-write, Stripe retries, and a second worker repeats work the first one half-finished. I hit all four of these in the first month of real billing volume at Loom. None of them showed up before that.
Why duplicate Stripe webhooks happen
This isn't a bug on Stripe's side, it's the contract. Stripe retries failed deliveries for up to three days and can deliver the same event more than once even after your handler already returned a 200 (Stripe webhook best practices). A few concrete triggers:
- Your endpoint times out before it finishes processing, so Stripe assumes failure and retries.
- A network blip drops the response even though your handler completed fine.
- Stripe emits more than one event type for a single business outcome. A checkout can trigger
payment_intent.succeededandinvoice.paidfor the same purchase. - Events arrive out of the order you'd assume from the dashboard. The failure mode isn't "Stripe sent it twice." It's building as if delivery were exactly-once when the API promises something weaker (Martin Fowler on the idempotent receiver pattern).
Definitions. Event idempotency: the same webhook event is processed once. Business idempotency: the same business outcome is applied once, regardless of which event or which code path triggered it. You need both. Event dedupe alone doesn't save you when Stripe sends two different events for one outcome.
If you're evaluating a boilerplate that claims "Stripe included," ask how long it takes to get a passing webhook replay test running. A calm demo hides a flaky lock every time.
The action ownership rule
One action, one owner.
For every side effect, email enqueue, subscription swap, session close, cache invalidation, exactly one function is allowed to execute it. Everything else can read state and reconcile, but it can't re-run the effect that function owns.
That's the whole rule. Here's what it looks like wired into a real handler, not because the names matter to you, but because seeing it mapped onto something concrete makes the abstract version stick:
| Business action | Owner | Non-owners |
|---|---|---|
| Checkout-driven plan swap + close session | invoicePaymentPaid.ts | paymentIntentSucceeded.ts only when an invoice is on the payload; customerSubscriptionUpdated.ts skips while checkout is open |
| Plan-changed email for checkout | maybeEnqueuePlanChangedEmailsForCheckoutSession in billingTransactionalEmails.ts | Called from invoicePaymentPaid and paymentIntentAmountCapturableUpdated; second call no-ops |
| One-off purchase finalize | finalizeOneOffCheckoutForSession in oneOffCheckout.ts | confirmPayment, paymentIntentSucceeded, invoicePaymentPaid all call the same finalizer |
If you're not on Refract, build the same table for your own handlers before you write a single line of locking code. It takes about twenty minutes and it's the cheapest bug prevention in this entire article. Without it, idempotency keys multiply across handlers and you still end up double-sending, because event-level dedupe does nothing when two different events both think they're allowed to finalize the same outcome.
How I dedupe at the event level
This is the section of Refract's billing architecture that handles webhook replay, if you want to see it in the context of the full billing system rather than pulled out into an article.
HTTP ingress, fast reject. apps/backend/src/routers/api/stripe.ts verifies the signature with Stripe.webhooks.constructEvent on the raw body, checks hasWebhookEventBeenProcessed, validates the payload with Zod, and enqueues to STRIPE_WEBHOOK. Duplicates return 200 without re-enqueueing.
Queue consumer, transaction + Redis NX + durable marker. stripeWebhookConsumer in apps/backend/src/tools/queue/consumers/stripeWebhookConsumer/index.ts runs inside a DB transaction:
hasWebhookEventBeenProcessedreads Redis andprocessed_webhooks_events.tools.cache.helpers.webhooksEvents.set({ resource: event.id, nx: true })acquires a 15-minute Redis lock.- If
NXfails, roll back and exit. Another worker already owns this event. - Run the typed handler. On
falseor a throw, delete the Redis key and roll back so Stripe can retry. - On success, insert
ProcessedWebhooksEventwith a uniqueevent_id, then commit. The table enforces uniqueness at the database level:
event_id: {
type: DataTypes.STRING,
allowNull: false,
unique: true,
},
Redis is the fast lock. Postgres is the durable ledger. If you're on a different stack, SQS instead of BullMQ, DynamoDB instead of Postgres, the roles stay the same: something fast that holds a lock for the duration of one attempt, and something durable that survives a crash. Here's why you need both and not just one:
Worker A acquires Redis NX lock for event.id
|
Worker A crashes mid-handler (no commit)
|
Redis TTL expires (15 min)
|
Stripe retries the same event
|
No committed processed-event row yet
|
Worker B acquires the lock and re-runs the handler
|
Duplicate side effect, no error anywhere
A unique constraint on the event id closes that gap. The Redis lock is a speed optimization for the common case where nothing crashes. It is not a guarantee, because a TTL is a timer, not a transaction.
Business-level idempotency on checkout
Event dedupe doesn't help when Stripe sends payment_intent.succeeded and invoice.paid for the same purchase. Two different event.id values, same business outcome.
finalizeOneOffCheckoutForSession loads the checkout session with LOCK.UPDATE, validates tenant, amount, and currency, then checks is_open:
if (!lockedSession.is_open) {
return { ok: true, outcome: 'skipped_idempotent_closed_session' };
}
Whichever event arrives first closes the session. The second one logs and returns success without touching state again. Recurring checkout follows the same shape: invoicePaymentPaid looks for an open CheckoutSession. If none exists, it logs and returns true. Something else already finalized.
Worked example. A customer upgrades through hosted checkout. Stripe fires customer.subscription.updated as the item changes, then invoice.paid when money settles, sometimes also payment_intent.succeeded if the payload carries an invoice reference. Without ownership, customerSubscriptionUpdated hydrates the new plan and enqueues a plan-changed email. Then invoicePaymentPaid swaps the subscription and enqueues the same email again. Replay either event and support gets duplicate mail.
With ownership: invoicePaymentPaid owns checkout finalization and is the only path that calls maybeEnqueuePlanChangedEmailsForCheckoutSession, which uses a Redis SET ... NX keyed per checkout session, so a second call no-ops. customerSubscriptionUpdated checks for an open checkout session first and skips its own transition emails if one exists:
if (openCheckoutSession) {
ftLogger.info(
'Skipping downgrade transition email in subscription.updated during active checkout',
{ checkoutSessionId: openCheckoutSession.id, eventId },
);
return { didEmitTransitionEmailFromDowngrade: false };
}
Reconciliation still runs. The side effect that belongs to checkout doesn't run twice.
The full pattern, end to end
Stripe sends event
|
Verify signature on raw body
|
Check durable event ledger (processed_webhooks_events)
|
Already processed? --YES--> Return 200, no-op
|
NO
|
Acquire Redis NX lock
|
Locked by another worker? --YES--> Roll back, exit
|
NO
|
Run handler inside DB transaction
|
Business finalizer checks is_open / existing session
|
Insert processed-event record
|
Commit
Seven steps: verify, check the ledger, lock, execute the owner, finalize, record, commit. That ordering is the whole pattern, independent of whichever event types your product happens to use.
The test that proves it
Specs should assert mechanism, not happy-path luck. The exact paths below are mine, the shape of the three specs is what's worth copying:
- Event replay (
stripeWebhookConsumer/__tests__/index.spec.ts): whenhasWebhookEventBeenProcessedis true, the handler is never called. On success,webhooksEvents.setis asserted called with{ nx: true }andProcessedWebhooksEvent.createreceives the sameevent_id. - Business dedupe (
billingTransactionalEmails.spec.ts): first call tomaybeEnqueuePlanChangedEmailsForCheckoutSessiongets'OK'from Redis and the mailer queues once. Second call with the same checkout session getsnullback and the mailer is not called. - Checkout race (
customerSubscriptionUpdated.spec.ts): the downgrade transition email is skipped when an open checkout session exists. Replay the same fixture twice in one test if you want certainty. The second run should produce zero new side effects.
Can your billing survive replay?
Five questions, answered honestly:
- Can the same
event.idrun twice and change anything? - Your frontend almost certainly calls a confirm endpoint the moment payment succeeds, and Stripe almost certainly sends a webhook for the same event around the same time. Can both of those paths try to finalize the same checkout independently?
- Can
invoice.paidandpayment_intent.succeededboth finalize the same purchase? - Can your email or notification queue dedupe per business event, not per delivery?
- Can you prove any of the above in a test, or only in your head? Two or more "I don't know" answers means you don't have replay safety. You have a webhook that hasn't been replayed yet.
The minimum bar
If you don't do anything else from this article, do these three things:
- Add a unique constraint on the event id in whatever table tracks processed events. One line of schema. It's the single highest-leverage fix here.
- Pick one function as the owner for every side effect that matters, email, plan swap, credit, refund, and make every other code path that touches the same outcome check state instead of re-running the effect.
- Write one test that replays the same event twice and asserts nothing happened the second time. If you don't have this test, you don't know your webhook is safe, you're just hoping it is. Everything else in this article, the Redis lock, the transaction boundary, the open-checkout guard, is a refinement on top of these three. They're also the three things I'd check first if I were auditing someone else's billing.
What it costs you to skip this
None of this is hard. It's roughly twenty extra lines of code and one schema constraint. What's expensive is finding out you needed it after a refund processes twice, or after finance asks why MRR doesn't match the Stripe dashboard. Billing bugs don't fail loudly. They fail quietly, for months, and then all at once when someone reconciles the numbers and the totals don't match.
If you're using Refract, this is already wired in and you don't have to think about it. If you're not, take the pattern and not the file paths: one durable ledger, one fast lock, one owner per side effect, one test that proves it. That's the whole thing.
The webhook that broke for me at Loom took about a day to fix and three weeks to notice. It's cheaper to build it right the first time.
FAQ
- Why do duplicate Stripe webhooks happen?
It's part of Stripe's contract. Endpoint timeouts, dropped responses, multiple event types for one business outcome, and out-of-order delivery are all documented, normal triggers for retries.
- What's the difference between event idempotency and business idempotency?
Event idempotency ensures the same webhook event runs once. Business idempotency ensures the same outcome applies once, even when two different event types both try to finalize it.
- Is returning a 200 status enough to guarantee idempotency?
No. A 200 only stops Stripe from retrying that specific delivery. It does nothing about duplicate work triggered by a different event type or a racing confirmation path.
- Should you rely on Redis alone for webhook deduplication?
No. Redis locks expire, so a durable store like Postgres with a unique constraint on event_id is needed to survive crashed workers and restarts.
- What's the minimum fix to make Stripe webhooks replay-safe?
Add a unique constraint on event_id, assign one owner function per side effect, and write one test that replays the same event twice and asserts nothing happens the second time.