Pluggable scheduler tooling: one clock, one contract
Why recurring jobs belong behind a scheduler contract instead of scattered cron entries, intervals, and delayed queue hacks.
The fastest way to make recurring work expensive to maintain is wiring it through the host crontab and calling it done.
The first version feels harmless:
0 3 * * * /app/scripts/reconcile-stale-sessions.sh
One entry. One script. Ship the sweep, move on.
TypeScript makes the script awkward fast. Compile to JavaScript and run with Node, and dev executes source while prod runs dist — two paths for the same job. Or run with ts-node in production and carry a second Node toolchain on the host just to fire a nightly sweep. Neither choice gets cleaner at month six.
So the timer moves into the process:
setInterval(() => reconcileStaleSessions(), 60_000);
Six months later: three processes firing overlapping ticks. One handler retries through the queue. Another logs to stdout. A deploy restarts the API and the interval fires twice. You want daily cleanup at 03:00 UTC but someone's laptop timezone leaked into the pattern. Heavy work runs inline on the tick because nobody drew a line between when to fire and what to execute.
The problem is not cron. The problem is that wall-clock triggers have no owner.
tools.scheduler gives them one. Same pattern Refract already uses for mailer, queues, payments, and AI: one contract in application code, vendor details behind it.
The shape of the problem
The trap is not in the first nightly job. The trap is everything that accumulates around it.
A handler grows to do the sweep inline because "it's only once a day." Then retries get bolted on. Then someone enqueues with sendToQueue({ delayMs }) in a loop to fake recurrence. Then a second schedule fires the same reconciliation from a different process. Then deploys leave orphan cron entries in Redis because patterns only lived in code. Then horizontal scaling duplicates ticks because the clock lived on the API tier.
That's usually the signal the abstraction boundary was placed too low.
Recurring work needs two separable concerns:
- When should this run? (wall-clock policy)
- What should happen? (business logic)
Collapsing both into a setInterval or a delayed queue message couples timing to execution and makes both harder to change.
The contract
tools.scheduler exposes a small surface:
tools.scheduler.registerSchedules(tools)
tools.scheduler.shutdown?.(tools) // optional, BullMQ client
tools.scheduler.runScheduleNow?.(name, tools) // test client only
The surface area is intentionally tiny. Scheduling is configuration, registration, and execution. Features like pause, dynamic registration, and schedule inspection remain implementation concerns — not application concerns.
Application code never imports BullMQ. Handlers never parse cron strings from environment variables scattered across modules. Schedule names, patterns, and time zones live in tools.scheduler.schedules[] — the same declarative shape as tools.queue.queues[].
On each tick, the registered handler runs with the same boundary as a queue consumer:
type ScheduleHandlerFn = (
tick: { scheduleName: string; firedAt: string },
tools: ToolsType,
) => Promise<string | void | boolean>;
The handler receives tools. It calls tools.queue.sendToQueue when the real work belongs on a business queue. It does not block the tick worker on a long Stripe sweep — that sweep gets enqueued and retried like any other job.
Unknown schedule names fail at reconcile time, not silently at runtime.
Scheduler and queue are independent
This detail matters more than it sounds.
tools.scheduler and tools.queue are separate tools. Production may run BullMQ for schedules and SQS for business queues. That combination is valid and intentional.
The scheduler is a thin clock. Queues still own heavy execution, retries, DLQs, and observability.
| Layer | Role |
|---|---|
tools.scheduler | Fire ticks on a cron pattern |
| Schedule handler | Decide what to do on this tick — often enqueue |
tools.queue | Execute, retry, and observe business jobs |
Anti-pattern: using repeated sendToQueue({ delayMs }) to simulate recurrence. Deferred queue messages are for one-off delays, not wall-clock policy. Recurring work belongs on tools.scheduler.
Where the clock runs
registerSchedules(tools) runs only in apps/backend/src/consumer.ts, after queue consumers start.
Not in the API process. Not from GraphQL. Not from a second consumer replica.
Run exactly one consumer replica for scheduler ticks. This is a deliberate tradeoff: operational simplicity over distributed coordination. There is no leader election, no distributed lock, no scheduler quorum. For Refract's workload — and almost certainly yours — that trade is worth making. The tools.scheduler contract doesn't change if a future design introduces leader election or a dedicated scheduler service. The constraint stays behind the seam.
The API backend service serves HTTP. The consumer service owns long-lived workers: queue consumers and scheduler registration. That split keeps request latency independent of cron reconciliation and tick processing.
Both tools are optional in configuration. Omit tools.queue and enqueue paths no-op through sendToQueueIfConfigured. Omit tools.scheduler and the consumer skips registration. A deployment can run API-only, consumer-with-queues-only, or both — without forking the codebase.
Production: reconciliation is the important behavior
The most important thing BullMQ does here is not run ticks. It is reconcile.
Refract uses BullMQ job schedulers (v5.16+), not the deprecated repeatable-jobs API. Cron metadata lives as durable scheduler entries in Redis; each tick is a job from the scheduler template.
On consumer startup, the BullMQ client (apps/tools/scheduler/bullmq/) compares Redis against config:
- List schedulers via
getJobSchedulersand remove orphans withremoveJobScheduler— including legacy repeatable keys from earlier deploys. - Call
upsertJobSchedulerwhen a schedule is missing or whenpattern/tzchanged. - Skip unchanged entries when the pattern and timezone fingerprint already match.
Configuration is the source of truth. A schedule removed from config disappears from Redis on the next consumer restart. A changed pattern takes effect on the next reconcile. No manual Redis cleanup required.
Stable IDs make this safe. Every entry uses id scheduler:${scheduleName} (buildStableScheduleJobId in shared). Without stable IDs, deploys leave duplicate or stale cron metadata in Redis — ticks drift from the config you thought you shipped.
On each tick, a worker resolves the handler by scheduleName from the job payload, records bounded metrics (scheduler.tick.*), and invokes the handler. Thrown errors fail the tick; BullMQ retries according to its policy.
Configuration
Schedules live in config alongside the rest of the tools definition:
tools: {
scheduler: {
client: SchedulerClientType.BULLMQ,
connection: { url: resolveSchedulerRedisUrl() },
schedulerQueueName: SCHEDULER_TICK_QUEUE_NAME,
schedulerTickJobName: SCHEDULER_TICK_JOB_NAME,
schedules: [
{
name: ScheduleName.DAILY_CLEANUP,
pattern: '0 0 3 * * *',
tz: 'UTC',
},
],
},
},
Each schedule name must match a handler registered in scheduleRegistry.ts. schedules: [] is valid — the consumer boots with no registered jobs until you add entries. Patterns, validation rules, and Redis connection separation are covered in the scheduler docs.
Tests: runScheduleNow, not fake cron
Jest uses SchedulerClientType.TEST — only in apps/backend/src/configuration/test.ts.
The test client has no Redis, no cron, no job schedulers. registerSchedules is a no-op. Handlers run on demand:
const result = await tools.scheduler.runScheduleNow(
ScheduleName.DAILY_CLEANUP,
tools,
);
Same handler code as production. Different clock adapter. Do not use the test client in dev, staging, or production.
Idempotency at the seam
Ticks can retry. Handlers must tolerate duplicate fires.
When a handler enqueues work, use deterministic job ids:
import { buildEnqueueJobId } from 'shared';
await tools.queue.sendToQueue(
{
queueName: QueueName.EMAIL,
message: { type: 'sweep', firedAt: tick.firedAt },
jobId: buildEnqueueJobId(tick.scheduleName, tick.firedAt),
},
tools,
);
buildEnqueueJobId buckets by schedule name and minute. A retried tick in the same minute does not spawn duplicate sweep jobs. Heavy logic still belongs in the queue consumer, where DLQ and retry policy already exist.
The important part is containment
| Concern | Owner |
|---|---|
| Cron patterns and time zones | tools.scheduler.schedules[] in config |
| Redis job schedulers and tick worker | apps/tools/scheduler/bullmq/ |
| Schedule handler business logic | apps/backend/src/tools/scheduler/schedules/ |
| Retries, DLQ, heavy execution | tools.queue consumers |
| Wall-clock bootstrap | registerSchedules in consumer.ts only |
Each concern has one place to live. That's what makes the system durable.
Adding a schedule: enum value, config entry, handler file, registry loader, consumer restart. Swapping the scheduler implementation: a loader change, not a handler rewrite. Pointing production at SQS for queues while keeping BullMQ for ticks: a config exercise.
The product layer keeps describing what should happen on a cadence. Infrastructure stays behind the seam.
The short version
- Cron patterns live in config —
tools.scheduler.schedules[]is the source of truth - Production uses BullMQ job schedulers (
upsertJobScheduler), not the deprecated repeatable-jobs API - Run exactly one consumer replica for ticks; this is an intentional simplicity tradeoff
- Scheduler and queue are independent tools — SQS + BullMQ is a valid combination
- Handlers enqueue heavy work; queues own retries, DLQs, and observability
- Reconciliation on startup removes orphans and updates changed patterns — stable IDs prevent drift
- Jest uses
runScheduleNow, not fake cron
Time becomes configuration.
Product code stays product code.
Shipped June 5, 2026. See the changelog for release notes and the scheduler overview for setup, handlers, and observability.
FAQ
- Why not use setInterval() for recurring jobs?
setInterval()couples scheduling and execution inside your application process. It also introduces problems around deployments, horizontal scaling, duplicate execution, and time zone handling that dedicated schedulers solve more reliably.- Why should cron schedules live in configuration?
Keeping schedules in configuration makes them the single source of truth. Adding, removing, or changing schedules becomes a configuration update instead of hunting through application code.
- Why separate the scheduler from the queue?
The scheduler should only decide when work runs. Queues should own retries, dead letter queues, observability, and long-running execution. Keeping these concerns separate makes each system easier to reason about.
- Why should only one scheduler instance run in production?
Running a single scheduler replica avoids duplicate ticks without introducing distributed locks or leader election. The architecture stays simple while remaining easy to evolve later if requirements change.
- How should recurring jobs be tested?
Instead of mocking cron, expose a method that executes schedules directly in tests. This allows production handlers to run unchanged while replacing only the scheduling implementation.