The Stack Behind Refract: Express, GraphQL, React, Astro — and Why
An honest, layer-by-layer walkthrough of every stack choice behind Refract, including which ones are architectural necessity and which are personal preference.
TL;DR
- Every boilerplate makes architectural bets. Most are implicit. Refract's are explicit — one typed config object is the entire infrastructure surface.
- The stack: Express backend, Apollo GraphQL, React + Vite portal (one SPA, two route zones), Astro marketing site, BullMQ or SQS for queues, Stripe billing.
- Every infrastructure tool sits behind a swappable typed interface. Changing queue provider is a config value and credentials. Domain code depends on the tool contract, not on the vendor package.
- This post is the honest reasoning behind each choice — including the ones that are personal familiarity rather than architectural necessity.
Every boilerplate makes bets. Most of them are implicit — the author picked what they knew, or what was trending, and the README calls it production-ready without explaining what problems the choices solve or what they cost.
Refract makes explicit bets. This post explains them layer by layer: what I picked, why, and where the honest answer is "I know this tool well" rather than "this is objectively superior."
Refract Boilerplate is the TypeScript SaaS boilerplate at userefract.io. The argument for Shape B over the default Next.js-shaped starter is in Why I chose a Node.js monorepo over Next.js. This post is just the stack.
Start here: the config object is the architecture made visible
The most important thing to understand about Refract is not any individual technology choice. It is that the entire infrastructure surface is declared in one typed config object.
The shape is ConfigType in apps/backend/src/configuration/type.ts. Each environment exports a concrete object from apps/backend/src/configuration/development.ts, staging.ts, production.ts, or test.ts. Startup validation runs through configSchema in apps/backend/src/configuration/validate.ts, which composes shared Zod pieces (for example queueConfigSchema from apps/shared/src/queue/schema.ts) with backend-specific rules.
const config: ConfigType = {
tools: {
logger: { client: LoggerClientType.PINO },
mailer: { client: MailerClientType.LOCAL },
queue: { client: MQType.BULLMQ, queues: [...] },
analytics: { client: AnalyticsClientType.LOCAL },
metrics: { client: MetricsClientType.STATSD },
paymentProcessor: { client: PaymentProcessorClientType.STRIPE },
cache: { client: CacheClientType.REDIS },
rds: { client: RDSClientType.SEQUELIZE },
},
};
MQType in backend config files is the same enum as QueueClientType from shared (apps/backend/src/configuration/type.ts re-exports it under that alias). The snippet above matches how development.ts is structured; field-level details differ by environment.
A new engineer — or an AI agent — reads this file and knows exactly what the system depends on. No hunting through scattered imports, no implicit dependencies discovered at runtime, no surprise third-party SDK buried three levels deep in a utility file. The system's shape is explicit, typed, and in one place.
Every tool loads lazily based on the client value. The queue loader is the cleanest example to read: apps/backend/src/tools/queue/loader.ts maps QueueClientType to a workspace package name, then dynamic-imports it so the unused implementation is not loaded.
const QUEUE_PACKAGE_MAP: Record<string, string> = {
[QueueClientType.SQS]: 'tooling-queue-sqs',
[QueueClientType.BULLMQ]: 'tooling-queue-bullmq',
};
export const buildQueue = async ({
configuration,
}: {
configuration: ConfigType;
}): Promise<QueueType> => {
const client = configuration.tools.queue.client;
const packageName = QUEUE_PACKAGE_MAP[client];
const mod = (await import(packageName)) as QueueModule;
return mod.buildQueue(queueConfig, registry);
};
The package that is not configured does not load. Both tooling-queue-bullmq and tooling-queue-sqs exist as workspace packages (apps/tools/queue/bullmq/, apps/tools/queue/sqs/). Only the one the config requests gets instantiated. Consumer routing stays on apps/backend/src/tools/queue/consumerRegistry.ts, which is adapter agnostic.
To verify adapter routing without digging through production, run:
make test module=backend path=apps/backend/src/tools/queue/__tests__/loader.spec.ts
What “swappable” means here
Swapping a tool swaps the implementation behind a stable contract (QueueType, payment processor interface, and so on). Product code still depends on GraphQL, domain utilities, and those contracts. You are not swapping those away with a config line.
If you add a job, you still register a consumer, keep idempotency rules honest, and respect billing ownership. The win is that Redis versus SQS, or one mail provider versus another, does not fork the domain layer.
This is the mechanism behind every "no vendor lock-in" claim Refract makes. Not a philosophy. A pattern you can read in the source and in the tooling system doc.
Express
The honest reason: maturity.
Express is fourteen years old. Its middleware model is understood by every Node.js engineer. Its failure modes are documented exhaustively. Its ecosystem is enormous. When something breaks in an Express app at 2am, the answer exists somewhere on the internet.
Newer frameworks — Hono, Fastify, Elysia — are genuinely interesting. Some are faster. Some have better TypeScript ergonomics out of the box. I chose Express because Refract is infrastructure someone else will run in production, extend with their own middleware, and debug under pressure. The right choice for that context is the one with the deepest paper trail, not the best benchmark numbers.
Express also has no opinions about application structure. That matters because Refract's structure comes from the domain architecture — onion layers, functional core, tool adapters — not from framework conventions. A framework with strong structural opinions would fight that. Express stays out of the way.
GraphQL + Apollo Server
The honest reason: explicit schema as contract, field-level control over data, and one schema serving every client.
Where the GraphQL schema lives
Operations are implemented under apps/backend/src/gql/ as co-located typeDefs plus resolvers per query or mutation, then merged in apps/backend/src/gql/index.ts with mergeSchemas. When you run make gql-codegen, apps/backend/generate_schema.sh builds tools, calls getExportedTypeDefs from that GraphQL entry, and writes the serialized schema to apps/shared/template.schema.graphql. Apollo Client codegen (apps/portal/codegen.yml) reads that file and generates apps/portal/src/gql/hooks.ts together with your portal operations.
graph TD
A["apps/backend/src/gql/"] --> B["apps/backend/src/gql/index.ts"]
B --> C["apps/backend/generate_schema.sh"]
C --> D["apps/shared/template.schema.graphql"]
D --> E["apps/portal/codegen.yml"]
E --> F["apps/portal/src/gql/hooks.ts"]
Rename a field, tighten an entitlement check, remove a resolver — every downstream consumer fails at compile time. Not at runtime. Not in production. At make build, before the deploy starts.
Field-level control matters specifically for billing and access control. A REST endpoint returns what it returns. A GraphQL resolver returns exactly what the query requests, and field-level resolvers can apply entitlement checks at the data layer rather than the presentation layer. When a second client surface appears — mobile app, partner API, internal tooling — it consumes the same schema. No new endpoints, no duplicated logic, no diverging type definitions.
Apollo Server integrates cleanly with Express middleware, handles subscriptions for real-time features, and ships the GraphQL server and client toolchain Refract uses for schema-first codegen end to end.
The maintenance caricature — edit DB, edit SDL, edit resolver, run two codegen passes — does not apply here. Refract exports schema from backend code. Resolver and schema builder live in one file per operation. make gql-codegen handles the rest. The maintenance surface is smaller than the reputation suggests.
React + Vite for the portal
The honest reason: the portal's public surface is auth and onboarding flows. Its protected surface is application UI. Neither has an SSR requirement worth the added complexity.
Route zones, layouts, and static hosting
The portal is a single React SPA — one Vite bundle, BrowserRouter, Apollo Client — that splits into two route zones:
| Zone | Paths | Auth enforcement |
|---|---|---|
| Gate (public) | /signin, /signup, /forgot-password, /reset-password/:token, /verify/:token, /invitation/:token | None — UnauthenticatedLayout |
| Admin | /admin/* | AdminLayout requires session + org membership |
| Super admin | /admin/super/* | SuperAdminLayout requires is_super_admin |
In production, Express serves index.html for all portal path prefixes — /admin, /signin, /signup, and so on — as defined in apps/backend/src/constants/browserStatic.ts. The client router handles navigation from there. Google OAuth runs through Express at /auth/google as a full page navigation, outside the React router entirely.
Auth is enforced in layout components, not at the router shell. AdminLayout redirects unauthenticated users to /signin?redirect=… and users without a current organization to /no-organization. There is no global ProtectedRoute wrapper. The public auth surface and the protected admin surface live in the same bundle because they share enough context — Apollo client, session state, organization resolution — that splitting them would create more seams than it solves.
SSR does not make sense for this shape. The gate routes are auth and onboarding flows — no SEO value, no cold-load performance requirement. The admin routes are application UI behind a session check. The marketing site handles all public content and Core Web Vitals. The portal's job is application UI, and for application UI a client-side SPA with layout-level auth guards is the right tool.
Astro for the marketing site
The honest reason: the marketing site is content. Astro is built for content. And structural isolation from the portal is not a convenience — it is a correctness property.
The marketing site is a separate Astro app in the monorepo. It shares no runtime with the portal. It does not download Apollo Client, MUI, or any portal dependency. A visitor to userefract.io gets a fast static site. A buyer who signs in gets the full portal application. Those are different products with different performance requirements and they deploy independently.
The isolation means a portal dependency upgrade — a new MUI major, an Apollo Client version — carries zero risk of affecting landing page performance or Lighthouse scores. That boundary is structural. It cannot be accidentally violated by an import.
BullMQ and SQS
The honest reason: two real queue providers, both shipping today, as live proof that the adapter pattern works.
This is not a roadmap item. BullMQ and SQS are both implemented and tested. The discriminated config union and Zod parsing live in apps/shared/src/queue/schema.ts (QueueConfigSchemaType). Switching between them is a config value and credentials:
// BullMQ — Redis-backed, runs locally and in production
{
client: QueueClientType.BULLMQ,
connection: { url: process.env.REDIS_URL },
queues: [...], // each entry matches queueConfigItemSchema in apps/shared/src/queue/schema.ts
}
// SQS — AWS-backed, same queue definitions, different credentials
{
client: QueueClientType.SQS,
region: process.env.AWS_REGION,
endpoint: process.env.SQS_ENDPOINT,
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
queues: [...],
}
Full shapes (including DLQ maxMessages and retry fields) mirror apps/backend/src/configuration/development.ts.
The Zod schemas validate the config shape at startup through apps/backend/src/configuration/validate.ts. The two providers have different credential shapes — BullMQ takes a Redis connection, SQS takes AWS credentials and a region — and the rest of the codebase sees neither. The product code talks to QueueType. The adapter handles the rest.
Queue setup in the docs: BullMQ and SQS.
The queue configuration encodes operational knowledge, not just connectivity:
// Illustrative — see development.ts for the full knobs (maxMessages, retryDelay, DLQ maxMessages, …)
{
name: QueueName.STRIPE_WEBHOOK,
maxRetries: 1, // idempotent — aggressive retries create more problems than they solve
errorRetryDelaySeconds: 60,
dlq: { name: QueueName.STRIPE_WEBHOOK_DLQ, maxMessages: 100 },
},
{
name: QueueName.EMAIL,
maxRetries: 3, // delivery matters, consumer is safe to retry
dlq: { name: QueueName.EMAIL_DLQ, maxMessages: 100 },
},
Retry policies are not defaults carried over from a template. They reflect the actual behavior of each queue's consumers.
Stripe
The honest reason: Stripe is what Refract ships today for real billing: subscriptions, checkout, and webhooks at the depth this codebase exercises. Merchant-of-record processors are a different trade. If you need one, the payment processor slot is designed like the queue slot: a new workspace adapter, not a rewrite of domain code. That path is not shipped yet.
Stripe gives direct control over checkout and subscription mechanics. Subscription schedules handle trials, phased pricing, and mid-cycle changes without a bespoke state machine in Refract. The webhook surface is rich enough to build event-driven billing with explicit ownership per action.
The Stripe adapter and consumers focus on the boring hard parts: idempotent webhook handling, subscription reconciliation, trial and invoice edge cases. The billing architecture doc maps money flow to real files. The changelog shows how much surface area that actually is.
Local development
make start boots the full stack in one command:
- Backend
- Consumers (queue workers)
- Portal (Vite dev server)
- Marketing (Astro dev server)
- Stripe CLI (webhook forwarding)
- PostgreSQL
- Redis
- LocalStack (SQS-compatible, if you need it)
What runs locally is structurally identical to what runs in production — different adapter values in the config, same code paths. LocalStack means SQS consumers can be tested locally without an AWS account. Stripe CLI means webhook flows work without tunneling. You can run the full billing lifecycle — subscription creation, trial expiry, payment failure, webhook retry — on your laptop before a deploy.
Production is lean: webapp, consumers, PostgreSQL, Redis. No proprietary platform dependencies, no mandatory cloud vendor. The local dev richness lives in the monorepo structure and the adapter layer — not in the runtime requirements.
The database layer: Sequelize
The honest reason: I have used Sequelize in production billing systems for years and I trust it completely. That is the primary reason.
But there are technical reasons it fits Refract's architecture specifically that are worth naming — not as an argument against Prisma, but as an explanation of why the choice compounds well with how Refract is built.
Types are your code, not a generated artifact. Sequelize with
InferAttributes and InferCreationAttributes means the model class is
the type definition. There is no intermediate schema file, no prisma generate step that can fall out of sync with your models. In a codebase
that already runs make gql-codegen to propagate GraphQL types, adding a
second required codegen step for the database layer is real friction.
With Sequelize, you rename a field and TypeScript tells you immediately —
no intermediate build step required.
Schema composition is native TypeScript. Prisma Schema Language has no import system and no mixins. If you have 20 models that all need audit fields (created_by, updated_by, deleted_at), you copy-paste those fields 20 times into a flat text file. In Sequelize you define a base model class and extend it. For a boilerplate where model patterns repeat across the domain layer, this is a meaningful maintainability difference — not a theoretical one.
Migrations and seeders are first-class TypeScript. Sequelize migrations
are plain TypeScript files under apps/backend/src/tools/rds/sequelize/migrations/. Models
live in apps/backend/src/tools/rds/sequelize/models/. You can fetch data from an API, run conditional
logic, execute within explicit transactions, and compose seed operations
programmatically. Prisma's seeding story is a separate script with limited
native tooling for complex seeding workflows.
The database docs entry for the shipped ORM: Sequelize (tooling).
The database adapter follows the same swappable interface as every other tool in the stack — rds.client: RDSClientType.SEQUELIZE is today's value in a typed config field. I have not built Prisma or Drizzle adapters yet because I have not had a business reason to. When that reason arrives, it is a new workspace package against an established interface. The slot is already there.
ORM choice is not the reason a SaaS succeeds or fails. Any competent
engineer can work with Sequelize, Prisma, or Drizzle. The architecture
does not care which one is in the config. But if you are going to use
Sequelize, use it correctly — InferAttributes, typed associations,
explicit model initialization, migrations as code. Refract shows you that
pattern and uses it throughout.
What this stack costs
More explicit structure than a single-framework app. A new contributor reads more files before they understand the system. The monorepo has more moving parts than a Next.js project. make start boots eight processes where a simpler setup boots one.
Those costs are real. What you get for them:
The infrastructure surface is visible in one config file. Adding a background job means writing the job — not first figuring out where domain logic lives or how to connect a new process to the existing domain layer. Changing a queue provider is a config value. The marketing site's performance is structurally isolated from portal complexity. Local dev runs the full billing lifecycle without external dependencies.
This stack is right for you if your product will have real billing complexity within the next year — subscriptions, trials, seat limits, webhook reconciliation — and you want that complexity to have explicit ownership boundaries from day one rather than discovering them under pressure later.
It is not right for you if you are building a prototype you might throw away, or if your primary risk is market risk and you need to move as fast as possible in the next eight weeks. In that case, pick the fastest starter you can find and come back when the architecture starts to matter.
If you want the live map after this inventory piece, read in this order: architecture overview (how deployables connect), tooling system (loaders and adapter contracts), GraphQL (schema spine in backend code), then billing architecture (money paths and file ownership).
FAQ
- What makes Refract's infrastructure surface explicit instead of implicit?
A single typed config object declares every tool dependency (logger, mailer, queue, payment processor, database), so a new engineer or AI agent can see the entire system's shape in one file.
- How does Refract make infrastructure vendors swappable?
Every tool loads lazily behind a typed contract based on the config value, so switching providers like BullMQ to SQS is a config change and credentials, not a rewrite of domain code.
- Why does Refract use Express instead of a newer framework like Hono or Fastify?
Express has fourteen years of documented failure modes and no strong opinions about application structure, which fits Refract's domain-driven architecture better than a framework with its own conventions.
- Why does the portal use a client-side SPA instead of server-side rendering?
The portal's routes are either auth flows with no SEO value or session-gated application UI, so SSR adds complexity without a real performance or discoverability benefit.
- Why did Refract choose Sequelize over Prisma?
Primarily long-standing familiarity from production billing systems, plus practical fit: Sequelize's typed model classes avoid a second codegen step and support native TypeScript schema composition.