How to Evaluate a SaaS Foundation Before You Bet a Quarter on It

A staff-engineer checklist for SaaS boilerplates, plus the kill criteria that should talk you out of buying one.

A SaaS foundation is not a theme pack. It's the set of defaults that decide how expensive the next quarter of commercial change will be.

Most founders evaluate boilerplates the way roundups teach them to evaluate boilerplates: auth screens, Stripe badges, how fast npm run dev looks in a launch video. Fair enough for a weekend prototype. The bet changes once you're putting real money, real seats, and a real roadmap on top of someone else's boundaries. Or lack of them.

I keep coming back to the same uncomfortable question when a founder asks me whether to buy a starter. Not "does it work today." What does a pricing rewrite cost in three months. What does a provider swap cost. What does hiring a second engineer cost when they have to learn why checkout, webhooks, and plan gates all disagree.

In 2026 that question got sharper. A lot of engineers don't hand-write the bulk of the change anymore. They plan, they steer agents, they review the diff. If the foundation is muddy, the agent still "ships", it just ships a huge PR with the wrong layer touched, SDKs leaking into product code, and permissions folklore copied from the nearest file. You're toast in review, not because you can't code, because the repo taught the model the wrong shape.

So evaluating a foundation is partly evaluating whether change stays small and reviewable when a human or a coding agent makes it.

That's the evaluation. Everything below is the framework I actually use, including the part where I try to talk you out of Refract.

What "betting a quarter" really means

A quarter, in founder terms, is not only the license fee. It's the weeks you spend learning the shape, the features you build into that shape, and the migration tax you eat if the shape was wrong.

If the foundation is weak, that tax shows up as:

  • a pricing experiment that touches more files than the experiment is worth
  • an authorization change that turns into a product freeze
  • a vendor move that rewrites half the domain because SDKs leaked everywhere
  • agent-generated diffs that look productive until you realize you're reviewing a rewrite dressed as a feature

If the foundation is solid, the same quarter buys product work. The boring infrastructure stays boring, and the diffs stay boring too.

Oh, and the demo is almost never the signal. Demos are optimized for the first hour. Foundations are optimized for month six, when Stripe retries an invoice.paid while a user is still staring at the success page and two handlers both think they own the transition.

The five checks (steal this even if you never evaluate Refract)

These map to how a staff engineer audits a SaaS codebase. Same themes as the architecture checklist: billing blast radius, authorization boundaries, dependency isolation, revenue-path regression coverage, ecosystem risk. Different job. Here you're deciding whether to put a quarter of roadmap on someone else's repo.

CheckWhat to inspectPass condition
Billing blast radiusPlan rename, seat limit, grandfather ruleCommercial rules live in a domain layer; file count is a tracer for why those files changed
Authorization boundariesrole === / isAdmin greps, plan × role mappingsScopes or capabilities checked as a layer, not string folklore in UI and handlers
Dependency isolationVendor SDK imports outside adaptersDomain talks to tools.*; SDKs stay in one implementation package
Revenue-path regression coverageDouble webhooks, downgrade scope loss, seat blocksContract/integration tests on money and entitlement paths, not only UI green
Ecosystem riskMaintenance, production usage, failure docs, hiring, directionYou can staff it, debug it at 2am, and survive a project pivot without rewriting your product

1. Billing blast radius

Open the repo and pretend you have to rename a plan, add a seat limit, and grandfather existing customers.

Count the files.

The number is a tracer, not a score. Look at why those files had to change. If the answer is "checkout component, a Stripe helper, a webhook route, a cron reconciler, and three UI gates," you don't have billing ownership. You have billing residue. A healthy foundation keeps commercial rules in a domain layer, maps internal product/price rows at the processor boundary, and gives each money event one owner path with idempotency. Stripe itself expects duplicate webhook deliveries; your foundation should too. The adapter may talk to Stripe. GraphQL and portal code should not import Stripe.Price.

Why it matters: pricing is usually the first serious commercial change after launch. If that change is a scavenger hunt, every later packaging experiment inherits the same tax.

Try this on the repo. Rename Pro to Growth, add a seat limit, and grant a partner role billing access. Watch the diff. If the blast fans into checkout UI, webhooks, cron, and resolvers for reasons that aren't the commercial rule itself, you've found the real cost.

2. Authorization boundaries

Search for role === and isAdmin.

If permissions live as string compares in components, resolvers, and random utilities, packaging will hurt. Plans tend to sell capabilities; roles are often just a convenience label on top. A foundation worth a quarter treats authorization as a boundary: scopes or capabilities, mappings from plan × role into those scopes, and checks that ask "can they invite members" rather than "are they admin folklore." That lines up with least privilege and keeping authz as an explicit layer (OWASP Authorization Cheat Sheet).

I've watched teams ship the admin screen in a day with role strings, then spend a sprint inventing a partner role that is neither admin nor member. The second role is where the architecture confesses.

3. Dependency isolation

Pick one vendor that is not optional in production. Mailer, queue, payment processor, AI provider, whatever you actually rely on.

Ask: can you change the implementation without editing product logic?

The pass condition is boring. Domain code talks to tools.*. SDKs live in one implementation package. Config selects the client. Bonus points if there are two adapters for the same contract, because one interface with one implementation is just a polite import.

If Stripe, Resend, and the queue client are imported across handlers, you are buying a feature list with lock-in included at no extra charge.

4. Revenue-path regression coverage

Coverage percentage is a weak proxy. Ask what fails when a webhook arrives twice, when a downgrade should remove a scope, when a seat limit should block an invite. Stripe will retry deliveries for days if your endpoint is slow or unhappy, so duplicate money events aren't a theoretical edge case.

You want contract or integration coverage around billing finalizers, entitlement checks, and the queue consumers that reconcile money events. UI tests are fine. They are not a substitute for revenue-path tests.

A foundation that looks green while money paths are untested is cosplay confidence.

5. Ecosystem risk

This one is softer and still load-bearing. Ask the concrete questions:

  • Is it actively maintained?
  • Is there substantial production usage, not just launch-week stars?
  • Are failure modes documented somewhere you can find at 2am?
  • Can you hire people who already know it?
  • What happens if the project changes direction?

You're not looking for fashion velocity. You're looking for whether you can staff, debug, and survive the stack. When those answers look good, a boring stack is usually the conclusion, not the premise. Express, GraphQL, React, boring queues tend to win that audit because production SaaS fails in specific ways and someone has written the postmortem before you. Maintenance cadence matters too, but direction beats velocity. A project that only chases framework minors is not the same as a project that hardens billing and boundaries.

How to run the audit in one afternoon

You don't need a consulting engagement. You need a ruthless pass with a notepad.

  1. Clone or open the source you can actually inspect.
  2. Trace one paid checkout from button click to subscription row to webhook finalizer.
  3. Trace one permission from plan mapping to a blocked action.
  4. Grep for vendor SDKs outside adapter folders.
  5. Run or read the tests that claim to cover billing and auth.
  6. Skim release notes for the last few months and ask what improved: product architecture, or framework fashion.
  7. Ask whether a coding agent (or a new hire) can find the money-path owner and the auth boundary from repo structure and agent docs without tribal Slack lore. If a packaging change would force a scattered rewrite, the agent will scatter too, only faster.

Write the blast radius in plain language. "Plan rename touches N packages." "Partner role requires rewriting M gates." If you can't estimate blast radius from the repo structure, that is itself a result. Same if an agent can't either.

Who this is wrong for (kill criteria)

Founders trust frameworks that talk them out of a purchase more than ones that only sell. So here are the kill criteria. If several of these are true, do not buy a heavy foundation this month. Including mine.

Kill the purchase if market risk still dominates. Sixty days of runway, zero paying users, an MVP that might pivot twice before anyone cares. Optimize for UI velocity. Accept architectural debt as a future-you problem. A contract-first monorepo is the wrong premium while you still don't know who pays.

Kill the purchase if you want the framework to be the product surface. If your bet is Next.js colocated everything, Server Actions as the domain, and ecosystem speed above sovereignty, buy a Next starter on purpose. Don't force a split backend into a plan that hates it.

Kill the purchase if you need schema-driven GraphQL more than contract-driven GraphQL. By contract-driven I mean the domain contract/schema is the product boundary. By schema-driven I mean Postgres-as-GraphQL-source tools in a PostGraphile-shaped mold. Different bet. Different win condition.

Kill the purchase if the repo fails the five checks above and the author won't show you the seams. Billing in route handlers. SDKs everywhere. Role-string auth. No webhook idempotency story. Tests that never touch money. Walk away. Feature checklists will not save you in month eighteen.

Kill Refract specifically if you want the lightest possible learning curve this week. Refract has more moving parts than a single Next app. make start boots a real local stack. That cost is worth it when billing, authorization, and async jobs are first-class within a year. It is not worth it if you only needed a prettier CRUD scaffold.

I mean that last one. I'd rather lose a tire-kicker than inherit a mismatched customer who needed ShipFast energy and got Shape B architecture.

Where Refract sits inside this framework

This is the Why Refract part of the series, so I'll be direct without turning the article into a landing page.

Refract is built so the five checks are not a speech, they're the default shape. Billing sits behind a payment contract with an internal catalog. Authorization is scope-based. Queues, mailer, metrics, AI, and friends go through tools with swappable clients. Webhook side effects have owner paths. The suite is pointed at regressions that cost money, not only at buttons. The stack is deliberately boring on purpose: Express, Apollo, React, Vite, Astro, BullMQ and SQS both live.

It's also laid out so agents aren't guessing. Root AGENTS.md plus package-scoped guides tell where billing, auth, and tools live before anyone starts generating a diff. That doesn't make bad prompts smart. It makes the audit cheaper for humans and coding agents, because the owner paths are written down instead of tribal.

If you run the afternoon audit on Refract and the kill criteria still fire, trust the kill criteria. The framework is doing its job.

If the audit passes and your roadmap already includes packaging experiments, seat logic, partner roles, or provider hedges, then the quarter you're about to bet is the kind of quarter this repo was designed to absorb.

The conversion ask, said plainly

Use the checklist on any foundation you're considering. Steal it. Argue with the weights. That alone is useful.

If you want the scored, side-by-side version against the options founders usually put next to Refract, start an evaluation with the Evaluation Framework and tell me where your constraints actually are: runway, billing complexity, Next commitment, team size. The honest outcome is sometimes "don't buy Refract." That outcome is still a successful evaluation.

The unsuccessful outcome is betting a quarter on a demo.

FAQ

How do you evaluate a SaaS foundation before buying one?

Treat it as a cost-of-change audit. Measure billing blast radius (count the files a pricing change touches, then ask why those files had to change), check whether authorization is a real boundary or scattered flags, confirm external deps can swap without rewriting product code, look for tests that catch billing and auth regressions, and judge ecosystem risk: maintenance, production usage, documented failure modes, hiring, and direction risk.

What are kill criteria when choosing a SaaS boilerplate?

Walk away if billing lives in route handlers, provider SDKs are imported across domain code, permissions are stringly-typed role checks, the suite barely covers webhooks and entitlements, or the project only ships framework-chasing updates. Also walk away if your real risk is still market fit, not architecture.

When should a founder not buy Refract?

Skip Refract if you have a few weeks of runway, zero paying users, and an MVP that might pivot hard before anyone pays. Skip it if you want a Next.js-shaped product surface first and are fine paying rewrite tax later. Skip it if you need a pure Postgres-schema-driven GraphQL stack rather than a contract-first backend.

Why does billing blast radius matter so much?

Pricing changes are the first serious commercial experiment most SaaS products run. The file count is a tracer, not a score: look at why those files had to change. If a plan rename, seat limit, or grandfather rule fans out across checkout UI, webhooks, cron jobs, and resolvers, you do not own billing. You own a scavenger hunt.

What do authorization boundaries mean in a SaaS foundation?

Permissions should be a first-class concern with scopes or capabilities that plans and roles map into. If authorization is a pile of if (role === 'admin') checks in components and handlers, every packaging experiment becomes a permissions rewrite.