Pluggable AI tooling: three providers, one contract

Why every AI SDK should live behind a single contract before it turns into architecture you cannot safely change.

The fastest way to make AI expensive to maintain is importing a provider SDK directly into product code.

The first version feels harmless:

import OpenAI from 'openai';

Call the SDK from a resolver, ship the feature, move on.

Six months later, changing models means touching resolvers, workers, quotas, metrics, retries, and observability code scattered across the repo. Model IDs end up hardcoded in places nobody remembers. Want to A/B Sonnet against GPT-4o? Now you're editing business logic instead of config.

I've seen the same pattern with billing and queues. AI just accelerates the failure mode because providers, pricing, and model policies move constantly.

The problem is not the SDK.

The problem is letting the SDK become load-bearing infrastructure with no boundary around it.

tools.ai is that boundary for Refract. Same pattern the boilerplate already uses for mailer, queues, and payments: one contract in application code, vendor details behind it.


The shape of the problem

The trap is not in the first integration. The trap is everything that accumulates around it.

Streaming gets a custom wrapper.

Then retries.

Then token counting for quota enforcement.

Then metrics per provider.

Then a second provider for fallback.

Then a third because one model got dramatically cheaper.

Every addition touches a different layer because there was never a single owner for AI orchestration. By the time you want to swap providers, it's a scavenger hunt.

That's usually the signal the abstraction boundary was placed too low.


The contract

tools.ai exposes four methods:

tools.ai.complete(...)
tools.ai.stream(...)
tools.ai.countTokens(...)
tools.ai.getAvailableModels(...)

That's the entire surface area application code touches.

Handlers call tools.ai.complete. They do not import Anthropic. They do not branch on provider names. They do not know whether the response came from Claude, GPT-4o, or Gemini.

Provider SDKs, retries, API keys, and transport quirks stay behind the boundary.

Same rule as billing:

product code should describe intent, not infrastructure.

Resolvers stay thin. Utilities or core interpret responses higher up the stack.


Calling it

Usage is identical regardless of provider:

const result = await tools.ai.complete(
  {
    prompts: [{ role: 'user', content: 'Summarize this invoice.' }],
    model: 'claude-sonnet-4-20250514',
  },
  tools,
);

The call site chooses a model. Routing happens internally.

Pass a Claude model ID and the request routes to Anthropic.

Pass gpt-4o and it routes to OpenAI.

No conditional logic leaks into handlers.

Unknown model IDs fail before a network request is sent.

In production, callers should always pass an explicit model. Defaults exist in config for internal tooling and local development, not as a substitute for deciding what workload actually needs.

Streaming uses the same contract:

const stream = await tools.ai.stream(
  {
    prompts: [{ role: 'user', content: 'Draft a reply to this email.' }],
    model: 'gpt-4o',
  },
  tools,
);

The caller consumes chunks the same way regardless of provider.

Switching models does not rewrite stream consumers.


Token counting exists for a reason

countTokens looks minor until quotas and context windows enter the picture.

Counting before sending lets the application reject requests cleanly when they would:

  • exceed a model context window
  • violate a quota ceiling
  • blow through a cost budget

That is much easier to reason about than catching a provider error after the request already left your system.

getAvailableModels reflects runtime config, not a hardcoded catalog. Remove a model from config and it disappears from the API surface automatically.

Nothing in the repo ships a canonical list of model IDs.


Configuration

tools.ai is optional.

If the config block is omitted, the client is never initialized.

Add one provider and AI becomes available anywhere tools is passed.

tools: {
  ai: {
    request: {
      timeoutMs: 60_000,
      maxAttempts: 3,
      retryDelayMs: 500,
      totalTimeoutMs: 120_000,
    },
    anthropic: {
      apiKey: process.env.ANTHROPIC_API_KEY,
      models: ['claude-sonnet-4-20250514'],
      defaultModel: 'claude-sonnet-4-20250514',
    },
    openai: {
      apiKey: process.env.OPENAI_API_KEY,
      models: ['gpt-4o'],
    },
  },
},

The shared shape:

type AiConfigType = {
  request?: AiRequestConfig;
  anthropic?: AiProviderConfig;
  google?: AiProviderConfig;
  openai?: AiProviderConfig;
};

Multiple providers can run simultaneously.

Routing follows the model field per request, not branching logic in handlers.

That detail matters because it keeps provider choice declarative.

The request config is shared across providers:

  • timeoutMs applies per attempt
  • totalTimeoutMs caps the entire retry window
  • retryDelayMs governs backoff between attempts

If maxAttempts: 3 conflicts with the total timeout budget, the total budget wins.

The goal is predictable request ceilings, not infinite retry enthusiasm.


One retry rule worth understanding

Retries on generation are not truly idempotent.

A timeout does not guarantee the provider stopped processing.

The model may have completed generation and the response was simply lost in transit. Retrying sends a second request and can incur a second billed completion.

This is different from queue jobs or database writes where retries are usually expected and safe.

For generation endpoints, maxAttempts should be sized conservatively for the user experience and cost profile you actually want.

Token counting and model listing are safe to retry aggressively. Text generation is not.


The important part is containment

The interesting architectural decision here is not supporting three providers.

It's isolating ownership.

ConcernOwner
Provider SDKs and authtools.ai internals
Model routingConfig + model IDs
Retry and timeout policyShared request config
Token estimationtools.ai.countTokens
Product behaviorHandlers and utilities

Each concern has one place to live.

That containment is what makes the system durable.

Adding a fourth provider, swapping a model, changing retry policy, or moving traffic between vendors becomes a configuration change instead of a refactor across application code.

The product layer keeps describing business intent.

Infrastructure stays behind the seam.


The short version

  • Import provider SDKs once, behind tools.ai
  • Application code only touches four methods
  • Model routing follows config, not branching logic
  • Multiple providers can run simultaneously
  • Retries on generation are not safely idempotent
  • Swapping models should be a config edit, not a refactor

Models become configuration.

Product code stays product code.

Shipped May 25, 2026. See the changelog for release notes and the AI tooling overview for setup and observability.

FAQ

Why shouldn't I import the OpenAI or Anthropic SDK directly into my application?

Direct SDK imports spread provider-specific logic throughout your codebase, making it harder to change models, switch providers, or centralize retries, metrics, and authentication.

What should an AI abstraction layer expose?

Keep the surface area intentionally small. Typical methods include:

tools.ai.complete(...)
tools.ai.stream(...)
tools.ai.countTokens(...)
tools.ai.getAvailableModels(...)

Everything provider-specific stays behind that interface.

Should model selection happen in application code?

Application code should specify which model it needs. Provider routing should happen inside the AI layer based on configuration, not through if statements or provider-specific branches.

Why is token counting important?

Counting tokens before sending a request helps enforce quotas, avoid context window errors, and control costs before an API call is made.

Are AI generation requests safe to retry?

Not always. A timeout does not guarantee generation stopped. Retrying may create a second successful completion and a second bill, so generation retries should be configured conservatively.