How We Implemented AGENTS.md in a Production TypeScript Monorepo
One canonical copy per instruction, package-boundary guides, and mirrors only where nesting fails.
How We Implemented AGENTS.md in a Production TypeScript Monorepo
One canonical copy per instruction, package-boundary guides, and mirrors only where nesting fails.
Cursor, Claude Code, and Codex all load project guidance differently. The obvious fix is to copy the same instructions into each system. That's also the fastest way to build documentation drift into your repo.
This is how we restructured guidance in Refract Boilerplate, our production TypeScript monorepo: 36 Cursor rule files (14 of them alwaysApply: true) replaced by a layered setup — one root AGENTS.md, six package-level indexes under apps/, one heavy domain guide for billing and webhooks, two cross-tool mirror pairs, and roughly 20 glob-scoped .cursor/rules/*.mdc files that still hold path-specific Cursor detail. Universal rules moved into AGENTS.md. Path-scoped detail stayed in scoped rules instead of getting duplicated everywhere. Not a theory piece. What we shipped, what broke on the way, and the file layout that's still running today.
The standard in 60 seconds
AGENTS.md started inside OpenAI's Codex tooling as a plain markdown file that sits next to README.md, except it's written for agents instead of humans: build commands, conventions, boundaries, the things you'd otherwise repeat in every chat session (agentsmd.io). The format is deliberately unopinionated. No required schema, no frontmatter, just headings and text that an agent parses directly (agents.md).
It's since moved past being "an OpenAI thing." In December 2025 the Linux Foundation launched the Agentic AI Foundation (AAIF) as a neutral home for the standard, with AGENTS.md donated alongside Anthropic's Model Context Protocol and Block's Goose framework, and Platinum members including AWS, Google, Microsoft, and Cloudflare (Linux Foundation). Adoption is past the "emerging" phase too: tens of thousands of open-source repositories already carry one, and every major coding agent (Codex, Claude Code, Cursor, Copilot, Gemini) reads it natively.
None of that is the hard part. The hard part is what happens once you have more than one AGENTS.md file and more than one tool reading them, which is exactly the situation a real monorepo puts you in.
The Refract Boilerplate monorepo shape
Refract — the TypeScript SaaS boilerplate at userefract.io — is organized as six packages under apps/: backend, portal, marketing, shared, tools, and documentation, all sharing a common tooling layer. Each package has its own build and test conventions, and in the marketing package's case, an entirely different rendering model (Astro SSR) from the rest of the stack (Express + Sequelize).
That heterogeneity is the whole reason a single root AGENTS.md was never going to work. "Run make test module=backend" is true for the backend and false for the marketing package's prerendering step. A rule about generated GraphQL types matters in the backend and is meaningless in the documentation package. A monorepo with genuinely different packages needs guidance that's true at the point where an agent is actually editing, not guidance that's true on average.
What we started with, and why it broke
Before AGENTS.md, guidance lived in .cursor/rules/*.mdc. By the time we started this migration, that directory held 36 files. Fourteen were marked alwaysApply: true, meaning they loaded into every single Cursor session regardless of what was being edited.
The moment an instruction existed in two places — a Cursor rule and a Claude-specific rule saying roughly the same thing — drift wasn't a risk anymore. It was scheduled. Someone updates the Cursor version after a postmortem, the Claude version doesn't get touched, and three weeks later an agent using Claude Code confidently does the thing you told it to stop doing.
We tried the obvious fixes before landing on the current structure, and all of them failed for specific, reproducible reasons:
- Copy everything into every tool. This is what we had. Guaranteed drift the moment one copy gets edited and the other doesn't, because nothing forces the two to stay in sync.
- One enormous AGENTS.md. We drafted this version. It stayed internally consistent, but locality died: an agent editing the marketing app's Astro config had to read past backend RBAC rules to get there, and "actionable" guidance turned into a wall of text nobody, human or agent, fully absorbed.
- One AGENTS.md per directory. We also tried this, briefly. Too many files, no clear ownership per file, and agents spent more effort figuring out which file governed a given edit than they saved by having the guidance at all.
- Cursor rules only, keep it tool-specific. This is closest to where we started. It works exactly as long as everyone on the team uses Cursor. The day someone runs Codex or Claude Code against the same repo, the guidance doesn't exist for them.
- Mirror everything everywhere. The inverse failure mode: instead of trading away sync, you trade away simplicity. Every rule now has two owners, which is the same drift problem with extra steps. Package boundaries were the smallest unit that balanced locality with maintainability without falling into any of the five failure modes above.
Worth being honest about the cost of that decision too: package boundaries are not free. They mean six index files (plus one heavy domain guide) to keep in sync with the packages they describe instead of one, and a new engineer, or a new agent session, has to know to check the right package file, not just the root one. We accept that cost because the alternative, one file trying to be true for six structurally different packages, produced guidance so hedged it stopped being actionable. "Run tests before committing, using the appropriate test command for your package" is not a rule an agent can execute. "Run make test module=backend" is. You only get to write the second version once you've committed to package-scoped files.
Per-app AGENTS.md vs root: when each wins
The rule we landed on: root holds universal invariants, package-level files hold locality. A universal invariant is something true regardless of which package an agent is touching — "run everything through make, never pnpm or tests directly on the host," "never commit .env files," "vendor SDKs live only inside implementation trees, never called directly from application code." Locality is anything that's only true inside one package's boundary — the marketing package's Astro build step, the backend's Sequelize migration conventions, the portal's component library.
The test we apply per rule is simple: if the instruction would be wrong or irrelevant in at least one other package, it doesn't belong at root. It goes in that package's AGENTS.md. This sounds obvious written down. In practice, about a third of what we'd originally have shoved into a "shared" AGENTS.md failed that test once we actually checked it against all six packages.
Root wins for: the Docker/Makefile workflow, monorepo-wide security constraints, the general "how this repo is organized" orientation an agent needs before it does anything else.
Package-level wins for: build and test commands specific to that package, framework-specific conventions (Astro vs Express), and anything that would actively mislead an agent working in a different package if it were promoted to root.
We didn't assume nested discovery would work. We tested it.
The whole architecture depends on one thing: nested AGENTS.md discovery actually has to work, reliably, across working directories and invocation patterns. That's not a safe assumption to make on faith, especially in Codex, where the discovery mechanism is more specific than "closest file wins."
Codex builds an instruction chain once per run. It starts at a global scope (your ~/.codex home directory), then walks from the project root down to your current working directory, checking each directory along the path for an AGENTS.override.md first and an AGENTS.md second. Files are concatenated in order from root to current directory, and the whole chain is capped by default at 32 KiB, after which Codex silently stops adding files (OpenAI Codex documentation). That cap detail matters: a monorepo with a fat root file and several nested package files can hit it without any error, and the nested guidance you were counting on just doesn't load.
So we didn't take working nesting on faith. We deliberately made root and nested guides disagree on a test instruction, then ran Codex from the repo root and from inside a package directory, on both single-file and multi-file edits, and checked what actually loaded using Codex's own instruction-summary command. This wasn't a sanity check we could skip. If nested discovery hadn't held up consistently, we'd have scrapped it and pushed everything into root as imperative "read before editing" directives instead, accepting the locality loss. It held. That result is what let us commit to the package-boundary structure instead of designing around a workaround.
This kind of hierarchical, closest-file-wins resolution isn't a Codex-only quirk, either. It's core to how the standard is meant to scale: OpenAI's own Codex repository uses the pattern at real scale, with 88 separate AGENTS.md files across its subcomponents (Agents.md best practices).
The pattern we landed on
Root AGENTS.md stays short on purpose. It's an index and a set of invariants, not a style guide:
# Refract — agent guide
Universal index for all agents. Read this first, then follow scoped
guidance in `apps/*/AGENTS.md` and tool-scoped rule files as needed.
## Repo layout
apps/backend, apps/portal, apps/marketing, apps/shared, apps/tools,
apps/documentation.
## Commands (Makefile only — never pnpm on host)
Use the root Makefile so commands run inside Docker.
- make test module=...
- make local-verification (completion gate for backend/cross-package work)
- make gql-codegen
## Universal invariants
- Use make from repo root; never run pnpm or tests on the host.
- Never commit .env files or secrets.
- Vendor SDKs live only in implementation trees (apps/tools/**,
apps/backend/src/tools/**).
- Before editing any package, read that package's AGENTS.md first.
## Where package-specific rules live
Each package under apps/ owns an AGENTS.md index. Root does not
duplicate them. Dense billing/webhook ownership:
apps/backend/src/tools/paymentProcessor/AGENTS.md.
Package-level files aren't full rule bodies. They're locality indexes that point an agent at the right scoped rule or nested guide instead of re-explaining it. apps/backend/AGENTS.md looks like this:
# apps/backend — agent guide
## Where to start
- GraphQL: apps/backend/src/gql/
- Utilities: apps/backend/src/utilities/
- Tooling adapters: apps/backend/src/tools/
## When you touch these areas
- Router routes → .cursor/rules/backend-router-conventions.mdc
- GraphQL mutations → .cursor/rules/graphql-mutations.mdc and
.cursor/rules/testing.mdc
- Test isolation → .cursor/rules/test-isolation.mdc
(.claude/rules/test-isolation.md for Claude Code)
- Configuration → .cursor/rules/no-silent-config-fallbacks.mdc
## Billing and Stripe webhooks
- Ownership and idempotency: apps/backend/src/tools/paymentProcessor/AGENTS.md
- Webhook consumers: apps/backend/src/tools/queue/consumers/stripeWebhookConsumer/
- RBAC in resolvers: canAccessScopes in apps/backend/src/utilities/rbac.ts
— not raw provider string checks
## Commands
- Test: make test module=backend
- Migrations: make migrate-up (Docker only — never against a shared database)
That index-not-body distinction matters more than it looks. The package file doesn't restate the RBAC rule or the webhook idempotency mechanism, it points at the file and the deeper guide that own those. One canonical location per rule holds all the way down, not just at the root-vs-package level.
Two rules never got a clean home in this structure, and that's what forced mirrors instead of nesting. CHANGELOG.md conventions apply to a single file that isn't scoped to any one app's directory. Test-isolation rules apply to a scattered glob pattern (**/__tests__/**) that cuts across every app. Neither nests cleanly under a package boundary, so without a mirror they'd split silently across tool-specific mechanisms with no shared source of truth. We accepted two small, deliberately boring mirror pairs instead:
.cursor/rules/
changelog.mdc
test-isolation.mdc
.claude/rules/
changelog.md
test-isolation.md
The contract on a mirror pair is strict: frontmatter can differ between the two files, but requirement text, examples, and any explicitly allowed or forbidden pattern must stay identical. If you find yourself editing the wording in one and not the other, you've reintroduced the exact drift this whole migration was meant to kill.
Verification: how we test that agents actually follow the rules
Writing an AGENTS.md file is not the same as confirming an agent reads it. We run two checks before trusting any new or edited rule.
First, a load-order check. For Codex specifically, this means asking it directly what instructions are currently active from a clean session, both from the repo root and from inside a package directory, and comparing that against what we expect to have loaded given the file layout above. If a rule we just added doesn't show up, either the file's in the wrong place or we've hit the size cap.
Second, a behavioral check. We pick a rule that's easy to violate by default — the "never call vendor SDKs directly, go through the implementation tree" rule is a good one, because it's exactly the kind of thing an agent will do the "normal" way unless told otherwise — and give it a task that would naturally tempt a direct SDK call. If the agent reaches for the adapter pattern without being told twice, the rule is doing its job. If it doesn't, the rule gets rewritten to be more explicit and less descriptive, since vague guidance ("be careful with vendor code") reliably gets ignored while concrete constraints ("go through tools.paymentProcessor, never import stripe directly outside apps/tools/** or apps/backend/src/tools/**") reliably don't.
Rules that fail this twice get deleted rather than reworded a third time. An AGENTS.md file with instructions nobody follows is worse than no file, because it creates false confidence that the constraint exists.
One scar from this process: our first draft of the webhook idempotency rule read "webhook handlers should be idempotent," which is true and useless. An agent asked to add a new subscription event handler read that line, nodded along conceptually, and still wrote a handler with no guard against replay. The rewrite named the actual mechanism — a Redis SET NX plus a processed_webhooks_events row keyed by the Stripe event.id, checked in hasWebhookEventBeenProcessed before any side effect runs — and pointed directly at apps/backend/src/tools/queue/consumers/stripeWebhookConsumer/ and apps/backend/src/tools/paymentProcessor/AGENTS.md. Same intent, different specificity. The second version got followed on the first try. That gap between "conceptually true" and "concretely actionable" is most of what separates an AGENTS.md file that works from one that just looks complete.
The precedence order, when guidance conflicts
The order we document, and the order shipped in root AGENTS.md, is: system and developer instructions in the moment, then root AGENTS.md, then the nearest nested AGENTS.md, then tool-scoped rules (.cursor/rules or .claude/rules matching the task), then local task or PR notes.
That's the reverse of a naive "most specific wins" read, and it's deliberate. Codex itself concatenates files root-to-cwd, so nested files physically appear later in the instruction chain — closer to where the model's attention lands. We document root above nested in our own precedence list for the opposite reason: so a package index can't quietly override a universal invariant just because it loads later. The two aren't in tension, one is about token position in the chain, the other is about which rule should win when a package file and root actually disagree (AGENTS.md Patterns).
Resulting layout
AGENTS.md
CLAUDE.md # @AGENTS.md — Claude Code import
apps/
backend/
AGENTS.md
CLAUDE.md
src/tools/paymentProcessor/
AGENTS.md # heavy domain guide (billing/webhooks)
CLAUDE.md
portal/
AGENTS.md
CLAUDE.md
marketing/
AGENTS.md
CLAUDE.md
shared/
AGENTS.md
CLAUDE.md
tools/
AGENTS.md
CLAUDE.md
documentation/
AGENTS.md
CLAUDE.md
.cursor/
rules/
changelog.mdc # mirror pair
test-isolation.mdc # mirror pair
… # ~20 glob-scoped rules (backend-patterns, portal-patterns, etc.)
.claude/
rules/
changelog.md
test-isolation.md
The CLAUDE.md files aren't a third copy of anything. Each one is a one-line @AGENTS.md import, so Claude Code reads the same canonical file everyone else does without us maintaining a second body of text under a different filename.
Closing principle
Supporting multiple AI coding assistants was never the hard part about their capabilities. It was treating guidance like code: one owner, explicit loading, no hidden overlap. Once every instruction had a single owner and a defined loading strategy, switching tools stopped being an architectural problem in Refract Boilerplate.
The full AGENTS.md guide, including the actual files referenced above, is at docs.userefract.io.
FAQ
- What is AGENTS.md and why use it instead of tool-specific rule files?
AGENTS.md is a plain markdown file that gives AI coding agents project context and instructions, readable by multiple tools instead of one. Tool-specific formats like
.cursor/ruleswork fine in isolation, but the moment more than one tool touches the repo, duplicating the same instructions across formats guarantees drift.- How do you avoid duplicating instructions across Cursor, Claude Code, and Codex?
Give every instruction exactly one canonical location and stop copying it anywhere else. Universal rules go in a root AGENTS.md, package-specific rules go in that package's own AGENTS.md, and only genuinely cross-cutting edge cases get a mirrored pair with a strict requirement-text contract.
- Should AGENTS.md be one file or nested per package in a monorepo?
Nested, scoped to package boundaries rather than every directory. A single root file loses locality once apps diverge in stack or convention, and a file per directory creates too many places to look with no clear ownership. Package boundaries were the smallest unit that kept guidance close to the code without turning the repo into a maze.
- Does Codex actually load nested AGENTS.md files reliably?
We didn't assume it would. We tested it directly: root and nested guides deliberately set to disagree, then Codex run from the repo root and from a package directory, across single-file and multi-file edits, checking what actually loaded. Codex does support hierarchical discovery by design, walking from the project root down to your working directory and concatenating files in order, but it's capped at 32 KiB by default and silently stops adding files past that limit, so verify in your own setup before building an architecture on top of it.
- What's the precedence order when instructions from different tools conflict?
In order: system and developer instructions, then root AGENTS.md, then the nearest nested AGENTS.md, then tool-scoped rules (.cursor/rules or .claude/rules matching the task), then local task or PR notes. Root sits above nested on purpose, so a package-level index can't quietly override a universal invariant just because it happens to load later in the instruction chain.
- How many AGENTS.md files is too many in a monorepo?
More than one full rule body per package is usually a sign you're duplicating rather than scoping — index files that point elsewhere are cheap, duplicated rule bodies aren't. We went from 36 Cursor rule files down to one root AGENTS.md, six package-level indexes, one heavy domain guide for billing and webhooks, and two mirrored pairs for the two scopes that couldn't nest cleanly. The roughly 20 remaining path-scoped Cursor rules stayed exactly where they were; they hold detail specific enough that promoting them to AGENTS.md would have meant re-fighting the locality problem all over again.
- Is AGENTS.md an official standard or just a convention OpenAI started?
It started as a practical Codex mechanism, but it's since moved past single-vendor ownership. In December 2025 it was donated to the Agentic AI Foundation, a directed fund under the Linux Foundation, alongside Anthropic's Model Context Protocol and Block's Goose framework, with governance and Platinum membership spanning AWS, Google, Microsoft, and Cloudflare.