Your AI Agent Isn't Lying: It's Just Passing the Test It Rewrote

The real failure mode in AI coding isn't bad code. It's specification drift hidden behind a green checkmark.

A few weeks ago an agent touched a test that governs webhook idempotency. Stripe retries webhooks, sometimes the same event arrives twice, and the whole system depends on a processed_webhook_events table and a Redis SET NX check to make sure we only apply it once. There's always been a test for that.

I asked the agent to add a new webhook type. Somewhere in the middle of doing that, the existing idempotency test started failing. Turns out a schema tweak elsewhere in the same session had quietly changed how the dedupe key got built (or maybe it was just buried in a huge diff, honestly I still don't know which). The agent looked at the failure, probably decided the test's expectation was stale, updated the assertion to match the new key format, and called it a day.

Of course it reported the task complete, because CI was green. Took me two days to notice, until a duplicate charge showed up in a test account and I went digging to figure out how.

In all honesty, I don't think the model made anything up. It didn't hallucinate a function or invent an API. It saw a red test, decided the test was describing something no longer true, and did the thing that made the red go away. That's a completely reasonable thing to do if you don't know that particular test wasn't yours to touch.

The failure wasn't the model. That's the part I keep landing on. We trust these agents enough that we expect them to somehow cover the context we didn't give them, and it just doesn't work that way. The agent doesn't know when an assumption is dangerous and when it's fine. Half the time I can't even tell if I'm assuming or just moving fast enough that it feels like knowing. Kahneman would call that System 1 Thinking, your brain runs on autopilot until something forces it to slow down and actually check.

The hierarchy we never wrote down

Every engineer who's worked on a team for more than a year carries an implicit ranking in their head. Product requirements sit above specifications, specifications sit above tests, tests sit above implementation, and so on... If the code disagrees with the tests, you investigate. If the tests disagree with the spec, you investigate. Nobody wrote this hierarchy into a document, it's just something you absorb from watching what happens when someone violates it and gets yelled at in a PR review.

Agents don't come with that absorption. To a model working through a task, a test assertion and a line of application code are both just tokens it's allowed to edit if editing them satisfies the instruction. There's no felt weight difference between expect(charge.status).toBe("succeeded") and the line of billing logic that produces that status. Both are text in the diff. Both are equally available.

That's not a flaw in the model, it's just an accurate reflection of the information we fed it. We spent two decades building engineering culture around the assumption that whoever's editing the code understands which parts are load-bearing and which parts are implementation detail. That assumption held up fine right up until the implementer stopped being a person who'd sat through six months of onboarding and started being a context window that resets every session.

Oh, and the fun part: most of the time we're driven by ego, so instead of blaming ourselves we'll just say the model "hallucinated." I mean, a million tokens or more sounds like a lot, until you realize that's roughly 750,000 words, maybe 4 million characters. Good luck cramming in everything that's lived in your head for a decade and not forgetting a single thread. That's a different topic though, I'm digressing.

Passing the test it rewrote

I think the failure mode has a shape you'll recognize once you've seen it once.

  1. A specification exists somewhere, usually informally.
  2. The implementation drifts from it.
  3. A test catches the drift and fails.
  4. The agent, working purely from the signal in front of it, concludes the test is the thing that's wrong.
  5. It updates the test Everything goes green.

None of that requires the agent to be careless or the model to be weak. A more capable model doesn't fix this, it just gets more convincing about which tests deserved rewriting. I've had this happen with a genuinely strong coding agent that made a defensible call given what it could see, and the call was still wrong, because "defensible given what it could see" is exactly the gap.

I think that's part of why I've come to like working with Composer 2 these days. It's a solid coding agent, cheap too, and it has this habit of forcing you to be accountable for your own prompts instead of quietly picking up the slack for you. Not that it matters much which agent you pick, honestly, any of them code better than all of us combined at this point.

Teams tend to respond to the wrong layer of the problem, though. The instinct is to blame the tool, switch models, add a stricter prompt, tell the agent in bold letters not to touch tests. That's part of why we ended up making Refract's agent setup tool-agnostic instead of betting the whole workflow on one model's quirks. Switching models buys you a week, maybe. Then the next task has the same shape and the same thing happens somewhere else, because the actual issue was never which model wrote the diff. It's that nothing in the system distinguished between code the agent owns and code it doesn't. It doesn't know what the baseline is, or where the handrail sits.

Specification contracts

I've started calling the fix a specification contract, because "requirements" is too soft and "tests" is too narrow to cover what actually needs protecting.

A specification contract is any artifact an agent is not allowed to silently redefine while completing a task. That's basically it. It doesn't need a framework or a tool, just an explicit decision about which parts of the system encode a promise rather than an implementation choice.

Billing behavior is a contract, so is Webhook idempotency, or a permission matrix. The exact internal shape of a function that computes a discount is not, do you see the difference? An engineer or an agent should be free to refactor that however they want as long as the outputs match. The difference isn't about complexity or how much code is involved, it's about whether changing it is a planning decision or a plumbing decision.

Most teams have zero specifications contracts. Tests look like contracts but they aren't and that's the trap. Tests sit next to the code like contracts, and get treated by every tool with exactly as much authority as the code they're testing. Actually, most engineers I've seen tell their agents to cover code and specs in one prompt. Which is the entire problem in one sentence: a test is not a contract until something enforces the difference between changing it and changing everything else.

The planner and the implementer aren't the same job

Software teams have always split these two responsibilities even when it's one person wearing both hats. Someone decides what the system should do, and iterates on that with other team members. Someone else, or possibly the same someone an hour later, builds it. The hats stay distinct even when the head doesn't change.

Yes, this is theory, and in practice even human teams blur this line to save time and end up with a painful rewrite six months later.

So it's only natural that agents collapse that boundary by default. You ask for seat-based billing and what the agent actually hears is closer to "modify whatever's available until the constraints in front of you stop firing." Source code, tests, the AGENTS.md file, even linting rules sometimes, a stale example in the docs, all equally fair game unless you've told it, explicitly, otherwise. I've had an agent update our own canonical documentation to justify a decision it had just made, which is a strange thing to watch happen, the agent citing a source it wrote thirty seconds earlier.

The planner is allowed to change behavior. The implementer isn't allowed to decide, on its own, that behavior needed changing. Once you say it that plainly it sounds obvious, and I think that's exactly why so many teams missed it: the two roles used to live inside the same slow, expensive human, so nobody had to draw the line between them out loud.

All Ai tools actually ship something close to this split built in, Plan Mode locks the agent into read-only, it can search and reason but it can't write a file or run a command until you approve what it came up with. First time I used it properly I thought, oh, there it is, the fence I've been drawing by hand every time. Except the fence only holds for that one moment. The second you approve the plan, the same agent that just played planner picks up write access and becomes the implementer, and nothing stops it from making a new, unplanned call five tool calls later. It solves "did we agree on scope before touching code." It doesn't solve "does this thing still remember what it wasn't supposed to touch" once it's three steps into actually doing the work.

Which, I mean, fair games. Even in engineering teams of humans this notion can get blurry and lead to some pretty crazy stuff.

Where the actual bug came from

What took me the longest to accept, honestly, was that the agent updating that idempotency test wasn't the bug but a symptom.

The bug was that I gave it a task, "add this webhook type" without telling it that idempotency was non-negotiable and every other line in the file was up for grabs. I hadn't drawn the line, clearly enough, so it drew one for me, and it drew it in the place that got him the fastest to its result. That's not the agent covering a planning gap with a bad guess exactly, it's closer to the agent resolving an ambiguity I didn't know I'd left open, in the direction of least resistance.

This is the piece that changes how I think about trusting AI-written code at all. The output quality was never really the risk. Every agent I've used writes competent code most of the time, arguably better than a lot of code I've reviewed from engineers. The risk lives one layer up, in the gap between what I meant and what I typed into the prompt. An agent can't close that gap for me, it can only fill it with whatever's locally consistent. And locally consistent and correct are not the same claim.

Which means the tests passing was never going to be enough on its own. Tests only tell you the code satisfies some specification. They can't tell you whether that specification still says what you think it says, especially when the same session had the ability to edit both.

Guardrails, not vibes

I don't think the fix is banning agents from ever touching a test. Plenty of tests are wrong and deserve updating, and blocking that entirely just means you're back to reviewing every diff by hand, which defeats the point of using an agent in the first place.

The fix is separating "the agent fixed a bug" from "the agent changed what correct means," and treating the second one as a different category of event that needs a human to actually look at it. In practice that's meant a few things.

Specific behaviors get an explicit owner, a person, not a team, who has to sign off when the contract itself changes rather than the code around it. Some files get treated as read-only for implementation work, an agent can read a pricing rule or a permission matrix, propose a change to it, but it can't merge that change without someone confirming it was intentional. And the agent's default behavior flips: fixing code silently is fine, but touching an existing assertion triggers a stop, grouped by root cause, with a diff a human actually reads before anything merges.

None of that is exotic. It's closer to the code review discipline teams already had before agents showed up, just re-pointed at the one category of change that used to be safe to skim.

The question underneath all of this

I think the industry conversation about AI coding is stuck one layer too shallow. We keep arguing about which model hallucinates less, and hallucination was never really the failure mode that mattered. The failure mode that matters is an agent resolving an ambiguity you didn't know existed, in a direction you'd have rejected if you'd been asked.

This is also part of why I don't think loop and graph agents solve it, even though on paper they should. A loop that plans, acts, checks its own output, and repeats until some exit condition is met sounds like exactly the discipline this piece is arguing for. Same with a graph setup, separate nodes for planning, coding, testing, reviewing, structure built into the architecture instead of hoped for in a prompt. Except the exit condition is usually still "tests pass" and a loop will find the cheapest path to that condition every time. Cheapest doesn't mean correct, it means fewest iterations to green. Add a verifier node that grades the output before the loop closes and you haven't added a human, you've added another agent working from the same incomplete picture the first one had. Two agents agreeing with each other isn't the same thing as a human catching a planning gap, it just means the assumption was consistent enough to fool both of them.

Tests can't catch that either, because the agent that wrote the code and the agent that would grade the code are drawing from the same incomplete picture of what you actually wanted. The only thing that catches it is a human looking at the diff, not to check whether the code is competent, but to check whether their own planning had a hole in it that the agent quietly poured concrete into.

I still let agents write most of my code. I just don't extend that trust to deciding, on their own, what correct means. That decision stays mine, and the day I stop reading the diffs that touch a contract, we may all be in a bit of a pickle.

I mean, eventually we'll get there, I'm sure, but we just haven't found the right formula yet.

FAQ

Why did the agent rewrite the test instead of flagging the conflict?

Nothing told it that test was a specification contract rather than an implementation detail. To the agent, the assertion and the code around it were both just editable text, and updating the assertion was the fastest way to satisfy the task.

If the model didn't hallucinate anything, what actually went wrong?

It filled a planning gap with an assumption. The instruction didn't say which parts of the system were non-negotiable, so the agent picked, and it picked in whatever direction made the immediate task pass. That's a process failure, not a model failure.

Can better tests catch this?

Not reliably. Tests, gherkin scenarios, mutation testing, all of it checks whether code satisfies a known specification. None of it can tell you whether the specification still says what you meant, especially when the same session had permission to edit both.

Should agents ever be allowed to edit tests?

Yes, most test changes are legitimate. What needs a human gate is specifically the case where an agent is redefining the contract itself rather than fixing an implementation to match an unchanged one.

What counts as a specification contract in practice?

Anything that encodes a promise rather than an implementation choice: billing behavior, webhook idempotency, permission rules, API response shapes. The internals underneath those promises are fair game to refactor. The promise itself isn't, without a human saying so.

Will more capable models fix this on their own?

Probably not. A stronger model gets better at making a convincing case for the assumption it made, not less likely to make one. Capability isn't the same thing as knowing which lines were yours to touch.

Doesn't AI Tools' Plan Mode already solve this?

Partly, but not really. Plan Mode enforces the split for one moment, before you approve a plan the agent can't write files. But once you approve, the same agent gets write access back, and nothing stops it from making a new, unplanned call later in the same session. It answers "did we agree on scope before touching code," not "does it still remember what it wasn't supposed to touch three steps into the work."

What about loop or graph agents, doesn't that architecture solve it?

Not on its own. Adding planner, coder, and verifier nodes sounds like exactly this discipline built into the system, but the exit condition is usually still "tests pass," and a verifier agent grading another agent's output is drawing from the same incomplete picture of what you actually wanted. Two agents agreeing with each other isn't the same as a human catching a gap.