Coada / writing

Stale Beats Broken

Your tests pass, your types check, your guards are green, and none of it describes the system you actually have.

Third post in the series. The first compared three approaches to agentic development and the second spent more time on OpenAI's. This one is about where most of my engineering budget actually goes, which is keeping things from quietly falling out of date.

When something in a system is broken, you find out. Tests go red, the compiler complains, a deploy fails, somebody gets paged at a bad hour. Those failures are annoying but they're the ones our whole industry is built to catch, and we're pretty good at them by now.

Stale is different. A stale test still passes, because it's correctly testing behavior you don't have anymore. A stale type still compiles. A stale runbook still reads like a runbook. A guard that stopped looking at anything still reports green. Nothing about the system looks wrong, and none of it is describing what you actually have running.

That's always been true, and for years it was survivable because a person writing code tends to read a fairly small slice of the repository, and they bring their own suspicion with them. They remember that the doc is old. Agents don't do either of those things. An agent reads much more of your repo than a person would, it takes what it finds at face value, and it produces work faster than anyone on your team can review carefully. Whatever it inherited from a stale artifact gets multiplied into a dozen files before you look up.

I wrote about one slice of this back in August, in Drift Is Not a Bug. The argument there was that spec drift isn't a defect anyone can patch, because in most spec-driven tooling the spec gets consumed at generation and then has no standing to object to anything afterward. I still think that's right. What I want to do here is widen it, because the spec is only one of the artifacts in your repository that can quietly stop being true.

That's why I stopped treating drift as cleanup work. I try to design it out of existence where I can, and where I can't, I try to make it fail loudly. Six places it shows up, and what I do about each of them.

Tests falling out of step with the code

Almost everyone doing spec-driven work has this one. You generate a test suite from a spec, commit the suite, and from that point forward you have two artifacts that only stay in agreement if somebody keeps them in agreement.

They don't stay in agreement. A regenerated suite lands in a large diff and nobody reads it closely, or the spec moves and the suite doesn't, and every test keeps passing because it's still an accurate test of last month's behavior. Signing the file doesn't rescue you either, since all a signature proves is that nobody has touched it since it was committed in whatever condition it was already in.

The worse version of this is a suite that was never derived from the spec at all. Most tests get written after the code exists, by a person or a model reading the implementation, which means the assertions describe what got built rather than what was asked for. That suite closes green forever around a system that stopped matching its specification several releases ago.

I got rid of the file. Generated suites are gitignored and CI compiles them from the current spec on every run, so there's no second artifact to keep current. The only thing a person edits is the spec, which means the only thing that can be wrong is the spec.

That's also where the spec stops being a document and starts being wired to the code. In Feature a spec declares which files its feature touches and which ones it's allowed to read, and that declaration is checked, not decorative. Delete a file that a spec still claims and CI tells you the spec is pointing at something dead. Let a domain file evolve underneath a spec that didn't evolve with it and the spec fails. The two can't wander off from each other quietly, because the spec has named exactly what it depends on and the build keeps checking whether those things still look the way the spec thinks they do.

That only works if generation is deterministic. Same spec and same resolved contracts have to produce a byte-identical file every time, with no timestamps or anything else that varies between runs, otherwise you can't trust a suite you threw away and rebuilt sixty seconds ago. Determinism is also why a model can't be the thing emitting your tests. Run a model twice and you get two suites, so you can never answer the only question that matters at release time, which is whether these tests still represent this specification.

In August I described that check as a byte comparison against the committed suite. We've since gone further and stopped committing the suite at all, for the reasons above. The property I wanted from the comparison, which is that tests and spec can never disagree, now comes from there being nothing to compare.

Specs falling out of step with the model

Every spec I write is a translation of something the domain model already said: the commands, the events, the contracts at each boundary, the branches where a flow refuses to proceed. When the model changes and the spec doesn't, both documents still look completely fine on their own, and there's nothing in either one that hints at the other being ahead.

So we run a diff in both directions. Every event and command a spec names gets checked against the model's vocabulary, and everything in the model gets checked against the specs. Anything that exists on one side with no home on the other fails the build. The first time I ran that on a real project it turned up six events and commands living in specs that the model had never heard of, plus a saga whose state list contradicted itself. I'd read those documents plenty of times before that. Reading them again was never going to find it.

Decisions falling out of step with other decisions

This is the one that cost me the most time this year and the one I almost never see anybody instrument.

You make a call in March. In June you make a different call that replaces it, maybe in a different document, maybe in a different repo, and possibly without ever saying out loud that the March version is dead. Both records are perfectly findable. Neither one knows about the other. Then somebody, or something, goes looking six months later and gets back whichever one happened to match the question better.

It happened to me while I was writing this series. I asked for a summary of my own execution protocol and got back a process I had retired weeks before, because the retirement was written in one place and the original in another. It read as entirely current, and I only caught it because I went hunting for supersession links specifically. Otherwise I'd have kept building against a rule I threw out in August.

What I use now is authored supersession. When a decision replaces another one, the replacement carries a typed link to the thing it killed, written at the moment the call gets made. Nothing is deleted, so you can still go back and see what you believed in March and why you changed your mind, but anything reading the record can tell which version is live.

Authored is the part that matters to me. A model going through your documents looking for things that seem out of date is guessing, and it's guessing about the one artifact you most need to be factual. The next post is entirely about this so I'll leave it alone here.

Contracts falling out of step with the things that use them

If the same fact lives in two places it will eventually live in two versions. A rate constant in a Go service and the TypeScript that meters against it. An event shape a producer emits and the schema its consumer validates. A config value copied across environments by somebody in a hurry.

Nothing fails when these diverge. Both sides keep working, just on different numbers, and you find out through a customer, a reconciliation report, or an invoice that doesn't add up.

Two things help. Shared shapes live in exactly one place and get pulled in by reference, so a change propagates when things compile rather than when somebody remembers. And where the same fact genuinely has to exist in two languages, I write a test whose only job is to fail when the two stop matching. There's one in my stack right now that locks a billing constant in Go to its TypeScript source. It's not elegant. It took five minutes and it has already paid for itself.

I found this class the hard way, in seven places at once, on a system I would have told you was in good shape.

Guards falling out of step with reality

This is the one that bothers me, because it's how the fixes above fail.

A check that runs and finds nothing looks exactly like a check that runs, finds something, and fails to report it. Same green tick, same silence in the log. I had a pipeline stage firing on every run and filing zero results for weeks, because the thing it read had moved and nothing in the design of that stage cared. Every run was green. The stage had stopped doing anything at all.

Two rules came out of that mess. Absence and satisfaction have to look different: a lane that legitimately found nothing has to say so, and a lane that produced no output whatsoever has to fail. And completeness has to mean reachability rather than a list, because a list is an artifact and artifacts go stale like everything else here. Derive the set of things to check from the system itself.

The second rule came out of a bad afternoon. A service had been declared complete twice. When I finally wrote a sweep that enumerated its functions from the deployed stack instead of from a document, 23 of 56 commands turned out to be unreachable. Not broken, unreachable, with nothing pointing at them, and every check we had was quietly working off the same list that was wrong in the first place.

Documentation falling out of step with the system

The instruction file is the artifact agents read most often and the one with the weakest feedback loop in the entire repo.

Everybody has hit this by now. OpenAI wrote about their instruction file turning into a pile of rules agents couldn't sort through. Anthropic recommends keeping theirs under a page for the same reason. The usual answer is to keep it short and have something come through periodically to tidy it up.

Shortening genuinely helps and a cleanup agent helps some too. But the reason this artifact rots faster than the others is that nothing about changing your system invalidates it. Change code and the tests notice. Change a type and the compiler notices. Change a rule and the paragraph describing the old rule sits there being perfectly grammatical and completely wrong.

The signal is already sitting in your repo, though. Every commit invalidates some slice of what you've written down somewhere. Connecting those two things, so that changing the system marks the guidance that covered it as suspect, gets you further than asking a model to guess, and most of what I've built into my memory layer comes down to that.

The part I haven't solved

My drift protection runs downhill. Design flows into specs and I can check that the specs still line up with the model. Carrying a change back up the other way is something I still can't do. Specs evolve during delivery, which is normal and usually a sign the work is going well, and when they do the model that produced them starts falling behind. The simulation keeps passing because it's simulating the model as written. The diagram the business signed off on in March still says what it said in March.

Today I handle that with discipline, which is a nicer way of saying I handle it by remembering, and I know exactly what that's worth given everything above.

The fix is probably a conformance check comparing what the specs assert against what the model describes, failing when they disagree in shape and not just in vocabulary. All the pieces are there. I haven't built it. Until I do, that's the softest part of my stack and anybody weighing up this way of working deserves to know that.

What it actually costs

Not every artifact can be deleted. Not every boundary can be computed. Sometimes duplication is the right answer and a lock test is as good as it gets.

So I make the trade per artifact, and the question is always the same. If this goes stale, what tells me? When the answer turns out to be a person noticing, then it isn't a control, and I either get rid of the artifact or put something in front of it that fails when it drifts.

That's the whole reason generated tests don't get committed, file scope is declared and enforced at build time, constants in two languages have a test sitting between them, checks derive their targets instead of trusting a list, and decisions carry links to the decisions that replaced them. None of it is clever work. It's the same question asked over and over about every durable thing in the system.

Garbage collection is a reasonable strategy when you can't avoid making garbage. I'd rather spend the afternoon it takes to stop making it.

Next

Next one goes deep on the memory layer: why recall and currency are not the same thing, what authored supersession looks like day to day, and why I think it's the piece most of this market is still missing.

If you want to poke at the tooling, Feature and Moment are open source, SiDD is documented in the open, Facet is free, and flmnt is the memory layer. It all fits together into a delivery chain called Complai.