Coada / writing

Marked Isn't Enforced

An agent that sees the path your decisions took can tell which way they're heading.

Most agent memory can find the decision you asked about. Far less of it can tell you whether that decision still stands, and even when it knows, it often hands you the old one anyway.

Fourth post in the series. The last one walked through six places things quietly go out of date, and I held one of them back for this post: decisions falling out of step with other decisions. I wrote about the problem itself back in July (here and here). This one is about the mechanics, where the rest of the field has landed since then, and the gap I think still matters most.

Found isn't the same as true

A quick version of the problem, for anyone who missed July.

In March your team picks MongoDB for a service. In August, after a rough quarter, you move it to PostgreSQL. The March decision lives in one thread. The August reversal lives in another. Neither one mentions the other, because nobody writing the August note thought to go back and annotate March.

Ask an agent what datastore that service uses and a retrieval system will happily find the March decision. It might find August too. Nothing in either text says which one is in force, and if the March note happens to be worded closer to the question, it's the one that ranks first.

Whether a decision still stands isn't something you can read off the decision. It's a fact about its relationship to later decisions, and if nobody recorded that relationship, there's nothing for retrieval to find.

What authoring looks like

In flmnt, when a decision replaces another one, the replacement carries a typed link to the thing it replaced. That link gets written at the moment the call is made, by whoever is making it, usually the agent I'm working with in the same session where we decided.

Nothing gets deleted. The March decision is still there with its reasons, and if you want to know why you once believed MongoDB was right, you can still read it. What changes is that anything reading the record now knows March has been replaced, and by what.

The link also doesn't care which stream the two decisions live in. Product decisions and operational decisions get recorded in different places, and an operational call made in the middle of a deploy regularly retires something decided months earlier in planning. If supersession only worked inside a single stream, it would miss most of the cases that actually bite.

Day to day it mostly looks like nothing, which is the point. An agent goes looking for how something works, and what comes back is the version that stands, with a note saying which older entry it replaced. The agent never has to notice a conflict and guess which side wins. The question was settled by the person who knew the answer, on the day they knew it.

Four places to put the judgment

When I started on this, very little in agent memory addressed it. That's no longer true, and I think the more useful way to read the field now isn't who noticed the problem but where each approach puts the judgment about what's stale.

A model decides at ingestion. Graphiti, the open source engine behind Zep, compares each new fact against existing ones as it comes in and asks a model whether the new one contradicts the old. If it does, the old fact gets an invalidation date. It's a real mechanism and a thoughtful one. The judgment is still a model reading two pieces of text and inferring a relationship, which means it works when the contradiction is visible in the words and has nothing to go on when it isn't.

A rule decides on extracted keys. MemStrata takes a deterministic route. It extracts facts as subject, relation and object, and when a new fact shares the same subject and relation as an old one, the new one replaces it. No model in the judgment, which I like. The catch is that it only works when both facts get extracted onto the same key, and by MemStrata's own count only a minority of real-world fixes arrive as clean one-for-one transitions like that.

The model learns it. Supersede, a research environment published this summer, trains models to prefer the updated fact when they see both. That moves the judgment into the weights. It helps when the update is stated somewhere the model can see it, and it gives you nothing to audit afterward.

A person decides at write. This is where flmnt sits, and it isn't alone. PROJECTMEM also lets you author a supersedes relationship directly, and credit to them for getting there. The judgment comes from whoever made the call, when they made it, recorded as a relation rather than inferred from text.

I land on the last one for the same reason that runs through everything else in this series. I'm comfortable with inference when something is being authored. I'm not comfortable with it at the point where the system decides what's true. Whether my March decision still stands isn't a question I want answered by a model's read of two paragraphs. I already know the answer. I knew it the day I made the August call, and the cheapest, most reliable moment to record it is right then.

Marked isn't enforced

Knowing a decision is stale turns out to be only half of the problem. The other half is what the system does with that knowledge when your agent asks a question.

A study published this month, the Vulcan revocation work, tested exactly that. The researchers flagged records as revoked and then watched what memory systems returned. By default, most of them still handed the flagged records back, often as the top result. In their agent tasks that led to unsafe actions 43.1% of the time. Adding warnings to the prompt only brought it down to 37.2%, because nothing in what the agent received made clear which fact had been replaced. When the old records were filtered out before they ever reached the agent, the unsafe actions disappeared entirely, zero across 1,620 cases.

That's the difference between marking and enforcing. A system that marks stale records knows the old decision was replaced and writes that down somewhere. A system that enforces makes sure the old decision never reaches your agent as an answer. By default, most of the systems in the study were doing the first.

To explain how flmnt enforces it, I need to give a quick picture of how it answers a question in the first place.

flmnt keeps two things. The first is a log of everything an agent records: decisions, explorations, mistakes, checkpoints. Entries are added and never edited. The second is a map of how those entries relate, which one led to which, and which one replaced which. The replacement links from earlier in this post live in that map.

When an agent asks a question, the answer gets assembled in three steps. flmnt first finds the entries whose wording is closest to the question. It then follows the map outward from those entries to pull in the context around them, what led up to a decision and what came after it. Finally it ranks everything it gathered and keeps as much as fits in a fixed amount of space, because an agent can only take in so much at once. Whatever ranks too low to fit gets left out.

The first step is where stale decisions do their damage. Go back to the March and August example. Your question probably uses the same words the March note used, since that's the note that defined the vocabulary everyone still uses. The August reversal was written later, by someone solving a different problem, in different words. If August had to earn its place by how closely it matched the question, it could easily rank too low to make the cut, and your agent would get March and nothing else.

So flmnt doesn't make it compete. When an entry that matched the question has been replaced, the replacement takes that entry's place in the ranking, at the same position and with the same claim on space. March is left out of the answer. The agent is told which entry was replaced, so it can still ask for the old decision if it wants the history, but the answer it gets is August.

Sometimes the replacement has been replaced as well. Maybe August was overturned again in November. In that case flmnt keeps following replacement links until it reaches a decision that nothing has replaced, and that's the one that takes the place. Stopping at August would only swap one stale answer for a slightly newer stale answer.

There's one kind of request where flmnt deliberately shows both. When an agent asks for recent history rather than an answer, it gets a record in the order things happened, and quietly removing entries from a history would misrepresent what happened. There the old decision appears with a short label saying which entry replaced it. The label only states that fact. It doesn't tell the model to ignore the entry, because once you put instructions in the context, whether the agent gets the right answer depends on whether it follows instructions, and I'd rather the answer depend on the record.

From history to direction

That label only exists because the replacement is a recorded relation rather than a guess. Everywhere else in agent memory the judgment about what's stale is still made by a model reading text, and of the approaches above, PROJECTMEM is the only other one I've found that records supersession as something a person authors. That difference in where the judgment sits also opens up something I haven't seen any memory system set out to offer.

Remember the three steps. When flmnt follows the map outward from a matched decision, it walks in both directions. It pulls in what led up to the decision and what came after it, including every decision that replaced another along the way. So an agent doesn't only receive the fact that stands today. It receives the path that got there, in order, with the reasons each step was taken.

A single fact tells you where you are, while a path also tells you which way you've been heading.

To make that concrete, take an illustrative example. Over a year, a team moves one service from a self-hosted database to a managed one, and months later replaces a cache it ran itself with a managed offering. The two decisions were made separately, by different people, for reasons written down at the time, and both cite the operational burden of running infrastructure themselves. Ask an agent that only sees current facts whether to self-host a new search index and it has nothing to push back with. An agent that sees the path can tell the proposal runs against a year of direction, point to the two decisions that set it, and suggest the managed option before anyone asks.

Handing the agent history in both directions was part of the original idea behind flmnt, and what agents do with it is still the part that impresses me most. I've watched an agent working with flmnt read a run of past decisions, name the direction they were moving in, and propose the next decision in that line before I'd gotten there myself. It was right, and more importantly it could show me which decisions it was reasoning from, so I could check it rather than take it on trust.

I want to be careful about what I'm claiming. I've observed this, repeatedly, in my own work. I haven't measured it the way I measured the supersession results below, and I'd treat it as an estimate the agent can show its work for, not a prediction to act on blindly. But it only works because the history is there, in order, with the replacements recorded. A memory that keeps only the latest version of each fact has already thrown away the thing that makes it possible.

What the evidence says

I built a test for this specifically, and I want to be clear that it's our own design.

Two decisions, the March and August case above. Neither text mentions the other, and neither uses words like "replaced" or "deprecated". The stale one is worded close to the question and the current one is worded differently, so a similarity search will reliably prefer the stale one. The only thing connecting them is the authored link.

A vector-only retriever served the stale decision every single time. Following the link served the current one every time. Those runs are published with the raw data in the open benchmark repository, so anyone can check the numbers or pick holes in the design.

The result I find more convincing comes from somewhere the graph didn't help. On LOCOMO, a public benchmark of long everyday conversations, my graph-based retrieval did no better than plain vector search, and slightly worse on the headline number. Those conversations are full of facts and almost empty of authored decisions, so there's nothing for the graph to follow. The graph structure on its own bought nothing there, which is what I'd expect. The advantage lives in the relations someone took the trouble to record.

There's also a check that involves no model at all. Take the same two decisions, run them through retrieval with the link and without it. With the link, the stale decision is replaced and marked. Remove the link and nothing is marked, even though the text is identical. Whatever the system does with these two records comes from the relation, not from anything in the wording.

What it costs

Authoring isn't free. Somebody has to say, at the moment of the decision, what it replaces. In practice that's cheap when the tooling is right there in the session, and agents do it without being asked once the habit is part of how they record decisions. It's still a discipline, and a team that doesn't keep it won't get the benefit.

And a reversal nobody records stays invisible. If you change your mind in a hallway and never write it down, no memory system will know, including mine. Authored supersession doesn't read minds. It makes sure that when you do make the call, the call is what your agents get.

I'm fine with both of those. The alternative is asking a model to work out, every time an agent asks a question, something a person already decided and could have written down in a sentence.

Where this fits

Every post in this series has ended up at the same question from a different direction. The first two were about who writes the checks your agents are graded against. The third was about artifacts that quietly stop describing your system. This one is about decisions, which are the artifact agents lean on hardest and the one most memory systems still treat as permanent.

If you want to see it working, flmnt is the memory layer. Feature and Moment are open source, SiDD is documented in the open, and Facet is free. Together they make up a delivery chain called Complai.