What the Harness Can't Prove
Harness engineering, reviewed: where agent-to-agent verification breaks, why drift should be impossible, and how architectural coherence survives.
This is the second post in a series. The first one looked at three approaches to agentic development from 10,000 feet: OpenAI's harness engineering, Anthropic's AI-native playbook, and the way I've been building for the past couple of years. Here I want to spend more time on OpenAI's.
Harness engineering: leveraging Codex in an agent-first world is probably the most useful thing I've read on this subject. Five months, roughly a million lines of code, nobody on the team writing any of it by hand, and then they went and published what it actually cost them to get there. Plenty of companies would have published the million-line number and stopped.
So this isn't a takedown. I've been building a delivery stack for agents in regulated industries for about two years now, after 15 or so years of building software and running the teams that build it, and I agree with almost everything they diagnosed. There are three places where I'd do it differently, and all three come down to the same thing: they optimized for throughput, and I can't.
The line I keep coming back to
The best part of their post has nothing to do with tooling. It's how their engineers reacted when Codex got something wrong. They didn't go rewrite the prompt. They asked what capability was missing from the environment, and how to make that capability legible and enforceable for the agent.
I've been asking some version of that for years. The only thing I'd add is a question about the second half of it. Enforceable by what, exactly?
The rule I work from is that a boundary has to be enforced inside the system it governs. Guidance in a markdown file is not a boundary, no matter how well written it is, because nothing happens when it's ignored. That's the thread running through all three of the disagreements below.
The reviewer problem
OpenAI has handed most of code review to agents, and their agents merge their own pull requests. Human attention was the thing they couldn't scale, so they moved review onto the thing they could.
The economics make sense to me. I still think it's the softest part of the whole setup, and I think it's the part they'll be paying for a year from now without knowing where the bill came from.
An agent reviewing another agent's work isn't really a second opinion. Same training distribution, same instincts about what good code looks like, same blind spots. When the author and the reviewer see the world the same way, agreement costs almost nothing, and what you get back looks a lot like verification without doing any of the work. The review passes, the PR merges, and nothing about that code was ever actually tested against something that could have disagreed.
I've been on the wrong side of this, so I'm not speaking hypothetically. Earlier this year I went looking into my own design simulator and found that its assertions couldn't fail. The walker was pulling the expected path out of the model, then checking it against that same expected path. Everything was green. None of it meant anything. We only caught it because we started mutating inputs deliberately to see whether the checks would break, and they didn't. Sixty-four of those mutations run in the suite now. If I can't make a check fail on purpose, I don't count it as a check anymore.
That's what I'd push back on here. It's not about whether the reviewing agent is any good. It's whether anything in the loop is capable of telling you no.
In my stack the agent never writes the assertions it gets graded against. A person writes the predictions for a feature, meaning the response, the events emitted, the data written, the messages published, and those predictions compile into the test suite. The agent has all the freedom it wants in how it implements. It has none at all in what counts as correct. If it produces something nobody predicted, that's a violation. If something predicted doesn't show up, also a violation.
This is the part where I'm out of step with most of the market, and I'm fine with that. A model grading a model will eventually agree with itself, and when it does you get a green check over broken behavior and no indication anything went wrong. If you work somewhere that mistakes are cheap, you can absorb that. I don't.
Cleaning up versus not making the mess
Their section on entropy is honest in a way I appreciated. Agents copy whatever's already in the repo, so a bad pattern spreads exactly as fast as a good one. They used to burn Fridays cleaning it up. Now they encode principles and run background agents that look for deviations and open cleanup PRs, which they compare to garbage collection.
Good metaphor, and that's sort of my issue with it. Garbage collection presumes garbage. You run it forever because whatever is making the mess is still making the mess.
Take generated tests, which is the clearest example I have. Most spec-driven setups generate the test suite once and commit it. That committed file is somewhere a stale artifact can sit quietly. It can fall out of step with the spec it came from, and nobody reviewing a large diff is going to notice. A signature check won't help either, since all it proves is that the file hasn't changed since it was committed in whatever state it was already in.
So I don't commit them at all. CI compiles the suite from the current spec on every run. There's no file to go stale, which means there's nothing to detect, nothing to clean up, and no background agent needed to keep two things in sync that are now one thing.
Scope works the same way. Rather than telling an agent to respect module boundaries and hoping for the best, each spec declares which files it's allowed to touch, and the build fails if anything else moved. Cheap to implement, and it takes a category of drift off the table instead of scheduling it for later.
I'll admit prevention isn't always the cheaper option. In these two cases it's so much cheaper that running a permanent cleanup crew seems like a strange trade.
They found the rot and then treated the symptom
The knowledge section was the one that had me nodding. They tried the single big instruction file and it fell apart on them. The way they describe it, the thing accumulates outdated rules until agents can't tell which ones still apply. Their fix is a structured docs directory plus a recurring agent that goes hunting for stale documentation and opens fix-up PRs.
They landed on the exact problem I've spent most of the last year on. A stale test fails. A stale type won't compile. A stale instruction just produces a very confident agent and no warning whatsoever, and that asymmetry is why these systems drift in ways that surprise the people running them.
But a gardening agent is a model guessing at which of your own rules are dead. That's inference pointed at the one artifact you need to be factual, which puts it in the same family as the reviewer problem.
What I ended up building is what I call decision currency. Decisions get authored, and when one replaces another, the new one supersedes the old with a typed link between them. Nothing gets deleted, so the history is all still there, but anything reading that record can tell what's current and what got replaced, even when the replacement was written somewhere else entirely, weeks later, by somebody else.
Here's the part that doesn't flatter me. While I was putting this series together, I went back through my own records and found an execution protocol I'd retired weeks earlier still being handed to me as if it were current, because the retirement lived in one place and the original lived in another. The typed links are what let me find it and clean it up in an afternoon. Without them I'd have kept building on a rule I'd already thrown out, and I probably wouldn't have noticed for months.
Both approaches admit that knowledge rots. The difference is whether a model has to notice or the record just tells you.
Their open question
At the end of the post they say they don't know yet how architectural coherence holds up over a few years when agents wrote most of the code. I respect them for putting that in writing. I also think the answer is hiding in how the question is asked.
Coherence isn't something a codebase keeps. It's something a design either had or didn't. If the domain model never converged, if the people who run the business never signed off on it, and if nobody ever tested the model itself, then coherence was always going to depend on how lucky you got with the agents.
Which is why that work happens in front of the code in my process. The domain gets modeled with the business in the room, then driven to convergence by working through every question the model can't answer yet, one at a time, until there aren't any left. Then the design gets tested. Not the implementation, the design. Do events land in the right contexts in the right order, do the contracts at every boundary hold, do the long-running processes actually move through the states they say they do, do the reactions fire.
By the time an agent opens a spec, the shape of the system has already been proven. Planning runs around 70% of my investment on a typical project and building plus validation is the other 30%, which sounds ridiculous until you see what happens on the back half. The agent moves fast because every boundary it could run into was measured before it started.
Where they're right
I don't think their tradeoffs are sloppy. For their situation I think they're correct, and I'd be wrong to drag mine into it.
The merge philosophy comes down to one line in their post: corrections are cheap, and waiting is expensive. For an internal product with tight feedback and reversible mistakes, that's just true. Blocking every merge behind a full gate would cost them more in throughput than the defects cost them in damage. They also say outright that this would be irresponsible in a lower-throughput environment, and that caveat is carrying more weight than people give it credit for.
My gates block because my mistakes aren't cheap. A defect in a healthcare platform is an incident, an audit finding, or something a patient runs into. A defect in a financial intelligence platform is a number somebody trades on. At that price, waiting is the cheaper purchase.
If I were running internal tooling with same-day feedback and no compliance surface, I'd loosen up too. Probably not to the point of agents merging their own work, but I'd stop paying for proof nobody was ever going to ask me for.
What I'd change if it were mine
Three things, and none of them cost throughput.
Make the check independent of whoever wrote the code. Keep the agents writing everything. Just move acceptance criteria into something a person authors and a compiler enforces, so the thing deciding "correct" wasn't produced by the thing being judged. That's human time once per feature, and it takes the agreeable-review problem off the board.
Get rid of the artifacts that can rot. Anything generated should be compiled at run time instead of committed, and every task should declare what it's allowed to touch so the build can fail when something else moves.
Give the knowledge base a real notion of supersession. Keep the docs directory, it's a good idea. Stop asking an agent to guess what's stale. When a decision replaces another one, write that link down so anybody reading it later, person or agent, doesn't have to infer anything.
How we'd know who's right
Nobody in this conversation is saying what would change their mind, so I'll go first.
I'm wrong if human-authored, compiled predictions let bad behavior ship at about the same rate as agent-reviewed code does. That would mean my rigor is buying confidence rather than correctness and that most of my planning is ceremony. It's measurable too. Run the same feature set through both and count what escapes to production, weighted by how bad it was.
I'm also wrong if agent reviewers turn out to catch things predictions structurally can't. I can predict that an event fires with a certain shape. I can't predict that an API is pleasant to use, or that a name will make sense to whoever inherits it, or that an interaction feels right. If that category turns out to drive most real defect cost, then the reviewing agent is doing work my compiler can't, and the answer is both instead of one.
They're wrong if the agreeable-reviewer thing shows up later as a defect class nobody caught in the moment. Architecture eroding slowly, contracts between services breaking quietly, behavior drifting away from intent with every individual review looking fine. That's the same coherence question they raised about themselves, and it's the one I think their setup is least able to answer.
Neither of us has run that experiment. I've got results I trust in regulated delivery and they've got a million lines and a shipped product, and those aren't the same claim. I'd rather say that than pretend either one settles it.
Next
Next post goes deep on drift, and why I'd rather delete a file than schedule a cleanup for it.
If you want to poke at any of the tooling, Feature and Moment are open source, SiDD is documented in the open, Facet is free, and flmnt is where decision currency lives. It all fits together into a delivery chain called Complai.
And if you work on Codex or on the harness itself, I'd like to compare notes. I suspect we're closer than this post makes it sound.
Get new posts in your inbox. No spam, unsubscribe anytime.