Coada / writing

Harnesses, Playbooks, and Proof

Two of the companies building the models have now told us how they think software should get built with agents.

In February, OpenAI published Harness engineering: leveraging Codex in an agent-first world. Ryan Lopopolo and his team spent five months building a product where no human wrote a line of code, and shared what it took. Six months later, in August, Anthropic published The AI-Native SDLC playbook, a stage-by-stage guide to rebuilding the whole development lifecycle around agents, from planning all the way through maintenance.

Different companies, different months, same purpose. Both are about making agentic development faster and more accurate.

That's been my purpose too. I've spent over 15 years building software and leading the teams that build it, and the last two years building a delivery stack for AI agents. When I read both posts, I recognized conclusions I had already written into my own methodology, in my own words. The tools that carry that methodology are recent, and their history is on record in my decision streams, my documentation, and my git history. The discipline behind them came from years of watching where software actually breaks.

We're all moving along the same path. We're just not following in each other's footsteps. This post is the view from 10,000 feet, and each section below gets its own deep dive in the series that follows.

Where all three of us agree

OpenAI's early progress was slow. The agent was capable. What it lacked was an environment that told it enough. When something failed, their engineers asked "what capability is missing, and how do we make it both legible and enforceable" for the agent.

Anthropic's playbook starts from a related observation. Code is no longer the bottleneck. The stages around it still move at human speed, and the old controls weren't built for agents writing most of the diff.

I agree with both. My own numbers tell the same story. Without structure, I was correcting roughly 40% of what an agent produced. With the full methodology in place, that dropped to around 5%. Same models. Better inputs.

We also agree on a lot of the mechanics. Knowledge belongs where the agent can read it. Each stage should leave behind an artifact the next stage can use. Guidance is helpful, but anything that must always hold needs a deterministic check behind it. And the agent that wrote the code should never be the one that approves it.

So this isn't a rebuttal of either post. It's a look at where our paths split, and why.

Where the paths split

OpenAI was clear about what they were optimizing for. They wanted engineering velocity by orders of magnitude, and the scarce resource they were protecting was human time and attention. Anthropic's playbook is aimed at the stages that still run at human speed, so that the whole lifecycle can keep up with the build.

Both are about throughput, and both do it well.

My goal is different. I build domain tooling so agents can work accurately and quickly, sure. But the real reason is that what we deliver has to be measurable, provable, and still true to the vision that was agreed on before any engineering started.

That changes who the process is for. Business decides what matters. Product defines the outcomes. Design shapes the experience. Engineering builds it. Testing proves it. Data depends on the shapes that move through it. Sales has to explain it to someone who writes a check. When the work is agent-driven, every one of those groups still needs to trust that what shipped is what they signed off on.

A lot of my work also lives in regulated industries like healthcare and financial services. That's a factual, ledgered world. Everything has to be proven. Most of the differences below come from those two facts.

Where the work starts

OpenAI's story begins with an empty git repository. Anthropic's begins with an intent.md, an idea written up with Claude and then turned into requirements and a design in a single working session.

Mine begins with an empty domain, and with the people who run the business in the room.

We start with domain-driven design and event storming, run as narratives that walk through what happens and when. From there I use Signal-Driven Development, or SiDD, an open methodology I wrote to help architects bring a domain model to convergence. SiDD takes a candidate model, produces a report of everything the model can't answer yet, resolves those gaps one at a time, and repeats until nothing is left. Across nine products, that process has resolved 275 gaps before a single line of code was written.

Then I test the design itself. Moment is an open source language I built for describing a domain model. Most modeling tools capture structure: the contexts, aggregates, commands, and events. Moment also captures time, meaning how events move from one part of the business to another and in what order.

Facet is a free visual simulator for Moment models. It plays the flows through the design and checks them as they go. Do events reach the contexts they should, in the right order? Do the contracts at each boundary hold? Do long-running processes move through their states the way they claim to? Do the reactions the model promises actually fire? Because Facet is visual, the people who approved the design can watch it run instead of reading about it.

OpenAI closes their post by saying they don't yet know how architectural coherence holds up over years in an agent-generated system. My answer is that coherence can't be something you hope survives. It has to be designed, agreed on, and tested before the agents start typing.

What a specification is

At OpenAI, humans prioritize work, turn user feedback into acceptance criteria, and validate outcomes. In Anthropic's playbook, the specification is a spec.md and a plan.md, readable markdown that both people and agents work from, and the agent generates the tests alongside the code.

This is where I hold my hardest opinion, and I'll be direct about it. An agent should never write the tests for its own work. It will agree with itself every time. That's the biggest flaw I see in spec-driven development in the AI space right now, and it's the one I've worked hardest to remove.

My specifications are written in Feature, an open source specification language I built for agent-driven delivery. A Feature spec has two parts, and they're treated very differently.

The first part is instruction. It's written in plain language and tells the agent what to build: where the handler lives, what it's allowed to touch, what constraints it has to respect. This part guides the agent, and the agent has room to decide how to meet it.

The second part is measurement. It defines the contracts and predicts every observable effect the feature will have: the response, the events, the data written, the messages sent. None of that is inferred. A person authors those predictions, a person agrees to them, and the compiler turns them into a complete test suite. The agent building the feature never writes a single assertion it will be graded against.

And here's the best part, if you've read this far. Those measurements aren't invented at the spec layer. They map directly to what the Moment model already produced and Facet already tested. The contracts at each boundary become the input shapes. The events a flow emits become the predicted events. The branches where a flow refuses become the rejection cases. The points where a flow ends become the final state the spec expects. The design that was agreed on and proven is the same design the tests measure.

The rule underneath it is simple. The spec declares everything that should happen. Anything that happens and wasn't predicted is a violation. Anything predicted that didn't happen is a violation too.

The scenarios come straight out of the narratives the business approved, so what stakeholders signed off on during design is exactly what the pipeline measures during delivery.

I won't sugarcoat the cost. This is the part of the process that demands the most human investment. Deciding exactly what a feature should produce, down to the last write, is real work, and it can't be handed off. If a product team lets its discipline slip here, the agent's output will slip right along with it. The quality of what gets built is capped by the quality of what gets predicted.

Who decides the work is right

This is the sharpest difference, so I want to be fair about it.

OpenAI has pushed almost all review to agents reviewing agents, and agents often merge their own pull requests. Anthropic's playbook uses layers of agent review, continuous evals, and in its most autonomous stage, an adversarial reviewing agent that can decide whether work moves forward. It still keeps human review for regulated and critical code, and a human has to authorize anything that reaches production.

None of that is wrong. For a lot of teams, it's exactly the right call. It just doesn't work for me.

In a ledgered world, inference isn't a luxury I get to spend at the point where something is judged correct. It isn't a trust I can place in any model, no matter how capable it is. So in my stack, inference is removed everywhere we can remove it. A model can help shape the instructions in a specification, because those instructions translate a design that was already tested. A model never writes the measurements, and it never decides whether the result passes. The compiled tests do, and the ledger records it.

Having to see things that way pushed me down different paths than the ones these posts describe. The upside surprised me. The feedback loop got tighter, and the product that came out the other side got better.

Keeping things from drifting

OpenAI is candid about drift. Agents copy whatever patterns already exist in the repo. Their answer is to encode principles and run background agents that find deviations and open cleanup pull requests, which they compare to garbage collection. Anthropic's answer is to keep a short CLAUDE.md that the team updates whenever the agent repeats a mistake, back important policies with hooks, and keep the plan in sync with the code.

I'd rather not produce the garbage in the first place.

In my stack, generated tests are never committed. The pipeline compiles them from the current spec on every single run. A committed test suite is a place where a stale file can hide, and where a reviewer or a signature check can miss it. Take the file out of the repository and there's nothing left to go stale.

That idea goes past tests. Every boundary has to be enforced inside the system it protects. A rule that lives in a document and nowhere else isn't a boundary. It's a suggestion.

Memory that knows what's no longer true

OpenAI tried the one giant instruction file and watched it rot. They describe it filling up with outdated rules that agents could no longer tell apart from current ones. Their fix is a structured docs directory plus a recurring agent that looks for stale documentation. Anthropic keeps CLAUDE.md under a page and treats the chain of committed artifacts as the record.

Both named the right problem. Staleness is the whole problem. A stale test fails. A stale type throws an error. A stale instruction file produces a confident agent with no warning at all.

My answer is flmnt, a memory layer for agents that took me months and hundreds of revisions to get right. In flmnt, decisions aren't just stored. They're authored, and when a decision changes, the new one explicitly supersedes the old one with a typed link. History is kept. Nothing is deleted. But an agent reading that memory can tell which decision is current and which one was replaced, even when the replacement was recorded somewhere else entirely.

I call this decision currency. Recall asks what was said. Currency asks what's still true. It's the piece I'm proudest of, because it's what keeps the original vision intact while the details keep changing underneath it.

Gates and proof

OpenAI runs with minimal blocking merge gates, and flaky tests often get rerun instead of blocking progress. When agents produce more than humans can watch, fixing a mistake later costs them less than holding work in a queue. Anthropic's playbook uses hooks as approval gates, and the agent can act all the way up to the production gate but never past it.

In regulated systems, a correction isn't cheap. It's an incident, an audit finding, or a bug a patient sees. So my gates block. A red build is a blocker. There are no skip flags.

Every pipeline run also produces an evidence bundle tied to the exact commit it came from. When Feature is wired into CI, those results are signed and written to an append-only ledger in the Feature dashboard, tracked per specification and per environment. Nothing in that ledger can be edited after the fact. An auditor doesn't have to take anyone's word for how the system behaves. They get a front row seat to the system proving itself, run after run, before anything reaches production.

That changes what a young company can say out loud. Certifications like SOC 2 and HITRUST need months of evidence that controls are working in production before an auditor will sign off. Most teams build first, start that clock later, and spend a year or more fixing what the audit finds. When compliance controls are part of the architecture from the first deploy, the clock starts on day one.

So the pitch stops being "we take security seriously." Two products I've owned this year show what it becomes instead. One is a health application that passed a third-party HIPAA audit within 90 days of its first line of code. The other is a financial intelligence platform that opened its SOC 2 observation window six months after the business opened its doors. An early-stage company with certifications in hand, months or years before its competitors can get there, is having a completely different conversation with buyers. Leadership and sales aren't describing what they hope the product does. They're pointing at proof.

We prove delivery. We don't hope for it.

Agents still get room

None of this means agents are on a short leash. They choose how to build, and they get creative about it. But every freedom an agent has sits inside a boundary the design process already measured.

The agent decides how. The design already decided what, and the ledger proves it.

This is leadership work

OpenAI makes a comparison I appreciated. They describe their approach as something like leading a large platform organization. Hold the boundaries firmly, and give teams room to decide how they work inside them. Anthropic ends their playbook in the same spirit: "The loop keeps running. Human judgement stays above it."

I'd take it one step further. Harness engineering is leadership, and leadership means extreme ownership.

When an agent ships something wrong, the agent isn't accountable. I am. That's why I front-load the design, make the specification something a person agrees to, keep decisions current, and block the build when something doesn't hold. Every one of those is a choice to own the outcome before it happens instead of explaining it afterward.

Ownership also means being honest about what tooling can't do. None of my tools can tell you whether your design matches what your business actually needs. That's a people problem. Traditional domain modeling sessions put business decision makers in the room, and I expect the same when we converge a design. If it's just an architect and an AI, the architect carries that risk. At a minimum, get buy-in from the wider organization before you put the design into motion.

Why I'm writing this

I've been heads down for almost two straight years building domain tooling for exactly this problem. Early on, I kept running into the same gaps in AI-driven development. Agents guessed where they should have asked. Boundaries lived in documents instead of in the system. Teams found out what was wrong after it was built.

So I went after those gaps one at a time, working alongside Claude the whole way. How do we limit what an agent is allowed to infer? How do we tighten boundaries so they enforce themselves? How much can we prove before we build anything at all?

That work has paid off in more ways than I expected, and one of them is watching the industry arrive at the same place from different directions. The market is converging fast on one idea. Speed isn't the hard part anymore. Accuracy is. I'm very much embedded in that part of the industry, and I plan to keep sharing what I learn from inside it.

What the rigor costs, and what it buys

I won't pretend this is the easy road. My way of working asks for a lot more discipline than either post describes. It takes more thought, more leadership, and more time before the first line of code. On a typical project, planning is around 70% of my investment. The remaining 30% goes to building and validation.

The harness and the playbook look like friendlier investments up front. They ask for enough discipline to lay a solid foundation for a long running project, and they let a team start moving quickly.

In my experience, the difference evens out over time, and then it flips. Every bit of rigor built into the environment before we build gives the agent more room to operate inside those boundaries. And because the boundaries were measured before the work started, far less human attention is needed to confirm what the agent did inside them. That's where the velocity catches up. Past that point, an agent working inside a proven design can keep pace with, and sometimes outrun, one working in an environment that asked for less up front.

You pay for the rigor once. You get the speed back on every feature after it.

What's next

Both posts confirmed something I've believed for a long time. The discipline didn't go away when agents started writing code. It just lives in different places now. Where you put it depends on where you're trying to go.

Over the next few posts I'll go deeper into each piece, starting with a closer, constructive look at OpenAI's harness engineering post. After that: drift, memory and supersession, specifications and prediction, domain design with SiDD, Moment, and Facet, and evidence you can hand to an auditor.

All of these tools fit together into one delivery chain called Complai, and most of it is open for you to use today. Feature, Moment, and SiDD are open source. Facet is free.

If you've landed on the same conclusions, or different ones, I want to hear about it.