(Humans) Yeah (What are they good for?) Absolutely everything.
Agents are extraordinary at building against a definition. They still can't produce one.
Anthropic documented where the software lifecycle is headed. I agree with most of it. The part I don't agree with is the part that decides whether the rest works.
This is the last post in the series, and the last from me for a while. Part 1 surveyed three approaches from a distance. Part 2 examined OpenAI's harness engineering post. This one turns to Anthropic's AI-native SDLC playbook. It's the right place to end because Anthropic is leading the market in agentic development, and its playbook is the clearest account yet of what the software development lifecycle looks like once agents write most of the code.
I should say up front what's mine and what isn't. I didn't invent the disciplines I use: domain-driven design, event storming, event sourcing, CQRS, or the SDLC. Other people established them long before I encountered them. What I've done is arrange proven disciplines into something that works for me, the way a painter doesn't invent color. Two tools are new: SiDD and Facet. Everything else is borrowed, with credit.
What AI did to the SDLC
The playbook's diagnosis is one I'd sign off on. Code stopped being the bottleneck. Build times collapsed from weeks to hours, while the stages on either side, planning, review, testing, and deployment, still move at the speed of people. Our controls were designed for a world in which one person wrote every line, and another read it. Those controls stop making sense once an agent writes the code.
The playbook keeps the six stages of the SDLC but changes how each one runs: plan, design, build, test, deploy, maintain. So do I. I want to be clear about that because "harness engineering" can sound like a break from the lifecycle everyone already knows. It isn't. It's a new way to run it. The SDLC is still my north star, and I'd be suspicious of any approach that claimed to have replaced it.
Where we part ways is in what to do about the bottleneck once it moves. The playbook compresses the human-speed stages. Requirements and design are handled in a single session with an agent. Review becomes layers of agent passes, with people at the gates. Intent gets captured by anyone with an idea and a chat window. Each of those moves aims to bring the slow stages closer to the speed of the fast one.
I do the opposite. I invest in the slow stages on purpose because that's where correctness is decided. On a typical project, about 70% of my effort goes into planning and design before an agent builds anything. That investment creates two distinct loops: one to make the design true and another to make the build match it. The 70% figure looks absurd next to the playbook's ambition, and it's the whole difference between us.
Where they're right
There's a lot in the playbook I'd tell any team to adopt tomorrow.
Plan before you build, and commit to the plan. The playbook starts every build in plan mode, iterates until an engineer who never saw the conversation could implement it, and checks the eventual diff against the plan. That's a smaller version of what I do with a whole design, and it's the right instinct.
For bug fixes, write the failing test first and block the agent from touching it. This is the closest the playbook comes to the rule I hold most tightly: the thing being measured shouldn't define the measurement. They apply it to bug fixes. I apply it to everything.
A skill is advice, and a hook is a wall. The playbook says it directly: the skill makes violations rare, and the hook makes them close to impossible. That's a boundary enforced inside the system it governs, and I've been arguing for exactly that since Part 2.
Detection stays deterministic. In the maintenance stage, a script monitors production, and a model is invoked only once a threshold is breached. No model decides whether something is wrong. A model reacts once something is. That's my inference rule in different clothes, and I was glad to see it.
And the agent acts up to the production gate and never past it. We agree on that completely.
What is a human for?
Both of us are one person plus AI. The playbook is built around an engineer who can steer multiple agent sessions simultaneously. My work is with one architect and an AI partner. The difference isn't how many people are in the room. It's what the people there are doing.
In the playbook, people review. Someone has an idea, and a model writes it up as intent. The person corrects it. A model turns the intent into a specification. A product owner reviews it. A model writes the plan. An engineer corrects it. A model reviews the pull request. A code owner approves. At every gate, there's a person reading something a model wrote and deciding whether it's close enough.
Below production, that review thins out by design. In development, the agent deploys on its own. In production, a release manager authorizes the deployment. Human attention gets spent where the playbook thinks it matters most, at the end.
In my work, people define. The business writes the requirements in the room before any model gets involved. When the design has converged, a person looks at Facet running it and decides whether to proceed to build. The design is translated into specifications, and people write the measurements, outputs, and contracts into those specs. Then an agent builds, and the compiled tests check what it built against what the people wrote.
The playbook's approach depends on sufficiently complete specifications, which means the model must understand the business well enough to write them. In my experience, models can't do that yet—not because they're poor writers, but because they're not psychic. A model can turn a paragraph of intent into a fluent, plausible specification for a business it has never met. Yet a reviewer can catch only what that paragraph represents. Whatever it leaves out remains invisible to both of them until the build runs into it.
I know the objection, so I'll meet it. Yes, there are people in the loop for the playbook. But what they're reviewing is a written description, and you can't run a test against a description. Put those same people on the contracts instead, and their time becomes something the pipeline can check on every run.
Two loops, not one
No approach gets every task right on the first pass. Something gets built, an unseen problem emerges, and the work comes back around. I call each pass a loop. How those loops are structured is the real difference between the playbook and my approach, and everything above follows from it.
There's a design loop. It starts with the people who run the business, working through narratives of what should happen and when. SiDD takes the resulting model, produces a report of every question the model can't answer, and I work through those questions one at a time with the people who own the answers. Then the model runs again. The number of open questions has to fall on every pass. If it doesn't, something underneath is wrong, and we stop and fix that instead of going around again. That loop ends when there are no open questions left, and Facet can run the design end-to-end without breaking a contract. Nothing in that loop builds anything. Nothing in it enforces anything. It exists to get the design right and to get the people who'll live with it to agree.
Then a different loop starts. The converged design gets translated into Feature specifications. Each measurement maps directly from the model: contracts at the boundaries become input shapes, emitted events become predicted events, and the conditions under which a flow refuses an action become rejection cases. Nothing the business agreed to gets reinterpreted along the way. What gets added are the build instructions—where the handler lives and what it may touch—and that's the part an agent and I write together. Once someone approves the spec, the agent builds it. On every run, the pipeline compiles tests from the spec, and the build either matches every prediction or it doesn't.
Those two loops have different reviewers, and that's deliberate. The design loop is reviewed by the business and by whoever owns each open question. The delivery loop is first reviewed by the person who approves the spec, then by a compiler. If a delivery problem turns out to be a design problem, it goes back to the design loop. It doesn't get patched into a spec by someone who wasn't in the room when the design was agreed upon.
The playbook presents one main loop where intent becomes a spec, then a plan, a build, a review, and a production release. A production breach creates a new intent and restarts the cycle. It's a clean model, and I understand why. One loop is easier to draw, automate, and hand off to a large organization. The cost is that every kind of correction has to travel the whole way around.
Where the loops come from
The playbook never says "fail fast, fail often," but its loop follows that pattern. A model drafts, a person adjusts, the model drafts again, and eventually the output matches what the person wanted. The playbook is honest about this. It expects two or three rounds of UI work. It says implementation is often a single pass when the plan is solid. When production reveals a problem, it says to write the next intent. Every stage is built to go around again, and I respect that more than I can say, because most of the industry still writes as if the first pass were the last.
I loop too. Anyone who tells you their process doesn't is selling something. The difference is how many times and what's being corrected.
My loops correct surprises. Things nobody predicted during design, because nobody predicts everything, and that will be true of every approach forever. When the contracts are right, I see about half a dozen loops in a feature before it's done. That's my experience, not a benchmark, and I'd hold it loosely.
The playbook's loops correct something else. They correct the distance between what a model understood and what the business meant. That distance is introduced at the first stage, when a model writes the intent, and each subsequent stage offers another chance to find it. Some of the mismatches surface in review, some in tests written by the same model, and some only in production. A production discovery creates a new intent, and it goes around the loop.
I spend human time early to close that distance before anything gets built. The playbook spends it late, one loop at a time. Both approaches are legitimate. In high-stakes work, however, the first is far cheaper when a mistake means an audit finding or a number someone trades on.
Faster code, slower feedback
This is the part I care about most, and it's the part I think the market hasn't caught up to yet.
Tight feedback loops are the core skill of a fast engineering organization. That was true before agents, and agents haven't changed it. What agents changed is where the loops hide.
An agent can write code faster than any person. Then, because nobody trusts what it wrote, the code goes through a self-check, a verifier, layered review passes, evaluation suites, and an adversarial gate before a person sees it. Each check is a sensible response to a real problem. Stacked inside one loop, however, they can delay the moment when a person learns whether the work was right. The typing gets faster, the feedback gets slower, and slow feedback is exactly what good engineering discipline exists to prevent.
The answer isn't fewer checks. It's smaller loops, each with the check that fits it. A design is checked by running it and by asking the people who own the business. A build is checked against the predictions that those people approved. Enforcement goes where it matters and nowhere else. Small chunks, checked hard, let a person learn the answer quickly.
Where their approach reaches further than mine
I'd be lying if I said the playbook has nothing to teach me. It does, in three places.
Their maintenance stage is more operable than mine—deterministic thresholds on production metrics, tiered responses, and findings that re-enter the pipeline as intent. I close the loop from production through live specifications running against deployed systems, and it works, but it's hard to run, and the playbook's version is easier to adopt.
Their controls on the agent itself are worth studying. Managed settings can deny access to credential files, allow only approved network destinations, require work to begin in a sandbox, and prevent local overrides. That's enforcement inside the system, applied to the agent's own working session, and it's more thorough than what most teams have.
And they work in parallel—several sessions per engineer, each on its own task. I work one thing at a time by rule, and I've accepted the throughput cost to avoid batching an unresolved decision. Inside a spec that touches nothing shared, their pattern would fit mine.
Where this series ends
Five posts, one question. When agents do the work, where does the discipline go?
OpenAI put the discipline in the environment and optimized for throughput. Anthropic put it in the lifecycle and optimized it for adoption. I put it in design and measurement, optimizing for proof because my work lives in places where being wrong isn't cheap.
I love where this is going. I've spent the year heads down on it, and I'm more convinced than ever that agents are the best thing to happen to software delivery in my career. But the market now has to adopt this in practice, and that means being honest about where models stand today. They're extraordinary at building against a definition. They're not good at producing one because the definition lives in the heads of the people who run businesses, and no amount of fluency can substitute for asking them.
Keep humans where they matter. Evolution in this market is happening with or without us. A decade ago, I thought my value was in authoring code to ship products. I was wrong. Let the agents build, and give them something true to build against. But most of all, let the humans create.
If you've been reading along, thank you. If you want to see this working, Feature and Moment are open source, SiDD is publicly documented, Facet is free, and flmnt provides the memory layer. Together, they form a delivery chain called Complai.
I'll be back when there's something worth saying.
Get new posts in your inbox. No spam, unsubscribe anytime.