Coada / writing

Slop Is Upstream

The AI slop you're looking at was created before anyone wrote a prompt.

My feed has decided the machines are bad at this.

The word is "slop." It gets applied to code, to prose, to pull requests, to entire products. The implied argument is always the same: the model produced garbage, therefore the model is garbage, therefore anyone shipping with it is a fraud. It's a satisfying take. It costs nothing to hold, and it flatters everyone who holds it.

I'm not going to tell you the models are flawless. I run an AI scheduling assistant in production that was built almost entirely by agents, and it has been wrong in ways that reached real people. That isn't the argument.

The argument is that almost every slop post I read is describing a loop, not a tool. And the loop belongs to the person posting.


What the complaint is actually describing

There are two failures stacked on top of each other, and neither one happens inside the model.

The first is the prompt. People love letting the model infer. They use natural language understanding to the absolute limit of its capacity and none of its strengths. A paragraph of intent. No inputs. No outputs. No contracts. No acceptance criteria. No domain vocabulary. No statement of what must be true when the thing is finished. And then they expect the model to be psychic — to reconstruct an entire unstated system from a sentence and land precisely on the picture in their head.

NLU is extraordinary at parsing an ambiguous instruction. That is not the same as the instruction being complete. Parseable and specified are different words.

Nobody would hand a contractor a single paragraph and expect back the building they were imagining. We understand this instinctively when the work takes six weeks and arrives from a person. We abandon it entirely when the work takes nine seconds and arrives from a machine — as though speed of delivery somehow implies completeness of instruction.

The second is the workflow. That loose prompt doesn't stay in a chat window. It goes into a real pipeline. The model produces work. The work gets a smell test — a glance, a nod, "looks okay" — and moves toward a customer-facing environment on the strength of that nod. There is no contract it was measured against, because nobody wrote one. There is nothing measurable at all, because nothing measurable was ever defined. The only quality signal in the entire chain is one person's aesthetic reaction to a diff they skimmed.

Then it breaks in front of a customer, and the model takes the blame.

That is not a model failure. That is a system with no specification, no acceptance criteria, and no evidence standard producing an unpredictable result — which is the only thing such a system has ever been capable of producing. It would behave identically with a junior engineer, an agency, or you at two in the morning.


The number that ends the argument

I've been keeping records on this for a while, across nine products taken from innovation to production.

Ungoverned, AI-mediated delivery in my own work ran roughly a 40% correction rate. Under the methodology I'll describe below, it runs 95%+ accuracy, and in its current version the correction rate is trending under 3% — because a whole class of former corrections has become a parse-time or audit-time refusal instead.

Same models. Same person. Two orders of difference in outcome.

The variable is input precision, not AI capability. That's not a slogan; it's the only conclusion the numbers support. If the difference between slop and rigor were a property of the model, that gap would not move when the only thing I changed was what I handed it.

There's a second finding underneath it that took me longer to accept. The two most expensive production failure classes in everything I've built were not wrong decisions. They were silent ones. Something proceeded plausibly where it should have stopped and asked.

That is a precise definition of slop, and notice where it lives. Slop isn't the model being stupid. Slop is a system continuing confidently past the point where it had enough information to continue. Every layer I'm about to describe exists to make that impossible — to force a halt where precision runs out, instead of letting plausibility carry the work forward.


The chain

The thing I'd most like people arguing about prompts to understand is that the prompt is not the top of the stack. It's near the bottom, and it's load-bearing only if everything above it is sound.

Seven layers, each reading only from the layers above it.

L0 — The palette. Domain-driven design and Event Storming as a binding grammar: commands, aggregate roots, domain events, policies, sagas, projections, invariants, value objects, bounded contexts, context relationships. Evans and Brandolini, extended with reactive concepts. Every layer speaks this vocabulary, every term has exactly one meaning inside its boundary, and the classification rules are enforced rather than suggested. A command targets exactly one aggregate. An aggregate emits events after a successful transition, or rejects with none. Events are past tense. State without rules is not an aggregate. Violating one is a defect.

L1 — Regulatory framing, when formal compliance is in scope. Skipped otherwise.

L2 — The converged domain model. The product's actual structure and behavior over time.

L3 — System architecture. Prescriptive defaults plus a corpus of architecture decision records.

L4 — Specification. Executable specs and schema contracts.

L5 — Execution. The build index, the agent, the closed boundary it works inside.

L6 — Verification. Automated comparison of what execution actually produced against what the specification said it would.

Architecture constrains specification. Specification constrains execution. Verification measures the distance between the two. Skip a layer, hand down an incomplete artifact, or let an ambiguity pass unchecked, and it doesn't stay local — it cascades into execution errors that look, from the bottom, exactly like the model being bad at its job.

Almost everyone posting about slop is operating at L5 with nothing above it and nothing below it. They're tuning the prompt because the prompt is the only layer they have.


L2: Converge the model before you write anything

This is where the work actually is, and it's the layer nobody builds.

I use Signal-Driven Development — domain-driven design with a feedback loop. Model, diagnose, resolve, repeat. It replaces the traditional multi-day workshop full of stakeholders with an AI-mediated convergence loop, which is the thing that makes rigorous DDD available to one person instead of a department.

Each pass produces a candidate domain specification, then a gap report — an enumeration of everything the model cannot answer about itself. Gaps come in four kinds. Structural ones are mechanical: an orphaned command, an aggregate that emits no events, an event nothing consumes. Heuristic ones are quality signals: an aggregate carrying too many commands, a context with too little policy coverage. Language ones are ubiquitous-language violations: synonyms, homonyms, undefined terms, events that aren't past tense.

The fourth kind is the important one. Decision gaps are places where multiple valid modeling choices exist and only the architect can choose. They're classified as errors, and they block convergence entirely. The system is not permitted to pick one and move on. That is the halt-loudly rule made structural at the modeling layer, and it is the single highest-leverage rule in the whole stack — because a decision gap silently resolved in pass one is a defect that reaches production in month four wearing a completely different costume.

Convergence is done at zero unresolved gaps, never at a pass count. And there's an invariant on the loop itself: the gap count must decrease across passes. If a pass finds more gaps than the previous one resolved, the resolutions are generating more uncertainty than they eliminate — stop, and go reassess the foundations rather than iterating harder.

Here's what that produces in practice, across nine products:

  • Pass one surfaces 16 to 34 gaps.

  • Pass two brings it to 6 to 9.

  • Pass three lands at 0 to 2.

  • Average across all nine: 3.2 passes, 275 total resolutions, and zero production-discovered modeling errors.

That last number is the one I'd put in front of anyone who thinks this is ceremony. Not zero bugs — I have plenty of bugs. Zero cases where production revealed that the domain model itself was wrong. Every structural surprise got caught while it was still cheap, by a process built to go looking for them.

Moment: the model as a language, not a whiteboard

The convergence loop needs somewhere to live, and sticky notes decay the moment the session ends.

Moment is an open-source domain specification language and toolchain. A single .moment file captures two things at once.

The spatial model is what you'd expect from any DDD tool: bounded contexts, aggregates, commands, events, value objects, invariants, policies, sagas, projections, and the relationships between contexts.

The temporal model is the part that doesn't exist anywhere else. Cross-context event flows. Moments ordered along a timeline. Context crossings with contract obligations attached to each one. The existing domain-modeling languages — ContextMapper, LEMMA, DomainLang — are spatial only. They describe what your system is. None of them describe what it does over time, which is where every interesting failure in a distributed system actually lives.

And all of it is executable as integration tests.

The toolchain runs a fixed pipeline: parse, derive, generate, emit. Parse enforces the grammar and roughly twenty structural validators. Derive walks every branch point and enumerates the paths. Generate emits Gherkin, specification documents, and an AsyncAPI description. Emit produces TypeScript types, aggregates, and test scaffolds.

On my scheduling product, that pipeline turned sixteen user journeys into 106 scenarios and over 1,300 events — exactly one happy path per flow, with everything else a variant, a failure, a negative, or a timeout. Plus the topology, an event catalog, an impact analysis, and the saga state machines. All of it derived, none of it hand-written, none of it free to drift.

Two properties matter more than the feature list.

The contract space is closed. Every schema name used in a scenario must be declared. An undeclared one is a parse error, not a warning. You cannot reference something that doesn't exist and find out at runtime.

There is no LLM anywhere in it. Moment is a local toolchain and a pure transformation pipeline. No API keys, no inference, nothing probabilistic. Given the same model, it produces byte-identical output forever. That is deliberate: the artifact your agent is held to cannot itself be something that might have hallucinated.

Before any agent was asked to build anything, I had a machine-verified enumeration of every path a user can take through my product, every point at which the system can refuse, and every contract crossing a context boundary. I could not confidently specify anything without that, and I don't believe anyone can.


L4: The spec is a translation, not an invention

Once the domain is converged, Feature takes over. One executable spec per command or query handler.

The scenarios aren't imagined. Happy-path moments become success scenarios. Branch conditions become rejection scenarios on the command that owns them. The invariants become rejection rules, corroborated in both directions against those scenarios. The spec's type is the palette classification, made executable.

The spec is not where I decide what the system does. It's where I formalize a decision the model already made and already tested, and bind it to something that runs.

That's the whole point, and it's the line I'd defend hardest: without a converged model, specs have nothing to translate. Without specs, the model leaves the decisions to the agent.

A .feat file has two zones, and the split is the reason the format works. The agent zone is freeform natural language — what to construct, what to enforce — where a parse error is impossible by design. The compiler zone is a strict typed grammar where malformation is a parse error and never a silent failure. One file, two audiences, no ambiguity about which half is prose and which half is law.

Four reference spaces are closed and validated at parse time: service keys, schema names, command names, actor names. An unknown name doesn't compile, and the error lists the valid ones.


L5: What the agent is actually allowed to do

Now, finally, the prompt.

Execution runs as four subtasks with checkpoints between them. Context review: read the spec, its contract registry, and the shared contracts it references — never another spec's internals. Test generation: a mechanical compiler step. Zero human- or agent-authored test code, byte-deterministic output, committed next to the spec. Implementation: make the generated suite green, touching only the file paths the spec declares. Verification: every scenario green, zero unpredicted, zero missing, zero schema violations.

Read that second step again, because it's the one people find surprising. The agent does not write its own tests. It cannot. Tests are compiled from the specification by a deterministic tool. The single most common way AI-generated code launders its own defects — writing an implementation, then writing a test that agrees with it — is structurally unavailable.

The boundary around execution is machine-enforced rather than requested. Only agreed specs get built. Only declared paths get touched. No architectural decisions that aren't already in the spec.

And two protocols do the heavy lifting:

Ambiguity. On an ambiguous directive, the agent halts that spec and asks. It does not guess, does not annotate the file with its interpretation, does not quietly downgrade the status and move on. Resolution is a spec edit made by me, after I answer.

Failure. If a generated test fails, the spec is correct and the code is wrong. If the agent believes the spec is wrong, it halts and flags it. It does not edit the spec. It does not edit the generated test. Those two sentences eliminate an entire genre of AI-assisted disaster, and they cost nothing except the discipline to mean them.


L6: Measuring the delta between what was specified and what was built

This is the layer that makes the rest of it verifiable rather than aspirational, and it rests on a single inversion.

A specification does not enumerate what shouldn't happen — that list is infinite. It declares exactly what should happen, which is finite. Everything else is a violation by default.

Every scenario predicts its complete footprint: the response, and the exact effects across every configured service. Then execution runs, effects are captured, and the two are diffed. Three verdicts:

  • Unpredicted — something happened that the spec never declared. Violation.

  • Missing — the spec declared it and it didn't happen. Violation.

  • Schema mismatch — the right event, in the wrong shape. Violation.

A rejection scenario predicts an error response and zero effects across every service. Which means a command that correctly refuses but still writes something is caught by construction. Not by a test somebody remembered to write — by the shape of the language. Every configured service must appear in every prediction; omitting one is a validation error, and explicit zero is a thing you write down.

The inversion runs at three levels. At run time, captured effects are diffed against predictions. At build time, the files actually touched are diffed against the paths the spec declared it would touch. At CI time, the committed test files are regenerated and byte-compared — any drift at all, whether a spec was edited without regenerating or a generated file was hand-edited, fails the build. Every CI run files a digest-anchored evidence bundle with per-scenario violation rows at spec coordinates.

And one law I'd tattoo on the pipeline: a lane that ran and filed nothing is a red gate, not a green run. Absence of evidence is treated as failure, because the most expensive class of defect I've ever shipped was a guard that had quietly stopped guarding anything while continuing to report success.

What this actually catches

Abstractions are easy to nod along with, so here is what the inversion has actually found in my own work — cases where the specification caught something neither I nor the agent knew was there.

An effect that was absent rather than wrong. A messaging policy recorded its outbound send into module-level state, but the handler and the adapter were importing that module by different specifiers — so the policy was writing into an array nothing read. The prediction inversion caught it immediately as a missing record. Note what makes this valuable: the effect wasn't incorrect, it was simply not there. In a conventional test suite that reads as a feature nobody built yet. In an inverted one it's a hard failure at a precise coordinate. The same run also caught the policy baking its own pre-rendered copy into a refusal notice, which put that wording out of the renderer's reach and out of translation entirely.

A closed schema silently dropping a fact. A send directive gained a new field, and the contract's closed schema quietly stripped it in transit. The value never arrived, no error was raised, and the message rendered without it. That defect is invisible to any system where schemas are advisory. It's a build failure where the reference space is closed.

A refusal a spec claimed but never declared. An audit run refused a spec because a predicted rejection had no matching rule in the enforce block — the law had been written with the wrong verb, so the audit couldn't read it. The scenarios were correct and ran green locally. The spec was asserting an enforcement it had never actually declared, and the two-way lint between rejections and rules is the only thing that could see it.

An event nobody emitted and a precondition with no law. The traceability sweep caught a projection spec consuming an event that no scenario anywhere ever produced, alongside a command precondition that cited no law. Both are the kind of quiet structural hole that ships happily and fails months later. The sweep, incidentally, caught its own author twice during the week it was being written.

And the one that justifies the entire approach. In a delivery product built on the same stack, a milestone that no contract had ever purchased could be demonstrated, accepted, and then invoiced — a money path that began with a name typed into a box. Closing that hole immediately turned eight scenarios red across three other specifications, because every one of them had been demonstrating milestones no instrument ever named.

Sit with that one. Those eight scenarios had been passing for weeks, and they had been passing because the hole existed. The specs encoded exactly the same wrong assumption as the code, so neither could catch the other. Nothing in a conventional test suite finds that, ever — the tests agree with the implementation, which is precisely what tests are usually for. It surfaced because a rule was tightened at the specification layer and the whole plane was forced to re-prove itself against the stricter rule.

That is what a specification layer buys that a test suite cannot: not confirmation that the code does what the code does, but a place to change your mind about what should be true and immediately learn everywhere that belief was already wrong.


The part that surprises people

There is no AI inference anywhere in this approach.

Not in the palette, which is a grammar. Not in Moment, which is a deterministic transformation pipeline with no model in it. Not in test generation, which is a compiler step. Not in verification, which is a diff. The specifications are authored and ratified. The decisions are documented, with causal chains linking each one to what caused it and explicit supersessions when a ruling replaces an earlier one. Nothing important in this system is arrived at by asking a model what it thinks.

We don't infer to create.

We infer in response — and almost always in response to a failure in the process. When something breaks, that's when the model earns its keep: reading the evidence, walking the causal chain, proposing where the fault lives. But by the time we get there, we already have the violation, the coordinate, the spec that predicted it, and the decision record that explains why the rule exists. The inference isn't generating truth from nothing. It's interpreting a body of evidence that already isolates the answer to a handful of possibilities.

That's the inversion of how most people are using these tools. The common pattern is to infer the design and verify by eye. Mine is to specify the design and infer only when the verification screams.


And it still fails

All of the above, and my system still produces defects that reach production. Convergence doesn't make you omniscient. Specs encode what you thought of. Gates catch the classes you've already learned to catch.

The hardest failure mode is one no toolchain detects: a spec that declares the wrong behavior, an implementation that matches it perfectly, and every test passing. Tests green, software wrong. The only defenses are that predictions are reviewable data sitting in one file, that every rejection must be justified in prose a human reads, and that changes to predictions surface explicitly at review. Review the contract, not just the code.

So the difference isn't that things stop breaking. It's what happens in the twenty minutes after they do.

When something lands wrong, I'm not standing in front of an opaque box wondering what the AI did. The question is narrow and answerable: which layer let this through? The domain model didn't capture the case. The spec didn't translate a branch the model had. The scenario set covered the transitions but missed a state-and-input combination. The prediction was complete but wrong. Or the thing was correct and simply never reachable.

Every one of those is a specific, findable, fixable defect in a specific artifact — and most of them are mine. Upstream of the prompt. In the modeling, the specification, or the rigor around the evidence.

That's what governance actually buys. Not the absence of failure — the ability to say precisely where failure entered, fix it at that layer, and encode the lesson so the class doesn't recur. A loop with a diagnosable failure mode gets better every week. A loop with a smell test in it has no failure mode at all, just an unbounded supply of surprise.


Where the critics have a point

I'd be doing the lazy version of this if I stopped here.

This is expensive. Everything I described is real infrastructure. A domain-specification language, a compiler, a derivation engine, a spec toolchain, an evidence discipline, a durable record. I built most of it. Telling a solo developer on a deadline that they "just need governance" is not far from telling an exhausted person they should exercise more. True, unhelpful, and a little smug.

Nobody sold them governance. The pitch is describe it and it appears. The demos are greenfield toys where nothing can be built-but-unreachable, because nothing is deployed and no journey has a fourth branch. The distance between that pitch and a real production system is enormous, and it isn't the user's fault that nobody mentioned it. A great deal of slop frustration is anger correctly aimed at marketing and misfiled as a verdict on capability.

And some slop is just slop. There's a real volume of low-effort output being pushed into the world by people who don't care about the result. Being annoyed by that is reasonable. I'm annoyed by it too.

So my position isn't "skill issue." It's narrower, and I think harder to argue with:

The slop you're seeing is real, and it is evidence about the loop that produced it — not about the ceiling of the tool. Generalizing from someone else's ungoverned loop to a claim about what's achievable is the same reasoning error I'd flag in any postmortem. A rule that explains most of the evidence is not a diagnosis.

If you want the tool to stop producing slop, the work is upstream of the prompt. It's in knowing your domain well enough to converge a model of it, converging it well enough to specify it, and specifying it well enough that inference was never required in the first place.

The design you agreed to should be the design that ships. Right now, for most people, there isn't a design — there's a paragraph, and a hope.


About this article

This article was generated by AI. I'm going to say that plainly, because the alternative is worse and because the entire point collapses if I'm cagey about it.

But understand what that actually involved.

I decided what this was going to argue, and I planned it section by section. I chose which layers led and which followed, and when an early draft opened with the wrong one, I said so and it was restructured. I chose the context it was allowed to draw from, and I pointed it at specific bodies of work rather than letting it improvise. I told it what to cut — the examples that read as blame, the jargon a reader wouldn't know, the passages that had drifted into somebody else's voice.

And the material underneath it exists because I built the infrastructure to record it. The decisions, the supersessions, the failures, the causal chains between them, the measured correction rates across nine products — those are queryable today because I decided a long time ago that they should be, and then did the work to make it so. My agent reviewed my recent work and cited it back to me because I made my recent work reviewable.

That is the whole argument, demonstrated instead of asserted. The output is exactly as good as the governance around it.

The governance is the part that was mine.