Coada / writing

Your tests can't see the waves

We can build far more complex systems than we could ten years ago. Our testing hasn't moved.


Ten years ago, if you wanted an event-sourced system with properly separated bounded contexts, sagas coordinating across service boundaries, projections rebuilding from an immutable log, and a real ubiquitous language holding it together — you needed a team. A big one. You needed an architect who'd done it before, several engineers who could be taught, a year, and an organization willing to fund all of that before anything shipped.

That was the actual barrier. Not the ideas. Evans published in 2003. The ideas have been sitting there, freely available, for twenty years. What stopped most people from using them was that the labor cost of doing it properly was enormous, and the payoff arrived far too late to survive a budget conversation.

That barrier is gone.

I'm one person. My current system has multiple bounded contexts, an event-sourced core, sagas with timeouts, projections, and a domain model that was converged over several formal passes before a line of implementation existed. Agents did nearly all of the typing. A meaningful number of people reading this have built something of comparable structural complexity in the past year, alone or nearly alone, and are quietly aware that it's more machinery than they've ever personally been responsible for before.

The complexity of what we can build went up by an order of magnitude. The way we verify it did not move at all.

We're still writing unit tests that assert return values, integration tests that assert a couple of collaborators got called, and end-to-end tests that assert the page says the right thing. That toolkit was designed for systems where the interesting behavior was the return value. We are no longer building those systems.


The splash and the waves

Every test you've ever written asserts the splash.

A function is called. It returns something. You check the something. Maybe you check that one collaborator got invoked with the right arguments. The assertion surface is shaped exactly like the thing you were thinking about when you wrote the code, because you wrote the test from the same mental model, usually minutes later, often in the same sitting.

That's fine as far as it goes. Return values matter. Response codes matter. I'm not here to tell you to stop checking them.

But in a modern system, the return value is the least interesting thing a command does.

Call one handler in an event-sourced system and the actual footprint is: rows written, events appended, projections updated, messages queued, webhooks dispatched, caches invalidated, analytics fired, audit rows inserted, a saga timer armed. Those are the waves. They travel outward, they hit other contexts, they arrive at other people's inboxes, and they persist long after the response has been serialized and forgotten.

Almost none of them are asserted.

And I want to be careful here, because a certain kind of reader is already typing. Yes, your framework can check some of this. Strict mocks exist. verifyNoMoreInteractions exists. Spies exist. You can assert that an unexpected call didn't happen.

But look at what that requires. It's opt-in. It's per-test. It's per-collaborator. And it only fires if someone already imagined the specific interaction worth forbidding. You have to have thought of the wave in advance in order to check that it didn't happen.

The default in every mainstream testing framework is permissive: anything you didn't assert against is allowed.

That default was harmless when a unit was a function and its effects were its return value. It's structurally dangerous when a unit is a command handler sitting on top of five services, and it's how we're all testing the most complex software any of us has ever personally shipped.


What this looks like when it bites

Take refunds. Everyone has built one. The domain is obvious to any reader and the stakes are obvious to any customer.

Here's the spec, in the language I use for this — a .feat file. The syntax is real; you can read it cold.

spec SPEC-BIL-014 "RefundPayment"
context   Billing
aggregate Payment
type      command
status    verified

touches handlers/refund-payment.ts

construct:
  handler at handlers/refund-payment.ts
  imports PaymentRepository  from repositories/payment-repository
  imports RefundGateway      from gateways/refund-gateway
  Refundable balance is folded from the aggregate's own PaymentRefunded
  history. There is no refunded boolean and no denormalised total to trust.

enforce:
  validate input against RefundPaymentInput
  load Payment by paymentId, at the version carried on the request
  fold refundedTotal from prior PaymentRefunded events on the aggregate
  compute refundableBalance as capturedAmount minus refundedTotal

  rejects PAYMENT_NOT_FOUND        when no Payment exists for paymentId
  rejects PAYMENT_NOT_CAPTURED     when the Payment has never reached captured
  rejects ALREADY_REFUNDED         when refundedTotal equals capturedAmount
  rejects AMOUNT_EXCEEDS_BALANCE   when amount is greater than refundableBalance
  rejects VERSION_CONFLICT         when the aggregate version has advanced

  replay the stored outcome when idempotencyKey matches a settled refund
  issue the refund through RefundGateway before appending any event
  append PaymentRefunded carrying refundId, amount, and gateway reference

contract:
  input     RefundPaymentInput from "schemas/refund-input"
  response  RefundResponse     from "schemas/refund-response"
  error     ErrorResponse      from "schemas/error"

scenario "refunds a captured payment in full":
  given: payment "pay_8812" captured at 4200, version 3, with no prior refund
  when: RefundPayment {
          paymentId: "pay_8812", amount: 4200,
          idempotencyKey: "idem_9f21", expectedVersion: 3
        }
  predict success:
    response 200 RefundResponse { refundId: "rfnd_5107", replayed: false }
    paymentGateway has [{ action: "REFUND", reference: "pay_8812", amount: 4200 }]
    eventStore     has [ PaymentRefunded { amount: 4200, version: 4 } ]
    database       has [{ action: "UPDATE", table: "payments", refundedTotal: 4200 }]
    email          has [{ template: "refund_confirmation", amount: 4200 }]
    analytics      has [{ event: "payment.refunded" }]

scenario "rejects a payment already refunded in full":
  given: payment "pay_8812" captured at 4200, version 4, carrying a settled
         PaymentRefunded of 4200 as "rfnd_5107" at "2026-08-14T09:12:04Z"
  when: RefundPayment {
          paymentId: "pay_8812", amount: 4200,
          idempotencyKey: "idem_c40b", expectedVersion: 4
        }
  predict rejection ALREADY_REFUNDED:
    response 409 ErrorResponse {
               refundedTotal:    4200,
               existingRefundId: "rfnd_5107",
               refundedAt:       "2026-08-14T09:12:04Z"
             }
    paymentGateway has []
    eventStore     has []
    database       has []
    email          has []
    analytics      has []

scenario "replays a settled refund under the same idempotency key":
  given: payment "pay_8812" refunded 4200 as "rfnd_5107" under key "idem_9f21"
  when: RefundPayment {
          paymentId: "pay_8812", amount: 4200,
          idempotencyKey: "idem_9f21", expectedVersion: 4
        }
  predict success:
    response 200 RefundResponse {
               refundId:   "rfnd_5107",
               refundedAt: "2026-08-14T09:12:04Z",
               replayed:   true
             }
    paymentGateway has []
    eventStore     has []
    database       has []
    email          has []
    analytics      has []

The language is called Feature. The toolchain is at github.com/mmmnt/feature, and there's a browser playground at feature.mmmnt.ai if you want to break the grammar yourself while you read the rest of this.

Now look at what those three scenarios pin down that a conventional suite would not.

Refunded-ness is derived, and the spec says where from. There is no refunded boolean. refundedTotal is a fold over prior PaymentRefunded events on the aggregate, and the construct block says so in a sentence the implementer has to read. That single line eliminates the most common refund bug in existence — a denormalised flag that disagrees with the event history after a partial refund, a replayed webhook, or a projection rebuild. The enforcement isn't "check if it's refunded." It's a stated derivation with a stated source.

Every rejection has its own condition, in the enforce block, individually. Not-found, not-captured, already-refunded, exceeds-balance, and version-conflict are five distinct refusals with five distinct triggers. feat audit refuses the spec if any declared rejection has no scenario demonstrating it, so you cannot claim a refusal you never proved.

The 409 is asserted on its contents, not its status. It echoes the total, the existing refund id, and the original timestamp — 2026-08-14T09:12:04Z, a literal from the given. If the handler stamps now() into that field, the test fails. That's the tell that it minted something instead of reading something. A status-code assertion cannot see the difference between a handler that correctly refused and a handler that created a duplicate refund record and then errored on the way out.

And the replay scenario asserts a 200 with zero effects everywhere. Success, and nothing happened. That combination is unrepresentable in most test suites — success is implicitly assumed to mean work occurred — and it's precisely how you prove idempotency rather than hoping for it.


Now the bug those specs catch, which I'd bet money exists in production somewhere right now.

Someone implements the refund. The double-refund guard is correct — it returns a 409 exactly when it should. But the notification dispatch was wired slightly upstream of the guard, or into middleware, or a policy subscribed to the wrong trigger. So on the rejection path, the customer receives an email that says your refund of $42.00 is on its way.

The refund did not happen. It is never going to happen. The response code was correct. The state was correct. Every conventional test on that handler passes, because every conventional test asserts the 409 and the 409 is right.

The waves went out anyway. Somebody is waiting for money that isn't coming, and the first person to learn about it will be a support agent, weeks later, with no idea why.


Queries are worse

Here's the second one, and it's the one that surprises people.

spec SPEC-BIL-021 "GetInvoice"
context Billing
type    query
status  verified

contract:
  input     GetInvoiceInput from "schemas/invoice-input"
  response  InvoiceResponse from "schemas/invoice-response"

scenario "returns an issued invoice by id":
  given: an issued invoice "inv_5501"
  when: GetInvoice { invoiceId: "inv_5501" }
  predict success:
    response 200 InvoiceResponse

There's a rule in the grammar: a spec declared as type query cannot express a write. Not "shouldn't." It's a parse error. You cannot author the file.

Which means declaring something a query silently predicts zero effects everywhere, permanently, without anyone writing a single assertion about it.

Now think about what actually accumulates on read endpoints in a real system over a couple of years. A last_viewed_at timestamp. A rate-limit counter increment. A cache warm. An analytics event. An access audit row — often added deliberately, for compliance, by someone who never told the original author.

Every one of those is a write on a path everybody in the building believes is read-only. They're the reason a report endpoint takes a row lock under load. They're the reason a "read" shows up in a GDPR export. They're the reason your replica lag spikes when someone opens a dashboard.

And no test anywhere fails, because the invoice body came back correct and the invoice body is the only thing anyone thought to check.

Side-effect freedom is treated as a convention that people honor. It should be something the grammar refuses to let you violate.


Inverting the assertion surface

The mechanism that makes all of this tractable is one inversion, and it's worth stating precisely because it sounds impossible until you see the trick.

You cannot enumerate everything that shouldn't happen. That list is infinite. Every test suite that approaches this from the "assert the absence" direction dies on that fact.

So you don't. You enumerate what should happen — which is finite and small — and everything outside it is a violation by construction.

That's the whole move. A scenario declares the complete observable footprint. Predicted-and-captured passes. Captured-but-not-predicted fails. Predicted-but-not-captured fails. Right event, wrong shape, fails. Nobody had to imagine the audit row in advance. Nobody had to know the email existed. The absence of prediction is the assertion.

A failure reads like this:

UNPREDICTED record PaymentAuditLogged in eventStore
  predicted  eventStore has [ PaymentRefunded { amount: 4200, version: 4 } ]
  captured   [0] PaymentRefunded    ✓ predicted
             [1] PaymentAuditLogged ✕ unpredicted
  anchor     refund-payment.feat › "refunds a captured payment in full" › eventStore[1]

You didn't write that assertion. You didn't have to.

Two supporting rules make it hold up in real use. Every configured service must appear in every prediction — omitting one is a validation error, and zero is written down explicitly as has []. You cannot quietly decline to look at a service, which is exactly how effects hide. And predictions compile deterministically into the suite with no model and no glue code in between, then get byte-compared in CI, so the spec you agreed to is provably the spec that ran.

We tested that claim adversarially: one planted defect per scenario, every tool configured the way its own docs recommend, every lane required to pass a clean implementation first, one binary question per cell. An unpredicted audit-row write and a leaky rejection both came back fully green in every conventional lane — not because those suites were bad, but because nothing in them was ever pointed at the thing that broke. Ten generations of the same spec produced one byte-identical suite; the agent lanes produced ten unique suites out of ten and none of them caught the planted write. Where conventional approaches win, they're published too. It's version 0.1.4, two lanes are still pending, and a benchmark I built is a benchmark I built — the harness regenerates from scripts, so go break it.

One limitation no inversion fixes: a spec can predict the wrong footprint, an implementation can match it perfectly, everything passes, and the software is wrong. Predictions being reviewable data in one file rather than assertions scattered across a suite makes that visible at review. That's a mitigation, not a guarantee, and anyone selling you a guarantee at that layer is lying.


The surface moved

Something changed in what we're capable of building, and it changed fast enough that most of us haven't stopped to account for it.

The architectural patterns that used to require a funded team and a year of runway are now reachable by one person with a converged model and an agent doing the typing. Event sourcing, bounded contexts, sagas, projections — the whole vocabulary that was enterprise-only a decade ago is now a weekend away for anyone who wants it. That's a genuine expansion, and I don't think it's reversible.

What expanded with it is the effect surface. Every one of those patterns multiplies the number of observable things a single command does. A handler that once returned a value now writes rows, appends events, updates projections, queues messages, notifies humans, and arms timers — six or eight waves per splash, across contexts that don't know about each other.

The verification did not expand. It is the same shape it was when a unit was a function: assert the return, assert a call or two, move on. We are measuring a system with eight effects using an instrument built to see one.

That gap isn't negligence. It's what happens when capability jumps a generation and the tooling around it doesn't. The assertion model most of us are using was correct for the software we were writing when we learned it, and it quietly stopped being sufficient somewhere in the last two years, without an announcement.

Go pick your highest-stakes command handler — the one that moves money, or provisions access, or sends something a customer will read. Write down every observable thing it does. Then count how many are asserted anywhere in your suite. Then do the same for its rejection path, which is the one nobody ever writes down.

The number is lower than you expect. It was in my system too, before I started predicting the footprint instead of checking the return.

Your code is making waves. Something should be watching them.