Command Palette

Search for a command to run...

Writing / Testing

If your correctness lives in SQL, mocking the database tests nothing

Aug 11, 2026·4 min read
PostgreSQLTestingBackend

There's a comfortable style of test that stubs the database out: the repository returns canned rows, the service does its thing, you assert on the result. Fast, isolated, disciplined-feeling. For a certain kind of system it also tests almost nothing — because if the behaviour you care about is enforced by the database, a mock handing back whatever you told it to will never reach the part that can break.

Where the correctness actually lives

In a data-heavy service, a surprising amount of the specification isn't in application code at all:

  • Uniqueness is a unique index.
  • "This can't exist without that" is a foreign key.
  • "At least one of these is required" is a constraint.
  • "You can only see rows you're allowed to" is a join in the query.

Mock the database and every one of those evaporates. The mock cheerfully returns two rows with the same key, a child with no parent, or a row the caller should never have been able to reach — it has no constraints to violate. The assertions pass. The invariant is still broken in production.

Mocks encode the assumption you're trying to test

A mocked database is a statement of what you believe the database would do. The bugs worth catching are precisely the cases where that belief is wrong: the constraint nobody added, the race two concurrent transactions hit, the access join that leaks a row under a filter you didn't picture.

A mock can't surface any of them, because it was built out of the same beliefs as the code it's checking. Green means "this code agrees with my assumptions," not "this system is correct." Those are different claims and only one of them is worth a merge.

The merge gate

So the rule is blunt: anything touching the data layer is tested against a real Postgres, and that's the gate a change clears to merge. Pure logic still gets fast unit tests — there's no reason to boot a database to test a formatting helper. But the moment an outcome depends on a constraint, a transaction, or a query, it runs against the real engine.

Coverage is a tool, not a score

The related decision, and the one that gets more argument: coverage is deliberately not comprehensive. It's aimed at the functionality where being wrong actually costs something — contract validity, access checks, the lifecycle transitions Customer Success can trigger — rather than spread evenly until a percentage looks respectable.

Chasing a number produces tests of getters. It also produces a suite slow enough that people stop running it locally, and a figure that reassures without corresponding to risk. A test exists to tell you when something important broke, and "important" is not evenly distributed across a codebase.

The honest version of that position is that it requires judgment, repeatedly, about what matters — which is harder to defend in a review than a threshold, and harder to fake.

What it costs

  • It's slower. A real database per run is seconds you didn't spend mocking. You buy them back the first afternoon you don't spend chasing a constraint violation that only appeared in production.
  • You need fixtures and isolation — a clean schema per run, transactional rollback or a fresh database between tests. Genuine setup, paid once and amortized over every test written after.
  • Untested paths are a decision, not an oversight — which only works if the decision is revisited. Targeted coverage degrades into accidental coverage the moment nobody is asking what moved into the "matters" column since last quarter.

Where a different shape fits

  • The database is a detail. If the schema is a bag of columns and every rule lives in application code, mocks test the thing that can break and a real database is ceremony. The question is always where the invariants are, not which style is more rigorous.
  • The suite outgrows the gate. Real-database tests scale in wall-clock time. Past a point the answer is parallel databases per worker or a subset gate on pull requests with the full suite on main — not a retreat to mocks.
  • Contract tests across a service boundary. Where correctness spans two services rather than one schema, neither a mock nor a real database catches the mismatch. That needs a shared contract both sides verify against, which is a different tool for a different failure.

The point of the suite is to fail when the database would. Anything that can't do that is measuring something, but it isn't measuring whether the thing works.