Writing / Architecture
No hard delete: a lifecycle you can walk back
Self-serve tooling for a non-engineering team means the person clicking "cancel" is not going to trace what that button does first. If cancel means a row disappears, the system's safety depends entirely on nobody ever clicking it by mistake, which is a bet you eventually lose.
Delete is a one-way door
A hard delete doesn't only remove a row. It removes the record of why the row existed and what state it was in beforehand.
When the deleted thing is a paying customer's contract, "restore from backup" is not the same as "undo." It's slower, it involves a different team, it recovers a point in time rather than a decision, and it's the kind of incident nobody wants to write up. The gap between those two words is the entire argument.
Cancellation as a transition, not a disappearance
A contract moves through an explicit, finite set of states — draft, active, suspended — with defined transitions between them, and there is no transition meaning "row no longer exists." Cancelling is suspending: reversible, visible, and exactly as findable afterwards as before.
draft ──▶ active ──▶ suspended
▲ │
└─────────────┘
(reactivation is just another transition)Reactivation isn't a recovery mechanism bolted on for emergencies. It's the same kind of operation as any other transition, which is why it can be trusted — it's exercised in normal use rather than only when something has gone wrong.
What it buys beyond undo
The reversibility is the obvious win. The one that mattered more in practice is that delete stops being special.
A destructive operation usually gets its own code path: its own permission check, its own confirmation, its own audit entry, often written later and separately from everything else. That path is used rarely, tested less, and is the one place a mistake is unrecoverable. Removing it means suspension is governed by the same authorization and audit logic as every other transition, because it is every other transition. There's no special case left to get wrong.
The history stays queryable too. Every transition is a row rather than a gap where a row used to be, so "what happened to this contract" is a question the system can answer.
What it costs
Nothing ever leaves. The table only grows, and every state a contract ever occupied is permanent history you're now storing and paging past. At this scale that's a rounding error; at a much larger one it's a partitioning and archival problem you've deferred rather than avoided.
Every query needs to care. Once "gone" isn't a state, no query can assume the rows it sees are live ones. Every read path has to filter on state, and the failure mode is a forgotten filter quietly showing suspended contracts as active — which is a subtler bug than a missing row and harder to notice.
"Deleted" still means something to someone. A customer asking for their data to be removed is not asking for a status change. Soft-delete-everything and a real erasure requirement will collide eventually, and the answer has to be a deliberate destructive path rather than the discovery that none exists.
Where a different shape fits
- Genuine erasure obligations. Where deletion is a legal requirement rather than a UI verb, the model needs a real removal path — ideally narrow, audited, and deliberately awkward to invoke, so it stays distinct from the everyday cancel button.
- High-volume, low-value rows. Lifecycle history earns its storage when each row represents a commercial relationship. For event or telemetry data it doesn't, and a retention window beats a state machine.
- The audit trail is the requirement. If you need to know not just the current state but who changed it and when, a status column isn't enough — that's an append-only transition log, with current state derived from it. Strictly better, and considerably more machinery.
Make the destructive-feeling action non-destructive by construction, and the question stops being whether someone will click the wrong thing. They will. It stops mattering.