Command Palette

Search for a command to run...

Writing / Infra

The topology that isn't there: what one box actually costs

Aug 28, 2026·4 min read
DockerAWSInfra

The default shape for "real infrastructure" is several managed services: containers on managed compute, a managed database, a managed cache, a load balancer in front. This site runs Next.js, a Go API, Redis, and TimescaleDB as Docker Compose services on one EC2 instance behind Caddy.

That's a deliberate choice rather than a compromise, but the case for it is usually made dishonestly — as though the alternative buys nothing. It buys plenty. It just doesn't buy anything this workload needs, and the cost of not having it is specific enough to write down.

Sizing to the traffic that exists

A personal site's traffic is small and predictable. Two vCPUs and two gigabytes of memory run all five services with room left over, and the marginal cost of the entire market-data pipeline is approximately zero, because the box was already there for the site.

Horizontal scaling solves two problems: surviving load one instance can't take, and isolating failure between services with genuinely different risk profiles. Neither is in evidence here. Provisioning for them anyway would mean more AWS surface to secure, more independently failing parts, and a monthly bill for headroom that sits idle — in exchange for nothing measurable.

What it costs

Every deploy is a small outage. docker compose up -d replaces containers in place. It's seconds, and nobody notices, but "zero-downtime deploys" is not a property this has, and claiming otherwise would be a lie about a real limitation.

One instance is one failure domain. An AZ event, a bad kernel update, or an instance retirement takes down the site, the API, the database and the cache simultaneously. Recovery is Terraform plus a deploy, which is fast and rehearsed, but it is not automatic and it is not instant.

The services contend. Postgres and Redis and a Node process share two gigabytes. Nothing enforces a boundary between them, so a memory-hungry query and a traffic spike can interfere in ways that separate services simply couldn't. At this scale that headroom is comfortable; it is not guaranteed.

Data durability is bounded rather than guaranteed. The database sits on a separate EBS volume, so it survives an instance replacement, and a nightly pg_dump goes to S3 from the host crontab using the instance's own role. That puts the recovery point objective at roughly twenty-four hours. It's an accepted bound rather than an oversight, but it is a bound: there's no point-in-time recovery and no standby, so losing the volume between two backups loses a day.

The stream fan-out can't scale past this instance. The SSE hub holds connections in one process. The coordination for running more than one ingesting replica exists and is tested, but the box it would scale onto doesn't.

What makes it survivable anyway

The failure modes above are bounded rather than open-ended, and that's mostly down to unglamorous choices. Every service is restart: unless-stopped. Timescale has a real healthcheck, and the services that depend on it wait for it rather than racing it. The migration step runs as a one-shot that must complete successfully before the API starts. The whole instance is described in Terraform, so recreating it is a command rather than an archaeology project, and what runs on it is a single Compose file in version control.

None of that prevents an outage. It makes the recovery boring, which at this scale is the property actually worth buying.

When I'd split it

Concretely, rather than as a principle:

  • When a deploy dropping requests starts mattering — that's two instances and a load balancer, and it's the first thing I'd add.
  • When the database's failure profile diverges from everything else's. A managed Postgres buys automated backups, point-in-time recovery, and a failover story that a cron job and a dump file do not. That's the upgrade I'd make first if this held data anyone depended on.
  • When one service's resource appetite becomes another's problem — that's the signal that shared memory has stopped being free.
  • When more than one person deploys. Single-box operations rely on nobody else needing to do something at the same time.

The argument for one box isn't that distributed systems are overkill. It's that the right topology is the smallest one whose failure modes you can name — and every one above is named, bounded, and priced.