Command Palette

Search for a command to run...

Writing / Infra

A box you can't log into: Terraform, SSM, and deploys without SSH

Aug 27, 2026·4 min read
AWSTerraformDockerInfra

Everything on this domain — the site, the market dashboard, the Go API behind it, Redis, and TimescaleDB — runs as Docker Compose services on a single Terraform-provisioned t4g.small. This post is about how that box is built and deployed to, rather than why it's one box.

Nothing listens but 443

The security group has exactly one ingress rule: TCP 443 from anywhere. No port 22, and deliberately no port 80 either. Operator access is aws ssm start-session, which works because Amazon Linux 2023 ships the SSM agent and the instance profile grants it — a shell without an open port, authenticated by IAM rather than a key file somebody has to hold.

Closing port 80 has a consequence worth knowing before you hit it. Automatic certificate issuance usually leans on the HTTP-01 challenge, which needs port 80. With only 443 open, TLS-ALPN-01 is the only challenge that can complete, and it works on 443 alone. Caddy negotiates that without any extra configuration — the requirement is simply not to map port 80, which was already the case.

Graviton, and the images that didn't match it

The instance is ARM64: an AL2023 arm64 AMI on a t4g instance, which is the cheapest way to buy this much memory on AWS.

This is also where the deploy path bit hardest. CI built container images on GitHub's runners, which are amd64, and pushed them to a registry an ARM box then pulled from. The containers crash-looped with exec format error — a failure that says nothing about architecture unless you already know that's what it means. The fix was QEMU plus Buildx and an explicit --platform linux/arm64 on the build, which is one line and entirely obvious in retrospect.

The instance boots empty on purpose

user_data installs Docker and the Compose plugin, formats and mounts a data volume, and stops. It starts no containers.

That's deliberate: what runs on the box is owned entirely by the deploy workflow, so there's one place that decides, not two racing each other at boot. It also sidesteps a trap — user_data only executes on an instance's first boot, so anything added to it later doesn't reach a box that's already running. The Terraform sets user_data_replace_on_change = false to make that explicit, because a destroy-and-recreate is never the right response to editing a boot script on a live instance. Changes to it apply to the next instance, and to this one only via a matching one-time command run through Session Manager.

Deploying to a box with no inbound path

The workflow assumes an AWS role over OIDC, uploads docker-compose.yml, the Caddyfile, and the backup script to an S3 release bucket, and then sends the box a command through SSM:

aws s3 cp …                       # pull the compose files down
aws ecr get-login-password | …    # authenticate to the registry
aws ssm get-parameter --with-decryption   # fetch the DB password
docker compose pull && docker compose up -d --remove-orphans

Three things about that are load-bearing. The instance is resolved by its Name tag rather than a hardcoded ID, so replacing the box needs no workflow change. The database password is read fresh from an SSM SecureString on every deploy and passed as an environment variable — it's never written to a file, never committed, and no human has ever seen it. And nothing in the sequence needs an inbound connection: SSM is an outbound agent poll, which is precisely why the security group can stay closed.

The data is not on the root volume

Timescale's data directory is a bind mount onto a separate gp3 EBS volume, attached and mounted at /mnt/data, so it survives an instance replacement instead of dying with the root disk. The formatting step in user_data is guarded by blkid, so re-running it can never reformat a volume that already has a filesystem on it.

The root volume is 30GB, which isn't a capacity decision — the AL2023 ARM64 AMI's snapshot is 30GB, and EBS won't create a volume smaller than the snapshot it's built from. The first attempt asked for 20 and RunInstances simply refused.

Where a different shape fits

  • More than one environment. The default VPC is fine for a single box with one public service. A second environment wants its own VPC and subnets, if only so a security-group edit can't reach across environments.
  • Deploys that can't drop requests. docker compose up -d replaces containers in place, which is a brief outage. Zero-downtime needs two instances and something in front of them, and that's the point where a load balancer stops being ceremony.
  • A team. SSM Session Manager scales to more operators than SSH keys do — that's an argument for it, not against — but the moment more than one person deploys, the release bucket wants versioning and the SSM command wants an approval gate rather than firing on every push to main.

None of the above is exotic. It's mostly the consequence of one decision — close port 22 — followed honestly to the places it leads.