Writing / Infra
A box you can't log into: Terraform, SSM, and deploys without SSH
Everything on this domain — the site, the market dashboard, the Go API behind it, Redis, and
TimescaleDB — runs as Docker Compose services on a single Terraform-provisioned t4g.small. This
post is about how that box is built and deployed to, rather than why it's one box.
Nothing listens but 443
The security group has exactly one ingress rule: TCP 443 from anywhere. No port 22, and
deliberately no port 80 either. Operator access is aws ssm start-session, which works because
Amazon Linux 2023 ships the SSM agent and the instance profile grants it — a shell without an open
port, authenticated by IAM rather than a key file somebody has to hold.
Closing port 80 has a consequence worth knowing before you hit it. Automatic certificate issuance usually leans on the HTTP-01 challenge, which needs port 80. With only 443 open, TLS-ALPN-01 is the only challenge that can complete, and it works on 443 alone. Caddy negotiates that without any extra configuration — the requirement is simply not to map port 80, which was already the case.
Graviton, and the images that didn't match it
The instance is ARM64: an AL2023 arm64 AMI on a t4g instance, which is the cheapest way to buy
this much memory on AWS.
This is also where the deploy path bit hardest. CI built container images on GitHub's runners,
which are amd64, and pushed them to a registry an ARM box then pulled from. The containers
crash-looped with exec format error — a failure that says nothing about architecture unless you
already know that's what it means. The fix was QEMU plus Buildx and an explicit
--platform linux/arm64 on the build, which is one line and entirely obvious in retrospect.
The instance boots empty on purpose
user_data installs Docker and the Compose plugin, formats and mounts a data volume, and stops.
It starts no containers.
That's deliberate: what runs on the box is owned entirely by the deploy workflow, so there's one
place that decides, not two racing each other at boot. It also sidesteps a trap — user_data only
executes on an instance's first boot, so anything added to it later doesn't reach a box that's
already running. The Terraform sets user_data_replace_on_change = false to make that explicit,
because a destroy-and-recreate is never the right response to editing a boot script on a live
instance. Changes to it apply to the next instance, and to this one only via a matching one-time
command run through Session Manager.
Deploying to a box with no inbound path
The workflow assumes an AWS role over OIDC, uploads docker-compose.yml, the Caddyfile, and the
backup script to an S3 release bucket, and then sends the box a command through SSM:
aws s3 cp … # pull the compose files down
aws ecr get-login-password | … # authenticate to the registry
aws ssm get-parameter --with-decryption # fetch the DB password
docker compose pull && docker compose up -d --remove-orphans
Three things about that are load-bearing. The instance is resolved by its Name tag rather than a
hardcoded ID, so replacing the box needs no workflow change. The database password is read fresh
from an SSM SecureString on every deploy and passed as an environment variable — it's never
written to a file, never committed, and no human has ever seen it. And nothing in the sequence
needs an inbound connection: SSM is an outbound agent poll, which is precisely why the security
group can stay closed.
The data is not on the root volume
Timescale's data directory is a bind mount onto a separate gp3 EBS volume, attached and mounted at
/mnt/data, so it survives an instance replacement instead of dying with the root disk. The
formatting step in user_data is guarded by blkid, so re-running it can never reformat a volume
that already has a filesystem on it.
The root volume is 30GB, which isn't a capacity decision — the AL2023 ARM64 AMI's snapshot is
30GB, and EBS won't create a volume smaller than the snapshot it's built from. The first attempt
asked for 20 and RunInstances simply refused.
Where a different shape fits
- More than one environment. The default VPC is fine for a single box with one public service. A second environment wants its own VPC and subnets, if only so a security-group edit can't reach across environments.
- Deploys that can't drop requests.
docker compose up -dreplaces containers in place, which is a brief outage. Zero-downtime needs two instances and something in front of them, and that's the point where a load balancer stops being ceremony. - A team. SSM Session Manager scales to more operators than SSH keys do — that's an argument for it, not against — but the moment more than one person deploys, the release bucket wants versioning and the SSM command wants an approval gate rather than firing on every push to main.
None of the above is exotic. It's mostly the consequence of one decision — close port 22 — followed honestly to the places it leads.