The honest question: what can you survive already
99.9% availability is 43 minutes of downtime per month. That level is reachable with a single server, correct restarts and fast recovery from code. Kubernetes earns its keep when you have dozens of services, dozens of deploys a day and a team ready to feed the cluster itself. For 2–5 services on one VPS, k8s adds more failure modes than it removes.
Level 1: containers heal themselves
Most incidents are not “the server burned down” but “a process hung”. A couple of compose lines cover this: a healthcheck catches the hang, restart: always revives crashes, autoheal restarts whatever went unhealthy.
services:
app:
image: registry.example.com/app:latest
restart: always
healthcheck:
test: ['CMD', 'curl', '-sf', 'http://localhost:8000/health']
interval: 15s
retries: 3
start_period: 30s
autoheal:
image: willfarrell/autoheal:latest
restart: always
environment: { AUTOHEAL_CONTAINER_LABEL: all }
volumes: ['/var/run/docker.sock:/var/run/docker.sock']
/health endpoint must check dependencies (database, cache), not just return 200. Otherwise the healthcheck is green while the app serves 500s — a classic in "monitoring said nothing" post-mortems.Level 2: a standby server and switchover
Only a second server saves you from losing the server itself. The cheap, cluster-free scheme: a small standby VPS receiving data every few minutes (restic or rsync), with the stack described by the same compose file. Switchover is an A-record change with a low TTL (60 s) or the provider’s floating IP if there is one.
monitor (external) ── catches down ──▶ standby
1. restore data from the latest backup
2. docker compose up -d
3. switch floating IP / DNS (TTL 60)
total: minutes of downtime, pennies for standby
The whole chain can be automated with a script — one of our case studies is a site that migrates itself to a fresh VPS with no human involved in about a minute.
Checklist
Want to survive a dead server?
We'll design failover for your budget: from healthcheck-driven restarts to automatic switchover to a standby server.