← Back to blog

2026-09-08 · 12 min read

Disposable EC2: How SSM and S3 Give You Kubernetes Resilience Without Kubernetes

A practical architecture pattern where a single EC2 instance can be destroyed and recreated with zero data loss, using SSM Parameter Store for secrets and S3 for backups. No orchestrator required.

#aws#ssm#s3#terraform#ec2#infrastructure#devops#disaster-recovery
Disposable EC2: How SSM and S3 Give You Kubernetes Resilience Without Kubernetes

Kubernetes popularised the idea of cattle-not-pets infrastructure: containers that can die and be replaced without anyone noticing. But running Kubernetes costs money, brainpower, and operational overhead that many startups and small teams can't justify.

What if you could get the same resilience property on a single EC2 instance running Docker Compose? Destroy it, recreate it, and your data comes back automatically. No orchestrator, no etcd, no control plane. Just Terraform, SSM Parameter Store, and S3.

I've been running this pattern in production for WeFoundIt, a 434-endpoint FastAPI platform serving users in Cameroon. Here's exactly how it works.

The Architecture

The core insight: separate state from compute. An EC2 instance is stateless if you can answer "yes" to: can I delete this box right now and lose nothing important?

                Terraform State (local or remote)
                         |
                    terraform apply
                         |
         +----- EC2 Instance (ephemeral) -----+
         |                                     |
         |  user_data.sh bootstraps on boot:   |
         |    1. Install Docker + awscli       |
         |    2. Clone repos from GitHub       |
         |    3. Fetch .env from SSM           |
         |    4. Restore DB from S3            |
         |    5. Run migrations                |
         |    6. Start services                |
         |    7. Setup backup cron             |
         |                                     |
         +-------------------------------------+
                    |             |
          SSM Parameter Store    S3 Backup Bucket
          (secrets persist)      (data persists)

Three things survive instance destruction:

  1. SSM Parameter Store holds the entire .env file as a SecureString (KMS-encrypted)
  2. S3 holds daily database backups (gzipped pg_dump) and uploaded files
  3. GitHub holds the application code

Everything else is derived from these three sources on boot.

The Bootstrap: user_data.sh

When Terraform creates the EC2 instance, the user_data template runs exactly once on first boot. Here's the flow with the important decisions annotated:

#!/bin/bash
set -euo pipefail
exec > /var/log/wefoundit-bootstrap.log 2>&1

# Install Docker, Compose, AWS CLI, git
apt-get update -y
apt-get install -y docker.io docker-compose-plugin awscli git curl jq
systemctl enable docker --now

# Clone application repos
git clone $REPO_URL /opt/wefoundit/wefoundit-api
git clone $WEB_REPO_URL /opt/wefoundit/wefoundit-web

# Fetch secrets from SSM with retry logic
for attempt in 1 2 3; do
  if aws ssm get-parameter --name "/wefoundit/env" --with-decryption \
    --query Parameter.Value --output text --region eu-west-1 > .env; then
    break
  fi
  sleep 10
done

# Check S3 for existing backup
if aws s3 ls "s3://$BUCKET/db/latest.sql.gz"; then
  aws s3 cp "s3://$BUCKET/db/latest.sql.gz" /tmp/restore.sql.gz
  RESTORE=true
fi

# Start database, wait for readiness
docker compose up -d db redis
for i in $(seq 1 30); do
  docker compose exec -T db pg_isready -U postgres && break
  sleep 2
done

# Restore if backup exists
if [ "$RESTORE" = "true" ]; then
  gunzip -c /tmp/restore.sql.gz | docker compose exec -T db psql -U postgres -d foundit_db
fi

# Apply schema migrations (handles backup-to-code drift)
docker compose up -d api
docker compose exec -T api alembic upgrade head

# Start everything
docker compose up -d

The key decisions here:

  • Retry on SSM fetch: SSM can be briefly unavailable during instance init. Three attempts with 10-second backoff handles transient failures.
  • Alembic after restore: The backup might be from yesterday, but the code was just cloned from main. Running alembic upgrade head applies any migrations committed since the backup.
  • Backup check before restore: If this is a truly fresh deployment (no prior data), skip the restore and start with an empty database.

Why SSM Parameter Store (Not Secrets Manager)

The entire .env file is stored as one SSM SecureString parameter at /wefoundit/env. One parameter, one get, one write-to-disk.

# Write (manual, when secrets change)
aws ssm put-parameter --name /wefoundit/env --type SecureString \
  --overwrite --value file://.env --region eu-west-1

# Read (automated at boot via IAM instance role)
aws ssm get-parameter --name /wefoundit/env --with-decryption \
  --query Parameter.Value --output text > .env

Why not Secrets Manager? Three reasons:

  1. Cost: Standard tier Parameter Store is free. Secrets Manager would be $0.40/month minimum.
  2. Simplicity: One blob parameter, one CLI command. No per-secret management overhead.
  3. No rotation needed: Our secrets are third-party API keys (Fapshi, Google, Firebase) that can't be auto-rotated anyway.

The IAM role is scoped to exactly one parameter:

resource "aws_iam_role_policy" "wefoundit" {
  policy = jsonencode({
    Statement = [{
      Effect   = "Allow"
      Action   = ["ssm:GetParameter"]
      Resource = ["arn:aws:ssm:eu-west-1:ACCOUNT:parameter/wefoundit/env"]
    }]
  })
}

The Backup Strategy

Daily at 3 AM, a cron job runs:

docker compose exec -T db pg_dump -U postgres -d foundit_db \
  --no-owner --no-acl --clean --if-exists | gzip > /tmp/backup.sql.gz

aws s3 cp /tmp/backup.sql.gz "s3://$BUCKET/db/latest.sql.gz"
aws s3 cp /tmp/backup.sql.gz "s3://$BUCKET/db/backup_$(date +%Y%m%d).sql.gz"

The S3 bucket has:

  • Versioning enabled: Even latest.sql.gz has full version history
  • 90-day lifecycle: Old timestamped backups auto-expire
  • prevent_destroy lifecycle rule: Terraform cannot accidentally delete the bucket

This last point is critical. When you run terraform destroy, it deletes the EC2 instance, the EIP, the security group, and the IAM role. But the S3 bucket survives because it contains objects and has prevent_destroy set.

The Destroy/Recreate Cycle

# Intentional rebuild
terraform destroy    # EC2 gone, S3 bucket survives
terraform apply      # New EC2 boots, user_data finds backup in S3, restores

# Update DNS A records to new EIP (only manual step)

Total data loss: zero. The database is restored from S3. Secrets come from SSM. Code comes from GitHub. The instance is truly disposable.

Day-2 Operations via SSM

No SSH needed. Everything is done through SSM Systems Manager:

# Deploy new code (no SSH, no open ports)
aws ssm send-command --instance-ids "i-xxxxx" \
  --document-name "AWS-RunShellScript" \
  --parameters 'commands=["cd /opt/wefoundit && sudo bash wefoundit-api/deploy/scripts/update.sh"]'

# Interactive shell (replaces SSH entirely)
aws ssm start-session --target i-xxxxx --region eu-west-1

# Emergency backup
aws ssm send-command --instance-ids "i-xxxxx" \
  --document-name "AWS-RunShellScript" \
  --parameters 'commands=["sudo bash /opt/wefoundit/wefoundit-api/deploy/scripts/backup_s3.sh"]'

SSH port 22 is open as a fallback but restricted to a single IP. Day-to-day operations are entirely SSM-based, which means IAM-authenticated, CloudTrail-logged, no key management.

Cost

ResourceMonthly
EC2 t3.small~$15.50
EBS 30 GB gp3~$2.40
S3 backups< $0.10
SSM Parameter Store$0.00
Total~$18/month

Compare that to managed Kubernetes (EKS: $73/month just for the control plane) or even a basic managed database (RDS: $13+/month for the smallest instance).

When This Pattern Breaks Down

Be honest about the limitations:

  • Multi-instance: If you need horizontal scaling, this pattern doesn't work. You need a load balancer and stateless app servers.
  • Sub-minute RTO: A full rebuild takes 5-7 minutes (install + clone + restore). If you need sub-minute recovery, you need a hot standby.
  • Large databases: A 50GB database takes too long to restore from S3 on boot. At that size, use RDS with automated backups.
  • Team operations: With multiple engineers deploying independently, you need proper CI/CD, not SSM send-command.

The Terraform

The full infrastructure is approximately 100 lines of Terraform:

# Core resources
resource "aws_instance" "wefoundit" {
  ami                  = data.aws_ami.ubuntu.id
  instance_type        = var.instance_type
  iam_instance_profile = aws_iam_instance_profile.wefoundit.name
  user_data            = templatefile("user_data.sh.tftpl", { ... })
  user_data_replace_on_change = true
  lifecycle { ignore_changes = [ami] }
}

resource "aws_s3_bucket" "backups" {
  bucket = "wefoundit-backups-${data.aws_caller_identity.current.account_id}"
  lifecycle { prevent_destroy = true }
}

user_data_replace_on_change = true means: if the bootstrap script changes in your Terraform, the instance is destroyed and recreated automatically. Combined with the S3 restore, this means infrastructure changes are applied by rebuilding from scratch rather than patching in place.

Key Takeaways

  1. Separate state from compute. If your secrets and data live outside your instance, the instance becomes disposable.
  2. SSM Parameter Store is underrated. Free encryption, IAM-scoped access, and zero infrastructure to manage.
  3. S3 is your persistence layer. Versioned, lifecycle-managed, and survives Terraform destroy.
  4. You don't need Kubernetes to get cattle-not-pets. A well-designed user_data script gives you the same rebuild guarantee.
  5. Daily backups + auto-restore-on-boot = zero-downtime rebuilds. The scariest Terraform operation (destroy) becomes routine.

This pattern has been running in production for months. The peace of mind of knowing I can terraform destroy && terraform apply without losing a single row of data is worth the engineering investment ten times over.


The full Terraform, bootstrap template, and backup scripts are in the wefoundit-api repo.

Share:LinkedInXWhatsApp

Related articles

Reactions & comments