← Back to blog

2026-08-27 · 9 min read

Disaster Recovery on AWS: RTO, RPO, and Multi-Region

A practical guide to AWS disaster recovery, defining RTO/RPO, choosing between backup-restore, pilot light, warm standby, and active-active, and what each really costs.

#aws#sre#disaster-recovery#reliability#architecture
Disaster Recovery on AWS: RTO, RPO, and Multi-Region

Disaster Recovery on AWS: RTO, RPO, and Multi-Region

"Make it highly available" is not a disaster-recovery plan. DR is a business decision dressed up as an architecture one: how much downtime and data loss can you tolerate, and how much are you willing to pay to shrink those numbers? Get the two target numbers right first, then pick a strategy.

RTO and RPO: the two numbers that decide everything

  • RTO: Recovery Time Objective: how long you can be down. "We must be back within 1 hour."
  • RPO: Recovery Point Objective: how much data you can afford to lose. "We can lose at most 5 minutes of writes."

These come from the business, not from engineering. A marketing site might be fine with RTO = 24h, RPO = 24h. A payments ledger might need RTO = minutes, RPO = zero. The tighter the numbers, the more the architecture costs, so don't gold-plate the brochure site.

The four AWS DR strategies

AWS frames DR as four tiers, trading cost against RTO/RPO:

1. Backup & Restore (RTO/RPO: hours)

Back up data and infra-as-code to another region. On disaster, redeploy from Terraform/CloudFormation and restore from backups. Cheapest. Slow recovery. Good for non-critical workloads.

  • AWS Backup with cross-region copy, S3 cross-region replication, AMI/EBS snapshot copies.
  • Your Terraform is your recovery plan, if the infra is in code, rebuilding is terraform apply in a new region.

2. Pilot Light (RTO: tens of minutes, RPO: minutes)

Core data is continuously replicated to the DR region, but compute is off (or minimal). On disaster, you "turn up the lights", scale compute from zero and point traffic over.

  • RDS cross-region read replica (promote on failover), data pre-replicated.
  • App tier defined but scaled to zero; scale up during failover.

3. Warm Standby (RTO: minutes, RPO: seconds)

A scaled-down but always-running copy of the full stack in the DR region. On disaster you scale it up and shift traffic. Faster than pilot light because everything's already running.

4. Active-Active / Multi-Region (RTO/RPO: near zero)

Traffic served from multiple regions simultaneously. A region failure just means routing away from it. Most expensive and most complex: you're paying for and operating full capacity in 2+ regions, plus solving cross-region data consistency.

A decision table

StrategyRTORPORelative costUse when
Backup & RestoreHoursHours$Non-critical, cost-sensitive
Pilot Light10s of minMinutes$$Important, some downtime OK
Warm StandbyMinutesSeconds$$$Business-critical
Active-Active~0~0$$$$Can't tolerate downtime at all

Most teams over-buy here. Pilot light covers a surprising number of real businesses: your data is safe and continuously replicated, and you can be back in well under an hour without paying to run a second full stack 24/7.

The hard part: data, not compute

Standing up compute in another region is easy. The genuinely hard problems are:

  • Database failover. RDS cross-region replicas have replication lag (your RPO). Aurora Global Database tightens it to ~1s. Promotion isn't instant, test it.
  • Data consistency in active-active. Two regions taking writes means conflict resolution. DynamoDB Global Tables handle this with last-writer-wins; relational stores are much harder.
  • Stateful dependencies. Caches, queues, search indexes: each needs a replication or rebuild story.

Routing the failover

  • Route 53 health checks + failover routing flip DNS from primary to DR region. Mind the TTL: low TTL = faster failover but more DNS queries.
  • Global Accelerator fails over faster than DNS (no client cache to wait out) for latency-sensitive apps.

The rule that actually matters: test it

A DR plan you've never executed is a hypothesis, not a plan. The number of teams whose first real failover is the disaster is alarming. Run game days:

  • Schedule regular DR drills, actually fail over, serve real traffic from DR, fail back.
  • Measure your real RTO/RPO against the targets. They're usually worse than you assumed.
  • Automate the runbook so failover isn't 40 manual steps under pressure, capture it like an incident runbook.

The short version

  • RTO/RPO come from the business; everything else follows.
  • Four tiers: backup-restore → pilot light → warm standby → active-active, increasing cost and decreasing RTO/RPO.
  • Pilot light is the sweet spot for many "important but not zero-downtime" systems.
  • Data replication and failover are the hard parts, not compute.
  • Test it with game days, or you don't really have a plan.

Designing for resilience without overspending is exactly the kind of work I do, see my services or get in touch.

Share:LinkedInXWhatsApp

Related articles

Reactions & comments