2026-08-27 · 9 min read
Disaster Recovery on AWS: RTO, RPO, and Multi-Region
A practical guide to AWS disaster recovery, defining RTO/RPO, choosing between backup-restore, pilot light, warm standby, and active-active, and what each really costs.

Disaster Recovery on AWS: RTO, RPO, and Multi-Region
"Make it highly available" is not a disaster-recovery plan. DR is a business decision dressed up as an architecture one: how much downtime and data loss can you tolerate, and how much are you willing to pay to shrink those numbers? Get the two target numbers right first, then pick a strategy.
RTO and RPO: the two numbers that decide everything
- RTO: Recovery Time Objective: how long you can be down. "We must be back within 1 hour."
- RPO: Recovery Point Objective: how much data you can afford to lose. "We can lose at most 5 minutes of writes."
These come from the business, not from engineering. A marketing site might be fine with RTO = 24h, RPO = 24h. A payments ledger might need RTO = minutes, RPO = zero. The tighter the numbers, the more the architecture costs, so don't gold-plate the brochure site.
The four AWS DR strategies
AWS frames DR as four tiers, trading cost against RTO/RPO:
1. Backup & Restore (RTO/RPO: hours)
Back up data and infra-as-code to another region. On disaster, redeploy from Terraform/CloudFormation and restore from backups. Cheapest. Slow recovery. Good for non-critical workloads.
- AWS Backup with cross-region copy, S3 cross-region replication, AMI/EBS snapshot copies.
- Your Terraform is your recovery
plan, if the infra is in code, rebuilding is
terraform applyin a new region.
2. Pilot Light (RTO: tens of minutes, RPO: minutes)
Core data is continuously replicated to the DR region, but compute is off (or minimal). On disaster, you "turn up the lights", scale compute from zero and point traffic over.
- RDS cross-region read replica (promote on failover), data pre-replicated.
- App tier defined but scaled to zero; scale up during failover.
3. Warm Standby (RTO: minutes, RPO: seconds)
A scaled-down but always-running copy of the full stack in the DR region. On disaster you scale it up and shift traffic. Faster than pilot light because everything's already running.
4. Active-Active / Multi-Region (RTO/RPO: near zero)
Traffic served from multiple regions simultaneously. A region failure just means routing away from it. Most expensive and most complex: you're paying for and operating full capacity in 2+ regions, plus solving cross-region data consistency.
A decision table
| Strategy | RTO | RPO | Relative cost | Use when |
|---|---|---|---|---|
| Backup & Restore | Hours | Hours | $ | Non-critical, cost-sensitive |
| Pilot Light | 10s of min | Minutes | $$ | Important, some downtime OK |
| Warm Standby | Minutes | Seconds | $$$ | Business-critical |
| Active-Active | ~0 | ~0 | $$$$ | Can't tolerate downtime at all |
Most teams over-buy here. Pilot light covers a surprising number of real businesses: your data is safe and continuously replicated, and you can be back in well under an hour without paying to run a second full stack 24/7.
The hard part: data, not compute
Standing up compute in another region is easy. The genuinely hard problems are:
- Database failover. RDS cross-region replicas have replication lag (your RPO). Aurora Global Database tightens it to ~1s. Promotion isn't instant, test it.
- Data consistency in active-active. Two regions taking writes means conflict resolution. DynamoDB Global Tables handle this with last-writer-wins; relational stores are much harder.
- Stateful dependencies. Caches, queues, search indexes: each needs a replication or rebuild story.
Routing the failover
- Route 53 health checks + failover routing flip DNS from primary to DR region. Mind the TTL: low TTL = faster failover but more DNS queries.
- Global Accelerator fails over faster than DNS (no client cache to wait out) for latency-sensitive apps.
The rule that actually matters: test it
A DR plan you've never executed is a hypothesis, not a plan. The number of teams whose first real failover is the disaster is alarming. Run game days:
- Schedule regular DR drills, actually fail over, serve real traffic from DR, fail back.
- Measure your real RTO/RPO against the targets. They're usually worse than you assumed.
- Automate the runbook so failover isn't 40 manual steps under pressure, capture it like an incident runbook.
The short version
- RTO/RPO come from the business; everything else follows.
- Four tiers: backup-restore → pilot light → warm standby → active-active, increasing cost and decreasing RTO/RPO.
- Pilot light is the sweet spot for many "important but not zero-downtime" systems.
- Data replication and failover are the hard parts, not compute.
- Test it with game days, or you don't really have a plan.
Designing for resilience without overspending is exactly the kind of work I do, see my services or get in touch.