“How available does this need to be?” is unanswerable as asked. “How long may it be down, and how much data may we lose?” produces two numbers, and those two numbers determine the architecture and its cost. Every disaster recovery conversation should start by converting the first question into the second.
Topic 1: RTO, RPO and What They Cost
RTO — Recovery Time Objective. The maximum acceptable time from failure to service restored.
RPO — Recovery Point Objective. The maximum acceptable data loss, measured in time. An RPO of one hour means you can lose the last hour of writes.
failure
│
←── RPO ──────┤────── RTO ──────→
last good │ service
recoverable │ restored
state │
Two independent dials, and they are set by different mechanisms: RPO is a function of replication and backup frequency, RTO is a function of how fast you can bring capacity up and repoint traffic. Continuous replication buys a near-zero RPO even if your RTO is hours; a warm standby buys a short RTO even if you only replicate every fifteen minutes.
Both are business decisions, not engineering preferences. Get them in writing from whoever owns the revenue, because the cost curve is steep: an RTO of four hours might be a backup and a runbook, while an RTO of one minute is two live regions and a multi-region data strategy. Without those two numbers, “make it highly available” reliably produces either too little or far too much.
Tiering is what keeps this affordable. A payments path may justify active/active; an internal reporting dashboard almost certainly does not. Assign a tier per workload and apply the matching strategy — a single organisation-wide RTO is always wrong somewhere.
Topic 2: The Four Strategies
Backup & restore — RTO hours to days, cheapest. Backups replicated to another region; on disaster, you rebuild infrastructure from code and restore data. Viable only if the infrastructure genuinely is code and someone has run the rebuild recently.
Pilot light — RTO tens of minutes. Data replicates continuously to the DR region; compute exists but is switched off (AMIs built, ASGs at zero, databases as small replicas). On disaster, scale up and promote.
Warm standby — RTO minutes. A scaled-down but running copy of the full stack in the second region, taking little or no traffic. On disaster, scale it up and shift DNS. Costs meaningfully more; recovers without needing capacity to appear on demand — which matters, because the control plane you would rely on is exactly what a large regional event degrades.
Active/active (multi-site) — RTO near zero. Both regions serve production. The infrastructure is straightforward; the data layer is where the difficulty lives, and it is substantial.
The uncomfortable truth about active/active: it is a data consistency problem wearing an infrastructure costume. Two regions accepting writes to the same records need conflict resolution, and conflict resolution is where correctness bugs are born. The patterns that work in practice avoid the problem rather than solving it — partition by tenant or geography so each record has one writing region, or designate a home region for writes and serve reads everywhere.
Topic 3: Multi-AZ First, and Usually Enough
Before any multi-region conversation, be honest about whether multi-AZ is done properly. Most “we need multi-region” requirements are actually “we have a single point of failure in one AZ”.
The multi-AZ checklist that catches real gaps:
□ Instances spread across ≥3 AZs, and actually running in all of them
□ NAT gateway per AZ, each private route table pointing at its own
□ Load balancer enabled in every AZ that has targets
□ RDS Multi-AZ, or Aurora with readers in other AZs
□ Any single-AZ EFS or One Zone-IA data identified and accepted
□ Quotas checked: can you run in 2 AZs the capacity you run in 3?
□ Tested: an AZ actually removed, on purpose, and the result observed
That last line is the one that is almost never done, and it is the only one that produces evidence. AWS Fault Injection Service can remove an AZ’s network reachability, stop instances, throttle an API, or inject latency — as a controlled experiment with a stop condition wired to your own alarms. Running one AZ-failure experiment per quarter, in production, with a rollback, is the difference between believing you are multi-AZ and knowing it.
The quota line matters more than it looks: losing an AZ means the remaining ones must run more instances than usual, and service quotas are per region. A quota sized for steady state becomes the reason the recovery stalls.
Topic 4: The Mechanisms
| Layer | Cross-region option | RPO | Notes |
|---|---|---|---|
| S3 | Cross-Region Replication | Minutes; 15 min with RTC | New objects only unless you run Batch Replication |
| RDS | Cross-region read replica | Seconds to minutes | Manual promotion |
| Aurora | Global Database | ~1 second | Managed promotion, RTO ~1 minute |
| DynamoDB | Global Tables | ~1 second | Multi-writer, last-writer-wins |
| EBS | Snapshot copy | Hours | Or Elastic Disaster Recovery for continuous block replication |
| AMIs | Copy to DR region | n/a | Must be part of the pipeline, not a manual step |
| Secrets | Multi-region secrets | Seconds | Easy to forget until the failover |
| ECR | Cross-region replication | Minutes | A DR region with no images cannot start anything |
| Route 53 | Health checks + failover records | n/a | This is the switch |
The last four rows are where drills fail. Everyone replicates the database; the region without AMIs, container images, secrets or a working TLS certificate is discovered at the worst moment. A DR region is a full environment or it is a plan on paper.
Route 53 is the traffic switch. Failover routing with health checks moves traffic automatically; the delay is the health check interval plus the record TTL. Keep the TTL low (60s) on records that participate in failover, and remember that some resolvers ignore short TTLs — which is why Global Accelerator, using anycast IPs that stay constant while the backend changes, gives a faster and more reliable switch when you can justify it.
Topic 5: Static Stability
The principle that separates designs that survive a large event from ones that fail during recovery: a recovery that depends on the control plane depends on the thing most likely to be degraded.
Concretely:
- Capacity that is already running is more reliable than capacity you plan to launch. Warm standby beats pilot light for this reason alone, independent of RTO.
- Pre-provision what the recovery needs: IPs, ENIs, load balancers, quotas. Creating them during an event is a dependency on an API that may be busy.
- Avoid recovery paths that require an AWS API you cannot test. If your runbook says “increase the quota”, the recovery includes a support case.
- Over-provision to survive the loss. Three AZs each running at 50% survive losing one; three running at 90% do not, no matter how fast the autoscaler is.
The same logic applies to your own dependencies: a failover that requires your CI system, your secret store, and your identity provider all to be healthy has three additional single points of failure, and they are rarely in the DR plan.
Topic 6: Drills, and the Runbook They Produce
An untested DR plan does not have a slow RTO. It has an unknown RTO, which is worse, because you cannot plan around a number you do not have.
A drill schedule that is actually sustainable:
Monthly Restore one backup to a scratch environment and verify the data.
Rotates through services. Catches the backup that has silently
been failing for six weeks.
Quarterly AZ failure experiment with FIS, in production, with a stop condition.
Annually Full regional failover drill. Announced, scheduled, with the
runbook open and a scribe recording every step and its duration.
Every drill produces two artifacts: an updated runbook and a list of what surprised you. The surprises are the value. Common ones, from real drills: the DR region lacked a container image; the secret existed but the KMS key it used did not; the certificate was regional and had never been issued in the DR region; the runbook named a person who left; the quota in the DR region was the default because nobody had ever launched anything there.
Write the runbook for the tired version of yourself. Exact commands, exact ARNs, explicit decision points (“if X, do Y; if not, escalate to Z”), and a communications plan naming who tells customers what. A runbook that says “restore the database” is a note, not a procedure.
Try it yourself: pick your least critical production service and restore it from backup into a scratch environment, end to end, while timing each step. The number you get is your real RTO for that workload — and it is almost always larger than the number in the document.
Common mistake: measuring RTO from “restore complete” rather than from “users are being served”. The database restore is often the fastest part; the time goes into finding the right backup, waiting for DNS, discovering that the application needs a configuration change to point at a new endpoint, and the twenty minutes of deciding whether to fail over at all. Time the whole thing, decision included, or the number is fiction.