Resilience and DR: RTO, RPO and the Four Strategies

Turning an availability conversation into two numbers, choosing the cheapest strategy that meets them, and why an untested plan has an unknown recovery time rather than a slow one.

advanced 22 min lesson hands-on task included

“How available does this need to be?” is unanswerable as asked. “How long may it be down, and how much data may we lose?” produces two numbers, and those two numbers determine the architecture and its cost. Every disaster recovery conversation should start by converting the first question into the second.


Topic 1: RTO, RPO and What They Cost

RTO — Recovery Time Objective. The maximum acceptable time from failure to service restored.

RPO — Recovery Point Objective. The maximum acceptable data loss, measured in time. An RPO of one hour means you can lose the last hour of writes.

             failure
                │
  ←── RPO ──────┤────── RTO ──────→
  last good     │              service
  recoverable   │              restored
  state         │

Two independent dials, and they are set by different mechanisms: RPO is a function of replication and backup frequency, RTO is a function of how fast you can bring capacity up and repoint traffic. Continuous replication buys a near-zero RPO even if your RTO is hours; a warm standby buys a short RTO even if you only replicate every fifteen minutes.

Both are business decisions, not engineering preferences. Get them in writing from whoever owns the revenue, because the cost curve is steep: an RTO of four hours might be a backup and a runbook, while an RTO of one minute is two live regions and a multi-region data strategy. Without those two numbers, “make it highly available” reliably produces either too little or far too much.

Tiering is what keeps this affordable. A payments path may justify active/active; an internal reporting dashboard almost certainly does not. Assign a tier per workload and apply the matching strategy — a single organisation-wide RTO is always wrong somewhere.


Topic 2: The Four Strategies

FOUR DR STRATEGIES — YOU ARE CHOOSING A PRICE FOR AN RTO RTO — how long until you are serving again → standing cost → Backup & restore RTO hours–days cheapest Pilot light RTO 10s of minutes data replicating, compute off Warm standby RTO minutes a small live copy, scaled up on failover Active/active RTO near zero two of everything, plus the hard part RPO IS A SEPARATE DIAL Continuous replication buys a near-zero RPO even with a slow RTO. AN UNTESTED PLAN HAS AN UNKNOWN RTO Not a slow one. Unknown. Restore a real backup on a schedule.
A straight trade between standing cost and recovery time. There is no clever fifth option — you are choosing a point on this line, and the honest version of that choice names the price.

Backup & restore — RTO hours to days, cheapest. Backups replicated to another region; on disaster, you rebuild infrastructure from code and restore data. Viable only if the infrastructure genuinely is code and someone has run the rebuild recently.

Pilot light — RTO tens of minutes. Data replicates continuously to the DR region; compute exists but is switched off (AMIs built, ASGs at zero, databases as small replicas). On disaster, scale up and promote.

Warm standby — RTO minutes. A scaled-down but running copy of the full stack in the second region, taking little or no traffic. On disaster, scale it up and shift DNS. Costs meaningfully more; recovers without needing capacity to appear on demand — which matters, because the control plane you would rely on is exactly what a large regional event degrades.

Active/active (multi-site) — RTO near zero. Both regions serve production. The infrastructure is straightforward; the data layer is where the difficulty lives, and it is substantial.

The uncomfortable truth about active/active: it is a data consistency problem wearing an infrastructure costume. Two regions accepting writes to the same records need conflict resolution, and conflict resolution is where correctness bugs are born. The patterns that work in practice avoid the problem rather than solving it — partition by tenant or geography so each record has one writing region, or designate a home region for writes and serve reads everywhere.


Topic 3: Multi-AZ First, and Usually Enough

Before any multi-region conversation, be honest about whether multi-AZ is done properly. Most “we need multi-region” requirements are actually “we have a single point of failure in one AZ”.

The multi-AZ checklist that catches real gaps:

□ Instances spread across ≥3 AZs, and actually running in all of them
□ NAT gateway per AZ, each private route table pointing at its own
□ Load balancer enabled in every AZ that has targets
□ RDS Multi-AZ, or Aurora with readers in other AZs
□ Any single-AZ EFS or One Zone-IA data identified and accepted
□ Quotas checked: can you run in 2 AZs the capacity you run in 3?
□ Tested: an AZ actually removed, on purpose, and the result observed

That last line is the one that is almost never done, and it is the only one that produces evidence. AWS Fault Injection Service can remove an AZ’s network reachability, stop instances, throttle an API, or inject latency — as a controlled experiment with a stop condition wired to your own alarms. Running one AZ-failure experiment per quarter, in production, with a rollback, is the difference between believing you are multi-AZ and knowing it.

The quota line matters more than it looks: losing an AZ means the remaining ones must run more instances than usual, and service quotas are per region. A quota sized for steady state becomes the reason the recovery stalls.


Topic 4: The Mechanisms

LayerCross-region optionRPONotes
S3Cross-Region ReplicationMinutes; 15 min with RTCNew objects only unless you run Batch Replication
RDSCross-region read replicaSeconds to minutesManual promotion
AuroraGlobal Database~1 secondManaged promotion, RTO ~1 minute
DynamoDBGlobal Tables~1 secondMulti-writer, last-writer-wins
EBSSnapshot copyHoursOr Elastic Disaster Recovery for continuous block replication
AMIsCopy to DR regionn/aMust be part of the pipeline, not a manual step
SecretsMulti-region secretsSecondsEasy to forget until the failover
ECRCross-region replicationMinutesA DR region with no images cannot start anything
Route 53Health checks + failover recordsn/aThis is the switch

The last four rows are where drills fail. Everyone replicates the database; the region without AMIs, container images, secrets or a working TLS certificate is discovered at the worst moment. A DR region is a full environment or it is a plan on paper.

Route 53 is the traffic switch. Failover routing with health checks moves traffic automatically; the delay is the health check interval plus the record TTL. Keep the TTL low (60s) on records that participate in failover, and remember that some resolvers ignore short TTLs — which is why Global Accelerator, using anycast IPs that stay constant while the backend changes, gives a faster and more reliable switch when you can justify it.


Topic 5: Static Stability

The principle that separates designs that survive a large event from ones that fail during recovery: a recovery that depends on the control plane depends on the thing most likely to be degraded.

Concretely:

  • Capacity that is already running is more reliable than capacity you plan to launch. Warm standby beats pilot light for this reason alone, independent of RTO.
  • Pre-provision what the recovery needs: IPs, ENIs, load balancers, quotas. Creating them during an event is a dependency on an API that may be busy.
  • Avoid recovery paths that require an AWS API you cannot test. If your runbook says “increase the quota”, the recovery includes a support case.
  • Over-provision to survive the loss. Three AZs each running at 50% survive losing one; three running at 90% do not, no matter how fast the autoscaler is.

The same logic applies to your own dependencies: a failover that requires your CI system, your secret store, and your identity provider all to be healthy has three additional single points of failure, and they are rarely in the DR plan.


Topic 6: Drills, and the Runbook They Produce

An untested DR plan does not have a slow RTO. It has an unknown RTO, which is worse, because you cannot plan around a number you do not have.

A drill schedule that is actually sustainable:

Monthly    Restore one backup to a scratch environment and verify the data.
           Rotates through services. Catches the backup that has silently
           been failing for six weeks.

Quarterly  AZ failure experiment with FIS, in production, with a stop condition.

Annually   Full regional failover drill. Announced, scheduled, with the
           runbook open and a scribe recording every step and its duration.

Every drill produces two artifacts: an updated runbook and a list of what surprised you. The surprises are the value. Common ones, from real drills: the DR region lacked a container image; the secret existed but the KMS key it used did not; the certificate was regional and had never been issued in the DR region; the runbook named a person who left; the quota in the DR region was the default because nobody had ever launched anything there.

Write the runbook for the tired version of yourself. Exact commands, exact ARNs, explicit decision points (“if X, do Y; if not, escalate to Z”), and a communications plan naming who tells customers what. A runbook that says “restore the database” is a note, not a procedure.

Try it yourself: pick your least critical production service and restore it from backup into a scratch environment, end to end, while timing each step. The number you get is your real RTO for that workload — and it is almost always larger than the number in the document.

Common mistake: measuring RTO from “restore complete” rather than from “users are being served”. The database restore is often the fastest part; the time goes into finding the right backup, waiting for DNS, discovering that the application needs a configuration change to point at a new endpoint, and the twenty minutes of deciding whether to fail over at all. Time the whole thing, decision included, or the number is fiction.