Managed databases remove the work of installing and patching, and leave you every decision that actually determines availability. This lesson is about those decisions: what failover does, what it costs in time, and which of the three overlapping “resilience” features solves which problem.
Topic 1: Managed vs Self-Hosted
Running a database on EC2 is a legitimate choice with a narrow set of reasons: an engine or version RDS does not offer, an extension RDS does not permit, OS-level access your tooling requires, or a licensing arrangement that only works that way.
What you take on: patching the OS and the engine, backups and their verification, failover automation, replication setup, monitoring, and scaling. All of which are solved problems that RDS solves for a price. The honest test is whether your team would do those jobs better than the managed service — usually not, and the exceptions are specific and articulable rather than general preference.
Topic 2: Multi-AZ vs Read Replicas
Multi-AZ (one standby) — a synchronous standby in another AZ that serves nothing. Every commit waits for the standby, so writes are slightly slower. On failure, AWS flips the endpoint’s DNS to the standby in 60–120 seconds. Your endpoint name never changes; your application must reconnect.
Multi-AZ DB cluster (two readable standbys) — a newer topology with two standbys that can serve reads, and failover typically under 35 seconds because it uses a different replication mechanism. Available for MySQL and PostgreSQL, and the better choice when it fits.
Read replicas — asynchronous copies that serve read traffic. Up to 15 per source, optionally in another region. Promotion is manual and permanently breaks replication.
| Multi-AZ standby | Read replica | |
|---|---|---|
| Replication | Synchronous | Asynchronous |
| Serves traffic | No (single-standby mode) | Reads only |
| Failover | Automatic, 60–120s | Manual promotion |
| Cross-region | No | Yes |
| Effect on write latency | Slightly higher | None |
| Data loss on failover | None | Possible — whatever had not replicated |
The application-level consequence of a read replica is replica lag, and it produces a specific class of bug: a user writes a record, is redirected to a page that reads it from a replica, and the record is not there yet. Route reads that must be consistent to the writer. Alarm on ReplicaLag, because a replica that has fallen ten minutes behind is serving answers that are wrong rather than merely slow.
Multi-AZ does not improve read capacity. This is the single most common misconception in RDS, and it leads to teams paying double for a standby while their read load still lands entirely on the primary.
Topic 3: Aurora Changes the Shape
Aurora separates compute from storage. One virtual volume, replicated six ways across three AZs, with up to fifteen readers attached to the same storage rather than to copies of it.
What that changes in practice:
- Replicas share storage, so replica lag is measured in milliseconds and adding a reader does not add write amplification.
- Failover to a reader takes ~30 seconds, because there is no data to catch up on.
- Storage grows automatically in 10 GB increments to 128 TB. No volume resizing, no running out at 2am.
- Endpoints do the routing for you: a cluster endpoint that always points at the writer, a reader endpoint that load-balances across readers, and custom endpoints for subsets.
That endpoint distinction is where the money is lost. Applications configured to use the cluster endpoint for everything send all reads to the writer, and every reader you pay for sits idle. Pointing reads at the reader endpoint is usually a one-line configuration change with a large effect.
Aurora also offers three things RDS does not:
- Backtrack (MySQL-compatible) — rewinds the cluster in place to a point in time, in seconds, without a restore. The undo button for a bad migration.
- Fast cloning — a copy-on-write clone of a multi-terabyte database in minutes, at almost no storage cost until it diverges. This is how you give a test environment production-shaped data.
- Serverless v2 — scales capacity in 0.5 ACU steps, in place, without dropping connections. Good for spiky and unpredictable load; more expensive than provisioned at steady high utilisation.
Global Database replicates to other regions with typical sub-second lag and a managed promotion path — an RPO around one second and an RTO around a minute, which is the strongest cross-region story AWS offers for a relational database.
Topic 4: Backups, PITR and What They Do Not Cover
Two mechanisms, frequently confused:
Automated backups — a daily snapshot plus continuous transaction logs, retained 1–35 days. They enable point-in-time recovery to any second in the window. They are deleted when you delete the instance, and the final-snapshot prompt at deletion is the last chance to keep anything.
Manual snapshots — taken by you, retained until you delete them, survive instance deletion, and can be copied to another region or shared with another account.
Both restore into a new instance with a new endpoint. Nothing is rewound in place (except Aurora Backtrack). That means recovery has an application step: repoint the connection string, or move a Route 53 CNAME you had the foresight to put in front of the endpoint. Practising that step is what turns a documented RTO into a real one.
What is not covered and needs its own answer:
- A logical error replicated to every replica instantly — replicas are not backups.
- A dropped table where PITR would mean losing every other write since. Extract the table from a restored copy instead of rolling the whole database back.
- Retention beyond 35 days for compliance — export snapshots to S3, or take periodic manual snapshots.
The maintenance window is not optional and is worth choosing. Engine patches, minor version upgrades and OS updates apply in it, and Multi-AZ makes them near-invisible: AWS patches the standby, fails over, then patches the former primary. That failover is a real connection drop, so “near-invisible” still requires an application that reconnects. Set the window to your genuine trough and set AutoMinorVersionUpgrade deliberately rather than by default.
Topic 5: Connections, Parameters and the Metrics That Matter
Connection limits are derived from instance memory by a formula in the parameter group, and applications that open a connection per request exhaust them long before CPU is a concern. RDS Proxy pools and multiplexes connections, which fixes that class of problem and additionally shortens failover for the application, because the proxy holds the client connection open while it reconnects underneath.
Parameter groups apply engine settings; some are dynamic and some require a reboot, and the console tells you which. The reboot requirement is why parameter changes belong in a change window rather than in a hurry.
Subnet groups decide which subnets the instance can live in — put them in your private data subnets, across at least the AZs you intend to fail over between.
The metrics worth an alarm, and why:
| Metric | Why it matters |
|---|---|
CPUUtilization | The obvious one, and the least often the cause |
DatabaseConnections | Against the max — exhaustion is a hard failure |
FreeableMemory | Falling toward zero means swapping, then an OOM restart |
ReadLatency / WriteLatency | Storage-level slowness before the application notices |
ReplicaLag | Correctness, not just performance |
FreeStorageSpace | A full disk stops the database and cannot be fixed quickly |
DiskQueueDepth | I/O saturation |
BurstBalance | On gp2 storage, the cliff before it arrives |
Performance Insights answers “what is the database actually doing” with per-query load attribution, and it is free at 7-day retention. Enabling it before an incident is worth far more than enabling it during one.
Topic 6: Security Defaults Worth Setting Once
✓ Never publicly accessible PubliclyAccessible=false, private subnets
✓ Encryption at rest must be set at creation — cannot be added later
✓ Encryption in transit enforce with rds.force_ssl / require_secure_transport
✓ IAM database authentication short-lived tokens instead of stored passwords
✓ Secrets Manager with rotation for the credentials that must exist
✓ Deletion protection on every production instance
✓ Enhanced monitoring + PI on before you need them
✓ Audit logs to CloudWatch Logs exported, retained, and actually searchable
Two of those deserve emphasis. Encryption at rest cannot be enabled on an existing instance — the path is snapshot, copy the snapshot with encryption, restore. Doing it later means a maintenance window; doing it at creation costs nothing.
Deletion protection is the cheapest control in this list and the one that most often justifies itself. It does not stop a DeleteDBCluster from a determined operator; it stops the accidental one.
Try it yourself: force a failover on a Multi-AZ instance while running a query loop. Time the gap. Then repeat with RDS Proxy in front and compare. The difference is the value of connection pooling stated in seconds of downtime rather than in architecture-diagram terms.
Common mistake: treating a read replica as a disaster recovery plan without ever testing promotion. Promotion is manual, it takes minutes, it breaks replication irreversibly, and the application still needs repointing. A cross-region replica is an excellent DR component and is not, by itself, a DR plan — the plan is the written sequence, and the proof is that you have executed it.