RDS and Aurora Operations

Why Multi-AZ adds no read capacity, what point-in-time recovery really restores, and the maintenance and failover behaviours that decide whether your application notices.

intermediate 22 min lesson hands-on task included

Managed databases remove the work of installing and patching, and leave you every decision that actually determines availability. This lesson is about those decisions: what failover does, what it costs in time, and which of the three overlapping “resilience” features solves which problem.


Topic 1: Managed vs Self-Hosted

Running a database on EC2 is a legitimate choice with a narrow set of reasons: an engine or version RDS does not offer, an extension RDS does not permit, OS-level access your tooling requires, or a licensing arrangement that only works that way.

What you take on: patching the OS and the engine, backups and their verification, failover automation, replication setup, monitoring, and scaling. All of which are solved problems that RDS solves for a price. The honest test is whether your team would do those jobs better than the managed service — usually not, and the exceptions are specific and articulable rather than general preference.


Topic 2: Multi-AZ vs Read Replicas

MULTI-AZ — AVAILABILITY primary AZ a · read + write standby AZ b · serves nothing synchronous Failover flips the DNS CNAME in ~60–120s. Your endpoint does not change. Your app must reconnect. It does NOT add read capacity. It also moves backup I/O off the primary, which is a real win. READ REPLICA — CAPACITY primary writes replica 1 reads only replica 2 reads only asynchronous → replica lag is a metric you must alarm on Promotion is manual and breaks replication permanently. Can live in another region — that is your DR read copy. A stale read after a write is the bug this design creates. AURORA CHANGES THE SHAPE: STORAGE IS SHARED, NOT COPIED One 6-way replicated volume across 3 AZs, up to 15 readers on it. A reader can be promoted in ~30s, and the reader endpoint load-balances for you. Write to the cluster endpoint, read from the reader endpoint. Pointing reads at the cluster endpoint wastes every replica you pay for. BACKUPS ARE A THIRD, SEPARATE THING Automated backups + transaction logs give point-in-time recovery to any second in the window. Manual snapshots persist after you delete the instance — automated ones do not.
Two features that both create a second database and solve entirely different problems. Multi-AZ buys availability and no capacity; a read replica buys capacity and no automatic failover.

Multi-AZ (one standby) — a synchronous standby in another AZ that serves nothing. Every commit waits for the standby, so writes are slightly slower. On failure, AWS flips the endpoint’s DNS to the standby in 60–120 seconds. Your endpoint name never changes; your application must reconnect.

Multi-AZ DB cluster (two readable standbys) — a newer topology with two standbys that can serve reads, and failover typically under 35 seconds because it uses a different replication mechanism. Available for MySQL and PostgreSQL, and the better choice when it fits.

Read replicas — asynchronous copies that serve read traffic. Up to 15 per source, optionally in another region. Promotion is manual and permanently breaks replication.

Multi-AZ standbyRead replica
ReplicationSynchronousAsynchronous
Serves trafficNo (single-standby mode)Reads only
FailoverAutomatic, 60–120sManual promotion
Cross-regionNoYes
Effect on write latencySlightly higherNone
Data loss on failoverNonePossible — whatever had not replicated

The application-level consequence of a read replica is replica lag, and it produces a specific class of bug: a user writes a record, is redirected to a page that reads it from a replica, and the record is not there yet. Route reads that must be consistent to the writer. Alarm on ReplicaLag, because a replica that has fallen ten minutes behind is serving answers that are wrong rather than merely slow.

Multi-AZ does not improve read capacity. This is the single most common misconception in RDS, and it leads to teams paying double for a standby while their read load still lands entirely on the primary.


Topic 3: Aurora Changes the Shape

Aurora separates compute from storage. One virtual volume, replicated six ways across three AZs, with up to fifteen readers attached to the same storage rather than to copies of it.

What that changes in practice:

  • Replicas share storage, so replica lag is measured in milliseconds and adding a reader does not add write amplification.
  • Failover to a reader takes ~30 seconds, because there is no data to catch up on.
  • Storage grows automatically in 10 GB increments to 128 TB. No volume resizing, no running out at 2am.
  • Endpoints do the routing for you: a cluster endpoint that always points at the writer, a reader endpoint that load-balances across readers, and custom endpoints for subsets.

That endpoint distinction is where the money is lost. Applications configured to use the cluster endpoint for everything send all reads to the writer, and every reader you pay for sits idle. Pointing reads at the reader endpoint is usually a one-line configuration change with a large effect.

Aurora also offers three things RDS does not:

  • Backtrack (MySQL-compatible) — rewinds the cluster in place to a point in time, in seconds, without a restore. The undo button for a bad migration.
  • Fast cloning — a copy-on-write clone of a multi-terabyte database in minutes, at almost no storage cost until it diverges. This is how you give a test environment production-shaped data.
  • Serverless v2 — scales capacity in 0.5 ACU steps, in place, without dropping connections. Good for spiky and unpredictable load; more expensive than provisioned at steady high utilisation.

Global Database replicates to other regions with typical sub-second lag and a managed promotion path — an RPO around one second and an RTO around a minute, which is the strongest cross-region story AWS offers for a relational database.


Topic 4: Backups, PITR and What They Do Not Cover

Two mechanisms, frequently confused:

Automated backups — a daily snapshot plus continuous transaction logs, retained 1–35 days. They enable point-in-time recovery to any second in the window. They are deleted when you delete the instance, and the final-snapshot prompt at deletion is the last chance to keep anything.

Manual snapshots — taken by you, retained until you delete them, survive instance deletion, and can be copied to another region or shared with another account.

Both restore into a new instance with a new endpoint. Nothing is rewound in place (except Aurora Backtrack). That means recovery has an application step: repoint the connection string, or move a Route 53 CNAME you had the foresight to put in front of the endpoint. Practising that step is what turns a documented RTO into a real one.

What is not covered and needs its own answer:

  • A logical error replicated to every replica instantly — replicas are not backups.
  • A dropped table where PITR would mean losing every other write since. Extract the table from a restored copy instead of rolling the whole database back.
  • Retention beyond 35 days for compliance — export snapshots to S3, or take periodic manual snapshots.

The maintenance window is not optional and is worth choosing. Engine patches, minor version upgrades and OS updates apply in it, and Multi-AZ makes them near-invisible: AWS patches the standby, fails over, then patches the former primary. That failover is a real connection drop, so “near-invisible” still requires an application that reconnects. Set the window to your genuine trough and set AutoMinorVersionUpgrade deliberately rather than by default.


Topic 5: Connections, Parameters and the Metrics That Matter

Connection limits are derived from instance memory by a formula in the parameter group, and applications that open a connection per request exhaust them long before CPU is a concern. RDS Proxy pools and multiplexes connections, which fixes that class of problem and additionally shortens failover for the application, because the proxy holds the client connection open while it reconnects underneath.

Parameter groups apply engine settings; some are dynamic and some require a reboot, and the console tells you which. The reboot requirement is why parameter changes belong in a change window rather than in a hurry.

Subnet groups decide which subnets the instance can live in — put them in your private data subnets, across at least the AZs you intend to fail over between.

The metrics worth an alarm, and why:

MetricWhy it matters
CPUUtilizationThe obvious one, and the least often the cause
DatabaseConnectionsAgainst the max — exhaustion is a hard failure
FreeableMemoryFalling toward zero means swapping, then an OOM restart
ReadLatency / WriteLatencyStorage-level slowness before the application notices
ReplicaLagCorrectness, not just performance
FreeStorageSpaceA full disk stops the database and cannot be fixed quickly
DiskQueueDepthI/O saturation
BurstBalanceOn gp2 storage, the cliff before it arrives

Performance Insights answers “what is the database actually doing” with per-query load attribution, and it is free at 7-day retention. Enabling it before an incident is worth far more than enabling it during one.


Topic 6: Security Defaults Worth Setting Once

✓ Never publicly accessible          PubliclyAccessible=false, private subnets
✓ Encryption at rest                 must be set at creation — cannot be added later
✓ Encryption in transit              enforce with rds.force_ssl / require_secure_transport
✓ IAM database authentication        short-lived tokens instead of stored passwords
✓ Secrets Manager with rotation      for the credentials that must exist
✓ Deletion protection                on every production instance
✓ Enhanced monitoring + PI           on before you need them
✓ Audit logs to CloudWatch Logs      exported, retained, and actually searchable

Two of those deserve emphasis. Encryption at rest cannot be enabled on an existing instance — the path is snapshot, copy the snapshot with encryption, restore. Doing it later means a maintenance window; doing it at creation costs nothing.

Deletion protection is the cheapest control in this list and the one that most often justifies itself. It does not stop a DeleteDBCluster from a determined operator; it stops the accidental one.

Try it yourself: force a failover on a Multi-AZ instance while running a query loop. Time the gap. Then repeat with RDS Proxy in front and compare. The difference is the value of connection pooling stated in seconds of downtime rather than in architecture-diagram terms.

Common mistake: treating a read replica as a disaster recovery plan without ever testing promotion. Promotion is manual, it takes minutes, it breaks replication irreversibly, and the application still needs repointing. A cross-region replica is an excellent DR component and is not, by itself, a DR plan — the plan is the written sequence, and the proof is that you have executed it.