Disaster recovery, cost control and monitoring are usually bolted on after an incident, an invoice or an outage. Encoded as part of the same module that creates the resource, they ship with it instead.
Topic 1: Multi-Region for DR and HA
The core strategy: deploy the same stack in two regions with aliased providers, and route traffic by health.
provider "aws" {
region = "us-east-1"
alias = "primary"
}
provider "aws" {
region = "us-west-2"
alias = "secondary"
}
Health-based DNS failover sends traffic to the primary and shifts to the secondary when health checks fail. Note that a DNS alias target must be a load balancer or distribution — pointing an alias record at a bare instance IP does not work, and it is a mistake that appears in a surprising amount of published example code.
The distinction worth keeping straight:
- High availability keeps the service up when part of it fails — redundancy within a region, across availability zones.
- Disaster recovery restores service when a whole region or dataset is lost — replication, backups, and a tested restore path.
They need different mechanisms, and most estates that claim DR have only implemented HA.
Topic 2: Database Resilience
resource "aws_db_instance" "app" {
allocated_storage = 20
engine = "mysql"
instance_class = "db.t3.micro"
username = var.db_username
password = data.aws_ssm_parameter.db_password.value
multi_az = true # synchronous standby, automatic failover
backup_retention_period = 7 # days
backup_window = "01:00-03:00"
skip_final_snapshot = false # keep a snapshot on destroy
lifecycle {
prevent_destroy = true
}
}
Four lines doing real work:
multi_azgives automatic failover to a standby on hardware or network failure.backup_retention_periodenables automated backups —0disables them, which is the default on some resources and a genuinely dangerous one.skip_final_snapshot = falsemeans a destroy leaves you a recovery point. Tutorials set this totrueto avoid clutter; production must not.prevent_destroymakes an accidental destroy an error rather than an incident.
Topic 3: Backup and Retention as Code
resource "aws_s3_bucket_lifecycle_configuration" "backups" {
bucket = aws_s3_bucket.backups.id
rule {
id = "archive-then-expire"
status = "Enabled"
transition {
days = 30
storage_class = "GLACIER"
}
expiration {
days = 365
}
}
}
This is simultaneously a retention policy and a cost lever — data moves to cheap archival storage after 30 days and is deleted after a year. Set the numbers from your actual retention obligation, not from an example.
Topic 4: Cost Levers Terraform Can Pull
| Lever | Mechanism |
|---|---|
| Right-sizing | Match instance classes to observed utilisation, not to guesses |
| Commitment discounts | Reserved capacity or savings plans for predictable baseline load |
| Scheduled shutdown | Stop non-production compute outside business hours |
| Autoscaling | Pay for capacity demand actually justifies |
| Storage tiering | Lifecycle transitions to archival classes |
| Budgets and alerts | Threshold notifications before the invoice arrives |
resource "aws_budgets_budget" "monthly" {
budget_type = "COST"
limit_amount = "500"
limit_unit = "USD"
time_unit = "MONTHLY"
notification {
comparison_operator = "GREATER_THAN"
threshold = 80
threshold_type = "PERCENTAGE"
notification_type = "ACTUAL"
subscriber_email_addresses = [var.finops_alert_email]
}
}
Scheduled shutdown wires a cron-style event rule to a function that stops tagged instances:
resource "aws_cloudwatch_event_rule" "nightly_shutdown" {
name = "shutdown-non-prod"
schedule_expression = "cron(0 18 * * ? *)" # 18:00 UTC daily
}
resource "aws_cloudwatch_event_target" "nightly_shutdown" {
rule = aws_cloudwatch_event_rule.nightly_shutdown.name
target_id = "StopInstances"
arn = aws_lambda_function.shutdown.arn
}
Note that commitment discounts are a billing construct, not a resource attribute. Reserved capacity and savings plans are purchased separately; there is no argument on a compute resource that makes it reserved. Published material occasionally suggests otherwise.
Cost-allocation tags are the precondition for all of this. Enforce a tagging standard through policy or none of the reporting works — which is the direct link back to the policy lesson.
Topic 5: Monitoring and Alerting
resource "aws_sns_topic" "alerts" {
name = "platform-alerts"
}
resource "aws_sns_topic_subscription" "email" {
topic_arn = aws_sns_topic.alerts.arn
protocol = "email"
endpoint = var.ops_alert_email
}
resource "aws_cloudwatch_metric_alarm" "high_cpu" {
alarm_name = "high-cpu-utilisation"
comparison_operator = "GreaterThanThreshold"
evaluation_periods = 2
metric_name = "CPUUtilization"
namespace = "AWS/EC2"
period = 300
statistic = "Average"
threshold = 80
alarm_description = "CPU above 80% for two consecutive periods"
alarm_actions = [aws_sns_topic.alerts.arn]
dimensions = {
InstanceId = aws_instance.web.id
}
}
resource "aws_cloudwatch_log_group" "app" {
name = "/platform/app"
retention_in_days = 7 # retention is a cost lever as well as a policy one
}
evaluation_periods = 2 is the difference between an alarm and a pager that cries wolf — a single 300-second window above threshold is often just a deployment.
Equivalents exist on other platforms: a metric alert bound to a resource scope, delivered through an action group with email or webhook receivers.
Log retention defaults to “forever” on several platforms, which is both a compliance problem and a slow-growing bill. Set it explicitly, every time.
Topic 6: Practices That Make This Real
- Ship monitoring with the resource. An alarm defined in the same module as the thing it watches cannot be forgotten, and is deleted when the resource is.
- Test the DR plan. A restore path that has never been executed is a hypothesis. Schedule a drill.
- Automate backups rather than documenting them. Documented backups are performed until the week someone is on holiday.
- Monitor DR spend. Idle standby capacity in a second region is expensive and invisible until the invoice.
- Alert on what predicts user impact, not on every metric available. An alarm nobody acts on trains people to ignore alarms.
Try it yourself: Add an alarm to a resource you already manage, then breach the threshold deliberately and confirm the notification arrives. An untested alarm is indistinguishable from no alarm, and roughly half of them are misconfigured on first write.
Common mistake: Treating skip_final_snapshot = true as a sensible default because it makes terraform destroy cleaner in testing. Carried into production, it means the destroy that was not supposed to happen also took your last recovery point.