DR, HA, Cost & Monitoring

Encoding disaster recovery and high availability as code, automating backups and failover, the cost levers Terraform can pull, and wiring alerting alongside the resource it watches.

advanced 21 min lesson hands-on task included

Disaster recovery, cost control and monitoring are usually bolted on after an incident, an invoice or an outage. Encoded as part of the same module that creates the resource, they ship with it instead.


Topic 1: Multi-Region for DR and HA

The core strategy: deploy the same stack in two regions with aliased providers, and route traffic by health.

provider "aws" {
  region = "us-east-1"
  alias  = "primary"
}

provider "aws" {
  region = "us-west-2"
  alias  = "secondary"
}

Health-based DNS failover sends traffic to the primary and shifts to the secondary when health checks fail. Note that a DNS alias target must be a load balancer or distribution — pointing an alias record at a bare instance IP does not work, and it is a mistake that appears in a surprising amount of published example code.

The distinction worth keeping straight:

  • High availability keeps the service up when part of it fails — redundancy within a region, across availability zones.
  • Disaster recovery restores service when a whole region or dataset is lost — replication, backups, and a tested restore path.

They need different mechanisms, and most estates that claim DR have only implemented HA.


Topic 2: Database Resilience

resource "aws_db_instance" "app" {
  allocated_storage       = 20
  engine                  = "mysql"
  instance_class          = "db.t3.micro"
  username                = var.db_username
  password                = data.aws_ssm_parameter.db_password.value

  multi_az                = true       # synchronous standby, automatic failover
  backup_retention_period = 7          # days
  backup_window           = "01:00-03:00"
  skip_final_snapshot     = false      # keep a snapshot on destroy

  lifecycle {
    prevent_destroy = true
  }
}

Four lines doing real work:

  • multi_az gives automatic failover to a standby on hardware or network failure.
  • backup_retention_period enables automated backups — 0 disables them, which is the default on some resources and a genuinely dangerous one.
  • skip_final_snapshot = false means a destroy leaves you a recovery point. Tutorials set this to true to avoid clutter; production must not.
  • prevent_destroy makes an accidental destroy an error rather than an incident.

Topic 3: Backup and Retention as Code

resource "aws_s3_bucket_lifecycle_configuration" "backups" {
  bucket = aws_s3_bucket.backups.id

  rule {
    id     = "archive-then-expire"
    status = "Enabled"

    transition {
      days          = 30
      storage_class = "GLACIER"
    }

    expiration {
      days = 365
    }
  }
}

This is simultaneously a retention policy and a cost lever — data moves to cheap archival storage after 30 days and is deleted after a year. Set the numbers from your actual retention obligation, not from an example.


Topic 4: Cost Levers Terraform Can Pull

LeverMechanism
Right-sizingMatch instance classes to observed utilisation, not to guesses
Commitment discountsReserved capacity or savings plans for predictable baseline load
Scheduled shutdownStop non-production compute outside business hours
AutoscalingPay for capacity demand actually justifies
Storage tieringLifecycle transitions to archival classes
Budgets and alertsThreshold notifications before the invoice arrives
resource "aws_budgets_budget" "monthly" {
  budget_type  = "COST"
  limit_amount = "500"
  limit_unit   = "USD"
  time_unit    = "MONTHLY"

  notification {
    comparison_operator        = "GREATER_THAN"
    threshold                  = 80
    threshold_type             = "PERCENTAGE"
    notification_type          = "ACTUAL"
    subscriber_email_addresses = [var.finops_alert_email]
  }
}

Scheduled shutdown wires a cron-style event rule to a function that stops tagged instances:

resource "aws_cloudwatch_event_rule" "nightly_shutdown" {
  name                = "shutdown-non-prod"
  schedule_expression = "cron(0 18 * * ? *)"   # 18:00 UTC daily
}

resource "aws_cloudwatch_event_target" "nightly_shutdown" {
  rule      = aws_cloudwatch_event_rule.nightly_shutdown.name
  target_id = "StopInstances"
  arn       = aws_lambda_function.shutdown.arn
}

Note that commitment discounts are a billing construct, not a resource attribute. Reserved capacity and savings plans are purchased separately; there is no argument on a compute resource that makes it reserved. Published material occasionally suggests otherwise.

Cost-allocation tags are the precondition for all of this. Enforce a tagging standard through policy or none of the reporting works — which is the direct link back to the policy lesson.


Topic 5: Monitoring and Alerting

resource "aws_sns_topic" "alerts" {
  name = "platform-alerts"
}

resource "aws_sns_topic_subscription" "email" {
  topic_arn = aws_sns_topic.alerts.arn
  protocol  = "email"
  endpoint  = var.ops_alert_email
}

resource "aws_cloudwatch_metric_alarm" "high_cpu" {
  alarm_name          = "high-cpu-utilisation"
  comparison_operator = "GreaterThanThreshold"
  evaluation_periods  = 2
  metric_name         = "CPUUtilization"
  namespace           = "AWS/EC2"
  period              = 300
  statistic           = "Average"
  threshold           = 80
  alarm_description   = "CPU above 80% for two consecutive periods"
  alarm_actions       = [aws_sns_topic.alerts.arn]

  dimensions = {
    InstanceId = aws_instance.web.id
  }
}

resource "aws_cloudwatch_log_group" "app" {
  name              = "/platform/app"
  retention_in_days = 7        # retention is a cost lever as well as a policy one
}

evaluation_periods = 2 is the difference between an alarm and a pager that cries wolf — a single 300-second window above threshold is often just a deployment.

Equivalents exist on other platforms: a metric alert bound to a resource scope, delivered through an action group with email or webhook receivers.

Log retention defaults to “forever” on several platforms, which is both a compliance problem and a slow-growing bill. Set it explicitly, every time.


Topic 6: Practices That Make This Real

  • Ship monitoring with the resource. An alarm defined in the same module as the thing it watches cannot be forgotten, and is deleted when the resource is.
  • Test the DR plan. A restore path that has never been executed is a hypothesis. Schedule a drill.
  • Automate backups rather than documenting them. Documented backups are performed until the week someone is on holiday.
  • Monitor DR spend. Idle standby capacity in a second region is expensive and invisible until the invoice.
  • Alert on what predicts user impact, not on every metric available. An alarm nobody acts on trains people to ignore alarms.

Try it yourself: Add an alarm to a resource you already manage, then breach the threshold deliberately and confirm the notification arrives. An untested alarm is indistinguishable from no alarm, and roughly half of them are misconfigured on first write.

Common mistake: Treating skip_final_snapshot = true as a sensible default because it makes terraform destroy cleaner in testing. Carried into production, it means the destroy that was not supposed to happen also took your last recovery point.