Load Balancing and Auto Scaling

Choosing between ALB and NLB on their real differences, the two independent health checks that decide whether a broken fleet ever heals, and scaling policies that react before users do.

intermediate 22 min lesson hands-on task included

A load balancer and an Auto Scaling group are two halves of one mechanism: the load balancer decides where traffic goes right now, the ASG decides how much capacity exists at all. They fail in ways that look like each other, which is why the boundary between them is worth drawing precisely.


Topic 1: Choosing a Load Balancer

Application LBNetwork LBGateway LB
Layer7 (HTTP/HTTPS/gRPC)4 (TCP/UDP/TLS)3 (traffic inspection)
Routing onPath, host, header, method, query, source IPPort onlyn/a — it forwards to appliances
Latency addedMillisecondsMicrosecondsDepends on the appliance
Static IPNo — use DNSYes, one per AZ, EIP-capablen/a
Preserves source IPNo — use X-Forwarded-ForYesYes
TLS terminationYes, with ACMYes, or pass throughn/a
WAFYesNoNo
TargetsInstance, IP, LambdaInstance, IP, ALBAppliances

Default to the ALB. Content-based routing, WAF integration, native gRPC and WebSocket support, and per-request logging are what most services need.

Reach for the NLB when you need a static IP (a partner firewall allow-lists addresses, not names), extreme throughput with minimal latency, a non-HTTP protocol, or the client’s real source IP without a header. Note the trade: an NLB preserving source IP means your instance’s security group must allow the client CIDRs, not the load balancer’s — a security group referencing the NLB does not work the way it does with an ALB.

Never use the Classic Load Balancer for new work. It is the previous generation and lacks nearly everything above.

Both are regional with zonal nodes: you enable the load balancer in specific subnets, one per AZ, and DNS returns the addresses of those nodes. That has a consequence people meet during an AZ event — if you enabled only two of three AZs, a third of your potential capacity is unreachable no matter how many instances the ASG runs there.

Cross-zone load balancing is on by default for ALB (free) and off by default for NLB (and billed as inter-AZ transfer when enabled). With it off, each zonal node distributes only to targets in its own AZ, so uneven target counts across AZs produce uneven load per instance. Even AZ counts matter more than they look.


Topic 2: The Two Health Checks

clients ALB listener :443 → rules target group HTTP GET /healthz ASG min 3 / desired 6 max 12 launch template v7 instance-1 target: healthy instance-2 target: healthy instance-3 target: unhealthy CHECK A — EC2 STATUS CHECKS Is the hypervisor reachable? Did the OS boot? A wedged app passes this happily. default ASG health source CHECK B — TARGET GROUP HEALTH Does the application answer correctly? Set the ASG health check type to ELB, or the ASG never replaces an instance the ALB has given up on. THE FAILURE MODE: A "HEALTHY" ASG SERVING 502s Six instances, all passing EC2 status checks, all failing /healthz. Traffic drains to zero targets and nothing is ever replaced.
Two independent checks answering different questions. Leaving the ASG on the EC2 default is how a fleet of technically-alive instances serves errors indefinitely without a single replacement.

EC2 status checks ask: is the host reachable, did the OS boot? A process that has deadlocked, run out of file descriptors, or is returning 500 to everything passes these happily.

Target group health checks ask: does the application answer correctly? This is the check that reflects reality.

The default ASG health check type is EC2. That default is the trap: the ALB removes failing targets from rotation, the ASG sees perfectly healthy instances, nothing is replaced, and traffic drains to zero healthy targets while the console shows a green Auto Scaling group.

aws autoscaling update-auto-scaling-group \
  --auto-scaling-group-name app-asg \
  --health-check-type ELB \
  --health-check-grace-period 300

The grace period is how long after launch the ASG ignores health checks. Set it slightly longer than your worst-case boot-to-service time. Too short and you get a replacement loop: instances are killed while still booting, forever, at full instance cost. Too long and a genuinely broken instance takes that long to be noticed.

Tune the health check itself rather than accepting defaults:

path                  /healthz   — cheap, no external dependencies
interval              10–30s
timeout               < interval
healthy threshold     2–3
unhealthy threshold   2–3
matcher               200 (be specific; 200-399 hides redirect bugs)

What /healthz should test is the argument worth having with your developers. A check that verifies the database connection turns a database blip into every instance being marked unhealthy and replaced simultaneously — a self-inflicted outage. A check that only proves the process is alive misses a broken instance. The workable rule: liveness for the load balancer check (can this process serve?), dependency checks on a separate endpoint used by monitoring and humans, never by the health check that controls replacement.

Deregistration delay (default 300s) is how long the load balancer keeps sending in-flight requests to a draining target. Set it to slightly above your longest normal request. Too high and deployments crawl; too low and you cut live connections during every scale-in.


Topic 3: Auto Scaling Groups

An ASG maintains a desired count across the subnets you give it, using a launch template.

min      3    ← never fewer, even at 3am
desired  6    ← what the policies move
max     12    ← the ceiling that bounds a runaway

Set max deliberately. It is the only thing between a traffic anomaly (or a metric bug, or a retry storm) and a five-figure surprise. Too low, though, and it silently caps you during the event you built the ASG for. Pick it from a capacity model, not from a round number, and alarm when desired reaches it — hitting max is an event you want to know about.

Availability Zones and rebalancing. An ASG spread across three subnets tries to keep zones balanced and will terminate an instance in an over-represented AZ to rebalance. If a subnet is out of IP addresses, the ASG quietly concentrates in the other two — which is how a “multi-AZ” deployment becomes single-AZ without any alarm firing. Watch for Failed to launch in scaling activities:

aws autoscaling describe-scaling-activities \
  --auto-scaling-group-name app-asg --max-items 10 \
  --query 'Activities[].[StartTime,StatusCode,StatusMessage]' --output table

Termination policy decides who dies on scale-in. The default picks the AZ with most instances, then the oldest launch template version, then the instance closest to the next billing hour. Override it when your workload has a reason — OldestInstance for forcing rotation, a custom Lambda for “not the leader”.

Lifecycle hooks pause an instance in Pending:Wait or Terminating:Wait so you can act: register with a service mesh, warm a cache, or on the way out, drain connections and ship the final logs. Without a terminating hook, scale-in loses whatever was in memory and whatever had not been flushed to disk.

Warm pools keep pre-initialised instances stopped and ready. You pay for their EBS only, and scale-out becomes a start rather than a launch — seconds instead of minutes. This is the right answer for workloads with long boot-to-service times and sharp traffic edges, and a needless complication for anything that boots in 40 seconds.


Topic 4: Scaling Policies

Target tracking — the default choice. Name a metric and a target; AWS manages the alarms.

aws autoscaling put-scaling-policy \
  --auto-scaling-group-name app-asg \
  --policy-name cpu-target --policy-type TargetTrackingScaling \
  --target-tracking-configuration '{
    "PredefinedMetricSpecification": {"PredefinedMetricType": "ASGAverageCPUUtilization"},
    "TargetValue": 60.0
  }'

It scales out aggressively and in conservatively, which is the correct asymmetry: being slow to add capacity costs availability, being slow to remove it costs a little money.

Step scaling — explicit thresholds and increments. Use it when the response should be non-linear: +1 instance at 60%, +4 at 85%.

Scheduled scaling — for known patterns. Business-hours workloads, a batch window, a marketing event with a date. Deterministic and free.

Predictive scaling — forecasts from history and provisions ahead of the curve. Genuinely useful for daily-cycle workloads whose scale-out is slower than their traffic ramp.

Choose the metric carefully. CPU is the default and often the wrong signal. A memory-bound service, an I/O-bound service, or one whose latency comes from a downstream dependency will not show the problem in CPU at all. Better targets: ALBRequestCountPerTarget (directly proportional to load, and immune to the “CPU looks fine” failure), or a custom queue-depth metric for worker fleets.

The physics you cannot policy your way around: detection (up to a metric period) + alarm evaluation + launch + boot-to-service. Three to five minutes is typical. If your traffic can double in ninety seconds, no scaling policy saves you — you need headroom, a warm pool, or a queue that absorbs the spike.


Topic 5: Instance Refresh — Deployments Without a Pipeline

Changing the launch template does nothing to running instances. Instance refresh rolls the fleet:

aws autoscaling start-instance-refresh \
  --auto-scaling-group-name app-asg \
  --preferences '{
    "MinHealthyPercentage": 90,
    "InstanceWarmup": 300,
    "CheckpointPercentages": [20, 50, 100],
    "CheckpointDelay": 600
  }'

MinHealthyPercentage bounds how much capacity can be missing at once. Checkpoints pause the rollout at a percentage so you can watch metrics before continuing — a canary without a deployment tool. If the new AMI is broken, the refresh stalls with the old instances still serving, which is the behaviour you want.

Maximum instance lifetime is the same mechanism on a timer: set it to 30 days and instances are replaced continuously, which forces you to keep AMIs current and guarantees that “we can rebuild any instance” is tested by default rather than asserted.


Topic 6: The Failure Modes Worth Recognising

SymptomUsual cause
503 from the ALBNo healthy targets in the AZ that received the request
502 from the ALBTarget closed the connection or sent a malformed response — often a keep-alive timeout shorter than the ALB’s 60s idle timeout
504 from the ALBTarget did not answer within the idle timeout
Healthy ASG, failing serviceHealth check type is EC2 instead of ELB
Instances launch and terminate in a loopGrace period shorter than boot-to-service time
Scale-out does nothingmax reached, or subnet IP exhaustion — check scaling activities
Uneven load per instanceCross-zone off with uneven AZ target counts

The 502-from-keep-alive one deserves a note because it is so often misdiagnosed as a network fault: your application’s keep-alive timeout must be longer than the load balancer’s idle timeout, or the ALB will reuse a connection the application has just closed. Nginx defaults to 75s and the ALB to 60s, which works; a service defaulting to 5s does not.

Try it yourself: with the ASG on health check type EC2, make every instance fail /healthz. Watch the target group go to zero healthy while the ASG reports full desired capacity, indefinitely. Then switch to ELB and watch the fleet heal itself. That contrast is the most valuable thing in this lesson.

Common mistake: setting the health check grace period from the boot time in the console rather than the boot-to-service time of the real application. An instance that boots in 40 seconds but takes four minutes to warm its cache and pass /healthz will be killed and relaunched forever under a 120-second grace period — at full cost, with no capacity ever entering service, and with scaling activities as the only place the loop is visible.